Unaligned infrared visible image fusion method based on information compensation and prediction correction
Through cross-modal information compensation and double deformation field prediction correction methods, the artifacts and offset problems caused by misalignment in infrared visible light image fusion are solved, and high-quality fusion images are generated, which improves the fusion performance and the effect of downstream tasks.
Patent Information
- Application Number
- CN202510763581.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing infrared visible light image fusion methods are prone to artifacts and offsets when not aligned, affecting the performance of downstream tasks and limiting their application in real scenarios.
The methods of cross-modal information compensation and double deformation field prediction and correction are adopted. Through feature encoding, cross-modal information compensation module, double deformation field prediction and correction module, and feature fusion module, the spatial misalignment of infrared visible image pairs are corrected, and artifacts and offsets in the fusion result are eliminated.
It effectively reduces modal differences, improves fusion performance, generates high-quality artifact-free and offset fusion images, and improves the effect of downstream tasks such as image recognition and target tracking.
Smart Images

Figure CN120339091A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an unaligned infrared and visible image fusion method based on information compensation and prediction correction, belonging to the technical field of image fusion. Background Art
[0002] Infrared and visible image fusion integrates information from infrared images and visible images to generate a more informative fused image. Infrared images capture thermal information but lack the ability to describe texture details of images; while visible images capture texture detail information on the surface of objects, but a large amount of information will be lost once affected by extreme weather, occlusion, and lighting. Therefore, the fused image obtained by integrating the information of these two modalities of images can retain both the texture details in the visible image and the thermal information in the infrared image. However, in actual application scenarios, due to the shooting of sensors in different environments, it is easy to have a situation where the infrared and visible image pairs are unaligned in space. Directly fusing the unaligned infrared and visible image pairs often results in a fused image full of artifacts and offsets, seriously affecting the performance of downstream tasks, such as image recognition, semantic segmentation, target tracking, etc., and greatly limiting the application of infrared and visible image fusion in real scenarios. Thus, in view of the above predicament, an unaligned infrared and visible image fusion method based on cross-modal information compensation and dual deformation field prediction correction is proposed. This method reduces the negative impact of the misalignment of the source image pairs on the fusion result and obtains a fused result with good visual effects without artifacts and offsets. Summary of the Invention
[0003] To solve the deficiencies of existing methods, the present invention provides an unaligned infrared and visible image fusion method based on information compensation and prediction correction. The present invention can correct the spatial misalignment of infrared and visible image pairs and eliminate artifacts and offsets in the fusion result, improving the fusion performance.
[0004] The technical solution of the present invention is: an unaligned infrared and visible image fusion method based on information compensation and prediction correction, the method comprising:
[0005] Step 1, obtaining a training data set for unaligned infrared and visible image fusion, wherein the fixed image is a visible image and the offset image is an infrared image;
[0006] Step 2, inputting the fixed image and the offset image into a feature encoder to obtain fixed features and offset features;
[0007] Step 3, inputting the fixed features and the offset features into a cross-modal information compensation module to respectively transform these two single-modal features into full-modal fixed features and full-modal offset features for reducing the modal difference between the fixed features and the offset features;
[0008] Step 4: Input the fixed features and offset features with reduced modal differences into the dual deformation field prediction and correction module to correct the offset features and obtain corrected features;
[0009] Step 5: Input the corrected features and fixed features into the feature fusion module and reconstruct the infrared-visible light fused image.
[0010] As a further solution of the present invention, in Step 1, each infrared-visible light image pair in the training dataset consists of an infrared image and a visible light image; the size of the images in each modality is 256×256;
[0011] Apply the same offset to the pre-aligned infrared-visible light image pair to obtain another pair of aligned offset infrared-visible light image pairs , and take as the unaligned infrared-visible light image pair, with the fixed image being the visible light image , and the offset image being the infrared image ; preprocess the training dataset for unaligned infrared-visible light image fusion, and the preprocessing includes: randomly flipping, randomly rotating and offsetting the data, and normalizing the processed images.
[0012] As a further solution of the present invention, Step 2 includes:
[0013] The fixed image is the visible light image , and the offset image is the infrared image . Input the visible light image and the infrared image into the fixed feature encoder and the offset feature encoder respectively to obtain the fixed features and the offset features ; the generation process of the fixed features and the offset features is expressed as:
[0014] ;
[0015] Both the fixed feature encoder and the offset feature encoder are connected in a dense connection manner by feature extraction blocks composed of 4 3×3 convolutional layers and ReLU activation functions, and finally composed of a Concat operation and a 1×1 convolutional layer.
[0016] As a further solution of the present invention, in Step 3, the specific operations of the cross-modal information compensation module are as follows:
[0017] The fixed image is a visible light image , and the offset image is an infrared image . The unaligned infrared image and the visible light image obtain features after CLIP and . And use the broadcasting mechanism to expand and into and , so that they have the same size as and . Apply and , the fixed feature and the offset feature through cross-modal compensation and 3×3 convolutional layer processing to obtain the full-modal fixed feature and the full-modal offset feature :
[0018] .
[0019] As a further solution of the present invention, step 4 includes:
[0020] Generate deformation fields and respectively from the full-modal fixed feature and the full-modal offset feature after reducing the modal difference, and the visible light image gradient and the infrared image gradient of the image to be aligned; The generation processes of the deformation fields and are represented by the following formulas:
[0021] ϕ f = DFPNet1 [ F vis,c , F ˜ ir,c ] ϕ g = DFPNet2 [ ∇ I vis , ∇ I ˜ ir ] ;
[0022] where DFPNet1 and DFPNet2 are U-Net-like networks, and DFPNet1 and DFPNet2 have the same structure but different parameters;
[0023] Apply the deformation fields and to the infrared image respectively to obtain the corrected image; Compare the gradient of the corrected image with the gradient of Perform splicing, and perform global average pooling and global max pooling on the channel and spatial positions respectively; add the results after pooling, and feed them into the Sigmoid activation function to obtain the corrected weight values of all pixels at the channel and spatial positions; through the broadcasting mechanism, form a matrix for dot product operation with the weight values on the channel and the weight values on the spatial position, and obtain the weight matrix for correcting and integrating the deformation field through the dot product operation of the matrix and ; the corrected deformation field after adjustment and integration is expressed as:
[0024] ;
[0025] Apply the corrected deformation field to the offset feature After the Warp operation, the corrected feature is obtained.
[0026] As a further solution of the present invention, in step 5, the processing flow of the feature fusion module includes:
[0027] Process the corrected feature and the fixed feature through the Sigmoid activation function to obtain the weight maps and ; splice the weight maps and with the corrected feature and the fixed feature and further extract features through a 3×3 convolutional layer; jointly use global average pooling and global max pooling operations on the features of each channel, add the results after pooling and feed them into a multi-layer perceptron; finally, after being processed by the Sigmoid activation function, obtain the weights and that can express the importance of each channel; add the outputs after global max pooling of the infrared and visible light branches, feed the added result into a multi-layer perceptron, and then through the Sigmoid activation function, obtain the weights of the infrared image and the visible light image at the same channel position ; add the weights to the original weights of the infrared and visible light branches and respectively, so as to obtain the final weights and of these two branches:
[0028] ;
[0029] According to and With calibration features and fixed features , finally obtaining fused features ; The fused features are expressed as:
[0030] ;
[0031] where and are broadcasted by and respectively, and represents element-wise multiplication of the expression matrices;
[0032] The fused features are fed into a decoder composed of a 3×3 convolutional layer and a Tanh activation function , obtaining a fused image .
[0033] As a further solution of the present invention, step 3 further includes:
[0034] Defining a consistency loss to optimize the 3×3 convolutional layer. The consistency loss is expressed as:
[0035] ;
[0036] where represents norm, is the fused feature, is the full-modal fixed feature, is the full-modal offset feature;
[0037] and are obtained by full-modal feature representation learning; the specific operations are as follows:
[0038] These two pairs of aligned infrared and visible light image pairs and are respectively input into the feature encoders and with shared parameters. The extracted features are fed into the feature fusion module in step 5 for fusion, finally obtaining the fused full-modal fused features and .
[0039] As a further solution of the present invention, step 5 further includes:
[0040] Using texture loss to retain the texture details of infrared images and visible light images; The texture loss The calculation formula is as follows:
[0041] ;
[0042] Among them, 、 、 respectively represent the gradients of the fused image 、infrared image 、visible light image , represents per-pixel maximum selection;
[0043] In terms of retaining the image content information, the content loss is defined to bring the fused image closer to the source image at the pixel level. The definition of the content loss is as follows:
[0044] ; ;
[0045] Among them, and are respectively the saliency matrices of the infrared image and the visible light image , and are respectively the weight maps of the infrared image and the visible light image .
[0046] The present invention also provides an unaligned infrared and visible image fusion system based on information compensation and prediction correction. The system includes: a module for executing the unaligned infrared and visible image fusion method based on information compensation and prediction correction.
[0047] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the unaligned infrared and visible image fusion method based on information compensation and prediction correction is implemented.
[0048] The beneficial effects of the present invention are:
[0049] 1. The present invention converts single-modal features into full-modal features through a cross-modal information compensation mechanism, effectively reducing the modal difference, thereby alleviating the challenges brought by the modal difference to cross-modal feature matching;
[0050] 2. The present invention predicts two deformation fields from two perspectives of features and source image gradients respectively through the method of dual deformation field prediction and correction, and uses the standard gradient to perceive the accuracy of the deformation field prediction, and integrates and generates a more accurate deformation field, improving the accuracy of feature correction;
[0051] 3. The present invention designs an adaptive perception feature fusion module to perceive the importance degree of each feature for the fusion result, and adaptively fuse the features according to these perception scores to reconstruct a high-quality fused image without artifacts and offsets, solving the difficulty that existing methods are difficult to handle the fusion of unaligned infrared and visible light image pairs;
[0052] 4. A large number of experimental results on public datasets show that the method proposed by the present invention can effectively fuse unaligned infrared and visible light image pairs and has better performance than existing advanced methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic flow structure diagram of the present invention;
[0054] Figure 2 is a schematic diagram of the cross-modal information compensation module of the present invention;
[0055] Figure 3 is a schematic diagram of the dual deformation field prediction and correction module of the present invention;
[0056] Figure 4 is a schematic diagram of the adaptive perception feature fusion module of the present invention;
[0057] Figure 5 is a comparison diagram of the test effects of the method of the present invention and existing methods;
[0058] Figure 6 is a quantitative comparison diagram of the fusion results of existing methods and the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Embodiment 1: As Figures 1 - 6 shown, in view of the spatial misalignment of infrared and visible light image pairs in the present invention, artifacts and offsets are introduced in the fusion result, thus seriously affecting the performance of downstream tasks and restricting its practical application. An unaligned infrared and visible image fusion method based on cross-modal information compensation and dual deformation field prediction and correction is proposed.
[0060] An unaligned infrared and visible image fusion method based on information compensation and prediction correction, the method comprising:
[0061] Step 1, obtain a training dataset for unaligned infrared and visible light image fusion, wherein the fixed image is a visible light image and the offset image is an infrared image;
[0062] As a further solution of the present invention, in step 1, each infrared and visible light image pair in the training dataset consists of an infrared image and a visible light image; the size of each modal image is 256×256;
[0063] For pre-aligned infrared and visible light image pairs Apply the same offset to obtain another pair of aligned offset infrared and visible light image pairs , take as the unaligned infrared and visible light image pair, with the fixed image being the visible light image , and the offset image being the infrared image ; Preprocess the training dataset for unaligned infrared and visible light image fusion, and the preprocessing includes: randomly flipping, randomly rotating and offsetting the data, and normalizing the processed images.
[0064] Step 2: Input the fixed image and the offset image into the feature encoder to obtain the fixed feature and the offset feature;
[0065] As a further solution of the present invention, the step 2 includes:
[0066] The fixed image is the visible light image , and the offset image is the infrared image , input the visible light image and the infrared image into the fixed feature encoder and the offset feature encoder respectively to obtain the fixed feature and the offset feature ; The generation process of the fixed feature and the offset feature is expressed as:
[0067] ;
[0068] Both the fixed feature encoder and the offset feature encoder are connected in a densely connected manner by feature extraction blocks composed of 4 3×3 convolutional layers and ReLU activation functions, and finally composed of a Concat operation and a 1×1 convolutional layer.
[0069] Step 3: Input the fixed feature and the offset feature into the cross-modal information compensation module to respectively convert these two single-modal features into full-modal fixed features and full-modal offset features for reducing the modal difference between the fixed feature and the offset feature;
[0070] As a further solution of the present invention, in the step 3, the schematic diagram of the cross-modal information compensation module is as shown in Figure 2 , and the specific operation of the cross-modal information compensation module is as follows:
[0071] The fixed image is the visible light image , and the offset image is the infrared image , input the unaligned infrared image and visible light images After passing through CLIP, features are obtained and , and the broadcast mechanism is used to and be expanded into and , so that it has the same size as and ; and , fixed features and offset features After cross-modal compensation and processing by a 3×3 convolutional layer, full-modal fixed features and full-modal offset features are obtained:
[0072] .
[0073] Step 4: Input the fixed features and offset features with reduced modal differences into the dual deformation field prediction and correction module to correct the offset features and obtain corrected features; as Figure 3 shown, the said step 4 includes:
[0074] The full-modal fixed features and full-modal offset features with reduced modal differences, as well as the visible light image gradient of the image to be aligned and the infrared image gradient are respectively used to generate deformation fields and ; The generation process of deformation fields and is expressed by the following formula:
[0075] ϕ f = DFPNet1 [ F vis,c , F ˜ ir,c ] ϕ g = DFPNet2 [ ∇ I vis , ∇ I ˜ ir ] ;
[0076] wherein, DFPNet1 and DFPNet2 are U-Net-like networks, and DFPNet1 and DFPNet2 have the same structure but different parameters;
[0077] The deformation fields and are respectively applied to the infrared image to obtain the corrected image; The gradient of the corrected image is compared with the gradient of Perform splicing, and perform global average pooling and global max pooling on the channel and spatial positions respectively; add the results after pooling and send them into the Sigmoid activation function to obtain the calibration weight values of all pixels in the channel and spatial positions; through the broadcast mechanism, form a matrix for dot multiplication operation with the weight values on the channel and the weight values on the spatial position, and obtain the weight matrix for calibrating and integrating the deformation field through the dot multiplication operation of the matrix and ; the calibrated deformation field after adjustment and integration is expressed as:
[0078] ;
[0079] Multiply the calibrated deformation field with the offset feature After the Warp operation, the calibrated feature is obtained.
[0080] Step 5: Input the calibrated feature and the fixed feature into the feature fusion module and reconstruct the infrared-visible light fusion image.
[0081] As a further solution of the present invention, in the said step 5, as Figure 4 shown, the processing flow of the feature fusion module includes:
[0082] Process the calibrated feature and the fixed feature through the Sigmoid activation function to obtain the weight maps and ; splice the weight maps and with the calibrated feature and the fixed feature and further extract features through a 3×3 convolutional layer; jointly use global average pooling and global max pooling operations on the features of each channel, add the results after pooling and send them into a multi-layer perceptron; finally, after being processed by the Sigmoid activation function, obtain the weights and that can express the importance of each channel; add the outputs after global max pooling of the infrared and visible light branches, send the added result into a multi-layer perceptron, and then through the Sigmoid activation function, obtain the weights of the infrared image and the visible light image at the same channel position ; add the weights to the original weights of the infrared and visible light branches respectively and to obtain the final weights and of these two branches:
[0083] ;
[0084] According to and with the calibration feature and the fixed feature , finally obtain the fused feature ; The fused feature is expressed as:
[0085] ;
[0086] wherein and are respectively broadcast by and , and the expression matrix is multiplied element by element; The fused feature
[0087] is fed into a decoder composed of a 3×3 convolutional layer and a Tanh activation function , and the fused image is obtained .
[0088] As a further solution of the present invention, step 3 further includes:
[0089] In order to make and effectively eliminate the influence brought by the modality difference and prevent and from losing complementary information, a consistency loss is defined to optimize the 3×3 convolutional layer, and the consistency loss is expressed as:
[0090] ;
[0091] wherein, represents norm, is the fused feature, is the full-modal fixed feature, is the full-modal offset feature;
[0092] and are obtained by full-modal feature representation learning; the specific operation is as follows:
[0093] These two pairs of aligned infrared and visible light image pairs and are respectively input into the feature encoders with shared parameters and Among them, the extracted features are sent to the feature fusion module in step 5 for fusion, and finally the fused features of the full modality after fusion are obtained. and .
[0094] As a further solution of the present invention, step 5 further includes:
[0095] Using texture loss to retain the texture details of infrared images and visible light images; the texture loss The calculation formula is as follows:
[0096] ;
[0097] Wherein, , , respectively represent the gradients of the fused image , infrared image , visible light image , represents per-pixel maximum value selection;
[0098] In terms of retaining the image content information, a content loss is defined to bring the fused image closer to the source image at the pixel level. The definition of the content loss is as follows:
[0099] ; ;
[0100] Wherein, and are respectively the saliency matrices of the infrared image and the visible light image , and are respectively the weight maps of the infrared image and the visible light image . In addition, in order to mine as much complementary information of the input multi-modal features as possible, a complementary information mining loss is defined:
[0101] ;
[0102] The present invention also provides an unaligned infrared and visible image fusion system based on information compensation and prediction correction. The system includes: a module for executing the unaligned infrared and visible image fusion method based on information compensation and prediction correction.
[0103] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for fusing misaligned infrared and visible images based on information compensation and prediction correction is implemented.
[0104] To verify the effectiveness of the method of the present invention, the present invention evaluated the performance of the proposed method on the publicly available RoadScene, MSRS, M3FD, and CVC-14 datasets. In this process, the corresponding model of the method of the present invention was trained on the training set of the RoadScene dataset and the results were tested on the test sets of the RoadScene, MSRS, M3FD, and CVC-14 datasets. The training set consists of 221 pairs of infrared and visible light image pairs from the RoadScene dataset; the test set consists of 20 images from RoadScene, 20 images from MSRS, 17 images from M3FD, and 21 images from the real-scene dataset CVC-14. The algorithm of the present invention was developed under the Pytorch 1.10.1 framework and trained on a single NVIDIA GTX3090 graphics card (with 24G of video memory). During training, we used the Adam optimizer to optimize the parameters of the model. In this process, the Batch size was set to 8. The learning rate decay strategy of "CosineAnnealing" was adopted, with an initial learning rate of 0.001, and a total of 1000 rounds of training were carried out.
[0105] Furthermore, the present invention was compared with IMF, IVFWSR, MURF, RFVIF, SemLA, SuperFusion, and UMF-CMGR in terms of the visual effect of the fusion results for misaligned infrared and visible light image pairs, as Figure 5 shown. As Figure 5 can be seen, the method proposed by the present invention can better correct the spatial misalignment and obtain high-fidelity fusion images with reduced artifacts and reduced offsets.
[0106] Six commonly used image quality evaluation metrics were selected to objectively evaluate the quality of the fusion results, including the correlation coefficient ( ), the gradient-based fusion performance metric ( ), the Chen-Blum metric ( ), the Chen-Varshney metric ( ), the sum of linear uncorrelations ( ), and the structural similarity ( ). Evaluated the linear correlation degree between the fused image and the source image to measure the similarity between the two. Evaluate the amount of edge information transmitted from the source image to the fused image. Measure the similarity of the fused image and the main features in the source image based on the human visual system. Then consider both the edge information in the fused image and human visual perception. Use the difference map to evaluate the correlation degree between the fused image and the source image. By comparing the fused image and the source image, quantify the information loss and distortion in the fused result image. Among these evaluation metrics, The lower the value of [[ ]] indicates the better the quality of the fused image, and the higher the values of the remaining metrics indicate the better the fusion quality.
[0107] Figure 6 It is a quantitative comparison graph of the fusion results of the existing method and the method of the present invention. The x-axis represents each method, and the y-axis represents the metric values. Figure 6 The y-axis in [[ ]] has no dimension and unit and is a constant. The up and down arrows after the metric indicate that the larger / smaller the metric value, the better the quality of the fused image. The upper and lower boundaries of each box represent the upper and lower quartiles respectively, the black line represents the median, and the red line represents the mean. From Figure 6 It can be seen that the method of the present invention achieves the best performance in 4 metrics and the sub-optimal performance in 2 metrics. However, in these 2 sub-optimal metrics, the span of the upper and lower boundaries of the method of the present invention is smaller, which indicates that our method is more robust than other methods.
[0108] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. An unaligned infrared and visible image fusion method based on information compensation and prediction correction, characterized in that: The method includes: Step 1: Obtain a training dataset for unaligned infrared and visible image fusion, where the fixed image is a visible image and the offset image is an infrared image; Step 2: Input the fixed image and the offset image into a feature encoder to obtain fixed features and offset features; Step 3: Input the fixed features and the offset features into a cross-modal information compensation module to convert these two single-modal features into full-modal fixed features and full-modal offset features respectively, for reducing the modal difference between the fixed features and the offset features; Step 4: Input the fixed features and the offset features with reduced modal difference into a dual deformation field prediction and correction module to correct the offset features and obtain corrected features; Step 5: Input the corrected features and the fixed features into a feature fusion module and reconstruct an infrared and visible fusion image.
2. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 1, wherein: In Step 1, each pair of infrared and visible images in the training dataset consists of an infrared image and a visible image; the size of each modal image is 256×256; The pre-aligned infrared and visible light image pair Apply the same offset to obtain another pair of aligned offset infrared and visible light image pairs , take as the unaligned infrared and visible light image pair, with the fixed image being the visible light image , and the offset image being the infrared image ; The preprocessing includes: randomly flipping, rotating and offsetting the data, and normalizing the processed images.
3. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 1, wherein: Step 2 includes: The fixed image is a visible light image , and the offset image is an infrared image . The visible light image and the infrared image are respectively input into the fixed feature encoder and the offset feature encoder to obtain the fixed feature and the offset feature ; The generation process of the fixed feature and the offset feature is expressed as: ; Fixed feature encoder and offset feature encoder are both composed of feature extraction blocks consisting of 4 3×3 convolutional layers and ReLU activation functions, which are connected in a densely connected manner, and finally composed of a Concat operation and a 1×1 convolutional layer.
4. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 1, wherein: In Step 3, the specific operations of the cross-modal information compensation module are as follows: The fixed image is a visible light image , and the offset image is an infrared image . The unaligned infrared image and the visible light image obtain features after CLIP and . Then, the broadcast mechanism is used to expand and into and to make them have the same size as and . The and , the fixed feature and the offset feature are processed through cross-modal compensation and a 3×3 convolutional layer to obtain the full-modal fixed feature and the full-modal offset feature : 。 5. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 1, wherein: Step 4 includes: Full-modal fixed features with reduced modal differences and full-modal offset features , and the visible-light image gradient and infrared image gradient of the image to be aligned are respectively used to generate deformation fields and ; The generation processes of the deformation fields and are expressed by the following formulas: ; Among them, DFPNet1 and DFPNet2 are U-Net-like networks, and DFPNet1 and DFPNet2 have the same structure but different parameters; The deformation field and Applied to infrared images On the rectified image, the gradient of the rectified image is compared with Gradient The pixels are spliced and global average pooling and global maximum pooling are performed on the channels and spatial positions respectively. The pooled results are added and sent to the Sigmoid activation function to obtain the correction weight values of all pixels on the channels and spatial positions. Through the broadcast mechanism, the weight values on the channels and the weight values on the spatial positions are formed into a matrix that can be multiplied by points, and after the dot multiplication operation of the matrix, the weight matrix for correcting and integrating the deformation field is obtained. and ; The corrected deformation field after adjustment and integration It is expressed as: ; Apply the correction deformation field and the offset feature After the Warp operation, the corrected feature is obtained .
6. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 1, wherein: In Step 5, the processing flow of the feature fusion module includes: The calibration features and the fixed features are processed by the Sigmoid activation function to obtain the weight maps and ; The weight maps and are concatenated with the calibration features and the fixed features and further extract features through a 3×3 convolutional layer; The global average pooling and global max pooling operations are jointly used for the features on each channel, and the results after pooling are added and then fed into a multi-layer perceptron; Finally, after being processed by the Sigmoid activation function, the weights and that can express the importance of each channel are obtained; The outputs after global max pooling of the infrared and visible light branches are added, and the added result is fed into a multi-layer perceptron, and then through the Sigmoid activation function, the weights of the infrared image and the visible light image at the same channel position are obtained; The weights are respectively added to the original weights and of the infrared and visible light branches, so as to obtain the final weights and : ; According to and with the calibration feature and the fixed feature , finally obtain the fused feature ; The fused feature is expressed as: ; Among them and are broadcasted by and respectively, and represents element-wise multiplication of matrices; The fused features are fed into a decoder composed of a 3×3 convolutional layer and a Tanh activation function to obtain the fused image .
7. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 4, characterized in that: Step 3 also includes: Defines the consistency loss to optimize the 3×3 convolutional layer. The consistency loss is expressed as: ; Among them, denotes the norm, is the fused feature, is the full-modal fixed feature, is the full-modal offset feature; and is obtained by learning full-modal feature representation; the specific operations are as follows: These two pairs of aligned infrared and visible light image pairs and are respectively input into the feature encoders and with shared parameters. The extracted features are sent to the feature fusion module in Step 5 for fusion, and finally the fused full-modal fusion features and .
8. The method for fusing misaligned infrared and visible images based on information compensation and prediction correction according to claim 6, characterized in that: Step 5 also includes: Use texture loss to preserve the texture details of infrared images and visible light images; texture loss The calculation formula is as follows: ; Among them, , , respectively represent the gradients of the fused image , the infrared image , and the visible light image . represents the per-pixel maximum value selection; Regarding the preservation of image content information, the content loss is defined to bring the fused image closer to the source image at the pixel level. The content loss is defined as follows: ; ; Among them, and are the saliency matrices of the infrared image and the visible light image respectively, and are the weight maps of the infrared image and the visible light image respectively.
9. An unaligned infrared and visible image fusion system based on information compensation and prediction correction, characterized in that, The system includes: a module for executing the method for unaligned infrared and visible image fusion based on information compensation and prediction correction as described in any one of claims 1 to 8.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for unaligned infrared and visible image fusion based on information compensation and prediction correction as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Infrared and visible light image fusion method and device
CN116757986A
Medical image fusion method based on double contrast learning and gradient channel attention mechanism
CN117788313A
Unregistered infrared visible image fusion method based on modal dictionary and feature matching
CN117934309A
Cross-modal pedestrian re-identification method based on superpixel patch mixing and information compensation
CN118447538A
System and a method for detecting computer-generated images
US20250078446A1
Cited By
Image distortion real-time correction method and system for industrial infrared camera
CN121837088A