Infrared and visible light image fusion method based on target perception

Through the target-perceptual area fusion sub-model and the space-frequency attention sub-model, the problem of insufficiently distinguishing the difference between the significant target area and the background area in the prior art is solved, and the high definition and rich details fusion effect of infrared and visible images are achieved.

CN120339094AActive Publication Date: 2025-07-18TIANJIN NORMAL UNIVERSITY

Patent Information

Application Number
CN202510811586.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing infrared and visible light image fusion methods fail to fully distinguish the differences between significant target areas and background areas, and ignore important structural information in the frequency domain, resulting in poor performance of fusion images in texture retention and structural reconstruction.

Method used

The target-aware area fusion sub-model is used to separate the significant target area and background area, combine the spatial-frequency attention sub-model and the fusion block sub-model. Through differentiated feature extraction and enhancement, the splicing and integration of regional information is achieved, and the fusion image of clear structure and significant targets is output.

Benefits of technology

Effectively extract significant target features in infrared images and preserve texture details in visible images, improving the clarity and detail retention capabilities of the fused image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339094A_ABST
    Figure CN120339094A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared and visible light image fusion method based on target perception, and the method comprises the following steps: S1, obtaining a training data set which comprises a plurality of input image pairs composed of infrared images and visible light images; s2, constructing an image fusion depth model; s3, training the image fusion depth model based on the training data set to obtain a target image fusion depth model; and S4, fusing the to-be-fused image pair by using the target image fusion depth model to obtain a fused image result. According to the invention, through the target sensing region fusion sub-model, the space-frequency attention sub-model and the fusion block sub-model, differential feature extraction and fusion are realized, and the definition, detail retention and downstream visual task performance of the fused image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing and deep learning, and particularly to an infrared and visible light image fusion method based on target perception. Background Art

[0002] In recent years, infrared and visible light image fusion technology has been widely applied in fields such as night video surveillance, remote sensing imaging, and target detection. Its purpose is to integrate complementary information obtained by different sensors, thereby generating a more informative fused image. Due to the limitations of single-modal sensors in perception ability, infrared and visible light image fusion can effectively make up for their respective deficiencies.

[0003] Existing methods are mainly divided into traditional image fusion methods and deep learning-based image fusion methods. Traditional methods mostly rely on mathematical modeling means such as sparse representation, subspace transformation, multi-scale transformation, and saliency analysis, and fuse source images through predefined activity metric indicators and fusion rules. However, due to the poor adaptability of these methods to scene changes, it is difficult to stably extract effective features in complex or dynamic environments, resulting in poor performance of fused images in terms of detail retention and structure preservation.

[0004] In recent years, deep learning-based image fusion methods have received extensive attention, mainly including fusion frameworks based on autoencoders, convolutional neural networks, and generative adversarial networks. These methods can learn more expressive image features driven by large-scale data, significantly improving the quality of image fusion. However, the above methods use convolutional neural networks for feature extraction, lacking the ability to model global dependencies and being difficult to capture long-distance feature correlations between source images. For this reason, some studies introduce the vision Transformer structure to achieve global modeling through multi-layer attention mechanisms, effectively enhancing the understanding ability of the overall semantics and spatial structure of images. For example, Tang et al. proposed an infrared and visible light image fusion method that combines dense convolution and a ViT encoder branch, which can extract local and global features simultaneously. Li et al. effectively model the differential information between source images by introducing a cross-attention module to extract irrelevant features. Yang et al. designed a semantic-aware fusion Transformer to model image semantic information through multiple ViT encoders, thereby improving the semantic expression ability of fused images.

[0005] In the process of implementing the present invention, the inventors found that there are at least the following disadvantages and deficiencies in the prior art:

[0006] For the image feature extraction stage, existing methods fail to fully distinguish the inherent differences between the significant target regions and the background regions in the source image. Usually, a unified feature extraction strategy is adopted, resulting in the difficulty of effectively highlighting the significant target information during the fusion process, which affects the expression ability of the fused image for key regions. For the fusion method, existing methods mainly focus on processing in the spatial domain, ignoring the important structural information and detailed features contained in the frequency domain, such as frequency domain attributes like amplitude and phase. This to some extent limits the performance of the fused image in texture preservation and structure reconstruction. Summary of the Invention

[0007] The present invention designs an infrared and visible light image fusion method based on target perception. The present invention uses a target perception region fusion sub-model to separate the significant target region and the background region, realizes differential feature extraction, adopts a spatial-frequency attention sub-model to enhance feature expression, introduces a fusion block sub-model to perform skip connection and fusion on enhanced features at different levels, realizes the splicing and integration of regional information, outputs the final fused image, and combines a hybrid fusion loss to improve the retention ability of effective information in the fused image. This image fusion method can generate a fused image result with a clear structure and significant targets.

[0008] To achieve the above object, an infrared and visible light image fusion method based on target perception proposed by the present invention includes the following steps:

[0009] Step S1, obtain a training data set, wherein the training data set includes a plurality of input image pairs composed of infrared images and visible light images, and the input image pairs include infrared images and visible light images ;

[0010] Step S2, construct an image fusion deep model, wherein the image fusion deep model includes a target perception region fusion sub-model, a spatial-frequency attention sub-model and a fusion block sub-model. The target perception region fusion sub-model includes a pre-trained binary image segmentation module, a region enhancement module and a parallel multi-scale convolution branch module, wherein: the pre-trained binary image segmentation module is used to receive an infrared image , and based on the infrared image generate a significant target mask and its anti-mask through a trained binary image segmentation network; the region enhancement module is used to use the mask and its anti-mask to perform differential enhancement processing on the significant target region and the background region of the infrared image and the visible light image to obtain a significant target region image pair and a background region image pair , where represents the salient target region, and represents the background region; the parallel multi-scale convolutional branch module is used to extract the multi-scale regional features of the salient target region and the background region based on the salient target region image pair

[0011] Step S3, training the image fusion depth model based on the training data set to obtain the target image fusion depth model;

[0012] Step S4, using the target image fusion depth model to fuse the to-be-fused image pair to obtain a fused image result.

[0013] Optionally, where:

[0014] The target-aware region fusion sub-model is used to separate the salient target region and the background region of the infrared image and the visible light image , and respectively extract the regional differentiation features of the salient target region and the background region:

[0015] The spatial-frequency attention sub-model is used to weight and enhance the regional differentiation features of the salient target region and the background region to obtain spatial-frequency enhanced features:

[0016] The fusion block sub-model is used to fuse the spatial-frequency enhanced features to obtain the final fused image.

[0017] Optionally, the region enhancement module uses the mask and its inverse mask to perform differential enhancement processing on the salient target region and the background region of the infrared image and the visible light image :

[0018]

[0019] Where is the salient target region image pair, is the background region image pair, represents element-wise multiplication, is the enhancement coefficient.

[0020] Optionally, the parallel multi-scale convolutional branch module includes a salient target region branch and a background region branch. Each branch includes three convolutional modules. Each convolutional module sequentially includes a convolutional layer, a batch normalization layer, and a PReLU activation function. The parallel multi-scale convolutional branch module is based on the salient target region image pair according to the following formula and the background region image pair Extract the multi-scale regional features of the significant target region and the background region, that is, the regional differentiation features :

[0021]

[0022]

[0023]

[0024] Among them, represents the regional feature output by the th convolutional module, , , represents the convolutional module, represents the max pooling operation, represents the channel dimension concatenation operation.

[0025] Optionally, the spatial-frequency attention sub-model includes a frequency integration double coordinate attention module and a spatial attention module, where:

[0026] The frequency integration double coordinate attention module is used to realize the interaction enhancement of infrared features and visible features in the frequency domain and generate frequency attention weights ;

[0027] The spatial attention module is used to use the frequency attention weights to enhance the regional differentiation features in the spatial domain respectively and generate spatial-frequency enhanced features , .

[0028] Optionally, the frequency integration double coordinate attention module includes a frequency integration module and a double coordinate attention module, where:

[0029] The frequency integration module is used to generate interaction features based on the regional differentiation features ;

[0030] The double coordinate attention module is used to generate frequency attention weights based on the interaction features .

[0031] Optionally, the frequency integration module generates interaction features based on the regional differentiation features according to the following formula :

[0032]

[0033] ​

[0034] Among them, represents the fast Fourier transform, represents the infrared amplitude information, represents the infrared phase information, represents the visible light amplitude information, represents the visible light phase information, is the regional differentiation feature, represents the phase information integration layer, represents the inverse fast Fourier transform, represents the significant target area, represents the background area;

[0035] The dual coordinate attention module generates frequency attention weights based on the interaction features according to the following formula : :

[0036]

[0037]

[0038] Among them, and respectively represent average pooling in the row and column directions, represents channel dimension concatenation, represents a combination layer of convolution, batch normalization and ELU activation function, represents a separation operation, represents a convolutional layer, represents the sigmoid function, represents element-wise multiplication.

[0039] Optionally, the spatial attention module enhances the regional differentiation features in the spatial domain respectively according to the following formula to generate spatial-frequency enhanced features :

[0040]

[0041]

[0042]

[0043] Among them, is the regional differentiation feature, represents the max pooling layer, represents the average pooling layer, represents channel dimension concatenation, represents a convolutional layer, represents the sigmoid function, Denotes element-wise multiplication.

[0044] Optionally, the fusion block sub-model sequentially includes fusion blocks to , each fusion block contains a convolutional module , and the fusion block sub-model fuses the spatio-frequency enhanced features using the following formula to obtain the final fused image :

[0045]

[0046]

[0047]

[0048]

[0049] where represents the spatio-frequency enhanced feature output by the spatio-frequency attention sub-model of the th layer, , represents the upsampling operation, represents the element-wise addition operation, represents the channel dimension concatenation.

[0050] The beneficial effects of the technical solution provided by the present invention are:

[0051] 1. The present invention can effectively extract the significant target features in the infrared image and retain the texture details in the visible light image, achieving an image fusion effect with high clarity and rich details.

[0052] 2. The present invention utilizes the target-aware region fusion sub-model, spatio-frequency attention sub-model, and fusion block sub-model introduced in the deep neural network to effectively achieve regional differential processing and hierarchical feature integration, improving the expression ability of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of a method for fusing infrared and visible light images based on target awareness according to an embodiment of the present invention;

[0054] Figure 2 is a schematic diagram of the region division and feature extraction structure according to an embodiment of the present invention;

[0055] Figure 3 is a schematic diagram of the spatio-frequency attention sub-model according to an embodiment of the present invention;

[0056] Figure 4Schematic diagram of the structure of the frequency integrated double coordinate attention module according to an embodiment of the present invention;

[0057] Figure 5 Schematic diagram of the comparison results of the average gradient, differential correlation sum, and standard deviation of different infrared and visible light image fusion methods according to an embodiment of the present invention. Detailed implementation manners

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the target perception infrared and visible light image fusion network method of the present invention will be further described in detail below in conjunction with the specific implementation manners and with reference to the accompanying drawings. It should be understood that these descriptions are only for illustrative purposes and are not used to limit the scope of the present invention; in the following descriptions, well-known structures and technologies are not specifically described to avoid unnecessarily confusing the content of the present invention.

[0059] Figure 1 Is a flowchart according to an embodiment of the present invention. Below, Figure 1 is taken as an example to illustrate the specific implementation process of the present invention. As Figure 1 shown, a target perception-based infrared and visible light image fusion method proposed by the present invention includes the following steps:

[0060] Step S1, obtaining a training data set, where the training data set includes a plurality of input image pairs composed of infrared images and visible light images, and the input image pairs include infrared images and visible light images .

[0061] Step S2, constructing an image fusion depth model;

[0062] In an embodiment of the present invention, the image fusion depth model includes a target perception region fusion sub-model, a spatial-frequency attention sub-model, and a fusion block sub-model.

[0063] Among them, the target perception region fusion sub-model is used to separate the significant target regions and background regions of the infrared image and the visible light image , and respectively extract the region differentiation features of the significant target regions and background regions.

[0064] Furthermore, the target perception region fusion sub-model sequentially includes a pre-trained binary image segmentation module, a region enhancement module, and a parallel multi-scale convolution branch module. Among them, the pre-trained binary image segmentation module is used to receive the infrared image , and based on the infrared image , generate a significant target mask and its inverse mask , for subsequent region enhancement operations, where the significant target mask is used to enhance the target region in the image, and the anti-mask is used to enhance the background region in the image. The enhancement operation is implemented based on the element-wise multiplication of the mask and the enhancement coefficient. Among them, the region enhancement module is used to utilize the mask and its anti-mask to perform differential enhancement processing on the significant target region and the background region of the infrared image and the visible light image to obtain the significant target region image pair and the background region image pair . The enhancement process is as follows:

[0065]

[0066] Among them, is the significant target region image pair, is the background region image pair, represents element-wise multiplication, is the enhancement coefficient.

[0067] Furthermore, the parallel multi-scale convolution branch module includes a significant target region branch and a background region branch, which are respectively used to extract the multi-scale region features of the significant target region and the background region based on the significant target region image pair and the background region image pair . The two branches have the same structure and are each composed of three convolution modules to obtain the multi-scale region features. Each sequentially includes a convolutional layer, a batch normalization layer BN, and a PReLU activation function. Among them, the input of the first convolution module is the region image pair , and the output is denoted as . The input of the second convolution module is the maximum pooling result of the concatenation of the region image pair and the output of the first convolution module , and the output is denoted as . The input of the third convolution module is the maximum pooling result of the concatenation of the outputs of the first and second convolution modules , and the output is denoted as . Among them, , and the max pooling process is used to extract multi-scale information. The region features extracted at each scale can be expressed as follows:

[0068]

[0069]

[0070]

[0071] Among them, represents the th convolutional module output regional feature, , , represents the max pooling operation, represents the channel dimension concatenation operation, such as Figure 2 shown.

[0072] The above multi-scale regional features will be used as the input of the subsequent spatial-frequency attention sub-model to further enhance the expression ability of the source image regional features.

[0073] When using the target-aware region fusion sub-model to extract regional features:

[0074] First, input the infrared image into the pre-trained binary image segmentation module to generate the saliency target mask and its anti-mask ;

[0075] Then, the region enhancement module uses the mask and its anti-mask and the enhancement coefficient to perform differential enhancement processing on the saliency target region and the background region of the infrared image and the visible light image to generate the saliency target region image pair and the background region image pair ;

[0076] Then, input the above saliency target region image pair and background region image pair into the saliency target region branch and background region branch of the parallel multi-scale convolutional branch module respectively, and extract regional features at different scales , for the weighted enhancement processing of the subsequent spatial-frequency attention sub-model, where the regional feature is expressed as:

[0077]

[0078] Among them, , , represents the multi-scale convolutional module.

[0079] Among them, the spatial-frequency attention sub-model is used to perform joint attention weighting on the regional differentiation features of the significant target region and the background region output by the target perception region fusion sub-model in the spatial domain and the frequency domain, so as to obtain spatial-frequency enhanced features, thereby enhancing the representation ability of the target and background features.

[0080] Further, the spatial-frequency attention sub-model sequentially includes a frequency integration double coordinate attention module (FICA) and a spatial attention module, as Figure 3 shown.

[0081] The frequency integration double coordinate attention module is used to realize the interactive enhancement of infrared features and visible light features in the frequency domain and generate frequency attention weights The frequency integration double coordinate attention module sequentially includes a frequency integration module and a double coordinate attention module, as Figure 4 shown.

[0082] The frequency integration module decomposes the input regional differentiation features into infrared amplitude information and infrared phase information, as well as visible light amplitude information and visible light phase information through fast Fourier transform (FFT). The infrared phase information and the visible light phase information are fused through a phase information integration layer to obtain fused phase information, and then inverse Fourier transform (IFFT) is respectively performed based on the infrared amplitude information, the visible light amplitude information and the fused phase information to generate interactive features . The double coordinate attention module performs average pooling on the interactive features along the row and column directions respectively, and after splicing, it is reduced in dimension through convolution, batch normalization and ELU activation, and then separated into attention vectors in the horizontal and vertical directions. Finally, frequency attention weights are generated through two convolutional layers and a sigmoid activation function .

[0083] Among them, the spatial attention module is used to enhance the regional differentiation features in the spatial domain by using the frequency attention weights Specifically, the spatial attention module performs maximum pooling and average pooling splicing on the input regional differentiation features respectively, and generates spatial enhanced features through convolution and sigmoid functions .

[0084] The specific process of the above processing can be expressed as:

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] Among them, represents the phase information integration layer, represents the combination layer of convolution, batch normalization, and ELU activation function, and respectively represent average pooling in the row and column directions, represents channel dimension concatenation, represents the reshape operation, which is used to match the feature dimensions, represents the split operation, which is used to split a vector into two or more sub-vectors, and respectively represent the attention features in the row and column directions, represents the sigmoid function, represents the convolutional layer, represents the max pooling layer, represents the average pooling layer.

[0092] Furthermore, the spatial attention module multiplies the spatial enhanced feature with the frequency attention weights and respectively, and performs concatenation in the channel dimension to obtain the spatial-frequency enhanced feature . The specific process can be expressed as:

[0093]

[0094] Among them, represents channel dimension concatenation, represents element-wise multiplication, is the spatial-frequency enhanced feature.

[0095] In summary, when using the spatial-frequency attention sub-model for feature enhancement:

[0096] First, input the infrared region feature and the visible light region feature into the spatial attention module and the frequency integration double-coordinate attention module respectively;

[0097] Then, use the spatial attention module to obtain the spatial enhanced feature of the infrared and visible light region features;

[0098] Then, the frequency integration module of the frequency integrated dual - coordinate attention module performs fast Fourier transform (FFT) on the infrared and visible - light region features to obtain infrared amplitude information, infrared phase information, visible - light amplitude information, and visible - light phase information. The phase information integration layer integrates the phase information and performs inverse transform (IFFT) to obtain the interaction features. , where, for the target region features inverse Fourier transform is performed using the infrared amplitude information, and for the background region features inverse Fourier transform is performed using the visible - light amplitude information;

[0099] Then, the dual - coordinate attention module of the frequency integrated dual - coordinate attention module calculates the frequency attention weights based on the interaction features using the dual - coordinate attention mechanism. , where the dual - coordinate attention module first performs average pooling on its input features along the height direction and width direction respectively to obtain statistical features in two directions. The obtained features are concatenated and passed through a convolutional layer, a batch normalization layer, and an ELU activation function to form a joint feature, and then through a separation operation to obtain attention vectors in two directions. Finally, direction attention maps are generated through two convolutional layers and a sigmoid activation function, and the frequency attention weights are generated by combining them in the form of matrix multiplication. ;

[0100] Then, the spatial attention module multiplies the spatial enhancement features of infrared and visible light with the frequency attention weights and respectively, and concatenates them to generate the spatial - frequency enhancement features, expressed as:

[0101]

[0102] where, represents concatenation in the channel dimension, represents element - wise multiplication, is the spatial - frequency enhancement feature.

[0103] Among them, the fusion block sub - model is used to fuse the spatial - frequency enhancement features output by the spatial - frequency attention sub - model to obtain the final fused image. The fusion block sub - model sequentially includes fusion blocks to , and each fusion block contains a convolutional module to integrate the input features through the convolutional module .

[0104] Among them, the fusion process is as follows: The first fusion block receives the third layer The spatial-frequency enhanced features output by the spatial-frequency attention sub-model , and generate the first-layer intermediate fusion result , the upsampled first-layer intermediate fusion result and the second layer The spatial-frequency enhanced features output by the spatial-frequency attention sub-model After concatenation, they are sent into the second fusion block , and then generate the second-layer intermediate fusion result , the upsampled second-layer intermediate fusion result and the first layer The spatial-frequency enhanced features output by the spatial-frequency attention sub-model After concatenation, they are sent into the third fusion block , and then generate the third-layer intermediate fusion result , the fourth fusion block receives the third-layer intermediate fusion result and generates the final fused image . The specific process can be expressed as:

[0105]

[0106]

[0107]

[0108]

[0109] wherein, represents the element-wise addition operation, which is used to splice the target region and the background region into a complete image, represents the upsampling operation.

[0110] In summary, when using the fusion block sub-model to fuse the enhanced features:

[0111] First, the features obtained by element-wise addition of the spatial-frequency enhanced features of the target region and the background region at the same scale are input into the corresponding fusion block. For the fusion blocks and , at the same time, the output features of the previous fusion block after upsampling are also input into the corresponding fusion block. Then, the two features input into the corresponding fusion block are spliced and integrated. For the fusion block , only the output features of the previous fusion block are input. Finally, the output of the fourth fusion block is the final fused image. The overall process of the fusion block sub-model fusing the spatial-frequency enhanced features can be expressed as:

[0112] Among them, represents the fusion block sub-model, represents the target region spatial-frequency enhanced feature, represents the background region spatial-frequency enhanced feature.

[0113] Step S3: Train the image fusion depth model based on the training data set to obtain the target image fusion depth model;

[0114] In an embodiment of the present invention, the training process uses a hybrid fusion loss function to optimize the model parameters to balance texture details, statistical correlation, and pixel intensity retention.

[0115] Among them, the hybrid fusion loss function is expressed as:

[0116]

[0117] Among them, represents the gradient loss function, represents the correlation loss function, represents the intensity loss function, , , are the corresponding weight coefficients.

[0118] Furthermore, the gradient loss function can be expressed as:

[0119]

[0120] Among them, and respectively represent the input visible light image and infrared image, represents the fused image output by the image fusion depth model, represents the Sobel gradient operator for calculating the image gradient information, represents the matrix absolute value calculation, represents the matrix 1-norm calculation, represents the pixel-level maximum value selection, and respectively represent the length and width of the image.

[0121] Furthermore, the correlation loss function can be expressed as:

[0122]

[0123] Among them,

[0124]

[0125] Among them, and respectively represent the input visible light image and infrared image, represents the fused image output by the image fusion depth model, represents the calculation of the correlation coefficient, and respectively represent the weight coefficients for the visible light image and infrared image, represents the calculation of covariance, represents the calculation of variance, represents a very small constant used to ensure that the denominator of the fraction is not zero.

[0126] Furthermore, the intensity loss function can be expressed as:

[0127]

[0128] Among them

[0129]

[0130] Among them, and respectively represent the input visible light image and infrared image, represents the fused image output by the image fusion depth model, represents the pixel mask used to compare the pixel point values of the source images, represents the calculation of the matrix 1-norm, and respectively represent the length and width of the image.

[0131] Step S4: Use the target image fusion depth model to fuse the to-be-fused image pair to obtain the fused image result.

[0132] Figure 5 lists the comparison results of the average gradient (AG), sum of contrast differences (SCD), and standard deviation (SD) obtained by different methods in the infrared and visible light image fusion task. The comparison algorithms include: Liu's method and Li's method. The larger the average gradient (AG), the richer the details of the fused image; the larger the sum of contrast differences (SCD), the more complementary information the fused image retains; the larger the standard deviation (SD), the higher the overall contrast of the fused image. By Figure 5It can be seen that, compared with the methods of Liu and Li, the method of the present invention has achieved significant improvements in the AG, SCD, and SD metrics. Specifically, in terms of the average gradient (AG) metric, the method of the present invention has increased by approximately 44% and 26% compared with the methods of Liu and Li, respectively, indicating that the detailed information in the fused image is more fully retained; in terms of the sum of contrast differences (SCD) metric, the method of the present invention is also much higher than the comparative methods, indicating that the fused result has a stronger ability in extracting and integrating complementary information of the source images; in terms of the standard deviation (SD) metric, the method of the present invention also obtains the highest value, indicating that the contrast of the fused image is more distinct and the visual effect is better. The main reason is that Liu's method focuses on multi-modal collaborative learning but pays insufficient attention to the differences between the salient target regions and the background regions; although Li's method introduces a cross-attention mechanism, it does not fully combine the multi-scale feature expressions in the spatial domain and the frequency domain, restricting the further improvement of the fusion performance. In contrast, the present invention designs a target-aware region fusion sub-model, a spatial-frequency attention sub-model, and a fusion block sub-model to extract and enhance features in the salient target regions and the background regions respectively, and combines the spatial domain and frequency domain attention mechanisms, significantly improving the overall performance of the fused image in terms of detail retention, complementary information integration, and contrast, and finally obtaining an excellent infrared and visible light image fusion effect.

[0133] In the embodiments of the present invention, except for those with special descriptions for the models of each device, the models of other devices are not limited, as long as the devices can perform the above functions.

[0134] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0135] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An infrared and visible light image fusion method based on target perception, characterized in that The method includes the following steps: Step S1, obtain a training data set, where the training data set includes a plurality of input image pairs composed of infrared images and visible light images, and the input image pairs include infrared images and visible light images ; Step S2, construct an image fusion depth model. The image fusion depth model includes a target perception region fusion sub-model, a spatial-frequency attention sub-model, and a fusion block sub-model. The target perception region fusion sub-model includes a pre-trained binary image segmentation module, a region enhancement module, and a parallel multi-scale convolution branch module, where: The pre-trained binary image segmentation module is used to receive an infrared image , and based on the infrared image , generate a significant target mask and its anti-mask through a trained binary image segmentation network; The region enhancement module is used to use the mask and its anti-mask to perform differential enhancement processing on the significant target region and the background region of the infrared image and the visible light image to obtain a significant target region image pair and a background region image pair , where represents the significant target region, represents the background region; The parallel multi-scale convolution branch module is used to extract multi-scale region features of the significant target region and the background region based on the significant target region image pair and the background region image pair . Step S3: Train the image fusion depth model based on the training data set to obtain a target image fusion depth model; Step S4: Use the target image fusion depth model to fuse the to-be-fused image pair to obtain a fused image result.

2. The method according to claim 1, wherein Wherein: The target perception area fusion sub-model is used to separate the infrared image and the visible light image into the salient target area and the background area, and respectively extract the regional differentiation features of the salient target area and the background area: The spatial-frequency attention sub-model is used to weight and enhance the regional differentiation features between the significant target region and the background region to obtain spatial-frequency enhanced features: The fusion block sub-model is used to fuse the spatial-frequency enhanced features to obtain a final fused image.

3. The method according to claim 2, wherein The described region enhancement module uses a mask based on the following formula and its inverse mask to perform differential enhancement processing on the infrared image and the visible light image for the differential enhancement of the significant target region and the background region: ; Among them, is the significant target region image pair, is the background region image pair, represents element-wise multiplication, is the enhancement coefficient.

4. The method according to claim 2 or 3, characterized in that, The parallel multi-scale convolution branch module includes a salient object region branch and a background region branch. Each branch includes three convolution modules, and each convolution module sequentially includes a convolutional layer, a batch normalization layer, and a PReLU activation function. The parallel multi-scale convolution branch module extracts multi-scale regional features of the salient object region and the background region, that is, regional differentiation features, based on the salient object region image pair and the background region image pair according to the following formula : ; ; ; Among them, represents the regional feature output by the th convolutional module, , , represents the convolutional module, represents the max pooling operation, represents the channel dimension concatenation operation.

5. The method according to claim 2, characterized in that, The spatial-frequency attention sub-model includes a frequency integration double coordinate attention module and a spatial attention module, wherein: The frequency integrated dual-coordinate attention module is used to achieve the interactive enhancement of infrared features and visible features in the frequency domain and generate frequency attention weights ; The spatial attention module is used to utilize the frequency attention weights to enhance the region-differentiated features in the spatial domain respectively, generating spatially-frequency enhanced features , .

6. The method according to claim 5, wherein The frequency integration double coordinate attention module includes a frequency integration module and a double coordinate attention module, wherein: The frequency integration module is used to generate interaction features based on the region differentiation features ; The dual - coordinate attention module is used to generate frequency attention weights based on the interaction features and .

7. The method according to claim 6, characterized in that, The frequency integration module generates interaction features based on the regional differentiation features according to the following formula : ; ; Among them, represents the fast Fourier transform, represents the infrared amplitude information, represents the infrared phase information, represents the visible light amplitude information, represents the visible light phase information, is the regional differentiation feature, represents the phase information integration layer, represents the inverse fast Fourier transform, represents the significant target area, represents the background area; The double coordinate attention module generates frequency attention weights based on the interaction features according to the following formula :​ ; ; Among them, and respectively represent average pooling in the row and column directions, represents concatenation in the channel dimension, represents a combined layer of convolution, batch normalization, and ELU activation function, represents a splitting operation, represents a convolutional layer, represents the sigmoid function, represents element-wise multiplication.

8. The method according to claim 5, characterized in that, The spatial attention module enhances the region-differentiated features in the spatial domain according to the following formula to generate spatial-frequency enhanced features :[[]]END]] ; ; ; Among them, is the regional differentiation feature, represents the max pooling layer, represents the average pooling layer, represents the channel dimension concatenation, represents the convolutional layer, represents the sigmoid function, represents the element-wise multiplication.

9. The method according to claim 2, characterized in that The fusion block sub-model sequentially includes fusion blocks to , and each fusion block contains a convolutional module . The fusion block sub-model fuses the spatial-frequency enhanced features using the following formula to obtain the final fused image : ; ; ; ; Among them, represents the spatial-frequency enhanced feature output by the -th layer spatial-frequency attention sub-model, represents the upsampling operation, represents the element-wise addition operation, represents the channel dimension concatenation.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on cross mode enhancement and multi-attention fusion strategy

    CN118096554A

  • Infrared and visible light fusion method based on multi-scale feature interaction enhancement

    CN119091269A

  • Visible light and infrared image fusion method based on space-frequency domain characteristics

    CN119963958A

  • Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network

    WO2024174488A1

Cited By

  • Infrared and visible light image fusion method based on degradation perception and frequency integration

    CN121095079A

  • Infrared and visible image fusion method based on degradation perception and frequency integration

    CN121095079B