A method of infrared and visible light image fusion based on target perception
Through the target-perceptual area fusion sub-model and the space-frequency attention sub-model, the problem of insufficient identification of significant target area and background area differences in infrared and visible light images is solved, and the image fusion effect with high definition and rich details is achieved.
Patent Information
- Application Number
- CN202510811586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing infrared and visible image fusion methods fail to effectively distinguish the differences between significant target areas and background areas, and ignore important structural information in the frequency domain, resulting in poor performance of fusion images in texture retention and structural reconstruction.
The target-aware area fusion sub-model is used to separate the significant target area and the background area, combine the spatial-frequency attention sub-model and the fusion block sub-model. Through differentiated feature extraction and enhancement, the splicing and integration of regional information is achieved, and the fusion image of clear structure and significant targets is output.
The clarity and detail retention ability of the fused image are improved, significant target information is effectively highlighted, and the texture and structure reconstruction effect is significantly improved.
Smart Images

Figure CN120339094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and deep learning, and in particular to a method for fusing infrared and visible light images based on target perception. Background Art
[0002] In recent years, infrared and visible light image fusion technology has been widely used in fields such as nighttime video surveillance, remote sensing imaging, and target detection. Its purpose is to integrate complementary information acquired by different sensors to produce a more informative fused image. Because single-modality sensors have limitations in their perception capabilities, infrared and visible light image fusion can effectively compensate for their respective shortcomings.
[0003] Existing image fusion methods are primarily categorized as traditional and deep learning-based. Traditional methods rely on mathematical modeling techniques such as sparse representation, subspace transformation, multiscale transformation, and saliency analysis, fusing source images using predefined activity metrics and fusion rules. However, these methods are poorly adaptable to scene changes and struggle to stably extract effective features in complex or dynamic environments, resulting in poor detail and structure preservation in the fused image.
[0004] In recent years, deep learning-based image fusion methods have garnered widespread attention, primarily including fusion frameworks based on autoencoders, convolutional neural networks, and generative adversarial networks. These methods are able to learn more expressive image features driven by large-scale data, significantly improving the quality of image fusion. However, these methods, which utilize convolutional neural networks for feature extraction, lack the ability to model global dependencies and struggle to capture long-range feature correlations between source images. To address this, some studies have introduced the visual Transformer architecture, which achieves global modeling through multi-layer attention mechanisms, effectively enhancing the understanding of the overall semantics and spatial structure of the image. For example, Tang et al. proposed a method for fusion of infrared and visible light images that combines dense convolution with a ViT encoder branch to simultaneously extract local and global features. Li et al. introduced a cross-attention module to extract non-correlated features, effectively modeling the differences between source images. Yang et al. designed a semantically aware fusion Transformer that models image semantic information through multiple ViT encoders, thereby improving the semantic expressiveness of the fused image.
[0005] In the process of realizing the present invention, the inventors found that the prior art has at least the following shortcomings and deficiencies:
[0006] In the image feature extraction stage, existing methods fail to fully distinguish the inherent differences between the salient target area and the background area in the source image. They usually adopt a unified feature extraction strategy, which makes it difficult to effectively highlight the salient target information during the fusion process, affecting the fused image's ability to express key areas. In terms of fusion methods, existing methods mainly focus on processing in the spatial domain, ignoring the important structural information and detailed features contained in the frequency domain, such as frequency domain attributes such as amplitude and phase. This, to a certain extent, limits the performance of the fused image in texture preservation and structure reconstruction. Summary of the Invention
[0007] This paper designs a method for fusion of infrared and visible light images based on target perception. It utilizes a target perception region fusion sub-model to separate salient target regions from background regions, achieving differentiated feature extraction. It employs a spatial-frequency attention sub-model to enhance feature representation. It introduces a fusion block sub-model to perform jump connections and fusion of enhanced features at different levels, achieving the splicing and integration of regional information. The final fused image is then outputted using a hybrid fusion loss to enhance the retention of effective information in the fused image. This image fusion method can produce fused images with clear structure and salient targets.
[0008] To achieve the above-mentioned purpose, the present invention proposes a method for fusion of infrared and visible light images based on target perception, which includes the following steps:
[0009] Step S1: obtaining a training data set, wherein the training data set includes a plurality of input image pairs consisting of infrared images and visible light images, wherein the input image pairs include infrared images and visible light images ;
[0010] Step S2, constructing an image fusion depth model, wherein the image fusion depth model includes a target perception region fusion sub-model, a space-frequency attention sub-model and a fusion block sub-model, and the target perception region fusion sub-model includes a pre-trained binary image segmentation module, a region enhancement module and a parallel multi-scale convolution branch module, wherein: the pre-trained binary image segmentation module is used to receive infrared images , and based on the infrared image Generate salient object masks through the trained binary image segmentation network and its inverse mask ; The region enhancement module is used to use the mask and its inverse mask For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area to obtain the salient target area image pair. and background area image pair ,in, Indicates the salient target area, Represents the background area; the parallel multi-scale convolution branch module is used to and background area image pair Extract multi-scale regional features of salient target areas and background areas;
[0011] Step S3, training the image fusion depth model based on the training data set to obtain a target image fusion depth model;
[0012] Step S4: fuse the to-be-fused image pair using the target image fusion depth model to obtain a fused image result.
[0013] Optionally, where:
[0014] The target perception area fusion sub-model is used to separate the infrared image and visible light images The salient target area and background area are extracted, and regional differentiation features of the salient target area and background area are extracted respectively:
[0015] The space-frequency attention sub-model is used to perform weighted enhancement on the regional differentiation features of the salient target area and the background area to obtain the space-frequency enhancement features:
[0016] The fusion block sub-model is used to fuse the space-frequency enhancement features to obtain a final fused image.
[0017] Optionally, the region enhancement module utilizes a mask based on the following formula: and its inverse mask For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area:
[0018]
[0019] in, is a pair of salient target region images, is a background region image pair, represents element-wise multiplication, is the enhancement factor.
[0020] Optionally, the parallel multi-scale convolution branch module includes a salient target region branch and a background region branch, each branch includes three convolution modules, each convolution module includes a convolution layer, a batch normalization layer and a PReLU activation function in sequence, and the parallel multi-scale convolution branch module is based on the salient target region image pair according to the following formula: and background area image pair Extract multi-scale regional features of salient target areas and background areas, namely regional differentiation features :
[0021]
[0022]
[0023]
[0024] in, Indicates the The regional features output by the convolution module, , , represents the convolution module, represents the maximum pooling operation, Represents a channel dimension concatenation operation.
[0025] Optionally, the space-frequency attention sub-model includes a frequency-integrated dual-coordinate attention module and a spatial attention module, wherein:
[0026] The frequency integrated dual-coordinate attention module is used to realize the interactive enhancement of infrared features and visible features in the frequency domain and generate frequency attention weights ;
[0027] The spatial attention module is used to utilize the frequency attention weight Enhance regional differentiation features in the spatial domain to generate space-frequency enhancement features , .
[0028] Optionally, the frequency integration dual-coordinate attention module includes a frequency integration module and a dual-coordinate attention module, wherein:
[0029] The frequency integration module is used to generate interactive features based on the regional differentiation features ;
[0030] The dual-coordinate attention module is used to Generate frequency attention weights .
[0031] Optionally, the frequency integration module generates an interactive feature based on the regional differentiation feature according to the following formula: :
[0032]
[0033]
[0034] in, represents the fast Fourier transform, Indicates infrared amplitude information, Indicates infrared phase information, Represents visible light amplitude information, Represents visible light phase information, For regional differentiation characteristics, represents the phase information integration layer, represents the inverse fast Fourier transform, Indicates the salient target area, Represents the background area;
[0035] The dual-coordinate attention module is based on the interaction feature according to the following formula Generate frequency attention weights :
[0036]
[0037]
[0038] in, and Represents row and column average pooling, respectively. represents channel dimension splicing, represents the combination layer of convolution, batch normalization and ELU activation function, Indicates a separation operation. represents the convolutional layer, represents the sigmoid function, Represents element-wise multiplication.
[0039] Optionally, the spatial attention module enhances the regional differentiation features in the spatial domain according to the following formula to generate space-frequency enhancement features :
[0040]
[0041]
[0042]
[0043] in, For regional differentiation characteristics, represents the maximum pooling layer, represents the average pooling layer, represents channel dimension splicing, represents the convolutional layer, represents the sigmoid function, Represents element-wise multiplication.
[0044] Optionally, the fusion block sub-model includes fusion blocks to , each fusion block contains a convolution module The fusion block sub-model uses the following formula to fuse the spatial-frequency enhancement features to obtain the final fused image :
[0045]
[0046]
[0047]
[0048]
[0049] in, Indicates the The spatial-frequency enhanced features output by the layer spatial-frequency attention sub-model, , represents the upsampling operation, represents the element-by-element addition operation, Indicates channel dimension splicing.
[0050] The beneficial effects of the technical solution provided by the present invention are:
[0051] 1. The present invention can effectively extract salient target features in infrared images and preserve texture details in visible light images, achieving an image fusion effect with high clarity and rich details.
[0052] 2. This paper utilizes the target perception region fusion sub-model, spatial-frequency attention sub-model, and fusion block sub-model introduced in deep neural networks to effectively implement regional differentiation processing and hierarchical feature integration, thereby improving the expressive power of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flow chart of a method for fusion of infrared and visible light images based on target perception according to an embodiment of the present invention;
[0054] Figure 2 is a schematic diagram of a region division and feature extraction structure according to an embodiment of the present invention;
[0055] Figure 3 is a schematic diagram of a space-frequency attention sub-model according to one embodiment of the present invention;
[0056] Figure 4is a schematic structural diagram of a frequency-integrated dual-coordinate attention module according to an embodiment of the present invention;
[0057] Figure 5 3 is a schematic diagram of comparison results of average gradient, difference correlation sum, and standard deviation of different infrared and visible light image fusion methods according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] To further clarify the objectives, technical solutions, and advantages of the present invention, the target-aware infrared-visible light image fusion network method of the present invention is further described in detail below, in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely illustrative and are not intended to limit the scope of the present invention. In the following description, well-known structures and technologies are not specifically described to avoid unnecessary obfuscation of the present invention.
[0059] Figure 1 This is a flow chart according to an embodiment of the present invention. Figure 1 Take the example to illustrate the specific implementation process of the present invention. Figure 1 As shown, the present invention proposes a method for fusion of infrared and visible light images based on target perception, which includes the following steps:
[0060] Step S1: obtaining a training data set, wherein the training data set includes a plurality of input image pairs consisting of infrared images and visible light images, wherein the input image pairs include infrared images and visible light images .
[0061] Step S2, constructing an image fusion depth model;
[0062] In one embodiment of the present invention, the image fusion depth model includes a target perception area fusion sub-model, a space-frequency attention sub-model and a fusion block sub-model.
[0063] The target perception area fusion sub-model is used to separate the infrared image and visible light images The salient target area and the background area are separated, and regional differential features of the salient target area and the background area are respectively extracted.
[0064] Furthermore, the target perception region fusion sub-model includes a pre-trained binary image segmentation module, a region enhancement module and a parallel multi-scale convolution branch module. , and based on the infrared image Generate salient object masks through the trained binary image segmentation network and its inverse mask , used for subsequent region enhancement operations, wherein the salient target mask Used to enhance the target area in the image, anti-mask It is used to enhance the background area in the image. The enhancement operation is based on the element-by-element multiplication of the mask and the enhancement coefficient. and its inverse mask For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area to obtain the salient target area image pair. and background area image pair , the enhancement process is as follows:
[0065]
[0066] in, is a pair of salient target region images, is a background region image pair, represents element-wise multiplication, is the enhancement factor.
[0067] Furthermore, the parallel multi-scale convolution branch module includes a salient target region branch and a background region branch, which are respectively used to and background area image pair Extract multi-scale regional features of salient target areas and background areas; the two branches have the same structure, each consisting of three convolutional modules Composition, used to obtain multi-scale regional features, each It includes convolutional layer, batch normalization layer BN and PReLU activation function in sequence. Among them, the first convolution module The input is a region image pair , the output is expressed as , the second convolutional module The input is a region image pair and the first convolutional module The maximum pooling result of the cascade of the output is expressed as , the third convolutional module The input to the first and second convolution modules The maximum pooling result of the cascade of the output is expressed as ,in, , the maximum pooling process is used to extract multi-scale information. The regional features extracted at each scale can be expressed as follows:
[0068]
[0069]
[0070]
[0071] in, Indicates the Convolutional modules The output regional features, , , represents the maximum pooling operation, Represents channel dimension splicing operation, such as Figure 2 shown.
[0072] The above multi-scale regional features will serve as the input of the subsequent space-frequency attention sub-model to further enhance the expressive power of the source image regional features.
[0073] When extracting regional features using the target perception region fusion sub-model:
[0074] First, the infrared image Input to the pre-trained binary image segmentation module to generate a salient object mask and its inverse mask ;
[0075] Then, the region enhancement module uses the mask and its inverse mask and enhancement factor For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area to generate the salient target area image pair and background area image pair ;
[0076] Then, the above-mentioned salient target region image pair and background region image pair are respectively input into the salient target region branch and background region branch of the parallel multi-scale convolution branch module, and the regional features at different scales are extracted through three convolution modules respectively. , used for weighted enhancement processing of the subsequent space-frequency attention sub-model, where the regional features are expressed as:
[0077]
[0078] in, , , Represents a multi-scale convolutional module.
[0079] Among them, the space-frequency attention sub-model is used to perform joint attention weighting on the regional differentiation features of the salient target area and the background area output by the target perception area fusion sub-model in the spatial domain and frequency domain to obtain space-frequency enhancement features to enhance the representation ability of the target and background features.
[0080] Furthermore, the space-frequency attention sub-model includes a frequency integrated dual-coordinate attention module (FICA) and a spatial attention module in sequence, such as Figure 3 shown.
[0081] The frequency integrated dual-coordinate attention module is used to realize the interactive enhancement of infrared features and visible light features in the frequency domain and generate frequency attention weights The frequency integration dual-coordinate attention module includes a frequency integration module and a dual-coordinate attention module in sequence. Figure 4 shown.
[0082] The frequency integration module transforms the input regional differentiation features into Decompose into infrared amplitude information and infrared phase information, as well as visible light amplitude information and visible light phase information, fuse the infrared phase information and visible light phase information through the phase information integration layer to obtain fused phase information, and then perform inverse Fourier transform (IFFT) based on the infrared amplitude information and visible light amplitude information and the fused phase information to generate interactive features The dual-coordinate attention module is used to analyze the interaction features. Average pooling is performed along the row and column directions respectively, and after splicing, convolution, batch normalization and ELU activation are performed to reduce the dimension, and then separated into attention vectors in the horizontal and vertical directions. Finally, two convolution layers and sigmoid activation functions are used to generate frequency attention weights. .
[0083] Among them, the spatial attention module is used to utilize the frequency attention weight The regional differentiation features are enhanced in the spatial domain. Specifically, the spatial attention module enhances the regional differentiation features of the input After performing maximum pooling and average pooling splicing respectively, spatial enhancement features are generated through convolution and sigmoid function .
[0084] The specific process of the above processing can be expressed as:
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091] in, represents the phase information integration layer, represents the combination layer of convolution, batch normalization and ELU activation function, and Represents row and column average pooling respectively, represents channel dimension splicing, Reshape is a deformation operation used to match feature dimensions. Represents a separation operation, which is used to separate a vector into two or more sub-vectors. and Represent the attention features in row and column directions respectively, represents the sigmoid function, represents the convolutional layer, represents the maximum pooling layer, represents an average pooling layer.
[0092] Furthermore, the spatial attention module enhances the spatial features and frequency attention weights respectively and Multiply and concatenate in the channel dimension to obtain space-frequency enhancement features The specific process can be expressed as:
[0093]
[0094] in, represents channel dimension splicing, represents element-wise multiplication, It is the space-frequency enhancement feature.
[0095] In summary, when using the spatial-frequency attention sub-model for feature enhancement:
[0096] First, the infrared region features and visible light region characteristics Input the spatial attention module and the frequency integrated dual-coordinate attention module respectively;
[0097] Then, the spatial attention module is used to obtain the spatial enhancement features of infrared and visible light region features. ;
[0098] Then, the frequency integration module of the frequency integration dual-coordinate attention module performs fast Fourier transform FFT on the infrared and visible light region features to obtain infrared amplitude information and infrared phase information, as well as visible light amplitude information and visible light phase information. The phase information integration layer is used to integrate the phase information, and the interaction features are obtained by inverse transform IFFT. , where for the target area features Use infrared amplitude information to perform inverse Fourier transform, for background area features Perform inverse Fourier transform using visible light amplitude information;
[0099] Then, the dual-coordinate attention module of the frequency integrated dual-coordinate attention module uses the dual-coordinate attention mechanism to calculate the frequency attention weight based on the interaction features , wherein the dual-coordinate attention module first performs average pooling on its input features along the height direction and width direction respectively to obtain statistical features in the two directions, and then forms a joint feature by splicing the obtained features through the convolution layer, batch normalization layer and ELU activation function, and then obtains the attention vectors in the two directions through separation operation. Finally, the directional attention map is generated through two convolution layers and sigmoid activation function, and the frequency attention weight is generated by combining them in the form of matrix multiplication. ;
[0100] Then, the spatial attention module combines the spatial enhancement features of infrared and visible light and frequency attention weights respectively and Multiply and concatenate to generate space-frequency enhancement features, expressed as:
[0101]
[0102] in, represents channel dimension splicing, represents element-wise multiplication, It is the space-frequency enhancement feature.
[0103] The fusion block sub-model is used to enhance the spatial-frequency features output by the spatial-frequency attention sub-model. The fusion is performed to obtain the final fusion image. The fusion block sub-model includes fusion blocks in sequence. to , each fusion block contains a convolution module , to pass through the convolution module Integrate the input features.
[0104] Among them, the fusion process is: the first fusion block Receive the third layer Spatial-frequency enhanced features output by the spatial-frequency attention sub-model , and generate the first layer of intermediate fusion results , the first layer of intermediate fusion result after upsampling With the second layer Spatial-frequency enhanced features output by the spatial-frequency attention sub-model After cascading, it is fed into the second fusion block , and then generate the second layer of intermediate fusion results , the second layer intermediate fusion result after upsampling With the first layer Spatial-frequency enhanced features output by the spatial-frequency attention sub-model After cascading, it is fed into the third fusion block , and then generate the third layer of intermediate fusion results , the fourth fusion block receives the third layer intermediate fusion result And generate the final fused image The specific process can be expressed as:
[0105]
[0106]
[0107]
[0108]
[0109] in, Represents an element-by-element addition operation, which is used to stitch the target area and the background area into a complete image. Represents an upsampling operation.
[0110] In summary, when the enhanced features are fused using the fusion block sub-model:
[0111] First, the spatial-frequency enhancement features of the target area and background area at the same scale are element-wise added and input into the corresponding fusion block. and At the same time, the output features of the previous fusion block after upsampling are input to the corresponding fusion block, and then the two features of the input corresponding fusion block are spliced and integrated. , only the output features of the previous fusion block are input, and finally the fourth fusion block Output That is the final fused image. The overall process of the fusion block sub-model fusing the space-frequency enhancement features can be expressed as:
[0112] in, represents the fusion block sub-model, Represents the spatial-frequency enhancement characteristics of the target area, Represents the spatial-frequency enhancement characteristics of the background area.
[0113] Step S3, training the image fusion depth model based on the training data set to obtain a target image fusion depth model;
[0114] In one embodiment of the present invention, the training process adopts a hybrid fusion loss function The model parameters are optimized to balance texture details, statistical correlation and pixel intensity preservation.
[0115] Among them, the hybrid fusion loss function Expressed as:
[0116]
[0117] in, represents the gradient loss function, represents the correlation loss function, represents the intensity loss function, 、 、 is the corresponding weight coefficient.
[0118] Furthermore, the gradient loss function It can be expressed as:
[0119]
[0120] in, and Represent the input visible light image and infrared image respectively, represents the fused image output by the image fusion depth model, Represents the Sobel gradient operator, which is used to calculate image gradient information. represents the calculation of the absolute value of the matrix, represents the matrix 1 norm calculation, Indicates pixel-level maximum value selection, and Represents the length and width of the image respectively.
[0121] Furthermore, the correlation loss function It can be expressed as:
[0122]
[0123] in,
[0124]
[0125] in, and Represent the input visible light image and infrared image respectively, represents the fused image output by the image fusion depth model, represents the correlation coefficient calculation, and Represent the weight coefficients of visible light image and infrared image respectively, represents the covariance calculation, represents variance calculation, Represents a very small constant used to ensure that the denominator of the fraction is not zero.
[0126] Furthermore, the intensity loss function It can be expressed as:
[0127]
[0128] in
[0129]
[0130] in, and Represent the input visible light image and infrared image respectively, represents the fused image output by the image fusion depth model, Represents a pixel mask, which is used to compare the pixel values of the source image. represents the matrix 1 norm calculation, and Represents the length and width of the image respectively.
[0131] Step S4: fuse the to-be-fused image pair using the target image fusion depth model to obtain a fused image result.
[0132] Figure 5 The comparison results of average gradient (AG), correlation sum of differences (SCD), and standard deviation (SD) obtained by different methods in the infrared and visible light image fusion task are listed. The comparison algorithms include: Liu's method and Li's method. The larger the average gradient (AG), the richer the details of the fused image; the larger the correlation sum of differences (SCD), the more complementary information is retained in the fused image; and the larger the standard deviation (SD), the higher the overall contrast of the fused image. Figure 5It can be seen that compared with Liu's method and Li's method, the method of the present invention achieves significant improvements in the AG, SCD, and SD indicators. Specifically, in terms of the average gradient (AG) indicator, the method of the present invention improves by approximately 44% and 26% compared with Liu's method and Li's method, respectively, indicating that the detail information in the fused image is more fully preserved. In terms of the difference correlation sum (SCD) indicator, the method of the present invention is also much higher than the comparison method, indicating that the fusion result has a stronger ability to extract and integrate complementary information from the source images. In terms of the standard deviation (SD) indicator, the method of the present invention also achieves the highest value, indicating that the contrast of the fused image is more distinct and the visual effect is better. The main reason is that Liu's method focuses on multimodal collaborative learning but pays insufficient attention to the differences between the salient target area and the background area. Although Li's method introduces a cross-attention mechanism, it does not fully combine the multi-scale feature expression of the spatial and frequency domains, which limits further improvement of fusion performance. In contrast, the present invention designs a target perception area fusion sub-model, a space-frequency attention sub-model and a fusion block sub-model to extract and enhance features in the salient target area and the background area respectively, and combines the spatial domain and frequency domain attention mechanisms to significantly improve the overall performance of the fused image in terms of detail preservation, complementary information integration and contrast, and ultimately achieves an excellent infrared and visible light image fusion effect.
[0133] Unless otherwise specified, the embodiments of the present invention do not limit the models of the components. Any component that can perform the above functions may be used.
[0134] Those skilled in the art will understand that the accompanying drawings are only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0135] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for fusion of infrared and visible light images based on target perception, characterized in that: The method comprises the following steps: Step S1: obtaining a training data set, wherein the training data set includes a plurality of input image pairs consisting of infrared images and visible light images, wherein the input image pairs include infrared images and visible light images ; Step S2, constructing an image fusion depth model, wherein the image fusion depth model includes a target perception region fusion sub-model, a space-frequency attention sub-model and a fusion block sub-model, and the target perception region fusion sub-model includes a pre-trained binary image segmentation module, a region enhancement module and a parallel multi-scale convolution branch module, wherein: the pre-trained binary image segmentation module is used to receive infrared images , and based on the infrared image Generate salient object masks through the trained binary image segmentation network and its inverse mask ; The region enhancement module is used to use the mask and its inverse mask For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area to obtain the salient target area image pair. and background area image pair ,in, Indicates the salient target area, Represents the background area; the parallel multi-scale convolution branch module is used to and background area image pair Extract multi-scale regional features of salient target areas and background areas; Step S3, training the image fusion depth model based on the training data set to obtain a target image fusion depth model; Step S4, using the target image fusion depth model to fuse the image pair to be fused to obtain a fused image result; The space-frequency attention sub-model includes a frequency-integrated dual-coordinate attention module and a spatial attention module, wherein: The frequency integrated dual-coordinate attention module is used to realize the interactive enhancement of infrared features and visible features in the frequency domain and generate frequency attention weights ; The spatial attention module is used to utilize the frequency attention weight Enhance regional differentiation features in the spatial domain to generate space-frequency enhancement features , ; The frequency integration dual-coordinate attention module includes a frequency integration module and a dual-coordinate attention module, wherein: The frequency integration module is used to generate interactive features based on the regional differentiation features ; The dual-coordinate attention module is used to Generate frequency attention weights .
2. The method according to claim 1, characterized in that in: The target perception area fusion sub-model is used to separate the infrared image and visible light images The salient target area and background area are extracted, and regional differentiation features of the salient target area and background area are extracted respectively: The space-frequency attention sub-model is used to perform weighted enhancement on the regional differentiation features of the salient target area and the background area to obtain the space-frequency enhancement features: The fusion block sub-model is used to fuse the space-frequency enhancement features to obtain a final fused image.
3. The method according to claim 2, characterized in that The region enhancement module utilizes a mask based on the following formula and its inverse mask For infrared images and visible light images Perform differential enhancement processing on the salient target area and the background area: ; in, is a pair of salient target region images, is a background region image pair, represents element-wise multiplication, is the enhancement factor.
4. The method according to claim 2 or 3, characterized in that The parallel multi-scale convolution branch module includes a salient target area branch and a background area branch. Each branch includes three convolution modules. Each convolution module includes a convolution layer, a batch normalization layer and a PReLU activation function in sequence. The parallel multi-scale convolution branch module is based on the salient target area image pair according to the following formula. and background area image pair Extract multi-scale regional features of salient target areas and background areas, namely regional differentiation features : ; ; ; in, Indicates the The regional features output by the convolution module, , , represents the convolution module, represents the maximum pooling operation, Represents a channel dimension concatenation operation.
5. The method according to claim 1, wherein The frequency integration module generates an interactive feature based on the regional differentiation feature according to the following formula: : ; ; in, represents the fast Fourier transform, Indicates infrared amplitude information, Indicates infrared phase information, Represents visible light amplitude information, Represents visible light phase information, For regional differentiation characteristics, represents the phase information integration layer, represents the inverse fast Fourier transform, Indicates the salient target area, Represents the background area; The dual-coordinate attention module is based on the interaction feature according to the following formula Generate frequency attention weights : ; ; in, and Represents row and column average pooling respectively, represents channel dimension splicing, represents the combination layer of convolution, batch normalization and ELU activation function, represents a separation operation, R represents a transformation operation, represents the convolutional layer, represents the sigmoid function, Represents element-wise multiplication.
6. The method according to claim 1, characterized in that The spatial attention module enhances the regional differentiation features in the spatial domain according to the following formula to generate space-frequency enhancement features : ; ; ; in, For regional differentiation characteristics, represents the maximum pooling layer, represents the average pooling layer, represents channel dimension splicing, represents the convolutional layer, represents the sigmoid function, Represents element-wise multiplication.
7. The method according to claim 2, characterized in that The fusion block sub-model includes fusion blocks in sequence to , each fusion block contains a convolution module The fusion block sub-model uses the following formula to fuse the spatial-frequency enhancement features to obtain the final fused image : ; ; ; ; in, Indicates the The spatial-frequency enhanced features output by the layer spatial-frequency attention sub-model, , represents the upsampling operation, represents the element-by-element addition operation, Indicates channel dimension splicing.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on cross mode enhancement and multi-attention fusion strategy
CN118096554A
Visible light and infrared image fusion method based on space-frequency domain characteristics
CN119963958A