Feature level fusion method, system and equipment for infrared and visible light images

By combining the feature extraction of Haar transform and convolutional neural network, and designing Haar cross attention mechanism and new dense connection forms, the problem of poor middle and low frequency and high frequency information processing of infrared and visible light images is solved, and efficient image fusion and feature extraction effects are achieved.

CN119942283AActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510028911.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

In the fusion process of infrared and visible light images, it is difficult to effectively pay attention to the differences in low-frequency and high-frequency information of the image, resulting in poor correlation and inconsistency processing of the fused image, and deepening the model leads to loss of shallow features and increasing the calculation amount.

Method used

By combining the feature extraction of Haar transform and convolutional neural network, a Haar cross attention mechanism is designed to process low-frequency shared information and high-frequency unique information of visible light and infrared images, and a new dense connection form is proposed, combining edge operators and convolutional block attention mechanisms to reduce the computational amount and improve the reusability of shallow features.

Benefits of technology

It realizes effective processing of low-frequency and high-frequency information during the fusion process of infrared and visible light images, improves the correlation and inconsistency of the fusion image, reduces the calculation amount, and maintains high-precision feature extraction and fusion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942283A_ABST
    Figure CN119942283A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and provides an infrared and visible light feature level fusion method, system and device. Haar transformation is combined with feature extraction of a convolutional neural network, so that feature extraction of infrared and visible light images has emphasis; a cross attention mechanism is designed, so that the model can respectively process low-frequency common information and high-frequency unique information of visible light and infrared images; and a new dense connection form is provided, and an edge operator and a convolution block attention mechanism are combined, so that the calculation amount is reduced, and the reusability of shallow layer features is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a feature-level fusion method, system and device for providing infrared and visible light fusion images for downstream target detection tasks. Background Art

[0002] In modern society, the importance of images has been highlighted in all walks of life. In the current field of image processing and research, single-modality imaging often fails to provide us with comprehensive and accurate information. The emergence of image fusion technology solves this problem. Its purpose is to fuse image information from different sensors to obtain a fused image that contains more information and better meets the requirements of downstream target detection tasks. Visible light images are formed by light reflected from objects, which can provide rich texture and color information, and have outstanding visual expression capabilities. It is the main way for human visual perception to obtain information. However, its imaging is highly dependent on lighting conditions. When encountering situations with low light intensity such as wind, snow, heavy rain, and night, it will cause it to lose a lot of image information. Infrared images are formed through thermal radiation imaging, which is not easily affected by lighting conditions. It can still complete imaging under conditions where visible light is difficult to image. At the same time, it can still image some targets that are blocked by other objects. Although the detailed information of the image is lost, it has more prominent information expression for targets such as pedestrians and vehicles in bad weather. In the field of security, since we need to conduct 24-hour all-weather monitoring, the fusion of visible light images and infrared images can comprehensively utilize the detailed texture information of visible light and the significant information of infrared, which can improve the performance of downstream target detection tasks.

[0003] At present, image fusion can be divided into three types from different stages of fusion: pixel-level fusion, feature-level fusion, and decision-level fusion; pixel-level fusion is to start directly from each pixel of the image and process the pixels by weighted average and other methods, but this method requires too much data to be processed, which will lead to a long fusion time and affect the real-time performance of the system; decision-level fusion is a high-level fusion, which is to analyze, process and make decisions on each source image separately, and then fuse these decision results to obtain the final fusion decision. This method is highly dependent on the accuracy of the decision results, so the accuracy is poor; feature-level fusion refers to extracting features from the source image, processing and making decisions on the extracted feature level, etc., and completing the fusion of different source images at the feature level, which is higher than the pixel level. This reduces the amount of calculation and the accuracy can still be maintained, but there are still the following three problems. First, the focus of infrared images and visible light images in the fusion process is different. The infrared image is affected by the imaging method, and its details, such as texture and other high-frequency information, are not clearly expressed. Therefore, the fused image is more likely to focus on the contour, shape and other low-frequency information of the infrared image, while the visible light has better detail expression ability, so it is more necessary to pay attention to its details and other high-frequency information. Second, during the fusion process, infrared and visible light have common low-frequency information, such as contours and shapes, but high-frequency signals are unique to each, such as visible light texture and infrared temperature information. How to improve the relevance of low-frequency information and the irrelevance of high-frequency information is a problem that needs to be solved. Third, deepening the model will lead to the loss of shallow features and an increase in the amount of calculation, which also needs to be solved.

[0004] In summary, it is necessary to design an image fusion method, system and equipment that can fully consider the differences between infrared and visible light images to solve the problem. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention proposes a method, system and device for feature-level fusion of infrared and visible light. By combining Haar transform with feature extraction of convolutional neural network, feature extraction of infrared and visible light images is focused. Haar cross-attention mechanism is designed so that the model can process low-frequency common information and high-frequency unique information of visible light and infrared images separately. A new dense connection form is proposed, combined with edge operator and convolution block attention mechanism, to reduce the amount of calculation and improve the reusability of shallow features.

[0006] In order to achieve the above object, the technical solution of the present invention is:

[0007] A method for characteristic-level fusion of infrared and visible light, comprising the following steps:

[0008] S1: Obtain infrared and visible light image datasets, preprocess the datasets, and obtain training sets;

[0009] S101: Obtain a sensor data set; use visible light and infrared cameras to shoot, create a data set D, divide the data set D in proportion, and obtain a training set, wherein the data set D contains both visible light and infrared data;

[0010] S2: Construct a feature-level infrared and visible light image fusion network model, including the shallow feature extraction module Haar Feature Extraction, the dense connection module SADense, the infrared image enhancement module Haar ImageEnhancement, and the cross attention module Haar CrossAttention;

[0011] The shallow feature extraction module Haar Feature Extraction uses Haar transform to decompose the image and divide the visible light and infrared images into high- and low-frequency sub-bands. The shallow feature extraction module includes an infrared feature extraction module, which consists of two 3*3 convolutional layers and one 1*1 convolutional layer; and a visible light feature extraction module, which consists of two 3*3 convolutional layers and one 1*1 convolutional layer.

[0012] The infrared image enhancement module Haar Image Enhancement decomposes the image through Haar transform, enhances the infrared background contour information, fits it with the detail information of visible light, and outputs an infrared feature enhanced image.

[0013] The densely connected module SADense realizes the reuse of shallow features by densely connecting shallow features. It includes a Sobel operator, a 1*1 convolutional layer, two 3*3 convolutional layers and a convolutional block attention mechanism.

[0014] Cross-Attention Module Haar CrossAttention, including a channel dimension attention module, consisting of a spatial average pooling layer, two 1*1 convolutional layers, two relu activation functions and a sigmoid activation function. A spatial dimension attention module, consisting of a channel average pooling layer, a 7*7 convolutional layer, and a sigmoid activation function. A transposed attention module, consisting of an image transposed layer and two 1*1 convolutional layers. Haar transform is used to process high and low frequency information separately.

[0015] S3: Set network parameters. The specific contents are as follows: set the initial learning rate Ir0, training rounds epochs, and batch size batch-size.

[0016] S4: According to the neural network model reconstructed in steps 2 and 3, the training set in step 1 is used for training to obtain the optimal weight file of the network model.

[0017] S5:,The obtained weight file is applied to image fusion, and finally the obtained fused image is input into the downstream object detection task.

[0018] Furthermore, the specific contents of step S2 are as follows:

[0019] Step 201: Construct a visible light shallow feature extraction module. The visible light image is transformed using Haar transform to obtain a low-frequency subband LL and three high-frequency subbands HL, LH, and HH. The visible light image pays more attention to high-frequency information such as details. The HH high-frequency subband is transformed by a second-order Haar transform. A 3*3 convolution operation is performed on the four subbands obtained by the second-order transform so that the model focuses on learning and using detail information. The obtained feature map is inversely transformed by Haar transform to restore the image. A 1*1 convolution is connected in series, and the number of channels is adjusted to obtain a visible light shallow feature extraction feature map F that focuses on high-frequency signals. vi1 , so that the next layer of network can process it.

[0020] Step 202: Construct an infrared shallow feature extraction module. Use Haar transform on the visible light image to obtain a low-frequency subband LL and three high-frequency subbands HL, LH, and HH. The infrared image pays more attention to low-frequency information such as contour shape. Perform a second-order Haar transform on the low-frequency subband LL. Perform a 3*3 convolution operation on the four subbands obtained by the second-order transform to make the model focus on learning and using information such as contours. Perform Haar inverse transform on the obtained feature map to restore the image. A 1*1 convolution is connected in series, and the number of channels is adjusted to obtain an infrared shallow feature extraction feature map F that focuses on low-frequency signals. ir1 , so that the next layer of network can process it.

[0021] Step 203: Construct a dense connection module SADense, and combine the feature maps F obtained in steps 201 and 202 above vi1 and F ir1 The features are reused on each branch, and connected to each subsequent layer on each single branch, and the first layer feature map F of size W*H*C is vi1 and F ir1 The weight information is added through 3*3 convolution respectively, as shown below:

[0022]

[0023] Where m, n represent the spatial dimensions of the convolution kernel, K(m,n,c,k) represents the learned weight information, and after the convolution operation, i, j represent the corresponding pixel coordinates on the feature map, 0≤i≤W, 0≤j≤H, and the weighted learned feature map F is obtained. vi2 and F ir2 , and then compared with the feature map F obtained in step 201 and step 202 vi1 and Fir1 Stitching:

[0024]

[0025] As shown above, we get a new feature map F of size W*H*(C+C1) viconcat1 and F irconcat1 , convolution weighted to obtain feature map F vi3 and F ir3 , and then reuse the feature map F obtained in step 201 and step 202 vi1 and F ir1 , and the feature map F vi3 and F ir3 Splice as follows:

[0026]

[0027] The size of the visible light and infrared branches is W*H*(C+C3) feature map F viconcat2 and F irconcat2 Input them into the convolution block attention mechanism respectively, and get the output feature map M vis and M irs , as follows:

[0028]

[0029] M vis =σ(Conv(F vic ))

[0030]

[0031] M irs =σ(Conv(F irc ))

[0032] Among them, F vic Yes F viconcat2 After the convolutional attention mechanism, the middle-level output feature map, F irc Yes F irconcat2 After the middle-layer output feature map of the convolutional attention mechanism, the nonlinear features are added through the activation function to obtain M vis and M irs ;σ represents the activation function, M vis and the feature map F obtained in step 201 vi1 The feature map F obtained by Sobel operator and convolution operation viSout Multiplication, M irs and the feature map F obtained in step 202 ir1 The feature map F obtained by Sobel operator and convolution operation irSout Multiply, as shown below:

[0033] F visout =Conv(Concat(Sobel x (F vi1 ),Sobel y (F vi1 )),K conv )

[0034] F vifinal =F SCout ·F sout

[0035] F irsout =Conv(Concat(Sobel x (F ir1 ),Sobel y (F ir1 )),K conv )

[0036] F irfinal =F SCout ·F sout

[0037] Sobel x Represents the Sobel operator in the horizontal direction. y Represents the Sobel operator in the vertical direction, concatenated convolution, and finally obtains the above SADense dense block output feature map F vifinal and F irfinal .

[0038] Step 204: Infrared image enhancement module Haar Image Enhancement, decomposes the image through Haar transformation, uses low-frequency signal LL to enhance infrared background contour information, and high-frequency signals HL, LH, HH to enhance visible light detail information. The enhancement module is the module described in step 205, and outputs an infrared feature enhanced image.

[0039] Step 205: Construct a cross-attention module Haar CrossAttention. The infrared and visible light feature maps are first transformed by Haar to obtain low-frequency sub-bands and high-frequency sub-bands. The low-frequency signal emphasizes relevance, and the low-frequency sub-band feature maps are added. The high-frequency signal emphasizes irrelevance, and the high-frequency sub-band feature maps are subtracted. The obtained feature maps are input into the cross-attention module. The input feature maps are processed in parallel, focusing on the spatial and channel dimensions respectively. The channel dimension is first average-pooled for the feature maps, as shown in the following formula:

[0040]

[0041] The feature map of size 1*1*C can be obtained. Each channel contains the global channel information of its own space. After 1*1 convolution, the linear transformation of each channel is completed, and the channel dimension is reduced and the redundant channel dimension is removed, as shown in the following formula:

[0042]

[0043] Introduce nonlinear transformation and implement it through the activation function Relu:

[0044]

[0045] Then, through a layer of 1*1 convolution operation and the introduction of nonlinear activation function, the dimension of the feature map is restored, and the channels are reintegrated and reconstructed. The subsequent Sigmoid normalization process is used to compress the weights to the range of [0,1].

[0046] In the spatial dimension, the channel dimension of the feature map is first average pooled, as shown in the following formula:

[0047]

[0048] A feature map of size w*h*1 is obtained. Each pixel in the space contains global channel information. The spatial dimension requires a larger receptive field to enhance the connection between features in different spaces. Therefore, a large kernel convolution of 7*7 is used, as shown in the following formula:

[0049]

[0050] Among them, in order to prevent the convolution process from crossing the boundary and still output the feature map of size w*h, padding is required. p represents the extra pixels added to the edge of the data. Here p is set to 3. The feature map processed in the spatial dimension is input into the Sigmoid activation function, the weight is normalized, and multiplied with the tensor in the channel dimension to generate a feature map that focuses on both the channel and spatial dimensions, as shown in the following formula:

[0051] Y(w,h,c)=A(1,1,c)×B(w,h,1)

[0052] Finally, we get the feature map F processed in both spatial and channel dimensions. SC , starting from its own global information, and then processing in parallel with the transposed attention module, by transposing the image itself, then splicing, and then convolving it with F in an appropriate weight ratio SC Stitching, weighting is completed, and then Haar inverse transform is performed to restore the image.

[0053] An infrared and visible light feature-level fusion system, comprising:

[0054] Data processing module: uses a camera to collect images and obtain the processed data.

[0055] Network building module: Construct a network model based on feature-level infrared and visible light image fusion, including shallow feature extraction module Haar Feature Extraction, dense connection module SADense, infrared image enhancement module Haarimage enhancement, cross attention module Haar CrossAttention

[0056] Network training module: used to perform training on the training set according to the set network parameters, output the training weight file, and obtain the required optimal image fusion weight file.

[0057] Output module: deploy the obtained image fusion model on the device, build a post-processing unit, and output the fused image to the downstream target detection model.

[0058] An infrared and visible light feature-level fusion device, comprising:

[0059] Machine executable instructions: a computer program storing the infrared and visible light feature-level fusion method, which is a computer-readable executable instruction;

[0060] Processor: A processor used to execute the computer program to implement the infrared and visible light feature-level fusion method.

[0061] A computer-readable storage medium: The computer-readable storage medium stores a computer program, and when the computer program is processed by a processor, the storage medium can store the infrared and visible light feature-level fusion method. The storage medium can be any physical storage medium and can include or store information, such as RAM, EMMC, ROM, SD, DVD, etc.

[0062] The beneficial effects of the present invention are mainly manifested in:

[0063] 1) The present invention designs a shallow feature extraction module, which combines Haar transform to focus on high-frequency information such as texture for visible light feature extraction and on low-frequency information such as contour shape for infrared feature extraction, so that the fused image can better combine the advantages of both.

[0064] 2) The present invention constructs a new type of dense connection module, which allows shallow features to be reused and uses edge operators to improve the expression of edge information, reducing the amount of calculation of traditional dense connections and reducing the occurrence of the problem of shallow features not being able to be learned during model training.

[0065] 3) The present invention proposes a cross-attention mechanism, which cross-processes the common information of low-frequency signals and the unique information of high-frequency signals of visible light images and infrared images, respectively, to improve the relevance of the common low-frequency information and the irrelevance of the unique high-frequency information. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a structural diagram of shallow feature extraction of the present invention;

[0067] Figure 2 It is the infrared image enhancement diagram of the present invention;

[0068] Figure 3 It is a cross-attention structure diagram of the present invention;

[0069] Figure 4 It is a parallel structure diagram of the cross-attention space and channel dimension of the present invention;

[0070] Figure 5 It is a diagram of the cross attention transposition attention structure of the present invention;

[0071] Figure 6 It is the overall network structure diagram of the present invention;

[0072] Figure 7 is a schematic diagram of the fused image of the present invention;

[0073] Figure 8 It is a schematic diagram of infrared image enhancement of the present invention;

[0074] Fig. 9 is a system block diagram of the present invention;

[0075] Fig.10 It is a hardware structure diagram of the electronic equipment of the present invention. DETAILED DESCRIPTION

[0076] The present invention will be further described below in conjunction with the accompanying drawings.

[0077] Reference Figure 1 to Figure 6 , a method for fusion of infrared and visible light feature level, comprising the following steps:

[0078] S1: Use visible light and infrared cameras to shoot, obtain infrared and visible light image datasets, preprocess the datasets, and obtain training sets.

[0079] S2: Construct a network model based on feature-level infrared and visible light image fusion, including the shallow feature extraction module Haar Feature Extraction, the dense connection module SADense, the infrared image enhancement module Haar ImageEnhancement, and the cross attention module Haar CrossAttention.

[0080] The shallow feature extraction module Haar Feature Extraction includes using Haar transform to decompose the image and divide the visible light and infrared images into high- and low-frequency sub-bands. The infrared feature extraction module consists of 2 3*3 convolutional layers and one 1*1 convolutional layer. The visible light feature extraction module consists of 2 3*3 convolutional layers and one 1*1 convolutional layer.

[0081] The densely connected module SADense realizes the reuse of shallow features by densely connecting shallow features. It includes a Sobel operator, a 1*1 convolutional layer, two 3*3 convolutional layers and a convolutional block attention mechanism.

[0082] The infrared image enhancement module Haar image enhancement decomposes the image through Haar transform, enhances the infrared background contour information, fits it with the detail information of visible light, and outputs an infrared feature enhanced image.

[0083] Cross-Attention Module Haar CrossAttention, including a channel dimension attention module, consisting of a spatial average pooling layer, two 1*1 convolutional layers, two relu activation functions and a sigmoid activation function. A spatial dimension attention module, consisting of a channel average pooling layer, a 7*7 convolutional layer, and a sigmoid activation function. A transposed attention module, consisting of an image transposed layer and two 1*1 convolutional layers. Haar transform is used to process high and low frequency information separately.

[0084] The specific process steps of S2 are as follows:

[0085] Step 201: Construct a visible light shallow feature extraction module. The visible light image is transformed using Haar transform to obtain a low-frequency subband LL and three high-frequency subbands HL, LH, and HH. The visible light image pays more attention to high-frequency information such as details. The HH high-frequency subband is transformed by a second-order Haar transform. A 3*3 convolution operation is performed on the four subbands obtained by the second-order transform so that the model focuses on learning and using detail information. The obtained feature map is inversely transformed by Haar transform to restore the image. A 1*1 convolution is connected in series, and the number of channels is adjusted to obtain a visible light shallow feature extraction feature map F that focuses on high-frequency signals. vi1 , so that the next layer of network can process it.

[0086] Step 202: Construct an infrared shallow feature extraction module. Use Haar transform on the visible light image to obtain a low-frequency subband LL and three high-frequency subbands HL, LH, and HH. The infrared image pays more attention to low-frequency information such as contour shape. Perform a second-order Haar transform on the low-frequency subband. Perform a 3*3 convolution operation on the four subbands obtained from the second-order transform to make the model focus on learning and using information such as contours. Perform an inverse Haar transform on the obtained feature map to restore the image. A 1*1 convolution is connected in series. The number of channels is adjusted to obtain an infrared shallow feature extraction feature map F that focuses on low-frequency signals. ir1 , so that the next layer of network can process it.

[0087] The Haar transform expression is:

[0088] Low-Low (LL) sub-band:

[0089]

[0090] Low-High (LH) sub-band:

[0091]

[0092] High-Low (HL) sub-band:

[0093]

[0094] High-High (HH) sub-band:

[0095]

[0096] The convolutional layer parameter settings are shown in the following table:

[0097] Table WT convolution kernel parameter settings

[0098] Table Convolution Kernel Parameter Settings

[0099]

[0100] Step 203: Construct a dense connection module SADense, and combine the feature maps F obtained in steps 201 and 202 above vi1 and F ir1 The features are reused on each branch, and connected to each subsequent layer on each single branch, and the first layer feature map F of size W*H*C is vi1 and F ir1 The weight information is added through 3*3 convolution respectively, as shown below:

[0101]

[0102]

[0103] Where m, n represent the spatial dimensions of the convolution kernel, K(m,n,c,k) represents the learned weight information, and after the convolution operation, i, j represent the corresponding pixel coordinates on the feature map, 0≤i≤W, 0≤j≤H, and the weighted learned feature map F is obtained. vi2 and F ir2 , and then compared with the feature map F obtained in step 201 and step 202 vi1 and F ir1 Stitching:

[0104]

[0105] As shown above, we get a new feature map F of size W*H*(C+C1) viconcat1 and F irconcat1 , convolution weighted to obtain feature map F vi3 and F ir3 , and then reuse the feature map F obtained in step 201 and step 202 vi1 and F ir1 , and the feature map F vi3 and F ir3 Splice as follows:

[0106]

[0107] The size of the visible light and infrared branches is W*H*(C+C3) feature map F viconcat2 and F irconcat2 Input them into the convolution block attention mechanism respectively, and get the output feature map M vis and M irs , as follows:

[0108]

[0109] M vis =σ(Conv(F vic ))

[0110]

[0111] M irs =σ(Conv(F irc ))

[0112] Where σ represents the activation function, M vis and the feature map F obtained in step 201 vi1 The feature map F obtained by Sobel operator and convolution operation viSout Multiplication, M irs and the feature map F obtained in step 202 ir1The feature map F obtained by Sobel operator and convolution operation irSout Multiply, as shown below:

[0113] F visout =Conv(Concat(Sobel x (F vi1 ),Sobel y (F vi1 )),K conv )

[0114] F vifinal =F SCout ·F sout

[0115] F irsout =Conv(Concat(Sobel x (F ir1 ),Sobel y (F ir1 )),K conv )

[0116] F irfinal =F SCout ·F sout

[0117] Sobel x Represents the Sobel operator in the horizontal direction. y Represents the Sobel operator in the vertical direction, concatenated convolution, and finally obtains the above SADense dense block output feature map F vifinal and F irfinal .

[0118] Step 204: Infrared image enhancement module Haar Image Enhancement, decomposes the image through Haar transformation, uses low-frequency signal LL to enhance the infrared background contour information, and high-frequency signals HL, LH, HH to enhance the detail information of visible light, and then fits. The enhancement method is the module described in step 205, and outputs an infrared feature enhanced image.

[0119] Step 205: Construct a cross attention module Haar Cross Attention. Infrared and visible light are first transformed by Haar to obtain low-frequency sub-bands and high-frequency sub-bands. Low-frequency signals emphasize relevance, and low-frequency sub-band feature maps are added. High-frequency signals emphasize irrelevance, and high-frequency sub-band feature maps are subtracted. The resulting feature maps are input into the cross module. The input feature maps are processed in parallel, focusing on the spatial and channel dimensions respectively. The channel dimension is first averaged and pooled, as shown in the following formula:

[0120]

[0121] The feature map of size 1*1*C can be obtained. Each channel contains the global channel information of its own space. After 1*1 convolution, the linear transformation of each channel is completed, and the channel dimension is reduced and the redundant channel dimension is removed, as shown in the following formula:

[0122]

[0123] Introduce nonlinear transformation and implement it through the activation function Relu:

[0124]

[0125] Then, through a layer of 1*1 convolution operation and the introduction of nonlinear activation function, the dimension of the feature map is restored, and the channels are reintegrated and reconstructed. The subsequent Sigmoid normalization process is used to compress the weights to the range of [0,1].

[0126] In the spatial dimension, the channel dimension of the feature map is first average pooled, as shown in the following formula:

[0127]

[0128] A feature map of size w*h*1 is obtained. Each pixel in the space contains global channel information. The spatial dimension requires a larger receptive field to enhance the connection between features in different spaces. Therefore, a large kernel convolution of 7*7 is used, as shown in the following formula:

[0129]

[0130] Among them, in order to prevent the convolution process from crossing the boundary and still outputting the feature map of size w*h, padding is required. Here, p is set to 3. The feature map processed in the spatial dimension is input into the Sigmoid activation function, the weight is normalized, and multiplied with the tensor in the channel dimension to generate a feature map that focuses on both the channel and spatial dimensions, as shown in the following formula:

[0131] Y(w,h,c)=A(1,1,c)×B(w,h,1)

[0132] Finally, we get the feature map F processed in both spatial and channel dimensions. SC , starting from its own global information, and then processing in parallel with the transposed attention module, by transposing the image itself, then splicing, and then convolving it with a weight ratio of 0.1 with F SC Stitching, weighting is completed, and then Haar inverse transform is performed to restore the image.

[0133] S3: Set network parameters: set the initial learning rate Ir0 = 0.001, training rounds epochs = 50, and batch size batch-size = 1.

[0134] S4: According to the neural network model reconstructed in steps 2 and 3, the training set in step 1 is used for training to obtain the optimal weight file of the network model.

[0135] S5:,The obtained weight file is applied to image fusion, and finally the obtained fused image is input into the downstream object detection task.

[0136] Reference Figure 8 , an infrared and visible light feature-level fusion system, comprising:

[0137] Data processing module: uses a camera to collect images and obtain the processed data.

[0138] Network building module: Build a network model based on feature-level infrared and visible light image fusion, including shallow feature extraction module Haar Feature Extraction, dense connection module SADense, infrared image enhancement module Haarimage enhancement, and cross attention module Haar CrossAttention.

[0139] Network training module: used to perform training on the training set according to the set network parameters, output the training weight file, and obtain the required optimal image fusion weight file.

[0140] Output module: deploy the obtained image fusion model on the device, build a post-processing unit, and output the fused image to the downstream target detection model.

[0141] Reference Fig. 9 , an infrared and visible light feature-level fusion device, comprising:

[0142] Machine executable instructions: a computer program storing the infrared and visible light feature-level fusion method, which is a computer-readable executable instruction;

[0143] Processor: A processor used to execute the computer program to implement the infrared and visible light feature-level fusion method.

[0144] A computer-readable storage medium: The computer-readable storage medium stores a computer program, and when the computer program is processed by a processor, the storage medium can store the infrared and visible light feature-level fusion method. The storage medium can be any physical storage medium and can include or store information, such as RAM, EMMC, ROM, SD, DVD, etc.

[0145] The present invention conducts experiments on the MSRS data set and compares it with the existing image fusion algorithm. Six indicators are selected for comparison, including SSIM (structural similarity between the fused image and the original image), AG (activity of the fused image), EN (the amount of information of the original image contained in the fused image), SD (contrast of the fused image), SCD (the degree of preservation of the details and texture information of the fused image), and Qabf (comprehensive performance), as shown in the following table:

[0146]

[0147] The present invention has achieved the best results in the four indicators of AG, EN, SCD, and Qabf, which shows that the fused image of the present invention has achieved more effective integration of information such as detail texture, more comprehensive utilization of information of visible light and infrared images, and the fused image has a higher visual contrast.

[0148] In this specification, the schematic description of the present invention does not necessarily refer to the same embodiment or example, and those skilled in the art may combine and combine different embodiments or examples described in this specification. In addition, the contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept, and the protection scope of the present invention should not be considered as limited to the specific forms described in the implementation cases. The protection scope of the present invention also includes equivalent technical means that those skilled in the art can think of based on the inventive concept.

Claims

1. A method for fusion of infrared and visible light feature levels, characterized in that: The following steps are involved: S1: Obtain infrared and visible light image datasets, preprocess the datasets, and obtain training sets; S2: Construct a network model based on feature-level infrared and visible light image fusion, including shallow feature extraction module HaarFeature Extraction, dense connection module SADense, infrared image enhancement module HaarImageEnhancement, and cross attention module WT CrossAttention; S3: Set network parameters, including: setting the initial learning rate Ir0, training rounds epochs, and batch size batch-size; S4: Based on the feature-level infrared and visible light image fusion network model constructed in S2, the training set in S1 is used for training to obtain the optimal weight file of the network model; S5: Apply the obtained weight file to image fusion, and finally input the obtained fused image into the downstream object detection task.

2. The method for fusion of infrared and visible light feature levels according to claim 1, characterized in that: The shallow feature extraction module Haar Feature Extraction uses Haar transformation to decompose the image and divide the visible light and infrared images into high and low frequency sub-bands; the shallow feature extraction module includes an infrared shallow feature extraction module and a visible light shallow feature extraction module, wherein the infrared feature extraction module includes two 3*3 convolutional layers and one 1*1 convolutional layer, and the visible light feature extraction module includes two 3*3 convolutional layers and one 1*1 convolutional layer; The dense connection module SADense realizes the reuse of shallow features by densely connecting shallow features. It includes a Sobel operator, a 1*1 convolution layer, two 3*3 convolution layers and a convolution block attention mechanism. The infrared image enhancement module Haar Image Enhancement decomposes the image through Haar transform, enhances the infrared background contour information, fits it with the detail information of visible light, and outputs an infrared feature enhanced image; The cross attention module Haar CrossAttention includes a channel dimension attention module, which consists of a spatial average pooling layer, two 1*1 convolution layers, two relu activation functions and a sigmoid activation function; a spatial dimension attention module, which consists of a channel average pooling layer, a 7*7 convolution layer, and a sigmoid activation function; and a transposition attention module, which consists of an image transposition layer and two 1*1 convolution layers. Wavelet transform is used to process high and low frequency information respectively.

3. The method for fusion of infrared and visible light feature levels according to claim 1, characterized in that: The specific contents of S1 are as follows: Step 101: Obtain a sensor data set; use visible light and infrared cameras to take pictures respectively to create a data set D.

4. The method for fusion of infrared and visible light feature levels according to claim 1, characterized in that: The specific contents of S2 are as follows: Step 201: Construct a visible light shallow feature extraction module. Use Haar transform on the visible light image to obtain a low-frequency subband LL and three high-frequency subbands HL, LH, and HH. Perform a second-order Haar transform on the HH high-frequency subband. Perform a 3*3 convolution operation on the four subbands obtained by the second-order transform. Perform an inverse Haar transform on the obtained feature map to restore the image. A 1*1 convolution is connected in series. The number of channels is adjusted to obtain a visible light shallow feature extraction feature map F that focuses on high-frequency signals. vi1 , so that the next layer of network can process; Step 202: Construct an infrared,feature extraction module. Use Haar transform on the visible light image to obtain a low-frequency subband LL, three high-frequency subbands HL, LH, HH. Perform a second-order Haar transform on the low-frequency subband LL. Perform a 3*3 convolution operation on the four subbands obtained by the second-order transform. Perform an inverse Haar transform on the obtained feature map to restore the image. A 1*1 convolution is connected in series. The number of channels is adjusted to obtain an infrared shallow feature extraction feature map F that focuses on low-frequency signals. ir1 , and then the next layer of network processing; Step 203: Construct a dense connection module SADense, and combine the feature maps F obtained in steps 201 and 202 above vi1 and F ir1 The features are reused on each branch, and connected to each subsequent layer on each single branch, and the first layer feature map F of size W*H*C is vi1 and F ir1 The weight information is added through 3*3 convolution respectively, as shown below: Where m, n represent the spatial dimensions of the convolution kernel, K(m,n,c,k) represents the learned weight information, and after the convolution operation, i, j represent the corresponding pixel coordinates on the feature map, 0≤i≤W, 0≤j≤H, and the weighted learned feature map F is obtained. vi2 and F ir2 , and then compared with the feature map F obtained in step 201 and step 202 vi1 and F ir1 Stitching: As shown above, we get a new feature map F of size W*H*(C+C1) viconcat1 and F irconcat1 , convolution weighted to obtain feature map F vi3 and F ir3 , and then reuse the feature map F obtained in step 201 and step 202 vi1 and F ir1 , and the feature map F vi3 and F ir3 Splice as follows: The size of the visible light and infrared branches is W*H*(C+C3) feature map F viconcat2 and F irconcat2 Input them into the convolution block attention mechanism respectively, and get the output feature map M vis and M irs , as follows: M vis =σ(Conv(F vi c)) M irs =σ(Conv(F irc )) Among them, F vic Yes F viconcat2 After the convolutional attention mechanism, the middle-level output feature map, F irc Yes F irconcat2 After the middle-layer output feature map of the convolutional attention mechanism, the nonlinear features are added through the activation function to obtain M vis and M irs ;σ represents the activation function, M vis and the feature map F obtained in step 201 vi1 The feature map F obtained by Sobel operator and convolution operation viSout Multiplication, M irs and the feature map F obtained in step 202 ir1 The feature map F obtained by Sobel operator and convolution operation irSout Multiply, as shown below: F visout =Conv(Concat(Sobel x (F vi1 ),Sobel y (F vi1 )),K conv ) F vifinal =F SCout ·F sout F irsout =Conv(Concat(Sobel x (F ir1 ),Sobel y (F ir1 )),K conv ) F irfinal =F SCout ·F sout Sobel x Represents the Sobel operator in the horizontal direction. y Represents the Sobel operator in the vertical direction, concatenated convolution, and finally obtains the above SADense dense block output feature map F vifinal and F irfinal . Step 204: Infrared image enhancement module Haar Image Enhancement, decomposes the image through Haar transformation, uses low-frequency signal LL to enhance the infrared background contour information, and high-frequency signals HL, LH, HH to enhance the detail information of visible light, and then fits. The enhancement method is the module described in step 205, and outputs an infrared feature enhanced image. Step 205: Construct a cross attention module Haar Cross Attention. The infrared and visible light feature maps are first transformed by Haar to obtain low-frequency sub-bands and high-frequency sub-bands. The low-frequency signal emphasizes the correlation, and the low-frequency sub-band feature maps are added. The high-frequency signal emphasizes the irrelevance, and the high-frequency sub-band feature maps are subtracted. The obtained feature maps are input into the cross module; the input feature maps are processed in parallel, focusing on the spatial and channel dimensions respectively. The channel dimension is first average pooled for the feature map, as shown in the following formula: A feature map of size 1*1*C can be obtained. Each channel contains the global channel information of its own space. After 1*1 convolution, the linear transformation of each channel is completed, and the channel dimension is reduced and the redundant channel dimension is removed, as shown in the following formula: Introduce nonlinear transformation and implement it through the activation function Relu: Then, through a layer of 1*1 convolution operation and the introduction of nonlinear activation function, the feature map dimension is restored, and the channels are reintegrated and reconstructed, which is used for the subsequent Sigmoid normalization processing to compress the weights into the range of [0,1]. In the spatial dimension, the channel dimension of the feature map is first average pooled, as shown in the following formula: A feature map of size w*h*1 is obtained. Each pixel in the space contains global channel information. The spatial dimension requires a larger receptive field to enhance the connection between features in different spaces. Therefore, a large kernel convolution of 7*7 is used, as shown in the following formula: Among them, in order to prevent the convolution process from crossing the boundary and still output the feature map of size w*h, padding is required. p represents the extra pixels added to the edge of the data. Here p is set to 3. The feature map processed in the spatial dimension is input into the Sigmoid activation function, the weight is normalized, and multiplied with the tensor in the channel dimension to generate a feature map that focuses on both the channel and spatial dimensions, as shown in the following formula: Y(w,h,c)=A(1,1,c)×B(w,h,1) Finally, we get the feature map F processed in both spatial and channel dimensions. SC , starting from its own global information, and then processing in parallel with the transposed attention module, by transposing the image itself, then splicing, and then convolving it with F in an appropriate weight ratio SC Stitching, weighting is completed, and then Haar inverse transform is performed to restore the image.

5. An infrared and visible light feature-level fusion system, characterized in that: include: Data processing module: using a camera to collect images and obtain the processed data; Network building module: builds a network model based on feature-level infrared and visible light image fusion, including shallow feature extraction module Haar Feature Extraction, dense connection module SADense, infrared image enhancement module Haar image enhancement, cross attention module Haar Cross Attention. Network training module: used to train on the training set according to the set network parameters, output the training weight file, and obtain the required optimal image fusion weight file; Output module: Deploy the feature-level infrared and visible light image fusion network model on the device, build a post-processing unit, and output the fused image to the downstream target detection model.

6. An infrared and visible light feature level fusion device, characterized in that: include: Machine executable instructions: a computer program storing a method for fusion of infrared and visible light feature levels as described in any one of claims 1 to 4, which is a computer-readable executable instruction; Processor: a processor used to execute the computer program to implement the infrared and visible light feature-level fusion method as described in any one of claims 1 to 4; A computer-readable storage medium: the computer-readable storage medium stores a computer program, and when the computer program is processed by a processor, the storage medium can store the infrared and visible light feature level fusion method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Infrared-visible light image target level fusion method based on full convolutional neural network

    CN116246138A

  • Infrared and visible light fusion method based on multi-feature extraction

    CN119205526A

  • Infrared and visible light image fusion method and system based on high and low frequency separation enhancement

    CN119206419A

  • Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network

    WO2024174488A1

Cited By

  • Infrared and visible light image fusion method and device based on lightweight model

    CN121329792A