Infrared and visible light image fusion method in low-illumination scene

Through Retinex theory and the cross-modal feature guidance module, the reflective component feature extraction is guided in low-illumination scenarios, and the texture information is enhanced, and the problem of poor fusion effect of infrared and visible light images in the prior art is solved, achieving high-quality image fusion at low-illumination.

CN120259097APending Publication Date: 2025-07-04NORTHWEST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510326010.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing infrared and visible light image fusion methods are not effective in low-illumination scenarios, affecting human eye perception and index evaluation.

Method used

The image fusion method is designed using Retinex theory, and the cross-modal feature guidance module is used to guide the extraction of reflected component features before the visible light image is fused with the infrared image, and the texture information in the mixed features is enhanced through the texture enhancement fusion module, discarding the illumination component features.

Benefits of technology

In low-illumination scenarios, the fused image has good visibility and texture details, which improves the visual effect and objective evaluation indicators of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259097A_ABST
    Figure CN120259097A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared and visible light image fusion method in a low-illumination scene, and the method comprises the steps: extracting a reflection component feature of a visible light sample image and an infrared image feature of an infrared sample image through an image processing module, and carrying out the weighting processing of the reflection component feature through the extracted infrared image feature; then connecting the reflection component features and the infrared image features output by the image processing module in series in a channel dimension through a texture enhancement fusion module to obtain mixed features, inputting the mixed features to a main stream and a residual stream, and multiplying output results of the main stream and the residual stream to obtain texture enhancement image features; and the fused image decoder reconstructs the texture enhanced image features to obtain a fused image. Therefore, before fusion of the visible light image and the infrared image, the cross-modal feature guidance module is utilized to guide extraction of reflection component features, texture information in mixed features is enhanced, and it is guaranteed that the reconstructed fusion image has good visibility in a low-illumination scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to an infrared and visible light image fusion method in low-light scenes. Background Art

[0002] At present, images captured by single-modal sensors cannot effectively and comprehensively describe the imaging scene; visible light sensors image through reflected light and can provide background details with high spatial resolution, but when the lighting or camouflage conditions are poor, they cannot clearly see the target; in contrast, infrared sensors image through the distinguishable thermal radiation emitted by objects, are not affected by harsh conditions, and can work all-weather. This complementary characteristic has given rise to the infrared-visible light image fusion. The fused image can highlight the target and show rich texture details. Therefore, infrared and visible light image fusion has been widely used as a preprocessing module for advanced vision tasks, such as target detection, pedestrian re-identification, and semantic segmentation.

[0003] Existing methods often are designed for normal lighting scenes, ignoring the problem of poor visibility in low-light scenes. If the low-light enhancement task and the image task are simply superimposed, that is, first using a low-light enhancement algorithm to enhance the visible light image and then fusing it with the infrared image, it often results in poor visual effects of the fusion result, affecting human eye perception and index evaluation. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an infrared and visible light image fusion method in low-light scenes to solve the technical problems mentioned in the prior art.

[0005] An infrared and visible light image fusion method in low-light scenes, the image fusion method includes the following steps:

[0006] S1. Obtain a data set, and arbitrarily select a pair of the visible light sample image and the infrared sample image from several pairs of visible light sample images and infrared sample images in the data set as a test set;

[0007] S2. An image processing module respectively extracts the reflection component feature of the visible light sample image and the infrared image feature of the infrared sample image from the test set;

[0008] S3. A cross-modal feature guidance module uses the infrared image feature extracted by the image processing module to perform weighted processing on the reflection component feature;

[0009] S4. The texture enhancement fusion module concatenates the reflection component features and the infrared image features output by the image processing module in the channel dimension to obtain hybrid features, inputs the hybrid features into the main stream and the residual stream respectively, and then multiplies the output results of the main stream and the residual stream to obtain texture enhanced image features;

[0010] And / or, the image processing module reconstructs the reflection component features or the infrared image features to obtain a reconstructed image;

[0011] S5. The fusion image decoder receives the texture enhanced image features output by the texture enhancement fusion module and performs reconstruction to obtain a fusion image.

[0012] Optionally, in S2, the image processing module is set as a reflection component codec, an illumination component codec, and / or an infrared image codec, where;

[0013] The reflection component codec is used to extract reflection component features from a visible light image and reconstruct a reflection image; the illumination component codec is used to extract illumination component features from a visible light image and reconstruct an illumination image; the infrared image codec is used to extract infrared image features from an infrared image and reconstruct an infrared image.

[0014] Optionally, the image processing module includes an encoder and a decoder, and both the encoder and the decoder are provided with several convolutional blocks; the cross-modal feature guidance module performs weighted processing on the reflection component features and the infrared image features input to each convolutional block in the encoder of the image processing module, and uses it as the input to the next convolutional block in the encoder corresponding to the reflection component features.

[0015] Optionally, in S3, the method by which the cross-modal feature guidance module performs weighted processing on the reflection component features and the infrared image features input to each convolutional block in the encoder of the image processing module includes:

[0016] S3-01. Downsample the reflection component features and the infrared image features input to the previous convolutional block respectively;

[0017] S3-02. The downsampled reflection component features and infrared image features sequentially pass through a convolutional layer and a reshaping operation, so that the reflection component features generate a query, and the infrared image features generate a key and a value;

[0018] S3-03. Multiply the query and the key in matrix form to obtain an attention map;

[0019] S3-04. Multiply the attention map after softmax by the value, then sequentially pass the multiplication result through a reshaping operation, upsampling, and a convolutional layer, and add it to the reflection component feature.

[0020] Optionally, the encoder is provided with five convolutional blocks, and each convolutional block is respectively set as a combined structure of a first convolutional layer and a normalization layer.

[0021] Optionally, the decoder is provided with four convolutional blocks, and the four convolutional blocks are sequentially set as:

[0022] A second convolutional layer, a batch normalization layer, a ReLu activation function, and a combined structure of a second convolutional layer and a Sigmod activation function.

[0023] Optionally, in S4, the main stream is sequentially provided with a dense connection block and a fourth convolutional layer along the data transmission direction;

[0024] The fourth convolutional layer is used to adjust the number of channels.

[0025] Optionally, in S4, the residual stream is sequentially provided with a Sobel operator, global average pooling, a fully connected layer, a ReLu activation function, a fully connected layer, and a Sigmod activation function along the data transmission direction.

[0026] Optionally, in S5, the fusion image decoder has four convolutional blocks, and the four convolutional blocks are sequentially set as:

[0027] A third convolutional layer, a batch normalization layer, a LeakyReLu activation function, and a combined structure of a third convolutional layer and a Sigmod activation function.

[0028] Optionally, the dataset is trained using a loss function to realize the update iteration of the image fusion method, and the loss function L total is:

[0029] L total = L Retinex + L fuse + α·L contrast ;

[0030] where L Retinex is the Retinex loss, L fuse is the fusion loss, L contrast is the contrast loss, and α is the first hyperparameter;

[0031] The Retinex loss is:

[0032] L Retinex = L ill + L de;

[0033]

[0034] Among them, L ill is the loss of the illumination component, I vi is the visible light image, is the illumination component of the visible light image, is the reflection component of the visible light image, L TV is the total variation loss, L de is the decoupling loss, SSIM(·) is the structural similarity calculation, I ir is the infrared image, is the reconstructed infrared image, and β is the second hyperparameter;

[0035] The fusion loss is:

[0036] L fuse = λ1·L int + λ2·L text + λ3·L corr ;

[0037]

[0038] Among them, λ1, λ2, and λ3 are the third hyperparameters respectively, L int is the intensity loss, L text is the texture loss, L corr is the correlation loss, ▽ is the Sobel operator, and corr(·) is the correlation calculation;

[0039] The contrast loss is:

[0040]

[0041] Among them, σ(·) is the standard deviation calculation, and μ is the fourth hyperparameter.

[0042] The beneficial effects that the present invention can produce include:

[0043] A method for fusing infrared and visible light images in a low-illumination scene provided by the present invention designs an image fusion method using the Retinex theory, extracts the reflection component features by using a cross-modal feature guidance module before fusing the visible light image and the infrared image, and discards the illumination component features, so that the texture enhancement fusion module fuses the reflection component features and the infrared image features to obtain hybrid features, and enhances the texture information in the hybrid features, thereby ensuring that the reconstructed fusion image can achieve the effect of enhancing the visible light image under low illumination. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1Schematic flow chart of an infrared and visible light image fusion method in low - illumination scenarios according to the present invention;

[0045] Figure 2 In the present invention Figure 1 Schematic framework diagram of the cross - modal feature guidance module;

[0046] Figure 3 In the present invention Figure 1 Schematic framework diagram of the texture enhancement fusion module;

[0047] Figure 4 In the present invention Figure 1 Schematic diagram of the comparison results between the image fusion method of the present invention and existing image fusion algorithms for the MSRS dataset;

[0048] Figure 5 In the present invention Figure 1 Schematic diagram of the comparison results between the image fusion method of the present invention and existing image fusion algorithms for the LLVIP dataset. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Please refer to Figure 1 As shown, the present invention provides an infrared and visible light image fusion method in low - illumination scenarios. The image fusion method includes the following steps:

[0051] Step 1: Obtain a dataset, and arbitrarily select a pair of visible light sample images and infrared sample images from several pairs of visible light sample images and infrared sample images in the dataset as a test set;

[0052] Step 2: The image processing module respectively extracts the reflection component features of the visible light sample image and the infrared image features of the infrared sample image from the test set;

[0053] Step 3: The cross - modal feature guidance module uses the infrared image features extracted by the image processing module to perform weighted processing on the reflection component features;

[0054] Step 4: The texture enhancement fusion module concatenates the reflection component features and the infrared image features output by the image processing module in the channel dimension to obtain mixed features, and inputs the mixed features into the main stream and the residual stream respectively, and then multiplies the output results of the main stream and the residual stream to obtain texture - enhanced image features;

[0055] And / or, the image processing module reconstructs the reflection component feature or the infrared image feature to obtain a reconstructed image;

[0056] Step Five: The fused image decoder receives the texture-enhanced image feature output by the texture enhancement fusion module and performs reconstruction to obtain a fused image.

[0057] In the above, the present invention designs an image fusion method using the Retinex theory, so as to guide the extraction of the reflection component feature by using the cross-modal feature guidance module before fusing the visible light image and the infrared image, and discard the illumination component feature, so that the texture enhancement fusion module fuses the reflection component feature and the infrared image feature to obtain a hybrid feature, and enhances the texture information in the hybrid feature, thereby ensuring that the reconstructed fused image has good visibility in low-illumination scenarios.

[0058] Further, the image processing module is set as a reflection component codec, an illumination component codec, and / or an infrared image codec, where; the reflection component codec is used to extract the reflection component feature in the visible light image and reconstruct the reflection image; the illumination component codec is used to extract the illumination component feature in the visible light image and reconstruct the illumination image; the infrared image codec is used to extract the infrared image feature in the infrared image and reconstruct the infrared image.

[0059] Further, the image processing module includes an encoder and a decoder, and both the encoder and the decoder are provided with several convolutional blocks; the cross-modal feature guidance module performs weighted processing on the reflection component feature and the infrared image feature input to each convolutional block in the encoder of the image processing module, and uses it as the input to the next convolutional block in the encoder corresponding to the reflection component feature.

[0060] Further, the method for the cross-modal feature guidance module to perform weighted processing on the reflection component feature and the infrared image feature input to each convolutional block in the encoder of the image processing module includes: respectively performing downsampling on the reflection component feature and the infrared image feature input to the previous convolutional block; sequentially passing the downsampled reflection component feature and infrared image feature through a convolutional layer and a reshaping operation, so that the reflection component feature generates a query, and the infrared image feature generates a key and a value; multiplying the query and the key to obtain an attention map; passing the attention map through softmax and then multiplying it with the value, and then sequentially passing the multiplication result through a reshaping operation, upsampling, and a convolutional layer, and adding it to the reflection component feature.

[0061] Further, the encoder is provided with five convolutional blocks, and each convolutional block is respectively set as a combined structure of a first convolutional layer and a normalization layer.

[0062] Further, the decoder is provided with four convolutional blocks, and the four convolutional blocks are sequentially set as:

[0063] The second convolutional layer, batch normalization layer, ReLu activation function, and the combined structure of the second convolutional layer and the Sigmod activation function.

[0064] Further, in S4, along the data transmission direction, a dense connection block and a fourth convolutional layer are sequentially arranged in the mainstream;

[0065] The fourth convolutional layer is used to adjust the number of channels.

[0066] Further, along the data transmission direction of the residual flow, a Sobel operator, global average pooling, fully connected layer, ReLu activation function, fully connected layer, and Sigmod activation function are sequentially arranged.

[0067] Further, the fusion image decoder has four convolutional blocks, and the four convolutional blocks are sequentially arranged as: the third convolutional layer, batch normalization layer, LeakyReLu activation function, and the combined structure of the third convolutional layer and the Sigmod activation function.

[0068] Further, the loss function is used to train the data set to realize the update and iteration of the image fusion method, and the loss function L total is:

[0069] L total = L Retinex + L fuse + α·L contrast ;

[0070] where, L Retinex is the Retinex loss, L fuse is the fusion loss, L contrast is the contrast loss, and α is the first hyperparameter;

[0071] The Retinex loss is:

[0072] L Rstinex = L ill + L de ;

[0073]

[0074] where, L ill is the illumination component loss, I vi is the visible light image, is the illumination component of the visible light image, is the reflection component of the visible light image, L TV is the total variation loss, L de is the decoupling loss, SSIM(·) is the structural similarity calculation, I ir is the infrared image, is the reconstructed infrared image, and β is the second hyperparameter;

[0075] The fusion loss is:

[0076] L fuse = λ1·L int + λ2·L text + λ3·L corr ;

[0077]

[0078] where λ1, λ2, and λ3 are the third hyperparameters respectively, L int is the intensity loss, L text is the texture loss, L corr is the correlation loss, ▽ is the Sobel operator, and corr(·) is the correlation calculation;

[0079] The contrast loss is:

[0080]

[0081] where σ(·) is the standard deviation calculation and μ is the fourth hyperparameter.

[0082] In the above, in order to verify the fusion effect of an infrared and visible light image fusion method in a low-light scene of the present invention, the following simulation experiment is used to verify its effectiveness:

[0083] First, in an embodiment of the present application, the infrared image and visible light image data in the LLVIP dataset are subjected to image fusion by the above image fusion method, and the specific process is as follows:

[0084] Select 100 pairs of sample images from the LLVIP dataset to train the image fusion method. Each sample image is cropped into 20 equal parts of size 256×256. Therefore, the sample images are expanded to 2000 pairs. The network optimizer uses Adam, epoch = 40, batchsize = 12, and the learning rate is 1×10-3. Here, epoch is for deep learning, and 1 epoch is equal to training once with all samples in the training set. Batchsize is the number of samples input to the model at one time during the training of the deep learning model; and linear attenuation starts in the middle of the training. The first hyperparameter α of the loss function is 2.5, the second hyperparameter β is 10, the third hyperparameters are set respectively as: λ1 = 1, λ2 = 20, λ3 = 3, and the fourth hyperparameter is 5; it should be noted that the optimizer Adam has the characteristics of small memory occupation and being suitable for dealing with sparse gradients, and it is currently a widely used and better-performing optimization algorithm in various fields; the test set selects the public datasets MSRS and LLVIP; in the experiment, all comparison algorithms use the relatively common image fusion algorithms in the existing technology and the above image fusion method to be trained on NVIDIA GeForce RTX 4090 (graphics card) and Intel Core i9-13900K (Intel Core 13th generation CPU product) using the Pytorch (an open-source Python machine learning library) framework.

[0085] Specifically, conduct an effect analysis on the fusion results of the simulated experimental images:

[0086] Compare the above image fusion method with the nine most advanced image fusion algorithms in the existing technology. The nine image fusion algorithms are BTSFusion (infrared and visible light image fusion network), LRRNet (novel representation learning-guided fusion network for infrared and visible light image fusion), FLFuse (a fast and lightweight infrared and visible light image fusion network), GANMcC (image fusion network), STDFusionNet (infrared and visible light image fusion network), FusionGAN (a generative infrared and visible light image fusion adversarial network), U2Fusion (multi-source image fusion network), LENFusion (night image fusion network), and DIVFusion (coupled mutual promotion low-light enhancement & image fusion network); as Figure 4 shown are the comparison experiment results of different image fusion methods on the MSRS dataset in the simulated experiment, as Figure 5The following shows the comparative experimental results of different image fusion methods on the LLVIP dataset in the simulation experiment; it should be noted that since visible light sensors rely on light reflection to form images, in low light conditions, the image detail information seriously degrades, and when the image part is completely submerged in darkness, the information in the dark area is completely lost. At the same time, infrared sensors form images by capturing thermal radiation and are not affected by light. Therefore, infrared images contain significant targets and extensive texture details, which can supplement the texture degradation and detail loss problems of visible light images.

[0087] In the above, according to the experimental comparison results, it can be seen that reasonably integrating the texture information of infrared images and visible light images is a challenge. For example, in Figure 4 the scene, LRRNet and U2Fusion cause the intensity weakening of significant targets, and BTSFusion, FLFuse, GANMcC, STDFusionNet, and FusionGAN all further cause the information loss in the dark area. Although LENFusion and DIVFusion reasonably enhance the dark part information of the image, they weaken the significant targets; while the fusion method of the present invention can reasonably enhance the dark area of the image and utilize the detail texture information between the two modalities to maximize the retention of detail information in the fusion result and highlight the significant targets. Similarly, in Figure 5 the scene, only the image fusion method of the present invention achieves the best contrast while realizing image fusion, has the most friendly visual effect, and this phenomenon can be specifically seen from the white fence in the upper left corner and the text area in the middle of the image, and there is no intensity weakening of significant targets.

[0088] Specifically, visual evaluation and comparison can provide a relatively intuitive understanding of the image fusion results. However, it is difficult to give the most accurate judgment on the image fusion results solely relying on subjective evaluation. It is necessary to combine objective indicators to jointly evaluate the image fusion results. Therefore, the present invention designs multiple key performance indicators to evaluate and quantify the effectiveness of the model. Specifically, six evaluation indicators are selected for quantitative evaluation, and the evaluation results are shown in Table 1. The six evaluation indicators respectively include information entropy EN, spatial frequency SF, average gradient AG, standard deviation SD, visual information fidelity VIF, and sum of differential correlation SCD. Among them, the information entropy EN calculates the amount of information contained in the fused image based on information theory; the spatial frequency SF reflects the richness of the image frequency by measuring the row frequency and column frequency of the fused image; the average gradient AG characterizes the texture details of the fused image by measuring the gradient information of the fused image; the standard deviation SD is an indicator reflecting the contrast and distribution of the fused image; the visual information fidelity VIF evaluates the information fidelity of the fused image from the perspective of the human visual system; the sum of differential correlation SCD characterizes the quality of the fusion algorithm by measuring the difference between the fused image and the source image. In addition, the indicators used are all positive indicators, that is, the larger the indicator value, the better the fusion performance.

[0089] Table 1

[0090]

[0091]

[0092] Specifically, as shown in Table 1, the comparison data of the fusion method of the present application and several other image fusion algorithms in various evaluation indicators are listed in detail, so as to comprehensively display the performance and advantages of the fusion method of the present application. It is worth noting that the fusion method of the present application has achieved great improvement in each indicator, which shows that the fusion method of the present application can obtain excellent fused image results. Combining visual evaluation and objective indicator evaluation, this fusion method can integrate detailed information from source images of two modalities, reasonably enhance the images, and generate fused images with rich details.

Claims

1. An infrared and visible light image fusion method in a low illumination scene, characterized in that, The described image fusion method includes the following steps: S1. Obtain a data set, and arbitrarily select a pair of the visible light sample image and the infrared sample image from several pairs of visible light sample images and infrared sample images in the data set as a test set; S2. The image processing module respectively extracts the reflection component features of the visible light sample image and the infrared image features of the infrared sample image from the test set; S3. The cross-modal feature guidance module uses the infrared image features extracted by the image processing module to perform weighted processing on the reflection component features; S4. The texture enhancement fusion module concatenates the reflection component features and the infrared image features output by the image processing module in the channel dimension to obtain mixed features, and respectively inputs the mixed features into the main stream and the residual stream, and then multiplies the output results of the main stream and the residual stream to obtain texture enhanced image features; and / or, the image processing module reconstructs the reflection component features or the infrared image features to obtain a reconstructed image; S5. The fusion image decoder receives the texture enhanced image features output by the texture enhancement fusion module and performs reconstruction to obtain a fusion image.

2. The infrared and visible light image fusion method in a low illuminance scenario according to claim 1, characterized in that, In S2, the image processing module is set as a reflection component codec, a lighting component codec, and / or an infrared image codec, where; The reflection component codec is used to extract the reflection component features in the visible light image and reconstruct the reflection image; the lighting component codec is used to extract the lighting component features in the visible light image and reconstruct the lighting image; the infrared image codec is used to extract the infrared image features in the infrared image and reconstruct the infrared image.

3. A method for fusing infrared and visible light images in a low illuminance scenario according to claim 1, characterized in that, The image processing module includes an encoder and a decoder, and both the encoder and the decoder are provided with several convolutional blocks; the cross-modal feature guidance module performs weighted processing on the reflection component features and the infrared image features input to each convolutional block in the encoder of the image processing module, and uses the result as the input to the next convolutional block in the encoder corresponding to the reflection component features.

4. The infrared and visible light image fusion method in a low illumination scene according to claim 3, wherein, In S3, the method by which the cross-modal feature guidance module performs weighted processing on the reflection component features and the infrared image features input to each convolutional block in the encoder of the image processing module includes: S3-01. Respectively perform downsampling on the reflection component features and the infrared image features input to the previous convolutional block; S3-02. The reflection component features and the infrared image features obtained by downsampling sequentially pass through a convolutional layer and a reshaping operation, so that the reflection component features generate queries, and the infrared image features generate keys and values; S3-03. Multiply the query and the key to obtain an attention map; S3-04. After passing the attention map through softmax, multiply it by the value, and then sequentially pass the multiplication result through a reshaping operation, upsampling, and a convolutional layer, and add it to the reflection component features.

5. The infrared and visible light image fusion method in a low illumination scene according to claim 3, characterized in that, The encoder is provided with five convolutional blocks, and each convolutional block is respectively set as a combined structure of a first convolutional layer and a normalization layer.

6. The infrared and visible light image fusion method in a low illuminance scenario according to claim 3, characterized in that, The decoder is provided with four layers of the convolutional blocks, and the four layers of the convolutional blocks are sequentially arranged as follows: The second convolutional layer, the batch normalization layer, the ReLu activation function, and the combined structure of the second convolutional layer and the Sigmod activation function.

7. A method for fusing infrared and visible light images in a low illumination scene according to claim 1, characterized in that, In the S4, a dense connection block and a fourth convolutional layer are sequentially arranged along the data transmission direction of the mainstream; The fourth convolutional layer is used to adjust the number of channels.

8. A method for fusing infrared and visible light images in a low illuminance scenario according to claim 1, characterized in that, In the S4, a Sobel operator, a global average pooling, a fully connected layer, a ReLu activation function, a fully connected layer, and a Sigmod activation function are sequentially arranged along the data transmission direction of the residual flow.

9. A method for fusing infrared and visible light images in a low illumination scene according to claim 1, characterized in that, In the S5, the fusion image decoder has four layers of convolutional blocks, and the four layers of the convolutional blocks are sequentially arranged as follows: The third convolutional layer, the batch normalization layer, the LeakyReLu activation function, and the combined structure of the third convolutional layer and the Sigmod activation function.

10. The infrared and visible light image fusion method in a low illumination scene according to claim 1, characterized in that, Training the dataset using a loss function to achieve iterative updates of the image fusion method, where the loss function L total is as follows: L total = L Retinex + L fuse + α·L contrast ; Among them, L Retinex is the Retinex loss, L fuse is the fusion loss, L contrast is the contrast loss, and α is the first hyperparameter; The Retinex loss is: L Retinex = L ill + L de ; Among them, L ill is the loss of the illumination component, I vi is the visible light image, is the illumination component of the visible light image, is the reflection component of the visible light image, L TV is the total variation loss, L de is the decoupling loss, SSIM(·) is the structural similarity calculation, I ir is the infrared image, is the reconstructed infrared image, and β is the second hyperparameter; The fusion loss is: L fuse = λ1·L int + λ2·L text + λ3·L corr ; Among them, λ1, λ2, and λ3 are the third hyperparameters respectively, and L int is the intensity loss, and L text is the texture loss, and L corr is the correlation loss, ▽ is the Sobel operator, and corr(·) is the correlation calculation; The contrast loss is: Wherein, σ(·) is the standard deviation calculation, and μ is the fourth hyperparameter.

Citation Information

Cited By

  • Extreme illumination-oriented visible light and infrared image fusion method, system and device based on text-guided spatial frequency domain interaction and medium

    CN121214128A