Low-illumination visible light and infrared image fusion method with color recovery capability
By combining a low-light enhancement network and a fusion network based on the Mamba architecture with contrastive learning and a hybrid attention mechanism, the problems of color shift and computational efficiency in low-light infrared-visible image fusion are solved, achieving high-quality image fusion results.
Patent Information
- Application Number
- CN202511845255.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-17
AI Technical Summary
Existing low-light infrared-visible image fusion methods are prone to introducing color shifts in low-light environments, and cannot balance global receptive field and computational efficiency, resulting in a decrease in the quality of the fused image.
We employ a low-light enhancement network and a fusion network based on the Mamba architecture. Through contrastive learning and hybrid attention mechanisms, we process the illumination and color components of visible light images separately, and fuse them with infrared images. We then use the Mamba module to decode and enhance the features.
It achieves the restoration of color fidelity and texture details in low-light environments, improves the quality of fused images, and meets the effect and efficiency requirements of practical applications.
Smart Images

Figure CN121685279A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision technology, and in particular relates to a method for fusing low-light visible light and infrared images with color recovery capabilities. Background Technology
[0002] In the fields of computer vision and multimodal image processing, sensors are often limited by physical performance, technical constraints, and harsh environmental factors, resulting in incomplete information capture in single-modal imaging. Image fusion technology, as a cutting-edge technology, can integrate complementary information from multimodal images to generate information-rich fused images, effectively solving the problem of incomplete information capture in single-modal sensor imaging.
[0003] Visible light images and infrared images are two widely used image modalities. In well-lit environments, visible light images accurately reproduce the texture details and color information of the environment; however, in dim environments, visible light images often suffer from detail degradation, color distortion, and low contrast. Infrared images, based on the thermal radiation of objects, can highlight salient targets in various environments, an advantage that visible light images cannot match; however, infrared images suffer from low resolution, blurred details, and a lack of color information. Therefore, infrared-visible image fusion (IVIF) technology combines the unique advantages of infrared and visible light images to generate a fused image that combines salient targets with detailed backgrounds.
[0004] Currently, numerous infrared-visible light image fusion methods have been proposed. However, most existing mainstream image fusion technologies are designed for well-lit scenes. In low-light environments, visible light images are prone to color distortion and texture blurring, which directly leads to a decrease in the quality of the fused image.
[0005] To address the degradation of visible light images during infrared-visible image fusion in low-light scenarios, some low-light infrared-visible image fusion methods incorporate Low-Light Image Enhancement (LLIE) technology into the infrared-visible image fusion model. This improves the quality of the fused image by correcting the illumination conditions of the visible light image. However, these existing low-light infrared-visible image fusion methods still have significant drawbacks: firstly, the unsupervised low-light enhancement technology they employ easily introduces color shift issues, affecting the color fidelity of the fused image; secondly, existing methods cannot balance the global receptive field with a small computational load, making it difficult to meet the dual requirements of fusion effect and computational efficiency in practical applications. Summary of the Invention
[0006] The purpose of this invention is to address the problems of low contrast, color distortion, and loss of detail in infrared-visible image fusion images under low-light conditions, and to propose a low-light visible light and infrared image fusion method with color restoration.
[0007] This invention proposes a method for fusing low-light visible light and infrared images with color restoration capabilities. The method includes: Construct a contrast learning dataset by acquiring well-exposed and low-light images; A low-light enhancement network based on the Mamba structure is constructed and trained using a low-light enhancement loss based on contrastive learning; the low-light enhancement loss includes a color distribution alignment loss. The low-light enhancement network is used to convert low-light images to YCbCr space, and the Mamba module is used to model the high-frequency and low-frequency features of the converted image separately, and the channel values of the reconstructed image are iteratively enhanced. A fusion network based on the Mamba structure is constructed and trained. The fusion network is used to extract modal features from infrared and visible light images. Modal features are obtained by a hybrid attention mechanism. The modal features are decoded into the luminance channel and the enhanced channel value of the fused image, and then converted into an RGB image. The low-light enhancement network and fusion network trained by the network are used to achieve color restoration by fusing low-light visible light and infrared images.
[0008] Furthermore, the construction of the contrastive learning dataset specifically includes: The CbCr channels of well-exposed images are selected as positive samples for contrastive learning, and the CbCr channels of overexposed / underexposed images are selected as negative samples for contrastive learning. The images of positive and negative samples are then sized and aligned.
[0009] Furthermore, the color distribution alignment loss is specifically as follows: ; In the above formula, This represents the color channels of the enhanced visible light image. The color channels represent the color channels of a normal light image. Indicates the color channels of an underexposed / overexposed image. The color channels represent low-light images. This represents the discrete cosine similarity.
[0010] Furthermore, the low-light enhancement loss is specifically obtained by weighting smoothing loss, spatial consistency loss, exposure loss, and color distribution alignment loss. The smoothing loss suppresses outliers in the enhancement. The spatial consistency loss calculates the difference between the low-light image and the enhanced image in each local region and its adjacent regions. The exposure loss calculates the difference between the average pixel intensity within the calculation window and a preset threshold to control the brightness of the local region.
[0011] Furthermore, the Mamba module specifically includes: First, discrete wavelet transform is used to extract the input features. Decomposed into low-frequency features and high frequency characteristics Remodeled as and ; Modeling low-frequency features: ; In the above formula, EinFFT It is the Einstein Fourier transform, Ln Presentation layer normalization operation, ASSM ( This indicates the use of a Mamba architecture based on non-causal unidirectional scanning; The mathematical model for modeling high-frequency features is as follows: ; In the above formula, Split ReLU represents the segmentation of features along the channels. This indicates the use of the ReLU activation function. This indicates the use of pointwise convolution. This indicates that a depthwise convolution operation can be separable using a kernel size of 5x5; The reconstructed low-frequency and high-frequency features are transformed into output features through inverse wavelet transform.
[0012] Furthermore, the iterative enhancement specifically includes: The output feature is divided into three sub-features along the channel. The number of iterations is determined by the dimension of the sub-features. The enhancement formula for the nth iteration is: ; ; ; ; in This represents the Y / Cb / Cr channels of an input low-light visible light image. Indicates the use for the nth time Features during channel iteration This represents the Y channel value of the image after n iterations. This represents the Cb / Cr value of the image after n iterations. This represents the Y / Cb / Cr values of the final enhanced visible light image after 8 iterations of enhancement.
[0013] Furthermore, the fusion network based on the Mamba structure consists of two encoders based on Mamba modules, feature fusion based on a hybrid attention mechanism, and a decoder based on Mamba modules. The Mamba-based encoder uses the Mamba module and pointwise convolutional layers to extract modal-specific features from infrared and visible light images. The feature fusion based on the hybrid attention mechanism first integrates the features extracted by the encoder. and Subtraction yields residual characteristics residual characteristics Channel weights are obtained through a hybrid attention mechanism. and spatial weights Modal weights are calculated based on channel weights and spatial weights, and then the visible light features and infrared features are fused using the modal weights to obtain fused features. The Mamba module-based decoder uses the Mamba module, pointwise convolutional layers, and Tanh activation function to decode the fusion features into the luminance channel values of the fused image. The Cb and Cr channels of the fused image directly adopt the Cb and Cr channel values corresponding to the enhanced visible light image to obtain the final fusion result.
[0014] Furthermore, the training of the fusion network based on the Mamba structure includes structural loss, strength loss, and gradient loss; The structural loss restores the structural information of the original image, the intensity loss causes the fused image to contain the most significant intensity information of each pixel, and the gradient loss ensures that the fused image retains more detail and texture information.
[0015] On the other hand, the specification also provides a low-light visible light and infrared image fusion device with color restoration capability, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the low-light visible light and infrared image fusion method with color restoration capability.
[0016] On the other hand, the specification also provides a computer-readable storage medium on which a program is stored, which, when executed by a processor, implements the described method for fusing low-light visible light and infrared images with color restoration capabilities.
[0017] The beneficial effects of this invention are: 1. Design a low-light enhancement loss function based on contrastive learning to align the color distribution of the enhanced visible light image with the color distribution of the ideal image, thus ensuring color fidelity.
[0018] 2. Design a low-light enhancement neural network based on the Mamba structure and convert low-light images into the YCbCr color space. Decouple the illumination component and color component of the low-light image and perform enhancement processing on them separately.
[0019] 3. A fusion network based on the Mamba structure is designed. Modal complementary information is obtained through a hybrid attention mechanism to fuse the enhanced visible light image and infrared image, which effectively enhances the texture details and color information of the low-light area in the fused image.
[0020] 4. The low-light enhancement network and fusion network adopt the Mamba structure, which has a global receptive field, meeting the dual requirements of fusion effect and computational efficiency in practical applications. Attached Figure Description
[0021] Figure 1 This invention provides a method for fusing low-light visible light and infrared images with color restoration capabilities, as part of an embodiment of the present invention. flow chart; Figure 2 This is a diagram of a low-light enhancement network structure provided in an embodiment of the present invention; Figure 3 This is a structural diagram of the Mamba module provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of YCbCr color space enhancement provided in an embodiment of the present invention; Figure 5 A fusion network structure diagram provided in an embodiment of the present invention; Figure 6 The encoder structure diagram based on the Mamba module provided in the embodiment of the present invention; Figure 7 This is a feature fusion structure diagram based on hybrid attention provided in an embodiment of the present invention; Figure 8 This is a channel attention structure diagram provided in an embodiment of the present invention; Figure 9 Spatial attention structure diagram provided for embodiments of the present invention; Figure 10 This is a diagram of a decoder structure based on a Mamba module provided in an embodiment of the present invention; Figure 11 This is a qualitative result of an embodiment of the present invention on the LLVIP dataset; Figure 12 This is a qualitative result of the embodiments of the present invention on the MSRS dataset; Figure 13 This is a schematic diagram of a low-light visible light and infrared image fusion device with color restoration capability, provided in an embodiment of the present invention. Detailed Implementation
[0022] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0023] The low-light visible light and infrared image fusion method with color restoration capability in this embodiment includes: S1, Design a low-light enhancement loss based on contrastive learning to align the color distribution of the enhanced visible light image with the color distribution of the ideal image; S2, Design a low-light enhancement neural network based on the Mamba structure and convert low-light images into the YCbCr color space. Decouple the illumination component and color component of the low-light image and perform enhancement processing on them separately. S3 designs a fusion network based on the Mamba structure, which obtains modal complementary information through a hybrid attention mechanism to fuse the enhanced visible light image and infrared image.
[0024] In S1, this embodiment selects the CbCr channels of 360 well-exposed images from the SCIE dataset as positive samples for contrastive learning. The CbCr channels of 360 overexposed / underexposed images were selected as negative samples for contrastive learning. The image sizes of both positive and negative samples were resized to 128x128. The visible light image to be enhanced in this embodiment... A low-light image is randomly selected from LLVIP and cropped to 128x128 size.
[0025] Lowlight Enhancement Loss Based on Contrast Learning The expression is: In the above formula, This represents the color channels of the enhanced visible light image. The color channels of the positive sample image. The color channels represent the original low-light image. Indicates the color channels of an underexposed / overexposed image. This represents the discrete cosine similarity.
[0026] In S2, the low-light enhancement neural network structure is as follows: Figure 2 As shown, the low-light enhancement neural network first converts the RGB image to the YCbCr color space, and the conversion method is defined as follows: ; ; ; in , , This represents the R, G, and B channel values of an image in the RGB color space. , , This represents the Y, Cb, and Cr channel values of the image in the YCbCr color space. After conversion to the YCbCr color space, the Y, Cb, and Cr channels of the visible light image are stitched together with the infrared image, and low-light enhancement is performed using a low-light enhancement neural network. In this embodiment, the infrared image is an infrared modal image from the LLVIP dataset, which has also been cropped to 128x128 pixels.
[0027] In this embodiment, the low-light enhancement neural network structure is as follows: Figure 2 As shown, the input and output feature maps of each module maintain a fixed length and width of 128, with only the channel dimension changing with the operation. The detailed structure of the low-light enhancement neural network is shown in Table 1, where the modules are arranged from top to bottom according to the actual execution order of the low-light enhancement neural network structure. Through processing by this low-light enhancement neural network, the output features can be obtained. .
[0028] Table 1. Low-light enhancement neural network structure In this embodiment, the Mamba module is as follows: Figure 3 As shown, the Mamba module first uses discrete wavelet transform to transform the input features Decomposed into low-frequency features and high frequency characteristics Subsequently, the Mamba module uses the Mamba architecture and a depthwise separable convolutional architecture to target low-frequency features. High frequency characteristics Remodeled as and The Mamba architecture is an efficient sequence modeling technique based on a selective state-space model. The mathematical model for modeling low-frequency features using the Mamba module is as follows: ; In the above formula, EinFFT This is a novel frequency domain channel mixing method employing Einstein matrix multiplication. Ln Presentation layer normalization operation. ASSM ( This indicates the use of the Mamba architecture based on non-causal unidirectional scanning.
[0029] The mathematical model for the Mamba architecture in modeling high-frequency features is as follows: ; In the above formula, Split ReLU represents the segmentation of features along the channels. This indicates the use of the ReLU activation function. This indicates the use of pointwise convolution. This indicates that a depthwise separable convolution operation is used with a kernel size of 5x5.
[0030] The Mamba module will eventually reconstruct the low-frequency characteristics. and reconstruction of high-frequency features The inverse wavelet transform is used to convert the output features. The mathematical model is: ; In the above formula, In this example, the illumination-color decoupling enhancement strategy based on the YCbCr color space is as follows: Figure 4 As shown, 24-dimensional output features The channel is divided into 3 dimensions of size 8 features. . The nth dimension feature will be used in the nth iteration enhancement, as shown in the formula: ; ; ; ; in This represents the Y / Cb / Cr channels of an input low-light visible light image. Indicates the use for the nth time Features during channel iteration This represents the Y channel value of the image after n iterations. This represents the Cb / Cr value of the image after n iterations. This represents the Y / Cb / Cr values of the final enhanced visible light image after 8 iterations of enhancement.
[0031] In this embodiment, the low-light enhancement neural network is trained using the Adam optimizer with a low-light enhancement loss function. To optimize the objective, training samples were derived from the LLVIP dataset. Visible-infrared image pairs under low-light conditions were selected, cropped to 128×128 pixels, resulting in 4000 pairs. During training, the batch size was set to 64, the number of epochs was 20, and the initial learning rate was 5e-5. Defined as: ; In this embodiment, smoothing loss Used to suppress enhanced outliers and make gradient changes between pixels smoother, spatial consistency loss. Constraining native low-light images Compared with the enhanced image The difference in characteristics between adjacent regions, exposure loss Used to control the brightness conditions of local areas, avoiding overexposure and underexposure problems caused by insufficient enhancement in certain areas, contrast learning loss. Used to constrain and enhance the color of an image. , , , They represent , , , The weight of the loss. In this embodiment, , , , Set them to 200, 1, 10, and 0.7 respectively.
[0032] In this embodiment, smoothing loss It suppresses enhanced outliers and makes gradient changes between pixels smoother. Defined as: ; In the above formula, Let i represent the i-th and j-th eigenvalues.
[0033] In this embodiment, spatial consistency loss The original low-light image was preserved during the low-light image enhancement process. Compared with the enhanced image The differences in characteristics between adjacent regions. Defined as: ; in This indicates the number of local regions, and the size of each local region is set to 4×4 based on experience. Indicated by It is the adjacent area to the center, including areas in the four directions of up, down, left, and right.
[0034] In this embodiment, exposure loss Used to control the brightness conditions of localized areas, avoiding overexposure and underexposure problems caused by insufficient enhancement in certain areas. Specifically, It can calculate the difference between the average pixel intensity within an 8×8 window and the threshold ε. Defined as: ; Where M represents the number of windows, This represents the average calculation within the window, and ε is set to 0.5 based on experience.
[0035] In S3, the fusion network used in this embodiment is as follows: Figure 5 As shown, the fusion network consists of two Mamba-based encoders, a feature fusion mechanism based on hybrid attention, and a Mamba-based decoder.
[0036] In this embodiment, the encoder structure based on the Mamba module is as follows: Figure 6 As shown, the encoder is used to extract modal unique features from infrared and visible light images. The input and output feature maps of each module in the encoder maintain a fixed length and width of 128, with only the channel dimension changing with the operation. The detailed structure of the encoder is shown in Table 2, where the modules are arranged from top to bottom according to the actual execution order of the encoder.
[0037] Table 2 Encoder Structure Based on Mamba Module like Figure 7 As shown, the feature fusion structure first integrates the features extracted by the encoder. and Subtraction yields residual characteristics residual characteristics Channel weights are obtained through a hybrid attention mechanism. and spatial weights The mathematical model is: ; ; In the above formula, .
[0038] In this embodiment, the channel attention structure diagram is as follows: Figure 8 As shown, channel attention first applies to residual features with dimensions of 128 (width and height) and 192 channels. Channel standard deviation calculation and global pooling are performed separately. Channel standard deviation calculation captures the standard deviation of pixel values within each channel, reflecting the dispersion of pixel values within each channel and indicating the contrast of the feature across channels. Global pooling calculates the average value of all pixel values across the spatial dimension for each channel, extracting the global mean information for each channel and reflecting the overall strength of the channel feature. The obtained channel standard deviation values are then added to the pooling result to obtain an output feature with dimensions 1. This is then sequentially connected to a pixel-wise convolutional layer, an LRelu activation layer, a pointwise convolutional layer, and a SoftMax activation layer, finally resulting in channel attention weights with dimensions 1 and 128. .
[0039] In this embodiment, the spatial attention structure diagram is as follows: Figure 9 As shown, spatial attention affects residual features Channel-dimensional average pooling and channel-dimensional maximization are performed. Channel-dimensional average pooling calculates the average value of all channel pixels at each spatial location to extract the average response intensity at that location. Channel-dimensional maximization calculates the maximum value of all channel pixels at each spatial location to extract the maximum response intensity at that location, highlighting locally salient regions. These two operations are then concatenated to fuse the two spatial statistical information before being sequentially connected to a 3x3 convolutional layer and a SoftMax layer, ultimately resulting in spatial attention weights with dimensions of 128 and a width of 1. .
[0040] In this embodiment, channel attention weights are obtained. Spatial attention weights Then, the modal weights with a width and height of 128 and a dimension of 192 can be calculated. The method is as follows: ; Based on modal weights The final fusion features are obtained. The formula is: ; In this embodiment, the decoder based on the Mamba module is as follows: Figure 10 As shown, the decoder is used to fuse features. Decoded into the luminance channel of the merged image The input and output feature maps of each module of the decoder maintain a fixed length and width of 128, with only the channel dimension changing with the operation. The detailed structure of the decoder is shown in Table 3, in which the modules are arranged from top to bottom according to the actual execution order of the decoder.
[0041] Table 3 Decoder Structure Based on Mamba Module In this embodiment, the fused image brightness map is obtained through a decoder. Subsequently, the Cb and Cr values of the fused image are directly adopted from the Cb and Cr channel values corresponding to the enhanced visible light image. The fused image, located in the YCbCr color space, is converted to the RGB color space to obtain the final fused result. ; in, This represents a blended graphic in the RGB color space. This represents the operation of converting a YCbCr image to an RGB image. The mathematical model for YCbCrtoRGB is: ; ; ; in, , , These represent the R / G / B values of the fused image in the RGB color space, respectively. , , These represent the Y / Cb / Cr values of the fused image in the YCbCr color space.
[0042] In this embodiment, the training of the fusion network uses the Adam optimizer to enhance the loss function with low light intensity. To optimize the objective, training samples were derived from the LLVIP dataset. Visible-infrared image pairs under low-light conditions were selected, cropped to 128×128 pixels, resulting in 2000 sample pairs. During training, the batch size was set to 24, the number of training epochs was 35, and the initial learning rate was 1e-4. Defined as: ; In this embodiment, structural loss Strength loss gradient loss This ensures that the merged image contains the most significant intensity information for each pixel, preserving more detail and texture information. , , They represent , , The weight of the loss. In this embodiment, , , Set them to 30, 120, and 10 respectively.
[0043] In this embodiment, structural loss , Defined as: ; In the above formula, The function calculates the structural consistency between images a and b by evaluating the brightness, contrast, and structure of images a and b.
[0044] In this embodiment, the strength loss This ensures that the fused image contains the most significant intensity information for each pixel. Defined as: ; In the above formula, H and W represent the height and width of the image, respectively. This represents the l1 norm.
[0045] In this embodiment, gradient loss Ensure that the merged image retains more detail and texture information. Defined as: ; In the above formula, This represents the Sobel operator.
[0046] Qualitative analysis: The low-light visible light and infrared image fusion method with color restoration provided in this invention (hereinafter referred to as the "invention method") was qualitatively tested on the LLVIP dataset and the MSRS dataset. The qualitative analysis results are consistent with... Figure 11 , 12 The results are compared with other methods in the present embodiment, where (a) represents a visible light image, (b) represents an infrared image, (c) represents the fusion result of the current SwinFusion method, (d) represents the fusion result of the current CDDFuse method, (e) represents the fusion result of the current Dif-Fusion method, (f) represents the fusion result of the current FusionMamba method, (g) represents the fusion result of the current EMMA method, (h) represents the fusion result of the current TCMOA method, (i) represents the fusion result of the current DIVFusion method, (j) represents the fusion result of the current LENFusion method, (k) represents the fusion result of the current DDBF method, and (l) represents the fusion result of the method in this embodiment. Figure 11 and Figure 12 As can be seen, the fused image obtained by the method of the present invention in a low-light environment has the advantages of high contrast, clear background, prominent target and rich color after low-light enhancement.
[0047] Quantitative analysis: The method of this invention underwent quantitative performance testing on the LLVIP and MSRS datasets, and its fusion performance was comprehensively evaluated using both datasets. The quantitative evaluation results of the method are shown in Table 4, and comparisons are made with other comparative methods. The results show that on the LLVIP dataset, the method of this invention outperforms other comparative methods in all key metrics—EN (entropy, measuring the richness of image information), AG (average gradient, reflecting image sharpness and rate of detail change), SD (standard deviation, reflecting the dispersion of image grayscale values and contrast), SF (spatial frequency, measuring the spatial domain variation characteristics and detail richness of the image), VIF (visual information fidelity, evaluating the information fidelity to the reference image), and NIQE (Natural Image Quality Evaluator, a quality evaluation metric without a reference image).
[0048] Furthermore, to verify the generalization ability of the method, the method of this invention was also tested on the untrained MSRS dataset. Table 5 presents the quantitative comparison results of the method of this invention with other comparative methods on the MSRS dataset, further confirming the effectiveness and generalization of the method of this invention.
[0049] Table 4 Comparison of the method of this invention with other methods in the LLVIP dataset Table 5 Comparison of the method of this invention with other methods in the MSRS dataset. Ablation experiment: The method of this invention was subjected to ablation experiments on the LLVIP dataset to verify the effectiveness of the module. To verify the contrast loss... middle , , The constraint effect was utilized, and the contrast loss was removed in the ablation experiment. In , , To verify the effectiveness of ASSM and EinFFT in the Mamba module, the ablation experiments replaced ASSM with a causal four-directional scanning Mamba architecture and EinFFT with a multilayer perceptron (MLP), respectively. Table 6 presents the ablation experiments of the method of this invention on the LLVIP dataset. Table 3 shows the results from the contrastive loss. Remove from , , All of these factors lead to a comprehensive decline in EN, AG, SD, SF, VIF, and NIQE indicators, highlighting the crucial role of positive and negative samples in contrastive learning. Furthermore, Replacing ASSM with SS2D improves AG, SF, and VIF metrics, but also significantly amplifies potential noise and increases NIQE. Therefore, using the ASSM mechanism to process low-frequency information enables global perception while avoiding amplification of local noise. While using MLP in EMEB improves AG, SF, and NIQE metrics, it also leads to a decrease in EN, SD, and VIF metrics, and increases computational complexity. Therefore, the ablation experiments effectively validated the superiority of the contrast loss and the Mamba module used.
[0050] Table 6 Ablation experiments of the method of this invention in the LLVIP dataset. Corresponding to the aforementioned embodiment of a low-light visible light and infrared image fusion method with color restoration capability, the present invention also provides an embodiment of a low-light visible light and infrared image fusion device with color restoration capability.
[0051] See Figure 13 The present invention provides a low-light visible light and infrared image fusion device with color restoration capability, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a low-light visible light and infrared image fusion method with color restoration capability as described in the above embodiment.
[0052] An embodiment of the low-light visible light and infrared image fusion device with color restoration capability provided by this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 13 The diagram shown is a hardware structure diagram of any data processing device, including the low-light visible light and infrared image fusion device with color restoration capability provided by the present invention. (Except for...) Figure 13 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0053] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0054] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0055] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a low-light visible light and infrared image fusion method with color restoration capability as described in the above embodiments.
[0056] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0057] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for fusing low-light visible light and infrared images with color restoration capabilities.
[0058] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0059] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. This application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A low-illumination visible and infrared image fusion method with color restoration capability, characterized in that, The method comprises: constructing a contrast learning dataset from well-exposed images and low-light images; constructing a low-light enhancement network based on a Mamba structure and training the network using a low-light enhancement loss based on contrast learning; the low-light enhancement loss includes a color distribution alignment loss; the low-light enhancement network is used to convert low-light images to YCbCr space, model high-frequency and low-frequency features of the converted images using a Mamba module, and iteratively enhance the channel values of the reconstructed images; constructing a fusion network based on a Mamba structure and training the network, the fusion network is used to extract modal features of infrared images and visible light images, obtain modal fusion features through a hybrid attention mechanism, decode the fusion features into a luminance channel of a fusion image, combine the luminance channel with enhanced channel values, and convert the fusion image to an RGB image; using the trained low-light enhancement network and fusion network to realize color restoration of low-light visible light and infrared image fusion.
2. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The construction of the contrast learning dataset specifically comprises: selecting the CbCr channel of a well-exposed image as a positive sample for contrast learning, selecting the CbCr channel of an overexposed / underexposed image as a negative sample for contrast learning, and aligning the sizes of the positive and negative samples.
3. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The color distribution alignment loss is specifically ; in the above formulae, denotes a color channel of the enhanced visible light image, denotes a color channel of the normal light image, denotes a color channel of the under / over exposed image, denotes a color channel of the low light image, denotes a discrete cosine similarity.
4. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The low-light enhancement loss is specifically obtained by weighting a smoothing loss, a spatial consistency loss, an exposure loss, and a color distribution alignment loss, the smoothing loss suppresses abnormal values of the enhancement, the spatial consistency loss calculates the difference characteristics of a low-light image and an enhanced image in each local region and adjacent regions, and the exposure loss calculates the difference between the average pixel intensity in a window and a preset threshold to control the brightness of the local region.
5. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The Mamba module specifically comprises: First, the input feature is decomposed into low frequency feature and high frequency feature using discrete wavelet transform ; Modeling low frequency features: ; In the above equation, EinFFT is an Einstein Fourier transform, Ln denotes a layer normalization operation, ASSM denotes the use of a Mamba architecture based on acausal one-way scan The mathematical model for modeling high-frequency features is: ; Split denotes splitting the feature along the channel, Relu denotes using a Relu activation function, denotes using a point-wise convolution operation, denotes using a depthwise separable convolution operation with a kernel size of 5x5; The reconstructed low-frequency features and the reconstructed high-frequency features are converted into output features through inverse wavelet transform.
6. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 5, characterized in that, The iterative enhancement specifically comprises: The output features are divided into three sub-features along the channel, the number of iterations is determined according to the dimensions of the sub-features, and the nth iteration enhancement formula is: ; ; ; ; ; wherein Y / Cb / Cr channels of the input low-light visible image, Y / Cb / Cr channels of the final enhanced visible image after n iterations of the enhancement process, characteristics at the channel iteration, Y channel values of the image after n iterations of the enhancement process, Cb / Cr values of the image after n iterations of the enhancement process, Y / Cb / Cr values of the final enhanced visible image after n iterations of the enhancement process.
7. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The fusion network based on the Mamba structure comprises two encoders based on the Mamba module, feature fusion based on a hybrid attention mechanism, and a decoder based on the Mamba module; The encoder based on the Mamba module uses the Mamba module and a point-wise convolution layer to extract modal unique features of infrared images and visible light images; The feature fusion based on the hybrid attention mechanism first subtracts the features extracted by the encoder and to obtain residual features , the residual features pass through the hybrid attention mechanism to obtain channel weights and spatial weights and , the modal weights are calculated according to the channel weights and the spatial weights, and the visible light features and the infrared features are fused by using the modal weights to obtain fused features; The decoder based on the Mamba module uses the Mamba module, a point-wise convolution layer, and a Tanh activation function to decode the fusion features into a luminance channel value of a fusion image, and directly uses the enhanced Cb and Cr channel values of the visible light image to obtain the final fusion result.
8. The low-illumination visible and infrared image fusion method with color restoration capability according to claim 1, characterized in that, The training of the fusion network based on the Mamba structure includes a structure loss, an intensity loss, and a gradient loss; The structure loss restores the structural information of the original image, the intensity loss causes the fusion image to contain the most prominent intensity information of each pixel, and the gradient loss ensures that the fusion image retains more detail and texture information.
9. A low-illumination visible and infrared image fusion device with color restoration capability, comprising a memory and one or more processors, wherein the memory stores executable code, and the executable code comprises the following steps: The processor executes the executable code to implement a low-illumination visible and infrared image fusion method with color restoration capability according to any one of claims 1-8. 10. A computer-readable storage medium having stored thereon a program, characterized in that, The program is executed by the processor to implement a low-illumination visible and infrared image fusion method with color restoration capability according to any one of claims 1-8.
Citation Information
Cited By
Tiny target detection method and device based on feature alignment and interaction and server
CN122067064A