Infrared and visible light image perception enhancement fusion method and system based on deep learning

Through a deep learning-based method, multi-scale decomposition and feature enhancement of infrared and visible images, combined with the illumination consistency loss function, the details and color consistency problems of image fusion in low-light environments are solved to generate high-quality fusion images.

CN119722492BActive Publication Date: 2025-05-16JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510215232.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-16
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In low-light environments, existing infrared and visible image fusion technologies are difficult to effectively combine the complementary characteristics of the two images, resulting in insufficient details and color consistency of the fusion image.

Method used

Deep learning-based method is adopted to perform multi-scale decomposition of infrared and visible light images through discrete wavelet transformation, combine adaptive sparse Transformer blocks to enhance and fusion the basic and detailed features, and design the lighting consistency loss function to optimize brightness and contrast.

Benefits of technology

The generated fusion images are significantly better than the existing technology in detail restoration and texture performance, and can better restore the brightness level of the real scene, reduce color distortion, and improve computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722492B_ABST
    Figure CN119722492B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for infrared and visible light image perception enhancement fusion based on deep learning. The method applies a multi-stage training strategy. First, the encoder and decoder are trained for single-mode reconstruction to ensure feature extraction and reconstruction capabilities. Based on the pre-trained encoder and decoder, discrete wavelet transform is applied for multi-scale decomposition to enhance and fuse basic features and detail features respectively. The illumination perception module and illumination consistency loss function are designed to optimize the brightness and contrast of the fused image. The present invention effectively improves the detail retention, brightness uniformity and color performance of images under low-light conditions, and is particularly suitable for application scenarios such as high perception requirements at night.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image processing and computer vision technology, and in particular to a method and system for enhancing the fusion of infrared and visible light image perception based on deep learning. Background Art

[0002] As an important research direction in the field of image processing and computer vision, infrared and visible light image fusion technology is widely used in scenes such as night monitoring, intelligent security, and aviation monitoring. By combining the complementary information of infrared and visible light images, images with rich details and clear targets can be generated, providing strong support for target detection, environmental perception, and visual analysis in complex environments. Infrared images have the characteristics of penetrating smoke and resisting strong light interference, and can still capture the thermal radiation information of the target even in completely dark conditions. However, infrared images usually lack texture and color details and cannot meet the needs of high-quality visual analysis. In contrast, visible light images can provide delicate texture and color information, which is consistent with human visual perception. However, in low-light environments, problems such as reduced contrast, loss of details, and increased noise seriously limit its application. Therefore, how to effectively combine the complementary characteristics of infrared and visible light images to generate high-quality images with both global information and detailed expression has become one of the core technical challenges in this field.

[0003] Existing infrared and visible light image fusion technologies are mainly divided into traditional methods and deep learning-based methods. Traditional methods usually rely on manually designed rules and feature extraction strategies, such as weighted averaging, pyramid decomposition, sparse representation, and wavelet transform. These methods achieve the extraction and fusion of local features by decomposing and reconstructing the low-frequency and high-frequency components of the image. For example, the weighted average method directly fuses two images using a fixed ratio, but it is easy to cause blurring effects and loss of details. The pyramid decomposition method better retains edge and detail information by performing multi-layer decomposition and fusion of images, but its rules are fixed and it is difficult to adapt to complex and changing scenes. Sparse representation and wavelet transform techniques finely model local features, but their robustness performance in dictionary construction and complex scenes is still limited.

[0004] With the rapid development of deep learning technology, image fusion methods based on deep learning have gradually become mainstream. Through an end-to-end learning framework, deep learning methods can automatically learn fusion rules, avoiding the limitations of traditional methods that rely on manually designed features. For example, convolutional neural networks (CNNs) extract deep features of multimodal images through multi-layer convolutions and combine attention mechanisms to achieve efficient interaction of features. Some methods also use generative adversarial networks (GANs) to generate high-quality fused images through adversarial training of generators and discriminators. In recent years, Transformer-based models have shown excellent performance in long-distance dependency modeling and multimodal feature interaction due to their powerful global attention mechanism. In addition, the Diffusion Model significantly improves the quality and consistency of fused images with its generation mechanism of gradual denoising.

[0005] Although the existing technology has made significant progress, it still faces many challenges under low-light conditions. First, the existing deep learning methods often assume that the input image quality is high, while visible light images in low-light environments often have problems such as reduced contrast and loss of details, which seriously restricts the fusion effect. Introducing low-light enhancement technology is a reasonable improvement direction, but direct coupling with image fusion often leads to incompatibility, and color distortion may be caused during the enhancement process, making it difficult to ensure the color consistency and visual naturalness of the fused image. In addition, the difference in modal characteristics between infrared and visible light images makes efficient feature interaction and fusion still a technical challenge. The existing methods lack balance in capturing global structure and local details, which easily leads to the loss of feature information. Finally, some deep learning methods have a high dependence on hardware resources and low computational efficiency, and their robustness under dynamic lighting changes or complex environments needs to be further improved. These problems limit the widespread application of current technologies under low-light conditions and also clarify the direction for further research. Summary of the invention

[0006] In view of the above situation, the main purpose of the present invention is to propose a method and system for infrared and visible light image perception enhancement fusion based on deep learning, so as to solve the problems of visible light image quality degradation in low-light environments, significant differences in infrared and visible light modal characteristics, and difficulty in ensuring color consistency.

[0007] The present invention proposes a method for enhancing the fusion of infrared and visible light image perception based on deep learning, and the method comprises the following steps:

[0008] Step 1, obtaining an infrared image and a visible light image, extracting a visible light image brightness channel of the visible light image, and preprocessing the infrared image and the visible light image brightness channel;

[0009] Step 2: Use the infrared image and visible light image brightness channels to perform single-modal reconstruction training on the encoder and decoder to obtain a pre-trained encoder and a pre-trained decoder; optimize the feature extraction and reconstruction capabilities of the encoder and decoder to provide high-quality features for the fusion task;

[0010] Step 3: Use the pre-trained encoder to extract the deep features of the infrared image and the visible light image brightness channel respectively, and apply discrete wavelet transform to the deep features of the infrared image and the visible light image brightness channel for multi-scale decomposition respectively; then perform enhancement and fusion operations on the multi-scale decomposition results in turn to obtain a fusion result;

[0011] The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image;

[0012] Step 4: Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image;

[0013] The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output.

[0014] Step 5: Merge the final luminance channel fusion image with the Cb and Cr chrominance channels of the visible light image and convert them into RGB color space to generate the final color fusion image.

[0015] The present invention also proposes a deep learning-based infrared and visible light image perception enhancement fusion system, wherein the system applies the deep learning-based infrared and visible light image perception enhancement fusion method as described above, and the system includes:

[0016] Data acquisition and preprocessing module for:

[0017] Acquire an infrared image and a visible light image, extract a visible light image brightness channel of the visible light image, and preprocess the infrared image and the visible light image brightness channels;

[0018] Encoder and decoder pre-training modules for:

[0019] The encoder and decoder are trained for single-modal reconstruction using infrared images and visible light image brightness channels to obtain pre-trained encoders and pre-trained decoders; the feature extraction and reconstruction capabilities of the encoder and decoder are optimized to provide high-quality features for fusion tasks;

[0020] Multi-scale feature fusion module, used for:

[0021] The deep features of the infrared image and the visible light image brightness channel are extracted using the pre-trained encoder, and the deep features of the infrared image and the visible light image brightness channel are decomposed by discrete wavelet transform at multiple scales. The multi-scale decomposition results are then enhanced and fused in turn to obtain the fused results.

[0022] The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image;

[0023] Lighting Optimization Module for:

[0024] Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image;

[0025] The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output.

[0026] Color image generation and output module for:

[0027] The final luminance channel fused image is merged with the Cb and Cr chrominance channels of the visible light image and converted to the RGB color space to generate the final color fused image.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The present invention performs multi-scale decomposition of infrared and visible light images through discrete wavelet transform, takes sub-bands as basic features and detail features, and enhances and fuses the two types of features respectively in combination with adaptive sparse Transformer blocks; the present invention can effectively balance the global structural information and the local detail expression ability, and the generated fused image is significantly superior to the existing technology in detail restoration and texture expression.

[0030] 2. The present invention designs an illumination consistency loss function, which significantly improves the contrast and clarity of the fused image by reducing the brightness histogram distribution and contrast difference between the fused image and the enhanced image. Compared with the prior art, the image generated by the present invention can better restore the brightness level of the real scene and effectively reduce the color distortion problem that occurs during low-light enhancement.

[0031] 3. The present invention adopts a lightweight adaptive sparse Transformer block (AST), which improves computational efficiency through sparse computing mechanism while maintaining high efficiency of multimodal feature interaction. Compared with the existing complex network structure, the present invention significantly improves the real-time processing capability, making it more suitable for real-time image processing tasks in complex scenarios such as low-resource devices and dynamic lighting changes.

[0032] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of the infrared and visible light image perception enhancement fusion method based on deep learning proposed by the present invention;

[0034] Figure 2 A schematic diagram of the infrared image and visible light image perception enhancement fusion network of the deep learning infrared and visible light image perception enhancement fusion method proposed in the present invention;

[0035] Figure 3 Schematic diagram of the adaptive sparse Transformer block of the infrared and visible light image perception enhancement fusion method based on deep learning proposed in the present invention;

[0036] Figure 4 Schematic diagram of the bimodal attention mechanism in the adaptive sparse Transformer block of the infrared and visible light image perception enhancement fusion method based on deep learning proposed in the present invention;

[0037] Figure 5 Schematic diagram of the feature refinement feedforward network in the adaptive sparse Transformer block of the infrared and visible light image perception enhancement fusion method based on deep learning proposed by the present invention;

[0038] Figure 6 This is a framework diagram of the infrared and visible light image perception enhancement fusion system based on deep learning proposed in the present invention. DETAILED DESCRIPTION

[0039] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0040] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0041] See also Figure 1 and Figure 2 , in the figure, Represents a conversion operation, Represents a merge operation. Represents a fusion operation.

[0042] This embodiment provides a method for enhancing the fusion of infrared and visible light image perception based on deep learning, and the method comprises the following steps:

[0043] Step 1, obtaining an infrared image and a visible light image, extracting a visible light image brightness channel of the visible light image, and preprocessing the infrared image and the visible light image brightness channel;

[0044] In step S1, the method for preprocessing the infrared and visible light images specifically includes the following steps:

[0045] Align and normalize the infrared and visible light image brightness channels to ensure that the data has consistent scale and distribution;

[0046] Perform data augmentation on the normalized images, including random cropping, rotation, etc., to improve the robustness of the model and increase the diversity of the training set;

[0047] The enhanced images are divided into training set and validation set in the ratio of 80% and 20%;

[0048] Step 2: Use the infrared image and visible light image brightness channels to perform single-modal reconstruction training on the encoder and decoder to obtain a pre-trained encoder and a pre-trained decoder; optimize the feature extraction and reconstruction capabilities of the encoder and decoder to provide high-quality features for the fusion task;

[0049] After completing the image preprocessing in step 1, the next step is to train the feature extraction and reconstruction of infrared images and visible light images. The specific operations are as follows:

[0050] Using infrared images from the training set and the visible light image brightness channel Use encoders separately Extracting deep features, the corresponding process has the following relationship:

[0051] ;

[0052] in, represents an encoder consisting of one convolutional layer and four adaptive sparse Transformer blocks, represents an infrared image; Represents the brightness channel of the visible light image; and Respectively represent the infrared image features and the visible light image brightness channel features;

[0053] The decoder is used to reconstruct the infrared image features and the visible light image brightness channel features to obtain the reconstructed infrared image and visible light image brightness channel. The corresponding process has the following relationship:

[0054] ;

[0055] in, Represents the decoder, whose structure is similar to that of the encoder; and Respectively represent the brightness channels of the infrared image and visible light image reconstructed by the decoder;

[0056] In order to minimize the difference between the reconstructed image and the original image, a reconstruction loss function is introduced. ; This loss function is used to quantify the difference between the reconstructed image and the original image, and guide the network to optimize during the training process; specifically, the infrared image structure similarity loss is constructed according to the difference between the reconstructed infrared image and the original infrared image, and the visible light image brightness channel structure similarity loss is constructed according to the difference between the reconstructed visible light image brightness channel and the original visible light image brightness channel. The corresponding process has the following relationship:

[0057] ;

[0058] ;

[0059] in, Represents the structural similarity loss function of the brightness channel of the visible light image, which is used to constrain the reconstruction loss function of the visible light image. represents the infrared image structural similarity loss function, which is used to constrain the reconstruction loss function of the infrared image. The structural similarity loss function is used to measure the structural similarity between the reconstructed image and the original image. It represents the square of the L2 norm, that is, the sum of the squares of the differences in the image pixel values ​​is calculated, and is often used to measure the Euclidean distance between the reconstructed image and the original image; and Represent the structural similarity loss of visible light images and infrared images, respectively, which are used to measure the structural similarity between the reconstructed image and the original image; Represents the hyperparameter that adjusts the weight of the structural similarity loss (SSIM) term;

[0060] The reconstruction loss function is constructed by combining the structural similarity loss of the brightness channel of the visible light image and the structural similarity loss of the infrared image. The corresponding process has the following relationship:

[0061] ;

[0062] in, represents the reconstruction loss function;

[0063] By minimizing the reconstruction loss function to optimize the parameters of the decoder and encoder, a pre-trained encoder and a pre-trained decoder are obtained.

[0064] Step 3: Use the pre-trained encoder to extract the deep features of the infrared image and the visible light image brightness channel respectively, and apply discrete wavelet transform to the deep features of the infrared image and the visible light image brightness channel for multi-scale decomposition respectively; then perform enhancement and fusion operations on the multi-scale decomposition results in turn to obtain a fusion result;

[0065] The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image;

[0066] In step 2, the pre-training of the encoder and decoder has been completed; the goal of this step is to combine two different types of image features through multimodal feature fusion, thereby enhancing the image's detail expression and information richness;

[0067] First, the infrared image and the visible light image brightness channel are input into the pre-trained encoder, which will extract deep features from the two images to obtain infrared image features and visible light image brightness channel features respectively;

[0068] Next, discrete wavelet transform is applied to the infrared image features and visible light image brightness channel features respectively; wavelet transform decomposes the image features into low-frequency and high-frequency sub-bands; low-frequency sub-bands usually contain the global structural information of the image, while high-frequency sub-bands mainly capture the details and texture information of the image; and After wavelet transform decomposition, four sub-bands are obtained: low-frequency sub-band and high frequency subband , the corresponding process has the following relationship:

[0069] ;

[0070] in, Represents discrete wavelet transform operation; the wavelet subband obtained after the infrared image features and the visible light image brightness channel features are subjected to discrete wavelet transform operation and , and Respectively represent the wavelet transform results extracted from the visible light image and the infrared image; LL represents the approximate feature, HL, LH and HH represent the vertical sub-band, horizontal sub-band and diagonal sub-band respectively; for the convenience of subsequent processing, the decomposed low-frequency sub-band is defined as the basic feature, and the high-frequency sub-band is defined as the detail feature. The corresponding process has the following relationship:

[0071] ;

[0072] ;

[0073] in, and Represent the basic features of infrared images and visible light images respectively, and Respectively represent the detail features of infrared images and visible light images;

[0074] Then, the basic and detailed features of the infrared image and the visible light image are input into the enhancement layer respectively to optimize the detail retention of the image; then, the enhanced basic features and detailed features are fused respectively using the fusion layer, and the corresponding process has the following relationship:

[0075] ;

[0076] ;

[0077] in, represents the basic fusion feature, represents the detail fusion feature, represents a fusion operation; represents the enhancement layer, which is used to enhance the important information in the features and suppress irrelevant information; both the enhancement layer and the fusion layer are built based on the adaptive sparse Transformer block, such as Figure 3 As shown, the adaptive sparse Transformer block consists of two normalization layers, a bimodal self-attention mechanism, and a feature refinement feed-forward network;

[0078] like Figure 4 As shown in the figure, represents matrix multiplication, Represents the fusion operation. Specifically, the characteristic of bimodal self-attention is to enhance the key feature representation by aggregating sparse and dense attention branches through adaptive weights; the workflow of this module is described as follows:

[0079] Given an input feature map ,in, represents the batch size, that is, the number of images processed at a time; Indicates the number of feature channels; in order to extract the cross-channel and spatial dependency information of the image, the query, key and value are calculated in the following way:

[0080] ;

[0081] in, express Point convolutional layer, designed to extract cross-channel contextual information; Represents query, key, and value respectively; Respectively The depth-wise separable convolution operation is used to extract queries, keys, and values. The depth-wise separable convolution first performs an independent convolution operation on each channel, and then fuses the information between channels through channel-by-channel convolution, which can effectively reduce the computational complexity and enhance the spatial representation capability of features.

[0082] Specifically, the sparse branch calculates the attention matrix by the following formula:

[0083] ;

[0084] in, represents the feature dimension, express The transpose of Represents the rectified linear unit activation function, which is used to increase nonlinear characteristics; represents the sparse attention branch; at the same time, the dense attention branch calculates the standard attention matrix through the Softmax function, and the corresponding process has the following relationship:

[0085] ;

[0086] in, represents the activation function, which normalizes the relationship between the query and the key to obtain the weight of each key; Represents the dense attention branch; in order to further improve the expressiveness of the model and give full play to the complementarity of sparse attention and dense attention mechanisms, an adaptive weighted fusion mechanism is introduced to obtain the final weighted fusion matrix. The corresponding process has the following relationship:

[0087] ;

[0088] in, and is the adaptive weight, which represents the weight coefficient of the sparse attention branch and the dense attention branch when they are fused; Represented as the final weighted fusion matrix; through this weighted fusion mechanism, the model can dynamically adjust the weights of sparse attention and dense attention according to different input images, thereby achieving better performance in various scenarios;

[0089] Although the bimodal self-attention mechanism eliminates redundancy to a large extent, there is still redundant information in the channel; Figure 5 As shown in the figure, Represents a merge operation. represents element-wise multiplication, Represents the fusion operation. In order to further improve the feature expression capability, a feature refinement feedforward network is introduced. The process of the feature refinement network is as follows;

[0090] Given an input feature map , first split it into two parts along the channel dimension, including the first part of the feature map And the second part of the feature map , the corresponding process has the following relationship:

[0091] ;

[0092] in, Represents a split operation, Indicates the dimension of the split; for the first part of the feature map , partial convolution is applied to capture local spatial information and enhance feature representation. The corresponding process has the following relationship:

[0093] ;

[0094] in, Indicates that the convolution kernel size is The partial convolution operation, Represents the enhanced features after partial convolution enhancement;

[0095] The second part of the feature map Keep unchanged, and then combine the enhanced features with the second part feature map Merge along the channel dimension to form a refined feature. The corresponding process has the following relationship:

[0096] ;

[0097] in, Represents a splicing operation; represents the refined features; subsequently, use Convolution projects the refined features to increase the feature dimension, obtain high-dimensional features, and improve the expressiveness of subsequent operations. The corresponding process has the following relationship:

[0098] ;

[0099] in, Indicates the use Convolutional projection operation maps the dimension of the feature map to a higher dimension; Represents high-dimensional features;

[0100] Next, the high-dimensional features are split into two parts along the channel dimension again to obtain the first high-dimensional feature and the second high-dimensional feature ; For the first high-dimensional feature , apply depthwise convolution Combined with the GELU activation function, nonlinear transformation is performed to capture deeper spatial dependencies and obtain nonlinear transformation features. The corresponding process has the following relationship:

[0101] ;

[0102] in, Indicates that the GELU activation function is used to increase nonlinear expression capabilities; Represents the nonlinear transformation features after GELU operation; the second high-dimensional feature remain unchanged;

[0103] Next, the nonlinear transformation features are combined with the second high-dimensional features By multiplying elements one by one, selective gating of information is performed to suppress redundant information and highlight key information. The corresponding process has the following relationship:

[0104] ;

[0105] in, Represents element-by-element multiplication, which can selectively activate key information and suppress redundant features, thereby optimizing feature expression; represents the features obtained through the gating mechanism;

[0106] Finally, the features obtained through the gating mechanism are passed through a Convolution, the gated features Mapping back to the original channel dimension , and fused with the input feature map to get the final output.

[0107] By combining selective convolution processing and gating mechanism, the feature refinement feedforward network can effectively enhance the feature representation ability and capture the global and local context information in the image while maintaining high computational efficiency;

[0108] After the feature fusion is completed, the fused features are restored to the spatial domain through inverse wavelet transform (IDWT) to obtain the spatial domain fusion features. The process of inverse wavelet transform can be expressed as:

[0109] ;

[0110] in, Represents the inverse wavelet transform operation, which is used to restore the fused features from the wavelet domain to the spatial domain; Represents the spatial fusion feature;

[0111] Then, the spatial fusion features are input into the pre-trained decoder to reconstruct the spatial fusion features into a brightness channel fusion image.

[0112] Step 4: Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image;

[0113] The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output.

[0114] Under low light conditions, the difference between the modalities of infrared images and visible light images is large, which may lead to insufficient brightness of the fused image. In order to optimize the image brightness and contrast, CLAHE (contrast limited adaptive histogram equalization) is used as an enhancement method to generate a reference image, and a light perception module is designed to dynamically adjust the enhancement parameters of CLAHE according to the brightness and contrast of the input image. According to the brightness channel of the visible light image, the average brightness and contrast of the brightness channel of the visible light image are first calculated. The corresponding process has the following relationship:

[0115] ;

[0116] in, Indicates that the image is i Row, No. j The pixel value of the column, represents the mean value of image brightness, Represents the contrast of the image, which measures the degree of fluctuation of the brightness of the image pixels. and Represents the height and width of the image respectively, Indicates the total number of pixels in the image;

[0117] Next, the coefficient of variation is calculated using the mean value of the image brightness and the image contrast, and then normalized. This allows the parameters of CLAHE to be adaptively adjusted to obtain the normalized coefficient of variation. The corresponding process has the following relationship:

[0118] ;

[0119] in, represents the normalized coefficient of variation; represents a constant used to prevent division by zero. When the contrast C is low, this nonlinear transformation will amplify the value of the coefficient of variation, thereby increasing the enhancement strength in low-contrast situations;

[0120] Then, according to the normalized coefficient of variation, the adaptive cliplimit value of CLAHE is calculated. The corresponding process has the following relationship:

[0121] ;

[0122] in, Indicates the strength of CLAHE enhancement. and Respectively represent the upper and lower limits of the enhancement intensity; through this formula, CL can be adaptively adjusted according to the brightness and contrast characteristics of the image, so that the image enhancement intensity can achieve the best effect;

[0123] By dynamically adjusting the parameters of CLAHE, an enhanced visible light image is generated. , and use it as a reference image for subsequent comparison and optimization;

[0124] In step 3, the brightness channel fused image has been processed; the goal at this stage is to further improve the visual quality of the fused image and ensure its brightness, contrast and color consistency; for this purpose, a comprehensive loss function is designed to simultaneously optimize texture details, lighting consistency and color consistency; the comprehensive loss function can be expressed as:

[0125] ;

[0126] in, and Represents texture loss, lighting consistency loss and color consistency loss respectively, , and Respectively represent the hyperparameters corresponding to texture loss, lighting consistency loss, and color consistency loss, which are used to adjust the importance of each loss term. represents the comprehensive loss function.

[0127] The goal of texture loss is to retain the high-frequency detail information in the fused image, enhance the sharpness and clarity of the image, and make the fused image more realistic visually. In order to extract the edge information in the image, the present invention uses a gradient operator to calculate the gradient of the image, and compares the gradient difference between the fused image and the original infrared and visible light images to ensure that the high-frequency details in the image are retained. The calculation process of texture loss has the following relationship:

[0128] ;

[0129] in, and Represent the brightness channel of the visible light image and the infrared image respectively; Represents the gradient operator, which is used to extract the edge information of the image. Represents the maximum value of the gradient, which is used to select stronger edge information. represents the L1 norm, Represents the luminance channel fused image;

[0130] To this end, the present invention designs an illumination consistency loss, which includes three sub-items: histogram matching loss, global contrast loss and local contrast loss, which are expressed as follows:

[0131] ;

[0132] in, represents the histogram matching loss, and Represent the global contrast loss and local contrast loss respectively, , and They represent the hyperparameters corresponding to the histogram matching loss, global contrast loss, and local contrast loss, respectively, and are used to adjust the hyperparameters of each loss item;

[0133] The histogram matching loss is used to ensure that the brightness distribution of the fused image is consistent with the brightness distribution of the enhanced visible light image, which is specifically defined as:

[0134] ;

[0135] in, Represents the operation of calculating the image histogram, which describes the distribution of brightness values ​​in the fused image; It represents the brightness channel of the visible light image after CLAHE processing. The L1 norm calculates the difference between the two histograms. The L1 norm calculates the sum of the absolute differences between the corresponding elements in the histogram. The histogram matching loss ensures that the brightness distribution of the fused image is consistent with the enhanced visible light image by comparing the brightness histograms.

[0136] The global contrast loss aims to optimize the consistency of brightness and contrast between the fused image and the reference image, which is defined as follows:

[0137] ;

[0138] in, Represents the average brightness of the image, that is, the average brightness of all pixels in the image; Represents the standard deviation of image brightness, which is used to measure the fluctuation of image brightness; Represents the brightness dynamic range of the image, which is calculated as the difference between the maximum and minimum brightness of the image. Through these calculations, the global contrast loss ensures the consistency of the overall brightness and contrast between the fused image and the reference image;

[0139] The local contrast loss ensures the brightness consistency of the fused image on details and edges and is defined as follows:

[0140] ;

[0141] in, Represents the standard deviation of the image gradient, which is used to measure the change in local contrast of the image; represents the gradient operator, which is used to extract the edge information of the image; the local contrast loss ensures that the contrast of the fused image at the details and edges is consistent with that of the reference image;

[0142] When adjusting the brightness and contrast of an image, color distortion may occur. In order to ensure that the color of the fused image is consistent with the original visible light image, the present invention introduces a color consistency loss. The loss measures the color consistency by calculating the angular difference between the brightness channels, which is defined as follows:

[0143] ;

[0144] in, represents the inner product operation, represents the L2 norm, i.e., the Euclidean distance of the image brightness channel; Represents the arccosine operation, which is used to calculate the angular difference between two vectors; color consistency loss ensures that the color of the fused image is consistent with the original visible light image, avoiding color cast caused by lighting adjustment;

[0145] Through the optimization of these three loss functions, the model can effectively fuse infrared images and visible light images, achieve a good balance in detail preservation, brightness contrast adjustment, and color consistency, and finally generate high-quality fused images;

[0146] Step 5: Merge the final luminance channel fusion image with the Cb and Cr chrominance channels of the visible light image and convert them into RGB color space to generate the final color fusion image;

[0147] Finally, the fused brightness channel image is merged with the chrominance channel of the visible light image to generate the final color fused image; the goal of this step is to ensure that the fused image is consistent with the original visible light image in color and can reflect the effective fusion of the infrared image and the visible light image;

[0148] The fused brightness channel image is merged with the chrominance channels of Cb and Cr of the original visible light image; the merging operation combines the brightness channel with the chrominance channel to obtain a complete image. The corresponding process has the following relationship:

[0149] ;

[0150] in, Represents the splicing operation of channels, and Corresponding to the original visible image and Through this operation, the brightness information and the chrominance information are combined to generate a complete fused image;

[0151] Since the image is in the YCbCr color space, the merged complete image needs to be converted to the RGB color space for visualization and storage; this conversion is done through the color space conversion function The implementation can be expressed as:

[0152] ;

[0153] in, Indicates that the image is from Operation of converting color space into RGB color space; Represents the final color image output by the model.

[0154] See also Figure 6 This embodiment further provides an infrared and visible light image perception enhancement fusion system based on deep learning, the system applies the infrared and visible light image perception enhancement fusion method based on deep learning as described above, and the system includes:

[0155] Data acquisition and preprocessing module for:

[0156] Acquire an infrared image and a visible light image, extract a visible light image brightness channel of the visible light image, and preprocess the infrared image and the visible light image brightness channels;

[0157] Encoder and decoder pre-training modules for:

[0158] The encoder and decoder are trained for single-modal reconstruction using infrared images and visible light image brightness channels to obtain pre-trained encoders and pre-trained decoders; the feature extraction and reconstruction capabilities of the encoder and decoder are optimized to provide high-quality features for fusion tasks;

[0159] Multi-scale feature fusion module, used for:

[0160] The deep features of the infrared image and the visible light image brightness channel are extracted using the pre-trained encoder, and the deep features of the infrared image and the visible light image brightness channel are decomposed by discrete wavelet transform at multiple scales. The multi-scale decomposition results are then enhanced and fused in turn to obtain the fused results.

[0161] The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image;

[0162] Lighting Optimization Module for:

[0163] Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image;

[0164] The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output.

[0165] Color image generation and output module for:

[0166] The final luminance channel fused image is merged with the Cb and Cr chrominance channels of the visible light image and converted to the RGB color space to generate the final color fused image.

[0167] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0168] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0169] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0170] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for enhancing the perception and fusion of infrared and visible light images based on deep learning, characterized in that: The method comprises the following steps: Step 1, obtaining an infrared image and a visible light image, extracting a visible light image brightness channel of the visible light image, and preprocessing the infrared image and the visible light image brightness channel; Step 2: Use the infrared image and visible light image brightness channels to perform single-modal reconstruction training on the encoder and decoder to obtain a pre-trained encoder and a pre-trained decoder; optimize the feature extraction and reconstruction capabilities of the encoder and decoder to provide high-quality features for the fusion task; Step 3: Use the pre-trained encoder to extract the deep features of the infrared image and the visible light image brightness channel respectively, and apply discrete wavelet transform to the deep features of the infrared image and the visible light image brightness channel for multi-scale decomposition respectively; then perform enhancement and fusion operations on the multi-scale decomposition results in turn to obtain a fusion result; The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image; Step 4: Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image; The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output. Step 5: Merge the final luminance channel fusion image with the Cb and Cr chrominance channels of the visible light image and convert them into RGB color space to generate the final color fusion image.

2. According to the method for enhancing the perception of infrared and visible light images based on deep learning in claim 1, it is characterized in that: In step 2, the encoder and decoder are trained for single-modal reconstruction using the infrared image and the visible light image brightness channel to obtain a pre-trained encoder and a pre-trained decoder, which specifically includes the following steps: For infrared images and the visible light image brightness channel Use encoders separately Extracting deep features, the corresponding process has the following relationship: ; in, represents an encoder consisting of one convolutional layer and four adaptive sparse Transformer blocks, represents an infrared image; Represents the brightness channel of the visible light image; and Respectively represent the infrared image features and the visible light image brightness channel features; The decoder is used to reconstruct the infrared image features and the visible light image brightness channel features to obtain the reconstructed infrared image and visible light image brightness channel. The corresponding process has the following relationship: ; in, Represents a decoder, and Respectively represent the brightness channels of the infrared image and visible light image reconstructed by the decoder; According to the difference between the reconstructed infrared image and the original infrared image, the infrared image structural similarity loss function is constructed. According to the difference between the reconstructed visible light image brightness channel and the original visible light image brightness channel, the visible light image brightness channel structural similarity loss function is constructed. The corresponding process has the following relationship: ; ; in, represents the structural similarity loss function of the brightness channel of the visible light image, represents the infrared image structural similarity loss function; represents the square of the L2 norm, represents the hyperparameter for adjusting the weight of the structural similarity loss term, and denote the structural similarity loss of visible light images and infrared images respectively; The reconstruction loss function is constructed by combining the structural similarity loss of the brightness channel of the visible light image and the structural similarity loss of the infrared image. The corresponding process has the following relationship: ; in, represents the reconstruction loss function; By minimizing the reconstruction loss function to optimize the parameters of the decoder and encoder, a pre-trained encoder and a pre-trained decoder are obtained.

3. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 2, characterized in that: In the step 3, the deep features of the infrared image and the visible light image brightness channel are respectively subjected to multi-scale decomposition by discrete wavelet transform; and the multi-scale decomposition results are then enhanced and fused in turn to obtain the fusion result, which specifically includes the following steps: Input the infrared image and visible light image brightness channel into the pre-trained encoder, which will extract deep features from the two images to obtain infrared image features and visible light image brightness channel features respectively; Discrete wavelet transform is applied to the infrared image features and the visible light image brightness channel features to obtain the wavelet transform results, which include low-frequency subbands. and high frequency subband , the corresponding process has the following relationship: ; in, represents the discrete wavelet transform operation, and They represent the wavelet transform results extracted from the visible light image and the infrared image respectively; LL represents the approximate feature, which is the low-frequency subband; HL, LH and HH represent the vertical subband, horizontal subband and diagonal subband, which are high-frequency subbands respectively; The decomposed low-frequency sub-band is defined as the basic feature, and the high-frequency sub-band is defined as the detail feature. The corresponding process has the following relationship: ; ; in, and Represent the basic features of infrared images and visible light images respectively, and Respectively represent the detail features of infrared images and visible light images; The basic features of the infrared image and the basic features of the visible light image are enhanced and fused to obtain the basic fusion features, and the detail features of the infrared image and the detail features of the visible light image are enhanced and fused to obtain the detail fusion features. The corresponding process has the following relationship: ; ; in, represents the basic fusion feature, represents the detail fusion feature, represents a fusion operation; Indicates the enhancement layer.

4. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 3 is characterized in that: In step 3, the fusion result is restored to the spatial domain by inverse wavelet transform operation, and the corresponding process has the following relationship: ; in, Represents the inverse wavelet transform operation, which is used to restore the fused features from the wavelet domain to the spatial domain; Represents the spatial fusion feature.

5. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 4, characterized in that: In step 4, adaptively enhancing the brightness channel of the visible light image to obtain the enhanced image specifically includes the following steps: Calculate the average brightness and contrast of the brightness channel of the visible light image. The corresponding process has the following relationship: ; in, Indicates that the image is i Row, No. j The pixel value of the column, represents the mean value of image brightness, represents the contrast of the image, and Represents the height and width of the image respectively, Indicates the total number of pixels in the image; The coefficient of variation CV is calculated using the mean value of the image brightness and the image contrast, and then normalized. This allows the parameters of CLAHE to be adaptively adjusted to obtain the normalized coefficient of variation. The corresponding process has the following relationship: ; in, represents the normalized coefficient of variation; Represents a constant used to prevent division by zero; According to the normalized coefficient of variation CV, the adaptive cliplimit value of CLAHE is calculated. The corresponding process has the following relationship: ; in, Indicates the strength of CLAHE enhancement. and Respectively represent the upper and lower limits of the enhancement strength; By dynamically adjusting the parameters of CLAHE, an enhanced visible light image is generated. .

6. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 5, characterized in that: In step 4, the comprehensive loss function is composed of texture loss, lighting consistency loss and color consistency loss, and the comprehensive loss function is defined as follows: ; in, , and Represents texture loss, lighting consistency loss and color consistency loss respectively, , and Respectively represent the hyperparameters corresponding to texture loss, lighting consistency loss, and color consistency loss, which are used to adjust the importance of each loss term. represents the comprehensive loss function.

7. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 6, characterized in that: The calculation process of texture loss has the following relationship: ; in, and Represent the brightness channel of the visible light image and the infrared image respectively; represents the gradient operator; represents the maximum value selection of the gradient, represents the L1 norm, Represents the fusion image of brightness channel; The calculation process of color consistency loss has the following relationship: ; in, represents the inner product operation, represents the L2 norm, Represents the arccosine operation.

8. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 7, characterized in that: The lighting consistency loss consists of histogram matching loss, global contrast loss and local contrast loss. The lighting consistency loss is defined as follows: ; in, represents the histogram matching loss, and Represent the global contrast loss and local contrast loss respectively, , and Respectively represent the hyperparameters corresponding to histogram matching loss, global contrast loss, and local contrast loss; The histogram matching loss calculation process has the following relationship: ; in, Represents the operation of calculating the image histogram, Represents the brightness channel of the visible light image after CLAHE processing; The global contrast loss calculation process has the following relationship: ; in, represents the average brightness of the image, represents the standard deviation of image brightness, Indicates the brightness dynamic range of the image; The local contrast loss calculation process has the following relationship: ; in, Represents the standard deviation of the image gradient, which is used to measure the change in local contrast of the image; Represents the gradient operator.

9. The method for enhancing the perception and fusion of infrared and visible light images based on deep learning according to claim 8, characterized in that: In step 5, the final luminance channel fusion image is merged with the chrominance channels of Cb and Cr of the visible light image and converted into the RGB color space to generate the final color fusion image, which specifically includes the following steps: Merge the fused brightness channel image with the Cb and Cr chrominance channels of the original visible light image to get the complete image: ; in, Represents the splicing operation of channels, and Corresponding to the original visible image and Through this operation, the brightness information and the chrominance information are combined to generate a complete fused image; Convert the complete image to RGB color space to get the final color image. The corresponding process has the following relationship: ; in, Indicates that the image is from Operation of converting color space into RGB color space; Represents the final color image output by the model.

10. A deep learning-based infrared and visible light image perception enhancement fusion system, characterized in that: The system applies the infrared and visible light image perception enhancement fusion method based on deep learning as described in any one of claims 1 to 9 above, and the system comprises: Data acquisition and preprocessing module for: Acquire an infrared image and a visible light image, extract a visible light image brightness channel of the visible light image, and preprocess the infrared image and the visible light image brightness channels; Encoder and decoder pre-training modules for: The encoder and decoder are trained for single-modal reconstruction using infrared images and visible light image brightness channels to obtain pre-trained encoders and pre-trained decoders; the feature extraction and reconstruction capabilities of the encoder and decoder are optimized to provide high-quality features for fusion tasks; Multi-scale feature fusion module, used for: The deep features of the infrared image and the visible light image brightness channel are extracted using the pre-trained encoder, and the deep features of the infrared image and the visible light image brightness channel are decomposed by discrete wavelet transform at multiple scales. The multi-scale decomposition results are then enhanced and fused in turn to obtain the fused results. The fusion result is restored to the spatial domain using the inverse wavelet transform operation, and then the image is reconstructed using the pre-trained decoder to obtain the brightness channel fusion image; Lighting Optimization Module for: Adaptively enhance the brightness channel of the visible light image to obtain an enhanced image, and use the enhanced image as a reference image; The comprehensive loss function is constructed by using the comprehensive difference between the brightness channel fusion image and the enhanced image. The difference between the brightness channel fusion image and the reference image is optimized by minimizing the comprehensive loss function. After the optimization is completed, the final brightness channel fusion image is output. Color image generation and output module for: The final luminance channel fused image is merged with the Cb and Cr chrominance channels of the visible light image and converted to the RGB color space to generate the final color fused image.

Citation Information

Patent Citations

  • Infrared and visible light image illumination optimization fusion method based on deep learning

    CN118799692A

  • Infrared and visible light fusion method

    US20220044374A1