Infrared and visible light image fusion method and system oriented to dark environment
Through the deep learning dark light enhancement network and fusion network, the problem of insufficient information utilization in the fusion of infrared and visible light images in dark environments is solved, high-quality fused images are generated, and the image details and visual effects are improved.
Patent Information
- Application Number
- CN202510788725.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing infrared and visible light image fusion technologies fail to fully utilize the rich scene content of visible light images in dark environments, ignore global features, and are prone to spectral pollution and noise, resulting in poor image quality.
A deep learning-based dark light enhancement network and fusion network are used to synchronously collect infrared and visible light images for image enhancement and feature extraction. Combined with the brightness feedback mechanism, a fused image with low noise and balanced brightness and chromaticity is generated.
It improves image quality at night or in low-light environments, enhances detail visibility and overall visual effects, and improves image clarity and information content.
Smart Images

Figure CN120708006A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to a method and system for fusing infrared and visible light images in a dark environment. Background Art
[0002] In fields such as security monitoring, autonomous driving, and military reconnaissance, it is often necessary to acquire image data in low-light or nighttime scenes. However, due to technical limitations and the influence of the shooting environment, a single image captured by the same device often cannot fully and effectively describe the entire scene. Therefore, image fusion technology has emerged. It can extract the most meaningful information from different source images (such as infrared and visible light images) and fuse it into a fused image.
[0003] Existing image fusion methods fall into two main categories: traditional methods and deep learning-based methods. Traditional methods typically involve the following three steps. First, features are extracted from the source image using a specific transformation (this is the feature extraction stage). These features are then fused using a fusion strategy in the feature fusion stage. Finally, in the feature reconstruction stage, a fused image is reconstructed from the combined features using the corresponding inverse transformation. Traditional infrared and visible image fusion methods can be further divided into five categories based on the mathematical transformations applied: multi-scale transformation-based methods, sparse representation-based methods, saliency-based methods, subspace-based methods, and hybrid-based methods. Deep learning-based methods can be further divided into three categories based on the network architecture: CNN (Convolutional Neural Network)-based methods, AE (Autoencoder)-based methods, and GAN (Generative Adversarial Network)-based methods. Deep learning-based methods can achieve feature extraction, feature fusion, and feature reconstruction through well-designed network structures and loss functions, resulting in a unique fusion result.
[0004] The inventors found that although current deep learning methods have achieved effective results in integrating the key information and complementary characteristics of visible light and infrared images, the following problems still exist: First, in low-light environments, previous image fusion techniques generally rely solely on infrared information to supplement the scene details lost in visible light images due to insufficient illumination. However, this method that relies on infrared data has obvious limitations. It fails to fully present the rich scene content contained in nighttime visible light images in the fusion results. This not only fails to achieve the original intention of fusion of infrared and visible light images, but also ignores the important information that visible light images can provide when shooting at night. Second, there may be multiple active colored light sources with low illumination intensity in nighttime scenes, which easily generate spectral pollution and cause image noise. Finally, the key to image fusion is to extract and reconstruct the most important information. Sufficient and effective information extraction is a prerequisite for good fusion. As for image detail information (such as edges and texture), this information is particularly rich in visible light images. However, visible light images also contain some global features such as background and shape information, and most existing methods ignore the extraction of global image features. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for fusing infrared and visible light images in dark environments. The solution effectively improves the image quality at night or in low-light environments, and enhances the detail visibility and overall visual effect of the image.
[0006] According to a first aspect of an embodiment of the present invention, a method for fusing infrared and visible light images in a dark environment is provided, comprising:
[0007] Synchronously collect infrared images and visible light images of the same scene;
[0008] Performing image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image;
[0009] Based on the collected infrared image and the enhanced visible light image, a fused image representation is obtained through a pre-trained deep learning-based fusion network; wherein, the fusion network specifically performs the following processing procedures: feature extraction is performed on the collected infrared image and the enhanced visible light image respectively to obtain an infrared image feature representation and a visible light image feature representation including local and global features; based on the infrared image feature representation and the visible light image feature representation, a fused image is obtained through feature fusion; and, during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
[0010] Furthermore, the dark light enhancement network specifically performs the following processing: taking the collected visible light image as input, obtaining a feature map through several convolution layers and dynamic spectrum filtering operations; decomposing the obtained feature map into several weighted feature maps; iteratively multiplying the visible light image and the obtained weighted feature map to obtain an enhanced visible light image.
[0011] Furthermore, the dark light enhancement network includes seven sequentially connected convolutional layers and three dynamic spectrum filtering modules, wherein the first convolutional layer, the second convolutional layer, the third convolutional layer, the fifth convolutional layer, the sixth convolutional layer and the seventh convolutional layer are all composed of sequentially connected convolution blocks and activation functions, and the fourth convolutional layer only includes convolution blocks; the input of the fifth convolutional layer is obtained by splicing the output result of the first convolutional layer after being processed by the first dynamic spectrum filtering module and the output result of the fourth convolutional layer; the input of the sixth convolutional layer is obtained by splicing the output result of the second convolutional layer after being processed by the second dynamic spectrum filtering module and the output result of the fifth convolutional layer; the input of the seventh convolutional layer is obtained by splicing the output result of the third convolutional layer after being processed by the third dynamic spectrum filtering module and the output result of the sixth convolutional layer.
[0012] Furthermore, the dynamic spectrum filtering module specifically performs the following processing: based on the output results of the convolution layer, different frequency domain information features are obtained through dynamic spectrum filtering operations; based on the obtained different frequency domain information features, spatial domain learning is performed locally through the self-attention mechanism to obtain the final nonlinear mapping output.
[0013] Furthermore, feature extraction is performed on the collected infrared image and the enhanced visible light image respectively, and the encoders used include a basic encoder for extracting global features and a detail encoder for extracting local detail features. The feature representation of the infrared image and the visible light image is obtained by splicing the outputs of the basic encoder and the detail encoder.
[0014] Furthermore, the feature fusion is specifically as follows: based on the infrared image feature representation and the visible light image feature representation, the channel features and spatial features corresponding to the infrared image and the visible light image are obtained respectively through average pooling and maximum pooling operations; the channel features and spatial features of the infrared image and the visible light image are respectively spliced to obtain aggregated channel features and spatial features; based on the aggregated channel features and spatial features, convolution processing and activation function processing are sequentially performed to obtain channel weights and spatial weights; based on the infrared image feature representation and the visible light image feature representation, the obtained channel weights and spatial weights are combined and weighted processing is performed to obtain a fused image.
[0015] Furthermore, the brightness probability of the fused image outputted in each round of iteration is calculated, and the following processing is specifically performed: based on the fused image outputted in each round of iteration, several convolution kernels of different sizes are used to extract brightness features at different resolutions in the fused image; and the brightness probability of the current fused image is determined based on the fusion results of the brightness features at different resolutions.
[0016] According to a second aspect of an embodiment of the present invention, a system for fusion of infrared and visible light images in a dark environment is provided, comprising:
[0017] A multimodal image acquisition unit, which is used to synchronously acquire infrared images and visible light images of the same scene;
[0018] a visible light image enhancement unit, configured to perform image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image;
[0019] An image fusion unit is configured to obtain a fused image representation based on a captured infrared image and an enhanced visible light image through a pre-trained deep learning-based fusion network; wherein the fusion network specifically performs the following processing procedures: feature extraction is performed on the captured infrared image and the enhanced visible light image, respectively, to obtain an infrared image feature representation and a visible light image feature representation including local and global features; based on the infrared image feature representation and the visible light image feature representation, a fused image is obtained through feature fusion; and, during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
[0020] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the infrared and visible light image fusion method for dark environments is implemented.
[0021] According to a fourth aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the infrared and visible light image fusion method for dark environments is implemented.
[0022] One or more of the above technical solutions have the following beneficial effects:
[0023] The present invention provides a method and system for fusing infrared and visible light images in dark environments. The solution is based on an independently constructed dark-light enhancement network and fusion network, effectively improving image quality at night or in low-light environments, enhancing the visibility of image details and the overall visual effect. At the same time, a brightness feedback mechanism is used to guide the dark-light enhancement network to produce images with better brightness, thereby revealing hidden details in the image, further improving image quality in dark environments.
[0024] The dark light enhancement network designed by the solution described in the present invention can generate nighttime visible light images with low noise and balanced brightness and chromaticity distribution, while maintaining image clarity and improving the overall image quality and visual effect;
[0025] The fusion network of the solution described in the present invention adopts a dual-branch encoder design, which can extract features of different scales or types through different branches, thereby improving the comprehensiveness of feature extraction. At the same time, by combining channel and spatial attention mechanisms, it can extract features more comprehensively, enhance important features and suppress noise, while maintaining the characteristics of lightweight and easy integration.
[0026] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0028] Figure 1 This is a flow chart of a method for fusing infrared and visible light images in a dark environment according to an embodiment of the present invention;
[0029] Figure 2 Schematic diagram of the dark light enhancement network structure described in an embodiment of the present invention;
[0030] Figure 3 Schematic diagram of the structure of the dynamic spectrum filtering module according to an embodiment of the present invention;
[0031] Figure 4 Schematic diagram of the encoder structure described in an embodiment of the present invention;
[0032] Figure 5 Schematic diagram of the fusion network structure described in an embodiment of the present invention;
[0033] Figure 6 4 is a flowchart of brightness probability calculation described in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0035] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0036] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0037] In one or more embodiments, Figure 1 As shown, a method for fusing infrared and visible light images in a dark environment is provided, comprising:
[0038] Synchronously collect infrared images and visible light images of the same scene;
[0039] Performing image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image;
[0040] Based on the collected infrared image and the enhanced visible light image, a fused image representation is obtained through a pre-trained deep learning-based fusion network; wherein, the fusion network specifically performs the following processing procedures: feature extraction is performed on the collected infrared image and the enhanced visible light image respectively to obtain an infrared image feature representation and a visible light image feature representation including local and global features; based on the infrared image feature representation and the visible light image feature representation, a fused image is obtained through feature fusion; and, during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
[0041] In a specific implementation, the dark light enhancement network specifically performs the following processing: taking the collected visible light image as input, obtaining a feature map through several convolutional layers and dynamic spectrum filtering operations; decomposing the obtained feature map into several weighted feature maps; iteratively multiplying the visible light image with the obtained weighted feature map to obtain an enhanced visible light image.
[0042] In a specific implementation, the dark light enhancement network includes seven sequentially connected convolutional layers and three dynamic spectrum filtering modules, wherein the first convolutional layer, the second convolutional layer, the third convolutional layer, the fifth convolutional layer, the sixth convolutional layer and the seventh convolutional layer are all composed of sequentially connected convolution blocks and activation functions, and the fourth convolutional layer only includes convolution blocks; the input of the fifth convolutional layer is obtained by splicing the output result of the first convolutional layer after being processed by the first dynamic spectrum filtering module and the output result of the fourth convolutional layer; the input of the sixth convolutional layer is obtained by splicing the output result of the second convolutional layer after being processed by the second dynamic spectrum filtering module and the output result of the fifth convolutional layer; the input of the seventh convolutional layer is obtained by splicing the output result of the third convolutional layer after being processed by the third dynamic spectrum filtering module and the output result of the sixth convolutional layer.
[0043] Specifically, the core goal of the dark light enhancement network is to generate nighttime visible light images with low noise and balanced brightness and chromaticity distribution, so as to improve the overall image quality and visual effect while maintaining image clarity. Figure 2 As shown, the network consists of seven 3×3 convolutional layers and three dynamic spectrum filtering modules (hereinafter referred to as SFII modules for ease of description). In each of these convolutional layers, the solution described in this embodiment uses ReLU as the activation function to introduce nonlinear factors and improve the network's representational capabilities. The Tanh activation function is specifically selected for the last convolutional layer of the entire structure to accelerate the network's convergence process.
[0044] In the specific implementation, the SFII module specifically performs the following processing: based on the output results of the convolution layer, different frequency domain information features are obtained through dynamic spectrum filtering operations; based on the obtained different frequency domain information features, spatial domain learning is performed locally through the self-attention mechanism to obtain the final nonlinear mapping output. Figure 3 As shown in FIG, the basic structure of the SFII module is shown, where D represents the feature representation of the input, and FS represents the dynamic spectrum filtering operation (adaptive Gaussian filtering is used in this embodiment). It represents the final output, and its processing process is as follows:
[0045] Q f =W Q *FS(D)+b Q ,K f =W K *FS(D)+b K ,V f =W V *FS(D)+b V
[0046] Among them, W Q , bQ , W K , b K , W V , b V is a learnable parameter, Q f , K f , V f It is the result of projecting the feature graph obtained after the FS operation into the query space, key space, and value space. The FS operation is specifically expressed as follows:
[0047]
[0048] Where D(x, y) is the input image, M is the radius of the filter, that is, the size of the filter window, which is set to 2 in this embodiment and can be set according to actual needs; is the Gaussian filter weight, σ x and σ y where i and j are the standard deviations of the Gaussian function in the x and y directions, respectively. They can be adaptively adjusted based on the local noise level or signal characteristics of the image. i and j are indices within the filter window, used to traverse all pixels within the window.
[0049] After obtaining the features, the spatial domain learning of the features is performed locally based on the self-attention mechanism, which is specifically expressed as follows:
[0050]
[0051] Among them, d and B represent the dimension and position deviation respectively, It represents matrix multiplication (MatMul), and sf() represents the self-attention mechanism.
[0052] D * =AT(Q f ,K f ,V f )+D
[0053] D fn =FS(D * )
[0054] D sn =FS(FS(D * ))
[0055] D fn , D sn Represent the extracted frequency domain features and spatial domain features respectively. Then D fn , D sn The fusion is the final nonlinear mapping output, which is specifically expressed as follows:
[0056]
[0057] As you can understand, the size of the dark image V input to the dark light enhancement network is defined by three dimensions: C represents the number of channels, H represents the height, and W represents the width. The entire processing process of the network can be expressed as follows:
[0058]
[0059] Among them, conv(·) is a convolution operation with kernel set to 3, padding and stride set to 1. The SFII operation is as follows Figure 3 As shown, (i.e., concat(·)) means concatenating two tensors on the channel.
[0060] Finally, Perform 8 iterations of segmentation on the channel. The definition is as follows:
[0061]
[0062] in, Indicates segmenting the image in the channel dimension, Split into 8 parts along the channel dimension.
[0063] The final enhanced image is shown below:
[0064]
[0065] Among them, V 0 =V, V n represents the visible light image at the nth iteration (the final enhanced image), represents an iterated function, is the weight graph at the nth iteration.
[0066] It should be noted here that: most traditional image enhancement methods are based on the sRGB color space, but its color and brightness are highly coupled, which can easily cause color cast and brightness artifacts during enhancement. For example, when enhancing dark light images, color distortion may occur, resulting in poor visual effects of the image; methods based on Retinex theory usually regard the reflection component as the enhancement result, but in practical applications, this assumption is not always true, especially under various complex lighting conditions, which may lead to unrealistic enhancement effects. In addition, in the Retinex model, noise is usually ignored, so noise may be retained or amplified in the enhancement results. The dark light enhancement network designed in the scheme described in this embodiment can generate night-time visible light images with low noise and balanced brightness and chromaticity distribution. While maintaining image clarity, it can effectively improve the overall quality and visual effects of the image.
[0067] In a specific implementation, feature extraction is performed on the collected infrared image and the enhanced visible light image respectively. The encoders used include a basic encoder for extracting global features and a detail encoder for extracting local detail features. The feature representation of the infrared image and the visible light image is obtained by splicing the outputs of the basic encoder and the detail encoder.
[0068] Specifically, such as Figure 4 As shown in the figure, the base encoder focuses on extracting global features from the image. Its structure mainly consists of a 3×3 convolution block and three Transformer blocks. The above design enables the encoder to efficiently capture the overall properties of the image and provide rich feature information for the subsequent fusion process. Secondly, the detail encoder is dedicated to extracting detailed features of the image. Its core structure includes two 3×3 convolution layers and three residual blocks (Resblock). The design of Resblock can cleverly combine skip connections and dense connections. This design not only extracts features but also integrates gradient information, thereby effectively maintaining the texture details of the image. Traditional encoders usually use a single branch, which can only extract features from a single perspective and has difficulty capturing local details and global context information of the image at the same time.
[0069] In a specific implementation, the feature fusion is specifically as follows: based on the infrared image feature representation and the visible light image feature representation, the channel features and spatial features corresponding to the infrared image and the visible light image are obtained respectively through average pooling and maximum pooling operations; the channel features and spatial features of the infrared image and the visible light image are respectively spliced to obtain aggregated channel features and spatial features; based on the aggregated channel features and spatial features, convolution processing and activation function processing are sequentially performed to obtain channel weights and spatial weights; based on the infrared image feature representation and the visible light image feature representation, the obtained channel weights and spatial weights are combined and weighted processing is performed to obtain a fused image.
[0070] Specifically, such as Figure 5 As shown, the feature fusion first uses average pooling and maximum pooling to aggregate channel features and spatial features:
[0071]
[0072] Among them, S c ,S s Represents the aggregated channel features and spatial features. Specifically, S c It is through and Φ I The pooling operation is performed in the channel dimension to capture the global information between channels; S s It is through and Φ I The pooling operation is performed in the spatial dimension (ie, height and width) to capture the global information of the spatial position. c (·) and Max c (·) respectively represent global average pooling and global maximum pooling in the channel dimension. Avg h,w (·) and Max h,w (·) represents global average pooling and global maximum pooling in the spatial dimensions (i.e., height and width), respectively. The aggregated channel features are passed to two one-dimensional convolutions, and then the Softmax operation is performed to produce the final output channel weights.
[0073] W c1 ,W c2 =Conv1(S c ),Conv1(S c )
[0074]
[0075] Among them, Conv1(·) represents one-dimensional convolution, W ' c1 ,W ' c2 represents the output channel weight. The aggregated spatial features are passed to two 2D convolutions, and finally the spatial weights are output for feature fusion.
[0076] W s1 ,W s2 =Conv2(S s ),Conv2(S s )
[0077]
[0078] Among them, Conv2(·) represents two-dimensional convolution, W ' s1 ,W ' s2 Represents spatial weight.
[0079] The final weight is obtained by summing up the channel weight and spatial weight to determine the important features of fusion. The final fusion process is as follows:
[0080]
[0081] in, is the visible light image feature representation, Φ I It is the feature representation of infrared image.
[0082] It should be noted that the feature fusion scheme used in this embodiment combines channel and spatial attention mechanisms to more comprehensively extract features, enhance important features, and suppress noise, while maintaining its lightweight and easy integration characteristics. In comparison, existing traditional fusion schemes have certain shortcomings in terms of feature expression flexibility, dynamic adjustment capabilities, and computational overhead.
[0083] In a specific implementation, the brightness probability of the fused image outputted based on each round of iteration is calculated, and the following processing is specifically performed: based on the fused image outputted from each round of iteration, a number of convolution kernels of different sizes are used to extract the brightness features at different resolutions in the fused image; and the brightness probability of the current fused image is determined based on the fusion results of the brightness features at different resolutions.
[0084] Specifically, such as Figure 6 As shown, the calculation of the brightness probability specifically adopts the BFN (Brightness Feedback Network) network model, which takes the fused image as input and outputs the probability that the image belongs to daytime (pd) or nighttime (pn). Among them, BFN uses four different sizes of convolution kernels to extract brightness features of different resolutions in the image, thereby capturing the brightness information in detail. Then, the spatial information of these features is compressed through a 3×3 convolution layer. In the final stage of the process, global average pooling (GAP) and fully connected (FC) layers are used to calculate the final brightness probability, and then the brightness probability is used to calculate the brightness loss function (L bri ) further guides the training of the dark-light enhancement network. During this process, the Leaky ReLU (LReLU) activation function is introduced to increase the network's nonlinearity and improve prediction accuracy. Furthermore, the ReLU activation function filters out negative values, ensuring that the predicted probability values are effectively constrained to a reasonable range between 0 and 1.
[0085] Furthermore, in the solution described in this embodiment, the following loss function is used:
[0086] The fusion network based on deep learning adopts the following loss function:
[0087] Structural similarity loss L ssim :
[0088]
[0089] Among them, F is the fused image, I is the infrared image, V enA visible light image showing the initial enhanced image after adaptive histogram equalization. Adaptive histogram equalization is an effective image enhancement technique that evenly distributes excess pixel probabilities to other pixels, avoiding sudden increases in maximum brightness. This effectively improves local contrast, reduces noise, and preserves more detail.
[0090] Gradient-preserving loss L grad :
[0091]
[0092] in, Represents the gradient using the Sobel operator.
[0093] Color consistency loss L color :
[0094]
[0095] in, Representing the average value of the input image I in the R channel, we use the L2 norm and average operation to suppress the influence of outliers in the value of one of the RGB channels.
[0096]
[0097] Among them, F is the fused image, V en represents the visible light image of the initial enhanced image after adaptive histogram equalization, I is the infrared image, and β1 and β2 are hyperparameters.
[0098] The loss functions used by the deep learning-based dark light enhancement network include:
[0099] Spatial consistency loss L spatial :
[0100]
[0101] Where Avg(·) performs an average pooling operation, K is the window size of the average pooling (K=4 in this embodiment), Ω(i) represents the adjacent region centered on i, including the top, bottom, left, and right sides, J is one of the pixel indices of the adjacent region, used to traverse each pixel in Ω(i), and Y i To enhance the pixel value of image Y at position i, I i is the pixel value of the original image I at position i.
[0102] Smoothness loss L tv :
[0103]
[0104] Where I is the input image, i and j represent the pixel positions, and H and W represent the height and width of the image.
[0105] Color balance loss L color :
[0106]
[0107] in, Represents the average value of the input image I in the R channel, represents the average value of the input image I in the G channel, Represents the average value of the input image I in the B channel. This embodiment uses the L2 norm and the average operation to suppress the influence of abnormal values in the values of one of the RGB channels.
[0108] Exposure control loss L expose :
[0109]
[0110] Wherein, M is the window size of average pooling. In this embodiment, M is set to 8, the threshold ε is set to 0.5, I is the input image, and Avg(·) performs the average pooling operation.
[0111] Reconstruction loss Lr:
[0112]
[0113] Where N is the total number of elements in the image tensor, J i and G i are the values of the enhanced image and the original image at the i-th element respectively.
[0114] Brightness loss L bri :
[0115]
[0116] Among them, ε is a hyperparameter, when P d = 0, the maximum value is limited to 10 and set to 4.5×10 -5 , H, W represent the height and width of the image, P d is the probability that the image is a daytime image, P n is the probability that the image is a night image.
[0117] The brightness feedback network based on deep learning adopts the following loss function:
[0118] Cross entropy loss function L cro :
[0119]
[0120] Where y is a one-hot label indicating whether the input image is day or night, Represents the predicted probability, defined as 0 to 1, P d is the probability that the image is a daytime image, P n is the probability that the image is a night image.
[0121] In one or more embodiments, corresponding to the above method, this embodiment provides an infrared and visible light image fusion system for dark environments, including:
[0122] A multimodal image acquisition unit, which is used to synchronously acquire infrared images and visible light images of the same scene;
[0123] a visible light image enhancement unit, configured to perform image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image;
[0124] An image fusion unit is configured to obtain a fused image representation based on a captured infrared image and an enhanced visible light image through a pre-trained deep learning-based fusion network; wherein the fusion network specifically performs the following processing procedures: feature extraction is performed on the captured infrared image and the enhanced visible light image, respectively, to obtain an infrared image feature representation and a visible light image feature representation including local and global features; based on the infrared image feature representation and the visible light image feature representation, a fused image is obtained through feature fusion; and, during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
[0125] It can be understood that the system described in this embodiment corresponds to the method described in the above embodiment, and its technical details have been described in detail in the above method embodiment, so they will not be repeated here.
[0126] It is understandable that the solution described in this embodiment can be widely applied to the following scenarios:
[0127] Military reconnaissance: Improve the accuracy and reliability of target identification at night or in low-light environments.
[0128] Autonomous driving: Enhance the vehicle's visual perception capabilities when driving at night, improving road safety.
[0129] Security monitoring: Improve the nighttime imaging effect of the monitoring system and enhance the details of the monitoring images.
[0130] Through the solution described in this embodiment, high-quality fused images can be generated at night or in low-light environments, significantly improving the visual effect and information content of the images, and providing strong technical support for applications in related fields.
[0131] In further embodiments, there is also provided:
[0132] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the above embodiment. For the sake of brevity, no further details are given here.
[0133] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0134] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0135] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in the above embodiment is completed.
[0136] The methods in the above embodiments can be directly implemented and executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in a memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above methods. To avoid repetition, a detailed description is not given here.
[0137] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0138] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.
Claims
1. A method for fusion of infrared and visible light images in dark environments, characterized in that: include: Synchronously collect infrared images and visible light images of the same scene; Performing image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image; Based on the captured infrared image and the enhanced visible light image, a fused image representation is obtained through a pre-trained deep learning-based fusion network. The fusion network specifically performs the following processing: feature extraction is performed on the captured infrared image and the enhanced visible light image, respectively, to obtain a feature representation of the infrared image and a feature representation of the visible light image including local and global features; Based on the feature representation of infrared images and the feature representation of visible light images, a fused image is obtained through feature fusion; and during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
2. The infrared and visible light image fusion method for dark environments according to claim 1, wherein: The dark light enhancement network specifically performs the following processing: taking the collected visible light image as input, obtaining a feature map through several convolutional layers and dynamic spectrum filtering operations; decomposing the obtained feature map into several weighted feature maps; and iteratively multiplying the visible light image with the obtained weighted feature map to obtain an enhanced visible light image.
3. The infrared and visible light image fusion method for dark environments according to claim 1, wherein: The dark light enhancement network includes seven sequentially connected convolutional layers and three dynamic spectrum filtering modules, wherein the first convolutional layer, the second convolutional layer, the third convolutional layer, the fifth convolutional layer, the sixth convolutional layer and the seventh convolutional layer are all composed of sequentially connected convolution blocks and activation functions, and the fourth convolutional layer only includes convolution blocks; the input of the fifth convolutional layer is obtained by splicing the output result of the first convolutional layer after being processed by the first dynamic spectrum filtering module and the output result of the fourth convolutional layer; the input of the sixth convolutional layer is obtained by splicing the output result of the second convolutional layer after being processed by the second dynamic spectrum filtering module and the output result of the fifth convolutional layer; the input of the seventh convolutional layer is obtained by splicing the output result of the third convolutional layer after being processed by the third dynamic spectrum filtering module and the output result of the sixth convolutional layer.
4. The infrared and visible light image fusion method for dark environments according to claim 3, wherein: The dynamic spectrum filtering module specifically performs the following processing: based on the output results of the convolution layer, different frequency domain information features are obtained through dynamic spectrum filtering operations; based on the obtained different frequency domain information features, spatial domain learning is performed locally through the self-attention mechanism to obtain the final nonlinear mapping output.
5. The infrared and visible light image fusion method for dark environments according to claim 1, wherein: The feature extraction is performed on the collected infrared image and the enhanced visible light image respectively. The encoders used include a basic encoder for extracting global features and a detail encoder for extracting local detail features. The feature representation of the infrared image and the visible light image is obtained by splicing the outputs of the basic encoder and the detail encoder.
6. The infrared and visible light image fusion method for dark environments according to claim 1, wherein: The feature fusion is specifically as follows: based on the infrared image feature representation and the visible light image feature representation, through average pooling and maximum pooling operations, the channel features and spatial features corresponding to the infrared image and the visible light image are obtained respectively; The channel features and spatial features of the infrared image and the visible light image are spliced separately to obtain aggregated channel features and spatial features; Based on the aggregated channel features and spatial features, the convolution and activation functions are processed sequentially to obtain the channel weights and spatial weights; Based on the feature representation of infrared images and visible light images, the obtained channel weights and spatial weights are combined and weighted processing is performed to obtain a fused image.
7. The infrared and visible light image fusion method for dark environments according to claim 1, wherein: The brightness probability of the fused image outputted in each round of iteration is calculated by specifically performing the following processing: based on the fused image outputted in each round of iteration, a number of convolution kernels of different sizes are used to extract brightness features at different resolutions in the fused image; The brightness probability of the current fused image is determined based on the fusion results of brightness features at different resolutions.
8. An infrared and visible light image fusion system for dark environments, characterized in that: include: A multimodal image acquisition unit, which is used to synchronously acquire infrared images and visible light images of the same scene; a visible light image enhancement unit, configured to perform image enhancement processing on the visible light image using a pre-trained dark light enhancement network based on deep learning to obtain an enhanced visible light image; An image fusion unit is configured to obtain a fused image representation based on the captured infrared image and the enhanced visible light image using a pre-trained deep learning-based fusion network. The fusion network specifically performs the following processing: extracting features from the captured infrared image and the enhanced visible light image to obtain a feature representation of the infrared image and a feature representation of the visible light image, each including local and global features; Based on the feature representation of infrared images and the feature representation of visible light images, a fused image is obtained through feature fusion; and during the iterative training process of the fusion network, the brightness probability of the fused image outputted in each round of iteration is calculated, and the obtained brightness probability is fed back to the dark light enhancement network to guide the brightness of the image outputted by the dark light enhancement network.
9. An electronic device comprising a memory, a processor, and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the infrared and visible light image fusion method for a dark environment as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the infrared and visible light image fusion method for dark environments as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Night infrared and visible image fusion method and system with cooperative illumination enhancement and modal balance
CN121437286A