Low visible light image processing method and processing device
Through multi-scale feature extraction and generation-discrimination competition mechanism, combined with image gradient and texture detail loss functions, the robustness and visual effect problems of low-visible-light images in complex environments are solved, and high-quality fused images are generated.
Patent Information
- Application Number
- CN202510703773.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing low-visible light image processing methods have low accuracy and poor robustness in complex environments. Traditional fusion methods fail to effectively utilize the complementary information of infrared and visible light sensors, resulting in low fusion performance and poor visual effects.
By acquiring multi-scale features of infrared images and low-visible light images, utilizing a generative-discriminative competitive mechanism and a mapping module based on a convolutional neural network, and combining them with high-quality visible light images to generate output images, the image is optimized through image gradient and texture detail loss functions to achieve image enhancement.
The robustness and visual effects of low-visible light images in complex environments are improved. The generated images retain the structural features of infrared images and the texture details of low-visible light images, improving the quality of image enhancement.
Smart Images

Figure CN120235774B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image enhancement technology, and particularly relates to a method and a device for processing low-visible-light images. Background Art
[0002] Low-visible-light images often suffer from low contrast, low brightness, and high noise. These issues severely impact human perception and downstream tasks such as object detection and image segmentation. Therefore, low-visible-light image enhancement technology holds significant practical value. Current object detection algorithms are mostly based on single-modal low-visible-light images as training data. These algorithms suffer from low accuracy and poor robustness in complex environments, such as poor lighting conditions and rainy or foggy conditions. Consequently, in recent years, a growing number of researchers have focused on combining low-visible-light and infrared imagery. By combining the advantages of low-visible-light images, which offer clear target texture features, with the advantages of infrared images, which are unaffected by lighting conditions and offer clear target outlines, these multimodal data can be used to train detection networks, thereby improving the robustness of detection algorithms. Fusion of infrared and low-visible-light images aims to combine the strengths of both sensor types. The resulting fused image offers enhanced target perception and scene representation, facilitating both human observation and subsequent computational processing. Infrared sensors are sensitive to thermal radiation and can capture prominent target regions, but the resulting infrared images often lack structural features and texture detail. In contrast, visible light sensors capture rich scene information and texture details through light reflection imaging. While visible light images have high spatial resolution and rich texture detail, they cannot effectively highlight target features and are easily affected by the external environment, particularly in low-light conditions, where information loss is severe. Due to the different imaging mechanisms of infrared and visible light, these two types of images possess highly complementary information. Only through the use of fusion technology can the collaborative detection capabilities of infrared and visible light imaging sensors be effectively enhanced. This has broad applications in remote sensing, medical diagnosis, intelligent driving, security monitoring, and other fields.
[0003] Currently, infrared and low-visible light image fusion technologies can be broadly categorized into traditional fusion methods and deep learning fusion methods. Traditional image fusion methods typically extract image features using the same feature transformation or representation, merge them using appropriate fusion rules, and then reconstruct the final fused image through inverse transformation. Due to the different imaging mechanisms of infrared and visible light sensors, infrared images represent target features using pixel brightness, while visible light images represent scene texture using edges and gradients. Traditional fusion methods ignore the inherently different characteristics of the source images and indiscriminately extract image features using the same transformation or representation model. This inevitably results in poor fusion performance and poor visual quality. Furthermore, fusion rules are manually set and increasingly complex, resulting in high computational costs, which limits the practical application of image fusion. Summary of the Invention
[0004] The purpose of the present invention is to propose a low-visible light image processing method by introducing infrared images that are not restricted by lighting conditions as an auxiliary modality, thereby achieving effective enhancement of exposure-deficient scene images, and providing a low-visible light image processing method and processing device.
[0005] The present invention is achieved through the following technical solutions:
[0006] A method for processing low-visible-light images comprises the following steps:
[0007] Step S1: acquiring and integrating multi-scale features of the infrared image and the low-visible-light image respectively, and then fusing the integrated multi-scale features and calculating the fused features;
[0008] Step S2: using the fused features as input to a generator, combined with a high-quality visible light image, and generating an output image using a generative-discriminative competitive mechanism. The generator includes a feature encoder and a mapping module based on a convolutional neural network.
[0009] Step S3: Calculate the structural loss function of the image gradient between the infrared image and the output image, use the VGG encoding network to extract the texture detail features in the low visible light image and construct a texture detail feature loss function, and generate the final image based on the structural loss function and the texture detail feature loss function.
[0010] As one of the preferred solutions, step S1 specifically includes the following steps:
[0011] Step S11: using two independent neural networks based on the UNet architecture to encode the infrared image and the low-visible light image respectively to obtain multi-scale features of the infrared image and the low-visible light image;
[0012] Step S12: performing parallel processing of the convolution layer and the attention layer on the multi-scale features of each infrared image and low-visible light image, and integrating them through an addition operation;
[0013] Step S13: Use a feature fusion module based on modal balance to fuse the integrated features. The steps of the feature fusion module include: first, obtaining a dual-mode combination feature by element-by-element addition, then processing the dual-mode combination feature separately by two parallel branches to obtain a global weight representing the overall relationship between the modalities and a local weight representing the local detail relationship between the modalities; finally, the global weight and the local weight are added together and then subjected to a Sigmoid activation function to obtain the final weight of the two-modal features and output the fused feature.
[0014] As one of the preferred solutions, in the branch for calculating the global weight in step S13, the global pooling operation is first performed and the spatial scale is converted to The global features of the two modal information are then processed by point-by-point convolution and ReLU activation layers to obtain the global weights of the two modal information. The branch for calculating the local weights does not undergo pooling operation but directly undergoes point-by-point convolution and pooling layer processing to obtain the local weights of the two modalities.
[0015] As one of the preferred solutions, in step S2, the loss functions of the generator and the discriminator are respectively , ,in, represents a high-quality visible light image with normal exposure, represents the input infrared image, represents the input low visible light image, is the discriminator network, The image generated by the generator, is the generator loss function, is the discriminator loss function.
[0016] As one of the preferred solutions, in step S3, the infrared image and the output image gradient are first calculated by the Sobel operator. and , then, through Paradigm calculation model structure loss function ,in ,in, is a loss function that measures the difference between the generated infrared image and the output image, To calculate the difference between two images norm, and Table infrared image and output image, and is the image gradient in the x direction, and is the image gradient in the y direction.
[0017] As one of the preferred solutions, in step S3, the image texture detail features are first extracted through the pre-trained VGG encoding network, and then the VGG feature distance between the generated image and the low visible light image is optimized, wherein the texture loss function based on the depth feature is mathematically expressed as:
[0018]
[0019] in, represents the VGG-16 network pre-trained on ImageNet, Euclid Norm, used to measure the length of a vector or the distance between two vectors, is used in this formula to measure and The degree of difference between them.
[0020] A low-visible-light image processing device includes a processor and a memory, wherein the memory stores program instructions, and the processor calls the program instructions in the memory to implement the following modules:
[0021] A feature acquisition module is used to acquire and integrate the multi-scale features of infrared images and low-visible light images respectively, and then fuse the integrated multi-scale features to calculate the fused features;
[0022] A feature fusion module is used to use the fused features as input to a generator, combined with a high-quality visible light image, and generate an output image using a generative-discriminative competitive mechanism. The generator includes a feature encoder and a mapping module based on a convolutional neural network;
[0023] The image reconstruction module is used to calculate the structural loss function of the image gradient between the infrared image and the output image, use the VGG encoding network to extract the texture detail features in the low visible light image, and generate the final image based on the structural loss function and the texture detail feature loss function.
[0024] The present invention has the following beneficial effects:
[0025] The present invention first realizes the extraction and fusion of multi-scale, high- and low-frequency features through a multi-scale coding network structure and a convolution-channel attention-based coding module; utilizes the generation-discrimination competition mechanism of normal exposure images to realize the mapping of low-light (low visible light) and infrared information to the high-quality visible light image domain; adopts the structure loss function based on image gradient and the texture loss function based on depth features to effectively integrate the saliency features of the infrared modality and the texture details of the low visible light modality, ensuring that the generated image retains the advantageous information of the original image, realizing image enhancement in exposure-poor scenarios, improving the robustness of image enhancement, and providing a new method for solving the image enhancement problem in exposure-poor scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention will be further described in detail below with reference to the accompanying drawings.
[0027] Figure 1 Flowchart of the method of the present invention.
[0028] Figure 2 1 is a principle block diagram of the device of the present invention. DETAILED DESCRIPTION
[0029] like Figure 1 and Figure 2 As shown, the low-visible-light image processing method and device include the following steps:
[0030] Step S1: Acquire and integrate multi-scale features of the infrared image and the low-visible light image respectively, then fuse the integrated multi-scale features and calculate the fusion features, such as through a convolution-channel attention-based encoding network and a modality balance-based feature fusion module to achieve effective extraction, integration, and fusion of high- and low-frequency features;
[0031] Specifically, first, two independent neural networks based on the UNet architecture are used to respectively analyze infrared images and low-visible light images. and Encode to obtain multi-scale features and multi-scale features ,in, is the total number of feature scales of the UNet network, represents the encoded image features, In the present invention, the infrared image and the low visible light image refer to two images taken at the same viewing angle and at a similar time, and are highly aligned with the low visible light image. In particular, in the encoding and subsequent decoding process, an attention-convolution joint structure is used as the basic module of the fusion network, and any intermediate feature is used. For example, the feature will be processed in parallel by the convolution layer and the attention layer, and reintegrated through the addition operation, that is: ,in, represents the convolution operation, It represents the attention operation, that is, through a heterogeneous dual-path architecture, it performs collaborative processing through convolutional layers and attention mechanisms, deeply integrating local spatial perception and global semantic association.
[0032] Subsequently, the encoded infrared and low-visible light features are fused using a feature fusion module based on modal balance. Specifically, given the features of an infrared image and a low-visible light image at a certain scale, and The feature fusion module first obtains the dual-mode combined feature by element-by-element addition , and then the combined features The two parallel branches process the modalities separately to obtain a global weight representing the overall relationship between the modalities. and local weights that represent the local detail relationships between modalities Finally, the global weight and the local weight are added together and then the final weight of the two modal features is obtained through the Sigmoid activation function. This weight is used to coordinate the two modal features. and The relationship between them and output fusion features ,Right now ,in, For broadcast addition, Represents the Sigmoid activation function. and is the eigenvalue of the infrared image and low visible light image at the jth scale.
[0033] In the branch used to calculate the global weight, the feature First, after the global pooling operation, it is converted into a spatial scale of The global feature of , then the feature is convolved with a kernel size of After the convolution layer (point-by-point convolution) and Relu activation layer processing, the global weight of the two modal information is obtained , the mathematical representation of this branch is: ,in, represents the global pooling operation, and are two independent point-by-point convolutional layers, is the Relu activation function. In the branch used to calculate the local weight, the combined features Without pooling operation, the local weights of the two modes are obtained by directly processing the point-by-point convolution and pooling layers. The mathematical expression is as follows: .
[0034] Step S2: Using the fused features as input to a generator, combined with the high-quality visible light image, and generating an output image using a generative-discriminative competitive mechanism. The generator includes a feature encoder and a mapping module based on a convolutional neural network. This step maps the low-visible light image and infrared image information to the high-quality visible light image domain.
[0035] The loss functions of the generator and discriminator are , ,in, A high-quality visible light image representing normal exposure. represents the input infrared image, is the discriminator network, represents the input low visible light image, The image generated by the generator, that is, the fusion of the two features {Fn}n=1N is used to obtain the generated image of the generator through the mapping module , is the generator loss function, is the discriminator loss function. After the generation-discrimination competition mechanism, the output enhanced image is , is the generator loss function, is the discriminator loss function.
[0036] like Figure 2 As shown, Receive normal exposure images for the discriminator The output when Receive the generator output image for the discriminator The output of the discriminator is To construct the generator loss function , and use the generator loss function Optimize the generator parameters to achieve the opposite effect on the generator, This is the generator loss function used to optimize the generator parameters.
[0037] That is, the fusion feature set After being processed by the decoding module of the fusion network and the mapping module based on the convolutional neural network, the output image of the generator is obtained. is the total number of layers in the fusion network. Using this generative-discriminative competitive mechanism, the enhanced image retains the excellent properties of high brightness and visibility of the normally exposed image. By constructing an adversarial loss function using images collaboratively generated from infrared and low-visible-light images and high-quality normally exposed images, we can fully leverage the complementarity of multimodal information, enhancing the detail and realism of the generated image. Furthermore, through the adversarial learning mechanism, we ensure that the generated image achieves optimal visual quality and exposure balance.
[0038] Step S3: Calculate the structure loss function for the image gradients of the infrared image and the output image. The VGG encoding network is used to extract texture detail features from the low-visible-light image and construct a texture detail feature loss function. The final image is generated based on the structure loss function and the texture detail feature loss function. The structure loss function works by calculating the difference in image gradients between the infrared image and the output image. Image gradients reflect changes in pixel values within an image and are closely related to structural information such as edges and contours. The structure loss function ensures that the output image maintains similarity to the infrared image in terms of edges and contours. The processed image preserves key structural information such as the shape and position of objects in the infrared image, avoiding structural distortion or loss during the fusion and generation process, and ensuring that the output image is more structurally consistent with the actual scene. The powerful feature extraction capabilities of the VGG encoding network are also utilized to extract texture detail features from the low-visible-light image. The VGG encoding network excels in image feature extraction and can capture rich texture information. The texture detail feature loss function works by constraining the consistency of texture detail features between the generated image and the low-visible-light image. The final image retains texture details in low-visible light images, such as surface lines and subtle patterns. This helps enhance the image's realism and detail, improving the visual quality of the processed image and facilitating subsequent image analysis and understanding. By optimizing the pattern at different levels using a structure loss function based on image gradients and a texture loss function based on deep features, the system effectively integrates the saliency of the infrared modality and the texture details of the low-visible light modality, ensuring that the generated image retains the valid original information of both the infrared and low-visible light images.
[0039] Specifically, the integration of the significant structural features of the infrared image is achieved through the loss function based on the Sobel operator. First, the gradient of the infrared image and the output image is calculated by the Sobel operator. and , then, through Paradigm calculation model structure loss ,in, is a loss function that measures the difference between the generated infrared image and the output image, To calculate the difference between two images norm, and Table infrared image and output image, and is the image gradient in the x direction, and is the image gradient in the y direction.
[0040] Secondly, the texture loss function based on deep features is used to extract and integrate the texture detail features of low-visible light images. Specifically, the image texture detail features are first extracted through the pre-trained VGG encoding network. Then, the VGG feature distance between the generated image and the low-visible light image is optimized to ensure that the color and complex texture of the generated image are consistent with the original information. The texture loss function based on deep features is mathematically expressed as:
[0041]
[0042] in, represents the VGG-16 network pre-trained on ImageNet, Euclid Norm, used to measure the length of a vector or the distance between two vectors, is used in this formula to measure and The degree of difference between them.
[0043] The loss function based on the Sobel operator can effectively capture the edge and detail features of the image by calculating the gradient information of the image, thereby improving the structural accuracy of the generated image; at the same time, combined with The norm can enhance the robustness of the loss function, reduce noise interference, and ensure that the generated image is highly consistent with the real image in terms of details. Since the pre-trained VGG model is insensitive to adjustments to the pixel intensity range of the input image, the introduction of the VGG feature distance can effectively reduce the impact of the low brightness problem of the original low-visible light image on the generated image. During the optimization process, multi-level feature matching is adopted to extract feature maps at different depths of the VGG network to ensure the precise alignment of texture details from local to global. At the same time, an adaptive weight mechanism is introduced through the texture loss function to dynamically adjust the contribution of different feature layers, focusing on texture-rich areas, thereby improving the color saturation and texture realism of the generated image, ensuring the consistency of the color and complex texture of the generated image with the original information.
[0044] Specifically, there are four types of images in the whole invention, namely, input low visible light image , input infrared image , input normal visible light image (high-quality visible light image) for adversarial training , and the resulting image (output image) The relationship between the generated output image and the other three input images is: and Both belong to the high-quality visible light image domain, compared with the input low-visible light image Infrared image with high brightness, high visibility and fusion input Excellent properties of structural information.
[0045] The structure loss function and texture detail loss function are used to preserve the structural features of infrared images and the texture detail features of low-visible-light images, respectively. A feature mapping network based on a UNet network is used to generate a high-quality final image. The ultimate goal of this invention is to output a high-brightness, high-visibility low-light image that effectively integrates the texture details in the original low-light image with the overall structural information of the infrared image.
[0046] At the same time, the present invention also discloses a low-visible light image processing device based on an adversarial mapping strategy and multimodal information preservation, which includes a processor and a memory, wherein the memory stores program instructions, and the processor calls the program instructions in the memory to implement the following modules:
[0047] Feature Acquisition Module: This module is used to separately acquire and integrate multi-scale features from infrared and low-visible-light images. The integrated multi-scale features are then fused and the fused features are calculated, yielding multi-scale, high- and low-frequency features for the infrared and low-visible-light images, respectively. The extracted features contain rich detail information as well as global semantic information. This is achieved through a combined attention-convolution architecture. Because the convolutional layer focuses only on the feature information of the local region within the convolution kernel during its sliding process, it can accurately encode high-frequency detail features. The attention mechanism, by calculating the relationship between each feature element and all other elements, can efficiently capture global low-frequency features.
[0048] A feature fusion module is used to use the fused features as input to a generator, combined with a high-quality visible light image, and generate an output image using a generative-discriminative competitive mechanism. The generator includes a feature encoder and a mapping module based on a convolutional neural network;
[0049] Image Reconstruction Module: This module calculates the structural loss function for the image gradient between the infrared image and the output image. It then uses the VGG encoding network to extract texture detail features from the low-visible-light image and generates the final image based on the structural loss function and texture detail features. Guided by the adversarial loss function, the saliency structure loss function, and the texture loss function, this module achieves high-quality visible light image reconstruction by fusing the complementary information of the two modal images.
[0050] It also includes constructing a feature memory storage device, which stores multi-level features of infrared and low-visible light images in the memory storage device, and fuses the stored features using a feature fusion module based on modal balance; it facilitates accelerated calculations, and the encoded features can be directly called from the memory to improve efficiency.
[0051] The above description is merely a preferred embodiment of the present invention and therefore cannot be used to limit the scope of the present invention. In other words, equivalent changes and modifications made according to the scope of the patent application and the contents of the specification should still fall within the scope of the patent of the present invention.
Claims
1. A method for processing low-visible-light images, characterized by: The following steps are included: Step S1: The generator includes a fusion network and a mapping network based on a convolutional neural network. The fusion network of the generator is used to obtain and integrate multi-scale features of the input infrared image and the low-visible light image, and then fuse the integrated multi-scale features to calculate the fusion feature; The fusion features are passed through the mapping network to obtain the generated image of the generator, which is combined with the high-quality visible light image and the output image is generated using the generation-discrimination competition mechanism; Step S2: Calculate the structural loss function of the image gradient between the infrared image and the output image, use the VGG encoding network to extract the texture detail features in the low visible light image and construct a texture detail feature loss function, and generate the final image based on the structural loss function and the texture detail feature loss function.
2. The low-visible-light image processing method according to claim 1, wherein: The step S1 specifically includes the following steps: Step S11: using two independent neural networks based on the UNet architecture to encode the infrared image and the low-visible light image respectively to obtain multi-scale features of the infrared image and the low-visible light image; Step S12: performing parallel processing of the convolution layer and the attention layer on the multi-scale features of each infrared image and low-visible light image, and integrating them through an addition operation; Step S13: Use a feature fusion network based on modal balance to fuse the integrated features. The steps of the feature fusion network include: first, obtaining a dual-mode combination feature by element-by-element addition, then processing the dual-mode combination feature separately by two parallel branches to obtain a global weight representing the overall relationship between the modalities and a local weight representing the local detail relationship between the modalities; finally, the global weight and the local weight are added together and then subjected to a Sigmoid activation function to obtain the final weight of the two-modal features and output the fused feature.
3. The low-visible-light image processing method according to claim 2, wherein: In the branch for calculating the global weight in step S13, the global pooling operation is first performed and then converted into a spatial scale of The global features of the two modal information are then processed by point-by-point convolution and ReLU activation layers to obtain the global weights of the two modal information. The branch for calculating the local weights does not undergo pooling operation but directly undergoes point-by-point convolution and pooling layer processing to obtain the local weights of the two modalities.
4. The low-visible-light image processing method according to claim 1, wherein: In step S1, the loss functions of the generator and the discriminator are , ,in, represents a high-quality visible light image with normal exposure, represents the input infrared image, represents the input low visible light image, The image generated by the generator, is the discriminator network, is the generator loss function, is the discriminator loss function.
5. The low-visible-light image processing method according to claim 1, wherein: In step S2, the infrared image and the output image gradient are first calculated by the Sobel operator. and , then, through Paradigm calculation model structure loss function ,in ,in, is a loss function that measures the difference between the generated infrared image and the output image, To calculate the difference between two images norm, and Table infrared image and output image, and is the image gradient in the x direction, and is the image gradient in the y direction.
6. The low-visible-light image processing method according to claim 5, wherein: In step S2, the image texture detail features are first extracted through the pre-trained VGG encoding network, and then the VGG feature distance between the generated image and the low visible light image is optimized, wherein the texture loss function based on the depth feature is mathematically expressed as: , in, represents the VGG-16 network pre-trained on ImageNet, Euclid Norm, used to measure the length of a vector or the distance between two vectors, is used in this formula to measure and The degree of difference between them.
7. A low-visible-light image processing device, characterized by: It includes a processor and a memory, wherein the memory stores program instructions, and the processor calls the program instructions in the memory to implement the following modules: An output image acquisition module, wherein the generator includes a fusion network and a mapping network based on a convolutional neural network. The fusion network of the generator is used to obtain and integrate multi-scale features of the input infrared image and low-visible light image respectively, and then fuse the integrated multi-scale features to calculate the fused features; the fused features are passed through the mapping network to obtain the generated image of the generator, which is combined with the high-quality visible light image and the output image is generated using a generation-discrimination competition mechanism; The image reconstruction module is used to calculate the structural loss function of the image gradient between the infrared image and the output image, use the VGG encoding network to extract the texture detail features in the low visible light image, and generate the final image based on the structural loss function and the texture detail feature loss function.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-scale generative adversarial network
CN111145131A
Infrared and visible light fusion method based on multi-scale feature interaction enhancement
CN119091269A