Weak light image enhancement method based on multi-scale feature fusion

Through the multi-scale feature fusion method, the bottom-level, middle-level and high-level features of low-light images are decomposed and enhanced, which solves the problem of feature confusion and semantic information attenuation in traditional methods and improves the visual quality and feature expression ability of low-light images.

CN120689226APending Publication Date: 2025-09-23CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510768992.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional low-light image enhancement methods suffer from feature confusion and detail degradation problems in complex scenes. Deep learning models suffer from insufficient representation of single-scale features, which leads to semantic information attenuation and affects image visual quality and machine vision parsing capabilities.

Method used

A multi-scale feature fusion method is used to decompose the low-light image into bottom-level, middle-level and high-level images, and four-level convolution and channel splicing are performed respectively. Through self-attention feature enhancement and multi-scale feature fusion, a multi-scale information feature map is generated, and finally the enhanced low-light image is obtained through comprehensive processing.

Benefits of technology

The visual quality and feature expression capability of low-light images are significantly improved, with PSNR and SSIM increased by 5.06dB and 0.22 respectively, and image feature expression capability improved by 1.4%, showing excellent performance in target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689226A_ABST
    Figure CN120689226A_ABST
Patent Text Reader

Abstract

The invention discloses a weak light image enhancement method based on multi-scale feature fusion. The method comprises the following steps: decomposing a weak light image into a bottom layer image, a middle layer image and a high layer image; respectively carrying out four-level convolution and splicing along a channel direction on the bottom layer image, the middle layer image and the high layer image to generate a multi-scale feature group consisting of a low layer feature map, a middle layer feature map and a high layer feature map; enhancing the high-level feature map to obtain a high-level enhanced feature map; performing fusion processing on a multi-scale feature group composed of the low-layer feature map, the middle-layer feature map and the high-layer enhanced feature map to obtain a multi-scale information feature map; and comprehensively processing the low-layer feature map, the middle-layer feature map and the multi-scale information feature map to obtain an enhanced weak light image. According to the weak light image enhancement method based on multi-scale feature fusion, through organic combination of a hierarchical feature decoupling strategy, a feature aggregation enhancement strategy and a multi-scale feature fusion mechanism, the visual quality of a weak light image is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for enhancing a low-light image. Background Art

[0002] Visual cognition, one of the key perceptual pathways for humans to understand the objective world, relies on images, a high-dimensional information carrier, to encode and transmit environmental information. Furthermore, image quality directly affects the entropy distribution and semantic fidelity of visual information. High-quality images effectively maintain information integrity, facilitating accurate observation and interpretation. Conversely, image degradation leads to a decrease in information channel capacity, resulting in reduced information transmission efficiency, which in turn can induce cognitive ambiguity and decision-making bias. Image degradation is particularly pronounced in low-light environments, manifesting as complex defects such as contrast loss, color shift, and noise contamination. This degradation not only increases the information decoding burden on the human visual system but also impairs the feature resolution capabilities of machine vision systems.

[0003] The rapid development of deep learning technology has led to breakthroughs in the field of low-light image enhancement. Compared to traditional methods based on hand-crafted designs and physical models, deep learning methods demonstrate significant advantages in performance enhancement, algorithmic robustness, and processing efficiency, attracting increasing attention. Therefore, leveraging deep learning methods and physical models to explore more effective low-light image enhancement methods, improving image visual quality and feature expression to meet the needs of both human visual perception and machine vision analysis, has become a critical and pressing issue in the field of computer vision.

[0004] Low-light image enhancement technology plays an indispensable role in key areas such as autonomous driving, security monitoring, and industrial manufacturing. In low-light environments, common visual defects such as insufficient contrast, increased noise, and loss of detail and texture can significantly reduce system reliability and safety. Summary of the Invention

[0005] In view of this, the present invention provides a low-light image enhancement method based on multi-scale feature fusion to overcome the limitations of traditional low-light image enhancement methods such as feature confusion and detail degradation in complex scenes, as well as the semantic information attenuation problem caused by insufficient single-scale feature representation in existing deep learning models, with the aim of improving the visual quality of low-light images.

[0006] The low-light image enhancement method based on multi-scale feature fusion of the present invention comprises the following steps:

[0007] 1) Decompose the low-light image into bottom-layer image, middle-layer image and high-layer image;

[0008] 2) Perform four-level convolution and channel-wise concatenation on the bottom-level image, middle-level image, and high-level image, respectively, to generate a multi-scale feature group consisting of low-level feature maps, middle-level feature maps, and high-level feature maps;

[0009] 3) Enhance the high-level feature map to obtain a high-level enhanced feature map;

[0010] 4) Fusing the multi-scale feature group consisting of the low-level feature map, the middle-level feature map, and the high-level enhanced feature map to obtain a multi-scale information feature map;

[0011] 5) Comprehensively process the low-level feature map, mid-level feature map and multi-scale information feature map to obtain an enhanced low-light image.

[0012] Furthermore, in step 1), the process of decomposing the low-light image is as follows:

[0013]

[0014] Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the reflection filling size; I is the low-light image, I b is the underlying image, I m is the middle layer image, I u For high-level images.

[0015] Furthermore, in step 2), the process of performing four-level convolution and channel-wise splicing on the bottom layer image, middle layer image, and high layer image is as follows:

[0016]

[0017] Where Conv is the conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size. The features are spliced ​​along the channel direction; the subscript x represents b, m and u; the generated multi-scale feature group is {F b ,F m ,F u}, where F b is the low-level feature map, F m is the middle-level feature map, F u is a high-level feature map.

[0018] Furthermore, in step 3), the process of enhancing the high-level feature map is as follows:

[0019]

[0020] Among them, Reshape is feature reshaping, Tr is feature transposition; Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size;

[0021]

[0022] Among them, Reshape is feature reshaping, is a matrix multiplication operation; Conv is a regular convolution operation.

[0023] Furthermore, in step 4), the process of fusing the multi-scale feature groups is as follows:

[0024]

[0025] Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the reflection padding size;

[0026]

[0027] Among them, ⊙ is the element-wise multiplication operation;

[0028]

[0029] Among them, Sum{·} is the dimension reduction summation function along the sequence dimension T, It is splicing along the sequence dimension T;

[0030]

[0031] Among them, Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size; F ms It is a multi-scale information feature map.

[0032] Furthermore, in step 5), the process of comprehensively processing the low-level feature map, the middle-level feature map, and the multi-scale information feature map is as follows:

[0033]

[0034] Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the reflection padding size; Up is the bilinear interpolation upsampling, sf is the upsampling scale factor; I local is the local image information, I global is the global image information; ms is the multi-scale image information, In m is the middle-level image information, In b is the underlying image information; ⊙ is the element-wise multiplication operation, α b is the underlying local coefficient, α m is the mid-layer local coefficient, β m is the middle-level global coefficient, β msis the multi-scale global coefficient; I e It is the enhanced low-light image.

[0035] Beneficial effects of the present invention:

[0036] The low-light image enhancement method based on multi-scale feature fusion in this paper effectively improves the visual quality of low-light images by organically combining hierarchical feature decoupling strategy, feature aggregation enhancement strategy and multi-scale feature fusion mechanism. The comparative test results show that on the LOL-v1 low-light image enhancement dataset, the PSNR and SSIM achieved by the method of this paper are improved by 5.06dB and 0.22 compared with the SCI algorithm, and the visual quality of the image is significantly improved; at the same time, the low-light image enhanced by the method of this paper has a mAP of 0.06dB under the RetinaNet detection framework. 50 Compared with the original low-light image, it is improved by 1.4%, and the image's feature expression ability is enhanced to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the network structure of the low-light image enhancement method based on multi-scale feature fusion provided by the present invention.

[0038] Figure 2 Schematic diagram of the network structure of the Feature Aggregation Module (FAM).

[0039] Figure 3 Schematic diagram of the network structure of the Self-attention Feature Enhancement Module (SFEM).

[0040] Figure 4 Schematic diagram of the experimental results of comparing the low-light image enhancement performance of 8 advanced low-light image enhancement models including PairLIE, IAT and Retinexformer with Multi-Enhancer for experimental verification on the LOL-v1 public low-light image enhancement dataset.

[0041] Figure 5 Schematic diagram of the experimental results of the ExDark dataset selected for target detection experiments in low-light environments, using the classic single-stage detection model RetinaNet and the two-stage detection model Sparse R-CNN as the baseline target detection model (implemented based on the MMDetection framework). DETAILED DESCRIPTION

[0042] The present invention will be further described below with reference to the accompanying drawings and examples.

[0043] The low-light image enhancement method based on multi-scale feature fusion in this embodiment includes the following steps:

[0044] 1) Decompose the low-light image into bottom-layer image, middle-layer image and high-layer image. This step uses the Image Decomposition Module (IDM) to decompose the input low-light image. The decomposition process is as follows:

[0045]

[0046] Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the reflection padding size, and no activation function is used; decomposition obtains the underlying image Middle-level image and high-level images The aforementioned IDM consists of three sets of reflection convolutions, which use standard 2D convolution operations with 2D reflection filling. They can efficiently decompose low-light images into three multi-level images with different resolutions, so as to extract and enhance the corresponding features in a targeted manner.

[0047] 2) The bottom-level image, middle-level image and high-level image are respectively subjected to four-level convolution and splicing along the channel direction to generate a multi-scale feature group consisting of low-level feature maps, middle-level feature maps and high-level feature maps.

[0048] This step uses the fully convolutional inter-scale feature aggregation module (FAM) to generate multi-scale feature groups, aiming to capture the key spatial and channel information in the image through the feature aggregation principle. The network structure of FAM is as follows: Figure 2 As shown. When FAM aggregates features of the input image, it optimizes the input image step by step by adopting a four-level progressive convolution structure. In the specific implementation, the input of each level of convolution operation is composed of all previous level convolution outputs and the input image spliced ​​along the channel dimension, and nonlinear mapping is achieved through the SiLU activation function. This dense connection structure not only maintains the geometric detail fidelity of the input data, but also realizes the continuous expression of deep and shallow semantics through the feature reuse mechanism, thereby significantly improving the expression integrity of the output features. The specific process of FAM performing four-level convolution and splicing along the channel direction on the bottom image, middle image and high-level image respectively is as follows:

[0049]

[0050] Where Conv is the conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the padding size, and the activation function is SiLU; The features are spliced ​​along the channel direction; the subscript x represents b, m and u; the generated multi-scale feature group is Among them Fb is the low-level feature map, F m is the middle-level feature map, F u is a high-level feature map.

[0051] 3) Enhance the high-level feature map to obtain a high-level enhanced feature map. This step uses the Self-attention Feature Enhancement Module (SFEM) to generate the high-level enhanced feature map. Compared with traditional convolution operations, the self-attention mechanism can capture long-range dependencies, thereby better modeling global context information. By introducing multi-scale attention weight calculation, SFEM can capture semantic associations in feature maps at different scales, further enhancing the expressive power of features. The process of SFEM enhancing high-level feature maps is as follows:

[0052] First, use convolution operations to input high-level feature maps Embed and get feature group In order to facilitate the subsequent matrix multiplication operation, F q and F v Reshape into Then F k After transposing the length and width dimensions, it is reshaped into The above process is expressed as follows through vertical expressions:

[0053]

[0054] Among them, Reshape is feature reshaping, Tr is feature transposition; Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the padding size, and there is no activation function.

[0055] Then, F′ q and After matrix multiplication, feature normalization is performed in the channel direction to obtain Once again, matrix multiplication is used to transform F m and F′ v After multiplication, the enhanced high-level feature map is reshaped In order not to lose the information of the input feature map, the feature map obtained after the convolution operation With the input feature map F u Perform skip connections to obtain enhanced feature maps The above process is expressed as follows through mathematical expressions:

[0056]

[0057] Among them, Reshape is feature reshaping, It is a matrix multiplication operation; Conv is a conventional convolution operation, and the activation function is SiLU. Although the feature map F′ enhanced by self-attention u With the input high-level feature map F u have the same size and number of channels, but F′ u It is rich in more light and dark area relationship information and global semantic information, and has stronger feature expression ability.

[0058] 4) The multi-scale feature group composed of the low-level feature map, the middle-level feature map and the high-level enhanced feature map is fused to obtain a multi-scale information feature map. This embodiment uses a multi-scale feature fusion module (MFFM) to generate a multi-scale information feature map. Specifically, to address the semantic information and spatial mismatch problem caused by direct fusion in MFFM, a reflective convolution is used to transform the multi-scale feature group. Perform cross-level spatial registration to generate feature sets that align both spatial scale and number of channels The number of feature fusion channels C is set f =128, the above processing process is expressed as follows:

[0059]

[0060] Among them, RefConv is the reflected convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the reflected padding size, and there is no activation function.

[0061] In order to improve computational efficiency, MFFM introduces a tensor decoupling strategy based on multi-head parallel processing: the input feature group {F b ,F′ m ,F″ u The channel dimension is decomposed into N sub-feature maps (N is 8), and the sequence dimension T is introduced to construct a five-dimensional tensor. Under the modular design framework, each subspace corresponds to an independent feature fusion module, and a lightweight attention mechanism is used to achieve interaction and adaptive cross-modal integration of subspace features. The mathematical expression for the above feature decomposition process is as follows:

[0062]

[0063] In a single feature fusion submodule, the three sub-feature maps are first and Perform element-wise multiplication between the two to obtain three cross-attention weights and By calculating the attention weight, the model can automatically learn the correlation between different feature maps and adjust the attention allocation to different feature maps according to the size of the attention weight. The mathematical expression of element multiplication is:

[0064]

[0065] where ⊙ is the element-wise multiplication operation.

[0066] Subsequently, all cross-attention weights are concatenated along the sequence dimension T, and the dimensionality reduction sum function is used to achieve global information fusion. Feature normalization is then performed along the channel dimension to dynamically generate the double-layer attention weights of the feature fusion submodule. The mathematical expression of attention weight is:

[0067]

[0068] Among them, Sum{·} is the dimension reduction summation function along the sequence dimension T, It is splicing along the sequence dimension T.

[0069] In order to integrate multi-scale feature information, the multi-scale attention weights are summed by weighted summation. Compared with the three sub-feature maps F before increasing the dimension T b n , F u n Fusion is performed on the grouping dimension of the submodules to obtain a feature map rich in multi-scale information Finally, F is adjusted by convolution operation mi The number of channels is , and the multi-scale feature map is obtained The above process is described by mathematical expressions as follows:

[0070]

[0071] Among them, Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the padding size, and there is no activation function; F ms It is a multi-scale information feature map.

[0072] 5) Comprehensively process the low-level feature map, the middle-level feature map and the multi-scale information feature map to obtain the enhanced low-light image. This step uses the Image Optimization Module (IOM) to perform image optimization on the feature group {F b ,F m ,F ms} perform comprehensive processing and finally output the enhanced low-light image During the image optimization process, low-level and mid-level features are mainly used to extract local information, while multi-scale features and mid-level features are used to extract global information of the image, thereby achieving comprehensive optimization of low-light images. The process of comprehensive processing of low-level feature maps, mid-level feature maps, and multi-scale information feature maps is as follows:

[0073]

[0074] Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the reflection padding size, and the activation function is Tanh; Up is the bilinear interpolation upsampling, sf is the upsampling scale factor; I local is the local image information, I global is the global image information; ms is the multi-scale image information, In m is the middle-level image information, In b is the underlying image information; ⊙ is the element-wise multiplication operation, α b is the underlying local coefficient, α m is the mid-layer local coefficient, β m is the middle-level global coefficient, β ms is a multi-scale global coefficient. In this embodiment, α b , α m , β m , β ms The values ​​are 1, 0.5, 0.5, and 1 respectively; I e It is the enhanced low-light image.

[0075] In particular, in this embodiment, the activation function adopts a differentiated configuration strategy: the Tanh function with saturation characteristics is used to constrain the pixel value range in the IOM, while the gradient-smoothed SiLU function is used in the feature extraction stage; this strategy enables the network to not only enhance the contextual association of local textures through feature aggregation, but also optimize the statistical distribution characteristics of global features with the help of activation function combination.

[0076] The performance of the low-light image enhancement method based on multi-scale feature fusion (named Multi-Enhancer) described in this embodiment is verified by experiments below.

[0077] Two experiments were designed: a low-light image enhancement experiment and a low-light object detection experiment. The former aims to evaluate Multi-Enhancer's effectiveness in improving the visual quality of low-light images, while the latter focuses on evaluating its ability to enhance the representation of low-light image features. These two experiments provide a comprehensive assessment of Multi-Enhancer's overall performance.

[0078] (1) Experimental environment configuration

[0079] All experiments were conducted on a local server. The experimental environment was built based on the following hardware and software configuration: the hardware system was equipped with an NVIDIA RTX A5000 GPU (24GB of video memory) and an Intel Core i7-12700KF multi-core processor (3.6GHz), supplemented by 128GB of DDR4 memory and 2TB of NVMe SSD storage. The software environment ran the Ubuntu 18.04 LTS operating system, integrated with the CUDA 11.3 acceleration library and the PyTorch 1.10 deep learning framework, using Python version 3.8.10. Detailed server configuration information is shown in Table 1.

[0080] Table 1 Server hardware and software configuration information

[0081]

[0082] (2) Parameter settings

[0083] ① Low-light image enhancement experiment

[0084] The low-light image enhancement experiments were conducted on three public low-light image enhancement datasets: LOL-v1, LOL-v2-real, and LOL-v2-synthetic. Eight advanced low-light image enhancement models, including PairLIE, IAT, and Retinexformer, were selected to compare their performance with Multi-Enhancer. In the training of the LOL-v1 dataset, the image size was set to 640×400, the total training epochs were 200, and the image batch size was 8. The Adam optimizer was used for parameter optimization, with an initial learning rate of 2×10 -4 , the weight decay coefficient is 1×10 -4 , and introduced a cosine annealing learning rate scheduling strategy to dynamically adjust the learning rate and balance the model convergence and overfitting risks. In terms of data enhancement, the training samples are geometrically transformed by random horizontal and vertical flipping to enhance data diversity and improve the generalization ability of the model. Peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are selected as image quality evaluation indicators to quantitatively evaluate the enhancement effect. In terms of loss function design, all enhancement models adopt a hybrid loss function composed of a smooth L1 loss function and a perceptual loss function (based on the VGG-16 network) to simultaneously optimize pixel-level features and high-scale semantic features, thereby improving the image enhancement effect. The mathematical expression of the loss function is:

[0085]

[0086] Among them, L s (I e,I) is the smooth L1 loss function, L p (I e ,I) is the perceptual loss function; represents the lth layer of the VGG network, ||·||2 is the L2 norm calculation, and h l w l , and c l are the height, width and number of channels of the feature map of the first layer of the VGG network; γ s is the L1 coefficient (value is 1), γ p is the perception coefficient (the value is 0.08).

[0087] For the training of LOL-v2-real and LOL-v2-synthetic datasets, the weight decay coefficient is increased to 5×10 while maintaining the same configuration as LOL-v1. -4 To enhance the model regularization effect.

[0088] ②Weak-light target detection experiment

[0089] The low-light target detection experiment selected the ExDark dataset to conduct target detection experiments in low-light environments, and selected the classic single-stage detection model RetinaNet and the two-stage detection model Sparse R-CNN as baseline target detection models (based on the MMDetection framework). In order to systematically evaluate the impact of the enhancement algorithm on the target detection performance, a standardized comparison benchmark was established: First, all low-light images in the ExDark dataset (including training sets and test sets) were enhanced using Zero DCE, IAT, MBLLEN and Multi-Enhancer respectively. Among them, each enhancement model uses the weight parameters pre-trained on the LOL-v1 and LOL-v2-real datasets to ensure the consistency of learning strategies between models. Subsequently, the enhanced image is passed into the target detection model for the second stage of training and performance verification. In addition, mAP and mAP are selected 50 As an evaluation indicator of target detection performance, it is used to quantitatively evaluate the detection effect.

[0090] In terms of model training configuration, RetinaNet uses the schedule_1x configuration file (using the SGD optimizer for 12 cycles, with an initial learning rate of 1×10-2, a momentum of 0.9, and a weight decay coefficient of 1×10 -4, introducing a step-down learning rate adjustment strategy to balance the convergence of the model and the risk of overfitting. In addition, the long side of the image is set to 608, the short side range is set from 320 to 512, and the image batch size is set to 8. In terms of data augmentation, the training samples are geometrically transformed by random horizontal and vertical flipping to improve the generalization ability of the model. For Sparse R-CNN, while maintaining the above basic settings, the optimizer is replaced with AdamW (initial learning rate 2.5×10 -5 ).

[0091] (3) Experimental results

[0092] ① Low-light image enhancement experiment

[0093] Table 2 shows the performance comparison of Multi-Enhancer with eight advanced low-light enhancement methods on the LOL-v1, LOL-v2-real, and LOL-v2-synthetic datasets. On the LOL-v1 dataset, its PSNR reached 21.73dB and SSIM was 0.81, which were 5.06dB and 0.22 higher than the SCI method, respectively. It also led the second-best method, Retinexformer, by 0.23dB and 0.01 in PSNR and SSIM, respectively. In the LOL-v2-real real scene, Multi-Enhancer surpassed the SCI method by 7.52dB and 0.31 with PSNR (22.24dB) and SSIM (0.80), and maintained a stable advantage of 1.26dB and 0.02 over Retinexformer. For LOL-v2-synthetic synthetic data, its PSNR (21.61dB) and SSIM (0.85) both reach suboptimal levels, only slightly lower than the optimal method by 0.40dB and 0.02.

[0094] Table 2 Comparison of low-light image enhancement performance of various enhancement methods on the LOL dataset

[0095]

[0096] Comparative experimental data demonstrates that Multi-Enhancer demonstrates significant advantages in both PSNR and SSIM metrics. Through its hierarchical feature aggregation enhancement mechanism and multi-scale feature fusion strategy, it achieves efficient cross-level information exchange, demonstrating strong cross-domain generalization adaptability and robustness across diverse datasets. Furthermore, by constructing a high-fidelity nonlinear mapping from low-light image space to standard illumination image space, Multi-Enhancer significantly improves visual reconstruction accuracy and topology preservation in low-light conditions.

[0097] ②Weak-light target detection experiment

[0098] Table 3 shows the experimental results of low-light image target detection on the ExDark dataset using various low-light image enhancement methods. In the LOL-v2-real dataset pre-training scenario, the low-light image enhanced by Multi-Enhancer is detected by RetinaNet.

[73] mAP on the model 50 The mAP and mAP reach 72.4% and 43.2% respectively, which are 1.4% and 0.7% higher than the baseline model, and significantly surpass MBLLEN.

[33] The model achieved 70.6% (+1.8%) and 42.4% (+0.8%). In addition, on the LOL-v1 dataset with Sparse R-CNN

[74] In the combination of frameworks, Multi-Enhancer still shows excellent performance, with its mAP 50 The mAP and mAP are 76.4% and 47.9% respectively, which are 1.3% and 1.1% higher than MBLLEN's 75.1% and 46.8% respectively.

[0099] Table 3 Comparison of target detection performance of various enhancement methods on the ExDark dataset

[0100]

[0101] Experimental results show that Multi-Enhancer can enhance the feature expression of low-light images to a certain extent under different data sets and detection frameworks, and the enhancement effect is more significant than the other three methods.

[0102] At the same time, the experimental results reveal the limitations of the simple cascade of enhancement-detection models (image enhancement first, then object detection). Although Multi-Enhancer achieves positive gains (mAP) in both RetinaNet and Sparse R-CNN frameworks, 50 1.3% and 0.1% respectively), but the performance improvement is still limited by the model architecture. More notably, low-light image enhancement methods such as Zero DCE and IAT have a negative performance offset (mAP 50 This phenomenon further reveals that the simple cascade of low-light image enhancement algorithms and object detection frameworks has problems such as incompatibility of feature spaces and mismatch of task requirements.

[0103] In summary, the low-light image enhancement algorithm based on multi-scale feature fusion in this embodiment effectively improves the visual quality of low-light images through the organic combination of hierarchical feature decoupling strategy, feature aggregation enhancement strategy and multi-scale feature fusion mechanism. The comparative test results show that: on the LOL-v1 low-light image enhancement dataset, the PSNR and SSIM achieved by the Multi-Enhancer algorithm are improved by 5.06dB and 0.22 compared with the SCI algorithm, and the visual quality of the image is significantly improved; at the same time, the low-light image enhanced by the Multi-Enhancer has a mAP of 0.06dB in the RetinaNet detection framework. 50 Compared with the original low-light image, it is improved by 1.4%, and the image's feature expression ability is enhanced to a certain extent.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A low-light image enhancement method based on multi-scale feature fusion, characterized by: It includes the following steps: 1) Decompose the low-light image into bottom-layer image, middle-layer image and high-layer image; 2) Perform four-level convolution and channel-wise concatenation on the bottom-level image, middle-level image, and high-level image, respectively, to generate a multi-scale feature group consisting of low-level feature maps, middle-level feature maps, and high-level feature maps; 3) Enhance the high-level feature map to obtain a high-level enhanced feature map; 4) Fusing the multi-scale feature group consisting of the low-level feature map, the middle-level feature map, and the high-level enhanced feature map to obtain a multi-scale information feature map; 5) Comprehensively process the low-level feature map, mid-level feature map and multi-scale information feature map to obtain an enhanced low-light image.

2. The low-light image enhancement method based on multi-scale feature fusion according to claim 1, characterized in that: In step 1), the process of decomposing the low-light image is as follows: Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, p is the reflection filling size; I is the low-light image, I b is the underlying image, I m is the middle layer image, I u For high-level images.

3. The low-light image enhancement method based on multi-scale feature fusion according to claim 2, characterized in that: In step 2), the process of performing four-level convolution and channel-wise splicing on the bottom layer image, middle layer image, and high layer image is as follows: Where Conv is the conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size. The features are spliced ​​along the channel direction; the subscript x represents b, m and u; the generated multi-scale feature group is {F b ,F m ,F u }, where F b is the low-level feature map, F m is the middle-level feature map, F u is a high-level feature map.

4. The low-light image enhancement method based on multi-scale feature fusion according to claim 3, characterized in that: The process of enhancing the high-level feature map in step 3) is as follows: Among them, Reshape is feature reshaping, Tr is feature transposition; Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size; Among them, Reshape is feature reshaping, is a matrix multiplication operation; Conv is a regular convolution operation.

5. The low-light image enhancement method based on multi-scale feature fusion according to claim 4, characterized in that: The process of fusing the multi-scale feature groups in step 4) is as follows: Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the reflection padding size; Among them, ⊙ is the element-wise multiplication operation; Among them, Sum{·} is the dimension reduction summation function along the sequence dimension T, It is splicing along the sequence dimension T; Among them, Conv is a conventional convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the padding size; F ms It is a multi-scale information feature map.

6. The low-light image enhancement method based on multi-scale feature fusion according to claim 5, characterized in that: In step 5), the process of comprehensively processing the low-level feature map, the middle-level feature map, and the multi-scale information feature map is as follows: Among them, RefConv is the reflection convolution operation, k is the convolution kernel size, s is the convolution kernel step size, and p is the reflection padding size; Up is the bilinear interpolation upsampling, sf is the upsampling scale factor; I local is the local image information, I global is the global image information; ms is the multi-scale image information, In m is the middle-level image information, In b is the underlying image information; ⊙ is the element-wise multiplication operation, α b is the underlying local coefficient, α m is the mid-layer local coefficient, β m is the middle-level global coefficient, β ms is the multi-scale global coefficient; I e It is the enhanced low-light image.