Low-light environment image enhancement and adaptive feature fusion target detection method

Through the improved YOLO-MFL network, combined with multi-scale attention-guided lighting estimation and damage recovery, the accuracy problem of small object detection in low-light environments is solved, and high-precision multi-scale object detection is achieved.

CN120355900AInactive Publication Date: 2025-07-22SOUTHWEST PETROLEUM UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510481070.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Small object detection of aerial images in low-light environments faces reduced contrast between target and background, blurred edges, loss of color information and complex background interference. The existing Retinexformer method cannot effectively capture the problems of heterogeneity of light distribution and insufficient feature fusion. The deep learning ASFF method has insufficient feature resolution during small object detection, resulting in a decrease in detection accuracy.

Method used

Using the improved YOLO-MFL network, combined with the MA-Retinexformer low-light enhancement layer and the FL-ASFF small object detection layer, we enhance image quality through multi-scale attention-guided illumination estimation and corruption recovery, and add feature detection levels based on the ASFF algorithm to achieve adaptive fusion of multi-scale objects.

Benefits of technology

It improves the detection accuracy of small targets under low light conditions, broadens the coverage range of target scales, enhances the model's ability to identify multi-scale targets, and improves detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355900A_ABST
    Figure CN120355900A_ABST
Patent Text Reader

Abstract

A low-light environment image enhancement and adaptive feature fusion target detection method relates to the technical field of image detection, and comprises the following steps: original image information is collected and input into an improved YOLO-MFL network, the improved YOLO-MFL network is composed of an MA-Retinexformer low-light enhancement layer, a trunk of YOLOv11, an FL-ASFF small target detection layer and a detection head of YOLOv11, and the detection head of the MA-Retinexformer low-light enhancement layer and the trunk of the YOLOv11 are connected with the FL-ASFF small target detection layer; carrying out dark light enhancement on the original image by utilizing an MA-Retinexformer low-illumination enhancement layer in the improved YOLO-MFL network, and continuously processing the image subjected to dark light enhancement through the improved YOLO-MFL network to obtain a detection result; according to the invention, illumination estimation and damage recovery are carried out on the low-illumination image through an MA-Retinexformer framework, so that the quality of the low-illumination image is improved, and the small target detection precision of the model when the processing illumination condition is poor is effectively improved. And meanwhile, an additional feature detection level is added on the basis of a conventional ASFF algorithm, so that the coverage range of the target scale is widened, the recognition capability of the model on the multi-scale target is enhanced, and the overall detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly relates to a method for low-light environment image enhancement and adaptive feature fusion object detection. Background Art

[0002] Small target detection in aerial images under low-light environments and complex backgrounds is a research direction with many challenges. The low-light environment causes the contrast between the target and the background to decrease, the target edges to be blurred, the color information to be lost, and the image noise to increase, seriously affecting the detection accuracy. The complex background makes small target detection face challenges such as small target size and background information interference.

[0003] Currently, in the latest Retinexformer method based on Retinex in traditional image enhancement technologies, although it can improve the brightness and contrast of low-light images to a certain extent, its single-scale convolution kernel is difficult to adapt to the spatial heterogeneity of light distribution, and it cannot effectively capture the detail changes under different lighting conditions. At the same time, in the feature fusion process, the contribution degree differences of different-scale features are not fully considered, resulting in inefficient utilization of feature information. When processing dark area details and light mutation regions, key information is easily missed, affecting the enhancement effect. And the conventional adaptively spatial feature fusion (ASFF) method based on deep learning can achieve a high accuracy rate under normal lighting conditions, but this algorithm will be limited by the feature resolution when processing small targets. Insufficient feature resolution makes it difficult to effectively capture the detail information of small targets, resulting in a decrease in detection accuracy. Summary of the Invention

[0004] In view of this, the present invention proposes a method for low-light environment image enhancement and adaptive feature fusion object detection, designs an improved YOLO-MFL (MFL, that is, the combination of MA-Retinexformer and FL-ASFF) network, and based on the multi-scale attention-guided Retinexformer (MA-Retinexformer) low-light enhancement layer included in the network, combines the multi-scale attention-guided illumination estimation and damage recovery in MA-Retinexformer to process the image, so as to improve the quality of low-light images, and further combines the four-layer adaptively spatial feature fusion (FL-ASFF) including the small object detection layer in the network to adaptively fuse features at different levels, enhances the algorithm's recognition ability for multi-scale objects, improves the accuracy of the model in small object detection in images, and enables it to effectively meet the detection requirements of current application fields with high requirements for the quality of small object detection, such as high-precision surveying and mapping, artificial intelligence image recognition, etc.

[0005] To solve the above at least one technical problem, the technical solution provided by the present invention is a method for low-light environment image enhancement and adaptive feature fusion object detection, including the following steps:

[0006] Step S1: Collect the original image information and input it into the improved YOLO-MFL network. The improved YOLO-MFL network is composed of a MA-Retinexformer low-light enhancement layer, the backbone of YOLOv11, a FL-ASFF small object detection layer, and the detection head of YOLOv11;

[0007] Step S2: Use the MA-Retinexformer low-light enhancement layer in the improved YOLO-MFL network to perform low-light enhancement on the original image;

[0008] Step S3: Continue to process the low-light enhanced image through the improved YOLO-MFL network to obtain the detection result.

[0009] The technical effects achieved by the present invention are:

[0010] 1. The present invention estimates the illumination and recovers the damage of low-light images through the MA-Retinexformer framework to improve the quality of low-light images, and effectively improves the small object detection accuracy of the model when dealing with poor illumination conditions.

[0011] 2. Based on the conventional ASFF algorithm, the present invention adds an additional feature detection level, thereby broadening the range of target scale coverage, effectively handling the scale changes of objects, ensuring that only relevant information is retained at each spatial position, enhancing the model's recognition ability for multi-scale targets, and improving the overall detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 Schematic diagram of the overall process of the present invention;

[0014] Figure 2 Schematic diagram of the structure of the improved YOLO-MFL network in the present invention;

[0015] Figure 3 Schematic diagram of the structure of the multi-scale attention-guided Retinexformer (MA-Retinexformer) in the present invention;

[0016] Figure 4 Schematic diagram of the structure of the multi-scale attention-guided illuminance estimator in the present invention;

[0017] Figure 5 Schematic diagram of the MGS structure in the present invention;

[0018] Figure 6 Schematic diagram of the structure of the damage repairer in the present invention;

[0019] Figure 7 Schematic diagram of the structure of the IGAB in the present invention;

[0020] Figure 8 Schematic diagram of the structure of the FL-ASFF detection layer in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention will be further described in detail below in conjunction with the embodiments and the drawings.

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention.

[0023] See Figure 1 、 Figure 2 , a method for low-light environment image enhancement and adaptive feature fusion object detection, comprising the following steps:

[0024] Step S1: Collect the original image information and input it into the improved YOLO-MFL network. The improved YOLO-MFL network is composed of a MA-Retinexformer low-light enhancement layer, the backbone of YOLOv11, an FL-ASFF small object detection layer, and the detection head of YOLOv11.

[0025] Step S2: Use the MA-Retinexformer low-light enhancement layer in the improved YOLO-MFL network to perform low-light enhancement on the original image.

[0026] Among them, as Figure 3 shown, MA-Retinexformer adopts a single-stage Retinex-based framework, which is composed of a multi-scale attention-guided illumination estimator and a damage repairer, and is located before the backbone of YOLOv11.

[0027] On this basis, the steps of low-light enhancement are as follows:

[0028] Step 1: Calculate the mean value of each pixel of the original image along the channel dimension to obtain an illumination prior map.

[0029] Step 2: Input the original image and the illumination prior map into the multi-scale attention-guided illumination estimator, and extract features of different scales by parallel branch convolutions.

[0030] The structure of the multi-scale attention-guided illumination estimator is shown in Figure 4, which includes a multi-scale parallel feature extraction module, a gate-controlled adaptive fusion module, a spatial attention enhancement module, and a triple convolutional network composed of two 1x1 convolutions and a 9x9 depthwise separable convolution. Among them, the combination of the multi-scale parallel feature extraction module, the gate-controlled adaptive fusion module, and the spatial attention enhancement module is Figure 5 the MGS structure shown in

[0031] The scale features are extracted through three parallel convolutional branches (3x3, 5x5, 7x7) in the multi-scale parallel feature extraction module, and then output to the next module together. The multi-scale illumination feature F is captured through the convolutional operation shown in formula (5) ms :

[0032] F ms = Concat(Conv(I in , 3), Conv(I in , 5), Conv(I in , 7)) (5)

[0033] In formula (5), F ms represents the multi-scale illumination feature; Conv(I in , k0), k0 = {3, 5, 7} represents the convolution operation on the input image I in .

[0034] Step 3: Adaptive fusion of features at different scales is performed through the gate-controlled adaptive fusion module to obtain the feature map after adaptive feature fusion.

[0035] Immediately afterwards, the features obtained above are input into the gate-controlled adaptive fusion module. This module has two parallel branches. One branch F compressed is responsible for channel conversion and GELU activation:

[0036] F compressed = GELU(W 1×1 × F ms + b 1×1 ) (6)

[0037] In formula (6), W 1×1 represents the weight matrix of the 1x1 convolution kernel, and b 1×1 represents the bias term of the 1x1 convolution.

[0038] Another branch is responsible for dynamically generating the fusion weights W of features at each scaleg :

[0039] W g = σ(W conv2 × GELU(W conv1 × F compressed + b conv1 ) + b conv2 ) (7)

[0040] In Equation (7), σ represents the final activation function; W conv1 and W conv2 represent the weight matrices of the first and second 1×1 convolutional kernels respectively, and b conv1 and b conv2 represent the bias terms of the first and second 1×1 convolutions respectively.

[0041] Then, through the feature weighting operation, the feature map F after adaptive feature fusion is obtained, as shown in Equation (1): fus as shown in Equation (1):

[0042] F fus = W g ⊙ F compressed (1)

[0043] In Equation (1), F fus represents the feature map after adaptive feature fusion; F compressed represents the branch of channel conversion and GELU activation in the gated adaptive feature fusion module; W g represents the branch in the gated adaptive feature fusion module responsible for dynamically generating the fusion weights of features at various scales.

[0044] Step 4: Introduce the feature map after adaptive feature fusion into the spatial attention enhancement module for spatial attention enhancement, and input the enhanced feature map into the triple convolutional network behind the MGS, and introduce the light-up map to light up the image.

[0045] Input the fused feature map F fus into the spatial attention enhancement module, generate the energy map E through the learnable energy function E(p), and then perform the Sigmoid normalization operation to map the energy map E into the probability distribution A spa (p):

[0046] E = Conv2(GELU(Conv1(F fus )) (8)

[0047]

[0048] In Equations (8) and (9), Conv1(F fus ) represents the first convolutional layer, which processes the input feature map Ffus Perform convolution operation; GELU (Conv1 (F fus ) represents the activation function layer, which applies the Gaussian error linear unit (GELU) activation function to the output of the first convolutional layer; Conv2(GELU(Conv1(F fus ))) represents the second convolutional layer, which performs convolution operation on the feature map processed by the GELU activation function; E(p) represents the energy function.

[0049] After passing the probability distribution A spa (p) Strengthen the characteristic response of key areas and improve the physical rationality of the illumination map:

[0050] F out =A spa (p)vF fus (10)

[0051] Next, the enhanced feature map F out Input such as Figure 4 After the triple convolutional network shown in , a light-up map is generated and light-up feature F lu , by introducing perturbations to account for real-world image corruption, such as noise, artifacts and color distortion A light-up map is introduced to light up the image. The specific operation is shown in formula (11):

[0052]

[0053] Use I lu represents the image after being lit, as shown in formula (2):

[0054]

[0055] Among them, I lu Indicates the image after being lit; Represents the image lit by the light-up map; R represents the reflectance map (according to the Retinex theory, the input low-light map is decomposed into the reflectance map (reflectance image) R and the illumination map (illumination map) L); C represents the overall damage.

[0056] Step 5: Input the illuminated image into the damage repairer, and use the illumination-guided Transformer (IGT) architecture of the damage repairer to repair image damage and enhance image quality.

[0057] The task of the damage repairer is to correct the problems caused by overexposure, noise amplification, and color distortion during the lighting process. Its structure is shown in Figure 6 , and the main operations performed are shown in Equation (11):

[0058]

[0059] In Equation (11), ε represents the light estimation, represents the damage recovery, and L p is the illumination prior map, and F lu represents the light-up feature, and I en represents the enhanced image.

[0060] See Figure 6 , Figure 7 , the IGT architecture contains an illumination-guided attention block (IGAB). The IGT architecture is supplemented by a key module in the IGAB - the illumination-guided multi-head self-attention (IG-MSA) mechanism. The IGT can capture the long-term dependencies between regions in the image that exhibit different lighting conditions. The IG-MSA mechanism enables regions with better lighting to provide valuable context information for darker regions, thus helping to more accurately recover underexposed regions. The Transformer architecture effectively models the non-local interactions in the image, overcoming the limitation of convolutional neural networks (CNNs) in dealing with long-range dependencies. The IGT optimizes the lit-up image by enhancing global features while retaining local details, thereby enhancing the overall image quality of the output.

[0061] Let the input feature F in ∈ R H×W×c , X ∈ R H×W×c , X is divided into k parts: X = (X1, X2,..., X k ), d k = c / k, i ∈ (1...k), c represents the total number of features. The fully connected layer projects X i onto the query element the key element and the value element among them. The projection process can be expressed as Equation (12):

[0062]

[0063] where represents the learnable parameters of the fully connected layer. Similarly, let the light-up feature Flu Represents the regional illumination information, which is consistent with the X shape and is expressed as: Y ∈ R H×W×c , and is also divided into k parts: Y = (Y1, Y2,..., Y k )

[0064] Then the final self-attention expression of IG-MSA is shown in Equation (3):

[0065]

[0066] In Equation (3), Q i represents the query element; K i represents the key element; V i represents the value element; Y i represents the regional illumination information; α i represents the learning parameter for adaptive scaling matrix multiplication; T represents matrix transpose.

[0067] Step S3: The image after low-light enhancement is further processed by the improved YOLO-MFL network to obtain the detection result.

[0068] The FL-ASFF small object detection layer is located in the neck of YOLOv11. The small object detection layer of FL-ASFF has a four-layer structure, specifically as Figure 8 shown. The enhanced image output by Step S2 is subjected to feature extraction by the backbone part of YOLOv11 and then input into the small object detection layer in Step S3 for detection. After the detection is completed, it continues to enter the detection head of YOLOv11 to perform subsequent detection processing.

[0069] The small object detection principle of FL-ASFF is as follows:

[0070] Use x l to represent the resolution feature of the l-th (l ∈ {1, 2, 3, 4}) level. For the l-th layer, set the feature x of another layer to the n-th (n ∈ {1, 2, 3, 4} and n ≠ l) level n scaled and adjusted to have the same shape as x l . Since the features of each level of YOLOv11 have different resolutions and numbers of channels (channels), the upsampling uses a 1×1 convolutional layer to compress the number of channels of the feature to 1 dimension, and then the resolution is increased by interpolation, while the downsampling uses a 3×3 convolutional layer with a stride of 2 to modify both the number of channels and the resolution at the same time.

[0071] Let represent the feature vector at the position (i, j) of the size of the l-th level scaled from the feature of the n-th level. Then the expression of the fused feature can be expressed as shown in Equation (4):

[0072]

[0073] In formula (4), represents the feature vector scaled from the first-level feature to the position (i, j) at the l-level size, represents the feature vector scaled from the second-level feature to the position (i, j) at the l-level size, represents the feature vector scaled from the third-level feature to the position (i, j) at the l-level size, represents the feature vector scaled from the fourth-level feature to the position (i, j) at the l-level size, where l = {1, 2, 3, 4}; represents the fused feature.

[0074] respectively represent the respective feature importance weights of the four levels at the position (i, j), and these weights are adaptively learned by the network. It should be noted that is normalized to the interval [0, 1], and their sum is equal to 1, that is:

[0075] To calculate these weights, the softmax function can be used in combination with specific parameters and calculated according to formula (13) as follows:

[0076]

[0077] where is the parameter obtained by calculating x 1→l , x 2→l , x 3→l , x 4→l respectively using a 1×1 convolutional layer.

[0078] The gradient is adjusted by dynamically adjusting to avoid the false alarm problem and improve the training efficiency. According to the chain rule, the gradient calculation formula is as shown in formula (14):

[0079]

[0080] In formula (14), is the loss function, which is used to measure the difference between the model prediction result and the actual one.

[0081] Example:

[0082] The YOLO-MFL method in the present invention is compared with the YOLO v11n method in terms of image detection effect. On the VisDrone dataset, mAP50 and mAP50-95 are increased by 12.6% and 8.2% respectively, and on the UAVDT dataset, mAP50 and mAP50-95 are increased by 8.1% and 6.2% respectively. Compared with the YOLO v11m method, the YOLO-MFL method reduces the number of parameters by 40.2%, and on the VisDrone dataset, mAP50 is increased by 1.5%, and on the UAVDT dataset, mAP50 is increased by 5.6%.

[0083] From the training results, it can be seen that the detection accuracy of the YOLO-MFL algorithm in the present invention is higher than that of the other algorithms. The FL-ASFF small target detection layer of the YOLO-MFL algorithm can effectively improve the small target detection ability of the algorithm from the perspectives of actual detection effect and accuracy indicators, and the MA-Retinexformer low-light enhancement layer of the YOLO-MFL algorithm can effectively improve the detection accuracy of the algorithm in low-light environments.

[0084] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for low-light environment image enhancement and adaptive feature fusion object detection, characterized in that It includes the following steps: Step S1: Collect the original image information and input it into the improved YOLO-MFL network. The improved YOLO-MFL network consists of a MA-Retinexformer low-light enhancement layer, the backbone of YOLOv11, an FL-ASFF small object detection layer, and the detection head of YOLOv11; Step S2: Use the MA-Retinexformer low-light enhancement layer in the improved YOLO-MFL network to enhance the low-light of the original image; Step S3: Continue to process the low-light enhanced image through the improved YOLO-MFL network to obtain the detection result.

2. The low-light environment image enhancement and adaptive feature fusion object detection method according to claim 1, wherein: The MA-Retinexformer low-light enhancement layer is located before the backbone of YOLOv11, and the FL-ASFF small object detection layer is located in the neck of YOLOv11. The FL-ASFF small object detection layer is a four-layer structure.

3. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 2, characterized in that: The MA-Retinexformer low-light enhancement layer adopts a single-stage Retinex-based framework, which consists of a multi-scale attention-guided illumination estimator and a damage repairer.

4. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 3, characterized in that: The steps of the low-light enhancement are as follows: Step 1: Calculate the mean value of each pixel of the original image along the channel dimension to obtain the illumination prior map; Step 2: Input the original image and the illumination prior map into the multi-scale attention-guided illumination estimator to extract features of different scales; Among them, the multi-scale attention-guided illumination estimator includes a multi-scale parallel feature extraction module, a gated adaptive feature fusion module, a spatial attention enhancement module, and a triple convolution network composed of two 1×1 convolutions and a 9×9 depthwise separable convolution; The extraction of features of different scales is implemented by the parallel branch convolution in the multi-scale parallel feature extraction module; Step 3: Adaptively fuse the features of different scales through the gated adaptive feature fusion module to obtain the feature map after adaptive feature fusion; Step 4: Enhance the spatial attention of the feature map after adaptive feature fusion through the spatial attention enhancement module, and input the enhanced feature map into the triple convolution network to introduce a lighting map to light up the image; Step 5: Input the lit image into the damage repairer, and use the IGT framework of the damage repairer to repair the image damage and enhance the image quality. Among them, the IGT architecture is supplemented by an illumination-guided multi-head self-attention mechanism.

5. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 4, characterized in that: The parallel branch convolution in Step 2 consists of three convolution branches with convolution kernels of 3×3, 5×5, and 7×7 respectively.

6. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 4, characterized in that: The adaptive feature fusion operation in Step 3 is executed by a gated adaptive feature fusion module including two parallel branches. The fusion process is specifically shown in Equation (1): F fus = W g ⊙F compressed (1) In formula (1), F fus represents the feature map after adaptive feature fusion; F compressed represents the branch of channel conversion and GELU activation in the gated adaptive feature fusion module; W g represents the branch in the gated adaptive feature fusion module responsible for dynamically generating the fusion weights of features at each scale.

7. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 4, characterized in that: The expression of the lit image in Step 5 is shown in Equation (2): In formula (2), I lu represents the image after being lit; represents the image lit by the light-up map; R represents the reflection map; C represents the overall damage.

8. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 4, characterized in that: The self-attention expression of the illumination-guided multi-head self-attention mechanism in Step 5 is: In formula (3), Q i represents the query element; K i represents the key element; V i represents the value element; Y i represents the regional illumination information; α i represents the learning parameter for adaptive scaling matrix multiplication; T represents matrix transpose.

9. A low-light environment image enhancement and adaptive feature fusion object detection method according to claim 2, characterized in that: The steps for the low-light enhanced image to continue to be processed by the improved YOLO-MFL network are: The image after low-light enhancement is input into the FL-ASFF small target detection layer for detection after feature extraction by the backbone part of YOLOv11, and then continues to enter the detection head of YOLOv11 for subsequent detection processing.

10. A method for low-light environment image enhancement and adaptive feature fusion object detection according to claim 9, characterized in that: The expression of the features fused in the FL-ASFF small target detection layer is shown in Equation (4): In formula (4), represents the feature vector scaled from the first-level feature to the feature at the position (i, j) of the l-level size, represents the feature vector scaled from the second-level feature to the feature at the position (i, j) of the l-level size, represents the feature vector scaled from the third-level feature to the feature at the position (i, j) of the l-level size, represents the feature vector scaled from the fourth-level feature to the feature at the position (i, j) of the l-level size, where l = {1, 2, 3, 4}; represents the fused feature; respectively represent the feature importance weights of the four levels at the position (i, j).

Citation Information

Cited By

  • Underground material accumulation early warning method and device based on low illumination enhancement and medium

    CN121302180A

  • Low-illumination small target safe wearing detection method and system

    CN121392752A