High-precision target detection method and system suitable for low-light condition
By adopting multiple filtering processing and feature fusion technology under low light conditions, combining dual adaptive filtering modules and dark feature pyramid networks, the problem of low target detection accuracy under low light is solved, and high-precision object detection is achieved.
Patent Information
- Application Number
- CN202411840084.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-11
AI Technical Summary
Under low light conditions, the performance of the object detection model is significantly reduced, feature extraction is difficult and noise affects the detection accuracy, resulting in a low target detection accuracy.
Multiple filtering processing and feature fusion technology are used to maximize pooling and average pooling in the channel and spatial dimensions through the dual adaptive filtering module. Combined with the dark feature pyramid network, the details and semantic information of the feature map are enhanced, and the jump connection is used to balance the feature map weights, and the feature extraction structure is optimized.
The feature map quality under low light conditions is improved, effective noise removal, and enhanced the accuracy and robustness of target detection, and improved the detection performance in low light environments.
Smart Images

Figure CN120298831A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of object detection, and particularly relates to a method and system for high-precision object detection applicable to low-light conditions. Background Art
[0002] Under challenging low-light conditions, the performance of object detection models significantly degrades, which affects their applications in fields such as autonomous driving, SLAM (Simultaneous Localization and Mapping), and surveillance. Low-light images are affected by low signal-to-noise ratio and poor contrast, which makes feature extraction difficult and unreliable. In addition, the mixing of object feature information and background noise affects detection. Noise may also obscure or distort the key features of the object, making it more difficult for the detector to accurately identify the object from the background. In recent years, detection algorithms based on deep learning have made advanced progress. In contrast, object detection in low light has not received the same attention as object detection in normal light.
[0003] Object detection is one of the core research directions in the field of computer vision, which aims to identify and locate one or more target objects from images or videos. This task is crucial for many applications such as autonomous driving, video surveillance, and image analysis. With the development of deep learning, especially the rise of convolutional neural networks (CNNs), the field of object detection has undergone revolutionary changes. The RCNN series (including R-CNN, Fast R-CNN, and Faster R-CNN) are early deep learning methods that use CNNs to extract features and then apply selective search or Region Proposal Network (RPN) to determine the regions in the image that may contain objects.
[0004] Due to the low light intensity, the clarity of objects in images and feature maps under low-light conditions is insufficient. This low-light environment not only reduces the overall quality of the image but also poses significant challenges to subsequent image processing and analysis tasks. Under low-light conditions, the noise level in the image usually increases, and at the same time, the detail information is weakened, resulting in blurred object edges and color distortion. These factors together affect the recognizability of the image, making it particularly difficult to obtain high-quality images in a low-light environment. Feature maps, as an important intermediate representation for deep learning models to process images, directly affect the performance of the model. However, the feature maps generated under low-light conditions often lack sufficient detail information and texture features, limiting the model's recognition and classification capabilities in complex scenarios. Summary of the Invention
[0005] The objective of the present invention is to provide a method and system for high-precision object detection applicable to low-light conditions to solve the problem of low object detection accuracy under low light.
[0006] The present invention adopts the following technical solutions: A method for high-precision target detection applicable to low-light conditions, including:
[0007] Step 1: Perform multiple filtering processes on the extracted features of the low-light image to obtain multiple filtering features from low to high;
[0008] Step 2: Element-wise add the extracted features of the low-level filtering features and the upsampled extracted features of the high-level filtering features to obtain multiple fusion features from high to low;
[0009] Step 3: Element-wise add the extracted features of each fusion feature, the filtering features of the corresponding layer, and the globally fused features strengthened by max pooling to obtain high-level features, and send each high-level feature to each detector for detection;
[0010] The globally fused feature is the feature obtained by element-wise adding the corresponding upsampled fused feature and the corresponding filtering feature;
[0011] The filtering process is specifically: using max pooling in the channel dimension and average pooling in the spatial dimension, multiplying the extracted feature after max pooling and the extracted feature of the low layer element-wise to obtain a pooled feature, and multiplying the extracted feature after average pooling and the extracted feature of the low layer element-wise to obtain an average feature, and then element-wise adding the pooled feature and the average feature to obtain a filtering feature.
[0012] Further, the relationship between the size of the convolutional kernel when extracting features after max pooling in Step 1 and the extracted feature map of the low-light image is:
[0013]
[0014] In the formula, k2 is the size of the convolutional kernel when extracting features after max pooling; c is the number of channels of the extracted feature map of the low-light image; b2 is a hyperparameter.
[0015] Further, the relationship between the size of the convolutional kernel when extracting features after average pooling in Step 1 and the extracted feature map of the low-light image includes:
[0016] k1 = t - log2(160 / h) - b1
[0017] In the formula, k1 is the size of the convolutional kernel when extracting features after average pooling; t is the size of the largest convolutional kernel used by the filter; h is the size of the extracted feature map of the low-light image; b1 is a hyperparameter.
[0018] A system for high-precision target detection applicable to low-light conditions, including:
[0019] The first dual adaptive filtering module, whose input is the extracted feature map of the low-light image after normalization processing, and whose output is the feature with rich details of the low-light image, is used to perform max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the extracted feature after max pooling and the extracted feature of the low-light image element by element to obtain the first pooled feature, and multiply the extracted feature after average pooling and the extracted feature of the low-light image element by element to obtain the first average feature. Then add the first pooled feature and the first average feature element by element to obtain the first filtered feature with rich details of the low-light image.
[0020] Furthermore, it further includes:
[0021] The second dual adaptive filtering module, whose input is the first filtered feature, and whose output is the feature with poor details of the low-light image, is used to perform max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the extracted feature after max pooling and the extracted feature of the first filtered feature element by element to obtain the second pooled feature, and multiply the extracted feature after average pooling and the extracted feature of the first filtered feature element by element to obtain the second average feature. Then add the second pooled feature and the second average feature element by element to obtain the second filtered feature with poor details of the low-light image.
[0022] Furthermore, it further includes:
[0023] The third dual adaptive filtering module, whose input is the second filtered feature, and whose output is the feature with poor semantics of the low-light image, is used to perform max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the extracted feature after max pooling and the extracted feature of the second filtered feature element by element to obtain the third pooled feature, and multiply the extracted feature after average pooling and the extracted feature of the second filtered feature element by element to obtain the third average feature. Then add the third pooled feature and the third average feature element by element to obtain the third filtered feature with poor semantics of the low-light image.
[0024] Furthermore, it further includes:
[0025] The fourth dual adaptive filtering module, whose input is the third filtered feature, and whose output is the feature with rich semantics of the low-light image, is used to perform max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the extracted feature after max pooling and the extracted feature of the third filtered feature element by element to obtain the fourth pooled feature, and multiply the extracted feature after average pooling and the extracted feature of the third filtered feature element by element to obtain the fourth average feature. Then add the fourth pooled feature and the fourth average feature element by element to obtain the fourth filtered feature with rich semantics of the low-light image.
[0026] The beneficial effects of the present invention are:
[0027] The present invention can improve the quality of the extracted feature map of low-light images. Two measures are taken to solve the problem of degradation of the extracted feature map, so as to denoise and enhance the extracted feature map. First, noise under low-light conditions is removed in the spatial dimension, that is, according to the spatial size of the extracted feature map, the core size is dynamically adjusted, which can denoise more accurately while maintaining important detail information. Second, in the channel dimension, the feature maps with rich information are highlighted through max pooling, and the channel interaction range is dynamically adjusted according to the interaction of the number of channels. Finally, the dark feature pyramid network is specifically improved to better compensate for the details of the feature map and highlight the target features. Description of the Drawings
[0028] Figure 1 It is a schematic structural diagram of the dual adaptive filtering module in the present invention;
[0029] Figure 2 It is the feature map sorted by signal-to-noise ratio (SNR) according to the present invention;
[0030] Figure 3 It is the image information according to the present invention, wherein (a) is a low-light image; (b) is the extracted feature map of the low-light image; (c) is a normal-light image; (d) is the extracted feature map of the normal-light image;
[0031] Figure 4 It is the detection result and heat map of YOLOX and the present invention;
[0032] Figure 5 It is a schematic structural diagram of the present invention. Detailed Embodiment
[0033] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0034] The present invention discloses a method for high-precision target detection applicable to low-light conditions, including:
[0035] Step 1: Perform multiple filtering processes on the extracted features of the low-light image to obtain multiple filtering features from low to high.
[0036] The global fusion feature is the feature obtained by upsampling the corresponding fusion feature and adding it element by element to the corresponding filtering feature; the filtering process is specifically: using max pooling in the channel dimension and average pooling in the spatial dimension, multiplying the extracted feature after max pooling and the extracted feature of the lower layer element by element to obtain the pooling feature, and multiplying the extracted feature after average pooling and the extracted feature of the lower layer element by element to obtain the average feature, and then adding the pooling feature and the average feature element by element to obtain the filtering feature.
[0037] Average pooling is used to filter the spatial information of the feature map, and max pooling is used to compress the channel information. Subsequently, it is processed through a convolutional layer with a variable kernel size, where the kernel size is adaptively adjusted according to the feature map and its channels. After the activation function processes each descriptor, a spatial attention map M s (x) and a channel attention map M c (x) are generated and they are multiplied element-wise by the original feature map to achieve the filtering effect.
[0038] In low-light image processing, noise is a major challenge, where dark noise and thermal noise may reduce the quality of the low-light feature map. The low signal-to-noise ratio in the feature map often results in incomplete target structures. Different from denoising on the original image, the present invention performs multi-stage denoising on the feature map to eliminate complex forms of dark noise. By averaging local regions to smooth image details, more context information is retained while suppressing noise to handle incomplete target structures.
[0039] To enhance the filter's understanding of the context, large convolutional kernels are used to capture long-range spatial features after filtering. The convolutional kernel size k is kept in a certain proportion to the size h of the feature map, effectively expanding the receptive field. To avoid the overfitting problem caused by an overly large convolutional kernel, the maximum size of the convolutional kernel is limited to 11. The correspondence between the size of the kernel and the feature map is as follows:
[0040] k1 = t - log2(160 / h) - b1
[0041] In the formula, k1 is the convolutional kernel size when extracting features after average pooling; t is the size of the largest convolutional kernel used by the filter, generally 10 or 11; h is the size of the extracted feature map of the low-light image; b1 is a hyperparameter.
[0042] Figure 2 The left side shows the feature maps sorted according to the signal-to-noise ratio (SNR), and the SNR is marked in red font in the upper left corner of each feature map; Figure 2 The right side shows the histograms after applying two different pooling methods to the feature maps.
[0043] Feature maps generated under low-light conditions usually contain complex and scattered high-intensity information or noise in the channel dimension. Feature maps with higher intensities usually contain more valuable information, such as Figure 2 as shown in the first row, these are called high-intensity feature maps. On the contrary, some feature maps are severely noisy, such as Figure 2As shown in the second row of . To emphasize these high-intensity feature maps, the feature maps in the channel dimension are compressed. The results of average pooling compression are concentrated between 2 and 5, indicating a low information entropy. However, after max pooling, the channel information is more evenly distributed, and high-intensity feature maps with a higher information entropy are given greater weights. Therefore, the present invention uses max pooling to assign greater weights to high-intensity regions. In addition, the present invention adjusts the convolutional kernel size according to the channels of the feature maps to enhance the interaction between channels. Specifically, the present invention generates a global feature channel weight map through global max pooling and performs local channel interaction through convolutional kernels of different sizes. The relationship between the convolutional kernel size k and the number of channels c is:
[0044]
[0045] In the formula, k2 is the convolutional kernel size when extracting features after max pooling; c is the number of channels of the extracted feature maps of the low-light image; b2 is a hyperparameter.
[0046] As Figure 3 shown, due to the low light intensity, the clarity of objects in images under low-light conditions is insufficient. Therefore, it is very difficult for the human eye to distinguish Figure 3 (a) the color information of the motorcycle. During the forward pass of the network, the detailed information on the feature maps is often not rich enough. As Figure 3 (b) shown, the blurred contour makes it difficult to distinguish how many motorcycles are in the picture, let alone the people hidden in the darkness. In contrast, Figure 3 (d) shows the normal-light feature maps after two downsampling processes, where the contour of the bicycle is clearly visible. It shows in the neural network that the detailed information in the low-light feature maps is insufficient, and the traditional neck network PANet is difficult to achieve the same effect as the normal-light feature maps when processing these feature maps.
[0047] As the network depth increases, the semantic information becomes more abundant, but the detailed information decreases, especially in low-light images. There are mainly two problems in the feature map during the feature fusion process, namely, the target information of the feature map is not rich enough and the information is lost during the transmission process. First, in order to enhance the information representation of the fusion, the present invention introduces lower-level feature maps and fuses four layers of feature maps to provide more comprehensive target information. Second, the present invention draws on the idea of skip connections to make up for the detailed information lost during the top-down transmission process. This method can effectively retain and restore important image details. Finally, the present invention optimizes the overall structure of the feature extraction module.
[0048] Step 2: Element-wise add the filtered features of the lower layer and the extracted features obtained by upsampling the filtered features of the higher layer to obtain multiple fusion features from high to low;
[0049] Step 3: Element-wise add the extracted features of each fusion feature, the filtering features of the corresponding layer, and the globally fused features strengthened by max pooling to obtain high-level features, and send each high-level feature to each detector for detection.
[0050] Steps 2 and 3 are specifically as Figure 5 shown, a dark feature pyramid network based on PANet is adopted, and three improvements are made:
[0051] 1. Introduce low-level feature maps
[0052] Only using feature maps from the top layer usually cannot retain important details such as color, texture, and contour. If only three layers of feature maps (L3, L4, L5) are used for fusion, it is difficult to effectively retain the key detail information of the low-light feature maps, such as color, texture, contour, etc. To solve this problem, the present invention introduces the feature layer of stage 2. As Figure 5 shown, an L2 feature layer of 160×160 is added to the bottom of the neck network, and the feature layer at the bottom is upsampled and fused with the 160×160 feature layer. During the process of passing from the bottom to the top, this compensation fully makes up for the loss of detail information caused by the gradually deepening feature extraction process.
[0053] 2. Skip connections
[0054] The PANet of YOLOX simply concatenates and fuses features at different levels by implementing a multi-scale feature fusion strategy. Although this method can capture the feature information of targets of different sizes, it may lead to unequal weight distribution of features of different sizes during the fusion process. In particular, the large-size feature maps have too much weight in the network, which has a significant adverse impact on the detection of low-level feature maps in low-light environments. Therefore, the present invention balances the information of low-level feature maps with smaller weights during the layer-by-layer feature fusion process by introducing skip connections from the input layer to the output layer, thereby enhancing the model's feature fusion ability.
[0055] 3. Structure optimization
[0056] The CSP module enhances the feature expression ability by fusing feature maps at multiple stages of the network, but this fusion method often increases the dimension of the generated feature maps. After feature fusion, the initial dimensionality reduction convolution may exacerbate the loss of information in the low-light feature maps. For the four-stage feature map fusion involved in the present invention, this dimensionality reduction further loses the detail information of the feature maps, and the additional dimensionality reduction operation is more likely to cause the loss of valuable information. Therefore, the present invention chooses to remove such dimensionality reduction convolutions.
[0057] In addition, to better highlight important features and reduce the interference of background noise on detection performance, the present invention adopts max pooling for downsampling, replacing the traditional convolutional downsampling method. This max pooling strategy helps to retain the most significant features in the feature map while suppressing background noise, further enhancing the detection ability of the model in low-light environments.
[0058] The present invention conducts qualitative experiments, comparing the method of the present invention with the advanced YOLOX detector to verify its beneficial effects. As Figure 4 shown, the superscript with * represents using COCO weights, and the superscript without * represents not using COCO weights; regardless of whether COCO weights are used, the model of the present invention is superior to YOLOX in detecting objects hidden in the dark. For Figure 4 (a) the scenario of overlapping objects, the model of the present invention performs better in distinguishing individuals and bicycles. Figure 4 (a) The heatmap reveals that YOLOX often ignores people in the background, while the present invention can effectively capture small objects by integrating detailed information. In addition, in Figure 4 (b) and Figure 4 (d) under low-light conditions, the noise filtering ability of the model of the present invention enhances the visibility of hidden targets (such as ships and cats), and can effectively detect even in the case of severe noise. The model of the present invention focuses on high-intensity regions, improving the detection of bright spots, such as Figure 4 (a) and Figure 4 the indoor scene and the people in the bus in (c). Therefore, it can be seen that the present invention significantly improves the robustness and accuracy of the model of the present invention in low-light environments.
[0059] Example 1
[0060] This example adopts the Stochastic Gradient Descent (SGD) algorithm and loads COCO pre-trained weights to train the model. The training parameters include a batch size of 8, a weight decay of 0.0004, a momentum of 0.9, and the initial learning rate is set to 0.01. All images are adjusted to a resolution of 640×640. To be consistent with the data augmentation strategy of YOLOX, this example implements conventional data augmentation techniques including random horizontal flipping, random scaling, random cropping, and random HSV augmentation. In addition, this example also follows other training strategies of YOLOX, such as Multi positives and SimOTA, etc.
[0061] First, the pre-trained weights with the minimum loss value are loaded, and the network parameters are fixed to ensure the stability of the model and the consistency of predictions. The coordinate and category information of the target are generated according to the output information of the detection head. These information will be plotted on the y in turn to visually display the localization and classification results of the target.
[0062] To verify the effectiveness of each module used in this embodiment, ablation experiments are conducted on the ExDark dataset. This dataset contains 7,363 images of 12 categories. In this embodiment, the dataset is divided into a training and validation set and a test set in a ratio of 8:2. It is worth noting that in the ablation experiment, this embodiment uses fine-tuning with pre-trained weights on ExDark. This approach prevents the prior knowledge learned during the day from affecting the model testing.
[0063] The entropy is selected as the test metric. Mean Average Precision (mAP for short) is usually used to evaluate the comprehensive performance of a detector, while recall measures the ability of the model to retrieve true positives among all true positives. Specifically, mean average precision @0.5 calculates the mAP at different intersection over union (IoU) thresholds from 0.50 to 0.95. To evaluate the complexity and speed of the model, this embodiment also calculates the network parameters (Params) in millions (M) and the number of floating-point operations (FLOPs) in billions (G).
[0064] The dual adaptive filtering module includes two parameters b1 and b2. k1 and k2 are the convolution kernel sizes when extracting features after average pooling and max pooling respectively. Specifically, filtering is performed when the feature map is downsampled to 160. This embodiment evaluates the effectiveness of adaptively adjusting the core size based on the size and number of the feature map. During training, this embodiment gradually increases b1 and b2 starting from 0 and 1. In Table 1, the values 11, 9, 9, and 7 correspond to using core sizes 160, 80, 40, and 20 respectively. As shown in Table 2, the values 3, 3, 5, and 5 correspond to feature maps with 128, 256, 512, and 1024 channels. First, for the hyperparameter b1. As shown in Table 1, a larger convolution kernel increases the perception ability and improves the detection accuracy by capturing more context information. Compared with manually setting the convolution kernel size, adaptive core adjustment provides greater flexibility. When b1 is set to 1, the performance is the best.
[0065] Table 1 Settings of Hyperparameter b1
[0066] b1 Nucleus size Average precision Average precision @ 0.5 Recall Parameter (M) Floating point number (G) None 3 0.648 0.378 0.468 54.16 155.72 0 11,9,9,7 0.645 0.373 0.467 54.16 155.75 1 9,9,7,7 0.653 0.379 0.468 54.16 155.75 2 9,7,7,5 0.650 0.381 0.468 54.16 155.74 3 7,7,5,5 0.650 0.375 0.466 54.16 155.74
[0067] As the size of the convolutional kernel gradually increases, the interaction between channels becomes more effective. In Table 2, a saturation phenomenon occurs as the convolutional kernel size increases. This is because there are significant differences in the number of channels between different feature maps, and the interaction ranges required for high-dimensional and low-dimensional feature maps are inconsistent. In contrast, the low-pass filter performs best when the hyperparameter b2 is set to 1, that is, when the dimension increases from 128 to 1024, the interaction range correspondingly expands from 3 to 7. The dual adaptive filtering module can more flexibly and effectively process feature maps of different sizes, thus improving the overall performance.
[0068] Table 2 Hyperparameter b2 Settings
[0069]
[0070]
[0071] The ablation experiments in Table 3 verify the effectiveness of the dual adaptive filtering module. In this embodiment, denoising and enhancement are separately performed on the basis of the baseline model YOLOX. Table 3 shows that whether it is denoising or enhancement, using them alone can improve the performance compared with the baseline model. This result indicates that this embodiment effectively removes thermal noise and dark noise in low-light images in the spatial dimension. In the channel dimension, it can more effectively utilize the light source information between channels. When the dual adaptive filtering module is used, not only the high-frequency noise is removed, but also the focusing on the light source is enhanced, and the impact on the model parameters is minimal.
[0072] Table 3 Ablation Experiments
[0073] Denoising Enhancement Average precision Average precision @ 0.5 Recall Parameter (M) Floating point number (G) 0.639 0.370 0.463 54.156 155.719 √ 0.653 0.379 0.468 54.159 155.746 √ 0.659 0.383 0.470 54.157 155.736 √ √ 0.661 0.389 0.480 54.159 155.746
[0074] Starting from the YOLOX baseline, the present invention gradually adds three components of the dark feature pyramid. Table 4 shows that the low-level feature maps significantly improve the accuracy, emphasizing their importance in low-light detection. The skip connections better balance the weights between the low-level and high-level feature maps. The structural adjustment reduces the parameters and computational complexity, while the max pooling enhances the feature emphasis during the downsampling process. Integrating these elements improves the balance of the model between details and significant features.
[0075] Table 4 Ablation Analysis of the Dark Feature Pyramid
[0076]
[0077] Without introducing pre-trained weights, as shown in Table 5, adding only the dual adaptive filtering module increased the average precision by 3.2% and the average precision @0.5 by 1.9%, demonstrating its effectiveness in filtering. Due to the significant compensation for details, the dark feature pyramid increased the average precision by 4.8%. When both the dark feature pyramid and the dual adaptive filtering module were integrated, the average precision reached 69.1%, enabling high-precision detection even without prior knowledge of normal lighting conditions.
[0078] Table 5 Effectiveness of the combination of the dual adaptive filtering module and the dark feature pyramid
[0079]
[0080] Table 6 compares the method of this embodiment with advanced low-light detection algorithms in the prior art. Table 6 is a comparison with low-light detectors on the ExDark dataset. The bold results indicate the best results, and the horizontal lines indicate the second-best results. The comparative experiments on the ExDark dataset show that image enhancement methods usually reduce the detection performance due to artifacts, blurring, or color distortion, which may have an adverse impact on downstream tasks. Therefore, this embodiment gives priority to maintaining image details. Although other method modules deal with simpler noise, complex dark noise and thermal noise are more difficult to remove. This embodiment adds a dual adaptive filtering module after each convolution, effectively optimizing the feature map. After using pre-trained weights, the accuracy of this embodiment exceeds 70% for each category, achieving the highest accuracy among 10 categories, with an overall detection accuracy of 83.2% and a frame rate of 42.9 FPS.
[0081] Table 6 Comparison with low-light detectors on the ExDark dataset
[0082]
[0083]
[0084] The specific detectors for each symbol in Table 6 are as follows:
[0085] [1]Zhang Y,Zhang J,Guo X.Kindling the darkness:A practical low-lightimage enhancer[C].In:Proceedings of the 27th ACM International Conference onMultimedia.2019:1632-1640.
[0086] [2]Lv F, Lu F, Wu J, et al. MBLLEN: Low-light image / video enhancement using CNNs[C]. In: British Machine Vision Conference (BMVC). 2018, 220(1):4.
[0087] [3]Cui Z, Li K, Gu L, et al. You only need 90k parameters to adapt light: A lightweight transformer for image enhancement and exposure correction[J]. arXiv preprint arXiv:2205.14871, 2022.
[0088] [4]Liu W, Ren G, Yu R, et al. Image-adaptive YOLO for object detection in adverse weather conditions[C]. In: Proceedings of the AAAI Conference on Artificial Intelligence. 2022, 36(2):1792-1800.
[0089] [5]Cui Z, Qi G J, Gu L, et al. Multitask AET with orthogonal tangent regularity for dark object detection[C]. In: Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV). 2021:2553-2562.
[0090] [6]Jiang Z, Shi D, Zhang S. FRSE-Net: Low-illumination object detection network based on feature representation refinement and semantic-aware enhancement[J]. The Visual Computer, 2024, 40(5):3233-3247.
[0091] [7] Sasagawa Y, Nagahara H. YOLO in the dark - domain adaptation method for merging multiple models[C]. In: Proceedings of the European Conference on Computer Vision (ECCV). Glasgow, UK, August 23 - 28, 2020, Part XXI: 16. Springer International Publishing, 2020.
[0092] The present invention also discloses a system for high - precision object detection applicable to low - light conditions, including:
[0093] A first dual - adaptive filtering module, whose input is the extracted feature map of the low - light image after normalization processing, and whose output is the feature with rich details of the low - light image. It is used to perform max - pooling in the channel dimension and average - pooling in the spatial dimension, multiply the extracted feature after max - pooling and the extracted feature of the low - light image element - by - element to obtain a first pooled feature, multiply the extracted feature after average - pooling and the extracted feature of the low - light image element - by - element to obtain a first average feature, and then add the first pooled feature and the first average feature element - by - element to obtain the first filtered feature with rich details of the low - light image.
[0094] The system of the present invention further includes:
[0095] A second dual - adaptive filtering module, whose input is the first filtered feature, and whose output is the feature with poor details of the low - light image. It is used to perform max - pooling in the channel dimension and average - pooling in the spatial dimension, multiply the extracted feature after max - pooling and the extracted feature of the first filtered feature element - by - element to obtain a second pooled feature, multiply the extracted feature after average - pooling and the extracted feature of the first filtered feature element - by - element to obtain a second average feature, and then add the second pooled feature and the second average feature element - by - element to obtain the second filtered feature with poor details of the low - light image.
[0096] The system of the present invention further includes:
[0097] The third dual adaptive filtering module, with the second filtering feature as its input and the feature of the low-light image with poor semantics as its output, is used to process by using max pooling in the channel dimension and average pooling in the spatial dimension, multiply the extracted feature after max pooling and the extracted feature of the second filtering feature element by element to obtain the third pooling feature, multiply the extracted feature after average pooling and the extracted feature of the second filtering feature element by element to obtain the third average feature, and then add the third pooling feature and the third average feature element by element to obtain the third filtering feature of the low-light image with poor semantics.
[0098] The system of the present invention further includes:
[0099] The fourth dual adaptive filtering module, with the third filtering feature as its input and the feature of the low-light image with rich semantics as its output, is used to process by using max pooling in the channel dimension and average pooling in the spatial dimension, multiply the extracted feature after max pooling and the extracted feature of the third filtering feature element by element to obtain the fourth pooling feature, multiply the extracted feature after average pooling and the extracted feature of the third filtering feature element by element to obtain the fourth average feature, and then add the fourth pooling feature and the fourth average feature element by element to obtain the fourth filtering feature of the low-light image with rich semantics.
[0100] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for high-precision target detection applicable to low-light conditions, characterized in that Including: Step 1: Perform multiple filtering processes on the extracted features of the low-light image to obtain multiple filtering features from low to high. Step 2: Element-wise add the upsampled extracted features of the low-level filtering features and the high-level filtering features to obtain multiple fusion features from high to low. Step 3: Element-wise add the extracted features of each fusion feature, the filtering features of the corresponding layer, and the globally fused features strengthened by max pooling to obtain high-level features, and send each high-level feature to each detector for detection. The globally fused feature is the feature obtained by upsampling the corresponding fusion feature and then element-wise adding it to the corresponding filtering feature. The filtering process is specifically as follows: Utilize max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the max-pooled extracted features and the extracted features of the low level element-wise to obtain pooled features, and multiply the average-pooled extracted features and the extracted features of the low level element-wise to obtain average features. Then, element-wise add the pooled features and the average features to obtain filtering features.
2. The method for high-precision target detection applicable to low-light conditions according to claim 1, wherein The relationship between the size of the convolutional kernel when extracting features after max pooling in Step 1 and the extracted feature map of the low-light image is: In the formula, k2 is the size of the convolutional kernel when extracting features after max pooling; c is the number of channels of the extracted feature map of the low-light image; b2 is a hyperparameter.
3. A method for high-precision target detection applicable to low-light conditions according to claim 1, characterized in that, The relationship between the size of the convolutional kernel when extracting features after average pooling in Step 1 and the extracted feature map of the low-light image includes: k1 = t - log2(160 / h) - b1 In the formula, k1 is the size of the convolutional kernel when extracting features after average pooling; t is the size of the largest convolutional kernel used by the filter; h is the size of the extracted feature map of the low-light image; b1 is a hyperparameter.
4. A system for high-precision target detection applicable to low-light conditions, characterized in that, Including: The first dual adaptive filtering module, whose input is the extracted feature map of the low-light image after normalization processing, and whose output is the feature with rich details of the low-light image. It is used to utilize max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the max-pooled extracted features and the extracted features of the low-light image element-wise to obtain the first pooled feature, and multiply the average-pooled extracted features and the extracted features of the low-light image element-wise to obtain the first average feature. Then, element-wise add the first pooled feature and the first average feature to obtain the first filtering feature with rich details of the low-light image.
5. The system for high-precision target detection applicable to low-light conditions according to claim 4, wherein, Also including: The second dual adaptive filtering module, whose input is the first filtering feature, and whose output is the feature with poor details of the low-light image. It is used to utilize max pooling in the channel dimension and average pooling in the spatial dimension. Multiply the max-pooled extracted features and the extracted features of the first filtering feature element-wise to obtain the second pooled feature, and multiply the average-pooled extracted features and the extracted features of the first filtering feature element-wise to obtain the second average feature. Then, element-wise add the second pooled feature and the second average feature to obtain the second filtering feature with poor details of the low-light image.
6. The system for high-precision target detection applicable to low-light conditions according to claim 5, characterized in that, Also including: The third dual adaptive filtering module, with its input being the second filtering feature and its output being the feature of the low-light image with poor semantics, is used to perform processing by using max pooling in the channel dimension and average pooling in the spatial dimension. The extracted feature after max pooling and the extracted feature of the second filtering feature are multiplied element-wise to obtain the third pooling feature, and the extracted feature after average pooling and the extracted feature of the second filtering feature are multiplied element-wise to obtain the third average feature. Then, the third pooling feature and the third average feature are added element-wise to obtain the third filtering feature of the low-light image with poor semantics.
7. The system for high-precision target detection applicable to low-light conditions according to claim 6, wherein It further includes: The fourth dual adaptive filtering module, with its input being the third filtering feature and its output being the feature of the low-light image with rich semantics, is used to perform processing by using max pooling in the channel dimension and average pooling in the spatial dimension. The extracted feature after max pooling and the extracted feature of the third filtering feature are multiplied element-wise to obtain the fourth pooling feature, and the extracted feature after average pooling and the extracted feature of the third filtering feature are multiplied element-wise to obtain the fourth average feature. Then, the fourth pooling feature and the fourth average feature are added element-wise to obtain the fourth filtering feature of the low-light image with rich semantics.
Citation Information
Cited By
Low-illumination target detection method and device based on multi-scale dynamic fusion
CN122156587A