Target detection system and method based on low-illumination image

Through the object detection system based on low-illumination images, multiple feature enhancement and feature fusion technology are used to solve the problem of poor object detection effect in low-illumination environments, and efficient detection of long-distance or weak targets is achieved.

CN120339640APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256523.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing deep learning object detection methods are difficult to effectively extract the detailed characteristics of the target in low-illumination environments, especially the detection effect of long-distance or weak targets is poor, and problems such as insufficient brightness, uneven light and blurred details have not been effectively solved.

Method used

The object detection system based on low-illumination images is adopted, including the first feature extraction module, an enhanced low-light feature extraction module, a coding module and a decoding module. Through the adjustment of spatial size and number of channels, multiple feature enhancement, encoding and fusion are enhanced, combined with convolutional gating, multi-scale convolutional attention and frequency domain self-attention mechanisms, the feature extraction capability is improved.

Benefits of technology

Effectively enhance the characteristic display of long-distance or weak targets, improves detection accuracy and stability in low-illumination environments, and improves the accuracy and robustness of target detection in conditions of uneven light and blurred details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339640A_ABST
    Figure CN120339640A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a target detection system and method based on a low-illumination image, and the method comprises the steps: carrying out the spatial size adjustment and channel number adjustment of a to-be-detected visible light image, and obtaining a first feature map; performing multiple times of feature enhancement on the first feature map to obtain a plurality of enhanced feature maps; encoding and fusing the plurality of enhanced feature maps to obtain an output feature map; and predicting the output feature map to obtain a target category and a target frame position in the to-be-detected visible light image. The enhanced feature map obtained by performing multiple feature enhancement on the first feature map can effectively enhance the features of the long-distance or dim and small target, so that the features of the long-distance or dim and small target are clearly displayed in the enhanced feature map; therefore, the problems of insufficient brightness, non-uniform illumination, fuzzy details and the like in a low-illumination environment are effectively solved, a plurality of enhanced feature maps are coded and fused to obtain an output feature map, the output feature map is predicted, and the detail features of the target can be effectively extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image target detection, and in particular, to a target detection system and method based on low-light images.

Background Art

[0002] The target detection system analyzes the input image through a deep learning model to identify the targets therein. However, when identifying and positioning targets in an environment with insufficient light, existing methods face many challenges. For example: 1) Low brightness. In a low-light environment, the overall brightness of the image is insufficient, making it difficult to distinguish the targets from the background. Since the target features in low-light images are not obvious, existing deep learning models may have difficulty effectively extracting useful features from low-quality images, resulting in a decrease in detection accuracy; 2) Uneven illumination. This causes a large brightness difference in different regions of the image, thereby affecting the detection performance of the model; 3) Blurred details. Due to insufficient light, the edges and details of the targets are difficult to identify, and the feature extraction ability of the detection network is insufficient, limiting the comprehensiveness of the detection.

[0003] Currently popular deep learning target detection methods, such as SSD, YOLO, and Faster R-CNN, rely on convolutional operations and multi-level feature fusion to achieve detection. Although they perform well in specific scenarios, they still have obvious deficiencies in low-light environments. These methods are difficult to effectively extract the detailed features of targets in low-light environments, especially for the detection of distant or small targets, with poor detection effects.

Summary of the Invention

[0004] In view of this, the present invention provides a target detection system and method based on low-light images.

[0005] The specific technical solution of the first embodiment of the present invention is: A target detection system based on low-light images, the system includes: a first feature extraction module, an enhanced low-light feature extraction module, an encoding module, and a decoding module; the output end of the first feature extraction module is connected to the input end of the enhanced low-light feature extraction module, and the output end of the enhanced low-light feature extraction module is connected to the input end of the encoding module; the output end of the encoding module is connected to the input end of the decoding module; the first feature extraction module is used to adjust the spatial size and the number of channels of the visible light image to be detected to obtain a first feature map; the enhanced low-light feature extraction module is used to perform multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; the encoding module is used to encode and fuse the multiple enhanced feature maps to obtain an output feature map; the decoding module is used to predict the output feature map to obtain the target category and the target bounding box position in the visible light image to be detected.

[0006] Preferably, the enhanced low-light feature extraction module includes a first enhanced low-light feature extraction unit, a second enhanced low-light feature extraction unit, and a third enhanced low-light feature extraction unit connected in sequence; the first enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first feature map to obtain a first enhanced feature map; the second enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first enhanced feature map to obtain a second enhanced feature map; the third enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the second enhanced feature map to obtain a third enhanced feature map.

[0007] Preferably, the performing spatial feature extraction at different scales and feature information fusion at different scales on the first feature map to obtain a first enhanced feature map includes: performing convolution addition on the first feature map and a preset convolutional layer to obtain a first convolved feature map; performing spatial information extraction on the convolved feature map and fusing the spatial information into the convolved feature map to obtain a first fused feature map; performing a sparse convolution operation on the first fused feature map to obtain a second convolved feature map; performing channel expansion on the second convolved feature map to obtain an expanded-channel feature map; dividing the expanded-channel feature map in the channel dimension to obtain a first sub-feature map and a second sub-feature map; the first sub-feature map and the second sub-feature map are the images after dividing the expanded-channel feature map; performing depthwise separable convolution at different scales on the first sub-feature map to obtain the spatial features of the first sub-feature map at different scales; using the second sub-feature map as a gating signal to perform weighted multiplication on the spatial features at different scales to obtain a first weighted feature map; restoring the number of channels of the weighted feature map and performing a residual connection between the feature map with the restored number of channels and the first feature map to obtain a residual feature map; performing feature information fusion at different scales on the residual feature map to obtain the first enhanced feature map.

[0008] Preferably, the obtaining of the first enhanced feature map by performing feature information fusion of different scales on the residual feature map includes: performing global max pooling operation and average pooling operation on the residual feature map, and splicing the results after pooling to obtain a spliced feature map; performing spatial relationship learning and Sigmoid function activation on the spliced feature map at different scales to obtain a spatial attention feature map; multiplying the spatial attention feature map by the residual feature map and performing weighted adjustment to obtain a second weighted feature map; using a preset multi-branch convolution structure to perform multi-scale feature extraction and multi-scale convolution on the second weighted feature map to obtain feature information of the residual feature map at different scales; performing feature information fusion on the feature information of different scales through a channel shuffle operation to obtain a second fusion feature map; adding the second fusion feature map to the convolved feature map to obtain the first enhanced feature map.

[0009] Preferably, the obtaining of the output feature map by encoding and fusing the multiple enhanced feature maps includes:

[0010] Performing feature extraction based on space and frequency domain on the third enhanced feature map to obtain a frequency domain feature output map; encoding and performing cross-scale feature fusion based on convolution on the first enhanced feature map, the second enhanced feature map, and the frequency domain feature output map to obtain the output feature map.

[0011] Preferably, the encoding of the first enhanced feature map, the second enhanced feature map, and the frequency-domain feature output map and the convolution-based cross-scale feature fusion to obtain the output feature map include: adjusting the dimension of the frequency-domain feature output map to obtain a dimension-adjusted feature map; performing an upsampling operation on the dimension-adjusted feature map to obtain a first upsampled feature map; the size of the first upsampled feature map is the same as the size of the second enhanced feature map; performing content-guided attention fusion on the first upsampled feature map and the second enhanced feature map to obtain a third fused feature map; using a preset RepC3 module and a convolutional layer to perform feature extraction on the third fused feature map to obtain a second feature map; performing an upsampling operation on the second feature map to obtain a second upsampled feature map; the size of the second upsampled feature map is the same as the size of the first enhanced feature map; performing content-guided attention fusion on the second upsampled feature map and the first enhanced feature map to obtain a fourth fused feature map; using a preset RepC3 module to perform feature extraction on the fourth fused feature map to obtain a third feature map; performing a downsampling operation on the third feature map to obtain a first downsampled feature map; the size of the first downsampled feature map is the same as the size of the second feature map; performing content-guided attention fusion on the first downsampled feature map and the second feature map to obtain a fifth fused feature map; using a preset RepC3 module to perform feature extraction on the fifth fused feature map to obtain a fourth feature map; performing a downsampling operation on the fourth feature map to obtain a second downsampled feature map; the size of the second downsampled feature map is the same as the size of the dimension-adjusted feature map; performing content-guided attention fusion on the second downsampled feature map and the dimension-adjusted feature map to obtain a sixth fused feature map; using a preset RepC3 module to perform feature extraction on the sixth fused feature map to obtain a fifth feature map; performing a concatenation operation on the third feature map, the fourth feature map, and the fifth feature map to obtain the output feature map.

[0012] Preferably, the feature extraction based on space and frequency domain for the third enhanced feature map to obtain a frequency domain feature output map includes: adjusting the number of channels of the third enhanced feature map to obtain the third enhanced feature map with adjusted channels; segmenting the third enhanced feature map with adjusted channels in the channel dimension to obtain a third sub-feature map and a fourth sub-feature map; performing local feature extraction of different scales on the third sub-feature map and the fourth sub-feature map to obtain a first local feature map and a second local feature map; splicing and convolving and dimension reduction of the first local feature map and the second local feature map along the channel dimension to obtain a feature map after dimension reduction; performing feature extraction of different dimensions on the feature map, and performing feature extraction based on space and frequency domain on the features of different dimensions to obtain a frequency domain feature output map.

[0013] Preferably, the spatial size of the first enhanced feature map is 2 times that of the second enhanced feature map, and the spatial size of the second enhanced feature map is 2 times that of the third enhanced feature map.

[0014] Preferably, the number of channels of the first enhanced feature map is 0.5 times that of the second enhanced feature map, and the number of channels of the second enhanced feature map is 0.5 times that of the third enhanced feature map.

[0015] The specific technical solution of the second embodiment of the present invention is: a target detection method based on a low-illumination image, the method includes: adjusting the spatial size and the number of channels of the visible light image to be detected to obtain a first feature map; performing multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; encoding and fusing the multiple enhanced feature maps to obtain an output feature map; predicting the output feature map to obtain the target category and the target bounding box position in the visible light image to be detected.

[0016] Implementing the embodiments of the present invention will have the following beneficial effects:

[0017] The present invention obtains a first feature map by adjusting the spatial size and the number of channels of the visible light image to be detected; performs multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; encodes and fuses the multiple enhanced feature maps to obtain an output feature map; and predicts the output feature map to obtain the target category and the target bounding box position in the visible light image to be detected. The enhanced feature maps obtained by performing multiple feature enhancements on the first feature map can effectively enhance the features of distant or small targets, so that the features of distant or small targets can be clearly displayed in the enhanced feature maps, thereby effectively coping with problems such as insufficient brightness, uneven illumination, and blurred details in low-illumination environments. Encoding and fusing the multiple enhanced feature maps to obtain an output feature map and making predictions based on the output feature map can effectively extract the detailed features of the target, especially improving the detection effect of distant or small targets.

BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic structural diagram of an object detection system based on a low-illumination image;

[0020] Figure 2 is a network structure diagram of a backbone network;

[0021] Figure 3 is a network structure diagram of an enhanced low-light feature extraction module;

[0022] Figure 4 is a network structure diagram of a feature extraction module based on space and frequency domain;

[0023] Figure 5 is a network structure diagram of a content-guided attention fusion module;

[0024] Figure 6 is a flowchart of the steps of a method for detecting an object based on a low-illumination image;

[0025] Among them, 201, a first feature extraction module; 202, an enhanced low-light feature extraction module; 203, an encoding module; 204, a decoding module.

DETAILED DESCRIPTION

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0027] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products or devices.

[0028] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0029] Please refer to Figure 1 , which is a schematic structural diagram of a target detection system based on low-light images in the first embodiment of the present application, so as to improve the detection effect of distant or small targets. The system includes: a first feature extraction module 201, an enhanced low-light feature extraction module 202, an encoding module 203, and a decoding module 204; the output end of the first feature extraction module is connected to the input end of the enhanced low-light feature extraction module, and the output end of the enhanced low-light feature extraction module is connected to the input end of the encoding module; the output end of the encoding module is connected to the input end of the decoding module; the first feature extraction module is used to adjust the spatial size and the number of channels of the visible light image to be detected to obtain a first feature map S2; the enhanced low-light feature extraction module is used to perform multiple feature enhancements on the first feature map S2 to obtain multiple enhanced feature maps; the encoding module is used to encode and fuse the multiple enhanced feature maps to obtain an output feature map F; the decoding module is used to predict the output feature map F to obtain the target category and the target bounding box position in the visible light image to be detected.

[0030] Specifically, please refer to Figure 2, a visible light detection system is used to collect the visible light image to be detected, and the visible light to be detected is input into the backbone network. The backbone network is constructed based on ResNet-18 and sequentially includes 3 convolutional layers, 1 max pooling layer, and 4 basic module layers (BasicBlock). The first three convolutional layers and the max pooling layer perform feature extraction and downsampling on the image. The spatial size of the image is halved twice, and the number of channels increases from the initial 3 channels to 64 channels, obtaining the first feature map S2. The first feature map is enhanced multiple times, such as 3 times, to obtain 3 enhanced feature maps; the 3 enhanced feature maps are encoded and fused to obtain the output feature map; the output feature map is predicted to obtain the target category and the target bounding box position in the visible light image to be detected.

[0031] In this embodiment, the system obtains the first feature map by adjusting the spatial size and the number of channels of the visible light image to be detected; the first feature map is enhanced multiple times to obtain multiple enhanced feature maps; the multiple enhanced feature maps are encoded and fused to obtain the output feature map; the output feature map is predicted to obtain the target category and the target bounding box position in the visible light image to be detected. The enhanced feature maps obtained by enhancing the first feature map multiple times can effectively enhance the features of distant or small targets, so that the features of distant or small targets can be clearly displayed in the enhanced feature maps, thereby effectively coping with problems such as insufficient brightness, uneven illumination, and blurred details in low-illumination environments. Encoding and fusing multiple enhanced feature maps to obtain the output feature map and predicting based on the output feature map can effectively extract the detailed features of the target, especially improving the detection effect of distant or small targets.

[0032] In a specific embodiment, the low-light feature enhancement module includes a first low-light feature enhancement unit, a second low-light feature enhancement unit, and a third low-light feature enhancement unit connected in sequence; the first low-light feature enhancement unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first feature map S2 to obtain the first enhanced feature map S3; the second low-light feature enhancement unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first enhanced feature map S3 to obtain the second enhanced feature map S4; the third low-light feature enhancement unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the second enhanced feature map S4 to obtain the third enhanced feature map S5.

[0033] Specifically, the first feature map S2 is input into the first basic module layer. Each basic module layer contains two convolutional layers, and the input features are added to the convolutional output through a skip connection to obtain the feature map S2'. In order to improve the feature extraction ability of the network in low-light environments, the second convolutional layer of the subsequent three basic module layers is replaced with an enhanced low-light feature extraction module. The enhanced low-light feature extraction module can effectively cope with the challenges under low-light conditions, improve the expressiveness of the target area, and reduce the interference of background noise on the network. The convolutional gating module adjusts the feature weights through a dynamic gating mechanism, while the multi-scale convolutional attention module combines multi-scale convolution and attention mechanisms to adaptively process problems such as low illuminance, noise, and detail loss, thereby improving the detection accuracy and stability of weak targets.

[0034] In a specific embodiment, the performing spatial feature extraction of different scales and feature information fusion of different scales on the first feature map S2 to obtain a first enhanced feature map S3 includes: performing convolutional addition on the first feature map S2 and a preset convolutional layer to obtain a first convolution feature map S2'; performing spatial information extraction on the convolution feature map S2' and fusing the spatial information into the convolution feature map S2' to obtain a first fusion feature map S2"; performing sparse convolution operation on the first fusion feature map S2" to obtain a second convolution feature map I; performing channel expansion on the second convolution feature map I to obtain an expanded-channel feature map; performing segmentation on the expanded-channel feature map in the channel dimension to obtain a first sub-feature map x and a second sub-feature map y; the first sub-feature map x and the second sub-feature map y are images obtained by segmenting the expanded-channel feature map; performing depthwise separable convolution of different scales on the first sub-feature map x to obtain spatial features of the first sub-feature map x at different scales; using the second sub-feature map y as a gating signal to perform weighted multiplication on the spatial features of different scales to obtain a first weighted feature map; restoring the number of channels of the weighted feature map and performing residual connection between the feature map with the restored number of channels and the first feature map S2 to obtain a residual feature map I'; performing feature information fusion of different scales on the residual feature map I' to obtain the first enhanced feature map S3.

[0035] Specifically, please refer to Figure 3, the first enhanced low-light feature extraction unit first performs a convolution operation on S2', extracts the spatial information of the feature map S2' after convolution, combines the feature map S2” with the extracted spatial information with a convolutional gating module and a multi-scale convolutional attention module, dynamically adjusts the feature weights, and optimizes the detection effect of weak targets through multi-scale feature fusion to obtain the first enhanced feature map S3. In the convolutional gating module, the input feature map S2” first undergoes feature extraction through PConv (Partial Convolution). PConv reduces the computational load through sparse convolution operations while retaining key feature information to obtain the second feature map I after convolution. Then, the second feature map I after convolution undergoes channel expansion through a 1×1 convolution, and the number of channels is doubled. The expanded feature map is evenly divided into two parts in the channel dimension, denoted as the first sub-feature map x and the second sub-feature map y respectively. Among them, the first sub-feature map x undergoes parallel 3×3 depthwise separable convolution and 5×5 depthwise separable convolution to extract spatial features of different scales and add the results; the second sub-feature map y serves as a gating signal to weight and adjust the extracted features. The two parts complete feature weighting through element-wise multiplication to implement the gating mechanism. The weighted feature map restores the number of channels through a 1×1 convolution and is added to the input feature map through a residual connection to obtain the feature map I' after residual. Different-scale feature information fusion is performed on the feature map I' after residual to obtain the first enhanced feature map S3. This gating mechanism can effectively suppress the interference of irrelevant information, strengthen the expression of the target area, reduce the influence of background noise, and enable the network to focus more on the key target features in low-illumination environments.

[0036] In a specific embodiment, the spatial size of the first enhanced feature map S3 is twice the spatial size of the second enhanced feature map S4, and the spatial size of the second enhanced feature map S4 is twice the spatial size of the third enhanced feature map S5.

[0037] In a specific embodiment, the number of channels of the first enhanced feature map S3 is 0.5 times the number of channels of the second enhanced feature map S4, and the number of channels of the second enhanced feature map S4 is 0.5 times the number of channels of the third enhanced feature map S5.

[0038] Specifically, the first enhanced feature map S3 passes through two improved basic module layers (the second convolutional layer is replaced by an enhanced low-light feature extraction module) to obtain two feature maps S4 and S5 of different scales, namely the second enhanced feature map S4 and the third enhanced feature map S5. Each time passing through a basic module layer, the spatial size of the feature map is reduced by half, and the number of channels is doubled, thereby providing multi-level feature representations for subsequent detection tasks.

[0039] In a specific embodiment, performing feature information fusion of different scales on the residual feature map I' to obtain the first enhanced feature map S3 includes: performing global max pooling operation and average pooling operation on the residual feature map I', and splicing the results after pooling to obtain a spliced feature map; performing spatial relationship learning and Sigmoid function activation on the spliced feature map at different scales to obtain a spatial attention feature map; multiplying the spatial attention feature map by the residual feature map I' and performing weighted adjustment to obtain a second weighted feature map; using a preset multi-branch convolution structure to perform multi-scale feature extraction and multi-scale convolution on the second weighted feature map to obtain feature information of the residual feature map I' at different scales; performing feature information fusion on the feature information at different scales through a channel shuffle operation to obtain a second fusion feature map; adding the second fusion feature map to the convolved feature map S2' to obtain the first enhanced feature map S3.

[0040] Specifically, the residual feature map I' enters the multi-scale convolutional attention module. First, the spatial distribution of the feature map is dynamically adjusted through spatial attention. The specific operation is as follows: The residual feature map I' respectively undergoes global max pooling and average pooling operations to capture the significant features and global information of the image. Then, the results after pooling are spliced and processed through a 7×7 convolution to learn the spatial relationships at different scales. The convolved features are then activated by the Sigmoid activation function to generate the final spatial attention feature map. Finally, the spatial attention feature map is multiplied by the residual feature map I' to perform weighted adjustment on the features, obtaining the feature map I”, enhancing the model's attention ability to important spatial regions and making the model focus on the regions where the target may exist. After that, the feature map I” undergoes multi-scale feature extraction through a multi-branch convolution structure, passing through 1×1, 3×3, and 5×5 convolutions in parallel, and fusing the feature information from different scales through a channel shuffle operation to obtain a second fusion feature map. This second fusion feature map is added to the convolved feature map S2' to obtain the feature map S3.

[0041] In a specific embodiment, the encoding and fusion of the multiple enhanced feature maps to obtain the output feature map includes: performing feature extraction based on space and frequency domain on the third enhanced feature map S5 to obtain a frequency domain feature output map F5; encoding and performing cross-scale feature fusion based on convolution on the first enhanced feature map S3, the second enhanced feature map S4, and the frequency domain feature output map F5 to obtain the output feature map F.

[0042] Specifically, the output features {S3, S4, S5} of the last three stages of the backbone network are used as the input of the encoder. The final feature F5 is obtained by processing S5 through a 1×1 convolution and a feature extraction module based on spatial and frequency domains. Encoding and convolution-based cross-scale feature fusion are performed according to the first enhanced feature map S3, the second enhanced feature map S4, and the frequency-domain feature output map F5 to obtain the output feature map F.

[0043] In a specific embodiment, the encoding and convolution-based cross-scale feature fusion of the first enhanced feature map S3, the second enhanced feature map S4, and the frequency-domain feature output map F5 to obtain the output feature map F includes: adjusting the dimension of the frequency-domain feature output map F5 to obtain a dimension-adjusted feature map F5'; performing an upsampling operation on the dimension-adjusted feature map F5' to obtain a first upsampled feature map; the size of the first upsampled feature map is the same as the size of the second enhanced feature map S4; performing content-guided attention fusion on the first upsampled feature map and the second enhanced feature map S4 to obtain a third fused feature map; using a preset RepC3 module and a convolutional layer to perform feature extraction on the third fused feature map to obtain a second feature map Y4; performing an upsampling operation on the second feature map Y4 to obtain a second upsampled feature map; the size of the second upsampled feature map is the same as the size of the first enhanced feature map S3; performing content-guided attention fusion on the second upsampled feature map and the first enhanced feature map S3 to obtain a fourth fused feature map; using a preset RepC3 module to perform feature extraction on the fourth fused feature map to obtain a third feature map F3; performing a downsampling operation on the third feature map F3 to obtain a first downsampled feature map; the size of the first downsampled feature map is the same as the size of the second feature map Y4; performing content-guided attention fusion on the first downsampled feature map and the second feature map Y4 to obtain a fifth fused feature map; using a preset RepC3 module to perform feature extraction on the fifth fused feature map to obtain a fourth feature map F4; performing a downsampling operation on the fourth feature map F4 to obtain a second downsampled feature map; the size of the second downsampled feature map is the same as the size of the dimension-adjusted feature map F5'; performing content-guided attention fusion on the second downsampled feature map and the dimension-adjusted feature map F5' to obtain a sixth fused feature map; using a preset RepC3 module to perform feature extraction on the sixth fused feature map to obtain a fifth feature map F5; performing a concatenation operation on the third feature map F3, the fourth feature map F4, and the fifth feature map F5 to obtain the output feature map F.

[0044] Specifically, {S3, S4, F5} are input into the convolutional cross-scale feature fusion module CCFM for operation to obtain the output feature map F. CCFM is a structure of FPN, which constructs a feature pyramid through top-down and bottom-up paths. In CCFM, the frequency-domain feature output map F5 will first pass through a 1×1 convolution to adjust the dimension to obtain the dimension-adjusted feature map F5'. Then, it passes through an upsampling module to adjust the size of the feature map to be the same as that of the second enhanced feature map S4, and then they are input into the content-guided attention fusion module simultaneously to form a new feature map. Then, after passing through the RepC3 module and a 1×1 convolution to extract features, the second feature map Y4 is obtained. Then, the second feature map Y4 passes through an upsampling module to adjust the size of the feature map to be the same as that of S3, and then they are input into the content-guided attention fusion module simultaneously to form a new feature map. Then, the feature map F3 is obtained after passing through the RepC3 module. After that, through a convolution module for downsampling operation to adjust the size of the feature map to be the same as that of the second feature map Y4, and then they are input into the content-guided attention fusion module simultaneously to form a new feature map. After passing through the RepC3 module and a 1×1 convolution to extract features, the fourth feature map F4 is obtained. After that, through a convolution module for downsampling operation to adjust the size of the feature map to be the same as that of the dimension-adjusted feature map F5', and then they are input into the content-guided attention fusion module simultaneously to form a new feature map. After passing through the RepC3 module and a 1×1 convolution to extract features, the fifth feature map F5 is obtained. After splicing F3, F4, and F5, the output feature map F is obtained.

[0045] In a specific embodiment, the extracting the frequency-domain feature output map F5 by performing spatial and frequency-domain feature extraction on the third enhanced feature map S5 includes: adjusting the number of channels of the third enhanced feature map S5 to obtain the third enhanced feature map after channel adjustment; splitting the third enhanced feature map after channel adjustment in the channel dimension to obtain the third sub-feature map x_1 and the fourth sub-feature map x_2; performing local feature extraction of different scales on the third sub-feature map x_1 and the fourth sub-feature map x_2 to obtain the first local feature map and the second local feature map; splicing and convolving and reducing the dimension of the first local feature map and the second local feature map along the channel dimension to obtain the dimension-reduced feature map x0; performing feature extraction of different dimensions on the feature map x0, and performing spatial and frequency-domain feature extraction on the features of different dimensions to obtain the frequency-domain feature output map F5.

[0046] Specifically, please refer to Figure 4 , the features of the third enhanced feature map S5 output in the backbone network are first processed by a 1×1 convolution to adjust the number of channels to 256 to obtain the third enhanced feature map S5'. The third enhanced feature map S5' is used as the input of the spatial and frequency-domain feature extraction module. The spatial and frequency-domain feature extraction module is as Figure 4As shown, by combining spatial-domain and frequency-domain feature extraction, the module can utilize both local details and global context information simultaneously, significantly enhancing the model's performance in low illumination and improving its ability to capture feature interactions within a scale. First, the third enhanced feature map S5' passes through a 1×1 convolutional layer to double the number of channels, and then is split into two parts along the channel dimension, namely the third sub-feature map x_1 and the fourth sub-feature map x_2. The third sub-feature map x_1 and the fourth sub-feature map x_2 respectively pass through 3×3 and 5×5 separable convolutional modules to extract local features at different scales, obtaining the first local feature map and the second local feature map. The first local feature map and the second local feature map are concatenated along the channel dimension and reduced in dimension through a 1×1 convolution to obtain the feature map x0 after dimensionality reduction. Then, the FSAS frequency-domain self-attention is introduced. The frequency-domain self-attention transforms the features into the frequency domain through the Fourier transform, enhancing the correlation between high-frequency and low-frequency features, enabling better capture of global context information, and helping to extract information such as textures or edges blurred due to low illumination. First, the feature map x0 after dimensionality reduction passes through a 1×1 convolution to increase the dimension, and then is split into three parts along the channel dimension. For these three parts, 1×1 convolution and 3×3 depth convolution are respectively used to extract Fq, Fk, and Fv. Then, Fq and Fk are transformed into the frequency domain through the Fourier transform to obtain the frequency-domain representations of Q and K. The dot product operation is performed on Q and K to obtain the feature A, and then these optimized frequency features are transformed back to the spatial domain through the inverse fast Fourier transform. The dot product operation is performed on the normalized feature and Fv to obtain the final frequency-domain feature representation, indicating the correlation between different feature components in the frequency domain, enhancing the correlation between high-frequency and low-frequency features. Finally, the dimension is restored through a 1×1 convolution to obtain the feature representation out_f that integrates frequency information. To avoid information loss and improve the training stability of the network, out_f and x0 are added together to obtain the final output map F5.

[0047] Replace the original fusion operation in the feature fusion module in the encoder with an optimized content-guided attention fusion module, as Figure 5As shown, after the module receives two feature maps, it first performs a channel concatenation operation to fuse the two feature maps. Then, the combined feature map is fed into a multi-path branch module. One path passes through a convolution operation, then uses mean and max operations respectively, and after a channel concatenation operation, global semantic features are extracted through convolution. The other path extracts the global average value of the feature map through global average pooling, and extracts deep features through 1×1 convolution, ReLU activation, and another 1×1 convolution operation. The outputs of the above branches are fused through an addition operation and then concatenated with the combined feature map in channels. The fusion result is rearranged in channels through a channel shuffle operation to improve feature interaction. Finally, channel attention weights are generated through a Sigmoid activation function. The attention weights are multiplied pixel by pixel with the input feature map to generate a weighted feature map. The attention-weighted feature map and the original input feature map are fused through an element-wise addition operation respectively. Finally, information is integrated through a concatenation operation and a convolution operation to generate an output feature map. This module effectively fuses features from different sources through a multi-path branch and channel concatenation strategy, and strengthens feature interaction between different channels using the channel shuffle operation, which helps capture fine-grained local information and global semantic information, and improves the recognition ability of weak targets.

[0048] The technique of selecting the initial query of the decoder through IoU-aware query is used to predict the detection box for the feature F and perform screening processing on it. The screened feature F is input into the decoder and the detection head module to process the selected queries to generate the final detection output, completing the detection of the target.

[0049] (1) Compared with existing object detection methods, the present invention can effectively address problems such as insufficient brightness, uneven illumination, and blurred details in low-illumination environments by introducing an enhanced low-light feature extraction module. By combining convolutional gating and multi-scale convolutional attention mechanisms, this method can suppress noise interference, strengthen the expression of the target area, and improve the detection accuracy of weak targets.

[0050] (2) Compared with existing object detection methods, the present invention improves the robustness of object detection in low-illumination environments by introducing a feature extraction module based on spatial and frequency domains. Different scales of local features are extracted through multi-scale local features, and the frequency domain self-attention enhances the correlation between high-frequency and low-frequency features through Fourier transform, helping to capture image details and enhancing the recognition ability of weak targets. Especially in complex lighting and backgrounds, the detection accuracy and stability are improved.

[0051] (3) Compared with existing object detection methods, the present invention adopts a content-guided attention fusion module, which can adaptively adjust the fusion method of features at different scales. Through a multi-path branch design, local details and global structure information are extracted, and feature fusion is optimized through an attention mechanism, significantly improving the ability to extract small and weak targets and detailed features, especially performing well in low-illumination environments.

[0052] In a specific embodiment, please refer to Figure 6 , which is a flowchart of the steps of a method for object detection based on low-illumination images according to the present application. The method includes:

[0053] Step 201: Adjust the spatial size and number of channels of the visible light image to be detected to obtain a first feature map;

[0054] Step 202: Perform multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps;

[0055] Step 203: Encode and fuse the multiple enhanced feature maps to obtain an output feature map;

[0056] Step 204: Perform predictions on the output feature map to obtain the target category and target bounding box position in the visible light image to be detected.

[0057] The method in this embodiment adjusts the spatial size and number of channels of the visible light image to be detected to obtain a first feature map; performs multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; encodes and fuses the multiple enhanced feature maps to obtain an output feature map; and performs predictions on the output feature map to obtain the target category and target bounding box position in the visible light image to be detected. The enhanced feature maps obtained by performing multiple feature enhancements on the first feature map can effectively enhance the features of distant or small and weak targets, so that the features of distant or small and weak targets can be clearly displayed in the enhanced feature maps, thereby effectively coping with problems such as insufficient brightness, uneven illumination, and blurred details in low-illumination environments. Encoding and fusing the multiple enhanced feature maps to obtain an output feature map and performing predictions based on the output feature map can effectively extract the detailed features of the target, especially improving the detection effect of distant or small and weak targets.

[0058] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

[0059] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A target detection system based on low-illumination images, characterized in that, The system includes: a first feature extraction module, an enhanced low-light feature extraction module, an encoding module, and a decoding module; the output end of the first feature extraction module is connected to the input end of the enhanced low-light feature extraction module, and the output end of the enhanced low-light feature extraction module is connected to the input end of the encoding module; the output end of the encoding module is connected to the input end of the decoding module; The first feature extraction module is used to adjust the spatial size and the number of channels of the visible light image to be detected, and obtain a first feature map; The enhanced low-light feature extraction module is used to perform multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; The encoding module is used to encode and fuse the multiple enhanced feature maps to obtain an output feature map; The decoding module is used to predict the output feature map to obtain the target category and the target bounding box position in the visible light image to be detected.

2. The object detection system based on low-illumination images according to claim 1, characterized in that, The enhanced low-light feature extraction module includes a first enhanced low-light feature extraction unit, a second enhanced low-light feature extraction unit, and a third enhanced low-light feature extraction unit connected in sequence; The first enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first feature map to obtain a first enhanced feature map; The second enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the first enhanced feature map to obtain a second enhanced feature map; The third enhanced low-light feature extraction unit is used to perform spatial feature extraction at different scales and feature information fusion at different scales on the second enhanced feature map to obtain a third enhanced feature map.

3. The object detection system based on low-light images according to claim 2, wherein Performing spatial feature extraction at different scales and feature information fusion at different scales on the first feature map to obtain a first enhanced feature map includes: Performing convolutional addition on the first feature map and a preset convolutional layer to obtain a first convolution feature map; Performing spatial information extraction on the convolution feature map and fusing the spatial information into the convolution feature map to obtain a first fusion feature map; Performing a sparse convolution operation on the first fusion feature map to obtain a second convolution feature map; Performing channel expansion on the second convolution feature map to obtain a channel-expanded feature map; Dividing the channel-expanded feature map in the channel dimension to obtain a first sub-feature map and a second sub-feature map; the first sub-feature map and the second sub-feature map are images obtained by dividing the channel-expanded feature map; Performing depthwise separable convolutions at different scales on the first sub-feature map to obtain spatial features of the first sub-feature map at different scales; Using the second sub-feature map as a gating signal to perform weighted multiplication on the spatial features at different scales to obtain a first weighted feature map; Performing channel number recovery on the weighted feature map and performing residual connection between the feature map after channel number recovery and the first feature map to obtain a residual feature map; Performing feature information fusion at different scales on the residual feature map to obtain the first enhanced feature map.

4. The object detection system based on low-illumination images according to claim 3, characterized in that, Performing feature information fusion of the residualized feature map at different scales to obtain the first enhanced feature map includes: Performing global max pooling operation and average pooling operation on the residualized feature map, and concatenating the pooled results to obtain a concatenated feature map; Performing spatial relationship learning and Sigmoid function activation on the concatenated feature map at different scales to obtain a spatial attention feature map; Multiplying and weighted adjusting the spatial attention feature map with the residualized feature map to obtain a second weighted feature map; Using a preset multi-branch convolution structure to perform multi-scale feature extraction and multi-scale convolution on the second weighted feature map to obtain feature information of the residualized feature map at different scales; Performing feature information fusion on the feature information at different scales through a channel shuffle operation to obtain a second fusion feature map; Adding the second fusion feature map to the convolved feature map to obtain the first enhanced feature map.

5. The object detection system based on low-illumination images according to claim 2, characterized in that, Encoding and fusing the multiple enhanced feature maps to obtain an output feature map includes: Performing feature extraction based on space and frequency domain on the third enhanced feature map to obtain a frequency domain feature output map; Encoding and performing cross-scale feature fusion based on convolution on the first enhanced feature map, the second enhanced feature map, and the frequency domain feature output map to obtain the output feature map.

6. The object detection system based on low-illumination images according to claim 5, characterized in that, Encoding and performing cross-scale feature fusion based on convolution on the first enhanced feature map, the second enhanced feature map, and the frequency domain feature output map to obtain the output feature map includes: Adjusting the dimension of the frequency domain feature output map to obtain a dimension-adjusted feature map; Performing an upsampling operation on the dimension-adjusted feature map to obtain a first upsampled feature map; the size of the first upsampled feature map is the same as that of the second enhanced feature map; Performing content-guided attention fusion on the first upsampled feature map and the second enhanced feature map to obtain a third fusion feature map; Using a preset RepC3 module and a convolution layer to perform feature extraction on the third fusion feature map to obtain a second feature map; Performing an upsampling operation on the second feature map to obtain a second upsampled feature map; the size of the second upsampled feature map is the same as that of the first enhanced feature map; Performing content-guided attention fusion on the second upsampled feature map and the first enhanced feature map to obtain a fourth fusion feature map; Using a preset RepC3 module to perform feature extraction on the fourth fusion feature map to obtain a third feature map; Performing a downsampling operation on the third feature map to obtain a first downsampled feature map; the size of the first downsampled feature map is the same as that of the second feature map; Performing content-guided attention fusion on the first downsampled feature map and the second feature map to obtain a fifth fusion feature map; Using a preset RepC3 module to perform feature extraction on the fifth fusion feature map to obtain a fourth feature map; Perform a downsampling operation on the fourth feature map to obtain a second downsampled feature map; the size of the second downsampled feature map is the same as the size of the dimension-adjusted feature map; Perform content-guided attention fusion on the second downsampled feature map and the dimension-adjusted feature map to obtain a sixth fused feature map; Use a preset RepC3 module to perform feature extraction on the sixth fused feature map to obtain a fifth feature map; Perform a concatenation operation on the third feature map, the fourth feature map, and the fifth feature map to obtain the output feature map.

7. The object detection system based on low-illumination images according to claim 5, wherein The performing feature extraction on the third enhanced feature map based on space and frequency domain to obtain a frequency domain feature output map includes: Adjust the number of channels of the third enhanced feature map to obtain a third enhanced feature map with adjusted channels; Segment the third enhanced feature map with adjusted channels in the channel dimension to obtain a third sub-feature map and a fourth sub-feature map; Perform local feature extraction at different scales on the third sub-feature map and the fourth sub-feature map to obtain a first local feature map and a second local feature map; Concatenate and perform convolution and dimensionality reduction on the first local feature map and the second local feature map along the channel dimension to obtain a feature map with reduced dimensions; Perform feature extraction on the feature map in different dimensions, and perform feature extraction based on space and frequency domain on the features in different dimensions to obtain a frequency domain feature output map.

8. The object detection system based on low-light images according to claim 5, wherein The spatial size of the first enhanced feature map is 2 times the spatial size of the second enhanced feature map, and the spatial size of the second enhanced feature map is 2 times the spatial size of the third enhanced feature map.

9. The target detection system based on low-illumination images according to claim 5, wherein The number of channels of the first enhanced feature map is 0.5 times the number of channels of the second enhanced feature map, and the number of channels of the second enhanced feature map is 0.5 times the number of channels of the third enhanced feature map.

10. A target detection method based on low-illumination images, characterized in that, The method includes: Adjust the spatial size and the number of channels of the visible light image to be detected to obtain a first feature map; Perform multiple feature enhancements on the first feature map to obtain multiple enhanced feature maps; Encode and fuse the multiple enhanced feature maps to obtain an output feature map; Perform prediction on the output feature map to obtain the target category and the target bounding box position in the visible light image to be detected.