Small target identification method for improving target detection precision under dark condition
By introducing attention aggregation fusion module and dynamic perception head in the YOLOv8 backbone network, the problem of low object detection accuracy under low light conditions is solved, and high-precision recognition of small objects in night environments is achieved, which is suitable for real-time detection of street light control systems.
Patent Information
- Application Number
- CN202510470331.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
Smart Images

Figure CN120339585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a target detection technology, and in particular to a small target recognition method for improving the target detection accuracy under dark conditions. Background Art
[0002] Existing target detection algorithms such as yolov8[1] simply fuse multi-scale features using AFPN and use a fixed detection head to frame objects, performing well under normal light conditions. However, under low light conditions, due to the loss of details and a large number of noise problems, the performance is not satisfactory and it is difficult to be used for night recognition projects. Some deep learning-based low light enhancement methods. RetinexNet[2] simulates the visual processing mechanism of the human eye and enhances low-light images by separating the reflection component and illumination component of the image. LLNet[3] denoises and enhances low-light images through an autoencoder, thereby improving the detection performance. Although these methods perform well in low light enhancement, due to their separation from the target detection model, they cannot be dynamically adjusted during the detection process. Compared with separating image enhancement from the detection task, existing research tends to jointly design image enhancement and the target detection network. For example, MAET[4] proposed to simultaneously train an image enhancement network and a target detection network through multi-task learning, improving the detection accuracy of the model under low light conditions. Although these methods have significantly improved the target detection performance in low light environments, due to the large number of parameters and large computational amount of the above models, it is difficult to ensure real-time performance in the street lamp control system. Summary of the Invention
[0003] To overcome the above defects, the present invention provides a small target recognition method for improving the target detection accuracy under dark conditions. By using a low illuminance target detection model, it can improve the accuracy of object detection in a dark environment. The present invention is deployed on street lamps and can significantly improve the detection accuracy when identifying small targets on the road at night.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A small target recognition method for improving the target detection accuracy under dark conditions, comprising the following steps:
[0006] Step 1, image acquisition, obtaining an image by capturing a panoramic view of the road through a camera;
[0007] Step 2, image feature extraction, inputting the image into the backbone network of YOLOv8, the backbone network of YOLOv8 being a CSPDarknet53 structure, and extracting multi-level features of the image through a number of convolutional layers and residual connections to obtain a multi-scale feature map;
[0008] Step 3: Upsample the feature map generated by p4 and fuse it with the feature map generated by p3 to generate S1 and send it to the first detection head. Upsample the attention fusion aggregation module and fuse it with p4 to generate S2. Downsample S1 and fuse it with S2 to generate S3 and input it into the second detection head. Downsample S3 and fuse it with the feature map generated by the attention aggregation fusion module to generate S4 and input it into the third detection head. Identify small targets through the above three detection heads.
[0009] Preferably, the attention fusion aggregation module uses pooling operations with three different kernel sizes to extract features of various sizes and generate a feature map Z. i (i = 1, 2, 3). After the third max pooling, a convolution operation with a large kernel is adopted, and the process is as follows:
[0010] Z i = MaxPool2d i (Z, k i , s). (i = 1, 2, 3)
[0011] For the generated feature map Z3, first perform a one-dimensional convolution operation on it in the vertical direction to generate a feature map Z C vert , and then perform a one-dimensional convolution in the horizontal direction on the basis of the paired Z C vert to obtain Z C horzs . Finally, generate an adaptive weight map A through a 1*1 convolution: C :
[0012]
[0013] Finally, use A C to perform adaptive weighting on Z3 to obtain an accurate and comprehensive representation of the features as follows:
[0014]
[0015] Preferably, the first detection head, the second detection head, and the second detection head are dynamic perception heads. The dynamic perception head includes dynamic convolution, and the adaptive weight W i of the dynamic convolution is represented as a function of generating the input feature X as follows:
[0016] W i = f(X), i = 1, 2,..., k
[0017] The output Y of the dynamic convolution is calculated from the input feature X and the adaptive weight W i as follows:
[0018]
[0019] Preferably, the loss function I-MPDIou of the dynamic perception head is as follows:
[0020]
[0021] where inter is the intersection of the ground truth bounding box and the predicted bounding box of the image, union is the union of the ground truth bounding box and the predicted bounding box of the image. h and w are the height and width of the box, and d i is the Euclidean distance between the top-left corner points of the predicted box and the ground truth box.
[0022] Preferably, in step 2, after the image passes through the backbone network of YOLOv8, it enters the initial stem layer p1. The convolutional kernel size of the p1 layer is 3x3, the stride is 2, and the number of output channels is 64. Then, the image features enter the p2, p3, p4, and p5 layers. Each layer consists of a convolutional layer and a cross-stage partial fusion layer. The convolutional kernel size of each layer is 3x3, the stride is 2. The cross-stage partial fusion layer divides the input feature map X into two parts X1 and X2. Among them, X1 directly passes through a 3x3 convolutional layer, and X2 passes through several residual blocks and then fuses with X1. The residual block contains two 3x3 convolutional layers and a skip connection.
[0023] Preferably, the operations of the residual block are as follows: The first 3x3 convolutional layer performs a convolution operation on X2 to obtain the feature map F1; the second 3x3 convolutional layer: performs a convolution operation on F1 to obtain the feature map F2; skip connection: adds X2 and F2 element-wise to obtain the output R of the residual block; adds X1 and R element-wise to obtain the feature map Y;
[0024] Send p5 to the attention aggregation and fusion module. First, use a 1x1 convolution to reduce the channel dimension of p5 to obtain the feature map Z. Then, use three pooling layers with different kernel sizes to extract features from Z to obtain Z i (i = 1, 2, 3). Subsequently, perform one-dimensional convolutions on Z3 in the horizontal and vertical directions to generate the attention weight A C and finally use the attention weight map A C to weight Z3 to obtain a comprehensive feature map.
[0025] Preferably, the camera is placed on the top of the street lamp, and the camera transmits the collected images in real time to the edge processing device Jetson TX2 NX deployed in the street lamp distribution box through a wired method.
[0026] Preferably, information such as the position and category of the identified small targets is transmitted through the Mqtt protocol, and the Mqtt messages are transmitted to the street lamp control server through the Lora wireless communication technology.
[0027] Advantages of the present invention: Under low illumination conditions, the present invention solves the problems of multi-noise and few details in target detection and improves the detection accuracy. The present invention is deployed on street lamps at night to identify small targets on the road, and can significantly improve the detection accuracy. Description of the Drawings
[0028] Figure 1 It is a schematic structural diagram of a low-illumination detection model.
[0029] Figure 2 It is a schematic diagram of the visualization comparison of the effects of the low-illumination target detection algorithm and the mainstream target detection algorithms. Detailed Implementation Modes
[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0031] This embodiment discloses a small target recognition method for improving the target detection accuracy under dark conditions. Aiming at the detection defects of existing target detection algorithms under low light conditions, an attention aggregation fusion module, a dynamic feature extraction perception head and an optimized loss function are adopted, which can effectively solve problems such as multi-noise and few details existing under dark conditions, so as to achieve high monitoring accuracy under low illumination. Specifically:
[0032] This embodiment has been applied in the field for the small target recognition project on the night road, such as Figure 1As shown in the figure, an algorithm deployed using the edge processing device Jetson TX2 NX. Jetson places the TX2 into the distribution box of the street lights at both ends of the road, places the external camera on the top of the street light, and aims it at the road to ensure that the entire panorama of the road can be covered. The resolution of the camera is configured according to actual needs. The camera transmits the captured images in real time to the edge processing device Jetson TX2 NX deployed in the distribution box of the street light through a wired connection. The captured images are first input into the network proposed in the present invention. The first part of the network is the detection backbone network of YOLOv8. The backbone network of YOLOv8 adopts the CSPDarknet53 structure, and extracts multi-level features of the image through multiple convolutional layers and residual connections. Specifically, the input image passes through convolutional layers, BatchNormalization (BN) layers, and LeakyReLU activation functions, gradually extracting low-level features (such as edges and textures) and high-level features (such as object shapes and semantic information). After the image passes through the backbone network of YOLOv8, it enters the initial stem layer p1. The convolutional kernel size of the p1 layer is 3x3, the stride is 2, and the number of output channels is 64. Through the convolutional operation, the size of the feature map is compressed to half of the input, and at the same time the number of channels increases to retain more feature information. Subsequently, the image features enter the p2, p3, p4, and p5 layers. Each layer consists of a convolutional layer and a cross-stage partial fusion layer (C2f layer). The convolutional kernel size of each layer is 3x3, the stride is 2, and the number of output channels increases layer by layer (for example, p2 is 128, p3 is 256, p4 is 512, p5 is 1024). The C2f layer divides the input feature map into two parts. One part directly passes through the convolutional layer, and the other part passes through multiple residual blocks and then fuses with the first part. Specifically, the input feature map X is first divided into two parts X1 and X2. Among them, X1 directly passes through a 3x3 convolutional layer, and X2 passes through multiple residual blocks (each residual block contains two 3x3 convolutional layers and a skip connection). The specific operations of each residual block are as follows: The first 3x3 convolutional layer: performs a convolutional operation on X2 to obtain the feature map F1; the second 3x3 convolutional layer: performs a convolutional operation on F1 to obtain the feature map F2; skip link: adds X2 and F2 element by element to obtain the output R of the residual block; feature fusion: adds X1 and R element by element to obtain the final output feature map Y; where p5 is sent to the attention aggregation fusion module. First, use a 1x1 convolution to reduce the channel dimension of p5 to obtain the feature map Z. Subsequently, use three pooling layers with different kernel sizes to extract features from Z to obtain Z i (i = 1, 2, 3). Subsequently, perform one-dimensional convolutions on Z3 in the horizontal and vertical directions to generate the attention weight A C , and finally use the attention weight map A C to weight Z3 to obtain a comprehensive feature map.
[0033] The multi-scale features extracted from the p2, p3, p4, and p5 layers are fed into the feature fusion module. This module further improves the representation ability of the feature maps by fusing the feature maps at different stages in the backbone network. Specifically, first, the feature maps at each scale are upsampled or downsampled to make their sizes consistent. Then, the feature maps at different scales are fused by weighted summation (F i represents the feature map of the i-th size, and w i is the corresponding weight) to fuse the feature maps at different scales. The feature map generated by p4 is upsampled and fused with the feature map generated by p3 to generate S1 and sent to the first detection head. The attention fusion aggregation module is upsampled and aggregated with p4 to generate S2. S1 is downsampled and aggregated with S2 to generate S3 and sent to the second detection head. S3 is downsampled and aggregated with the feature map generated by the attention aggregation fusion module to generate S4 and sent to the third detection head. The first detection head is used to detect small objects, the second detection head is used to detect medium-sized objects, and the third detection head is used to detect large objects. Each perception head contains a bounding box prediction branch and a class prediction branch. The bounding box prediction branch predicts the bounding box of the target object through a regression algorithm. Specifically, the network outputs the center coordinates (x, y), width (w), and height (h) of each bounding box. To optimize the prediction accuracy of the bounding box, the present invention proposes an optimized loss function that combines innerIou and MpDIou and can better constrain the prediction of the bounding box during the training process. The class prediction branch outputs the class probability of each bounding box through the softmax function. The network uses the cross-entropy loss function to optimize the accuracy of class prediction during the training process. Information such as the position and class of the identified small targets is transmitted through the Mqtt protocol. The Mqtt protocol is a lightweight publish / subscribe message transmission protocol suitable for use in low-bandwidth and unstable network environments. Mqtt messages are transmitted to the street lamp control server through Lora wireless communication technology. Lora is a low-power and long-distance wireless communication technology suitable for use in street lamp control systems. After receiving the small target information, the street lamp control server calculates the street lamps that the small target is about to pass by according to the position, class, and speed of the small target. The server lights up the street lamps that the small target is about to pass by through the control circuit. Specifically, the server sends a control signal to the corresponding street lamp controller, and the controller lights up the street lamp after receiving the signal. When no small target is detected, the server does not send a lighting signal, and the street lamps remain off, thus achieving the effect of energy conservation and environmental protection.
[0034] Attention Aggregation Fusion Module AFAM: Compared with existing feature pyramid networks, such as FPN, ASPP, PAFPN (PANET+FPN) [5, 6, 7], they perform well in target detection under normal lighting conditions, but not in low-light scenes. This is because their way of fusing multi-scale features is too simple. In view of this, the AFAM module can effectively capture the key feature areas in low-light images by introducing a large-core attention mechanism, thereby enhancing the accuracy and robustness of low-light target detection.
[0035] AFAM uses three pooling operations with different kernel sizes to extract features of various sizes and generate feature maps Z i (i=1,2,3), after the third maximum pooling, a convolution operation with a large kernel is used. The specific process is as follows:
[0036] Z i =MaxPool2d i (Z, k i , s).(i=1,2,3)
[0037] For the generated feature map Z3, we first perform a one-dimensional convolution operation on it in the vertical direction to generate the feature map Z C vert , followed by the combination of Z C vert Based on the one-dimensional convolution in the horizontal direction, we get Z C horzs , and finally generate an adaptive weight map A through 1*1 convolution C , which indicates the importance of different regions in the feature map:
[0038]
[0039] Finally, using A C Adaptively weight Z3 to obtain an accurate and comprehensive representation of the features:
[0040]
[0041] Dynamic Perception Head (DPH): The core of the DFH module lies in its dynamic convolution mechanism. Unlike traditional convolution, the weights of the dynamic convolution kernel are adaptively generated, which means that the kernel can be dynamically adjusted in different scenarios, enhancing the flexibility and adaptability of feature extraction. In dynamic convolution, the adaptive weight W i is represented as a function that generates input features X:
[0042] W i =f(X), i=1, 2, ..., k
[0043] The output Y of the dynamic convolution is determined by the input feature X and the adaptive weight W i jointly:
[0044]
[0045] Optimized loss function (I-MPDIoU): In this embodiment, the idea of innerIou is borrowed and combined with MpDIou to propose an optimized loss function I-MPDIou.
[0046] MPDIoU introduces a minimum point distance correction term and an auxiliary bounding box, enabling the model to focus more on the shape of the target and the changes in the relative position under low illumination conditions, ensuring good detection performance. The minimum distance is obtained from the Euclidean distance between the upper left and lower right corners of the predicted box and the ground truth box. Let box1 and box2 represent the predicted box and the ground truth box respectively, where each box represents [x, y, h, w]. We have:
[0047]
[0048] The introduction of the auxiliary bounding box is achieved through the scaling factor ratio, enabling the flexible application of different strategies between high-IoU samples and low-IoU samples. For the ground truth box B gt and the predicted box B, we have:
[0049]
[0050] The intersection and union of the ground truth box and the predicted box can be calculated from the obtained boundary information:
[0051]
[0052] union=(w gt ×h gt )×(raio) 2 +(w×h)×(raio) 2 -inter
[0053] In summary, the calculation formula of I-MPDIou is:
[0054]
[0055] As Figure 2 shown, under low illumination conditions, the problems of multi-noise and few details in object detection are solved, and the detection accuracy is improved. In the visual comparison with mainstream algorithms, the present invention can better detect objects in low illumination scenarios, and can achieve higher accuracy under the map@0.5 metric.
[0056] Table 1 shows the comparison between the present invention and mainstream detection algorithms on the Ex-Dark dataset based on MAP@0.5.
[0057] Table 1
[0058]
[0059] In Table 1, ADHI-YOLO is the algorithm adopted by the present invention. It can be seen from Table 1 that when compared with mainstream detection algorithms on the Ex-Dark dataset, the present invention has the highest precision.
[0060] Citation index:
[0061] [1] https: / / github.com / ultralytics / ultralytics。
[0062] [2] Wei, C., Wang, W., Yang, W., & Liu, J. (2018). Deep Retinex Decomposition for Low-Light Enhancement.arXiv preprint arXiv:1808.04560 . https: / / arxiv.org / abs / 1808.04560。
[0063] [3] Lore,K.G.,Akintayo,A.,&Sarkar,S.(2017).LLNet:A deep autoencoder approach to natural low-light image enhancement.Pattern Recognition,61,650- 662.https: / / doi.org / 10.1016 / j.patcog.2016.06.008。
[0064] [4] Cui,Z.,Qi,G.-J.,Gu,L.,You,S.,Zhang,Z.,&Harada,T.(2022) . Multitask AET with Orthogonal Tangent Regularity for Dark Object Detection . arXiv preprint arXiv:2205.03346.https: / / arxiv.org / abs / 2205.03346。
[0065] [5] Chen S,Zhao J,Zhou Y,et al.Info-FPN:An Informative Feature Pyramid Network for object detection in remote sensing images[J].Expert Systems with Applications,2023,214:119132。
[0066] [6] Nie L,Li B,Jiao F,et al.ASPP-YOLOv5:A study on constructing pig facial expression recognition for heat stress[J].Computers and Electronics in Agriculture,2023,214:108346。
[0067] [7]W.Xiao,M.Xu,and Y.Lin.Global feature pyramid net-work,2024。
[0068] Certainly, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A small target recognition method for improving the accuracy of target detection in dark conditions, characterized in that, It includes the following steps: Step 1, image acquisition, capturing a panoramic view of the road through a camera to obtain an image; Step 2, image feature extraction, inputting the image into the backbone network of YOLOv8, the backbone network of YOLOv8 being a CSPDarknet53 structure, extracting multi-level features of the image through a number of convolutional layers and residual connections to obtain a multi-scale feature map; Step 3, upsampling the feature map generated by p4 and fusing it with the feature map generated by p3 to generate S1 and sending it to the first detection head, upsampling the attention fusion aggregation module and aggregating it with p4 to generate S2, downsampling S1 and aggregating it with S2 to generate S3 and inputting it into the second detection head, downsampling S3 and aggregating it with the feature map generated by the attention aggregation fusion module to generate S4 and inputting it into the third detection head, and identifying small targets through the above three detection heads.
2. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, characterized in that, The attention fusion and aggregation module uses pooling operations with three different kernel sizes to extract features of various sizes and generate the feature map Z i (i = 1, 2, 3). After the third max pooling, a convolution operation with a large kernel is adopted, and the process is as follows: Z i = MaxPool2d i (Z, k i , s).(i = 1, 2, 3) For the generated specific graph Z3, first perform a one-dimensional convolution operation on it in the vertical direction to generate the feature map Z C vert , and then, based on the combination of Z C vert , perform a one-dimensional convolution in the horizontal direction to obtain Z C horzs . Finally, generate the adaptive weight map A through a 1*1 convolution C : Finally, use A C to perform adaptive weighting on Z3, and obtain an accurate and comprehensive representation of the features as follows: 。 3. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, wherein The first detection head, the second detection head, and the second detection head are dynamic perception heads. The dynamic perception head includes a dynamic convolution, and the adaptive weight W of the dynamic convolution i is expressed as a function of generating the input feature X as follows: W i = f(X), i = 1, 2, …, k The output Y of the dynamic convolution is determined by the input feature X and the adaptive weight W i which is calculated as follows: 。 4. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 3, characterized in that, The loss function I-MPDIou of the dynamic perception head is as follows: Among them, inter is the intersection of the true border and the predicted border of the image, union is the union of the true border and the predicted border of the image, h and w are the height and width of the box, and d i is the Euclidean distance between the upper left corner points of the predicted box and the true box.
5. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, characterized in that, In Step 2, after the image passes through the backbone network of YOLOv8, it enters the initial stem layer p1. The convolutional kernel size of the p1 layer is 3x3, the stride is 2, and the number of output channels is 64. Then, the image features enter the p2, p3, p4, and p5 layers. Each layer consists of a convolutional layer and a cross-stage partial fusion layer. The convolutional kernel size of each layer is 3x3, the stride is 2. The cross-stage partial fusion layer divides the input feature map X into two parts X1 and X2. Among them, X1 directly passes through a 3x3 convolutional layer, and X2 is fused with X1 after passing through a number of residual blocks. The residual block contains two 3x3 convolutional layers and a skip connection.
6. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, wherein, The operations of the residual block are as follows: The first 3x3 convolutional layer performs a convolutional operation on X2 to obtain a feature map F1; the second 3x3 convolutional layer: performs a convolutional operation on F1 to obtain a feature map F2; skip connection: adding X2 and F2 element-wise to obtain the output R of the residual block; adding X1 and R element-wise to obtain the feature map Y; Send p5 into the attention aggregation and fusion module. First, use a 1x1 convolution to reduce the channel dimension of p5 to obtain the feature map Z. Then, use three pooling layers with different kernel sizes to extract features from Z to obtain Z i (i = 1, 2, 3). Subsequently, perform one-dimensional convolutions on Z3 in the horizontal and vertical directions to generate the attention weight A C , and finally use the attention weight map A C to weight Z3 to obtain a comprehensive feature map.
7. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, wherein The camera is placed on top of the street lamp, and the camera transmits the collected image in real time to the edge processing device Jetson TX2 NX deployed in the street lamp distribution box through a wired method.
8. The small target recognition method for improving the target detection accuracy under dark conditions according to claim 1, characterized in that, Information such as the position and category of the identified small targets is transmitted through the Mqtt protocol, and the Mqtt message is transmitted to the street lamp control server through the Lora wireless communication technology.
Citation Information
Cited By
Detection method suitable for small target in road image
CN121074633A