Small Target Detection Method for Autonomous Driving Scenarios

By adopting a learnable anchor strategy and an improved feature fusion module in the small-object detection method, the problem of poor small-object detection performance in autonomous driving scenarios is solved, the accuracy and real-time detection are improved, and the generalization of the method is enhanced.

CN119723047BActive Publication Date: 2025-06-20JIANGSU ANZIDA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411794510.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-06-20
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

The existing small-objective detection methods perform poorly in autonomous driving scenarios, especially in the low detection performance of small-objectives. The method of manually setting up anchors leads to a lack of generalization, especially when the target size distribution between data sets has a large difference, the detection performance will decline.

Method used

A small object detection method based on a single-stage detection method is proposed. Using a learnable anchor strategy, the size and size of anchor is learned through the network, combined with the improved feature fusion module and the decoupled prediction branch, improving the accuracy and real-time detection.

Benefits of technology

By adopting a learnable anchor strategy and an improved feature fusion module, the performance of small object detection is improved, the generalization of the method is enhanced, and the small object detection indicator APs is improved, while reducing the calculation amount of the algorithm and improving the detection speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723047B_ABST
    Figure CN119723047B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of object detection in deep learning, and specifically relates to a small object detection method for autonomous driving scenarios. The present invention innovatively adopts an improved structural design. Through a bidirectional feature fusion module, high-level semantic information and low-level feature information are better fused. At the same time, a learnable anchor strategy is adopted. Through randomly initialized anchor sizes and scales, during the training phase, the anchor sizes and scales are learned according to the error loss. Compared with the traditional manually designed anchor hyperparameters, this method has better generalization. Experiments prove that while achieving better performance with fewer anchor numbers, the performance index APs for small object detection is also improved. Fewer anchors also reduce the computational complexity of the algorithm and improve the detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection in deep learning, and particularly relates to a small object detection method for autonomous driving scenarios. Background Art

[0002] At present, deep learning small object detection methods can be divided into single-stage and two-stage detection methods. Single-stage detection methods include Yolo series, SSD, RetinaNet, etc., which extract image features through the network and directly predict the category and location information of the target on the feature map. Two-stage detection methods mainly focus on the Faster R-CNN series. First, the neural network extracts features from the image, then the Region Proposal Network (RPN) is used to initially extract potential target location information, and finally the final prediction result is obtained based on the initially predicted proposal boxes. Single-stage methods have a faster detection speed. Compared with two-stage detection methods, they omit the step of generating potential target regions and are an end-to-end detection method.

[0003] In real life, due to the inconsistent size distributions of targets in different scenarios, single-stage detection methods such as SSD and RetinaNet that adopt the anchor mechanism set the size and dimensions of the anchor through clustering or manual experience, resulting in the lack of generalization of the object detection method. If the size distributions of targets in the datasets vary greatly, it may lead to a decline in the performance of the detection method. And in the self-driving scenario, traffic signs are mostly small objects, occupying fewer image pixels, yet most algorithms have low performance in detecting small objects. Summary of the Invention

[0004] In order to reduce the problems of poor generalization and low small object detection performance caused by manually setting anchors, the present invention proposes a small object detection method for autonomous driving scenarios based on a single-stage detection method, providing a learnable anchor strategy that can meet high accuracy and real-time performance.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A small object detection method for autonomous driving scenarios specifically includes the following steps:

[0007] Step0: Initialization of the detection method: including the initialization of the network and the preprocessing of the image. Among them, a learnable strategy is adopted for the size and dimensions of the anchor, and it is randomly set in the initial stage, and the network is used to learn the size and dimensions of the anchor; in the image preprocessing part, during the training stage, it mainly includes randomly scaling, color jittering, and normalizing the image;

[0008] Step1: The input is image data. Feature extraction is performed on the preprocessed image to obtain feature maps C3, C4, and C5 after 8x, 16x, and 32x downsampling respectively.

[0009] Step2: An improved feature fusion module is adopted to perform bidirectional feature fusion on the features of C3, C4, and C5 to obtain F3, F4, and F5. Among them, two downsampling operations are performed on the F5 feature map to further extract image features to obtain F6 and F7.

[0010] Step3: Through a shared decoupled prediction branch, predictions are made on the F3, F4, F5, F6, and F7 feature maps to obtain predictions of the target's class and location information.

[0011] Step4: Positive and negative sample assignment: By calculating the intersection over union (IoU) between the prediction and the anchor, when the IoU is greater than 0.5, the prediction is a positive sample; when the IoU is less than 0.4, the prediction is a negative sample; when the IoU is in the middle value, the prediction is an ignored sample and the loss is not calculated.

[0012] Step5: Calculate the loss: By calculating the loss between the predicted target and the ground truth, the loss includes the class loss and the coordinate loss. Among them, the class loss is calculated using Focal loss, and the calculation formula is shown in (1-1), and the coordinate loss uses Generalized IoU, and the calculation formula is shown in (1-2).

[0013]

[0014] Among them, α and γ are balance factors, p t is the class prediction probability, A C is the area of the rectangle enclosing the minimum bounding target box and the prediction box, and U is the overlapping area between the target box and the prediction box.

[0015] Step6: Update the network parameters: According to the error, through gradient backpropagation, learn and adjust the network weight parameters, including the size and dimensions of the anchor.

[0016] Furthermore, in Step0, the random scaling range of the image is 640, 672, 704, 736, 768, 800, and the maximum side length after scaling does not exceed 1333 pixels.

[0017] Furthermore, in Step1, ResNet50 is used as the backbone network to perform feature extraction on the preprocessed image.

[0018] Further, the improved feature fusion module in Step 2 includes a dimensionality reduction layer, a dimensionality increase layer, and upsampling and downsampling modules; the dimensionality increase layer and the dimensionality reduction layer use 1×1 convolutions to change the size of the feature channel dimension, the upsampling and downsampling modules use residual structures, the upsampling module uses a transposed convolution with a stride of 2 to expand the feature map, and the downsampling module mainly uses a 3×3 convolution to scale the feature map by 2 times.

[0019] Further, the decoupled class and location information prediction branches in Step 3 respectively use three convolutional layers in series.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0021] The present invention innovatively adopts an improved structural design. Through the bidirectional feature fusion module, the high-level semantic information and the low-level feature information are better fused. At the same time, a learnable anchor strategy is adopted. Through the randomly initialized anchor size and scale, the anchor size and scale are learned according to the error loss during the training phase. Compared with the traditional manually designed anchor hyperparameters, this method has better generalization. Experiments prove that while using fewer anchor numbers can achieve better performance, the performance index APs for small object detection is also improved. Fewer anchors also reduce the computational complexity of the algorithm and improve the detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0023] Figure 1 is the algorithm training flow chart of the present invention.

[0024] Figure 2 is the structural diagram of the feature fusion module of the present invention.

[0025] Figure 3 are the detection results of the baseline network and the algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific embodiments, structures, features, and their effects of the present invention as follows.

[0027] A small target detection method for autonomous driving scenarios proposed by the present invention is an improved algorithm based on the learnable anchor strategy (hereinafter referred to as the LAtinaNet algorithm), which is a single-stage detection method. After feature extraction and feature fusion of the input image, the class and location information of the target are obtained by a decoupled prediction branch. Initialization processing is performed before image processing, the size and dimensions of the anchor are learned through the network, and during the training stage, the error loss between the predicted target and the true target is calculated, and the anchor is adjusted through gradient backpropagation. The specific detection method is as follows Figure 1 shown, including the following steps:

[0028] Step0: Initialization of the detection method. It includes network parameter initialization and image preprocessing. Among them, a learnable strategy is adopted for the size and dimensions of the anchor, and it is randomly set in the initial stage, and the size and dimensions of the anchor are learned through the network. In the image preprocessing part, during the training stage, it mainly includes randomly scaling the image, color jittering, and normalization processing. Among them, the random scaling of the image size has a scaling size range of (640, 672, 704, 736, 768, 800), and the maximum side length after scaling does not exceed 1333 pixels.

[0029] Step1: The input is image data. The present invention uses ResNet50 as the backbone network to extract features from the preprocessed image, and obtains feature maps C3, C4, and C5 after 8-fold, 16-fold, and 32-fold downsampling respectively.

[0030] Step2: An improved feature fusion module (as shown in Figure 2 shown) is used to perform bidirectional feature fusion on the features of C3, C4, and C5 to obtain F3, F4, and F5. Among them, the F5 feature map is further downsampled 2 times to further extract image features to obtain F6 and F7.

[0031] Step3: Through a shared decoupled prediction branch, predictions are made on the F3, F4, F5, F6, and F7 feature maps to obtain target class predictions and coordinate predictions. The decoupled class and coordinate prediction branches are each composed of 3 convolutional layers in series.

[0032] Step4: Positive and negative sample assignment. By calculating the intersection over union (IoU) between the prediction and the anchor, when the IoU is greater than 0.5, the prediction is a positive sample; when the IoU is less than 0.4, the prediction is a negative sample; when the IoU is in the middle value, the prediction is an ignored sample and the loss is not calculated.

[0033] Step 5: Calculate the loss. By calculating the loss between the predicted target and the ground truth, the loss includes the class loss and the coordinate loss. Among them, the class loss is calculated using Focal loss, and the calculation formula is shown in (1-1), and the coordinate loss uses Generalized IoU, and the calculation formula is shown in (1-2).

[0034]

[0035] Among them, α and γ are balance factors, and p t is the class prediction probability, A C is the area of the rectangle that encloses the target box and the predicted box, and U is the overlapping area between the target box and the predicted box.

[0036] Step 6: Update the network parameters. According to the error, through gradient backpropagation, learn and adjust the network weight parameters, including the size and dimensions of the anchor.

[0037] Step 7: After one round of parameter iteration, determine whether the preset number of training epochs is reached until the training ends. This method is trained for a total of 30 epochs.

[0038] Example: Detection experiment of a small target detection method (LAtinaNet) for autonomous driving scenarios of the present invention

[0039] 1) Comparison of detection results

[0040] This experiment was trained and tested on the TT100K traffic sign dataset. The dataset includes approximately 100,000 images in total, among which 10,592 images are annotated, and the annotated images include 232 traffic sign categories. As Figure 3 shown, the left column is the detection result of the baseline network, the middle column is the detection result after improving the network structure, and the right column is the LAtinaNet detection structure using the learnable anchor strategy. It can be seen that there are misdetection cases in the baseline network, the method after improving the network structure has no misdetection but there are missed detections, while LAtinaNet can correctly detect the traffic signs appearing in the images without false alarms and missed detections.

[0041] At the same time, it was tested on the VisDrone dataset to further verify the improvement of this method in small target detection performance. The specific detection results are shown in Table 1. Compared with the other five methods, LAtinaNet has the smallest number of network parameters while ensuring the detection performance, and also ensures the accuracy of small target detection.

[0042] Table 1 Performance comparison of five detection methods

[0043] Method AP <![CDATA[AP 50 > <![CDATA[AP 75 > APs <![CDATA[AP M > <![CDATA[AP L > Params(M) Cascade R-CNN 23.2 39.9 23.4 16.5 36.8 39.4 273.2 Libra RCNN 24.3 41.2 24.9 16.8 34.0 36.8 185.4 HawkNet 25.6 44.3 25.8 19.9 36.0 39.1 130.9 VFNet 25.9 42.1 27.0 16.8 37.3 41.4 296.2 LAtinaNet 26.6 46.1 24.3 18.7 34.9 38.3 45.9

[0044] Table 2 shows the ablation experiments with different anchor parameters. 'LA' represents the learnable anchor strategy, 'anchorscale' represents the number of anchor sizes adopted, and 'aspect ratio' represents the number of anchor aspect ratios adopted. As can be seen from Table 2, the method with 1 size and 1 aspect ratio (a total of 1 anchor) improves the APs (a performance metric for small object detection) by 2.1% compared to the traditional method with 3 sizes and 3 aspect ratios (a total of 9 anchors) while ensuring performance. It uses fewer anchors, reduces the computational load of the model, and speeds up the algorithm inference speed.

[0045] Table 2 Ablation experiments with different anchor parameters

[0046]

[0047]

[0048] As described above, it is only the preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A small target detection method for autonomous driving scenarios, characterized by: The specific steps include: Step 0: Initialization of the detection method: including network initialization and image preprocessing. A learnable strategy is used for the size and size of the anchor. In the initial stage, it is randomly set and the size and size of the anchor is learned through the network. In the image preprocessing part, during the training stage, it mainly includes random scaling, color jittering, and normalization of the image. Step 1: The input is image data, and the features are extracted from the preprocessed image to obtain the feature maps C3, C4 and C5 after 8, 16 and 32 times downsampling respectively; Step 2: Using the improved feature fusion module, the features of C3, C4 and C5 are bidirectionally fused to obtain F3, F4 and F5. The F5 feature map is downsampled twice to further extract image features to obtain F6 and F7. Step 3: Through the shared decoupled prediction branch, prediction is performed on the F3, F4, F5, F6 and F7 feature maps to obtain the category and location information prediction of the target; Step 4: Positive and negative sample allocation: By calculating the intersection over union (IoU) of the prediction and the anchor, when IoU is greater than 0.5, the prediction is a positive sample; when IoU is less than 0.4, the prediction is a negative sample; when IoU is an intermediate value, the prediction is to ignore the sample and no loss is calculated; Step 5: Calculate the loss: Calculate the loss of the predicted target and the true value. The loss includes category loss and coordinate loss. The category loss is calculated using Focal loss. The calculation formula is shown in (1-1). The coordinate loss is calculated using GeneralizedIoU. The calculation formula is shown in (1-2). Where α and γ are balance factors, p t is the category prediction probability, A C is the area of ​​the smallest rectangular box that encloses the target box and the prediction box, and U is the overlapping area of ​​the target box and the prediction box; Step 6: Update network parameters: Based on the error, learn and adjust the network weight parameters, including the size and size of the anchor, through gradient back propagation.

2. The small target detection method for autonomous driving scenarios according to claim 1, characterized in that: In the step Step 0, the image is randomly scaled to a size ranging from 640, 672, 704, 736, 768, and 800, and the maximum side length after scaling does not exceed 1333 pixels.

3. The small target detection method for autonomous driving scenarios according to claim 1, characterized in that: In Step 1, ResNet50 is used as the backbone network to extract features from the preprocessed image.

4. The small target detection method for autonomous driving scenarios according to claim 1, characterized in that: The improved feature fusion module of Step 2 includes a dimensionality reduction layer, a dimensionality increase layer, and up- and down-sampling modules; the dimensionality increase layer and the dimensionality reduction layer adopt 1×1 convolution to change the dimension size of the feature channel, the up- and down-sampling modules adopt a residual structure, the up-sampling module adopts a transposed convolution with a step size of 2 to achieve the expansion of the feature map, and the down-sampling module mainly uses 3×3 convolution to achieve a 2-fold scaling of the feature map.

5. The small target detection method for autonomous driving scenarios according to claim 1, characterized in that: The Step 3 decouples the category and location information prediction branches by connecting three convolutional layers in series.

Citation Information

Patent Citations

  • Small target detection method based on regional nomination

    CN108830280A

  • An automatic driving scene key target detection and extraction method based on deep learning

    CN109784190A