A Small Target Enhancement and Optimization Method for Traffic Sign Detection
By adopting a priority-based small-objective enhancement strategy and optimal anchor box width and height clustering method in traffic sign detection, the shortcomings of deep learning models in small-objective detection and data set distribution imbalance are solved, and the detection accuracy is significantly improved.
Patent Information
- Application Number
- CN202111505215.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The deep learning-based object detection algorithm lacks detection capabilities in traffic sign detection, especially the detection accuracy of small targets, and the uneven distribution of the data set leads to poor model performance.
Data augmentation strategy based on priority is adopted for small-objective enhancement strategy, and the training data is optimized through optimal anchor box width and height clustering to generate more reasonable anchor box initial values to improve the detection accuracy of the model.
The detection accuracy and overall detection accuracy of small objects are improved, and the model's detection ability of small objects is improved, especially when the data set is unevenly distributed.
Smart Images

Figure CN114187576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and traffic detection, and particularly relates to a small target enhancement and optimization method for traffic sign detection. Background Art
[0002] The traffic sign recognition system (TSR) is an important part of the intelligent transportation system and the advanced driver assistance system. Improving the accuracy of traffic sign detection and recognition algorithms is a key issue to be solved in the process of moving towards practical applications. The accuracy of the algorithm is a very important factor in traffic sign recognition research. Incorrect recognition results not only cannot play an auxiliary driving role, but also lead to serious safety accidents. In the face of the increasing number of automobiles and the high incidence of traffic safety accidents, and the realistic pressure of continuously improving the driving intelligence of automobiles, carrying out research on traffic sign detection and recognition technology is of great significance for increasing driving safety.
[0003] In recent years, with the booming development of deep learning-based object detection algorithms, deep learning-based object detection algorithms have been widely applied in the field of traffic sign detection. Traffic sign detection requires detecting the area and type of traffic signs from a photo taken in a natural environment. The difference between traffic sign detection and ordinary object detection tasks is that the absolute size of traffic signs occupies a very small area in the picture. Usually, the original picture has a relatively high resolution, while the resolution of the input picture of the deep learning-based object detection algorithm model is lower than that of the original picture. The original picture needs to be reduced when input into the algorithm model, resulting in a further reduction in the size of traffic signs in the input picture of the algorithm model. In addition, there is also a problem of uneven distribution in the datasets of different types of traffic signs. These problems lead to insufficient detection ability of the deep learning-based object detection model for traffic sign targets. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a small target enhancement and optimization method for traffic sign detection.
[0005] The present invention performs two parts of optimization:
[0006] (1) Adopt a strategy of small target enhancement based on priority for data enhancement
[0007] Aiming at the problem that the previous small target enhancement ignored the distribution differences of various types of targets and unified enhancement resulted in poor effects, a strategy of small target enhancement based on priority is adopted for data enhancement, which improves the detection accuracy of the model to a certain extent.
[0008] (2) Adopt optimal anchor box width-height clustering to optimize training data
[0009] Regarding the problem that only large samples are focused on when obtaining positive samples for the model, while small targets are ignored, and the problem that simply clustering based on the width and height of the target to obtain the initial values of the anchor boxes results in unreasonable training data. The optimal anchor box width and height clustering is used to optimize the training data, thereby ultimately improving the detection accuracy of the model.
[0010] The specific steps of the method of the present invention are as follows:
[0011] Step 1: Determine the types of small targets that need to be enhanced using the enhancement priority. According to the distribution of each type in the dataset, the size distribution of small targets, and the distribution of feature diversity, and comprehensively obtaining the enhancement priority index based on the detection accuracy of the benchmark detection model, determine that the type of small target with the highest enhancement priority is the type of small target that needs to be enhanced for small targets. Such small targets are called biased small targets.
[0012] Step 2: Extract masked enhanced small target samples. The following operations are performed on the original dataset: For the pictures containing biased small targets, copy the biased small targets and fuse them in random areas of the original pictures using masks; for the original pictures without biased small targets, perform unified oversampling. The newly generated picture dataset is called the small target enhanced dataset. The original data and the small target enhanced dataset together are used as the dataset for model training, called the enhanced dataset.
[0013] Step 3: Generate the initial values of the anchor boxes by optimal anchor box width and height clustering. The following operations are performed on the enhanced dataset: First, abandon the original method of simply using the width and height of the target as the width and height of the anchor box, but calculate to obtain the approximate optimal anchor box width and height for each target. Then, perform K-nearest neighbor clustering on the optimal anchor box width and height of all targets to obtain the initial values of the model anchor boxes, thereby constraining the range of target sizes that the model focuses on when obtaining positive samples.
[0014] Step 4: Optimize the training data generation strategy for model training and detection. Improve the training data generation strategy of the algorithm model to use the clustered width and height as the final anchor box width and height. Input the enhanced dataset into the improved model for training and detection, and compare the detection accuracy of the benchmark model and the improved model of this patent on the validation set.
[0015] The beneficial effects of the present invention: In view of the problem of insufficient small target detection ability of the benchmark model, the present invention proposes a small target enhancement and optimization method for traffic sign detection. This method adopts a strategy of small target enhancement based on priority for data enhancement, and uses optimal anchor box width and height clustering to cooperate with the enhanced dataset to optimize the training data, thereby ultimately improving the small target detection accuracy and the overall detection accuracy. Description of the Drawings
[0016] Figure 1 It is a flowchart of the biased small target enhancement and optimization method for traffic sign detection.
[0017] Figure 2 It is a comparison chart of the detection accuracy of the benchmark detection model for various types and sizes of targets.
[0018] Figure 3 It is a distribution chart of various small targets in the dataset.
[0019] Figure 4 It is a schematic diagram of the width and height of the optimal anchor box.
[0020] Figure 5 It is a comparison of the detection accuracy of each model for various types.
[0021] Figure 6 It is a comparison of the detection accuracy of each model for various sizes. Specific implementation manner
[0022] The present invention will be further described below with reference to the accompanying drawings and in combination with the preferred implementation manners.
[0023] The present invention uses the GTSDB dataset for traffic sign detection research, uses the SSD model as the benchmark model, and defines the targets with a size less than 64*64 in the original image as small targets for research. The GTSDB dataset is divided into four major categories, namely prohibited signs, warning signs, mandatory signs, and others. Among them, the first three categories of targets have unified and obvious features and sufficient target numbers, while the fourth category of targets has diverse features, insufficient target numbers, and uneven distribution. The SSD model network is a model of the One-Stage architecture, without a region extraction stage, and uses preset anchor boxes as the initial values for regression. Taking the GTSDB dataset and the SSD object detection model as an example: the input image size of SSD is 300*300, while the original image size of the GTSDB dataset is 1360*800. The original data needs to be reduced to be input into the model, resulting in the reduction of the original targets. The size of the reduced targets is much smaller than the target size that can be detected in the original SSD model, and the original model cannot be well applied to the detection of small targets in high-resolution images.
[0024] As Figure 1 shown, the specific steps of this embodiment are as follows:
[0025] Step 1: In this step, the augmented priority is comprehensively obtained according to the type distribution score S_dist, the detection accuracy score S_prec, the size distribution score S_size, and the intra-class difference score S_diff. The target type with the highest augmented priority is the target to be enhanced.
[0026] The first step: Analyze the distribution ratio of small targets of each type in the dataset. Each type of small target obtains a ratio of the number of all small targets, and this is used as the type distribution score. See Figure 3 .
[0027] Step 2: Train a baseline model using the original dataset. The AP value of the baseline model for detecting each type of small target is used as the detection accuracy score, as shown in Figure 2 .
[0028] Step 3: Calculate the average width and height of each type of small target and the average width and height of all small targets. Calculate the Euclidean distance between the average width and height of each type of small target and the average width and height of all small targets, and normalize this difference to obtain the size distribution score.
[0029] Step 4: For each type, first scale the target size to the average width and height. Then use clustering to calculate a clustering center for each type with the structural similarity measure (SSIM) as the distance function. Finally, calculate the mean of the structural similarity measures between the remaining targets and the clustering center. This is used as the intra-class difference score S_diff. The value of Augmented Priority is set as follows:
[0030] Augmented Priority = ((1 - S_dist) * (1 - S_prec) * S_size * (1 - S_diff))
[0031] Step 2: First, obtain the mask of the small targets of the fourth type based on the rectangular boxes cropped from all the small targets of the fourth type in the image. The mask is required to be segmented along the edge of the small target, with the target area being pure white and the non-target area being pure black.
[0032] Then, use the mask of the small targets of the fourth type to fuse the small targets of the fourth type into the original image. The fusion position is random but does not touch the original targets.
[0033] Let the initial point in the original image closer to the fusion position of the small target be (x 0 , y 0 ). The pixel value at position (x, y) in the original image is f(x, y). The pixel value at position (i, j) closer to the rectangular box of the small target is r(i, j), and the pixel value at position (i, j) in the mask of the small target is m(i, j). The size of the original image is (w 0 , h 0 ), and the size of the rectangular box closer to the small target is (w, h). The pixel value at position (x, y) in the enhanced image of the small target closer to it is g(x, y).
[0034]
[0035] Finally, oversample one copy of the original image without small targets of the fourth type. The small target enhanced dataset and the original dataset are used as the enhanced dataset for the final input model.
[0036] Step 3: The optimization of the method for the model to obtain positive samples is divided into the following steps:
[0037] S1: Optimal Anchor Box Width and Height Calculation
[0038] Let the width of the original image be W, the height be H, the width of the feature layer for detection be w, the height be h, and all the original target objects ui(x, y, w, h) biased towards small targets. First, transform the center point of each target into a region block Block with width W / w and height H / h to obtain the new target center point coordinates bi(x, y, w, h) for calculating the optimal anchor box width and height. Then, taking the center point C(W / 2w, H / 2h) of the region block Block as the center of the anchor box, calculate the IOU between the anchor box and the actual target at different widths and heights, and find the optimal anchor box width anchorW and height anchorH when the IOU is the largest. Obtain the optimal anchor box width and height for each target in the small-target enhanced dataset, as shown in Figure 4 .
[0039]
[0040]
[0041] S2: Width and Height Dimension Clustering Based on the Optimal Anchor Box Width and Height
[0042] Based on the optimal anchor box width and height calculated in S1, perform K-nearest neighbor clustering, and use the Euclidean distance as the distance function to cluster the center values of three width and height values. These three width and height values are used as the initial regression values of the width-to-height ratios of the priorbox layer of the SSD model.
[0043] Step 4: Train the improved SSD model applicable to traffic sign detection. Use the SSD model implemented by the caffe framework to optimize the anchor box generation strategy in the priorbox layer of the SSD model.
[0044] In the original SSD model, the initial values of the anchor boxes are calculated based on settings such as min_size, max_size, aspect_ratio, and flip. The randomness of generating the initial values of the anchor boxes in this way is relatively large, and the anchor boxes cannot be customized according to the characteristics of the targets in the dataset. For data with a relatively single aspect ratio of target length and width such as traffic signs, using anchor boxes with multiple width-to-height ratios does not achieve good results, and it will also increase the computational load of the model. Moreover, the original SSD model is for detecting relatively large targets and is not applicable to small-target traffic signs.
[0045] The optimized anchor box generation strategy in the priorbox layer generates the initial values of the anchor boxes only based on the width-to-height ratio, so that the clustered width-to-height ratio obtained in Step 3 can be used to initialize the anchor boxes.
[0046] The following are the detection effects of the present invention:
[0047] The experimental results are as shown in Figure 5 、 Figure 6As shown. From Figure 5 It can be seen that the final improved model has improved both the overall detection accuracy and the detection accuracy of various categories. And it can be seen that before optimizing the way to obtain positive samples of the model, the improvement effect is relatively small (an increase of 0.8%), while after optimization, the enhancement of small targets has a more obvious effect (an increase of 2.9%). From Figure 6 It can be seen that compared with the baseline model, the detection accuracy of small targets of the final improved model has been improved (an increase of 19%). At the same time, it can be seen that after optimizing the way to obtain positive samples of the model, although the detection accuracy of large targets has decreased slightly, the improvement effect of the detection accuracy of small targets has increased from 0 to 3.5%, indicating that the enhancement of small targets has an obvious effect at this time.
Claims
1. A small target enhancement and optimization method for traffic sign detection, characterized in that the method comprises the following steps: Step 1: Determine the type of small target to be enhanced by using the enhancement priority; Comprehensively obtain the enhancement priority index according to the distribution of each type in the data set, the size distribution of small targets, the distribution of feature diversity, and the detection accuracy of the benchmark detection model, and determine the type of small target with the highest enhancement priority as the type of small target to be enhanced, and call this type of small target the biased small target; specifically: The first step: Analyze the distribution ratio of small targets of each type in the data set, and each type of small target gets a ratio of the total number of all small targets, which is used as the type distribution score; The second step: Train a benchmark detection model with the original data set, and the ap value of the benchmark detection model for detecting each type of small target is used as the detection accuracy score; The third step: Calculate the average width and height of each type of small target and the average width and height of all small targets, calculate the Euclidean distance between the average width and height of each type of small target and the average width and height of all small targets, and normalize this difference, which is used as the size distribution score; The fourth step: For each type of small target, first scale the target size to the average width and height, and then use clustering, with the structural similarity measure SSIM as the distance function to calculate a clustering center for each type of small target; finally calculate the mean of the structural similarity measure between the remaining targets and the clustering center; this is used as the within-class difference score S_diff; Finally, the value of the augmented priority is set as follows: Augmented Priority = ((1 - S_dist) * (1 - S_prec) * S_size * (1 - S_diff)) where S_dist represents the type distribution score, S_prec represents the detection accuracy score, S_size represents the size distribution score, and S_diff represents the within-class difference score; Step 2: Extract masked enhanced small target samples; Perform the following operations on the original data set: For the pictures containing biased small targets, copy the biased small targets and fuse them in random areas of the original pictures using masks; perform unified oversampling on the original pictures that do not contain biased small targets; The newly generated picture data set is called the small target enhanced data set; The original data and the small target enhanced data set together are used as the data set for model training, called the enhanced data set; Step 3: Generate the initial value of the anchor box by optimal anchor box width and height clustering; Perform the following operations on the enhanced data set: First, abandon the original method of simply using the target width and height as the anchor box width and height, but calculate to obtain the approximate optimal anchor box width and height of each target; then perform K-nearest neighbor clustering on the optimal anchor box width and height of all targets to obtain the initial value of the model anchor box, so as to constrain the target size range that the model focuses on when obtaining positive samples; Step 4: Optimize the training data generation strategy and perform model training and detection; Improve the training data generation strategy of the SSD model for traffic sign detection, and use the clustered width and height as the final anchor box width and height; input the enhanced data set into the improved model for training and detection, and compare the detection accuracy of the benchmark model and the improved model on the validation set.
Citation Information
Patent Citations
Method for detecting small target of high-resolution image of any scale
CN111222474A
Traffic sign detection and identification method based on YOLOv4 improvement
CN113239753A