A traffic sign detection method based on small target feature enhancement
By performing data augmentation, Anchor Box clustering optimization, feature enhancement and loss function design on the traffic sign detection algorithm, the problem of small targets and positive and negative samples is solved, and the accuracy and practicality of the detection are improved.
Patent Information
- Application Number
- CN202111428028.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-26
AI Technical Summary
The existing traffic sign detection algorithm has the problem of poor detection results when dealing with small targets and positive and negative samples imbalances, especially on the TT100K dataset, Yolov5's detection performance is poor.
By constructing traffic sign data sets and performing data augmentation, Anchor Box's clustering algorithm is optimized, fine-grained traffic signs and channel characteristics are enhanced, loss functions for positive and negative sample imbalances are designed, and the effect of the improved detection algorithm is evaluated.
The accuracy of traffic sign detection is improved. Compared with the original Yolov5 algorithm, the detection results are more accurate, which can better perceive the signs around the vehicle, reduce decision-making errors and reduce the frequency of accidents.
Smart Images

Figure CN114120280B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a traffic sign detection method based on small target feature enhancement. Background Art
[0002] With the rapid development of artificial intelligence technology, the intelligent driving industry has also developed rapidly. In particular, the development of deep learning has enabled vehicles to achieve great success in perception, positioning and other aspects.
[0003] Traffic signs contain important road traffic information and are an important part of vehicle environmental perception. The detection accuracy of traffic signs is an important measure of whether the algorithm can be actually deployed on vehicles.
[0004] Among the existing detection methods, they are mainly divided into two categories. The first category is the two-stage detection algorithm represented by Faster-Rcnn. This type of algorithm divides the entire detection process into two parts. The first step is to train the RPN network, and the second step is to train the network for target area detection. The second category is the one-stage detection algorithm represented by SSD and Yolo. This type of detection algorithm directly gives the category and location information through the backbone network. Although the two-stage algorithm has higher accuracy than the one-stage algorithm, its speed does not meet the real-time requirements. As the latest version of the Yolo series, Yolov5 has good performance in detection speed and accuracy. However, due to the large number of small targets in the traffic sign dataset and the serious imbalance of positive and negative samples, Yolov5's detection effect on the traffic sign dataset TT100K is poor. Summary of the invention
[0005] The purpose of the present invention is to remedy the defects of the prior art and provide a traffic sign detection method based on small target feature enhancement.
[0006] The present invention is achieved through the following technical solutions:
[0007] A traffic sign detection method based on small target feature enhancement specifically comprises the following steps:
[0008] Build a traffic sign dataset and perform data augmentation;
[0009] Optimize the clustering algorithm for building Anchor Box;
[0010] Optimize the network structure to enhance the fine-grained features of traffic signs;
[0011] Optimize the network structure to enhance the characteristics of traffic sign channels;
[0012] Design loss function to address the serious imbalance between positive and negative samples;
[0013] The effectiveness of the improved traffic sign detection algorithm is evaluated.
[0014] The specific contents of constructing a traffic sign dataset and performing data enhancement are as follows: selecting a public dataset TT100K as the research object, analyzing the dataset, including statistics on the target size in the dataset and the number of targets in each image, and obtaining the characteristics of the imbalance of positive and negative samples in the dataset; in view of the serious imbalance problem of positive and negative samples in the dataset obtained by analysis, the target replication method is used to perform data enhancement.
[0015] The method of using target replication to perform data enhancement is described as follows:
[0016] First, all objects in the dataset with a size less than 50*50 are cropped according to the label file of the dataset; secondly, the number of objects of each category is counted according to the category to obtain n 1 , n 2 , n 3 ,......,n 45 , that is, the number of each type of target, and the total number of targets cut out n; further calculation to obtain n i The difference between m and n i , that is, m i =nn i , for all m i Normalize, that is, calculate m = Σm i , p i =m i / m, each probability p i They all occupy an interval in (0, 1). When selecting a target for replication, a random number r between (0, 1) is selected. The probability interval in which r falls determines which type of target is selected for replication.
[0017] The optimized clustering algorithm for constructing the Anchor Box is specifically as follows: First, the Anchor Box in the Yolo algorithm can be understood as a multi-scale sliding window, that is, the shapes and sizes of the longest-appearing boxes found from all the ground truth in the training set. This prior knowledge is added to the model to constrain the shape and size of the predicted object to achieve the purpose of multi-scale learning. In the present invention, the selection of the Anchor Box is optimized for the characteristics that the target size in the data set is small and the size is relatively concentrated.
[0018] The K-means++ clustering algorithm is used to cluster the target sizes in the annotation file and obtain 9 anchor boxes. The clustering distance formula is:
[0019]
[0020] Among them, S center Represents the area of the current cluster center, S box Indicates the area of the current box to be classified.
[0021] The optimized network structure enhances the fine-grained features of traffic signs, specifically as follows: by introducing the BiFPN structure to replace the original PANet structure in the network structure, when fusing feature maps of different sizes, considering that the information contained in features of different sizes has different degrees of influence on the final result, the weighted fusion method is introduced to fuse the underlying feature maps with obvious fine-grained features.
[0022] The optimized network structure enhances the channel characteristics of traffic signs, specifically as follows: the SE structure is introduced to enable the network to better learn the correlation information between channels. First, the feature map input of the layer with a size of W*H*C is compressed, that is, a global average pooling is performed to obtain a 1*1*C vector, and then the vector is excited. First, a fully connected layer is connected to the obtained 1*1*C vector, and then an activation function layer, and then another fully connected layer to restore the number of input channels, and finally an activation function layer is superimposed, and finally a 1*1*C vector is output, representing the weight vector of each channel of this layer; finally, the input feature map is scaled, that is, the 1*1*C weight vector obtained in the previous step is multiplied by the channel weight of the input feature map to obtain the output of this layer.
[0023] The loss function is designed to address the serious imbalance between positive and negative samples. Specifically, the loss function is divided into positioning loss, confidence loss, and classification loss. Positioning loss measures the difference between the predicted box and the real box. Confidence loss measures the accuracy of judging whether the predicted box has an object. Classification loss measures whether the algorithm correctly classifies the object in the image. By weakening the weight of negative samples when calculating confidence loss, the influence of a large number of negative samples on the loss function is weakened, and the network is finally able to better learn the features of positive samples.
[0024] The calculation formula of the positioning loss is as follows:
[0025]
[0026] in A represents the real target box, B represents the predicted target box, Distance_2 represents the center point of the jth predicted box of the ith grid and the center point of the target real box in the predicted box, and Distance_c represents the diagonal length of the minimum bounding box formed by the two boxes. wgt represents the width of the target real box, hgt represents the height of the real box, wp represents the width of the predicted box, hp represents the height of the predicted box, S*S represents the number of grids of the predicted feature map, B represents the number of boxes predicted for each grid, and λ iou is the weight of the defined positioning loss in the entire loss function, It means that the current box is predicted as a positive sample.
[0027] The calculation formula of the classification loss is as follows:
[0028]
[0029] In the formula, classes represents 45 detection targets, λ c Represents the weight of classification loss in the entire loss function, and p i (c) represents the predicted probability and true probability that the target in the jth prediction box of the i-th grid belongs to the c-th class.
[0030] The calculation formula of the confidence loss is as follows:
[0031]
[0032] Among them C i is the confidence that there is a positive sample in the j-th prediction box of the i-th grid, is the confidence that there are actually positive samples in the prediction box, P Ci It is in C i The confidence interval to which this confidence value belongs is obtained by transforming the sample density in the previous batch through a function.
[0033] The improved traffic sign detection algorithm is evaluated as follows: In the detection algorithm, when judging the effect of the target detection model, four categories are mainly classified: true positive (TP), that is, the true label is a positive sample, and the predicted label is also a positive sample; false positive (FP), that is, the true label is a negative sample, and the predicted label is a positive sample; true negative (TN), that is, the true label is a negative sample, and the predicted label is a negative sample; false negative (FN), that is, the true label is a positive sample, and the predicted label is a negative sample.
[0034] Table 1 Classification model label table
[0035]
[0036] The calculation method of the model evaluation index precision is as follows:
[0037]
[0038] The calculation method of recall rate Recall is as follows:
[0039]
[0040] The advantages of the present invention are: the traffic sign detection method based on small target feature enhancement proposed by the present invention aims at the serious problem of positive and negative sample imbalance in the small target detection process. By improving the network structure, loss function and data amplification technology of Yolov5s, the final detection result is improved by 3% compared with the original Yolov5s algorithm, which can enable the vehicle to perceive the surrounding signs more accurately during driving, reduce decision-making errors, and reduce the frequency of accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is the overall flow chart of the traffic sign detection algorithm based on small target feature enhancement of the present invention;
[0042] Figure 2 It is a specific network structure diagram of a traffic sign detection algorithm based on small target feature enhancement of the present invention;
[0043] Figure 3 It is a schematic diagram of the result of the improved algorithm image detection of the present invention. DETAILED DESCRIPTION
[0044] To make the technical solutions and details in the embodiments of the present invention more clearly described, of course, the examples described are only part of the embodiments of the present invention, not all of the embodiments. The technical solutions of the present invention will be described in detail and completely in conjunction with the drawings in the embodiments of the present invention.
[0045] like Figure 1 As shown, the present invention proposes a traffic sign detection algorithm based on small target feature enhancement, comprising the following steps:
[0046] The present invention is implemented by the following technical solution, and the specific steps are as follows:
[0047] (1) Construct a traffic sign dataset and perform offline data augmentation. When constructing a traffic sign dataset, attention should be paid to the diversity of the dataset to enhance the generalization ability of the model.
[0048] (2) Optimize the clustering algorithm for constructing the Anchor Box. In order to obtain a more suitable Anchor Box, the selection of the initial cluster center and the calculation of the cluster distance are optimized to obtain a suitable Anchor Box.
[0049] (3) Optimize the network structure to enhance the fine-grained features of traffic signs. By analyzing the data set, it is found that small targets account for a large proportion. Coarse-grained features contain more detailed information, and high-level coarse-grained features contain more information about the location and contour of the target. Obviously, in the field of target detection, improving the detection effect of small targets depends more on fine-grained features. In order to obtain more fine-grained features, the present invention uses the idea of Bi-FPN to perform weighted fusion of feature maps during the fusion of the underlying feature map and the high-level feature map.
[0050] (4) Optimize the network structure to enhance the channel characteristics of traffic signs. Color is the basic attribute of traffic signs. After analysis, it is found that the colors of traffic signs in the dataset are mainly red, yellow and blue. In order to solve the problem that different channels of the feature map have different proportions in the convolution process, the channel attention mechanism SE (Squeeze-and-Excitation Module) is introduced.
[0051] (5) Designing a loss function for the serious imbalance between positive and negative samples. The loss function of Yolov5 mainly includes three parts, namely positioning loss, confidence loss and classification loss. Among them, in the confidence loss, due to the existence of a large number of negative samples, the cumulative confidence loss of all negative samples has a greater impact on the loss function. In view of this situation, the present invention adopts a dynamic weighting method in training to reduce the weight of negative samples and reduce the impact of negative samples on the loss function.
[0052] (6) Evaluate the effectiveness of the improved traffic sign detection algorithm and calculate the detection accuracy and recall after algorithm optimization.
[0053] The specific contents are as follows:
[0054] (1) Construct a traffic sign dataset and perform data enhancement:
[0055] The public dataset TT100k is used as the object of this study. By analyzing the dataset, including the size of the objects in the dataset and the number of objects in each image, the characteristics of the imbalance of positive and negative samples in the dataset are obtained. In view of the serious imbalance of positive and negative samples in the dataset obtained by analysis, the target replication data enhancement method is used. First, all objects with a size less than 80*80 in the dataset are cropped according to the label file of the dataset; secondly, the number of objects of each category is counted according to the category to obtain n 1 , n 2 , n 3 ,......,n 45 , that is, the number of each type of target, and the total number of targets cut out n; further calculation to obtain n iThe difference between m and n i , that is, m i =nn i , for all m i Normalize, that is, calculate m = Σm i , p i =m i / m, each probability p i They all occupy an interval in (0, 1). When selecting a target for replication, a random number r between (0, 1) is selected. The probability interval in which r falls determines which type of target is selected for replication.
[0056] (2) Optimize the clustering algorithm for building Anchor Box:
[0057] In order to make the clustered anchor box better fit the real box of the target, the K-means++ clustering algorithm is used to cluster the target sizes in the annotation file to obtain 9 anchor boxes. The clustering distance formula is:
[0058]
[0059] Among them, S center Represents the area of the current cluster center, S box Indicates the area of the current box to be classified.
[0060] (3) Optimizing the network structure to enhance the fine-grained features of traffic signs:
[0061] The Bi-FPN module is introduced. When fusing the three feature maps of different sizes of the backbone network, the weighted fusion method is introduced, taking into account the different degrees of influence of the information contained in the feature maps of different sizes on the final result.
[0062] (4) Optimize the network structure to enhance the characteristics of traffic sign channels:
[0063] In order to enable the network to learn the correlation between channels, the SE structure is introduced. First, the feature map input of the layer with a size of W*H*C is squeezed, that is, a global average pooling is performed to obtain a 1*1*C vector. Next, the vector is stimulated. First, an FC (fully connected) layer is connected after the obtained 1*1*C vector to reduce the number of channels and thus reduce the amount of calculation. Then an activation function layer is connected to increase nonlinearity, and then another FC layer is connected to restore the number of input channels. Finally, an activation function layer is superimposed, and finally a 1*1*C vector is output, representing the weight vector of each channel of this layer. Finally, the input feature map is scaled, that is, the 1*1*C weight vector obtained in the previous step is multiplied by the channel weight of the input feature map to obtain the output of this layer.
[0064] (5) Design loss function to address the serious imbalance between positive and negative samples:
[0065] The loss function mainly consists of three parts, including positioning loss, confidence loss and classification loss.
[0066] The calculation formula of positioning loss is as follows:
[0067]
[0068] in A represents the real target box, and B represents the predicted target box. w gt Represents the width of the target real box, h gt Represents the height of the real frame, w p Represents the width of the prediction box, h p Represents the height of the prediction box.
[0069] The calculation formula for classification loss is as follows:
[0070]
[0071] The calculation formula of confidence loss is as follows:
[0072]
[0073] Among them C i is the confidence of the Anchor box prediction, Is there really a target in the Anchor box? Ci It is in C i The confidence interval to which this confidence value belongs is obtained by transforming the sample density in the previous batch through a function.
[0074] (6) Evaluate the effectiveness of the improved traffic sign detection algorithm:
[0075] In the detection algorithm, the calculation method of the model evaluation index precision is as follows:
[0076]
[0077] Among them, as shown in Table 2, TP represents a true positive example, FP represents a false positive example, TN represents a true negative example, and FN represents a false negative example;
[0078] Table 2 Classification model label table
[0079]
[0080] The calculation method of recall rate Recall is as follows:
[0081]
[0082] Figure 2 This is the network structure diagram of the present invention: In order to better learn the fine-grained features and channel features of the image, the SE module is added to the original Yolov5s network structure and the idea of BiFPN is introduced to perform weighted fusion of feature maps of different sizes.
[0083] Figure 3 This is the result of image detection using the improved algorithm of the present invention. A mark of the category "w59" is detected in the image, and the target is framed with a box. At the same time, the probability value of the target being detected is displayed in the figure.
Claims
1. A traffic sign detection method based on small target feature enhancement, characterized in that: The specific steps include: Build a traffic sign dataset and perform data augmentation; Optimize the clustering algorithm for building Anchor Box; Optimize the network structure to enhance the fine-grained features of traffic signs; Optimize the network structure to enhance the characteristics of traffic sign channels; Design loss function to address the serious imbalance between positive and negative samples; Evaluate the effectiveness of the improved traffic sign detection algorithm; The construction of the traffic sign dataset and data enhancement are as follows: the public dataset TT100K is selected as the research object, and the characteristics of the imbalance of positive and negative samples in the dataset are obtained by analyzing the dataset, including the statistics of the target size in the dataset and the number of targets in each image; in view of the serious imbalance of positive and negative samples in the dataset obtained by analysis, the small target replication method is used to perform data enhancement; The method of using small target replication to perform data enhancement is described as follows: First, all objects in the dataset with a size less than 50*50 are cropped according to the label file of the dataset; secondly, the number of objects in each category is counted according to the category to obtain n1, n2, n3, ..., n 45 , that is, the number of each type of target, and the total number of targets cut out n; further calculation to obtain n i The difference between m and n i , that is, m i =nn i , for all m i Normalize, that is, calculate m = Σm i , p i =m i / m, each probability p i They all occupy an interval between (0, 1). When selecting a target for replication, a random number r between (0, 1) is selected. The probability interval in which r falls determines which type of target is selected for replication. The loss function is designed for the serious imbalance of positive and negative samples, as follows: the loss function is divided into positioning loss, confidence loss and classification loss; the positioning loss measures the difference between the predicted box and the real box; the confidence loss measures the accuracy of judging whether the predicted box has an object; the classification loss measures whether the algorithm correctly classifies the object in the image; The calculation formula of the positioning loss is as follows: in A represents the real target box, B represents the predicted target box, Distance_2 represents the center point of the jth predicted box of the ith grid and the center point of the target real box in the predicted box, and Distance_c represents the diagonal length of the minimum bounding box formed by the two boxes. w gt Represents the width of the target real box, h gt Represents the height of the real frame, w p Represents the width of the prediction box, h p represents the height of the prediction box, S*S represents the number of grids in the predicted feature map, B represents the number of Boxes predicted for each grid, and λ iou is the weight of the defined positioning loss in the entire loss function, It means that the current box is predicted to be a positive sample; The calculation formula of the classification loss is as follows: In the formula, classes represents 45 detection targets, λ c Represents the weight of classification loss in the entire loss function, and p i (c) represents the predicted probability and true probability that the target in the jth prediction box of the i-th grid belongs to the c-th class; The calculation formula of the confidence loss is as follows: Among them C i is the confidence that there is a positive sample in the j-th prediction box of the i-th grid, is the confidence that there are actually positive samples in the prediction box, It is in C i The confidence interval to which this confidence value belongs is obtained by transforming the sample density in the previous batch through a function.
2. The traffic sign detection method based on small target feature enhancement according to claim 1, characterized in that: The optimized clustering algorithm for constructing the Anchor Box is specifically as follows: the target sizes in the annotation file are clustered using the K-means++ clustering algorithm to obtain 9 Anchor boxes, and the clustering distance formula is: Among them, S center Represents the area of the current cluster center, S box Indicates the area of the current box to be classified.
3. The traffic sign detection method based on small target feature enhancement according to claim 1, characterized in that: The optimized network structure enhances the fine-grained features of traffic signs, specifically as follows: by introducing the BiFPN structure, when fusing feature maps of different sizes, considering that the information contained in feature maps of different sizes has different degrees of influence on the final detection results, a weighted fusion method is introduced to perform weighted fusion on feature maps of different sizes, aiming to include more fine-grained information in the finally calculated feature map.
4. The traffic sign detection method based on small target feature enhancement according to claim 1, characterized in that: The optimized network structure enhances the channel characteristics of traffic signs, specifically as follows: the SE structure is introduced, first, the feature map input with an input size of W*H*C is compressed, that is, a global average pooling is performed to obtain a 1*1*C vector, and then an excitation operation is performed on the vector, first a fully connected layer is connected to the obtained 1*1*C vector, then an activation function layer, and then another fully connected layer to restore the number of input channels, and finally an activation function layer is superimposed, and finally a 1*1*C vector is output, representing the weight vector of each channel of this layer; finally, a scale operation is performed on the input feature map, that is, the 1*1*C weight vector obtained in the previous step is multiplied by the channel weight of the input feature map to obtain the output of this layer.
5. The traffic sign detection method based on small target feature enhancement according to claim 1, characterized in that: The effect evaluation of the improved traffic sign detection algorithm is as follows: In the detection algorithm, when judging the effect of the target detection model, four categories are classified: true positive examples TP, that is, the true label is a positive sample, and the predicted label is also a positive sample; false positive examples FP, that is, the true label is a negative sample, and the predicted label is a positive sample; true negative examples TN, that is, the true label is a negative sample, and the predicted label is a negative sample; false negative examples FN, that is, the true label is a positive sample, and the predicted label is a negative sample; The calculation method of the model evaluation index precision is as follows: The calculation method of recall rate Recall is as follows:
Citation Information
Patent Citations
Remote traffic sign detection and recognition method based on F-RCNN
CN110163187A
Traffic sign board identification method based on multi-level fusion multi-scale prediction
CN110414417A