An Improved Lightweight SCB-YOLOv5 Traffic Sign Detection Method and System Based on ShuffleNet

By using ShuffleNet v2 to replace YOLOv5's backbone network with ShuffleNet v2 in the traffic sign detection algorithm, adding a CA attention mechanism and designing BCS-FPN, the problem of traffic sign detection in complex road environments is solved, and efficient and accurate traffic sign detection is achieved.

CN118918567BActive Publication Date: 2025-06-17TANGSHAN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411163276.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-06-17
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

The existing traffic sign detection algorithm is difficult to effectively extract target features in complex road environments, and the models are complex and computationally large, making it difficult to deploy on automotive processors, especially when the target scales vary greatly.

Method used

ShuffleNet v2 is used to replace YOLOv5's backbone network to simplify the computational complexity and reduce the amount of parameters; add CA attention mechanism to the backbone network to improve the significance of targets; design BCS-FPN to fuse multi-scale features and improve the representation ability of small-scale objects.

Benefits of technology

Real-time accurate detection of traffic signs in complex road environments is achieved, which reduces the calculation complexity and parameter volume of the model, is suitable for deployment on automotive processors, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918567B_ABST
    Figure CN118918567B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic sign detection method and system based on an improved lightweight YOLOv5 of ShuffleNet, including: installing a camera in combination with the vehicle's perspective to collect road traffic sign images; after collecting the traffic sign images, performing annotation of various traffic sign targets to produce a traffic sign detection data set; constructing a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet, and training the detection model using the traffic sign detection data set; using the improved lightweight YOLOv5 detection model to detect road traffic signs on an in-vehicle processor, obtaining the types of road traffic signs ahead, and outputting the traffic sign types to subsequent functional modules. The present invention designs a traffic sign detection system to improve the efficiency and accuracy of traffic sign detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent driving perception systems, and particularly relates to an improved lightweight SCB-YOLOv5 traffic sign detection method and system based on ShuffleNet. Background Art

[0002] With the rapid development of autonomous driving technology, the intelligent perception technology of intelligent vehicles is also constantly updated and iterated. Among them, road traffic sign detection is a key task of the intelligent driving perception system. The effective recognition of road traffic signs is the basis of intelligent transportation systems and unmanned driving technologies, and also provides convenient conditions for the accuracy of subsequent unmanned intelligent decision-making.

[0003] In recent years, more and more traffic sign detection frameworks have used CNN, and many achievements have been made in object detection algorithms based on CNN. All along, object detection has been the most basic and challenging branch in computer vision. Object detection frameworks based on CNN are mainly divided into two categories. One is the two-stage object detection algorithm that pursues accuracy, and the other is the single-stage object detection algorithm that pursues speed. The difference between these two algorithms depends on whether the proposed region is further divided. Among them, the two-stage object detection algorithm will filter the proposed region and then match the prediction box. The two-stage object detection algorithms mainly include the R-CNN series, Mask-RCNN, and Cascade-RCNN. The single-stage object detection algorithms mainly include the YOLO series and SSD.

[0004] The network frameworks of existing algorithms are complex, the calculations are complex, and the running memory occupied by the model during deployment is also very large, which requires the device to have strong computing power support. Generally speaking, the computing power of the processors in cars is often poor and the running memory is relatively small. Therefore, the above algorithms are not suitable for direct application to road traffic detection. In addition, in complex road environments, the above algorithms cannot effectively extract target features, the detection effect on traffic signs with large scale differences is not good, and the detection accuracy is not very high. Therefore, it is necessary to develop a lightweight and efficient road traffic sign detection method. Summary of the Invention

[0005] Aiming at the problems of large differences in the target scales of traffic signs, complex detection models with large computational amounts, and model deployment being restricted by devices in complex road environments, the present invention proposes an improved lightweight SCB-YOLOv5 traffic sign detection method and system based on ShuffleNet. First, ShuffleNet v2 is used to replace the backbone network of YOLOv5, simplifying the computational complexity and reducing the parameters of the backbone network. Secondly, aiming at the problem of unclear traffic sign features in complex road environments, a CA attention mechanism is adopted to improve the saliency of the target. Finally, aiming at the problems of large differences in the target scales of traffic signs and a large proportion of small-scale targets, BCS-FPN is designed to fuse multi-scale features and improve the representation ability of small-scale objects. The present invention improves the efficiency and accuracy of traffic sign detection.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] An improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet, comprising the following steps:

[0008] Step 1, install a camera in combination with the perspective of the vehicle to collect road traffic sign images;

[0009] Step 2, after collecting the traffic sign images, perform annotation of various traffic sign targets to produce a traffic sign detection data set;

[0010] Step 3, construct a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet, and use the traffic sign detection data set to train the detection model;

[0011] Step 4, use the improved lightweight SCB-YOLOv5 detection model to detect road traffic signs on the vehicle-mounted processor, obtain the types of road traffic signs ahead, and output the traffic sign types to subsequent functional modules.

[0012] Preferably, in Step 1, install the camera according to the structure and perspective inside the vehicle cab. The installation of the camera should not affect the driver's vision and should completely capture the road view ahead. Then drive the vehicle to collect road traffic sign images at different times and road surface conditions.

[0013] Preferably, in Step 3, the structure improvement method of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the backbone network of YOLOv5, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN.

[0014] Preferably, the implementation method of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolutional unit module; designing the C2f-SCConv structure, where SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. The SRU uses a separation-reconstruction method to suppress spatial redundancy, and the CRU uses a split-transform-fusion strategy to reduce channel redundancy.

[0015] The present invention also provides an improved lightweight SCB-YOLOv5 traffic sign detection system based on ShuffleNet, including: an acquisition module, an annotation module, a construction module, and a detection module;

[0016] The acquisition module is used to install a camera in combination with the perspective of the vehicle to collect road traffic sign images;

[0017] The annotation module is used to perform annotation of various traffic sign targets after collecting traffic sign images to produce a traffic sign detection data set;

[0018] The construction module is used to construct a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet and train the detection model using the traffic sign detection data set;

[0019] The detection module is used to detect road traffic signs on the vehicle-mounted processor using the improved lightweight SCB-YOLOv5 detection model, obtain the types of road traffic signs ahead, and output the traffic sign types to subsequent functional modules.

[0020] Preferably, in the acquisition module, the camera is installed according to the structure and perspective inside the vehicle cab. The installation of the camera should not affect the driver's vision and should completely capture the road view ahead. Then, the vehicle is driven to collect road traffic sign images at different times and road conditions.

[0021] Preferably, in the construction module, the structural improvement process of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the backbone network of YOLOv5, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN.

[0022] Preferably, the implementation process of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolutional unit module; designing the C2f-SCConv structure, where SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. SRU uses the separation-reconstruction method to suppress spatial redundancy, and CRU uses the split-transform-fusion strategy to reduce channel redundancy.

[0023] Due to the adoption of the above technical solutions, the technical progress achieved by the present invention is as follows:

[0024] 1. To solve the problems of complex model, large number of parameters, and model deployment restricted by devices, ShuffleNetv2 is used to replace the YOLOv5 backbone network for feature extraction, and SimSPPF is used to replace the SPPF structure, which greatly reduces the number of network parameters and calculations.

[0025] 2. To solve the problem of difficult to effectively extract target features in complex road environments, a lightweight CA attention module is added to the backbone network to increase the smaller computational cost and improve the saliency of target detection.

[0026] 3. To address the problem of large differences in target scales, BCS-FPN is designed to replace the FPN+PAN structure of YOLOv5, which improves the feature fusion ability of multi-scale targets and further reduces the number of network parameters while ensuring accuracy.

[0027] 4. The present invention can achieve real-time and accurate detection of traffic signs in front of the road and input the detection results into subsequent functional modules. Description of the Drawings

[0028] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is the overall flowchart of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0030] Figure 2 It is the SCB-YOLOv5 algorithm structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0031] Figure 3It is the ShuffleNet v2 network structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0032] Figure 4 It is the CA attention mechanism structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0033] Figure 5 It is the SimSPPF structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0034] Figure 6 It is the BCS-FPN structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0035] Figure 7 It is the SCConv structure diagram of the improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet in the embodiments of the present invention;

[0036] Figure 8 It is the scaling statistical chart of the TT-100K dataset in the embodiments of the present invention;

[0037] Figure 9 It is the schematic diagram of the mAP50 curve of SCB-YOLOv5 and YOLOv5s with the number of epochs in the embodiments of the present invention. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0040] The traffic sign detection method is a system formed by the operation of a traffic sign detection algorithm on an in-vehicle processor, which can realize real-time traffic sign detection for the front road image input by the camera.

[0041] Embodiment 1

[0042] As Figure 1As shown in the figure, an improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet includes the following steps:

[0043] Step 1: Install a camera in combination with the vehicle's perspective to collect road traffic sign images.

[0044] Specifically: In Step 1, install the camera according to the structure and perspective inside the vehicle cab. The installation of the camera should not affect the driver's vision and can completely capture the front road vision. Then drive the vehicle to collect road traffic sign images at different times and road conditions.

[0045] Step 2: After collecting the traffic sign images, perform the annotation of various traffic sign targets to produce a traffic sign detection dataset.

[0046] Specifically: In Step 2, after collecting the traffic sign images, perform the annotation of various traffic sign targets to produce a traffic sign detection dataset.

[0047] Step 3: Build a traffic sign detection model based on the improved lightweight SCB-YOLOv5 of ShuffleNet, and use the traffic sign detection dataset to train the detection model.

[0048] Specifically: In Step 3, the structure improvement method of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the backbone network of YOLOv5, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN.

[0049] Furthermore, the implementation method of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolutional unit module; designing the C2f-SCConv structure, where SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. SRU uses the separation-reconstruction method to suppress spatial redundancy, and CRU uses the split-transform-fusion strategy to reduce channel redundancy.

[0050] Step 4: Use the improved lightweight SCB-YOLOv5 detection model to detect road traffic signs on the vehicle-mounted processor, obtain the types of the front road traffic signs, and output the traffic sign types to the subsequent functional modules.

[0051] Specifically: In step 4, an improved lightweight SCB-YOLOv5 detection model is used to detect road traffic signs on the vehicle-mounted processor, obtain the types of road traffic signs ahead, and output the traffic sign types to subsequent functional modules.

[0052] Furthermore, the traffic sign detection method is a system formed by the operation of a traffic sign detection algorithm on a vehicle-mounted processor, which can realize real-time traffic sign detection for the front road image input by the camera.

[0053] In this embodiment, a camera and a vehicle-mounted processor are installed inside the vehicle cab. The installation of the camera should not affect the driver's vision and can completely capture the front road vision. Then, the vehicle is driven to collect road traffic sign images at different times and road conditions. Frames are extracted from the collected videos to obtain pictures, and about 3000 pictures are selected as the dataset in total. The LabelImg annotation software is used for image annotation. When annotating, try to exactly frame the target with the annotation box. After annotating one picture, the LabelImg software will generate an ".xml" file, which contains information such as the annotated category, the size and coordinates of the annotation box. Since the training set of the SCB-YOLOV5 algorithm needs to use the VOC dataset format, the dataset also needs to be converted.

[0054] The network of the SCB-YOLOV5 algorithm needs to generate anchor boxes of different sizes to infer the size of the target box. The K-means clustering algorithm is adopted in the present invention.

[0055] The structural diagram of the SCB-YOLOV5 algorithm is as Figure 2 shown, and its process is as follows:

[0056] (1) First, the input picture is adaptively scaled to 640*640 size after Mosaic data augmentation, and then used as the input of the convolutional network according to the set minibatch.

[0057] (2) Feature extraction is carried out through the backbone network in sequence, feature fusion is carried out through the multi-scale feature fusion network, and classification prediction is carried out through the output layer to obtain the position, size and category of the target prediction box.

[0058] (3) The loss function is used to calculate the error between the prediction box and the labeled ground truth.

[0059] (4) The weight matrix and bias parameters in the network are iteratively updated through the gradient descent algorithm to reduce the error between the prediction box and the actual box.

[0060] (5) The weight matrix and bias parameters when the loss function takes the minimum value under the preset number of iterations are obtained.

[0061] (6) Use the obtained weight matrix and bias parameters as the network parameters in the detection stage to obtain the prediction information of the image to be detected.

[0062] The input end of the SCB-YOLOV5 algorithm adopts methods such as Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling.

[0063] The backbone network of the SCB-YOLOV5 algorithm uses ShuffleNet v2 for feature extraction. The network structure of ShuffleNet v2 is as Figure 3 shown. For Figure 3 the basic unit of the ShuffleNet v2 network shown in (a), the channels of the input feature matrix are divided into two branches. First, in accordance with the principle of simplifying network complexity, the degree of fragmentation is reduced in the subsequent network design. No operation is added to the left branch, and the number of input and output channels of the three Convs in the right branch is the same, which also conforms to the principle of equal input and output channel numbers. After Conv, these two branches are connected using Concat, which also makes the number of channels in the front and back of the entire unit consistent. Then channel shuffling is performed at the end of the unit. The addition operation is no longer performed in the entire basic unit, and the ReLU activation function and DW Conv only exist in one branch, minimizing element-wise operations as much as possible. For Figure 3 the downsampling unit of ShuffleNet v2 shown in (b), the operation of channel splitting is cancelled, and finally the channels of the output feature matrix are doubled after Concat. The 3×3 average pooling in one branch becomes 3×3 DW Conv, which can be regarded as a DW Conv with a weight of one-ninth, which can increase more possibilities, and finally a 1×1 convolution is added. In particular, in the two units of ShuffleNet v2, there is only a BN layer after DW Conv, offsetting the ReLU layer.

[0064] After the SCB-YOLOV5 algorithm extracts features from the backbone network, in order to enhance the saliency ability of feature expression, a CA attention mechanism is added behind the backbone network. The CA attention mechanism is a lightweight attention mechanism that can effectively improve the expression ability of the network to learn features. Its implementation process is as Figure 4 shown. The CA attention mechanism transfers the position information to channel attention and encodes the channel relationship and long-term dependence through precise position information. In order to capture attention and encode the position information in the width and height of the image, first divide the input feature map into two directions of width and height, and perform global average pooling respectively to obtain the feature maps in the width and height directions, as shown in the following formula:

[0065]

[0066] Among them, is the output of the c-th channel with height h, is the output of the c-th channel with width w. These two conversions aggregate features along two spatial directions respectively, thus generating a pair of feature maps C / r with direction perception ability. Then, the two feature maps are concatenated together and input into a shared convolution module to reduce it to the original dimension, as shown in the following formula:

[0067] f = δ(F1([z h , z w ))

[0068] where [·,·] is the Concat operation along the spatial dimension, δ is the non-linear activation function, and f is the intermediate feature map that encodes spatial information horizontally and vertically. Then f is decomposed into two independent tensors f h ∈R C / r×H , f w ∈R C / r×W . Then it is converted into the same tensors f h , f w by using another two 1×1 convolution transforms, and we get:

[0069] g h = σ(F h (f h )),

[0070] g w = σ(F w (f w ))

[0071] where σ is the sigmoid activation function. Finally, the weights g h in the height direction and the weights g w in the width direction of the obtained input feature map are jointly weighted on the original feature map to obtain a feature map with weights in the height direction and width direction, as shown in the following formula:

[0072]

[0073] The SCB-YOLOV5 algorithm uses the SimSPPF structure to extract features by using pooling kernels of different sizes at different scales, thus improving the performance of the detector. The specific structure of SimSPPF is as Figure 5 shown.

[0074] A lightweight cross-scale feature fusion mechanism BCS-FPN is designed in the SCB-YOLOv5 algorithm, as Figure 6 shown. In Figure 6Among them, C2f-SCConv is a designed lightweight module, where the SCCBL module is an effective convolutional module that uses SCConv as the basic unit of convolution. SCConv compresses CNN by leveraging the features between spatial and channel redundancies, reducing the representative features of redundant computations and making it easy to learn. The implementation process of BCS-FPN is as follows: First, connect the feature map output by the C1 layer to the prediction end of the F1 layer, and connect the feature map output by the C2 layer to the prediction end of the F2 layer. Add edges from the original input nodes to the output nodes, and as many target features as possible can be fused on the premise of increasing a small amount of computational effort. Secondly, replace the convolutional module with the SCCBL module to further reduce the computational effort while ensuring accuracy. Finally, design the C2f-SCConv structure for the feature fusion network to further reduce the number of parameters and improve the detection speed.

[0075] The SCConv structure in the SCB-YOLOv5 algorithm is as Figure 7 shown. SCConv consists of a spatial reconstruction unit (SRU) and a channel reconstruction unit (CRU). The SRU uses a separation-reconstruction method to suppress spatial redundancy, while the CRU uses a split-transform-fusion strategy to reduce channel redundancy. Specifically, for the intermediate input features in the bottleneck residual block, first obtain the spatially refined feature X through the SRU operation, and then obtain the channel-refined feature X w . By leveraging the spatial and channel redundancies between features in the SCConv module, it can be seamlessly integrated into any CNN architecture to reduce the redundancy between intermediate feature maps and improve the feature representation Y of CNN. The role of the SRU is to utilize the redundant features of space, as Figure 7 shown in the SRU structure in, and perform separation and reconstruction operations. The purpose of the separation operation is to have rich information content, with the feature graph and space separated. Then, use the scale factor in the normalization layer to evaluate the information content of different feature maps. The role of the CRU is to utilize the redundant channel features, as Figure 7 shown in the CRU structure, the split-transform-fusion strategy, which further reduces the spatially refined feature map redundant along the channel dimension. In addition, the CRU can extract rich representative features X w , and the redundant features can be processed through low-cost operations and lightweight convolutional operations.

[0076] Example Two

[0077] The lightweight traffic sign detection framework SCB-YOLOv5 based on YOLOv5 is a traffic sign detection model of an improved lightweight YOLOv5 based on ShuffleNet

[0078] The process of the detection framework in this example is as follows:

[0079] (1) Analyze the TT-100K dataset and divide it into a training set, a validation set, and a test set.

[0080] (2) By using the training set and the validation set divided by the method to train the network, the optimal weights of the model are obtained.

[0081] (3) Use the obtained model training weights to verify and test the test set.

[0082] In the experiment of the present invention, a comprehensive analysis of the SCB-YOLOv5 model was carried out using the TT-100K dataset. The experiment was carried out under the Pytourch framework. The operating system of the server is Linux Ubuntu 18.04, the CPU model is Intel(R) Xeon(R) Gold 6248R, the CPU frequency is 3.00GH, and the server memory is 128GB DDR4. The GPU is RTX A6000 with a memory of 48GB.

[0083] In this section, the present invention first analyzes the TT-100K dataset. The TT-100K dataset is a common traffic sign dataset jointly produced by Tsinghua University and Tencent. Among them, the training set contains 6,105 images, the test set contains 3,071 images, and the dataset contains a total of 232 traffic signs. The present invention first calculates the number of targets and the scale of the targets in the TT-100K dataset, and the statistical results are as Figure 8 shown. Then, the present invention classifies these objects into three categories according to their pixel sizes. Small-scale objects, whose pixel size is less than 32×32. Medium-scale objects, whose pixel size is between 32×32 and 96×96. Pixel size greater than 96×96.

[0084] Since the number of datasets in some quantity categories is small, it is easy to cause underfitting during the training of the network. Therefore, in order to ensure the effectiveness of the model, the present invention only selects a target quantity of more than 100 categories to continue training. Finally, the categories and their numbers of the filtered dataset are shown in Table 1, the objects and quantities after screening the TT-100-K dataset.

[0085] Table 1

[0086]

[0087] Before training the model, the ratio of the training set to the validation set is divided into 7:3. The image is input into the model for preprocessing, and the image size is adjusted to 640x640. This training method uses SGD, the momentum parameter size is 0.937, the initial learning rate is 0.01, and 16 images are processed in each batch. According to these parameters, all models are trained for 300 epochs.

[0088] To objectively evaluate the advantages of the proposed SCB-YOLOv5 algorithm of the present invention, the present invention selects precision, recall rate, mean average precision mAP@50, and average inference time as evaluation indicators, and the calculation formulas are as follows.

[0089]

[0090] Among them, TP is the number of correctly detected objects, FP is the number of falsely detected objects, FN is the number of missed detected objects, AP is the integral of the precision-recall curve, m is the number of detection categories, and the average inference time of the present invention is the average time calculated after selecting 500 images for testing.

[0091] The SCB-YOLOv5 traffic signal detection algorithm proposed by the present invention is compared with TT-1OLOv3, FCOS, YOLOv3, YOLOv3, YOLOv5-GhostNet, YOLOv5-MobileNet, and YOLOv5s. The experimental results are shown in Table 2. It can be seen from Table 2 that the proposed SCB-YOLOv5 algorithm has little difference from the YOLOv5 algorithm in terms of accuracy and recall rate. However, the mAP@50 of SCB-YOLOv5 is 0.2% lower than that of YOLOv5, and the inference speed is 20.8% faster than that of YOLOv5. Compared with other algorithms, SCB-YOLOv5 has obvious advantages in terms of speed and accuracy.

[0092] Table 3 shows the comparison of SCB-YOLOv5 with other algorithms in terms of the number of parameters, computational complexity, and model size. Among them, SCB-YOLOv5 has the smallest model size, the fewest number of parameters, and the least computational complexity. Compared with YOLOv5s, the number of parameters of SCB-YOLOv5 is reduced by 50.8%, the computational complexity is reduced by 59.8%, and the model size is reduced by 48.8%.

[0093] Table 2

[0094]

[0095] It can be seen from Table 2 and Table 3 that SCB-YOLOv5 has obvious advantages over other mainstream detection algorithms in terms of the number of model parameters, computational accuracy, accuracy, and speed indicators.

[0096] During the training process, the mAP50 curves of SCB-YOLOv5 and YOLOv5s with the number of epochs are as Figure 9 shown. It can be seen from Figure 9 this that the training convergence speed of SCB-YOLOv5 is relatively fast.

[0097] Table 3

[0098]

[0099] Finally, to more intuitively observe the comparison of the detection effects of SCB-YOLOv5 and YOLOv5, Figure 9 The comparison of the detection effects of SCB-YOLOv5 and YOLOv5 is shown.

[0100] The present invention proposes an improved lightweight traffic sign detection framework based on YOLOv5. First, ShuffleNet v2 is used to replace the backbone network of YOLOv5, simplifying the computational complexity and reducing the parameters of the backbone network. Second, aiming at the problem that traffic sign features are not obvious in complex road environments, the present invention adopts the CA attention mechanism to improve the saliency of targets. Finally, aiming at the problems of large differences in traffic sign target scales and a large proportion of small-scale targets, the present invention designs BCS-FPN to fuse multi-scale features and improve the representation ability of small-scale objects. The TT-100K dataset is analyzed and sorted out. The improved YOLOv5 is tested on the TT-100K dataset sorted out in the present invention. The results show that compared with YOLOv5s, the mAP of the algorithm of the present invention is comparable to that of YOLOv5s, and the speed is increased by 20.8%. The present invention also conducts experiments on embedded devices, and the experimental results show that the algorithm of the present invention has better effects on embedded devices with poor computing capabilities.

[0101] Embodiment III

[0102] The present invention also provides an improved lightweight SCB-YOLOv5 traffic sign detection system based on ShuffleNet, including: a collection module, an annotation module, a construction module, and a detection module;

[0103] The collection module is used to install a camera in combination with the perspective of the vehicle to collect road traffic sign images;

[0104] The annotation module is used to perform annotation of various traffic sign targets after collecting traffic sign images and make a traffic sign detection dataset;

[0105] The construction module is used to construct a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet and train the detection model using the traffic sign detection dataset;

[0106] The detection module is used to detect road traffic signs on the vehicle-mounted processor using the improved lightweight SCB-YOLOv5 detection model, obtain the types of road traffic signs ahead, and output the traffic sign types to subsequent functional modules.

[0107] In this embodiment, in the acquisition module, cameras are installed according to the structure and perspective inside the vehicle cab. The installation of the cameras should not affect the driver's vision and should capture the entire front road vision. Then, the vehicle is driven at different times and on different road conditions to collect road traffic sign images.

[0108] In this embodiment, in the construction module, the structural improvement process of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the backbone network of YOLOv5, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN.

[0109] In this embodiment, the implementation process of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolutional unit module; designing the C2f-SCConv structure, where SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. SRU uses a separation-reconstruction method to suppress spatial redundancy, and CRU uses a split-transform-fusion strategy to reduce channel redundancy.

[0110] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet, characterized in that: The following steps are involved: Step 1: Install a camera based on the vehicle's viewing angle to collect road traffic sign images; Step 2: After collecting the traffic sign images, various traffic sign targets are labeled to create a traffic sign detection dataset; Step 3, build a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet, and use the traffic sign detection dataset to train the detection model; Step 4: Use the improved lightweight SCB-YOLOv5 detection model to detect the road traffic sign on the vehicle processor, obtain the type of the road traffic sign ahead, and output the traffic sign type to the subsequent functional module; In step 3, the structural improvement method of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the YOLOv5 backbone network, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN; The implementation method of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolution unit module; designing the C2f-SCConv structure, in which SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. SRU adopts a separation-reconstruction method to suppress spatial redundancy, and CRU adopts a split transform-fusion strategy to reduce channel redundancy.

2. The improved lightweight SCB-YOLOv5 traffic sign detection method based on ShuffleNet according to claim 1, characterized in that: In step 1, the camera is installed according to the structure and viewing angle of the interior of the vehicle cab. The camera should not affect the driver's field of view and should completely capture the road ahead. Then the vehicle is driven to collect road traffic sign images at different time periods and road conditions.

3. An improved lightweight SCB-YOLOv5 traffic sign detection system based on ShuffleNet, characterized in that: include: Acquisition module, annotation module, construction module and detection module; The acquisition module is used to install a camera in combination with the vehicle's viewing angle to acquire road traffic sign images; The labeling module is used to label various traffic sign targets after collecting traffic sign images and to produce a traffic sign detection data set; The building module is used to build a traffic sign detection model based on the improved lightweight YOLOv5 of ShuffleNet, and train the detection model using the traffic sign detection dataset; The detection module is used to detect road traffic signs on the vehicle processor using an improved lightweight SCB-YOLOv5 detection model, obtain the type of the road traffic sign ahead, and output the traffic sign type to a subsequent functional module; In the building module, the structural improvement process of the improved lightweight SCB-YOLOv5 traffic sign detection model based on ShuffleNet includes: using the ShuffleNet v2 network to replace the YOLOv5 backbone network, adding a CA attention mechanism after the ShuffleNet v2 network, using SimSPPF to replace the SPP structure, and designing a lightweight cross-scale feature fusion mechanism BCS-FPN; The implementation process of BCS-FPN includes: connecting the feature map output by the C1 layer to the prediction end of the F1 layer, and connecting the feature map output by the C2 layer to the prediction end of the F2 layer; using the SCCBL module as the basic convolution unit module; designing the C2f-SCConv structure, in which SCConv consists of a spatial reconstruction unit SRU and a channel reconstruction unit CRU. SRU uses a separation-reconstruction method to suppress spatial redundancy, and CRU uses a split transform-fusion strategy to reduce channel redundancy.

4. The improved lightweight SCB-YOLOv5 traffic sign detection system based on ShuffleNet according to claim 3, characterized in that: In the acquisition module, the camera is installed according to the structure and viewing angle of the vehicle's cab. The installation of the camera should not affect the driver's field of view and should completely capture the road ahead. Then the vehicle is driven to collect road traffic sign images at different time periods and road conditions.

Citation Information

Patent Citations

  • Lightweight multi-target detection method based on improved YOLOv4

    CN117409355A

  • Power transmission line tower bird nest lightweight detection method fusing ShuffleNet and CBAM

    CN117611970A