Traffic sign detection method, device and system, and storage medium

Through the improved YOLOv5s model, combined with lightweight design, feature prediction scale optimization, feature fusion enhancement and EIoU Loss loss function, the existing traffic sign detection algorithm has solved the problem of low accuracy and poor real-time performance when detecting small targets, and achieved higher detection accuracy and real-time performance.

CN120236266AActive Publication Date: 2025-07-01UBISOFT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510326310.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-01
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing traffic sign detection algorithm has low accuracy and poor real-time performance when detecting small targets, and is prone to missed and missed detection.

Method used

The improved YOLOv5s model is adopted to improve the detection accuracy and real-timeness of small-target traffic signs by lightweight design, feature prediction scale optimization, feature fusion enhancement, and the EIoU Loss loss function is used.

Benefits of technology

It effectively reduces the missed detection rate of traffic signs in small targets, improves the accuracy and real-timeness of detection, and enhances the model's perception of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236266A_ABST
    Figure CN120236266A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic sign detection method, device and system, and a storage medium. The method comprises the following steps: S1, obtaining a traffic sign data set; s2, training an improved YOLOv5s model according to the traffic sign data set; and S3, inputting a real-time video stream of vehicle-mounted monitoring into the trained improved YOLOv5s model to carry out traffic sign detection. By adopting the technical scheme of the invention, the problems of low detection accuracy and poor real-time performance of the small target of the traffic sign are solved, and the omission factor of the small target traffic sign is effectively reduced on the premise of not introducing new parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a traffic sign detection method, device, system, and storage medium. Background Art

[0002] Traffic sign detection is one of the key technologies in intelligent transportation systems. By accurately identifying traffic signs, driverless vehicles can better understand the road environment, make corresponding decisions and driving plans, thereby improving driving safety and efficiency. Traffic sign detection is of great significance for promoting the development of driverless vehicle technology, having important theoretical significance and broad application prospects. Traffic signs guide vehicles to drive in a standardized manner, improve driving efficiency, and promote road traffic safety. For example, dangerous pre-judgment can be made through prohibition signs, obstacle avoidance can be processed through warning signs, and control pre-processing can be carried out through indication signs. However, the environment where traffic signs are located is very complex, and the detection process is easily affected by weather conditions, environmental illumination, and angle changes.

[0003] Since traffic signs usually have distinct color features (red, yellow, blue) and regular shape structures (triangle, circle, square), traffic sign detection methods based on traditional handcrafted features usually detect based on the features of the color or shape attributes of different traffic signs. Traditional detection methods are easily affected by illumination changes and detection backgrounds, with low detection accuracy, poor real-time performance, and weak generalization ability.

[0004] With the emergence of algorithms such as Fast R-CNN, Faster R-CNN, and the YOLO series, general object detection methods based on deep learning have developed rapidly, and many traffic sign detection algorithms have emerged based on these general object detection algorithms. Although traffic sign detection technology based on deep learning has been continuously developing, in the images collected by actual on-vehicle cameras, small target traffic signs often account for less than 0.1% of the total image area, and objects similar to the detection targets are likely to appear in the detection background, which will interfere with the detection, easily resulting in false detections. Moreover, current traffic sign detection algorithms usually require a large amount of computing resources, and the detection real-time performance needs to be improved.

[0005] Currently, for the problem of small target detection of traffic signs, mainly due to the lack of a complete traffic sign dataset, the small target area is too small and the detection difficulty is large, there are many types of traffic signs, light changes, obstacle occlusion, and current detection algorithms have problems such as difficulty in simultaneously ensuring detection speed and detection accuracy and weak generalization ability of the detection model. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a traffic sign detection method, device, system, and storage medium, which solve the problems of low accuracy and poor real-time performance in detecting small target traffic signs, and effectively reduce the missed detection rate of small target traffic signs without introducing new parameters.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A traffic sign detection method includes:

[0009] Step S1, obtaining a traffic sign data set;

[0010] Step S2, training an improved YOLOv5s model according to the traffic sign data set;

[0011] Step S3, inputting the real-time video stream of vehicle-mounted monitoring into the trained improved YOLOv5s model for traffic sign detection.

[0012] Preferably, in step S2, the YOLOv5s model is subjected to lightweight design, feature prediction scale optimization, and feature fusion enhancement to obtain an improved YOLOv5s model; wherein, channel pruning is performed on the YOLOv5s model, and the 3x3 convolution is replaced with a depthwise separable convolution to achieve lightweight design; in multi-scale fusion, 160×160 shallow features are introduced for feature prediction scale optimization; the 5C module is used to perform fusion processing on the shallow and deep feature maps for feature enhancement.

[0013] Preferably, the improved YOLOv5s model uses the loss function EIoU Loss to replace the loss function GIoU Loss of the YOLOv5s model.

[0014] The present invention also provides a traffic sign detection device, including:

[0015] An acquisition module, configured to acquire a traffic sign data set;

[0016] A training module, configured to train an improved YOLOv5s model according to the traffic sign data set;

[0017] A detection module, configured to input the real-time video stream of vehicle-mounted monitoring into the trained improved YOLOv5s model for traffic sign detection.

[0018] Preferably, a lightweight design is carried out on the YOLOv5s model, the feature prediction scale is optimized, and feature fusion is enhanced to obtain an improved YOLOv5s model. Among them, channel pruning is performed on the YOLOv5s model, and the 3x3 convolution is replaced with a depthwise separable convolution to achieve lightweight design. In multi-scale fusion, shallow features of 160×160 are introduced to optimize the feature prediction scale. The 5C module is used to fuse the shallow and deep feature maps for feature enhancement.

[0019] Preferably, the improved YOLOv5s model uses the EIoU Loss function to replace the GIoU Loss function of the YOLOv5s model.

[0020] The present invention also provides a traffic sign detection system, including: a memory and a processor. A computer program is stored on the memory and run by the processor. The computer program executes the traffic sign detection method when run by the processor.

[0021] The present invention also provides a storage medium. A computer program is stored on the storage medium. The computer program executes the traffic sign detection method when running.

[0022] The present invention prunes unimportant and redundant modules in YOLOv5s, introduces depthwise separable convolution, reduces the model size, reduces the computing resources required by the model, and improves the real-time performance of traffic sign detection. The optimization of the feature prediction scale makes full use of shallow features, improves the perception ability of small targets, and at the same time reduces unnecessary computational overhead. The 5C module is used in the feature fusion part to weaken the mutual interference between shallow features and deep features and avoid inaccurate target positioning. The EIoU loss function is introduced to effectively solve the problem that the penalty term for proportional changes in aspect ratio fails, and effectively improve the accuracy of small target detection. The present invention has better performance in traffic sign detection. Through the accurate recognition of traffic signs, driverless cars can better understand the road environment, make corresponding decisions and driving plans, thereby improving driving safety and efficiency. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0024] Figure 1 It is a flowchart of the traffic sign detection method according to the embodiment of the present invention;

[0025] Figure 2Schematic diagram of the improved YOLOv5s model;

[0026] Figure 3 Schematic diagram of the structure for optimizing the feature prediction scale. Specific implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0029] Embodiment 1:

[0030] As Figure 1 shown, the embodiment of the present invention provides a traffic sign detection method, including:

[0031] Step S1: Obtain a traffic sign data set;

[0032] Step S2: Train an improved YOLOv5s model according to the traffic sign data set;

[0033] Step S3: Input the real-time video stream of in-vehicle monitoring into the trained improved YOLOv5s model for traffic sign detection.

[0034] As an implementation manner of the embodiment of the present invention, Chinese Traffic Sign Detection Benchmark 2021 (CCTSDB 2021) is selected. It is an improved version of Chinese Traffic Sign Detection Benchmark 2017 (CCTSDB 2017), with more than 4,000 additional images, simple samples removed, complex samples added, and a complete test data set made. CCTSDB 2021 has a total of 16,356 training images and 2,000 test images, and a total of three types of traffic signs are included. The training set contains 13,876 prohibition signs, 4,598 warning signs, and 8,363 mandatory signs, including road traffic scenarios under different weather conditions, such as foggy days, snowy days, rainy days, nights, cloudy days, and sunny days. Among the 2,000 test images, there are 1,500 positive sample images and 500 negative sample images. The test set is divided into five categories according to the size of the traffic signs. CCTSDB 2021 better restores the traffic sign detection scenario in the real road environment and can effectively improve the robustness of the model.

[0035] As an implementation manner of the embodiment of the present invention, in step S2, the YOLOv5s model is subjected to lightweight design, feature prediction scale optimization, feature fusion enhancement, and loss function optimization to obtain an improved YOLOv5s model, as Figure 2 shown below:

[0036] 1. Lightweight design

[0037] Traffic sign detection systems are usually deployed on driverless vehicles or other transportation tools with limited computing resources. Therefore, when pursuing high detection accuracy, the computational amount of the model should be minimized as much as possible to make the model easy to be deployed on the device side with limited computing resources. YOLOv5 includes four models of different sizes: YOLOv5s, YOLOv5x, YOLOv5m, and YOLOv5l. In the embodiment of the present invention, the smallest-sized YOLOv5s is selected as the baseline model. To meet the high real-time requirements of traffic sign detection, channel pruning is performed on the YOLOv5s model, and depthwise separable convolution is applied to reduce the model size and computational amount. When pruning the model, first load the network model, evaluate the neurons in the network, and then delete the unimportant and large-sized network layers. Pruning has the effect of improving the model detection speed, but to obtain better detection real-time performance, depthwise separable convolution is introduced to replace some 3x3 convolutions with depthwise separable convolutions. Depthwise separable convolution is mainly used to reduce network parameters and improve the computational efficiency of the model.

[0038] Depthwise separable convolution consists of two processes: depthwise convolution and pointwise convolution. In depthwise convolution, a convolution kernel is assigned to each channel of the feature map for convolution. If a convolution of size K×K×M is performed on an image of size H×W×N, the computational power consumed is H×W×N×K×K×M. In depthwise separable convolution, a depthwise convolution of size K×K×N is first performed on the feature map, followed by a pointwise convolution of size 1×1×M. The total computational power consumed is H×W×N×K×K + N×M×H×W. Given the same input features and obtaining a feature map with the same number of output channels after convolution, the computational cost of depthwise separable convolution is 1 / M + 1 / K that of conventional convolution. 2 After the model is lightweighted, the model size becomes smaller, the computational cost decreases, and the detection speed is greatly improved.

[0039] 2. Feature prediction scale optimization

[0040] In the CCTSDB 2021 dataset, small target traffic signs account for the highest proportion. In the actual traffic sign detection scenario images, most traffic signs are small targets. However, YOLOv5s only uses features at three scales, 20×20, 40×40, and 80×80, for detection, using less shallow information and prone to losing the position information of small target traffic signs. Small target detection relies on the larger feature maps at the bottom layer for detection, and rich spatial position information can be captured from the shallow features to fully extract the features of small target traffic signs. To make full use of shallow features for detecting small target traffic signs, two strategies are adopted to improve the network prediction scale. The first strategy is to introduce shallow features of 160×160 in the multi-scale fusion part, increasing the number of detection feature scales of YOLOv5s. The YOLOv5 prediction scales become 20×20, 40×40, 80×80, and 160×160, respectively responsible for detecting traffic targets of four different size levels: large, medium, small, and extremely small. The newly added detection branch consists of shallow features (the feature layer before P3) and deep features that have undergone more convolutions, making full use of the rich fine-grained information in the shallow features and the rich semantic information in the deep features, which can theoretically effectively improve the detection effect of small target traffic signs. The second improvement strategy is to introduce shallow features in the multi-scale fusion part without increasing the number of detection feature scales of YOLOv5s. Considering that the number of large targets in general traffic sign datasets is very small, while introducing the feature scales responsible for detecting small and extremely small targets, the feature scales responsible for detecting large target traffic signs are deleted. This is theoretically not expected to affect the network's detection effect, and ultimately, unnecessary computational expenses are reduced without affecting the network's detection effect. This strategy aims to improve the network's ability to detect small targets. The network frameworks of the two improved prediction scales are as Figure 3As shown, it is the simple structure of lightweight YOLOv5s. F1-F4 represent the fusion of backbone features and depth features at each layer, and P1-P4 represent further processing of the fused features. F1 is a newly added feature layer that fuses shallow and deep information. P1 further extracts the information of the F1 feature layer, and processes the features of P1 to obtain a detection head of 160×160. By introducing features of 160×160 and integrating them with deeper features, the network can better capture spatial location information and extract necessary features to detect small target traffic signs.

[0041] 3. Feature Fusion Enhancement

[0042] YOLOv5s consists of four parts: the input end, Backbone, Neck, and Prediction. The main function of Backbone is to extract features, and Neck is responsible for fusing features from Backbone. A large number of C3 structures and convolutional layers are used in the Backbone and neck parts. C3 is stacked by convolutions, and feature information can be effectively extracted through the convolutional module. However, using a large number of convolutions will cause the loss of key image information in the deep feature map. To solve this problem, a more efficient module is constructed to improve the quality of feature extraction, increase the information flow, and enhance the network's expression ability. This module consists of 5 convolutions, so this model is named 5C. 5C can be split into three branches. The first branch first passes through a 1×1 convolution to halve the number of channels, and then passes through a 3×3 convolution for further feature extraction. The second branch first passes through a 3×3 convolution and then through a 1×1 convolution to halve the number of channels. The third branch inputs the feature map through a 1×1 convolution for dimensionality reduction processing, and the number of channels becomes half of the original number of channels, and the size of the feature map remains unchanged. Their expressions are as follows:

[0043] Y1 = f Conv1×1 (f Conv3×3 (I)) (1)

[0044] Y2 = f Conv3×3 (f Conv1×1 (I)) (2)

[0045] Y3 = f Conv1×1 (I) (3)

[0046] In the formula, I is the input feature map, and Y1, Y2, and Y3 are the feature maps output by the three branches respectively. f Conv3×3 represents a 3×3 two-dimensional convolution, and f Conv1×1 represents a 1×1 two-dimensional convolution.

[0047] Y1, Y2, and Y3 are of the same size. An add operation is performed on Y1 and Y2 to obtain Y4. The number of channels in the Y4 feature map remains unchanged, and the amount of information in each dimension of the feature map is increased. Then, Y3 and Y4 are concatenated to fuse the features and increase the feature dimension to ensure that the number of channels in the output feature map is the same as that in the input feature map.

[0048]

[0049] Y = Y3 ⊙ Y4 (5)

[0050] In the formula, represents the addition of feature maps with the number of channels remaining unchanged, ⊙ represents the addition of the number of channels with the feature map remaining unchanged, Y4 represents the feature map output after ⊙ of Y1 and Y2, and Y is the feature map output by the 5C module. Y3 and Y4 have different receptive fields, enabling the network to learn rich features and making up for the accuracy decline caused by insufficient feature learning ability. This module further extracts the input features, retains the characteristics of the input features, and enriches the information level of the features. Since a non-linear activation function is added after the 1×1 convolution, this module increases the non-linear characteristics of the entire model and improves the overall expression ability of the model. To explore the performance of the 5C module, two modules similar to 3C were constructed and compared with 3C. Through ablation experiments, it was proved that using the 3C module in the feature fusion part of YOLOv5s can improve the detection performance of the model.

[0051] 4. Loss Function Optimization

[0052] To make full use of the target information in the dataset, the model is further optimized from the aspect of the loss function. The EIoU Loss is used to replace the original loss function GIoU Loss in the model to improve the detection accuracy of the model by obtaining more feature information. When there are two prediction boxes with the same width and height and in the same horizontal plane, the GIoU Loss degrades to the IoU Loss, resulting in problems such as slow convergence and inaccurate regression. The DIoU Loss further solves the problems existing in the GIoU Loss. The CIoU Loss adds the consideration of the aspect ratio on the basis of the DIoU Loss, but there are problems such as the description of the aspect ratio being fuzzy and the balance of easy and difficult samples not being considered. The EIoU Loss addresses the problems existing in the CIoU Loss. On the basis of the CIoU Loss, the difference values of the width and height are calculated separately to replace the aspect ratio, and the Focal Loss is introduced to solve the problem of sample imbalance. The EIoU Loss minimizes the width and height differences between the target box and the anchor box, has a faster convergence speed and better positioning effect, and its formula is as follows:

[0053]

[0054] Among them, c w and c h respectively represent the width and height of the minimum bounding rectangles of the target box and the predicted box, ρ 2 represents the Euclidean distance between two points, b and b gt respectively represent the center point coordinates of the target box and the predicted box, d is the distance between the center points of the target box and the predicted box, and c is the diagonal distance of the minimum bounding rectangle. By introducing the EIoU Loss, the model can extract target features more thoroughly, further improving the model detection performance.

[0055] Furthermore, in step S2, the traffic sign dataset is input into the improved YOLOv5s model for training; the experiment is carried out on a platform with NVIDIA GeForce GTX 3080ti and 16GB RAM; the input image size is set to 640×640, and the model is trained using the Stochastic Gradient Descent (SGD) optimizer, with a batch size of 32, a weight decay of 0.0005, and a learning rate of 0.01.

[0056] Example 2:

[0057] The embodiment of the present invention also provides a traffic sign detection device, including:

[0058] An acquisition module, used to acquire the traffic sign dataset;

[0059] A training module, used to train the improved YOLOv5s model according to the traffic sign dataset;

[0060] A detection module, used to input the real-time video stream of vehicle-mounted monitoring into the trained improved YOLOv5s model for traffic sign detection.

[0061] As an implementation manner of the embodiment of the present invention, a lightweight design, feature prediction scale optimization, and feature fusion enhancement are performed on the YOLOv5s model to obtain the improved YOLOv5s model; among them, channel pruning is performed on the YOLOv5s model, and the 3x3 convolution is replaced with a depthwise separable convolution to achieve lightweight design; 160×160 shallow features are introduced in multi-scale fusion for feature prediction scale optimization; the 5C module is used to perform fusion processing on the shallow and deep feature maps for feature enhancement.

[0062] As an implementation manner of the embodiment of the present invention, the improved YOLOv5s model uses the loss function EIoU Loss to replace the loss function GIoU Loss of the YOLOv5s model.

[0063] Example 3:

[0064] An embodiment of the present invention further provides a traffic sign detection system, including: a memory and a processor, where a computer program run by the processor is stored on the memory, and the computer program executes a traffic sign detection method when being run by the processor.

[0065] Embodiment 4:

[0066] An embodiment of the present invention further provides a storage medium, where a computer program is stored on the storage medium, and the computer program executes a traffic sign detection method when running.

[0067] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A traffic sign detection method, characterized in that: include: Step S1, obtaining a traffic sign dataset; Step S2: training an improved YOLOv5s model based on the traffic sign dataset; Step S3: input the real-time video stream of the vehicle monitoring into the trained improved YOLOv5s model for traffic sign detection.

2. The traffic sign detection method according to claim 1, characterized in that: In step S2, the YOLOv5s model is subjected to lightweight design, feature prediction scale optimization, and feature fusion enhancement to obtain an improved YOLOv5s model; wherein, the YOLOv5s model is subjected to channel pruning, and the 3x3 convolution is replaced with a depthwise separable convolution to achieve lightweight design; 160×160 shallow features are introduced in multi-scale fusion to optimize the feature prediction scale; and the 5C module is used to fuse the shallow and deep feature maps for feature enhancement.

3. The traffic sign detection method according to claim 2, characterized in that: The improved YOLOv5s model uses the loss function EIoU Loss instead of the loss function GIoU Loss of the YOLOv5s model.

4. A traffic sign detection device, characterized in that: include: An acquisition module, used to acquire traffic sign datasets; Training module, used to train the improved YOLOv5s model based on the traffic sign dataset; The detection module is used to input the real-time video stream of vehicle monitoring into the trained improved YOLOv5s model for traffic sign detection.

5. The traffic sign detection device according to claim 4, characterized in that: The YOLOv5s model is subjected to lightweight design, feature prediction scale optimization, and feature fusion enhancement to obtain an improved YOLOv5s model. Among them, the YOLOv5s model is subjected to channel pruning and the 3x3 convolution is replaced by a depthwise separable convolution to achieve lightweight design. In multi-scale fusion, 160×160 shallow features are introduced to optimize the feature prediction scale. The 5C module is used to fuse the shallow and deep feature maps for feature enhancement.

6. The traffic sign detection device according to claim 5, characterized in that: The improved YOLOv5s model uses the loss function EIoU Loss instead of the loss function GIoU Loss of the YOLOv5s model.

7. A traffic sign detection system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the traffic sign detection method according to any one of claims 1 to 3 is executed.

8. A storage medium, characterized in that: The storage medium stores a computer program, which executes the traffic sign detection method according to any one of claims 1 to 3 when running.

Citation Information

Patent Citations

  • Traffic sign detection method based on YOLOv5

    CN115661788A

  • Lightweight traffic sign detection method, storage medium and system

    CN116453091A

  • Real-time high-precision traffic sign small target detection method based on improved YOLOv7-tiny

    CN116778455A

  • Traffic sign detection method based on improved YOLOv8

    CN118314551A