Traffic signal sign detection method based on improved yov5

By improving the yolov5 model, using the C2f module, parallel Mamba module and SimAM attention mechanism, the problem of traffic signal sign detection accuracy and calculation amount is solved, and efficient small object detection and accurate recognition in complex scenarios is achieved.

CN120279510APending Publication Date: 2025-07-08云南公路联网收费管理有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432838.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing traffic signal sign detection methods have shortcomings in accuracy, parameter quantity and calculation quantity, which are difficult to meet the needs of real-time application.

Method used

By improving the yolov5 model, the improved C2f module, parallel Mamba module, RFB module and SimAM attention mechanism are adopted to enhance the network's perception of traffic signal targets and improve detection accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of small-objective detection of traffic signal signs, improves the model's processing ability for complex scenarios, and reduces calculation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279510A_ABST
    Figure CN120279510A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic signal sign detection method based on improved yolov5, and belongs to the field of traffic signal sign detection, and the method comprises the following steps: collecting a traffic signal sign data set, and carrying out the marking, preprocessing and data enhancement operation of the data set; the method comprises the following steps of: improving a yolov5 model from three parts of a feature extraction network, spatial pyramid pooling and an attention mechanism, extracting features by using an improved C2f convolution module and a parallel Mama module in a yolov5 backbone network, namely a feature extraction stage, and selecting RFB as a new spatial pyramid pooling module; in the yolov5 feature processing stage, a SimAM attention mechanism is added; sending the data into the optimized and improved yolov5 model for training; and storing the optimal model in the training process, and performing target detection on the traffic signal sign image or video by using the model to obtain a prediction result. By improving a backbone network and increasing an attention mechanism, the perception capability of the network to a traffic signal target is enhanced, and the detection precision of the traffic signal sign target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic signal sign detection, and particularly relates to a traffic signal sign detection method based on improved YOLOv5. Background Art

[0002] Object detection is a core technology in the field of computer vision, aiming to determine the position of specific objects in an image and classify them. Traditional object detection methods, such as multi-stage methods based on sliding windows and region proposals, usually rely on complex feature engineering and lengthy processing flows. These methods often face a trade-off between computational efficiency and detection accuracy when dealing with real-time applications or high-resolution images, and it is difficult to meet actual requirements. When applied to traffic signal sign detection, the accuracy is poor.

[0003] With the rapid development of deep learning technology, especially the widespread application of convolutional neural networks (CNNs), object detection technology has been significantly improved. For example, methods such as R-CNN, Fast R-CNN, and Faster R-CNN use deep convolutional neural networks for feature extraction and region classification, which not only simplifies the process but also greatly improves the detection accuracy. At the same time, single-stage detection methods (such as YOLO and SSD) have significantly improved the detection speed by optimizing the detection process and are more suitable for real-time application scenarios. However, the limitations of traditional CNN architectures still exist, such as insufficient ability to capture long-range dependencies and limited model scalability. In recent years, the Transformer architecture has gradually been introduced into the field of computer vision and has shown superior performance to traditional CNNs in many tasks. However, although the Transformer has strong advantages in long-range modeling, the resulting large number of parameters and computational requirements do not achieve a balance between accuracy and lightness. Although many methods seek to improve lightweight Transformer architectures, their parameter and computational requirements always far exceed those of CNNs. Summary of the Invention

[0004] The purpose of the present invention is to provide a traffic signal sign detection method based on improved YOLOv5 to address the problems of poor detection accuracy and large number of parameters and computational requirements for current traffic signal sign object detection. By replacing the backbone network and adding an attention mechanism, the network's perception ability of traffic signal targets is enhanced, and the detection accuracy of small traffic signal signs is improved.

[0005] The technical solution of the present invention is as follows:

[0006] A traffic signal sign detection method based on improved YOLOv5 includes the following steps:

[0007] Obtain a traffic signal sign image or video dataset, and perform annotation, preprocessing, and data augmentation operations on the dataset;

[0008] Improve the YOLOv5 model from the feature extraction stage, spatial pyramid pooling, and attention mechanism parts; in the feature extraction stage, use an improved C2f convolutional module and a parallel Mamba module for feature extraction; use the RFB module as the new spatial pyramid pooling module; in the feature processing stage, add the SimAM attention mechanism;

[0009] Input the dataset after preprocessing and data augmentation operations into the improved YOLOv5 model for training;

[0010] Use the improved YOLOv5 model with the best training results to perform object detection on traffic signal sign images or videos to obtain detection results.

[0011] Furthermore, an EMA attention mechanism is set in the improved C2f module to weight the importance of different channels or regions in the feature map and adaptively allocate attention; the improved C2f module specifically includes the following steps:

[0012] x = Conv(x),

[0013] x1, x2 = Sp(x),

[0014] y0 = x1,

[0015] y = EMA(y),

[0016] x = Conv(Concat(x2, y)),

[0017] where the operation of the Bottleneck module in the i-th layer is f i (), where i = 1, 2,..., n, the input feature map x passes through n layers of Bottleneck modules in sequence, and the output of each layer is denoted as y i , Concat represents the concatenation operation in the channel dimension, Sp represents the separation operation in the channel dimension, and EMA represents the EMA attention mechanism.

[0018] Furthermore, the parallel Mamba module for feature extraction includes the following steps:

[0019]

[0020] Y i ′ C / 4 = VSS(Y i C / 4 ) + θ·Y i C / 4 , i = 1, 2, 3, 4,

[0021] X out = Cat(Y1′ C / 4 ,Y2′ C / 4 ,Y3′ C / 4 ,Y4′ C / 4 ),

[0022] Out = Projection[LN(X out )],

[0023] Among them, the feature map X has C channels. First, through a LayerNorm layer, the feature map is split into four feature maps and Each feature map has C / 4 channels; subsequently, each feature map is separately input into the VSS module and scaled, and then the global spatial information is enhanced through residual connections; finally, the four feature maps are merged into an output feature map X with C channels through a splicing operation out , and LayerNorm and projection processing are performed on the result.

[0024] Furthermore, the SimAM attention mechanism evaluates the importance of each neuron through an energy function, and this energy function measures the linear separability between the target neuron and other neurons; the energy function is:

[0025]

[0026] Using binary labels and adding a regularization term, the final energy function is:

[0027]

[0028] The analytical solution is:

[0029] Calculate the mean and variance of the input features in the H and W dimensions:

[0030]

[0031] The whole process can be expressed as:

[0032]

[0033] Furthermore, the EMA attention mechanism includes:

[0034] Parallel branch setting: EMA splits the 1×1 convolution part in the original CA into a separate branch and places a 3×3 convolution branch in parallel beside it;

[0035] Feature grouping: EMA divides the input feature map into GGG sub - feature groups along the channel dimension, and each sub - feature group undergoes feature extraction and fusion through parallel 1×1 and 3×3 branches respectively;

[0036] 1×1 branch: This branch aggregates channel information in two spatial directions through 1D global average pooling and uses 1×1 convolution to learn the cross - channel attention distribution; at the output end, the obtained attention vectors in the two directions are non - linearly activated and multiplied to obtain the channel - level attention map;

[0037] 3×3 branch: This branch uses a 3×3 convolutional kernel to capture local spatial interaction information, complementing the 1×1 branch; in the cross - spatial learning stage, global average pooling is also used in combination with functions such as Softmax to generate pixel - level attention maps to retain richer spatial location information;

[0038] Cross - spatial learning: EMA performs global pooling and information aggregation in the spatial dimension on the features obtained from the 1×1 branch and 3×3 branch in the output stage to form two complementary spatial attention maps; finally, these two spatial attention weights are combined through methods such as dot - product or weighted fusion to obtain a more refined pixel - level attention distribution.

[0039] Furthermore, the RFB module simulates the receptive field of human vision to enhance the feature extraction ability of the model. The RFB module introduces a multi - scale parallel convolution structure in the feature extraction process and captures and fuses the input features through convolutional kernels of different scales.

[0040] Furthermore, the pre - processing and data augmentation operations include Mosaic, scaling, rotation, translation, cropping, and HSV color space enhancement.

[0041] Furthermore, the traffic signal sign image or video dataset includes prohibitory signs, warning signs, and indicative signs.

[0042] Furthermore, the dataset is the publicly available dataset CCTSDB2021, and the data annotation process annotates the dataset in YOLO format through the labeling software labelImage. The labeling categories include: prohibitory, warning, mandatory.

[0043] Furthermore, the improved yolov5 model is evaluated by mean average precision (mAP), average precision (AP), precision, and recall. The calculation formulas for mean average precision (mAP), average precision (AP), precision, and recall are as follows:

[0044]

[0045]

[0046] The beneficial effects of the present invention compared with the existing technologies are as follows:

[0047] 1. A traffic signal sign detection method based on improved YOLOv5 adopts an improved C2f module. Compared with the traditional C3 module, the improved C2f module can weight the importance of different channels or regions in the feature map and adaptively allocate attention, thus significantly enhancing the detection network's ability to capture details of traffic signal signs and robustness. The C2f module introduces the EMA attention mechanism in the feature fusion stage, dynamically weights the input features, calculates the response degrees of each scale and each channel, strengthens the attention to key information, and simultaneously suppresses background or interference features. The C2f module based on the EMA attention mechanism can more effectively fuse multi-scale information, capture local details, and maintain global consistency compared with the traditional C3 module, which enables the traffic signal sign detection network to exhibit more excellent performance in small target recognition, complex scene interference processing, and multi-class target discrimination, etc.;

[0048] 2. A traffic signal sign detection method based on improved YOLOv5 adds a parallel Mamba module to the backbone network, and the Mamba module adopts a lightweight design, which can exchange a small number of parameters for good feature extraction ability;

[0049] 3. A traffic signal sign detection method based on improved YOLOv5 adopts a new feature extraction module RFB. The RFB module strengthens the network's feature extraction ability for low-resolution images by simulating the receptive field of human vision;

[0050] 4. A traffic signal sign detection method based on improved YOLOv5 adds the attention mechanism SimAM, and proposes an efficient multi-scale attention mechanism, which can capture both channel and spatial information and effectively improve the feature representation ability without increasing too many parameters and computational costs; by reshaping some channels into the batch dimension and grouping the channel dimension into multiple sub-features, the spatial semantic features can be well distributed within each feature group, improving the expression ability of features; adopting a parallel sub-network design helps to capture cross-dimensional interactions and dependencies between proposed dimensions, and improves the model's ability to model long-range dependencies. Brief Description of the Drawings

[0051] Figure 1 It is a flowchart of a traffic signal sign detection method based on improved YOLOv5.

[0052] Figure 2 It is an overall network architecture diagram of a traffic signal sign detection method based on improved YOLOv5.

[0053] Figure 3 It is a structural diagram of the RFB module for a traffic signal sign detection method based on improved YOLOv5.

[0054] Figure 4 It is the architecture of the EMA attention for a traffic signal sign detection method based on improved YOLOv5 Figure 1 .

[0055] Figure 5 It is the architecture of the EMA attention for a traffic signal sign detection method based on improved YOLOv5 Figure 2 .

[0056] Figure 6 It is a structural diagram of the parallel Mamba module for a traffic signal sign detection method based on improved YOLOv5.

[0057] Figure 7 It is a detection result diagram for a traffic signal sign detection method based on improved YOLOv5. Specific implementation manners

[0058] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0059] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.

[0060] Please refer to Figures 1-7 , a traffic signal sign detection method based on improved YOLOv5, as Figure 1 and Figure 2 shown, which includes the following steps:

[0061] Obtain traffic signal sign image or video datasets, and perform annotation, preprocessing and data augmentation operations on the datasets; the preprocessing and data augmentation operations include Mosaic, scaling, rotation, translation, cropping, and HSV color space enhancement. The traffic signal sign image or video datasets include prohibition signs, warning signs, and indication signs.

[0062] Improve the YOLOv5 model from the feature extraction stage, spatial pyramid pooling, and attention mechanism parts; in the feature extraction stage, use the improved C2f convolution module and parallel Mamba module for feature extraction; use the RFB module as the new spatial pyramid pooling module; in the feature processing stage, add the SimAM attention mechanism;

[0063] The improved C2f module is provided with an EMA attention mechanism, which weights the importance of different channels or regions in the feature map and adaptively allocates attention; by integrating the EMA attention in the C2f structure, the network further enhances the attention to local and global information while maintaining the efficient feature extraction ability, thereby improving the detection accuracy and robustness of traffic signal signs. The specific steps of the improved C2f module are as follows:

[0064] x = Conv(x),

[0065] x1, x2 = Sp(x),

[0066] y0 = x1,

[0067] y = EMA(y),

[0068] x = Conv(Concat(x2, y)),

[0069] Among them, the operation of the i-th layer Bottleneck module is f i (), where i = 1, 2,..., n, the input feature map x passes through n layers of Bottleneck modules in sequence, and the output of each layer is denoted as y i , Concat represents the concatenation operation in the channel dimension, Sp represents the separation operation in the channel dimension, and EMA represents the EMA attention mechanism.

[0070] The improved C2f convolution module replaces the C3 module of the traditional YOLOv5. Among them, the improved C2f module adds an EMA attention mechanism on the basis of the original network structure. This mechanism can weight the importance of different channels or regions in the feature map and adaptively allocate attention, thereby further improving the ability of the detection network to capture details of traffic signal signs and robustness. Specifically, the C2f module dynamically weights the input features through the EMA attention mechanism in the feature fusion stage, calculates the response degree of each scale and each channel, strengthens the attention to key information, and suppresses background or interference features. Compared with the traditional C3 module, the C2f module based on EMA attention can better fuse multi-scale information, capture local details and maintain global consistency, making the traffic signal sign detection network perform better in small target recognition, complex scene interference, and multi-class target discrimination.

[0071] As Figure 4 and Figure 5 shown, the EMA attention mechanism includes:

[0072] Parallel branch setting: The MA module uses parallel substructures to avoid the overhead and gradient transmission problems caused by overly deep network layers. EMA splits the 1×1 convolution part in the original CA (Coordinate Attention) into a separate branch (1×1 branch) and places a 3×3 convolution branch (3×3 branch) in parallel beside it; this parallel structure can capture local and global multi-scale features simultaneously within the same stage.

[0073] Feature grouping: EMA divides the input feature map into GGG sub-feature groups along the channel dimension, and each sub-feature group undergoes feature extraction and fusion through parallel 1×1 and 3×3 branches respectively;

[0074] 1×1 branch: This branch aggregates channel information separately in two spatial directions (H and W) through 1D global average pooling and uses 1×1 convolution (without dimensional compression) to learn the cross-channel attention distribution; at the output end, the two-direction attention vectors obtained are non-linearly activated (Sigmoid) and multiplied to obtain the channel-level attention map;

[0075] 3×3 branch: This branch uses a 3×3 convolution kernel to capture local spatial interaction information, complementing the 1×1 branch; during the cross-space learning stage, global average pooling is also used in conjunction with functions such as Softmax to generate pixel-level attention maps to retain richer spatial location information;

[0076] Cross-space learning: EMA performs global pooling and information aggregation in the spatial dimension on the features obtained from the 1×1 branch and 3×3 branch during the output stage to form two complementary spatial attention maps; finally, these two spatial attention weights are merged through methods such as dot product or weighted fusion to obtain a more refined pixel-level attention distribution.

[0077] The output of EMA maintains the same resolution as the input, and under the combined action of parallel branches and cross-space aggregation, a richer and more accurate attention map is obtained. Compared with single-path or pure channel attention, EMA can simultaneously focus on multi-scale information of channels and space, capture more complete global context and finer-grained local features, and thus achieve better performance in tasks such as object detection and semantic segmentation.

[0078] As Figure 6As shown in the figure, the parallel Mamba module for feature extraction includes the following steps:

[0079]

[0080] Y i ′ C / 4 = VSS(Y i C / 4 ) + θ·Yi i C / 4 , i = 1, 2, 3, 4,

[0081] X out = Cat(Y1′ C / 4 , Y2′ C / 4 , Y3′ C / 4 , Y4′ C / 4 ),

[0082] Out = Projection[LN(X out )],

[0083] Among them, the feature map X has C channels. First, through a LayerNorm layer, the feature map is divided into four feature maps and each feature map has C / 4 channels; subsequently, each feature map is respectively input into the VSS module and undergoes scaling processing, and then the global spatial information acquisition is enhanced through residual connection; finally, the four feature maps are merged into an output feature map X with C channels through a concatenation operation out , and LayerNorm and projection processing are performed on the result.

[0084] The parallel Mamba module adopts a lightweight design, which can exchange a small number of parameters for good feature extraction ability. Cooperating with the improved C2f module, it can enhance the feature extraction ability of the backbone network, thus preparing for the subsequent feature processing stage and improving the detection accuracy.

[0085] As Figure 3 shown in the figure, the RFB module simulates the receptive field of human vision to enhance the feature extraction ability of the model. The RFB module introduces a multi-scale parallel convolution structure in the feature extraction process, and captures and fuses the input features through convolution kernels of different scales. Thus, richer context information and a more flexible receptive field range are obtained. Compared with the traditional convolution module, the RFB module can focus on the local details and global features of the target at multiple resolutions, further strengthening the network's detection and recognition performance of traffic signal signs.

[0086] The SimAM attention mechanism evaluates the importance of each neuron through an energy function, which measures the linear separability between the target neuron and other neurons; the energy function is:

[0087]

[0088] Using binary labels and adding a regularization term, the final energy function is:

[0089]

[0090] The analytical solution is:

[0091] Calculate the mean and variance of the input features in the H and W dimensions:

[0092]

[0093] The whole process can be expressed as:

[0094]

[0095] The SimAM attention mechanism is an attention mechanism designed based on neuroscience theory, aiming to overcome the limitations of traditional 1D or 2D attention mechanisms in feature discrimination ability. Its core idea is to assign unique 3D weights to each neuron, thereby achieving finer-grained feature selection and enhancement. SimAM evaluates the importance of each neuron by defining an energy function, which measures the linear separability between the target neuron and other neurons. To efficiently calculate the weights, SimAM provides a closed-form solution, avoiding the high computational cost of iterative solutions, assuming that all pixels within the same channel follow the same distribution, thus simplifying the calculation process. In addition, SimAM not only captures the information between channels but also retains the precise spatial structure. Through global average pooling and cross-space information aggregation, it can model long-range dependencies and retain local position information. This module is designed simply and is easy to integrate into existing deep learning frameworks such as PyTorch without significantly increasing the computational overhead.

[0096] Compared with traditional attention modules such as CA and CBAM, EMA can achieve better performance while maintaining the model size and computational efficiency.

[0097] Input the dataset after preprocessing and data augmentation operations into the improved yolov5 model for training;

[0098] Use the improved yolov5 model with the best training results to perform object detection on traffic signal sign images or videos to obtain the detection results.

[0099] Experimental verification:

[0100] This experiment was conducted on Hengyuan Cloud. The operating system was Ubuntu 20.04, the GPU was Nvidia GeForce GTX 3080 Ti with 12G video memory, the CPU was Intel(R) Xeon(R) Silver 4214R CPU @ 2.40 GHz, the Pytorch version was 1.13.0, the Python language environment was 3.8, and the CUDA version was 11.7.

[0101] The dataset was sourced from the CCTSDB2021 (China Road Traffic Sign Database 2021) traffic sign dataset, which was produced by relevant scholars and teams at Changsha University of Science and Technology. It had nearly 20,000 traffic sign sample images, containing nearly 40,000 traffic signs in total. It also included an image dataset classified by size and scene, with rich road background information. The images were sourced from roads with normal traffic in China, and the shooting perspectives were mostly street shots and dash cams inside the vehicle, mainly from the driver's and passenger's seats inside the vehicle. After random screening, a total of 17,856 images were obtained, including 16,356 training images and 1,500 test images.

[0102] In this experiment, the epoch value was 200, the learning rate was 0.001, the momentum was 0.937, the weight_decay was 0.0005, the batch size was 16, the number of processes was 1, the Adam optimizer was used, and the input image size of the model was 640X640 pixels. The same hyperparameters were used for each group of experiments.

[0103] The evaluation metrics used were the three relatively common evaluation metrics in object detection algorithms: average precision (AP, Average Precision), mean average precision (mAP, mean Average Precision), and frames per second (FPS, Frames Per Second) to evaluate the performance of the algorithm in this paper. Average precision is related to precision and recall. Precision refers to the number of correctly predicted positive samples in the predicted dataset divided by the number of samples predicted as positive by the model; recall refers to the number of correctly predicted positive samples in the predicted dataset divided by the number of actual positive samples.

[0104] The calculation formulas for average precision, mean average precision, precision, and recall are as follows:

[0105]

[0106] To verify the effectiveness of the algorithm, the algorithm was tested on the same test set as the original YOLOv5 algorithm, and a series of ablation experiments were carried out. The comparison results of various performance indicators are shown in the following table.

[0107] Table 1 Comparison of Various Performance Indicators

[0108]

[0109]

[0110] As Figure 7 shown, it can be seen from the experimental results in the table that the introduction of each improvement strategy has a certain impact on the model performance. Based on the baseline model YOLOv5s, the C2f module was first introduced, and the accuracy and recall rate of the model were significantly improved by optimizing the network structure. Among them, the P index increased by 2.9%, the R index increased by 0.9%, the mAP@0.5 increased by 1.7%, and the mAP@0.5-0.95 increased by 0.2%. Further adding the SimAM module optimized the feature representation ability through the spatial adaptive attention mechanism, increasing the recall rate R to 71.8%, the mAP@0.5 to 80.7%, and the mAP@0.5-0.95 to 50.5%, showing a stronger ability to capture fine-grained features. On this basis, after adding the Mamba framework, although the P index decreased slightly (by 1.3%), the R index and the mAP index still remained at a high level, indicating that the introduction of Mamba has a positive effect on the stability of the detection performance. It is worth noting that due to the addition of the feature enhancement module, the detection effect of small targets was significantly optimized. After further introducing the RFB module, the model performed well in multi-scale feature fusion, and the recall rate increased to 72.1%, but the mAP index decreased slightly, indicating that the RFB module is helpful for feature enhancement, but may cause certain interference to small target detection under specific conditions.

[0111] Finally, after adding the EMA attention mechanism module, the model performance reached the best state. The P index decreased slightly to 87.5%, but the R index increased to 73.9%, the mAP@0.5 increased to 81.1%, and the mAP@0.5-0.95 increased to 51.2%. The introduction of the EMA module effectively smoothed the weight update during training and improved the generalization ability of the model. Compared with the original YOLOv5s model, the improved algorithm achieved significant improvements in the recall rate, mAP@0.5, and mAP@0.5-0.95 indicators, especially performing well in the small target detection task. In addition, the organic combination of each module enhanced the robustness of the model and played an important role in reducing the false alarm and miss rate.

[0112] The above-described embodiments merely represent specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application.

Claims

1. A traffic signal sign detection method based on improved YOLOv5, characterized in that, It includes the following steps: Obtain a traffic signal sign image or video dataset, and perform annotation, preprocessing, and data augmentation operations on the dataset; Improve the yolov5 model from the feature extraction stage, spatial pyramid pooling, and attention mechanism parts; in the feature extraction stage, use an improved C2f convolution module and a parallel Mamba module for feature extraction; Adopt the RFB module as the new spatial pyramid pooling module; Add the SimAM attention mechanism in the feature processing stage; Input the dataset after preprocessing and data augmentation operations into the improved yolov5 model for training; Perform object detection on the traffic signal sign image or video through the improved yolov5 model with the optimal training result to obtain the detection result.

2. The traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that An EMA attention mechanism is set in the improved C2f module, and the importance of different channels or regions in the feature map is weighted through the EMA attention mechanism and attention is adaptively allocated; the improved C2f module specifically includes the following steps: x = Conv(x), x1, x2 = Sp(x), y = EMA(y), x = Conv(Concat(x2, y)), Among them, the operation of the Bottleneck module in the $i$-th layer is $f$ i (), where $i = 1, 2, \ldots, n$. The input feature map $x$ passes through $n$ layers of Bottleneck modules in sequence, and the output of each layer is denoted as $y$ i , Concat represents the concatenation operation in the channel dimension, Sp represents the separation operation in the channel dimension, and EMA represents the EMA attention mechanism.

3. A traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that The parallel Mamba module for feature extraction includes the following steps: Out=Projection[LN(X out )], Among them, the feature map X has C channels. First, through a LayerNorm layer, the feature map is divided into four feature maps and Each feature map has C / 4 channels; subsequently, each feature map is respectively input into the VSS module and undergoes scaling processing, and then the global spatial information acquisition is enhanced through residual connection; finally, the four feature maps are merged into an output feature map X with C channels through a concatenation operation out , and LayerNorm and projection processing are performed on the result.

4. A traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that, The SimAM attention mechanism evaluates the importance of each neuron through an energy function, and this energy function measures the linear separability between the target neuron and other neurons; the energy function is: Using binary labels and adding a regularization term, the final energy function is: The analytical solution is as follows: Calculate the mean and variance of the input features in the H and W dimensions: The whole process can be expressed as:

5. A traffic signal sign detection method based on improved YOLOv5 according to claim 2, characterized in that, The EMA attention mechanism includes: Parallel branch setting: EMA splits the 1×1 convolution part in the original CA into a separate branch and places a 3×3 convolution branch in parallel beside it; Feature grouping: EMA divides the input feature map along the channel dimension into GGG sub-feature groups, and each sub-feature group undergoes feature extraction and fusion through parallel 1×1 and 3×3 branches respectively; 1×1 branch: This branch aggregates channel information in two spatial directions through 1D global average pooling and uses 1×1 convolution to learn the cross-channel attention distribution; at the output end, the two-direction attention vectors obtained are non-linearly activated and multiplied to obtain the channel-level attention map; 3×3 branch: This branch uses a 3×3 convolution kernel to capture local spatial interaction information, which complements the 1×1 branch; in the cross-spatial learning stage, global average pooling and the Softmax function are also used to generate a pixel-level attention map to retain richer spatial position information; Cross-spatial learning: EMA performs global pooling and information aggregation in the spatial dimension on the features obtained from the 1×1 branch and the 3×3 branch in the output stage to form two complementary spatial attention maps; finally, these two spatial attention weights are combined through dot product or weighted fusion to obtain a more refined pixel-level attention distribution.

6. The traffic signal sign detection method based on improved YOLOv5 according to claim 1, wherein The RFB module simulates the feature extraction ability of the receptive field enhancement model of human vision. The RFB module introduces a multi-scale parallel convolution structure in the feature extraction process, and captures and fuses the input features through convolution kernels of different scales.

7. A traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that The preprocessing and data augmentation operations include Mosaic, scaling, rotation, translation, cropping, and HSV color space enhancement.

8. A traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that, The traffic signal sign image or video dataset includes prohibition signs, warning signs, and indication signs.

9. A traffic signal sign detection method based on improved YOLOv5 according to claim 8, characterized in that, The dataset is the publicly available dataset CCTSDB2021. The data annotation process annotates the dataset in YOLO format through the labeling software labelImage. The labeling categories include: prohibitory, warning, mandatory.

10. A traffic signal sign detection method based on improved YOLOv5 according to claim 1, characterized in that, The improved yolov5 model is evaluated by mean average precision (mAP), average precision (AP), precision, and recall. The calculation formulas for mAP, AP, precision, and recall are as follows: