Traffic sign detection method

Through the LTS-YOLOv10 model that integrates full-dimensional dynamic convolution and attention-guided features, the traffic sign detection problem in small targets and complex scenarios is solved, the detection accuracy and efficiency are improved, and it is suitable for autonomous driving systems.

CN120339995APending Publication Date: 2025-07-18CENT SOUTH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510415423.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing traffic sign detection technology has missed and missed detection problems in small-object detection and complex scenarios, and the calculation complexity is high, making it difficult to meet the real-time and accuracy requirements of autonomous driving.

Method used

The traffic sign detection model LTS-YOLOv10, which integrates full-dimensional dynamic convolution and attention-guided feature fusion, extracts multi-scale features through the ODConv module, optimizes feature fusion through EMA-BiFPN, and introduces MPDIoU loss function optimization detection box regression to design a lightweight multi-scale detection head.

Benefits of technology

It significantly improves the detection accuracy of small objects and detection performance in complex scenarios, reduces the computational complexity, maintains the efficient and real-time nature of the model, and is suitable for autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339995A_ABST
    Figure CN120339995A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic sign detection method, which performs traffic sign detection based on full-dimensional dynamic convolution and attention guidance feature fusion, and specifically comprises the following steps: A1, image acquisition: acquiring image data containing traffic signs through an image acquisition device; a2, feature extraction: enabling the image data to pass through an improved LTS-YOLOv10 backbone network; a3, feature fusion: adopting an attention-guided bidirectional feature pyramid network (EMA-BiFPN) to optimize a feature fusion process; a4, traffic sign detection and classification; and A5, result output. And the accuracy of traffic sign detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to traffic sign detection technology in intelligent transportation systems, and particularly to an object detection model based on deep learning, which is applicable to traffic sign recognition and detection in complex environments. Background Art

[0002] Traffic sign detection is a key task in autonomous driving and intelligent transportation systems, and is crucial for improving driving safety and traffic efficiency. Traditional detection methods mainly rely on technologies such as sliding windows and edge computing, but have high computational complexity, slow detection speed, and poor accuracy in complex backgrounds. Although deep learning methods, especially those based on convolutional neural networks (CNNs) and YOLO models, have improved detection accuracy and speed, there are still deficiencies in small object detection and complex scenarios. There is still room for improvement in feature extraction, fusion, and loss function optimization in existing networks, resulting in unsatisfactory detection effects for small objects and objects with complex shapes in practical applications. In addition, traffic sign detection also faces problems such as small sign sizes, environmental light changes, occlusion, weather conditions, and sign wear, further increasing the complexity of detection.

[0003] To address these challenges, researchers have proposed various improvement methods. To address the problem that traffic signs with small pixel areas are difficult to detect, some detectors choose to focus on improving the quality of feature extraction. For example, Kamal et al. combined the feature extraction characteristics of Seg-Net and U-Net to propose SegU-Net, which completed the detection of small signs in videos and improved the detection performance of traffic signs. To address the impact of weather changes and complex background environments on traffic sign detection, complete detection tasks in real scenarios, and improve the robustness of detectors, some scholars choose to enhance the feature representation ability of the neck network. Shen et al. constructed an adaptive pyramid convolution to reduce the inconsistency between features. Liu et al. proposed an attention-driven bilateral FPN, which emphasizes the use of foreground awareness and context information to provide occlusion compensation, enabling the learning of foreground features in complex environments. To make the model adapt to different environments and lighting conditions, Wang et al. proposed a filtering strategy to achieve targeted processing in multiple scenarios by fusing different filters. Karibasappa and Singh proposed an adaptive RBFNN network based on differential evolution, which is a hybrid model that can improve the recognition rate and enhance the robustness of traffic sign recognition. However, there are still problems of missed detection and false detection of small objects in the prior art. Summary of the Invention

[0004] Aiming at the defects of the prior art, the present invention provides a traffic sign detection model LTS-YOLOv10 that can significantly improve the detection accuracy of small targets and the detection performance in complex scenarios, effectively solve the common problems of missed detection and false detection of small targets in the prior art, and improve the detection accuracy while maintaining a lightweight detection model.

[0005] A traffic sign detection method based on all-dimensional dynamic convolution and attention-guided feature fusion for traffic sign detection, specifically including the following steps: A1 Image acquisition: Obtain image data containing traffic signs through an image acquisition device; A2 Feature extraction: The image data passes through an improved LTS-YOLOv10 backbone network, introducing an all-dimensional dynamic convolution (ODConv) module, and using a C2f-ODConv module to replace the original C2f module in YOLOv10 to extract multi-scale feature information; A3 Feature fusion: Use an attention-guided bidirectional feature pyramid network (EMA-BiFPN) to optimize the feature fusion process, add an additional detection head to handle extremely small targets, and form four detection heads; A4 Traffic sign detection and classification: Use an optimized detection head to predict the specific category of traffic signs and their positions in the image. At the same time, introduce the MPDIoU loss function to optimize the regression process of the detection box, reduce the influence of bad gradients while ensuring the quality of the anchor box, and finally generate accurate detection results; A5 Output result: Transmit the detection results to the driving system in real time, provide accurate and timely traffic sign information for driving decisions, and assist the driving system to make correct decisions.

[0006] Optionally, in step A2, after the image is input, it passes through the backbone structure of the LTS-YOLOv10 network for feature extraction. At this stage, first, the ODConv module is used to extract and process the underlying features of the image. The ODConv module includes all-dimensional dynamic convolution; the dynamic convolution operation can be expressed by the following formula:

[0007] y = (α w1 W1 + … + α wn W n ) * x

[0008] where W n represents different convolution kernels, and is the weight corresponding to each convolution kernel. For different inputs, different convolution kernels will be used, and these different convolution kernels are attention-weighted; the feature maps are connected together after passing through two all-dimensional dynamic convolution modules, and then processed by another all-dimensional dynamic convolution module. Each all-dimensional dynamic convolution module consists of all-dimensional dynamic convolution, a batch normalization layer, and a SiLU activation function; the ODConv can be described by the following formula

[0009]

[0010] Wherein: among which α Wi ∈R represents the convolution kernel, Wi represents the attention weight scalar, and α si ∈R k×k , respectively represent three new attention coefficients calculated along the spatial dimension, input channel dimension, and output channel dimension; ⊙ represents an element-wise multiplication operation performed on different dimensions of the convolution kernel space.

[0011] Furthermore, in the step A3, the feature fusion process is completed by EMA - BiFPN, and an additional C2f module for feature fusion is introduced, as well as paths for fusing more features between the P4 layer and the P6 layer and the back end of the feature pyramid network. The attention - guided bidirectional feature pyramid network is adopted, which efficiently extracts features from the shallow network through an efficient multi - scale attention mechanism and guides subsequent network modules to perform feature fusion through a top - down path, thereby generating more discriminative features; in the step A3, EMA - BiFPN introduces learnable weights to determine the importance of different input features. By adding weights to each input feature, the network learns the importance of each input feature, enabling the network to pay more attention to features with more information. At the same time, EMA - BiFPN optimizes cross - scale connections by removing nodes with only one input edge, adding additional edges between input and output nodes at the same level, and treating each bidirectional path as a feature network layer; in the step A3, LTS - YOLOv10 also introduces a small - object feature map of 160×160; in the step A4, the LTS - YOLOv10 model constructs four detection heads, corresponding to the detection of extremely small, small, medium, and large objects respectively; in the step A4, LTS - YOLOv10 introduces the MPDIoU loss function. The MPDIoU loss function comprehensively considers the overlapping area, center - point distance, and shape difference between the predicted box and the ground - truth box. By minimizing this loss function, the model can more accurately adjust the position and shape of the predicted box to make it closer to the real target;

[0012] In the method described above, in the step A4, based on MPDIoU, the loss function is defined as:

[0013] L MPDIoU = 1 - MPDIoU

[0014] In the step A4 described above, the calculation formula of MPDIoU is as follows:

[0015]

[0016] Wherein, A and B respectively represent two arbitrary convex shapes, and w and h respectively represent the width and height of the image. and respectively represent the coordinates of the upper left and lower right corner points of A. and respectively represent the square of the Euclidean distance between the upper left corner of A and the corresponding point of B.

[0017] The beneficial effects of the present invention are as follows: The present invention discloses a traffic sign detection model LTS-YOLOv10 based on full-dimensional dynamic convolution and attention-guided feature fusion, aiming to improve the detection accuracy of small targets and the detection performance in complex scenarios. Image data is acquired through A1 image acquisition, and multi-scale features are extracted by using an improved backbone network and ODConv module through A2 feature extraction. Then, through A3 feature fusion, the feature fusion process is optimized by EMA-BiFPN to enhance the small target detection and complex scene processing capabilities and reduce the computational complexity. Then, through A4 traffic sign detection and classification, an optimized detection head is used to predict the category and position, and the MPDIoU loss function is introduced to optimize the detection box regression. Finally, through A5 output result, the detection result is transmitted to the driving system in real time. Through multi-scale feature fusion, lightweight design, and small target detection optimization, the present invention effectively solves the problems of missed detection and false detection in the prior art, while maintaining the high efficiency and real-time performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is the improved structure diagram of LTS-YOLOv10 of the present invention;

[0019] Figure 2 is the network structure comparison diagram of YOLOv10 and LTS-YOLOv10;

[0020] Figure 3 is the integrated structure diagram of the C2f module and ODConv of the present invention and the internal structure diagram of ODConv;

[0021] Figure 4 is the attention-guided bidirectional feature pyramid network diagram of the present invention;

[0022] Figure 5 is the schematic diagram of the efficient multi-scale attention (EMA) module of the present invention;

[0023] Figure 6 is the bounding box regression diagram of the present invention;

[0024] Figure 7 is the detection effect diagram of LTS-YOLOv10 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings, making the above and other objects, features, and advantages of the present invention clearer. The same reference numerals indicate the same parts in all the drawings. The drawings are not deliberately drawn to scale, with the emphasis on showing the gist of the present invention.

[0026] The traffic sign detection model LTS-YOLOv10 based on full-dimensional dynamic convolution and attention-guided feature fusion proposed by the present invention aims to solve the problem of the decline in detection accuracy of traffic signs due to factors such as weather, light changes, and complex backgrounds in actual scenarios. Specifically, traffic signs may be blurred, unclear, or even completely invisible in rainy, foggy, or strong light conditions, which will seriously affect the performance of detection algorithms. In addition, traditional object detection algorithms often have difficulty accurately identifying small objects (such as distant traffic signs) due to insufficient feature information, thus unable to meet the high requirements for small object detection in autonomous driving. At the same time, complex background interference (such as buildings, pedestrians, other vehicles, etc.) may lead to false detections or missed detections, further affecting the safety of the system.

[0027] The traffic sign detection method based on full-dimensional dynamic convolution and attention-guided feature fusion provided by the present invention is mainly implemented through the following steps:

[0028] Step A1 Image acquisition: In this implementation, the first step in traffic sign detection is image acquisition.

[0029] Image data containing traffic signs is obtained in real time through a camera or other image acquisition devices installed on the vehicle. These image data will be used as the input for subsequent processing and contain traffic sign images under different environmental conditions, such as sunny, rainy, or haze weather.

[0030] Step A2 Feature extraction: After the image data is preprocessed, it is input into the backbone structure of the LTS-YOLOv10 network of the present invention for feature extraction.

[0031] LTS-YOLOv10 is an object detection network based on the YOLO (You Only Look Once) series, combining the characteristics of lightweight design, temporal information, and spatial information, aiming to improve detection accuracy and efficiency, especially in object detection tasks for videos or temporal data. The present invention uses an improved LTS-YOLOv10 network, and its structure is as Figure 1 shown, Figure 2 This is a comparison diagram of the network structures of YOLOv10 and LTS-YOLOv10 of the present invention.

[0032] At this stage, the image is processed through the ODConv module. ODConv (Omni-Dimensional Convolution) is an improved convolution operation designed to enhance the performance of convolutional neural networks (CNNs). It enhances the feature extraction ability by dynamically adjusting multiple dimensions of the convolutional kernel, such as spatial, channel, and input / output dimensions.

[0033] This module combines a full-dimensional dynamic convolution mechanism, which can effectively enhance the model's ability to capture multi-scale targets and complex background features. Compared with traditional convolution modules, the ODConv module can dynamically adjust the weights of the convolutional kernel according to the characteristics of the input image, thus optimizing the feature extraction process, especially having a higher recognition ability for small targets.

[0034] Specifically, as Figure 3 shown, the ODConv module first processes the underlying features of the image through per-channel convolution and pointwise convolution. Then, after being processed by multiple ODConv modules, multi-scale feature information of the image is obtained, which provides rich inputs for subsequent feature fusion. The dynamic convolution operation can be represented by the following formula:

[0035] y = (α w1 W1 + … + α wn W n ) * x

[0036] where W n represents different convolutional kernels, and are the weights corresponding to each convolutional kernel. For different inputs, different convolutional kernels are used, and these different convolutional kernels are attention-weighted.

[0037] Compared with traditional convolution modules, the ODConv module introduces full-dimensional dynamic convolution, which not only improves the detection accuracy but also effectively reduces the computational amount and the number of parameters. Each ODConv module consists of full-dimensional dynamic convolution, a batch normalization layer (Batch Normalization), and a SiLU activation function. Through this modular structure, the feature map is optimized after passing through multiple ODConv modules, and at the same time, the computational complexity is effectively controlled.

[0038] Starting from the definition of dynamic convolution, ODConv can be described by the following formula

[0039]

[0040] where: where α Wi ∈R represents the convolutional kernel, Wi represents the attention weight scalar, α si ∈R k×k , respectively represent three new attention coefficients calculated along the spatial dimension, input channel dimension, and output channel dimension; ⊙ represents an element-wise multiplication operation performed on different dimensions of the convolutional kernel space.

[0041] The ODConv module extends the convolutional mechanism in multiple dimensions, specifically including the dynamic characteristic expansion of the spatial dimension, input channel dimension, output channel dimension, and convolutional kernel dimension. This parallel strategy enables ODConv to learn complementary attention patterns, thereby significantly improving the model's expressive ability and detection performance. In this way, the LTS-YOLOv10 model can more effectively capture the features of small targets while maintaining high detection accuracy and efficiency in complex backgrounds.

[0042] Step A3 Feature Fusion: In the feature fusion process, the present invention introduces the EMA-BiFPN module, which is an attention-guided bidirectional feature pyramid network.

[0043] Refer to Figure 4 and Figure 5 , EMA-BiFPN combines an efficient multi-scale feature fusion mechanism, guiding feature fusion through top-down and bottom-up paths, ensuring the ability to detect small targets in complex backgrounds. EMA-BiFPN (Efficient Multi-scale Attention Bidirectional Feature Pyramid Network) is an efficient multi-scale feature fusion mechanism that combines the attention mechanism and the bidirectional feature pyramid network (BiFPN), aiming to further enhance the feature fusion ability in object detection and segmentation tasks. It enhances the expressive ability of multi-scale features and reduces the computational overhead by introducing the channel attention mechanism and efficient cross-scale connections. In addition, this implementation scheme designs a multi-level feature pyramid structure (including the P4 layer and the P6 layer), optimizes the feature representation, and determines the importance of each input feature through learnable weights, enabling the network to focus on more informative features. The use of EMA-BiFPN greatly improves the robustness and detection accuracy of the model in complex environments, especially when dealing with traffic signs of different scales and shapes.

[0044] Since multi-scale feature fusion can improve the performance and robustness of the network. However, YOLOv8 only uses the PANet structure and improves feature fusion by adding secondary fusion, resulting in a simple bidirectional fusion, while many convolutions lead to the degradation of semantic details related to small targets. To enable the model to better detect targets, the present invention proposes an attention-guided bidirectional feature pyramid network, the structure of which is as Figure 4As shown. Compared with the feature pyramid network in the original YOLOv8, we additionally introduced a C2f module for feature fusion, as well as paths for the P4 layer and P6 layer to fuse more features with the backend of the feature pyramid network. An attention-guided bidirectional feature pyramid network was designed to efficiently extract features from the shallow network through an efficient multi-scale attention mechanism and guide subsequent network modules for feature fusion through a top-down path, thereby generating more discriminative features. Since the attention-guided bidirectional feature pyramid network has a more efficient feature generation ability, we maintained the same advantage in the number of model parameters by reducing the number of feature maps generated by the C2f module.

[0045] Step A4 Traffic Sign Detection and Classification: After feature extraction and fusion, the input data enters multiple detection heads for object detection and classification.

[0046] LTS-YOLOv10 designed four detection heads for detecting extremely small, small, medium, and large objects respectively. This multi-scale detection head design ensures wide coverage of traffic signs of different sizes.

[0047] As Figure 6 shown, for the bounding box regression problem, in order to further improve the detection performance of the model, LTS-YOLOv10 introduced the MPDIoU loss function. This loss function plays a key role in optimizing the regression process of the detection box. The MPDIoU loss function comprehensively considers the overlapping area, center point distance, and shape difference between the predicted box and the ground truth box. By minimizing this loss function, the model can more accurately adjust the position and shape of the predicted box to make it closer to the real target. This optimization not only improves the detection accuracy but also significantly enhances the recall rate of the model for small objects and objects with complex shapes. The MPDIoU loss function avoids the interference of outliers during the training process by reducing the gradient influence of low-quality samples, thereby ensuring the effectiveness of each feature focus and enabling the network to more accurately focus on traffic signs in the image.

[0048] The calculation formula of MPDIoU is as follows:

[0049]

[0050] where A and B represent two arbitrary convex shapes respectively, w and h represent the width and height of the image respectively, and represent the coordinates of the upper left and lower right points of A respectively, and respectively represent the square of the Euclidean distance between the upper left corner of A and the corresponding point of B. In this way, MPDIoU can more precisely adjust the position of the bounding box, especially when dealing with objects of complex shapes, thereby improving the detection accuracy.

[0051] Based on MPDIoU, the loss function is defined as:

[0052] L MPDIoU = 1 - MPDIoU

[0053] This loss function further optimizes the regression process of the detection box by minimizing the difference between the predicted box and the ground truth box, ensuring that the model can converge stably and achieve the best performance during training.

[0054] Output result of step A5: After the processing of the foregoing steps, the finally generated detection results will be transmitted in real time to the driving decision-making system, providing accurate traffic sign information for the autonomous driving system. Based on this information, the driving decision-making system can timely adjust the driving strategy to ensure driving safety.

[0055] Experimental result analysis

[0056] 1) Datasets

[0057] To evaluate the performance of the traffic sign detection model proposed in the present invention, we used three publicly available traffic sign datasets: CCTSDB, TT100K, and DFG. CCTSDB was created by Changsha University of Science and Technology and contains 17,856 data samples, covering six extreme weather conditions and various environmental changes. TT100K was jointly released by Tsinghua University and Tencent University and contains approximately 100,000 images with 30,000 traffic signs annotated, covering 45 categories, showing street scenes under various lighting and weather conditions. The DFG dataset contains 7,000 images, covering 200 categories of traffic signs, showing traffic scenes in Slovenian cities and villages. These datasets provide diverse scenarios and conditions, ensuring the detection ability of the model in different environments.

[0058] 2) Experimental platform and parameters

[0059] All experiments were conducted on a unified platform equipped with an Intel i5-12490F CPU at 3 GHz, and the graphics processing was handled by an RTX 3090 GPU, providing 24 GB of video memory. The deep learning environment was built based on PyTorch 1.12.1. The proposed model adopted an end-to-end training method, using the SGD optimizer with a momentum set to 0.937 and a weight decay of 0.0005. The experiments were carried out on the TT100K, CCTSDB, and DFG datasets. For all datasets, the model was trained for 200 epochs using a single GPU with the same hyperparameters. The image size was set to 640×640 pixels, the initial learning rate was 0.02, and a cosine annealing learning rate scheduler was used to adjust the learning rate to 0.0001 at the end of training. The experimental results were evaluated using the COCO evaluation metrics.

[0060] 3) Comparative experiments of traffic sign detection models

[0061] LTS-YOLOv10 performs excellently under multiple datasets and conditions. On the CCTSDB, TT100K, and DFG datasets, it outperforms existing advanced models in terms of the mean average precision (mAP), especially in the detection of small objects and under extreme weather conditions. For example, on the CCTSDB dataset, its mAP is 83.6%, which is 11.4% higher than that of CAAD-3Net. On the TT100K dataset, its mAP is 4.3% higher than that of CAAD-Net, while maintaining a faster inference speed and reducing the model parameters. On the DFG dataset, LTS-YOLOv10 demonstrates strong generalization ability and exceeds UCN-YOLOv5 in all metrics. Although it is slightly inferior in terms of precision (P), its recall rate (R) has increased significantly, showing its robustness in detecting traffic signs in complex environments. Overall, LTS-YOLOv10 provides significant advantages in terms of performance and efficiency for practical traffic sign detection applications.

[0062] Table 1 Comparison of the detection evaluation results of different models on the TT100K, CCTSDB, and DFG test sets.

[0063]

[0064]

[0065] 4) Ablation experiments

[0066] We conducted ablation experiments on the CCTSDB dataset to analyze the contributions of different components in the LTS-YOLOv10 model. The results are shown in Table 2. Specifically, we conducted individual and combined tests on three modules: ODConv, EMA-BiFPN, and MPDIoU. The results are shown in Table 2. Compared with the baseline YOLOv10 (mAP of 79.8%, F1 of 78.3%, and FPS of 48.2), after introducing ODConv (Model 1), the mAP increased to 81.9% and the F1 increased to 79.5%, but the computational cost increased slightly; replacing it with EMA-BiFPN (Model 2) significantly reduced the number of parameters while maintaining the accuracy basically unchanged and increased the FPS to 60.3, showing good lightweight characteristics; the addition of MPDIoU (Model 3) effectively improved the regression accuracy, making the F1 reach 79.1%. In the comparison of multi-module combinations, Model 4 and Model 5 integrated ODConv+EMA-BiFPN and the three-module joint scheme respectively, both achieving further improvements in accuracy. Among them, the mAP of Model 5 reached 82.0%. The final model (Model 7) achieved the best performance after integrating all modules, with the mAP increasing to 83.6% and the F1 increasing to 81.0%, while maintaining a high inference speed (52.8 FPS) and a moderate number of parameters (25.6M), verifying the significant improvement of each module in maintaining efficiency and detecting accuracy.

[0067] Table 2 Comparison of ablation experiment results on the CCTSDB dataset.

[0068]

[0069] 5) Comparison of detection results

[0070] To demonstrate the detection ability of LTS-YOLOv10, Figure 7 The detection effect diagram of LTS-YOLOv10 is shown. The green border represents the detection results of LTS-YOLOv10 on three datasets, while the blue box shows the details of a specific area. It can be clearly seen that LTS-YOLOv10 successfully detected most traffic signs, especially showing excellent performance in detecting small-scale targets. However, there are still some cases of detection failure, which are marked in red and yellow. There are multiple reasons for the errors: signs at a long distance appear as very small areas in the image, resulting in insufficient feature information; long-term exposure to the environment causes signs to fade, distort, or be blocked, increasing the difficulty of detection; the visual similarity between some categories causes confusion in complex backgrounds; data imbalance limits the generalization ability of the model for minority categories. Although LTS-YOLOv10 shows strong overall performance, these challenges indicate that there is still room for improvement in enhancing the robustness against environmental changes and improving the detection accuracy of small targets.

[0071] From the above analysis, it can be seen that the present invention has the following advantages:

[0072] Multi-scale feature fusion: Traditional traffic sign detection algorithms are prone to missed detection or false detection when facing complex backgrounds, lighting changes, and small target detection. The present invention improves the robustness of the model by using multi-scale feature fusion and generating high-resolution feature layers, enabling it to work stably in complex real-world scenarios.

[0073] Lightweight and high efficiency: Traffic sign detection algorithms in the prior art often have a large number of parameters and high computational complexity, making it difficult to implement in embedded devices and resource-constrained environments. The present invention significantly reduces the number of parameters and computational amount of the model and improves the computational efficiency by introducing the ODConv module and using full-dimensional dynamic convolution, while maintaining high detection accuracy.

[0074] Optimized design for small target detection: Aiming at the problem of small target traffic sign detection, the present invention specifically optimizes the detection performance of small targets by adding a 160×160 prediction feature map. This design effectively improves the accuracy of the model in detecting small traffic signs and solves the common problems of missed detection and false detection in the prior art.

[0075] Improved detection performance in complex scenarios: EMA-BiFPN optimizes the feature fusion process, enabling the model to perform well in complex background and multi-scale target detection. It efficiently extracts features from the shallow network through an efficient multi-scale attention mechanism and guides subsequent network modules for feature fusion through a top-down path, thereby generating more discriminative features.

[0076] Optimized detection box regression: The introduced MPDIoU loss function optimizes the regression process of the detection box by comprehensively considering the overlapping area, center point distance, and shape difference between the predicted box and the ground truth box, improving the accuracy of the detection results.

[0077] Real-time detection ability: While maintaining high detection accuracy, the computational efficiency of the model is improved, meeting the requirements of real-time detection and being applicable to practical applications in autonomous driving and intelligent transportation systems.

[0078] Numerous specific details have been set forth in the above description to facilitate a full understanding of the present invention. However, the above description is only a preferred embodiment of the present invention, and the present invention can be implemented in many other ways different from those described herein. Therefore, the present invention is not limited by the specific implementations disclosed above. At the same time, any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention. All simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A traffic sign detection method, characterized in that, Traffic sign detection based on full-dimensional dynamic convolution and attention-guided feature fusion specifically includes the following steps: A1 Image acquisition: Image data containing traffic signs is obtained through an image acquisition device; A2 Feature extraction: The image data passes through an improved LTS-YOLOv10 backbone network, introducing an Omni-Dimensional Dynamic Convolution (ODConv) module, and using a C2f-ODConv module to replace the original C2f module in YOLOv10 to extract multi-scale feature information; A3 Feature fusion: An attention-guided bidirectional feature pyramid network (EMA-BiFPN) is used to optimize the feature fusion process, adding an additional detection head to handle extremely small targets, forming four detection heads; A4 Traffic sign detection and classification: Optimally designed detection heads are used to predict the specific categories of traffic signs and their positions in the image. At the same time, the MPDIoU loss function is introduced to optimize the regression process of the detection box, reducing the influence of adverse gradients while ensuring the quality of the anchor box, and finally generating accurate detection results; A5 Output result: The detection results are transmitted to the driving system in real time, providing accurate and timely traffic sign information for driving decisions to assist the driving system in making correct decisions.

2. The traffic sign detection method according to claim 1, characterized in that, In step A2, after the image is input, it undergoes feature extraction through the backbone structure of the LTS-YOLOv10 network. At this stage, the underlying features of the image are first extracted and processed by the ODConv module, and the ODConv module contains Omni-Dimensional Dynamic Convolution.

3. The traffic sign detection method according to claim 2, wherein The dynamic convolution operation can be represented by the following formula: y = (α w1 W1 + … + α wn W n ) * x Among them, W n represents different convolutional kernels, and is the weight corresponding to each convolutional kernel. For different inputs, different convolutional kernels will be used, and these different convolutional kernels are attention-weighted.

4. The traffic sign detection method according to claim 3, wherein The feature map is connected together after passing through two Omni-Dimensional Dynamic Convolution modules, and then processed by another Omni-Dimensional Dynamic Convolution module. Each Omni-Dimensional Dynamic Convolution module consists of Omni-Dimensional Dynamic Convolution, a batch normalization layer, and a SiLU activation function.

5. The traffic sign detection method according to claim 4, wherein The ODConv can be described by the following formula where: where α Wi ∈R represents the convolution kernel, Wi represents the attention weight scalar, α si ∈R k×k , respectively represent three new attention coefficients calculated along the spatial dimension, the input channel dimension, and the output channel dimension; ⊙ represents the element-wise multiplication operation performed on different dimensions of the convolution kernel space.

6. The traffic sign detection method according to claim 1, characterized in that In step A3, the feature fusion process is completed by EMA-BiFPN. An additional C2f module for feature fusion is introduced, as well as paths for P4 layer and P6 layer to fuse more features with the backend of the feature pyramid network. The attention-guided bidirectional feature pyramid network is used to efficiently extract features from the shallow network through an efficient multi-scale attention mechanism and guide subsequent network modules to perform feature fusion through a top-down path, thereby generating more discriminative features.

7. The traffic sign detection method according to claim 6, wherein In step A3, EMA-BiFPN introduces learnable weights to determine the importance of different input features. By adding weights to each input feature, the network learns the importance of each input feature, enabling the network to pay more attention to features with more information. At the same time, EMA-BiFPN optimizes cross-scale connections by removing nodes with only one input edge, adding additional edges between input and output nodes at the same level, and treating each bidirectional path as a feature network layer.

8. The traffic sign detection method according to claim 7, wherein, In step A3, LTS-YOLOv10 also introduces a small target feature map of 160×160.

9. The traffic sign detection method according to claim 7, wherein In the step A4, the LTS-YOLOv10 model constructs four detection heads, corresponding to the detection of extremely small, small, medium, and large targets respectively.

10. The traffic sign detection method according to claim 9, wherein, In the step A4, the LTS-YOLOv10 introduces the MPDIoU loss function. The MPDIoU loss function comprehensively considers the overlapping area, the distance between the center points, and the shape difference between the predicted bounding box and the ground truth bounding box. By minimizing this loss function, the model can more accurately adjust the position and shape of the predicted bounding box to make it closer to the real target. In the method described above, in the step A4, based on MPDIoU, the loss function is defined as: L MPDIoU = 1 - MPDIoU In the step A4, the calculation formula of MPDIoU is as follows: Wherein, A and B respectively represent two arbitrary convex shapes, and w and h respectively represent the width and height of the image, and respectively represent the coordinates of the upper left corner and the lower right corner points of A, and respectively represent the square of the Euclidean distance between the upper left corner of A and the corresponding point of B.

Citation Information

Cited By

  • Real-time traffic sign target detection method based on YOLOv10 improvement

    CN121074849A