Complex environment-oriented multi-scale adaptive dynamic feature fusion traffic sign detection method
By improving the YOLOv8 model, the introduction of CPAM, TFE and DZSF modules and Wise_MPDIoU loss function, the performance bottleneck of detection of small objects and dense objects in complex environments is solved, and efficient and accurate traffic sign detection is achieved.
Patent Information
- Application Number
- CN202510619976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
AI Technical Summary
Existing object detection models have performance bottlenecks when dealing with small object detection, dense object separation and multi-scale object recognition in complex environments, especially in edge computing devices with limited computing resources, which are difficult to meet real-time detection requirements.
By improving the YOLOv8 model, the CPAM module was introduced to replace the C2f module, the TFE and DZSF modules were introduced to reconstruct the neck network, and the Wise_MPDIoU loss function was used to optimize the feature fusion and loss function, and improve the multi-scale feature fusion capability and robustness of the model.
It significantly improves the accuracy and robustness of traffic sign target detection, especially in the detection of small objects and dense objects, and is suitable for real-time object detection tasks in complex industrial environments.
Smart Images

Figure CN120510592A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and target detection, and in particular to a multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments. Background Art
[0002] With the advancement of industrial automation and intelligence, accurate and efficient object detection technology has become increasingly important in various fields, especially in production lines, warehouse logistics, and intelligent monitoring environments. In these application scenarios, object detection must address various complex challenges, such as varying lighting, object occlusion, diverse object scales, and the recognition of densely populated objects. These challenges make traditional detection methods based on manual features difficult to adapt to the demands of these complex environments.
[0003] In recent years, deep learning techniques, particularly convolutional neural networks (CNNs), have made significant progress in the field of object detection. Despite this, existing detection models still face challenges in detecting small objects, separating densely packed objects, and recognizing multi-scale objects. While some advanced object detection models, such as the YOLO series, have achieved promising results in some applications, they still face significant performance bottlenecks when handling multi-scale, multi-object detection tasks, as well as those with complex backgrounds. This is particularly true on edge computing devices with limited computing resources, where performance often fails to meet real-time detection requirements.
[0004] In complex industrial environments, existing object detection methods often struggle to balance accuracy and computational efficiency, especially when detecting small, densely packed components and markers. In particular, in real-time monitoring and predictive maintenance scenarios, efficient, accurate detection methods that can run on resource-constrained devices are needed to ensure rapid system response and decision-making.
[0005] This invention aims to address these technical challenges by introducing a multi-scale adaptive feature fusion mechanism, an innovative dynamic sampling method, and an optimized loss function to propose an efficient object detection method. This method can effectively improve the model's detection capabilities for small objects, multi-scale objects, and densely packed objects, and is particularly suitable for real-time object detection tasks in complex industrial environments. By improving on the YOLOv8 model, this invention provides a highly accurate object detection solution that can run efficiently on edge computing devices, providing strong technical support for fields such as industrial automation and intelligent monitoring. Summary of the Invention
[0006] Purpose of the invention: In response to the problems in the background technology, the present invention provides a multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments. By improving the YOLOv8 model, including introducing the CPAM module to improve the C2f module, and introducing the TFE module and DZSF module into the neck network, the performance of traffic sign target detection is significantly improved.
[0007] Technical solution: The present invention discloses a multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments, comprising the following steps:
[0008] Step (1) Select a suitable traffic sign dataset and preprocess it, then divide the dataset into training set, test set and validation set according to a certain ratio;
[0009] Step (2) constructing a CPADM-YOLO model, which is improved on the basis of the YOLOv8 model. The backbone network proposes a C2f-CPAM module to replace the C2f module in YOLOv8; the neck network introduces the TFE module and the DZSF module to reconstruct a new neck network to enhance feature fusion and multi-scale target detection capabilities;
[0010] Step (3) uses the processed data set to train the CPADM-YOLO model. After the training is completed, the performance of the model is evaluated using the test set, the test results are obtained, and the performance of the model is comprehensively analyzed through various evaluation indicators;
[0011] Step (4) uses the trained CPADM-YOLO model to detect traffic signs on the image to be processed.
[0012] Furthermore, the preprocessing process of the data set in step (1) includes data cleaning and formatting. The specific operation is to perform data cleaning through the Python language, first deleting the image files without labels, and then removing other labels in the original data set except for category and location information; then, converting the processed data into COCO format, and finally converting the COCO format data into txt file format.
[0013] Furthermore, the C2f-CPAM module replaces the Bottleneck module in the C2f module by designing a CPAM module. The specific operation of the CPAM module is as follows:
[0014] First, the input feature map is preliminarily processed by a 3×3 conv1 layer; then, the feature map output by the conv1 layer is divided into two parts, each containing an equal number of channels; one part is processed by a 5×5 conv2 layer and divided into two parts again; a part of the conv2 layer is processed by a 7×7 conv3 layer; the feature map processed by the conv3 layer is fused with the other part of the conv2 layer and the other part of the conv1 layer through a splicing operation to form a multi-scale feature representation; the spliced feature map is passed through a 1×1 conv4 layer for channel adjustment and is residually connected with the original input feature map through an addition operation to form the final output feature map.
[0015] Furthermore, the new neck network includes a pair of TFE modules and a DZSF module as well as multiple C2f modules and multiple Conv modules; the features output by the second, third, and fourth C2f-CPAM modules of the backbone network are P3, P4, and P5 layer feature maps;
[0016] The output feature map of the SPPF module of the backbone network is processed by the first Conv module, and the feature map of the P4 layer and the feature map of the P3 layer are processed by the second Conv module and then enter the first TEE module;
[0017] The features output by the first TEE module are processed by the first C2f module, the fourth Conv module, and the output feature map. The feature map output by the first C2f-CPAM module of the backbone network is processed by the third Conv module and the P3 layer feature map is then input into the second TEE module.
[0018] The output features of the second TEE module pass through the second C2f module and the fifth Conv module in sequence, and are concat-processed with the output features of the fourth convolution module before being input into the third C2f module and the sixth Conv module in sequence.
[0019] The output features of the sixth Conv module and the output features of the first Conv module are concat-processed and then input into the fourth C2f module;
[0020] The feature maps of the P3, P4, and P5 layers are input to the DZSF module and then ADD processed with the second C2f module and the feature maps output by the third and fourth C2f modules. Figure 1 Enter the detection head module for detection.
[0021] Furthermore, the specific operations of the DZSF module are as follows:
[0022] Adjust the number of channels of feature maps of each scale through convolution operations so that they have the same number of channels;
[0023] A dynamic sampling mechanism is introduced to spatially adjust the feature maps of P4 and P5 at different scaling ratios to align them with the feature map of P3. After dynamic sampling, the feature maps of the three scales will be expanded into 3D tensors;
[0024] By canceling the squeezing operation to expand it into a 4D feature map, splicing it along the depth dimension, a 3D fusion feature map containing multi-scale information is synthesized;
[0025] Finally, the feature map is processed by 3D convolution, batch normalization and activation function to output the final fused feature map, completing the effective perception and detection accuracy enhancement of targets of different scales.
[0026] Furthermore, the dynamic sampling mechanism redefines the upsampling process from the perspective of point sampling and introduces content-aware sampling position generation as follows:
[0027] Given an input feature map And upsampling ratio s, construct the initial sampling grid Follow the "bilinear initialization" principle:
[0028]
[0029] Among them, the meshgrid operation generates a regular grid, and the transpose operation ensures that the coordinates are arranged correctly;
[0030] Second, content-aware offsets are generated by linear projection.
[0031]
[0032] Or use the dynamic range factor for even greater flexibility:
[0033]
[0034] Among them, σ represents the sigmoid activation function, and the range factor prevents the sampling points from overlapping and ensures the correct expression of the boundary area; then, the sampling set S is composed of the grid and offset composition:
[0035]
[0036] Finally, the input features are resampled by the sampling set to obtain the upsampled features
[0037]
[0038] Furthermore, the dynamic sampling mechanism also introduces a grouping processing mechanism to divide the features into g groups, so that each group of features uses an independent sampling position.
[0039] Furthermore, the CPADM-YOLO model uses Wise_MPDIoU as the loss function, which is expressed as follows:
[0040]
[0041] The definition of MPDIoU is as follows:
[0042]
[0043] in, Indicates the points at the upper left and lower right corners of the predicted box and the real box, Indicates the distance between corresponding points, w and h represent the width and height of the bounding box, and WIoU is defined as follows:
[0044]
[0045] Among them, b i Represents the coordinates of i object boxes, g i represents the coordinates of the ground truth box of the i-th object, w i Indicates the weight value.
[0046] Beneficial effects:
[0047] The present invention significantly improves the performance of traffic sign target detection by introducing the CPAM module, TFE module, DZSF module and Wise-MPDIoU loss function. Specifically: 1) Backbone network part: The CPAM module not only improves the computational efficiency but also significantly enhances the feature expression ability of the model by introducing multi-scale feature extraction and enhanced feature fusion. Through partial convolution operations, redundant calculations are reduced, especially when processing complex targets, showing higher efficiency. The module adopts a fusion mechanism of 1x1 convolution and residual connection, which not only retains the rich information of the original features but also effectively fuses multi-scale features, thereby further improving the target detection accuracy. 2) The introduction of TFE and DZSF modules can adaptively adjust the spatial scale of feature maps of different scales and optimize the spatial matching between scales. With the help of the dynamic sampling mechanism, the feature fusion ability of multi-scale targets is enhanced, especially when detecting small objects, which significantly improves the detection accuracy. At the same time, through the effective fusion of multi-scale features and dynamic spatial adjustment, the robustness of the model is improved in complex backgrounds and multi-object scenes. This module effectively balances the accuracy and inference speed in real-time detection tasks, especially in small object detection tasks such as traffic signs, showing stronger detection capabilities and higher accuracy, and has broad practical application value. 3) Loss function part: The Wise_MPDIoU loss function optimizes the limitations of the traditional loss function of YOLOv8 by introducing a new multi-object box metric. This loss function not only improves the accuracy of bounding box positioning, but also optimizes the accuracy of multi-object detection by introducing weights, enhances the model's adaptability to different targets, and further improves the model's performance in complex scenes. Through the above innovations, the present invention significantly improves the accuracy, robustness and adaptability of traffic sign target detection to multi-scale, small objects and dense objects, demonstrating its excellent performance in complex environments, and has broad practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a target detection flow chart of the present invention;
[0049] Figure 2 This is a diagram of the C2f network structure after the backbone network of the present invention is improved;
[0050] Figure 3 This is a network structure diagram of the CPAM module in the present invention;
[0051] Figure 4 This is the overall network structure diagram of the improved YOLOv8 of the present invention;
[0052] Figure 5 FIG1 is a structural diagram of the TFE module in the AFSDM structure in the neck network of the present invention;
[0053] Figure 6 FIG1 is a structural diagram of the DZSF module in the AFSDM structure in the neck network of the present invention;
[0054] Figure 7 Schematic diagram of the dynamic sampling mechanism in the present invention
[0055] Figure 8 This is a display diagram of part of the detection of the validation set data in the embodiment;
[0056] Figure 9 This is a dataset category picture in the embodiment;
[0057] Figure 10 This is a detection example of YOLOv8 in the traditional conventional technology test set;
[0058] Figure 11 This is a detection example in the test set of the present invention. DETAILED DESCRIPTION
[0059] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0060] The embodiment of the present invention provides a multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments, comprising the following steps:
[0061] Step 1: Build and preprocess the dataset;
[0062] Step 1.1: This example uses the open-source TT100K traffic sign dataset provided by Tsinghua University as the basis for data construction. Because the original dataset contained unlabeled images and a severely uneven distribution of categories, we used a Python script to clean the dataset, removing unlabeled images and filtering for images with more than 100 instances per category. Ultimately, we selected 45 categories and a total of 9,738 images.
[0063] Step 1.2: The labels for the original dataset are stored in a JSON file. This example requires the data to be in a txt format, containing the category labels and location information. Therefore, we first use a Python script to parse the JSON file, iterating through each sublist, extracting the corresponding category labels, and numbering them. Next, we generate an XML file and extract the bounding box information from it. Finally, we convert this information into the YOLO format and generate a corresponding txt file based on the box coordinates for each image.
[0064] Step 1.3: Based on the traffic sign dataset processed in step 1.1, divide it into training set, validation set, and test set in a ratio of 7:2:1, containing 6793 images for training, 1949 images for validation, and 996 images for testing.
[0065] Step 2: Model building: The enhanced data set is input into the improved YOLOv8 network model (i.e., CPADM-YOLO model) of this embodiment for training. The improved complete model structure is as follows: Figure 4 As shown in the figure, the target detection network model is established. The CPADM-YOLO model is an improvement on YOLOv8. The backbone network adopts the C2f-CPAM module to replace the C2f module in YOLOv8. The C2f-CPAM module introduces the CPAM module to replace the Bottleneck structure in the C2f module, improving the network's feature extraction capabilities. The neck network introduces the TFE and DZSF modules and redesigns the neck network structure to enhance feature fusion and multi-scale target detection capabilities. The loss function introduces a new loss function to replace the traditional loss function in YOLOv8, further optimizing the model's training results and target detection performance.
[0066] Step 2.1: The backbone network is responsible for feature extraction. Drawing on the innovative ideas of CVPR2020-GhostNet and CVPR2024-FasterNet, a new structure called CPAM is designed to replace the Bottleneck module in YOLOv8. The specific structure is as follows: Figure 3 shown.
[0067] The CPAM module utilizes efficient partial convolution operations, a method that significantly reduces redundant computation. Compared to the traditional approach of performing the same operation on all channels, partial convolution reduces unnecessary computation, significantly improving computational speed. Furthermore, the final 1x1 convolution layer of the CPAM module is used to fuse feature information from different scales, adding the input features to the processed features via a residual connection. This approach not only preserves the key information of the original features but also introduces new multi-scale features, significantly enhancing the model's expressive power and object perception capabilities.
[0068] Specifically, the CPAM module first performs preliminary processing on the input feature map through a 3×3 convolutional layer (conv1). The feature map output by conv1 is then divided into two parts, each containing an equal number of channels. Next, one part is processed by a 5×5 convolutional layer (conv2) and divided into two parts again. Then, one part of conv2 is processed by a 7×7 convolutional layer (conv3) to extract features with a larger receptive field. Finally, the processed feature map from conv3 is fused with the other part of conv2 and the other part of conv1 through a splicing operation to form a multi-scale feature representation. Finally, the spliced feature map passes through a 1×1 convolutional layer (conv4) for channel adjustment and is residually connected to the original input feature map through an addition operation. This process retains the key information of the input while effectively preventing gradient vanishing, ensuring the effective transmission of information and enhancing the feature expression capability, thereby generating the final output feature map.
[0069] In step 2.2, during the construction of the neck structure of the CPADM-YOLO model, the ASF-YOLO network structure was innovatively introduced and cleverly improved, employing a dynamic sampling mechanism. This component primarily consists of two modules: the TFE module and the DZSF module. The TFE module significantly improves the model's feature representation capabilities by fusing feature maps from three different scales. Specifically, the low-resolution feature maps are first downsampled using max pooling and average pooling, while the high-resolution feature maps are resized to an intermediate scale through interpolation. This allows the TFE module to generate a fused feature map containing multi-scale information, effectively fusing information from different scales and improving the comprehensive representation of multi-scale features. The DZSF module then deeply fuses feature maps from three different scales (P3, P4, and P5). This module incorporates a dynamic sampling mechanism, spatially adjusting the P4 and P5 feature maps at different scaling ratios (e.g., 2x and 4x) to align them with the P3 feature map. The dynamic sampling mechanism ensures spatial alignment between feature maps of different scales by optimizing the spatial sampling of the feature maps. Finally, the feature maps are spliced in the depth dimension to form a 3D feature map, which is further fused and processed through 3D convolution, thereby effectively integrating multi-scale feature information and improving the model's perception and feature representation capabilities for targets of different scales.
[0070] The features output by the second, third, and fourth C2f-CPAM modules of the backbone network are the P3, P4, and P5 layer feature maps. The feature map output by the SPPF module of the backbone network is processed by the first Conv module, and the P4 layer feature map and the P3 layer feature map are processed by the second Conv module and then enter the first TEE module.
[0071] The features output by the first TEE module are processed by the first C2f module, the fourth Conv module, the feature map output by the first C2f-CPAM module of the backbone network, the feature map processed by the third Conv module, and the P3 layer feature map, which are then input into the second TEE module.
[0072] The output features of the second TEE module pass through the second C2f module and the fifth Conv module in sequence, and are concat-processed with the output features of the fourth convolution module before being input into the third C2f module and the sixth Conv module in sequence.
[0073] The output features of the sixth Conv module and the output features of the first Conv module are concat-processed and then input into the fourth C2f module.
[0074] The feature maps of the P3, P4, and P5 layers are input to the DZSF module and then ADD processed with the second C2f module and the feature maps output by the third and fourth C2f modules. Figure 1 Enter the detection head module for detection.
[0075] The TFE module integrates information of different scales by downsampling and upsampling large-scale and small-scale feature maps, thereby improving the network's ability to fuse multi-scale features. Figure 5 As shown. First, the number of channels of each scale feature map is adjusted through the convolution operation to make it consistent with the number of channels of the main scale feature map. For the large-scale feature map (Large), the number of channels is adjusted through the convolution module and then a hybrid pooling structure (including maximum pooling and average pooling) is applied for downsampling, which helps to retain the effectiveness and details of the high-resolution feature map. For the small-scale feature map (Small), the number of channels is also adjusted through the convolution module, and the nearest neighbor interpolation method is used to upsample it to the size of the larger feature map to ensure that the local features in the low-resolution image are retained and prevent the loss of small target feature information. Finally, after downsampling and upsampling operations, the feature maps of the three scales are spliced in the channel dimension to form a fused multi-scale feature map, thereby improving the network's feature expression ability and multi-scale target perception ability. The output results of the TFE module are as follows:
[0076] F TFE =Concat(F large ,F medium ,F small )
[0077] Among them, F TFE Represents the feature map output by the TFE module, F large ,F medium ,F smallRepresent feature maps of large, medium and small sizes respectively.
[0078] The DZSF module enhances the network’s ability to perceive multi-scale information by processing and fusing feature maps of different scales. Figure 6 As shown in the figure, first, the number of channels of the feature maps of each scale is adjusted through the convolution operation so that they have the same number of channels; secondly, the DZSF module introduces a dynamic sampling mechanism to spatially adjust the feature maps of P4 and P5 with different scaling ratios (such as 2 times and 4 times) so that they are aligned with the feature map of P3. After dynamic sampling, the feature maps of the three scales (P3, P4, P5) will be expanded into 3D tensors, and then, they are expanded into 4D feature maps by unsqueezing the operation, and then they are spliced along the depth dimension to synthesize a 3D fused feature map containing multi-scale information; finally, the feature map is processed by 3D convolution, batch normalization and activation function to output the final fused feature map, completing the effective perception and detection accuracy enhancement of targets of different scales.
[0079] Traditional upsampling methods (such as nearest neighbor interpolation and bilinear interpolation) use fixed rules to reconstruct features, ignoring the semantic information of features. The dynamic sampling mechanism proposed in this paper redefines the upsampling process from the perspective of point sampling and introduces content-aware sampling position generation, such as Figure 7 shown.
[0080] Given an input feature map and upsampling ratio s, the dynamic sampling process can be expressed as follows: First, construct the initial sampling grid Follow the "bilinear initialization" principle:
[0081]
[0082] The meshgrid operation generates a regular grid, and the transpose operation ensures that the coordinates are correctly arranged. This initialization ensures that when the offset is zero, the result is equivalent to standard bilinear interpolation.
[0083] Second, content-aware offsets are generated by linear projection.
[0084]
[0085] Or use the dynamic range factor for even greater flexibility:
[0086]
[0087] Here, σ represents the sigmoid activation function, and the range factor (0.25 or a dynamically generated value) prevents the sampling points from overlapping and ensures the correct representation of the boundary area.
[0088] Then, the sampling set S is composed of the grid and offset composition:
[0089]
[0090] Finally, the input features are resampled by the sampling set to obtain the upsampled features
[0091]
[0092] Furthermore, we introduce a grouping mechanism, dividing features into g groups (typically g = 4), with each group using independent sampling locations to further improve expressiveness. This dynamic sampling process learns the optimal sampling point at each location, adapting upsampling to different semantic content, thereby generating more accurate high-resolution feature representations. This is particularly suitable for detecting detailed objects such as traffic signs. The entire process utilizes optimized tensor operations to ensure efficient computation while maintaining low memory usage.
[0093] In step 2.3, the YOLOv8n model uses the CIoU loss function by default. The CIoU loss function improves on the traditional IoU loss function. It not only considers the overlapping area between the predicted box and the ground-truth box, but also introduces the distance between the center points of the predicted box and the ground-truth box, as well as the aspect ratio information between the two. This allows for a more comprehensive evaluation of the degree of match between the predicted box and the ground-truth box in the bounding box regression task. Its calculation formula is as follows:
[0094]
[0095] Among them, v is a parameter used to measure the consistency of aspect ratio, α is a weight function, B and B gt They represent the center points of the predicted box and the real box respectively, ρ represents the Euclidean distance, c represents the diagonal distance of the minimum outer rectangle of the predicted box and the real box, and w gt Indicates the width of the real box (real target bounding box), h gt Indicates the height of the real box, w pred and h pred Indicates the width and height of the prediction box.
[0096] However, the CIoU loss function has certain limitations. The definition of aspect ratio is relatively ambiguous, making it difficult to further optimize some high-quality regression samples during training. This limitation also leads to an imbalance between positive and negative samples, where the contributions of positive samples (predicted boxes with a high degree of match to the ground-truth box) and negative samples (predicted boxes with a low degree of match to the ground-truth box) differ significantly during training, affecting the overall performance of the model.
[0097] To address these issues, this method replaces the CIoU loss function with the Wise_MPDIoU loss function. The Wise_MPDIoU loss function effectively addresses the potential bias of traditional IoU by weighting the area between the predicted box and the ground-truth box, thereby alleviating the imbalance between positive and negative samples.
[0098] Wise_MPDIoU is used as the loss function, and its expression is as follows:
[0099]
[0100] The definition of MPDIoU is as follows:
[0101]
[0102] in, Indicates the points at the upper left and lower right corners of the predicted box and the real box, Indicates the distance between corresponding points, w and h represent the width and height of the bounding box. The definition of WIoU is as follows:
[0103]
[0104] Among them, b i Represents the coordinates of i object boxes, g i represents the coordinates of the ground truth box of the i-th object, w i Indicates the weight value.
[0105] Step 3: Using a computer running Ubuntu 18.04, equipped with an NVIDIA GeForce RTX-4090 graphics card with 24GB of video memory, we selected the PyTorch deep learning framework, version 2.2.2, as the primary development tool. We also used Python interpreter version 3.10 and the SGD optimizer to tune model parameters. Table 1 provides detailed configuration of key experimental parameters.
[0106] Table 1 Hyperparameter configuration
[0107]
[0108] The present invention adopts the evaluation indicators commonly used in the field of target detection, including precision (P), recall (R) and mean average precision (mAP). The specific calculation method is as follows: True Positive (TP) refers to the target in the image being correctly identified as the target. False Positive (FP) refers to the target position being correctly identified, but the target category is misclassified. False Negative (FN) refers to the target that originally existed but was not detected, but was mistakenly classified as another category, which is manifested as missed detection. N represents the number of categories. In addition, AP i The area under the precision-recall curve for each category. A larger area indicates better classifier performance. To demonstrate the advantages of this method on small devices, in addition to mAP, we also include the number of model parameters, computational complexity, and model file size as additional evaluation metrics.
[0109]
[0110] To verify the effectiveness of the CPADM-YOLO model, we used YOLOv8 as the baseline model and conducted ablation experiments on the same experimental conditions and dataset. For convenience, the redesigned neck network structure is referred to as the AFSDM module. The results are shown in Table 2.
[0111] Table 2 Ablation experiment
[0112]
[0113] In Method 2, replacing the C2f module in the backbone network with the C2f-CPAM module reduced the model's parameter count to 2.61M and the computational overhead to 7.2 GFLOPs. Simultaneously, mAP@0.5 and mAP@0.5:0.95 each increased by 1%. This demonstrates that the introduction of the CPAM module effectively reduces computational overhead while maintaining high performance. In Method 3, replacing the neck network with the AFSDM module improved mAP@0.5 and mAP@0.5:0.95 by 2.1% and 2.4%, respectively. However, the model's parameter count and computational overhead increased compared to YOLOv8, indicating that while performance improved significantly, computational overhead also increased. Method 4 further replaces the neck network with the module from Method 3, significantly improving mAP@0.5 and mAP@0.5:0.95 compared to Method 2 at the expense of computational overhead. Compared to Method 3, it improves mAP@0.5 and mAP@0.5:0.95 while reducing computational overhead. Ultimately, Method 5, the proposed CPADM-YOLO model, achieves improvements in both mAP and mAP@0.5:0.95, achieving improvements of 4.1% and 3.6% over YOLOv8, respectively. This result demonstrates that the CPADM-YOLO model achieves significant performance improvements through a series of module optimizations, fully validating its effectiveness.
[0114] From a practical application perspective, with the reduced computational complexity and hardware requirements of the model, this method can run smoothly on a wider range of hardware devices. This is particularly true for complex traffic sign detection and recognition tasks, which can reduce hardware requirements and broaden its application scope. Detection speed is also improved compared to similar methods. Furthermore, due to the significant reduction in parameters, this method has the ability to be deployed in real time. These outstanding advantages give this method broad application prospects in various fields such as intelligent transportation. It can provide more efficient, accurate, and convenient technical support for traffic sign detection and recognition, and promote the development of related industries.
Claims
1. A multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments, characterized by: The following steps are involved: Step (1) selects a traffic sign dataset and preprocesses it, then divides the dataset into a training set, a test set, and a validation set according to a certain ratio; Step (2) constructing a CPADM-YOLO model, which is improved on the basis of the YOLOv8 model. The backbone network proposes a C2f-CPAM module to replace the C2f module in YOLOv8; the neck network introduces the TFE module and the DZSF module to reconstruct a new neck network to enhance feature fusion and multi-scale target detection capabilities; Step (3) uses the processed data set to train the CPADM-YOLO model. After the training is completed, the performance of the model is evaluated using the test set, the test results are obtained, and the performance of the model is comprehensively analyzed through various evaluation indicators; Step (4) uses the trained CPADM-YOLO model to detect traffic signs on the image to be processed.
2. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 1 is characterized in that: The preprocessing process of the data set in step (1) includes data cleaning and formatting. The specific operation is to perform data cleaning in Python language, first delete the image files without labels, and then remove other labels in the original data set except for category and location information; then, convert the processed data into COCO format, and finally convert the COCO format data into txt file format.
3. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 1 is characterized in that: The C2f-CPAM module replaces the Bottleneck module in the C2f module by designing a CPAM module. The specific operations of the CPAM module are as follows: First, the input feature map is preliminarily processed by a 3×3 conv1 layer; then, the feature map output by the conv1 layer is divided into two parts, each containing an equal number of channels; one part is processed by a 5×5 conv2 layer and divided into two parts again; one part of the conv2 layer is processed by a 7×7 conv3 layer; the feature map processed by the conv3 layer is fused with the other part of the conv2 layer and the other part of the conv1 layer through a splicing operation to form a multi-scale feature representation; The concatenated feature map is passed through a 1×1 conv4 layer for channel adjustment and is residually connected with the original input feature map through an addition operation to form the final output feature map.
4. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 1 is characterized in that: The new neck network includes a pair of TFE modules and a DZSF module as well as multiple C2f modules and multiple Conv modules; the features output by the second, third, and fourth C2f-CPAM modules of the backbone network are P3, P4, and P5 layer feature maps; The output feature map of the SPPF module of the backbone network is processed by the first Conv module, and the feature map of the P4 layer and the feature map of the P3 layer are processed by the second Conv module and then enter the first TEE module; The features output by the first TEE module are processed by the first C2f module, the fourth Conv module, and the output feature map. The feature map output by the first C2f-CPAM module of the backbone network is processed by the third Conv module and the P3 layer feature map is then input into the second TEE module. The output features of the second TEE module pass through the second C2f module and the fifth Conv module in sequence, and are concat-processed with the output features of the fourth convolution module before being input into the third C2f module and the sixth Conv module in sequence. The output features of the sixth Conv module and the output features of the first Conv module are concat-processed and then input into the fourth C2f module; The feature maps of the P3, P4, and P5 layers are input into the DZSF module, and then ADD processed with the second C2f module. Then, they enter the detection head module together with the feature maps output by the third C2f module and the fourth C2f module for detection.
5. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 4 is characterized in that: The specific operations of the DZSF module are as follows: Adjust the number of channels of feature maps of each scale through convolution operations so that they have the same number of channels; A dynamic sampling mechanism is introduced to spatially adjust the feature maps of P4 and P5 at different scaling ratios to align them with the feature map of P3. After dynamic sampling, the feature maps of the three scales will be expanded into 3D tensors; By canceling the squeezing operation to expand it into a 4D feature map, splicing it along the depth dimension, a 3D fusion feature map containing multi-scale information is synthesized; Finally, the feature map is processed by 3D convolution, batch normalization and activation function to output the final fused feature map, completing the effective perception and detection accuracy enhancement of targets of different scales.
6. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 5, characterized in that: The dynamic sampling mechanism redefines the upsampling process from the perspective of point sampling and introduces content-aware sampling position generation as follows: Given an input feature map And upsampling ratio s, construct the initial sampling grid Follow the "bilinear initialization" principle: Among them, the meshgrid operation generates a regular grid, and the transpose operation ensures that the coordinates are arranged correctly; Second, content-aware offsets are generated by linear projection. Or use the dynamic range factor for even greater flexibility: Among them, σ represents the sigmoid activation function, and the range factor prevents the overlapping of sampling points and ensures the correct expression of the boundary area; Then, the sampling set S is composed of the grid and offset composition: Finally, the input features are resampled by the sampling set to obtain the upsampled features 7. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 6, characterized in that: The dynamic sampling mechanism also introduces a grouping processing mechanism to divide the features into g groups, so that each group of features uses an independent sampling position.
8. The multi-scale adaptive dynamic feature fusion traffic sign detection method for complex environments according to claim 1 is characterized in that: The CPADM-YOLO model uses Wise_MPDIoU as the loss function, which is expressed as follows: The definition of MPDIoU is as follows: in, Indicates the points at the upper left and lower right corners of the predicted box and the real box, Indicates the distance between corresponding points, w and h represent the width and height of the bounding box, and WIoU is defined as follows: Among them, b i Represents the coordinates of i object boxes, g i represents the coordinates of the ground truth box of the i-th object, w i Indicates the weight value.
Citation Information
Cited By
Mutton sheep body size measuring method and device based on target detection network model SVW-YOLO
CN121214484A