Traffic sign target detection method based on improved CPBM-YOLO

By improving the backbone and neck network structure of the YOLOv8 model and using CSP-PMSFA and MAF-YOLO networks, the problem of insufficient feature extraction and fusion in traffic sign detection is solved, the detection accuracy and speed are improved, and it is suitable for a variety of devices.

CN120298997APending Publication Date: 2025-07-11HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510166276.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing YOLOv8 model has problems such as insufficient feature extraction capability and insufficient feature fusion in traffic sign detection, resulting in low detection accuracy and robustness, especially in the face of traffic signs of diverse shapes and scales.

Method used

The CSP-PMSFA module is used to replace the C2f module of YOLOv8 as the backbone network, and the improved MAF-YOLO network structure is introduced as the neck network, including the BIFPN module, forming a CPBM-YOLO model, and improving the detection performance of the model through efficient feature fusion and adaptive feature weight adjustment.

Benefits of technology

The model's detection ability of traffic signs of different scales is improved, the calculation complexity and parameter quantity is reduced, the ability to perceive details is enhanced, and the detection accuracy and faster detection speed is achieved. It is suitable for devices with a variety of computing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298997A_ABST
    Figure CN120298997A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic sign target detection method based on improved CPBM-YOLO, and belongs to the field of traffic sign detection. A CPBM-YOLO model is constructed, improvement is carried out based on a YOLOv8 model, and a C2f module of the YOLOv8 model is replaced by a CSP-PMSFA module in a backbone network part; the neck network part is introduced into an MAF-YOLO network structure, an SAF module and an AAF module of the MAF-YOLO network structure are replaced with BIFPN modules, a RepHELAN module is replaced with a C2f module, and the improved MAF-YOLO network structure is formed to serve as a neck network. Compared with the prior art, the performance and adaptability of the model in the aspect of traffic sign target detection are greatly enhanced, the improved method is practically applied to roadside traffic sign detection scenes, compared with an original YOLOv8 method, the improved method has all obvious advantages, the detection precision is improved, the parameter quantity is effectively reduced, and the calculated amount is remarkably reduced; the calculation speed is greatly increased, strong competitiveness is shown in the industry, and further development of the traffic sign detection technology is expected to be promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and particularly to a traffic sign object detection method based on an improved CPBM-YOLO. Background Art

[0002] With the increasing complexity and intelligent development of modern traffic systems, the accurate detection of traffic signs plays a crucial role in ensuring road traffic safety and smoothness. Traffic signs contain rich information, such as speed limits, no-entry signs, turning instructions, etc. These pieces of information can provide necessary guidance and restrictions for drivers and reduce the occurrence of traffic accidents. In the intelligent transportation system (ITS), traffic sign detection is the basis and key link for realizing functions such as autonomous driving, assisted driving, and traffic monitoring and management.

[0003] Traditional traffic sign detection methods mainly rely on handcrafted features and machine learning algorithms. For example, handcrafted features such as Histogram of Oriented Gradients (HOG), Local Binary Pattern (LBP), etc. are used to describe the appearance of traffic signs, and then classifiers such as Support Vector Machine (SVM), Adaboost, etc. are used for classification and detection. However, these methods have obvious limitations:

[0004] 1) The design of handcrafted features requires a large amount of professional knowledge and experience, and the generalization ability of the features is poor. For traffic signs in different scenarios and of different types, features often need to be redesigned.

[0005] 2) When dealing with complex backgrounds and diverse traffic sign morphologies, the detection accuracy and robustness of these methods are relatively low, and they are easily affected by factors such as illumination, weather, and occlusion.

[0006] In recent years, deep learning techniques have achieved great success in the field of object detection. Object detection algorithms based on Convolutional Neural Networks (CNNs) such as the R-CNN series, SSD, and YOLO have become the current mainstream object detection methods. These methods can automatically learn features in images without the need for handcrafted feature design, and have greatly improved in terms of detection accuracy and speed.

[0007] Among them, the YOLO series of algorithms have received extensive attention for their fast detection speed and good detection accuracy. As a relatively new version of the YOLO series, YOLOv8 has been further optimized in terms of structure design and performance. However, in the specific application scenario of traffic sign object detection, YOLOv8 still has some deficiencies: in terms of feature extraction, in the backbone network, although the C2f module of YOLOv8 can extract features, for objects like traffic signs with diverse shapes and scales, its feature extraction ability still needs to be improved. In practical applications, there may be situations where the feature extraction of some small-sized or peculiarly shaped traffic signs is insufficient, thus affecting the detection accuracy. In terms of feature fusion, in the neck layer, the feature fusion mechanism of YOLOv8 is not efficient enough for the fusion of multi-scale features of traffic signs. Traffic signs may exist in different sizes and resolutions in the image, and the network needs to be able to better fuse features at different levels to accurately detect objects of different scales. The existing feature fusion methods may lead to information loss or insufficient utilization of features during the fusion process, thereby affecting the detection performance of traffic signs of different sizes.

[0008] In summary, although the existing object detection methods based on deep learning have achieved certain results in traffic sign detection, there are still deficiencies in accuracy and performance in practical applications. Especially for objects like traffic signs with specific shape and scale characteristics, the existing YOLOv8 model needs to be further improved and optimized to meet the requirements of accurate and fast detection of traffic signs in intelligent transportation systems. Therefore, developing a traffic sign object detection method based on the improved CPBM-YOLO has important practical significance, can effectively improve the accuracy and reliability of traffic sign detection, and thus provide strong technical support for the development of intelligent transportation systems. Summary of the Invention

[0009] Object of the Invention: The object of the present invention is to propose a traffic sign target monitoring method based on the improved CPBM-YOLO. By referring to the idea of CVPR2024-FasterNet, an efficient CSP-PMSFA module is used to replace the C2f module of the original YOLOv8 model, the MAF-YOLO network structure is introduced, and it is improved by referring to the idea of the BIFPN network structure to make it a neck module, aiming to improve the accuracy and reduce the computational amount while reducing the number of parameters, effectively meeting the requirements of real-time monitoring of traffic signs, and having the advantages of high accuracy, with a wide application prospect.

[0010] Technical Solution: A traffic sign object detection method based on the improved CPBM-YOLO of the present invention includes the following steps:

[0011] Step (1): Select a dataset, preprocess the dataset, and then divide the dataset into a training set, a test set, and a validation set according to a ratio.

[0012] Step (2): Construct a CPBM-YOLO model. The CPBM-YOLO model is improved based on the YOLOv8 model. In the backbone network part, the CSP-PMSFA module is used to replace the C2f module of the YOLOv8 model. In the neck network part, the MAF-YOLO network structure is introduced. The SAF module and the AAF module of the MAF-YOLO network structure are replaced with the BIFPN module, and the RepHELAN module is replaced with the C2f module to form the improved MAF-YOLO network structure as the neck network.

[0013] Step (3): Use the dataset to perform object detection training on the CPBM-YOLO model for traffic signs. After the training is completed, use the test set to detect the model performance, obtain the recognition results, and evaluate the model performance using various metrics.

[0014] Furthermore, the preprocessing of the dataset in step (1) includes data cleaning and format unification. Data cleaning is implemented through the Python language to delete unlabeled pictures and remove other labels in the original dataset file except for the category information and location information. Subsequently, the remaining information is converted into a COCO-format dataset, and finally, the COCO data type is converted into a txt file format.

[0015] Furthermore, the specific operation of the CSP-PMSFA module is as follows:

[0016] The CSP-PMSFA module first passes the input feature map into conv1. conv1 is a 3×3 convolutional layer for preliminary processing. Then the feature map is split into two parts. One part is successively processed by a 5×5 convolutional layer conv2 with the number of groups being half of the input channels and then split into two parts again and processed by a 7×7 convolutional layer conv3 with the number of groups being one-fourth of the input channels. The part passing through conv1, conv2, and the other part of conv2 is retained. Then the feature processed by conv3 and the retained feature are concatenated through a splicing operation, and then processed by a 1×1 convolutional layer conv4. Finally, a residual connection is added to add the original input and the processed feature.

[0017] Furthermore, the improved MAF-YOLO network structure includes 6 BIFPN modules, 6 C2f modules, and two upsampling modules. The backbone network takes an image with a size of 640×640×3 as input. The CSP-PMSFA module in the P2 layer processes the features, and its output is used as the input of the first BIFPN module. The features output by the CSP-PMSFA module in the P3 layer of the backbone network are respectively input into the first and second BIFPN modules; the features output by the CSP-PMSFA module in the P4 layer of the backbone network are respectively used as the inputs of the second and third BIFPN modules; the features output by the SPPF module in the P5 layer of the backbone network are used as the input of the third BIFPN.

[0018] Furthermore, the improved MAF-YOLO network structure includes a bottom-up path and a top-down path:

[0019] Bottom-up path: Starting from the features processed by the SPPF module in the P5 layer, the SPPF module extracts more global and abstract features. The features of the P4 layer are downsampled to match the scale of the P5 layer and then concatenated, and then sent into the third BIFPN module. The features of the two are fused through bidirectional cross-scale connection and weighted fusion, and then the integrated features are deeply extracted by the first C2f module; the processed features are upsampled for the first time and concatenated with the features of the P4 layer and the downsampled P3 layer, and then processed by the second BIFPN module and the second C2f module in turn; continue the second upsampling in this way and concatenate with the relevant features of the P3 layer and P2 layer and process through the first BIFPN module and the third C2f module to transfer the low-level detailed information to the high level;

[0020] Top-down path: Starting from the features output by the third C2f module, first concatenate with the features processed by the second upsampling, and then send them into the fourth BIFPN module for fusion, and then further extract features through the fourth C2f module; the processed features are downsampled and concatenated with the features processed by the second C2f module and then downsampled, and the features processed by the third C2f module and then downsampled, and then processed by the fifth BIFPN module and the fifth C2f module in turn to fuse multi-level semantic information; continue downsampling in this mode and concatenate with the features processed by the first C2f module and the second C2f module and then downsampled, and process through the sixth BIFPN module and the sixth C2f module.

[0021] Beneficial effects:

[0022] (1) The present invention improves the backbone network of the YOLOv8n network model by replacing the original C2f module with the CSP-PMSFA module. Among them, the efficient PartialConv operation adopted by the CSP-PMSFA module is very distinctive. This operation method does not extract multi-scale feature information on all channels, but selectively (i.e., partially) processes. Through the way of partial channel processing, unnecessary computational complexity is effectively reduced. Compared with the traditional method of performing the same operation on all channels, the computational speed can be significantly improved. At the same time, when processing features, the input features are added to the processed features using residual connections. This method retains the key information of the original input while introducing new multi-scale information. The retention of the original information ensures that the model does not lose important basic features, while the new multi-scale information enriches the model's representation ability for targets of different scales, thereby greatly improving the model's expression ability and enabling it to capture and identify targets such as traffic signs more accurately.

[0023] (2) The present invention has also achieved remarkable results in the improvement of the neck layer. By introducing the MAF-YOLO network structure and improving it by referring to the idea of the BIFPN network structure, and embedding it into the CPBM-YOLO model as the NECK module, the following important effects are brought: 1) Improvement of multi-resolution feature processing ability: On feature maps of different resolutions, the improved BIFPN module shows powerful functions. It can ensure the efficient and accurate transmission and fusion of features at each resolution level. This ability enables the model to make full use of the feature information at each resolution when facing targets of different sizes, avoiding detection difficulties caused by differences in target sizes, and thus significantly improving the detection ability for targets of different sizes. 2) Reduction of model complexity and improvement of performance: The introduction of BIFPN not only does not increase the burden on the model, but effectively reduces the number of parameters and computational complexity. On the premise of ensuring the model performance, the demand for hardware resources of the model is reduced, enabling it to be more widely applied to devices with different computing capabilities. 3) Adaptive feature weight adjustment and enhanced detail perception: The integration of the MAF-YOLO structure and BIFPN is another highlight of the present invention. This integration enables the model to adaptively adjust its weights according to the importance of features at different scales. When processing targets with rich details such as traffic signs, this adaptive weight adjustment mechanism can highlight important detail features, enhance the model's perception ability of details, and thus enable the model to generate diverse and accurate detection outputs at multiple resolution levels, comprehensively improving the detection performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is the flowchart of object detection of the present invention;

[0025] Figure 2 The C2f network structure diagram after the improvement of the backbone network of the present invention;

[0026] Figure 3 The network structure diagram of the CSP-PMSFA module in the present invention;

[0027] Figure 4 The overall network structure diagram of the improved YOLOv8 of the present invention;

[0028] Figure 5 The partial detection display diagram of the validation set data in the embodiment;

[0029] Figure 6 The category pictures of the data set in the embodiment;

[0030] Figure 7 The detection example of YOLOv8 in the traditional conventional technology test set;

[0031] Figure 8 The detection example in the test set of the present invention. Detailed implementation manners

[0032] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0033] The embodiment of the present invention provides a traffic sign target detection method based on the improved CPBM-YOLO, including the following steps:

[0034] Step 1: Construct a data set and preprocess the data set;

[0035] Step 1.1: In this embodiment, for the data set in the field of traffic signs, the open-source data set TT100K manufactured by Tsinghua University is used as a representative to construct the data set; because there are problems such as pictures without labels and serious imbalance in the number of different types of pictures in the original data set, the TT100K data set is cleaned by the Python language, pictures without labels are deleted, and pictures with the number of category instances greater than 100 are screened out, totaling 9,738 pictures in 45 categories

[0036] Step 1.2: Since the label information of the original data set is recorded in a json file, while the txt file required in this embodiment, which contains category information and location information, needs to be converted in format; first, the data set is parsed under the python script, each sub-list is traversed, then the category labels corresponding to the numbers are written, and finally an XML file is generated. The bounding box information is extracted from the XML file, and then the information is converted into a YOLO-format data set. By reading the coordinates of the boxes, the corresponding TXT file of the picture is generated.

[0037] Step 1.3: Divide the traffic sign dataset in Step 1.1 into a training set, a validation set, and a test set according to a ratio of 7:2:1. That is, the training set has 6,793 images, the validation set has 1,949 images, and the test set has 996 images.

[0038] Step 2. Model establishment: Send the augmented dataset to the improved YOLOv8 network model (CPBM-YOLO model) of this embodiment for training. The overall improved model is as Figure 4 shown, and the target network model is established.

[0039] The CPBM-YOLO model is improved based on the YOLOv8 model. In the backbone network part, the CSP-PMSFA module is used to replace the C2f module of the YOLOv8 model; in the neck network part, the MAF-YOLO network structure is introduced, the SAF module and AAF module of the MAF-YOLO network structure are replaced with the BIFPN module, and the RepHELAN module is replaced with the C2f module to form the improved MAF-YOLO network structure as the neck network.

[0040] Step 2.1: The backbone network is responsible for extracting feature information. Referring to the idea of CVPR2024-FasterNet, a new structure called CSP-Partial Multi-Scale Feature Aggregation (CSP-PMSFA) is designed to replace the C2f module in YOLOv8. The structure is as Figure 3 shown. The efficient PartialConv operation adopted by the CSP-PMSFA module is very characteristic. This operation method does not perform multi-scale feature information extraction on all channels, but selectively (i.e., partially) processes. By the way of partial channel processing, the unnecessary computational amount is effectively reduced. Compared with the traditional method of performing the same operation on all channels, the computational speed can be significantly improved. At the same time, the final 1x1 convolutional layer of this module fuses features of different scales together, and uses residual connection to add the input feature and the processed feature, effectively retaining the original information and introducing new multi-scale information, thereby improving the expression ability of the model.

[0041] The CSP-PMSFA module first passes the input feature map into conv1, which is a 3×3 convolutional layer for preliminary processing. Then the feature map is split into two parts. One part is successively processed by the 5×5 convolutional layer conv2 with the number of groups being half of the number of input channels and then by the 7×7 convolutional layer conv3 with the number of groups being one-fourth of the number of input channels after being split in half again. After conv1, conv2, and the other part of conv2 are retained. Then the features processed by conv3 and the retained features are concatenated, and then processed by the 1×1 convolutional layer conv4. Finally, a residual connection is added to add the original input and the processed features.

[0042] Step 2.2: During the construction of the neck structure of the CPBM-YOLO model, the MAF-YOLO network structure was innovatively introduced, and the advanced idea of the BIFPN network structure was cleverly borrowed to improve it. Then the improved structure was embedded into the whole model as the NECK module. In the first bottom-up path of this neck structure, the shallow auxiliary fusion (BIFPN) module extracts multi-scale features from the backbone network and performs preliminary fusion to enhance the model's perception ability of details. The second top-down path then uses the advanced auxiliary fusion (BIFPN) module for denser connections to integrate the gradient information of each layer, and then generates diverse detection outputs at multiple resolution levels.

[0043] The improved MAF-YOLO network structure includes 6 BIFPN modules, 6 C2f modules, and two upsampling modules. The backbone network takes an image with a size of 640×640×3 as the input. The CSP-PMSFA module in the P2 layer processes the features, and its output is used as the input of the first BIFPN module. The features output by the CSP-PMSFA module in the P3 layer of the backbone network are respectively input into the first and second BIFPN modules; the features output by the CSP-PMSFA module in the P4 layer of the backbone network are respectively used as the inputs of the second and third BIFPN modules; the features output by the SPPF module in the P5 layer of the backbone network are used as the input of the third BIFPN.

[0044] The improved MAF-YOLO network structure includes a bottom-up path and a top-down path:

[0045] Bottom-up path: Starting from the features processed by the SPPF module at the P5 layer, the SPPF module extracts more global and abstract features. After the P4 layer is downsampled to match the scale of the P5 layer and then concatenated, it is fed into the third BIFPN module. The features of the two are fused through bidirectional cross-scale connections and weighted fusion, and then the integrated features are deeply extracted by the first C2f module. The processed features are upsampled for the first time and concatenated with the features of the P4 layer and the downsampled P3 layer, and then processed by the second BIFPN module and the second C2f module in sequence. Continuing in this way, the second upsampling is performed and concatenated with the relevant features of the P3 layer and the P2 layer, and then processed by the first BIFPN module and the third C2f module, transferring the low-level detail information to the high level, enhancing the high-level's perception ability of details such as the target position, and facilitating the accurate positioning of the target.

[0046] Top-down path: Starting from the features output by the third C2f module, first concatenate with the features processed by the second upsampling, and then feed them into the fourth BIFPN module for fusion. Then, further extract features through the fourth C2f module. The processed features are downsampled and concatenated with the features processed by the second C2f module after downsampling and the features processed by the third C2f module after downsampling, and then processed by the fifth BIFPN module and the fifth C2f module in sequence to fuse multi-level semantic information. Continuing in this mode, the downsampling is performed and concatenated with the features processed by the first C2f module and the second C2f module after downsampling, and then processed by the sixth BIFPN module and the sixth C2f module, transferring the high-level semantic information to the low level, enriching the low-level semantic content, helping to judge the target category, and improving the detection accuracy.

[0047] Specifically, introducing the BIFPN bidirectional feature pyramid network structure into the fusion module in the MAF-YOLO network structure and adding the information extracted by the backbone network to the aggregation path has significant advantages:

[0048] a) Advantage of bidirectional connection: Breaking the limitation of unidirectional information transfer in the traditional feature pyramid network, constructing bidirectional connections, realizing the full interaction and fusion of feature information at different levels, enabling high-level and low-level features to complement each other, and comprehensively capturing feature information at various scales and positions, laying a solid foundation for accurate target detection. In traffic sign detection, it can better capture the feature expressions of signs at different scales.

[0049] b) Advantage of adaptive feature adjustment: Relying on the adaptive weighting mechanism, dynamically allocating fusion weights according to the importance of features, making reasonable and full use of each feature, avoiding the problem of unreasonable weight allocation, improving the quality of feature fusion, enhancing the representativeness and discriminability of the fused features, and helping to improve the accurate recognition ability of traffic signs.

[0050] c) Advantages of modular design: Its modular design offers high flexibility, enabling seamless embedding into different network architectures, adapting and integrating into the neck structure of the CPBM-YOLO model to enhance performance, and being easily expandable and customizable as needed, facilitating the optimization of traffic sign detection performance for different traffic scenarios, dataset characteristics, etc.

[0051] d) Advantages of high efficiency: Focus on optimizing computational efficiency, streamlining the computational process and reducing the number of parameters. While ensuring high-quality feature fusion and high-performance detection, it reduces the model complexity, not only improving the operation speed to meet the real-time detection requirements but also reducing the consumption of hardware resources, facilitating the wide deployment of the model on various types of devices and promoting the implementation of traffic sign detection applications.

[0052] Step 3: After training is completed, use the test set to detect the performance of the model, obtain the recognition results, and evaluate the model performance using various metrics.

[0053] Step 4: The YOLOv8 model uses the CIoU loss when calculating the regression loss. The formula for the CIoU loss during the calculation process is as follows:

[0054]

[0055] Among them, v is a parameter used to measure the aspect ratio consistency, and α is a weight function, defined as:

[0056]

[0057] Among them, b and b gt respectively represent the center points of the predicted bounding box and the ground truth bounding box, ρ represents the Euclidean distance, c represents the diagonal distance of the smallest enclosing rectangle of the predicted bounding box and the ground truth bounding box, w gt represents the width of the ground truth bounding box, h gt represents the height of the ground truth bounding box, and w and h represent the width and height of the predicted bounding box.

[0058] Table 1 Hyperparameter configuration

[0059]

[0060] The present invention adopts the commonly used evaluation metrics in the field of object detection, including precision (P), recall (R), and mean average precision (mAP). The specific calculation formulas for each metric are as follows: TP (true positive) refers to the case where the target in the image is accurately recognized as the correct target; FP (false positive) is a false positive example, which means that although the location of the target can be recognized, the category of the target is misrecognized; TPFN (false negative) is a false negative example, that is, the correct target is not recognized and is instead regarded as other targets, resulting in missed detection. Here, N represents the number of categories. Among them, AP i represents the area under the precision-recall curve for each category. The larger this area is, the better the performance of the classifier. In addition, to highlight the advantages of the present invention when applied to small devices compared with other models, the number of parameters, the amount of computation, and the size of the model file are additionally included in the scope of evaluation metrics.

[0061]

[0062] To verify the effectiveness of the method in this embodiment, this method and the traditional YOLOv8 method are tested on the same test set, and the comparison results of various performance metrics are shown in the following table.

[0063] Table 2 Experimental Results

[0064]

[0065] In the table, compared with the traditional YOLOv8 method, the method proposed in this embodiment shows more excellent detection performance in the task of road test traffic sign detection and recognition. Among them, the mean average precision (mAP value) can reach 74.2%, which fully demonstrates the advantage of this method in detection accuracy. Not only that, this method has also achieved a substantial optimization in the number of model parameters. Compared with the YOLOv8 method, the reduction in the number of parameters is about 12%. At the same time, the floating-point operation amount is also lower than that of the basic YOLOv8 method. This means that in practical applications, when the device performs complex calculation tasks such as traffic sign detection and recognition, the requirements for hardware are reduced, and thus it can run smoothly on more types of hardware devices, broadening the hardware adaptation range of its application. This method also has an improvement in detection speed compared with other similar methods, and the substantial reduction in the number of parameters even makes it meet the conditions for real-time deployment. With these outstanding advantages, this method has a broad application prospect in many fields such as intelligent transportation, and is expected to provide more efficient, accurate and convenient technical support for traffic sign detection and recognition work, helping the relevant industries to develop better.

Claims

1. A traffic sign target detection method based on improved CPBM-YOLO, characterized in that, It includes the following steps: Step (1): Select a dataset, preprocess the dataset, and then divide the dataset into a training set, a test set, and a validation set according to a ratio; Step (2): Construct a CPBM-YOLO model. The CPBM-YOLO model is improved based on the YOLOv8 model. In the backbone network part, the CSP-PMSFA module is used to replace the C2f module of the YOLOv8 model; in the neck network part, the MAF-YOLO network structure is introduced. The SAF module and the AAF module of the MAF-YOLO network structure are replaced with the BIFPN module, and the RepHELAN module is replaced with the C2f module to form the improved MAF-YOLO network structure as the neck network; Step (3): Use the dataset to perform object detection training on the CPBM-YOLO model. After the training is completed, use the test set to detect the model performance, obtain the recognition results, and use various indicators to evaluate the model performance.

2. A traffic sign target detection method based on improved CPBM-YOLO according to claim 1, characterized in that, The preprocessing of the dataset in step (1) includes data cleaning and format unification. Data cleaning is implemented through the Python language, unlabeled pictures are deleted, and other labels except the category information and location information in the original dataset file are removed. Subsequently, the remaining information is converted into a COCO format dataset, and finally the COCO data type is converted into a txt file format.

3. A traffic sign target detection method based on improved CPBM-YOLO according to claim 1, characterized in that, The specific operation of the CSP-PMSFA module is as follows: The CSP-PMSFA module first passes the input feature map into conv1. conv1 is a 3×3 convolutional layer for preliminary processing. Then the feature map is split into two parts. One part is successively processed by a 5×5 convolutional layer conv2 with the number of groups being half of the input channels and then split into two parts again and processed by a 7×7 convolutional layer conv3 with the number of groups being one-fourth of the input channels. The part passing through conv1, conv2, and the other part of conv2 is retained. Then the features processed by conv3 and the retained features are concatenated and then processed by a 1×1 convolutional layer conv4. Finally, a residual connection is added to add the original input and the processed features.

4. A traffic sign target detection method based on improved CPBM-YOLO according to claim 1, characterized in that The improved MAF-YOLO network structure includes 6 BIFPN modules, 6 C2f modules, and two upsampling modules. The backbone network takes an image with a size of 640×640×3 as the input. The CSP-PMSFA module of the P2 layer processes the features, and its output is used as the input of the first BIFPN module. The features output by the CSP-PMSFA module of the P3 layer of the backbone network are respectively input into the first and second BIFPN modules; the features output by the CSP-PMSFA module of the P4 layer of the backbone network are respectively used as the inputs of the second and third BIFPN modules; the features output by the SPPF module of the P5 layer of the backbone network are used as the input of the third BIFPN.

5. The traffic sign target detection method based on the improved CPBM-YOLO according to claim 4, characterized in that The improved MAF-YOLO network structure includes a bottom-up path and a top-down path: Bottom-up path: Starting from the features processed by the SPPF module at the P5 layer, the SPPF module extracts more global and abstract features. The features of the P4 layer are downsampled to match the scale of the P5 layer and then concatenated, and then fed into the third BIFPN module. The features of the two are fused through bidirectional cross-scale connections and weighted fusion, and then the integrated features are deeply extracted by the first C2f module; the processed features are concatenated with the features of the P4 layer and the downsampled P3 layer after passing through the first upsampling module, and then processed by the second BIFPN module and the second C2f module in sequence; continue in this way to splice the second upsampling module with the relevant features of the P3 layer and the P2 layer and process them by the first BIFPN module and the third C2f module, transferring the low-level detail information to the high level; Top-down path: Starting from the output features of the third C2f module, first concatenate with the features processed by the second upsampling module, feed them into the fourth BIFPN module for fusion, and then further extract features through the fourth C2f module; the processed features are downsampled and concatenated with the features processed by the second C2f module after downsampling and the features processed by the third C2f module after downsampling, and then processed by the fifth BIFPN module and the fifth C2f module in sequence to fuse multi-level semantic information; continue in this mode to downsample and concatenate with the features processed by the first C2f module and the second C2f module after downsampling and process them by the sixth BIFPN module and the sixth C2f module.

6. A traffic sign target detection method based on an improved CPBM-YOLO according to any one of claims 1 to 5, characterized in that, CIoU loss is used when calculating the regression loss, and the formula of CIoU loss in the calculation process is as follows: Among them, v is a parameter used to measure the aspect ratio consistency, and α is a weight function, defined as: Among them, b and b gt respectively represent the center points of the predicted bounding box and the ground truth bounding box, ρ represents the Euclidean distance, c represents the diagonal distance of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, w gt represents the width of the ground truth bounding box, h gt represents the height of the ground truth bounding box, and w and h represent the width and height of the predicted bounding box.

Citation Information

Cited By

  • Endoscope bourdon tube defect detection method based on improved YOLOv8n

    CN120997198A

  • Endoscope spring tube defect detection method based on improved YOLOv8n

    CN120997198B