Lightweight traffic sign detection method based on yov8

By improving the lightweight Adown module, WTConv module, HS-FPN module and context-anchored attention CAA mechanism of the yolov8 model, the problems of accuracy reduction and computing resource consumption in lightweight traffic sign detection are solved, and efficient real-time detection on low-power devices are achieved.

CN120299006APending Publication Date: 2025-07-11YANCHENG INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363101.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing lightweight traffic sign detection methods are often accompanied by a significant decrease in accuracy while reducing the amount of parameters. Training deep learning models requires a large amount of labeled data, which is difficult to obtain. The traditional methods are low in accuracy and poor in real environments.

Method used

Using the improved yolov8 model, the ordinary downsampled Conv module is replaced by using a lightweight Adown module in the Backbone network, combined with the WTConv module to alleviate the size of the convolution layer core, the HS-FPN module is introduced in the Neck network for feature fusion, and the context-anchored attention CAA mechanism and the cross-stage context-guided feature fusion module are introduced to optimize the model structure.

Benefits of technology

While reducing the amount of model parameters and calculations, maintain or improve detection accuracy and speed. It is suitable for low-power equipment and real-time processing scenarios, such as autonomous driving and drone monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299006A_ABST
    Figure CN120299006A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network model detection, in particular to a light-weight traffic sign detection method based on yov8, and the method comprises the steps: obtaining traffic sign image data through a high-definition camera, and obtaining a data set which comprises a test set, a training set and a verification set; a YOLO model of the improved yolov8 is built; the YOLO model comprises a Backbone network, a Neck network and a Head network, the Backbone network is used for feature extraction, and the Neck network is used for feature fusion; the Head network is used for completing a final prediction task; inputting the pictures in the data set into a YOLO model for training; and the performance of the obtained YOLO model is evaluated and compared, so that the detection precision and speed of the model can be maintained on the premise of reducing the parameter quantity and the calculation quantity of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network model detection, and specifically to a lightweight traffic sign detection method based on YOLOv8. Background Art

[0002] Currently, the mainstream methods for traffic sign recognition and detection are traditional object detection methods and deep learning-based object detection methods.

[0003] Traditional machine learning object detection methods usually adopt a sliding window detection framework, taking the partial features collected by the sliding window as candidate regions, and then classifying the collected features through a classifier to obtain the classification result. However, after a long period of development, compared with deep learning methods, traditional methods have defects such as low accuracy and poor practicability in the real natural environment. Therefore, in this paper, the deep learning methods are mainly studied in depth.

[0004] In the field of deep learning, there are mainly two-stage object detection methods based on candidate box extraction, such as R-CNN and Faster R-CNN, and single-stage object detection methods mainly based on the YOLO series and SSD. Representative methods of two-stage object detection are Faster R-CNN, Mask R-CNN, Cascade R-CNN, etc. Since two-stage object detection methods require two independent stages, they need more computational and time resources, and the inference speed is slow. In real situations, they may not achieve good real-time traffic sign recognition effects. Single-stage object detection methods can well meet the requirements of detection speed and real-time performance. Currently, YOLOv8 has better frames per second and accuracy in traffic sign object detection than other mainstream methods, so it has been more widely used. To ensure the accuracy and real-time performance of the algorithm, a series of lightweight networks have been proposed, including MobileNet, ShuffleNet, etc. These networks have low complexity and are suitable for running on mobile devices or other embedded hardware. However, the computational amount and memory occupancy of these lightweight networks are relatively large in actual applications in fatigue driving detection and still need to be improved. Moreover, training deep learning models requires a large amount of labeled data, but the acquisition of labeled data is difficult and often requires a large amount of manual labor.

[0005] Currently, widely used lightweight networks often show a significant decrease in accuracy while reducing the number of parameters. Compared with general lightweight networks, this patent shows good performance in reducing the number of parameters, and at the same time, the accuracy has not changed significantly compared with the original network. This patent makes good use of the ideas of other algorithms and applies them to the YOLO network, such as the RTDETR object detection model. Summary of the Invention

[0006] The object of the present invention is to provide a lightweight traffic sign detection method based on YOLOv8 to solve the problems raised in the above-mentioned background technology.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] A lightweight traffic sign detection method based on YOLOv8, the method comprising:

[0009] S100. Obtain traffic sign image data by using a high-definition camera to obtain a data set, where the data set includes a test set, a training set, and a validation set;

[0010] S200. Build a YOLO model that improves YOLOv8; the YOLO model includes a Backbone network, a Neck network, and a Head network. The Backbone network is used for feature extraction, and the Neck network is used for feature fusion; the Head network is used to complete the final prediction task;

[0011] S300. Input the pictures in the data set into the YOLO model for training;

[0012] S400. Evaluate and compare the performance of the obtained YOLO model.

[0013] Preferably, the YOLO model that improves YOLOv8 in S200 includes:

[0014] S201. In the Backbone network, replace the ordinary downsampling Conv module with a lightweight Adown module;

[0015] S202. Use the WTConv module to alleviate the change in the number of kernel sizes in the convolutional layer;

[0016] S203. In the Neck network, use the HS-FPN module to filter the low-level feature information with the high-level features as weights, and then merge the filtered information with the high-level features to enhance the feature expression ability of the model;

[0017] S204. Introduce the context anchor attention CAA mechanism to capture the potential relationship between long-distance pixels in the image, thereby enhancing the feature learning ability within the central key area;

[0018] S205. Use DConv and Conv in the cross-stage context-guided feature fusion module to extract surrounding and local features from the target area respectively. After batch normalization and PReLU activation, the two are merged to form a preliminary joint feature, and then the above joint feature is dynamically adjusted and weighted.

[0019] Preferably, the lightweight Adown module in S201 includes:

[0020] Perform average pooling operation on the input feature map to reduce the size of the feature map and the computational amount of the subsequent convolutional layer;

[0021] Process in two parts in the channel dimension. One part first undergoes max pooling and then convolution, and the other part only undergoes convolution operation. The parameters of the two convolutional layers are shared, reducing the number of model parameters and improving the parameter efficiency of the model;

[0022] Concatenate the output results to obtain the output;

[0023] In this paper, the downsampling Adown module is used to replace the ordinary downsampling Conv module, effectively improving the model's perception ability of the target and enhancing the accuracy of traffic sign detection.

[0024] Preferably, the WTConv module in S202 includes:

[0025] S202-1. Use wavelet transform to filter and downsample the input low-frequency and high-frequency content;

[0026] S202-2. Before constructing the output using IWT, perform a small kernel depth convolution on different frequency maps according to the formula: Y = IWT(Conv(W, WT(X)));

[0027] Among them, X represents the input tensor, and W represents the weight tensor of the k×k depth kernel, whose input channel number is four times that of X. This operation not only separates the convolution between frequency components but also allows smaller kernels to operate in a larger area of the original input;

[0028] The specific operation includes: performing 3×3 convolution on the low-frequency band in the second-level wavelet domain to obtain a 9-parameter convolution, responding to the low frequency in the 12×12 receptive field of the input X; adopting this 1-level combined operation, according to the formula:

[0029]

[0030] Obtain the aggregated output Z starting from the i-th layer (i) , where the sum of the convolution outputs of two different sizes is used as the output;

[0031] Among them, W (i) , is the input of this layer, representing all three high-frequency maps of the i-th layer.

[0032] Preferably, the HS-FPN module in S203 includes: a feature selection module and a feature fusion module;

[0033] S203-1. In the feature selection module, the input feature map f is obtained. in ∈R C×H×W , and the features obtained after the feature map passes through two pooling layers, namely MaxPool and Avg Pool, are integrated together; and the Sigmoid activation function is used to determine the weight value of each channel to obtain the weight f of each channel. CA ∈R C×1×1 ;

[0034] S203-2. In the feature fusion module, the high-level features are used as weights to fuse features strategically: Given the input high-level features and low-level features, first, the dimensions of the high-level features and low-level features are unified using transposed convolution and bilinear interpolation. Then, the channel feature self-filtering module is applied to convert the high-level features into corresponding weights for filtering the low-level features. Finally, the filtered low-level features and high-level features are fused.

[0035] The specific operations include:

[0036] Given the input high-level feature f high ∈R C×H×W and the input low-level feature The high-level features are initially processed using transposed convolution with a stride of 2 and a convolution kernel of 3×3 to obtain features of size f high ∈R C×2H×2W ; To unify the dimensions of the high-level features and low-level features, first, bilinear interpolation is used to upsample or downsample the high-level features to match their dimensions with those of the low-level features, obtaining the feature Then, the attention CA module is used to convert the high-level features into corresponding attention weights to filter the low-level features and ensure they have consistent dimensions; finally, the filtered low-scale features and high-level features are fused to enhance the model's feature representation ability and produce the final output features.

[0037] Preferably, the CAA mechanism in S204 includes:

[0038] Performing spatial compression on the input feature map using global average pooling operation;

[0039] Further extracting and integrating the feature information of the local region through 1×1 convolution;

[0040] Adopting two lightweight depthwise strip convolutions, H-Conv2d and W-Conv2d, to efficiently capture and integrate pixel information along the height and width directions respectively, and identify and extract the features of objects with slender or special shapes.

[0041] Preferably, the cross-stage context-guided feature fusion module in S205 includes:

[0042] Define local, nearby, and global operations as foc, fsur, and fglo respectively, and define the combination of local and nearby information as fjoi;

[0043] Calculate foc and fsur, and learn the features of local and surrounding respectively. foc is a 3*3 convolution to learn features from 8 adjacent points; fsur is a 3*3 dilated convolution to learn surrounding information; then perform the fioi operation, where fioi is a concat+bn+PRelu operation to fuse local and surrounding features;

[0044] S205-3: fglo extracts global features and introduces channelwise weights.

[0045] A computer-readable storage medium has a computer program stored thereon, and when the program is executed by a processor, it implements the steps in the above-mentioned lightweight traffic sign detection method based on yolov8.

[0046] A computer device includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the steps in the above-mentioned lightweight traffic sign detection method based on yolov8.

[0047] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0048] The present invention can maintain the detection accuracy and speed of the model on the premise of reducing the number of model parameters and the amount of computation;

[0049] (1) Reduce computational resource consumption: By reducing the number of model parameters, the requirements for storage and memory are reduced, making the model easier to deploy on low-power and resource-constrained devices, such as embedded systems and mobile devices;

[0050] (2) Improve detection speed: The optimized lightweight YOLO model can significantly improve the detection speed while maintaining high accuracy, and is suitable for application scenarios with real-time processing and high-speed response, such as autonomous driving and drone monitoring;

[0051] (3) Optimize the model structure: By means of lightweight components, etc., the model structure is simplified and the complexity is reduced, making the model easier to understand and optimize. Through reasonable design and optimization (such as introducing attention mechanisms, improving loss functions, etc.), the model accuracy can often be improved or maintained at a basically unchanged level. Description of the Drawings

[0052] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:

[0053] Figure 1 is a flowchart of a lightweight traffic sign detection method based on yolov8 of the present invention;

[0054] Figure 2 is the overall structure diagram of the network model after improvement of the present invention;

[0055] Figure 3 is a flowchart of the working process of the Adown module of the present invention;

[0056] Figure 4 is a schematic diagram of performing convolution in the wavelet domain of the present invention;

[0057] Figure 5 is a schematic diagram of the working process principle of the HS-FPN module of the present invention;

[0058] Figure 6 is a schematic diagram of the working process principle of the CA module of the present invention;

[0059] Figure 7 is a schematic diagram of the working process principle of the SFF module of the present invention;

[0060] Figure 8 is a schematic diagram of the working process principle of the improved CAA module of the present invention;

[0061] Figure 9 is the ContextGuided structure diagram of the present invention;

[0062] Figure 10 is a schematic diagram of the detection effect of the network model of the present invention. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0064] Please refer to Figures 1-10 , the present invention provides a technical solution:

[0065] Embodiment 1: A lightweight traffic sign detection method based on yolov8, the method includes:

[0066] S100. Obtain traffic sign image data using a high-definition camera to get a dataset, where the dataset includes a test set, a training set, and a validation set;

[0067] Data acquisition: The TT100K (Tsinghua-Tencent 100K) dataset is a traffic sign dataset jointly established by Tsinghua University and Tencent. This dataset was released in 2016 and covers five cities in China. The image resolution of this dataset is 2048×2048, with a large number of images and rich semantic information. The images are collected by high-definition cameras in the streets, restoring the first perspective of drivers while obtaining real street scenes. In this project, 9738 images were selected from the TT100K dataset, and 45 traffic sign categories were screened out according to the criterion that the number of samples within a single category exceeds 100. Among them, the test set contains 996 images, the training set contains 6793 images, and the validation set contains 1949 images.

[0068] S200. Build a YOLO model that improves yolov8; the YOLO model includes a Backbone network, a Neck network, and a Head network. The Backbone network is used for feature extraction, the Neck network is used for feature fusion; the Head network is used to complete the final prediction task;

[0069] Preferably, the YOLO model that improves yolov8 in S200 includes:

[0070] S201. In the Backbone network, replace the ordinary downsampling Conv module with a lightweight Adown module;

[0071] Preferably, the lightweight Adown module in S201 includes:

[0072] Perform average pooling operation on the input feature map to reduce the size of the feature map and the computational amount of subsequent convolutional layers;

[0073] Process in two parts in the channel dimension. One part first undergoes max pooling and then convolution, and the other part only undergoes convolution operation. The parameters of the two convolutional layers are shared, reducing the number of model parameters and improving the parameter efficiency of the model;

[0074] Concatenate the output results to get the output;

[0075] Among them, downsampling is a commonly used operation in convolutional neural networks, which is used to reduce the spatial size of feature maps, thereby reducing the amount of computation and the number of parameters. Ordinary downsampling convolution is simple to implement, easy to understand and use, but it will lose some feature information. Currently, there are mainly two implementation methods of ordinary downsampling convolution. The first one is to skip some pixels in each convolution operation by setting the stride of the convolution kernel; the second one is to use pooling operations (such as max pooling or average pooling) to reduce the size of the output features. The original model uses the first method of ordinary downsampling, while the main purpose of the downsampling Adown module used in this paper is to retain as much feature information as possible during the implementation process, improve the quality of feature maps and the detection performance of the model.

[0076] The Adown module first performs an average pooling operation on the input feature map, effectively reducing the size of the feature map and the amount of computation of the subsequent convolutional layer. Then, it is divided into two parts for processing in the channel dimension. One part first undergoes max pooling and then convolution, and the other part only undergoes convolution. The parameters of these two convolutional layers are shared, which can reduce the number of model parameters and improve the parameter efficiency of the model. Finally, the output results are concatenated to obtain the output. In this paper, the downsampling Adown module is used to replace the ordinary downsampling Conv module, effectively improving the model's perception ability of the target and enhancing the accuracy of traffic sign detection.

[0077] S202. Use the WTConv module to alleviate the change in the number of kernel sizes in the convolutional layer;

[0078] Preferably, the WTConv module in S202 includes:

[0079] WTConv module: Increasing the kernel size of the convolutional layer will increase the number of parameters in a quadratic manner. To alleviate this situation, the following suggestions are put forward.

[0080] First, use wavelet transform to filter and downsample the input low-frequency and high-frequency content. Then, before constructing the output using IWT, perform a small-kernel depth convolution on different frequency maps. This process is represented by Y =

[0081] IWT(Conv(W, WT(X)). Among them, X is the input tensor, W is the weight tensor of the k×k depth kernel, and its number of input channels is four times that of X. This operation not only separates the convolution between frequency components, but also allows smaller kernels to operate in a larger area of the original input. As Figure 4 shown, performing convolution in the wavelet domain can obtain a larger receptive field. In this example, a 3×3 convolution is performed on the low-frequency band in the second-level wavelet domain to obtain a 9-parameter convolution, responding to the low frequency of the 12×12 receptive field in the input X. Adopt this 1-level combined operation and use the formula The same cascading principle further enhances it, and the process is given by the following formula: where is the input of this layer, representing all three high-frequency mappings of the i-th layer. To combine the outputs of different frequencies, we use the fact that the WT and its inverse are linear operations, which means IWT(X + Y) = IWT(X) + IWT(Y). Therefore, perform where Z (i) is the aggregated output starting from the i-th layer, where the convolution outputs of two different sizes are summed as the output. We cannot normalize each of because their individual normalizations do not correspond to the normalization in the original domain. Instead, we find that it is sufficient to perform only channel scaling to measure the contribution of each frequency component.

[0082] S203. In the Neck network, the HS-FPN module uses high-level features as weights to filter low-level feature information, and then merges the filtered information with the high-level features to enhance the feature expression ability of the model;

[0083] Preferably, the HS-FPN module in S203 includes: a feature selection module and a feature fusion module;

[0084] HS-FPN module: HS-FPN mainly uses high-level features as weights to filter low-level feature information, and then merges the filtered information with the high-level features to enhance the feature expression ability of the model. HS-FPN consists of a feature selection module and a feature fusion module.

[0085] The HS-FPN network structure diagram is as Figure 5 shown. Among them, S3, S4, and S5 in the Backbone part represent feature maps of three different scales; CA (Channel Attention) in the Feature Selection part is channel attention, and P3, P4, and P5 represent feature maps of three different scales; N3, N4, and N5 in the Feature Fusion part represent feature maps of three different scales, and SFF represents the Selective Feature Fusion module.

[0086] Feature maps of different scales are first screened in the feature selection module, and after screening, the selective feature fusion mechanism is used to finally integrate the information of different layers in the obtained feature maps. The features generated by this fusion have rich semantic content, which helps to detect the subtle features of fabric defects, thereby enhancing the detection ability of the model.

[0087] (1) Feature selection module: In this process, the CA module plays a key role. The CA module first processes the input feature map f in ∈R C×H×W, the features obtained after the feature map is processed by two pooling layers, namely Max Pool and AvgPool, are integrated together; then the Sigmoid activation function is used to determine the weight value of each channel to obtain the weight f of each channel CA ∈R C×1×1 .

[0088] Feature Fusion Module: The multi-scale feature maps generated by the Backbone layer network have rich semantic information, but the target localization is relatively rough. On the contrary, the low-scale features provide accurate target position information, but the semantic information is limited. To solve this problem, a common method is to directly add the upsampled results of the high-level feature map and the low-scale feature map to enhance the semantic information of each layer. However, this technique does not perform feature selection and simply adds the pixel values of multiple feature layers. To address this limitation, the SFF module strategically fuses features by using high-level features as weights to filter out the basic semantic information embedded in the low-level features.

[0089] (2) Spatial Feature Fusion (SFF) uses high-level features as weights to strategically fuse features. Given the input high-level features and low-level features, first, the dimensions of the high-level features and low-level features are unified using transposed convolution and bilinear interpolation. Then, the channel feature self-filtering module is applied to convert the high-level features into corresponding weights for filtering the low-level features. Finally, the filtered low-level features are fused with the high-level features. The schematic diagram of SFF is as Figure 7 shown. Among them, TransposedConvolution represents transposed convolution, Bilinear Interpolation represents bilinear interpolation, and different cubes represent different feature maps. The input high-level feature f high ∈R C×H×W and the input low-level feature The high-level features are initially processed using transposed convolution with a stride of 2 and a convolution kernel of 3×3 to obtain features of size f high ∈R C×2H×2W . To unify the dimensions of the high-level features and low-level features, bilinear interpolation is first used to upsample or downsample the high-level features to match their dimensions with those of the low-level features, obtaining the feature Then, the attention CA module is used to convert the high-level features into corresponding attention weights to filter the low-level features and ensure that they have consistent dimensions. Finally, the filtered low-scale features are fused with the high-level features to enhance the model's ability to represent features and generate the final output features

[0090] S204. Introduce the Context Anchor Attention (CAA) mechanism to capture the potential relationships between distant pixels in the image, thereby enhancing the feature learning ability within the central key region;

[0091] Preferably, the CAA mechanism in S204 includes:

[0092] Performing spatial compression on the input feature map using global average pooling operation;

[0093] Further extracting and integrating the feature information of the local area through 1×1 convolution;

[0094] Adopting two lightweight depth strip convolutions, H-Conv2d and W-Conv2d, to efficiently capture and integrate pixel information along the height and width directions respectively, and identify and extract the features of objects with slender or special shapes.

[0095] Through the carefully designed HS-FPN architecture, a new type of lightweight network structure is successfully built above. This innovation significantly reduces the number of model parameters and computational volume, achieving higher computational efficiency, but inevitably brings a slight decrease in detection accuracy to a certain extent. In order to effectively recover and optimize the detection accuracy, the context-anchored attention CAA mechanism is ingeniously introduced in this paper. The core of the CAA mechanism lies in the combined application of its unique global average pooling and one-dimensional strip convolution. This combined strategy enables the network to accurately capture the potential relationships between distant pixels in the image, thereby enhancing the feature learning ability within the central key area. The CAA mechanism aims to master the context interdependence between pixels at different positions in the image, especially learning the distant pixel information that has a significant impact on the detection accuracy, to strengthen the feature representation of the central area.

[0096] In the specific implementation, CAA applies the global average pooling operation to perform spatial compression on the input feature map, and then further extracts and integrates the feature information of the local area through 1×1 convolution. In order to more effectively simulate and approximate the function of traditional large-kernel depth convolution, CAA also adopts two lightweight depth strip convolutions, H-Conv2d and W-Conv2d. This strip convolution structure can efficiently capture and integrate pixel information along the height and width directions, and is very suitable for identifying and extracting the features of objects with slender or special shapes (such as garbage like water bottles), so it also better meets the requirements of water garbage detection. This paper ingeniously integrates and optimizes CAA into HS-FPN, thereby effectively alleviating the problem of accuracy decline that the network structure may encounter in the detection task. The new CAA-HS-FPN module not only inherits the powerful feature extraction ability of the original HS-FPN, but also further enhances the network's attention to the key feature area, achieving a significant improvement in detection accuracy.

[0097] S205. Use DConv and Conv in the cross-stage context-guided feature fusion module to extract surrounding and local features from the target region respectively. After batch normalization and PReLU activation, the two are merged to form a preliminary combined feature, and then the above combined feature is dynamically adjusted and weighted accordingly.

[0098] Preferably, the cross-stage context-guided feature fusion module in S205 includes:

[0099] S205-1. Define local, nearby, and global operations as foc, fsur, and fglo respectively, and define the combination of local and nearby information as fjoi;

[0100] Calculate foc and fsur to learn the features of local and surrounding respectively. Foc is a 3*3 convolution to learn features from 8 adjacent points; fsur is a 3*3 dilated convolution to learn surrounding information; then perform the fioi operation, where fioi is a concat+bn+PRelu operation to fuse local and surrounding features;

[0101] S205-3. Fglo extracts global features and introduces channelwise weights.

[0102] ContextGuided (cross-stage context-guided feature fusion) module: Use DConv (dilated convolution) and Conv (standard convolution) in this module to extract surrounding and local features from the target region respectively. After BN (batch normalization) and PReLU (Parametric Rectified Linear Unit) activation, the two are merged to form a preliminary combined feature, and then the above combined feature is dynamically adjusted and weighted accordingly. As Figure 9 shown, define local, nearby, and global operations as foc, fsur, and fglo respectively, and define the combination of local and nearby information as fjoi. This module contains two main steps: (1) Calculate foc and fsur to learn the features of local and surrounding respectively. Foc is a 3*3 convolution to learn features from 8 adjacent points; fsur is a 3*3 dilated convolution to learn surrounding information. Subsequently, fioi is a concat+bn+PRelu operation to fuse local and surrounding features. (2) Fglo extracts global features and introduces channelwise weights (which is equivalent to connecting an SE layer to the foi feature).

[0103] S300. Input the pictures in the dataset into the YOLO model for training;

[0104] Model Training: First, create a configuration file (yaml file) to specify various parameters required for training, including data paths, the number of classes, class names, etc.; then load the YOLO model and call its training method, where you can choose to train from scratch or fine-tune based on a pre-trained model. Key parameters involved in the training process include the number of epochs, batch size, etc.; finally, use the trained model to predict the images in the validation set and save the prediction results for further analysis.

[0105] S400. Evaluate and compare the performance of the obtained YOLO model.

[0106] Obtain experimental data on the performance of the product or result:

[0107]

[0108] To prove the improved detection performance of the YOLOv8n - new model, on the TT100K dataset, YOLOv8n - new was compared and evaluated with YOLOv8. The Precision, Recall, mAP50, and mAP50 - 95 of the original YOLOv8 model were 75.6%, 67.7%, 75.5%, and 58.2% respectively; those of the improved model were 76.8%, 61.8%, 71.8%, and 54.2% respectively. In terms of the number of model parameters, it decreased from the original 3014423 to the current 1596507, and in terms of the GFLOPs metric, it decreased from the original 8.1 to 6.0.

[0109] This paper proposes a lightweight traffic sign detection algorithm based on YOLOv8, aiming to achieve accurate and real - time detection of traffic signs. This algorithm stabilizes the detection performance of the algorithm while meeting the real - time requirement.

[0110] Specifically, in the backbone part, this paper introduces the Adown module to improve the original Conv module. The main purpose of the Adown module is to retain as much feature information as possible during implementation, improve the quality of the feature map and the detection performance of the model, and reduce the parameters while maintaining the detection accuracy.

[0111] In the Neck part, HS - FPN is used for optimization. The high - level features are used as weights to filter the low - level feature information, and the filtered information is then merged with the high - level features to enhance the feature expression ability of the model, significantly reducing the number of parameters and improving the detection speed.

[0112] Experimental results show that, in terms of the number of parameters, the improved YOLOv8 algorithm is reduced to 59% of the original, and at the same time, it can also meet the accuracy of traffic sign detection. In the future, further research will be carried out to improve the detection accuracy and speed of the model, especially for complex environments and other situations, so that it can show higher reliability and efficiency in traffic sign detection tasks.

[0113] Example 2: The computer-readable storage medium of this example stores a computer program, and when this program is executed by a processor, it implements the steps in a lightweight traffic sign detection method based on YOLOv8 in Example 1.

[0114] The computer-readable storage medium of this example can be the internal storage unit of the terminal, such as the hard disk or memory of the terminal; the computer-readable storage medium of this example can also be the external storage device of the terminal, such as the plug-in hard disk, smart memory card, secure digital card, flash card, etc. equipped on the terminal; further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the terminal.

[0115] The computer-readable storage medium of this example is used to store the computer program and other programs and data required by the terminal, and the computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0116] Example 3: The computer device of this example includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a lightweight traffic sign detection method based on YOLOv8 in Example 1.

[0117] In this example, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc.; the memory can include a read-only memory and a random access memory, and provides instructions and data to the processor. A part of the memory can also include a non-volatile random access memory. For example, the memory can also store information about the device type.

[0118] Those skilled in the art should understand that the content disclosed in the examples can be provided as a method, a system, or a computer program product. Therefore, this solution can be implemented in the form of a hardware example, a software example, or a form combining software and hardware examples. Moreover, this solution can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0119] This solution is described with reference to the flowcharts and / or block diagrams of methods, and computer program products, according to embodiments of the solution. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions; these computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or block diagram Figure 1 or multiple blocks.

[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or block diagram Figure 1 or multiple blocks.

[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or block diagram Figure 1 or multiple blocks.

[0122] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0123] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A lightweight traffic sign detection method based on yolov8, characterized in that: The method includes: S100. Obtain traffic sign image data by using a high-definition camera to get a data set, where the data set includes a test set, a training set, and a validation set; S200. Build a YOLO model that improves yolov8; the YOLO model includes a Backbone network, a Neck network, and a Head network. The Backbone network is used for feature extraction, and the Neck network is used for feature fusion; the Head network is used to complete the final prediction task; S300. Input the pictures in the data set into the YOLO model for training; S400. Evaluate and compare the performance of the obtained YOLO model.

2. The lightweight traffic sign detection method based on YOLOv8 according to claim 1, wherein, The YOLO model that improves yolov8 in S200 includes: S201. In the Backbone network, replace the ordinary downsampling Conv module with a lightweight Adown module; S202. Use the WTConv module to relieve the change in the number of kernel sizes in the convolutional layer; S203. In the Neck network, use the HS-FPN module to filter the low-level feature information with the high-level features as weights, and then merge the filtered information with the high-level features to enhance the feature expression ability of the model; S204. Introduce the context anchor attention CAA mechanism to capture the potential relationship between distant pixels in the image, thereby enhancing the feature learning ability within the central key area; S205. Use DConv and Conv in the cross-stage context-guided feature fusion module to extract surrounding and local features from the target area respectively. After batch normalization and PReLU activation, the two are merged to form a preliminary joint feature, and then the above joint feature is dynamically adjusted and weighted.

3. The lightweight traffic sign detection method based on YOLOv8 according to claim 2, characterized in that, The lightweight Adown module in S201 includes: Perform an average pooling operation on the input feature map to reduce the size of the feature map and the computational amount of the subsequent convolutional layer; Process in two parts in the channel dimension. One part first goes through max pooling and then convolution, and the other part only performs convolution operations, where the parameters of the two convolutional layers are shared; Concatenate the output results to get the output.

4. The lightweight traffic sign detection method based on YOLOv8 according to claim 2, characterized in that, The WTConv module in S202 includes: S202-1. Use wavelet transform to filter and downsample the input low-frequency and high-frequency content; S202-2. Before constructing the output using IWT, perform a small-kernel depth convolution on different frequency maps according to the formula: Y = IWT(Conv(W, WT(X))); where X represents the input tensor, and W represents the weight tensor of the k×k depth kernel, and its input channel number is four times that of X; The specific operations include: in the low-frequency band of the second-level wavelet domain perform a 3×3 convolution to obtain a 9-parameter convolution, which responds to the low frequency in the 12×12 receptive field of the input X; adopt this first-level combined operation, according to the formula: Obtain the aggregated output Z starting from the i-th layer (i) , where the summation of two convolution outputs with different sizes is used as the output; Among them, W (i) , is the input of this layer, representing all three high-frequency maps of the i-th layer.

5. The lightweight traffic sign detection method based on YOLOv8 according to claim 2, characterized in that, The HS-FPN module in S203 includes a feature selection module and a feature fusion module; S203-1. In the feature selection module, the input feature map f is obtained in ∈R C×H×W , the features obtained after the feature map passes through two pooling layers, namely Max Pool and Avg Pool, are integrated together; and the Sigmoid activation function is used to determine the weight value of each channel to obtain the weight f of each channel CA ∈R C×1×1 ; S203-2. In the feature fusion module, high-level features are used as weights to strategically fuse features. Given the input high-level and low-level features, first, transposed convolution and bilinear interpolation are used to unify the dimensions of the high-level and low-level features. Then, the channel feature self-filtering module is applied to convert the high-level features into corresponding weights for filtering the low-level features. Finally, the filtered low-level features are fused with the high-level features. The specific operations include: Given the input high-level feature f high ∈R C×H×W and the input low-level feature f low ∈R C×H1×W1 The high-level feature is initially obtained using a transposed convolution with a stride of 2 and a 3×3 convolutional kernel, resulting in a feature of size f high ∈R C×2H×2W ; To unify the dimensions of the high-level and low-level features, the high-level feature is first upsampled or downsampled using bilinear interpolation to match its dimensions with those of the low-level feature, obtaining the feature f att ∈R C×H1×W1 ; Then, the attention CA module is used to convert the high-level feature into corresponding attention weights to filter the low-level feature and ensure they have consistent dimensions; Finally, the filtered low-scale feature is fused with the high-level feature to enhance the model's ability to represent features and produce the final output feature f out ∈R C×H1×W1 .

6. The lightweight traffic sign detection method based on YOLOv8 according to claim 2, characterized in that, The CAA mechanism in S204 includes: Using global average pooling operation to perform spatial compression on the input feature map; Further extracting and integrating the feature information of the local region through 1×1 convolution; Adopting two lightweight depthwise strip convolutions, H-Conv2d and W-Conv2d, to efficiently capture and integrate pixel information along the height and width directions respectively, and identify and extract the features of objects with slender or special shapes.

7. The lightweight traffic sign detection method based on YOLOv8 according to claim 2, wherein The cross-stage context-guided feature fusion module in S205 includes: S205-1. Define local, nearby, and global operations as foc, fsur, and fglo respectively, and define the combination of local and nearby information as fjoi; Calculate foc and fsur to learn the features of local and surrounding respectively. foc is a 3*3 convolution to learn features from 8 adjacent points; fsur is a 3*3 dilated convolution to learn surrounding information. Then perform the fioi operation, where fioi is a concat+bn+PRelu operation to fuse local and surrounding features; S205-3. fglo extracts global features and introduces channelwise weights.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the steps in a lightweight traffic sign detection method based on yolov8 as described in any one of claims 1-7.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps in a lightweight traffic sign detection method based on yolov8 as described in any one of claims 1-7.