A traffic sign detection method under foggy weather
By introducing a hollow selection nuclear attention mechanism and a lightweight asymptotic feature pyramid network based on YOLOv8, the problems of low accuracy and complex calculations in foggy traffic sign detection are solved, and high-precision and lightweight detection effects are achieved.
Patent Information
- Application Number
- CN202411192629.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-08-28
AI Technical Summary
The prior art has problems such as low detection accuracy, complex calculations, bloated models and poor characteristics in traffic sign detection under foggy conditions, which is difficult to meet the practical application needs.
The foggy traffic sign detection method based on YOLOv8 is adopted, combined with the hollow selection nuclear attention mechanism (ASKA) and the lightweight asymptotic feature pyramid network (C-AFPN), the receptive field range of the feature map is adjusted, and the detection accuracy is improved through the multi-head detection structure.
It realizes high-precision and lightweight traffic sign detection under foggy conditions, improves detection accuracy and generalization capabilities, and is suitable for practical applications of multiple devices.
Smart Images

Figure CN119181076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly, to a method for detecting traffic signs in foggy weather. Background Art
[0002] The research on traffic sign detection under foggy conditions aims to enable the driverless system to operate effectively in fog, thereby improving driving safety and convenience. Currently, there are mainly two methods for traffic sign detection: traditional methods and deep learning-based methods. In traditional methods, an exhaustive search strategy is mainly adopted, and classical techniques such as sliding windows and image scaling are used for object detection. Subsequently, object features are extracted by combining information about color and shape. Finally, methods such as support vector machines are used to classify the detected objects. However, traditional object detection methods still face some challenges, including the uncertainty of object positions and the variation of object sizes and shapes. To address these problems, rectangular bounding boxes are usually used to define objects. However, due to the different aspect ratios of objects, this method introduces new problems. The detection model requires high time and space costs, resulting in a complex and cumbersome detection process, which is not conducive to widespread adoption. In addition, traditional methods have low detection and recognition accuracy, complex calculations, bloated models, and poor feature robustness, all of which hinder practical applications. Therefore, they are not suitable for meeting the current requirements of traffic sign detection under foggy conditions.
[0003] With the progress of technology and the improvement of the computing power of devices such as computers, the deep learning-based method for traffic sign detection in fog has gradually become the focus of attention and a hot topic in current detection tasks. This method eliminates the cumbersome steps involved in defining features and achieves end-to-end object detection. It can be mainly divided into two categories: one-stage models and two-stage models. For foggy traffic sign detection based on two-stage models, such as R-CNN, Faster R-CNN, etc., it relies on a complete convolutional neural network for the process of target detection and extracts features through convolution. Its most obvious feature is that, during the training process, it is divided into two parts. The first part is to train the RPN network, and the second part is to train the network for the target detection region. This requires generating candidate boxes through an algorithm and then using a convolutional neural network to classify these samples. Compared with traditional methods, the advantage of this type of target detection algorithm is simplicity and higher accuracy. However, the disadvantage is slower speed. On the other hand, one-stage detection methods, such as SSD and YOLO, do not require the participation of the RPN network. They directly obtain the position and category information of objects through the main network. The main advantage brought by this difference is faster algorithm implementation speed, but the disadvantage is that the accuracy is not as good as that of two-stage models.
[0004] Therefore, in order to better identify traffic signs in road scenes under foggy weather, it is necessary to propose an object detection method that can reduce the interference caused by background noise, has high detection accuracy, and a lightweight and uncomplicated model. The YOLO series of object detection algorithms have high detection efficiency, but perform poorly when detecting noisy images. There are still some defects in the original model that can be improved, such as false alarms, missed detections, large model size, and small detection scale. Summary of the Invention
[0005] According to the technical problems mentioned in the above background art, a traffic sign detection method under fog is provided. The present invention can accurately and quickly detect traffic signs under fog. The method of the present invention not only ensures the lightweight of the model, but also improves the detection accuracy and generalization ability, and is more suitable for the traffic sign detection task under foggy conditions.
[0006] In the deep learning method of the present invention, the Atrous Selective Kernel Attention mechanism (ASKA) is further incorporated to adjust the receptive field range in the feature map, and the lightweight Progressive Feature Pyramid Network C-AFPN is added to alleviate the conflict of information fusion between different scales.
[0007] The technical means adopted by the present invention are as follows: A traffic sign detection method under fog based on YOLOv8, including the following steps:
[0008] S1. Construct a foggy traffic sign dataset: Obtain a public dataset, preprocess the images in the public dataset, and divide the preprocessed dataset into a training set and a test set according to a certain ratio;
[0009] S2. Construct a network model for foggy traffic sign detection;
[0010] S3. Train the network model constructed in S2: Input the training set data into the network model, train the network model, and obtain a trained model;
[0011] S4. Complete the test by testing the trained model in S3 with the test set.
[0012] Further, S1 includes the following steps:
[0013] S11: Download the traffic sign dataset CCTSDB 2021 from the GitHub website;
[0014] S12: Use the method of Mosaic data augmentation to augment the traffic sign dataset;
[0015] S13: The dataset CCTSDB 2021 contains foggy images and other weather images; using the OpenCV library in Python, fog is added to the non-foggy images in the dataset CCTSDB 2021, and the foggy images are not processed; a new dataset containing the fog-added images and foggy images is obtained;
[0016] S14: Randomly divide all the pictures in the new dataset into a training set and a test set in a ratio of 9:1 to obtain the final foggy weather traffic sign dataset.
[0017] Further, the S2 includes the following steps:
[0018] S21: Design the Atrous Selective Kernel Attention mechanism ASKA, and add the attention mechanism ASKA to the last layer of the backbone network Backbone; the Atrous Selective Kernel Attention mechanism ASKA includes: large kernel convolution, atrous convolution, standard convolution, pooling operation and Sigmoid activation function;
[0019] S22: Replace some convolutional blocks in the backbone network Backbone layer with Group Shuffle Convolution GSConv to accelerate the model training speed;
[0020] S23: Introduce the lightweight asymptotic feature pyramid network C-AFPN of the neck network Neck with the idea of lightweight asymptotic feature fusion;
[0021] S24: Implement the multi-head detection structure of the Head layer, that is, use four convolutional layers with scale sizes of 20×20, 40×40, 80×80 and 160×160 to detect traffic signs of different scale sizes;
[0022] S25: Set the SIoU loss function as the regression loss function of the YOLOv8 network; the calculation formula of the SIoU loss function is:
[0023]
[0024] where, Δ represents the distance loss, and Ω represents the shape loss;
[0025] S26: Combine the Backbone layer, Neck layer and Head layer and the SIoU loss function to form a complete foggy weather traffic sign detection model.
[0026] Further, the S3 includes the following steps:
[0027] S31: Set the size of the input network model picture to a fixed size of: 640×640 pixels;
[0028] S32: Select the Stochastic Gradient Descent (SGD) function as the optimization function in the network model; set the number of training epochs Epoch to 300, the batch size Bachsize to 16, the initial learning rate to 0.001, and the strategy for adjusting the learning rate. After each round of training, the network model automatically saves the corresponding weight information;
[0029] S33: Determine the size of the training epoch Epoch;
[0030] When the training epoch Epoch is less than 50, freeze the Backbone layer, that is, set the learning rate of the Backbone layer to 0 and stop updating it;
[0031] When the epoch Epoch is greater than or equal to 50, unfreeze the Backbone layer and restore it to the initial learning rate;
[0032] When the epoch Epoch reaches 300, save the final model weights, end the model training, and output the evaluation of the model's performance, specifically including the number of parameters, accuracy P, recall R, and mean average precision mAP. The specific calculation formulas are as follows:
[0033]
[0034]
[0035]
[0036] Among them, TP represents the number of positive examples predicted as positive examples, FP represents the number of negative examples predicted as positive examples, and FN represents the number of positive examples predicted as negative examples.
[0037] Furthermore, the following steps are included in S4:
[0038] S41: Set the image resolution size to 640 pixels × 640 pixels;
[0039] S42: Perform target prediction on the images in the test set; during the prediction process, the model will generate target boxes to capture the targets. The target boxes represent the range of the target objects in the image, and at the same time, a class and a target confidence are assigned to the target boxes;
[0040] S43: Output the prediction results, that is, the class, confidence, and target box coordinate position information of the traffic signs in the image, and the detection ends.
[0041] Furthermore, the following steps are included in S21:
[0042] S211: Obtain the background information features of different regions through large kernel convolution and dilated convolution, and reduce the number of parameters of the model;
[0043] S212: Coordinate the channel dimensions of various feature maps through standard convolution, and perform average pooling and max pooling operations to extract spatial relationships;
[0044] S213: Obtain the weights from different convolutions through the Sigmoid activation function; perform weighted processing on the weights and the feature maps obtained above to obtain the attention feature Y; the calculation formula is:
[0045]
[0046] where, represents convolution, represents the features from different convolutions, represents the spatial pooling features.
[0047] Furthermore, in S23, the following steps are included:
[0048] S231: First, obtain two features extracted from the image by the third layer and the fifth layer in the Backbone layer; perform upsampling and downsampling operations on these two features respectively;
[0049] S232: Perform adaptive spatial fusion on the two features after the sampling operation, and the fusion calculation formula is:
[0050]
[0051] where, and represent the initial feature vectors, and represent the spatial weights of the two levels of features at different levels, represents the fused feature vector;
[0052] S233: Extract visual features from the spatially fused features through a fused phantom convolution block;
[0053] S234: Perform adaptive spatial fusion on the features obtained in S233 and the visual features of the seventh layer in the Backbone layer, and extract visual features through a phantom convolution block;
[0054] S235: Perform adaptive spatial fusion on the features obtained in S234 and the visual features of the tenth layer in the Backbone layer, and extract visual features through a phantom convolution block.
[0055] Compared with the prior art, the present invention has the following advantages:
[0056] 1. The present invention proposes a hole selection kernel attention mechanism ASKA. This mechanism can dynamically adjust the range of the receptive field in the feature map, reducing the impact of fog noise on traffic signs to be detected. It uses dilated convolution to expand the receptive field and a selective attention mechanism to enhance the feature extraction ability.
[0057] 2. The present invention proposes a lightweight progressive feature pyramid network (C-AFPN), which can solve the suboptimal fusion result problem caused by non-adjacent scales in the feature pyramid network. Features are gradually fused from low scale to high scale, alleviating the semantic conflict between non-adjacent levels. The subsequent multi-head detection structure can adapt to the diversity of different traffic sign scales and improve the detection accuracy.
[0058] 3. The present invention proposes a foggy weather traffic sign detection model improved based on YOLOv8. While having high precision, it also maintains a lightweight effect, can be deployed in practical applications, and has excellent detection effects on traffic signs in foggy weather.
[0059] For the above reasons, the present invention can be widely applied to drones, cars, monitors and other devices in foggy weather. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0061] Figure 1 It is a flowchart of a foggy weather traffic sign detection method based on improved YOLOv8 of the present invention;
[0062] Figure 2 It is a schematic diagram of the YOLOv8 structure of the present invention;
[0063] Figure 3 It is a schematic diagram of the ASKA attention mechanism structure of the present invention;
[0064] Figure 4 It is a schematic diagram of the C-AFPN network structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0066] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0067] As Figures 1-4 shown, the present invention provides a foggy weather traffic sign detection method based on YOLOv8, including the following steps:
[0068] S1. Construct a foggy weather traffic sign dataset: Obtain a public dataset, preprocess the images in the public dataset, and divide the preprocessed dataset into a training set and a test set according to a certain ratio.
[0069] Preferably, in the present application, S1 includes the following steps:
[0070] S11: Download the traffic sign dataset CCTSDB 2021 from the GitHub website; it can be understood that in other embodiments, different public datasets can also be downloaded to implement the detection of traffic signs.
[0071] S12: Use the Mosaic data augmentation method to augment the traffic sign dataset; first randomly select four pictures, and perform random scaling and random cropping on these four pictures. The random scaling is controlled by the scaling ratio scale, which is a random number between 0.6 and 0.8. The random cropping is to perform a cropping operation at a random position, and crop the image into a fixed size of 640×640 pixels. Then, splice the four processed pictures at the upper left, upper right, lower left, and lower right positions into a new picture.
[0072] S13: The dataset contains foggy images and images of other weather conditions. Using the OpenCV library in Python, fog is added to the non-foggy images in the dataset, and the foggy images are not processed.
[0073] S14: All the images in the dataset are divided into a training set and a test set in a ratio of 9:1 to obtain the final foggy traffic sign dataset. It can be understood that in other implementation manners, the specific division ratio can be determined according to the actual number of data in the dataset, and training with accurate and clear line-of-sight data and achieving testing are required.
[0074] Further, S2: Construct a network model for foggy traffic sign detection. The S2 includes the following steps:
[0075] S21: Design a dilated selective kernel attention mechanism ASKA and add the attention mechanism ASKA to the last layer of the backbone network Backbone; the dilated selective kernel attention mechanism ASKA includes: large kernel convolution, dilated convolution, standard convolution, pooling operation, and Sigmoid activation function; in the S21, the following steps are included:
[0076] S211: The image data obtains background information features of different regions through large kernel convolution and dilated convolution, reducing the number of parameters of the model;
[0077] S212: Coordinate the channel dimensions of various feature maps through standard convolution, and perform average pooling and max pooling operations to extract spatial relationships;
[0078] S213: Through the Sigmoid activation function, obtain the weights from different convolutions; perform weighted processing on the weights and the feature maps obtained above to obtain the attention feature Y. The calculation formula is:
[0079]
[0080] where, represents convolution, represents the features from different convolutions, represents the spatial pooling features.
[0081] S22: Replace some convolution blocks in the backbone network Backbone layer with group shuffle convolution GSConv to accelerate the model training speed;
[0082] S23: With the idea of lightweight progressive feature fusion, design a lightweight progressive feature pyramid network C-AFPN for the neck network Neck; in the S23, the following steps are included:
[0083] S231: First, obtain the two features extracted from the image by the third and fifth layers in the Backbone layer. The features mainly include basic visual elements such as color, texture, shape, and edge. Then, perform upsampling and downsampling operations on these two features respectively;
[0084] S232: Perform adaptive spatial fusion on the two features after the sampling operation. The fusion calculation formula is:
[0085]
[0086] Among them, and represent the initial feature vectors, and represent the spatial weights of the two levels of features at different levels, represents the fused feature vector;
[0087] S233: Extract visual features from the spatially fused features through a fused phantom convolution block;
[0088] S234: Perform adaptive spatial fusion on the features obtained in S233 with the visual features of the seventh layer in the Backbone layer, and extract visual features through a phantom convolution block.
[0089] S235: Perform adaptive spatial fusion on the features obtained in S234 with the visual features of the tenth layer in the Backbone layer, and extract visual features through a phantom convolution block.
[0090] S24: Implement the multi-head detection structure of the Head layer, that is, use four convolutional layers with scale sizes of 20×20, 40×40, 80×80, and 160×160 to detect targets of different scale sizes;
[0091] S25: Set the SIOU loss function as the regression loss function of the improved YOLOv8 network; The calculation formula of the SIOU loss function is:
[0092]
[0093] Among them, Δ represents the distance loss, and Ω represents the shape loss;
[0094] S26: Combine the Backbone layer, Neck layer, Head layer, and SIoU loss function to form a complete foggy-weather traffic sign detection model. The image data first enters the Backbone layer for feature extraction, then enters the Neck layer for feature fusion, and finally enters the Head layer. The Head layer generates candidate bounding boxes to predict the objects in the image. During the process of predicting the object's position, size, and category, the SIoU loss function adjusts the prediction results through distance loss and shape loss to achieve the best balance in prediction.
[0095] As a preferred implementation, S3: Train the network model constructed in S2: Input the training set data into the network model and train the network model. The following steps are included in S3:
[0096] S31: Set the input image to the network to a fixed size of 640×640 pixels.
[0097] S32: Select the Stochastic Gradient Descent (SGD) function as the optimization function in the network. Set the number of epochs Epoch to 300, the batch size Bachsize to 16, the initial learning rate to 0.001, and the strategy for adjusting the learning rate. The model will save the model weights once every 1 round of training.
[0098] S33: Start training the network. When the number of epochs Epoch is less than 50, freeze the Backbone layer, that is, set the learning rate of the Backbone layer to 0 and stop updating it.
[0099] S34: When the number of epochs Epoch is greater than or equal to 50, unfreeze the Backbone layer and restore the learning rate.
[0100] S35: When the number of epochs Epoch reaches 300, save the final model weights, end the model training, and output the evaluation of the model's performance, specifically including the number of parameters, accuracy P, recall R, and mean average precision mAP. The specific calculation formulas are as follows:
[0101]
[0102]
[0103]
[0104] Among them, TP represents the number of positive examples predicted as positive examples, FP represents the number of negative examples predicted as positive examples, and FN represents the number of positive examples predicted as negative examples.
[0105] Preferably, S4: Complete the test by testing the model trained in S3 with the test set. The following steps are included in S4:
[0106] S41: Set the picture resolution size to 640×640 pixels;
[0107] S42: During the training process, the higher the mean average precision (mAP) value of the model, the better the detection effect of the model, that is, more targets are accurately detected and there are fewer cases of missed detection and false detection. Select the best model during the training process to predict the targets in the pictures in the test set. During the prediction process, the model will generate target boxes to capture the targets, use the target boxes to represent the range of the target objects in the image, and at the same time assign a class and a target confidence to the target boxes;
[0108] S43: Output the predicted results, that is, the class, confidence, and target box coordinate position information of the traffic signs in the picture, and the detection ends.
[0109] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0110] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0111] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0112] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0113] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0114] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A traffic sign detection method in foggy weather based on YOLOv8, characterized in that: The following steps are involved: S1. Construct a foggy traffic sign dataset: obtain a public dataset, preprocess the images in the public dataset, and divide the preprocessed dataset into a training set and a test set according to a certain ratio; S2, constructing a network model for foggy traffic sign detection; S2 comprises the following steps: S21: Design a hole selection kernel attention mechanism ASKA, and add the attention mechanism ASKA to the last layer of the backbone network Backbone; the hole selection kernel attention mechanism ASKA includes: large kernel convolution, hole convolution, standard convolution, pooling operation and Sigmoid activation function; S22: Replace some convolution blocks in the backbone network Backbone layer with group shuffle convolution GSConv to speed up model training; S23: Based on the idea of lightweight asymptotic feature fusion, a lightweight asymptotic feature pyramid network C-AFPN of the neck network Neck is introduced; S23 includes the following steps: S231: first obtain two features extracted from the image by the third and fifth layers in the Backbone layer; perform upsampling and downsampling operations on the two features respectively; S232: Adaptively spatially fuse the two features after the sampling operation. The fusion calculation formula is: in, and represents the initial eigenvector, and Represents the spatial weights of the features at two levels at different levels, Represents the fused feature vector; S233: extract visual features from the spatially fused features by fusing the phantom convolutional blocks; S234: Adaptively spatially fuse the features obtained in S233 with the seventh layer visual features in the Backbone layer, and extract visual features through the phantom convolution block; S235: Adaptively spatially fuse the features obtained in S234 with the tenth layer visual features in the Backbone layer, and extract visual features through the phantom convolution block; S24: Implement the multi-head detection structure of the Head layer, that is, use four convolutional layers with scales of 20×20, 40×40, 80×80, and 160×160 to detect traffic signs of different scales; S25: Set the SIOU loss function as the regression loss function of the YOLOv8 network; the calculation formula of the SIOU loss function is: Among them, Δ represents distance loss, Ω represents shape loss; S26: The foggy traffic sign detection model is represented by Backbone layer, Neck layer, Head layer and SIOU loss function; S3, training the network model constructed in S2; inputting the data of the training set into the network model, training the network model, and obtaining a trained model; S4. Test the model trained in S3 using a test set, complete the test, and obtain the final model.
2. The method for detecting traffic signs in foggy weather based on YOLOv8 according to claim 1, characterized in that: The S1 comprises the following steps: S11: Download the traffic sign dataset CCTSDB 2021 from the GitHub website; S12: Use the Mosaic data enhancement method to enhance the traffic sign dataset; S13: The dataset CCTSDB 2021 includes foggy images and other weather images; fogging the fog-free images in the dataset CCTSDB 2021 using the OpenCV library in Python, while leaving the foggy images unprocessed; obtaining a new dataset including fogged images and foggy images; S14: All the images in the new data set are randomly divided into a training set and a test set in a ratio of 9:1 to obtain a final foggy traffic sign data set.
3. The method for detecting traffic signs in foggy weather based on YOLOv8 according to claim 1, characterized in that: The S3 includes the following steps: S31: Set the image input to the network model to a fixed size of 640×640 pixels; S32: Selecting the stochastic gradient descent function SGD as the optimization function in the network model; setting the training round Epoch to 300, the batch size Bachsize to 16, the initial learning rate to 0.001 and the strategy for adjusting the learning rate, and the network model automatically saves the corresponding weight information after each training round; S33: Determine the size of the training epoch; When the training epoch is less than 50, freeze the Backbone layer, that is, set the learning rate of the Backbone layer to 0 and no longer update it; When the Epoch is greater than or equal to 50, the Backbone layer is no longer frozen and the initial learning rate is restored; When the epoch reaches 300, the final model weight is saved, the model training is completed, and the evaluation of the model effect is output, including the number of parameters, accuracy P, recall rate R and average precision mAP. The calculation formula is as follows: Among them, TP represents the number of positive examples predicted as positive examples, FP represents the number of negative examples predicted as positive examples, and FN represents the number of positive examples predicted as negative examples.
4. The method for detecting traffic signs in foggy weather based on YOLOv8 according to claim 1, characterized in that: The S4 includes the following steps: S41: set the image resolution to 640 pixels × 640 pixels; S42: Predict the target for the images in the test set. During the prediction process, the model generates a target box to capture the target, uses the target box to represent the range of the target object in the image, and assigns a category and target confidence to the target box. S43: Output the predicted result, i.e., the category, confidence level, and target frame coordinate position information of the traffic sign in the image, and the detection is completed.
5. The method for detecting traffic signs in foggy weather based on YOLOv8 according to claim 1, characterized in that: The S21 includes the following steps: S211: Obtain background information features of different regions through large kernel convolution and hole convolution to reduce the number of model parameters; S212: Standard convolution is used to coordinate the channel dimensions of various feature maps, and average pooling and maximum pooling operations are performed to extract spatial relationships; S213: Obtain weights from different convolutions through the Sigmoid activation function; perform weighted processing on the weights and the feature map obtained above to obtain the attention feature Y; the calculation formula of the attention feature Y is: in, represents convolution, represents features from different convolutions, Represents spatial pooling features.
Citation Information
Patent Citations
Foggy day traffic sign detection method and device, storage medium and electronic equipment
CN117437615A
Foggy day traffic sign detection method based on deep learning
CN117765507A
Cited By
Method and system for detecting traffic signs in severe weather based on multistage optimization and meteorological classification
CN122454534A