Unmanned vehicle traffic cone detection method based on dark channel defogging and improved YOLOv8

By using dark channel defog and improving the YOLOv8 model in the traffic cone detection system of unmanned vehicles, the problem of insufficient accuracy and real-time performance of traditional detection algorithms in foggy environments is solved, and higher detection accuracy and speed are achieved.

CN120088756APending Publication Date: 2025-06-03HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510081822.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional traffic cone detection algorithms perform poorly under complex environments and specific tasks, especially in foggy environments, and the complex calculation of high-precision detection algorithms lead to slow detection speeds and cannot meet the real-time requirements of unmanned vehicles.

Method used

Pretreatment method based on dark channel defog is adopted to restore image information under the influence of fog, and an attention mechanism is added to the YOLOv8 model, Slim-neck is introduced to replace the "neck" layer of the original model, and the CLLAHead detection head is used to replace the original detection head to improve detection accuracy and speed.

Benefits of technology

It significantly improves the accuracy of autonomous vehicles in foggy scenes, ensures the reliability of autonomous driving system in judgment of road conditions, and maintains efficient detection speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088756A_ABST
    Figure CN120088756A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned vehicle traffic cone detection method based on dark channel defogging and improved YOLOv8, and the method comprises the steps: carrying out the preprocessing of a foggy scene traffic cone image obtained by a camera of an unmanned vehicle through a dark channel defogging algorithm, and obtaining a defogged traffic cone image through an atmospheric scattering model; a ShuffleAttention attention mechanism is added to a Backbone network layer of an original YOLOv8 model, the ShuffleAttention firstly divides a channel dimension into a plurality of sub-features, and then the sub-features are subjected to parallel processing; adopting Slim-check to replace the structure of the Neck part in the original YOLOv8 model; a CLLAHead detection head is adopted to replace an original YOLOv8 model detection head; according to the method, the remote traffic cone is detected, the rate of missed alarm and false alarm is reduced, and compared with an original YOLOv8n model, the detection precision and the detection speed are improved, and the traffic cone detection accuracy of an unmanned vehicle in a foggy day scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting an unmanned vehicle, and in particular to a method for detecting traffic cones of an unmanned vehicle based on dark channel defogging and improved YOLOv8. Background Art

[0002] With the rapid development of science and technology, driverless technology has become a research hotspot in the automotive industry and the field of artificial intelligence. Driverless vehicles are designed to achieve autonomous navigation, obstacle recognition and avoidance through various sensors and intelligent algorithms, thereby improving the safety and efficiency of transportation. In a complex traffic environment, accurately identifying various traffic elements is one of the key tasks of the driverless system. Traffic cones are a common traffic warning sign. Rapid and accurate detection of them is extremely important for the path planning and safe driving of driverless vehicles.

[0003] Traditional traffic cone detection algorithms have certain limitations. Early traffic cone detection algorithms mostly rely on traditional image processing techniques, such as edge detection and color segmentation. These methods may achieve certain results under ideal lighting conditions and simple backgrounds. However, in actual road scenes, lighting changes and weather factors (such as fog) will affect the detection results. For example, in a foggy environment, the contrast and clarity of the image are greatly reduced, and the color and outline of the traffic cone become blurred, making it difficult for traditional algorithms based on color and edge features to accurately extract effective information about the traffic cone, resulting in a sharp decrease in detection accuracy.

[0004] In recent years, deep learning target detection algorithms such as the YOLO series and Faster R-CNN have achieved great success in the field of target detection and have been widely used in traffic cone detection tasks. Although these algorithms can effectively detect traffic cones under normal circumstances, there are still some problems. For example, they are less adaptable to severe weather conditions such as foggy days. Due to the interference of fog, it is difficult for the network to learn the accurate feature representation of traffic cones, which affects the detection performance; the balance between detection accuracy and speed is also a problem. Some high-precision target detection algorithms are often complex in calculation, resulting in slow detection speed and unable to meet the strict real-time requirements of unmanned vehicles. For example, although some algorithms based on complex feature extraction networks perform well in detection accuracy, in actual unmanned driving systems, they will miss key traffic cone information because the processing speed cannot keep up with the vehicle's driving speed, causing safety hazards.

[0005] Therefore, it is urgent to propose a method for unmanned vehicle to detect traffic cones to overcome the limitations of traditional traffic cone detection algorithms under complex environments and specific task requirements, and to improve the safety and reliability of unmanned vehicles under various road conditions. Summary of the invention

[0006] Objective of the Invention: Aiming at the deficiencies in the existing technology, the present invention proposes a traffic cone detection method for driverless vehicles based on dark channel defogging and improved YOLOv8. In complex weather conditions, especially in foggy environments, the dark channel defogging algorithm is used to preprocess the image first, restore the image information blurred and attenuated due to fog, and add an attention mechanism to the YOLOv8 model, introduce Slim-neck to replace the "neck" layer of the original model, and use the CLLAHead detection head to replace the original detection head, so as to improve the detection accuracy and speed, significantly improve the accuracy of traffic cone detection by driverless vehicles in foggy scenarios, and ensure the reliability of the driverless system's judgment of road conditions.

[0007] Technical Solution: The traffic cone detection method for driverless vehicles based on dark channel defogging and improved YOLOv8 of the present invention includes the following steps:

[0008] Step (1), preprocess the foggy scene traffic cone image obtained by the driverless vehicle camera through the dark channel defogging algorithm, and use the atmospheric scattering model to obtain the defogged traffic cone image; the defogging process is as follows:

[0009] x is the spatial coordinate of the image, I(x) represents the foggy traffic cone image, J(x) represents the fog-free traffic cone image, A represents the global atmospheric light value, and t(x) represents the transmittance;

[0010] I(x) = J(x)t(x) + A(1 - t(x)) (1)

[0011] In the local non-sky area, the dark channel of the fog-free traffic cone image J(x) is expressed as follows:

[0012]

[0013] Among them, Ω(x) represents the local area centered on x, and J c (y) represents the result of minimum filtering of the grayscale image with the same size as the stored original traffic cone image.

[0014] The dark channel of the fog-free traffic cone image J(x) in the non-sky area tends to 0, which is expressed as:

[0015] J dark →0 (3)

[0016] The formula (1) atmospheric scattering model is transformed to obtain formula (4):

[0017]

[0018] Assume that the transmittance t(x) of each window is a constant, denoted as Calculate the dark channel on both sides of formula (4) to obtain the point with the minimum pixel value:

[0019]

[0020] Formula (5) represents taking the point with the minimum pixel value in the three-channel regions covered by a window area in the traffic cone image. From the fact that the dark channel intensity of the traffic cone image in the non-sky area tends to 0, we get:

[0021]

[0022] It is deduced that:

[0023]

[0024] Substitute equation (7) to obtain the estimated transmittance value As follows:

[0025]

[0026] I c (y) represents taking the dark channel map on the foggy traffic cone image; using formula (8), the estimated transmittance value under the window is obtained The transmittance map t(x) of the entire foggy traffic cone image is obtained, and at the same time, is introduced to control the degree of defogging, that is, formula (9), represents the degree of defogging, and 0.95 is taken in the embodiment.

[0027]

[0028] Record the positions of the top 0.1% of the pixels in the dark channel map according to the brightness, and find the value of the highest brightness point at the position in the original foggy traffic cone map as the estimate of the A value. After calculation by formula (1), the defogged traffic cone image is obtained:

[0029]

[0030] Step (2), add the ShuffleAttention attention mechanism to the Backbone network layer of the original YOLOv8 model. ShuffleAttention first divides the channel dimension into multiple sub-features, and then performs parallel processing on the sub-features. The steps are as follows:

[0031] Step (2.1): Group the input features. Assume the input feature is X ∈ R C×H×W , and split the input feature X along the channel dimension into G groups: X = [X 1 , ……, X G , R C / G×H×W . The sub-feature Xk ∈R C×H×W , each sub - feature X k is split into two branches along the channel dimension: X k1 , X k2 ∈R C / 2G×H×W , one branch X k1 learns channel attention features, and the other branch X k2 learns spatial attention features.

[0032] Step (2.2): When calculating channel attention, the combination of GAP + Scale + Sigmoid is adopted. Among them, GAP is global average pooling, which reduces model parameters and avoids overfitting; Scale is to scale the data, and by adjusting the numerical range of the data, the data is distributed within the interval; the Sigmoid function is a non - linear activation function, and the process is as follows:

[0033]

[0034] where s is the result obtained after the global average pooling operation, represents the function corresponding to global average pooling, X k1 is the input data, H is the height of X k1 , W is the width of X k1 , is the normalization factor for calculating the average value, X' k1 is the result after processing the input data, σ represents the Sigmoid activation function, represents a linear transformation of the result, W 1 is the weight matrix, b 1 is the bias term;

[0035] Step (2.3): When calculating spatial attention, first use Group Norm to process X k2 to obtain statistical data at the spatial level as the input data, and then use the F c (·) function to enhance the input data, and the process is as follows:

[0036] X' k2 = σ(W 2 ·GN(X k2 ) + b 2 )·X k2

[0037] where X k2 is the input data, X' k2 is the result after processing X k2 , σ represents the Sigmoid activation function, GN represents Group Norm normalization, W 2is the weight matrix, b 2 is the bias term;

[0038] Step (2.4): After calculating the above two kinds of attention, integrate the calculation results X' k1 and X' k2 First, fuse them through Concat to get X' k = [X' k1 , X' k2 ∈ R C / 2G×H×W , and adopt the channel permutation operation for inter-group communication; the output of ShuffleAttention has the same size as the input, enabling ShuffleAttention to be embedded into the existing CNN architecture, as shown in Figure 1 . This attention mechanism not only reduces the complexity of the YOLOv8 model but also maintains high accuracy when the driverless vehicle detects traffic cones.

[0039] Step (3): Replace the structure of the Neck part in the original YOLOv8 model with Slim-neck. The process is as follows:

[0040] Step (3.1): Perform ordinary convolution operations on the input data to achieve downsampling;

[0041] Step (3.1.1): Determine the parameters of the downsampling convolution, including the number of input channels, the number of output channels, the convolution kernel size, the stride, and the padding;

[0042] Step (3.1.2): Apply ordinary convolution operations, create an ordinary convolution layer using the selected parameters, and perform convolution operations on the input data to achieve downsampling;

[0043] Step (3.2): Use DWConv depth convolution to process the data; the process is as follows:

[0044] Step (3.2.1): Determine the parameters of the depth convolution, including the number of input channels, the number of output channels, the convolution kernel size, the stride, and the padding;

[0045] Step (3.2.2): Apply depth convolution operations, create a depth convolution layer using the selected parameters, and perform convolution on the input data;

[0046] Step (3.3): Use the concatenation function to concatenate SC_output obtained through ordinary convolution and DSC_output obtained through depth convolution;

[0047] Step (3.3.1): Perform ordinary convolution and depth convolution operations, and use ordinary convolution and depth convolution modules to process the input data;

[0048] Step (3.3.2): Check the consistency of the output dimensions of height, width, and number of channels in ordinary convolution and depthwise convolution;

[0049] Step (3.3.3): Use a concatenation function to concatenate SC_output obtained through ordinary convolution and DSC_output obtained through depthwise convolution.

[0050] Step (3.4): Perform a shape transformation operation on the data tensor and return the output tensor. Divide the channels into two groups through grouped convolution and perform a shuffle operation between these two groups so that the corresponding number of channels of the previous two convolutions are together, thus arranging the corresponding number of channels of the results produced by the previous two different convolutions together, and completing the entire GSConv processing process.

[0051] Step (3.5): Introduce the GSbottleneck module and the VoV-GSCSP module in the Neck layer. The GSbottleneck module enhances the network's feature processing ability by stacking GSConv modules, forming a GSbottleneck structure by serially combining multiple GSConv modules. In the custom forward function, use the traffic cone feature map processed by the GSbottleneck module as the input of the VoV-GSCSP module for feature extraction and fusion. See the diagrams of the GS bottleneck module and the VoV-GSCSP module in Figure 2 。

[0052] Step (4): Replace the detection head in the original YOLOv8 model with the CLLAHead detection head, and the process is as follows:

[0053] Step (4.1): Extract multi-scale values from the traffic cone feature maps of different layers. The shallow network captures local details for detecting small-sized traffic cones, and the deep network is used to detect large-sized traffic cones; and enhance the feature expression through the channel attention mechanism. First, calculate the importance of each channel, use global average pooling to compress the spatial dimension of each channel into a single value to obtain a vector of the channel dimension; then pass the vector through a fully connected network to learn the weight of each channel; finally, multiply the obtained weight by each channel of the traffic cone feature map, enhance the important channels, and suppress the unimportant channels, thereby enhancing the feature expression.

[0054] Step (4.2): Focus on the regions on the traffic cone feature map through the self-attention mechanism.

[0055] Step (4.2.1): Obtain the query, key, and value matrices by linearly transforming the traffic cone feature map:

[0056] Q = F'W QK = F'W K V = F'W V

[0057] Wherein, F' is the traffic cone feature map, Q is the query, K is the key, V is the value matrix, and W Q is the first linear transformation matrix; W K is the second linear transformation matrix, and W V is the third linear transformation matrix;

[0058] Step (4.2.2): Calculate the dot product between the query matrix and the key matrix:

[0059] S = QK T

[0060] Wherein, S is the attention score matrix;

[0061] Step (4.2.3): Normalize the attention score matrix:

[0062]

[0063] Wherein, S ij represents the attention score between the i-th query and the j-th key, represents the normalized result;

[0064] Step (4.2.4): Multiply by the value matrix to obtain the result of weighted summation:

[0065]

[0066] Wherein, A is the output of the self-attention mechanism, and this output enables the network to focus on the area where the traffic cones are located, thereby reducing background interference.

[0067] Step (4.3): Integrate the hierarchical and focal information and then output the final bounding box and class prediction. The multi-scale information extracted from different layers and the traffic cone feature map processed by the self-attention mechanism are further fused through convolution operations. Perform multiple convolution operations on the fused traffic cone feature map to extract higher-level feature representations, and at the same time use normalization and activation functions to stabilize the training and enhance the non-linear expression ability; the focal information obtained by the attention mechanism indicates the important areas in the traffic cone feature map. Multiply the focal information element-wise with the above-mentioned fused traffic cone feature map to make the network pay more attention to the important areas, and input the final traffic cone feature map into the bounding box and class prediction branches.

[0068] Optimize the predicted bounding box using the regression loss:

[0069]

[0070] Among them, B is the real border, is the predicted border;

[0071] Use cross-entropy loss to optimize class prediction:

[0072]

[0073] Among them, y is the real class, is the predicted class probability;

[0074] According to the set confidence threshold, filter out the bounding boxes with low confidence and remove the overlapping bounding boxes, and finally output the bounding boxes and classes of the detected traffic cones. The structure diagram of the CLLAead distribution focus detection head is shown in Figure 3 .

[0075] In step (2.1), importance coefficients are generated for each group of input features through the Spatial and Channel attention modules.

[0076] In step (3), when using Slim-neck to replace the structure of the Neck part in the original YOLOv8 model, GSConv is introduced to replace the SC module in the original YOLOv8 model.

[0077] In step (3.1), the parameters of the downsampling convolution include the number of input channels, the number of output channels, the kernel size, the stride, and the padding.

[0078] The process of step (3.3) is as follows: Step (3.3.1): Perform ordinary convolution and depthwise convolution operations, and use the ordinary convolution and depthwise convolution modules to process the input data; the process is, according to the ordinary convolution operation in step (3.1), create an ordinary convolution layer with the determined parameters, input the input data into the ordinary convolution layer, and this ordinary convolution layer performs convolution operations on each channel of the input data; similarly, establish a depthwise convolution layer according to step (3.2), and the depthwise convolution layer performs convolution operations on each channel of the input data separately, the number of channels of the input data remains unchanged, and the number of output channels is determined according to the settings of the depthwise convolution layer;

[0079] Step (3.3.2): Check the output dimension consistency of the height, width, and number of channels in the ordinary convolution and depthwise convolution;

[0080] In step (3.3.1), according to step (3.2), a depthwise convolution layer is established, and the depthwise convolution layer performs convolution operations on each channel of the input data. The number of channels of the input data remains unchanged, and the number of output channels is determined by the depthwise convolution layer.

[0081] In step (3.3.3), in PyTorch, the torch.cat function is used to concatenate the SC_output obtained through ordinary convolution and the DSC_output obtained through depth convolution.

[0082] In step (4.1), the obtained weights are multiplied by each channel of the traffic cone feature map to enhance the important channels.

[0083] In step (4.3), the non-maximum suppression method is used to remove overlapping bounding boxes.

[0084] Working principle: The present invention adopts the dark channel dehazing algorithm. By estimating the dark channel of the image and combining with the atmospheric scattering model, a haze-free image is restored, and the atmospheric light and transmittance are accurately estimated. The calculation efficiency is higher, and good dehazing effects are achieved in different degrees of foggy environments, and the edge and texture detail information of the traffic cone are retained; when processing traffic cone image data, the ShuffleAttention attention mechanism added in the Backbone network layer effectively combines spatial attention and channel attention using the Shuffle unit; the introduced Slim-neck is a structure for optimizing the "neck" part in the convolutional neural network; the "neck" is the part connecting the backbone network and the head network of the CNN, responsible for feature fusion and processing to improve the accuracy and efficiency of detection; the used CLLAHead detection head improves the recognition and localization ability of the targets in the image through multi-level feature extraction and integration, using the distribution focal loss function and the attention mechanism, meeting the high requirements of accurate and real-time detection of traffic cones in the driverless scenario.

[0085] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0086] (1) In the context of driverless traffic cone detection, the present invention is a method for detecting traffic cones of a driverless vehicle based on dark channel dehazing and improved YOLOv8. Compared with the detection results before dehazing, the average precision and accuracy of the YOLOv8 model are significantly improved after dark channel dehazing; the YOLOv8 model with an added attention mechanism, an improved neck network, and a replaced original detection head can effectively detect traffic cones at a long distance, improves false positives and false negatives, and has a good edge recognition effect, with an improvement in the mean average precision and accuracy compared to the original YOLOv8n model;

[0087] (2) The present invention significantly improves the quality of foggy images, enhances the contrast between the traffic cone and the background, making the features of the traffic cone more obvious, thus providing a clearer and more accurate input image for subsequent target detection algorithms and improving the detection accuracy of traffic cones in foggy environments.

[0088] (3) The ShuffleAttention attention mechanism added in the present invention enhances the network's ability to focus on the features of traffic cones, making the recognition of traffic cones by driverless vehicles more efficient and effective.

[0089] (4) The Slim-neck structure adopted in the present invention reduces the complexity of the YOLOv8n model while maintaining the accuracy of traffic cone recognition by driverless vehicles.

[0090] (5) The present invention uses the CLLAHead distributed focus detection head, adding a hierarchical focus attention mechanism to the detection head design, strengthening the capture of traffic cones by driverless vehicles, with less computational complexity and being more conducive to real-time detection. Description of the Drawings

[0091] Figure 1 It is the ShuffleAttention module diagram added to the improved YOLOv8 model of the present invention;

[0092] Figure 2 It is the GS bottleneck module and VoV-GSCSP module diagram introduced in the improved YOLOv8 model of the present invention;

[0093] Figure 3 It is the structural diagram of the CLLAHead distributed focus detection head used in the improved YOLOv8 model of the present invention;

[0094] Figure 4 It is the YOLOv8 network structure diagram adopted by the present invention;

[0095] Figure 5 It is the flowchart of the method for detecting traffic cones by driverless vehicles based on dark channel dehazing and improved YOLOv8 of the present invention;

[0096] Figure 6 It is the diagram of the added position of the ShuffleAttention attention mechanism added by the present invention;

[0097] Figure 7 It is the design diagram of the Slim-neck for the YOLO series adopted by the present invention;

[0098] Figure 8 It is the schematic diagram of the structural principle of the CLLAHead detection head adopted by the present invention;

[0099] Figure 9 It is the LabelImg operation interface diagram when labeling the traffic cone dataset of the present invention;

[0100] Figure 10 It is the training process diagram of the present invention;

[0101] Figure 11This is the training result diagram of traffic cones in foggy days for the present invention;

[0102] Figure 12 This is the training result diagram of traffic cones after defogging for the present invention;

[0103] Figure 13 This is the detection accuracy diagram of traffic cones before defogging for the present invention;

[0104] Figure 14 This is the detection accuracy diagram of traffic cones after defogging for the present invention;

[0105] Figure 15 This is the effect diagram of part of the validation set after defogging for the present invention;

[0106] Figure 16 This is the result diagram of part of the test set after defogging for the present invention;

[0107] Figure 17 This is the improved network structure diagram of the present invention;

[0108] Figure 18 This is the process diagram of data augmentation for the present invention;

[0109] Figure 19 This is the training result diagram after improvement for the present invention;

[0110] Figure 20 This is the accuracy diagram after improvement for the present invention;

[0111] Figure 21 This is the effect diagram of part of the validation set after improvement for the present invention;

[0112] Figure 22 This is the result diagram of part of the test set after improvement for the present invention. Detailed implementation manners

[0113] As Figures 1 to 22 shown, the traffic cone detection method for driverless vehicles based on dark channel defogging and improved YOLOv8 of the present invention preprocesses the foggy scene traffic cone images obtained by the driverless vehicle camera through the dark channel defogging algorithm, and uses the atmospheric scattering model to obtain clear images of traffic cones after defogging.

[0114] The present invention divides the preprocessed dataset into three folders: a training set, a validation set, and a test set, named train, val, and test respectively. Inside each folder path, an images folder and a labels folder are created respectively. The total number of images is 1963, which is divided into 1570 for the training set, 196 for the validation set, and 197 for the test set according to the ratio of 8:1:1. The images are saved in the images folders of each respective folder. The annotation tool used is LabelImg. The save location of the annotation files is selected as the labels folders of each respective folder. The annotation format uses the txt format of YOLO. The annotation box is selected as a rectangle box, and the annotation name is named "traffic cone". The LabelImg operation interface is as Figure 9 .

[0115] The original YOLOv8 model is used to detect the traffic cone dataset. The process is as follows:

[0116] Firstly, the original YOLOv8 model is used for training, validation, and testing. The selected YOLOv8n model belongs to a lightweight model, which not only ensures the test speed but also the detection efficiency. The weight file named yolov8n.pt is selected. Secondly, a dataset parameter file is created. According to the relevant folder locations, the locations of the training set, validation set images, and labels are input, and the class value is set to 1, and the detected object is named traffic cone. Finally, the hyperparameter settings are adjusted in the relevant files, including the training weight file, training model file, dataset parameter file, number of training epochs, number of batch processing files, and input size. To ensure the test efficiency and speed, the training device uses GPU training. After multiple trainings, it is found that the effect is best when the number of training epochs is 300. Figure 10 Screenshot during training.

[0117] The training results of traffic cones in foggy weather are as Figure 11 shown, and the training results of traffic cones after dark channel defogging are as Figure 12 shown. The first row is the bounding box loss, object detection loss, classification loss, precision, and recall rate of the training set. The second row is the bounding box loss, object detection loss, classification loss, and two mean average precisions mAP of the validation set. The results show that although the various curves decrease and increase regularly, the overall fluctuation is large and unstable.

[0118] The change curve of the detection accuracy of traffic cones before defogging is as Figure 13 shown. The average accuracy of the dataset before defogging on the original YOLOv8 model is 91.9%, and the average accuracy of the dataset after dark channel defogging on the original YOLOv8 model is 92.0%, which is improved to a certain extent as Figure 14 shown.

[0119] The verification effect of the dehazed validation set is as follows Figure 15 shown. Configure the relevant hyperparameter settings, select the best weight file trained by the original YOLOv8 model, adjust the dataset parameter file and the image size, and start the test. Some of the test results are shown in Figure 16 shown. The original YOLOv8 model cannot achieve an ideal accuracy for detecting traffic cones in some far scenes. There are problems in edge processing, and it does not completely enclose the traffic cones. At the same time, there are phenomena of missed detection and false detection.

[0120] After modifying the YOLOv8 model, the traffic cone dataset is detected. The process is as follows:

[0121] In the original YOLOv8 model, add the ShuffleAttention attention mechanism to the backbone network layer, replace the neck network layer in the original model with Slim-neck, and replace the original detection head with the CLLAHead detection head. First, create the core code files for adding the attention mechanism, Slim-neck, and CLLAHead in the modules directory of the original YOLOv8 respectively. At the same time, add the ShuffleAttention, GSConv, VoVGSCSP, and CLLAHead modules to the parse_model function. Secondly, create a training model file. This file adds the ShuffleAttention attention mechanism to the last layer of the backbone grid, adds GSConv and VoVGSCSP to the head grid, and replaces the last layer of Detect with the CLLAHead distribution focus detection head. After successful addition, it is as shown in Figure 17 shown.

[0122] After successful modification, select the yolov8n.pt weight file and the created training model file for the training model. Similar to the hyperparameter setting adjustment of the original YOLOv8 model, first perform model training, and perform data augmentation in the last ten rounds. The process is as shown in Figure 18 shown, and the various curves after training are as shown in Figure 19 shown. Compared with the original model, the fluctuations of various curves are significantly reduced, and the trend basically tends to be stable. The accuracy rate is as shown in Figure 20 shown. The accuracy rate of the improved model is as high as 93.7%, which is improved compared with the original model. The verification effect of the improved validation set is as shown in Figure 21 shown. Use the best file in the improved trained model for testing. Some of the test result pictures are as shown in Figure 22 shown, and the test results are good.

[0123] The final results show that the YOLOv8 model with the addition of the ShuffleAttention attention mechanism, replacement with Slim-neck, and replacement with the CLLAHead detection head has a higher accuracy in detecting traffic cones in the driverless scenario. Moreover, the attention mechanism is lightweight and flexible enough to consider not only channel information but also spatial information, which can improve the efficiency in traffic cone detection, is more excellent in detecting small traffic cones at a long distance, and has higher accuracy and efficiency in traffic cone detection than the original model, reducing the missed detection and false detection rates, and also improving the recognition ability of various traffic cones under different lights and different angles.

Claims

1. A traffic cone detection method for unmanned vehicles based on dark channel defogging and improved YOLOv8, characterized by: The following steps are involved: Step (1) preprocesses the traffic cone image in the foggy scene acquired by the camera of the driverless car through the dark channel defogging algorithm, and obtains the defogged traffic cone image using the atmospheric scattering model; the process is as follows: x is the spatial coordinate of the image, I(x) represents the foggy traffic cone image, J(x) represents the fog-free traffic cone image, A represents the global atmospheric light value, and t(x) represents the transmittance; I(x)=J(x)t(x)+A(1-t(x)) (1) In the local area of ​​non-sky, the dark channel of the fog-free traffic cone image J(x) is as follows: Among them, Ω(x) represents the local area centered on x, J c (y) represents the result of minimum filtering of a grayscale image of the same size as the original traffic cone image stored; The dark channel of the fog-free traffic cone image J(x) in the non-sky area tends to 0, which is expressed as: J dark →0 (3) The atmospheric scattering model of equation (1) is transformed into equation (4): Assume that the transmittance t(x) of each window is a constant, denoted as Calculate the dark channel on both sides of equation (4) and get the point with the minimum pixel value: The dark channel intensity of the traffic cone image in the non-sky area tends to 0, so: Derived: Substituting equation (7) into the formula, we can get the transmittance estimate: Ic ( y) represents the dark channel image taken on the foggy traffic cone image; using formula (8), the transmittance estimate under the window is obtained And the transmittance map t(x) of the entire foggy traffic cone image is obtained, and the Control the degree of defogging: is the degree of defogging; The first 0.1% of pixels in the dark channel image are recorded in the image according to the brightness, and the value of the highest brightness point in the original foggy traffic cone image is used as the estimate of the A value. The defogging traffic cone image is calculated by formula (1): Step (2): Add the ShuffleAttention mechanism to the Backbone network layer of the original YOLOv8 model. ShuffleAttention first divides the channel dimension into multiple sub-features, and then processes the sub-features in parallel. The steps are as follows: Step (2.1): Group the input features. Assume that the input features are X∈R C×H×W , split the input feature X into G groups along the channel dimension: x = [X1, ..., X G ], R C / G×H×W ; Sub-feature X k ∈R C×H×W is split into the first branch X along the channel dimension k1 and the second branch X k2 , X k1 , X k2 ∈R C / 2G×H×W , the first branch learns the channel attention features, and the second branch learns the spatial attention features; Step i2.2): Use GAP+Scale+Sigmoid combination to calculate channel attention: Among them, s is the result after the global average pooling operation, represents the function corresponding to the global average pooling, X k1 is the input data, H is X k1 The height of W is X k1 The width of is the normalization factor for calculating the mean value, X′ k1 is the result of processing the input data, σ represents the Sigmoid activation function, Indicates a linear transformation of the result, W1 is the weight matrix, and b1 is the bias term; Step (2.3): When calculating the spatial attention, use Group Norm to calculate X k2 The statistical data at the spatial level are processed as input data, and then F c (·) function enhances the input data, the process is as follows: X′ k2 =σ(W2·GN(X k2 )+b2)·X k2 Among them, X k2 is the input data, X′ k2 Yes X k2 The processed result, σ represents the Sigmoid activation function, GN represents GroupNorm normalization, W2 is the weight matrix, and b2 is the bias term; Step (2.4): Calculate the result X′ k1 , X′ k2 Integration, first through Concat fusion to get X′ k =[X′ k1 , X′ k2 ]∈R C / 2G×H×W ,channel replacement is used for inter-group communication; Step (3) uses Slim-neck to replace the structure of the Neck part in the original YOLOv8 model. The process is as follows: Step (3.1): Create a normal convolution layer using the parameters of the downsampling convolution, and perform a normal convolution operation on the input data for downsampling; Step (3.2): taking the number of input channels, the number of output channels, the convolution kernel size, the step size and the padding as parameters of the depthwise convolution; creating a depthwise convolution layer according to the parameters to perform depthwise convolution on the input data; Step (3.3): Use the concatenation function to concatenate the SC_output obtained by the ordinary convolution and the DSC_output obtained by the deep convolution; Step (3.4): perform a shape transformation operation on the data tensor and return the output tensor; Step (3.5): Introduce the GSBottleneck module and the VoV-GSCSP module into the Neck layer. In the forward function, connect the GSConv modules in series to form the GSBottleneck module. Use the traffic cone feature map processed by the GSBottleneck module as the input of the VoV-GSCSP module for feature extraction and fusion. Step (4): Use CLLAHead detection head to replace the detection head in the original YOLOv8 model. The process is as follows: Step (4.1): Extract multi-scale values ​​from the traffic cone feature map and enhance feature expression through the channel attention mechanism. First, calculate each channel and use global average pooling to compress the spatial dimension data of each channel into a single value to obtain a channel-dimensional vector; then pass the vector through a fully connected network to learn the weight of each channel; multiply the obtained weight by each channel of the traffic cone feature map; Step (4.2): Focus on the area on the traffic cone feature map through the self-attention mechanism; Step (4.2.1) transforms the traffic cone feature graph linearly to obtain the query, key and value matrices: Q=F'W Q K=F'W K V=F'W V Among them, F' is the traffic cone feature map, Q is the query, K is the key, V is the value matrix, and W Q is the first linear transformation matrix; W K is the second linear transformation matrix, W V is the third linear transformation matrix; Step (4.2.2): Calculate the dot product between the query matrix and the key matrix: S=QK T Among them, S is the attention score matrix; Step (4.2.3): Normalize the attention score matrix: Among them, S ij represents the attention score between the i-th query and the j-th key, represents the normalized result; Step (4.2.4): Multiplying with the value matrix gives the weighted sum: Among them, A is the output of the self-attention mechanism; Step (4.3) uses regression loss to optimize the predicted bounding box: Among them, B is the real border, To predict the bounding box; Use cross entropy loss to optimize category prediction: Among them, y is the true category, is the predicted category probability; According to the set confidence threshold, low-confidence bounding boxes are filtered, overlapping bounding boxes are removed, and the bounding boxes and categories of the detected traffic cones are output.

2. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (2.1), the importance coefficient is generated for each set of input features through the Spatial and Channel attention modules.

3. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (3), when Slim-neck is used to replace the structure of the Neck part in the original YOLOv8 model, GSConv is introduced to replace the SC module in the original YOLOv8 model.

4. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (3.1), the parameters of the downsampling convolution include the number of input channels, the number of output channels, the convolution kernel size, the step size and the padding.

5. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: The process of step (3.3) is: Step (3.3.1): Perform ordinary convolution and deep convolution operations, and use ordinary convolution and deep convolution modules to process the input data; Step (3.3.2): Check the consistency of the output dimensions of height, width, and number of channels in ordinary convolution and depthwise convolution; Step (3.3.3): Use the concatenation function to concatenate the SC_output obtained by ordinary convolution and the DSC_output obtained by deep convolution.

6. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (3.3.1), a deep convolutional layer is established according to step (3.2). The deep convolutional layer performs a convolution operation on each channel of the input data. The number of channels of the input data remains unchanged, and the number of output channels is determined by the deep convolutional layer.

7. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (3.3.3), in PyTorch, the torch.cat function is used to concatenate the SC_output obtained by ordinary convolution and the DSC_output obtained by deep convolution.

8. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (4.1), the obtained weight is multiplied by each channel of the traffic cone feature map to enhance the important channels.

9. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (4.3), the non-maximum suppression method is used to remove overlapping bounding boxes.

10. The method for detecting traffic cones for unmanned vehicles based on dark channel defogging and improved YOLOv8 according to claim 1, characterized in that: In step (4.3), the extracted multi-scale values ​​and the traffic cone feature map processed by the self-attention mechanism are fused through a convolution operation.

Citation Information

Cited By

  • Landslide identification method and system for collapsible loess slope, terminal and storage medium

    CN121170569A