A Pedestrian Crosswalk Detection Method Based on Improved YOLOv5
By improving the YOLOv5 network model, the C2f module, CBAM, DE-head detection head module, FPN-HALF feature extraction module and Alpha-IoU loss function are used to solve the accuracy of pedestrian crosswalk detection in complex scenarios, realizing timely detection and identification in harsh environments, and reducing the risk of traffic accidents.
Patent Information
- Application Number
- CN202310693824.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-12
AI Technical Summary
In complex scenarios, it is difficult for the existing technology to accurately detect and identify crosswalks, which makes it difficult for drivers to detect in time in severe weather or complex environments, increasing the risk of traffic accidents.
A pedestrian crosswalk detection method based on improved YOLOv5 is proposed. By constructing the CDE-YOLOv5 network model, the C2f module, the spatial channel attention mechanism module (CBAM), the DE-head detection head module, the FPN-HALF feature extraction module and the new bounding box regression loss function, the Alpha-IoU loss function, is used to improve the detection head structure and feature extraction capability.
In complex scenarios, timely detection and identification of pedestrian crosswalks can be achieved, and voice broadcasts can be performed in real time, reminding drivers to slow down and reduce the risk of traffic accidents.
Smart Images

Figure CN116883963B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic image processing, and specifically relates to a crosswalk detection method based on improved YOLOv5. Background Technique
[0002] In recent years, with the progress of technology, the vehicle industry and road systems have developed rapidly. As the number of vehicles increases year by year, the resulting traffic accidents have become more frequent. The application of autonomous driving technology can effectively reduce the accident rate of urban traffic. As an important traffic sign, it is particularly important for autonomous driving technology to accurately detect and identify crosswalks. In complex scenarios such as urban roads, due to the influence of various factors, such as bad weather, strong light, low illuminance or high occlusion, etc., drivers need to be highly concentrated to accurately identify, or they need to be close to identify, which is prone to traffic accidents. Therefore, detecting crosswalks at a long distance to reduce the burden on drivers and improve driving safety is of great significance.
[0003] The existing visual perception road detection algorithms mainly include methods based on road features, methods based on road models, and methods based on convolutional neural networks. The first two methods belong to traditional methods and require designing mathematical models to extract target features. For different types of targets, or even the same type of target but with different backgrounds, the mathematical models are also different. Therefore, the traditional target detection methods have poor adaptability and strong dependence on the background. With the rapid development of deep learning, using deep neural networks to automatically learn high-level features from a large amount of data and using the learning results for target detection and recognition, this technology has been increasingly applied to visual perception. Therefore, a crosswalk detection method based on the CDE-YOLOv5 model is proposed. Summary of the Invention
[0004] The present invention aims to at least solve one of the technical problems existing in the prior art; for this reason, the present invention proposes a crosswalk detection method based on improved YOLOv5.
[0005] A crosswalk detection method based on improved YOLOv5, the method specifically includes the following steps:
[0006] S1. Construct a crosswalk detection data set, and the detection data set is several training set pictures and test set pictures of crosswalks;
[0007] S2. Perform data augmentation processing, and use fog-adding data augmentation processing on the labeled crosswalk detection data set to expand the data set;
[0008] S3. Construct a CDE-YOLOv5 network model, specifically as follows:
[0009] S31. Construct the C2f module to replace the original C3 module structure;
[0010] S32. Construct the Spatial Channel Attention Mechanism module (CBAM) and add it to the corresponding network layer;
[0011] S33. Construct the DE-head detection head module and replace the original detection head structure;
[0012] S34. Construct the FPN-HALF feature extraction module;
[0013] S35. Add the FPN-HALF module to the corresponding network layer in the neck module;
[0014] S36. Construct a new bounding box regression loss function, the Alpha-IoU loss function, and replace the original CIOU loss function; use the Alpha-CIOU loss function to replace the original CIOU bounding box regression loss function;
[0015] S37. Use the labeled dataset to complete the training of the model.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0017] Through the detection of images, the present invention can detect and analyze videos in real time. In complex scenarios, it can perform timely voice broadcasts for the crosswalks that appear. When a crosswalk target is detected, it can timely remind the driver to slow down and avoid accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is the network structure diagram of the crosswalk detection model CDE-YOLOv5 and the DE-head network structure diagram of the technical solution of the present invention;
[0019] Figure 2 It is the C2f structure diagram of the technical solution of the present invention;
[0020] Figure 3 It is the structure diagram of the DE-head module of the technical solution of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to Figures 1 - 3As shown in the figure, this application provides a crosswalk detection method based on improved YOLOv5. The method specifically includes the following steps:
[0023] S1. Construct a crosswalk detection dataset;
[0024] Use an in-vehicle camera to save the video of the road conditions frame by frame in complex scenarios of urban roads, such as rainy days and nights. Manually screen the saved pictures to obtain pictures containing crosswalk detection targets. Use the labelimg tool to label the crosswalks in the pictures to obtain 3,080 crosswalk training set pictures and 354 test set pictures;
[0025] S2. Perform data augmentation processing. Use fog-adding data augmentation processing on the labeled crosswalk detection dataset. This is an existing technology, so no specific description will be given; Expand the dataset; enable the model to obtain better detection effects in complex foggy scenarios;
[0026] S3. Construct a CDE-YOLOv5 network model, specifically as follows:
[0027] S31. Construct a C2f module to replace the original C3 module structure;
[0028] The C2f module is mainly composed of a CBL module, a split operation, a Bottleneck module, and a concat operation; among them, the CBL module is composed of a 1×1 convolution with a stride of 1, a BN layer, and a SiLU activation function. The input passes through the CBL module and the split operation respectively to obtain two feature outputs with the same number of channels. One of the outputs passes through a parallel Bottleneck operation and then is concatenated with the output of the other split operation for channel fusion, and then passes through the CBL module to obtain the final output feature matrix;
[0029] S32. Construct a convolutional block attention module (CBAM) and add it to the corresponding network layer;
[0030] The convolutional block attention module (CBAM) is composed of a spatial attention module and a channel attention module respectively. The calculation formula of the channel attention module is as follows:
[0031] F1 = σ(W2f(W1x1))
[0032] Where σ represents the Sigmoid activation function, W1 and W2 represent linear transformation matrices, f represents the global pooling function, x1 represents the input feature matrix, and F1 represents the output feature matrix.
[0033] The formula of the spatial attention module is as follows:
[0034] F2 = σ(f(x2))
[0035] f represents pooling in the spatial dimension, and the output F2 of the obtained spatial attention module is the weight in the channel dimension of the feature map. Multiplying F2 by the input x1 can obtain the enhanced feature map.
[0036] S33. Construct a DE-head detection head module and replace the original detection head structure;
[0037] The improved new DE-head detection head is mainly composed of two parallel branches. The outputs of the C2f structures of the 22nd, 26th, and 30th layers of the network are respectively used as the inputs of the detection heads for large, medium, and small targets. First, the number of channels is changed to 256 layers through a 1×1 convolutional layer, and two identical parallel branches are used to output for the classification task, confidence task, and localization task respectively. The parallel branches are composed of 3×3 convolutional modules and 1×1 convolutional modules.
[0038] S34. Construct an FPN-HALF feature extraction module;
[0039] The FPN-HALF module is composed of a CBL module, an upsampling module, a concatenation (concat) operation, and a C2f module, and is inserted after the FPN structure of the neck module. First, the input from the upper layer passes through the CBL module and the upsampling structure and then performs a concatenation (concat) operation with the shallow feature map of the second layer to enhance the information interaction level between the shallow and deep networks, and then outputs through the C2f structure.
[0040] S35. Add an FPN-HALF module to the corresponding network layer in the neck module;
[0041] S36. Construct a new bounding box regression loss function, the Alpha-IoU loss function, and replace the original CIOU loss function; use the Alpha-CIOU loss function to replace the original CIOU bounding box regression loss function.
[0042] Alpha-C IOU The loss function is as follows:
[0043]
[0044]
[0045]
[0046] Among them, IoU is the intersection over union of the ground truth box and the detection box, w gt , h gtDenote the width and height of the ground truth box, and \(w, h\) represent the width and height of the predicted box. The \(v\) calculated using formula (2) is a parameter for measuring the aspect ratio consistency. \(b\) and \(b_{gt}\) represent the center points of the predicted box and the target box respectively, and \(p\) represents the Euclidean distance between the two calculated center points. \(c\) represents the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box. The main role of \(\beta\) obtained through formula (3) is to promote the optimization of the loss function towards an increase in the overlapping region. The \(\alpha\) exponent is generally taken as 3 to increase the gradient and accelerate convergence.
[0047] S37. Complete the training of the model using the labeled dataset.
[0048] S4. Extract the features of the target crosswalk and give voice prompts.
[0049] Input the crosswalk test set into the final crosswalk network model to test the detection accuracy and detection speed of the network model; it is found that by training several different networks and the original YOLOv5 model using the same labeled dataset, the improved CDE - YOLOv5 model has obvious advantages in the detection accuracy of the crosswalk and can meet the requirements of real - time detection.
[0050] Table 1 Comparison table of model detection accuracy and speed:
[0051]
[0052] As Figure 2 shown, the CDE - YOLOv5 object detection network uses the C2f structure to replace the original C3 structure. While ensuring the efficient extraction of network features, its structure is more lightweight and can obtain richer gradient flow information.
[0053] The CDE - YOLOv5 object detection network further extracts the feature map of the output feature map of the C2f structure in the neck module using the channel - spatial attention mechanism (CBAM). This attention mechanism is mainly composed of a channel attention structure and a spatial attention structure. The channel attention module performs attention weighting on the channel dimension of the input feature map, reduces the dependence on useless information in the feature map, and enhances the feature extraction ability for important information. The spatial attention mechanism enhances the network's attention to local details by strengthening the spatial dimension of the input feature map. At the same time, the channel - spatial attention mechanism module (CBAM) can keep the dimensions of the input and output feature maps the same.
[0054] As Figure 1As shown in the figure, the FPN-HALF module of CDE-YOLOv5 mainly consists of a convolutional layer, an upsampling layer, a concatenation (concat) operation, and a C3 module. By performing a concatenation operation on the output of the upsampling layer and the shallow feature map of the backbone network, the interaction level between the deep feature information and the shallow feature information is improved, and feature information is better extracted. The output of the C3 module is input into the DE-head module to introduce a new head detection head, enhancing the multi-scale learning ability of the network.
[0055] As Figure 3 shown, the CDE-YOLOv5 model uses a new detection head module, the DE-head module. Since the feature information emphasized by the classification and localization tasks is different, the classification and localization tasks are regressed and output separately using two parallel structures, thereby improving the detection accuracy of the model.
[0056] The CDE-YOLOv5 model uses the Alpha-IoU bounding box regression loss function to replace the original CIOU loss function, which better handles the loss of high IOU targets and gradient adaptive weighting, improving the accuracy of bounding box regression.
[0057] Through the detection of images, the present invention can perform real-time detection and analysis of videos. In complex scenarios, it can perform timely voice announcements for the pedestrian crossings that appear. When a pedestrian crossing target is detected, the driver is timely reminded to slow down to avoid accidents.
[0058] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A crosswalk detection method based on improved YOLOv5, characterized in that, The method specifically includes the following steps: S1. Construct a crosswalk detection dataset, which consists of several training set images and test set images of crosswalks; S2. Perform data augmentation processing. Apply fogging data augmentation processing to the labeled crosswalk detection dataset to expand the dataset; S3. Construct a CDE-YOLOv5 network model, specifically as follows: S31. Construct a C2f module to replace the original C3 module structure; S32. Construct a channel attention mechanism module CBAM and add it to the corresponding network layer; S33. Construct a DE-head detection head module and replace the original detection head structure; The DE-head detection head consists of two parallel branches. The outputs of the C2f structures of the 22nd, 26th, and 30th layers of the network are used as the inputs of the detection heads for large, medium, and small targets respectively. First, the number of channels is changed to 256 layers through a 1×1 convolutional layer, and two identical parallel branches are used to output the classification task, confidence task, and localization task respectively. The parallel branches consist of a 3×3 convolutional module and a 1×1 convolutional module; S34. Construct an FPN-HALF feature extraction module; The FPN-HALF feature extraction module consists of a CBL module, an upsampling module, a concatenation (concat) operation, and a C2f module. Insert this module after the FPN structure of the neck module, that is, insert the previous components after the FPN structure of the neck to form the FPN-HALF feature extraction module; First, the input from the upper layer passes through the CBL module and the upsampling structure, and then a concatenation (concat) operation is performed with the shallow feature map of the second layer to enhance the information interaction level between the shallow and deep network layers, and then the output is obtained through the C2f structure; S35. Add the FPN-HALF module to the corresponding network layer in the neck module; S36. Construct a new bounding box regression loss function, the Alpha-IoU loss function, and replace the original CIOU loss function; use the Alpha-CIOU loss function to replace the original CIOU bounding box regression loss function; S37. Use the labeled dataset to complete the training of the model; Apply the trained model. When a crosswalk is detected, a voice prompt is given.
2. The method for detecting crosswalks based on improved YOLOv5 according to claim 1, wherein, The specific method for constructing the crosswalk detection dataset in step S1 is as follows: Use an in-vehicle camera to save the road condition videos frame by frame in complex scenarios on urban roads, manually screen the saved pictures to obtain pictures containing crosswalk detection targets, and use the labelimg tool to label the crosswalks in the pictures to obtain 3080 crosswalk training set pictures and 354 test set pictures.
3. The pedestrian crossing detection method based on improved YOLOv5 according to claim 2, characterized in that, The complex scenarios include rainy days and nights.
4. A crosswalk detection method based on improved YOLOv5 according to claim 1, characterized in that In step S31, the C2f module consists of a CBL module, a split operation, a Bottleneck module, and a concatenation (concat) operation; Among them, the CBL module consists of a 1×1 convolution with a stride of 1, a BN layer, and a SiLU activation function. The input passes through the CBL module and the split operation respectively to obtain two feature outputs with the same number of channels. One of the outputs undergoes a parallel Bottleneck operation and then is concatenated with the output of the other split operation through a concat operation, and finally passes through the CBL module to obtain the final output feature matrix.
5. A crosswalk detection method based on improved YOLOv5 according to claim 1, characterized in that, In step S32, the spatial channel attention mechanism CBAM is composed of a spatial attention mechanism module and a channel attention mechanism module respectively. The calculation formula of the channel attention module is as follows: F1 = σ(W2f(W1x1)); where σ represents the Sigmoid activation function, , represents the linear transformation matrix, f represents the global pooling function, represents the input feature matrix, represents the output feature matrix; The formula of the spatial attention module is as follows: F2 = σ(f(x2)) f represents pooling in the spatial dimension, and the output of the obtained spatial attention module is the weight in the channel dimension of the feature map. Multiply with the input to obtain the enhanced feature map.
6. The method for detecting crosswalks based on improved YOLOv5 according to claim 1, wherein, Alpha-C in step S36 IOU The loss function is as follows: Among them, IoU is the intersection over union of the ground truth box and the detection box, , represents the width and height of the ground truth box, w and h represent the width and height of the predicted box, and v calculated using formula (2) is a parameter for measuring the aspect ratio consistency. b, respectively represent the center points of the predicted box and the target box, and p represents the Euclidean distance between the two center points; c represents the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box; β obtained through formula (3) serves to promote the optimization of the loss function towards an increase in the overlapping region, and the α exponent takes a value of 3 to increase the gradient and accelerate convergence.
Citation Information
Patent Citations
Railway wagon brake shoe fault detection method based on deep learning
CN114399672A
Real-time target detection method suitable for embedded platform
CN114898171A