A method for nighttime vehicle recognition based on improved YOLOv5
By improving the YOLOv5 network model and combining BiFPN and C3DRSN modules, the problem of low accuracy and efficiency of night vehicle detection is solved, and more efficient night vehicle identification and detection is achieved.
Patent Information
- Application Number
- CN202210783690.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-05
AI Technical Summary
The existing vehicle detection algorithms are not accurate and efficient when used at night, and cannot be effectively applied to night scenes.
The improved YOLOv5 network model is adopted, combined with the BiFPN network and the C3DRSN module to perform night vehicle identification and detection. Through improved network structure and modules, this method enhances the extraction and fusion of night vehicle features, and improves detection accuracy and speed.
It improves the accuracy and speed of night vehicle detection, and can be more effectively applied to night traffic scenes and meets real-time detection needs.
Smart Images

Figure CN115311529B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition and detection, and particularly relates to a method for nighttime vehicle recognition based on improved YOLOv5. Background Art
[0002] Traditional vehicle detection algorithms are feature-classifier combination methods that use manual extraction and selection of features and classify them through a classifier. In traditional vehicle detection algorithms, it is particularly important whether the feature extraction is sufficient. However, there are often errors in manual extraction and selection of features, which affect the results of the classifier and deteriorate the detection effect, making it impossible to meet the detection requirements of actual applications. Nowadays, with the continuous development of convolutional neural networks, convolutional neural networks can automatically extract features. Compared with manual feature extraction, convolutional neural networks have the advantages of small error and fast speed.
[0003] For nighttime scenarios, current vehicle detection algorithms are not ideal, and the main reasons are as follows:
[0004] 1) Due to the poor lighting in the nighttime environment, it is difficult to distinguish the features of various types of vehicles;
[0005] 2) The shooting angles of images are different, and the pixel scale of vehicles in some images is small, making it more difficult for the network to extract their features in the nighttime environment and increasing the detection difficulty;
[0006] 3) In practical applications, the real-time performance of vehicle detection and recognition is very important, and the nighttime environment will cause the detection speed of the model to fail to meet the real-time requirements of actual nighttime traffic scenarios. Summary of the Invention
[0007] The purpose of the present invention is to solve the technical problem that the existing vehicle detection algorithms have low accuracy and efficiency when used at night and cannot be effectively applied to nighttime scenarios, and to propose a method for nighttime vehicle recognition based on improved YOLOv5.
[0008] A method for nighttime vehicle recognition based on improved YOLOv5 includes the following steps:
[0009] Step 1: Collect several images of nighttime vehicles;
[0010] Step 2: Perform preprocessing on the images to obtain a training set, a validation set, and a test set;
[0011] Step 3: Obtain a network model for nighttime vehicle recognition;
[0012] Step 4: Use the obtained network model for nighttime vehicle recognition to perform nighttime vehicle recognition and detection.
[0013] In step three, the network model for night vehicle recognition is the YOLOv5 network model, and its structure is as follows:
[0014] Input layer → Focus layer → First convolution module → First C3 module → Second convolution module → Second C3 module → Third convolution module → Third C3 module → Fourth convolution module → SPP module → Fourth C3 module → Fifth convolution module → First upsampling, and the feature map formed by the first upsampling is fused with the feature map formed by the third C3 module in terms of channels → First C3DRSN module → Sixth convolution module → Second upsampling, and the feature map formed by the second upsampling is fused with the feature map formed by the second C3 module in terms of channels → Second C3DRSN module → Seventh convolution module, and the feature map formed by the seventh convolution module is fused with the feature map formed by the third C3 module and the feature map formed by the convolution module after the first C3DRSN module in terms of channels → Third C3DRSN → Eighth convolution module, and the feature map formed by the eighth convolution module is fused with the feature map formed by the convolution module after the fourth C3 module in terms of channels → Fourth C3DRSN → Third prediction map;
[0015] Among them, the second C3DRSN → First prediction map;
[0016] Among them, the third C3DRSN → Second prediction map;
[0017] The third prediction map is obtained after the fourth C3DRSN passes through the third ordinary convolution; the first prediction map is obtained after the second C3DRSN passes through the first ordinary convolution; the second prediction map is obtained after the third C3DRSN passes through the second ordinary convolution.
[0018] When the YOLOv5 network model is used, it includes the following sub-steps:
[0019] Step 1) Extract vehicle features in the night vehicle image;
[0020] Input the night vehicle image to be detected into the Backbone for feature extraction; the Backbone uses CSPdarknet. Input the night vehicle image into CSPdarknet, and use the Focus module to slice the input feature map to increase the number of channels. Then, through four convolution operations and three C3 module processes, increase the network depth and improve the network's learning ability for useful features in the feature map. Then, input the above processed feature map into the SPP module for multi-scale feature fusion through the maximum pooling method, and finally, after another C3 module process, obtain the night vehicle features;
[0021] Step 2) Fuse the extracted night vehicle features; input the night vehicle features extracted in step 1) into the BiFPN network for feature fusion;
[0022] First, the input night vehicle features are subjected to operations of the fifth convolutional module and upsampling. The resulting feature map is subjected to channel fusion processing with the feature map formed by the third C3 module in step 1). The resulting feature map after processing is subjected to the first C3DRSN module. The resulting feature map after processing is subjected to the sixth convolutional module and upsampling. The resulting feature map after processing is subjected to channel fusion processing with the feature map formed by the second C3 module in step 1). The resulting feature map after processing is subjected to the second C3DRSN module and the seventh convolutional module. After the second C3DRSN module processes, the first prediction map is obtained through the first ordinary convolution. The feature map after the seventh convolutional module processes is subjected to channel fusion with the feature map formed by the third C3 module in step 1) and the feature map formed by the sixth convolutional module. After passing through the third C3DRSN module, the second prediction map is obtained through the second ordinary convolution and the eighth convolutional module respectively. The feature map formed after passing through the eighth convolutional module is subjected to channel fusion processing with the feature map formed by the fifth convolutional module. The resulting feature map after processing is subjected to the fourth C3DRSN module and the third ordinary convolution to obtain the third prediction map.
[0023] The structure of the C3DRSN module is as follows:
[0024] Input layer → Ninth convolutional module → Tenth convolutional module → First soft thresholding module → Second soft thresholding module → Add layer → Output layer
[0025] Among them, input layer → Add layer;
[0026] Among them, tenth convolutional module → Channel attention module → Eleventh convolutional module → First soft thresholding module;
[0027] Among them, first soft thresholding module → Spatial attention module → Twelfth convolutional module → Second soft thresholding module.
[0028] When the C3DRSN module is used, it includes the following sub-steps:
[0029] Step 1) Enhance useful features;
[0030] The feature map is input into two convolutional modules for processing. The resulting feature map is subjected to the channel attention module and the eleventh convolutional module. The resulting feature map is subjected to soft thresholding processing with the feature map formed after the tenth convolutional module processes. The resulting feature map after processing is subjected to the spatial attention module and the twelfth convolutional module. The resulting feature map is subjected to the second soft thresholding module with the feature map formed after the first soft thresholding module processes;
[0031] Step 2) Fuse the enhanced night vehicle features; perform an Add module process on the night vehicle feature map output after the second soft thresholding process in Step 1) and the original input layer features to obtain an enhanced night vehicle feature map.
[0032] Compared with the prior art, the present invention has the following technical effects:
[0033] 1) The present invention uses YOLOv5 to perform the night vehicle recognition task. YOLOv5 belongs to a single-stage network in convolutional neural networks. Compared with traditional machine vision methods, YOLOv5 is superior to traditional methods in terms of detection accuracy and speed. The present invention uses the improved YOLOv5 to introduce the BiFPN network in the Neck layer. Compared with the FPN+PANet of the original YOLOv5, BiFPN speeds up the network detection speed by deleting nodes with only one input, and adds an edge from the input end of the deleted node to the next layer node, so that more features can be fused, realizing cross-scale connection and improving the detection accuracy.
[0034] 2) The present invention proposes an improved C3DRSN; after two convolutional modules, compared with the original Bottleneck, C3DRSN has two more branches. The first branch: uses a channel attention module and a soft thresholding module to learn the channel weights of the image, and reduces useless channel information through the soft thresholding module to enhance useful channel information; the second branch: uses a spatial attention module and a soft thresholding module to learn the position information of the image, and reduces useless position information through the soft thresholding module to enhance useful position information. Since the soft thresholding adaptively sets the threshold, and the input images to C3DRSN are different, the enhancement degrees are different, which can more flexibly enhance useful information, reduce redundant information, and improve the vehicle recognition accuracy. Description of the Drawings
[0035] The following further describes the present invention with reference to the drawings and embodiments:
[0036] Figure 1 is a schematic structural diagram of the C3DRSN module in the present invention;
[0037] Figure 2 is a schematic structural diagram of the improved YOLOv5 network model in the present invention;
[0038] Figure 3 is a flowchart of the present invention. Detailed Embodiments
[0039] As Figure 3 shown, a night vehicle recognition method based on improved YOLOv5 includes the following steps:
[0040] Step 1: Collect several images of vehicles at night;
[0041] Step 2: Perform preprocessing on the images to obtain a training set, a validation set, and a test set;
[0042] Step 3: Obtain a network model for night vehicle recognition;
[0043] Step 4: Use the obtained network model for night vehicle recognition to perform night vehicle recognition and detection.
[0044] In Step 3, the network model for night vehicle recognition is the YOLOv5 network model,
[0045] As Figure 2 shown, the improved YOLOv5 network model structure is:
[0046] Input layer → Focus layer → Convolution module → First C3 module → Convolution module → Second C3 module → Convolution module → Third C3 module → Convolution module → SPP module → Fourth C3 module → Convolution module → First upsampling, the feature map formed by the first upsampling is fused with the feature map formed by the third C3 module → First C3DRSN module → Convolution module → Second upsampling, the feature map formed by the second upsampling is fused with the feature map formed by the second C3 module → Second C3DRSN module, where a branch is separated: After ordinary convolution, a prediction map of 80×80 is obtained → Convolution module, the feature map formed by the convolution module is fused with the feature map formed by the third C3 module and the feature map formed by the convolution module after the first C3DRSN module → Third C3DRSN, where a branch is separated: After ordinary convolution processing, a prediction map of 40×40 is obtained → Convolution module, the feature map formed by the convolution module is fused with the feature map formed by the convolution module after the fourth C3 module → Fourth C3DRSN → Ordinary convolution → 20×20 prediction map;
[0047] When the YOLOv5 network model is used, its working process is:
[0048] Step 3-1) Extract vehicle features from night-time vehicle images; input the night-time vehicle images to be detected into the Backbone for feature extraction. The Backbone uses CSPdarknet. Input the night-time vehicle images into CSPdarknet, use the Focus module to slice the input feature map to increase the number of channels, then perform four convolutional operations and three C3 module processes to increase the network depth and improve the network's learning ability for useful features in the feature map. Then input the processed feature map into the SPP module for multi-scale feature fusion through max pooling, and finally, after another C3 module process, obtain the night-time vehicle features;
[0049] Step 3-2) Fuse the extracted night-time vehicle features; input the night-time vehicle features extracted in 3-1) into the BiFPN network for feature fusion. First, perform the fifth convolutional module operation and upsampling on the input night-time vehicle features. The processed feature map is subjected to channel fusion processing with the feature map formed by the third C3 module in step 1). The processed feature map is then processed by the first C3DRSN module. The formed feature map is subjected to the sixth convolutional module and upsampling processing. The processed feature map is subjected to channel fusion processing with the feature map formed by the second C3 module in step 1). The processed feature map is processed by the second C3DRSN module and the seventh convolutional module. After the second C3DRSN module processes, the first prediction map is obtained through the first ordinary convolution. The feature map processed by the seventh convolutional module is subjected to channel fusion with the feature map formed by the third C3 module in step 1) and the feature map formed by the sixth convolutional module processing. Then, after being processed by the third C3DRSN module, it passes through the second ordinary convolution and the eighth convolutional module respectively. The second prediction map is obtained through the second ordinary convolution. The feature map formed after passing through the eighth convolutional module is subjected to channel fusion processing with the feature map formed by the fifth convolutional module. The processed feature map passes through the fourth C3DRSN module and the third ordinary convolution to obtain the third prediction map.
[0050] As Figure 1 shown, the structure of the C3DRSN module is:
[0051] Input layer → Ninth convolutional module → Tenth convolutional module → First soft thresholding module → Second soft thresholding module → Add layer → Output layer
[0052] Among them, input layer → Add layer;
[0053] Among them, tenth convolutional module → Channel attention module → Eleventh convolutional module → First soft thresholding module;
[0054] Among them, first soft thresholding module → Spatial attention module → Twelfth convolutional module → Second soft thresholding module.
[0055] When the C3DRSN module is in use, its working process is as follows:
[0056] Step 5-1) Enhance useful features; input the feature map into two convolutional modules for processing. The processed feature map is then processed by a channel attention module and an eleventh convolutional module. The resulting feature map is subjected to soft thresholding processing with the feature map formed after processing by the tenth convolutional module. The processed feature map is then processed by a spatial attention module and a twelfth convolutional module. The resulting feature map and the feature map formed after processing by the first soft thresholding module are processed by a second soft thresholding module;
[0057] Step 5-2) Fuse the enhanced night vehicle features; perform Add module processing on the night vehicle feature map output after the second soft thresholding in 5-1) and the original input layer features to obtain an enhanced night vehicle feature map.
[0058] Example:
[0059] The night vehicle recognition method based on the improved YOLOv5 includes the following steps:
[0060] (1) Collect the Berkeley Deep Drive (BDD) 100K dataset. A total of 10,703 night vehicle images are obtained, among which 8,701 night vehicle images are used as the training set and validation set of the network, and 2,002 night vehicle images are used as the test set;
[0061] (2) Name the night vehicle images in the BDD dataset according to the format of the Pascal VOC dataset. At the same time, divide the images and annotation files into a training set, a validation set, and a test set, with the training set, validation set, and test set accounting for 60%, 20%, and 20% respectively. Import the dataset into the server and import the dataset path into the model;
[0062] (3) Set the improved YOLOv5 network model to the YOLOv5s version, and set training parameters such as the size of the input image of the network, the number of types of recognition tasks, and the number of iteration times according to the size of the computer's memory and video memory, etc., and start model training;
[0063] (4) In this paper, the training image size is set to 640*640. The training set is input into the network, and the network automatically modifies the image size to 640*640. The 640*640 training images are input into YOLOv5. The Yolov5 network includes: Focus module: Input a three-channel color picture of 3*640*640. After downsampling, it outputs a feature map of 32*320*320; Conv module: Performs convolution, SiLU activation function, and normalization calculation on the input feature map, and the output feature map size is half of the input; C3 module: Divides the input feature map into two parts for operation. One part passes through Conv convolution and multiple Bottleneck modules, and the other part passes through Conv convolution. Finally, the two parts are concatenated, and the resulting output is further subjected to Conv convolution operation to ensure that the output and input features remain unchanged; SPP module: Spatial pyramid pooling layer. The number of input channels becomes half after passing through the standard convolution module. Max pooling operations with kernel sizes of 5, 7, and 9 are performed on it, and Concat feature fusion operation is performed on the pooling results;
[0064] (5) The 3*640*640 images are input into the Backbone network. Three kinds of night vehicle feature maps of 80*80*256, 40*40*512, and 20*20*1024 are respectively output to the Neck network at the second C3 module, the third C3 module, and the fourth C3 module;
[0065] (6) In the Neck network, in BiFPN, a node from upsampling to Concat is deleted, and a C3 and Conv to Concat node is added to achieve bidirectional cross-scale connection to enhance multi-scale feature fusion of night vehicles and reduce the network training time. At the same time, the C3DRSN module is introduced to enhance vehicle information features and improve accuracy. The three branches output by the Backbone network are input into BiFPN. After each branch passes through the Conv convolution module, the upsampling module, and the C3DRSN module, three kinds of night vehicle feature maps of 80*80*256, 40*40*512, and 20*20*1024 are respectively output;
[0066] (7) The three enhanced branches obtained through multi-scale feature fusion of the Neck network are input into the standard convolution and then sent to the Head network to obtain the detection results. The training model is obtained through the above operations;
[0067] (8) After obtaining a network model with good results, it is deployed to the mobile terminal, and the video is obtained through the camera and input into the mobile terminal. The mobile terminal performs real-time vehicle detection on the vehicles in the obtained video, identifies the types, quantities, and density of the vehicles in the video, and takes corresponding measures quickly for the crowded sections and accident-occurring sections of the road according to these data.
Claims
1. A method for nighttime vehicle recognition based on improved YOLOv5, characterized in that, It includes the following steps: Step 1: Collect several images of vehicles at night; Step 2: Perform preprocessing on the images to obtain a training set, a validation set, and a test set; Step 3: Obtain a network model for vehicle recognition at night; Step 4: Use the obtained network model for vehicle recognition at night to perform vehicle recognition and detection at night; In Step 3, the network model for vehicle recognition at night is the YOLOv5 network model, and its structure is: Input layer → Focus layer → First convolution module → First C3 module → Second convolution module → Second C3 module → Third convolution module → Third C3 module → Fourth convolution module → SPP module → Fourth C3 module → Fifth convolution module → First upsampling, the feature map formed by the first upsampling is fused with the feature map formed by the third C3 module in channels → First C3DRSN module → Sixth convolution module → Second upsampling, the feature map formed by the second upsampling is fused with the feature map formed by the second C3 module in channels → Second C3DRSN module → Seventh convolution module, the feature map formed by the seventh convolution module is fused with the feature map formed by the third C3 module and the feature map formed by the sixth convolution module in channels → Third C3DRSN module → Eighth convolution module, the feature map formed by the eighth convolution module is fused with the feature map formed by the fifth convolution module in channels → Fourth C3DRSN module → Third prediction map; among them, the second C3DRSN module → First prediction map; among them, the third C3DRSN module → Second prediction map; The structure of the C3DRSN module is: Input layer → Ninth convolution module → Tenth convolution module → First soft thresholding module → Second soft thresholding module → Add layer → Output layer; Among them, input layer → Add layer; Among them, tenth convolution module → Channel attention module → Eleventh convolution module → First soft thresholding module; Among them, first soft thresholding module → Spatial attention module → Twelfth convolution module → Second soft thresholding module.
2. The method according to claim 1, characterized in that, The fourth C3DRSN module obtains the third prediction map after passing through the third ordinary convolution; the second C3DRSN module obtains the first prediction map after passing through the first ordinary convolution; the third C3DRSN module obtains the second prediction map after passing through the second ordinary convolution.
3. The method according to claim 1, characterized in that, When the YOLOv5 network model is used, it includes the following sub-steps: Step 1) Extract vehicle features in the vehicle image at night; Input the vehicle image at night to be detected into the Backbone for feature extraction; the Backbone uses CSPdarknet. Input the vehicle image at night into CSPdarknet, use the Focus module to slice the input feature map to increase the number of channels, and then through four convolution operations and three C3 module processes to increase the network depth and improve the network's learning ability for useful features in the feature map. Then input the processed feature map into the SPP module for multi-scale feature fusion through the maximum pooling method, and finally pass through one more C3 module process to obtain the vehicle features at night; Step 2) Fuse the extracted night vehicle features; input the night vehicle features extracted in Step 1) into the BiFPN network for feature fusion; First, perform the fifth convolution module operation and the first upsampling on the input night vehicle features. The feature map obtained after processing is subjected to channel fusion processing with the feature map formed by the third C3 module in Step 1). The obtained feature map is processed by the first C3DRSN module. The formed feature map is processed by the sixth convolution module and the second upsampling. The obtained feature map is subjected to channel fusion processing with the feature map formed by the second C3 module in Step 1). The obtained feature map is processed by the second C3DRSN module and the seventh convolution module. After the second C3DRSN module processes, the first prediction map is obtained through the first ordinary convolution. The feature map processed by the seventh convolution module is subjected to channel fusion with the feature map formed by the third C3 module in Step 1) and the feature map formed by the sixth convolution module processing, and then processed by the third C3DRSN module, and respectively passed through the second ordinary convolution and the eighth convolution module. The second prediction map is obtained through the second ordinary convolution. The feature map formed after passing through the eighth convolution module is subjected to channel fusion processing with the feature map formed by the fifth convolution module. The obtained feature map is processed by the fourth C3DRSN module and the third ordinary convolution to obtain the third prediction map.
4. The method according to claim 1, characterized in that The feature map formed by the input layer is input to the Add layer; the feature map formed by the tenth convolution module is input to the first soft thresholding module after passing through the channel attention module and the eleventh convolution module; The feature map formed by the first soft thresholding module is input to the second soft thresholding module after passing through the spatial attention module and the twelfth convolution module.
5. The method according to claim 1, wherein When the C3DRSN module is used, it includes the following sub-steps: Step 1) Enhance useful features; Input the feature map into two convolution modules for processing. The processed feature map is processed by the channel attention module and the eleventh convolution module. The obtained feature map is subjected to soft thresholding processing with the feature map formed by the tenth convolution module processing. The obtained feature map is processed by the spatial attention module and the twelfth convolution module. The obtained feature map is processed by the second soft thresholding module with the feature map formed by the first soft thresholding module processing; Step 2) Fuse the enhanced night vehicle features; perform Add module processing on the night vehicle feature map output after the second soft thresholding processing in Step 1) and the original input layer features to obtain the enhanced night vehicle feature map.
Citation Information
Patent Citations
YOLOv5 neural network vehicle detection method added with attention mechanism
CN114092764A