An Infrared Vehicle Detection Method Based on Improved YOLOv7 Algorithm

The improved YOLOv7 algorithm with a new backbone network and multi-scale feature fusion addresses the accuracy and speed challenges in infrared vehicle detection, achieving a 2.47% increase in accuracy with comparable speed.

CN116129327BActive Publication Date: 2025-07-15XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310175297.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-07-15
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

The existing infrared vehicle detection methods have shortcomings in detection speed and accuracy, especially when infrared image feature extraction is difficult, it is difficult to meet the actual needs of real-time and accuracy.

Method used

Improve the backbone feature extraction network of the YOLOv7 algorithm, build a new backbone feature extraction network Conv31, which contains 31 convolutional blocks, and is connected to the YOLOv7 prediction network to form a Conv31-YOLOv7 model, and improve detection accuracy by increasing network depth and multi-scale feature fusion.

Benefits of technology

On the premise of ensuring a higher detection speed, the accuracy of infrared vehicle detection is significantly improved, the network's feature extraction capability and detection capability of vehicle targets at different scales are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129327B_ABST
    Figure CN116129327B_ABST
Patent Text Reader

Abstract

The present invention discloses an infrared vehicle detection method based on an improved YOLOv7 algorithm, which includes the following steps; Step 1: Collect vehicle videos on traffic roads for frame extraction and image preprocessing to obtain an infrared vehicle image dataset; Step 2: Construct a new backbone feature extraction network Conv31 containing 31 convolutional blocks; Step 3: Connect the new backbone feature extraction network with the prediction network of the original YOLOv7 to form a new network model Conv31-YOLOv7; Step 4: Feed the training dataset obtained in Step 1 into the network model Conv31-YOLOv7 of the new Step 3, and use the mini-batch stochastic gradient descent algorithm for training to obtain a trained infrared vehicle detection model; Step 5: Feed the infrared vehicle videos on traffic roads collected in real time by the infrared thermal imaging device into the trained infrared vehicle detection model frame by frame to obtain the real-time position information, scale information and confidence of the vehicle. The present invention significantly improves the detection accuracy on the premise of ensuring a high detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle detection, and particularly relates to an infrared vehicle detection method based on an improved YOLOv7 algorithm. Background Art

[0002] Infrared target detection technology refers to automatically extracting the position information of targets from infrared images. Given the advantages of infrared thermal imaging, infrared target detection technology can be applied to vehicle detection scenarios on traffic roads and can adapt to situations such as night, strong light, and extreme weather. Therefore, the breakthrough of this technology has important theoretical significance and practical value for fields such as autonomous driving and intelligent transportation.

[0003] Traditional infrared vehicle detection methods usually first extract the features of targets using methods such as histogram of oriented gradients, and then use positive and negative samples to train classifiers such as support vector machines to classify the target features. This method has a slow detection speed, cannot meet the requirements of timeliness, and has problems such as limited application scenarios, poor robustness, and weak generalization ability.

[0004] In recent years, with the rapid development of artificial intelligence technology, infrared vehicle detection methods based on convolutional neural networks have been widely applied. It can automatically perform feature abstraction and feature extraction on images through convolutional neural networks, and has high detection accuracy and strong robustness.

[0005] Currently, object detection algorithms based on deep learning mainly include two categories. One is the two-stage detection algorithm, which divides the detection process into two stages. In the first stage, candidate regions of the image to be detected are generated, and in the second stage, classification and regression are performed on the generated candidate regions to obtain the final detection results. The first stage of this type of algorithm is time-consuming, and generally has high detection accuracy but slow detection speed, and generally cannot meet the real-time requirements. Representative algorithms include R-CNN, Fast R-CNN, Faster R-CNN, etc. The other is the single-stage detection algorithm, which unifies the above two-stage detection process into an end-to-end regression process, combining the two steps of region selection and detection judgment into one. It has lower detection accuracy but faster detection speed. Representative algorithms include YOLO and SSD, etc.

[0006] Object detection algorithms based on deep learning have good detection effects in visible light image object detection scenarios. However, in infrared object detection scenarios, due to infrared images being single-channel images with unclear features, it is difficult to extract the features of infrared vehicle targets, resulting in generally low detection accuracy of current mainstream object detection algorithms and being difficult to meet actual requirements. Summary of the Invention

[0007] In order to overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide an infrared vehicle detection method based on an improved YOLOv7 algorithm, which can significantly improve the detection accuracy on the premise of ensuring a high detection speed.

[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0009] An infrared vehicle detection method based on an improved YOLOv7 algorithm, comprising the following steps;

[0010] Step 1: Collect vehicle videos on the traffic road for frame extraction and image preprocessing to obtain an infrared vehicle image dataset;

[0011] Step 2: Improve the backbone feature extraction network of the YOLOv7 algorithm, that is, discard the backbone feature extraction network in the YOLOv7 algorithm, construct a new backbone feature extraction network Conv31 containing 31 convolutional blocks, and replace the backbone feature extraction network in the YOLOv7 algorithm;

[0012] Step 3: Connect the new backbone feature extraction network with the prediction network of the original YOLOv7 to form a new network model Conv31-YOLOv7;

[0013] Step 4: Feed the training dataset obtained in Step 1 into the network model Conv31-YOLOv7 in Step 3, and use the mini-batch stochastic gradient descent algorithm for training to obtain a trained infrared vehicle detection model;

[0014] Step 5: Feed the infrared vehicle videos on the traffic road collected in real time by the infrared thermal imaging device into the trained infrared vehicle detection model frame by frame to obtain the real-time position information, scale information and confidence of the vehicle.

[0015] The steps of frame extraction and image preprocessing in Step 1 are specifically as follows:

[0016] (1.1) Collect infrared vehicle videos on the traffic road, read the first 10,000 frames of the video, set the resolution of the image to be output as 640×640, output each frame in sequence in image format to obtain 10,000 infrared vehicle images, and mark the position information of the vehicle targets in the obtained infrared vehicle images to make an infrared vehicle image dataset. This dataset has 10,000 infrared vehicle images with a resolution of 640×640;

[0017] (1.2) Divide the infrared vehicle image dataset into a training dataset and a test dataset according to a ratio of 9:1, that is, randomly select 9,000 infrared images from the dataset to form a training set, and the remaining 1,000 infrared images form a test set.

[0018] Step 2 is specifically as follows:

[0019] (2.1) Discard the backbone feature extraction network in the YOLOv7 algorithm and construct a new backbone feature extraction network Conv31 to replace the backbone feature extraction network in the YOLOv7 algorithm. The new backbone feature extraction network Conv31 contains 31 convolutional blocks, and the structure of each convolutional block is as follows:

[0020] The first convolutional block: contains a convolutional layer with an input channel number of 1, an output channel number of 32, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0021] The second convolutional block: contains a convolutional layer with an input channel number of 32, an output channel number of 16, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 16, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0022] The third convolutional block: contains a convolutional layer with an input channel number of 16, an output channel number of 32, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0023] The fourth convolutional block: contains a convolutional layer with an input channel number of 32, an output channel number of 32, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0024] The fifth convolutional block and the seventh convolutional block: contain a convolutional layer with an input channel number of 32, an output channel number of 64, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0025] The sixth convolutional block: contains a convolutional layer with an input channel number of 64, an output channel number of 32, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0026] The 8th convolutional block: It contains a convolutional layer with an input channel number of 64, an output channel number of 64, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0027] The 9th and 11th convolutional blocks: They contain a convolutional layer with an input channel number of 64, an output channel number of 128, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0028] The 10th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 64, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0029] The 12th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 128, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0030] The 13th and 15th convolutional blocks: They contain a convolutional layer with an input channel number of 128, an output channel number of 256, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0031] The 14th convolutional block: It contains a convolutional layer with an input channel number of 256, an output channel number of 128, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0032] The 16th, 19th, 21st, and 23rd convolutional blocks: They contain a convolutional layer with an input channel number of 256, an output channel number of 512, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a negative slope of 0.1;

[0033] The 17th convolutional block, the 20th convolutional block, and the 22nd convolutional block: contain a convolutional layer with an input channel number of 512, an output channel number of 256, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0034] The 18th convolutional block: contains a convolutional layer with an input channel number of 256, an output channel number of 256, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0035] The 24th convolutional block, the 27th convolutional block, the 29th convolutional block, and the 31st convolutional block: contain a convolutional layer with an input channel number of 512, an output channel number of 1024, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0036] The 25th convolutional block, the 28th convolutional block, and the 30th convolutional block: all contain a convolutional layer with an input channel number of 1024, an output channel number of 512, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0037] The 26th convolutional block: contains a convolutional layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0038] (2.2) Connect the above 31 convolutional blocks in sequence to obtain a new backbone feature extraction network Conv31 with the following structure:

[0039] The 1st convolutional block -> the 2nd convolutional block -> the 3rd convolutional block -> the 4th convolutional block -> the 5th convolutional block -> the 6th convolutional block -> the 7th convolutional block -> the 8th convolutional block -> the 9th convolutional block -> the 10th convolutional block -> the 11th convolutional block -> the 12th convolutional block -> the 13th convolutional block -> the 14th convolutional block -> the 15th convolutional block -> the 16th convolutional block -> the 17th convolutional block -> the 18th convolutional block -> the 19th convolutional block -> the 20th convolutional block -> the 21st convolutional block -> the 22nd convolutional block -> the 23rd convolutional block -> the 24th convolutional block -> the 25th convolutional block -> the 26th convolutional block -> the 27th convolutional block -> the 28th convolutional block -> the 29th convolutional block -> the 30th convolutional block -> the 31st convolutional block;

[0040] (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with Conv31.

[0041] Specifically, the step 3 is as follows:

[0042] Connect the 16th convolutional block in the new backbone feature extraction network Conv31 obtained in step 2 to the 1st prediction branch of the YOLOv7 prediction network;

[0043] Connect the 24th convolutional block in the new backbone feature extraction network Conv31 to the 2nd prediction branch of the YOLOv7 prediction network;

[0044] Connect the 31st convolutional block in the new backbone feature extraction network Conv31 to the 3rd prediction branch of the YOLOv7 prediction network.

[0045] The internal module connection relationship of the YOLOv7 prediction network is as follows:

[0046] The connection relationship of each module in the 1st prediction branch is as follows:

[0047] The 16th convolutional block -> the branch convolutional block 1 -> Multi_Concat_Block1 -> RepConv1 -> the detection head 1;

[0048] The connection relationship of each module in the 2nd prediction branch is as follows:

[0049] The 24th convolutional block -> the branch convolutional block 2 -> Multi_Concat_Block2 -> Multi_Concat_Block3 -> RepConv2 -> the detection head 2;

[0050] The connection relationship of each module in the 3rd prediction branch is as follows:

[0051] The 31st convolutional block -> Multi_Concat_Block4 -> RepConv3 -> Detection head 3;

[0052] The connection relationships between the respective prediction branch modules are as follows:

[0053] The 31st convolutional block -> Upsampling convolutional block 1 -> Upsampling layer 1 -> Multi_Concat_Block2 -> Upsampling convolutional block 2 -> Upsampling layer 2 -> Multi_Concat_Block1;

[0054] Multi_Concat_Block1 -> TransitionBlock1 -> Multi_Concat_Block2 -> TransitionBlock2 -> Multi_Concat_Block4.

[0055] Specifically, step 4 is as follows:

[0056] (4.1) Set the training parameters: the number of training epochs is 200, the number of infrared vehicle images selected for one training is set to 16, the learning rate is set to 0.001, and both the confidence threshold and the IOU ignore threshold are set to 0.5;

[0057] (4.2) Input 9000 infrared vehicle images in the training set into the model Conv31 - YOLOv7 in batches of 16 each time, and each time the output obtains the offset values (t x , t y , t w , t h ) of the target bounding box relative to the labeled box and the target confidence p, where t x is the offset value of the target bounding box relative to the labeled box in the x - direction, t y is the offset value of the target bounding box relative to the labeled box in the y - direction, t w is the offset value of the target bounding box relative to the width of the labeled box, and t h is the offset value of the target bounding box relative to the height of the labeled box;

[0058] (4.3) Calculate the position and width - height of the predicted box from the offset values (t x , t y , t w , t h ) through the following coordinate offset formula:

[0059]

[0060]

[0061]

[0062]

[0063] Among them, b x , b y is the position of the prediction box, c x , c y is the position of the annotation box, b w , b h is the width and height of the prediction box, p w , p h is the width and height of the annotation box;

[0064] (4.4) Substitute the position, width, height, and confidence of the target of the prediction box (b x , b y , b w , b h , p) and the position, width, height, and confidence of the annotation box into the loss function to calculate the loss value, and use the mini-batch stochastic gradient descent algorithm to update its weights;

[0065] (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain the trained infrared vehicle detection model.

[0066] The specific steps of step 5 are as follows:

[0067] Use an infrared thermal imaging device to collect infrared vehicle videos on the traffic road in real time, and feed them into the trained infrared vehicle detection model frame by frame to obtain the real-time position information and confidence of the vehicles.

[0068] Advantages of the present invention:

[0069] In the present invention, since the backbone feature extraction network of the YOLOv7 algorithm is replaced with a new backbone feature extraction network Conv31, Conv31 contains 31 convolutional blocks, each convolutional block contains a convolutional layer, a total of 31 convolutional layers. Increasing the network depth by increasing the number of convolutional layers can effectively improve the detection accuracy to a certain extent. This method can deepen the network depth and strengthen the feature extraction ability by stacking convolutional blocks to greatly increase the number of convolutional layers, effectively improving the detection accuracy; the 16th convolutional block, the 24th convolutional block, and the 31st convolutional block of Conv31 can extract the shallow layer features, middle layer features, and deep layer features of the infrared vehicle target respectively. By connecting these 3 convolutional blocks to 3 prediction branches, the fusion of multi-scale features can be realized, improving the ability of the network model to detect infrared vehicle targets of different scales, thereby further improving the detection accuracy. The test results show that compared with other vehicle detection methods based on convolutional neural networks, the present invention can significantly improve the detection accuracy on the premise of ensuring a high detection speed. Description of the Drawings

[0070] Figure 1 is the implementation flowchart of the present invention.

[0071] Figure 2 is the structural diagram of the Conv31-YOLOv7 network constructed in the present invention.

[0072] Figure 3 is the detection schematic diagram of the present invention in the actual scenario. Detailed Embodiment

[0073] The present invention will be further described in detail below with reference to the accompanying drawings.

[0074] As Figure 1 shown:

[0075] Step 1: Construct an infrared vehicle dataset.

[0076] (1.1) Collect infrared vehicle videos on traffic roads, read the first 10,000 frames of the video, set the resolution of the image to be output as 640×640, output each frame in sequence in image format to obtain 10,000 infrared vehicle images, and mark the position information of the vehicle targets in the obtained infrared vehicle images to make an infrared vehicle image dataset. This dataset has a total of 10,000 infrared images with a resolution of 640×640.

[0077] (1.2) Divide the infrared vehicle image dataset into a training dataset and a test dataset according to a ratio of 9:1, that is, randomly select 9,000 infrared images from the dataset to form a training set, and the remaining 1,000 infrared images form a test set.

[0078] Step 2: Construct a new backbone feature extraction network.

[0079] In this step, a new backbone feature extraction network is constructed based on the improvement of the backbone feature extraction network of the existing YOLOv7 algorithm. The network model in the YOLOv7 algorithm includes a backbone feature extraction network and a prediction network. In this step, only its backbone feature extraction network is improved, and the specific implementation is as follows:

[0080] (2.1) Discard the backbone feature extraction network in the YOLOv7 algorithm and construct a new backbone feature extraction network Conv31 to replace the backbone feature extraction network in the YOLOv7 algorithm. Construct a new backbone feature extraction network Conv31, which contains 31 convolutional blocks, and the structure of each convolutional block is as follows:

[0081] The first convolutional block: It contains a convolutional layer with an input channel number of 1, an output channel number of 32, a convolutional kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0082] The second convolutional block: It contains a convolutional layer with an input channel number of 32, an output channel number of 16, a convolutional kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 16, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0083] The third convolutional block: It contains a convolutional layer with an input channel number of 16, an output channel number of 32, a convolutional kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0084] The fourth convolutional block: It contains a convolutional layer with an input channel number of 32, an output channel number of 32, a convolutional kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0085] The fifth convolutional block and the seventh convolutional block: It contains a convolutional layer with an input channel number of 32, an output channel number of 64, a convolutional kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0086] The sixth convolutional block: It contains a convolutional layer with an input channel number of 64, an output channel number of 32, a convolutional kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0087] The eighth convolutional block: It contains a convolutional layer with an input channel number of 64, an output channel number of 64, a convolutional kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0088] The ninth convolutional block and the eleventh convolutional block: It contains a convolutional layer with an input channel number of 64, an output channel number of 128, a convolutional kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0089] The 10th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 64, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0090] The 12th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 128, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0091] The 13th and 15th convolutional blocks: It contains a convolutional layer with an input channel number of 128, an output channel number of 256, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0092] The 14th convolutional block: It contains a convolutional layer with an input channel number of 256, an output channel number of 128, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0093] The 16th, 19th, 21st, and 23rd convolutional blocks: It contains a convolutional layer with an input channel number of 256, an output channel number of 512, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 512 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0094] The 17th, 20th, and 22nd convolutional blocks: It contains a convolutional layer with an input channel number of 512, an output channel number of 256, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0095] The 18th convolutional block: It contains a convolutional layer with an input channel number of 256, an output channel number of 256, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0096] The 24th convolutional block, the 27th convolutional block, the 29th convolutional block, and the 31st convolutional block: Each contains a convolutional layer with an input channel number of 512, an output channel number of 1024, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0097] The 25th convolutional block, the 28th convolutional block, and the 30th convolutional block: Each contains a convolutional layer with an input channel number of 1024, an output channel number of 512, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0098] The 26th convolutional block: contains a convolutional layer with an input channel number of 512, an output channel number of 512, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a negative slope of 0.1.

[0099] (2.2) Connect the above 31 convolutional blocks in sequence to obtain a new backbone feature extraction network Conv31 with the following structure:

[0100] The 1st convolutional block -> The 2nd convolutional block -> The 3rd convolutional block -> The 4th convolutional block -> The 5th convolutional block -> The 6th convolutional block -> The 7th convolutional block -> The 8th convolutional block -> The 9th convolutional block -> The 10th convolutional block -> The 11th convolutional block -> The 12th convolutional block -> The 13th convolutional block -> The 14th convolutional block -> The 15th convolutional block -> The 16th convolutional block -> The 17th convolutional block -> The 18th convolutional block -> The 19th convolutional block -> The 20th convolutional block -> The 21st convolutional block -> The 22nd convolutional block -> The 23rd convolutional block -> The 24th convolutional block -> The 25th convolutional block -> The 26th convolutional block -> The 27th convolutional block -> The 28th convolutional block -> The 29th convolutional block -> The 30th convolutional block -> The 31st convolutional block.

[0101] (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with Conv31.

[0102] Step 3: Construct a new network model Conv31 - YOLOv7.

[0103] Refer to Figure 2 , connect the new backbone feature extraction network and the prediction network of YOLOv7 according to the following structural relationship to form a new network model Conv31 - YOLOv7:

[0104] Connect the 16th convolutional block in the new backbone feature extraction network Conv31 to the first prediction branch of the YOLOv7 prediction network.

[0105] Connect the 24th convolutional block in the new backbone feature extraction network Conv31 to the second prediction branch of the YOLOv7 prediction network.

[0106] Connect the 31st convolutional block in the new backbone feature extraction network Conv31 to the third prediction branch of the YOLOv7 prediction network.

[0107] Step 4: Train the new network model Conv31 - YOLOv7.

[0108] (4.1) Set the training parameters: the number of training epochs is 200, the number of infrared vehicle images selected for one training is set to 16, the learning rate is set to 0.001, and both the confidence threshold and the IOU ignore threshold are set to 0.5.

[0109] (4.2) Input 9000 infrared vehicle images in the training set into the model Conv31 - YOLOv7 in batches of 16 each time, and each time the output obtains the offset values (t x , t y , t w , t h ) of the target bounding box relative to the labeled box and the target confidence p, where t x is the offset value of the target bounding box relative to the labeled box in the x - direction, t y is the offset value of the target bounding box relative to the labeled box in the y - direction, t w is the offset value of the target bounding box relative to the width of the labeled box, and t h is the offset value of the target bounding box relative to the height of the labeled box.

[0110] (4.3) Calculate the position and width - height of the predicted box from the offset values (t x , t y , t w , t h ) through the following coordinate offset formula:

[0111]

[0112]

[0113]

[0114]

[0115] where b x , b y are the positions of the predicted box, and cx , c y is the position of the annotation box, b w , b h are the width and height of the prediction box, p w , p h are the width and height of the annotation box.

[0116] (4.4) Substitute the position, width and height of the prediction box, and the confidence of the target (b x , b y , b w , b h , p) and the position, width and height of the annotation box, and the confidence of the target into the loss function to calculate the loss value, and use the mini-batch stochastic gradient descent algorithm to update its weights.

[0117] (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain the trained infrared vehicle detection model.

[0118] Step 5: Use the trained model for infrared vehicle detection.

[0119] Use an infrared thermal imaging device to collect infrared vehicle videos on the traffic road in real time, and feed them into the trained infrared vehicle detection model frame by frame to obtain the real-time position information and confidence of the vehicles.

[0120] The effects of the present invention are further illustrated by the following simulation experiments and measured data:

[0121] I. Simulation and measurement environment

[0122] For the simulation and measurement of the present invention, the Windows 10 operating system is used, an NVIDIA GeForce GTX2060 GPU is used for acceleration, and the deep learning framework used is pytorch 1.8.1.

[0123] II. Simulation content

[0124] Simulation 1: Use the same training set and parameters as the present invention to train other object detection models based on convolutional neural networks to obtain their respective trained infrared vehicle detection models.

[0125] Send 1000 test set images into the trained model of the present invention one by one for testing to obtain the accuracy and speed of the infrared vehicle detection of the present invention with an IOU threshold of 0.5.

[0126] Use the same test set as the present invention to test the accuracy and speed of the infrared vehicle detection of other methods with an IOU threshold of 0.5.

[0127] The simulation experiment comparison between the method of the present invention and the infrared vehicle detection method based on YOLOv7 is as shown in Table 1 below:

[0128] Table 1

[0129]

[0130] Comparing the present invention with YOLOv7, it is found that the method of the present invention can detect 31 images per second, which is slightly lower than the detection speed of 33 images per second of YOLOv7. The accuracy rate of the method of the present invention for infrared vehicle detection with an IOU threshold of 0.5 is 94.36%, and the accuracy rate of YOLOv7 for infrared vehicle target detection is 91.89%. The method of the present invention has a significant improvement in detection accuracy compared to YOLOv7, ensuring a relatively high detection accuracy.

[0131] III. Measured Content

[0132] Use an infrared thermal imaging device to collect infrared vehicle videos on the traffic road in real time, and feed them into the infrared vehicle detection model that has been trained by the present invention frame by frame to obtain the real-time position information and confidence level of the vehicles, as Figure 3 shown.

[0133] Figure 3 The large rectangular box represents the prediction box that encloses the vehicle in the infrared image, and the small rectangular box above the large rectangular box shows the confidence level of the vehicle target.

[0134] The specific embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An infrared vehicle detection method based on an improved YOLOv7 algorithm, characterized in that It includes the following steps; Step 1: Collect vehicle videos on traffic roads for frame extraction and image preprocessing to obtain an infrared vehicle image dataset; Step 2: Improve the backbone feature extraction network of the YOLOv7 algorithm, that is, discard the backbone feature extraction network in the YOLOv7 algorithm, construct a new backbone feature extraction network Conv31 containing 31 convolutional blocks, and replace the backbone feature extraction network in the YOLOv7 algorithm; Step 3: Connect the new backbone feature extraction network with the prediction network of the original YOLOv7 to form a new network model Conv31-YOLOv7; Step 4: Feed the training dataset obtained in Step 1 into the network model Conv31-YOLOv7 in Step 3, and use the mini-batch stochastic gradient descent algorithm for training to obtain a trained infrared vehicle detection model; Step 5: Feed the infrared vehicle videos on the traffic road collected in real time by the infrared thermal imaging device into the trained infrared vehicle detection model frame by frame to obtain the real-time position information, scale information and confidence of the vehicle; The specific content of Step 3 is as follows: Connect the 16th convolutional block in the new backbone feature extraction network Conv31 obtained in Step 2 with the first prediction branch of the YOLOv7 prediction network; Connect the 24th convolutional block in the new backbone feature extraction network Conv31 with the second prediction branch of the YOLOv7 prediction network; Connect the 31st convolutional block in the new backbone feature extraction network Conv31 with the third prediction branch of the YOLOv7 prediction network; The internal module connection relationship of the YOLOv7 prediction network is as follows: The connection relationship of each module in the first prediction branch is as follows: The 16th convolutional block -> branch convolutional block 1 -> Multi_Concat_Block1 -> RepConv1 -> detection head 1; The connection relationship of each module in the second prediction branch is as follows: The 24th convolutional block -> branch convolutional block 2 -> Multi_Concat_Block2 -> Multi_Concat_Block3 -> RepConv2 -> detection head 2; The connection relationship of each module in the third prediction branch is as follows: The 31st convolutional block -> Multi_Concat_Block4 -> RepConv3 -> detection head 3; The connection relationship between the modules of each prediction branch is as follows: The 31st convolutional block -> upsampling convolutional block 1 -> upsampling layer 1 -> Multi_Concat_Block2 -> upsampling convolutional block 2 -> upsampling layer 2 -> Multi_Concat_Block1; Multi_Concat_Block1 -> TransitionBlock1 -> Multi_Concat_Block2 -> TransitionBlock2 -> Multi_Concat_Block4.

2. The infrared vehicle detection method based on the improved YOLOv7 algorithm according to claim 1, wherein, The specific steps of frame extraction and image preprocessing in Step 1 are as follows: (1.1)Collect the infrared vehicle videos on the traffic road, read the first 10,000 frames of the videos, set the resolution of the image to be output as 640×640, output each frame in sequence in image format to obtain 10,000 infrared vehicle images, and label the position information of the vehicle targets in the obtained infrared vehicle images to make an infrared vehicle image dataset. This dataset has a total of 10,000 infrared vehicle images with a resolution of 640×640; (1.2)Divide the infrared vehicle image dataset into a training dataset and a test dataset according to a ratio of 9:1, that is, randomly select 9,000 infrared images from the dataset to form a training set, and the remaining 1,000 infrared images form a test set.

3. An infrared vehicle detection method based on the improved YOLOv7 algorithm according to claim 1, characterized in that, (2.1)Discard the backbone feature extraction network in the YOLOv7 algorithm and construct a new backbone feature extraction network Conv31 to replace the backbone feature extraction network in the YOLOv7 algorithm. Construct the new backbone feature extraction network Conv31, which contains 31 convolutional blocks. The structure of each convolutional block is as follows: (2.1)Discard the backbone feature extraction network in the YOLOv7 algorithm and construct a new backbone feature extraction network Conv31 to replace the backbone feature extraction network in the YOLOv7 algorithm. Construct the new backbone feature extraction network Conv31, which contains 31 convolutional blocks. The structure of each convolutional block is as follows: The first convolutional block: contains a convolutional layer with an input channel number of 1, an output channel number of 32, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1; The second convolutional block: contains a convolutional layer with an input channel number of 32, an output channel number of 16, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 16, and a LeakyReLU activation function layer with a complex slope of 0.1; The third convolutional block: contains a convolutional layer with an input channel number of 16, an output channel number of 32, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1; The fourth convolutional block: contains a convolutional layer with an input channel number of 32, an output channel number of 32, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1; The fifth and seventh convolutional blocks: contain a convolutional layer with an input channel number of 32, an output channel number of 64, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a complex slope of 0.1; The sixth convolutional block: contains a convolutional layer with an input channel number of 64, an output channel number of 32, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1; The 8th convolutional block: It contains a convolutional layer with an input channel number of 64, an output channel number of 64, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 9th and 11th convolutional blocks: They contain a convolutional layer with an input channel number of 64, an output channel number of 128, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 10th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 64, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 12th convolutional block: It contains a convolutional layer with an input channel number of 128, an output channel number of 128, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 13th and 15th convolutional blocks: They contain a convolutional layer with an input channel number of 128, an output channel number of 256, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 14th convolutional block: It contains a convolutional layer with an input channel number of 256, an output channel number of 128, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 16th, 19th, 21st, and 23rd convolutional blocks: They contain a convolutional layer with an input channel number of 256, an output channel number of 512, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 512 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 17th, 20th, and 22nd convolutional blocks: They contain a convolutional layer with an input channel number of 512, an output channel number of 256, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 18th convolutional block: It contains a convolutional layer with an input channel number of 256, an output channel number of 256, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a negative slope of 0.1; The 24th convolutional block, the 27th convolutional block, the 29th convolutional block, and the 31st convolutional block: Each contains a convolutional layer with an input channel number of 512, an output channel number of 1024, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a negative slope of 0.1; The 25th convolutional block, the 28th convolutional block, and the 30th convolutional block: Each contains a convolutional layer with an input channel number of 1024, an output channel number of 512, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a negative slope of 0.1; The 26th convolutional block: contains a convolutional layer with an input channel number of 512, an output channel number of 512, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a negative slope of 0.1; (2.2) Connect the above 31 convolutional blocks in sequence to obtain a new backbone feature extraction network Conv31 with the following structure: The 1st convolutional block -> The 2nd convolutional block -> The 3rd convolutional block -> The 4th convolutional block -> The 5th convolutional block -> The 6th convolutional block -> The 7th convolutional block -> The 8th convolutional block -> The 9th convolutional block -> The 10th convolutional block -> The 11th convolutional block -> The 12th convolutional block -> The 13th convolutional block -> The 14th convolutional block -> The 15th convolutional block -> The 16th convolutional block -> The 17th convolutional block -> The 18th convolutional block -> The 19th convolutional block -> The 20th convolutional block -> The 21st convolutional block -> The 22nd convolutional block -> The 23rd convolutional block -> The 24th convolutional block -> The 25th convolutional block -> The 26th convolutional block -> The 27th convolutional block -> The 28th convolutional block -> The 29th convolutional block -> The 30th convolutional block -> The 31st convolutional block; (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with Conv31.

4. An infrared vehicle detection method based on an improved YOLOv7 algorithm according to claim 1, characterized in that, The specific content of step 4 is as follows: (4.1) Set the training parameters: The number of training epochs is 200, the number of infrared vehicle images selected for one training is set to 16, the learning rate is set to 0.001, and both the confidence threshold and the IOU ignore threshold are set to 0.5; (4.2) Input 9000 infrared vehicle images in the training set into the model Conv31 - YOLOv7, 16 images at a time. Each time, the output obtains the offset values (t x , t y , t w , t h ) of the target bounding box relative to the annotation box and the target confidence p, where t x is the offset value of the target bounding box relative to the annotation box in the x - direction, t y is the offset value of the target bounding box relative to the annotation box in the y - direction, t w is the offset value of the target bounding box relative to the width of the annotation box, t h is the offset value of the target bounding box relative to the height of the annotation box; (4.3) Calculate the position, width, and height of the predicted bounding box by using the following coordinate offset formula with the offset values (t x , t y , t w , t h ): Among them, b x , b y is the position of the prediction box, c x , c y is the position of the annotation box, b w , b h are the width and height of the prediction box, p w , p h are the width and height of the annotation box; (4.4) Substitute the position, width and height of the predicted bounding box and the confidence of the target (b x , b y , b w , b h , p) and the position, width and height of the labeled bounding box and the confidence of the target into the loss function to calculate the loss value, and use the mini-batch stochastic gradient descent algorithm to update its weights; (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain a trained infrared vehicle detection model.

Citation Information

Patent Citations

  • Infrared vehicle rapid detection method based on improved YOLOv3 algorithm

    CN113705423A

  • Method and system for detecting vehicle on enhanced thermal infrared image based on improved SSD (Solid State Disk)

    CN115171001A