A fast infrared vehicle detection method based on improved YOLOv7 algorithm

By improving the YOLOv7 algorithm, a new lightweight trunk feature extraction network lite is built, which solves the problem of slow infrared vehicle detection speed and realizes efficient and real-time infrared vehicle detection.

CN116189059BActive Publication Date: 2025-05-23XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310218118.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-05-23
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

The existing infrared vehicle detection methods are slow to meet the real-time requirements, especially when the infrared image resolution is low and the characteristics are not obvious.

Method used

Improve the YOLOv7 algorithm, abandon the original backbone feature extraction network, build a new backbone feature extraction network lite containing 18 ordinary convolution blocks and 8 grouped convolution blocks, replace the original network, form a new network model lite-YOLOv7, and train it through the small batch stochastic gradient descent algorithm.

Benefits of technology

It significantly improves the speed of infrared vehicle detection, while maintaining a high detection accuracy, and can detect infrared vehicle targets in real time to meet real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189059B_ABST
    Figure CN116189059B_ABST
Patent Text Reader

Abstract

The invention discloses a method for rapid infrared vehicle detection based on an improved YOLOv7 algorithm, comprising the following steps: step 1: collecting vehicle videos on a traffic road for frame extraction and image preprocessing to obtain an infrared vehicle image data set; step 2: constructing a new backbone feature extraction network lite to replace the backbone feature extraction network in the YOLOv7 algorithm; step 3: connecting the new backbone feature extraction network with the prediction network of the original YOLOv7 to form a new network model lite-YOLOv7; step 4: sending the obtained training data set into the new network model lite-YOLOv7, training with a small batch random gradient descent algorithm to obtain a trained infrared vehicle detection model; step 5: sending the infrared vehicle video on the traffic road collected in real time by an infrared thermal imaging device to the trained model frame by frame to obtain real-time position information and confidence of the vehicle. The invention significantly improves the detection speed under the premise of ensuring a high detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of infrared vehicle rapid detection, and in particular relates to an infrared vehicle rapid detection method based on an improved YOLOv7 algorithm. Background Art

[0002] Infrared target detection technology refers to the automatic extraction of target location information from infrared images. Given the advantages of infrared thermal imaging, infrared target detection technology can be applied to vehicle detection scenarios on traffic roads and can adapt to conditions such as darkness, strong light, and extreme weather. Therefore, the breakthrough of this technology has important theoretical significance and practical value in fields such as autonomous driving and intelligent transportation.

[0003] Traditional infrared vehicle detection methods usually first use methods such as gradient directional histograms to extract target features, and then use positive and negative samples to train classifiers such as support vector machines to classify target features. This method has a slow detection speed and cannot meet the requirements of timeliness. It also has problems such as limited application scenarios, poor robustness, and weak generalization ability.

[0004] In recent years, with the rapid development of artificial intelligence technology, infrared vehicle detection methods based on convolutional neural networks have been widely used. It can automatically abstract and extract features from images through convolutional neural networks, and has high detection accuracy and strong robustness. At present, there are mainly two types of target detection algorithms based on deep learning. One is a two-stage detection algorithm, which divides the detection process into two stages. The first stage generates candidate regions of the image to be detected, and the second stage classifies and regresses the generated candidate regions to obtain the final detection results. The first stage of this type of algorithm is relatively time-consuming, and the overall detection accuracy is high, but the detection speed is slow, and generally cannot meet the real-time requirements. Representative algorithms include R-CNN, Fast R-CNN, Faster R-CNN, etc. The other is a single-stage detection algorithm, which unifies the above two-stage detection process into an end-to-end regression process, combining the two steps of region selection and detection judgment into one. Although the detection speed is fast, the detection accuracy is low. Representative algorithms include YOLO and SSD, etc.

[0005] The target detection algorithm based on deep learning has a good detection effect in the visible light image target detection scene, but in the infrared target detection scene, due to the low resolution and unclear features of infrared images, a more complex neural network is required for processing, resulting in a slow detection speed. The above-mentioned traditional infrared vehicle detection method and the infrared vehicle detection method based on convolutional neural network cannot meet the real-time requirements due to their slow detection speed.

[0006] The Conv31_YOLOv7 model constructs a backbone feature extraction network by stacking 31 convolutional blocks, which enhances the model's feature extraction capability and has a higher detection accuracy. However, the model has a complex network structure and a large number of parameters, so the detection speed is slow and cannot meet the needs of real-time detection. Summary of the invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide an infrared vehicle rapid detection method based on an improved YOLOv7 algorithm, which significantly improves the detection speed while ensuring a high detection accuracy.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is:

[0009] An infrared vehicle rapid detection method based on an improved YOLOv7 algorithm comprises the following steps:

[0010] Step 1: Collect vehicle videos on traffic roads, perform frame extraction and image preprocessing, and obtain infrared vehicle image datasets;

[0011] Step 2: Improve the backbone feature extraction network of the YOLOv7 algorithm, that is, discard the backbone feature extraction network in the YOLOv7 algorithm, build a new backbone feature extraction network lite containing 18 ordinary convolution blocks and 8 grouped convolution blocks, and replace the backbone feature extraction network in the YOLOv7 algorithm;

[0012] Step 3: Connect the new backbone feature extraction network with the original YOLOv7 prediction network to form a new network model lite-YOLOv7;

[0013] Step 4: Send the training data set obtained in step 1 to the new network model lite-YOLOv7, and use the small batch stochastic gradient descent algorithm for training to obtain a trained infrared vehicle detection model;

[0014] Step 5: The infrared vehicle video on the traffic road collected by the infrared thermal imaging device in real time is sent frame by frame to the trained model to obtain the real-time location information and confidence of the vehicle.

[0015] Step 1 of performing frame extraction and image preprocessing:

[0016] (1.1) Collect infrared vehicle videos on traffic roads, read the videos, set the resolution of the images to be output, output each frame in sequence in image format, obtain infrared vehicle images, and annotate the location information of vehicle targets in the obtained infrared vehicle images to create an infrared vehicle image dataset;

[0017] (1.2) The infrared vehicle image dataset is divided into a training dataset and a test dataset. In other words, infrared images are randomly selected from the dataset to form the training set, and the remaining infrared images form the test set.

[0018] The step 2 is specifically as follows:

[0019] (2.1) Abandon the backbone feature extraction network in the YOLOv7 algorithm and construct a new lightweight backbone feature extraction network lite to replace the backbone feature extraction network in the YOLOv7 algorithm. The new backbone feature extraction network lite is constructed, which contains 18 ordinary convolution blocks and 8 grouped convolution blocks, where the structure of each convolution block is as follows:

[0020] The first ordinary convolution block: contains a convolution layer with 1 input channel, 1 output channel, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 1 channel, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0021] The second ordinary convolution block: contains a convolution layer with an input channel of 1, an output channel of 32, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0022] The third normal convolution block: contains a convolution layer with 32 input channels, 32 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 32 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0023] The fourth normal convolution block: contains a convolution layer with 32 input channels, 64 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0024] The fifth normal convolution block: contains a convolution layer with 64 input channels, 64 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0025] The sixth normal convolution block: contains a convolution layer with 64 input channels, 128 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0026] The seventh normal convolution block: contains a convolution layer with 128 input channels, 128 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0027] The 8th normal convolution block: contains a convolution layer with 128 input channels, 256 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0028] The 9th and 12th normal convolution blocks: contain a convolution layer with 256 input channels, 512 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 512 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0029] The 10th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 256, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0030] The 11th normal convolution block: contains a convolution layer with 256 input channels, 256 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0031] The 13th normal convolution block, the 16th normal convolution block, and the 18th normal convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 1024, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0032] The 14th and 17th normal convolution blocks: contain a convolution layer with an input channel number of 1024, an output channel number of 512, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0033] The 15th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0034] The first grouped convolution block: contains a convolutional layer with 32 input channels, 32 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a grouping of 32, a batch normalization layer with 32 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0035] The second grouped convolution block: contains a convolution layer with 64 input channels, 64 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 64, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0036] The third grouped convolution block: contains a convolution layer with 128 input channels, 128 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 128, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0037] The 4th and 5th grouped convolution blocks: contain a convolution layer with 256 input channels, 256 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 256, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1;

[0038] The 6th group convolution block, the 7th group convolution block, and the 8th group convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 3×3, a step size of 1, a padding of 1, a group number of 512, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1. (2.2) The above 18 ordinary convolution blocks and 8 group convolution blocks are combined and connected to obtain a new backbone feature extraction network lite with the following structure:

[0039] The 1st normal convolution block -> the 2nd normal convolution block -> the 3rd normal convolution block -> the 1st grouped convolution block -> the 4th normal convolution block -> the 5th normal convolution block -> the 2nd grouped convolution block -> the 6th normal convolution block -> the 7th normal convolution block -> the 3rd grouped convolution block -> the 8th normal convolution block -> the 4th grouped convolution block -> the 9th normal convolution block -> the 10th normal convolution block -> the 11th normal convolution block -> the 5th grouped convolution block -> the 12th normal convolution block -> the 6th grouped convolution block -> the 13th normal convolution block -> the 14th normal convolution block -> the 15th normal convolution block -> the 7th grouped convolution block -> the 16th normal convolution block -> the 17th normal convolution block -> the 8th grouped convolution block -> the 18th normal convolution block.

[0040] (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with lite.

[0041] The step 3 is specifically as follows:

[0042] Connect the 9th common convolutional block in the new backbone feature extraction network lite to the 1st prediction branch of the YOLOv7 prediction network;

[0043] Connect the 13th common convolutional block in the new backbone feature extraction network lite to the second prediction branch of the YOLOv7 prediction network;

[0044] Connect the 18th common convolutional block in the new backbone feature extraction network lite to the 3rd prediction branch of the YOLOv7 prediction network;

[0045] The connection relationship between the internal modules of the YOLOv7 prediction network is:

[0046] The connection relationship between the modules of the first prediction branch is as follows:

[0047] The 9th normal convolution block->branch convolution block 1->Multi_Concat_Block1->RepConv1->detection head 1; the connection relationship of each module of the second prediction branch is as follows:

[0048] The 13th normal convolution block->branch convolution block 2->Multi_Concat_Block2->Multi_Concat_Block3->RepConv2->Detection head 2;

[0049] The connection relationship between the modules of the third prediction branch is as follows:

[0050] 18th normal convolution block->Multi_Concat_Block4->RepConv3->Detection head 3;

[0051] The connection relationship between the prediction branch modules is as follows:

[0052] 18th normal convolution block->Upsampling convolution block 1->Upsampling layer 1->Multi_Concat_Block2->Upsampling convolution block 2->Upsampling layer 2->Multi_Concat_Block1->TransitionBlock1->

[0053] Multi_Concat_Block3->TransitionBlock2->Multi_Concat_Block4.

[0054] The step 4 is specifically as follows:

[0055] (4.1) Set the training parameters: the number of training rounds is 180, the number of images selected for one training is set to 1, the learning rate is set to 0.001, and the confidence threshold and IOU ignore threshold are both set to 0.5;

[0056] (4.2) The 9000 infrared images in the training set are input into the model lite-YOLOv7 one at a time, and the offset value of the target bounding box relative to the annotation box is obtained each time. x , t y , t w , t h ) and target confidence p, where t x is the offset value of the target bounding box relative to the annotation box in the x direction, t y is the offset value of the target bounding box relative to the annotation box in the y direction, t w is the offset value of the target bounding box relative to the width of the annotation box, t h It is the offset value of the target bounding box relative to the height of the annotation box.

[0057] (4.3) The offset value (t x , t y , t w , t h ) The position, width and height of the predicted box are calculated using the following coordinate offset formula:

[0058]

[0059]

[0060]

[0061]

[0062] Among them, b x , by is the position of the prediction box, c x , c y is the position of the annotation box, b w , b h is the width and height of the prediction box, p w , p h The width and height of the annotation box;

[0063] (4.4) The position, width and height of the prediction box and the confidence of the target (b x , b y , b w , b h , p) and the position, width and height of the annotation box and the confidence of the target are substituted into the loss function to calculate the loss value, and the mini-batch stochastic gradient descent algorithm is used to update its weight;

[0064] (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain a trained infrared vehicle detection model.

[0065] Step 5: Use the trained model to perform infrared vehicle detection.

[0066] Infrared thermal imaging equipment is used to collect infrared vehicle videos on traffic roads in real time, and the videos are fed frame by frame into the trained infrared vehicle detection model to obtain the real-time location information and confidence of the vehicle.

[0067] Beneficial effects of the present invention:

[0068] The present invention replaces the main feature extraction network of the YOLOv7 algorithm with a new main feature extraction network lite, which includes 18 common convolution blocks and 8 grouped convolution blocks. The deep separable convolution operation can be realized by combining the grouped convolution blocks and the common convolution blocks, which can greatly reduce the parameter amount of the network model, realize the lightweight of the network structure, and significantly improve the detection speed, so that the infrared vehicle target can be detected in real time; the 9th common convolution block, the 13th common convolution block and the 18th common convolution block of the lite can respectively extract the shallow features, the deeper features and the deep features of the infrared vehicle target, and by connecting the three common convolution blocks with the three prediction branches, multi-scale feature fusion detection can be realized, which is conducive to the network model to detect infrared vehicle targets of different scales, thereby ensuring the high detection accuracy of the network model. Compared with the Conv31_YOLOv7 model, the lite_YOLOv7 model of the present method has the characteristics of lightweight, its structure is more streamlined, the parameter amount is less, the detection speed is faster, and it can meet the needs of real-time detection of infrared vehicle targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1It is a flow chart for implementing the present invention.

[0070] Figure 2 It is a diagram of the lite-YOLOv7 network structure constructed in the present invention.

[0071] Figure 3 It is a detection schematic diagram of the present invention in an actual scenario. DETAILED DESCRIPTION

[0072] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0073] like Figure 1 As shown:

[0074] Step 1: Build an infrared vehicle dataset.

[0075] (1.1) Collect infrared vehicle videos on traffic roads, read the first 10,000 frames of the video, set the resolution of the output image to 640×640, output each frame in sequence in image format, obtain 10,000 infrared vehicle images, and annotate the location information of the vehicle targets in the obtained infrared vehicle images to create an infrared vehicle image dataset. The dataset contains 10,000 infrared images with a resolution of 640×640.

[0076] (1.2) The infrared vehicle image dataset is divided into a training dataset and a test dataset in a ratio of 9:1. That is, 9000 infrared images are randomly selected from the dataset to form the training set, and the remaining 1000 infrared images form the test set.

[0077] Step 2: Build a new backbone feature extraction network.

[0078] This step constructs a new backbone feature extraction network based on a lightweight improvement of the backbone feature extraction network of the existing YOLOv7 algorithm. The network model in the YOLOv7 algorithm includes a backbone feature extraction network and a prediction network. This step only improves its backbone feature extraction network, which is specifically implemented as follows:

[0079] (2.1) Abandon the backbone feature extraction network in the YOLOv7 algorithm and construct a new lightweight backbone feature extraction network lite to replace the backbone feature extraction network in the YOLOv7 algorithm. Construct a new backbone feature extraction network lite, which contains 18 ordinary convolution blocks and 8 grouped convolution blocks. Among them, the 3rd, 5th, 7th, 11th, and 15th ordinary convolution blocks are used to downsample the input feature map, which can not only reduce the network parameters and prevent network overfitting, but also increase the receptive field, so that the network can learn more global information; the role of other ordinary convolution blocks is to deepen the network depth and improve the nonlinear expression ability of the network, thereby extracting more effective features of the input image; the role of the grouped convolution block is to group the input feature map by channel, and each group is convolved separately, which can effectively reduce the network parameters compared with the ordinary convolution block. The structure of each convolution block is as follows:

[0080] The first ordinary convolution block: contains a convolution layer with an input channel of 1, an output channel of 1, a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with a channel number of 1, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0081] The second ordinary convolution block: contains a convolution layer with an input channel of 1, an output channel of 32, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0082] The third ordinary convolution block: contains a convolution layer with 32 input channels, 32 output channels, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 32 channels, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0083] The fourth ordinary convolution block: contains a convolution layer with 32 input channels, 64 output channels, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0084] The fifth ordinary convolution block: contains a convolution layer with 64 input channels, 64 output channels, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0085] The sixth normal convolution block: contains a convolution layer with 64 input channels, 128 output channels, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0086] The seventh ordinary convolution block: contains a convolution layer with an input channel number of 128, an output channel number of 128, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0087] The 8th normal convolution block: contains a convolution layer with an input channel number of 128, an output channel number of 256, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0088] The 9th and 12th ordinary convolution blocks contain a convolution layer with an input channel number of 256, an output channel number of 512, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0089] The 10th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 256, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0090] The 11th normal convolution block: contains a convolution layer with an input channel number of 256, an output channel number of 256, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0091] The 13th ordinary convolution block, the 16th ordinary convolution block, and the 18th ordinary convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 1024, a convolution kernel size of 1×1, a step size of 1, and a padding of 0, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0092] The 14th and 17th ordinary convolution blocks contain a convolution layer with an input channel number of 1024, an output channel number of 512, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0093] The 15th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0094] The first grouped convolution block: contains a convolution layer with 32 input channels, 32 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 32, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0095] The second grouped convolution block: contains a convolution layer with 64 input channels, 64 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 64, a batch normalization layer with a channel number of 64, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0096] The third grouped convolution block: contains a convolution layer with 128 input channels, 128 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 128, a batch normalization layer with a channel number of 128, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0097] The 4th and 5th grouped convolution blocks contain a convolution layer with 256 input channels, 256 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 256, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0098] The 6th grouped convolution block, the 7th grouped convolution block, and the 8th grouped convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 3×3, a step size of 1, a padding of 1, and a group number of 512, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1.

[0099] (2.2) The above 18 ordinary convolution blocks and 8 grouped convolution blocks are combined and connected to obtain a new backbone feature extraction network lite with the following structure. Lite contains 18 ordinary convolution blocks and 8 grouped convolution blocks. Among them, the 3rd, 5th, 7th, 11th, and 15th ordinary convolution blocks are used to downsample the input feature map, which can not only reduce the network parameters and prevent network overfitting, but also increase the receptive field, so that the network can learn more global information; the role of other ordinary convolution blocks is to deepen the network depth, improve the nonlinear expression ability of the network, and then extract more effective features of the input image; the role of the grouped convolution block is to group the input feature map by channel, and convolve each group separately, which can effectively reduce the network parameters compared with the ordinary convolution block. By combining the grouped convolution block and the ordinary convolution block, the depth separable convolution operation can be realized, which can greatly reduce the parameters of the network model, realize the lightweight network structure, and significantly improve the detection speed, so that infrared vehicle targets can be detected in real time.

[0100] The 1st normal convolution block -> the 2nd normal convolution block -> the 3rd normal convolution block -> the 1st grouped convolution block -> the 4th normal convolution block -> the 5th normal convolution block -> the 2nd grouped convolution block -> the 6th normal convolution block -> the 7th normal convolution block -> the 3rd grouped convolution block -> the 8th normal convolution block -> the 4th grouped convolution block -> the 9th normal convolution block -> the 10th normal convolution block -> the 11th normal convolution block -> the 5th grouped convolution block -> the 12th normal convolution block -> the 6th grouped convolution block -> the 13th normal convolution block -> the 14th normal convolution block -> the 15th normal convolution block -> the 7th grouped convolution block -> the 16th normal convolution block -> the 17th normal convolution block -> the 8th grouped convolution block -> the 18th normal convolution block.

[0101] (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with lite.

[0102] Step 3: Build a new network model lite-YOLOv7.

[0103] Reference Figure 2 , connect the new backbone feature extraction network and the prediction network of YOLOv7 according to the following structural relationship to form a new network model lite-YOLOv7:

[0104] Connect the 9th convolutional block in the new backbone feature extraction network lite to the 1st prediction branch of the YOLOv7 prediction network.

[0105] Connect the 13th convolutional block in the new backbone feature extraction network lite to the 2nd prediction branch of the YOLOv7 prediction network.

[0106] Connect the 18th convolutional block in the new backbone feature extraction network lite to the 3rd prediction branch of the YOLOv7 prediction network.

[0107] Step 4: Train the new network model lite-YOLOv7.

[0108] (4.1) Set the training parameters: the number of training rounds is 180, the number of images selected for one training is set to 1, the learning rate is set to 0.001, and the confidence threshold and IOU ignore threshold are both set to 0.5.

[0109] (4.2) The 9000 infrared images in the training set are input into the model lite-YOLOv7 one at a time, and the offset value of the target bounding box relative to the annotation box is obtained each time. x , t y , t w , t h ) and target confidence p, where t x is the offset value of the target bounding box relative to the annotation box in the x direction, t y is the offset value of the target bounding box relative to the annotation box in the y direction, t w is the offset value of the target bounding box relative to the width of the annotation box, t h It is the offset value of the target bounding box relative to the height of the annotation box.

[0110] (4.3) The offset value (t x , t y , t w , t h ) The position, width and height of the predicted box are calculated using the following coordinate offset formula:

[0111]

[0112]

[0113]

[0114]

[0115] Among them, b x , b y is the position of the prediction box, c x , c y is the position of the annotation box, b w , b h is the width and height of the prediction box, p w , p h The width and height of the label box.

[0116] (4.4) The position, width and height of the prediction box and the confidence of the target (bx , b y , b w , b h , p) and the position, width and height of the annotation box and the confidence of the target are substituted into the loss function to calculate the loss value, and the small batch stochastic gradient descent algorithm is used to update its weight.

[0117] (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain a trained infrared vehicle detection model.

[0118] Step 5: Use the trained model to perform infrared vehicle detection.

[0119] Infrared thermal imaging equipment is used to collect infrared vehicle videos on traffic roads in real time, and the videos are fed frame by frame into the trained infrared vehicle detection model to obtain the real-time location information and confidence of the vehicle.

[0120] The effect of the present invention is further illustrated by the following simulation experiments and measured data:

[0121] 1. Simulation and test environment

[0122] The simulation and actual measurement of the present invention use the Windows 10 operating system, an NVIDIA GeForce GTX1050Ti GPU for acceleration, and the deep learning framework used is pytorch 1.8.1.

[0123] 2. Simulation Content

[0124] Simulation 1: Use the same training set and parameters as the present invention to train other target detection models based on convolutional neural networks to obtain respective trained infrared vehicle detection models.

[0125] 1000 test set images are sent one at a time to the trained model of the present invention for testing, and the accuracy and speed of the infrared vehicle detection of the present invention with an IOU threshold of 0.5 are obtained.

[0126] Using the same test set as the present invention, the accuracy and speed of infrared vehicle detection of other methods are tested with an IOU threshold of 0.5.

[0127] The method of the present invention is compared with the infrared vehicle detection method based on YOLOv7 through simulation experiments. The comparison results are shown in Table 1:

[0128] Table 1

[0129]

[0130] By comparing the present invention with YOLOv7, it is found that the accuracy of the method of the present invention for infrared vehicle detection with an IOU threshold of 0.5 is 91.57%, and the accuracy of YOLOv7 for infrared vehicle target detection is 91.89%. The detection accuracy of the method of the present invention is similar to that of YOLOv7, ensuring a high detection accuracy; and the detection speed of the method of the present invention is more than twice that of other methods, with obvious real-time advantages.

[0131] 3. Test content

[0132] Infrared thermal imaging equipment is used to collect infrared vehicle videos on traffic roads in real time, and the videos are fed into the trained infrared vehicle detection model of the present invention frame by frame to obtain the real-time position information and confidence of the vehicle, such as Figure 3 shown.

[0133] Figure 3 The large rectangular box in the middle represents the predicted box surrounding the vehicle in the infrared image, and the small rectangular box above the large rectangular box shows the confidence of the vehicle target.

[0134] The specific implementation methods described above are only descriptions of the preferred methods of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A fast infrared vehicle detection method based on improved YOLOv7 algorithm, It is characterized in that The steps include: Step 1: Collect vehicle videos on traffic roads, perform frame extraction and image preprocessing, and obtain infrared vehicle image datasets; Step 2: Improve the backbone feature extraction network of the YOLOv7 algorithm, that is, discard the backbone feature extraction network in the YOLOv7 algorithm, build a new backbone feature extraction network lite containing 18 ordinary convolution blocks and 8 grouped convolution blocks, and replace the backbone feature extraction network in the YOLOv7 algorithm; Step 3: Connect the new backbone feature extraction network with the original YOLOv7 prediction network to form a new network model lite-YOLOv7; Step 4: Send the training data set obtained in step 1 to the new network model lite-YOLOv7, and use the small batch stochastic gradient descent algorithm for training to obtain a trained infrared vehicle detection model; Step 5: The infrared vehicle video on the traffic road collected by the infrared thermal imaging device in real time is sent frame by frame to the trained model to obtain the real-time location information and confidence of the vehicle.

2. According to claim 1, a method for rapid infrared vehicle detection based on an improved YOLOv7 algorithm, It is characterized in that Step 1 of performing frame extraction and image preprocessing: (1.1) Collect infrared vehicle videos on traffic roads, read the videos, set the resolution of the images to be output, output each frame in sequence in image format, obtain infrared vehicle images, and annotate the location information of vehicle targets in the obtained infrared vehicle images to create an infrared vehicle image dataset; (1.2) The infrared vehicle image dataset is divided into a training dataset and a test dataset. In other words, infrared images are randomly selected from the dataset to form the training set, and the remaining infrared images form the test set.

3. According to claim 1, a method for rapid infrared vehicle detection based on an improved YOLOv7 algorithm, It is characterized in that The step 2 is specifically as follows: (2.1) Abandon the backbone feature extraction network in the YOLOv7 algorithm and construct a new lightweight backbone feature extraction network lite to replace the backbone feature extraction network in the YOLOv7 algorithm. The new backbone feature extraction network lite is constructed, which contains 18 ordinary convolution blocks and 8 grouped convolution blocks, where the structure of each convolution block is as follows: The first ordinary convolution block: contains a convolution layer with 1 input channel, 1 output channel, a kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization layer with 1 channel, and a LeakyReLU activation function layer with a complex slope of 0.1; The second ordinary convolution block: contains a convolution layer with an input channel of 1, an output channel of 32, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 32, and a LeakyReLU activation function layer with a complex slope of 0.1; The third normal convolution block: contains a convolution layer with 32 input channels, 32 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 32 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The fourth normal convolution block: contains a convolution layer with 32 input channels, 64 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The fifth normal convolution block: contains a convolution layer with 64 input channels, 64 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The sixth normal convolution block: contains a convolution layer with 64 input channels, 128 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The seventh normal convolution block: contains a convolution layer with 128 input channels, 128 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The 8th normal convolution block: contains a convolution layer with 128 input channels, 256 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The 9th and 12th normal convolution blocks: contain a convolution layer with 256 input channels, 512 output channels, a kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with 512 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The 10th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 256, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 256, and a LeakyReLU activation function layer with a complex slope of 0.1; The 11th normal convolution block: contains a convolution layer with 256 input channels, 256 output channels, a kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The 13th normal convolution block, the 16th normal convolution block, and the 18th normal convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 1024, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 1024, and a LeakyReLU activation function layer with a complex slope of 0.1; The 14th and 17th normal convolution blocks: contain a convolution layer with an input channel number of 1024, an output channel number of 512, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1; The 15th normal convolution block: contains a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 1×1, a stride of 2, and a padding of 0, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1; The first grouped convolution block: contains a convolutional layer with 32 input channels, 32 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a grouping of 32, a batch normalization layer with 32 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The second grouped convolution block: contains a convolution layer with 64 input channels, 64 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 64, a batch normalization layer with 64 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The third grouped convolution block: contains a convolution layer with 128 input channels, 128 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 128, a batch normalization layer with 128 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The 4th and 5th grouped convolution blocks: contain a convolution layer with 256 input channels, 256 output channels, a kernel size of 3×3, a stride of 1, a padding of 1, and a group number of 256, a batch normalization layer with 256 channels, and a LeakyReLU activation function layer with a complex slope of 0.1; The sixth group convolution block, the seventh group convolution block, and the eighth group convolution block: contain a convolution layer with an input channel number of 512, an output channel number of 512, a convolution kernel size of 3×3, a step size of 1, a padding of 1, and a group number of 512, a batch normalization layer with a channel number of 512, and a LeakyReLU activation function layer with a complex slope of 0.1; (2.2) The above 18 ordinary convolution blocks and 8 group convolution blocks are combined and connected to obtain a new backbone feature extraction network lite with the following structure: 1st normal convolution block -> 2nd normal convolution block -> 3rd normal convolution block -> 1st grouped convolution block -> 4th normal convolution block -> 5th normal convolution block -> 2nd grouped convolution block -> 6th normal convolution block -> 7th normal convolution block -> 3rd grouped convolution block -> 8th normal convolution block -> 4th grouped convolution block -> 9th normal convolution block -> 10th normal convolution block -> 11th normal convolution block -> 5th grouped convolution block -> 12th normal convolution block -> 6th grouped convolution block -> 13th normal convolution block -> 14th normal convolution block -> 15th normal convolution block -> 7th grouped convolution block -> 16th normal convolution block -> 17th normal convolution block -> 8th grouped convolution block -> 18th normal convolution block; (2.3) Replace the backbone feature extraction network in the YOLOv7 algorithm with lite.

4. According to claim 3, a method for rapid infrared vehicle detection based on an improved YOLOv7 algorithm, It is characterized in that The step 3 is specifically as follows: Connect the 9th common convolutional block in the new backbone feature extraction network lite to the 1st prediction branch of the YOLOv7 prediction network; Connect the 13th common convolutional block in the new backbone feature extraction network lite to the second prediction branch of the YOLOv7 prediction network; Connect the 18th common convolutional block in the new backbone feature extraction network lite to the 3rd prediction branch of the YOLOv7 prediction network; The connection relationship between the internal modules of the YOLOv7 prediction network is: The connection relationship between the modules of the first prediction branch is as follows: The 9th normal convolution block->branch convolution block 1->Multi_Concat_Block1->RepConv1->detection head 1; the connection relationship of each module of the second prediction branch is as follows: The 13th normal convolution block->branch convolution block 2->Multi_Concat_Block2->Multi_Concat_Block3->RepConv2->Detection head 2; The connection relationship between the modules of the third prediction branch is as follows: 18th normal convolution block->Multi_Concat_Block4->RepConv3->Detection head 3; The connection relationship between the prediction branch modules is as follows: 18th normal convolution block -> upsampling convolution block 1 -> upsampling layer 1 -> Multi_Concat_Block2 -> upsampling convolution block 2 -> upsampling layer 2 -> Multi_Concat_Block1 -> TransitionBlock1 -> Multi_Concat_Block3 -> TransitionBlock2 -> Multi_Concat_Block4.

5. According to claim 1, a method for rapid infrared vehicle detection based on an improved YOLOv7 algorithm, It is characterized in that The step 4 is specifically as follows: (4.1) Set the training parameters: the number of training rounds is 180, the number of images selected for one training is set to 1, the learning rate is set to 0.001, and the confidence threshold and IOU ignore threshold are both set to 0.5; (4.2) The 9000 infrared images in the training set are input into the model lite-YOLOv7 one at a time, and the offset value of the target bounding box relative to the annotation box is obtained each time. x ,t y ,t w ,t h ) and target confidence p, where t x is the offset value of the target bounding box relative to the annotation box in the x direction, t y is the offset value of the target bounding box relative to the annotation box in the y direction, t w is the offset value of the target bounding box relative to the width of the annotation box, t h is the offset value of the target bounding box relative to the height of the annotation box; (4.3) The offset value (t x ,t y ,t w ,t h ) The position and width and height of the predicted box are calculated using the following coordinate offset formula: Among them, b x , b y is the position of the prediction box, c x , c y is the position of the annotation box, b w , b h is the width and height of the prediction box, p w , p h The width and height of the annotation box; (4.4) The position, width and height of the prediction box and the confidence of the target (b x ,b y ,b w ,b h ,p) Substitute the position, width, height and confidence of the target with the annotation box into the loss function to calculate the loss value, and use the small batch stochastic gradient descent algorithm to update its weight; (4.5) Repeat (4.2)-(4.4) until the loss value stabilizes and no longer decreases, then stop training to obtain a trained infrared vehicle detection model.

Citation Information

Patent Citations

  • Feature fusion and dense connection-based method for infrared plane object detection

    US20210174149A1

  • Lightweight network-based traffic sign recognition method

    WO2022205685A1