Vehicle Detection Method, System and Device for Traffic Monitoring Videos Based on Improved CenterNet
Through GridMask data enhancement and multi-scale feature fusion, the CenterNet network is improved to solve the problem of reduced vehicle detection accuracy in traffic surveillance video, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202310195185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-03
AI Technical Summary
In traffic surveillance video, the problems of large changes in vehicle target scale and reduction in detection accuracy caused by partial occlusion are difficult to effectively solve in the prior art.
The GridMask data enhancement strategy is used to simulate occlusion scenarios, combining multi-scale feature fusion and adaptive spatial feature fusion methods to perform vehicle detection by improving the CenterNet network.
It improves the accuracy and robustness of vehicle detection in traffic surveillance videos, effectively alleviating the problem of reducing detection accuracy caused by vehicle target scale changes and occlusion.
Smart Images

Figure CN116343138B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method, a system and a device for vehicle detection in traffic monitoring videos based on improved CenterNet. Background Art
[0002] Object detection is an important branch in the field of computer vision. Its basic task is to find all the objects of interest in an image or a video and obtain the category and location information of the objects. As a specific application field in object detection, vehicle detection plays a key role in constructing an intelligent transportation system and is a prerequisite for a series of operations such as vehicle counting, anomaly detection, and fine vehicle recognition.
[0003] Traditional vehicle detection algorithms are mainly divided into three parts: region selection, feature extraction, and classifier classification. First, the image is scanned through a sliding window, and regions of interest are selected according to the position and size features of the vehicle in the image. Then, features are extracted from the candidate regions, and finally, the features are classified by a classifier such as a support vector machine to complete vehicle detection. Although traditional object detection algorithms have achieved good results in the field of vehicle detection, traditional vehicle detection algorithms need to manually construct features, and the manually established features have poor robustness and generalization to environmental changes, light intensity, and object shape changes. At the same time, based on the sliding window region selection strategy, the time complexity is high, the windows are redundant, and it cannot meet the real-time detection requirements of vehicles. Therefore, traditional vehicle detection algorithms are used less and less. In the past decade, in the field of object detection, deep learning methods have achieved great success. Compared with traditional detection methods, vehicle detection methods based on deep learning can automatically extract features through convolutional neural networks, have good robustness, and have been widely used.
[0004] Patent application CN110298227A discloses a method for vehicle detection in drone aerial images based on deep learning. Aiming at problems such as environmental interference and light influence in the process of vehicle detection by traditional image processing algorithms, the network structure of Faster RCNN is improved and applied to vehicle detection in drone aerial images. However, for vehicles with variable scales and partially occluded vehicles in traffic monitoring videos, they still cannot be accurately detected. Summary of the Invention
[0005] In order to overcome the above deficiencies of the prior art, the purpose of the present invention is to provide a method, a system and a device for vehicle detection in traffic monitoring videos based on improved CenterNet. Aiming at problems such as large changes in the scale of vehicle targets and partial occlusion caused by vehicle movement in traffic monitoring videos, the GridMask data augmentation strategy is used to simulate vehicle occlusion scenarios, and methods of multi-scale feature fusion and adaptive spatial feature fusion are adopted to improve the accuracy of vehicle detection in traffic monitoring scenarios.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A traffic surveillance video vehicle detection method based on improved CenterNet, comprising the following steps:
[0008] Step 1, collect experimental data and divide the traffic surveillance video into image sequences, preprocess the input images to obtain the images after data augmentation;
[0009] Step 2, in the training stage, use the feature extraction network to extract features of different scales for the input images, fuse the extracted features of different scales and output;
[0010] Step 3, in the training stage, for the feature map output in Step 2, through the detection head convolutional network branch, generate the prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset, and calculate the loss value;
[0011] Step 4, in the test stage, divide the test video into image sequences, input them into the network model trained in Steps 1 to 3, and generate the test results.
[0012] Further, the image preprocessing in Step 1 includes the following steps:
[0013] Step 1.1, perform Random Flipping operation on the input images;
[0014] Step 1.2, perform Random Scaling operation on the input images;
[0015] Step 1.3, perform GridMask operation on the input images.
[0016] Further, the feature extraction network and feature fusion in Step 2 include the following steps:
[0017] Step 2.1, input the images into the feature extraction network to obtain features C0, C1, C2, C3, C4;
[0018] Step 2.2, expand the receptive field of the feature C4 through the dilated encoder to obtain P4;
[0019] Step 2.3, obtain the feature P4 through deformable convolution and transposed convolution ′ , the feature C4 passes through deformable convolution and the feature P4 ′ are added to obtain the feature P3, and the feature P3 is further processed through deformable convolution and transposed convolution to obtain the feature P3 ′ is added to the C3 feature passing through deformable convolution to obtain P2; and so on, P2 and C1 features are fused to obtain the feature P1;
[0020] Step 2.4: Input the features P3, P2, and P1 obtained in Step 2.3 into the adaptive spatial feature fusion module to obtain the feature F.
[0021] Furthermore, the dilation encoder described in Step 2.2 is implemented according to the following steps:
[0022] Step 2.2.a: Sequentially pass the input feature C4 through a 1x1 convolution, a BatchNorm layer, and a Relu layer to reduce the channel dimension.
[0023] Step 2.2.b: Pass the feature obtained in Step 2.2.a through a 3x3 convolution, a BatchNorm layer, and a Relu layer to refine the context semantic features.
[0024] Step 2.2.c: In order to enable the receptive field of the feature output by the encoder to cover targets at various scales, in the present invention, the feature obtained in Step 2.2.b is passed through 4 consecutive stacked residual network structures to obtain the feature P4; among them, the residual network structure consists of 3 consecutive convolutions: the first 1x1 convolution reduces the number of channels at a reduction rate of 4 times, then a 3x3 dilation convolution is used to expand the receptive field, and finally, a 1x1 convolution is used to restore the number of channels; the dilation rates of the 4 consecutive dilation convolutions are [2, 4, 6, 8].
[0025] Furthermore, the specific method of the different feature fusion methods described in Step 2.3 is as follows:
[0026] Step 2.3.a: Sequentially pass the feature P4 obtained in Step 2.2.c through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer to further extract features.
[0027] Step 2.3.b: Sequentially pass the feature obtained in Step 2.3.a through a 4x4 transposed convolution, a BatchNorm layer, and a Relu layer to increase the resolution of the feature map to obtain the feature P4 ′ ;
[0028] Step 2.3.c: Sequentially pass the feature C3 obtained in Step 2.2 through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer to adjust the number of feature channels.
[0029] Step 2.3.d: Perform an addition operation on the feature P4 obtained in Step 2.3.b ′ and the feature obtained in Step 2.3.c, and after passing through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer, obtain the fused feature P3.
[0030] Step 2.3.e, similar to the feature fusion methods in Steps 2.3.a, 2.3.b, 2.3.c, and 2.3.d, fuse Feature P3 and Feature C2 to obtain Feature P2, as Figure 5 shown;
[0031] Step 2.3.f, similar to the feature fusion methods in Steps 2.3.a, 2.3.b, 2.3.c, and 2.3.d, fuse Feature P2 and Feature C1 to obtain Feature P1;
[0032] Furthermore, the adaptive spatial feature fusion method described in Step 2.4 is specifically as follows:
[0033] Step 2.4.a, adjust the size of Feature P3 to be the same as that of Feature P1 through 1x1 convolution and an upsampling layer to obtain Feature F3;
[0034] Step 2.4.b, adjust the size of Feature P2 to be the same as that of Feature P1 through 1x1 convolution and an upsampling layer to obtain Feature F2;
[0035] Step 2.4.c, according to the fusion strategy, perform weighted fusion on F3, F2, and P1 to obtain the fused feature F;
[0036] Among them, the fusion strategy is:
[0037] F ij = α ij ·F3 ij + β ij ·F2 ij + γ ij ·P1 ij
[0038] F ij is the value of Feature F at position (i, j), and F3 ij , F2 ij and P1 ij have similar meanings; α ij , β ij and γ ij are the weights of their respective feature maps at position (i, j), and satisfy the constraint conditions of α ij + β ij + γ ij = 1 and α ij , β ij , γ j ∈ [0, 1]. This constraint condition is achieved by performing softmax calculation on F3 ij , F2 ij and P1 ij after 1x1 convolution.
[0039] Further, the convolutional network branches in step 3 are as follows: After further extracting features through a 3x3 convolution, a BatchNorm layer, and a Relu layer, a 1x1 convolution, a BatchNorm layer, and a Relu layer are used to adjust the number of feature channels;
[0040] Further, step 3 specifically includes:
[0041] Step 3.1, Detection head branch 1 predicts the heatmap of vehicle targets on feature F and calculates the center point loss;
[0042] Step 3.2, Detection head branch 2 predicts the width and height information of vehicle targets on feature F and calculates the width and height loss;
[0043] Step 3.3, Detection head branch 3 predicts the center point offset caused by downsampling on feature F and calculates the center point offset loss.
[0044] The specific steps of step 4 are as follows:
[0045] Step 4.1, Divide the test video into a sequence of pictures;
[0046] Step 4.2, Perform Flip data augmentation on the test data and input it into the model trained in steps 1 to 3 together with the original image;
[0047] Step 4.3, During testing, according to the prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset generated in step 3, calculate the loss value to obtain a detection box with offset.
[0048] Step 4.4, Combine the test results of the original image and the test results of the Flip-enhanced image to obtain the final detection box.
[0049] A traffic surveillance video vehicle detection system based on improved CenterNet, comprising:
[0050] A data processing module, used to divide the input traffic surveillance video into an image sequence, preprocess the input images, expand the data set, and at the same time simulate the situation where the vehicle is blocked to enhance the robustness of the model;
[0051] A feature extraction module, used to extract multi-scale features from the data processed by the data processing module;
[0052] A feature fusion module, first uses a dilated encoder to expand the receptive field for the deepest feature; then, fuses the feature after passing through the dilated convolutional layer with the shallow features extracted by the feature extraction module to improve the detection accuracy of the model for targets of different scales; finally, uses an adaptive spatial feature fusion module for the features of the last three layers to filter out the useless information of the last three layers in space and only retain the useful information for fusion;
[0053] A detection module, which is used to predict the center point of the vehicle, the width and height information of the vehicle, and the offset of the vehicle center point, and generate a final vehicle detection frame.
[0054] A vehicle detection device for traffic monitoring video based on improved CenterNet, comprising:
[0055] A video image collector, which is used to collect traffic monitoring video images;
[0056] A program processor, which is used to store a computer program and execute the computer program to implement the vehicle detection method for traffic monitoring video based on improved CenterNet described in steps 1 to 4;
[0057] A display, which is used to display the vehicle detection results of traffic monitoring video.
[0058] Compared with the prior art, the present invention has the following beneficial technical effects:
[0059] 1. Considering that in the traffic monitoring video scenario, partial occlusion between vehicles often occurs, the present invention uses GridMask data augmentation in the data preprocessing stage to simulate the occurrence of occlusion. At the initial stage of the training stage, occlusion is performed with a random grid at a probability of 0%. As the number of training times increases, the probability of performing GridMask augmentation on the pictures gradually increases until 70%, which improves the robustness and generalization of the model.
[0060] 2. Considering that in the traffic monitoring video scenario, the scale of vehicles changes greatly under the camera view, after using ResNet101 to extract features, the present invention introduces a dilated encoder to expand the receptive field while extracting features; at the same time, a feature fusion module is designed to fuse multi-scale spatial features to further improve the detection accuracy.
[0061] 3. The present invention uses Flip data augmentation in the test stage to further improve the accuracy of vehicle detection in traffic monitoring video.
[0062] The vehicle target detection method for monitoring traffic video of the present invention effectively alleviates the problem of reduced detection accuracy caused by large scale changes of vehicle targets and partial occlusion of vehicles in traffic monitoring video, and effectively solves the vehicle detection task in traffic monitoring video. Description of the Drawings
[0063] Figure 1 It is a schematic diagram of the overall network structure adopted by the vehicle target detection method for traffic monitoring video.
[0064] Figure 2 It is a network training flow chart.
[0065] Figure 3It is a schematic diagram of the structure of the dilation encoder.
[0066] Figure 4 It is a flowchart of the test phase.
[0067] Figure 5 It is a schematic diagram of the structure for fusing features of different scales.
[0068] Figure 6 It is a schematic diagram of the input picture.
[0069] Figure 7 It is a schematic diagram of GridMask data augmentation. Specific implementation manners
[0070] The present invention will be further described in detail below with reference to the accompanying drawings.
[0071] This embodiment provides a method for detecting vehicle targets in traffic videos. The overall framework of the network structure specifically adopted is as Figure 1 shown:
[0072] First, input a traffic monitoring picture with a size of [512, 512, 3], and use the ResNet101 network with COCO as the pre-trained model to extract feature maps C0, C1, C2, C3, C4 of different levels. The sizes of the features are [256, 256, 64], [128, 128, 128], [64, 64, 256], [32, 32, 512], [16, 16, 2048] respectively; further, use the dilation encoder to expand the receptive field of the C4 feature to obtain the feature P4; then, fuse the feature P4 and C3 to generate the P3 feature with a size of [32, 32, 256]; similarly, fuse the feature P3 and C2 to generate the P2 feature with a size of [64, 64, 128], and fuse the feature P2 and C1 to generate the P1 feature with a size of [128, 128, 64]; use the adaptive spatial feature fusion module to fuse P3, P2, and P1 to obtain the feature F; finally, on the feature F, use 3 detection head convolutional branches to respectively output the heat map Target size and the center point offset The final sizes of the network outputs are [128, 128, 4], [128, 128, 2], and [128, 128, 2].
[0073] The specific training process is as Figure 2 shown, including the following steps:
[0074] Step 1, before the start of the training phase, collect experimental data and divide the traffic monitoring video into picture sequences, and perform picture preprocessing on the input pictures to obtain the pictures after data augmentation.
[0075] In this embodiment, all experimental dataset pictures are as shown in Figure 6 shown below
[0076] Specific picture preprocessing operations are as follows:
[0077] Step 1.1: Perform a Flipping operation on the input picture, randomly flipping the input picture horizontally with a probability of 0.5;
[0078] Step 1.2: Perform a Scaling operation on the input picture, randomly scaling the input picture with a probability of 0.5. The scaled picture is at least 0.6 times the original size and at most 1.3 times the original size;
[0079] Step 1.3: Perform a GridMask operation on the input picture. At the beginning of training, randomly mask the picture with a probability of 0. As the number of training times increases, the probability of GridMask enhancement for the picture gradually increases and finally becomes 0.7. Among them, the ratio r of the input image in each grid is 0.5, and the size d of each grid ranges from 96 to 224. The results of GridMask data augmentation are as shown in Figure 7 shown below
[0080] Step 2: In the training stage, use the deep feature extraction network ResNet101 to extract features of different scales from the input pictures, and fuse the features of different scales extracted by the network ResNet101. The specific steps are as follows:
[0081] Step 2.1: Input the preprocessed picture into the ResNet101 feature extraction network to obtain features C0, C1, C2, C3, and C4 at different levels. Among them, the size of the input picture is [512, 512, 3], and the sizes of the output features of different scales are [256, 256, 64], [128, 128, 128], [64, 64, 256], [32, 32, 512], and [16, 16, 2048] respectively;
[0082] Step 2.2: Use a dilated encoder to expand the receptive field of the C4 feature. The specific schematic diagram of the dilated encoder structure is as shown in Figure 3 shown below
[0083] Step 2.2.a: Pass the input feature C4 through a 1x1 convolution layer, a BatchNorm layer, and a Relu layer in sequence to reduce the channel dimension;
[0084] Step 2.2.b: Pass the feature obtained in Step 2.2.a through a 3x3 convolution layer, a BatchNorm layer, and a Relu layer to refine the context semantic features;
[0085] Step 2.2.c. To enable the feature receptive field output by the encoder to cover targets of various scales, the present invention obtains feature P4 by passing the features obtained in Step 2.2.b through four consecutive stacked residual network structures; among them, the residual network structure consists of three consecutive convolutions: the first 1x1 convolution reduces the number of channels at a reduction rate of 4 times, then a 3x3 dilated convolution is used to expand the receptive field, and finally, a 1x1 convolution is used to restore the number of channels; the dilation rates of the four consecutive dilated convolutions are [2, 4, 6, 8].
[0086] Step 2.3. Obtain feature P4 by passing feature P4 through deformable convolution and transposed convolution ′ , feature C4 passes through deformable convolution and feature P4 ′ are added to obtain feature P3. Feature P3 further passes through deformable convolution and transposed convolution to obtain feature P3 ′ and is added to the C3 feature that has passed through deformable convolution to obtain P2; and so on, until P2 and C1 features are fused to obtain feature P1. The schematic diagram of the specific feature fusion structure at different scales is shown in Figure 5.
[0087] Step 2.3.a. Pass the feature P4 obtained in Step 2.2.c through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer in sequence to further extract features;
[0088] Step 2.3.b. Pass the features obtained in Step 2.3.a through a 4x4 transposed convolution, a BatchNorm layer, and a Relu layer in sequence to increase the resolution of the feature map to obtain feature P4 ′ ;
[0089] Step 2.3.c. Pass the feature C3 obtained in Step 2.2 through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer in sequence to adjust the number of feature channels;
[0090] Step 2.3.d. Add the feature P4 obtained in Step 2.3.b ′ and the feature obtained in Step 2.3.c, and after passing through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer, the fused feature P3 is obtained;
[0091] Step 2.3.e. Similar to the feature fusion methods in Step 2.3.a, Step 2.3.b, Step 2.3.c, and Step 2.3.d, fuse feature P3 and feature C2 to obtain feature P2;
[0092] Step 2.3.f. Similar to the feature fusion methods in Step 2.3.a, Step 2.3.b, Step 2.3.c, and Step 2.3.d, fuse feature P2 and feature C1 to obtain feature P1;
[0093] Step 2.4, input the features P3, P2, and P1 obtained in Step 2.3 into the adaptive spatial feature fusion module to obtain the feature F.
[0094] Step 2.4.a, adjust the size of the feature P3 to be the same as that of the feature P1 through 1x1 convolution and an upsampling layer to obtain the feature F3;
[0095] Step 2.4.b, adjust the size of the feature P2 to be the same as that of the feature P1 through 1x1 convolution and an upsampling layer to obtain the feature F2;
[0096] Step 2.4.c, according to the fusion strategy, perform weighted fusion on F3, F2, and P1 to obtain the fused feature F;
[0097] Among them, the fusion strategy is:
[0098] F ij = α ij ·F3 ij + β ij ·F2 ij + γ ij ·P1 ij
[0099] F ij is the value of the feature F at the position (i, j), and F3 ij , F2 ij and P1 ij have similar meanings; α ij , β ij and γ ij are the weights of their respective feature maps at the position (i, j), and satisfy the constraints of α ij + β ij + γ ij = 1 and α ij , β ij , γ ij ∈ [0, 1]. This constraint condition is achieved by performing softmax calculation after 1x1 convolution on F3 ij , F2 ij and P1 ij ;
[0100] Step 3, in the training stage, for the feature map F output in Step 2, generate the prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset through 3 detection head convolutional network branches;
[0101] Specifically, the convolutional network branch in Step 3 is to first further extract features through 3x3 convolution, BatchNorm layer, and Relu layer, and then use 1x1 convolution, BatchNorm layer, and Relu layer to adjust the number of feature channels;
[0102] Step 3.1, predict the heatmap of the vehicle target Calculate the center point loss.
[0103] The center point loss function adopts Focal Loss, and the formula is as follows:
[0104]
[0105] Among them, by map the target on the Ground Truth to the heatmap, Y ∈ [0, 1] 128×128×4 , p is the position of the center point of the target on the Ground Truth, α and β are hyperparameters, which are set to 2 and 4 in specific experiments, and N is the number of targets in the heatmap.
[0106] Step 3.2, predict the width and height information of the vehicle target, and calculate the width and height loss.
[0107] Suppose is the k-th bounding box of class c, and the target size is The width and height loss function is L1Loss, and the formula is:
[0108]
[0109] In the formula, N is the number of targets, is the width and height information of the predicted vehicle target, s k is the Ground Truth of the width and height information of the vehicle target.
[0110] Step 3.3, predict the center point offset caused by downsampling, and calculate the center point offset loss.
[0111] The center point offset loss function is L1 Loss, and the specific formula is:
[0112]
[0113] In the formula, N is the number of targets, is the offset information of the center point of the predicted vehicle target, p is the position of the center point of the vehicle target on the Ground Truth.
[0114] Step 4, in the test stage, divide the test video into a sequence of pictures, input it into the above-mentioned converged network model, and generate the detection box of the vehicle target. The specific process is as Figure 4 shown.
[0115] Step 4.1, divide the test video into a sequence of pictures;
[0116] Step 4.2, perform Flip data augmentation on the test data and input it into the trained model together with the original image;
[0117] Step 4.3, during testing, obtain the top 100 peaks of different classes on the heat map through 3x3 MaxPool operation; assume the peak point coordinates are Combine the sizes of the vehicle targets predicted by the other two branches and the center point offset to obtain the detection box with offset. The formula is as follows:
[0118]
[0119] Step 4.4, synthesize the test results of the original image and the test results of the Flip-enhanced image to obtain the final detection box.
Claims
1. An improved CenterNet-based vehicle detection method for traffic surveillance videos, characterized in that, It includes the following steps: Step 1, before the training phase starts, collect experimental data and divide the traffic monitoring video into a sequence of pictures, and preprocess the input pictures to obtain pictures with enhanced data; Step 2, in the training phase, use a deep feature extraction network for the input pictures to extract features at different scales, and fuse the features at different scales extracted; it includes the following steps: Step 2.1, input the pictures into the feature extraction network to obtain features C0, C1, C2, C3, C4; Step 2.2, pass the feature C4 through a dilated encoder to output features with multiple receptive fields, obtaining P4; the dilated encoder is implemented according to the following steps: Step 2.2.a, successively pass the input feature C4 through a 1x1 convolution, a BatchNorm layer, and a Relu layer to reduce the channel dimension; Step 2.2.b, pass the feature obtained in Step 2.2.a through a 3x3 convolution, a BatchNorm layer, and a Relu layer to refine the context semantic features; Step 2.2.c, in order to enable the receptive field of the features output by the encoder to cover targets at various scales, in the present invention, the feature obtained in Step 2.2.b is passed through 4 consecutive stacked residual network structures to obtain the feature P4; among them, the residual network structure consists of 3 consecutive convolutions: the first 1x1 convolution reduces the number of channels at a reduction rate of 4 times, then uses a 3x3 dilated convolution to expand the receptive field, and finally, uses a 1x1 convolution to restore the number of channels; the dilation rates of the 4 consecutive dilated convolutions are [2, 4, 6, 8]; Step 2.3, obtaining feature P4 through deformable convolution and deconvolution of feature P4 ′ , feature C4 undergoes deformable convolution and feature P4 ′ are added to obtain feature P3. Feature P3 further undergoes deformable convolution and deconvolution to obtain feature P3 ′ is added to the C3 feature that has undergone deformable convolution to obtain P2; and so on, until P2 and C1 features are fused to obtain feature P1; Step 2.4, input the features P3, P2, P1 obtained in Step 2.3 into an adaptive spatial feature fusion module to obtain the feature F; Step 3, in the training phase, for the feature map output in Step 2, through the detection head convolutional network branch, generate prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset, and calculate the loss value; Step 4, in the test phase, divide the test video into a sequence of pictures, input them into the above-mentioned network model that has converged, and generate test results.
2. The traffic surveillance video vehicle detection method based on the improved CenterNet according to claim 1, wherein, The picture preprocessing described in Step 1 is carried out according to the following steps: Step 1.1, perform a Random Flipping operation on the input pictures; Step 1.2, perform a Random Scaling operation on the input pictures; Step 1.3, perform a GridMask operation on the input pictures.
3. The vehicle detection method for traffic monitoring video based on the improved CenterNet according to claim 1, characterized in that, The specific method of the different feature fusion methods described in Step 2.3 is as follows: Step 2.3.a, successively pass the feature P4 obtained in Step 2.2.c through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer to further extract features; Step 2.3.b, passing the features obtained in Step 2.3.a through a 4x4 transposed convolution layer, a BatchNorm layer, and a Relu layer in sequence to increase the resolution of the feature map and obtain feature P4 ′ ; Step 2.3.c, successively pass the feature C3 obtained in Step 2.2 through a 3x3 deformable convolution, a BatchNorm layer, and a Relu layer to adjust the number of feature channels; Step 2.3.d, add the feature P4 obtained in Step 2.3.b ′ and the feature obtained in Step 2.3.c, and after performing a summation operation, passing through a 3x3 deformable convolution layer, a BatchNorm layer, and a Relu layer, the fused feature P3 is obtained; Step 2.3.e, similar to the feature fusion methods of Step 2.3.a, Step 2.3.b, Step 2.3.c, and Step 2.3.d, fuse the feature P3 and the feature C2 to obtain the feature P2; Step 2.3.f, similar to the feature fusion methods in Step 2.3.a, Step 2.3.b, Step 2.3.c, and Step 2.3.d, fuses feature P2 and feature C1 to obtain feature P1.
4. The traffic monitoring video vehicle detection method based on the improved CenterNet according to claim 1, characterized in that, The specific adaptive spatial feature fusion method described in Step 2.4 is as follows: Step 2.4.a, adjust the size of feature P3 to be the same as that of feature P1 through a 1x1 convolution and an upsampling layer to obtain feature F3; Step 2.4.b, adjust the size of feature P2 to be the same as that of feature P1 through a 1x1 convolution and an upsampling layer to obtain feature F2; Step 2.4.c, according to the fusion strategy, perform weighted fusion on F3, F2, and P1 to obtain the fused feature F; Among them, the fusion strategy is: F ij = α ij · F3 ij + β ij · F2 ij + γ ij · P1 ij F ij is the value of feature F at position (i, j), F3 ij , F2 ij and P1 ij are the value meanings of features F3, F2, and P1 at position (i, j); α ij , β ij and γ ij are the weights of their respective feature maps at position (i, j), and satisfy the constraints of α ij + β ij + γ ij = 1 and α ij , β ij , γ ij ∈ [0, 1], and this constraint is achieved by performing a 1x1 convolution on F3 ij , F2 ij and P1 ij followed by a softmax calculation.
5. The vehicle detection method for traffic monitoring video based on improved CenterNet according to claim 1, characterized in that The specific method for using the detection head convolution branch network described in Step 3 to generate prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset, and calculate the loss value is as follows: Step 3.1, the detection head branch 1 predicts the heat map of the vehicle target on feature F and calculates the center point loss; Step 3.2, the detection head branch 2 predicts the width and height information of the vehicle target on feature F and calculates the width and height loss; Step 3.3, the detection head branch 3 predicts the center point offset caused by downsampling on feature F and calculates the center point offset loss.
6. The traffic surveillance video vehicle detection method based on the improved CenterNet according to claim 1, characterized in that, The specific method for generating the test result described in Step 4 is: Step 4.1, divide the test video into a sequence of pictures; Step 4.2, perform Flip data augmentation on the test data and input it into the trained model together with the original image; Step 4.3, during testing, according to Step 3, generate prediction results of the vehicle center point, vehicle width and height information, and vehicle center point offset, calculate the loss value, and obtain the detection box with offset; Step 4.4, comprehensively combine the test results of the original image and the test results of the Flip-enhanced image to obtain the final detection box.
7. A vehicle detection system for the traffic monitoring video vehicle detection method based on the improved CenterNet according to claim 1, characterized in that Including: A data processing module, which is used to divide the input traffic monitoring video into an image sequence, preprocess the input image, expand the data set, and at the same time simulate the situation where the vehicle is blocked to enhance the robustness of the model; A feature extraction module, which is used to extract multi-scale features from the data processed by the data processing module; A feature fusion module, first uses a dilated encoder on the deepest feature to expand the receptive field; then, fuses the feature after passing through the dilated convolutional layer with the shallow features extracted by the feature extraction module to improve the detection accuracy of the model for targets of different scales; finally, uses an adaptive spatial feature fusion module on the features of the last three layers to filter out the useless information of the last three layers in the space and only retain the useful information for fusion; A detection module, which is used to predict the vehicle center point, vehicle width and height information, and vehicle center point offset, and generate the final vehicle detection box.
8. A vehicle detection device for the traffic monitoring video vehicle detection method based on the improved CenterNet according to claim 1, characterized in that Including: A video image collector, which is used to collect traffic monitoring video images; A program processor, which is used to store a computer program and execute the computer program to implement the traffic monitoring video vehicle detection method based on the improved CenterNet described in Steps 1 to 4; A display, which is used to display the traffic monitoring video vehicle detection results.
Citation Information
Patent Citations
Vehicle detection method in aerial images of unmanned aerial vehicle based on deep learning
CN110298227A
Rapid vehicle detection method based on deep learning
CN111461083A
Road vehicle detection system and method based on YOLOv4
CN114821492A