A traffic flow statistics method based on improved YOLO V7 and Deep-Sort
By improving the YOLOv7 model and Deep-Sort algorithm, combining feature enhancement and region of interest technology, the accuracy and efficiency of vehicle detection and tracking in complex traffic environments are solved, and efficient traffic flow statistics are achieved.
Patent Information
- Application Number
- CN202310392311.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-13
AI Technical Summary
The existing traffic flow statistics methods are difficult to accurately detect vehicles and effectively track them in complex traffic environments, resulting in low detection efficiency, high resource consumption and low real-time performance.
The improved YOLOv7 model is used to add the SE-Net module after feature extraction, combined with the Deep-Sort algorithm, by adding the area of interest to the detection video, focusing on vehicle tracking on the main roads, ignoring the interference on both sides of the city streets, and using statistical methods based on motion trajectory and detection lines.
Improves the accuracy and tracking speed of vehicle detection, ensuring the accuracy and real-time traffic counting in complex contexts.
Smart Images

Figure CN116434159B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic flow statistics, and particularly relates to a traffic flow statistics method based on improved YOLO V7 and Deep-Sort. Background Art
[0002] In the intelligent transportation system, traffic flow statistics based on computer vision using surveillance videos is a research field that has received much attention. It helps traffic management departments understand the traffic flow on the road in a timely manner, rationally allocate traffic resources, improve the traffic efficiency of roads, effectively prevent and respond to urban traffic congestion problems, and provides strong support for urban traffic management. Traffic flow statistics generally includes two parts: vehicle target detection and tracking.
[0003] In a complex traffic environment, how to accurately detect vehicles is the primary condition for ensuring the accuracy rate of traffic flow statistics. Early motion target detection algorithms based on vision technology mainly include background subtraction method, frame difference method, and optical flow method. For example, in the invention patent of a road monitoring traffic flow simulation method, device, equipment, and medium based on machine learning, the frame difference method is used to separate the moving foreground and background of each frame of image, and according to the pixel points corresponding to the moving foreground in each frame of image, it is judged whether there are moving vehicles in the surveillance video. However, due to the defects of its principle, it cannot extract the complete area of the object, but only the boundary; at the same time, it depends on the selected inter-frame time interval, resulting in such methods generally having disadvantages such as low real-time performance, low detection efficiency, and large resource consumption.
[0004] In addition to detecting the target vehicle, it is also necessary to track the identified vehicle by using a target tracking algorithm to establish the connection of the target in adjacent frames of the video and ensure the accuracy of counting. According to the order of generating the target trajectory, target tracking can be divided into online tracking and offline tracking, and the main difference between the two lies in the data processing method. Online tracking needs to use the target information of the current frame and all previous frames to calculate the matching degree between the target and the existing trajectory, so it has better real-time performance; while offline tracking uses the target information of the entire video image for processing, so it has higher accuracy. Therefore, it is necessary to select a suitable target tracking algorithm according to the requirements of the application scenario. Summary of the Invention
[0005] In order to overcome the defects existing in the above prior art, the purpose of the present invention is to provide a traffic flow statistics method based on improved YOLOV7 and Deep-Sort. The backbone network of YOLOv7 is improved, and an attention mechanism SE-Net module that can enhance the important information channels in the feature map is added after each feature extraction to improve the detection accuracy of target vehicles. And it is proposed to add regions of interest in the detection video, focus on tracking the vehicles on the main road, reduce the number of target detection frames by ignoring the sidewalks and ramps on both sides of the urban streets, and improve the vehicle tracking speed.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A traffic flow statistics method based on improved YOLO V7 and Deep-Sort, comprising the following steps;
[0008] (1) Prepare a vehicle dataset;
[0009] (2) Build an improved YOLOv7 model and use the vehicle dataset for training and detection; the improved YOLOv7 model refers to adding an SE-Net module after each feature extraction network of its backbone network on the basis of the original YOLOv7 model;
[0010] (3) Build a Deep-Sort model to track the detected vehicles; the Deep-Sort model includes a target detection module, a position prediction module, a feature matching module and an update module;
[0011] (4) Use a traffic flow statistics method based on motion trajectories and detection lines to obtain a traffic monitoring video, draw a virtual detection line, set regions of interest, and enter the detection process and tracking process, thereby completing traffic flow statistics.
[0012] In the step (1), the vehicle dataset required for training includes; UA-DETRAC dataset and self-made dataset; they are divided into a training set and a test set according to a ratio;
[0013] First, convert the xml format of the UA-DETRAC dataset into the xml format of the VOC dataset, and then convert the xml format of the VOC dataset into the txt format of the YOLOv7 dataset to complete the conversion of the dataset format;
[0014] The self-made dataset consists of multiple videos taken, with scenarios including the driving conditions of road vehicles during peak hours, off-peak hours, and night hours. Each folder contains a sequence of images taken every 5 frames from a video, which is used to effectively reduce the similarity of the images, prevent the training network from being overly redundant, and use the LabelImg tool to annotate the collected images. They are divided into a training set and a test set according to a certain proportion. The UA-DETRAC dataset and the self-made dataset are combined to finally obtain a preprocessed vehicle dataset.
[0015] The specific content of step (2) is as follows:
[0016] 1) Build the input end of the improved YOLOv7 model, including:
[0017] (1) Mosaic data augmentation: Take every four images from the image sequence in step (1) as a group, and splice them together in one image through flipping, scaling, and color gamut changes within the region;
[0018] (2) Adaptive image scaling: Specify that the size of the image for training is 640×640, and scale the length x and width y; calculate the sizes of the scaled x and y, which are represented as x1 and y1 respectively, where x1 = x × min{x / 640, y / 640}, y1 = y × min{x / 640, y / 640}; if x1 < 640, then add black borders with a height of [(640 - x1) % 64] / 2 to the upper and lower heights of the corresponding x, and finally make up an image of 640×640 size; perform the same operation in the y direction, where the min operation represents taking the minimum value within the curly brackets, and the % operation represents taking the remainder operation;
[0019] 2) Build the feature extraction network of the improved YOLOv7 model, including:
[0020] Introduce the SE-Net module to improve the feature extraction network of YOLOv7:
[0021] The SE-Net module introduces an attention mechanism using the channel dimension attribute, enabling the network model to automatically learn features and dynamically obtain the weights of each feature channel; the SE-Net module sequentially performs squeezing operation, activation operation, and weighting operation on the input feature map, automatically obtains the importance of each feature channel through learning, and enhances useful features and suppresses features that are not very useful for the current task according to this importance, and finally outputs a feature vector with multiple feature channels to improve the network expression ability;
[0022] The feature extraction network of the YOLOv7 model includes the CBS module, E-ELAN module, MP Conv module, and SPPCSPC module;
[0023] Among them, the CBS module includes a convolutional layer, a normalization layer, and the activation function SiLU. There are a total of three CBS modules, which are used to change the number of channels, feature extraction, and downsampling respectively; the tasks of the E-ELAN module and the MP Conv module in the backbone network are to aggregate images. The E-ELAN module uses an attention network to control the longest and shortest paths of the gradient, and performs operations such as expansion, shuffling, and merging elements to increase the network depth and prevent phenomena such as information loss and information over-expansion; the MP Conv module is used to expand the receptive field of the current feature layer, and then fuse it with the feature information processed by normal convolution to improve the generalization performance of the network; the SPPCSPC module obtains different receptive fields through max pooling, solving the problem of image distortion caused by scaling and cropping operations.
[0024] The E-ELAN module and the MP Conv module are the key factors determining the network's feature extraction ability, while the SPPCSPC module plays an auxiliary role. Embed three SE-Net modules between the E-ELAN module and the MP Conv module. The feature map output by the E-ELAN module is used as the input of the SE-Net module, and the feature map output by the SE-Net module is used as the input of the MP Conv module. The last SE-Net is embedded between the E-ELAN module and the SPPCSPC module. The feature map output by the E-ELAN module is used as the input of the SE-Net module, and the feature map output by the SE-Net module is used as the input of the SPPCSPC module, finally obtaining the feature extraction network of the improved YOLOv7 model; it can effectively improve the effective flow of feature information in the network, focus on useful features while suppressing useless features, so as to enhance the network's detection ability for vehicle targets of different scaling scales and improve the detection accuracy of the network.
[0025] And add a four-fold downsampling process. The feature map with a size of 640*640 after being processed at the input end undergoes a four-fold downsampling operation to obtain a feature map with a size of (640 / 4)*(640 / 4), that is, a size of 160*160. Since its number of layers is relatively shallow and the receptive field is small, the features contained in this feature map tend to be local and detailed, which can effectively improve the detection effect of the network for occluded vehicles.
[0026] 3) Build the feature fusion network of the improved YOLOv7 model, including:
[0027] Adopt the FPN and PAN structures to fuse the features output by the feature extraction network of the improved YOLOv7 model to obtain the feature fusion network of the improved YOLOv7 model.
[0028] First, feature extraction is performed on the 640*640-sized image generated at the input end to obtain feature maps of 160*160, 80*80, 40*40, and 20*20; the FPN network transmits the semantic information of the feature maps from high dimensions to low dimensions, performs multiple upsamplings and channel concatenations to generate a feature map containing the semantic information of the vehicle target, the PAN network transmits the semantic information from low dimensions to high dimensions one more time, performs multiple downsamplings and channel concatenations to generate a feature map containing the position information of the vehicle target, and finally fuses the feature maps generated by the two networks, so that feature maps of different sizes all contain image semantic information and image feature information, ensuring accurate prediction of images of different sizes;
[0029] 4) Build the output end of the improved YOLOv7 model, including:
[0030] The output end of YOLOv7 includes confidence loss, localization loss, and classification loss; the confidence loss is used to calculate the credibility of the prediction box, the localization loss is used for the error between the prediction box and the calibration box, and the classification loss is used to calculate whether the anchor box and the corresponding calibrated classification are correct; the output of YOLOv7 is not limited to single output, and by introducing an auxiliary head (auxiliaryHead) to perform auxiliary training on the intermediate layer and perform deep supervision on the training of the model, the overall performance of the model is improved;
[0031] In actual detection, first, the prediction confidence of each prediction box is judged. If it exceeds the set threshold, it is considered that there is a target in the prediction box and its approximate position is determined. Then, the non-maximum suppression algorithm is used to screen the prediction boxes with targets and remove the duplicate detection boxes of the same target. Finally, according to the classification probability of the screened prediction boxes, the index corresponding to the maximum probability is taken as the classification index number of the target to obtain the category of the target.
[0032] The specific content of step (3) is as follows:
[0033] 1) Use the improved YOLOv7 model as the target detection module of the Deep-Sort model;
[0034] 2) The position prediction module uses the Kalman filter algorithm to predict the position information of the vehicle at the next moment: when the vehicle moves, according to the speed and position information of the vehicle in the previous frame, the speed and position information of the vehicle in the current frame are predicted. The position coordinates (x', y') of the vehicle are: x' = x + w / 2, y' = y, where (x, y, w, h) are the vehicle frame diagram information recognized by the improved YOLOv7 model, (x, y) are the coordinates of the lower left corner of the vehicle frame diagram, and (w, h) are the width and height of the vehicle frame respectively;
[0035] 3) The feature matching module uses the improved YOLOv7 model to obtain the position information of each vehicle in the video frame at the next moment, and then associates the detected vehicle information with the vehicle information obtained from the Deep-Sort prediction and tracking part; among them, the Hungarian algorithm is used for data association, and the cost matrix is constructed through the appearance information distance of the vehicle and the Mahalanobis distance of the vehicle position, and the optimal vehicle matching scheme is calculated. The Mahalanobis distance of the vehicle position is:
[0036] d (1) (i,j) = (d j -y i ) T S -1 (d j -y i )
[0037] In the formula, i is the serial number of the predicted and tracked vehicle frame, j is the serial number of the detected vehicle frame, d and y are the distributions of the detected vehicle and the predicted and tracked vehicle respectively, and S is the covariance matrix between the two distributions;
[0038] The vehicle appearance information distance is to use a ReID network trained offline on a vehicle re-identification dataset to extract the appearance feature description vector from the vehicle picture, and then for each predicted and tracked vehicle, retain the last 100 sets of appearance feature descriptors R associated successfully with the detection frame and calculate their minimum cosine distance from the detected vehicle frame:
[0039]
[0040] In the formula, is the appearance feature of the j-th detected vehicle, is the k-th appearance feature of the i-th predicted and tracked vehicle, and R i is the appearance feature set of the i-th predicted and tracked vehicle;
[0041] The total cost matrix c i,j is the weighted result of the Mahalanobis distance d (1) (i,j) of the vehicle position and the appearance information distance d (2) (i,j) of the vehicle appearance:
[0042] c i,j = αd (1) (i,j) + (1 - α)d (2) (i,j)
[0043] In the formula, α is the weighting ratio.
[0044] 4) The update module transfers the vehicle ID information at this moment to the corresponding vehicle at the next moment according to the optimal matching scheme of the vehicle in the matching part, and then continues to repeat the prediction, matching, and update processes of Deep-Sort.
[0045] The specific steps of step (4) are as follows:
[0046] 1) Obtain traffic surveillance videos of highways and urban environments respectively, and loop through each frame of the video to read the pictures;
[0047] 2) Select a place closer to the shooting angle of the surveillance camera to draw a virtual detection line, maximizing the capture of the feature information of the target vehicle; improve the reliability of traffic flow statistics; in the actual traffic road scene, the camera is placed at a relatively high position from the road surface, with a large shooting range, complex background, and blurred and highly overlapping distant vehicles. If the virtual detection line is placed far from the surveillance camera, the detected vehicle target size is small, which causes difficulties in traffic flow statistics;
[0048] 3) Set the region of interest: In the traffic surveillance video of the urban environment with a complex background, since there are often parked vehicles by the roadside, many vehicle bounding boxes appear outside the area where non-traffic flow is detected, resulting in a slower vehicle tracking speed in the target area. Adding the region of interest focuses on tracking the vehicles on the main road, ignoring the sidewalks and ramps on both sides of the urban streets, reducing the number of vehicle bounding boxes, thereby solving the problem of the slower vehicle tracking speed in the target area when the Deep-Sort object tracking algorithm is used for traffic flow detection and tracking on urban roads with a complex background.
[0049] The specific steps are as follows: First, determine the coordinates of the two outermost lanes in the video and connect them into two straight lines to form two regions that fit the urban traffic lanes. The selected region is the region of interest, and traffic flow statistics are only implemented in this region;
[0050] 4) Combine the improved YOLO-V7 model and the Deep-Sort algorithm: YOLO-V7 identifies vehicle targets to obtain the detection box information of the targets and imports it into the Deep-Sort framework to generate the tracking bounding box of the target vehicle;
[0051] 5) According to the bounding boxes generated in 4), determine the center point of the target vehicle bounding box in each frame. When the center point crosses the virtual detection line set in 2), the traffic counter is incremented by 1 to complete the counting of the traffic flow on the road during the specified period.
[0052] The beneficial effects of the present invention:
[0053] First: By adding an SE-Net module, an attention mechanism that can enhance the important information channels in the feature map, after each feature extraction in the YOLOv7 backbone network, the detection accuracy when there is overlap or occlusion between vehicle targets is improved, meeting the performance requirements of vehicle tracking and statistical algorithms.
[0054] Second: By adding regions of interest in the detection video, it focuses on tracking vehicles on the main road while ignoring the sidewalks and ramps on both sides of the urban streets, reducing the number of target detection boxes, thus solving the problem that the vehicle tracking speed in the target area slows down when the Deep-Sort target tracking algorithm is used for traffic flow detection and tracking on urban roads under complex backgrounds.
[0055] Third: Through a statistical method based on motion trajectories and detection lines, the real-time performance of vehicle detection at high speeds is ensured and the accuracy of traffic flow counting under complex backgrounds is met. Description of the Drawings
[0056] Figure 1 is the traffic flow statistical flowchart based on the improved YOLOv7 and Deep-Sort in the present invention.
[0057] Figure 2 is the schematic diagram of the model structure of the YOLOv7 network in the present invention.
[0058] Figure 3 is the schematic diagram of the model structure of the improved YOLOv7 network in the present invention.
[0059] Figure 4 is the schematic diagram of the designed virtual detection line in the present invention.
[0060] Figure 5 is the schematic diagram of the traffic flow counting process in the present invention.
[0061] Figure 6 is the comparison of traffic flow statistical counting between the present invention and existing detection algorithms in different scenarios. Detailed Embodiments
[0062] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0063] As Figure 1 shown: A traffic flow statistical method based on the improved YOLO V7 and Deep-Sort includes the following specific steps;
[0064] Step 1: Prepare the vehicle data set required for training;
[0065] (1) The vehicle datasets used are the UA-DETRAC dataset and the self-made dataset. The UA-DETRAC dataset was filmed in different locations in two cities, Beijing and Tianjin, and was divided into a training set and a test set in the ratio of 8:2. Since this patent uses the YOLOv7 algorithm and requires the dataset format to be txt, while the UA-DETRAC dataset is in xml format. First, convert the xml format of the UA-DETRAC dataset to the xml format of the VOC dataset, and then convert the xml format of the OC dataset to the txt format of the YOLOv7 dataset to complete the conversion of the dataset format. The self-made dataset is composed of multiple videos filmed at the overpass on Qintang Avenue in Lintong District, Xi'an and the Taibai Impression City Overpass on the west section of the Second Ring Road in Yanta District. The scenarios include the driving conditions of road vehicles during peak hours, off-peak hours, and night hours, improving the generalizability of the dataset. Each folder contains a sequence of pictures taken every 5 frames from a video clip, which can effectively reduce the similarity of the pictures and prevent the training network from being too redundant. And use the LabelImg tool to annotate the collected pictures, and divide them into a training set and a test set in the ratio of 8:2. Combine the UA-DETRAC dataset and the self-made dataset to finally obtain the preprocessed vehicle dataset;
[0066] Step 2: Build and train an improved YOLOv7 model for vehicle detection: Based on the YOLOv7 model, improve it to address the problem of low detection accuracy, and obtain an improved YOLOv7 model. Its structure includes an input end, a feature extraction network, a feature fusion network, and an output end;
[0067] (2.1) Build the input end of the improved YOLOv7 model, including:
[0068] 1) Mosaic data augmentation: Stitch four pictures together in one picture through flipping, scaling, and gamut changes within the region;
[0069] 2) Adaptive image scaling: Specify that the size of the pictures for training is 640×640, and scale the length x and width y; Calculate the sizes of the scaled x and y, denoted as x1 and y1 respectively, where x1 = x×min{x / 640, y / 640}, y1 = y×min{x / 640, y / 640}; If x1 < 640, add black borders with a height of [(640 - x1) % 64] / 2 to the top and bottom of the corresponding x height, and finally make up a picture of 640×640 size; Similarly for the y direction operation, where the min operation represents taking the minimum value within the curly brackets, and % represents the remainder operation;
[0070] (2.2) Build the feature extraction network of the improved YOLOv7 model, including:
[0071] Introducing the SE-Net attention mechanism to improve the feature extraction network of YOLOv7: The feature extraction network of the YOLOv7 model mainly includes the CBS module, E-ELAN module, MP Conv module, and SPPCSPC module, as Figure 2 shown. Among them, the CBS module is mainly composed of a convolutional layer, a normalization layer, and the activation function SiLU, and there are a total of three CBS modules, which are used to change the number of channels, extract features, and downsample respectively; the tasks of the E-ELAN module and the MP Conv module in the backbone network are to aggregate images. The E-ELAN module uses an attention network to control the longest and shortest paths of the gradient, and performs operations such as expansion (Expand), shuffling (Shuffle), and merging elements (Merge cardinality) to increase the network depth and prevent phenomena such as information loss and information over-inflation; the role of the MP Conv module is to expand the receptive field of the current feature layer, and then fuse it with the feature information processed by normal convolution to improve the generalization performance of the network; the SPPCSPC module obtains different receptive fields through max pooling, solving the problem of image distortion caused by scaling and cropping operations.
[0072] From the above analysis, it can be seen that the E-ELAN module and the MP Conv module are the key factors determining the network's feature extraction ability, while the SPPCSPC module plays an auxiliary role. Therefore, if we want to optimize the algorithm process of the entire network, we can embed the SE-Net module between the E-ELAN module and the MP Conv module, as Figure 3 shown, to improve the effective flow of feature information in the network, while focusing on useful features and suppressing useless features, so as to enhance the network's detection ability for vehicle targets at different scaling scales and improve the detection accuracy of the network;
[0073] (2.3) Build the feature fusion network of the improved YOLOv7 model, including:
[0074] Adopt the FPN and PAN structures to fuse the features output by the feature extraction network: FPN transfers semantic information from high dimensions to low dimensions, so that the underlying feature map contains more semantic information of vehicle targets; while the PAN structure transfers the semantic information from low dimensions to high dimensions again, so that the top-level features contain more position information of vehicle targets;
[0075] (2.4) Build the output end of the improved YOLOv7 model, including:
[0076] The original output end of the YOLOv7 model predicts results by outputting three feature maps of different sizes. The improved YOLOv7 adds a four-fold downsampling process to obtain a feature map with the largest size. Since its number of layers is relatively shallow and the receptive field is small, the features contained in this feature map tend to be local and detailed, which can effectively improve the detection effect of the network for occluded vehicles.
[0077] Step 3: Build the Deep-Sort model for vehicle tracking: The Deep-Sort model includes an object detection module, a position prediction module, a feature matching module, and an update module;
[0078] (3.1) Use the improved YOLOv7 model as the object detection module of the Deep-Sort model;
[0079] (3.2) The position prediction module mainly uses the Kalman filter algorithm to predict the position information of the vehicle at the next moment: When the vehicle moves, based on the speed and position information of the vehicle in the previous frame, the speed and position information of the vehicle in the current frame are predicted. The position coordinates (x', y') of the vehicle are: x' = x + w / 2, y' = y, where (x, y, w, h) are the vehicle bounding box information recognized by the improved YOLOv7 model, (x, y) are the coordinates of the lower left corner of the vehicle bounding box, and (w, h) are the width and height of the vehicle bounding box respectively.
[0080] (3.3) The feature matching module uses the improved YOLOv7 model to obtain the position information of each vehicle in the video frame at the next moment, and then associates the detected vehicle information with the vehicle information obtained from the prediction and tracking part of Deep-Sort. Among them, the Hungarian algorithm is used for data association, and the cost matrix is constructed through the appearance information distance of the vehicle and the Mahalanobis distance of the vehicle position, and the optimal vehicle matching scheme is calculated. The Mahalanobis distance of the vehicle position is:
[0081] d (1) (i,j) = (d j -y i ) T S -1 (d j -y i )
[0082] In the formula, i is the serial number of the predicted and tracked vehicle bounding box, j is the serial number of the detected vehicle bounding box, d and y are the distributions of the detected vehicle and the predicted and tracked vehicle respectively, and S is the covariance matrix between the two distributions;
[0083] The vehicle appearance information distance uses a ReID network pre-trained offline on a vehicle re-identification dataset to extract appearance feature description vectors from vehicle images. Then, for each predicted and tracked vehicle, the last 100 sets of appearance feature descriptors R successfully associated with the detection box are retained, and the minimum cosine distance between them and the detected vehicle box is calculated:
[0084] d (2) (i,j)=min{1-r j T r k (i) |r k (i) ∈R i}
[0085] In the formula, is the appearance feature of the j-th detected vehicle, is the k-th appearance feature of the i-th predicted and tracked vehicle, and R i is the appearance feature set of the i-th predicted and tracked vehicle;
[0086] The total cost matrix c i,j is the weighted result of the Mahalanobis distance d (1) (i,j) of the vehicle position and the vehicle appearance information distance d (2) (i,j):
[0087] c i,j =αd (1) (i,j)+(1-α)d (2) (i,j)
[0088] In the formula, α is the weighting ratio.
[0089] (3.4) The update module transfers the vehicle ID information at this moment to the corresponding vehicle at the next moment according to the optimal matching scheme of the vehicles in the matching part. Then, the prediction, matching, and update processes of Deep-Sort are continued.
[0090] Step 4: Use the traffic flow statistics method based on the motion trajectory and the detection line. Its algorithm process includes: obtaining the traffic monitoring video, drawing the virtual detection line, setting the region of interest, entering the detection process and the tracking process, and completing the traffic flow statistics.
[0091] (4.1) Obtain the traffic monitoring videos of highways and urban environments respectively, and read each frame of the video in a loop;
[0092] (4.2) Draw the virtual detection line at a location close to the surveillance camera: In actual traffic road scenes, the camera is usually placed at a high altitude from the road surface, with a large shooting range, complex background, and blurred and overlapping display of distant vehicles. If the virtual detection line is placed far away from the surveillance camera, the detected vehicle target size will be small, which will cause difficulties in traffic flow statistics. Set the virtual detection line at a location close to the surveillance camera to maximize the capture of the target vehicle's feature information and improve the reliability of traffic flow statistics. Figure 4 As shown;
[0093] (4.3) Setting the region of interest: In traffic monitoring videos in urban environments with complex backgrounds, since there are often parked vehicles on the roadside, many vehicle blocks appear outside the non-traffic flow detection area, resulting in a slowdown in vehicle tracking in the target area. Adding a region of interest focuses on tracking vehicles on the main roads, ignoring the sidewalks and gates on both sides of the city streets, and reducing the number of vehicle blocks, thereby solving the problem of slow tracking of vehicles in the target area when the Deep-Sort target tracking algorithm performs traffic flow detection and tracking on urban roads with complex backgrounds. The specific steps are: first determine the coordinates of the two outer lanes in the video and connect them into two straight lines to form two sections that fit the urban traffic lanes. The selected area is the region of interest, and traffic flow statistics are only implemented in this area.
[0094] (4.4) Combine the improved YOLO-V7 model with the Deep-Sort algorithm: YOLO-V7 identifies the vehicle target to obtain the detection box information of the target object, and imports it into the Deep-Sort framework to generate the target vehicle tracking block diagram.
[0095] (4.5) According to the block diagram generated by (4.4), the center point of the target vehicle block diagram in each frame is determined. When the center point crosses the virtual detection line set by (4.2), the vehicle counter is incremented by 1 to complete the counting of the vehicle flow on the road during the specified period. Figure 5 shown.
[0096] (6) Output the traffic flow statistics in the traffic monitoring video.
[0097] The effects of the present invention are further described below in conjunction with simulation results.
[0098] 1. The simulation conditions are shown in Table 1;
[0099] Table 1 Experimental hardware and software parameters
[0100]
[0101] 2.Simulation content and result analysis;
[0102] Simulation 1. To verify the superiority of the improved YOLOv7 vehicle detection algorithm, three mainstream object detection algorithms were selected in this paper for comparative experiments to study the mean average precision (mAP), detection rate (FPS), number of parameters (Params), and computational complexity (GFLOPs) of the YOLOv7 vehicle detection algorithm with the addition of the SE-Net attention mechanism. The results are shown in Table 2.
[0103] Table 2 YOLO Object Detection Performance Analysis
[0104]
[0105] As can be seen from Table 1, compared with YOLOv4, mAP@0.5 it increased by 7.6%, Params and GFLOPs decreased by 26.8M and 37.3G respectively, and FPS increased by 9; compared with YOLOv5L, mAP@0.5 it increased by 4.4%, Params and GFLOPs decreased by 8.9M and 3.6G respectively, and FPS increased by 5; compared with YOLOv7, mAP@0.5 it increased by 1.3%. Since four SE-Net attention mechanism networks were added to the backbone network, the Params and GFLOPs of the improved network increased slightly, resulting in a small decrease in the detection speed, but the impact is not significant. Compared with the other two algorithms, the model size and computational complexity of the improved network have been greatly reduced, thus improving the speed of the algorithm.
[0106] Simulation 2 studied the real-time detection of vehicles using the improved YOLOv7 network, as well as the tracking and statistical effects of vehicles in high-speed and urban environments using the traffic flow statistical algorithms of the improved YOLOv7 and Deep-Sort. Two traffic surveillance videos were selected for this experiment, namely the highway surveillance video and the urban traffic surveillance video. Figure 6 (a) and Figure 6 (b) are the video capture frames using the two detection algorithms in the highway environment respectively, Figure 6 (c) and Figure 6 (d) are the video capture frames using the two detection algorithms in the urban traffic environment respectively.
[0107] From Figure 6 it can be seen that using YOLOv7 with the added attention mechanism SE-Net has a certain degree of improvement in the statistical accuracy compared with YOLOv7, which can enhance the target extraction ability of vehicle detection, reduce the occurrence of missed detections or false detections, and overall improve the performance of YOLOv7 vehicle detection.
[0108] For Figure 6The comparison method in is used to statistically analyze the real traffic flow, track the traffic flow and accuracy rate, and the results are represented by 0.
[0109] Table 3 Statistical results of traffic flow under highways and urban streets
[0110]
Claims
1. A traffic flow statistics method based on improved YOLO V7 and Deep-Sort, characterized in that, Including the following steps; Step (1): Prepare a vehicle dataset; Step (2): Build an improved YOLOv7 model and use the vehicle dataset for training and detection; the improved YOLOv7 model refers to adding an SE-Net module after each feature extraction network in the backbone network on the basis of the original YOLOv7 model; Step (3): Build a Deep-Sort model to track the detected vehicles; the Deep-Sort model includes an object detection module, a position prediction module, a feature matching module, and an update module; Step (4): Use a traffic flow statistics method based on the motion trajectory and the detection line to obtain a traffic surveillance video, draw a virtual detection line, set an area of interest, and enter the detection process and the tracking process, thereby completing the traffic flow statistics; The specific content of step (2) is as follows: 1) Build the input end of the improved YOLOv7 model, including: (1) Mosaic data augmentation: Take every four pictures in the picture sequence in step (1) as a group, and splice them together in one picture through flipping, scaling, and color gamut change within the area; (2) Adaptive picture scaling: Specify that the size of the picture for training is 640×640, and scale the length x and width y; calculate the sizes of the scaled x and y, which are respectively represented as x1 and y1, where x1 = x×min{x / 640, y / 640}, y1 = y×min{x / 640, y / 640}; if x1 < 640, add black edges with a height of [(640 - x1) % 64] / 2 above and below the corresponding x height, and finally make up a picture of 640×640 size; the same operation is performed in the y direction, where the min operation represents taking the minimum value within the curly brackets, and the % operation represents taking the remainder operation; 2) Build the feature extraction network of the improved YOLOv7 model, including: Embed three SE-Net modules between the E-ELAN module and the MP Conv module. The feature map output by the E-ELAN module is used as the input of the SE-Net module, and the feature map output by the SE-Net module is used as the input of the MP Conv module. The last SE-Net is embedded between the E-ELAN module and the SPPCSPC module. The feature map output by the E-ELAN module is used as the input of the SE-Net module, and the feature map output by the SE-Net module is used as the input of the SPPCSPC module, and finally obtain the feature extraction network of the improved YOLOv7 model; 3) Build the feature fusion network of the improved YOLOv7 model, including: Adopt the FPN and PAN structures to fuse the features output by the feature extraction network of the improved YOLOv7 model to obtain the feature fusion network of the improved YOLOv7 model; First, feature extraction is performed on the 640*640-sized image generated at the input end to obtain feature maps of 160*160, 80*80, 40*40, and 20*20; the FPN network transfers the semantic information of the feature maps from high dimensions to low dimensions, performs multiple upsamplings and channel concatenations to generate a feature map containing the semantic information of the vehicle target, the PAN network transfers the semantic information from low dimensions to high dimensions once again, performs multiple downsamplings and channel concatenations to generate a feature map containing the position information of the vehicle target, and finally fuses the feature maps generated by the two networks; 4) Build the output end of the improved YOLOv7 model, including: The output end of YOLOv7 includes confidence loss, localization loss, and classification loss; the confidence loss is used to calculate the credibility of the prediction box, the localization loss is used for the error between the prediction box and the calibration box, and the classification loss is used to calculate whether the anchor box and the corresponding calibrated classification are correct; the output of YOLOv7 is not limited to single output, and auxiliary training is performed on the intermediate layer by introducing an auxiliary head to perform deep supervision on the training of the model; In actual detection, first judge the prediction confidence of each prediction box. If it exceeds the set threshold, it is considered that there is a target in the prediction box and its approximate position is determined. Then, use the non-maximum suppression algorithm to screen the prediction boxes with targets and remove the duplicate detection boxes of the same target. Finally, take the index corresponding to the maximum probability according to the classification probability of the screened prediction box as the classification index number of the target to obtain the category of the target.
2. The traffic flow statistics method based on improved YOLO V7 and Deep-Sort according to claim 1, characterized in that In the step (1), the vehicle data set required for training includes; UA-DETRAC data set and self-made data set; they are divided into training set and test set according to a certain proportion; First, convert the xml format of the UA-DETRAC data set into the xml format of the VOC data set, and then convert the xml format of the VOC data set into the txt format of the YOLOv7 data set to complete the conversion of the data set format; The self-made data set is multiple segments of videos taken, and the scenarios include the driving conditions of road vehicles during peak hours, off-peak hours, and night hours. Each folder contains a sequence of images taken every 5 frames intercepted from a video, and the collected images are labeled using the LabelImg tool, divided into training set and test set according to a certain proportion, and the UA-DETRAC data set and the self-made data set are combined to finally obtain the preprocessed vehicle data set.
3. A traffic flow statistics method based on improved YOLO V7 and Deep-Sort according to claim 1, characterized in that The step (3) is specifically: 1) Use the improved YOLOv7 model as the target detection module of the Deep-Sort model; 2) The position prediction module uses the Kalman filter algorithm to predict the position information of the vehicle at the next moment: when the vehicle moves, according to the speed and position information of the vehicle in the previous frame, predict the speed and position information of the vehicle in the current frame. The position coordinates (x', y') of the vehicle are: x' = x + w / 2, y' = y, where (x, y, w, h) is the vehicle frame diagram information recognized by the improved YOLOv7 model, (x, y) is the lower left corner coordinates of the vehicle frame diagram, and (w, h) are the width and height of the vehicle frame respectively; 3) The feature matching module uses the improved YOLOv7 model to obtain the position information of each vehicle in the next video frame, and then associates the detected vehicle information with the vehicle information obtained from the Deep-Sort prediction and tracking part; the Hungarian algorithm is used for data association, and the cost matrix is constructed through the appearance information distance of the vehicle and the Mahalanobis distance of the vehicle position, and the optimal vehicle matching scheme is calculated. The Mahalanobis distance of the vehicle position is: d (1) (i,j) = (d j -y i ) T S -1 (d j -y i ) In the formula, i is the serial number of the predicted and tracked vehicle box, j is the serial number of the detected vehicle box, d and y are the distributions of the detected vehicle and the predicted and tracked vehicle respectively, and S is the covariance matrix between the two distributions; The vehicle appearance information distance is to use a ReID network offline trained on a vehicle re-identification dataset to extract the appearance feature description vector from the vehicle picture. Then, for each predicted and tracked vehicle, the last 100 sets of appearance feature descriptors R associated successfully with the detection box are retained, and the minimum cosine distance between them and the detected vehicle box is calculated: In the formula, is the appearance feature of the j-th detected vehicle, is the k-th appearance feature of the i-th predicted tracking vehicle, and R i is the appearance feature set of the i-th predicted tracking vehicle; The total cost matrix c i,j is the Mahalanobis distance d of the vehicle position (1) (i, j) and the distance d of the vehicle appearance information (2) The weighted result of (i, j): c i,j = αd (1) (i, j) + (1 - α)d (2) (i, j) In the formula, α is the weighting ratio; 4) The update module transfers the vehicle ID information at this moment to the corresponding vehicle at the next moment according to the optimal vehicle matching scheme in the matching part, and then continues to repeat the prediction, matching and update processes of Deep-Sort.
4. A traffic flow statistics method based on improved YOLO V7 and Deep-Sort according to claim 1, characterized in that The specific steps of step (4) are as follows: 1) Obtain the traffic monitoring videos of highways and urban environments respectively, and loop to read each frame of the video pictures; 2) Select a place closer to the shooting angle of the monitoring camera to draw a virtual detection line to maximize the capture of the feature information of the target vehicle; 3) Set the region of interest: First, determine the coordinates of the two outermost lanes in the video and connect them into two straight lines to form two regions fitting the urban traffic lanes. The selected region is the region of interest, and traffic flow statistics are only realized in this region; 4) Combine the improved YOLO-V7 model and the Deep-Sort algorithm: YOLO-V7 identifies vehicle targets to obtain the detection box information of the target objects and imports it into the Deep-Sort framework to generate the target vehicle tracking block diagram; 5) According to the block diagram generated in 4), determine the center point of the target vehicle block diagram in each frame. When the center point crosses the virtual detection line set in 2), the vehicle counter is incremented by 1 to complete the counting of the vehicle flow on the road during the specified period.