A Monocular Multi-Target Tracking and Speed Measurement Method Based on Deep Learning
By using the target detection model optimized by YOLOv5 and SimAM modules, combined with the kalman algorithm, the problem of target appearance changes and occlusion in traffic scenarios is solved, and high-precision multi-objective tracking and speed measurement is achieved, which is suitable for lightweight hardware deployment.
Patent Information
- Application Number
- CN202411543176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-31
AI Technical Summary
In actual traffic scenarios, the significant appearance changes of the target object and the mutual occlusion problems lead to poor tracking effects of existing discriminatory models, affecting the accuracy of speed measurement analysis.
YOLOv5 is used as the backbone network architecture, combined with SimAM module and wiou loss function, the object detection model is trained, the kalman algorithm is used for target tracking and speed measurement, and attention weight is generated by calculating the local self-similarity of the feature map, solving the problem of data quality inconsistency, and the model is optimized by IOU.
It realizes high-precision multi-objective tracking and speed measurement in complex traffic scenarios. The model runs fast, has strong adaptability and low hardware requirements, and is suitable for lightweight deployment.
Smart Images

Figure CN119579645B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a monocular multi-object tracking and speed measurement method based on deep learning. Background Art
[0002] With the rapid development of artificial intelligence, deep learning technology is constantly updated and iterated, which can better empower frontier industrial applications. Among them, the urban road monitoring system plays an important role in public security prevention and control. Further optimizing the monitoring system and expanding its various functions can not only reduce its construction and maintenance costs, but also provide technical support for various fields such as transportation and public security, greatly improving the level and efficiency of law enforcement management. Currently, in the field of intelligent traffic monitoring, when people solve the problems of (vehicle, personnel) target tracking and speed measurement, most of them adopt discriminative models and use the principle of correlation filtering to distinguish the target objects and background environment in the monitoring video sequence, that is, learning a classifier or a filter to discriminate targets and non-targets. In the video frame, the learned classifier or filter is used as a discriminant model to search for and locate the target. In subsequent new frames, multiple candidate regions are evaluated, and the region with the highest score output by the discriminant model is selected as the new position of the target, so as to achieve the tracking effect. Subsequently, mathematical analysis is carried out to calculate the moving speed of the tracked target.
[0003] When using a discriminative model and the principle of correlation filtering for target tracking, the effect of target tracking depends to a large extent on correctly predicting the approximate position of the target. Fine-tuning the discriminant model is estimated to provide a more accurate tracking effect. If the appearance of the tracked object is very simple and does not change during movement and there is no object occlusion, the tracking effect will be very good. However, in actual traffic scenarios, the appearance of the tracked object is very likely to change significantly, and there will also be a large number of problems where tracked objects block each other. The discriminant model is easily interfered, resulting in poor tracking effects, and the accuracy of subsequent speed measurement analysis is also affected.
[0004] Therefore, it is necessary to provide a monocular multi-object tracking and speed measurement method based on deep learning to solve the above technical problems. Summary of the Invention
[0005] The present invention provides a monocular multi-object tracking and speed measurement method based on deep learning, which solves the problems that in actual traffic scenarios, the appearance of the tracked object is very likely to change significantly, and there will also be a large number of problems where tracked objects block each other. The discriminant model is easily interfered, resulting in poor tracking effects, and the accuracy of subsequent speed measurement analysis is also affected.
[0006] To solve the above technical problems, a monocular multi-object tracking and speed measurement method based on deep learning provided by the present invention includes the following steps:
[0007] S1. Prepare a training dataset. Under real traffic scenarios, use a camera to collect real vehicle and pedestrian data in real time. At the same time, add open-source datasets, which can effectively improve the robustness and generalization of subsequent model inferences.
[0008] S2. Adopt YOLOv5 as the backbone network architecture for model training, and train the vehicle and pedestrian datasets obtained in S1 to obtain an object detection model.
[0009] S3. Use the trained vehicle and pedestrian object detection model in S2 to quickly and accurately detect all traffic objects in the video frame, and then pass these detection results to the tracking algorithm.
[0010] S4. Measure the moving speed of traffic objects.
[0011] Preferably, the open-source datasets in S1 include (COCO2017, KITTI, VisDrone2019, OpenImagesV7, etc.). Screen, clean, and annotate all data to produce a large-scale training dataset suitable for vehicle and pedestrian object detection.
[0012] Preferably, YOLOv5 is selected as the deep learning object detection network architecture in S2 because YOLOv5 builds on the success of previous YOLO versions in object detection tasks, further improving performance and flexibility, and having cutting-edge performance in terms of accuracy and speed. At the same time, the YOLOv5 network architecture model is better adapted to the hardware rk3588 development board to achieve a more lightweight deployment. In this invention, a lightweight and parameter-free convolutional neural network attention mechanism SimAM module is added to the basis of the YOLOv5 network framework, and the wiou loss function is also introduced.
[0013] Preferably, the SimAM module generates attention weights by calculating the local self-similarity of the feature map, and adding the SimAM module does not require introducing any additional parameters, which can effectively improve the convolutional learning ability during model training. The core idea of the SimAM module is based on the local self-similarity of the image. Usually, adjacent pixels in the image have strong similarity, while the similarity between distant pixels is weak. SimAM utilizes this feature to generate attention weights by calculating the similarity between each pixel in the feature map and its adjacent pixels.
[0014] Preferably, in the operation process of SimAM, the global average pooling operation is first used to extract the global features of the feature map, then each pixel point is calculated to obtain the local features of the pixel point, and then the global features and local features are used for calculation to obtain the attention weight of each pixel point. The local features are weighted and fused using the attention weight to obtain the final feature representation. The calculation expression is as follows:
[0015] A(x) = σ(W × G(x) + b)
[0016] In the expression, A(x) represents the attention weight, x is the local feature, G(x) is the global feature, W is the weight matrix, b is the bias term, and σ is the activation function;
[0017] Since the currently trained dataset is large and contains various open-source data, there are inconsistencies in the data annotation standards and quality. When improving the model, the wiou loss function is adopted to solve the balance problem between samples with good data quality and poor data quality, which helps to improve the generalization ability of the model;
[0018] The wiou loss function needs to use the IOU to obtain the intersection-over-union value of the predicted anchor box and the ground-truth anchor box, L IOU = 1 - IOU. Here, L IOU represents the loss of IOU. Then, the outlier degree β of the target anchor box is defined. The outlier degree β is represented by the ratio of the current IOU loss to the mean value of L IOU ;
[0019]
[0020] The smaller the outlier degree β of the target anchor box, the higher the quality of the anchor box. A small gradient gain is assigned to make the bounding box regression focus on the anchor boxes of ordinary quality. When the outlier degree β of the target anchor box is large, a smaller gradient gain is assigned, which can effectively prevent the influence of low-quality label data.
[0021] Preferably, the process of tracking traffic objects in S3 is divided into the following steps:
[0022] S31. Initialize data information: First, create an empty trajectory set to store the tracking information of each traffic object in the video;
[0023] S32. Data object detection: Perform traffic object detection on the video to be tracked and analyzed (here, the object detection uses the model trained in S2 for inference), and output the position anchor box of the predicted target and the confidence information of its predicted target anchor box;
[0024] S33. Classify the object detection results: Set a specific threshold for the confidence information of the object anchor boxes obtained in S2 to distinguish between high-confidence detections and low-confidence detections. Here, the threshold is selected as 0.6;
[0025] S34. Predict the trajectories of traffic objects: Use the Kalman algorithm to predict the new positions of each existing trajectory in the current frame;
[0026] S35. First data association: Use a similarity metric to associate the existing trajectories with the high-confidence detections;
[0027] S36. Update the remaining unmatched detections and trajectories: After S5, find the remaining unmatched detection anchor boxes and trajectory information;
[0028] S37. Second data association: Use a similarity metric to associate the remaining trajectories in S6 with the low-confidence detections in S3;
[0029] S38. Clean up the trajectory information: Delete all the trajectory information that was not matched in S7 from the trajectory set;
[0030] S39. Initialize the new trajectory information: For the remaining unmatched detections after the second association, initialize a new trajectory for each detection and add it to the trajectory set;
[0031] S310. Return the traffic object trajectory results: Return the updated trajectory set, which is the tracking path information of all traffic objects.
[0032] Preferably, the Kalman algorithm in S34 is used to predict the target positions during multi-object tracking. To implement the Kalman algorithm calculation, first, initialize the target state to be tracked and its covariance matrix, then predict the state of the target to be tracked at the next moment, and update the target state and covariance matrix information. The mathematical expressions are as follows:
[0033]
[0034] Where represents the predicted state of the tracked target, represents the predicted covariance matrix, F t is the state transition matrix, B t is the control input matrix, u t is the control input vector, Q t is the process noise covariance matrix. Here, t represents the t-th moment, and t - 1 represents the (t - 1)-th moment.
[0035] Preferably, in S4, it is first necessary to determine that the image and video collector is fixedly placed and has clear position information. Then, using S2 and S3 above, traffic target objects in the image are detected in real time and tracked. The detection anchor box of each target is extracted, and for each detected traffic target anchor box, it is further mathematically converted into its pixel coordinates in the image, and the actual distance between each traffic target and the image and video collector is calculated;
[0036] Then, start calculating the speed of the tracked target. Perform continuous frame analysis on the real-time video. After successfully detecting and tracking the traffic target, record the position changes of each time step and the target detection anchor box. Its mathematical expression is:
[0037]
[0038] That is, the instantaneous speed Δv is equal to the ratio of distance to time. In the formula, Δs is the actual distance change value between the same-source traffic target and the image and video collector in two adjacent frames. The calculation formula of Δt is:
[0039]
[0040] Δt represents the time difference between two adjacent frames, and fps is the frequency at which the image continuously appears on the display, that is, the frame rate of the image and video collector. Here, Δt is converted to seconds for subsequent measurement and recording, realizing the speed measurement function of the tracked target.
[0041] Compared with the related technology, a monocular multi-target tracking and speed measurement method based on deep learning provided by the present invention has the following beneficial effects:
[0042] The present invention provides a monocular multi-target tracking and speed measurement method based on deep learning. The proposed technical solution first establishes a detection model based on a deep learning model, obtains prior target information based on the powerful detection model, and uses an association strategy to achieve high-precision tracking and speed measurement functions for traffic objects. It runs fast in actual application reasoning, can achieve real-time monitoring, has very strong adaptability in hardware deployment, the model encapsulation is simple and small, and has low requirements for the computing power performance of the hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic structural diagram of a preferred embodiment of a monocular multi-target tracking and speed measurement method based on deep learning provided by the present invention;
[0044] Figure 2 It is a schematic diagram of image and video collection;
[0045] Figure 3 Schematic diagram of the target tracking execution strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0047] First Embodiment
[0048] Please refer to Figure 1 、 Figure 2 and Figure 3 , wherein, Figure 1 is a schematic structural diagram of a preferred embodiment of a monocular multi-object tracking and speed measurement method based on deep learning provided by the present invention; Figure 2 is a schematic diagram of image and video acquisition; Figure 3 Schematic diagram of the target tracking execution strategy. A monocular multi-object tracking and speed measurement method based on deep learning includes the following steps:
[0049] S1. Prepare a training data set. Under traffic real scenes, use a camera to collect real vehicle and pedestrian data in real time, and at the same time add open-source data sets, which can effectively improve the robustness and generalization of subsequent model inference;
[0050] S2. Use YOLOv5 as the backbone network architecture for model training, train the vehicle and pedestrian data sets obtained in S1, and obtain an object detection model;
[0051] S3. Use the trained vehicle and pedestrian object detection model in S2 to quickly and accurately detect all traffic objects in the video frame, and then transfer these detection results to the tracking algorithm;
[0052] S4. Measure the moving speed of traffic objects.
[0053] The open-source data sets in S1 include (COCO2017, KITTI, VisDrone2019, OpenImagesV7, etc.). All data are screened, cleaned, and labeled to produce a large-scale training data set suitable for vehicle and pedestrian object detection.
[0054] In S2, YOLOv5 is selected as the deep learning object detection network architecture. Because YOLOv5 builds on the success of previous YOLO versions in object detection tasks, it further improves performance and flexibility, and has cutting-edge performance in terms of accuracy and speed. At the same time, the YOLOv5 network architecture model is better adapted to the hardware rk3588 development board to achieve a more lightweight deployment. In the present invention, a lightweight and parameter-free convolutional neural network attention mechanism SimAM module is added to the basis of the YOLOv5 network framework, and the wiou loss function is also introduced.
[0055] Here, YOLOv5 is selected as the deep learning object detection network architecture. Because YOLOv5 builds on the success of previous YOLO versions in object detection tasks, it further improves performance and flexibility, and has cutting-edge performance in terms of both accuracy and speed. At the same time, the YOLOv5 network architecture model is better adapted to the hardware rk3588 development board to achieve a more lightweight deployment. In this invention, a lightweight and parameter-free convolutional neural network attention mechanism, the SimAM module, is added to the YOLOv5 network framework, and the wiou loss function is also introduced.
[0056] The SimAM module generates attention weights by calculating the local self-similarity of the feature map, and adding the SimAM module does not require introducing any additional parameters, which can effectively improve the convolutional learning ability during model training. The core idea of the SimAM module is based on the local self-similarity of images. Usually, adjacent pixels in an image have strong similarity, while the similarity between distant pixels is weak. SimAM utilizes this characteristic to generate attention weights by calculating the similarity between each pixel in the feature map and its adjacent pixels.
[0057] In the running process of the SimAM, first, the global average pooling operation is used to extract the global features of the feature map, then each pixel point is calculated to obtain the local features of that pixel point, and then the global features and local features are used for calculation to obtain the attention weights of each pixel point. The local features are weighted and fused using the attention weights to obtain the final feature representation. Its calculation expression is:
[0058] A(x) = σ(W × G(x) + b)
[0059] In the expression, A(x) represents the attention weight, x is the local feature, G(x) is the global feature, W is the weight matrix, b is the bias term, and σ is the activation function;
[0060] Since the currently trained dataset is large and contains various open-source data, there are inconsistencies in the data annotation standards and quality. When the model is improved, the wiou loss function is adopted to solve the balance problem between samples with good data quality and those with poor data quality, which helps to improve the generalization ability of the model;
[0061] The wiou loss function needs to use the IOU to obtain the intersection-over-union value of the predicted anchor box and the true anchor box, L IOU = 1 - IOU, where L IOU represents the loss of IOU. Then, the outlier degree β of the target anchor box is defined. The outlier degree β is expressed as the ratio of the current IOU loss to the mean value of L IOU
[0062]
[0063] The smaller the outlier degree β of the target anchor box, the higher the quality of the anchor box, and a small gradient gain is assigned to focus the bounding box regression on the anchor boxes of ordinary quality. When the outlier degree β of the target anchor box is large, a smaller gradient gain is assigned to effectively prevent the influence of low-quality labeled data.
[0064] The process of tracking traffic objects in S3 is divided into the following steps:
[0065] S31. Initialize data information: First, create an empty trajectory set to store the tracking information of each traffic object in the video.
[0066] S32. Data target detection: Perform traffic object target detection on the video to be tracked and analyzed (here, the target detection uses the model trained in S2 for inference), and output the position anchor box of the predicted target and the confidence information of its predicted target anchor box.
[0067] S33. Classify the target detection results: Set a specific threshold for the confidence information of the target anchor box obtained in S2 to distinguish between high-confidence detection and low-confidence detection. Here, the threshold is selected as 0.6.
[0068] S34. Traffic object trajectory prediction: Use the Kalman algorithm to predict the new position of each existing trajectory in the current frame.
[0069] S35. First data association: Use similarity measurement to associate the existing trajectories with the high-confidence detections.
[0070] S36. Update the remaining unmatched detections and trajectories: After S5, find the remaining unmatched detection anchor boxes and trajectory information.
[0071] S37. Second data association: Use similarity measurement to associate the remaining trajectories in S6 with the low-confidence detections in S3.
[0072] S38. Clean up trajectory information: Delete all trajectory information that is not matched in S7 from the trajectory set.
[0073] S39. Initialize new trajectory information: For the remaining unmatched detections after the second association, initialize a new trajectory for each detection and add it to the trajectory set.
[0074] S310. Return the traffic object trajectory results: Return the updated trajectory set, which is the tracking path information of all traffic objects.
[0075] The Kalman algorithm in S34 is used to predict the target position during multi-target tracking. To implement the Kalman algorithm calculation, first, the target state to be tracked and its covariance matrix need to be initialized. Then, the state of the target to be tracked at the next moment is predicted, and the target state and covariance matrix information are updated. The mathematical expression is as follows:
[0076]
[0077] Where represents the predicted state of the tracked target, represents the predicted covariance matrix, and F t is the state transition matrix, B t is the control input matrix, u t is the control input vector, and Q t is the process noise covariance matrix. Here, t represents the t-th moment, and t - 1 represents the (t - 1)-th moment.
[0078] In S4, first, it is necessary to determine that the image video collector is fixedly placed and has clear position information. Then, using S2 and S3 above, traffic target objects in the image are detected and tracked in real time. The detection anchor box of each target is extracted, and for each detected traffic target anchor box, it is further mathematically converted into its pixel coordinates in the image, and the actual distance between each traffic target and the image video collector is calculated;
[0079] Then, the speed of the tracked target is calculated. By analyzing consecutive frames of the real-time video, after successfully detecting and tracking the traffic target, the position changes of each time step and the target detection anchor box are recorded. The mathematical expression is as follows:
[0080]
[0081] That is, the instantaneous speed Δv is equal to the ratio of distance to time. In the formula, Δs is the actual distance change value between the same-source traffic target and the image video collector in two adjacent frames, and the calculation formula for Δt is:
[0082]
[0083] Δt represents the time difference between two adjacent frames, and fps is the frequency at which the image continuously appears on the display, that is, the frame rate of the image video collector. Here, Δt is converted to seconds for subsequent measurement and recording, realizing the speed measurement function of the tracked target.
[0084] Compared with the related technologies, a monocular multi-target tracking and speed measurement method based on deep learning provided by the present invention has the following beneficial effects:
[0085] The present invention provides a monocular multi-target tracking and speed measurement method based on deep learning. The proposed technical solution first establishes a detection model based on the deep learning model, obtains prior target information based on the powerful detection model, and references the association strategy to achieve high-precision tracking and speed measurement functions for traffic objects. The present invention runs fast in actual application reasoning, can realize real-time monitoring, has strong adaptability in hardware deployment, has a simple and compact model package, and does not require high hardware computing power performance.
[0086] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A monocular multi-object tracking and speed measurement method based on deep learning, characterized in that, It includes the following steps: S1. Prepare a training dataset. Under real traffic scenarios, use a camera to collect real vehicle and pedestrian data in real time, and also add open-source datasets; S2. Adopt YOLOv5 as the backbone network architecture for model training, and train the vehicle and pedestrian datasets obtained in S1 to obtain an object detection model; S3. Use the trained vehicle and pedestrian object detection model in S2 to quickly and accurately detect all traffic objects in the video frame, and then transfer these detection results to the tracking algorithm; S4. Measure the moving speed of traffic objects; The process of tracking traffic objects in S3 is divided into the following steps: S31. Initialize data information: First, create an empty trajectory set to store the tracking information of each traffic object in the video; S32. Data object detection: Perform traffic object detection on the video to be tracked and analyzed, and output the position anchor box of the predicted object and the confidence information of its predicted object anchor box; S33. Classify the object detection results: Set a specific threshold for the confidence information of the object anchor box obtained in S2 to distinguish between high-confidence detections and low-confidence detections. Here, the threshold is selected as 0.6; S34. Traffic object trajectory prediction: Use the Kalman algorithm to predict the new position of each existing trajectory in the current frame; S35. First data association: Use a similarity metric to associate the existing trajectories with high-confidence detections; S36. Update the remaining unmatched detections and trajectories: After S5, find the remaining unmatched detection anchor boxes and trajectory information; S37. Second data association: Use a similarity metric to associate the remaining trajectories in S6 with the low-confidence detections in S3; S38. Clean up trajectory information: Delete all trajectory information that was not matched in S7 from the trajectory set; S39. Initialize new trajectory information: For the remaining unmatched detections after the second association, initialize a new trajectory for each detection and add it to the trajectory set; S310. Return traffic object trajectory results: Return the updated trajectory set, which is the tracking path information of all traffic objects.
2. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 1, characterized in that, The open-source datasets in S1 include COCO2017, KITTI, VisDrone2019, and OpenImagesV7. All data are screened, cleaned, and labeled to produce a large-scale training dataset suitable for vehicle and pedestrian object detection.
3. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 1, characterized in that In S2, YOLOv5 is selected as the deep learning object detection network architecture because YOLOv5 builds on the success of previous YOLO versions in object detection tasks, further improving performance and flexibility, and having cutting-edge performance in terms of accuracy and speed. At the same time, the YOLOv5 network architecture model is better adapted to the hardware rk3588 development board to achieve a more lightweight deployment. A lightweight and parameter-free convolutional neural network attention mechanism SimAM module is added to the basis of the YOLOv5 network framework, and the wiou loss function is also introduced.
4. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 3, characterized in that, The SimAM module generates attention weights by calculating the local self-similarity of the feature map, and adding the SimAM module does not require the introduction of any additional parameters, which can effectively improve the convolution learning ability during model training. The core idea of the SimAM module is based on the local self-similarity of the image. In the image, adjacent pixels usually have strong similarities, while the similarity between distant pixels is weaker. SimAM uses this feature to generate attention weights by calculating the similarity between each pixel in the feature map and its adjacent pixels.
5. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 4, characterized in that, The SimAM running process first uses the global average pooling operation to extract the global features of the feature map, then calculates each pixel to obtain the local features of the pixel, and then uses the global features and local features to calculate to obtain the attention weight of each pixel. The attention weight is used to perform weighted fusion on the local features to obtain the final feature representation, and its calculation expression is: In the expression, represents the attention weight, is the local feature, is the global feature, W is the weight matrix, is the bias term, and σ is the activation function; Since the current training data set is large and contains various open source data, there are inconsistencies in data annotation standards and quality. When improving the model, the wiou loss function is used to solve the balance problem between samples with good data quality and samples with poor data quality, which helps to improve the generalization ability of the model; The wiou loss function needs to use the IOU to obtain the intersection over union value of the predicted anchor box and the true anchor box. , where the is expressed as the loss of the IOU, and then the outlier degree of the target anchor box is defined . The outlier degree is represented by the ratio of the current IOU loss to the mean value. Target anchor box outlier degree The smaller it is, the higher the quality of the anchor box. A small gradient gain is assigned to focus the bounding box regression on the anchor boxes of ordinary quality. When the target anchor box outlier degree is large, a smaller gradient gain is assigned, which can effectively prevent the influence of low-quality labeled data.
6. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 1, characterized in that, The Kalman algorithm in S34 is used to predict the target position during multi-target tracking. To implement the Kalman algorithm calculation, the target state and its covariance matrix to be tracked must first be initialized, and then the state of the tracked target at the next moment is predicted, and the target state and covariance matrix information are updated. The mathematical expression is: where represents the state of the predicted tracking target, represents the predicted covariance matrix, is the state transition matrix, is the control input matrix, is the control input vector, is the process noise covariance matrix, where t represents the t-th moment and t - 1 represents the (t - 1)-th moment.
7. The monocular multi-object tracking and speed measurement method based on deep learning according to claim 1, characterized in that In S4, it is first necessary to determine that the image video collector is fixed and has clear position information, and then use the above S2 and S3 to detect and track the traffic target objects in the image in real time, extract the detection anchor frame of each target, and further mathematically convert each detected traffic target anchor frame into its pixel coordinates in the image to infer the actual distance between each traffic target and the image video collector; Then we start to calculate the speed of the tracked target and analyze the continuous frames of the real-time video. After successfully detecting and tracking the traffic target, we record the position change of each time step and the target detection anchor frame. The mathematical expression is: That is, the instantaneous speed is equal to the ratio of distance to time, where is the actual distance change value between the homologous traffic target and the image video collector in two adjacent frames, and the calculation formula is: represents the time difference between two adjacent frames, and fps is the frequency at which images continuously appear on the display, that is, the frame rate of the image video collector. Here it is converted to seconds for subsequent measurement and recording, and the function of measuring the speed of the tracking target is realized.
Citation Information
Patent Citations
Method and device for determining automobile driving speed
CN101187671A
Vehicle tracking method based on video stream deep learning
CN116778224A
Driver safety belt detection method based on deep learning
CN117333852A