A Method for Dynamic Traffic Target Detection, Tracking and Forward Collision Warning

By improving the YOLOv8 network and DeepSORT algorithm, the target detection and tracking problems in complex traffic environments are solved, and high-precision and low-latency vehicle collision warning is achieved, which is suitable for intelligent traffic and autonomous driving.

CN119540886BActive Publication Date: 2025-07-22SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411596908.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-22
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing YOLO and DeepSORT algorithms have problems such as insufficient target detection accuracy, poor tracking stability and insufficient real-time performance in complex traffic environments. Especially when detecting long-distance small targets and fast moving targets, it is easy to detect and miss detection, which is difficult to meet the high-precision and low-latency requirements of autonomous driving and intelligent transportation systems.

Method used

By improving the feature fusion module, attention module and detection loss function of the YOLOv8 network, combined with the parameterless attention mechanism, the target detection accuracy is improved; the complete crossover loss function is introduced in the DeepSORT algorithm, the matching of the tracking bounding box and the real bounding box is optimized, the stability of multi-objective tracking is enhanced, and the vehicle spacing and collision time are calculated through the Mobileye model for three-stage risk assessment.

Benefits of technology

It significantly improves the target detection accuracy and tracking stability in complex traffic environments, realizes efficient real-time processing capabilities, provides high-precision vehicle collision warning, and is suitable for intelligent traffic systems and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540886B_ABST
    Figure CN119540886B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dynamic traffic target detection, tracking and forward collision warning, which solves the problems of low target detection accuracy, poor tracking stability and high real-time requirements in complex traffic environments. By improving the YOLOv8 network, the detection accuracy of small targets and the lightweight performance of the network are enhanced; the complete intersection over union is used to optimize the tracking process of DeepSORT, improving the tracking effect of vehicles in fast-moving and occluded scenarios. In addition, based on the Mobileye model, the vehicle distance and collision time are calculated in real time, realizing a three-stage risk warning. The present invention has the advantages of high detection accuracy, strong real-time performance, wide application, etc., and is applicable to vehicle collision warning tasks in autonomous driving and intelligent transportation systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving and intelligent transportation, and in particular to a method for dynamic traffic target detection, tracking and forward collision warning. Background Art

[0002] With the development of autonomous driving technology and the popularization of intelligent transportation systems, vehicle target detection and real-time tracking technologies have become key components to ensure driving safety. The existing YOLO series of target detection algorithms are widely used in vehicle detection due to their real-time performance and high efficiency. As a classic multi-target tracking algorithm, DeepSORT can perform real-time vehicle tracking based on features such as the spatial position and motion trajectory of the target.

[0003] However, in actual application scenarios, especially in complex traffic environments such as highways and congested urban roads, the existing YOLO and DeepSORT algorithms have some deficiencies. For example, insufficient target detection accuracy: When detecting small targets at a long distance and occluded vehicles, the traditional YOLO network has deficiencies in feature extraction and target localization, resulting in a decrease in detection accuracy. Tracking stability issues: In multi-target tracking scenarios, DeepSORT is limited by the traditional intersection over union (IoU) matching algorithm and is prone to target loss or mis-matching, especially when the target moves quickly or the vehicle size changes. Real-time requirements: Existing algorithms are difficult to balance detection accuracy and processing speed in complex traffic scenarios and cannot meet the requirements of efficient real-time processing, affecting the actual application effect.

[0004] Therefore, it is an urgent need in the current technical field to propose a method for dynamic traffic target detection, tracking and collision warning that can maintain high-efficiency detection and stable tracking in complex environments and has a real-time warning function. Summary of the Invention

[0005] The object of the present invention is to solve the problems of complex background, high vehicle movement speed, difficult detection of small targets and high real-time requirements in vehicle target detection and tracking in dynamic traffic scenarios, and propose a method for dynamic traffic target detection, tracking and forward collision warning. This method can effectively improve the detection accuracy of small targets at a long distance and fast-moving targets in complex traffic environments, and avoid false detection and missed detection problems. At the same time, combined with lightweight model design and parameter-free attention mechanism, the real-time processing ability and detection efficiency are greatly improved, making this method show robust performance in dynamic traffic scenarios and meeting the requirements of high precision and low latency for autonomous driving and intelligent transportation systems.

[0006] To achieve the above object, the technical solution provided by the present invention is: a dynamic traffic target detection, tracking and forward collision warning method, which realizes the accurate detection and real-time tracking of the target vehicle in front of the driving vehicle by improving the YOLOv8 network and the DeepSORT algorithm; wherein, the improvement of the YOLOv8 network lies in enhancing the feature fusion module, the attention module and the detection loss function module. In the feature fusion module, lightweight fusion is performed on the feature maps of different scales extracted by the backbone network of the YOLOv8 network. In the attention module, a parameter-free attention mechanism is introduced to dynamically adjust the weights of each channel, emphasize important features and suppress irrelevant and redundant features. In the detection loss function module, weighted intersection over union is introduced as the detection loss function; the improvement of the DeepSORT algorithm lies in enhancing the tracking loss function module, and using complete intersection over union as the tracking loss function to optimize the distance between the tracking bounding box and the ground truth bounding box, so as to reduce the mis-matching phenomenon;

[0007] The specific implementation of this method includes the following steps:

[0008] 1) Obtain the video image data in the driving vehicle's dashboard camera;

[0009] 2) Use the trained improved YOLOv8 network and improved DeepSORT algorithm to process the video image data as follows:

[0010] Input the video image data into the backbone network of the improved YOLOv8 network to obtain feature map A; input feature map A into the feature fusion module to obtain multi-scale feature map B; input multi-scale feature map B into the attention module to obtain weighted feature map C; input weighted feature map C into the detection head of the improved YOLOv8 network to obtain the target vehicle bounding box;

[0011] Input the target vehicle bounding box into the tracking feature extraction module of the improved DeepSORT algorithm to output the identity feature vector corresponding to the target vehicle bounding box; input the target vehicle bounding box and the identity feature vector into the motion prediction module of the improved DeepSORT algorithm to obtain the predicted bounding box position E; input the predicted bounding box position E into the matching module of the improved DeepSORT algorithm to obtain the continuous tracking result of the target vehicle, that is, the vehicle continuous tracking bounding box;

[0012] 3) Based on the pixel size of the vehicle continuous tracking bounding box and the physical size of the actual vehicle, calculate the real-time distance between the target vehicle and the driving vehicle through the Mobileye model, and then calculate the collision time based on the real-time distance and the relative speed between the target vehicle and the driving vehicle. Finally, according to the size relationship between the collision time and the pre-set safety time, the risk level is divided into three levels: high, medium, and no risk, and real-time warning information is output.

[0013] Further, in step 1), video image data of the road ahead of the vehicle is collected in real time by the driving recorder of the vehicle, and the video image data is preprocessed, including denoising, picture size adjustment, and frame rate normalization, to meet the computational requirements of the improved YOLOv8 network and the improved DeepSORT algorithm.

[0014] Further, the improved YOLOv8 network includes:

[0015] A feature fusion module, which includes a lightweight convolutional layer for extracting multi-scale features and upsampling and downsampling units for scale conversion;

[0016] An attention module, which consists of a non-parametric calculation unit and a non-parametric attention mechanism, including a convolutional layer and a Sigmoid activation layer;

[0017] A detection loss function module, which includes a weighted intersection over union loss function for measuring the weights of various targets; a backbone network, which consists of 5 convolutional layer groups and residual structure groups

[0018] A detection head, which consists of 3 convolutional layers, a Sigmoid activation function, and batch normalization.

[0019] Further, the feature fusion module specifically performs the following operations:

[0020] Lightweight convolutional operation:

[0021] F conv = DepthwiseConv(A) (1)

[0022] Upsampling and downsampling unit operation:

[0023] F up = Up(F Conv ), F down = Down(F Conv ) (2)

[0024] Multi-scale feature fusion:

[0025] B = F up + F Conv + F Down (3)

[0026] In the formula, DepthwiseConv represents the lightweight convolutional operation, Up and Down represent the upsampling and downsampling operations respectively; F up , F Conv , F Down respectively represent the multi-scale feature maps B obtained after upsampling, lightweight convolutional operation, and downsampling;

[0027] The attention module takes the multi-scale feature map B as input and, through calculation, obtains the weighted feature map C. The calculation process is as follows:

[0028] C = B·σ(Conv 1×1 (Concat(AvgPool(B), MaxPool(B)))) (4)

[0029] In the formula, AvgPool(B) represents average pooling of the multi-scale feature map B along the channel dimension; MaxPool(B) represents max pooling of the multi-scale feature map B along the channel dimension; Concat(·) represents feature concatenation; Conv 1×1 (·) represents performing a 1×1 convolution on the concatenated feature map, and σ(·) represents the Sigmoid activation function;

[0030] The calculation process of the weighted intersection over union loss function TotalLoss of the detection loss function module is as follows:

[0031]

[0032]

[0033] In the formula, N represents the total number of targets, w class_i represents the weight factor of the i-th target class class_i, X pi , X ti respectively represent the predicted bounding box and the ground truth bounding box of the i-th target; |X pi ∩X ti | represents the intersection area of the predicted bounding box and the ground truth bounding box of the i-th target; |X pi ∪X ti | represents the union area of the predicted bounding box and the ground truth bounding box of the i-th target, and L IOU_i represents the intersection over union loss of the i-th target.

[0034] Furthermore, the improved DeepSORT algorithm includes:

[0035] A tracking feature extraction module, which consists of lightweight convolutional layers and is used to extract identity features from the target vehicle bounding box;

[0036] A motion prediction module, which consists of a Kalman filter and calculates the predicted bounding box position E by estimating the motion state of the target;

[0037] A matching module, which includes the Hungarian matching algorithm and a cost matrix unit to associate the detection results of the current frame with the tracks of the previous frame;

[0038] The tracking loss function module, which consists of the complete intersection over union loss function, ensures the consistency of the target position in consecutive frames.

[0039] Furthermore, the target vehicle bounding box is input into the tracking feature extraction module to obtain the identity feature vector d corresponding to each bounding box. i , and the specific steps are as follows:

[0040] d i = Conv(R i )(7)

[0041] In the formula, represents the vector space, represents the identity feature vector of target i, D is the dimension of the feature vector, and R i represents the vehicle bounding box of target i;

[0042] The target vehicle bounding box is input into the motion prediction module, and the motion prediction module uses the Kalman filter to estimate the predicted bounding box position E of the target in the next frame. The expression formula is as follows:

[0043]

[0044] In the formula, F represents the state transition matrix, X t-1 represents the target vehicle bounding box at the previous moment, K represents the Kalman gain, Z t represents the detected bounding box at the current moment, H represents the observation matrix, P t represents the updated covariance matrix, I represents the identity matrix, P t-1 represents the covariance matrix at the previous moment, F T represents the transpose of the state transition matrix, and Q represents the process noise covariance matrix;

[0045] The predicted bounding box position E and the identity feature vector of the detected bounding box are input into the matching module. The matching module calculates the cost matrix between the two and selects the detected bounding box with the minimum cost matrix as the tracking bounding box. The expression formula is as follows:

[0046] cost = βcost IOU + βcost dis (9)

[0047]

[0048] In the formula, cost is the total cost, cost IOU is the intersection over union cost, cost dis is the distance cost, b t and respectively represent the detected bounding box at the current moment and the predicted bounding box at the previous moment, d t-1 and dt respectively represent the identity feature vectors of the detected bounding box at the previous moment and the target vehicle bounding box at the current moment;

[0049] Input the tracking bounding box obtained by improving the DeepSORT algorithm into the tracking loss function module, and this module adopts the complete intersection over union loss function L CIoU to calculate the error. By calculating the overlapping area between the tracking bounding box and the ground truth bounding box, and introducing the consistency of the center point position deviation and the aspect ratio of the box body for a more comprehensive matching; among them, the complete intersection over union loss function L CIoU The specific calculation formula is:

[0050]

[0051] In the formula, N represents the total number of targets, ρ 2 (b, b gt ) is the Euclidean distance between the center points of the tracking bounding box and the ground truth bounding box, b represents the coordinates of the center point of the tracking bounding box, and b gt represents the coordinates of the center point of the ground truth bounding box; c is the diagonal length of the smallest circumscribed matrix that can accommodate both the tracking bounding box and the ground truth bounding box; α is the weight function; v is the metric aspect ratio consistency index, and the specific calculation method is as follows:

[0052]

[0053] In the formula, ω gt and h gt are the width and height of the ground truth bounding box respectively, and ω and h are the width and height of the tracking bounding box respectively.

[0054] Furthermore, in step 3), combining the pixel size of the continuous tracking bounding box of the vehicle and the physical size of the actual vehicle, use the Mobileye model to calculate the real-time distance Dis and the collision time t pre between the target vehicle and the vehicle being driven, and the specific calculation is as follows:

[0055] The real-time distance Dis is calculated through the focal length F of the camera in the driving recorder, the actual height W of the vehicle, and the pixel size P occupied by the vehicle in the image, and its formula is:

[0056]

[0057] Based on the real-time distance Dis between the vehicles and the relative speed V rel between the vehicle being driven and the target vehicle, calculate the collision time t pre , where the collision time refers to the time required for the vehicle being driven to possibly collide with the target vehicle in front under the condition of maintaining the current vehicle speed and distance, and the calculation formula is:

[0058]

[0059] By using the kinematic formula, calculate the prefabricated safety distance Dis_safe required to avoid collision with the following vehicle during emergency braking when the vehicle is accelerating. The calculation formula is as follows:

[0060]

[0061] In the formula, T W is the time delay budget required for the vehicle in which the driver is driving to issue a collision warning in a general risk scenario, V host is the vehicle speed of the vehicle in which the driver is driving; T hW is the time delay budget required for the vehicle to issue an emergency collision warning in a high-risk scenario; Tpr is the boost delay, that is, the time required for the pressure to build up when the braking system is in emergency braking; A W is the acceleration during emergency braking, and Safe is the safe stopping distance;

[0062] Using the relative speed V rel and the prefabricated safety distance Dis_safe, the prefabricated safety time t can be calculated as follows:

[0063]

[0064] Adopt a three-stage risk judgment method to compare the magnitude relationship between the collision time t pre and the prefabricated safety time t, so as to evaluate the rear-end collision risk and make the following warning judgments:

[0065]

[0066] When the collision time t pre is less than half of the prefabricated safety time t, it is determined to be a high risk, and the target vehicle is marked with a red frame for warning; when the collision time t pre is between the prefabricated safety time t and the prefabricated safety time , it is determined to be a low risk, and the target vehicle is marked with a yellow frame for warning; when the collision time t pre is greater than the prefabricated safety time t, it is determined to be risk-free, and the target vehicle is marked with a green frame.

[0067] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0068] 1. The present invention combines the most advanced object detection technology, neural network optimization technology, and real-time tracking technology, significantly improving the object detection and tracking performance in complex dynamic traffic environments. By improving the feature fusion module of the YOLOv8 network and introducing a parameter-free attention mechanism, the detection ability for vehicles at long distances and small sizes is enhanced. The attention mechanism can dynamically adjust the weights of each feature channel, effectively capturing the key features of vehicles, especially improving the detection accuracy in complex traffic scenarios.

[0069] 2. The present invention optimizes the matching accuracy between the predicted bounding box and the ground truth bounding box in the training of the YOLOv8 network by introducing the weighted intersection over union. Especially in scenarios with fast-moving and large variations in vehicle size, the false detection and missed detection phenomena are significantly reduced. In addition, the complete intersection over union is introduced into the DeepSORT algorithm, improving the stability of multi-object vehicle tracking and being able to effectively handle complex scenarios such as vehicle occlusion and speed changes.

[0070] 3. While optimizing the detection accuracy, the present invention also conducts a network lightweight design for real-time requirements. By adopting lightweight convolution operations and multi-scale feature map fusion technology, the computational efficiency of the YOLOv8 network is significantly improved. On the basis of ensuring the detection and tracking accuracy, high-efficiency real-time processing is achieved, meeting the high-performance requirements in complex traffic environments.

[0071] 4. The present invention successfully applies the improved YOLOv8 network and DeepSORT algorithm to the dynamic traffic object detection and real-time tracking scenarios, effectively improving the accuracy of vehicle collision warning in the traffic safety system. By calculating the real-time vehicle distance and collision time through the Mobileye model and combining the three-stage risk assessment method, an efficient and accurate warning system is provided, which is applicable to multiple fields such as intelligent transportation systems and autonomous driving, and has a wide application prospect and is worthy of promotion. Description of the Drawings

[0072] Figure 1For the structural diagram of the improved YOLOv8 network; in the figure, the Backbone part represents the backbone network module, the Neck part represents the feature fusion module, the Head part represents the detection head, CBS represents the basic convolution module, s represents the convolution stride, k represents the convolution kernel size, n represents the number of convolution repetitions; C2f represents multiple convolutional layers; SPPF represents the fast spatial pyramid pooling layer; SimAM represents the parameter-free attention mechanism; VoVGSCSP represents a feature fusion structure; Concat represents the concatenation operation; Upsample represents the upsampling operation; GSConv represents the grouped spatial convolution; 1x1 Conv-4*reg_max represents the 1×1 convolutional layer for predicting the bounding box regression information, and 1x1 Conv-num_class represents the 1×1 convolutional layer for predicting the target class information.

[0073] Figure 2 For the schematic diagram of the principle of the improved DeepSORT algorithm.

[0074] Figure 3 For the schematic diagram of the physical model for the Mobileye model to measure the inter-vehicle distance.

[0075] Figure 4 For the flowchart of the method of the present invention. Detailed implementation manners

[0076] The present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0077] As Figures 1 to 4 shown, this embodiment discloses a method for dynamic traffic target detection, tracking and forward collision warning. This method realizes the precise detection and real-time tracking of the target vehicle in front by the driving vehicle through improving the YOLOv8 network and the improved DeepSORT algorithm; among them, the improvement of the YOLOv8 network lies in enhancing the feature fusion module, the attention module and the detection loss function module. In the feature fusion module, lightweight fusion is performed on the feature maps of different scales extracted by the backbone network of the YOLOv8 network. In the attention module, a parameter-free attention mechanism is introduced to dynamically adjust the weights of each channel, emphasize important features and suppress irrelevant and redundant features. In the detection loss function module, the weighted intersection over union is introduced as the detection loss function; the improvement of the DeepSORT algorithm lies in enhancing the tracking loss function module, and the complete intersection over union is used as the tracking loss function to optimize the distance between the tracking bounding box and the true bounding box, so as to reduce the mis-matching phenomenon;

[0078] The specific implementation of this method includes the following steps:

[0079] 1) Obtain the video image data in the driving vehicle's dashboard camera;

[0080] Real-time video image data of the road ahead of the vehicle is collected by the driving recorder of the vehicle in motion, and the video image data is preprocessed, including denoising, picture size adjustment, and frame rate normalization, to meet the computational requirements of the improved YOLOv8 network and the improved DeepSORT algorithm.

[0081] 2) The following processing is performed on the video image data using the trained improved YOLOv8 network and the improved DeepSORT algorithm:

[0082] The video image data is input into the backbone network of the improved YOLOv8 network to obtain feature map A; feature map A is input into the feature fusion module to obtain multi-scale feature map B; multi-scale feature map B is input into the attention module to obtain weighted feature map C; weighted feature map C is input into the detection head of the improved YOLOv8 network to obtain the bounding box of the target vehicle;

[0083] The bounding box of the target vehicle is input into the tracking feature extraction module of the improved DeepSORT algorithm, and the identity feature vector corresponding to the bounding box of the target vehicle is output; the bounding box of the target vehicle and the identity feature vector are input into the motion prediction module of the improved DeepSORT algorithm to obtain the predicted bounding box position E; the predicted bounding box position E is input into the matching module of the improved DeepSORT algorithm to obtain the continuous tracking result of the target vehicle;

[0084] 3) Based on the pixel size of the continuous tracking bounding box of the vehicle and the physical size of the actual vehicle, the real-time distance between the target vehicle and the vehicle in motion is calculated through the Mobileye model, and then the time to collision is calculated based on the real-time distance and the relative speed between the target vehicle and the vehicle in motion. Finally, according to the magnitude relationship between the time to collision and the pre-set safety time, the risk level is divided into three levels: high, medium, and no risk, and real-time warning information is output.

[0085] Specifically, the improved YOLOv8 network includes:

[0086] A feature fusion module, which includes lightweight convolutional layers for extracting multi-scale features and upsampling and downsampling units for scale conversion;

[0087] An attention module, which consists of non-parametric calculation units and non-parametric attention mechanisms, including convolutional layers and Sigmoid activation layers;

[0088] A detection loss function module, which includes a weighted intersection over union loss function for measuring the weights of various types of targets;

[0089] A backbone network, which consists of 5 convolutional layer groups and residual structures;

[0090] The detection head is composed of 3 convolutional layers, a Sigmoid activation function, and batch normalization.

[0091] The feature fusion module specifically performs the following operations:

[0092] Lightweight convolution operation:

[0093] F conv = DepthwiseConv(A) (1)

[0094] Upsampling and downsampling unit operations:

[0095] F up = Up(F Conv ),F down = Down(F Conv ) (2)

[0096] Multi-scale feature fusion:

[0097] B = F up + F Conv + F Down (3)

[0098] In the formula, DepthwiseConv represents the lightweight convolution operation, Up and Down represent the upsampling and downsampling operations respectively; F up , F Conv , F Down respectively represent the multi-scale feature maps B obtained after upsampling, lightweight convolution operation, and downsampling;

[0099] The attention module takes the multi-scale feature map B as the input and, through calculation, obtains the weighted feature map C. The calculation process is as follows:

[0100] C = B·σ(Conv 1×1 (Concat(AvgPool(B), MaxPool(B)))) (4)

[0101] In the formula, AvgPool(B) represents average pooling of the multi-scale feature map B along the channel dimension; MaxPool(B) represents max pooling of the multi-scale feature map B along the channel dimension; Concat(·) represents feature concatenation; Conv 1×1 (·) represents 1×1 convolution of the concatenated feature map, and σ(·) represents the Sigmoid activation function;

[0102] The calculation process of the weighted intersection over union loss function TotalLoss of the detection loss function module is as follows:

[0103]

[0104] In the formula, N represents the total number of targets, and w class_i represents the weight factor of the i-th target category class_i, and X pi , X ti respectively represent the predicted bounding box and the ground truth bounding box of the i-th target; |X pi ∩X ti | represents the intersection area of the predicted bounding box and the ground truth bounding box of the i-th target; |X pi ∪X ti | represents the union area of the predicted bounding box and the ground truth bounding box of the i-th target, and L IOU_i represents the intersection over union loss of the i-th target.

[0105] The improved DeepSORT algorithm includes:

[0106] A tracking feature extraction module, which consists of lightweight convolutional layers and is used to extract identity features from the target vehicle bounding box;

[0107] A motion prediction module, which consists of a Kalman filter and calculates the predicted bounding box position E by estimating the motion state of the target;

[0108] A matching module, which includes a Hungarian matching algorithm and a cost matrix unit to associate the detection results of the current frame with the tracks of the previous frame;

[0109] A tracking loss function module, which consists of a complete intersection over union (CIoU) loss function to ensure the consistency of the target position in consecutive frames.

[0110] Input the target vehicle bounding box into the tracking feature extraction module to obtain the identity feature vector d i corresponding to each bounding box. The specific steps are as follows:

[0111] d i = Conv(R i )(7)

[0112] In the formula, represents the vector space, represents the identity feature vector of target i, D is the dimension of the feature vector, and R i represents the vehicle bounding box of target i;

[0113] Input the target vehicle bounding box into the motion prediction module. The motion prediction module uses a Kalman filter to estimate the predicted bounding box position E of the target in the next frame. The specific steps are as follows:

[0114] Input the target vehicle bounding box into the motion prediction module. The motion prediction module uses a Kalman filter to estimate the predicted bounding box position E of the target in the next frame. The expression formula is as follows:

[0115]

[0116] In the formula, F represents the state transition matrix, X t-1 represents the bounding box of the target vehicle at the previous moment, K represents the Kalman gain, Z t represents the detected bounding box at the current moment, H represents the observation matrix, P t represents the updated covariance matrix, I represents the identity matrix, P t-1 represents the covariance matrix at the previous moment, F T the transpose of the state transition matrix, Q represents the process noise covariance matrix.

[0117] Input the predicted bounding box position E and the identity feature vector of the detected bounding box into the matching module. The matching module calculates the cost matrix between the two and selects the detected bounding box with the minimum cost matrix as the tracking bounding box. The formula is as follows:

[0118] cost = βcost IOU +βcost dis (9)

[0119]

[0120] In the formula, cost is the total cost, cost IOU is the intersection over union cost, cost dis is the distance cost, b t and respectively represent the detected bounding box at the current moment and the predicted bounding box at the previous moment, d t-1 and d t respectively represent the identity feature vectors of the detected bounding box at the previous moment and the bounding box of the target vehicle at the current moment;

[0121] Input the tracking bounding box obtained by the improved DeepSORT algorithm into the tracking loss function module. This module uses the complete intersection over union loss function L CIoU to calculate the error. By calculating the overlapping area between the tracking bounding box and the ground truth bounding box, and introducing the consistency of the center point position deviation and the aspect ratio of the box body for more comprehensive matching; among them, the specific calculation formula of the complete intersection over union loss function L CIoU is:

[0122]

[0123] In the formula, N represents the total number of targets, ρ 2 (b, b gt ) is the Euclidean distance between the center points of the tracking bounding box and the ground truth bounding box, b represents the coordinates of the center point of the tracking bounding box, b gtIndicates the coordinates of the center point of the true bounding box; c is the diagonal length of the smallest circumscribed matrix that can accommodate both the tracking bounding box and the true bounding box; α is the weight function; v is the metric aspect ratio consistency index, and the specific calculation method is as follows:

[0124]

[0125] In the formula, ω gt and h gt are the width and height of the true bounding box respectively, and ω and h are the width and height of the tracking bounding box respectively.

[0126] Combining the pixel size of the vehicle's continuous tracking bounding box and the physical size of the actual vehicle, the Mobileye model is used to calculate the real-time distance Dis and collision time t between the target vehicle and the vehicle being driven pre . In the appendix Figure 3 , the relative position relationship between the vehicle being driven, Car1, and the target vehicles in front, Car2 and Car3, is shown to illustrate the distance calculation method between vehicles and the determination of the collision time. In the appendix Figure 3 :

[0127] I: Represents the optical center position of the camera, installed in front of the vehicle being driven, Car1, for capturing images of the vehicles in front.

[0128] Dis1 and Dis2: Represent the actual distances between the vehicle being driven and the target vehicles in front, Car2 and Car3, respectively.

[0129] F: Represents the focal length of the camera in the dashcam, which is a known camera parameter used to calculate the relationship between the pixel size on the image and the actual distance.

[0130] H: Represents the actual height of the vehicle being driven.

[0131] W: Represents the actual height of the target vehicle.

[0132] P1 and P2: Represent the vertical coordinates of the imaging planes of different target vehicles on the image plane.

[0133] Specifically, the distances Dis1 and Dis2 between vehicles can be calculated through the focal length F of the camera, the actual height W of the vehicle, and the pixel sizes P1 and P2 occupied by the vehicle in the image. The formula is:

[0134]

[0135] Taking the vehicle being driven, Car1, and the target vehicle in front, Car2, as an example, calculate their relative speed V rel , and calculate the collision time t pre, where the collision time refers to the time required for the vehicle being driven to potentially collide with the target vehicle ahead under the conditions of maintaining the current vehicle speed and distance. The calculation formula is as follows:

[0136]

[0137] Using the kinematic formula, calculate the prefabricated safety distance Dis_safe required to avoid colliding with the following vehicle during emergency braking when the vehicle is accelerating. The calculation formula is as follows:

[0138]

[0139] In the formula, T W is the time delay budget required for the vehicle being driven to issue a collision warning in a general risk scenario, V host is the vehicle speed of the vehicle being driven; T hW is the time delay budget required for the vehicle to issue an emergency collision warning in a high-risk scenario; Tpr is the boost delay, that is, the time required for the pressure to build up when the braking system is in emergency braking; A W is the acceleration during emergency braking, and Safe is the safe stopping distance;

[0140] Using the relative speed V rel and the prefabricated safety distance Dis_safe, the prefabricated safety time t can be calculated as follows:

[0141]

[0142] Adopt a three-stage risk determination method to compare the magnitude relationship between the collision time t pre and the prefabricated safety time t, so as to evaluate the rear-end collision risk and make the following warning judgments:

[0143]

[0144] When the collision time t pre is less than half of the prefabricated safety time t, it is determined to be a high risk, and the target vehicle is marked with a red frame for warning; when the collision time t pre is between the prefabricated safety time t and the prefabricated safety time , it is determined to be a low risk, and the target vehicle is marked with a yellow frame for warning; when the collision time t pre is greater than the prefabricated safety time t, it is determined to be risk-free, and the target vehicle is marked with a green frame.

[0145] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for dynamic traffic target detection, tracking and forward collision warning, characterized in that, This method realizes the precise detection and real-time tracking of the target vehicle in front by the vehicle in driving by improving the YOLOv8 network and the DeepSORT algorithm. Among them, the improvement of the YOLOv8 network lies in enhancing the feature fusion module, the attention module and the detection loss function module. In the feature fusion module, lightweight fusion is performed on the feature maps of different scales extracted by the backbone network of the YOLOv8 network. In the attention module, a parameter-free attention mechanism is introduced to dynamically adjust the weights of each channel, emphasize important features and suppress irrelevant and redundant features. In the detection loss function module, weighted intersection over union is introduced as the detection loss function. The improvement of the DeepSORT algorithm lies in enhancing the tracking loss function module, and using complete intersection over union as the tracking loss function to optimize the distance between the tracking bounding box and the ground truth bounding box to reduce the mis-matching phenomenon. The specific implementation of this method includes the following steps: 1) Obtain the video image data in the driving recorder of the vehicle in driving. 2) Use the trained improved YOLOv8 network and improved DeepSORT algorithm to process the video image data as follows: Input the video image data into the backbone network of the improved YOLOv8 network to obtain feature map A; input feature map A into the feature fusion module to obtain multi-scale feature map B; input multi-scale feature map B into the attention module to obtain weighted feature map C; input weighted feature map C into the detection head of the improved YOLOv8 network to obtain the bounding box of the target vehicle. Input the bounding box of the target vehicle into the tracking feature extraction module of the improved DeepSORT algorithm to output the identity feature vector corresponding to the bounding box of the target vehicle; input the bounding box of the target vehicle and the identity feature vector into the motion prediction module of the improved DeepSORT algorithm to obtain the predicted bounding box position E; input the predicted bounding box position E into the matching module of the improved DeepSORT algorithm to obtain the continuous tracking result of the target vehicle, that is, the continuous tracking bounding box of the vehicle. 3) Based on the pixel size of the continuous tracking bounding box of the vehicle and the physical size of the actual vehicle, calculate the real-time distance between the target vehicle and the vehicle in driving through the Mobileye model, and then calculate the collision time based on the real-time distance and the relative speed between the target vehicle and the vehicle in driving. Finally, according to the size relationship between the collision time and the pre-set safety time, the risk level is divided into three levels of high, medium and none risks, and real-time warning information is output as follows: Combined with the pixel size of the vehicle's continuous tracking bounding box and the physical size of the actual vehicle, the Mobileye model is used to calculate the real-time distance Dis and the collision time t between the target vehicle and the vehicle being driven, and the specific calculation is as follows: pre , specifically calculated as follows: The real-time distance Dis is calculated through the focal length F of the camera in the driving recorder, the actual height W of the vehicle, and the pixel size P occupied by the vehicle in the image. The formula is: Based on the real-time distance Dis between vehicles and the relative speed V between the vehicle being driven and the target vehicle rel , calculate the time to collision t pre , where the time to collision refers to the time required for the vehicle being driven to potentially collide with the target vehicle ahead under the condition of maintaining the current vehicle speed and distance, and the calculation formula is: Through the kinematic formula, calculate the pre-set safe distance Dis_safe required to avoid collision with the vehicle behind when braking emergently under the condition of vehicle acceleration. The calculation formula is as follows: where T W is the time delay budget required for the vehicle in motion to issue a collision warning in a general risk scenario, V host is the vehicle speed of the vehicle in motion; T hW is the time delay budget required for the vehicle to issue an emergency collision warning in a high-risk scenario; Tpr is the boost delay, that is, the time required for the pressure to build up when the braking system performs emergency braking; A W is the acceleration during emergency braking, and Safe is the safe stopping distance; Using the relative speed V rel and the prefabricated safety distance Dis_safe, the prefabricated safety time t can be calculated as follows: Adopt a three-stage risk determination method to compare the collision time t pre and the prefabricated safety time t to evaluate the rear-end collision risk and make the following early warning judgments: When the collision time t pre is less than half of the pre-set safety time t, it is determined as a high risk, and the target vehicle is marked with a red box for warning; when the collision time t pre is between the pre-set safety time t and the pre-set safety time it is determined as a low risk, and the target vehicle is marked with a yellow box for warning; when the collision time t pre is greater than the pre-set safety time t, it is determined as no risk, and the target vehicle is marked with a green box.

2. The dynamic traffic target detection, tracking and forward collision warning method according to claim 1, wherein, In step 1), video image data of the road ahead of the vehicle are collected in real time by the driving recorder of the vehicle, and the video image data are preprocessed, including denoising, picture size adjustment and frame rate normalization, to meet the computational requirements of the improved YOLOv8 network and the improved DeepSORT algorithm.

3. A dynamic traffic target detection, tracking and forward collision warning method according to claim 1, characterized in that, The improved YOLOv8 network includes: A feature fusion module, which includes lightweight convolutional layers for extracting multi-scale features and upsampling and downsampling units for scale conversion; An attention module, which consists of non-parametric calculation units and non-parametric attention mechanisms, including convolutional layers and Sigmoid activation layers; A detection loss function module, which includes a weighted intersection over union loss function for measuring the weights of various types of targets; A backbone network, which consists of 5 convolutional layer groups and residual structure groups A detection head, which consists of 3 convolutional layers, a Sigmoid activation function and batch normalization.

4. A dynamic traffic target detection, tracking and forward collision warning method according to claim 3, characterized in that The feature fusion module specifically performs the following operations: Lightweight convolutional operation: F conv = DepthwiseConv(A) (1) Upsampling and downsampling unit operation: F up = Up(F Conv ), F down = Down(F Conv )(2) Multi-scale feature fusion: B = F up +F Conv +F Down (3) Wherein, DepthwiseConv represents a lightweight convolution operation, Up and Down respectively represent upsampling and downsampling operations; F up , F Conv , F Down respectively represent the multi-scale feature maps B obtained after upsampling, lightweight convolution operation, and downsampling; The attention module takes the multi-scale feature map B as input and obtains the weighted feature map C through calculation. The calculation process is as follows: C = B·σ(Conv 1×1 (Concat(AvgPool(B), MaxPool(B)))) In Equation (4), AvgPool(B) represents average pooling of the multi-scale feature map B along the channel dimension; MaxPool(B) represents max pooling of the multi-scale feature map B along the channel dimension; Concat(·) represents feature concatenation; Conv 1×1 (·) represents 1×1 convolution of the concatenated feature map, and σ(·) represents the Sigmoid activation function; The calculation process of the weighted intersection over union loss function TotalLoss of the detection loss function module is as follows: Where N represents the total number of targets, w class_i represents the weight factor of the i-th target category class_i, and X pi , X ti represent the predicted bounding box and the ground truth bounding box of the i-th target respectively; |X pi ∩X ti | represents the intersection area of the predicted bounding box and the ground truth bounding box of the i-th target; |X pi ∪X ti | represents the union area of the predicted bounding box and the ground truth bounding box of the i-th target, and L IOU_i represents the intersection over union loss of the i-th target.

5. A method for dynamic traffic target detection, tracking and forward collision warning according to claim 1, characterized in that The improved DeepSORT algorithm includes: A tracking feature extraction module, which consists of lightweight convolutional layers and is used to extract identity features from the bounding box of the target vehicle; A motion prediction module, which consists of a Kalman filter and calculates the predicted bounding box position E by estimating the motion state of the target; A matching module, which includes the Hungarian matching algorithm and a cost matrix unit, and associates the detection results of the current frame with the trajectories of the previous frames; A tracking loss function module, which consists of a complete intersection over union loss function to ensure the consistency of the target position in consecutive frames.

6. The dynamic traffic target detection, tracking and forward collision warning method according to claim 5, characterized in that, Input the target vehicle bounding box into the tracking feature extraction module to obtain the identity feature vector d corresponding to each bounding box i , and the specific steps are as follows: d i = Conv(R i ) (7) In the formula, represents a vector space, represents the identity feature vector of target i, D is the dimension of the feature vector, and R i represents the vehicle bounding box of target i; The bounding box of the target vehicle is input into the motion prediction module, and the motion prediction module uses the Kalman filter to estimate the predicted bounding box position E of the target in the next frame. The expression formula is as follows: In the formula, F represents the state transition matrix, X t-1 represents the bounding box of the target vehicle at the previous moment, K represents the Kalman gain, Z t represents the detected bounding box at the current moment, H represents the observation matrix, P t represents the updated covariance matrix, I represents the identity matrix, P t-1 represents the covariance matrix at the previous moment, F T represents the transpose of the state transition matrix, and Q represents the process noise covariance matrix; The predicted bounding box position E and the identity feature vector of the detection bounding box are input into the matching module. The matching module calculates the cost matrix between the two and selects the detection bounding box with the smallest cost matrix as the tracking bounding box. The expression formula is as follows: cost = βcost IOU + βcost dis (9) Where cost is the total cost, cost IOU is the intersection over union cost, cost dis is the distance cost, b t and represent the detected bounding box at the current moment and the predicted bounding box at the previous moment respectively, d t-1 and d t represent the identity feature vectors of the detected bounding box at the previous moment and the target vehicle bounding box at the current moment respectively; The tracking bounding boxes obtained by improving the DeepSORT algorithm are input into the tracking loss function module, which adopts the complete intersection over union loss function L CIoU to calculate the error. By calculating the overlapping area between the tracking bounding box and the ground truth bounding box, and introducing the consistency of the center point position deviation and the aspect ratio of the box body for a more comprehensive match; among them, the complete intersection over union loss function L CIoU The specific calculation formula is: Where N represents the total number of targets, and ρ 2 (b,b gt ) is the Euclidean distance between the center points of the tracking bounding box and the ground truth bounding box. b represents the coordinates of the center point of the tracking bounding box, and b gt represents the coordinates of the center point of the ground truth bounding box; c is the diagonal length of the smallest circumscribed matrix that can accommodate both the tracking bounding box and the ground truth bounding box; α is a weight function; v is a metric aspect ratio consistency index, and the specific calculation method is as follows: where ω gt and h gt are the width and height of the true bounding box, respectively, and ω and h are the width and height of the tracking bounding box, respectively.

Citation Information

Patent Citations

  • Target detection and tracking system based on improved YOLOv7 and DeepSORT

    CN117423031A

  • Method for short-term traffic risk prediction of road sections using roadside observation data

    US20220383738A1