A traffic behavior determination method based on multi-target tracking
By adopting a traffic behavior judgment method based on multi-objective tracking in the intelligent traffic system, using OpenCV tools and Kalman filters and other technologies, the problem of low efficiency in existing systems when dealing with pedestrians and non-motor vehicles is solved, and efficient and accurate traffic behavior judgment is achieved, assisting traffic police in maintaining traffic order.
Patent Information
- Application Number
- CN202210369314.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The existing intelligent transportation system is inefficient in handling traffic behaviors between pedestrians and non-motor vehicles, and it is difficult to meet the timeliness requirements, making it difficult to achieve accurate traffic behavior judgments in traffic monitoring.
Using a traffic behavior determination method based on multi-object tracking, the traffic monitoring video stream is read through OpenCV tool, and using interval detection and a tracking strategy that matches visual and motion information, a single-object tracker based on visual information and a Kalman filter based on motion information is established to realize the tracking of pedestrians and non-motor vehicle targets and the determination of traffic behavior.
It improves the operation efficiency and accuracy of multi-target tracking, and can quickly and accurately determine the traffic behavior of pedestrians and non-motor vehicles in traffic monitoring scenarios, assist traffic police in maintaining traffic order and saving manpower and material resources.
Smart Images

Figure CN114694078B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision and intelligent recognition, and in particular relates to a traffic behavior determination method based on multi-target tracking. Background Art
[0002] With the vigorous development of data communication technology, sensor technology, electronic control technology, and computer vision technology in recent years, urban intelligent transportation systems have been widely covered in cities and play a role all day and all-round. By improving the capacity of the traffic network, investigating and punishing traffic violations, and intelligent dispatching, it can alleviate traffic congestion, reduce air pollution, and reduce the occurrence of traffic accidents. The current existing intelligent transportation system focuses on the management and development of motor vehicles, while ignoring pedestrians, non-motor vehicles, and control. According to statistics, the proportion of traffic accidents involving pedestrians and non-motor vehicles in my country is not high. However, it cannot be ignored that due to the high proportion of motor vehicle accidents, the proportion of serious consequences for pedestrians and non-motor vehicles remains high. For the handling of accidents of the above two types of traffic entities, a large amount of police force is needed to maintain order, but the number of intersections far exceeds the coverage of manpower. Since traffic violations occur from time to time, they are also potential sources of traffic accidents, so further control and governance are needed.
[0003] Multi-target tracking is an important visual technology that can detect the positions of multiple categories of targets of interest and generate trajectories by continuously locating and identifying targets in video sequences. It can be used in video surveillance, human-computer interaction, autonomous driving, etc. With the rapid development of deep learning, the performance of multi-target tracking models has been greatly improved. However, due to the deep network-based target detection and complex data association mechanism, the multi-target tracking model is still relatively lacking in efficiency and is difficult to be used in practical scenarios with high timeliness requirements such as traffic monitoring. Summary of the invention
[0004] The purpose of the present invention is to address the deficiencies of the prior art and provide a traffic behavior determination method based on multi-target tracking. By tracking pedestrians and non-motor vehicle targets in traffic monitoring scenarios, their traffic behaviors can be determined quickly and accurately to assist traffic police in maintaining traffic order.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] Step 1: Input a video of a traffic monitoring scene;
[0007] Step 2: Given a traffic surveillance video, extract a frame from it and track the initial region for pedestrians and non-motor vehicles. s To delineate;
[0008] Step 3: Use OpenCV to read the RTSP protocol address of the traffic monitoring server, obtain the real-time video stream, and track the initial region in the picture. s The specific steps are as follows:
[0009] Step 3.1: Create a temporarily lost target set
[0010] Step 3.2: Read a frame of image from the traffic monitoring video stream in sequence. When the first frame is input, perform pedestrian and non-motor vehicle target detection on the entire image. s A single target tracker based on visual information and a Kalman filter based on motion information are established for the target, and the appearance feature vector of the target is stored;
[0011] Step 3.3: Perform target detection at intervals of a fixed number of frames in the subsequent input frame images. If target detection is not performed in the current frame:
[0012] Step 3.3.1: Traverse all current tracking targets, use the single target tracker based on visual information and the Kalman filter based on motion information to predict the position, and obtain the predicted position Box of the target sot 、Box kal ;
[0013] Step 3.3.2: If Box sot 、Box kal If the overlap coverage (IoU) of is higher than the threshold to achieve a match, then use Box sot Update the state of the Kalman filter as the observation value and obtain the observation value Box of the Kalman filter at this time out The position of the target in the current frame.
[0014] If a target is successfully tracked for a time exceeding the threshold T track , or reaches the boundary of the image, the tracking of the target is stopped.
[0015] Step 3.3.3: If Box sot 、Box kal If the IoU is lower than the threshold and the position cannot be matched, the target is considered to be temporarily lost and added to the set {Target lost}, suspend the operation of the single target tracker of the target and continue to predict using the Kalman filter;
[0016] Step 3.3.4: Loop through all {Target lost}, use the prediction result of the Kalman filter to update its own state. If the loss time exceeds the threshold, the target is deleted.
[0017] Step 3.4: When target detection is required in the current frame:
[0018] Step 3.4.1: Perform target detection and obtain all target box sets {Box det};
[0019] Step 3.4.2: Traverse all current tracking targets, use the single target tracker based on visual information and the Kalman filter based on motion information to predict the position, and obtain the predicted position Box of the target sot 、Box kal , take the average of the two as the predicted value Box of the current target pred ;
[0020] Step 3.4.3: Traverse all Boxes pred , find the detection result {Box det} and Box pred The IoU of the bounding box is greater than the threshold to achieve a matching bounding box. If a match can be achieved, the matching detection result is used as the observation value to update the state of the Kalman filter and obtain the observation value Box of the Kalman filter at this time. out As the position of the target in the current frame, and update the appearance feature vector of the target.
[0021] If a target is successfully tracked for a time exceeding the threshold T track , or reaches the boundary of the image, the tracking of the target is stopped.
[0022] Step 3.4.4: If the Box in step ③ pred If there is no matching detection result, the target is considered temporarily lost and {Target lost}, suspend the single target tracker of the target and retain the Kalman filter of the target;
[0023] Step 3.4.5: Traverse {Target lost}, if they can match the remaining detection results after steps ③ and ④ in terms of position and appearance feature measurement, then the single target tracker of the target is re-established, starting from {Target lost} and treat it as a normal tracking target. If {Target lost If the target in} is lost for more than the threshold, the target is deleted;
[0024] Step 3.4.6: If the test result {Box det} There are still regions in the tracking initial region s If a bounding box is found, a new tracking target is established, a single target tracker based on visual information and a Kalman filter based on motion information are established for the target, and the appearance feature vector of the target is stored.
[0025] Step 3.5: Continue to process the remaining frame images of the traffic monitoring video stream, repeating steps 2.3 and 2.4 until the video ends.
[0026] Step 4: For the target tracked in step 2, determine the traffic behavior based on its starting position and tracking end position, and output the tracking ID, traffic behavior, occurrence time and image of the drawn trajectory of the target. The determination method is as follows:
[0027] Step 4.1: Set a unit vector in a certain direction as a reference vector The unit vector from the target's initial position to the tracking end position is taken as Calculate the value used to determine the direction
[0028] Step 4.2: For pedestrian targets, there are only two traffic behaviors: crossing the road to the left and crossing the road to the right; while for non-motor vehicle targets, there are three traffic behaviors: turning left, going straight, and turning right. d The value of can be used to infer the target motion direction vector With reference vector The angle of the vehicle is determined to determine the traffic behavior.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention is a method for determining traffic behavior based on multi-target tracking, which improves the accuracy of tracking while improving the operational efficiency of multi-target tracking by using a tracking strategy that matches interval detection, vision, and operation prediction. Traffic behavior determination can be completed by tracking pedestrians and non-motor vehicle targets that pass through the starting area for a certain period of time. The present invention optimizes the multi-target tracking process, making it timely in traffic monitoring scenarios, and at the same time, the automated analysis of the traffic behavior of pedestrians and non-motor vehicle targets can save manpower and material resources in traffic control, and better help traffic police investigate and deal with traffic violations. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of the traffic behavior determination method based on multi-target tracking of the present invention;
[0031] Figure 2 This is a schematic diagram of traffic behavior judgment of the present invention;
[0032] Figure 3 It is a network architecture diagram of the target detection module of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. On the contrary, the present invention encompasses any substitutions, modifications, equivalent methods and schemes made on the essence and scope of the present invention as defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. For those skilled in the art, the present invention can be fully understood without the description of these details.
[0034] refer to Figure 1 , is a flow chart of a traffic behavior determination method based on multi-target tracking according to an embodiment of the present invention:
[0035] Step 1: Input a video of a traffic monitoring scene;
[0036] Step 2: Reference Figure 2 , extract a frame of traffic monitoring video and define the initial tracking area region s ={left starting area for pedestrians, right starting area for pedestrians, starting area for non-motor vehicles};
[0037] Step 3: Use OpenCV to read the RTSP protocol address of the traffic monitoring server, read the traffic monitoring video stream, and track the initial region in the picture. s The specific steps are as follows:
[0038] Step 3.1: Create a temporarily lost target set
[0039] Step 3.2: Read a frame of image from the traffic monitoring video stream in sequence. When the first frame is input, perform pedestrian and non-motor vehicle target detection on the entire image. s A twin network single target tracker based on visual information and a Kalman filter based on motion information are established for the target, and the appearance feature vector of the target is stored;
[0040] refer to Figure 3, target detection uses a model based on a deep network. After the image is input, it passes through the feature extraction network to obtain visual features. The visual features are enhanced by multi-scale combination to predict the target position. It includes classification and box regression branches. The appearance feature vector of the target is obtained by aligning the region of interest (ROI Align) and performing global average pooling operations on the visual feature map. The loss function loss of the target detection network includes the cross entropy loss of the classification branch and the IoU loss of the regression branch, specifically:
[0041]
[0042] l reg =1-IoU(box gt ,box pred )
[0043] Among them, l cls is the cross entropy loss of the classification branch, H and W are the width and height of the feature map, and p i is the true value probability of a point on the feature map being a foreground, is the predicted probability that a point on the feature map is a foreground point; l reg is the IoU loss of the regression branch, where box gt is the true value bounding box position, box pred is the location of the predicted bounding box.
[0044] Step 3.3: Perform target detection every 5 frames in the subsequent input frames. If target detection is not performed in the current frame:
[0045] Step 3.3.1: Traverse all current tracking targets, use the single target tracker based on visual information and the Kalman filter based on motion information to predict the position, and obtain the predicted position Box of the target sot 、Box kal ;
[0046] Step 3.3.2: If Box sot 、Box kal If the IoU (Intersection over Union) is higher than the threshold to achieve a match, Box sot Update the state of the Kalman filter as the observation value and obtain the observation value Box of the Kalman filter at this time out The position of the target in the current frame.
[0047] If a target is successfully tracked for a time exceeding the threshold T track , or reaches the boundary of the image, the tracking of the target is stopped.
[0048] Step 3.3.3: If Box sot、Box kal If the IoU is lower than the threshold and the position cannot be matched, the target is considered to be temporarily lost and added to the set {Target lost}, suspend the operation of the single target tracker of the target and continue to predict using the Kalman filter;
[0049] Step 3.3.4: Loop through all {Target lost}, use the prediction result of the Kalman filter to update its own state. If the loss time exceeds the threshold, the target is deleted.
[0050] Step 3.4: When target detection is required in the current frame:
[0051] Step 3.4.1: Perform target detection and obtain all target box sets {Box det};
[0052] Step 3.4.2: Traverse all current tracking targets, use the single target tracker based on visual information and the Kalman filter based on motion information to predict the position, and obtain the predicted position Box of the target sot 、Box kal , take the average of the two as the predicted value Box of the current target pred ;
[0053] Step 3.4.3: Traverse all Boxes pred , find the detection result {Box det} and Box pred The IoU of the bounding box is greater than the threshold to achieve a matching bounding box. If a match can be achieved, the matching detection result is used as the observation value to update the state of the Kalman filter and obtain the observation value Box of the Kalman filter at this time. out As the position of the target in the current frame, and update the appearance feature vector of the target:
[0054] feature new =(1-θ)*feature old +θ*feature curr
[0055] Among them, feature old is the old appearance feature vector of the target, feature curr is the appearance feature vector of the detection result that matches the current frame, feature new is the updated feature vector, and θ is the updated learning rate. If a target is successfully tracked for more than the threshold T track , or reaches the boundary of the image, the tracking of the target is stopped.
[0056] Step 3.4.4: If the Box in step ③ pred If there is no matching detection result, the target is considered temporarily lost and {Target lost}, suspend the single target tracker of the target and retain the Kalman filter of the target;
[0057] Step 3.4.5: Traverse {Target lost}, if they can match the remaining detection results after steps ③ and ④ in terms of position and appearance feature measurement, then the single target tracker of the target is re-established, starting from {Target lost} and treat it as a normal tracking target. If {Target lost If the target in} is lost for more than the threshold, the target is deleted;
[0058] Step 3.4.6: If the test result {Box det} There are still regions in the tracking initial region s If a bounding box is found, a new tracking target is established, a single target tracker based on visual information and a Kalman filter based on motion information are established for the target, and the appearance feature vector of the target is stored.
[0059] Step 3.5: Continue to process the remaining frame images of the traffic monitoring video stream, repeating steps 2.3 and 2.4 until the video ends.
[0060] Step 4: Reference Figure 2 Schematic diagram of traffic behavior determination. For the target tracked in step 2, the traffic behavior is determined according to its starting position and tracking end position, and the tracking ID, traffic behavior, occurrence time and image of the drawn trajectory of the target are output. The determination method is as follows:
[0061] Step 4.1: Set the unit vector rotated 30° clockwise from the direction of the road facing the camera as the reference vector The unit vector from the target's initial position to the tracking end position is taken as Calculate the value used to determine the direction
[0062] Step 4.2: For pedestrian targets, there are only two traffic behaviors: crossing the road to the left and crossing the road to the right: If value d ≥0, the traffic behavior is right, if value d <0, the traffic behavior is left;
[0063] Step 4.3: For non-motorized vehicles, there are three traffic behaviors: left turn, straight ahead, and right turn: If -0.707≤value d<0.707, the traffic behavior is straight ahead. If value d ≥0.707, the traffic behavior is right turn, if value d <-0.707, the traffic behavior is left turn.
[0064] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A traffic behavior determination method based on multi-target tracking, characterized by , specifically including the following steps: (1) First, a traffic surveillance video is given; (2) Extract a frame of traffic surveillance video from this traffic surveillance video and track the initial region of pedestrians and non-motor vehicles in the video. s To delineate; (3) Use OpenCV tools to request the RTSP protocol address of the traffic monitoring server, read the real-time video stream, and use multi-target tracking technology to track the initial region in the picture. s Pedestrians and non-motor vehicles are tracked. When the target is tracked for more than the time threshold T track or when the target reaches the border of the screen, the tracking of the target is terminated; the multi-target tracking technology is used to track the initial region region in the screen. s Pedestrians and non-motor vehicles are tracked. When the target is tracked for more than the time threshold T track Or when the screen boundary is reached, the tracking of the target is ended; specifically: (3.1) Read a frame of image from the traffic monitoring video stream in sequence. When the first frame is input, detect pedestrians and non-motor vehicles. s The target is modeled based on visual information and motion information, and the appearance feature representation of the target is stored; wherein the modeling based on visual information is to establish the correlation between the appearance feature representation of the target and the spatial position, and the modeling based on motion information is to establish a prediction model based on the target movement speed and size change; (3.2) Target detection is performed at a fixed interval of frames in the frames input thereafter. If target detection is not performed in the current frame, the position of the target being tracked in the previous frame is predicted based on the visual information and the motion information, and the position predictions of the two are fused as the tracking result; The position prediction of the target being tracked in the previous frame is performed based on the visual information and the motion information, and the position predictions of the two are fused as the tracking result, specifically: Traverse all current tracking targets, use the single target tracker based on visual information and the Kalman filter based on motion information to predict the position, and obtain the predicted position of the target , ; Step 3.3.2: If , If the overlap coverage IoU of is higher than the threshold to achieve a match, then use Update the state of the Kalman filter as the observation value and obtain the observation value of the Kalman filter at this time As the position of the target in the current frame; (3.3) When target detection is required in the current frame, the position prediction of the target being tracked in the previous frame based on visual information and motion information is obtained, and the prediction result is corrected based on the target bounding box intersection over union (IoU) using the tracking result of the target detection. If there is no tracked target in the detection result, the operation in step (2) is performed again; (3.4) Continue to process the remaining frame images of the traffic monitoring video stream, repeating steps (3.2) and (3.3) until the video ends; (4) For the target tracked in step (3), the traffic behavior is determined based on its starting position and tracking end position, and the tracking ID, traffic behavior, occurrence time and image of the drawn trajectory of the target are output.
2. The traffic behavior determination method based on multi-target tracking according to claim 1 is characterized in that: In addition to outputting the location of the target of the class of interest, the target detection model also outputs a representation of the appearance features of each target.
3. According to claim 1, a traffic behavior determination method based on multi-target tracking is characterized in that: The traffic behavior determination method described in step (4) is based on the target motion direction vector With reference vector The angle is obtained, and the specific determination method is as follows: (4.1) Set the unit vector in a certain direction as the reference vector , with the unit vector from the target's initial position to the tracking end position as , calculate the value used to determine the direction ; (4.2) According to the value of the direction d , infer the target motion direction vector With reference vector The angle between the two sides is used as the basis for determining the target traffic behavior.
Citation Information
Patent Citations
Vehicle illegal act detection method and device as well as computer equipment
CN110459064A
Online multi-pedestrian detection tracking method in complex scene
CN111739053A