Multi-target tracking method based on traditional and deep learning algorithms
By introducing an improved optical flow estimation-region matching method in the multi-objective tracking method, combining deep learning and Kalman filtering, the problem of Kalman filtering trajectory prediction deviation is solved, and the accuracy of tracking results is significantly improved.
Patent Information
- Application Number
- CN202111349149.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-15
AI Technical Summary
Existing multi-objective tracking methods are prone to major deviations in the prediction of Kalman filtering trajectory, resulting in a decrease in the accuracy of the result.
The accuracy of trajectory prediction is improved by adding an improved optical flow estimation-region matching method to the multi-objective tracking method, combining deep learning algorithms and traditional Kalman filtering.
It effectively avoids the deviation of Kalman filter trajectory prediction and improves the accuracy of multi-objective tracking results.
Smart Images

Figure CN114092517B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-target tracking method based on traditional and deep learning algorithms. Background Art
[0002] Computer vision has always been a hot topic of research. As a part of computer vision, multi-target tracking has also been studied by scholars at home and abroad. By studying the papers in top conference journals, it is found that most scholars currently use Kalman filtering and Hungarian algorithm to establish tracking models. The advantage of this combination is that the detection speed is improved, but the target tracking phenomenon is not effectively improved. The present invention effectively avoids the situation where the extended Kalman filter has a serious deviation in the predicted trajectory by adding a traditional algorithm - streamer estimation. Summary of the invention
[0003] 1. Technical issues to be resolved
[0004] In view of the shortcomings of the prior art, the present invention provides a multi-target tracking method based on traditional and deep learning algorithms. By adding an improved optical flow algorithm, a solution is provided to prevent significant deviations in Kalman filter trajectory prediction, which would lead to a decrease in the accuracy of the results.
[0005] (II) Technical solution
[0006] To achieve the above objectives, the present invention is implemented by the following technical solutions: a multi-target tracking method based on traditional and deep learning algorithms, specifically including:
[0007] The target detector is used to detect the target in the video frame and extract the detected target features. The appearance features can avoid missing the target and the ability to handle obstacles. The motion features mainly rely on algorithms such as streamer and Kalman to predict the target trajectory, and then use cascade matching and IOU matching to assign an ID to the target.
[0008] The method specifically comprises:
[0009] S1, collect data and pre-process it;
[0010] S2, use yolov5 structure network to train the above data set to obtain the optimal weight model;
[0011] S3. Improve and optimize the tracking model;
[0012] S4, combining the model obtained in step 2 with the model in step 3 to generate a multi-tracking model;
[0013] S5. Perform real-time verification on the tracking model.
[0014] Preferably, the step 1 specifically includes:
[0015] By collecting complex intersection videos, they are preprocessed as a dataset, and the dataset is divided into training, verification and testing according to a certain ratio.
[0016] Preferably, in step 2, the yolov5 structure network is used to train the above data set to obtain the optimal weight model.
[0017] Preferably, in step 3, deepsort is improved, and the specific method is as follows:
[0018] In terms of trajectory prediction, an improved optical flow estimation-region matching method is added. This method is to define the speed Vm as the disparity d = (dx, dy) T , so that the image area at two moments matches best. In order to obtain the accuracy of sub-pixels, a surface is used to fit the similarity around the obtained d to find the maximum value. To solve the problem of large motion, a strategy from coarse to fine is adopted;
[0019] Consider the image at pixel M = (x, y) T , the gray value at time t is I(x,y,t), let the velocity of point m be Vm=(Vx,Vy) T ,Assuming that the grayscale of point m remains unchanged, when the time interval dt is very short, the constraint equation for optical flow estimation is:
[0020] I(x,y,t)=I(x+Vxdx,y+Vydy,t+dt);
[0021] Transform the equation:
[0022]
[0023] in is the gradient of point m;
[0024] Area matching method: velocity Vm is defined as parallax d = (dx, dy) T In addition, the extended Kalman filter is still used in deepsort. The state estimation at a certain moment is (u, v, r, h, u1, v1, r1, h1), where (u, v, r, h) are the variables at the current moment, and (u1, v1, r1, h1) are the predicted variables. The Hungarian algorithm is used to solve the allocation problem. A group of detection frames and Kalman predicted frames are allocated, so that the Kalman predicted frame can find the detection frame that best matches itself to achieve the tracking effect.
[0025] Preferably, in step 4, the model obtained in step 2 is combined with the improved deepsort model to generate a multi-tracking model. The specific process is as follows:
[0026] 1) Create Tracks, initialize the motion variables of optical flow and Kalman filter, predict its frame through optical flow and Kalman filter, calculate IOU between the two frames, if it is greater than the threshold, proceed to step 5), otherwise proceed to step 2);
[0027] 2) Perform IOU matching between the frame and the frame predicted by Tracks in the previous frame, and then calculate the cost matrix based on the result of IOU matching;
[0028] 3) All the cost matrices obtained in 2) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch, and the mismatched Tracks are directly deleted; the second is Detections mismatch, and such Detections are initialized as a new Track; the third is that the detection box and the predicted box are successfully paired, indicating that the previous frame and the next frame are tracked successfully, and the corresponding Detections are updated with their corresponding Tracks variables through Kalman filtering;
[0029] 4) Repeat steps 2)-3) until a confirmed track appears or the video frame ends;
[0030] 5) Cascade matching is performed between the two algorithms’ predicted boxes whose IOU is greater than the threshold and the boxes of confirmed Tracks and Detections;
[0031] 6) There are three possible results after cascade matching. The first is Tracks matching. Such Tracks update their corresponding Tracks variables through optical flow and Kalman filtering. The second and third are Detections and Tracks mismatch. At this time, the previously unconfirmed Tracks and mismatched Tracks are combined with the Unmatched Detections for IOU matching one by one, and then the cost matrix is calculated based on the results of IOU matching.
[0032] 7) All the cost matrices obtained in 6) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch. We directly delete the mismatched Tracks. The second is Detections mismatch. Such Detections are initialized as a new Track. The third is that the detection box and the predicted box are successfully paired, indicating that the previous frame and the next frame are tracked successfully. The corresponding Detections are updated with their corresponding Tracks variables through Kalman filtering.
[0033] 8) Repeat steps 5)-7) until the video ends.
[0034] (III) Beneficial effects
[0035] The present invention provides a multi-target tracking method based on traditional and deep learning algorithms. It has the following beneficial effects:
[0036] The present invention adopts a robust optical flow algorithm to avoid the significant deviation of Kalman filter trajectory prediction, which leads to a decrease in the accuracy of the results. Through the trained yolov5 detection model, the real-time verification is carried out through the deepsort tracking model, and it is found that the accuracy of tracking the target is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is the overall flow chart of the present invention;
[0038] Figure 2 is a flow chart of the tracking model of the present invention;
[0039] Figure 3 It is a detection flow chart of the present invention. DETAILED DESCRIPTION
[0040] The following will be described clearly and completely in conjunction with the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Example:
[0042] refer to Figure 1-3 As shown, the embodiment of the present invention provides a multi-target tracking method based on traditional and deep learning algorithms, specifically including:
[0043] The target detector is used to detect the target in the video frame and extract the detected target features. The appearance features can avoid missing the target and the ability to handle obstacles. The motion features mainly rely on algorithms such as streamer and Kalman to predict the target trajectory, and then use cascade matching and IOU matching to assign an ID to the target.
[0044] The method specifically includes:
[0045] S1, collect data and pre-process it;
[0046] Specifically include:
[0047] By collecting complex intersection videos, preprocessing them as data sets, and dividing the data sets into training, verification and testing according to a certain ratio;
[0048] S2, use yolov5 structure network to train the above data set to obtain the optimal weight model;
[0049] S3. Improve and optimize the tracking model;
[0050] The specific method is as follows:
[0051] In terms of trajectory prediction, an improved optical flow estimation-region matching method is added. This method is to define the speed Vm as the disparity d = (dx, dy) T , so that the image area at two moments matches best. In order to obtain the sub-pixel accuracy, a surface is used to fit the similarity around the obtained d to find the maximum value. To solve the problem of large motion, a coarse-to-fine strategy is adopted;
[0052] Consider the image at pixel M = (x, y) T , the gray value at time t is I(x,y,t), let the velocity of point m be Vm=(Vx,Vy) T ,Assuming that the grayscale of point m remains unchanged, when the time interval dt is very short, the constraint equation for optical flow estimation is:
[0053] I(x,y,t)=I(x+Vxdx,y+Vydy,t+dt);
[0054] Transform the equation:
[0055]
[0056] in is the gradient of point m;
[0057] Area matching method: velocity Vm is defined as parallax d = (dx, dy) T In addition, the extended Kalman filter is still used in deepsort. The state estimation at a certain moment is (u, v, r, h, u1, v1, r1, h1), where (u, v, r, h) is the current moment variable and (u1, v1, r1, h1) is the prediction variable. The Hungarian algorithm is used to solve the allocation problem. A group of detection frames and Kalman predicted frames are allocated, so that the Kalman predicted frame can find the detection frame that best matches itself to achieve the tracking effect.
[0058] S4, combining the model obtained in step 2 with the model in step 3 to generate a multi-tracking model;
[0059] The specific process is:
[0060] 1) Create Tracks, initialize the motion variables of optical flow and Kalman filter, predict its frame through optical flow and Kalman filter, calculate IOU between the two frames, if it is greater than the threshold, proceed to step 5), otherwise proceed to step 2);
[0061] 2) Perform IOU matching between the frame and the frame predicted by Tracks in the previous frame, and then calculate the cost matrix (the calculation method is 1-IOU) based on the result of IOU matching;
[0062] 3) All the cost matrices obtained in 2) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch (Unmatched Tracks), and the mismatched Tracks are directly deleted (because this Track is in an uncertain state. If it is in a certain state, it must be deleted after a certain number of consecutive times (the default is 30 times)); the second is Detections mismatch (Unmatched Detections), and such Detections are initialized as a new Tracks (new Tracks); the third is that the detection box and the predicted box are paired successfully, indicating that the previous frame and the next frame are tracked successfully, and the corresponding Detections are updated with their corresponding Tracks variables through Kalman filtering;
[0063] 4) Repeat steps 2)-3) until confirmed tracks appear or the video frame ends;
[0064] 5) Cascade matching is performed between the two algorithm prediction boxes whose IOU is greater than the threshold and the boxes of confirmed Tracks and Detections (previously, the appearance features and motion information of Detections are saved every time Tracks are matched. By default, the first 100 frames are saved, and the appearance features and motion information are used to perform cascade matching with Detections. This is done because confirmed Tracks and Detections are more likely to match).
[0065] 6) There are three possible results after cascade matching. The first is Tracks matching. Such Tracks update their corresponding Tracks variables through optical flow and Kalman filtering. The second and third are Detections and Tracks mismatch. At this time, the previously unconfirmed Tracks and mismatched Tracks are matched with the Unmatched Detections one by one by IOU matching, and then the cost matrix (the calculation method is 1-IOU) is calculated based on the results of IOU matching.
[0066] 7) All the cost matrices obtained in 6) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch (Unmatched Tracks). We directly delete the mismatched Tracks (because this Track is in an uncertain state. If it is in a certain state, it must be deleted after a certain number of consecutive times (the default is 30 times); the second is Detections mismatch (Unmatched Detections). Such Detections are initialized as a new Tracks (new Tracks); the third is that the detection box and the predicted box are paired successfully, indicating that the previous frame and the next frame are tracked successfully, and the corresponding Detections are updated through the Kalman filter to their corresponding Tracks variables;
[0067] 8) Repeat steps 5)-7) until the video ends;
[0068] S5. Perform real-time verification on the tracking model.
[0069] The present invention adopts a robust optical flow algorithm to avoid the significant deviation of Kalman filter trajectory prediction, which leads to a decrease in the accuracy of the results. Through the trained yolov5 detection model, the real-time verification is carried out through the deepsort tracking model, and it is found that the accuracy of tracking the target is greatly improved.
[0070] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. Multi-target tracking methods based on traditional and deep learning algorithms, with specific characteristics include: The target detector is used to detect the target in the video frame and extract the detected target features. The appearance feature can avoid the loss of the target and the ability to handle obstacles. The motion feature relies on the streamer and Kalman algorithm to predict the target trajectory, and then cascade matching and IOU matching are used to finally assign an ID to the target. The method specifically comprises: S1, collect data and pre-process it; S2, use yolov5 structure network to train the above data set to obtain the optimal weight model; S3. Improve and optimize the tracking model; S4, combining the model obtained in step 2 with the model in step 3 to generate a multi-tracking model; S5. Verify the real-time performance of the tracking model; In step 3, deepsort is improved, and the specific method is as follows: In terms of trajectory prediction, an improved optical flow estimation-region matching method is added. This method is to define the speed Vm as the disparity d = (dx, dy) T , so that the image area at two moments matches best. In order to obtain the accuracy of sub-pixels, a surface is used to fit the similarity around the obtained d to find the maximum value. To solve the problem of large motion, a strategy from coarse to fine is adopted; Consider the image at pixel M = (x, y) T , the gray value at time t is I(x,y,t), let the velocity of point m be Vm=(Vx,Vy) T ,Assuming that the grayscale of point m remains unchanged, when the time interval dt is very short, the constraint equation for optical flow estimation is: I(x,y,t)=I(x+Vxdx,y+Vydy,t+dt); Transform the equation: in is the gradient of point m; Area matching method: velocity Vm is defined as parallax d = (dx, dy) T In addition, the extended Kalman filter is still used in deepsort. The state estimation at a certain moment is (u, v, r, h, u1, v1, r1, h1), where (u, v, r, h) are the variables at the current moment, and (u1, v1, r1, h1) are the predicted variables. The Hungarian algorithm is used to solve the allocation problem. A group of detection frames and Kalman predicted frames are allocated, so that the Kalman predicted frame can find the detection frame that best matches itself to achieve the tracking effect.
2. The multi-target tracking method based on traditional and deep learning algorithms according to claim 1, characterized in that: The step 1 specifically includes: By collecting complex intersection videos, they are preprocessed as a dataset, and the dataset is divided into training, verification and testing according to a certain ratio.
3. The multi-target tracking method based on traditional and deep learning algorithms according to claim 1, characterized in that: In step 2, the YOLOV5 structure network is used to train the above data set to obtain the optimal weight model.
4. The multi-target tracking method based on traditional and deep learning algorithms according to claim 1, characterized in that: In step 4, the model obtained in step 2 is combined with the improved deepsort model to generate a multi-tracking model. The specific process is as follows: 1) Create Tracks, initialize the motion variables of optical flow and Kalman filter, predict its frame through optical flow and Kalman filter, calculate IOU between the two frames, if it is greater than the threshold, proceed to step 5), otherwise proceed to step 2); 2) Perform IOU matching between the frame and the frame predicted by Tracks in the previous frame, and then calculate the cost matrix based on the result of IOU matching; 3) All the cost matrices obtained in 2) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch, and the mismatched Tracks are directly deleted; the second is Detections mismatch, and such Detections are initialized as a new Track; the third is that the detection box and the predicted box are successfully paired, indicating that the previous frame and the next frame are tracked successfully, and the corresponding Detections are updated with their corresponding Tracks variables through Kalman filtering; 4) Repeat steps 2)-3) until a confirmed track appears or the video frame ends; 5) Cascade matching is performed between the two algorithms’ predicted boxes whose IOU is greater than the threshold and the boxes of confirmed Tracks and Detections; 6) There are three possible results after cascade matching. The first is Tracks matching. Such Tracks update their corresponding Tracks variables through optical flow and Kalman filtering. The second and third are Detections and Tracks mismatch. At this time, the previously unconfirmed Tracks and mismatched Tracks are combined with the Unmatched Detections for IOU matching one by one, and then the cost matrix is calculated based on the results of IOU matching. 7) All the cost matrices obtained in 6) are used as the input of the Hungarian algorithm to obtain the linear matching results. There are three types of results. The first is Tracks mismatch. We directly delete the mismatched Tracks. The second is Detections mismatch. Such Detections are initialized as a new Track. The third is that the detection box and the predicted box are successfully paired, indicating that the previous frame and the next frame are tracked successfully. The corresponding Detections are updated with their corresponding Tracks variables through Kalman filtering. 8) Repeat steps 5)-7) until the video ends.
Citation Information
Patent Citations
Target tracking method and device, storage medium and intelligent video system
CN112926410A
Multi-target tracking detection method, system and device and storage medium
CN113158995A