Long-Term Video Object Tracking Method Based on Siamese Network for Joint Tracking and Detection
Through the method of joint tracking and detection of twin networks, the TDS module is used to determine whether the target disappears or reappears, and the tracking and detector are used alternately, the problem of target disappearing in long-term video target tracking is solved, and the tracking accuracy and success rate are improved.
Patent Information
- Application Number
- CN202310546720.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-05-16
AI Technical Summary
In long-term video target tracking tasks, the target cannot be discovered in time and repositioned accurately after disappearing, resulting in tracking failure and reducing the accuracy and success rate of the tracking algorithm.
Using a joint tracking and detection strategy based on twin networks, the TDS module is used to determine whether the target disappears or reappears, and alternately use the tracker and detector to avoid using both in the same frame to ensure the tracking speed.
It improves the accuracy and robustness of long-term tracking of video targets, avoids tracking failure caused by the disappearance of targets, and improves the success rate and accuracy of the tracker.
Smart Images

Figure CN116664623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video object tracking, and specifically to a video object tracking method based on the joint tracking and detection of a Siamese network. Background Art
[0002] Video object tracking technology predicts the position and scale of the bounding box of the same object in subsequent frames based on the bounding box information of any object to be tracked given in the first frame of the video sequence, and is widely used in fields such as autonomous driving, video surveillance, and human-computer interaction. Traditional correlation filter-based methods use handcrafted features to establish a filtering template and update it online, such as Histogram Of Oriented Gradient (HOG), Haar-like features, and Local Binary Pattern (LBP). First, a series of candidate boxes are given, and then all candidate boxes are correlated with the filtering template to obtain the confidence of each candidate box. The candidate box with the highest confidence is the target position.
[0003] In recent years, with the rapid improvement of computer performance and the rapid development of deep learning technology, deep features have been applied to the field of object tracking. Although the tracking accuracy has been improved, the computational complexity of the backpropagation process of the deep network is huge, resulting in a sharp increase in computational complexity and unable to meet the real-time requirements. The emergence of the Siamese network has well balanced the tracking accuracy and speed. The object tracking algorithm based on the Siamese network has become an important research direction in the field of video object tracking. In deep learning-based object tracking algorithms, the object tracking algorithm based on the Siamese network uses the Siamese network structure to establish an object tracking model, adjusts parameters and trains a suitable tracker to complete the object tracking task, and has good performance. The Siamese network refers to two parallel networks with the same or similar structure, having a template and a search branch in a Y-shaped structure. The biggest feature is sharing weights, that is, when the two networks are reasoning, different inputs pass through the same feature extraction network. The higher the similarity between the inputs, the more similar the obtained features. Related algorithms based on the Siamese network have been applied to tasks such as template matching and similarity measurement as early as the 1990s. Since the video object tracking task can be regarded as a continuously executed template matching task, it becomes possible to use the Siamese network to perform the object tracking task.
[0004] As the short-term target tracking task gradually reaches saturation, the long-term target tracking task has gradually come into view. In practical applications, long-term target tracking tasks are more common than short-term target tracking tasks. The challenge of long-term tracking lies in that, due to the long video sequence, the appearance features of the target will constantly change. After tracking for a period of time, the features of the target will be very different from the initial state. In addition, the constantly moving target will frequently be occluded, disappear, have a cluttered background, or undergo large deformations. The occurrence of these situations has significantly increased the difficulty of long-term target tracking. Among these challenges, target disappearance is one of the most challenging problems. Target disappearance means that within a certain period of time, ① the target is occluded by other objects or ② moves out of the camera's field of view. When the target reappears, the tracker needs to promptly detect and re-locate the target's position. If the target disappears for a long time, the features such as the position, scale, shape, and color of the target will generally change significantly, resulting in the tracker being unable to effectively update and thus tracking failure. In this case, re-locating the target will be a huge challenge. In actual tracking tasks, the situation of target disappearance often occurs in long-term tracking tasks. Compared with short-term tracking tasks, long-term tracking is more likely to encounter challenges such as occlusion, deformation, and cluttered backgrounds. When performing a long-term tracking task, when the target disappearance situation occurs, if no countermeasures are taken in time, the target cannot be re-tracked to its position when it reappears, which will lead to a large amount of incorrect tracking data, significantly reducing the accuracy and success rate of the tracking algorithm.
[0005] Currently, a technical problem that urgently needs to be solved by those skilled in the art is: how to promptly detect and reasonably respond when the target disappears, and accurately locate the target's position in a timely manner when the target reappears, so as to improve the accuracy and robustness of the tracking algorithm. Summary of the Invention
[0006] Aiming at the defects of the prior art, the present invention provides a long-term visual tracking method for video targets based on Siamese network joint tracking and detection (ALong-Term Visual Tracking Algorithm Based on Siamese NetworkUsing Joint Tracking and Detection Strategy, SiamTD). Using the above method can avoid the problem of tracking failure caused by target disappearance and improve the tracking accuracy and success rate.
[0007] To achieve the above object, the technical solution adopted by the present invention is: a long-term visual tracking method for video targets based on Siamese network joint tracking and detection, including the following steps:
[0008] S1. Crop the template Z according to the first-frame input image I of the video sequence and the bounding box information B, and according to the second-frame input image I iCrop out the search area X i , i ∈ [2, n], where n is the total number of frames in the video sequence;
[0009] S2. Send Z and X i into the pre-trained offline Siamese network to extract features, obtaining feature φ(Z) and φ(X i ); Pass the features through the RPN network respectively. The classification branch and the regression branch output two response maps of 17×17×10 and 17×17×20 respectively, denoted as S1 and S2;
[0010] S3. Input S1 and S2 into the TDS module to determine whether the target has disappeared; if the target exists, output the T signal, indicating that the tracker will continue to be used in the next frame; if the target has disappeared, output the D signal, indicating that the detector will be used in the next frame.
[0011] S4. If it is determined in step S3 that the target exists, add a cosine window and scale penalty to the response map S1 to limit large displacements, take the index at the position with the maximum response value, and find the data corresponding to the same index in S2, and convert it into the position and scale of the new prediction box, which is the tracking result of the current frame.
[0012] S5. If the tracking signal is D at the start of tracking in a certain frame, it means that the target disappeared in the previous frame. To determine whether the target reappears, the detector needs to be used in the current frame; zoom the image and divide it into 480×640×3 and then input it into the detector to obtain three features; input the three features into the TDS module. If the target appears, output the T signal and perform non-maximum suppression (NMS) operation on the three features to eliminate redundant candidate boxes, output the detection result, calculate the cosine similarity with the template, and take the one with the maximum similarity as the target to be tracked, output the T signal and start using the tracker from the next frame; if the target does not appear, output the D signal and skip the NMS operation, directly enter the next frame and continue to use the detector.
[0013] Furthermore, the Siamese network mentioned in step S2 has two major branches: a template branch and a search branch. The network structures of both major branches adopt the modified AlexNet, and the network parameters are shared.
[0014] Furthermore, the TDS module in step S3 is used to determine whether the target in the current frame has disappeared and use different tracking methods for different situations. The specific implementation steps are as follows:
[0015] Input the S1 and S2 obtained in S2 into the TDS module. Extract the response map storing the target probability from the output results of each anchor box in S1 to obtain a score map of 17×17×5. Perform a Global MaxPooling (GMP) operation on the score map to find the part with the maximum response as the region of interest. If the target score in this region exceeds the threshold, it is considered that there is a target to be tracked in the current frame. The TDS module will output a T signal and execute step S4; if the result is less than the threshold, it is determined that there is no target in the current frame. After outputting a D signal, directly go to the next frame and use the detector to find the target.
[0016] Among them, the tracker threshold in step S3 is used to judge whether there is a target in the current frame. It is set as a 5-dimensional column vector, and the specific values are [0.648, 0.523, 0.5, 0.523, 0.648].
[0017] Furthermore, the steps of the TDS module in step S5 when using the detector are as follows:
[0018] S5.1. When extracting the response map of the detector, only extract the feature layers recording the responses and confidences of specific categories. For example, if the target of the current sequence is the "bird" category, only extract the feature layer recording the classification scores of "bird" and the confidence layer recording the probability of the existence of a target in the current region; since the target category does not change in the same sequence, multiply the classification score matrix and the corresponding confidence matrix element by element as the final score.
[0019] S5.2. Perform a GMP operation on the extracted response map. If the obtained result is less than the threshold, it is determined that there is no target, output a D signal and start from the next frame; if it is greater than the threshold, use the NMS operation to obtain the detection result, calculate the cosine similarity between all suspected targets and the template image, and select the detection with the maximum similarity as the tracking result, output a T signal, indicating that the tracker will be used in the next frame.
[0020] The calculation process of the cosine similarity is as follows. Calculate the similarity S between the i-th detection and the template i , scale the i-th detection to the size of the template and convert it to a grayscale image, and flatten it into a one-dimensional column vector, denoted as D i , R is the one-dimensional column vector obtained by converting the template image to a grayscale image and flattening it, ||D i || and ||R|| are the two-norms of the two; cosine similarity:
[0021]
[0022] Beneficial effects: The video object tracking method provided by the present invention is based on a twin network combined tracking algorithm and a detection algorithm. In order to avoid tracking failure caused by the disappearance of the object in a long-term tracking task, a target tracking strategy combining tracking and detection is proposed. The TDS module is used to determine whether the object disappears or reappears, and the tracker and the detector are alternately used. To ensure the tracking speed, the tracker and the detector are not used simultaneously in the same frame. Using the method of the present invention can accurately and timely determine whether the object in the video sequence disappears and take corresponding countermeasures, improving the accuracy and robustness of long-term video object tracking. Description of the Drawings
[0023] Figure 1 It is a schematic diagram of the network structure of the tracking algorithm in the present invention;
[0024] Figure 2 It is a schematic diagram of the structure of the tracker used in the present invention;
[0025] Figure 3 It is a schematic diagram of the structure of the TDS module used in the present invention;
[0026] Figure 4 It is an explanation of the parameters of each layer in the twin network;
[0027] Figure 5 It is a comparison chart of the long-term tracking performance of the method (SiamTD) of the present invention and some methods provided by the official in the OxUvA dataset simulation experiment;
[0028] Figure 6 It is a comparison chart of the accuracy and success rate of the method (SiamTD) of the present invention and some other methods in the UAV20L dataset simulation experiment. Detailed Embodiment
[0029] The following further describes the present invention in detail with reference to the drawings and specific embodiments.
[0030] The present invention provides a long-term visual tracking method for video objects based on Siamese network using joint tracking and detection strategy (SiamTD). When the object to be tracked in the video sequence is completely occluded or leaves the field of view, that is, when the object disappears, traditional object tracking algorithms based on Siamese network cannot locate the reappearing object. After the Tracker or Detector Switch Module (TDS) of the SiamTD of the present invention determines that the object has disappeared, it selects to use an object detector to perform full-image detection. When the object reappears, the detector gives all objects of the same class, and the object to be tracked is found by comparing the similarity with the template, and the tracker is restarted. Using this method can avoid the problem of tracking failure caused by object disappearance and improve the tracking accuracy and success rate.
[0031] As Figure 1-4 shown, a long-term visual tracking method for video objects based on Siamese network using joint tracking and detection strategy specifically includes the following steps S1 to S5.
[0032] S1. Crop the template Z according to the first-frame input image I and the bounding box information B of the video sequence, and crop the search region X according to the second-frame input image I i ; for i ∈ [2, n], where n refers to the total number of frames of the video, that is, starting from the second frame, the search region is cropped in the same way for each frame; i
[0033] S2. Feed Z and X i into the pre-trained Siamese network offline to extract features, obtaining the features φ(Z) and φ(X i ); pass the features through the RPN network respectively, and the classification branch and the regression branch output two response maps of 17×17×10 and 17×17×20 respectively, denoted as S1 and S2;
[0034] In step S2, the Siamese network has two major branches: a template branch and a detection branch. The network structures of the two major branches both adopt the modified AlexNet (the Alex network is a convolutional neural network structure proposed by Alex Krizhevsky et al. in 2012. We modified it on this basis, removed the fully connected layer and padding operation in the original network structure, and adjusted the network stride to 8 to obtain a larger receptive field to meet the requirements of this method), and the network parameters are shared. The network structure is shown in Figure 1 and the network parameters can be referred to Figure 4 . The specific training steps are as follows:
[0035] S2.1. Preprocess the ILSRVC2015 dataset. Take two frames with an interval of t from the same video sequence, where t ranges from 1 to 5. According to the annotation information, crop the two frames of pictures centered on the target to sizes of 127×127 and 255×255 respectively, denoted as Z and X, which are used as the inputs of the template branch and the search branch.
[0036] S2.2. Feed the two processed frames of pictures Z and X obtained from S2.1 into the Siamese network for feature extraction to obtain feature maps φ(Z) and φ(X) of 6×6×256 and 22×22×256, and send them into the RPN network.
[0037] S2.3. The RPN network is divided into two parts: a classification branch and a regression branch. In the classification branch, after convolution operation on φ(Z), the number of channels increases from 256 to 256×2k, and the size becomes 4×4×(2k×256), where k is the number of anchor boxes. In this paper, k = 5 is selected, and the aspect ratios are (0.33, 0.5, 1, 2, 3). 2k represents that the anchor boxes only distinguish the target from the background. After passing through the convolution kernel, the number of channels of φ(X) remains unchanged, and the size of the feature map becomes 20×20×256. As for the regression branch, after convolution operation, the size of the φ(X) feature map becomes 20×20×256, and φ(Z) becomes 4×4×(4k×256). Among them, 4k represents that each anchor box requires four data, namely the center offsets x and y and the scales w and h. The trained regression branch can more accurately describe the position of the target.
[0038] S2.4. Respectively use the template feature maps in the classification and regression branches as convolution kernels to convolve with the search feature maps. Perform a softmax operation in the classification branch to obtain a response map of 17×17×2k, which represents whether each anchor box in each area of the search area is the target or the background, denoted as S1. The regression branch can obtain a response map of 17×17×4k, which represents the relative position and size of each anchor box and the bounding box in the search area, denoted as S2.
[0039] S2.5. Calculate the classification loss L cls : Generate a matrix of size 17×17 as the sample label G1 according to the marking information of the input picture. Each element in the matrix is {+1, -1}, representing positive and negative samples. Those within a certain distance from the target center are set as positive samples, and vice versa. Normalize the response map S1 obtained in step S2.4 to S1′, and use G1 and S1′ as the two inputs of the Binary Cross Entropy loss function. The loss function is defined as follows:
[0040] l(y, x) = l + g(1 + exp(-yx))
[0041]
[0042] Where y is the sample label, which is an element in the label matrix G1 of size 17×17, and its value is {+1, -1}; x represents an element in the response map S1′; D represents the overall sample space included in the normalized response map S1′; u represents the position index of x in S1′; ((y, x) represents the loss function for a single sample, which here refers to the cross-entropy loss function;
[0043] L cls (G1, S1′) represents the loss function of the overall sample, which here refers to the average of the losses of individual samples, and L2 regularization is adopted to prevent overfitting, where w is the network weight of each layer and λ is the regularization coefficient with a value of 0.01.
[0044] S2.6. Calculate the regression loss L r9g : Obtain G2(g cx , g cy , g w , g h ) from the annotation data of the input image, which respectively represent the actual position (g cx , g cy ) and size (g w , g h ) of the current image target, and define A cB , A cy , A w , A h as the predicted center coordinates and size. The normalized loss of a single sample is expressed as:
[0045]
[0046] The regression loss uses smooth L1 loss, as follows:
[0047]
[0048]
[0049] S2.7. Calculate the overall loss, and perform a weighted sum of the classification loss and the regression loss. The weight is denoted as λ, and the overall loss is:
[0050] L = L cls + λL r9g .
[0051] S2.8. Randomly initialize the network parameters to follow a normal distribution, set the batch size to 32, the learning rate to 0.01, and use the Stochastic Gradient Descent (SGD) algorithm to iteratively train 30 times to optimize the network parameters and save the results of each iteration.
[0052] S2.9. Test the results of iterations 10 - 30 on the OxUvA dataset and select the optimal model as the final training result.
[0053] S3. Input S1 and S2 into the TDS module to determine whether the target has disappeared; if the target exists, output the T signal, indicating that the tracker will be used in the next frame; if the target has disappeared, output the D signal, indicating that the detector will be used in the next frame.
[0054] The TDS module in step S3 is used to determine whether the target in the current frame has disappeared and use different tracking methods for different situations. The specific implementation steps are as follows:
[0055] Input S1 and S2 obtained in S2 into the TDS module. Extract the response map storing the target probability from the output results of each anchor box in S1 to obtain a score map of 17×17×5. Perform a Global MaxPooling (GMP) operation on the score map to find the part with the maximum response as the region of interest. If the target score in this region exceeds the threshold, it is considered that there is a target being tracked in the current frame. The TDS module will output the T signal and execute step S4; if the result is less than the threshold, it is determined that there is no target in the current frame. After outputting the D signal, directly enter the next frame and use the detector to find the target.
[0056] Among them, the tracker threshold in step S3 is used to determine whether the target in the current frame exists. It is set as a 5 - dimensional column vector, and the specific values are [0.648, 0.523, 0.5, 0.523, 0.648].
[0057] S4. If it is determined in step S3 that the target exists, add a cosine window and scale penalty to the response map S1 to limit large displacements, take the index of the location with the maximum response value, and find the corresponding data at the same index in S2, and convert it into the position and scale of the new prediction box, which is the tracking result of the current frame.
[0058] S5. If the tracking signal is D at the start of tracking for a certain frame, it indicates that the target disappeared in the previous frame. To determine whether the target reappears, the detector needs to be used in the current frame. The image is scaled and segmented into 480×640×3 and then input into the detector to obtain three features. The three features are input into the TDS module. If the target appears, a T signal is output and non-maximum suppression (NMS) operation is performed on the three features to eliminate redundant candidate boxes, and the detection result is output. By calculating the cosine similarity with the template, the one with the maximum similarity is used as the tracked target, a T signal is output and the tracker is used starting from the next frame. If the target does not appear, a D signal is output and the NMS operation is skipped, directly proceeding to the next frame and continuing to use the detector. Step S5 is specifically as follows:
[0059] S5.1. When extracting the response map of the detector, only the feature layers that record the responses and confidences of specific categories are extracted. For example, if the target of the current sequence is the "bird" category, only the feature layer that records the classification score of "bird" and the confidence layer that records the probability of the existence of the target in the current region are extracted. Since the target category does not change in the same sequence, the classification score matrix and the corresponding confidence matrix are multiplied element by element as the final score.
[0060] S5.2. Perform GMP operation on the extracted response map. If the result is less than the threshold, it is determined that there is no target, a D signal is output and the next frame is started. If it is greater than the threshold, the NMS operation is used to obtain the detection result, calculate the cosine similarity between all suspected targets and the template image, and select the detection with the maximum similarity as the tracking result, output a T signal, indicating that the tracker is used in the next frame.
[0061] The threshold requirement in step S5.2 is specifically set as a 3D column vector, with specific values being [0.64, 0.64, 0.64].
[0062] The cosine similarity calculation process in step S5.2 is as follows. Calculate the similarity S between the i-th detection and the template i , scale the i-th detection to the template size and convert it to a grayscale image, and flatten it into a one-dimensional column vector, denoted as D i , R is the one-dimensional column vector obtained by converting the template image to a grayscale image and flattening it, ||D i || and ||R|| are the two-norms of the two, then
[0063]
[0064] The above S1 - S4 are the target tracking processes using the tracker, and S5 is the tracking process using the detector. Through the judgment of the TDS module, the tracker or the detector is selected according to the situation, achieving the effect of avoiding tracking failure caused by the disappearance of the target, and constituting a complete target tracking process. In the actual target tracking process, the entire target tracking is completed by repeating steps S1 - S5. The bounding box information of the target tracking is obtained from step S4 and step S5 therein.
[0065] The following validates the effect of the present invention through simulation experiments. The simulation experiments use the OxUvA and UAV20L datasets and are compared with some open - source methods provided by the OxUvA official.
[0066] Among them, SiamTD is the method of the present invention. The official methods used in the simulation experiments of the present invention include the following 9 kinds:
[0067] 1. TLD (Online Learning for Tracking from Detection in Videos), see reference [1]. Kalal Z, Mikolajczyk K, Matas J. Tracking - learning - detection[J]. IEEE Transactions on Software Engineering, 2011, 34(7): 1409 - 1422.
[0068] 2. SiamFC (Fully - Convolutional Siamese Networks for Object Tracking), see reference [2]. Bertinetto L, Valmadre J, Henriques J F, et al. Fully - Convolutional Siamese Networks for Object Tracking[J]. 2016.
[0069] 3. LCT (Long - Term Correlation Tracking), see reference [3]. Chao M, Yang X, Zhang C, et al. Long - term correlation tracking[C], 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2015.
[0070] 4. MDNet (Visual Object Tracking Algorithm Based on Multi-Video Sequence Learning), see reference [4]. Nam H, Han B. Learning Multi-Domain Convolutional Neural Networks for Visual Tracking[J]. IEEE, 2016.
[0071] 5. SINT (Video Object Tracking Method Based on Siamese Network Template Matching), see reference [5].[1] Ran Tao, Efstratios Gavves, Arnold W.M. Smeulders. Siamese Instance Search for Tracking.[J]. CoRR, 2016, abs / 1605.05863.
[0072] 6. ECO-HC (Object Tracking Method Based on Efficient Convolutional Network), see reference [6]. Danelljan M, Bhat G, Khan F S, et al. ECO: Efficient Convolution Operators for Tracking: IEEE Computer Society, 10.1109 / CVPR.2017.733[P]. 2016.
[0073] 7. EBT (Video Object Tracking Method Based on Fast Global Detection), see reference [7]. Zhu G, Porikli F, Li H. Beyond Local Search: Tracking Objects Everywhere with Instance-Specific Proposals: IEEE, 10.1109 / CVPR.2016.108[P]. 2016.
[0074] 8. BACF (Video Object Tracking Method Based on Context Information), see reference [8]. Galoogahi H K, Fagg A, Lucey S. Learning Background-Aware Correlation Filters for Visual Tracking[J]. IEEE Computer Society, 2017.
[0075] 9. Staple (Object Tracking Method Based on Complementary Learners), see reference [9]. Bertinetto L, Valmadre J, Golodetz S, et al. Staple: Complementary Learners for Real-Time Tracking[C] / / Computer Vision&Pattern Recognition. IEEE, 2016.
[0076] 10. SiamRPN (Object Tracking Method Based on Siamese Network and Region Proposal Network), see reference
[10] . Bo L, Yan J, Wei W, et al. High Performance Visual Tracking with Siamese Region Proposal Network[C] / / 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2018.
[0077] 11. DaSiamRPN (Object Tracking Method Based on Siamese Network and Distractor Awareness), see reference
[11] . Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu. Distractor-aware siamese networks for visual object tracking. In European Conference on Computer Vision, pages 101–117, 2018.
[0078] 12. SiamBAN (Adaptive Bounding Box Object Tracking Method Based on Siamese Network), see reference
[12] . Chen Z, Zhong B, Li G, et al. Siamese Box Adaptive Network for Visual Tracking[J]. 2020.
[0079] The results of the simulation experiment are referred to in Appendix Figure 5 and Appendix Figure 6 , Figure 5 which is a comparison chart of prediction accuracy and success rate on the UAV20L dataset, Figure 5For the left-middle figure, the abscissa represents the distance threshold between the center point of the target position (bounding box) estimated by the algorithm and the center point of the target of the manual annotation (ground truth), and the ordinate represents the percentage of the number of frames less than this threshold in the total number of frames, that is, the prediction accuracy; Figure 5 For the right-middle figure, the abscissa represents the overlap rate threshold between the area of the target bounding box estimated by the algorithm and the bounding box of the target of the manual annotation (ground truth), and the ordinate represents the proportion of the number of frames greater than this threshold in the total number of frames, that is, the success rate. From Figure 5 It can be seen that on the UAV20L dataset, the prediction accuracy and success rate of the method described in the present invention are generally better than those of other algorithms.
[0080] Figure 6 The figure is a comparison chart of the evaluation results of SiamTD and other algorithms on the OxUvA dataset. The upper figure evaluates the judgment accuracy of the tracker for the target disappearance situation. It is set that when the target exists and the tracker judges correctly, it is True Positive (TP), and when the target disappears and the tracker judges correctly, it is True Negative (TN). The ordinate and abscissa are respectively the proportion of the number of frames marked as TP and TN in the total number of frames, denoted as TPR and TNR. The higher the TNR and TPR, the higher the tracking quality. MaxGM takes both of these two indicators into account at the same time; the values in the legend represent the comprehensive evaluation results, and the calculation formula is as follows:
[0081]
[0082] Figure 6 The other two figures (the middle figure and the lower figure) evaluate the long-term tracking performance of the tracker. The abscissa represents (0, x) minutes and (x, 10) minutes respectively, and the ordinate represents the tracking accuracy of the target during this period. According to the illustrated results, the long-term tracking performance of the tracker of the method of the present invention is significantly better than that of other algorithms participating in the comparison.
[0083] According to the results of the simulation experiment, it can be proved that on the UAV20L and OxUvA datasets, the prediction accuracy and success rate of the method (SiamTD) of the present invention are better than those of several other algorithms participating in the performance comparison, and this method also has certain advantages in long-term tracking tasks; in addition, the present invention does not use the tracker and the detector in the same frame at the same time, ensuring that the tracking speed is at least 74fps. In summary, the present invention solves the problem of tracking failure caused by target disappearance in long-term tracking on the premise of ensuring the tracking speed, and improves the success rate and accuracy of the tracker.
[0084] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention by using the above-disclosed technical content. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A long-term video object tracking method based on Siamese network for joint tracking and detection, characterized in that It includes the following steps: S1. Crop the template Z according to the first-frame input image I of the video sequence and the bounding box information B, and crop the search region X according to the second-frame input image I i i , i ∈ [2, n]; S2. Send Z and X i into the pre-trained offline Siamese network to extract features, obtaining feature φ(Z) and φ(X i ); send the features into the RPN network respectively, and output two response maps of 17×17×10 and 17×17×20 through the classification branch and the regression branch respectively, denoted as S1 and S2; S3. Input S1 and S2 into the TDS module to determine whether the target has disappeared. If the target exists, output the T signal, indicating that the tracker will be used in the next frame. If the target has disappeared, output the D signal, indicating that the detector will be used in the next frame. S4. If it is determined in step S3 that the target exists, add a cosine window and scale penalty to the response map S1 to limit large displacements, take the index at the position with the maximum response value, and find the data corresponding to the same index in S2, and convert it into the position and scale of the new prediction box, which is the tracking result of the current frame. S5. If the tracking signal is D at the start of tracking in a certain frame, it means that the target disappeared in the previous frame. To determine whether the target reappears, the detector needs to be used in the current frame. Scale the image and divide it into 480×640×3 and then input it into the detector to obtain three features. Input the three features into the TDS module. If the target appears, output the T signal and perform non-maximum suppression (NMS) operation on the three features to eliminate redundant candidate boxes, output the detection result, calculate the cosine similarity with the template, and take the one with the maximum similarity as the target to be tracked, output the T signal and start using the tracker from the next frame. If the target does not appear, output the D signal and skip the NMS operation, directly enter the next frame and continue to use the detector.
2. The long-term video object tracking method based on the combined tracking and detection of twin networks according to claim 1, wherein The Siamese network described in step S2 has two major branches: a template branch and a detection branch. The network structures of both major branches adopt the modified AlexNet, and the network parameters are shared.
3. The long-term video object tracking method based on twin network combined tracking and detection according to claim 1, characterized in that, The TDS module in step S3 is used to determine whether the target in the current frame has disappeared and use different tracking methods for different situations. Specifically, input S1 and S2 obtained in S2 into the TDS module, extract the response map storing the target probability from the output results of each anchor box in S1 to obtain a score map of 17×17×5, perform global max pooling (GMP) operation on the score map, find the part with the maximum response as the region of interest. If the target score in this region exceeds the threshold, it is considered that there is a target to be tracked in the current frame. The TDS module will output the T signal and execute step S4. If the result is less than the threshold, it is determined that there is no target in the current frame, output the D signal and directly enter the next frame, and use the detector to find the target.
4. The long-term video object tracking method based on twin network combined tracking and detection according to claim 3, characterized in that, The tracker threshold in step S3 is used to determine whether the target in the current frame exists. It is set as a 5-dimensional column vector, and the specific value is [0.648, 0.523, 0.5, 0.523, 0.648].
5. The long-term video object tracking method based on twin network combined tracking and detection according to claim 3, characterized in that, The steps of the TDS module in step S5 when using the detector are as follows: S5.
1. When extracting the response map of the detector, only extract the feature layer that records the response and confidence of a specific category. Since the target category does not change in the same sequence, multiply the classification score matrix and the corresponding confidence matrix element by element as the final score. S5.
2. Perform GMP operation on the extracted response map. If the obtained result is less than the threshold, it is determined that there is no target, output the D signal and start from the next frame. If it is greater than the threshold, use the NMS operation to obtain the detection result, calculate the cosine similarity between all suspected targets and the template image, select the detection with the maximum similarity as the tracking result, and output the T signal, indicating that the tracker will be used in the next frame.
6. The long-term video object tracking method based on the combined tracking and detection of twin networks according to claim 5, characterized in that, The cosine similarity calculation process in step S5 is as follows. Calculate the similarity S between the i-th detection and the template i , scale the i-th detection to the size of the template and convert it into a grayscale image, and flatten it into a one-dimensional column vector, denoted as D i , R is the one-dimensional column vector obtained by converting the template image into a grayscale image and flattening it, ||D i || and ||R|| are the two-norms of the two; then the cosine similarity is:
7. A long-term video object tracking method based on twin network joint tracking and detection according to claim 5, characterized in that The threshold requirement in step S5.2 is specifically set as a 3D column vector with specific values [0.64, 0.64, 0.64].
Citation Information
Patent Citations
Fast video target tracking method based on twin network
CN112164094A
Football player tracking method based on deep learning
CN112308013A