A video target tracking method and system

Through the historical template fusion and feature enhancement mechanism, combined with the Transformer encoder and the anchor-free box object detection algorithm, the problems of appearance changes and interference of similar targets in video target tracking are solved, achieving higher robustness and real-timeness.

CN119579655BActive Publication Date: 2025-06-06CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510135346.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-06
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

When facing changes in target appearance and interference from similar targets, existing video target tracking methods are prone to accumulate errors, resulting in tracking offsets or target loss, and consume large computing resources and poor real-time performance.

Method used

A video target tracking method is adopted to improve the adaptability to appearance changes through historical template fusion and feature enhancement and fusion mechanisms, and the timing relationship between appearance features and search domain features is optimized through the Transformer encoder and anchor-free box object detection algorithm to reduce interference from similar targets.

Benefits of technology

It improves the adaptability to target appearance changes, reduces the impact of similar target interference, enhances the robustness and real-timeness of the tracking method, and improves the success rate of the tracking task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579655B_ABST
    Figure CN119579655B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for tracking a video target, including video data acquisition, video data preprocessing, tracking model training, tracking model testing and verification, and obtaining a tracking result after the tracking model is verified by the input of a video to be processed. The tracking model uses a Transformer encoder to enhance the significance of appearance template features and search domain features, and calculates the similarity weight to obtain a memory feature; then the Transformer decoder fuses the enhanced search domain features and memory features to obtain the final feature, and the target bounding box is output by an anchor-free regression mechanism. In addition, the tracking model estimates the position of the target through a trajectory prediction model to avoid interference from similar objects nearby and correct the tracking result. The present invention also discloses a video target tracking system, which solves the problem of error accumulation and interference from similar targets caused by changes in target appearance, and has high tracking accuracy and success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recognition, and in particular to a method and system for tracking a video target. Background Art

[0002] In complex scenes, changes in the target's appearance will affect the detection performance of the tracking method, and errors are easily accumulated, resulting in tracking offsets, which is not suitable for long-term tracking. In addition, general tracking methods are interfered by similar objects, causing the target to be lost and unable to be retrieved. Traditional tracking methods based on twin networks only use the appearance features of the first frame as a matching template, which is not enough to discriminate targets and cannot handle similar targets, which easily leads to the failure of the overall tracking task.

[0003] In response to appearance changes, some tracking methods are equipped with complex online template update mechanisms, which are more robust. However, this type of online template update method consumes too many computing resources and reduces real-time performance. In addition, many devices do not support the gradient backpropagation algorithm and cannot modify the gradient value, which limits the implementation and promotion of such algorithms. Other methods require the selection of appropriate appearance features, but the accuracy is insufficient due to the training method, and reliable appearance features cannot be accurately obtained.

[0004] In response to the interference of similar objects, some tracking methods distinguish similar objects based on more significant appearance features. However, such methods are still easily disturbed by similar objects when faced with problems such as posture changes and angle changes of the original target. Summary of the invention

[0005] In view of the above defects of the prior art, the present invention provides a method for tracking a video target, aiming to solve the problem of error accumulation caused by target appearance changes and interference from similar targets. The present invention also provides a tracking system for implementing the method.

[0006] The technical solution of the present invention is as follows: a video target tracking method, comprising the following steps: video data acquisition, video data preprocessing, tracking model training, tracking model testing and verification, and obtaining tracking results after the tracking model is verified by the video input to be processed.

[0007] The video data preprocessing is to divide the video frame image of the video data into an appearance template and a search domain according to the bounding box of Groundtruth, wherein the appearance template includes target appearance and background features;

[0008] The process of processing the input video data during the tracking model training includes the following steps:

[0009] Feature extraction and feature enhancement: The appearance template and the label map corresponding to the appearance template are mapped, added and reduced in dimension by the feature extraction backbone network and the additional convolutional layer to obtain the appearance template features, and the search domain is mapped and reduced in dimension by the feature extraction backbone network to obtain the search domain features; the appearance template features and the search domain features are respectively sent to the Transformer encoder to obtain the enhanced appearance template features and the enhanced search domain features;

[0010] Historical template fusion: Calculate the similarity weights of the enhanced appearance template features and the enhanced search domain features and multiply and fuse them with the enhanced search domain features to obtain the memory features;

[0011] Feature enhancement and bounding box regression: The Transformer decoder fuses the enhanced search domain features and memory features to obtain the final features, and a one-stage anchor-free object detection algorithm obtains the object bounding box based on the final features;

[0012] Trajectory prediction: record the target bounding box generated by the tracking process, obtain the coordinates of the center point of the target bounding box, and put together a coordinate sequence of length L to form a historical trajectory. The coordinates of the historical trajectory are converted into world coordinates and input into the trajectory prediction model to obtain the predicted bounding box. The distance between the target bounding box and the predicted bounding box and the center point of the previous frame is calculated to determine whether it is interfered by similar objects. If there is no interference, the target bounding box is output as the result. If there is interference, the predicted bounding box is used as a new search domain for tracking and the target bounding box is output as the result.

[0013] Furthermore, the feature extraction backbone network for mapping all the appearance templates and the label images corresponding to the appearance templates and for mapping the search domain has the same structure but different weights, and the feature extraction backbone network is GoogleNet;

[0014] The dimensionality reduction for obtaining the appearance template features and the search domain features uses linear convolution with the same structure but different weights.

[0015] Furthermore, the Transformer encoder includes adding sinusoidal position coding to the input feature to form a first feature, flattening the input feature to obtain a series of feature vectors, inputting the input feature into the AttninAttn of the AiA module, and then summing and normalizing the feature vector to form a second feature, and inputting the second feature into the feedforward network of the AiA module and then summing and normalizing the second feature with the first feature to form an enhanced input feature.

[0016] Furthermore, the similarity weights of the enhanced appearance template features and the enhanced search domain features are calculated by calculating the similarity matrix w between the enhanced appearance template features and the enhanced search domain features, and then using SoftMax to normalize w to obtain the similarity weights. w is calculated by the following formula:

[0017] ,

[0018] is the enhanced appearance template feature, is the enhanced search domain feature, is the appearance template feature, is the search domain feature, i is the pixel subscript of the enhanced appearance template feature, j is the pixel subscript of the enhanced search domain feature, k is the subscript coefficient of the summation function, is the dot product, is the scale factor, and C is the dimension of the appearance template feature.

[0019] Furthermore, the distance between the target bounding box and the predicted bounding box and the center point of the previous frame is calculated to determine whether it is interfered by similar objects. The weight is calculated according to the following formula

[0020] ,

[0021] Among them, k1 and k2 are scale parameters, dist1 is the distance between the center point of the target bounding box or the predicted bounding box and the center point of the previous frame, dist_max is the maximum distance of the search domain, and dist_near is the set safety distance. When the calculated target bounding box weight is greater than the predicted bounding box weight, there is no interference, otherwise there is interference.

[0022] Another technical solution of the present invention is a video target tracking system, comprising the following modules:

[0023] Acquisition module, used for video data acquisition;

[0024] Preprocessing module, used for video data preprocessing;

[0025] Training module, used to train the tracking model;

[0026] A test and verification module is used to test and verify the tracking model;

[0027] An output module is used to obtain the tracking result after the tracking model is verified by the video input to be processed;

[0028] The video data preprocessing is to divide the video frame image of the video data into a plurality of appearance templates and search domains according to the bounding box of the Groundtruth, wherein the appearance template includes the target appearance and background features;

[0029] The tracking model includes modules:

[0030] A feature extraction and feature enhancement module, which is used to map, add and reduce the dimension of all the appearance templates and the label images corresponding to the appearance templates by a feature extraction backbone network and an additional convolutional layer to obtain appearance template features, and map and reduce the dimension of the search domain by a feature extraction backbone network to obtain search domain features; and send the appearance template features and the search domain features to a Transformer encoder to obtain enhanced appearance template features and enhanced search domain features respectively;

[0031] A history template fusion module, used to calculate the similarity weight of the enhanced appearance template feature and the enhanced search domain feature and multiply and fuse them with the enhanced search domain feature to obtain a memory feature;

[0032] A feature enhancement and bounding box regression module is used to obtain a final feature by fusing the enhanced search domain features and memory features by a Transformer decoder, and output a target bounding box according to the final feature by a one-stage anchor-free target detection algorithm;

[0033] And, a trajectory prediction module is used to record the target bounding box generated by the tracking process, obtain the coordinates of the center point of the target bounding box, put together a coordinate sequence of length L to form a historical trajectory, convert the coordinates of the historical trajectory into world coordinates and input the trajectory prediction model to obtain a predicted bounding box, calculate the distance between the target bounding box and the predicted bounding box and the center point of the previous frame to determine whether it is interfered by similar objects, output the target bounding box as the result when there is no interference, and use the predicted bounding box as a new search domain for tracking and output the target bounding box as the result when there is interference.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] (1) To address the problem of error accumulation caused by changes in target appearance, the present invention introduces a historical template fusion mechanism and a feature enhancement and fusion mechanism, uses multiple appearance features as matching basis and increases feature saliency. Compared with the prior art, the present invention improves the adaptability to appearance changes.

[0036] (2) Secondly, the present invention aims to solve the problem that the tracking method lacks the ability to deal with interference from similar targets. The tracking method proposes a Transformer-based framework to optimize appearance features and establish a temporal relationship between template features and search domain features to reduce the interference from similar targets. A framework based on spatiotemporal graph convolutional neural network is used to optimize the prediction of future trajectories. Selecting targets that conform to historical motion laws can further reduce such interference and improve tracking progress. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1This is a network structure diagram of the present invention.

[0038] Figure 2 This is a comparison chart of the feature enhancement effects used in the present invention. DETAILED DESCRIPTION

[0039] The present invention will be further described below in conjunction with the embodiments, but are not intended to limit the present invention.

[0040] The video target tracking system of the embodiment of the present invention includes a collection module for collecting video data;

[0041] Preprocessing module, used for video data preprocessing;

[0042] Training module, used to train the tracking model;

[0043] A test and verification module is used to test and verify the tracking model;

[0044] An output module is used to obtain the tracking result after the tracking model is verified by the video input to be processed;

[0045] The video data preprocessing is to divide the video frame image of the video data into several appearance templates and search domains according to the bounding box of the Groundtruth;

[0046] The tracking model includes modules:

[0047] The feature extraction and feature enhancement module is used to map, add and reduce the dimension of all appearance templates and label images corresponding to the appearance templates by the feature extraction backbone network and the additional convolution layer to obtain the appearance template features, and map and reduce the dimension of the search domain by the feature extraction backbone network to obtain the search domain features; the appearance template features and the search domain features are respectively sent to the Transformer encoder to obtain the enhanced appearance template features and the enhanced search domain features;

[0048] A history template fusion module is used to calculate the similarity weights of the enhanced appearance template features and the enhanced search domain features and multiply and fuse them with the enhanced search domain features to obtain the memory features;

[0049] The feature enhancement and bounding box regression module is used to obtain the final features by fusing the enhanced search domain features and memory features by the Transformer decoder, and the one-stage anchor-free object detection algorithm outputs the target bounding box according to the final features;

[0050] And, a trajectory prediction module is used to record the target bounding box generated by the tracking process, obtain the coordinates of the center point of the target bounding box, put together a coordinate sequence of length L to form a historical trajectory, convert the coordinates of the historical trajectory into world coordinates and input the trajectory prediction model to obtain a predicted bounding box, calculate the distance between the target bounding box and the predicted bounding box and the center point of the previous frame to determine whether it is interfered by similar objects, output the target bounding box as the result when there is no interference, and use the predicted bounding box as a new search domain for tracking and output the target bounding box as the result when there is interference.

[0051] The specific tracking method implemented by the video target tracking system includes the following steps:

[0052] Step 1: Data Collection

[0053] Download the public data set from the official website to obtain the training set, test set, and validation set as samples for tracking model training and testing.

[0054] Step 2: Data preprocessing

[0055] The input of the tracking model is mainly the video frames collected by the camera. The step of video data preprocessing is to divide the video image of the video data into several appearance templates and search domains based on the bounding box of Groundtruth, where the appearance template contains the target appearance and background features. The foreground and background information is obtained from the appearance template, and a label map is generated, which consists of 0 and 1 values, with the background value being 0 and the target value being 1. Also based on Groundtruth, the center point (x, y) of each target is obtained, and the center point sequence is collected as the training data of the trajectory prediction model.

[0056] Step 3: Track model training. Figure 1 and Figure 2 As shown, the processing process of the tracking template on the input video frame is:

[0057] Step 3.1: Feature extraction

[0058] The feature extraction network of the tracking model includes a memory branch and a search domain branch. Different from the traditional twin network-based tracking method, the memory branch has T appearance templates. m , and a corresponding label map is required c , used to distinguish the foreground from the background. In addition to the same size as the memory branch, the tracking method has only one search domain branch and does not require a label map. The specific steps are as follows:

[0059] 1) Extraction of appearance template features: As shown in formula (1), the appearance template m The format is not a traditional tracking template, and contains both target appearance and background features, so a label map is required. c Distinguish foreground from background. Label image cEnsure memory backbone network Learn the real appearance features and avoid the interference of background appearance. In the specific method, the appearance of the target is set as the label image c The value of 1 is the background value, and the value of 0 is the backbone network. And additional convolutional layers g Will m and c Map to the same spatial scale and add, then perform linear convolution Reduce the feature dimension to the specified dimension to obtain the feature set , where the features in the set are .

[0060] (2) Extraction of search domain features: As shown in formula (2), the backbone network structure corresponding to the search domain branch and the memory branch is the same, and the search domain feature input q Search Domain Backbone Network and convolutional networks Degrade to the specified dimension to obtain the search domain features . But the search domain backbone network and memory backbone network The weights are different, so it is not the same network.

[0061] (1)

[0062] (2)

[0063] Step 3.2: Feature Enhancement

[0064] (1) The Transformer encoder is used to enhance the appearance features extracted by the backbone network. The Transformer encoder consists of a multi-layer stack, each of which includes a multi-head self-attention module and a feedforward network. The self-attention module is used to capture the correlation between all feature vectors to enhance the original features, and is equipped with an AiA module () to exploit the information clues between the correlations in the features. As shown in formula (3), the AiA module seeks the consistency of the correlation around each keyword to enhance the correct correlation of relevant query keywords and suppress the incorrect correlation of irrelevant query keywords. If the multi-head attention mechanism embedded in the traditional Transformer is used, the AiA module can be obtained, as shown in formula (4).

[0065] (3)

[0066] (4)

[0067] (2) Transformer encoder. The encoder aims to focus on more useful features and remove useless interfering features. In order to better explore the extracted features f t ( t for m or q ), the encoder uses two multi-head attention layers including a feed-forward layer. The encoder extracts f t The effective information in the .

[0068] Specifically, the tracking method combines the features f t Flatten to get a sequence of eigenvectors and add sinusoidal position encoding Then, the encoder converts the new feature vector sequence Input AttninAttn of AiA module and get Finally, the encoder converts the new feature vector sequence Input the feedforward network FFN of the AiA module and get .

[0069] In summary, the feature enhancement network transforms the template features f m and search domain characteristics f q Input into formula (5), formula (6) and formula (7), we get and , the output features highlight the target and weaken the background, that is, the feature saliency is enhanced.

[0070] (5)

[0071] (6)

[0072] (7)

[0073] Step 3.3: Fusion History Template

[0074] The history template fusion mechanism is used to find and The relationship between To search The target position information in . First calculate and The similarity matrix between Then use SoftMax to standardize w . Refer to formula (8) for details. For example, i yes The pixel index of j yes The pixel index of is the dot product, is the proportional coefficient to prevent the function from overflowing, where C is the characteristic The dimension of is, k is the subscript coefficient of the sum function. Referring to formula (9), the algorithm finally converts As a weight map, use take Get the fused memory features .because Stores historical appearance information related to the target. According to the needs of the search domain, the algorithm can adaptively retrieve the information stored in Target information in.

[0075] (8)

[0076] (9)

[0077] Step 3.4: Feature Fusion

[0078] Fusion using Transformer decoder and , and introduces a cross attention mechanism to find the temporal relationship between the two and promote information propagation. The cross attention mechanism has the same structure as the AiA attention mechanism of the encoder, only the input features are different. In order to better integrate the current memory features Search domain features ,The feature fusion network uses a multi-head attention layer with feedforward to generate an attention map and extract and integrate the effective information in , and obtain the final feature .

[0079] Specifically referring to formulas (10) and (11), the decoder converts the feature vector sequence and Input the AiA module and get Finally, the decoder converts the new feature vector sequence Input the feedforward network of the AiA module and get .

[0080] (10)

[0081] (11)

[0082] In summary, the encoder used by the template feature set and the search domain feature in the Transformer encoder and decoder used in the tracking method is the same network with shared weights. After obtaining the enhanced features, the algorithm inputs the two into the historical template fusion mechanism and calculates the weights to obtain the memory features and search domain features. Finally, it enters the decoder to obtain the fused features. Unlike traditional target tracking methods based on twin networks, the tracking method does not use cross-correlation operations to calculate the similarity between template features and search domain features, but uses the historical template fusion mechanism to enhance the robustness of the tracking method.

[0083] Step 3.5: Classification and regression of bounding boxes

[0084] The one-stage anchor-free object detection algorithm (FCOS) has good performance and fewer network parameters, so the tracking method introduces an anchor-free head task network to generate bounding boxes. The network contains: a classification branch that distinguishes background and target, a branch that suppresses low-quality bounding boxes, and a regression branch for anchor-free boxes that directly estimates the target bounding box.

[0085] First, we enter the classification branch. The head task network uses a lightweight classification convolutional network. right Then, the head task network uses a linear convolution layer with a 1×1 kernel to encode The output dimension is reduced to 1, generating the final classification response map Then we enter the branch that suppresses low-quality bounding boxes (center-ness), which can select positive samples near the target boundary and suppress low-quality target bounding boxes. Then, the sub-branches are separated to generate the centrality response graph. In the forward process, Multiply , to suppress the classification confidence of pixels far from the target center. In the regression branch (Regression), the head task network will Passed to another lightweight regression convolutional network , then reduce the dimension of the output feature to 4 and generate the regression response graph Estimate object bounding box.

[0086] Step 4: Trajectory prediction and result correction, including:

[0087] The tracking method predicts future trajectories based on the Social-STGCNN model. The Social-STGCNN model includes a spatiotemporal graph convolutional neural network (STG-CNN) and a temporal extrapolation convolutional neural network (TXP-CNN). STG-CNN performs spatiotemporal convolution operations on the graphical representation of the target trajectory to extract features. These features are a concise representation of the observed historical trajectories of the target. TXP-CNN takes these features as input and predicts the future trajectories of all targets as a whole. Social-STGCNN can predict the motion trajectories of two objects at the same time, so in addition to the trajectory of the center point of the object, another trajectory can be predicted, that is, the trajectory is defined as the temporal change of the length and width (w, h) of the bounding box. In this way, Social-STGCNN can be used to predict the future bounding box of the target. For the coordinates (x, y) and coordinates (w, h) of the trajectory, they are defined as graphs. .

[0088] Spatiotemporal Graph Convolutional Neural Network (STG-CNN): STG-CNN defines a new graph The spatial graph convolution is extended to the spatiotemporal graph convolution. The properties are The attribute set of . Combined with the spatiotemporal information of the target trajectory. It is worth noting that arrive The topological structure is the same, but when t changes, it becomes Assigned different attributes. Therefore, the tracking method will Defined as ,in , . Midpoint The properties are A collection of In addition, The corresponding weighted adjacency matrix yes The embedding generated by STG-CNN is represented as .

[0089] Temporal Extrapolation Convolutional Neural Network (TXP-CNN): The function of STG-CNN is to extract spatiotemporal node embeddings from the input graph. However, the purpose of tracking methods is to predict further steps in the future, hoping to be a stateless system, and here TXP-CNN comes into play. TXP-CNN directly operates on the temporal dimension of the graph embedding V and expands it as a necessary condition for prediction. Since TXP-CNN relies on convolution operations on the feature space, it has a smaller parameter size compared to recurrent units. It is important to note that TXP-CNN layers are not permutation invariant, as changes in the graph embedding before TXP-CNN will lead to different results.

[0090] In summary, given a set of N targets in a scene, in the time period The corresponding historical location , , tracking methods require predictions for future time horizons The upcoming track For a target n, the tracking method writes the corresponding trajectory to be predicted as ,in is a random variable describing the probability distribution of the position of target n in 2D space at time t. Specifically, the tracking method expresses the trajectory as , which follows a bivariate distribution .

[0091] The trajectory prediction model is trained to minimize the negative log-likelihood, which is defined as:

[0092] (11)

[0093] in This includes all trainable parameters of the model, is the mean of the distribution, is the variance, It's correlation.

[0094] Based on the bounding box predicted by the tracking algorithm and the bounding box predicted by the trajectory, the distance weight is introduced to select a more appropriate bounding box. When the center point of the new bounding box is closer to the center point of the previous frame, it means that the position of the target is more in line with the law of motion. Otherwise, it is likely to be interfered by similar objects. Therefore, the distance weight is defined as a value between 0 and 1. When the distance is smaller, the weight value is closer to 1. In summary, the formula for the distance weight is as follows:

[0095] ,

[0096] Among them, k1 and k2 are scale parameters, dist1 is the distance between the center point of the bounding box and the center point of the previous frame, dist_max is the maximum distance of the search domain, and dist_near is the set safety distance, which is usually half of dist_max. The weight score and displacement offset are used to determine whether the tracking method is interfered by similar objects. When the calculated target bounding box weight is greater than the predicted bounding box weight, there is no interference, otherwise there is interference. If there is interference, the tracking method obtains the possible position of the current target based on the target historical trajectory through the trajectory prediction model, turns it into a new search domain and tracks it again.

[0097] A qualitative and quantitative analysis was conducted on the effect of the video target tracking method of the present invention.

[0098] The tracking model training set of the present invention includes TrackingNet, LaSOT, GOT-10k, ILSVRCVID, ILSVRCDET and COCO. The experiment first synthesizes training samples based on ILSVRCDET and COCO: within an interval of 100 frames, T (T = 3 in this paper) frames are sampled, and random affine transformation is used to increase the number of samples. Specifically, the translation parameters of the experiment are randomly executed from −0.2S to 0.2S, and the ratio of the resize is adjusted in 1 1+r and 1 1-r The algorithm starts from the center of the target, cuts a square slice with a side length of S, and resizes the slice to 289×289 as the search domain or template.

[0099] Regarding the adjustment of the training set, taking GOT-10k as an example, the number of training samples per epoch is set to 150,000 and the batch size is set to 32. The momentum and weight decay rates are set to 0.9 and 1×10 respectively. −1 In the first 10 epochs, the experiment freezes the tracking network m and q All weight parameters of the layer do not participate in training. At this time, the experiment uses the warm-up technique and the learning rate is increased from 1×10 −2 Increased to 8×10 −1 After 10 epochs, the experiment was unfrozen m and q The parameters of the third and fourth stages of the layer. At this time, the experiment uses a cosine annealing learning rate schedule, and the learning rate is increased from 8×10 –2 Reduced to 1×10 –6 .

[0100] In terms of training, the tracking model was trained for 20 epochs using the SGD optimizer, and the loss value required about 27 hours to converge. Each epoch had 150,000 samples, and the batch size was set to 64. The tracking method of the present invention was implemented on the pytorch 1.1 framework of python3.7. The tracking method was run on an Intel(R) Xeon(R) Gold 6132 CPU @ 2.60GHz CPU and an NVIDIA Corporation GP100GL [Tesla P100 PCIe 16GB] graphics card, and the tracking method had an average speed of 21 frames per second (FPS).

[0101] The trajectory prediction model training set of the present invention is the trajectory sequence of LaSOT. Because LaSOT is a target tracking dataset, it includes video images and manually annotated bounding box labels. Because the experiment requires modifying the manually annotated bounding box labels into trajectory sequences. For each bounding box label , the experiment changes it to the center point coordinates , transforming it from pixel coordinates to world coordinates through the matrix H According to the dataset format of trajectory prediction, the center point coordinates plus the frame number and target ID, the training label of the final trajectory model is In terms of actual training, the trajectory prediction model was trained for 150 epochs using the Adam optimizer, and the loss value took about 40 hours to converge.

[0102] From the perspective of qualitative analysis, Figure 2 The comparison of heat maps before and after the feature is shown. Faced with the interference of similar objects such as pedestrians, the tracking method can better improve the significance of the appearance features and improve the detection ability of the tracking method. In the video example of the UAV123 data set, the target appearance changes significantly, there are similar objects interfering, and the scale changes greatly. However, the tracking method of the present invention is robust and has strong robustness under challenging conditions. From the qualitative results, the present invention successfully combines the advantages of feature enhancement network and feature fusion network.

[0103] From the perspective of quantitative analysis, the present invention demonstrates the superiority of the tracking method through experiments. The experiments are based on the following benchmark datasets: OTB-2015, LaSOT, TrackingNet, UAV123 and GOT-10k. The proposed tracking method is verified and evaluated from the real world, standard datasets and drone environments. OTB-2015: OTB-2015 is a common visual tracking benchmark dataset, including 100 video sequences. Each video is marked with certain attributes, such as occlusion, out of view and crowd challenges, which can comprehensively evaluate the performance of the tracking method.

[0104] Table 1 OTB-2015 test table

[0105]

[0106] Table 1 lists the comparison of the tracking method with other state-of-the-art tracking methods in terms of success rate. Among them, Globaltrack can accurately regress the scale of the target and is good at dealing with target loss. Ocean and DiMP series can automatically update weights, modify filters, and improve adaptability to appearance changes. Stark uses a self-attention mechanism to fuse the temporal information between the search domain and the template, effectively dealing with changes in target appearance. The present invention surpasses them in success rate, proving the superior performance of the tracking method, strong robustness and applicability to various complex tracking scenarios.

[0107] TrackingNet: TrackingNet is a large-scale long-term tracking dataset that provides a large number of videos of real environments for training and testing.

[0108] Table 2 TrackingNet test table

[0109]

[0110] As shown in Table 2, the experiment selected Success rate, Precision rate and standard Precision rate as evaluation indicators, and made a comprehensive comparison with other mainstream algorithms. The data shows that in terms of Success rate and Precision rate, the present invention is 1.4% and 2.3% higher than the second-ranked PrDiMP50 tracking method, respectively, with strong performance, and is superior to mainstream advanced real-time tracking methods. It is worth noting that TrackingNet has a wide diversity in the number of categories and scenarios, and is more prone to special situations such as target appearance changes and interference from similar objects. Therefore, the remarkable performance of the present invention demonstrates its strong generalization ability in the real world.

[0111] UAV123: UAV123 is designed to evaluate tracking methods in UAV applications and includes 123 low-altitude aerial videos with an average of 915 frames per video. In addition to short-term tracking, the dataset also contains 20 long videos, namely UAV20L, which are used to evaluate the performance of tracking methods in long-term tasks. Due to the characteristics of UAVs, this dataset has many challenging elements, including partial and complete occlusion, targets out of view, and small target tracking. Therefore, many tracked targets in UAV123 have low resolution.

[0112] Table 3 UAV123 test table

[0113]

[0114] However, as shown in Table 3, our method achieves an Area Under the Circumference (AUC) score of 0.594, which significantly outperforms recent competitive twin network tracking methods such as ECO, CCOT, and SiamRPN++ while running at real-time speeds.

[0115] GOT-10k: GOT-10k is a recently released large-scale general target tracking dataset with a total of 10,000 videos, of which the test set contains 180 videos and has multiple challenges. The dataset promotes the development of general target tracking methods by following the one-shot rule of zero overlap between the training set and the test set. Because the test set is not public, all algorithms evaluate their performance on the official server.

[0116] Table 4 GOT-10k test table

[0117]

[0118] Table 4 lists the comparison of the average overlap (AO) and success rate (SR) of the present invention and other state-of-the-art tracking methods at thresholds of 0.5 and 0.75. The present invention is 0.41% higher than the second-ranked PrDiMP50 tracking method in terms of SR0.75. The performance of SR0.75 proves that the tracking method can estimate the scale of the target more accurately than the ordinary algorithm, that is, it is easier to obtain an excellent appearance template, which further improves the accuracy of the tracking method. The two complement each other and jointly improve the performance of the overall tracking method.

Claims

1. A method for tracking a video target, comprising the following steps: After video data acquisition, video data preprocessing, tracking model training, tracking model testing and verification, and after the video input to be processed is verified, the tracking model obtains the tracking result, which is characterized in that: The video data preprocessing is to divide the video frame image of the video data into an appearance template and a search domain according to the bounding box of Groundtruth, wherein the appearance template includes target appearance and background features; The process of processing the input video data during the tracking model training includes the following steps: Feature extraction and feature enhancement: The appearance template and the label map corresponding to the appearance template are mapped, added and reduced in dimension by the feature extraction backbone network and the additional convolutional layer to obtain the appearance template features, and the search domain is mapped and reduced in dimension by the feature extraction backbone network to obtain the search domain features; the appearance template features and the search domain features are respectively sent to the Transformer encoder to obtain the enhanced appearance template features and the enhanced search domain features; Historical template fusion: Calculate the similarity weights of the enhanced appearance template features and the enhanced search domain features and multiply and fuse them with the enhanced search domain features to obtain the memory features; Feature enhancement and bounding box regression: The Transformer decoder fuses the enhanced search domain features and memory features to obtain the final features, and a one-stage anchor-free object detection algorithm obtains the object bounding box based on the final features; Trajectory prediction: record the target bounding box generated by the tracking process, obtain the center point coordinates of the target bounding box and the length and width of the target bounding box, and put together a coordinate sequence of length L to form a historical trajectory. The coordinate sequence includes the center point coordinates and the length and width. The coordinates of the historical trajectory are converted into world coordinates and input into the trajectory prediction model to obtain a predicted bounding box. The trajectory prediction model includes a spatiotemporal graph convolutional neural network and a time extrapolation convolutional neural network. The distance between the target bounding box and the predicted bounding box and the center point of the previous frame is calculated to determine whether it is interfered by similar objects. If there is no interference, the target bounding box is output as the result. If there is interference, the predicted bounding box is used as a new search domain for tracking and the target bounding box is output as the result. The weight is calculated according to the following formula when calculating the distance between the target bounding box and the predicted bounding box and the center point of the previous frame to determine whether it is interfered by similar objects. , Among them, k1 and k2 are scale parameters, dist1 is the distance between the center point of the target bounding box or the predicted bounding box and the center point of the previous frame, dist_max is the maximum distance of the search domain, and dist_near is the set safety distance. When the calculated target bounding box weight is greater than the predicted bounding box weight, there is no interference, otherwise there is interference.

2. The video target tracking method according to claim 1, characterized in that: The feature extraction backbone network for mapping all the appearance templates and the label images corresponding to the appearance templates and mapping the search domain has the same structure but different weights, and the feature extraction backbone network is GoogleNet; The dimensionality reduction for obtaining the appearance template features and the search domain features uses linear convolution with the same structure but different weights.

3. The video target tracking method according to claim 1, characterized in that: The Transformer encoder includes adding sinusoidal position coding to the input feature to form a first feature, flattening the input feature to obtain a series of feature vectors and then inputting them into the AttninAttn of the AiA module, which are then summed with the feature vectors and normalized to form a second feature, and inputting the second feature into the feedforward network of the AiA module and then summing and normalizing them with the first feature to form an enhanced input feature.

4. The video target tracking method according to claim 1, characterized in that: The similarity weights of the enhanced appearance template features and the enhanced search domain features are calculated by calculating the similarity matrix w between the enhanced appearance template features and the enhanced search domain features, and then using SoftMax to normalize w to obtain the similarity weights. w is calculated by the following formula , is the enhanced appearance template feature, is the enhanced search domain feature, is the appearance template feature, is the search domain feature, i is the pixel subscript of the enhanced appearance template feature, j is the pixel subscript of the enhanced search domain feature, k is the subscript coefficient of the summation function, is the dot product, is the scale factor, and C is the dimension of the appearance template feature.

5. A video target tracking system, characterized in that: Includes the following modules: Acquisition module, used for video data acquisition; Preprocessing module, used for video data preprocessing; Training module, used to train the tracking model; A test and verification module is used to test and verify the tracking model; An output module is used to obtain the tracking result after the tracking model is verified by the video input to be processed; The video data preprocessing is to divide the video frame image of the video data into a plurality of appearance templates and search domains according to the bounding box of the Groundtruth, wherein the appearance template includes the target appearance and background features; The tracking model includes modules: A feature extraction and feature enhancement module, configured to map, add, and reduce the dimensions of all the appearance templates and the label images corresponding to the appearance templates by a feature extraction backbone network and an additional convolutional layer to obtain appearance template features, and map and reduce the dimensions of the search domain by a feature extraction backbone network to obtain search domain features; The appearance template feature and the search domain feature are respectively sent to a Transformer encoder to obtain an enhanced appearance template feature and an enhanced search domain feature; A history template fusion module, used to calculate the similarity weight of the enhanced appearance template feature and the enhanced search domain feature and multiply and fuse them with the enhanced search domain feature to obtain a memory feature; A feature enhancement and bounding box regression module is used to obtain a final feature by fusing the enhanced search domain features and memory features by a Transformer decoder, and output a target bounding box according to the final feature by a one-stage anchor-free target detection algorithm; And, a trajectory prediction module, which is used to record the target bounding box generated during the tracking process, obtain the center point coordinates of the target bounding box and the length and width of the target bounding box, and put together a coordinate sequence of length L to form a historical trajectory, wherein the coordinate sequence includes the center point coordinates and the length and width, and the coordinates of the historical trajectory are converted into world coordinates and input into a trajectory prediction model to obtain a predicted bounding box. The trajectory prediction model includes a spatiotemporal graph convolutional neural network and a time extrapolation convolutional neural network, and calculates the distance between the target bounding box and the predicted bounding box and the center point of the previous frame to determine whether they are interfered by similar objects. When there is no interference, the target bounding box is output as a result. When there is interference, the predicted bounding box is used as a new search domain for tracking and the target bounding box is output as a result. The weight is calculated according to the following formula when calculating the distance between the target bounding box and the predicted bounding box and the center point of the previous frame to determine whether it is interfered by similar objects. , Among them, k1 and k2 are scale parameters, dist1 is the distance between the center point of the target bounding box or the predicted bounding box and the center point of the previous frame, dist_max is the maximum distance of the search domain, and dist_near is the set safety distance. When the calculated target bounding box weight is greater than the predicted bounding box weight, there is no interference, otherwise there is interference.

6. The video target tracking system according to claim 5, characterized in that: The feature extraction backbone network for mapping all the appearance templates and the label images corresponding to the appearance templates and mapping the search domain has the same structure but different weights, and the feature extraction backbone network is GoogleNet; The dimensionality reduction for obtaining the appearance template features and the search domain features uses linear convolution with the same structure but different weights.

7. The video target tracking system according to claim 5, characterized in that: The Transformer encoder includes adding sinusoidal position coding to the input feature to form a first feature, flattening the input feature to obtain a series of feature vectors and then inputting them into the AttninAttn of the AiA module, which are then summed with the feature vectors and normalized to form a second feature, and inputting the second feature into the feedforward network of the AiA module and then summing and normalizing them with the first feature to form an enhanced input feature.

8. The video target tracking system according to claim 5, characterized in that: The similarity weights of the enhanced appearance template features and the enhanced search domain features are calculated by calculating the similarity matrix w between the enhanced appearance template features and the enhanced search domain features, and then using SoftMax to normalize w to obtain the similarity weights. w is calculated by the following formula , is the enhanced appearance template feature, is the enhanced search domain feature, is the appearance template feature, is the search domain feature, i is the pixel subscript of the enhanced appearance template feature, j is the pixel subscript of the enhanced search domain feature, k is the subscript coefficient of the summation function, is the dot product, is the scale factor, and C is the dimension of the appearance template feature.