A method for tracking unmanned aerial vehicle targets
By adopting a feature-related network based on a self-attention mechanism in drone target tracking, using global information and time update strategies, the problem of insufficient accuracy and robustness of drone target tracking in complex environments is solved, and higher tracking performance and more direct structure are achieved.
Patent Information
- Application Number
- CN202210695272.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-17
AI Technical Summary
UAV target tracking has problems with insufficient accuracy and robustness in complex and uncertain environments, especially when target shapes are small, blurred boundaries and low resolution.
A feature-related network based on a self-attention mechanism is adopted, and global information is fully utilized, combining the characteristics between the search area and template, reducing external interference, and obtaining global spatio-temporal features through learning query embedding and time update strategies to improve the accuracy and robustness of tracking.
It improves the accuracy and robustness of drone target tracking, enhances adaptability to the rapid changes in the appearance of target objects, and simplifies the tracking pipeline, avoids post-processing steps, and meets the requirements of onboard running speed.
Smart Images

Figure CN114972439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) intelligent design, and in particular to a UAV target tracking method. Background Art
[0002] Unmanned aerial vehicles (UAVs) are one of the main types of unmanned aerial vehicle systems, and their mission range is diverse and expanding. In recent years, UAVs have been widely used in military and civilian fields in areas such as environmental monitoring, search and rescue, and autonomous positioning. Target tracking is one of the important topics in UAV research. There are currently two main methods, correlation filter (CF)-based methods and deep learning-based methods. However, target tracking from the perspective of UAVs suffers from small target shape, blurred boundaries, and low resolution due to its inherent altitude and flight jitter, which seriously affects the integration and matching degree between the template and the search area. Most existing trackers use correlation to integrate the information in the template and search area into the region of interest, which is a local linear matching process that mainly processes local areas. Due to the lack of global information in most cases, the feature fusion between the template and the search area is poor, affecting the accuracy and robustness of target tracking. Therefore, stable and effective tracking of agile targets in complex and uncertain environments has always been a challenge for UAV target tracking.
[0003] CF-based trackers treat tracking as a classification problem. They convert the computation of cyclic correlation and convolution in the spatial domain into element-by-element multiplication of elements in the frequency domain through discrete Fourier transform. This strategy significantly improves the running speed of CF-based trackers on a single CPU to around hundreds of fps, thus meeting the real-time requirements of drones. Therefore, CF-based trackers on drones have been rapidly expanded and widely used in the past few years. Huang Z proposed using the response map generated in the detection stage to form the learning limit. This strategy enhances the robustness and accuracy of the tracked object. Fu C designed a new multi-core correlation tracking framework that uses image quality metrics and contextual information to construct adaptive interference sources to improve robustness. Although many scholars have made outstanding contributions to improving the robustness and accuracy of CF trackers, most tracking methods only use information from the latest frame to update the model, resulting in limited understanding of historical information. In addition, the correlation-based operation only evaluates the similarity with local features, and the appearance of the target object changes over time, resulting in tracking drift when fast motion or occlusion occurs.
[0004] With the development of convolutional neural networks (CNNs) and the improvement of computer computing power, learning-based methods have become more and more common in the field of tracking. Siamese networks pioneered end-to-end learning and paved the way for deep learning methods to gradually surpass CF. SiamFC is a pioneering work that combines the original feature association with the Siamese framework and uses an offline end-to-end training strategy to avoid online parameter updates. Q. Guo et al. proposed a dynamic Siamese network (DSiam) with an anchor-free structure based on SiamFC, in which the fast transformation learning model can effectively handle changes in object appearance. B. Li et al. proposed a Siamese region proposal network (SiamRPN), which combines the Siamese network with the region proposal network (RPN) and uses deep correlation for feature fusion to obtain more accurate tracking results. In addition, many researchers have made further improvements, such as adding additional branches, building deeper architectures, and developing anchor-free architectures. Although Siamese-based trackers have achieved amazing results, UAVs do not support the high complexity of convolution operations in the network and various post-processing to select the best bounding box as the tracking result. Commonly used post-processing includes cosine window, scale or aspect ratio penalty, bounding box smoothing, etc. While post-processing provides better results, it leads to performance sensitivity to hyper-parameters. Some trackers try to simplify the tracking pipeline, but their performance degrades significantly. Summary of the invention
[0005] The purpose of the present invention is to provide a method for UAV target tracking, which solves this problem by fully utilizing global information through a feature correlation network based on a self-attention mechanism. The method effectively combines the features between the search area and the template, reduces the impact of external interference, and thus improves the accuracy and robustness of the tracking method. Global spatiotemporal features are obtained by learning query embedding and time update strategies for prediction, enhancing adaptability to rapid changes in the appearance of the target object. In order to meet the requirements of airborne operation speed, the proposed method has no suggested or predefined anchors, does not require post-processing steps, and the entire method is end-to-end to overcome the shortcomings of the prior art.
[0006] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a UAV target tracking method, which utilizes the principle of self-attention mechanism; the self-attention mechanism is used to scan each element in the sequence and update it by aggregating the information in the entire sequence, so as to pay attention to the relationship between global information, thereby achieving realization in remote interaction problems.
[0007] As a further solution of the present invention: the method includes a tracking framework based on an attention mechanism; the tracking framework includes a feature extraction network, a feature correlation network, a time update strategy and a prediction head;
[0008] The feature extraction network is used to obtain the change of target appearance over time and provide additional time information;
[0009] The feature correlation network uses an attention mechanism to learn the relationship between inputs and predict the spatial location of the target object, effectively capturing templates and regions of interest with more feature map associations;
[0010] The time update strategy is used to capture the changes of the target object over time and enhance robustness;
[0011] The prediction head is used to estimate the target object in the current frame.
[0012] As a further solution of the present invention: the specific method of the feature extraction network framework is to use three groups of image blocks as input, namely, template image Search area image Dynamically update the template image with other inputs It is used to obtain the changes of target appearance over time and provide additional time information. ResNet is used as the backbone for feature extraction, and the last level and fully connected layers of ResNet are removed. After passing through the backbone, template z, z d and the search region image x is mapped to three feature maps and
[0013] As a further solution of the present invention: the specific method of the feature correlation network is to expand the scope of the feature correlation network based on the attention mechanism, obtain global feature information, avoid falling into the local optimum, thereby enhancing the remote feature capture capability;
[0014] The feature correlation network includes an encoder, a decoder and a feature cross correlation (FCC);
[0015] As a further solution of the present invention: the input of the attention mechanism is the query matrix Q, the key matrix K and the value matrix V, and the attention function based on the scaled dot product is defined as equation (1);
[0016]
[0017] Among them, D k Indicates the dimension of the construction; softmax multiplies by row; the dot product of the query and the key is divided by To alleviate the gradient vanishing problem of the softmax function; when Q=K=V, the attention mechanism becomes a self-attention mechanism.
[0018] As a further solution of the present invention: the self-attention mechanism uses multiple attention heads to focus on different aspects of information, that is, in order to obtain information in different representation subspaces at different positions, Q, K, V are projected into multiple linear spaces for attention calculation (1), and the attention results in each linear space are connected in series to obtain multi-head attention (MA), which is defined as equation (2);
[0019]
[0020] in, represents the parameter matrix of the projection, and h is the number of heads.
[0021] As a further solution of the present invention: the encoder and the decoder are both stacks consisting of N identical layers;
[0022] Each layer of the encoder includes two sub-modules: a multi-head self-attention (MSA) module and a feed-forward network (FFN), and the output of each module uses residual connection and layer normalization (LN) operations;
[0023] The decoder includes a third submodule that takes as input the enhanced feature sequence from the encoder to perform multi-head self-attention; residual cascading and layer normalization are used around each layer;
[0024] As a further solution of the present invention: the input of the encoder is preprocessed by using a bottleneck layer to reduce the number of channels from C to d, and then compressing and flattening the three sets of feature maps along the spatial dimension to generate a length of And the feature sequence of dimension is d.
[0025] As a further solution of the present invention: the mathematical method process of the encoder and decoder is:
[0026] I 1 =Flatten(f z ,f d ,f x )
[0027] I n =I n +PE
[0028] I n′ =LN(I n +MSA(I n ))
[0029] I n+1 =LN(I n′ +FFN(I n′ ))
[0030] ………
[0031] T n =T n +PE
[0032] T n′ =LN(T n +MSA(T n ))
[0033] I N′ =I N +PE
[0034] T n′ =LN(T n′ +MA(T n′ ,I N′ ,I N ))
[0035] T n+1 =LN(T n′ +FFN(T n′ ))
[0036] ………
[0037] Output:T N
[0038] Where n represents the nth layer, n∈[0,1,2,…,N], N represents the total number of layers, I N and T N They represent the outputs of the last layer of the encoder and decoder respectively; the sinusoidal position is embedded into the input sequence, denoted as position encoding (PE), as shown in (3);
[0039]
[0040] Where pos represents each position in the sequence, i is the dimension of the vector.
[0041] The above method captures features of all elements in the sequence and enhances the original features with global context information to perform global reasoning on the target, so that the target query can focus on all locations on the template and search for regional features for the final bounding box prediction.
[0042] As a further solution of the present invention: the specific method of the feature cross correlation (FCC) is: the template sequence output by the encoder in the feature correlation network and search region sequence As two branches, they are input into the feature cross-correlation (FCC) module respectively. The feature cross-correlation (FCC) module receives two inputs at the same time, and uses multi-head self-attention to adaptively focus on useful information and cross as feature maps. After fusing the two feature maps through multi-head cross-attention, the template cross-map is output after repeating M times. Cross-mapping with search area Then add the spatial position encoding to the input, and the mathematical description process is described as follows:
[0043]
[0044]
[0045] The feature correlation network uses cross-attention operation to fuse the template and the search area, focusing on the boundary information of the target object and deepening the understanding of the features between the template and the search area.
[0046] As a further solution of the present invention: the time update strategy is to add time information to spatial information to obtain the latest state of the target object for tracking; specifically, a dynamic update template is added to the encoder within the feature correlation network to capture updates from intermediate frames as input, and the encoder extracts spatiotemporal features by modeling the global relationship between all elements in spatial and temporal dimensions, thereby effectively fusing the two types of information and enhancing robustness.
[0047] As a further solution of the present invention: a fractional head is used to control time update. The fractional head is a three-level perceptron, which uses sigmoid activation and sets a threshold τ. When the score is higher than the threshold, the search area is considered to contain the target, and the dynamics are updated by setting the number of frame intervals.
[0048] As a further solution of the present invention: the prediction head is designed using a corner point-based method; that is, the output of the feature correlation network is multiplied by the output of the feature correlation network re-decoder, and the output is multiplied by the corresponding elements of the feature correlation network to enhance the important area and weaken the low-resolution area; the new feature sequence is reshaped into a feature map Then feed it into a fully convolutional network (FCN); FCN consists of L Conv-BN-ReLU layers, which output two mapping probabilities P for the upper left corner and lower right corner of the object bounding box respectively. tl (x,y) and P br (x, y); Finally, the predicted box coordinates are obtained by calculating the expected value of the probability distribution based on the corner points and As shown in Equation 4:
[0049]
[0050] As a further solution of the present invention: the entire tracking framework is trained by adopting a loss function method, and the training process includes two steps;
[0051] Step 1: The entire network except the fractional head in the time update strategy is trained end-to-end, that is, the model is allowed to learn the localization ability. Combining the l1 loss and the generalized IoU loss, the loss function can be written as Equation 5;
[0052]
[0053] Among them, b i , Represents the true bounding box value and the predicted bounding box value; λ iou , is a hyperparameter;
[0054] Step 2: Use binary cross entropy loss to optimize the fractional head, as shown in Equation 6;
[0055] L ce =y i log(P i )+(1-y i )log(1-P i ) (6)
[0056] Among them, y i is the true label of the bounding box, P i is the confidence score, all other parameters are frozen to avoid affecting the localization ability;
[0057] The final model can learn both localization and classification capabilities after two stages of training.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] 1. The framework in this invention combines spatial and temporal dimensions to solve the challenging problems in aerial tracking. The framework mainly consists of feature extraction network, feature correlation network, temporal update strategy and prediction head; it has higher performance and more direct structure;
[0060] 2. The present invention adds a feature correlation network, including an attention module that captures global information and a feature enhancement fusion module based on cross-attention. The network pays more attention to valuable information such as edges and similar targets, and better understands the relationship between global features rather than correlations; in order to capture the changes of targets over time, we design an updated template that increases the attention to temporal information. At the same time, the score head is learned to control the update of the updated template image;
[0061] 3. The present invention adopts a prediction head based on corner points. The whole method is end-to-end and does not require any post-processing steps such as cosine window and bounding box smoothing. This strategy greatly simplifies the existing tracking pipeline and allows the operation of drone trackers.
[0062] 4. Test results on multiple benchmarks show that the proposed tracker has significant performance advantages, especially on the large-scale aerial datasets UAV123 and UAVDT. In addition, our tracker runs at about 42FPS on a GPU, running at real-time speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic diagram of the workflow of the tracking framework in the present invention;
[0064] Figure 2 It is a schematic diagram of the structure of the encoder-decoder in the present invention;
[0065] Figure 3 Schematic diagram of the structure of the feature cross correlation (FCC) module in the present invention;
[0066] Figure 4 It is a structural schematic diagram of the fractional head in the present invention;
[0067] Figure 5(a) shows the tracking performance of TransUAV in a real scene. Figure 1 (occlusion and motion jitter);
[0068] Figure 5(b) shows the tracking performance of the TransUAV in a real scene. Figure 2 (stable tracking);
[0069] Figure 5(c) shows the tracking performance of TransUAV in a real scene. Figure 3 (Stable tracking). DETAILED DESCRIPTION
[0070] The present invention will be described below in detail and completely in conjunction with the accompanying drawings and table parameters; obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0071] This embodiment proposes a Transformer-based UAV target tracking architecture, called TransUAV; Figure 1 As shown, the architecture consists of four modules: feature extractor network, feature correlation network, temporal update strategy and prediction head; Figure 1The workflow of TransUAV during tracking is shown.
[0072] 1. Feature Extraction Network
[0073] Different from the traditional Siamese-based tracker, we use three sets of image patches as input, namely the template image Search area image Dynamically update the template image with other inputs This method uses ResNet as the backbone for feature extraction and removes the last level and fully connected layers of ResNet. d and the search region image x is mapped to three feature maps and
[0074] 2. Feature Correlation Network
[0075] Target tracking from the perspective of drones is seriously affected by the inherent height and flight jitter, small target shape, blurred boundaries, low resolution and other issues, which seriously affect the integration and matching degree between the template and the search area. We use the attention mechanism to expand the scope of feature fusion (i.e., feature correlation network), obtain global feature information, avoid falling into local optimality, and thus enhance the long-range feature capture capability. This part consists of three main modules, namely encoder-decoder and feature cross correlation (FCC).
[0076] The attention mechanism is a basic component of feature fusion, and its input is the query matrix Q, the key matrix K and the value matrix V. The scaled dot product based attention function is defined as Equation (1).
[0077]
[0078] Among them, D k Indicates the dimension of the construction; softmax multiplies by row. The dot product of the query and the key is divided by To alleviate the gradient vanishing problem of the softmax function. When Q = K = V, the attention mechanism becomes a self-attention mechanism.
[0079] As mentioned in, the model cannot jointly focus on multiple aspects of information with one attention head. In order to obtain information in different representation subspaces at different positions, Q, K, V are projected into multiple linear spaces for attention calculation (1), and the attention results in each linear space are concatenated to obtain multi-head attention (MA), defined as equation (2).
[0080]
[0081] in, represents the parameter matrix of the projection, and h is the number of heads. In this work, we use h = 8, d = 256,
[0082] 1. Encoder-Decoder
[0083] Both the encoder and decoder are stacks consisting of N identical layers. The encoder includes two sub-modules in each layer, a multi-head self-attention (MSA) module and a feed-forward network (FFN). The output of each module adopts residual connection and layer normalization (LN) operations. In addition to the two sub-modules in each encoder layer, the decoder has a third sub-module that takes the enhanced feature sequence from the encoder as input to perform multi-head attention. Similar to the encoder, residual cascades and layer normalization are adopted around each sub-layer. In addition, we input a target query in the decoder to predict the bounding box of the target object. The encoder-decoder structure is as follows Figure 2 shown.
[0084] The input to the encoder needs to be preprocessed by first using a bottleneck layer to reduce the number of channels from C to d, and then compressing and flattening the three sets of feature maps along the spatial dimension to generate a length of and a feature sequence of dimension d. This operation merges multiple branches containing spatiotemporal features. This is an intuitive self-attention mechanism that can be performed on the features in each branch to complete the feature extraction step. Compared with a single branch operation, this operation saves computational operations and reduces model parameters through weight sharing.
[0085] The whole process of Encoder-Decoder can be described as:
[0086] I 1 =Flatten(f z ,f d ,f x )
[0087] I n =I n +PE
[0088] I n′ =LN(I n +MSA(I n ))
[0089] I n+1 =LN(I n′ +FFN(I n′ ))
[0090] …
[0091] T n =T n +PE
[0092] T n′ =LN(T n +MSA(T n ))
[0093] I N′ =I N +PE
[0094] T n′ =LN(T n′ +MA(T n′ ,I N′ ,I N ))
[0095] T n+1 =LN(T n′ +FFN(T n′ ))
[0096] …
[0097] Output:T N
[0098] Among them, n represents the nth layer, n∈[0,1,2,…,N], and N represents the total number of layers. Therefore, I N and T N denote the outputs of the last layer of the encoder and decoder, respectively. Due to the permutation invariance of the transformer, it cannot understand the input order, so we embed the sinusoidal position into the input sequence, denoted as position encoding (PE), as shown in (3).
[0099]
[0100] Where pos represents each position in the sequence, pos∈[0,max_sequence_length], i is the dimension of the vector.
[0101] The Encoder-Decoder mechanism captures features of all elements in the sequence and enhances the original features with global context information to perform global reasoning on the object, allowing the object query to attend to all locations on the template and search for regional features for the final bounding box prediction.
[0102] 2. Feature Cross Correlation (FCC)
[0103] In the study of target tracking, the correlation and matching degree between the template and the search area determine the output position of the bounding box. In order to deal with the problems of small target objects, blurred boundaries, and low resolution from the perspective of drones, this work designs a feature correlation network to replace the traditional correlation operation, process and fuse the features of the template and the search area, such as Figure 1 As shown in the figure, the output of the Encoder has completed the feature extraction. In order to strengthen the feature correlation dependency between the template and the search area, we convert the template sequence output by the Encoder into and search region sequence As two branches, they are input into the feature cross-correlation (FCC) module respectively. The two feature cross-correlation (FCC) modules receive two inputs at the same time, firstly use multi-head self-attention to adaptively focus on useful information, and cross as feature maps, and then use multi-head cross-attention to fuse the two feature maps, repeat M times and output the template cross-map Cross-mapping with search area The structure of feature cross correlation (FCC) is as follows Figure 3 As shown in the figure, similar to Encoder, it is necessary to add spatial position encoding to the input, and feature cross-correlation (FCC) is used to enhance the fitting ability of the model. The process of the entire feature correlation network is described as follows:
[0104]
[0105]
[0106] The feature correlation network uses cross-attention operation to fuse the template and the search area, focusing on the boundary information of the target object and deepening the understanding of the features between the template and the search area.
[0107] 3. Time Update Strategy
[0108] Since the appearance of the target object may change significantly over time, the temporal information must be added to the spatial information to obtain the latest state of the target object for tracking. Therefore, a temporal information branch is designed to update the temporal information to obtain the appearance of the target. Figure 2 As shown, a dynamically updated template is added to the encoder to capture updates from intermediate frames as input, and the encoder extracts spatiotemporal features by modeling the global relationships between all elements in both spatial and temporal dimensions, thereby effectively fusing two types of information and enhancing robustness.
[0109] 1. Score Head During the tracking process, the appearance of the template does not fluctuate significantly over several consecutive frames, and there are some situations where the dynamic template should not be updated. For example, when the target is completely occluded or the tracker drifts, the cropped template is unreliable to judge whether the current state is reliable. We recommend applying a score head to control the update of the dynamic template, such as Figure 4 As shown in Figure 2. It is a three-level perceptron followed by a sigmoid activation. In order to simplify the framework to save computational cost, we set a threshold τ. When the score is higher than the threshold, the search area is considered to contain the target, and the frame interval is set to more than 50 frames before the dynamic update template can be updated.
[0110] 4. Prediction Head
[0111] Corner-based prediction head: Considering the computing power of the drone and avoiding the computational cost brought by post-processing of the predicted target box, the structure of the prediction head must be simple and robust. In order to improve the quality of target box estimation, we adopt a corner-based approach to design the prediction head. As shown in Table 1, first, the similarity between the output of the feature correlation network and the output of the decoder is calculated. Since the output of the feature correlation network is the result of the fusion of the search area and the template, and the output of the decoder is a global information feature extraction containing temporal changes, the result calculated between the two contains a large amount of target boundary information and spatial information. Then, the output is element-wise multiplied with the feature correlation network to enhance important areas and weaken low-resolution areas. The new feature sequence is reshaped into a feature map It is then fed into a simple fully convolutional network (FCN). The FCN consists of L Conv-BN-ReLU layers, which output two mapping probabilities P for the upper left corner and lower right corner of the object bounding box, respectively. tl (x,y) and P br (x,y). Finally, the predicted box coordinates are obtained by calculating the expected value of the probability distribution based on the corner points and As shown in Equation 4.
[0112]
[0113] 5. Loss Function
[0114] We divide the training process into two steps. In the first stage, the entire network except the score head is trained end-to-end. Regarding the first step, localization is a primary task to ensure that all search images contain the target object and allow the model to learn localization capabilities. The loss function can be written as Equation 5 by combining l1 loss and generalized IoU loss.
[0115]
[0116] Among them, bi , Represents the true bounding box value and the predicted bounding box value; λ iou , is a hyperparameter.
[0117] In the second stage, the fractional head is optimized using binary cross entropy loss as shown in Equation 6.
[0118] L ce =y i log(P i )+(1-y i )log(1-P i ) (6)
[0119] Among them, y i is the true label of the bounding box, P i is the confidence score, and all other parameters are frozen to avoid affecting the localization ability. The final model can learn both localization and classification capabilities after two stages of training.
[0120] Experimental process and effect comparison of this embodiment
[0121] First, we introduce the implementation details and results of TransformerUAV on multiple benchmarks and compare them with multiple state-of-the-art methods. Then we evaluate TransformerUAV on three challenging UAV benchmarks and quantitatively and qualitatively analyze the tracking performance of the tracker under UAV motion. Finally, we conduct an ablation study to analyze the impact of key components in the proposed network and comprehensively verify the effectiveness and superiority of the proposed tracker.
[0122] 2. Results and Comparison
[0123] We compare the TransUAV approach with multiple representative trackers on six benchmarks, including the standard aerial tracking benchmarks UAV123, UAV123@10fps, UAVDT, two short-term benchmarks GOT-10K, TrackingNet, and one long-term benchmark LaSOT.
[0124] In order to ensure fairness and objectivity, we selected a total of 32 latest and classic tracking methods for evaluation, including correlation filter-based methods and deep learning-based methods. The results were obtained by running the official code with corresponding hyperparameters. For the UAV123, UAV123@10fps, UAVDT, LaSOT, and TrackingNet benchmarks, the experiments were based on one-pass evaluation (OPE), involving two indicators: precision (P) and success rate (AUC). The GOT-10K benchmark uses two evaluation indicators: average overlap AO and SR success rate.
[0125] UAV123 is a large-scale aerial tracking benchmark involving 123 challenging aerial video sequences. It is one of the most authoritative and comprehensive datasets in the field of UAV tracking. The tracking performance of 19 trackers under the most common aerial tracking conditions was evaluated using UAV123, and the results are shown in Table 1. Our proposed TransUAV ranks first in both success rate and accuracy. In terms of accuracy, TransUAV obtained 0.904, which is 2.6% higher than the second-ranked PrDiMP50 and 2.8% higher than TransT, which is also based on the self-attention mechanism. In terms of success rate, TransUAV obtained 0.692, which is 2.3% and 1.1% higher than PrDiMP50 and TransT. This fully demonstrates the effectiveness of feature correlation networks and temporal information acquisition. In addition, we also compared the performance of some trackers on each attribute, as shown in Table 2. TransUAV outperforms other trackers on each attribute, with good robustness and anti-interference ability, especially under similar targets and partial occlusion attributes, with AUC of 0.680 and 0.641, which are 3.9% and 2.2% higher than the second place.
[0126] Table 1: Area Under the Circumference (AUC) and Precision (P) of TransUAV and multiple trackers tested on the datasets UAV123, UAV123@fps, and UAVDT.
[0127]
[0128] Table 2: AUC-based attribute scores of TransUAV and multiple trackers on the UAV123 dataset. Red indicates the top-ranked tracker.
[0129]
[0130] Table 3: AUC-based attribute scores of TransUAV and multiple trackers on the UAV123@10fps dataset. Red indicates the top-ranked tracker.
[0131]
[0132] In order to rigorously evaluate the TransUAV method, we also used the more challenging UAV123@10fps. This benchmark uses an image rate of 10 frames per second. Between consecutive frames, the movement and changes of the target are more drastic and severe, which greatly increases the difficulty of tracking. As shown in Table 1, it can be clearly seen from the comparison with other trackers that TransUAV ranks first in accuracy (0.904) and success rate (0.694), 2.7% and 2.1% higher than the second PrDiMP50. Compared with UAV123, our tracking performance has not declined and maintained excellent robustness. And it performs well in most attributes, and the overall performance has not dropped significantly, as shown in Table 3, proving that TransUAV is able to handle large changes in target appearance over time.
[0133] UAVDT consists of 50 challenge sequences, focusing on aircraft tracking in various unconstrained complex scenarios such as high density, small targets, and weather conditions. We tested and evaluated 13 trackers on the UAVDT benchmark. The results are shown in Table 1. The TransUAV tracker achieved the best accuracy score (83.4) and success rate (60.7), which is 2.5% and 1.5% higher than the second-ranked PrDiMP50. In addition, our evaluation on each attribute performed well, as shown in Table 4, especially in the case of target blur and target motion, proving that TransUAV is suitable for complex aircraft tracking scenarios.
[0134] Table 4: AUC-based attribute scores of TransUAV and multiple trackers on the UAVDT dataset. Red indicates the top-ranked tracker.
[0135]
[0136] LaSOT is a large-scale high-quality target tracking dataset released in 2019, which contains 1,400 challenging videos, 120 for training and 280 for testing. We compared 16 trackers on the test set, and Table VI shows that TransUAV achieved the best performance, with AUC and P scores of 0.649 and 0.693. Although TransUAV's advantage over TransT is not significant, it performs better than other trackers in the evaluation based on all attributes, as shown in Table 5, indicating that the method proposed in this paper is not only applicable to tracking scenarios under aerial perspective, but also effective for general large-scale tracking benchmarks.
[0137] Table 5: AUC-based attribute scores of TransUAV and multiple trackers on the LaSOT dataset. Red indicates the top-ranked tracker.
[0138]
[0139] Table 6: Test success rate (AUC) and precision (P) of TransUAV and multiple trackers on the LaSOT dataset. Red, green, and blue represent the first, second, and third ranked trackers respectively.
[0140]
[0141] TrackingNet is a large-scale short-term tracking dataset. The test set contains 511 sequences covering various object classes and scenes. We submit the output of TransUAV to the official evaluation server. The results are shown in Table VII. TransUAV has the highest performance in AUC, P, and P. Norm In terms of performance, they obtained 81.2%, 77.7%, and 85.1% respectively, second only to SiamRCNN. However, SiamRCNN runs at less than 5fps on our device, which cannot meet the real-time requirements, while TransUAV runs at 42FPS.
[0142] The GOT-10k dataset contains 10k training sequences and 180 test sequences. We use the test sequences for model testing and submit the output results to the official evaluation server. The AO and SR metric scores of multiple trackers are reported in Table VII. TransUAV has the best performance, AO, SR 0.5 ,SR 0.75 The scores of the proposed method are 3.9%, 5.1%, and 3.6% higher than those of SiamRCNN, respectively.
[0143] Table 7: Overall performance of TrackingNet and GOT-10k.
[0144]
[0145]
[0146] 3. Physical Experiment
[0147] We implement our tracker on a UAV to verify its performance in a real-world environment. This experiment uses NVIDIA's Jetson AGX Xavier as the onboard computer and Pixhawk4 as the flight controller. In the test, we observed that the average utilization of the GPU and CPU was 37% and 52% respectively.
[0148] Figure 5 shows the tracking performance of TransUAV in real-world scenarios. The center position error (CLE) is used to evaluate the tracking performance, involving scenarios such as occlusion, low resolution, motion blur, and scale changes. As shown in Figure 5(a), TransUAV can still maintain excellent stability and robustness when encountering partial occlusion and motion jitter. Figures 5(b) and (c) show the stable tracking performance of TransUAV, which effectively reduces redundancy and corrects the tracked target, and can achieve satisfactory tracking even under the interference of similar targets. In addition, without TensorRT3 acceleration, we maintained an airborne running speed of more than 12FPS, which shows the practicality and feasibility of TransUAV in complex aerial tracking conditions.
[0149] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for tracking a target of an unmanned aerial vehicle, characterized in that: This method uses the principle of self-attention mechanism; the self-attention mechanism is used to scan each element in the sequence and update it by aggregating information in the entire sequence to focus on the relationship between global information, thus achieving the goal of long-range interaction problem; The method includes a tracking framework based on an attention mechanism; the tracking framework includes a feature extraction network, a feature correlation network, a time update strategy and a prediction head; The feature extraction network is used to obtain feature information and reduce data dimensions; the attention mechanism is used to learn the relationship between inputs and predict the spatial position of the target object, effectively capturing templates and regions of interest with more feature map associations; The time update strategy is used to capture the changes of the target object over time and enhance robustness; The prediction head is used to estimate the target object in the current frame; The specific method of the feature extraction network framework is to use three groups of image blocks as input, namely, the template image Search area image Dynamically update the template image with other inputs It is used to obtain the changes of target appearance over time and provide additional time information. ResNet is used as the backbone for feature extraction, and the last level and fully connected layers of ResNet are removed. After feature extraction using ResNet as the backbone, template z, z d and the search region image x is mapped to three feature maps and The specific method of the feature correlation network is to expand the scope of the feature correlation network based on the attention mechanism, obtain global feature information, avoid falling into the local optimum, and thus enhance the long-range feature capture capability; The feature correlation network includes an encoder, a decoder and a feature cross correlation (FCC); The specific method of the feature cross correlation (FCC) is: the template sequence output by the encoder in the feature correlation network and search region sequence As two branches, they are input into the feature cross-correlation (FCC) module respectively. The feature cross-correlation (FCC) module receives two inputs at the same time, and uses multi-head self-attention to adaptively focus on useful information and cross as feature maps. After fusing the two feature maps through multi-head cross-attention, the template cross-map is output after repeating M times. Cross-mapping with search area Then add the spatial position encoding to the input, and the mathematical description process is described as follows: The feature correlation network uses cross-attention operation to fuse the template and the search area, focusing on the boundary information of the target object and deepening the understanding of the features between the template and the search area; The entire tracking framework is trained using a loss function. The training process includes two steps. Step 1: The entire network except the score head in the time update strategy is trained end-to-end, that is, the model is allowed to learn the positioning ability; the loss function can be written as Equation 5 by combining the l1 loss and the generalized IoU loss; Among them, b i , Represents the true bounding box value and the predicted bounding box value; λ iou , is a hyperparameter; Step 2: Use binary cross entropy loss to optimize the fractional head, as shown in Equation 6; L ce =y i log(P i )+(1-y i )log(1-P i ) (6) Among them, y i is the true label of the bounding box, P i is the confidence score, all other parameters are frozen to avoid affecting the localization ability; The final model can learn both localization and classification capabilities after two stages of training.
2. The method for tracking a target by using an unmanned aerial vehicle according to claim 1, characterized in that: The input of the attention mechanism is the query matrix Q, the key matrix K and the value matrix V. The scaled dot product based attention function is defined as Equation (1); Among them, D k Indicates the dimension of the construction; softmax multiplies by row; the dot product of the query and the key is divided by To alleviate the gradient vanishing problem of the softmax function; when Q=K=V, the attention mechanism becomes a self-attention mechanism.
3. A method for tracking a target by using an unmanned aerial vehicle according to claim 2, characterized in that: The self-attention mechanism uses multiple attention heads to focus on different aspects of information, that is, in order to obtain information in different representation subspaces at different positions, Q, K, V are projected into multiple linear spaces for attention calculation (1), and the attention results in each linear space are concatenated to obtain multi-head attention (MA), defined as equation (2); in, represents the parameter matrix of the projection, and h is the number of heads.
4. The method for tracking a target by using an unmanned aerial vehicle according to claim 1, characterized in that: The encoder and decoder are both stacks consisting of N identical layers; Each layer of encoder consists of two sub-modules: multi-head self-attention (MSA) module and feed-forward network (FFN). The output of each module adopts residual connection and layer normalization (LN) operation; The decoder includes a third submodule that takes the enhanced feature sequence from the encoder as input to perform multi-head self-attention; residual cascading and layer normalization are used around each layer.
5. A method for tracking a target by using an unmanned aerial vehicle according to claim 4, characterized in that: The input of the encoder is preprocessed by using a bottleneck layer to reduce the number of channels from C to d, and then compressing and flattening the three sets of feature maps along the spatial dimension to generate a length of And the feature sequence of dimension is d.
6. A method for tracking a target of an unmanned aerial vehicle according to any one of claims 4 or 5, characterized in that The mathematical process of the encoder and decoder is: I 1 =Flatten(f z ,f d ,f x ) I n =I n +PE I n′ =LN(I n +MSA(I n )) I n+1 =LN(I n′ +FFN(I n )) ……… T n =T n +PE T n′ =LN(T n +MSA(T n )) I N′ =I N +PE T n′ =LN(T n′ +MA(T n′ ,I N′ ,I N )) T n+1 =LN(T n′ +FFN(T n′ )) ……… Output:T N Where n represents the nth layer, n∈[0,1,2,…,N], N represents the total number of layers, I N and T N They represent the outputs of the last layer of the encoder and decoder respectively; the sinusoidal position is embedded into the input sequence, denoted as position encoding (PE), as shown in (3); PE(pos,2i)=sin(pos / 10000 2i / d ) (3) PE(pos,2i+1)=cos(pos / 10000 2i / d ) Where pos represents each position in the sequence, pos∈[0,max_sequence_length], i is the dimension of the vector; The above method captures features of all elements in the sequence and enhances the original features with global context information to perform global reasoning on the target, so that the target query can focus on all locations on the template and search for regional features for the final bounding box prediction.
7. The method for tracking a target by using an unmanned aerial vehicle according to claim 1, characterized in that: The temporal update strategy is to add temporal information to spatial information to obtain the latest state of the target object for tracking; specifically, a dynamic update template is added to the encoder within the feature correlation network to capture updates from intermediate frames as input. The encoder extracts spatiotemporal features by modeling the global relationship between all elements in spatial and temporal dimensions, thereby effectively fusing the two types of information and enhancing robustness.
8. The method for tracking a target by using an unmanned aerial vehicle according to claim 7, characterized in that: A score head is used to control time update. The score head is a three-level perceptron with sigmoid activation. A threshold τ is set. When the score is higher than the threshold, the search area is considered to contain the target, and the dynamics are updated by setting the number of frame intervals.
9. The method for tracking a target by using an unmanned aerial vehicle according to claim 1, characterized in that: The prediction head is designed based on the corner point method; that is, the output of the feature correlation network is multiplied by the output of the feature correlation network re-decoder, and the output is multiplied by the corresponding elements of the feature correlation network to enhance the important areas and weaken the low-resolution areas; the new feature sequence is reshaped into a feature map Then feed it into a fully convolutional network (FCN); FCN consists of L Conv-BN-ReLU layers, which output two mapping probabilities P for the upper left corner and lower right corner of the object bounding box respectively. tl (x,y) and P br (x, y); Finally, the predicted box coordinates are obtained by calculating the expected value of the probability distribution based on the corner points and As shown in Equation 4: