Unmanned aerial vehicle target long-time tracking method based on mixed attention mechanism and hierarchical discriminator
By introducing a hybrid attention mechanism and a layered discriminator, the problems of robustness and template update of drone target tracking in complex scenarios are solved, and efficient and accurate long-term tracking effect is achieved.
Patent Information
- Application Number
- CN202510528726.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-15
AI Technical Summary
The existing drone target tracking methods are weak in complex scenarios, especially when the target appearance changes, and insufficient template updates and reliability assessments lead to poor long-term tracking results.
A drone target long-term tracking method based on a hybrid attention mechanism and a hierarchical discriminator is adopted. Through a conjoined network structure of a multi-attention mechanism, combined with ResNet50 and SiamRPN++, a multiple hybrid attention module and a hierarchical discriminator are introduced, and feature weights are adjusted adaptively, reliability evaluation criteria are generated, and the re-detection module update template is activated when unreliable.
It significantly improves the robustness and accuracy of drone target tracking, can adapt to complex scenarios and target appearance changes, achieve efficient long-term tracking performance, and meet real-time tracking requirements.
Smart Images

Figure CN120495336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a long-term tracking method for UAV targets, and more specifically to a long-term tracking method for UAV targets based on a hybrid attention mechanism and a hierarchical discriminator. Background Art
[0002] In recent years, drones (drones) have been widely used in military, transportation, logistics, and security fields due to their high efficiency and low cost. However, the illegal and unregulated use of drones has also brought numerous security risks, such as interference with civil aviation and invasion of privacy, posing a significant threat to public safety. To ensure public safety, there is an urgent need to develop effective counter-drone systems, particularly those that utilize target tracking technology to monitor drone activity in real time.
[0003] Target tracking technologies are primarily categorized into correlation filtering-based and deep learning-based methods. Correlation filtering-based methods employ filters to identify and track targets, offering the advantages of high computational efficiency but limited robustness in complex scenarios. In contrast, deep learning methods, which learn characteristic patterns from large amounts of data to establish models that distinguish targets from backgrounds, have become a hot topic of research. However, deep learning methods still face challenges in real-time performance and online updates, particularly in adapting to changes in the appearance of drone targets.
[0004] In recent years, trackers based on Siamese networks have made significant progress. These methods treat object tracking as a matching problem, with the core idea being to learn a similarity mapping between a target template and a search region. By transforming the tracking problem into a similarity calculation, Siamese networks have demonstrated excellent accuracy and real-time performance. However, existing Siamese network models (such as CFNet and SiamRPN) are typically trained entirely offline, preventing templates from being updated online and making them difficult to adapt to changes in target appearance. Furthermore, existing methods only generate feature maps based on the target during template updates, ignoring the contextual information between positive and negative samples, making it difficult to achieve high-quality performance in practical tracking scenarios.
[0005] Long-term tracking tasks are more complex, especially when the target is severely occluded or out of sight, the tracking results may become unreliable. Therefore, methods for tracking and re-detection through detection become crucial. However, existing long-term tracking methods usually combine correlation filtering with image detection. Although more robust in practical applications, the cost of each frame detection is high. In addition, existing methods often rely on correlation filters of manually created features for reliability assessment, failing to fully utilize the feature representation capabilities of deep learning. As a result, accurately identifying whether a tracker has failed remains a difficult problem. Summary of the Invention
[0006] The present invention provides a long-term tracking method for UAV targets based on a hybrid attention mechanism and a hierarchical discriminator, the purpose of which is to be able to.
[0007] The above objectives are achieved through the following technical solutions:
[0008] A long-term tracking method for UAV targets based on hybrid attention mechanism and hierarchical discriminator.
[0009] Step 1: A Siamese network structure with a multi-attention mechanism for target tracking, including a ResNet50 deep neural network model and SiamRPN++ as a baseline for drone target tracking. The baseline consists of a feature extraction network, a region proposal network, and an attention module.
[0010] Step 2: The Siamese network consists of a template branch and a search branch. The template branch extracts the target feature from the first frame of the input video as the template z. The search branch extracts the image feature x from the search area of subsequent frames and uses the cross-correlation function to measure its similarity with the template z.
[0011] Step 3: According to SiamRPN++, the multi-level convolutional features of the last three residual blocks in ResNet50 are aggregated and input into three RPN modules respectively, and then weighted summation is directly applied on the RPN output;
[0012] Step 4: Introduce a multi-hybrid attention mechanism to adaptively adjust the weight of each channel feature. The multi-hybrid attention module includes channel and spatial attention modules, as well as a contextual attention module.
[0013] Step 5: Channel attention: Through average pooling and maximum pooling operations, channel information is summarized from the feature map to obtain two one-dimensional global context descriptors, which are then input into a fully connected network with shared parameters to generate global channel attention weights.
[0014] Step 6: Spatial attention uses average pooling and maximum pooling techniques to compress the feature map along the channel dimension to generate two two-dimensional feature maps, which are then convolved through a single-layer convolutional network to generate a local spatial attention map.
[0015] Step 7: The channel-spatial attention module is run in parallel with the contextual attention module. The dot product of the contextual attention matrix and the feature map matrix is performed to obtain the output of the attention module.
[0016] Step 8. Select k anchor points. The classification branch will generate 2k response maps, representing the negative activation and positive activation of each anchor point respectively. The anchor point with the highest probability of being the correct sample will be regarded as the target position.
[0017] The beneficial effects of the long-term tracking method of UAV targets based on the hybrid attention mechanism and hierarchical discriminator of the present invention are as follows:
[0018] The attention mechanism and multi-scale feature fusion mechanism are introduced into the neural network, and the attention distribution fit is improved through three series-parallel attention sub-networks. The hybrid attention module enhances the feature learning ability of the Siamese network through training. The feature maps extracted from the network contain richer semantic information, which enhances the model's discriminative performance and generates more discriminative object representations.
[0019] To address the challenges of relocalizing and updating templates for UAV targets during long-term tracking, a hierarchical discriminator generates a response map for target localization based on the output of the Siamese network and establishes a reliability criterion to assess the credibility of the response map. When the output indicates low credibility, the algorithm activates the re-detection module and updates the template. This effectively addresses the challenges of existing methods in template updating, reliability assessment, and long-term tracking, significantly improving the robustness and accuracy of UAV target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The basic framework diagram of the anti-UAV long-term target tracking method is shown;
[0021] Figure 2 The diagram of the Siamese network structure based on the multi-attention mechanism is shown;
[0022] Figure 3 Shows the architecture diagram of the multiple hybrid attention mechanism;
[0023] Figure 4 A comparison of SiamAD with 9 SOTA trackers is shown. DETAILED DESCRIPTION
[0024] A method for long-term tracking of UAV targets based on a hybrid attention mechanism and a hierarchical discriminator includes the following steps:
[0025] Step 1: A Siamese network architecture with multiple attention mechanisms for target tracking, including the ResNet50 deep neural network model and SiamRPN++ (Siamese Region Proposal Networks) as a baseline for drone target tracking. The baseline consists of a feature extraction network, a region proposal network (RPN), and an attention mechanism (AM).
[0026] Step 2: The Siamese network consists of a template branch and a search branch. The template branch extracts the target feature from the first frame of the input video as the template z, and the search branch extracts the image feature x from the search area of the subsequent frames and uses the cross-correlation function to measure its similarity with the template z. That is:
[0027] f(z,x)=f(φ(x),φ(z))+b·I;
[0028] b represents the offset of the model, I is the identity matrix, and φ(·) represents the semantic embedding space.
[0029] Step 3: According to SiamRPN++, the multi-level convolutional features of the last three residual blocks in ResNet50 are aggregated and input into the three RPN modules respectively. Then, the weighted summation is directly applied to the RPN output, i.e.:
[0030]
[0031] α l , β l Represents the data weight value, R l , C l Represents the multi-level convolution eigenvalue of the residual block, R out , C out Indicates the output value.
[0032] Step 4: Introduce a multi-hybrid attention mechanism to adaptively adjust the weight of each channel feature. The multi-hybrid attention module includes the Channel and Spatial Attention Module (CSAM) and the Context Attention Module (CAM).
[0033] Step 5: Channel Attention Through average pooling and maximum pooling operations, channel information is summarized from the feature map to obtain two one-dimensional global context descriptors: u a and u b These two descriptors are then fed into a fully connected network with shared parameters to generate global channel attention weights.
[0034] u a =Avgpool(F),u b =Maxpool(F)
[0035] M Global (X)=σ(F1(F0(u a )))+σ(F1(F0(u b )))
[0036] Spatial attention uses average pooling and maximum pooling techniques to compress the feature map along the channel dimension to generate two two-dimensional feature maps: a and f m Then, the feature map is convolved through a single-layer convolutional network to generate a local spatial attention map M Local :
[0037]
[0038] Here, σ represents the sigmoid function.
[0039] Step 6: Use the context attention module based on the self-attention mechanism to The feature vector is input into two parallel 1×1 convolutional layers to obtain:
[0040]
[0041] Denote the two output feature vectors corresponding to the input signal X. U and V are then reshaped into Get the relationship matrix
[0042] Then, the weights are normalized using average pooling and sigmoid function:
[0043] M Context =σ(Avgpool(M));
[0044] Step 7: The channel-spatial attention module is parallel to the contextual attention module, and the contextual attention matrix M Context and the feature map matrix M out Perform dot product to get the output of the attention module:
[0045] T AM =M Context ⊙M out ;
[0046] Step 8. Select k anchor points. The classification branch will generate 2k response maps, representing the negative activation and positive activation of each anchor point. The anchor point with the highest probability of being the correct sample will be considered the target position. The probability score calculation method is as follows:
[0047] S p =max(Softmax(z0,z1));
[0048] Where z0,z1 represents the positive activation and negative activation of each anchor point. A threshold is assigned to each video. Specifically, the history S p Values are grouped into a set M sDefine a coefficient O s ,when , the tracking result is determined to be unreliable.
[0049] Step 9. The peak-to-sidelobe ratio (APCE) is an indicator that quantifies the sharpness of the correlation peak. If the APCE value is low, the correlation between the search and the template may not be strong. APCE is introduced into SiamRPN++. Its calculation formula is:
[0050]
[0051] Among them F max and F min are the maximum response and minimum response of the current frame respectively. w,h Represents the element value in the wth row and hth column of the response matrix.
[0052] Compute the weighted sum of the response maps generated by the RPN:
[0053]
[0054] Among them, M i ,M j Response maps representing foreground and background activations, respectively. and Represents the adaptive weight used for normalization, that is, the maximum activation value of each channel.
[0055] With S p Scores are similar, APCE scores are calculated and recorded in represents the APCE score of the i-th frame, M a Defined as C a The average value of the pool. If The current APCE score is considered If the tracking result is obviously too low, such as unsafe, the re-detection module should be started.
[0056] Step 10: When tracking fails, enable the re-detection module. YOLOv5 is selected as the re-detection network, which mainly includes (Convolution + Batch Normalization + LeakyReLU) CBL and (Cross Stage Partial Network) CSPNet modules. CBL consists of two-dimensional convolution, batch normalization, and LeakyReLU activation functions. CSPNet, consisting of two-dimensional convolution, res unit, and CBL modules, is primarily responsible for enhancing the learning capabilities of convolutional neural networks (CNNs), ensuring accuracy while being lightweight. The network input is an RGB image of size 3×608×608, and the output is a feature map of size 32×304×304.
[0057] Step 11: Use the update discriminator to update the template, employing an adaptive update strategy based on the drone's historical information. When the re-detection module is activated, it indicates a significant appearance change, and the template must be updated to match the new tracking system. The YOLO network outputs a detection score, and when the YOLO score exceeds a certain threshold, the template is updated.
[0058] Adopting adaptive weight template calculation model:
[0059]
[0060] in represents the updated template of the current frame, is the target template of the initial frame, It is a template generated based on historical information. i Represents the template of the current frame. and It is an adaptive weight determined based on the APCE score and the YOLO score. When the APCE value exceeds the specified value, the weight is adjusted to:
[0061]
[0062] in, is the APCE value of the current frame, set Contains the confidence information of all historical templates. The activation of the re-detection module indicates that the appearance has changed fundamentally compared with the historical template, and T needs to be reduced. i-1 Weight:
[0063]
[0064] Among them, S YOLO is the detection score of YOLO.
[0065] Comparison: As a benchmark for dynamic environments, the anti-drone dataset proposed at the First International Workshop on Anti-drone Challenge at CVPR2020 is used. This dataset consists of more than 300 video pairs (including RGB and infrared videos), annotated with more than 580k bounding boxes, covering multiple occurrences of multi-scale drones (i.e., large, small, and micro drones), mainly from DJI and Parrot.
[0066] In the field of anti-UAV tracking, the proposed algorithm SiamAD is compared with nine recently launched SOTA trackers: MDNet, SiamDW, SiamFC, SiamRPN, SiamRPN++, TransT, SiamCAR, SiamRPN++LT and CLNet. In addition, two indicators are introduced to evaluate the performance of all methods: distance precision (DP) under a threshold of 20 pixels and the area under the curve (AUC) of the success map. The score in the precision map is defined as the percentage of frames whose center position error is less than a predetermined threshold. At the same time, the score in the success map represents the percentage of frames whose overlap rate between the tracking area and the boundary frame is greater than the threshold. In the training and testing stages, equal-sized patches of 127 pixels are used as templates and patches of 255 pixels are used as search areas.
[0067] Compared to other state-of-the-art trackers for anti-UAV tracking, SiamAD achieves the best results in terms of AUC (67.7%) and DP (88.4%). Compared to the second-ranked SiamRPN++LT, SiamAD improves AUC and DP by 7.3% and 9.0%, respectively. Furthermore, the tracker achieves significant improvements over the baseline SiamRPN++, with a relative increase of 13.7% in success rate and 16.5% in accuracy. Notably, SiamAD achieves a speed of 38.8 FPS, meeting real-time tracking requirements. These results demonstrate that the proposed multi-scale feature extraction network, hybrid attention module, and hierarchical discriminator can jointly adaptively adapt to environmental changes and improve tracking performance in complex scenarios. This is because the re-detection module provides high-quality templates for algorithm updates, and SiamAD utilizes all reliable features to generate robust templates with rich information, resulting in efficient and accurate UAV target tracking.
Claims
1. A long-term tracking method for UAV targets based on a hybrid attention mechanism and a hierarchical discriminator. Step 1: A Siamese network structure with a multi-attention mechanism for target tracking, including a ResNet50 deep neural network model and SiamRPN++ as a baseline for drone target tracking. The baseline consists of a feature extraction network, a region proposal network, and an attention module. Step 2: The Siamese network consists of a template branch and a search branch. The template branch extracts the target feature from the first frame of the input video as the template z. The search branch extracts the image feature x from the search area of subsequent frames and uses the cross-correlation function to measure its similarity with the template z. Step 3: According to SiamRPN++, the multi-level convolutional features of the last three residual blocks in ResNet50 are aggregated and input into three RPN modules respectively, and then weighted summation is directly applied on the RPN output; Step 4: Introduce a multi-hybrid attention mechanism to adaptively adjust the weight of each channel feature. The multi-hybrid attention module includes channel and spatial attention modules, as well as a contextual attention module. Step 5: Channel attention: Through average pooling and maximum pooling operations, channel information is summarized from the feature map to obtain two one-dimensional global context descriptors, which are then input into a fully connected network with shared parameters to generate global channel attention weights. Step 6: Spatial attention uses average pooling and maximum pooling techniques to compress the feature map along the channel dimension to generate two two-dimensional feature maps, which are then convolved through a single-layer convolutional network to generate a local spatial attention map. Step 7: The channel-spatial attention module is run in parallel with the contextual attention module. The dot product of the contextual attention matrix and the feature map matrix is performed to obtain the output of the attention module. Step 8. Select k anchor points. The classification branch will generate 2k response maps, representing the negative activation and positive activation of each anchor point respectively. The anchor point with the highest probability of being the correct sample will be regarded as the target position.
2. According to the long-term tracking method for UAV targets based on the hybrid attention mechanism and hierarchical discriminator in claim 1, the region candidate network is RPN, the attention module is AM; the spatial attention module is CSAM, and the contextual attention module is CAM.
3. According to the method for long-term tracking of UAV targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 1, the probability score calculation method is as follows: S p =max(Softmax(z0,z1); in, z0,z1 represent the positive activation and negative activation of each anchor point, and a threshold is assigned to each video. The history S p Values are grouped into a set M s Define a coefficient O s ,when , the tracking result is determined to be unreliable.
4. The method for long-term tracking of UAV targets based on hybrid attention mechanism and hierarchical discriminator according to claim 1: Step 9: If the APCE value is low, introduce APCE into SiamRPN++; if The current APCE score is considered If the tracking result is obviously too low, such as unsafe, the re-detection module should be started.
5. The method for long-term tracking of unmanned aerial vehicle targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 1 uses YOLO as the re-detection network, including CBL and CSPNet modules. CBL includes two-dimensional convolution, batch normalization, and LeakyReLU activation function, and CSPNet includes two-dimensional convolution, res unit, and CBL.
6. According to the method for long-term tracking of drone targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 5, the input of the network is an RGB image of size 3×608×608, and the output is a feature map of size 32×304×304.
7. The method for long-term tracking of drone targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 6: Step 11: Use the updated discriminator to update the template, and adopt an adaptive update strategy based on the drone's historical information.
8. In the method for long-term tracking of unmanned aerial vehicle targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 7, when the re-detection module is activated, it indicates that the appearance has changed significantly, so the template is updated to match the new tracking system, and the YOLO network can output a score of the detection result. When the YOLO score is greater than a specified threshold, the template is updated.
9. The method for long-term tracking of unmanned aerial vehicle targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 7, wherein the adaptive weight template calculation model is used: represents the updated template of the current frame, is the target template of the initial frame, It is a template generated based on historical information. i Represents the template of the current frame, and It is an adaptive weight determined by the APCE score and the YOLO score. When the APCE value exceeds the specified value, the weight is adjusted to: is the APCE value of the current frame, set Contains confidence information for all historical templates.
10. The method for long-term tracking of unmanned aerial vehicle targets based on a hybrid attention mechanism and a hierarchical discriminator according to claim 7, wherein when the APCE value exceeds a specified value, the weight is adjusted to: is the APCE value of the current frame, set Contains the confidence information of all historical templates. The activation of the re-detection module indicates that the appearance has changed fundamentally compared with the historical template, and T needs to be reduced. i-1 Weight: S YOLO is the detection score of YOLO.