Unmanned aerial vehicle target tracking method based on MixFormer
Through the MixFormer-based drone target tracking method, the hybrid attention module and the online sample confidence prediction module are used to solve the precise tracking problem of small drones in complex environments, and achieve fast and stable target positioning and tracking, which is suitable for a variety of drone types and imaging conditions.
Patent Information
- Application Number
- CN202510468075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
Existing drone target tracking methods are difficult to achieve accurate and fast target positioning and tracking in complex environments, especially the challenges caused by small size, variable flight speed and complex environments.
The MixFormer-based target tracking method is adopted, including a hybrid attention module, a prediction head and an online sample confidence prediction module. Features are extracted through self-attention and cross-attention operations, pixel-level positioning is combined with a full convolutional network, and reliable templates are updated through the online sample confidence prediction module to enhance robustness.
Stabilize tracking of targets in complex flight environments, adapt to different types of drones, realize fast and accurate target box output and confidence calculation, avoid long-term tracking drift, and is suitable for a variety of imaging conditions and target characteristics.
Smart Images

Figure CN120339334A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft tracking and monitoring, and particularly relates to a method for tracking unmanned aerial vehicle (UAV) targets based on MixFormer. Background Art
[0002] Due to the characteristics of small size, variable flight speed, and complex flight environment of small UAV targets, the actual UAV monitoring and countermeasure systems have high requirements for the accuracy and speed of target positioning and tracking methods. As a key part of the UAV monitoring and countermeasure systems, by studying the detection and tracking methods of small UAV moving targets, the flight area and flight motivation of UAVs can be monitored in real time, providing accurate target positioning for the subsequent interference and interception part. This technology has significant economic and social benefits in many aspects.
[0003] Visual object tracking has been a fundamental task in the field of computer vision for decades, aiming to estimate the state of an arbitrary target with a given initial state in a video sequence. It has been successfully deployed in various applications such as visual surveillance. In real-world scenarios, designing a simple and effective end-to-end tracker remains challenging. The main challenges come from aspects such as scale changes, object deformations, occlusions, and interference from complex backgrounds.
[0004] The popular single-object tracking paradigm in the current visual tracking field is usually a multi-stage pipeline, mainly including three components: (1) a CNN backbone network to extract features of the target and the search region; (2) an independent fusion module to achieve information exchange between the features of the target and the search region, realizing feature fusion and cross-correlation operations; (3) a tracking prediction head to accurately locate the target and estimate the target bounding box. Among them, the feature fusion module is usually the key to the design of tracking algorithms. In traditional methods, operations based on cross-correlation (such as SiamFC, SiamRPN, SiamFC++) and online update methods (such as KCF, ATOM, DiMP, FCOT, etc.) are mainly used. Recently, inspired by the Transformer in the field of computer vision, information fusion methods based on the attention mechanism have been introduced and achieved very effective results (such as TransT, STARK, TrDiMP, TREG, etc.). However, these Transformer-based trackers still rely on CNN for general feature extraction and only perform attention operations in the high-level abstract representation space of the latter. CNN is usually pre-trained for general object recognition and may ignore finer structural information for tracking. In addition, the feature representation of CNN uses local convolutional kernels and lacks global modeling ability. Therefore, the CNN feature representation is still the bottleneck of the above trackers, hindering the full release of the self-attention ability of the entire tracking pipeline. Summary of the Invention
[0005] The object of the present invention is to overcome one or more deficiencies of the prior art and provide a UAV target tracking method based on MixFormer.
[0006] The object of the present invention is achieved by the following technical solutions:
[0007] A UAV target tracking method based on MixFormer, comprising the following steps:
[0008] Step 1. Establish a target tracking model; the tracking model includes a mixed attention module, a prediction head, and an online sample confidence prediction module;
[0009] Step 2. Train the model: Clean and label image data sets such as UAV types and flight postures, and then train the model through data augmentation, model hyperparameter adjustment, an optimizer, and a loss function;
[0010] Step 3. Combine the tracking model with a target detection algorithm: Input the initial frame image of the video into the target detection algorithm to obtain the target position of the UAV, and then input the obtained target position into the tracking model, and output the target box and confidence of the UAV through the tracking model.
[0011] Further, the tracking model includes: a mixed attention module, a prediction head, and a confidence prediction module;
[0012] The backbone network of the Mixed Attention Module (MAM) is used to extract and fuse features, and at the same time perform feature extraction and the interaction between the target template and the search area. In MAM, self-attention and cross-attention operations of tokens from the target template and the search area are used; self-attention is responsible for extracting the self-features of the target or the search area, while cross-attention realizes feature interaction to mix the target and search area information;
[0013] The prediction head outputs a bounding box for target pixel-level positioning;
[0014] The Online Sample Confidence Prediction Module (SPM) selects reliable online templates according to the predicted confidence scores to enhance the robustness of long-term target tracking.
[0015] Further, in the tracking model, in the mixed attention module, the target template and the search area are input, and the long-distance features of the target template and the search area are extracted at the same time, and the distance features are interacted.
[0016] Furthermore, in the hybrid attention module, dual attention operations are performed on two separate token sequences of the target template and the search region, and by concatenating the token sequences, cross-attention operations are performed on the tokens in the two sequences for communication between the target template and the search region.
[0017] Furthermore, the prediction head passes through a fully convolutional network (FCN), which consists of L stacked Conv-BN-Relu layers, and outputs two probability maps corresponding to the bounding boxes. The predicted box coordinates are obtained by calculating the expectation of the corner probability distribution based on the diagonals of the bounding boxes. The calculation formula is:
[0018] ,
[0019] ,
[0020] where and are the predicted box coordinates, and are the probability maps, H is the height of the predicted box boundary, and W is the width of the predicted box boundary.
[0021] Furthermore, in the online sample confidence prediction module, it includes two attentions and a three-layer MLP. The online sample confidence prediction module selects the online template according to the predicted confidence score, inputs a learnable score token, calculates the attention with the Search ROI token to encode the target information mined in the search graph, then calculates the attention between the score token and the template token of the first frame, compares the mined target with the initial target, and finally predicts the confidence score through an MLP. If it is less than 0.5, it is judged as unreliable.
[0022] Furthermore, in step 2, during the training of the tracking model, the image data is cleaned and labeled, the training and validation data sets are divided, and data augmentation, hyperparameters, optimizers, and loss functions are selected for training. The specific process is as follows: The backbone and head are trained through 500 epochs, and finally the online sample confidence prediction module is trained separately for 40 epochs, freezing the parameters of other parts.
[0023] Furthermore, in step 2, the loss function of the target box is through and joint loss, and the training loss function of the online sample confidence prediction module adopts cross-entropy loss:
[0024] ;
[0025] ;
[0026] Among them, , are the weights of the two losses, is the ground-truth bounding box, is the predicted bounding box of the target, is the true label, is the predicted confidence score.
[0027] Furthermore, in step 3, according to the output predicted confidence score, calculate the confidence for every 100 frames, and take the frame with the highest score to update the template of the UAV target and input it into the tracking model for updated prediction.
[0028] The beneficial effects of the present invention are:
[0029] (1) The UAV target tracking method based on MixFormer can stably track when the target appearance and background change violently through the cooperation of the hybrid attention module, prediction head and online sample confidence prediction module, quickly output the target box and confidence, and adapt to complex flight environments;
[0030] (2) This method has good tracking effects on different types of UAVs such as visible light helicopters, fixed-wing UAVs and infrared small targets, and can accurately locate under various imaging conditions and target characteristics, not limited to specific scenarios;
[0031] (3) The prediction head can achieve pixel-level positioning, and the online sample confidence prediction module filters reliable templates for update according to the scores, avoiding long-term tracking drift and ensuring accurate positioning and long-term stability of tracking. Description of the Drawings
[0033] Figure 1 is the network structure diagram of the Mixfomer tracking model of this embodiment;
[0034] Figure 2 is the hybrid attention module diagram of this embodiment;
[0035] Figure 3 is the online template update module diagram of this embodiment;
[0036] Figure 4 is the method step flow chart of this embodiment;
[0037] Figure 5 is the visible light helicopter tracking effect diagram of this embodiment;
[0038] Figure 6Visible light fixed-wing UAV tracking effect diagram of this embodiment;
[0039] Figure 7 Infrared small target tracking effect diagram of this embodiment. Specific implementation manner
[0041] Next, the technical solution of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0042] A UAV target tracking method based on MixFormer is provided, which is characterized by including the following steps:
[0043] Step 1. Establish a target tracking model; the tracking model includes a hybrid attention module, a prediction head, and an online sample confidence prediction module;
[0044] Step 2. Train the model: Clean and annotate the UAV image dataset, and then train the model through data augmentation, model hyperparameter adjustment, an optimizer, and a loss function;
[0045] Step 3. Combine the tracking model with a target detection algorithm: Input the initial frame image of the video into the target detection algorithm to obtain the target position of the UAV, and then input the obtained target position into the tracking model, and output the target box and confidence of the UAV through the tracking model.
[0046] As Figure 1 shown, the target tracking model mainly consists of three parts: The backbone network of the Mixed Attention Module (MAM) is used to extract and fuse features, and at the same time perform feature extraction and the interaction between the target template and the search area. In MAM, self-attention and cross-attention operations of tokens from the target template and the search area are used. Self-attention is responsible for extracting the self-features of the target or the search area, while cross-attention realizes feature interaction to mix the target and search area information. The prediction head performs target pixel-level positioning. The Online Sample Confidence Prediction Module (SPM) selects reliable online templates according to the prediction scores to enhance the robustness of long-term target tracking.
[0047] As Figure 2As shown in the figure, the hybrid attention module (MAM) is designed as a concise and compact end-to-end tracker. The input of MAM is the target template and the search region. Its purpose is to simultaneously extract the long-range features of the target template and the search region respectively, and fuse the mutual information between the target template and the search region. Unlike the original Multi Head Attention, MAM can handle multiple targets at the same time and improve the tracking accuracy through multi-scale feature fusion. MAM performs dual attention operations on two separate token sequences of the target template and the search region. MAM performs self-attention operations on the tokens in each sequence to capture target or search specific information. At the same time, the hybrid attention mechanism can cross-attention operations on the tokens in the two sequences by splicing the token sequences to allow communication between the target template and the search region. Each feature map of the target template and the search region is flattened and linearly projected to obtain the query, key, and value of the attention operation. , and Indicates the goal, , and represents the search area. The mixed attention is defined as:
[0048] ;
[0049] ;
[0050] ;
[0051] in represents the dimension of the key, and They are the attention maps of the target template and the search area, respectively. MAM includes both self-attention and cross-attention, combining feature extraction and information integration. Finally, the target token and the search token are connected and linearly projected.
[0052] The bounding box regression network uses a corner-based positioning prediction head, which uses a simple fully convolutional network (FCN). FCN consists of L layers of stacked Conv-BN-Relu layers, outputting two probability maps and They correspond to the upper left corner and lower right corner of the target bounding box respectively. Finally, the predicted box coordinates are obtained based on the expected probability distribution of the corner points. and , as shown below:
[0053] ;
[0054] ;
[0055] The online update template makes good use of temporal information to handle target deformation and appearance changes. However, low-quality update templates may lead to tracking drift. SPM selects reliable online templates based on the predicted confidence scores. As Figure 3 shown, SPM consists of two attentions and a three-layer MLP. This module is connected after the last stage of the backbone and is parallel to the prediction head. First, a learnable score token is input, which calculates attention with the Search ROI token to encode the target information mined from the search graph. Then, the score token calculates attention with the template token of the first frame, implicitly comparing the mined target with the initial target. Finally, a confidence score is predicted through an MLP. If the score is less than 0.5, it is judged as unreliable.
[0056] The method process is as Figure 4 shown, divided into a training stage and an inference stage. In the model training stage, first, a UAV image dataset with various complex backgrounds, various UAV types, various flight postures, etc. is collected. The image dataset is cleaned and labeled, and the training and validation datasets are divided. Appropriate data augmentation techniques, hyperparameters, optimizers, loss functions, etc. are selected to train and validate to obtain the tracking model. The training process is divided into two steps. First, the backbone and head are trained for 500 epochs; finally, SPM is trained alone for 40 epochs, freezing the parameters of other parts. The target box loss function adopts and joint loss. The SPM training loss function adopts cross-entropy loss:
[0057] ;
[0058] ;
[0059] where , are the weights of the two losses, is the ground-truth bounding box, is the predicted bounding box of the target, is the true label, is the predicted confidence score.
[0060] During the inference stage, the tracking algorithm is combined with the object detection algorithm. The initial frame image of the video is input into the object detection algorithm to obtain the position of the UAV target in the initial frame image of the video, and a target box is given, that is, the pixel coordinates of the upper left corner, the width and height of the target box. The UAV target within this initial box is used as the input of the tracking model template, and the feature representation of the template is extracted and fused. The network model tracks the UAV target in subsequent infrared images and outputs the target box and confidence of the UAV. According to the predicted confidence score, the template is updated every 100 frames during the inference stage, and the template with the highest score in the interval is selected to replace the previous template. The Mixformer tracking framework allows any number of templates to be input, and by default, it only contains two templates, one initial template and one online updated template.
[0061] Refer to Figures 5 - 6 , and the tracking effects of helicopters under visible light and fixed-wing UAVs under visible light can be clearly obtained from the figure.
[0062] Such as Figure 7 shown, the target tracking effect under weak infrared.
[0063] Through this method, it can be used to track UAVs within a range and output the target box in real time.
[0064] The method provided in this embodiment has good tracking robustness and real-time performance. In the face of drastic changes in the target appearance and background, such as in a complex flight environment, the UAV target may experience rapid attitude changes, partial occlusion, and background interference. Thanks to the collaborative work of the hybrid attention module, prediction head, and online sample confidence prediction module, the method can still stably track the target and accurately output the target box and confidence.
[0065] It has good tracking effects on different types of UAVs, including visible light helicopters, visible light fixed-wing UAVs, and infrared small and weak targets. As can be intuitively seen from Figures 5 to 7 in the document, in various scenarios, this method can effectively locate the target, adapt to different imaging conditions and target characteristics, indicating that it can maintain stable tracking performance in multiple situations and is not limited to specific UAV types or scenarios.
[0066] The online sample confidence prediction module selects reliable online templates according to the predicted confidence score, enhancing the robustness of long-term target tracking. During the long-term tracking process, the appearance of the target will change. By screening reliable templates for updating, the problem of tracking drift caused by low-quality templates can be avoided, ensuring the accuracy and stability of tracking.
[0067] The prediction head calculates the coordinates of the prediction box by outputting a probability map through a fully convolutional network, enabling target pixel-level positioning. This positioning method improves the accuracy of target positioning. In practical applications, for scenarios that require accurate acquisition of the UAV's position information, such as UAV monitoring and countermeasure systems, it can provide accurate data support.
[0068] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A drone target tracking method based on MixFormer, characterized in that It includes the following steps: Step 1. Establish a target tracking model; the tracking model includes a hybrid attention module, a prediction head, and an online sample confidence prediction module; Step 2. Train the model: Clean and annotate the UAV image dataset, and then train the model through data augmentation, model hyperparameter adjustment, optimizer, and loss function; Step 3. Combine the tracking model with the target detection algorithm: Input the initial frame image of the video into the target detection algorithm to obtain the target position of the UAV, and then input the obtained target position into the tracking model, and output the target box and confidence of the UAV through the tracking model.
2. The method for tracking an unmanned aerial vehicle target based on MixFormer according to claim 1, wherein The tracking model includes: a hybrid attention module, a prediction head, and a confidence prediction module; The backbone network of the hybrid attention module is used to extract and fuse features, and at the same time perform feature extraction and the interaction between the target template and the search area. In the hybrid attention module, self-attention and cross-attention operations of tokens from the target template and the search area are used; self-attention is responsible for extracting the self-features of the target or search area, while cross-attention realizes feature interaction to mix target and search area information; The prediction head outputs a bounding box for target pixel-level positioning; The online sample confidence prediction module selects a reliable online template according to the predicted confidence score.
3. The method for tracking an unmanned aerial vehicle target based on MixFormer according to claim 2, wherein, In the tracking model, in the hybrid attention module, the target template and the search area are input, and the distance features of the target template and the search area are extracted at the same time, and the distance features are interacted.
4. The method for tracking an unmanned aerial vehicle target based on MixFormer according to claim 3, wherein, In the hybrid attention module, dual attention operations are performed on two separate token sequences of the target template and the search area, and through concatenating the token sequences, cross-attention operations are performed on the tokens in the two sequences for communication between the target template and the search area.
5. A method for tracking an unmanned aerial vehicle target based on MixFormer according to claim 1, characterized in that, The prediction head outputs two probability maps corresponding to the bounding box through a fully convolutional network, and calculates the predicted box coordinates according to the expectation of the corner probability distribution of the diagonal of the bounding box. The calculation formula is: ; ; Among them, and are the coordinates of the prediction box, and is the probability map, H is the height of the prediction box boundary, and W is the width of the prediction box boundary.
6. A method for tracking an unmanned aerial vehicle target based on MixFormer according to claim 1, characterized in that, In the online sample confidence prediction module, it includes two attentions and a three-layer MLP. Input a learnable score token, calculate the attention with the Search ROI token, encode the target information mined in the search graph, and then calculate the attention between the score token and the template token of the first frame, compare the mined target with the initial target, and finally predict the confidence score through an MLP. If it is less than 0.5, it is judged as unreliable.
7. A method for drone target tracking based on MixFormer according to claim 1, characterized in that In Step 2, during the training of the tracking model, the image data is cleaned and annotated, the training and validation datasets are divided, and data augmentation, hyperparameters, optimizers, and loss functions are selected for training. The specific process is: train the backbone and head through 500 epochs, and finally train the online sample confidence prediction module alone for 40 epochs.
8. A method for drone target tracking based on MixFormer according to claim 1, characterized in that In the said step 2, the loss function is obtained through and the combined loss, and the loss function for training the online sample confidence prediction module adopts the cross-entropy loss: ; ; Among them, , are the weights of the two losses, is the ground-truth bounding box, is the predicted bounding box of the target, is the true label, is the predicted confidence score.
9. A method for drone target tracking based on MixFormer according to claim 1, characterized in that, In step 3, according to the output prediction confidence score, calculate the confidence for every 100 frames, and update the template of the UAV target of the frame with the highest score and input it into the tracking model for updated prediction.