Unmanned aerial vehicle target tracking method, system and device and storage medium

By using the improved ShuffleNetV2 network model and self-interference fusion technology in drone tracking, the problem of drone tracking decreases in accuracy and stability under fast motion and occlusion phenomena is solved, achieving high-precision, robustness and real-time target tracking effects.

CN120047859APending Publication Date: 2025-05-27CHONGQING COLLEGE OF ELECTRONICS ENG +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510123632.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Drone tracking faces the problem of difficulty in tracking targets under fast motion and occlusion, resulting in a decrease in tracking accuracy and stability, making it difficult to meet real-time processing needs.

Method used

The ShuffleNetV2 network model is adopted, by dividing the network layer into multiple sub-feature groups, the features of the target image and the search area image are extracted, and the horizontal and vertical self-attention fusion is carried out to generate the target semantic features and regional semantic features, and finally predict the position and scale of the target through classification regression.

Benefits of technology

It improves the accuracy and robustness of drone target tracking, reduces background interference, improves spatial scale details, meets real-time processing needs, and can achieve stable and efficient target tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047859A_ABST
    Figure CN120047859A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target tracking method, system and device and a storage medium, and the method comprises the steps: extracting a target image and a search region image from a first frame image of a video sequence, carrying out the feature extraction of the target image and the search region image, and obtaining a target template feature set and a search region feature set; performing transverse and longitudinal self-mutual attention fusion on features in the target template feature set and the search region feature set to obtain target semantic features and region semantic features; and performing classification regression according to the target semantic features and the regional semantic features to obtain the predicted position and scale of the target. According to the method, the problem that in the prior art, when a traditional visual tracking task tracks a target, the target is difficult to track due to rapid movement and a shielding phenomenon is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, system, device and storage medium for tracking an unmanned aerial vehicle target. Background Art

[0002] Visual tracking is a basic and in-demand technology in the field of computer vision and pattern recognition, and has wide application value in many fields. In recent years, with the rapid development of drone technology and its wide application in remote sensing, smart agriculture, intelligent navigation, public safety, disaster relief and public transportation, drone tracking has become an important challenge in the field of computer vision.

[0003] The core task of drone tracking is to infer and predict the position and scale of the target in the subsequent aerial frames in real time based on the initial state of the target in the first frame of the video sequence. Although drone tracking has shown great potential in practical applications, its unique technical challenges often exceed the problems faced by traditional visual tracking tasks. First, the high-speed flight of drones causes a variety of complex problems, such as motion blur caused by rapid motion, significant scale changes of targets, and drastic changes in perspective during flight. In addition, when drones operate in extreme environments, the frequent occurrence of occlusion further increases the difficulty of continuous target tracking. The superposition of these factors significantly affects the accuracy and stability of the tracking method. At the same time, the operation of drones is limited by battery capacity and computing resources, which puts higher requirements on tracking methods: the method must not only have high accuracy, but also meet the needs of real-time processing to ensure stable and efficient target tracking at a rate of at least 30 frames per second. Therefore, designing drone tracking methods that are both efficient and accurate in complex environments has become a key issue that needs to be solved in this field. Although drone tracking faces many technical difficulties, its huge potential and significant advantages in multiple application fields fully highlight the importance and necessity of continued research and technological innovation in this direction. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides a UAV target tracking method, system, device and storage medium, which solves the problem of difficulty in target tracking due to rapid motion and occlusion in traditional visual tracking tasks in the prior art.

[0005] According to an embodiment of the present invention, a method for tracking a drone target includes:

[0006] Extracting the target image and the search area image from the first frame image of the video sequence, and performing feature extraction on the target image and the search area image respectively to obtain the target template feature set and the search area feature set;

[0007] Perform horizontal and vertical self - mutual attention fusion on the features in the target template feature set and the search region feature set respectively to obtain the target semantic feature and the regional semantic feature;

[0008] Perform classification regression based on the target semantic feature and the regional semantic feature to obtain the predicted position and scale of the target.

[0009] Preferably, the method for feature extraction of the target image and the search region image includes:

[0010] Divide the network layers of the ShufflenetV2 network model into multiple sub - feature groups of different scales to obtain an improved extraction network;

[0011] Use two identical improved extraction networks to perform feature extraction on the target image and the search region image respectively, and extract the target template features or search region features extracted by each sub - feature group, and combine them into the target template feature set and the search region feature set.

[0012] Preferably, the sub - feature group includes four module layers and an attention layer arranged in sequence. After feature extraction, the target template features or search region features of different scales extracted by the second module layer, the third module layer and the attention layer are added to the corresponding target template feature set or search region feature set.

[0013] Preferably, the horizontal and vertical self - mutual attention fusion includes horizontal attention enhancement and vertical attention fusion;

[0014] Combine the target template features and search region features extracted from the same sub - feature group into feature pairs, then perform horizontal attention enhancement on each feature pair, and then perform vertical attention fusion on the target template features or search region features in the target template feature set or search region feature set that have undergone horizontal attention enhancement.

[0015] Preferably, the method for horizontal attention enhancement includes:

[0016] Extract the query vector, key vector, value vector and scale of the target template feature or search region feature in the feature pair to construct the corresponding attention weight matrix;

[0017] Use a pooling module to reduce the dimension of the attention weight matrix, and then combine all the attention weight matrices to obtain a multi - head attention matrix;

[0018] Import the multi - head attention matrix into the PAB module to enhance the target template feature or search region feature.

[0019] Preferably, the method for vertical attention fusion includes:

[0020] A1: Use the query vector corresponding to the target template feature or the search area feature corresponding to the third module layer as the attention value;

[0021] A2: Import all the features and attention values that have undergone horizontal attention enhancement in the target template feature set or the search area feature set into the PAB module to obtain the scale feature corresponding to each feature, and then directly sum all the scale features to obtain the fused feature;

[0022] A3: Repeat step A2 multiple times, and input the fused feature obtained by the last summation into the PAB module again to obtain the target semantic feature or the regional semantic feature.

[0023] Preferably, the method for performing classification regression based on the target semantic feature and the regional semantic feature to obtain the predicted position and scale of the target includes:

[0024] B1: Calculate the depth cross-correlation between the target semantic feature and the regional semantic feature to obtain a multi-channel correlation map;

[0025] B2: Construct two classification heads and prediction heads with the same structure, and input the multi-channel correlation map into the classification head and the prediction head respectively to obtain the classification result and the regression result;

[0026] B3: Repeat step B3 until the loss function converges, and then obtain the scale of the target according to the classification result and the predicted position of the target according to the regression result.

[0027] On the other hand, according to an embodiment of the present invention, there is also provided a UAV target tracking system, which uses the above-mentioned UAV target tracking method, including:

[0028] An acquisition module, which is used to extract an image and a search area image;

[0029] A processing module, which is used to extract features from the extracted image and the search area image, and perform horizontal and vertical self-mutual attention fusion on the extracted features to obtain the target semantic feature and the regional semantic feature;

[0030] A prediction module, which is used to perform classification regression based on the target semantic feature and the regional semantic feature to obtain the predicted position and scale of the target.

[0031] On the other hand, according to an embodiment of the present invention, there is also provided a computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the above-mentioned UAV target tracking method.

[0032] On the other hand, according to an embodiment of the present invention, there is also provided a computer storage medium storing a computer program, which when executed by a processor causes the processor to execute the above-mentioned unmanned aerial vehicle target tracking method.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention uses the ShuffleNetV2 network model as the backbone, divides multiple sub-feature groups in the network layer of the network model, extracts features from the target image and the search area image at different scales through different features, and then integrates the self-cross attention interaction in the horizontal and vertical dimensions of the extracted features to enhance the semantic information of the high-level features, obtaining the target semantic features and the regional semantic features. The present invention cleverly combines the multi-scale feature maps of the target template features and the search area features, thereby optimizing the features, reducing background interference, improving the spatial scale details, and finally predicting the target position and scale according to the target semantic features and the regional semantic features after semantic enhancement, greatly improving the accuracy and robustness of the tracker. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is the architecture diagram of the target tracking method according to the embodiment of the present invention.

[0036] Figure 2 It is the architecture diagram of the prediction head according to the embodiment of the present invention.

[0037] Figure 3 It is the comparison diagram of the tracking effects according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions in the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0039] As Figure 1 shown, the embodiment of the present invention proposes an unmanned aerial vehicle target tracking method, including:

[0040] Extract the target image and the search area image from the first frame image of the video sequence, and perform feature extraction on the target image and the search area image respectively to obtain the target template feature set and the search area feature set;

[0041] In the present invention, a lightweight network model ShuffleNetV2 is introduced and improved simultaneously. Based on the architecture of ShufflenetV2, the network layers of the ShufflenetV2 network model are reconstructed into batch dimensions, and the network layers are divided into multiple sub-feature groups that can extract features at different scales according to the channel dimension, resulting in an improved extraction network. This enables the spatial semantic features to be evenly distributed within each feature group, aiming to retain the information of each channel and reduce the computational overhead. In addition, this network layer not only recalibrates the channel weights of each parallel branch by encoding global information but also further aggregates the output features of the two parallel branches through cross-dimensional interaction, thereby effectively capturing pixel-level pairwise relationships.

[0042] As Figure 1 shown, two ShuffleNetV2 network models with the same structure and weight sharing are used to extract features from the target image and the search region image respectively. The sub-feature groups include four module layers (Module 1, Module 2, Module 3, and Module 4) arranged in sequence from left to right and an attention layer. The target template features or search region features extracted by Module 2, Module 3, and the attention layer are combined into a target template feature set and a search region feature set. At this time, the target template feature set includes target template features at different scales. The search region feature set includes search region features at different scales.

[0043] The features in the target template feature set and the search region feature set are respectively subjected to horizontal and vertical self-mutual attention fusion to obtain target semantic features and regional semantic features;

[0044] The horizontal and vertical self-mutual attention fusion includes horizontal attention enhancement and vertical attention fusion;

[0045] The target template features and search region features extracted from the same sub-feature group are combined into feature pairs, and then each feature pair is subjected to horizontal attention enhancement. After that, the target template features or search region features that have undergone horizontal attention enhancement in the target template feature set or the search region feature set are subjected to vertical attention fusion.

[0046] For attention calculation, different from the traditional attention mechanism, the present invention processes feature maps of different scales from the search region and the template image patch within the same frame and runs with the smallest time span. Therefore, we omit the position encoding commonly used in the traditional attention mechanism. The method for calculating the attention weight matrix is as follows:

[0047]

[0048] Among them, Q, K, and V represent the query vector, key vector, and value vector of the target template feature or search area feature respectively, and c is the scale used to normalize the attention. Expanding the attention mechanism into a multi-head method can further enrich the semantic information of the features.

[0049] The calculation method of multi-head attention is as follows:

[0050] H i = Attention(QW i Q , KW i K , VW i V )

[0051] MHA(Q, K, V) = Concat(H 1 , H 2 ,..., H N )W o

[0052] Here, N represents the number of attention heads, which is also the total number of target template features and search area features, where i ∈ {1, 2,..., N}, and are the weight matrices of the query vector, key vector, and value vector respectively, and W O represents the linear projection parameter, and Concat represents the concatenation operation.

[0053] In the horizontal direction, the target template features and search area features extracted from the same sub-feature group are combined into feature pairs. To reduce the computational cost, K and V are processed by a pooling module to effectively reduce their spatial dimensions. Then, the results of multi-head attention are merged, and the final attention output is generated through a multi-layer perceptron (MLP). The calculation process of the PAB module is as follows:

[0054] F = Norm(Q + PA(R)(Q, K, V))

[0055] PAB(Q, K, V, R) = Norm(F + MLP(F))

[0056] Among them, R represents the pooling kernel size and stride, MLP represents a fully connected feed-forward neural network, and Norm represents LayerNorm used to smooth the input features. On this basis, we first perform dimensionality reduction and flattening on the feature maps of three different scales from the search image patch and the template image patch respectively. Then, we horizontally connect these flattened feature maps at the same ratio and use the PAB module to calculate the attention feature map, thereby obtaining the enhanced target template feature or search area feature. The calculation formula is as follows:

[0057]

[0058] Among them, R represents the pooling kernel size and stride, represents the feature map of the target image, represents the feature map of the search region image, and i ∈ (1, 2, 3) represents three different scales. After obtaining the three different-scale enhanced feature maps, we input these maps into a central path.

[0059] In the vertical direction, the feature maps of three different scales in the target template feature set or search region feature set are input into the PAB module. The attention values of the template feature map and the search region feature map are respectively calculated using and as the query vectors for all feature layers. Then, the obtained outputs are directly summed to obtain the fused feature. After performing two attention calculations, the fused feature is input into the PAB module again to generate the final feature map. The entire calculation process is as follows:

[0060]

[0061] F tem = {PAB(F′ tem , F′ tem , F′ tem , R = 2)} n=2

[0062] F se = {PAB(F′ se , F′ se , F′ se , R = 2)} n=2

[0063] Here, n represents the number of repetitions of the module, and R represents the pooling kernel size and stride. After performing two attention calculations, we will obtain the target semantic feature F tem and the regional semantic feature F se . Through the fusion of multi-scale feature maps, the present invention cleverly combines the multi-scale feature maps of the target template feature and the search region feature, thereby optimizing the features, reducing background interference, and improving the spatial scale details. We can obtain richer semantic information.

[0064] According to the target semantic feature and the regional semantic feature, classification regression is performed to obtain the predicted position and scale of the target.

[0065] Such as Figure 2As shown, the present invention adopts a prediction head based on a fully convolutional network to extract the classification and bounding box information of the target from the input target semantic features or regional semantic features. Specifically, the prediction head consists of two main branches: a classification head and a regression head. Both branches are composed of network layers constructed by multiple convolutional-batch normalization-ReLU (Conv-BN-ReLU) blocks, which are respectively responsible for predicting the target category and the bounding box coordinates.

[0066] First, the depth cross-correlation between the feature maps of the search image and the template image is calculated to generate a multi-channel correlation map. This correlation map is then input into the two branches to produce the classification results at each spatial position and the regression results Here, represents the predicted category, while corresponds to the corner coordinates of the predicted bounding box. For the tracking task, we adopt the cross-entropy loss function for classification and combine the L1 loss and the generalized intersection over union (GIoU) loss for bounding box regression. The calculation formula of the overall loss function is as follows:

[0067]

[0068] where the constant λ cls = 5, λ iou = 2, Here, λ cls represents the cross-entropy loss function for classification, and λ iou represents the generalized intersection over union (GIoU) loss, which is used to measure the consistency between the predicted position and the ground truth bounding box position. corresponds to the L1 loss for the regression task. When the overall loss function converges, the scale of the target can be obtained according to the classification results, and the predicted position of the target can be obtained according to the regression results.

[0069] For performance evaluation, the present invention conducted comparative experiments on the above-mentioned proposed method with nine other state-of-the-art trackers. The experimental data is based on four mainstream UAV datasets. As shown in Table 1, the proposed method of the present invention has achieved the best results in terms of both average precision and success rate. Specifically, the average precision of our method is 0.837, which is 0.017 higher than that of SIFTrack ranked second in terms of precision; its average success rate is 0.688, which is 0.017 higher than that of SIFTrack ranked second in terms of success rate. These results strongly verify the excellent tracking performance of the proposed method on UAV datasets.

[0070] Table 1 Comparison with other trackers

[0071]

[0072]

[0073] In addition, Figure 3 Some tracking results of the method proposed in the present invention and nine other tracking methods on different challenging video sequences are shown. These sequences cover the common challenges in UAV target tracking, including occlusion, in-plane rotation, surrounding similar objects, scale change, motion blur, aspect ratio change, and view change. The results show that most of the comparison methods are severely affected by these challenges, resulting in target loss or drift in many cases. In contrast, the method proposed in the present invention can effectively cope with these challenges. It is worth noting that in the "Animal4_1" sequence, the present method is the only tracker that can successfully track the target, achieving accurate and robust fitting results. In the "car4_1" sequence, due to severe occlusion, only the present method and SIFTrack can stably locate the target. A similar trend is also observed in the "car6_2_1" sequence, where most tracking methods either cannot track the target or exhibit large fitting errors, while the present method and SIFTrack maintain successful tracking. In the "Truck1_1" sequence, occlusion and scale change cause most trackers to fail or drift, but the present method can always accurately track the target. Compared with other methods, these results highlight the robustness of the present method in dealing with complex tracking scenarios.

[0074] On the other hand, an embodiment of the present invention also provides a UAV target tracking system that uses the above-mentioned UAV target tracking method, including:

[0075] An acquisition module, which is used to extract images and search area images;

[0076] A processing module, which is used to extract features from the extracted images and search area images, and perform horizontal and vertical self-mutual attention fusion on the extracted features to obtain target semantic features and regional semantic features;

[0077] A prediction module, which is used to perform classification regression based on the target semantic features and regional semantic features to obtain the predicted position and scale of the target.

[0078] On the other hand, an embodiment of the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the above-mentioned UAV target tracking method.

[0079] On the other hand, an embodiment of the present invention also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the above-mentioned UAV target tracking method.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for tracking a target of an unmanned aerial vehicle, characterized in that: include: Extracting the target image and the search area image from the first frame image of the video sequence, and performing feature extraction on the target image and the search area image respectively to obtain the target template feature set and the search area feature set; The features in the target template feature set and the search area feature set are fused horizontally and vertically to obtain the target semantic features and the area semantic features. Classification regression is performed based on the target semantic features and regional semantic features to obtain the predicted position and scale of the target.

2. A method for tracking a target by using an unmanned aerial vehicle according to claim 1, characterized in that: Methods for extracting features from target images and search area images include: The network layer of the ShufflenetV2 network model is divided into multiple sub-feature groups of different scales to obtain an improved extraction network; Two identical improved extraction networks are used to perform feature extraction on the target image and the search area image respectively, and the target template features or search area features extracted by each sub-feature group are extracted and combined into a target template feature set and a search area feature set.

3. A method for tracking a target by using an unmanned aerial vehicle according to claim 2, characterized in that: The sub-feature group includes four module layers and an attention layer arranged in sequence. After feature extraction, the target template features or search area features of different scales extracted by the second module layer, the third module layer and the attention layer are added to the corresponding target template feature set or search area feature set.

4. A method for tracking a target by using an unmanned aerial vehicle according to claim 2, characterized in that: Horizontal and vertical mutual attention fusion includes horizontal attention enhancement and vertical attention fusion; The target template features and search area features extracted from the same sub-feature group are combined into feature pairs, and then each feature pair is horizontally enhanced. After that, the target template features or search area features that have been horizontally enhanced in the target template feature set or the search area feature set are vertically fused with attention.

5. A method for tracking a target by using an unmanned aerial vehicle according to claim 4, characterized in that: Methods for increasing lateral attention include: Extract the query vector, key vector, value vector and scale of the target template feature or search area feature in the feature pair to construct the corresponding attention weight matrix; Use the pooling module to reduce the dimension of the attention weight matrix, and then merge all the attention weight matrices to obtain a multi-head attention matrix; The multi-head attention matrix is ​​imported into the PAB module to enhance the target template features or search area features.

6. A method for tracking a target by using an unmanned aerial vehicle according to claim 4, characterized in that: Methods for vertical attention fusion include: A1: The target template features corresponding to the third module layer or the query vector corresponding to the search area features are used as the attention value; A2: Import all the features and attention values ​​that have been enhanced with horizontal attention in the target template feature set or the search area feature set into the PAB module to obtain the scale feature corresponding to each feature, and then directly sum all the scale features to obtain the fusion feature; A3: Repeat step A2 multiple times, and input the fused features obtained by the last summation into the PAB module again to obtain the target semantic features or regional semantic features.

7. The method for tracking a target by using an unmanned aerial vehicle according to claim 1, wherein: Methods for performing classification regression based on target semantic features and regional semantic features to obtain the predicted position and scale of the target include: B1: Calculate the deep cross-correlation between the target semantic features and the regional semantic features to obtain a multi-channel correlation map; B2: Construct two classification heads and prediction heads with the same structure, and input the multi-channel correlation graphs into the classification head and prediction head respectively to obtain the classification results and regression results; B3: Repeat step B3 until the loss function converges, then obtain the scale of the target based on the classification result, and obtain the predicted position of the target based on the regression result.

8. An unmanned aerial vehicle target tracking system, characterized in that: The system uses a drone target tracking method as described in any one of claims 1 to 7, comprising: An acquisition module, the acquisition module is used to extract images and search area images; A processing module, the processing module is used to extract features from the extraction image and the search area image, and perform horizontal and vertical mutual attention fusion on the extracted features to obtain target semantic features and regional semantic features; The prediction module is used to perform classification regression according to the target semantic features and the regional semantic features to obtain the predicted position and scale of the target.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes a method for tracking a target of an unmanned aerial vehicle as claimed in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes a drone target tracking method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic target tracking method based on LBP features and semantic features

    CN120411176A