A method and device for unmanned aerial vehicle target tracking

By introducing spatial position perception module and multi-frequency domain target positioning method in the UAV target tracking technology, the problem of position information retention and adaptability of scale changes in the drone target in complex environments is solved, and the accuracy and success rate of tracking are improved.

CN119579649BActive Publication Date: 2025-05-30CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411639712.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-05-30
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing drone target tracking technology is difficult to retain precise position information in complex flight environments, and is not robust enough for scale changes, resulting in a decrease in tracking accuracy.

Method used

The spatial position perception module is used to capture long-distance dependence while retaining precise position information, and the network is adaptable to targets of different scales through the multi-frequency domain target positioning method.

Benefits of technology

This improves the accuracy and success rate of drone target tracking, so that drone targets can be extracted better in complex backgrounds and adapts to changes at different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579649B_ABST
    Figure CN119579649B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for unmanned aerial vehicle (UAV) target tracking. The method includes: collecting UAV flight videos using a color camera; determining whether the current frame is the first frame of the tracking video; if so, manually annotating the UAV target to obtain a template region image; if not, determining a search region image for UAV target tracking according to the tracking result of the previous frame; inputting the obtained template region and the search region image into a UAV target tracking network; displaying the tracking result output by the network on a display module; obtaining the next frame data collected by the camera, and repeating the above steps until the video processing is completed. The present invention aims to capture long-distance dependencies while retaining accurate position information, effectively extract target features from complex backgrounds. At the same time, a multi-frequency domain target positioning method is adopted to make the network adaptable to targets of different scales, ultimately improving the accuracy and success rate of UAV target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target tracking, and particularly relates to a method and device for unmanned aerial vehicle (UAV) target tracking. Background Art

[0002] With the rapid development of deep learning and automation technologies, UAVs have been widely used in many fields such as aerial photography and environmental monitoring. However, with the popularization of UAVs, their potential public safety threats have become increasingly apparent. Therefore, it is crucial to be able to monitor and track the position information of UAVs in real time for maintaining public safety. However, due to the scale changes, sudden movements during the flight of UAVs, and complex flight environments, there are still many challenges in achieving precise UAV tracking.

[0003] In the field of target tracking, due to the lack of sufficient UAV-specific datasets, the research on tracking algorithms for UAVs is relatively limited. Traditional general target tracking algorithms, such as SiamFC (Fully-Convolutional Siamese Networks for Object Tracking), SiamRPN (High Performance Visual Tracking with Siamese Region Proposal Network), SiamMask (High Performance Visual Tracking with Siamese Region Proposal Network), etc., these traditional target tracking algorithms usually use linear correlation operations to match the target with the search area, which limits the non-linear interaction between the template and the search area, resulting in insufficient representation ability of the model. In recent years, with the development of large models, tracking algorithms based on Transformer, such as STARK (Learning Spatio-Temporal Transformer for Visual Tracking) and ODTrack (Online Dense Temporal Token Learning for Visual Tracking), have gradually emerged, but these algorithms still mainly rely on the training of large targets in general datasets and have poor adaptability to small targets or small-scale datasets.

[0004] In summary, the problems existing in the existing technologies available for UAV target tracking are as follows:

[0005] In view of the problem of the complex flight environment of UAV targets, although the existing technology extracts features by establishing long-distance dependencies, it ignores the position information, which is crucial for the result of target tracking. At the same time, current target tracking algorithms are mostly trained on large targets in general datasets, which is less robust for UAV target tracking with continuously changing scales. Summary of the Invention

[0006] The present invention aims to solve the technical problems in the existing technology and provides a UAV target tracking method and device. The present invention adopts a spatial position perception module to capture long-distance dependencies while retaining accurate position information, and effectively extract target features from complex backgrounds. At the same time, a multi-frequency domain target positioning method is adopted to make the network adaptable to targets of different scales, and finally improve the accuracy and success rate of UAV target tracking.

[0007] To solve the above technical problems, the technical solution of the present invention is specifically as follows:

[0008] A UAV target tracking device, comprising: a UAV target tracking network structure;

[0009] The UAV target tracking network structure includes: a multi-frequency domain feature fusion module, and two spatial position perception modules connected to the multi-frequency domain feature fusion module, and each spatial position perception module is also respectively connected to a feature extraction network; the multi-frequency domain feature fusion module is also respectively connected to a classification branch and a regression branch;

[0010] The feature extraction network is used for feature extraction of the template image and the search image;

[0011] The spatial position perception module is used to retain the accurate position information of the features while capturing long-distance dependencies, enhance the object features, and enable UAV targets to be better extracted from complex backgrounds;

[0012] The multi-frequency domain feature fusion module is used to fuse the features extracted by the feature extraction network and the spatial position perception module in sequence using cross-attention;

[0013] The classification branch is used to distinguish the target area and the non-target area during tracking, that is, to determine which position in the search area has a target;

[0014] The regression branch is used to fine-tune the coordinates, size and aspect ratio of the bounding box at the possible position of the target during tracking, so as to generate an accurate target box;

[0015] The tracking result output by the UAV target tracking network structure is displayed on the display module.

[0016] A UAV target tracking method applicable to the above UAV target tracking device, comprising the following steps:

[0017] S1. Use a camera to collect UAV flight videos;

[0018] S2. Determine whether the current frame of the image collected by the camera is the first frame of the tracking video;

[0019] S3. If so, perform manual annotation of the UAV target to obtain a template area image;

[0020] S4. If not, determine the search area image for UAV target tracking according to the tracking result of the previous frame;

[0021] S5. Input the obtained template area and the search area image into the UAV target tracking network;

[0022] S6. Display the tracking result output by the UAV target tracking network on the display module;

[0023] S7. Obtain the next frame data of the image collected by the camera, and repeat steps S2 - S6 until the video processing is completed.

[0024] In the above technical solution, the construction process of the template area image in step S3 is as follows:

[0025] S301. Obtain the marked target coordinate positions (x, y, w, h), where (x, y) represents the upper left corner coordinates of the tracking frame covering the target, and (w, h) represents the width and height of the tracking frame;

[0026] S302. Calculate the template cropping interval where W Z and H Z The values are calculated as follows:

[0027]

[0028] where, W Z represents the width of the template area, and H Z represents the height of the template area;

[0029] S303. Transform the cropped area image of the template into Obtain the template image;

[0030] where, Z represents the template image, represents a tensor, that is, transform the cropped template area into a tensor with a size of 128 * 128 and 3 channels.

[0031] In the above technical solution, the construction process of the search area image in step S4 is as follows:

[0032] S401. Obtain the target coordinate position (x, y, w, h) of the previous frame, where (x, y) represents the upper left corner coordinates of the tracking box covering the target, and (w, h) represents the width and height of the tracking box;

[0033] S402. Calculate the search clipping interval Where W X and H X The values are calculated as follows:

[0034]

[0035] Where, W X represents the width of the search area, and H X represents the height of the search area;

[0036] S403. Transform the cropped search area image into Obtain the search image;

[0037] Where, X represents the search area image.

[0038] In the above technical solution, in step S5, load the training weights in the UAV target tracking network and perform target tracking.

[0039] In the above technical solution, further, in step S5, loading the training weights in the UAV target tracking network and performing target tracking includes the following steps:

[0040] S501. When processing the first frame of image, input the template image obtained in step S303 into the feature extraction network and the spatial position perception module to obtain the feature map of the template area image

[0041] S502. When processing subsequent frames other than the first frame, input the search image obtained in step S403 into the feature extraction network and the spatial position perception module to obtain the feature map of the search area image

[0042] S503. The feature map generated in step S501 and the feature map generated in step S502 Respectively, through two-dimensional convolution, flattening, and channel transformation operations, generate the feature map and the feature map

[0043] S504. Input the feature map F Z1 and the feature map F X1 Generated in step S503 into the multi-frequency domain feature fusion module respectively;

[0044] First, obtain high-frequency features through global multi-head self-attention. Then, downsample the feature map and input it into the multi-head self-attention module to obtain low-frequency features. After that, concatenate the high-frequency features and low-frequency features to generate features and features Then, through the residual connection method, generate features and features

[0045] The whole process is summarized by the formula:

[0046] F Z3 = F Z1 + F Z2 (5)

[0047] F X3 = F X1 + F X2 (6)

[0048] Among them,

[0049] F Z2 = MultiHead(F Z1 ) + MultiHead(AvgPool(F Z1 )) (7)

[0050] F X2 = MultiHead(F X1 ) + MultiHead(AvgPool(F X1 )) (8)

[0051] Among them, MultiHead(*) is the multi-head self-attention operation;

[0052] S505. Fuse the features F Z3 and F X3 obtained in step S504 using cross-attention. In this layer, two multi-head attentions are used. For the first multi-head attention, its matrix Q is composed of F Z3 and the spatial position encoding generated by the sine function, K is composed of the feature F X3 and the spatial position encoding generated by the sine function, and V is the feature F X3 , and a fusion matrix is generated through the attention mechanism For the second multi-head attention, its matrix Q is composed of the feature F X3 and the spatial position encoding generated by the sine function, K is composed of the feature F Z3 and the spatial position encoding generated by the sine function, and V is the feature F Z3 , and a fusion matrix is generated through the attention

[0053] S506. The output result of step S505 is further passed through a multi-head attention mechanism. The matrix Q is the feature F generated in step S505 2 combined with the spatial position encoding generated by the sine function. The matrix K is the feature F 1 combined with the spatial position encoding generated by the sine function. The matrix V is the feature F 1 , and a fused feature is generated through the attention mechanism It is input into the classification branch and the regression branch networks to obtain the target tracking box output by the UAV target tracking network

[0054] The beneficial effects of the present invention are as follows

[0055] The UAV target tracking method of the present invention includes a spatial position perception module, which can capture long-range dependencies while retaining accurate position information, enhance object features, and enable UAV targets to be better extracted from complex backgrounds. Ultimately, the accuracy and success rate of UAV target tracking are improved

[0056] The UAV target tracking method of the present invention is a multi-frequency domain target positioning method, which enables the network to simultaneously focus on high-frequency and low-frequency information and has adaptability to targets of different scales. It can effectively solve the problem of reduced tracking accuracy caused by scale changes during UAV tracking BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The present invention will be further described in detail below with reference to the drawings and specific embodiments

[0058] Figure 1 It is a flowchart of the UAV target tracking method of the present invention

[0059] Figure 2 It is a network structure diagram of the UAV target tracking device of the present invention

[0060] Figure 3 It is a comparison chart of the accuracy and success rate between the UAV target tracking method of the present invention and other tracking algorithms on the self-made dataset UAV-Tracking-390 DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The inventive concept of the present invention is as follows: With the increasing threat of drone intrusion, the perception ability for drones has become crucial in anti-drone missions. In the field of target tracking, although there are many excellent tracking algorithms, due to the scale changes, sudden movements during the flight of drones, and complex flight environments, etc., achieving precise drone target tracking still faces many challenges. To solve this problem, we designed a drone target tracking method and device. This method designed a spatial position perception module, which can capture long-range dependencies while retaining precise position information to enhance object features in complex background environments. In addition, we designed a multi-frequency domain target localization method to make the network adaptable to targets of different scale sizes to cope with the scale changes during the flight of drone targets, and adopted a Siamese network structure to reduce the number of model parameters.

[0062] The following will describe the present invention in detail with reference to the accompanying drawings.

[0063] Combined with Figure 2 Specifically describe the drone target tracking device of the present invention, including: a drone target tracking network structure; the drone target tracking network structure includes: a multi-frequency domain feature fusion module, and two spatial position perception modules connected to the multi-frequency domain feature fusion module, and each spatial position perception module is also respectively connected to a feature extraction network; the multi-frequency domain feature fusion module is also respectively connected to a classification branch and a regression branch;

[0064] The feature extraction network is used for feature extraction of the template image and the search image;

[0065] The spatial position perception module is used to capture long-range dependencies while retaining the precise position information of the features, enhance object features, and enable drone targets to be better extracted from complex backgrounds;

[0066] The multi-frequency domain feature fusion module is used to fuse the features extracted successively through the feature extraction network and the spatial position perception module using cross-attention;

[0067] The classification branch is mainly used for distinguishing the target area and the non-target area during tracking, that is, judging which position in the search area has a target;

[0068] The regression branch is used to finely adjust the coordinates, size, and aspect ratio of the bounding box at the position where the target may exist during tracking, so as to generate an accurate target box;

[0069] The tracking result output by the drone target tracking network structure is displayed on the display module.

[0070] The flowchart of the drone target tracking method applicable to the drone target tracking device of the present invention is shown in Figure 1, including the following steps:

[0071] S1. Use a color camera to collect UAV flight videos;

[0072] S2. Determine whether the current frame of the image collected by the camera is the first frame of the tracking video;

[0073] S3. If it is, perform manual annotation of the UAV target to obtain the template area image;

[0074] S4. If not, determine the search area image for UAV target tracking according to the tracking result of the previous frame;

[0075] S5. Input the obtained template area and the search area image into the UAV target tracking network;

[0076] S6. Display the tracking result output by the UAV target tracking network on the display module;

[0077] S7. Obtain the next frame data of the image collected by the camera, and repeat steps S2 - S6 until the video processing is completed.

[0078] The construction process of the template area image in step S3 is as follows:

[0079] S301. Obtain the marked target coordinate positions (x, y, w, h), where (x, y) represents the upper left corner coordinates of the tracking box covering the target, and (w, h) represents the width and height of the tracking box;

[0080] S302. Calculate the template cropping interval where W Z and H Z The values are calculated as follows:

[0081]

[0082] where, W Z represents the width of the template area, and H Z represents the height of the template area;

[0083] S303. Change the cropped area image of the template to Obtain the template image;

[0084] where, Z represents the template image, represents the tensor. That is, change the cropped template area to a tensor with a size of 128 * 128 and 3 channels.

[0085] The construction process of the search area image in step S4 is as follows:

[0086] S401. Obtain the target coordinate position (x, y, w, h) of the previous frame, where (x, y) represents the upper left corner coordinates of the tracking box covering the target, and (w, h) represents the width and height of the tracking box;

[0087] S402. Calculate the search cropping interval where W X and H X are calculated as follows:

[0088]

[0089] where, W X represents the width of the search area, and H X represents the height of the search area;

[0090] S403. Transform the cropped area image of the template into obtain the search image;

[0091] where, X represents the search area image.

[0092] In step S5, load the training weights in the UAV target tracking network and perform target tracking. The structure of the UAV target tracking network is as Figure 2 shown, including the following steps:

[0093] S501. When processing the first frame of image, input the template image obtained in step S303 into the feature extraction network and the spatial position perception module to obtain the feature map of the template area image

[0094] S502. When processing subsequent frames other than the first frame, input the search image obtained in step S403 into the feature extraction network and the spatial position perception module to obtain the feature map of the search area image

[0095] S503. Respectively perform two-dimensional convolution, flattening, and channel transformation operations on the feature map and the feature map generated in step S501 to generate the feature map and the feature map

[0096] S504. Input the feature map F Z1 and the feature map F X1 generated in step S503 into the multi-frequency domain feature fusion module respectively. First, obtain the high-frequency features through global multi-head self-attention, then downsample the feature map, and then input it into the multi-head self-attention module to obtain the low-frequency features. Then, splice the high-frequency features and the low-frequency features to generate the feature and the feature Generate features through the residual connection method and features

[0097] The whole process can be summarized by the formula:

[0098] F Z3 = F Z1 + F Z2 (5)

[0099] F X3 = F X1 + F X2 (6)

[0100] where

[0101] F Z2 = MultiHead(F Z1 ) + MultiHead(AvgPool(F Z1 )) (7)

[0102] F X2 = MultiHead(F X1 ) + MultiHead(AvgPool(F X1 )) (8)

[0103] MultiHead(*) is the multi-head self-attention operation.

[0104] S505. Fuse the features F Z3 and the feature F X3 obtained in step S504 using cross-attention. That is, in this layer, two multi-head attentions are used. For the first multi-head attention, its matrix Q is composed of the feature F Z3 and the spatial position encoding generated by the sine function. K is composed of the feature F X3 and the spatial position encoding generated by the sine function. V is the feature F X3 , and a fusion matrix is generated through the attention mechanism For the second multi-head attention, its matrix Q is composed of the feature F X3 and the spatial position encoding generated by the sine function. K is composed of the feature F Z3 and the spatial position encoding generated by the sine function. V is the feature F Z3 , and a fusion matrix is generated through the attention

[0105] S506. Pass the output result of step S505 through another multi-head attention. Its matrix Q is the feature F generated in step S505 2Combined with the spatial position encoding generated by the sine function, where K is the feature F 1 Combined with the spatial position encoding generated by the sine function, where V is the feature F 1 , and the fused feature is generated through the attention mechanism Input it into the classification branch and the regression branch network to obtain the target tracking box output by the UAV target tracking network

[0106] In summary, the UAV target tracking method of the present invention can capture long-range dependencies while retaining accurate position information through the spatial position perception module to enhance object features in complex background situations. This enables the UAV target to be better extracted from the complex background. Ultimately, it improves the accuracy and success rate of UAV target tracking (see Figure 3 ).

[0107] The UAV target tracking method of the present invention is a multi-frequency domain target positioning method, which enables the network to be adaptable to targets of different scales and adopts a siamese network structure to reduce the number of model parameters

[0108] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention

Claims

1. A drone target tracking device, characterized in that: include: UAV target tracking network structure; The UAV target tracking network structure includes: a multi-frequency domain feature fusion module, and two spatial position perception modules connected to the multi-frequency domain feature fusion module, each spatial position perception module is also connected to a feature extraction network; the multi-frequency domain feature fusion module is also connected to a classification branch and a regression branch; The feature extraction network is used for feature extraction of template images and search images; The spatial position perception module is used to retain the precise position information of the features while capturing long-range dependencies, enhancing the object features, and better extracting the drone target from the complex background; The multi-frequency domain feature fusion module is used to fuse the features extracted by the feature extraction network and the spatial position perception module in sequence using cross attention; The classification branch is used to distinguish between the target area and the non-target area in tracking, that is, to determine where the target exists in the search area; The regression branch is used in tracking to fine-tune the coordinates, size, and aspect ratio of the bounding box at the location where the target may exist, thereby generating an accurate target box; The tracking result output by the UAV target tracking network structure is displayed on the display module.

2. A drone target tracking method applicable to the drone target tracking device according to claim 1, characterized in that: The steps include: S1. Use the camera to collect drone flight video; S2, determining whether the current frame of the image captured by the camera is the first frame of the tracking video; S3. If yes, manually mark the drone target and obtain the template area image; S4. If not, determine the search area image for drone target tracking based on the tracking result of the previous frame; S5, inputting the obtained template area and search area images into the UAV target tracking network; S6, displaying the tracking result output by the drone target tracking network on the display module; S7, obtaining the next frame of image data captured by the camera, and repeating steps S2-S6 until the video processing is completed.

3. The method for tracking a target by using an unmanned aerial vehicle according to claim 2, wherein: The construction process of the template region image in step S3 is: S301, obtaining the marked target coordinate position (x, y, w, h), where (x, y) represents the coordinates of the upper left corner of the tracking box covering the target, and (w, h) represents the width and height of the tracking box; S302: Calculate template clipping interval Where W Z and H Z The value is calculated as follows: Among them, W Z Represents the template area width, H Z Represents the template area is high; S303, convert the cropped area image of the template into Get a template image; Among them, Z represents the template image, Represents a tensor, that is, the cropped template area is converted into a tensor of size 128*128 and the number of channels is 3.

4. The method for tracking a target by using an unmanned aerial vehicle according to claim 3, characterized in that: The construction process of the search area image in step S4 is: S401, obtaining the target coordinate position (x, y, w, h) of the previous frame, where (x, y) represents the coordinates of the upper left corner of the tracking frame covering the target, and (w, h) represents the width and height of the tracking frame; S402: Calculate search and trim interval Where W X and H X The value is calculated as follows: Among them, W X Represents the width of the search area, H X It represents a high search area; S403, the cropped search area image is transformed into Obtaining a search image; Where X represents the search area image.

5. The method for tracking a target by using an unmanned aerial vehicle according to claim 2, wherein: In step S5, the training weights are loaded into the drone target tracking network, and target tracking is performed.

6. The method for tracking a target by using an unmanned aerial vehicle according to claim 4, characterized in that: In step S5, training weights are loaded into the drone target tracking network and target tracking is performed, including the following steps: S501: When processing the first frame image, the template image obtained in step S303 is input into the feature extraction network and the spatial position perception module to obtain a feature map of the template area image. S502: When processing subsequent frames except the first frame, the search image obtained in step S403 is input into the feature extraction network and the spatial position perception module to obtain a feature map of the search area image. S503: The feature map generated in step S501 and the feature map generated in step S502 Generate feature maps through two-dimensional convolution, flattening and channel transformation operations respectively and feature map S504: The feature graph F generated in step S503 is Z1 and feature map F X1 Input into the multi-frequency domain feature fusion module respectively; First, high-frequency features are obtained through global multi-head self-attention, and then the feature map is downsampled and input into the multi-head self-attention module to obtain low-frequency features; then the high-frequency features and low-frequency features are concatenated to generate features. and Features Then, through the residual connection method, the features are generated and Features The whole process can be summarized as follows: F Z3 =F Z1 +F Z2 (5) F X3 =F X1 +F X2 (6) in, F Z2 =MultiHead(F Z1 )+MultiHead(AvgPool(F Z1 )) (7) F X2 =MultiHead(F X1 )+MultiHead(AvgPool(F X1 )) (8) Among them, MultiHead(*) is a multi-head self-attention operation; S505: The feature F obtained in step S504 is Z3 and F X3 Use cross attention for fusion. In this layer, two multi-head attentions are used. For the first multi-head attention, its matrix Q is the feature F Z3 and the spatial position encoding generated by the sine function, K is the feature F X3 and the spatial position encoding generated by the sine function, V is the feature F X3 , generate the fusion matrix through the attention mechanism For the second multi-head attention, its matrix Q is the feature F X3 and the spatial position encoding generated by the sine function, K is the feature F Z3 and the spatial position encoding generated by the sine function, V is the feature F Z3 , generate the fusion matrix through attention S506: The output result of step S505 is passed through a multi-head attention mechanism, where the matrix Q is a combination of the feature F2 generated in step S505 and the spatial position code generated by the sine function, K is a combination of the feature F1 and the spatial position code generated by the sine function, and V is the feature F1. The fusion feature is generated through the attention mechanism. Input it into the classification branch and regression branch network to obtain the target tracking box output by the drone target tracking network.

Citation Information

Patent Citations

  • Target tracking method based on residual channel attention and multilevel classification regression

    CN113706581A

  • Image multi-classification network structure based on feature remapping and training method

    CN115439681A