Target tracking method for aerial video of unmanned aerial vehicle

By using ShuffleNet V2 to build a lightweight twin network in the SiamCAR network and introducing feature fusion and dynamic template updates, the SiamCAR network has solved the problem of high model complexity and weak anti-interference ability in drone aerial photography tracking, achieving more efficient feature extraction and tracking effects.

CN119942381APending Publication Date: 2025-05-06中船智控科技(武汉)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057060.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

SiamCAR network has high model complexity and weak anti-interference ability in drone aerial photography and tracking scenarios.

Method used

ShuffleNet V2 is used as feature extraction network to build a lightweight twin network framework, and introduce feature fusion and dynamic update mechanisms for target templates.

Benefits of technology

The number of parameters and calculations of the model is reduced, the feature extraction ability is enhanced, and the anti-interference ability and tracking effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942381A_ABST
    Figure CN119942381A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle aerial video target tracking method, which comprises the following steps of: constructing a twin network framework by using ShuffleNet V2 as a feature extraction network to obtain a SiamCAR target tracker comprising a feature extraction sub-network and a classification regression sub-network, calculating an initial template after initialization, initializing a dynamic template, calculating search graph features after reading a next frame of image, and obtaining a target tracking result; performing feature fusion, calculating a response diagram, outputting a tracking result, selecting the tracking result, calculating a peak sidelobe ratio, judging whether the peak sidelobe ratio is greater than a threshold value or not and whether an updating interval is reached or not, updating the dynamic template, and finally judging whether all frames are processed or not. According to the method, the feature maps output in each stage are fused in the twin network, so that the feature extraction capability of the network is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to a method for tracking targets in unmanned aerial vehicle aerial videos based on a twin network. Background Art

[0002] Drones are widely used because they are unmanned, highly maneuverable and easy to operate. In recent years, with the development of drone flight control systems and video imaging technology, aerial video target tracking has gradually become an essential function of drones. The drone aerial video target tracking system first uses a visual tracking algorithm to obtain the position offset information of the target in the image, and then the offset information is sent to the drone's servo control system to control the camera's visual axis to point to the target position to achieve target tracking. In the whole process, using computer vision technology to detect the position of the target of interest is a prerequisite for obtaining the target offset, which is critical for successful tracking.

[0003] Visual target tracking technology first marks the target of interest on the initial frame of the video sequence, and then runs the tracking algorithm to detect the tracked target on each subsequent frame. The target tracking algorithm based on deep learning has greatly improved the performance of the tracker compared to traditional methods and has become the mainstream development direction in the current target tracking field. The twin network is one of the representatives. It uses two parallel branches with the same weight to calculate the similarity of the outputs of the two branches.

[0004] As a target tracker based on twin networks, the SiamCAR network uses cross-correlation operations after parallel branches to calculate the response map of the output results of the two branches, and finally sends the response map to the classification regression subnetwork to predict the size and position of the target. Finally, a good tracking effect was achieved on the drone aerial photography test set, but the use of ResNet 50 to build a twin network increased the number of model parameters and calculations, and there were still deficiencies in dealing with complex scenarios such as target deformation, target occlusion, and interference from similar targets. Summary of the invention

[0005] Aiming at the problems of high model complexity and weak anti-interference ability of SiamCAR network in UAV aerial tracking scenes, the present invention proposes an aerial video target tracking method with lightweight, feature fusion and dynamic update of target template.

[0006] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention to solve the technical problem is: a method for tracking a target in an unmanned aerial video, comprising the following steps:

[0007] S1, using ShuffleNet V2 as the feature extraction network to build a twin network framework, and obtain the SiamCAR target tracker consisting of a feature extraction subnetwork and a classification regression subnetwork;

[0008] S2, with template image Z∈R 3×127×127 and search image X∈R 3×255×255 As input, the image features are extracted through the feature extraction sub-network respectively. After convolution and maximum pooling, three image features of different depths are finally obtained: Stage2, Stage3, and Stage4. The feature fusion module is used to combine the feature maps output by the template branch and the search branch. The feature map output by the template branch is Output. 32 (Z) and Output 43 (Z) Output of the search branch 32 (X) and Output 43 (X) Feature fusion is performed. The fused template branch feature map and the search branch feature map are cross-correlated by channel to obtain the response map. The response map of Stage 2 is added element by element with the response maps of Stage 3 and Stage 4 and then sent to the classification and regression subnetwork. Finally, the feature maps f are output through the classification branch. cls ∈R 1×25×25 , output feature map f through the center branch cen ∈R 2×25×25 , output feature map f through regression branch reg ∈R 4×25×25 , where the output results of the classification branch and the centrality branch are used to locate the target, and the result of the regression branch is used to estimate the scale of the target;

[0009] S3, calculate the initial template and initialize the dynamic template: After starting tracking, first obtain the position and scale of the tracked target in the first frame of the aerial video sequence through initialization, then crop the target area and send it to the twin network to obtain the feature map of the target initial template, and use the feature map to initialize the target dynamic template;

[0010] S4, then read the next frame of image, and cut out the search area of ​​the image according to the tracking result of the previous frame of image, and then send it to the twin network to calculate the deep features of the search image;

[0011] S5, feature fusion: feature fusion is performed on the template branch and search branch of the twin network respectively to obtain a fused feature map with both deep abstract information and shallow appearance information;

[0012] S6, calculate the response map: obtain the response map on the search area by performing channel-by-channel cross-correlation operation on the template branch and the search branch;

[0013] S7, output tracking results: the response graph is then sent to the classification and regression sub-network to output two tracking results, namely the target initial template and the target dynamic template, and then the tracking results are screened;

[0014] S8, finally calculate the peak sidelobe ratio of the tracking result, and determine whether the peak sidelobe ratio is greater than the threshold; that is, whether the target dynamic template update conditions are met. If the update conditions are met, the area of ​​the current frame target is cropped and sent to the twin network to calculate the deep features. If not, determine whether to process the next frame: if so, further determine whether the update interval is reached; if so, update the dynamic template: use the current tracking result to update the dynamic template features and end the target tracking after processing all frames; if it is determined that the peak sidelobe ratio is not greater than the threshold or the update interval is not reached, jump directly to determine whether all frames have been processed; otherwise repeat step S4, and if so, end the target tracking.

[0015] Furthermore, the feature fusion module first changes the number of channels of the shallow feature map and the deep feature map to 255 through a 1×1 convolution, and then adds the shallow feature map and the deep feature map element by element to obtain a single feature map, and then the feature map is calculated through the attention module to calculate the fusion factor; after the attention module outputs the fusion factor ω, the deep and shallow features after the channel change are weighted and summed to finally obtain the output after feature fusion.

[0016] Going a step further, the attention module consists of two branches, a left branch and a right branch: the left branch performs global average pooling on the feature map, and then changes the number of channels of the feature map through point-by-point convolution, and finally obtains the weight of each channel on the feature map; the right branch does not change the size of the feature map, and directly obtains a feature map consistent with the input shape; then the channel weight calculated on the left is multiplied by the feature map output on the right, adding channel attention to the feature map, so that the network pays more attention to the feature channel of interest; finally, the output of the attention module is normalized by the sigmoid operation.

[0017] Furthermore, in step S3, if the input channel number is M, the feature map is obtained by using M channels of size D K The convolution operation is performed channel-by-channel by using a single-channel convolution kernel. Each convolution kernel convolves a channel of the input feature map separately. Point-by-point convolution operations are performed by N convolution kernels of size 1×1 and number of channels M. Since the channel-by-channel convolution changes the size of the feature map, the point-by-point convolution changes the number of channels of the feature map. The number of output channels is N and the size is D. F feature map.

[0018] Furthermore, the template image Z∈R 3×127×127 After inputting the twin network, the twin network built using ShuffleNet V2 can output feature maps of the same size but different numbers of channels in Stage2, Stage3, and Stage4 respectively. and According to the tracking results of the previous frame, the position and size of the search area X of the current frame are determined, where the size is selected according to the following formula After selecting the search area X, it is cropped to a size of 255×255 and sent to the search branch. Finally, Stage 2, Stage 3, and Stage 4 of the twin network output and feature map.

[0019] Furthermore, the peak sidelobe ratio is used to describe the sharpness of data distribution, through the formula Calculate the ratio of the main lobe peak to the side lobe peak, where Represents the maximum response value in the response graph of the tth frame, and Represents the number of is the mean and standard deviation of the response values ​​in the region r at the center.

[0020] Furthermore, when the tracking target is blocked by the background or disturbed by other targets, by calculating The historical average value is used to establish the dynamic threshold where α psr is the threshold parameter; during tracking, if the PSR value of the response graph jointly output by the classification branch and the centrality branch is greater than If the interval from the last target dynamic template update exceeds the set update period, the target dynamic template is updated using the current tracking result.

[0021] Furthermore, during the tracking process, the continuously updated target dynamic template and the target initial template are sent to the tracker, and the initial template and the dynamic template tracking results are compared and the tracking result is selected. The present invention uses the intersection over union ratio IoU and the response peak difference PD to evaluate the prediction effect, which is defined as:

[0022]

[0023] In the formula represents the tracking result obtained using the target initial template and the target dynamic template, R max (z i ), R max (z s ) represents the maximum value of the target initial template and the target dynamic template on the response graph; if the intersection over union ratio IoU and the response peak difference PD calculated by the target dynamic template and the target initial template are both greater than the set threshold, the tracking result of the target dynamic template is selected as the final tracking result B, otherwise the tracking result of the target initial template is selected. The screening process is expressed as:

[0024]

[0025] Where α iou and α pdThey represent the intersection-over-combination ratio threshold and the response peak difference threshold respectively.

[0026] The beneficial effects of the present invention are as follows: the present invention adopts deep learning technology, and uses a modified lightweight feature extraction network ShuffleNet V2 to replace the original ResNet50 based on the SiamCAR network. ShuffleNet V2 uses deep separable convolution to replace conventional convolution, which decomposes conventional convolution into two parts: channel-by-channel convolution and point-by-point convolution. After decomposition, the number of parameters and the amount of calculation of convolution can be greatly reduced. The present invention fuses the feature maps output at each stage in the twin network, thereby enhancing the feature extraction capability of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flow chart of the target tracking method of the present invention;

[0028] Figure 2 It is a schematic diagram of the structure of the SiamCAR target tracker of the present invention;

[0029] Figure 3 It is a structural schematic diagram of the feature fusion module of the present invention;

[0030] Figure 4 This is the network structure after the present invention introduces dynamic template update. DETAILED DESCRIPTION

[0031] The present invention is further described as follows with reference to the accompanying drawings and embodiments:

[0032] Reference Figure 1 As shown, a method for tracking a target in an unmanned aerial vehicle aerial video disclosed in the present invention includes the following steps.

[0033] S1, using ShuffleNet V2 as the feature extraction network to build the twin network framework, and obtain the SiamCAR target tracker consisting of two parts: the feature extraction subnetwork and the classification regression subnetwork; Appendix Figure 2 The figure shows the basic structure of the target tracker of the present invention.

[0034] S2, Initialization: Get the initialization parameters of the target.

[0035] The target tracker is based on the template image Z∈R 3×127×127 and search image X∈R 3×255×255 As input, the image features are extracted through the feature extraction sub-network, and after convolution and maximum pooling, three image features of different depths are finally obtained: Stage2, Stage3, and Stage4. These features are then added to the Figure 3The feature fusion module shown in the figure performs feature fusion, that is, the feature map output by the template branch and the search branch. 32 (Z) and Output 43 (Z) Output of the search branch 32 (X) and Output 43 (X) Feature fusion is performed. The fused template branch feature map and the search branch feature map are cross-correlated by channel to obtain a response map. The response map of Stage 2 is added element by element with the response maps of Stage 3 and Stage 4 and then sent to the classification and regression subnetwork. Finally, the classification branch outputs the feature map f cls ∈R 1×25×25 , the centrality branch outputs the feature map f cen ∈R 2×25×25 , the regression branch outputs the feature map f reg ∈R 4×25×25 , where the output results of the classification branch and the centrality branch are used to locate the target, and the result of the regression branch is used to estimate the scale of the target.

[0036] Attached Figure 3 The feature fusion module used in the present invention is shown. The module first changes the number of channels of the shallow feature map and the deep feature map to 255 through a 1×1 convolution, and then adds the shallow feature map and the deep feature map element by element to obtain a single feature map, and then the feature map is passed through the attention module to calculate the fusion factor. After the attention module outputs the fusion factor ω, the deep and shallow features after the channel change are weighted and summed, and finally the output after feature fusion is obtained.

[0037] Figure 3 The attention module in the paper consists of two branches, where the left branch performs global average pooling on the feature map, and then changes the number of channels of the feature map through point-by-point convolution, and finally obtains the weight of each channel on the feature map. The right branch does not change the size of the feature map, and directly obtains a feature map with the same shape as the input. After that, the channel weight calculated on the left is multiplied by the feature map output on the right, adding channel attention to the feature map, so that the network pays more attention to the feature channel of interest. Finally, the output of the attention module is normalized by the sigmoid operation.

[0038] The previous SiamCAR network used ResNet50 as the feature extraction network to build the twin network framework, which greatly increased the model space complexity and computational complexity. The present invention uses the modified lightweight feature extraction network ShuffleNetV2 to replace the original ResNet50. ShuffleNet V2 uses depthwise separable convolution to replace conventional convolution, which decomposes conventional convolution into two parts: channel-by-channel convolution and point-by-point convolution. After decomposition, the number of convolution parameters and the amount of computation can be greatly reduced. Therefore, the SiamCAR framework of this patent application uses huffleNet V2 as the backbone network.

[0039] If the number of channels of the input feature map is M, the output feature map requires the number of channels to be N and the size to be D F The channel-by-channel convolution uses M single channels of size D K Convolution is performed with N convolution kernels, where each convolution kernel convolves one channel of the input feature map separately. Point-wise convolution uses N convolution kernels of size 1×1 and M channels to perform the convolution operation.

[0040] Since channel-by-channel convolution changes the size of the feature map and point-by-point convolution changes the number of channels in the feature map, any standard convolution can be replaced by depthwise separable convolution, which greatly reduces the number of parameters and computation. The ratio of the number of parameters α and the ratio of the amount of computation β between depthwise separable convolution and standard convolution are respectively

[0041]

[0042] In order to make ShuffleNet V2 more suitable for target tracking tasks, the ShuffleNet V2 used in the present invention removes the last convolution layer, the global pooling layer and the fully connected layer in the original network, modifies the network step size, and makes Stage 3 and Stage 4 not downsample the feature map. At the same time, in order to increase the receptive field of the convolution layer, atrous convolutions with dilation rates of 2 and 4 are used in Stage 3 and Stage 4 respectively.

[0043] After using ShuffleNet V2 to build a twin network, the twin network can output feature maps of the same size but different numbers of channels in Stage2, Stage3, and Stage4. 3×127×127 After inputting the twin network, Stage2, Stage3, and Stage4 will output and feature map.

[0044] S3, calculate the initial template and initialize the dynamic template: obtain the initial template of the labeled target through the twin network. After starting tracking, first obtain the position and scale of the tracked target in the first frame of the aerial video sequence through initialization, then crop the target area and send it to the twin network to obtain the feature map of the target initial template, and use the feature map to initialize the target dynamic template; otherwise, determine whether all frames have been processed, and if so, end.

[0045] S4, then reads the next frame of image, and crops the search area of ​​the image according to the tracking result of the previous frame of image, and then sends it to the twin network to calculate the deep features of the search graph.

[0046] According to the tracking results of the previous frame, the position and size of the search area X of the current frame are determined, where the size is selected according to the following formula After selecting the search area, X is cropped to 255×255 and sent to the search branch. Finally, the network's Stage2, Stage3, and Stage4 output and feature map.

[0047] S5, feature fusion: The template branch and search branch of the twin network are fused separately to obtain a fused feature map with both deep abstract information and shallow appearance information.

[0048] In a deep convolutional neural network, the feature map output by the shallow structure of the network mainly contains feature information such as target color and shape that are helpful for target positioning, while the feature map output by the deep structure of the network contains more abstract semantic information of the target. These high-level semantic information are more conducive to target classification and detection. The present invention fuses the feature maps output at each stage in the twin network to enhance the feature extraction capability of the network.

[0049] The feature maps output by Stage2, Stage3, and Stage4 are first transformed into 1×1 convolution layers to transform the number of channels. After that, Stage4 and Stage3 are first added element by element, and then the weight factors 1-ω and ω of the output feature maps of Stage4 and Stage3 are obtained through an attention module Att. Finally, the weight factors are used for the weighted fusion of the feature maps of the two stages: Output 43 On the one hand, it is used to perform channel-by-channel cross-correlation in the template branch and the search branch. On the other hand, it is used to perform feature fusion in the same way as above with the feature map output by the shallow Stage 2 of the twin network, and finally outputs Output 32 .

[0050] S6, calculate the response map: obtain the response map on the search area by performing channel-by-channel cross-correlation operation between the template branch and the search branch.

[0051] Output of the feature map output by the search branch 32 (X) and Output 43 (X), feature map output by the template branch Output 32 (Z) and Output 43 (Z), and correspondingly perform channel-by-channel cross-correlation operations to obtain response graphs R1 and R2:

[0052]

[0053] Then the response graphs R1 and R2 are weighted to obtain the final output response graph R of the twin network:

[0054] R=μ×R1+(1-μ)×R2,

[0055] Where μ represents the weight parameter, which is learned during the network training process.

[0056] S7, output tracking results: send the response map to the classification regression subnetwork, output two tracking results of the target initial template and the target dynamic template, and then screen the tracking results.

[0057] The response map R is sent to the classification regression subnetwork to predict the size and position of the target. The classification regression subnetwork contains the classification branch cls, the center branch cen and the regression branch reg. The cls and cen branches output the category score and quality assessment score of a single channel to jointly determine the position of the tracked target. The reg branch outputs the distance value l from the current position (x, y) to the four sides of the target bounding box. * ,t * 、r * and b * , and finally the coordinates of the upper left corner (x0, y0) and the lower right corner (x1, y1) of the target in the search area are obtained through coordinate mapping, where s = 8 represents the step size of the network:

[0058]

[0059] S8, finally calculate the peak sidelobe ratio of the tracking result, and determine whether the peak sidelobe ratio is greater than the threshold; that is, whether the target dynamic template update conditions are met. If the update conditions are met, the area of ​​the current frame target is cropped and sent to the twin network to calculate the deep features. If not, determine whether to process the next frame: if so, further determine whether the update interval is reached; if so, update the dynamic template: use the current tracking result to update the dynamic template features and end the target tracking after processing all frames; if it is determined that the peak sidelobe ratio is not greater than the threshold or the update interval is not reached, jump directly to determine whether all frames have been processed; otherwise repeat step S4, and if so, end the target tracking.

[0060] The SiamCAR network only uses the target features extracted on the initial frame as the cross-correlated target template output response map. As the tracking time goes by, the tracking background and target appearance may change greatly, resulting in a large difference in the target features at this time compared to the initial frame, making it difficult for the tracker to handle challenging scenarios such as changes in the target appearance state. Based on the SiamCAR network, the present invention introduces a tracking target dynamic template and uses an evaluation strategy based on the peak sidelobe ratio to determine whether the target dynamic template is updated.

[0061] The peak-to-sidelobe ratio is the ratio of the main lobe peak to the sidelobe peak, which can be used to describe the sharpness of the data distribution and is defined as in Represents the maximum response value in the response graph of the tth frame, and Represents the number of is the mean and standard deviation of the response values ​​in the region r at the center.

[0062] When the tracking target is blocked by the background or interfered by other targets, The value of will suddenly drop, the present invention calculates The historical average value is used to establish the dynamic threshold where α psr is the threshold parameter. During tracking, if the PSR value of the response graph jointly output by the classification branch and the centrality branch is greater than If the interval between the last target dynamic template update and the set update period exceeds the set update period, the current tracking result is used to update the target dynamic template. The continuously updated target dynamic template and the target initial template are sent to the tracker during the tracking process, and finally two tracking results, the initial template and the dynamic template, are obtained. The two tracking results need to be compared and the best tracking result is selected.

[0063] The present invention uses the intersection over union (IoU) and the response peak difference (PD) to evaluate the prediction effect, which is defined as In the formula represents the tracking result obtained using the target initial template and the target dynamic template, Rmax (z i ), R max (z s ) represents the maximum value of the target initial template and the target dynamic template on the response graph.

[0064] If the IoU and PD differences calculated between the target dynamic template and the target initial template are both greater than the set threshold, the tracking result of the target dynamic template is selected as the final tracking result B, otherwise the tracking result of the target initial template is selected. The screening process is expressed as Where α iou and α pd They represent the intersection-over-combination ratio threshold and the response peak difference threshold respectively.

[0065] Attached Figure 4 The tracking framework after the twin network is added to the target dynamic template update is demonstrated. Due to the addition of the target dynamic template, the network becomes a three-branch structure, so it can simultaneously output two tracking results using the target initial template and the target dynamic template. The screening process formula is used as the screening condition to select the best tracking result. After selecting the tracking result, the peak sidelobe ratio is calculated on the response graph of the result. Then, the value is judged whether it is greater than the set threshold and whether the update interval is reached to decide whether to update the target dynamic template. If the target dynamic template needs to be updated, the target dynamic template image is cropped using the current final tracking result, and then sent to the twin network to extract deep features. The feature is used as the feature of the target dynamic template to calculate the cross-correlation with the search area.

[0066] The above embodiments are only illustrative of the principles and effects of the present invention, as well as some embodiments of its application. For those skilled in the art, several modifications and improvements may be made without departing from the creative concept of the present invention, and all of these belong to the protection scope of the present invention.

Claims

1. A method for tracking a target in an unmanned aerial video, characterized in that: The following steps are included S1, using ShuffleNet V2 as the feature extraction network to build a twin network framework, and obtain the SiamCAR target tracker including the feature extraction subnetwork and the classification regression subnetwork; S2, with template image Z∈R 3×127×127 and search image X∈R 3×255×255 As input, the image features are extracted through the feature extraction sub-network to obtain image features of different depths Stage2, Stage3, and Stage4. The feature fusion module is used to fuse the feature maps output by the template branch and the search branch. The fused template branch feature map and the search branch feature map are cross-correlated by channel to obtain the response map. The response map of Stage2 is added element by element with the response maps of Stage3 and Stage4 and then sent to the classification and regression sub-network. The feature maps f are output by the classification branch. cls ∈R 1×25×25 , output feature map f through the center branch cen ∈R 2×25×25 , output feature map f through regression branch reg ∈R 4 ×25×25 , the output results of the classification branch and the centrality branch are used to locate the target, and the result of the regression branch is used to estimate the scale of the target; S3, after starting tracking, first obtain the position and scale of the tracked target in the first frame of the aerial video, then crop the target area and send it to the twin network to obtain the feature map of the target initial template, and use the feature map to initialize the target dynamic template; S4, read the next frame of image, and cut out the search area of ​​the image according to the tracking result of the previous frame of image, and then send it to the twin network to calculate the deep features of the search image; S5, perform feature fusion on the template branch and search branch of the twin network respectively to obtain a fused feature map; S6, performing channel-by-channel cross-correlation operation on the template branch and the search branch to obtain a response map on the search area; S7, sending the response graph to the classification regression sub-network to output two tracking results, namely the target initial template and the target dynamic template, and then screening the tracking results; S8, calculate the peak sidelobe ratio of the tracking result, and determine whether the peak sidelobe ratio is greater than the threshold; if so, further determine whether the update interval is reached; if so, update the dynamic template: use the current tracking result to update the dynamic template feature and end the target tracking after processing all frames; if it is determined that the peak sidelobe ratio is not greater than the threshold or the update interval is not reached, directly jump to determine whether all frames have been processed; otherwise repeat step S4, and if so, end the target tracking.

2. The method for tracking a target in an unmanned aerial video according to claim 1, characterized in that: The feature fusion module first changes the number of channels of the shallow feature map and the deep feature map to 255 through a 1×1 convolution, and then adds the shallow feature map and the deep feature map element by element to obtain a single feature map, which is then input into the attention module. After the attention module outputs the fusion factor ω, the deep and shallow features after the channel change are weighted and summed to finally obtain the output after feature fusion.

3. The method for tracking a target in an unmanned aerial video according to claim 2, characterized in that: The attention module consists of a left branch and a right branch: the left branch performs global average pooling on the feature map and then changes the number of channels through point-by-point convolution to finally obtain the weight of each channel on the feature map; the right branch directly obtains a feature map with the same shape as the input; then the channel weight of the left branch is multiplied to the feature map output by the left branch, and finally the output of the attention module is normalized by a sigmoid operation.

4. The method for tracking a target in an unmanned aerial video according to claim 3, characterized in that: In the step S3, the input channel number is M, and the feature map is obtained by M channels of size D. K The single-channel convolution kernel performs channel-by-channel convolution operations. Each convolution kernel convolves a channel of the input feature map separately. The point-by-point convolution operation is performed through N convolution kernels of size 1×1 and channel number M. The output channel number is N and the size is D. F feature map.

5. The method for tracking a target in an unmanned aerial video according to claim 4, characterized in that: The template image Z∈R 3×127×127 Input the twin network, Stage2, Stage3 and Stage4 output feature maps of the same size but different numbers of channels respectively and According to the tracking result of the previous frame, the position of the search area X of the current frame is determined, and the size is calculated according to the following formula: After selecting the search area X, crop it to 255×255 and send it to the search branch. Finally, Stage2, Stage3 and Stage4 output it respectively. and feature map.

6. The method for tracking a target in an unmanned aerial video according to claim 5, characterized in that: The peak-to-sidelobe ratio is given by the formula Calculate the ratio of the main lobe peak to the side lobe peak, where Represents the maximum response value in the response graph of the tth frame, and Represents the number of is the mean and standard deviation of the response values ​​in the region r at the center.

7. The method for tracking a target in an unmanned aerial video according to claim 6, characterized in that: When the tracking target is blocked by the background or interfered by other targets, the The historical average value is used to establish the dynamic threshold where α psr is the threshold parameter; during tracking, if the PSR value of the response graph jointly output by the classification branch and the centrality branch is greater than If the interval from the last target dynamic template update exceeds the set update period, the target dynamic template is updated using the current tracking result.

8. The method for tracking a target in an unmanned aerial video according to claim 7, characterized in that: During the tracking process, the continuously updated target dynamic template and the target initial template are sent to the tracker. The initial template and the dynamic template tracking results are compared and screened. The intersection over union (IoU) and response peak difference (PD) are used to evaluate the prediction effect: In the formula represents the tracking result obtained using the target initial template and the target dynamic template, R max (z i ), R max (z s ) represents the maximum value of the target initial template and the target dynamic template on the response graph; If the intersection over union (IoU) and response peak difference (PD) calculated by the target dynamic template and the target initial template are both greater than the set threshold, the tracking result of the target dynamic template is selected as the final tracking result B, otherwise the tracking result of the target initial template is selected: Where α iou and α pd They represent the intersection-over-combination ratio threshold and the response peak difference threshold respectively.