A method, system, medium, device and terminal for unmanned aerial vehicle target tracking

The SiamTCP model with pixel-level correlation and cascaded matching addresses the challenges of occlusions and background interference in drone tracking, providing stable and precise target localization.

CN115619829BActive Publication Date: 2025-07-15ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211384704.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-07-15
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

The target tracking algorithm of the existing twin network from the perspective of drones is affected by obstacles, changes in the appearance of tracking objects and complex background interference, resulting in unstable tracking. The target information in the first frame contains background information, affecting the matching effect.

Method used

Using a drone target tracking method based on target-aware cascading pixel matching, a SiamTCP model is built through the improved GoogLeNet network and PyTorch framework, feature decomposition and cascading pixel matching are performed, background information is eliminated, and fine-grained matching is performed.

Benefits of technology

It achieves more accurate target tracking in unmanned airport scenes, can cope with target occlusion and deformation in complex scenarios, improves tracking accuracy and speed, and is suitable for applications such as drone navigation and smart agricultural seeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115619829B_ABST
    Figure CN115619829B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of object tracking in computer vision, and discloses a drone object tracking method, system, medium, device and terminal. The backbone network is pre-trained using the ImageNet dataset; a SiamTCP drone tracking model is created, and the SiamTCP drone tracking model is trained on multiple video datasets. The drone object tracking method based on target-aware cascaded pixel matching provided by the present invention determines the parameters of the neural network model by training on multiple datasets. The trained model can be applied to the object tracking task in the drone scenario. Specifically, the present invention proposes pixel-by-pixel cross-correlation and performs a cascaded operation, replacing the global-to-global matching with a part-to-part matching operation, enhancing the result of feature fusion. At the same time, the target-aware module proposed by the present invention can eliminate unnecessary background information and can cooperate well with the cascaded pixel matching process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object tracking in computer vision, and particularly relates to a method, system, medium, device and terminal for unmanned aerial vehicle (UAV) object tracking. Background Art

[0002] Visual object tracking is a basic research topic in computer vision. The task of visual object tracking is usually to select the object of interest at the beginning of the video and locate the target in the subsequent sequence. In daily life, visual object tracking has many application scenarios, such as video surveillance, intelligent human-computer interaction, autonomous driving, etc. In recent years, with the progress of tracking algorithms, the field of object tracking has developed rapidly. However, due to reasons such as occlusion by obstacles, appearance changes of the tracked object, interference from similar objects near the target, and rapid movement of the tracked object, the tracker often fails to track stably. Therefore, designing a tracking algorithm with faster speed and more accurate positioning is still challenging.

[0003] Currently, considering the superior tracking effect of convolutional neural networks, the present invention applies them to the field of object tracking. Among them, the tracking algorithm based on Siamese networks has received great response and attention. However, if applied to real-time object tracking from the perspective of an aerial UAV, these algorithms still have some limitations. Since Siamese networks usually directly use the first frame as the template image, the target information in the first frame has a great impact on the matching of subsequent frames. Different from general Siamese network object trackers, the present invention selects to perform masking processing on the template features obtained after convolution, only retaining the feature information of the target, and avoiding the interference of complex backgrounds on the target object. Considering the uniqueness of aerial tracking, that is, the objects appearing in the scene are usually relatively small, the present invention designs a more fine-grained matching method, that is, performing cascaded pixel correlation on the template and the search area. Therefore, the present invention aims to propose an effective UAV object tracking method.

[0004] Through the above analysis, the problems and defects existing in the prior art are as follows:

[0005] (1) Due to reasons such as occlusion by obstacles and appearance changes of the tracked object, the global matching used in Siamese networks lacks discriminative representation of the target, making the tracker often unable to track stably.

[0006] (2) In existing Siamese network-based tracking algorithms, the first-frame target area is usually used as the template, but it also contains background information, which has a certain negative interference on UAV tracking. Summary of the Invention

[0007] In view of the problems existing in the prior art, the present invention provides a method, system, medium, device and terminal for unmanned aerial vehicle (UAV) target tracking, and in particular, relates to a method, system, medium, device and terminal for UAV target tracking based on target-aware cascaded pixel matching.

[0008] The present invention is implemented as follows. A method for UAV target tracking, the method for UAV target tracking includes: pre-training the backbone network using the ImageNet dataset; based on the Siamese network structure, building a SiamTCP UAV tracking model using the deep learning framework PyTorch; training the SiamTCP UAV tracking model on multiple video datasets, and updating the network parameters through the loss function.

[0009] Further, the method for UAV target tracking includes the following steps:

[0010] Step 1, initialize the backbone network;

[0011] Step 2, construct a UAV target tracking model based on target-aware cascaded pixel matching;

[0012] Step 3, train the UAV target tracking model.

[0013] Further, the backbone network in Step 1 adopts an improved GoogLeNet; the backbone network is pre-trained on the ImageNet image classification dataset to initialize the weights of the backbone network.

[0014] Further, the construction of the UAV target tracking model based on target-aware cascaded pixel matching in Step 2 includes: building a SiamTCP UAV tracking model using the deep learning framework PyTorch, and the SiamTCP consists of an improved backbone network GoogLeNet, pixel-level correlation, a target-aware module, a cascaded pixel matching module, and classification and regression sub-networks with a fully convolutional neural network structure. The input of the SiamTCP tracking model adopts a dual-branch structure, namely a template branch and a search region branch, and the two branches are sent to the CNN for feature extraction; the target-aware cascaded pixel matching module includes a target-aware module and a cascaded pixel-level correlation module for feature fusion; the subsequent classification and regression sub-networks are used for target localization.

[0015] For the pixel-level correlation, the template feature is decomposed into multiple spatial kernels to achieve high-quality feature representation, and the spatial kernel size is 1×1; the template feature T t is divided along the width and height to obtain n t small kernels;

[0016] n t = w t × h t;

[0017] Among them, w t and h t represent the width and height of the template feature. The template feature is divided into a group with a size of 1×1×C, where C is the channel of the template feature.

[0018] During the part-to-part matching process, w s and h s are represented as the width and height of the search region feature S f , and is represented as the j-th row and the i-th column in S f ;

[0019]

[0020] Among them, is the similarity between the n-th position of T t and in the spatial dimension.

[0021] Furthermore, the target perception module generates a variable template through the supervision of the marked bounding box B t , projects B t onto the template feature Z t to obtain the region of interest R t . The new template feature T t can be expressed as:

[0022]

[0023] The processed template feature T t is decomposed into many kernels with a size of 1×1×C, where C is the channel of the template feature T t . The obtained spatial kernels can be used to calculate the per-pixel correlation.

[0024] The cascaded pixel matching module performs multiple cascades on the feature fusion between the template region and the search region. S f represents the search feature map after feature extraction, and T t represents the template feature map after feature extraction, and the template feature map is obtained after being processed by the target perception module; The SiamTCP model performs a pixel-level cross-correlation operation on the search feature map S f and the template feature map T t to obtain F1, and then combines F1 obtained from the pixel-level cross-correlation operation with the search feature map S fPerform a splicing operation to obtain F2; continuously perform pixel-level cross-correlation operations between the newly obtained features and the template features; splice the features obtained from all pixel-level correlation operations to obtain new features; the SiamTCP model repeats the operation three times.

[0025] Further, the training of the drone target tracking model in step three includes: using stochastic gradient descent to train for 20 iterations on five datasets of the constructed models COCO, ImageNet DET, ImageNet VID, YouTube-BB, and GOT-10k, and setting the training batch size to 76; fixing the parameters of the first 10 iterations in the improved GoogLeNet for training the head network of the model. In the last 10 iterations, fix the parameters of the first and second stages in the backbone network and fine-tune the parameters of the third and fourth stages; for the first 5 iterations, set a warm-up learning rate that linearly increases from 0.005 to 0.01; for the last 15 iterations, use a learning rate that exponentially decays to 0.0005; throughout the training process, use 127 pixels as the template patch and 287 pixels as the search area.

[0026] Another object of the present invention is to provide a drone target tracking system based on target-aware cascaded pixel matching as described above. The drone target tracking system includes: testing and deploying the trained SiamTCP network model for target tracking tasks in drone scenarios.

[0027] Another object of the present invention is to provide a computer device. The computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the drone target tracking method.

[0028] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the drone target tracking method.

[0029] Another object of the present invention is to provide an information data processing terminal for implementing the drone target tracking system as described above.

[0030] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are:

[0031] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving such problems, closely combining with the technical solution to be protected by the present invention, as well as the results and data during the R & D process, etc., analyze in detail and profoundly how the technical solution of the present invention solves the technical problems and the creative technical effects brought about after solving the problems. The specific description is as follows:

[0032] The object of the present invention is to provide a UAV target tracking method based on target-aware cascaded pixel matching. By training on multiple data sets, the parameters of the neural network model are determined. After the training is completed, the model can be applied to the target tracking task in the UAV scenario.

[0033] Most Siamese network-based trackers fuse the two-way features using the original cross-correlation or depthwise separable correlation operations. However, this method blurs the spatial and detailed information. Because their kernels are a complete template patch for cross-correlation calculation with the search area. Since UAV tracking requires obtaining as much spatial and detailed information as possible, traditional cross-correlation or depthwise cross-correlation is not applicable to the aerial tracking scenario. Therefore, the present invention proposes pixel-by-pixel cross-correlation and performs a cascaded operation. By replacing the global-to-global matching with a part-to-part matching operation, more accurate results can be obtained. At the same time, considering the target tracking in the UAV scenario where there is more interference information around, the target-aware module proposed by the present invention eliminates unnecessary background information and can cooperate well with the cascaded pixel matching process.

[0034] Second, regarding the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are specifically described as follows:

[0035] Considering the unique characteristics of aerial tracking, that is, the objects appearing in the scene are usually relatively small, the present invention provides a UAV target tracking model based on target-aware cascaded pixel matching, which adopts a different cross-correlation method from the traditional one, that is, a cascaded pixel-level correlation operation to perform a fine-grained matching process.

[0036] The present invention mainly proposes a real-time tracking model of target-aware cascaded pixel matching. By performing large-scale offline training on multiple public video data sets, the parameters of the network model are determined; after the training is completed, the model can be better applied to the target tracking in the UAV scenario.

[0037] Third, as the creative auxiliary evidence of the claims of the present invention, it is also reflected in the following important aspects:

[0038] (1) The expected benefits and commercial value after the transformation of the technical solution of the present invention are as follows: The UAV target tracking model proposed by the present invention can better capture the discriminative representations of the target through the target perception and pixel matching modules, and can complete the UAV target tracking task with variable perspectives, thereby serving UAV operations (such as UAV navigation, intelligent agricultural seeding and other application scenarios) more effectively.

[0039] (2) The technical solution of the present invention solves the technical problems that people have always been eager to solve but have never been successful: The UAV target tracking model proposed by the present invention benefits from cascaded pixel-level matching and can, to a certain extent, cope with challenging factors such as partial occlusion and deformation of the target in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0041] Figure 1 It is a flowchart of the UAV target tracking method provided by the embodiment of the present invention;

[0042] Figure 2 It is a basic architecture diagram of the SiamTCP target tracking model provided by the embodiment of the present invention;

[0043] Figure 3 It is a schematic diagram of the decomposition of the template feature in pixel-level correlation provided by the embodiment of the present invention;

[0044] Figure 4 It is a schematic diagram of similarity matching in pixel-level correlation provided by the embodiment of the present invention;

[0045] Figure 5 It is a schematic diagram of pixel-level correlation combined with the target perception module provided by the embodiment of the present invention;

[0046] Figure 6 It is a schematic diagram of the cascaded pixel-level matching process provided by the embodiment of the present invention;

[0047] Figure 7 It is a schematic diagram of the performance and speed comparison with other tracking algorithms on the UAV123 dataset provided by the embodiment of the present invention;

[0048] Figure 8 It is a schematic diagram of the accuracy comparison with other tracking algorithms on the OTB100 dataset provided by the embodiment of the present invention;

[0049] Figure 9It is a schematic diagram showing the comparison of the success rate of the present invention's embodiment with other tracking algorithms on the OTB100 dataset. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] Aiming at the problems existing in the prior art, the present invention provides a method, system, medium, device and terminal for unmanned aerial vehicle (UAV) target tracking. The following describes the present invention in detail with reference to the accompanying drawings.

[0052] I. Explanation of embodiments. In order to enable those skilled in the art to fully understand how the present invention is specifically implemented, this part is an explanatory embodiment that expands and explains the technical solution of the claims.

[0053] As Figure 1 shown, the UAV target tracking method provided by the embodiment of the present invention includes the following steps:

[0054] S101, performing pre-training on the backbone network using the ImageNet dataset;

[0055] S102, creating a SiamTCP UAV tracking model;

[0056] S103, training the SiamTCP model on multiple video datasets.

[0057] As a preferred embodiment, the UAV target tracking method provided by the embodiment of the present invention specifically includes the following steps:

[0058] Step 1, pre-training of the backbone network;

[0059] The backbone network of the present invention uses an improved GoogLeNet (Inception v3), and the weights of the backbone network are initialized by pre-training the backbone network on the ImageNet image classification dataset;

[0060] Step 2, constructing a UAV target tracking model based on target-aware cascaded pixel matching;

[0061] Use the deep learning framework Pytorch to build the SiamTCP model. This model consists of an improved backbone network GoogLeNet (Inception v3), pixel-level correlation, target-aware module, cascaded pixel matching module, and classification and regression sub-networks. The basic framework provided by the embodiment of the present invention is as Figure 2As shown, the input of this algorithm adopts a dual-branch structure, namely the template patch branch and the search area branch. These two branches are sent to the CNN for feature extraction. The target-aware cascaded pixel matching module is mainly used for feature fusion, which includes a target-aware module and a cascaded pixel-level correlation module. The subsequent classification and regression sub-network is used for target localization. In the embodiment of the present invention, the template feature is decomposed into multiple spatial kernels to achieve high-quality feature representation, and the size of the spatial kernel is 1×1. As Figure 3 shown, the template feature T t is divided along the width and height to obtain n t kernels of size 1×1.

[0062] n t = w t ×h t

[0063] where w t and h t represent the width and height of the template feature. In the embodiment of the present invention, the template feature is divided into a group of size 1×1×C, where C is the number of channels of the template feature.

[0064] The part-to-part matching process is as Figure 4 shown. In the embodiment of the present invention, w s and h s are represented as the width and height of the search area feature S f . In the embodiment of the present invention, is represented as the j-th row and the i-th column in S f .

[0065]

[0066] is the similarity between the n-th position of T t and in the spatial dimension.

[0067] Traditional cross-correlation calculations usually have a lot of background interference in the response map because they match the search feature with the target feature cropped from the center of the template. Different from the above traditional method, in the embodiment of the present invention, a variable template is generated through the supervision of the marked bounding box B t , and B t is projected onto the template feature Z t to obtain a region of interest R t . Through this simple operation, in the embodiment of the present invention, a new template feature T t can be obtained, which can be expressed as:

[0068]

[0069] The processed template feature T t is decomposed into a number of kernels of size 1×1×C, where C is the number of channels of the template feature T t Then, the embodiments of the present invention use the obtained spatial kernels to calculate the per-pixel correlation. The illustration between the target-aware and pixel-level correlation module proposed by the embodiments of the present invention is as shown in Figure 5 shown

[0070] The cascaded pixel matching module proposed by the embodiments of the present invention is, specifically, the feature fusion between the cascaded template region and the search region. The cascaded pixel matching module is as shown in Figure 6 shown, where S f represents the search feature map after feature extraction, T t represents the template feature map after feature extraction, and this template feature map is obtained after being processed by the target-aware module. The SiamTCP model first performs a pixel-level cross-correlation operation on the search feature map S f and the template feature map T t to obtain F1, and then performs a concatenation operation on F1 obtained from the pixel-level cross-correlation operation and the search feature map S f to obtain F2. By continuously performing pixel-level cross-correlation operations on the newly obtained features and the template features, and then concatenating the features obtained from all previous pixel-level correlation operations to obtain new features. The SiamTCP model repeats this operation three times in total. Through this cascaded operation, the SiamTCP model can capture more detailed information of feature fusion, which is of great significance for subsequent positioning

[0071] Step 3, model training;

[0072] The constructed model COCO, ImageNet DET, ImageNet VID, YouTube-BB and GOT-10k five datasets are trained using stochastic gradient descent for 20 iterations. The embodiments of the present invention set the training batch size to 76. In order to train the head network of the model, the present invention fixes the parameters of the first 10 iterations in the improved GoogLeNet. In the last 10 iterations, the present invention fixes the parameters of the first stage and the second stage in the backbone network, and then fine-tunes the parameters of the third stage and the fourth stage. For the first 5 iterations, the present invention sets a warm-up learning rate that linearly increases from 0.005 to 0.01. For the last 15 iterations, the present invention uses a learning rate that exponentially decays to 0.0005. Throughout the training process, the present invention uses 127 pixels as the template patch and 287 pixels as the search region

[0073] For the process and effect of processing images provided by the embodiments of the present invention, please refer toFigures 1 to 9 as shown

[0074] The UAV target tracking system provided by the embodiment of the present invention includes: testing and deploying the trained SiamTCP network model for target tracking tasks in the UAV scenario.

[0075] II. Application embodiments. In order to prove the creativity and technical value of the technical solution of the present invention, this part is an application embodiment of the technical solution of the claims on specific products or related technologies.

[0076] The UAV target tracking model provided by the application embodiment of the present invention can complete the UAV target tracking task with variable perspectives. Benefiting from the cascaded pixel-level matching, it can, to a certain extent, cope with challenging factors such as target occlusion, deformation, and interference from similar objects in complex scenarios. At the same time, it can serve actual UAV operations more effectively, such as UAV navigation, intelligent agricultural seeding, and other application scenarios.

[0077] III. Evidence of related effects of the embodiments. Some positive effects have been achieved during the research and development or use of the embodiments of the present invention, and it indeed has great advantages compared with the prior art. The following content is described in combination with data, charts, etc. in the test process.

[0078] Table 1 Performance comparison of the target tracking algorithm of the present invention and other algorithms on the UAV123 dataset

[0079]

[0080] To better illustrate the effectiveness of the embodiments of the present invention in target tracking in the UAV scenario, performance comparisons were made with some other siamese network algorithms on the UAV123 dataset, as shown in Table 1. SiamTCP of the present invention exceeds the 6 compared tracking methods in both accuracy and precision metrics. In addition, Figure 7 It further demonstrates the superiority of the performance and speed of the present invention patent on the UAV123 dataset. The tracking speed can reach 50 FPS, and it is faster and more accurate than the tracking speed of SiamRPN++.

[0081] Figure 8 and Figure 9 Further, the one-time evaluation curve graphs of the precision and accuracy metrics of the embodiments of the present invention on the OTB100 dataset are respectively shown. SiamTCP of the present invention performs significantly better than other target tracking methods including SiamRPN and SiamDW.

[0082] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.

[0083] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for unmanned aerial vehicle target tracking, characterized in that, The described UAV target tracking method includes: pre-training the backbone network using the ImageNet dataset; building a SiamTCP UAV tracking model based on the Siamese network structure using the deep learning framework PyTorch; training the SiamTCP UAV tracking model on multiple video datasets and updating the network parameters through a loss function. The described UAV target tracking method includes the following steps: Step 1, initialize the backbone network. Step 2, build a UAV target tracking model based on target-aware cascaded pixel matching. Step 3, train the UAV target tracking model. The construction of the UAV target tracking model based on target-aware cascaded pixel matching in Step 2 includes: building a SiamTCP target tracking model using the deep learning framework PyTorch. The SiamTCP target tracking model consists of an improved backbone network GoogLeNet, pixel-level correlation, a target-aware module, a cascaded pixel matching module, and classification and regression sub-networks with a fully convolutional neural network structure. The input of the SiamTCP target tracking model adopts a dual-branch structure, namely the template branch and the search region branch. The two branches are fed into a CNN for feature extraction. The cascaded pixel matching module for target perception includes a target perception module and a cascaded pixel-level correlation module for feature fusion. The subsequent classification and regression sub-network is used for target localization. The template features are decomposed into many spatial kernels to achieve high-quality feature representation, and the size of the spatial kernel is 1×1; the template feature T t is divided along the width and height to obtain n t small kernels; n t = w t × h t ; Among them, w t and h t represent the width and height of the template feature; the template feature is divided into a group with a size of 1×1×C, where C is the number of channels of the template feature; During the part-to-part matching process, w s and h s are represented as the width and height of the search area feature S f , and is represented as the j-th row and the i-th column in S f ; Among them, is the similarity between t T and the nth position in the spatial dimension; The target perception module generates a variable template through the supervision of the marked bounding box B t and projects B t onto the template feature Z t to obtain the region of interest R t and obtain a new template feature T t which is expressed as: The processed template feature T t is decomposed into a number of kernels of size 1×1×C, where C is the number of channels of the template feature T t and the obtained spatial kernels are used to calculate the per-pixel correlation; The cascaded pixel matching module cascades the feature fusion between the template region and the search region multiple times, S f represents the search feature map after feature extraction, T t represents the template feature map after feature extraction, and the template feature map is obtained after being processed by the target perception module; the SiamTCP model uses the search feature map S f and the template feature map T t to perform pixel-level cross-correlation operation to obtain F1, and then splice F1 obtained by the pixel-level cross-correlation operation with the search feature map S f to obtain F2; by continuously performing pixel-level cross-correlation operations on the newly obtained features and the template features; splicing the features obtained by all pixel-level correlation operations to obtain new features; the SiamTCP model repeats the operation three times in total.

2. The drone target tracking method according to claim 1, characterized in that, The backbone network in Step 1 uses an improved GoogLeNet; pre-train the backbone network on the ImageNet image classification dataset to initialize the weights of the backbone network.

3. The drone target tracking method according to claim 1, wherein, The training of the UAV target tracking model in Step 3 includes: using stochastic gradient descent to train for 20 iterations on five datasets, namely the constructed model COCO, ImageNet DET, ImageNet VID, YouTube-BB, and GOT-10k, and setting the training batch size to 76; fixing the parameters of the first 10 iterations in the improved GoogLeNet for training the head network of the model; in the last 10 iterations, fixing the parameters of the first and second stages in the backbone network and fine-tuning the parameters of the third and fourth stages; for the first 5 iterations, set a warm-up learning rate that linearly increases from 0.005 to 0.01; for the last 15 iterations, use a learning rate that exponentially decays to 0.0005; during the entire training process, use 127 pixels as the template patch and 287 pixels as the search area.

4. A drone target tracking system applying the drone target tracking method according to any one of claims 1 to 3, characterized in that, The described UAV target tracking system includes: testing and deploying the SiamTCP network model for target tracking tasks in UAV scenarios.

5. A computer device, characterized in that, The described computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the UAV target tracking method described in any one of claims 1 to 3.

6. A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the UAV target tracking method described in any one of claims 1 to 3.

7. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the UAV target tracking system described in claim 4.