A UAV target tracking method, system, device and medium based on UAVTransT network
By introducing frequency domain and spatial information enhancement modules and spatiotemporal feature fusion modules into the drone target tracking network, combined with the mixed loss function, the problems of low accuracy and high computational complexity of drone target tracking under complex scenes and lighting changes are solved, and the target tracking effect with high accuracy and high efficiency is achieved.
Patent Information
- Application Number
- CN202411301936.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-18
AI Technical Summary
The existing drone target tracking algorithm has low accuracy and high computational complexity under complex scenarios and lighting changes, making it difficult to apply to the drone platform.
The UAVTransT network is designed to track the UAV target of tracking method, by introducing the frequency domain and spatial information enhancement module FSIE into the feature extraction subnet, and introducing the spatiotemporal feature fusion module STFF into the feature fusion subnet, combining L1 loss, CIoU loss and regression loss as the loss function of the network.
It improves the target tracking accuracy of the drone under the light changes, enhances the target tracking accuracy in background interference and occlusion scenarios, and reduces the computational complexity, and improves the target tracking efficiency of the model.
Smart Images

Figure CN119229322B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and further relates to deep learning and digital image processing technology, and is specifically a method, system, device and medium for tracking unmanned aerial vehicle targets based on a UAVTransT network, wherein the unmanned aerial vehicle target tracking network UAVTransT (UAV Transformer Tracking) involved is improved on the basis of the existing unmanned aerial vehicle target tracking network SwinTransformer. The present invention can be used for automatic tracking of targets in unmanned aerial vehicle optical images. Background Art
[0002] UAV target tracking technology is an important research direction in the field of UAVs and is widely used in civil and military fields, such as search and reconnaissance, environmental monitoring, disaster relief and other tasks. With the rapid development of computer vision technology, vision-based target tracking has become an intuitive tracking method that is close to human behavior. This technology achieves continuous tracking of the target by extracting target features from video sequences and predicting the position and size of the target in future frames. The challenges faced by UAV target tracking include the complexity of the scene environment, the variability of the target, the poor distinguishability of the target from the background, scale changes, camera shake and perspective changes. These factors increase the difficulty of target feature extraction and model building, and pose challenges to tracking performance.
[0003] In terms of target tracking algorithms, they are mainly divided into generative and discriminative methods. Generative methods focus on the target itself and ignore background information, while discriminative methods improve tracking accuracy and speed by extracting more useful information. Tracking algorithms based on correlation filtering and tracking algorithms based on deep learning are currently hot topics. They adapt to changes in the target through online learning and updating models to achieve real-time tracking effects.
[0004] Nanjing University of Aeronautics and Astronautics disclosed a UAV target tracking method based on improved YoLov7 and DeepSort in the patent document applied for (application number: CN202211428837.7, application publication number: CN115984319A). The implementation steps of this method are: constructing a target data set; optimizing the Yo Lov7 detector, and training the optimized YoLov7 detector using the target data set; building a DeepSort tracker, and training the DeepSort tracker using the target data set; using the target data set as an input item, using the trained YoLov7 detector and DeepSort tracker to track the target and obtain the target position; calculating the miss amount according to the target position, adjusting the UAV's posture according to the miss amount, and tracking the target. However, the invention still has shortcomings in that the UAV target tracking accuracy of the invention is still low in complex scenes, and the computational complexity of the invention is still high, making it difficult to apply to UAV platforms.
[0005] The patent application document with publication number CN117456196A discloses a method and system for tracking drone targets based on transformer multi-layer feature fusion. The implementation steps of the method are: using the Mobilenetv3 network to extract features from the target area in the search image and the template image, and obtaining feature maps of three different scales respectively; using the multi-layer feature fusion method to fuse the feature maps of three different scales, and obtain a fused search feature map and a fused template feature map; using a transformer-based feature matching method to deeply fuse the fused search feature map and the fused template feature map, and obtain a response map; inputting the response map into the multi-head attention module to obtain the target bounding box. However, the invention still has the disadvantage that the drone target tracking accuracy of the invention is still low in scenes with changing lighting.
[0006] In summary, the prior art has the following defects and deficiencies:
[0007] 1. In complex natural scenes, such as forests, cities, and oceans, the colors and textures of the target and the background are often very similar, making it difficult for existing drone target tracking algorithms to accurately separate and track the target.
[0008] 2. Under different lighting conditions (such as day, night, shadow, etc.), the appearance characteristics of the target will change significantly, and existing methods will suffer from missed tracking and incorrect tracking.
[0009] 3. UAV platforms usually have limited computing power and find it difficult to support complex deep learning models or other computationally intensive algorithms, resulting in a trade-off between real-time performance and accuracy of target tracking. Summary of the invention
[0010] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method, system, device and medium for unmanned aerial vehicle target tracking based on the UAVTransT network. A frequency domain and spatial information enhancement module FSIE is designed in the feature extraction subnetwork part. The module captures and combines the frequency and spatial characteristics of the target, improves the accuracy of target feature representation and positioning, and can improve the target tracking accuracy of the unmanned aerial vehicle under the condition of changing illumination; at the same time, a spatiotemporal feature fusion module STFF is designed in the feature fusion subnetwork part. The module uses three multi-head sub-attention mechanism layers to integrate the spatial and temporal state information of the target. The design effectively models the tracking state and improves the target tracking accuracy in challenging scenes with background interference and occlusion; in addition, by mixing L1 loss, CIoU loss and regression loss as the loss function of the network, the convergence of the model can be accelerated and the target tracking efficiency of the model can be improved; the present invention solves the problems of low tracking accuracy of the unmanned aerial vehicle target tracking method under complex background conditions, and missed tracking and incorrect tracking of targets under changing illumination.
[0011] In order to achieve the above object, the technical solution adopted by the present invention is:
[0012] A UAV target tracking method based on UAVTransT network includes the following steps:
[0013] Step 1: Build a UAV target tracking network based on UAVTransT, including feature extraction subnetwork, feature fusion subnetwork, and classification and regression subnetwork;
[0014] Step 2: Use the training set in the public data set of drones to train the drone target tracking network based on UAVTransT constructed in step 1. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public data set of drones, and the training weight file with the highest accuracy is selected as the optimal training weight file;
[0015] Step 3: Use the test set in the UAV public dataset and the optimal training weight file obtained in step 2 to track the UAV target tracking network based on UAVTransT constructed in step 1 and obtain the target tracking result.
[0016] The UAV target tracking network based on UAVTransT in step 1 includes a feature extraction subnetwork, a feature fusion subnetwork, and a classification and regression subnetwork;
[0017] The feature extraction subnetwork includes a Backbone1 network, a Backbone2 network, and two frequency domain and spatial information enhancement modules FSIE; wherein the Backbone1 network and the Backbone2 network both include a Swin Transformer stage1 layer, a Swin Transformer stage2 layer, and a Swin Transformer stage3 layer, the Backbone1 network and the Backbone2 network together form a Backbone network, and the Backbone1 network and the Backbone2 network share weights with each other; the frequency domain and spatial information enhancement module FSIE includes a fast Fourier transform layer FFT, a graph convolution network layer GCN, a multi-head sub-attention mechanism layer, two feedforward network layers FFN, two sum normalization layers, a concatenated convolution layer, and a global average pooling layer, and the frequency domain and spatial information enhancement module FSIE is expressed as follows:
[0018] F'=Norm(F4+FFN(F4))
[0019]
[0020]
[0021] F2=Conv(Cat(Norm(F1+SF),FF))
[0022] F1=MHSA((SF) Q ,(SF+FF) K ,(SF+FF) V )
[0023] SF=GCN(F)
[0024] FF=FFT(F)
[0025] Among them, GCN stands for graph convolutional network; FFT stands for fast Fourier transform; MHSA stands for multi-head self-attention mechanism; Norm stands for normalization operation; Cat stands for concatenation operation; Conv stands for convolution operation; FFN stands for feedforward network; GAP stands for global average pooling operation; Indicates channel multiplication; represents element addition; F' represents the output feature map; F represents the input feature map;
[0026] The feature fusion subnetwork includes a spatiotemporal feature fusion module STFF and a summation layer; the spatiotemporal feature fusion module STFF includes three multi-head self-attention mechanism layers, two feedforward network layers FFN and five summation normalization layers;
[0027] The classification and regression subnetwork includes a classification network and a regression network; the loss function of the classification and regression subnetwork adopts a mixed loss function of L1 loss, CIoU loss and regression loss, which is expressed as follows:
[0028] L=L cls +λ L1 L L1 +λ CIoU L CIoU
[0029]
[0030]
[0031]
[0032]
[0033] L cls =(1-P t ) γ ·log(P t )
[0034] Among them, L L1 represents L1 loss; L CIoU represents CIoU loss; L cls represents the regression loss; L1 and λ CIoU is the balance coefficient; P t represents the probability of the network predicting the positive category; γ represents the equalization coefficient; P represents the probability of the true category; n represents the number of samples; w gt and h gt The width and height of the label box; w pr and h pr The width and height of the predicted box, IoU means intersection over union; ρ is the distance between the center points of the rectangular box, and c is the diagonal length of the outer rectangular box.
[0035] In step 2, set the training rounds to ≥ 200, the batch size batch_size ≥ 16, and the learning rate ≤ 10 -5 , loss threshold ≤ 0.001, correlation coefficient conf-thres ≤ 0.5, intersection-over-union coefficient iou-thres ≤ 0.5.
[0036] In the step 3, the batch size batch_size ≥ 8, the correlation coefficient conf-thres ≤ 0.5, and the intersection-and-union ratio coefficient iou-thres ≤ 0.5.
[0037] The present invention also provides a UAV target tracking system based on the UAVTransT network, comprising:
[0038] Network construction module: used to build a UAV target tracking network based on UAVTransT, including feature extraction subnetwork, feature fusion subnetwork, and classification and regression subnetwork;
[0039] Network training module: used to train the UAV target tracking network built based on UAVTransT using the training set in the public UAV data set. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public UAV data set, and the training weight file with the highest accuracy is selected as the optimal training weight file;
[0040] Target tracking module: It is used to track the target of the UAV target tracking network based on UAVTransT using the test set and optimal training weight file in the public UAV data set to obtain the target tracking result.
[0041] The present invention also provides a UAV target tracking device based on the UAVTransT network, comprising:
[0042] Memory: a computer program storing the above-mentioned UAV target tracking method based on the UAVTransT network, which is a computer-readable device;
[0043] Processor: used to implement the UAVTransT network-based UAV target tracking method when executing the computer program.
[0044] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the UAVTransT network-based unmanned aerial vehicle target tracking method.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. The present invention designs two frequency domain and spatial information enhancement modules FSIE in the feature extraction subnetwork part. The modules capture and combine the frequency and spatial characteristics of the target. The frequency domain and spatial information enhancement module FSIE improves the accuracy of target feature representation and positioning, and can improve the target tracking accuracy of UAVs under changing lighting conditions.
[0047] 2. The present invention designs a spatiotemporal feature fusion module (STFF) in the feature fusion subnetwork part. This module uses three multi-head self-attention mechanism layers to integrate the spatial and temporal state information of the target. This design effectively models the tracking state and improves the target tracking accuracy in challenging scenes with background interference and occlusion.
[0048] 3. The present invention can accelerate the convergence of the model and improve the target tracking efficiency of the model by mixing L1 loss, CIoU loss and regression loss as the loss function of the UAV target tracking network based on UAVT ransT.
[0049] In summary, the present invention improves the UAV target tracking network Swin Transformer by introducing the frequency domain and spatial information enhancement module, the spatiotemporal feature fusion module STFF and the hybrid loss function method into the UAV target tracking network Swin Transformer, constructs a UAV target tracking network based on UAVTransT, and trains and predicts the UAV target tracking network based on UAVTransT by using a public UAV data set, which can effectively improve the UAV target tracking accuracy, and has the advantages of strong adaptability and high target tracking accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of the principle flow of an embodiment of the present invention.
[0051] Figure 2 It is an improved network structure diagram of an embodiment of the present invention.
[0052] Figure 3 It is a structural diagram of a frequency domain and spatial information enhancement module FSIE according to an embodiment of the present invention.
[0053] Figure 4 It is a structural diagram of the spatiotemporal feature fusion module STFF of an embodiment of the present invention.
[0054] Figure 5 This is a comparison chart of the drone target tracking accuracy between the present invention and the existing method.
[0055] Figure 6 This is a comparison chart of the UAV target tracking efficiency between the present invention and the existing method. DETAILED DESCRIPTION
[0056] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings.
[0057] like Figure 1 As shown, a UAV target tracking method based on the UAVTransT network includes the following steps:
[0058] like Figure 2 As shown, step 1: construct a UAV target tracking network based on UAVTransT consisting of a feature extraction subnetwork, a feature fusion subnetwork, and a classification and regression subnetwork;
[0059] Step 101: the feature extraction subnetwork includes Backbone1 network, Backbone2 network, and two frequency domain and spatial information enhancement modules FSIE; wherein Backbone1 network and Backbone2 network both include Swin Transformerstage1 layer, Swin Transformer stage2 layer and Swin Transformer stage3 layer, Backbone1 network and Backbone2 network together constitute Backbone network, and Backbone1 network and Backbone2 network share weights with each other;
[0060] like Figure 3 As shown in the figure, the frequency domain and spatial information enhancement module FSIE includes a fast Fourier transform layer FFT, a graph convolution network layer GCN, a multi-head attention mechanism layer, two feedforward network layers FFN, two sum normalization layers, a concatenated convolution layer, and a global average pooling layer. The frequency domain and spatial information enhancement module FSIE is used to combine the frequency and spatial features of the target, which can improve the accuracy of target feature representation and positioning, and thus improve the target tracking accuracy of the drone under changing lighting conditions. The frequency domain and spatial information enhancement module FSIE is expressed as follows:
[0061] F'=Norm(F4+FFN(F4))
[0062]
[0063]
[0064] F2=Conv(Cat(Norm(F1+SF),FF))
[0065] F1=MHSA((SF) Q ,(SF+FF) K ,(SF+FF) V )
[0066] SF=GCN(F)
[0067] FF=FFT(F)
[0068] Among them, GCN stands for graph convolutional network; FFT stands for fast Fourier transform; MHSA stands for multi-head self-attention mechanism; Norm stands for normalization operation; Cat stands for concatenation operation; Conv stands for convolution operation; FFN stands for feedforward network; GAP stands for global average pooling operation; Indicates channel multiplication; represents element addition; F' represents the output feature map; F represents the input feature map.
[0069] Step 102: The feature fusion subnetwork includes a spatiotemporal feature fusion module STFF and a summation layer; the spatiotemporal feature fusion module STFF includes three multi-head self-attention mechanism layers, two feedforward network layers FFN and five summation normalization layers;
[0070] like Figure 4 As shown in Figure 1, the spatiotemporal feature fusion module STFF includes three multi-head sub-attention mechanism layers, two feedforward network layers FFN, and five sum normalization layers. The spatiotemporal feature fusion module STFF uses multiple multi-head self-attention mechanisms to integrate the spatial and temporal state information of the target, effectively modeling the tracking state and improving the target tracking accuracy in challenging scenes with background interference and occlusion.
[0071] Step 103: the classification and regression sub-network includes a classification network and a regression network;
[0072] The mixed loss function of L1 loss, CIoU loss and regression loss is used as the loss function of the classification and regression sub-networks to evaluate the convergence of the network. The loss function is expressed as follows:
[0073] L=L cls +λ L1 L L1 +λ CIoU L CIoU
[0074]
[0075]
[0076]
[0077]
[0078] L cls =(1-P t ) γ ·log(P t )
[0079] Among them, L L1 represents L1 loss; L CIoU represents CIoU loss; L cls represents the regression loss; L1 and λ CIoU is the balance coefficient. t represents the probability of the network predicting the positive category; γ represents the equalization coefficient; P represents the probability of the true category; n represents the number of samples; w gt and h gt The width and height of the label box; w pr and hpr The width and height of the prediction box, IoU stands for intersection over union; ρ is the distance between the center points of the rectangular box, and c is the diagonal length of the outer rectangular box. When performing gradient descent optimization, L1 loss will not cause gradient explosion problems due to large errors, making the gradient of the model more stable during training. CIoU loss combines the effects of IoU, center point distance, and scale. During training, it can guide the bounding box to converge to the optimal position and size more quickly.
[0080] Step 2: Set the training rounds to ≥ 200, batch size batch_size ≥ 16, and learning rate ≤ 10 -5 , loss threshold ≤ 0.001, correlation coefficient conf-thres ≤ 0.5, intersection-over-union coefficient iou-thres ≤ 0.5, use the training set in the public data set of drones to train the drone target tracking network based on UAVTransT constructed in step 1. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public data set of drones, and the training weight file with the highest accuracy is selected as the optimal training weight file;
[0081] Step 3: Batch size batch_size ≥ 8, correlation coefficient conf-thres ≤ 0.5, intersection-over-intersection coefficient iou-thres ≤ 0.5, use the test set D in the public drone dataset test The optimal training weight file UAVTransT.pt obtained in step 2 is used to track the UAV target tracking network based on UAVTransT constructed in step 1 to obtain the target tracking result.
[0082] like Figure 2 As shown, the present invention also provides a UAV target tracking system based on the UAVTransT network, comprising:
[0083] Network construction module: used to implement the construction of the UAV target tracking network based on UAVTransT in step 1, including feature extraction subnetwork, feature fusion subnetwork, and classification and regression subnetwork;
[0084] Network training module: used to implement the training set in the public data set of drones in step 2 to train the drone target tracking network based on UAVTransT built in step 1. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public data set of drones, and the training weight file with the highest accuracy is selected as the optimal training weight file;
[0085] Target tracking module: It is used to implement the target tracking of the UAV target tracking network based on UAVTransT constructed in step 1 using the test set in the UAV public data set and the optimal training weight file obtained in step 2 in step 3 to obtain the target tracking result.
[0086] The present invention also provides a UAV target tracking device based on the UAVTransT network, comprising:
[0087] Memory: a computer program storing the above-mentioned UAV target tracking method based on the UAVTransT network, which is a computer-readable device;
[0088] Processor: used to implement the UAVTransT network-based UAV target tracking method when executing the computer program.
[0089] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the UAVTransT network-based unmanned aerial vehicle target tracking method.
[0090] The effect of the present invention is further described below in conjunction with simulation experiments:
[0091] 1. Simulation Experiment Conditions
[0092] The hardware platform of the simulation experiment of the present invention is: the processor is Intel i9-13900K, the main frequency is 3.0GHz, and the running memory is 64G.
[0093] The software platform of the simulation experiment platform of the present invention is: Windows 11 operating system and PyCharm, PyTorch2.1.0, and CUDA 12.1.
[0094] 2. Simulation steps
[0095] The training set in the public data set of drones is input into the drone target tracking network based on UAVTransT for optimization training. The training process is: the input image will be extracted through the feature extraction subnetwork to obtain feature maps of different scales, and then these feature maps will be classified and regressed. The regression results will be reconstructed to obtain more refined feature maps. On this basis, classification and regression operations are performed again, and the loss is calculated to complete the drone target tracking based on the present invention. All target tracking uses single-scale training, the image input size is 512×512 pixels, and the number of iterations epoch is set to 200.
[0096] 3. Simulation content and results analysis
[0097] The simulation experiment of the present invention is to perform target tracking processing on three UAV images, and the results are as follows: Figure 5 shown.
[0098] Combine the following Figure 5 The simulation effect of the present invention is further described.
[0099] Figure 5 This is a comparison chart of the target tracking accuracy of drones of the present invention and the existing method. The comparison algorithm is SwinTrack (LinL, Fan H, Zhang Z, et al. Swintrack: A simple and strong baseline for transformer tracking [J]. Advances in Neural Information Processing Systems, 2022, 35: 16743-16754.), the test image size is 512×512, and the accuracy evaluation indicators are target tracking success rate (Suc) and target tracking accuracy (Pre). The target tracking success rate (Suc) is usually used to measure the performance of the tracking algorithm in the entire video sequence or multiple video sequences. The larger its value, the better the target tracking effect; the target tracking accuracy (Pre) is used to measure the accuracy of the algorithm in locating the target during the tracking process. The larger its value, the better the target tracking effect.
[0100] like Figure 5 As shown, it can be seen that the target tracking success rate (Suc) and target tracking accuracy (Pre) of the present invention are higher than SwinTrack, which are 65.6 and 80.5 respectively. The experimental results show that the frequency domain and spatial information enhancement module FSIE designed in the present invention significantly improves the accuracy of target feature representation and positioning, thereby improving the target tracking accuracy of the UAV under conditions of changing illumination. At the same time, the spatiotemporal feature fusion module STFF can effectively model the tracking state and enhance the target tracking accuracy in challenging scenarios such as background interference and occlusion. In general, the UAV target tracking method based on the UAVTransT network proposed in the present invention significantly improves the accuracy of target tracking and meets the requirements of high-precision tracking.
[0101] Figure 6This is a comparison chart of the drone target tracking efficiency of the present invention and the existing method. The comparison algorithm is SwinTrack, the test image size is 512×512, and the efficiency evaluation indicators are the model parameter quantity (Params) and frame rate (FPS). The parameter quantity (Params) is one of the important indicators to measure the complexity of a model, which directly affects the computing requirements, memory usage and reasoning speed of the model. The smaller the value, the lower the model complexity; the frame rate (FPS) is used to evaluate the real-time performance and efficiency of the algorithm. The larger the value, the better the target tracking efficiency.
[0102] like Figure 6 As shown, it can be seen that the frame rate (FPS) of the present invention is 57.3, which is higher than SwinTrack; the parameter amount (Params) is 2.3M, which is less than SwinTrack. The experimental results show that the present invention significantly accelerates the convergence speed of the model and improves the efficiency of target tracking by mixing L1 loss, CIoU loss and regression loss as the loss function of the UAV target tracking network based on UAVTransT. Therefore, the UAV target tracking method based on the UAVTransT network proposed in the present invention can meet the real-time requirements of the UAV platform.
Claims
1. A UAV target tracking method based on UAVTransT network, characterized in that: The following steps are involved: Step 1: Build a UAV target tracking network based on UAVTransT, including feature extraction subnetwork, feature fusion subnetwork, and classification and regression subnetwork; The feature extraction subnetwork includes a Backbone1 network, a Backbone2 network, and two frequency domain and spatial information enhancement modules FSIE; wherein the Backbone1 network and the Backbone2 network both include a Swin Transformer stage1 layer, a Swin Transformer stage2 layer, and a Swin Transformer stage3 layer, the Backbone1 network and the Backbone2 network together form a Backbone network, and the Backbone1 network and the Backbone2 network share weights with each other; the frequency domain and spatial information enhancement module FSIE includes a fast Fourier transform layer FFT, a graph convolution network layer GCN, a multi-head sub-attention mechanism layer, two feedforward network layers FFN, two sum normalization layers, a concatenated convolution layer, and a global average pooling layer, and the frequency domain and spatial information enhancement module FSIE is expressed as follows: F'=Norm(F4+FFN(F4)) F2=Conv(Cat(Norm(F1+SF),FF)) F1=MHSA((SF) Q ,(SF+FF) K ,(SF+FF) V ) SF=GCN(F) FF=FFT(F) Among them, GCN stands for graph convolutional network; FFT stands for fast Fourier transform; MHSA stands for multi-head self-attention mechanism; Norm stands for normalization operation; Cat stands for concatenation operation; Conv stands for convolution operation; FFN stands for feedforward network; GAP stands for global average pooling operation; Indicates channel multiplication; represents element addition; F' represents the output feature map; F represents the input feature map; The feature fusion subnetwork includes a spatiotemporal feature fusion module STFF and a summation layer; the spatiotemporal feature fusion module STFF includes three multi-head self-attention mechanism layers, two feedforward network layers FFN and five summation normalization layers; The classification and regression subnetwork includes a classification network and a regression network; the loss function of the classification and regression subnetwork adopts a mixed loss function of L1 loss, CIoU loss and regression loss, which is expressed as follows: in, represents L1 loss; represents CIoU loss; represents the regression loss; L1 and λ CIoU is the balance coefficient; P t represents the probability of the network predicting the positive category; γ represents the equalization coefficient; P represents the probability of the true category; n represents the number of samples; w gt and h gt The width and height of the label box; w pr and h pr The width and height of the prediction box, IoU means intersection over union; ρ is the distance between the center points of the rectangular box, and c is the diagonal length of the outer rectangular box; Step 2: Use the training set in the public data set of drones to train the drone target tracking network based on UAVTransT constructed in step 1. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public data set of drones, and the training weight file with the highest accuracy is selected as the optimal training weight file; Step 3: Use the test set in the UAV public dataset and the optimal training weight file obtained in step 2 to track the UAV target tracking network based on UAVTransT constructed in step 1 and obtain the target tracking result.
2. The method for tracking a target by using a UAVTransT network according to claim 1, characterized in that: In step 2, set the training rounds to ≥ 200, the batch size batch_size ≥ 16, and the learning rate ≤ 10 -5 , loss threshold ≤ 0.001, correlation coefficient conf-thres ≤ 0.5, intersection-over-union coefficient iou-thres ≤ 0.
5.
3. The method for tracking a target by using a UAVTransT network according to claim 1, characterized in that: In the step 3, the batch size batch_size ≥ 8, the correlation coefficient conf-thres ≤ 0.5, and the intersection-and-union ratio coefficient iou-thres ≤ 0.
5.
4. A UAV target tracking system based on a UAVTransT network based on the method according to any one of claims 1 to 3, characterized in that: include: Network construction module: used to build a UAV target tracking network based on UAVTransT, including feature extraction subnetwork, feature fusion subnetwork, and classification and regression subnetwork; Network training module: used to train the UAV target tracking network built based on UAVTransT using the training set in the public UAV data set. After each round of training, a training weight file is obtained; the training weight file is verified by the verification set in the public UAV data set, and the training weight file with the highest accuracy is selected as the optimal training weight file; Target tracking module: It is used to track the target of the UAV target tracking network based on UAVTransT using the test set in the public UAV data set and the optimal training weight file obtained in step 2 to obtain the target tracking result.
5. A UAV target tracking device based on UAVTransT network, characterized in that: include: Memory: a computer program storing a method for tracking a target by a UAV based on a UAVTransT network as described in any one of claims 1 to 3, which is a computer-readable device; Processor: used to implement the UAV target tracking method based on the UAVTransT network as described in any one of claims 1-3 when executing the computer program.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a UAV target tracking method based on a UAVTransT network as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Unmanned aerial vehicle target tracking method based on improved YoLov7 and DeepSort
CN115984319A
Transform multilayer feature fusion-based unmanned aerial vehicle target tracking method and system
CN117456196A
Vehicle-mounted liquid crystal screen light guide plate defect visual detection method based on target detection network
CN113421230A
Improved unmanned aerial vehicle aerial image dense and small target identification method, system and device based on YOLOv5s and medium
CN117456389A