Unmanned aerial vehicle photoelectric tracking method based on improved SiamDT

Through adaptive high and low frequency image enhancement and improved twin network model, combined with motion constraint dynamic screening module, the accuracy and stability problems of UAV optoelectronic tracking in complex environments are solved, and efficient target tracking is achieved.

CN120708107APending Publication Date: 2025-09-26CHENGDU KONGYU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510804955.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing UAV optoelectronic tracking methods lack tracking accuracy and stability when the target is moving rapidly, obscured, or in complex backgrounds, and existing twin networks fail to effectively utilize motion prior information.

Method used

An adaptive high- and low-frequency image enhancement algorithm is used to preprocess UAV images. Combined with the improved Siamese network model and motion constraint dynamic screening module, the tracking results are optimized through background noise suppression, similarity evaluation, and regional target prior probability prediction branches.

Benefits of technology

The accuracy and stability of UAV optoelectronic tracking have been significantly improved, as well as its robustness and real-time performance in complex environments. The tracking success rate has reached 66.64%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708107A_ABST
    Figure CN120708107A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle photoelectric tracking, in particular to an unmanned aerial vehicle photoelectric tracking method based on improved SiamDT, and the method comprises the steps: carrying out the preprocessing of an unmanned aerial vehicle image through a self-adaptive high and low frequency image enhancement algorithm (HLE), effectively enhancing the edge and texture features of a target, and suppressing the background noise; the processed image is input into an improved twin network model (SiamXC), the model works cooperatively through a background noise suppression branch, a similarity evaluation branch and a region target prior probability prediction branch, and network training is optimized in combination with a gradient balanced Xavier initialization strategy; and finally, performing multi-level confidence threshold screening and continuous frame verification on the candidate targets by using a motion constraint dynamic screening module (DFMC) to ensure the tracking accuracy and stability. The problems of insufficient tracking precision and poor stability under the conditions of rapid target movement, shielding and complex backgrounds in the prior art are obviously solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) photoelectric tracking, and in particular to a UAV photoelectric tracking method based on improved SiamDT. Background Art

[0002] Unmanned aerial vehicles (UAVs), as flexible and efficient mobile platforms, have been widely used in military reconnaissance, disaster relief, logistics, and other fields. However, tracking UAVs in complex environments still faces many challenges. Traditional UAV tracking technologies are mostly based on target detection and tracking algorithms, such as Kalman filters, particle filters, and vision-based tracking methods. These methods mostly focus on single visual information or target motion models and have poor adaptability to environmental changes, illumination variations, and target occlusion.

[0003] In recent years, Siamese networks have made significant progress in the field of target tracking. The Siamese network can extract the deep features of the target by sharing weights, and calculate the similarity between the current frame and the target template to perform template matching. However, the existing Siamese network is easily affected by noise in the case of fast motion, occlusion or complex background, resulting in a decrease in tracking accuracy. The Region Proposal Network (RPN) is mainly used in target detection to predict the location information of the target by generating candidate regions. The introduction of RPN makes the target positioning more accurate, but when it is used independently, it is prone to too many redundant region proposals, which increases the computational complexity, and it is difficult to update the target position in time when the UAV target moves quickly.

[0004] Therefore, the existing UAV optoelectronic tracking methods have the following problems: tracking accuracy is difficult to guarantee when the target moves rapidly or is occluded; tracking stability is poor under the influence of complex backgrounds and noise; the existing target tracking methods based on twin networks fail to utilize motion prior information, resulting in limited target tracking accuracy and robustness. Summary of the Invention

[0005] The purpose of the present invention is to provide a UAV optoelectronic tracking method based on improved SiamDT to solve the problem of insufficient tracking accuracy and stability of existing UAV tracking technology when the target moves rapidly, is blocked, and is affected by background noise in complex environments.

[0006] To achieve the above objectives, the present invention provides a UAV optoelectronic tracking method based on an improved SiamDT, comprising the following steps:

[0007] Adaptive high and low frequency image enhancement algorithm is used to pre-process the collected UAV images;

[0008] The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch;

[0009] Dynamically screen candidate targets output by the twin network, and optimize tracking results by combining motion constraints and multi-frame verification mechanisms;

[0010] Output the final target tracking information.

[0011] Among them, the adaptive high and low frequency image enhancement algorithm is used to pre-process the collected drone images. The specific steps include:

[0012] Perform Gaussian blur processing on the input image to extract the low-frequency information of the image;

[0013] Subtract each pixel value of the original image from the maximum value to generate a high-frequency information map;

[0014] Enhance the high-frequency information graph to highlight the edge and texture features of the image;

[0015] Suppress background and smooth areas in low-frequency information;

[0016] Dynamically adjust the brightness of the enhanced image according to preset brightness parameters to optimize the image contrast and visibility.

[0017] The candidate targets output by the twin network are dynamically screened, and the tracking results are optimized by combining motion constraints and a multi-frame verification mechanism. The specific steps include:

[0018] Calculate the Euclidean distance motion constraint of the center point of the candidate target box;

[0019] Set multi-level confidence thresholds for target screening, with a low threshold of 0.3 used for preliminary screening and a high threshold of 0.8 used for final confirmation;

[0020] The final tracking target is determined by stability verification of 3 consecutive frames.

[0021] The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch.

[0022] The twin network model uses the gradient-balanced Xavier initialization strategy, and the weight distribution satisfies:

[0023]

[0024] In the formula, fan in is the number of input nodes, fan out is the number of output nodes.

[0025] The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch.

[0026] The background noise suppression branch includes generating interference region candidate frames through the DS-RPN Head and performing feature matching and interference frame filtering using the VR-CNN Head.

[0027] The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch.

[0028] The similarity evaluation branch generates Top-N target candidate boxes by calculating the similarity of deep features between the template image and the search image.

[0029] The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch.

[0030] The regional target prior probability prediction branch extracts multi-scale features from the target template and predicts the probability of target existence in the candidate region.

[0031] The collected drone images are preprocessed using an adaptive high- and low-frequency image enhancement algorithm, and the steps include:

[0032] An infrared camera is used to collect image data;

[0033] After data collection, the DarkLabel annotation tool was used to manually annotate the drone targets.

[0034] This paper presents a method for UAV optoelectronic tracking based on an improved SiamDT. The method preprocesses UAV images using an adaptive high- and low-frequency image enhancement algorithm (HLE), effectively enhancing target edge and texture features and suppressing background noise. The processed images are then fed into an improved SiamXC network model. This model uses a background noise suppression branch, a similarity assessment branch, and a regional target prior probability prediction branch to work together, optimizing network training with a gradient-balanced Xavier initialization strategy. Finally, a motion-constrained dynamic screening module (DFMC) performs multi-level confidence threshold screening and continuous frame verification on candidate targets to ensure tracking accuracy and stability. This method significantly addresses the issues of insufficient tracking accuracy and poor stability associated with traditional technologies in the face of rapid target motion, occlusion, and complex backgrounds. By leveraging motion prior information and multi-level feature fusion, it improves the robustness and real-time performance of UAV tracking in complex environments. Experiments show that this method achieves a tracking success rate (SA) of 66.64% on the Anti-UAV500 dataset, outperforming existing mainstream algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.

[0036] Figure 1 It is a schematic diagram of the SiamXC network structure of the present invention.

[0037] Figure 2 It is a flow chart of the motion constraint dynamic screening algorithm of the present invention.

[0038] Figure 3 Schematic diagram of the image enhancement algorithm of the present invention.

[0039] Figure 4 It is a flowchart of the steps of the UAV photoelectric tracking method based on the improved SiamDT of the present invention. DETAILED DESCRIPTION

[0040] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be understood as limiting the present invention.

[0041] See also Figures 1 to 4 ,in, Figure 1 It is a schematic diagram of the SiamXC network structure of the present invention. Figure 2 It is a flow chart of the motion constraint dynamic screening algorithm of the present invention. Figure 3 Schematic diagram of the image enhancement algorithm of the present invention. Figure 4It is a flowchart of the steps of the UAV photoelectric tracking method based on the improved SiamDT of the present invention.

[0042] The present invention provides an unmanned aerial vehicle (UAV) photoelectric tracking method based on an improved SiamDT, comprising the following steps:

[0043] S101: Preprocessing the collected UAV images using an adaptive high- and low-frequency image enhancement algorithm;

[0044] S102: Input the preprocessed image into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch;

[0045] S103: Dynamically screen the candidate targets output by the twin network, and optimize the tracking results by combining motion constraints and multi-frame verification mechanism;

[0046] S104: Output the final target tracking information.

[0047] Specifically, it integrates adaptive high and low frequency image enhancement (HLE), dynamic filtering module with motion constraints (DFMC), and improves the initialization strategy of network weights. The improved method is named SiamXC. Its model structure is as follows Figure 1 shown.

[0048] The decision subnetwork of the SiamXC method consists of three specially designed branches, each performing a different function:

[0049] (1) Background noise suppression branch (right): This branch aims to improve target detection accuracy by suppressing background noise and enhance the clarity and recognizability of target areas in complex scenes;

[0050] Step 1: Input the search image of the current frame;

[0051] Step 2: Use Swin Transformer to extract deep semantic features of the interference image;

[0052] Step 3: Input DS-RPN Head to generate multiple candidate boxes (1K proposals);

[0053] Step 4: Determine whether there is overlap between the background candidate boxes and the true boundary, and retain up to 50 candidate boxes that have no overlap at all;

[0054] Step 5: The candidate box is input into the query convolution to generate query features and guide the aggregation of interference area features;

[0055] Step 6: Use VR-CNN Head to extract features and perform matching scores on all candidate regions;

[0056] Step 7: Input them into the DFMC module in the order of scores as interference boxes.

[0057] (2) Similarity evaluation branch (middle): This branch performs similarity evaluation on the detection image and the template to further improve the accuracy of target matching and ensure that the model can maintain high tracking accuracy under target morphological changes or environmental interference.

[0058] Step 1: Input the search image of the current frame;

[0059] Step 2: Use SwinTransformer to extract deep semantic features of the interference image;

[0060] Step 3: The template image and the search image are input into the query convolution to generate query features;

[0061] Step 4: Use VR-CNN Head to extract features and perform matching scores on all candidate regions to generate Top-N Proposals;

[0062] Step 5: Enter the scores into the DFMC module in order.

[0063] (3) Regional target prior probability prediction branch (left): This branch is used to predict the prior probability of the existence of a target in the candidate region, thereby accurately locating the target and ensuring accurate recognition when the target features are blurred or occluded.

[0064] Step 1: Input the target template cropped from the first frame image;

[0065] Step 2: Use the HLE enhancement module to enhance the brightness and edges of the template image to highlight the distinguishable features of the target area;

[0066] Step 3: Input the enhanced image into SwinTransformer to extract multi-scale hierarchical visual features;

[0067] Step 4: The feature map is input into the DS-RPN Head to generate a series of candidate boxes;

[0068] Step 5: The features of the candidate boxes are input into the VR-CNN Head to further extract local area features and perform matching scores;

[0069] Step 6: Sort all candidate regions by matching scores to obtain the Top-M Proposals and their corresponding scores;

[0070] Step 7: Match Top-M Proposals with GroundTruth and calculate IoU Loss for training optimization.

[0071] Through the collaboration of the three branches, SiamXC can improve detection and tracking accuracy and enhance network performance in complex scenarios and under conditions of rapidly changing targets. This innovative design provides a new approach to target tracking technology, with broad application prospects and technical advantages. To further improve the tracking performance of SiamXC, this paper optimizes and improves SiamXC from the following four aspects:

[0072] (1) Improved network weight initialization: gradient-balanced Xavier initialization strategy.

[0073] In order to solve the gradient vanishing / exploding problem during network training and improve training stability and convergence speed, the present invention adopts the Xavier weight initialization strategy (as shown in Formula 1) to dynamically adjust the network weight distribution and balance the gradient propagation between network layers, thereby improving the target tracking accuracy. Specifically, Xavier initialization is performed based on the number of input nodes (fan in ) and the number of output nodes (fan out ) to adjust the distribution of weights, thereby ensuring that the gradients of each layer remain balanced during propagation.

[0074]

[0075] In the formula, fan in is the number of input nodes, fan out is the number of output nodes.

[0076] The Xavier method is used to initialize the weights of the similarity learning module, which optimizes the network weight propagation mechanism and improves the network training efficiency and target tracking accuracy. This improvement has the following advantages:

[0077] A. Improved training stability: Xavier initialization dynamically adapts the input and output scale, balances signal propagation, and avoids gradient explosion or vanishing;

[0078] B. Performance Enhancement: In the process of drone detection in infrared images, the improved network converges faster, and both accuracy and robustness are improved;

[0079] C. Strong versatility: Xavier initialization can be extended to other deep learning models, providing new ideas for weight initialization design and having wide application value.

[0080] (2) Target screening: Dynamic Screening Module for Motion Constraints (DFMC).

[0081] The present invention proposes a dynamic screening algorithm based on motion constraints (DFMC) to improve the accuracy and robustness of target tracking in complex scenarios. The core idea of ​​the algorithm is to accurately calculate the dynamic distance between the center points of the target frames of the previous and next frames by introducing a regional association mechanism, thereby screening out the real target. In order to further reduce false detections and missed detections, the present invention designs a dynamic adaptive filtering strategy that integrates target confidence thresholds and motion constraints. This strategy sets multi-level confidence thresholds and dynamically adjusts the detection logic to ensure that the model can maintain high performance in different scenarios. The algorithm flow chart is as follows Figure 2 shown.

[0082] Step 1: Tracking flag determination: Detect all potential targets in the pre-processed image and output a list of candidate targets (including location, size, confidence, etc.). Check the tracking flag status:

[0083] If the tracking flag is True: Go to step 4 for low threshold screening and range filtering.

[0084] If the tracking flag is False: Skip the screening step and go directly to step 5 to confirm the target result.

[0085] Step 2: Low threshold screening and range filtering (only performed when the tracking flag is True):

[0086] Low confidence threshold screening: Use a lower confidence threshold (such as 0.3) to filter targets and retain more potential candidates (avoid missing targets).

[0087] Spatial range filtering: further filter targets based on preset position and size ranges (e.g., targets must be in the center of the image, or wider than 50 pixels, etc.).

[0088] Keep the screening results: Output a list of candidate targets that meet the conditions and proceed to the next step of judgment.

[0089] Step 3: Confirmation of target results:

[0090] Determine whether there is a target result in the current frame: If there is a target, clear the tracking result of the previous frame, force the tracking flag to False (to avoid interference from historical results), and proceed to step 6 (high threshold continuous frame verification). If there is no target: end the process directly.

[0091] Step 4: High Threshold Consecutive Frame Verification:

[0092] High confidence threshold screening: Use a high confidence threshold (such as 0.8) to strictly screen the current frame target.

[0093] Consecutive frame stability check: Checks whether the target meets the high threshold condition for N consecutive frames (e.g., 3 frames). If so, the tracking flag is set to True, indicating stable tracking. If not, the tracking flag remains False (to prevent false tracking due to transient noise).

[0094] Step 5: Process termination:

[0095] After all status updates and judgments are completed, the process ends and the final tracking results are output.

[0096] Specifically, the algorithm first calculates the center point of the detection box, as shown in the following formula (2):

[0097]

[0098] Among them, x min and x max is the left and right boundaries of the detection box, y min and y max In order to determine whether the target is within a reasonable range of motion, the algorithm further calculates the Euclidean distance between the center point of the current frame detection frame and the center point of the previous frame detection frame, as shown in formula (3):

[0099]

[0100] Among them, (x cur ,y cur ) is the center point of the current frame target, (x prev ,y prev ) is the center point of the previous frame. When the calculated Euclidean distance exceeds the preset threshold, the target is considered invalid, thereby effectively suppressing false detections and dynamically updating the target's tracking status.

[0101] Compared to traditional single-frame detection methods, the region association mechanism introduced in this invention maintains detection continuity even when the target is slightly obscured or subject to noise interference, significantly improving target tracking reliability. This continuous frame target verification mechanism effectively reduces false alarms caused by occasional noise, improving response speed and detection accuracy for new targets.

[0102] This invention successfully achieves a balance between lightweightness and precision by integrating single-frame detection with multi-frame tracking strategies: single-frame detection is used to capture new targets, while multi-frame tracking ensures target continuity. This significantly improves the accuracy and robustness of target detection while maintaining low computational complexity.

[0103] In summary, the dynamic motion constraint screening module (DFMC) of the present invention can effectively improve the stability and reliability of target detection and tracking in dynamic and complex scenarios, and has broad application prospects.

[0104] (3) Input image enhancement: adaptive high and low frequency image enhancement algorithm (HLE).

[0105] This paper proposes an adaptive high- and low-frequency image enhancement algorithm, systematically optimizing test data to improve tracking accuracy for small targets. The method's core technical solutions include high-frequency detail enhancement and low-frequency information compression, effectively highlighting the edges and textures of small targets while suppressing background noise. To address the challenges of low illumination and blurred target features in infrared images, the present invention introduces a brightness adjustment mechanism to improve overall image quality and enhance the model's ability to capture targets.

[0106] The specific implementation steps are as follows: First, use Gaussian blur to extract the low-frequency information of the image, and then generate a high-frequency information map by subtracting the color value of each pixel from the maximum value of 255. This operation is equivalent to inverting the brightness of the image, making the details of the image (such as edges and textures) more prominent. Next, the high-frequency information is enhanced to improve the clarity and detail of the image. Low-frequency information usually contains background or large-scale smooth areas and is usually suppressed to avoid redundant information affecting the fineness of the image. Finally, according to the preset brightness parameters, the enhanced image is dynamically adjusted to ensure that the image details are effectively visualized while the brightness and contrast conform to the cognitive habits of natural vision to avoid overexposure or too dark conditions.

[0107] This method performs well in low-contrast scenes and high dynamic range image processing, and can effectively overcome the problems of insufficient detail enhancement and uneven brightness optimization in traditional methods. Figure 3 As shown, the image change results of each step clearly demonstrate the detail enhancement effect. Comprehensive analysis shows that the image enhancement method of the present invention has significant advantages and achieves a balance between retaining small target features and natural brightness distribution. Compared with the existing technology, the image enhancement method provided by the present invention can significantly improve the detection accuracy of small targets in complex environments, enhance the robustness of image processing, and has strong application prospects. It is particularly suitable for scenarios requiring high-precision target detection.

[0108] (4) Sample expansion: Improve model tracking robustness.

[0109] To expand the Anti-UAV410 dataset to include small targets and urban building scenes, this paper uses a high-performance infrared camera for image data acquisition. To improve data quality and adaptability, a high-performance infrared camera with a resolution of 640×512 and a frame rate of 25Hz is used for image acquisition. The camera supports pan / tilt control and zoom adjustment, enabling flexible acquisition of infrared image data from multiple angles and distances.

[0110] After data collection, the DarkLabel annotation tool was used to accurately and manually annotate the drone targets to ensure data accuracy and practicality. Based on the original Anti-UAV410 dataset, this paper added 90 video clips, including 51 training set videos, 16 validation set videos, and 23 test set videos. Each video is approximately one minute long, generating a total of 107,854 target annotation boxes.

[0111] The expanded dataset is named Anti-UAV500. It significantly enhances the diversity and richness of the data in terms of scenes, target size, background complexity, etc., especially in the coverage of urban building backgrounds and small targets, thereby effectively enhancing the generalization ability and robustness of the target detection model.

[0112] Technical effects and advantages:

[0113] Object tracking experiments use Success Accuracy (SA) as an evaluation metric. SA represents the degree of overlap between the predicted object bounding box and the ground-truth bounding box. Intersection over Union (IoU) is often used as a metric, calculated by counting the proportion of frames with an IoU greater than a certain threshold to the total number of +--+96 frames.

[0114]

[0115] Among them B p represents the bounding box predicted by the algorithm, B g represents the ground-truth bounding box.

[0116] This paper compares and analyzes various target tracking algorithms to verify the effectiveness of the expanded Anti-UAV500 dataset in practical applications. The experiments were trained on the Anti-UAV410 and Anti-UAV500 datasets, and evaluated on the Anti-UAV500 test set.

[0117] Table 1 Anti-UAV500 dataset test comparison

[0118]

[0119]

[0120] (1) Algorithm performance comparison:

[0121] SiamXC: When trained on Anti-UAV410, SiamXC achieves a significantly improved SA value compared to SiamDT, demonstrating its potential advantages in target tracking tasks.

[0122] Performance Improvement After Dataset Expansion: When trained on the Anti-UAV500 dataset, all algorithms saw performance improvements. SiamXC maintained its leading position, further demonstrating its superiority in complex scenes and diverse target tracking tasks.

[0123] (2) Algorithm Advantage Analysis:

[0124] Generalization and Robustness: The SiamXC algorithm's performance on the Anti-UAV500 dataset demonstrates significant advantages in generalization and robustness. In particular, SiamXC demonstrates exceptional performance in tracking complex scenes (such as urban building backgrounds) and diverse targets (such as small objects).

[0125] Adaptability to complex scenes: The expansion of the Anti-UAV500 dataset significantly enhances the algorithm's adaptability to complex backgrounds, and SiamXC can effectively address these challenges, providing new technical ideas for target tracking research.

[0126] Practical engineering application significance:

[0127] In real-world environments, drone countermeasure systems face numerous challenges, including complex electromagnetic interference, drastically changing environmental conditions, rapid drone flight, and complex background interference. These factors significantly impact the effectiveness and accuracy of traditional countermeasure systems. To address these challenges, the proposed SiamXC model significantly improves the detection and tracking accuracy of drone tracking systems in complex environments by combining advanced adaptive high- and low-frequency image enhancement (HLE) technology with a dynamic motion constraint filtering module (DFMC).

[0128] Specifically, adaptive high- and low-frequency image enhancement (HLE) technology automatically optimizes image quality and enhances image detail in complex lighting conditions such as low light, rain, snow, and strong reflections, ensuring that the countermeasure system can effectively identify and locate drones even with a restricted field of view. The dynamic screening module with motion constraints (DFMC) considers the drone's motion characteristics and dynamically screens areas with irregular or abnormal motion, effectively suppressing background noise and interference to ensure the system focuses on the target drone, thereby improving target tracking accuracy and stability.

[0129] Therefore, the present invention not only enhances the adaptability of optoelectronic equipment to drone countermeasure systems in complex environments, but also provides more stable and reliable technical guarantees for practical applications, effectively improving the overall performance and safety of drone defense systems.

[0130] The above disclosure is merely one or more preferred embodiments of the present application and is not intended to limit the scope of the present application. A person skilled in the art will understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present application are still within the scope of the present application.

Claims

1. A UAV optoelectronic tracking method based on improved SiamDT, characterized in that: The following steps are involved: Adaptive high and low frequency image enhancement algorithm is used to pre-process the collected UAV images; The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch; Dynamically screen candidate targets output by the twin network, and optimize tracking results by combining motion constraints and multi-frame verification mechanisms; Output the final target tracking information.

2. The UAV optoelectronic tracking method based on improved SiamDT according to claim 1, characterized in that: The collected UAV images are preprocessed using an adaptive high- and low-frequency image enhancement algorithm. The specific steps include: Perform Gaussian blur processing on the input image to extract the low-frequency information of the image; Subtract each pixel value of the original image from the maximum value to generate a high-frequency information map; Enhance the high-frequency information graph to highlight the edge and texture features of the image; Suppress background and smooth areas in low-frequency information; Dynamically adjust the brightness of the enhanced image according to preset brightness parameters to optimize the image contrast and visibility.

3. The UAV optoelectronic tracking method based on improved SiamDT according to claim 2, characterized in that: Dynamically screen candidate targets output by the twin network, and optimize tracking results by combining motion constraints and multi-frame verification mechanisms. The specific steps include: Calculate the Euclidean distance motion constraint of the center point of the candidate target box; Set multi-level confidence thresholds for target screening, with a low threshold of 0.3 used for preliminary screening and a high threshold of 0.8 used for final confirmation; The final tracking target is determined by stability verification of 3 consecutive frames.

4. The UAV optoelectronic tracking method based on improved SiamDT according to claim 2, characterized in that: The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch. The twin network model uses the gradient-balanced Xavier initialization strategy, and the weight distribution satisfies: In the formula, fan in is the number of input nodes, fan out is the number of output nodes.

5. The UAV optoelectronic tracking method based on improved SiamDT according to claim 4, characterized in that: The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch. The background noise suppression branch includes generating interference region candidate frames through the DS-RPN Head and performing feature matching and interference frame filtering using the VR-CNN Head.

6. The UAV optoelectronic tracking method based on improved SiamDT according to claim 5, characterized in that: The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch. The similarity evaluation branch generates Top-N target candidate boxes by calculating the similarity of deep features between the template image and the search image.

7. The UAV optoelectronic tracking method based on improved SiamDT according to claim 6, characterized in that: The preprocessed image is input into the improved Siamese network model, which includes a background noise suppression branch, a similarity evaluation branch, and a regional target prior probability prediction branch. The regional target prior probability prediction branch extracts multi-scale features from the target template and predicts the probability of target existence in the candidate region.

8. The UAV optoelectronic tracking method based on improved SiamDT according to claim 1, characterized in that: The collected UAV images are preprocessed using an adaptive high- and low-frequency image enhancement algorithm, and the steps include: An infrared camera is used to collect image data; After data collection, the DarkLabel annotation tool was used to manually annotate the drone targets.