Attached raindrop perception unmanned aerial vehicle target tracking method based on state monitoring Transfomer architecture

Through the attached raindrop perception method based on the status monitoring Transformer architecture, combined with the Transformer visual feature encoding backbone network and local raindrop removal module, the image degradation problem caused by raindrop interference in rainy days is solved, and high-precision and robust real-time tracking is achieved, which improves the tracking performance of the drone in rainy days.

CN120495343APending Publication Date: 2025-08-15HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510590263.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The image degradation caused by the lens's attachment of raindrops in rainy scenes and the existing rain removal technology are slow to process, making it difficult to achieve high-precision and robust real-time tracking.

Method used

The attached raindrop perception method based on the status monitoring Transformer architecture is adopted, combined with the Transformer visual feature encoding backbone network, CornerHead parameter regression network, real-time status monitoring module and local raindrop removal module, the target interference strength is evaluated through the peak side lobe ratio and confidence peak, and the start and stop of the local raindrop removal module is dynamically controlled to achieve high robust real-time tracking.

Benefits of technology

High robust real-time tracking is achieved in harsh rainy environments, which significantly improves the tracking accuracy and success rate of the drone, reduces information loss and computing overhead during rain removal, and achieves real-time processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495343A_ABST
    Figure CN120495343A_ABST
Patent Text Reader

Abstract

The invention relates to an attached raindrop perception unmanned aerial vehicle target tracking method based on a state monitoring Transform architecture, which integrates a deep learning image rain removal technology into a Transform backbone network so as to improve the quality of target coding features, realizes closed-loop feedback and result correction of a whole attached raindrop perception target tracking framework by using a real-time state monitoring module, and improves the target tracking accuracy. The information loss and the operation overhead in the rain removal process are reduced to the maximum extent; a peak sidelobe ratio and a confidence peak value are introduced as core evaluation indexes of a target state, a raindrop sensing strategy dynamically controls starting and stopping of a local raindrop removal module according to the dynamic state, dynamic changes of disturbed states of a target are coped with according to situations, high-robustness real-time tracking of the unmanned aerial vehicle in the view angle of attaching raindrops is achieved, real-time processing performance is kept, and meanwhile, the real-time performance of the unmanned aerial vehicle is improved. According to the method, the tracking accuracy and success rate of the unmanned aerial vehicle in rainy days are remarkably improved, and reliable technical support is provided for practical application of the unmanned aerial vehicle under all-weather conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) video target tracking, and in particular to a method for tracking UAV targets using attached raindrops and sensing based on a state monitoring Transformer framework. Background Art

[0002] Drones (UAVs) have attracted widespread attention in recent years due to their exceptional bird's-eye view, convenient remote control, and superior maneuverability. UAV video target tracking technology, particularly in rainy conditions, plays an important role in both civilian and military applications, such as sports analysis, flood rescue, rainy-weather reconnaissance, and all-weather combat. However, when UAVs perform tracking missions in rainy scenes, the imaging lens is often subject to strong interference from falling rain, wind-blown mist, or raindrops adhering to the tracking target during flight through humid air. This can easily cause trackers with limited perception range to frequently lose the correct target, leading to tracking drift or even tracking failure. Therefore, how to implement a real-time tracking algorithm that combines high precision and strong robustness in rainy conditions has become a key technical bottleneck in promoting the practical application of UAVs in all weather conditions.

[0003] Conventional drone tracking algorithms are designed based on ideal weather conditions, and their performance degrades significantly when operating in rainy scenarios. This is because in dynamic weather conditions like rain, the scene images captured by drone imaging equipment are severely distorted, and attached raindrops easily obscure the tracked target. This strong interference factor can partially or completely obscure the designated target. This is especially true for drone visual tracking of small or micro targets, which is more likely to cause key information loss and global feature distribution confusion, severely limiting the resolution capabilities of conventional trackers in rainy scenarios. Unlike traditional drone video sequences, video sequences in rainy scenarios are not only subject to the interference of attached raindrops, but also face the dual challenges of large target scale variations and high background complexity, further exacerbating the difficulty of drone visual target tracking in the rain.

[0004] Although traditional physical methods (such as wipers, vibrations, airflow, etc.) can remove raindrops from the lens, they will not only add additional load to the drone, but the hardware rain wiping process will inevitably introduce additional interference sources. With the rapid development of deep learning and the improvement of computer computing power, algorithm-based rain-free target reconstruction has become a feasible solution for anti-interference feature extraction. However, a derivative problem is the time cost of the combined image processing and visual tracking framework. Despite the current achievements in raindrop removal technology, the processing speed of the most advanced rain removal algorithm is still slow, and the frame-by-frame tracking process itself is relatively time-consuming. Therefore, how to ensure accurate and effective raindrop removal while achieving real-time tracking of 25 frames per second or more is another difficult problem that needs to be overcome.

[0005] The technical problem that technicians in this field urgently need to solve is: how to achieve real-time tracking tasks with high precision and robustness when using drones in rainy scenes, facing two major challenges: image degradation caused by the attachment of raindrops of different densities and scales, and the slow processing speed of existing deraining technologies that can be used for tracking. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the present invention provides a method for tracking UAV targets with attached raindrop perception based on a state-monitoring Transformer architecture.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a method for tracking drone targets with attached raindrops based on a state-monitoring Transformer architecture, comprising the following steps:

[0008] S1, based on the first frame input image I1 of the video sequence and the true value bounding box (x0, y0, w0, h0), the basic template z1 is cut out, based on the subsequent frame input image I i Crop out the target search area x i ,i∈[2,n];

[0009] S2, template image z1 and search image x i The data is fed into the offline trained Transformer visual feature encoding backbone network, and feature extraction and fusion are performed simultaneously to obtain the initial attention map.

[0010] S3. Decouple the attention map and reconstruct it into a 256×760×1 search area feature map A, which is then fed into the offline trained CornerHead parameter regression network to obtain the target confidence score. Position offset and discrete error O i and S i The bounding box regression result of the target is formed, but the final output result is controlled by the real-time status monitoring module;

[0011] S4. Confidence score map The target state estimator first calculates and saves the peak sidelobe ratio (PSR) of the current frame. i and confidence peak F i According to the attached raindrop sensing strategy, if the target is in weak interference, the initial O containing the target position information i and S i As the tracking result output; conversely, the real-time status monitoring module will emit a feedback signal source;

[0012] S5, after receiving the feedback signal, trigger z1 and x iThe image is sent to the offline trained local raindrop removal module for clarity processing to obtain a clear rain-free template image z1. free and the target search image x i free ;

[0013] S6. Get a clear rain-free template image z1 free and the target search image x i free As input, repeat steps S2 and S3 to obtain the new target confidence score Position offset and discrete error

[0014] S7, the new confidence score map Input the real-time state monitoring module, then the target state estimator is activated again to calculate and store the peak sidelobe ratio of the current frame and confidence peak F i free ; According to the attached raindrop perception strategy, if the target is in a state of strong raindrop interference, there will be no rain and Parameter regression is used as the current frame tracking result; otherwise, the initial O i and S i The parameters are the tracking results, and the local raindrop removal module and target state estimator are turned off for the next 10 frames.

[0015] Furthermore, the appearance Transformer visual feature encoding backbone network in step S2 includes a data preprocessor and an iterative encoder layer, and the specific implementation steps are as follows:

[0016] S2.1、Data preprocessor will be rainless template and search area Split into non-overlapping patches of consistent resolution, including N z Template patches x p and N x Search patches z p After flattening and linearization mapping operations, the generated C-dimensional patch embedding is positionally memorized, specifically by adding two learnable one-dimensional position embeddings to generate the final template word embedding and search region word embeddings Finally, concatenate these two tokens into N z +N x Single-stream sequence of length;

[0017] S2.2. The concatenated sequence is fed into the iterative encoder layer in the form of a single information stream. Each layer updates the input label through the Multi-Head Attention Network (MHA) and the Multi-Layer Perceptron (MLP) feedforward network. The internal operation can be expressed as: Where [;] represents the splicing operation, Q i , K i and V i to represent the query, key, and value fed to the multi-head attention block in layer i, respectively, which can be obtained by the formula: Here W is the trained model parameter, and the three-dimensional attention map of the target is obtained after iterating the encoder 12 times.

[0018] The appearance Transformer localization backbone is trained using the COCO, LaSOT, and TrackingNet datasets. The sizes of the template image and the search area image are set to 128×128×3 and 256×256×3 respectively, the batch size is 16, and the training cycle is 300 rounds. The model is trained offline using the AdamW optimizer with a weight decay rate of 10 -4 The initial learning rates of the backbone network and other parts are 4×10 -5 and 4×10 -4 ; Note that after training to the 240th round, the learning rate will decay to 1 / 10 of the original.

[0019] Furthermore, the CornerHead parameter regression network in step S3 is a fully convolutional network composed of several stacked convolutional layers. The four key parameters representing the tracking results are calculated using the following formula: (x, y, w, h) = (x d +O(0,x d ,y d ),y d +O(1,x d ,y d ),S(0,x d ,y d ),S(1,x d ,y d )); where P, O, and S are the target’s confidence score, position offset, and discrete error, respectively. (x d ,y d ) represents the coordinates of the highest value in the confidence score map P.

[0020] The fully convolutional network uses weighted focal loss for classification, l1 loss and IOU loss for regression during the training phase; the total loss function L is: L = L cls +L reg =Lcls +α1L iou +α2L1, where L represents the total loss, L cls represents the classification loss, L reg represents the regression loss, L iou Denotes the intersection-over-union loss, L1 denotes the l1 loss, and the specific calculation of the branch loss is the sum of the l1 loss and the IOU loss. The constants α1 and α2 control the weight ratio of the classification branch and the bounding box regression branch loss. We set them to 2 and 5 respectively based on experience.

[0021] Furthermore, the peak-to-sidelobe ratio in step S4 is an indicator that measures the ratio of the signal peak to its side area, which can indicate the reliability of the target area in the response graph. The specific mathematical calculation principle is: In the formula μ i ,σ i They represent the peak value, sidelobe mean, and standard deviation in the confidence score map of the i-th frame respectively.

[0022] When the target experiences weak interference, the response graph is sharp and the PSR value is high. When the target experiences strong interference, the response graph is flat and blunt and the PSR value is small. Given that the PSR value can effectively quantify the degree of interference on the target, we choose it as the key indicator for PSR perception of the target tracking status in the current frame. The raindrop perception strategy relies on the peak sidelobe ratio as a measurement indicator to judge strong and weak interference. The specific distinction process is as follows: λ is the threshold constant.

[0023] Furthermore, the local raindrop removal module described in step S5 uses the improved AttentiveGAN as the core structure of the raindrop removal module, which consists of two sub-networks: the generative sub-network and the discriminative sub-network. The present invention adopts a dual optimization strategy of the autoencoder-discriminator (A+D) joint architecture and local image processing to significantly improve the feature recovery effect and computational efficiency in the anti-interference tracking task. The specific training steps are as follows:

[0024] S5.1. Template image z1 and search area x i It is input into the generative network, which uses a contextual autoencoder with jump connections as the main architecture; the designed deep autoencoder contains 16 conversion blocks, which generate images that are as realistic and rain-free as possible. and The training of the autoencoder is based on the multi-scale loss L M and perceptual loss L P Two loss functions;

[0025] To obtain context information at different scales, L M Extract features from different decoder layers to form outputs of different sizes. The specific calculation formula is: Among them S i represents the i-th output extracted from the decoder layer, T i Indicates that S i have true values of the same scale; are weight factors at different scales.

[0026] L P It is used to measure the global difference between the output features of the autoencoder and the corresponding real image features. These features are extracted from the trained VGG16 convolutional neural network. The specific estimation formula is: L P =L MSE (VGG(O), VGG(T)). O=G(I) is the output image of the generative network, T is the true image without raindrops; L G =10 -2 log(1-D(O))+L M +L P It is the calculation formula of the total loss function of the final generated network, where G and D represent the generation network and the identification network respectively.

[0027] S5.2. The main task of the discriminator network is to measure the difference between the image generated by the generator network and the real image. Its specific structure includes 7 convolutional layers with a convolution kernel size of 3×3 and a fully connected layer with 1024 neurons. The neurons use the Sigmoid activation function to determine the authenticity of the generated image. The training loss function of the discriminator is L D Defined as: L D = -log(D(T))-log(1-D(O)).

[0028] The entire local raindrop removal module is specially trained using the improved RainDrop dataset, and the initial learning rate of the network is set to 5×10 -4 , the batch size is 16, and 600 rounds of training are performed to achieve the optimal effect.

[0029] Furthermore, when performing secondary raindrop sensing in step S7, the target often experiences continuous strong interference due to large-scale occlusion by raindrops or other environmental interference, making it particularly difficult to distinguish between the two types of interference. To achieve specific attached raindrop sensing, the confidence peak F is introduced as another key indicator of the raindrop sensing strategy, and the change in the indicator before and after rain removal in the current frame is used to assist in the state prediction of subsequent dynamic frames. The differentiation formula is: where ΔPSR i and ΔF i They represent the difference between the peak sidelobe ratio and the confidence peak before and after rain removal in the frame.

[0030] If the real-time state monitoring module senses that the target is disturbed by raindrops, it promptly retracts the target based on the feature map of the rain-free search area. Conversely, since the entry and exit of the target is usually a continuous process, the original search area attention map is used for bounding box regression, and the target state estimator and local rain remover are disabled within the next 10 frames.

[0031] Beneficial effects:

[0032] To address the problem of high-robust real-time tracking of drones in rainy scenes due to raindrops adhering to the lens, this paper designs an attached raindrop-aware drone video target tracking method based on a state-monitoring Transformer architecture from the perspective of algorithm design. Deep learning image deraining technology is integrated into the Transformer backbone network to improve the quality of target coding features. A real-time state monitoring module is used to achieve closed-loop feedback and result correction of the entire attached raindrop-aware target tracking framework, thereby minimizing information loss and computational overhead in the deraining process.

[0033] The present invention introduces peak-to-sidelobe ratio and confidence peak as core evaluation indicators of the target state. The raindrop sensing strategy dynamically controls the start and stop of the local raindrop removal module based on the strength of the perceived target interference and the type of strong interference, and responds to the dynamic changes of the target interference state according to different situations. It realizes highly robust real-time tracking of drones from the perspective of attached raindrops, effectively solving the key problem of video tracking in harsh rainy environments.

[0034] The present invention uses the appearance Transformer visual feature encoding backbone network and removes the candidate elimination module to increase tracking speed; adopts the improved AttentiveGAN as the core structure of the raindrop removal module, adopts the autoencoder discriminator (A+D) joint architecture, removes part of the structure in the original network and retrains it, while retaining good rain removal efficiency, ultimately achieving a lightweight effect and enabling the algorithm to run in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the overall structure of the tracking algorithm network in the present invention;

[0036] Figure 2 This is a schematic diagram of the Transformer backbone structure in the present invention;

[0037] Figure 3 This is a schematic diagram of the principle of the real-time status monitoring module in the present invention;

[0038] Figure 4 This is a schematic diagram of the principle of the local raindrop removal module in the present invention;

[0039] Figure 5This is a comparison chart of the accuracy, success rate, and attribute radar of the proposed method (ARATrack) and other mainstream methods on the RainDropUAV123 test benchmark, which simulates a raindrop tracking scenario attached to a drone.

[0040] Figure 6 This is a comparison chart of the accuracy, success rate, and attribute radar of the proposed method (ARATrack) and other methods on the RainDropUAVTrack112 test benchmark for simulated drone-attached raindrop tracking scenarios;

[0041] Figure 7 This is a comparison chart of the accuracy, success rate, and attribute radar of the proposed method (ARATrack) and other mainstream methods on the RainDropUAV20L long-term test benchmark for simulating drone-attached raindrop tracking scenarios;

[0042] Figure 8 This is a comparison chart of the accuracy, success rate, and attribute radar of the proposed method (ARATrack) and other mainstream methods on the RainDropDTB70 short-term test benchmark that simulates drone-attached raindrop tracking scenarios;

[0043] Figure 9 This is a speed demonstration experiment on four simulated drone-attached raindrop tracking scene test benchmarks. The method of the present invention (ARATrack) meets the real-time requirements (Running Speed>25FPS). DETAILED DESCRIPTION

[0044] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0045] like Figure 1 As shown, an Adherent Raindrop-Aware Visual Object Tracking Based on StatusMonitoring Transformer Architecture in UAV Videos (ARATrack) method is described, which specifically includes the following steps S1 to S7.

[0046] S1, based on the first frame input image I1 of the video sequence and the true value bounding box (x0, y0, w0, h0), cut out the basic template z1, based on the subsequent frame input image I i Crop out the target search area x i ,i∈[2,n];

[0047] S2, such as Figure 2As shown, the template image and the search area image are fed into the offline trained Transformer visual feature encoding backbone network, and feature extraction and fusion are performed simultaneously to obtain the search area feature response map A. The specific implementation method is as follows;

[0048] S2.1, the data preprocessor converts the target template and search area Split into non-overlapping patches of consistent resolution, including N z Template patches x p and N x Search patches z p After flattening and linearization mapping operations, the generated C-dimensional patch embedding is positionally memorized, specifically by adding two learnable one-dimensional position embeddings to generate the final template word embedding and search region word embeddings Finally, concatenate these two tokens into N z +N x Single-stream sequence of length;

[0049] S2.2. The concatenated sequence is fed into the iterative encoder layer in the form of a single information stream. Each layer updates the input label through the Multi-Head Attention Network (MHA) and the Multi-Layer Perceptron (MLP) feedforward network. The internal operation can be expressed as: Where [;] represents the splicing operation, Q i , K i and V i to represent the query, key, and value fed to the multi-head attention block in layer i, respectively, which can be obtained by the formula: Here W is the trained model parameter, and the three-dimensional attention map of the target is obtained after iterating the encoder 12 times.

[0050] In this paper, the appearance Transformer localization backbone uses the COCO, LaSOT, and TrackingNet training datasets. The template image and search area image sizes are set to 128×128×3 and 256×256×3 respectively, the batch size is 16, and the training cycle is 300 rounds. The model is trained offline using the AdamW optimizer with a model decay rate of 10 -4 The initial learning rates of the backbone and other parts are set to 4×10 -5 and 4×10 -4 It should be noted that after 240 rounds of training, the learning rate is reduced by a factor of 10.

[0051] S3. Decouple the attention map and reconstruct it into a 256×760×1 search area feature map A, which is then fed into the offline trained CornerHead parameter regression network to obtain the target confidence score. Position offset and discrete error O i and S i The parameter regression results of the target are composed, but the output of the position result, that is, the four parameters (x, y, w, h) of the rectangular box, are controlled by the real-time status monitoring module; (x, y, w, h) represent the horizontal and vertical coordinates of the predicted target in the frame image and the width and height of the rectangular box respectively.

[0052] In step S3, the CornerHead parameter regression network is a fully convolutional network composed of several stacked convolutional layers. The four key parameters (x, y, w, h) representing the tracking results are calculated by the following formula: (x, y, w, h) = (x d +O(0,x d ,y d ),y d +O(1,x d ,y d ),S(0,x d ,y d ),S(1,x d ,y d )); P, O, S represent the confidence score, position offset and discrete error of the target respectively, (x d ,y d ) represents the coordinates of the highest value in the confidence score map P.

[0053] The fully convolutional network uses weighted focal loss for classification during the training phase and L1 loss and IOU loss for regression. The total loss function L is: L = L cls +L reg =L cls +α1L iou +α2L1, where L represents the total loss, L cls represents the classification loss, L reg represents the regression loss, L iou Denotes the intersection-over-union loss, L1 denotes the l1 loss, and the specific calculation of the branch loss is the sum of the l1 loss and the IOU loss. The constants α1 and α2 control the weight ratio of the classification branch and the bounding box regression branch loss. We set them to 2 and 5 respectively based on experience.

[0054] S4, such as Figure 3 As shown, the confidence score map of the frame The target state estimator first calculates and saves the peak sidelobe ratio (PSR) of the current frame.i and confidence peak F i According to the attached raindrop sensing strategy, if the target is in weak interference, the initial O containing the target position information i and S i It is output as the tracking result; conversely, the real-time status monitoring module will emit a feedback signal source.

[0055] The peak-to-sidelobe ratio (PSR) is an indicator that measures the ratio of the signal peak to its side area, and can indicate the reliability of the target area in the response graph. Its calculation formula is: In the formula μ i ,σ i They represent the peak value, sidelobe mean, and standard deviation in the confidence score map of the i-th frame respectively. When the target is weakly interfered, the response graph is sharp and the PSR value is high. When the target is strongly interfered, the response graph is flat and blunt and the PSR value is small. The raindrop sensing strategy relies on the peak-to-sidelobe ratio as a measurement indicator to judge strong and weak interference. The specific distinction process is as follows: λ is a set threshold constant, and in this embodiment, λ=2.40.

[0056] S5, after receiving the feedback signal, trigger z1 and x i The image is sent to the offline trained local raindrop removal module for clarity processing to obtain a clear rain-free template image z1. free and the target search image x i free .

[0057] like Figure 4 As shown in Figure 2, the local raindrop removal module uses an improved AttentiveGAN as its core architecture, consisting of two sub-networks: a generative network and a discriminative network. This paper employs a dual optimization strategy: a combined autoencoder-discriminator (A+D) architecture and local image processing. This significantly improves feature recovery and computational efficiency in interference-free tracking tasks. The specific training process is as follows.

[0058] Template image z1 and search area x i It is input into the generative network, which uses a contextual autoencoder with jump connections as the main architecture. The deep autoencoder contains 16 transformation blocks, which generate images that are as realistic and rain-free as possible. and The training of the autoencoder is based on the multi-scale loss L M and perceptual loss L P Two-term loss function L P .

[0059] To obtain context information at different scales, L MExtract features from different decoder layers to form outputs of different sizes. The specific calculation formula is: Among them S i represents the i-th output extracted from the decoder layer, T i Indicates that S i The true value of the same scale, L MSE represents the mean square error loss between Si and Ti; are weight factors at different scales.

[0060] L P It is used to measure the global difference between the output features of the autoencoder and the corresponding real image features. These features are extracted from the trained VGG16 convolutional neural network. The specific estimation formula is: L P =L MSE (VGG(O), VGG(T)); O = G(I) is the output image of the generative network, T is the true value image without raindrops; L G =10 -2 log(1-D(O))+L M +L P It is the calculation formula of the total loss function of the final generated network, where G and D represent the generation network and the identification network respectively.

[0061] The main task of the discriminator network is to evaluate the difference between the image generated by the generator network and the real image. Its specific structure includes 7 convolutional layers with a convolution kernel size of 3×3 and a fully connected layer with 1024 neurons. The neurons use the Sigmoid activation function to determine the authenticity of the generated image. The training loss function of the discriminator is L D For: L D = -log(D(T))-log(1-D(O)).

[0062] The entire local raindrop removal module is specially trained using the improved RainDrop dataset, and the initial learning rate of the network is set to 5×10 -4 , the batch size is 16, and 600 rounds of training are performed to achieve the optimal effect.

[0063] As one of the most balanced raindrop removal algorithms currently available, AttentiveGAN cannot be applied to tracking algorithms due to its poor real-time performance. Therefore, we have made improvements based on this. First, we adaptively clarify the local search area based on the center value of the prediction result. We also remove structures such as the recurrent attention network from the original network and retrain it. While maintaining good raindrop removal efficiency, we ultimately achieve a lightweight effect and enable the algorithm to run in real time.

[0064] S6, the obtained clear rain-free template z1 free and search area xi free As input, repeat steps S2 and S3 to obtain the interference-free target confidence map Position offset and discrete error

[0065] S7, the new confidence score map Input the real-time state monitoring module, then the target state estimator is activated again to calculate and store the peak sidelobe ratio of the current frame and confidence peak F i free Since strong interference includes large-area occlusion by raindrops or other environmental interference, the attached raindrop sensing strategy further distinguishes the categories of strong interference using the following distinction formula: where ΔPSR i and ΔF i They represent the changes in the peak sidelobe ratio and confidence peak before and after raindrop removal in the i-th frame, respectively, and F i 、F free They are the confidence peaks before and after raindrop removal, η1, η2, η3, η4 represent the corresponding thresholds respectively; in order to strictly avoid discrimination deviation during training, these four thresholds η i are set to 0, 0, 0.4, and 0.4 respectively.

[0066] According to the attached raindrop perception strategy, if the target is in a state of strong raindrop interference, there will be no rain. and Parameter regression is used as the tracking result of the current frame; otherwise, if the target is in other strong interference states, the initial O i and S i The parameters are the tracking results, and the local raindrop removal module and target state estimator are turned off for the next 10 frames.

[0067] In other words, if the real-time state monitoring module senses that the target is strongly interfered by raindrops, it will promptly retract the target based on the feature map of the rain-free search area; conversely, if the target is in other strong interference states, since the entry and exit of the target is usually a continuous process, the original search area attention map is used for bounding box regression, and the target state estimator and local raindrop removal module are disabled in the next 10 frames to speed up the algorithm.

[0068] S1-S3 above represent the feature extraction and fusion process, S4 represents the result parameter regression process, S5 represents the local raindrop removal process, and S6-S7 represents the real-time status monitoring process. Together, these steps provide closed-loop feedback and result correction for the entire raindrop-aware target tracking framework, minimizing information loss and computational overhead during rain removal. In the actual tracking process, steps S1-S7 are repeated until the entire video sequence is tracked.

[0069] The present invention is equipped with This paper was implemented on a computer platform based on a Silver 4110 2.1GHz CPU and dual NVIDIA GeForce RTX 2080 GPUs, developed using the PyTorch framework and Python. The comparison algorithm was replicated based on the original paper using open-source code, and performance testing and evaluation were conducted on the same hardware configuration.

[0070] The implementation effect of the present invention is verified by quantitative analysis experiments. The test experiments are carried out on four synthetic drone-attached raindrop tracking test sets with high target diversity, namely RainDropUAV123 with 10 challenge attributes, RainDropUAVTrack112 with 11 challenge attributes, RainDropUAV20L with 10 challenge attributes, and RainDropDTB70 with 10 challenge attributes.

[0071] The performance of the proposed method ARATrack is compared with mainstream drone tracking algorithms, including PVT++ based on MobileNet, SiamPW-BRO based on ResNet50, and GRM and SMAT based on the Transformer architecture. The overall and attribute tracking performance is evaluated through comprehensive tests.

[0072] 1. SMAT, see literature[1]Gopal GY,Amer M A.Separable self and mixedattention transformers for efficient object tracking[C] / / Proceedings of the IEEE / CVF winter conference on applications of computer vision.2024:6708-6717.

[0073] 2. PVT++, see literature [2] Li B, Huang Z, Ye J, et al. Pvt++: a simple end-to-endlatency-aware visual tracking framework [C] / / Proceedings of the IEEE / CVFinternationalconference on computer vision.2023:10006-10016.

[0074] 3. GRM, see literature [3] Gao S, Zhou C, Zhang J. Generalized relation modeling for transformer tracking [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023: 18686-18695.

[0075] 4. SiamPW-BRO, see literature [4] Tang F, Ling Q. Ranking-based Siamese visual tracking [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:8741-8750.

[0076] Figures 5 to 8 Comparisons of the accuracy (left), success rate (center), and attribute radar (right) of our proposed method, ARATrack, with other tracking methods, using simulation experiments on four simulated drone-attached raindrop benchmarks. The results show that our proposed method significantly outperforms other compared methods in overall tracking accuracy, success rate, and attribute classification performance across various rainy scenarios.

[0077] Figure 9 This is a speed demonstration experiment on four simulated drone-attached raindrop tracking scenario test benchmarks. The tracking speed of the method (ARATrack) of the present invention meets the real-time requirement (Running Speed>25FPS).

[0078] This demonstrates that the method described in this paper achieves highly robust real-time tracking of drones from the perspective of attached raindrops, effectively resolving the key challenge of video tracking in harsh rainy environments. While maintaining real-time processing performance, this method significantly improves the accuracy and success rate of drone tracking in rainy conditions, providing reliable technical support for practical drone applications in all weather conditions.

[0079] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture, characterized by: The steps include: S1, based on the first frame input image I1 of the video sequence and the true value bounding box, cut out the basic template z1, based on the subsequent frame input image I i Crop out the target search area x i ,i∈[2,n]; S2, z1 and x i The data is fed into the offline trained Transformer visual feature encoding backbone network, and feature extraction and fusion are performed simultaneously to obtain the initial attention map. S3, decouple the initial attention map and reconstruct it into the search area feature map, and then feed it to the offline trained CornerHead parameter regression network to obtain the target confidence score P i , position offset O i and the discrete error S i ; S4, the confidence score map P of the current frame i The target state estimator first calculates and saves the peak sidelobe ratio (PSR) of the current frame. i and confidence peak F i , the attached raindrop sensing strategy senses the interference strength. If the target is in weak interference, the initial O i and S i As the tracking result output; conversely, the real-time status monitoring module will emit a feedback signal source; S5, after receiving the feedback signal, trigger z1 and x i The image is sent to the offline trained local raindrop removal module for clarity processing to obtain a clear rain-free template image z1. free and the target search image x i free ; S6, z1 obtained by the clearing process free and x i free As input, repeat steps S2 and S3 to obtain a new target confidence score P i free , position offset and discrete error S7, P i free Input the real-time state monitoring module, then the target state estimator is activated again to calculate and store the peak sidelobe ratio of the current frame and confidence peak F i free , the attached raindrop perception strategy determines the category of strong interference. If the target is in a state of strong raindrop interference, and As the tracking result of the current frame; Otherwise, the initial O will be output i and S i To obtain the tracking results, the local raindrop removal module and target state estimator are turned off for the next 10 frames.

2. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 1 is characterized in that: The appearance Transformer visual feature encoding backbone network in step S2 consists of two parts: the data preprocessor and the iterative encoder layer. The specific steps of feature extraction and fusion are as follows: S2.

1. The data preprocessor performs patch segmentation and embedding mapping, and transforms the target template z1 and the search area x i Segment into non-overlapping patches of consistent resolution, consisting of N z Template patches x p and N x Search patches z p After flattening and linearization mapping operations, the generated C-dimensional patch embedding is positionally memorized and two learnable one-dimensional position embeddings are added to generate the final template word embedding E z ∈R Nz×C and search region word embedding E x ∈R Nx×C , and finally E z and E x Splice to N z +N x Single-stream sequence of length; S2.

2. Feature encoding: Use the offline trained Transformer encoder to perform hierarchical feature extraction on the embedded sequences of the template and search area, and perform feature interaction and fusion. The spliced sequence is fed into the iterative encoder layer in the form of a single information stream. Each layer updates the input label through a multi-head attention network and a multi-layer perceptron feedforward network. After 12 layer operations of the iterative encoder, a three-dimensional attention map of the target is obtained.

3. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 1 is characterized in that: In step S3, the CornerHead parameter regression network is a fully convolutional network composed of several stacked convolutional layers. The target rectangle parameters are calculated by the following formula: (x, y, w, h) = (x d +O(0,x d ,y d ),y d +O(1,x d ,y d ),S(0,x d ,y d ),S(1,x d ,y d )); Among them, (x, y, w, h) represent the horizontal and vertical coordinates of the predicted target in the current frame image and the width and height of the rectangular frame respectively. d ,y d ) represents the coordinates of the highest value in the confidence score map.

4. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 3 is characterized in that: The fully convolutional network uses weighted focal loss for classification during the training phase and L1 loss and intersection-over-union loss for regression. The total loss function L is: L = L cls +L reg =L cls +α1L iou +α2L1; Where L cls represents the classification loss, L reg represents the regression loss, L iou represents the intersection-over-union loss, L1 represents the l1 loss, and α1 and α2 are constants representing the weights.

5. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 1 is characterized in that: In step S4, the specific process of the attached raindrop sensing strategy sensing strong and weak interference is as follows: if PSR i ≥λ, it is strong interference, otherwise it is weak interference, and λ is the set threshold constant.

6. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 1 is characterized in that: In step S7, since strong interference includes large-area occlusion of raindrops or other environmental interference, the distinction formula for attached raindrop perception is: where ΔPSR i and ΔF i They represent the changes in the peak sidelobe ratio and confidence peak before and after raindrop removal in the i-th frame, respectively, and F i 、F free are the confidence peaks before and after raindrop removal, and η1, η2, η3, and η4 represent the corresponding thresholds.

7. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 5 or 6, characterized in that: Peak-to-sidelobe ratio In the formula μ i ,σ i They represent the peak value, sidelobe mean, and standard deviation in the confidence score map of the i-th frame respectively.

8. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 1 is characterized in that: In step S5, the local raindrop removal module uses the improved AttentiveGAN as the core structure of the raindrop removal module. It consists of two sub-networks, generative and discriminative. It adopts a dual optimization strategy of autoencoder discriminator joint architecture and local image processing to significantly improve the feature recovery effect and computational efficiency in the anti-interference tracking task.

9. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 8 is characterized in that: Initial template image z1 and search area x i It is fed into the generative network, which generates an image that is as realistic as possible without raindrops. and The difference, the generation network adopts the contextual autoencoder with skip connection as the main architecture, and the training of the autoencoder is based on the multi-scale loss L M and perceptual loss L P Two loss functions, the loss function L of the generated network G =10 -2 log(1-D(O))+L M +L P , D(O) represents the probability that the identification network judges that O is from real data.

10. The method for tracking UAV targets with attached raindrops based on a state-monitoring Transformer architecture according to claim 8 is characterized in that: The discriminant network measures the difference between the images generated by the generative network and the real images during training. The discriminant network contains 7 convolutional layers with a convolution kernel size of 3×3 and a fully connected layer with 1024 neurons. The loss function of the discriminant network is L D Defined as: L D = -log(D(T))-log(1-D(O)); O = G(I) is the output image of the generative network, and T is the true image without raindrops.