Low-small-slow unmanned aerial vehicle target tracking method based on multi-scale background dynamic query fusion perception detection under space-time constraint
By using a multi-scale background dynamic query fusion perception and detection method, the problems of environmental noise sensitivity, incomplete target area extraction, and poor multi-scale adaptability in the tracking of "low, small, and slow" UAVs are solved, achieving high-precision and robust real-time tracking, which is suitable for high-risk scenarios such as low-altitude security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU UNIV
- Filing Date
- 2025-12-02
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are sensitive to environmental noise, have incomplete target area extraction, poor multi-scale adaptability, insufficient utilization of cross-domain features, difficulty in anomaly tracking and recovery, and weak coordination between identification, tracking and prediction in tracking "low, small and slow" UAVs, making it difficult to achieve high-precision and robust real-time tracking.
A multi-scale background dynamic query fusion perception detection method is adopted. By combining multi-scale background dynamic modeling, cross-domain feature fusion, spatiotemporal constraint localization, and anomaly recovery mechanism with trajectory analysis, the integrated design of target detection, tracking and prediction is realized, thus solving the above problems.
It achieves high-precision and robust real-time tracking in complex environments, enhances the reliability of target recognition and continuous tracking, reduces hardware resource requirements, and meets the engineering needs of high-risk scenarios such as low-altitude security.
Smart Images

Figure BDA0005718117660000031 
Figure BDA0005718117660000045 
Figure BDA0005718117660000048
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning, particularly the field of deep neural network interpretability research. In the academic field, it can be used to build infrared "low, small, and slow" target tracking models, and it can also be applied to the field of anti-drone scenarios in military settings. Background Technology
[0002] In recent years, small, low-altitude, and slow-moving drones have rapidly gained popularity in civilian fields (such as aerial photography, logistics, and agricultural plant protection) and some professional fields due to their small size, low cost, and flexible operation. However, this has also brought serious low-altitude security risks—unauthorized small, low-altitude, and slow-moving drones may intrude into sensitive areas such as airport runways, military restricted areas, and large venues, disrupting normal order and even causing safety accidents. Therefore, real-time and stable tracking of small, low-altitude, and slow-moving drone targets has become one of the core requirements in the field of low-altitude security.
[0003] Current mainstream target tracking technologies are mainly divided into two categories: traditional methods and deep learning-based methods. However, both have significant limitations in tracking "low, small, and slow" UAVs.
[0004] The shortcomings of traditional tracking methods: Traditional methods, represented by optical flow, Meanshift algorithm, and inter-frame difference division, rely on the target's grayscale features, color histogram, or simple motion information for tracking. Among them, inter-frame difference division was a commonly used early moving target extraction method, which segments the moving region by calculating the pixel difference between adjacent frames. However, this method is extremely sensitive to environmental noise (such as sudden changes in illumination, cloud cover, and ground clutter)—slight changes in illumination can lead to a large number of false moving regions being detected. Furthermore, the small size and low grayscale difference between "low, small, and slow" UAVs and the background (such as the sky and trees) can lead to incomplete target region extraction and "tracking interruption" phenomena. At the same time, traditional methods lack modeling of the target's long-term motion patterns and cannot cope with the UAV's maneuvering flight (such as turning, hovering, and acceleration), resulting in extremely poor tracking robustness.
[0005] Limitations of existing deep learning tracking methods: While deep learning-based tracking methods (such as the Siamese series and Transformer-based trackers) improve target representation capabilities through deep features, they still have shortcomings in "low, small, and slow" scenarios.
[0006] Background interference problem: Most methods use fixed-scale background modeling, which cannot adapt to the scale changes of "low, small and slow" drones at different distances (e.g., the target occupies only a few pixels at a distance, but the scale increases significantly at a close distance), and it is easy to misjudge similar targets in the background (such as birds and kites) as the tracking object.
[0007] Insufficient feature fusion: Existing methods mostly rely on visual features (such as RGB image features) and ignore the complementarity of cross-domain information such as motion features, infrared features, and depth features. This leads to tracking interruption when visual features fail in complex environments (such as nighttime or hazy days).
[0008] Weak anomaly recovery capability: When the drone is briefly obscured (such as by buildings or trees) or the tracking confidence decreases, the existing methods lack effective memory correction and global re-examination mechanisms, making it difficult to retrieve the target and ensuring tracking continuity.
[0009] The unique challenges of industry scenarios: Tracking "small, slow, and low-end" drones requires extremely high real-time performance (usually above 25fps), but existing high-precision tracking methods often rely on complex network structures and large amounts of computing resources, making them difficult to deploy on embedded devices (such as security cameras and drone-borne detectors). At the same time, industry demands not only require "tracking" but also "accurate identification" (i.e., distinguishing drones from other interfering targets) and "accurate prediction" (predicting drone flight trajectories), but existing methods mostly focus only on the single function of "tracking" and lack an integrated design for target recognition and trajectory prediction.
[0010] In summary, the core challenges currently faced by target tracking technology for low-altitude, small, and slow UAVs are: sensitivity to environmental noise, incomplete target region extraction, poor multi-scale adaptability, insufficient utilization of cross-domain features, difficulty in anomaly tracking and recovery, and weak coordination between identification, tracking, and prediction. To address these issues, this invention proposes a target tracking method for low-altitude, small, and slow UAVs based on multi-scale background dynamic query fusion perception detection. Through multi-scale background dynamic modeling, cross-domain feature fusion, spatiotemporal constraint localization, and integrated design of anomaly recovery and trajectory prediction, this method overcomes existing technological bottlenecks and achieves high-precision, highly robust real-time tracking of low-altitude, small, and slow UAVs. Summary of the Invention
[0011] Purpose of the Invention: To address the problems of existing methods (especially inter-frame difference division) in tracking low-altitude, small- and slow-moving unmanned aerial vehicles (UAVs), such as sensitivity to environmental noise, incomplete target region extraction, poor multi-scale adaptability, and difficulty in recovering from tracking anomalies, this invention provides a target tracking method for low-altitude, small- and slow-moving UAVs based on multi-scale background dynamic query fusion perception detection. This method achieves accurate target detection through multi-scale background dynamic query, enhances target representation through cross-domain feature fusion, optimizes tracking accuracy through spatiotemporal constraint positioning, ensures tracking continuity through an anomaly recovery mechanism, and achieves secondary target judgment by combining trajectory analysis, ultimately solving the problem of stable tracking of low-altitude, small- and slow-moving UAVs in complex environments.
[0012] Technical solution:
[0013] The "low, small, and slow" UAV target tracking method based on multi-scale background dynamic query fusion perception detection described in this invention is generally divided into four major modules: target detection stage, spatiotemporal optimization constraint stage, memory trajectory judgment stage, and tracking processing stage. The specific steps are as follows:
[0014] 1. A method for tracking "low, small, and slow" UAV targets based on multi-scale background dynamic query fusion perception detection under spatiotemporal constraints, characterized by the following steps:
[0015] Step 1.1: Construct a multi-scale background dynamic query fusion perception detection module. Input the first frame of the target tracking video stream to generate the bounding box, detection confidence, and initial spatial coordinates of the candidate targets. By setting a detection confidence threshold, filter out low-confidence targets below the threshold and retain the set of high-potential candidate targets to provide an effective target basis for the subsequent initial cross-domain feature extraction and tracking process.
[0016] Step 1.2: Construct a spatiotemporal optimization constraint module. First, it calls up the compressed feature maps of the last 5 frames to output the initial value of the target position. Then, it performs attention matching between the current purified feature and the features of the last 12 frames in the distillation memory pool, and corrects the initial position value by combining uniform speed, uniform acceleration, and coordinated turning motion models to obtain the optimized position value. Finally, it performs cross-segment association optimization, using a fixed number of frames as tracking segments to solve the background interference problem, realize target association, and ensure the continuity of cross-segment tracking.
[0017] Step 1.3: Construct a memory trajectory judgment module. First, based on the memory of the current frame and the target features of the current frame, calculate the updated memory data according to a specific weight coefficient; then, judge the similarity between the current target features and the existing features to generate a positive and negative tree; finally, simultaneously update the trajectory cache, record the latest motion trajectory information of the target, and predict the target's motion in the next frame based on the information.
[0018] Step 1.4: Construct a tracking and judgment module to perform multi-scale background dynamic query fusion perception detection on the current frame, adjust the features of the current frame and generate purified features; calibrate the detected bounding boxes to ensure that the bounding boxes can accurately reflect the position and size of the target in the current frame.
[0019] 2. The "low, small, and slow" UAV target tracking method based on multi-scale background dynamic query fusion perception detection under spatiotemporal constraints as described in claim 1, characterized in that the method of the multi-scale background dynamic query fusion perception detection module in step 1.1 is as follows:
[0020] Step 2.1: Select the first frame image of the video stream and input it into the feature extraction layer of the detection module to obtain the original multi-scale feature map;
[0021] Step 2.2: Perform a dynamic query generation operation on the original multi-scale feature map to generate a window-level dynamic query vector and divide the multi-scale feature window;
[0022] Step 2.3: By fusing perceptual attention calculations, the relationships between the whole and its parts in images of different scales are mapped, using the following formula:
[0023]
[0024] Where Sim is the query-feature similarity, β is the background suppression coefficient, and M is the feature similarity. bg Use a background mask to suppress background interference. F represents the query feature in perceptual attention computation. st Sim represents the target features to be matched. bg The final similarity after background suppression is achieved;
[0025] Step 2.4: Perform non-maximum suppression on the fused perception results and calculate the candidate box detection confidence. (N win (To cover the number of candidate boxes), output the candidate target bounding boxes, detection confidence scores, and initial spatial coordinates, and filter C. det Low confidence targets with a confidence level <0.72 are retained, while the set of high-potential candidate targets is preserved.
[0026] 3. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that the method for constructing the spatiotemporal optimization constraint module in step 1.4 is as follows:
[0027] Step 3.1: Call the historical features of the last 5 frames {F t-5 ,F t-4 ,…,F t-1 The 6D feature map of the current frame (3-channel RGB image + 1-channel infrared image + 1-channel depth map + 1-channel foreground mask) is input into the encoder for compression processing, compressing it to its original size. Obtain the compressed feature F enc ;
[0028] F enc Backbone feature F bb Residual fusion is performed, and the fusion formula is as follows:
[0029] F fuse =F enc +λ f ·F bb
[0030] Where, λ f=0.65 is the fusion coefficient, used to balance the weights of compressed features and backbone features;
[0031] The fused features F fuse Input the CBAM attention module, focus on the key region of the target through the attention mechanism, and output the initial value P of the target position. t ;
[0032] Step 3.2: Obtain the purified features of the current frame Combine it with distillation memory cell M D Attention matching is performed on the features of the last 12 frames stored in the database, and the matching score is calculated using the following formula:
[0033]
[0034] in, D is the transpose of the characteristic matrix of the distillation memory cell. v =256 is the visual feature dimension. This is the normalization term; combined with the motion model, the initial position value P is... t After correction, the motion model formula is as follows:
[0035]
[0036] Among them, P t-1 v represents the target position in the previous frame. t-1 a represents the target velocity in the previous frame. t-1 Δt represents the target acceleration in the previous frame, and Δt represents the time interval between adjacent frames.
[0037] By matching score S attn We adjust the initial position value using weighted correction to obtain the optimized position value. The weighted formula is:
[0038]
[0039] Where α is the weighting coefficient (from S) attn The maximum value is dynamically adjusted, ranging from [0.4, 0.6];
[0040] Step 3.3: Using 35 frames as a fixed tracking segment, background interference is suppressed within the segment through local 4D correlation calculation. The correlation calculation formula is as follows:
[0041]
[0042] in, Extract the top k features (k < 35) from the fragment, where i and j are the row and column indices of the feature map, respectively. 4D A higher value indicates less background interference;
[0043] The association between segments is achieved through association scores, and the formula for calculating the association score is as follows:
[0044] Score clip =0.7S attn +0.3IOU(B t B t-35 )
[0045] Among them, IOU(B t B t-35 ) represents the current frame bounding box B. t Bounding box B of the previous 35 frames t-35 The intersection-union ratios are 0.7 and 0.3, respectively, which are the weighting coefficients for attention matching score and IOU.
[0046] When Score clip When the value is ≥0.5, it is determined to be the same target, realizing cross-segment tracking association and ensuring tracking continuity.
[0047] 4. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that the method for constructing the memory trajectory judgment module in step 1.4 is as follows:
[0048] Step 4.1: Long-term memory data update. Based on the current frame memory and the current frame target features, the updated long-term memory data is calculated using fixed weighting coefficients. The formula is:
[0049]
[0050] in, For the current frame's long-term memory data, This is the long-term memory data from the previous frame. The target feature of the current frame is represented by λ = 0.7 (the weighting coefficient during normal tracking), which is used to balance the contribution of historical memory and current features to ensure the continuity and timeliness of memory updates.
[0051] Step 4.2: Positive and Negative Tree Generation and Feature Update. First, calculate the similarity between the current target feature and the latest feature in the positive branch of the positive and negative tree. Cosine similarity is used for calculation, and the formula is:
[0052]
[0053] Among them, F curr For the current target features, F last This represents the latest feature of the positive branch in a positive-negative tree. Let ||·|| be the target feature of the current frame, and ||·|| be the L2 norm.
[0054] If Sim < 0.8, then the current target feature Fcurr Add to the deepest level of the positive branch of the positive-negative tree to complete the generation and update of the positive-negative tree; if Sim≥0.8, then keep the existing structure of the positive-negative tree unchanged.
[0055] Step 4.3: Trajectory Cache Update and Motion Prediction First, the trajectory cache is updated synchronously, setting the target position P of the current frame... t Add to the trajectory cache list to form a historical trajectory set Traj = {P1, P2, ..., P} t}, P1 is the initial position of the first frame, P t Position of the current frame;
[0056] Subsequently, based on the motion patterns of historical trajectories, the target's position in the next frame is predicted. A quadratic polynomial is used to fit the motion model, with the following formula:
[0057] P t+1 =a·t 2 +b·t+c
[0058] Where t is the current frame number, a, b, and c are coefficients obtained by least squares fitting based on the historical trajectory Traj, and P t+1 It predicts the target's position in the next frame, providing a position reference for target search in subsequent tracking frames.
[0059] 5. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that the method for constructing the tracking processing module in step 1.5 is as follows:
[0060] Step 5.1: Multi-scale background dynamic query fusion perception detection and feature adjustment for the current frame. The current frame from the tracking stage is input into the multi-scale background dynamic query fusion perception detection submodule. The original features of the current frame are adjusted through feature alignment to ensure that the features meet the requirements of subsequent processing in terms of dimensionality and spatial distribution. The feature alignment formula is:
[0061]
[0062] Among them, F t The original features of the current frame, For aligned features, W a To align the weight matrix (for adapting to feature dimensions and distribution), b a This is a bias term (used to correct feature offset);
[0063] Then the aligned features The input foreground-background separation network uses depthwise separable convolutions to suppress background noise and a CBAM attention module to highlight foreground information, generating refined features focused on the target. The formula is as follows:
[0064]
[0065] SepConv(·) is a depthwise separable convolution operation, and CBAM(·) is an attention mechanism operation. Refine features for the target in the current frame;
[0066] Step 5.2: Obtain the initial bounding box B of the current frame output by the multi-scale background dynamic query fusion perception detection submodule. t Using the lens distortion parameter matrix K (including radial and tangential distortion coefficients) obtained beforehand through camera intrinsic parameter calibration, the initial bounding box is geometrically calibrated to correct positional deviations caused by lens optical characteristics. The calibration formula is as follows:
[0067]
[0068] Among them, B t The initial bounding box for the current frame is not calibrated (parameters include the coordinates of the top left corner (x)). t ,y t Width w t Height h t ), This is the final bounding box after calibration; through this calibration operation, we can ensure... It can accurately reflect the actual spatial position and size of "low, small and slow" UAV targets in the current frame image, providing an accurate basis for subsequent trajectory recording and target positioning.
[0069] The beneficial effects of this invention are:
[0070] 1. In the field of academic research, this method can be used to construct a deep neural network model related to the tracking of "low, small and slow" UAV targets. It can realize the understanding and interpretation of the multi-scale feature interaction and spatiotemporal constraint decision-making process in the tracking network. Through positive and negative tree structured storage (positive branch target features, negative branch interference templates) and distillation memory pool key feature screening, a clear boundary representation of target feature matching and tracking state determination can be obtained, which gives the deep neural network controllability and safety in the learning process of small target tracking scenarios.
[0071] 2. In engineering applications, this method can be applied to the monitoring and tracking of "low, small, and slow" UAVs in high-risk scenarios such as low-altitude security, airport airspace control, and border patrol. This enhances the reliability of UAV target identification and continuous tracking in high-risk areas. Because this method uses confidence filtering (C... det <0.72) Eliminating invalid target calculations and feature compression (compressing 6D feature maps to 1 / 20 size) reduces redundant calculations. The compressed tracking model reduces hardware computing power dependence, which can reduce equipment and maintenance costs in industrial deployments, thereby improving engineering benefits in the field of low-altitude safety monitoring. Attached Figure Description
[0072] Figure 1 A schematic diagram of the tracking method steps;
[0073] Figure 2 Detailed diagram of the structure and technology of the "small, slow" target detection method Detailed Implementation
[0074] The invention will now be further described with reference to the accompanying drawings.
[0075] A schematic diagram of the steps of a "low, small, and slow" UAV target tracking method based on multi-scale background dynamic query fusion perception detection under spatiotemporal constraints is shown below. Figure 1 As shown, the overall process includes the following steps:
[0076] (1) Construct a multi-scale background dynamic query fusion perception and detection module, such as Figure 2 As shown, a multi-branch downsampling structure is first used, with each branch equipped with a task-adaptive weight learning module to model discrete weight channels, providing sparse and diverse discrete channel decision paths. The calculation path is dynamically selected according to the weights, and the feature vectors of each feature layer are dynamically inferred. Key vector features with large weights are selected and sent to the subsequent context-based fusion module. Then, the features of each branch enter the context-based fusion module with the number decreasing with the scale, and multiple rounds of feature interaction and context information fusion are performed. Finally, all branch features are merged into the decision module. After result fusion and final decision-making, multi-scale and multi-task feature information is integrated, and the detection result is output.
[0077] (2) The detection results are brought into the spatiotemporal optimization constraint module, the feature maps detected in the most recent frames are called to purify the features, and attention matching is performed with the features in the distillation memory pool. The motion trajectory of the UAV is judged by combining the uniform speed, uniform acceleration and coordinated turning motion models. Finally, cross-segment association optimization is performed with a fixed number of frames as the tracking segment to achieve target association and ensure the continuity of cross-segment tracking.
[0078] (3) Construct a memory trajectory judgment module based on the results of the spatiotemporal optimization constraint module, draw a trajectory diagram based on frame memory, judge the similarity between the current target features and existing features, generate a positive and negative tree, and judge the target's next frame action based on the positive and negative tree;
[0079] (4) Based on the results of the memory trajectory judgment module, predict the action of the next frame, perform multi-scale background dynamic query fusion perception detection near the position of the next frame, adjust the current frame features to generate purified features; at the same time, calibrate the detected bounding box to ensure that it accurately reflects the position and size of the target in the current frame, and optimize the memory trajectory judgment model to achieve real-time tracking.
[0080] Assuming we take an embedded device (such as NVIDIA Jetson Xavier NX) deployed in a low-altitude security scenario as an example, we will verify the feasibility of the "low, small, and slow" UAV target tracking system of this invention. The device has a computing power of 21 TOPS, which meets the computing power requirements of low-altitude security front-end devices. The input video stream resolution is set to 1280×720, and the frame rate is 25fps. We simulate a mixed scenario of "low, small, and slow" UAVs (wingspan 0.5m, flight altitude 50-100m) and background interference (birds, clouds, ground buildings) commonly found in airport airspace. The formula for this step is as follows:
[0081]
[0082] In the initialization phase, the first frame of the video stream (resolution 1280×720, 3 channels) is selected and input into the multi-scale background dynamic query fusion perception detection module. The module outputs original multi-scale feature maps (scales of 1280×720, 640×360, and 320×180) after feature extraction. A dynamic query generation operation is then performed on the original feature maps to generate 32 window-level dynamic query vectors, dividing the data into feature windows corresponding to the scales. Finally, fusion perception attention is calculated (setting the background suppression coefficient β = 0.5 and the background mask M). bg (Generated by statistical analysis of the sky region in the first frame), suppressing interference from cloud and ground building backgrounds; non-maximum suppression is performed on the result, the formula for this step is as follows:
[0083]
[0084] Calculate the confidence score using the candidate box detection confidence score formula, and filter C. det For targets with a low confidence level of <0.72, two high-potential candidate targets (one drone and one bird) are retained. The bounding box coordinates (drone: (320,240,40,30), bird: (512,180,20,15)) and the initial spatial coordinates are output to complete the initialization verification, proving that the module can effectively filter targets.
[0085] During the normal tracking phase, the historical features from the last 5 frames and the 6D feature map of the current frame are retrieved. The 6D feature map is then input into the Encoder and compressed to 1 / 20 of its original size (64×36), resulting in the compressed feature F. enc The residual fusion is performed according to the fusion formula, which is shown below:
[0086] F fuse =F enc +λ f ·F bb
[0087] The initial position value P of the UAV is output after inputting into the CBAM attention module. t= (325, 242); Obtain the purified features F of the current frame. pt Attention matching is performed on the features of the last 12 frames in the distillation memory pool (information entropy threshold Tent = 0.6), and the matching score S is calculated. attn =0.82 (drones), 0.35 (birds), the calculation formula is as follows:
[0088]
[0089] Combining the motion model (Δt = 0.04s, previous frame velocity v) t-1 = (2,1) pixels / frame, acceleration at-1 = (0,0)) calculate position correction value P corrt = (327, 243), according to the weighted formula P2_t = 0.55·325 + 0.45·327 = 326, 0.55·242 + 0.45·243 = 242.45, the optimized position (326, 242) is obtained; taking 35 frames as the tracking segment, the 4D correlation within the segment is calculated to be Corr4D = 0.85 (UAV) and 0.42 (bird), the formula is as follows:
[0090]
[0091] Based on the formula for the correlation score between segments:
[0092] Score clip =0.7×0.82+0.3×IOU((326,242,40,30),(280,210,40,30))
[0093] =0.7×0.82+0.3×0.65=0.779≥0.5
[0094] The target was identified as the same UAV, and no "tracking interruption" occurred during 300 consecutive frames (12 seconds) of tracking, proving that the spatiotemporal optimization constraint module can guarantee tracking continuity.
[0095] During the operation of the memory trajectory judgment module, the long-term memory is updated ( For the memory of drone features from the previous frame, The formula for updating long-term memory (using the features of the current frame) is as follows:
[0096]
[0097] Calculate the cosine similarity Sim = 0.85 ≥ 0.8 between the current feature and the latest feature of the positive branch of the positive and negative trees, keeping the positive and negative tree structure unchanged; update the trajectory cache to form the historical trajectory set Traj = {P1, P2, ..., P...} t The predicted position P is obtained by fitting a quadratic polynomial (a = 0.002, b = 1.8, c = 320).t+1 = (328,244), the actual detection position in the next frame is (329,243), with an error of 2 pixels, proving that the trajectory prediction is effective.
[0098] During the verification of the tracking processing module, an alignment operation is performed on the original features of the current frame (W). a The weight matrix is 3×3, b a =0), thus obtaining the alignment feature F. a_t Features F are generated and refined through depthwise separable convolution and CBAM attention. p_t To highlight the potential of drones, the formula for generating purified features is as follows:
[0099]
[0100] By combining the camera intrinsic distortion matrix K (radial distortion K1 = -0.01, K2 = 0.002, tangential distortion p1 = 0.001, p2 = -0.0005), the final bounding box (326, 242, 40, 30) was obtained after calibrating the initial bounding box. The error between the initial bounding box and the manually labeled position is ≤3 pixels, which proves that the bounding box calibration is accurate.
[0101] In the anomaly handling verification, a simulated drone was briefly obscured by a building (frames 150-155). The correlation score was [not specified]. clip =0.42<0.5, triggering exception handling; invoking the global re-examination mechanism, based on the positive branch features of the positive and negative trees and the key features of the distillation memory pool, re-examining is performed near the prediction region (prediction position (410,280) in frame 150), and tracking resumes in frame 156. attn =0.78, proving that the abnormal recovery ability is effective.
[0102] After being deployed and tested on embedded devices, the entire system achieved an average inference speed of 28fps (higher than the real-time requirement of 25fps), a single frame computing power consumption of 1.2TOPS, a hardware resource utilization rate of ≤70%, and ran continuously for 2 hours without failure. It can also accurately distinguish between drones and birds (recognition accuracy rate of 92%), meeting the needs of low-altitude security scenarios.
[0103] The above text, through specific scenarios, parameters, and test results, comprehensively verifies the feasibility of the system of the present invention, from initialization, tracking, memory prediction, anomaly recovery to hardware deployment. The descriptions listed are only specific examples of feasible implementation methods and do not depart from the equivalent implementation or modification of the technology of the present invention, and all fall within the protection scope of the present invention.
Claims
1. A method for tracking "low, small, and slow" UAV targets based on multi-scale background dynamic query fusion perception detection under spatiotemporal constraints, characterized in that, Includes the following steps: Step 1.1: Construct a multi-scale background dynamic query fusion perception detection module, input the first frame of the target tracking video stream, and generate the bounding box, detection confidence and initial spatial coordinates of the candidate target; By setting a detection confidence threshold, low-confidence targets below the threshold are filtered out, while a set of high-potential candidate targets is retained, providing an effective target basis for subsequent initial cross-domain feature extraction and tracking processes. Step 1.2: Construct a spatiotemporal optimization constraint module. First, it calls up the compressed feature maps of the last 5 frames to output the initial value of the target position. Then, it performs attention matching between the current purified feature and the features of the last 12 frames in the distillation memory pool, and corrects the initial position value by combining uniform speed, uniform acceleration, and coordinated turning motion models to obtain the optimized position value. Finally, it performs cross-segment association optimization, using a fixed number of frames as tracking segments to solve the background interference problem, realize target association, and ensure the continuity of cross-segment tracking. Step 1.3: Construct a memory trajectory judgment module. First, based on the memory of the current frame and the target features of the current frame, calculate the updated memory data according to a specific weight coefficient; then, judge the similarity between the current target features and existing features, and generate a positive and negative tree. Finally, the trajectory cache is updated synchronously to record the latest motion trajectory information of the target and predict the motion of the target in the next frame based on the information. Step 1.4: Construct a tracking and judgment module to perform multi-scale background dynamic query fusion perception detection on the current frame, adjust the features of the current frame and generate refined features; The detected bounding boxes are calibrated to ensure that they accurately reflect the position and size of the target in the current frame.
2. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection as described in claim 1, characterized in that, The method of the multi-scale background dynamic query fusion perception detection module in step 1.1 is as follows: Step 2.1: Select the first frame image of the video stream and input it into the feature extraction layer of the detection module to obtain the original multi-scale feature map; Step 2.2: Perform a dynamic query generation operation on the original multi-scale feature map to generate a window-level dynamic query vector and divide the multi-scale feature window; Step 2.3: By fusing perceptual attention calculations, the relationships between the whole and its parts in images of different scales are mapped, using the following formula: Where Sim is the query-feature similarity, β is the background suppression coefficient, and M is the feature similarity. bg Use a background mask to suppress background interference. F represents the query feature in perceptual attention computation. st Sim represents the target features to be matched. bg The final similarity after background suppression is achieved; Step 2.4: Perform non-maximum suppression on the fused perception results and calculate the candidate box detection confidence. (N win (To cover the number of candidate boxes), output the candidate target bounding boxes, detection confidence scores, and initial spatial coordinates, and filter C. det Low confidence targets with a confidence level <0.72 are retained, while the set of high-potential candidate targets is preserved.
3. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that, The method for constructing the spatiotemporal optimization constraint module in step 1.4 is as follows: Step 3.1: Call the historical features of the last 5 frames {F t-5 ,F t-4 ,…,F t-1 The 6D feature map of the current frame (3-channel RGB image + 1-channel infrared image + 1-channel depth map + 1-channel foreground mask) is input into the encoder for compression processing, compressing it to its original size. Obtain the compressed feature F enc ; F enc Backbone feature F bb Residual fusion is performed, and the fusion formula is as follows: F fuse =F enc +λ f ·F bb Where, λ f =0.65 is the fusion coefficient, used to balance the weights of compressed features and backbone features; The fused features F fuse Input the CBAM attention module, focus on the key region of the target through the attention mechanism, and output the initial value P of the target position. t ; Step 3.2: Obtain the purified features of the current frame Combine it with distillation memory cell M D Attention matching is performed on the features of the last 12 frames stored in the database, and the matching score is calculated using the following formula: in, D is the transpose of the characteristic matrix of the distillation memory cell. v =256 is the visual feature dimension. This is the normalization term; combined with the motion model, the initial position value P is... t After correction, the motion model formula is as follows: Among them, P t-1 v represents the target position in the previous frame. t-1 a represents the target velocity in the previous frame. t-1 Δt represents the target acceleration in the previous frame, and Δt represents the time interval between adjacent frames. By matching score S attn We adjust the initial position value using weighted correction to obtain the optimized position value. The weighted formula is: Where α is the weighting coefficient (from S) attn The maximum value is dynamically adjusted, ranging from [0.4, 0.6]; Step 3.3: Using 35 frames as a fixed tracking segment, background interference is suppressed within the segment through local 4D correlation calculation. The correlation calculation formula is as follows: in, Extract the top k features (k < 35) from the fragment, where i and j are the row and column indices of the feature map, respectively. 4D A higher value indicates less background interference; The association between segments is achieved through association scores, and the formula for calculating the association score is as follows: Score clip =0.7S attn +0.3IOU(B t ,B t-35 ) Among them, IOU(B t B t-35 ) represents the current frame bounding box B. t Bounding box B of the previous 35 frames t-35 The intersection-union ratios are 0.7 and 0.3, respectively, which are the weighting coefficients for attention matching score and IOU. When Score clip When the value is ≥0.5, it is determined to be the same target, realizing cross-segment tracking association and ensuring tracking continuity.
4. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that, The method for constructing the memory trajectory judgment module in step 1.4 is as follows: Step 4.1: Long-term memory data update. Based on the current frame memory and the current frame target features, the updated long-term memory data is calculated using fixed weighting coefficients. The formula is: in, For the current frame's long-term memory data, This is the long-term memory data from the previous frame. The target feature of the current frame is represented by λ = 0.7 (the weighting coefficient during normal tracking), which is used to balance the contribution of historical memory and current features to ensure the continuity and timeliness of memory updates. Step 4.2: Positive and Negative Tree Generation and Feature Update. First, calculate the similarity between the current target feature and the latest feature in the positive branch of the positive and negative tree. Cosine similarity is used for calculation, and the formula is: Among them, F curr For the current target features, F last This represents the latest feature of the positive branch in a positive-negative tree. Let ||·|| be the target feature of the current frame, and ||·|| be the L2 norm. If Sim < 0.8, then the current target feature F curr Add to the deepest level of the positive branch of the positive-negative tree to complete the generation and update of the positive-negative tree; if Sim≥0.8, then keep the existing structure of the positive-negative tree unchanged. Step 4.3: Trajectory Cache Update and Motion Prediction First, the trajectory cache is updated synchronously, setting the target position P of the current frame... t Add to the trajectory cache list to form a historical trajectory set Traj = {P1, P2, ..., P} t }, P1 is the initial position of the first frame, P t Position of the current frame; Subsequently, based on the motion patterns of historical trajectories, the target's position in the next frame is predicted. A quadratic polynomial is used to fit the motion model, with the following formula: P t+1 =a·t 2 +b·t+c Where t is the current frame number, a, b, and c are coefficients obtained by least squares fitting based on the historical trajectory Traj, and P t+1 It predicts the target's position in the next frame, providing a position reference for target search in subsequent tracking frames.
5. The method for tracking "low, small, and slow" UAV targets under spatiotemporal constraints based on multi-scale background dynamic query fusion perception detection according to claim 1, characterized in that, The method for constructing the tracking processing module in step 1.5 is as follows: Step 5.1: Multi-scale background dynamic query fusion perception detection and feature adjustment for the current frame. The current frame from the tracking stage is input into the multi-scale background dynamic query fusion perception detection submodule. The original features of the current frame are adjusted through feature alignment to ensure that the features meet the requirements of subsequent processing in terms of dimensionality and spatial distribution. The feature alignment formula is: Among them, F t The original features of the current frame, For aligned features, W a To align the weight matrix (for adapting to feature dimensions and distribution), b a This is a bias term (used to correct feature offset); Then the aligned features The input foreground-background separation network uses depthwise separable convolutions to suppress background noise and a CBAM attention module to highlight foreground information, generating refined features focused on the target. The formula is as follows: SepConv(·) is a depthwise separable convolution operation, and CBAM(·) is an attention mechanism operation. Refine features for the target in the current frame; Step 5.2: Obtain the initial bounding box B of the current frame output by the multi-scale background dynamic query fusion perception detection submodule. t Using the lens distortion parameter matrix K (including radial and tangential distortion coefficients) obtained beforehand through camera intrinsic parameter calibration, the initial bounding box is geometrically calibrated to correct positional deviations caused by lens optical characteristics. The calibration formula is as follows: Among them, B t The initial bounding box for the current frame is not calibrated (parameters include the coordinates of the top left corner (x)). t ,y t Width w t Height h t ), This is the final bounding box after calibration; through this calibration operation, we can ensure... It can accurately reflect the actual spatial position and size of "low, small and slow" UAV targets in the current frame image, providing an accurate basis for subsequent trajectory recording and target positioning.