A UAV target tracking method

Through confidence-based feature extraction and judgment, lightweight feature re-identification network and adaptive Kalman filtering algorithm, the UAV target tracking method is optimized, the target detection accuracy and tracking stability problems in complex environments are solved, and efficient and accurate multi-target tracking is achieved.

CN119887855BActive Publication Date: 2025-10-28WEIYUAN SHENGXIANG COMPOSITE MATERIAL CO LTD

Patent Information

Application Number
CN202510332138.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-10-28
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing multi-target tracking technologies suffer from problems such as low target detection accuracy, poor tracking stability, and insufficient adaptability to resource-constrained environments in complex environments.

Method used

We employ a confidence-based feature extraction and decision-making strategy, a lightweight feature re-identification network Rep-OSNet, an adaptive cost function that dynamically fuses motion and appearance similarity, and a noise-adaptive Kalman filter algorithm to optimize the target detection and tracking process.

Benefits of technology

It improves the accuracy and stability of target tracking, reduces computational overhead, adapts to complex scenarios and resource-constrained environments, and enhances system real-time performance and tracking success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887855B_ABST
    Figure CN119887855B_ABST
Patent Text Reader

Abstract

This invention discloses a target tracking method for unmanned aerial vehicles (UAVs), relating to the field of computer vision, aiming to solve the problems of target occlusion, rapid movement, and efficiency in dynamic scenes. The method includes: acquiring the target's bounding box and confidence score through a detector, and predicting the target's position and state using a noisy adaptive Kalman filter; proposing a confidence-based feature extraction and judgment strategy, performing IoU matching between the trajectory and the detection box; for targets with a matching degree below a threshold and a confidence score change rate above a threshold, extracting appearance features using a lightweight feature re-identification network (Rep-OSNet); otherwise, reusing features from the previous frame; designing a cost function that fuses motion direction and appearance similarity to perform cascaded matching between the target and the confirmed trajectory; performing secondary association matching on targets and trajectories that mismatch in the cascaded matching; updating the matched trajectory state using a noisy adaptive Kalman filter, deleting long-term lost mismatched trajectories, and outputting the target trajectory prediction box and ID.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and UAV application technology, specifically relating to a UAV target tracking method, which is suitable for real-time multi-target tracking and trajectory management in dynamic scenarios. Background Technology

[0002] The widespread application of drones in key areas such as modern security and military defense has brought numerous potential risks and challenges. Drone target tracking technology has become an important research direction that urgently needs development and improvement.

[0003] Multi-object tracking comprises two parts: multi-object detection and multi-object tracking. Classic deep learning object detection networks are divided into two-stage networks and single-stage networks. Two-stage detection algorithms offer high detection accuracy but are slow; single-stage detection algorithms are fast but have a high false positive rate. Two-stage networks, such as R-CNN, FastR-CNN, FasterR-CNN, and CascadeR-CNN, first generate candidate regions, then classify and locate them, making them suitable for applications requiring higher detection accuracy. Single-stage networks, such as SSD, YOLO series, and CenterNet, directly generate coordinate positions and class probabilities, making them faster than two-stage networks. Deep learning-based multi-object tracking methods are mainly divided into two categories: Tracking Based Detection (TBD) and Joint Detection Tracking (JDT). JDT algorithms attempt to fuse detection and tracking modules to improve inference speed, but in practical applications, they face difficulties in co-training the modules, leading to unstable overall performance. The TBD strategy, with its clear multi-stage design structure, allows for separate optimization of detection and tracking, demonstrating good adaptability to complex scenes. The classic TBD tracker SORT proposed a simple and real-time data concatenation method. DeepSORT, building upon the SORT framework, adds appearance information to improve algorithm performance, enabling tracking of targets with prolonged occlusion. StrongSORT further upgrades it in terms of detection, embedding, and concatenation, employing the latest components and training techniques to optimize DeepSORT. It proposes an appearance-free linking model (AFLink) that concatenates short trajectories into complete trajectories using only spatiotemporal information, and Gaussian smooth interpolation (GSI) to compensate for missing detections. BoT-SORT achieves better bounding box localization through camera motion compensation and more accurate Kalman filter state vectors, as well as a novel fusion method based on IoU and re-id cosine distance. However, when faced with complex problems such as blurred background interference, object occlusion, and disappearance, it still suffers from inaccurate tracking and difficulty in re-locking targets after loss.

[0004] In summary, existing multi-target tracking technologies still have many problems, especially in terms of target detection accuracy, tracking stability, and adaptability to resource-constrained environments in complex conditions. Therefore, developing a new UAV target tracking method is an urgent need. Summary of the Invention

[0005] To address the shortcomings and deficiencies of the above methods, this invention proposes a UAV target tracking method, comprising the following steps:

[0006] Step S1, UAV target detection: Obtain the bounding box and confidence score of the UAV target in the current frame through the UAV target detector;

[0007] Step S2, State Prediction: Predict the position and motion state of the target in the current frame using the noise adaptive Kalman filter algorithm;

[0008] Step S3, Feature Extraction Judgment: A confidence-based feature extraction judgment strategy is proposed. The trajectory and detection box are matched by IoU, and the confidence change rate of the target is calculated. Targets with IoU matching degree below the threshold and confidence change rate above the threshold are marked. Targets with IoU matching degree above the threshold or confidence change rate below the threshold are reused from the previous frame.

[0009] Step S4, extract appearance features: By constructing a lightweight feature re-identification network Rep-OSNet, extract appearance features for targets in the current frame whose IoU matching degree is lower than the threshold and whose confidence change rate is higher than the threshold.

[0010] Step S5, Cascaded Matching: By designing an adaptive cost function that integrates motion direction and appearance similarity, the detected target and the confirmed trajectory are cascaded matched.

[0011] Step S6, Secondary Association Matching: Perform secondary association matching on targets and trajectories that have mismatched in the cascaded matching;

[0012] Step S7, State Update: Update the matching trajectory state using noise adaptive Kalman filter, delete trajectories that still mismatch after secondary association matching and whose loss time is greater than the threshold; output the updated target trajectory prediction box and ID.

[0013] A further preferred embodiment of the UAV target detection network in step S1 includes the following steps:

[0014] S11, Constructing a drone target detection network:

[0015] Construct a feature extraction backbone network;

[0016] Construct a multi-scale feature fusion network;

[0017] Construct the detection head.

[0018] S12, use the NMS algorithm to process the redundant detection boxes output by the network and output the detection results.

[0019] More preferably, step S2 uses a noise-adaptive Kalman filter algorithm to predict the position and motion state of the target in the current frame, specifically including the following steps:

[0020] S21, based on the target's motion model, predict the target's state at the current moment according to the state estimate of the previous moment.

[0021] S22, calculate the correlation metric between the predicted state and the currently detected target position, set the correlation metric threshold to 0.7, compare the correlation metric with the threshold, and determine whether the trajectory is in a confirmed state or an unconfirmed state. If the correlation metric is less than the threshold, it is in a confirmed state; otherwise, it is in an unconfirmed state.

[0022] Further preferably, step S3 proposes a confidence-based feature extraction and determination strategy, which performs IoU matching between the trajectory and the detection box, calculates the target's confidence change rate, marks targets with an IoU matching degree below a threshold and a confidence change rate above a threshold, and reuses the features from the previous frame for targets with an IoU matching degree above a threshold or a confidence change rate below a threshold; specifically, it includes the following steps:

[0023] S31, Calculate the IoU between the trajectory prediction bounding box and the detection bounding box in the current frame, mark targets with an IoU matching degree lower than the threshold, and define the judgment condition as follows:

[0024]

[0025] in, Let i be the predicted bounding box of the i-th trajectory. To detect bounding boxes, a threshold is used. If this condition is met, the target motion state is considered stable, the appearance features of the previous frame are directly reused, the feature extraction of the current frame is skipped, and the target with an IoU matching degree lower than the threshold is marked.

[0026] S32, for targets with an IoU matching degree below the threshold, further calculate their confidence change rate, specifically defined as follows:

[0027]

[0028] in, Detect the confidence level for the current frame. The threshold is the confidence level of the previous frame. ;

[0029] If satisfied If the condition is not met, the target's appearance is considered to have changed significantly, and it is marked; otherwise, the target's appearance is considered not to have changed significantly, and there is no need to re-extract features.

[0030] Further preferably, in step S4, a lightweight feature re-identification network, Rep-OSNet, is constructed to extract target appearance features in the current frame where the IoU matching degree is lower than a threshold and the confidence change rate is higher than a threshold. This specifically includes the following steps:

[0031] S41 introduces the reparameterized structure SAG into the OSNET network to construct Rep-OSNet.

[0032] S42 and SAG have a multi-branch convolutional structure during the training phase.

[0033] Branch 1 generates spatial attention weights using 3×3 convolution, average pooling, multilayer perceptron (MLP), and sigmoid.

[0034] Branch 2 is a 3×3 depthwise separable convolution;

[0035] The output is the weight of branch 1 × the feature of branch 2 + the original input.

[0036] During the inference phase, the multi-branch structure is equivalently transformed into a single 3×3 convolution, reducing computational complexity and resource consumption. The parameter fusion formula is as follows:

[0037]

[0038] in, The weight matrix for the 3×3 depthwise separable convolution in branch 2. The weight parameters of the MLP in branch 1, This is the attention weight matrix after the Sigmoid activation function, used to spatially weight the features of branch 2.

[0039] In a further preferred embodiment, step S5 involves designing an adaptive cost function that integrates motion direction and appearance similarity to perform cascade matching between the detected target and the confirmed trajectory. This specifically includes the following steps:

[0040] S51, normalize the calculation for multiple matching items, calculate the IoU between the trajectory prediction box and the detection box, and define the normalized IoU distance as:

[0041]

[0042] in, Let i be the predicted bounding box of the i-th trajectory. To detect bounding boxes;

[0043] Extract the appearance feature vector of the detected target, calculate its cosine similarity with the orbital history feature database, and define the appearance distance as:

[0044]

[0045] in, To detect the appearance features of the target, The k-th feature in the trajectory history feature database;

[0046] Based on the historical trajectory movement direction and the relative position of the detection box, the direction difference angle is calculated and normalized as follows:

[0047]

[0048] in, To predict the direction angle for the trajectory, The azimuth angle of the detection box relative to the position of the previous frame on the trajectory.

[0049] S52, sets a dynamic weight adaptive mechanism based on trajectory loss time. Frames and maximum allowable loss time Calculate the appearance feature weighting coefficient α:

[0050]

[0051] in, Let i be the time when trajectory i was lost. To adapt the time loss threshold, it is dynamically adjusted according to the scenario. Let be the velocity variance of trajectory i in the most recent 5 frames; γ=0.05 is the velocity variance adjustment factor.

[0052] By fusing motion and appearance matching terms, the final matching cost function is generated.

[0053]

[0054] in, , And all distance items take values ​​in the range [0,1].

[0055] S53, Set the conditions for enabling direction matching. To reduce computational overhead, the direction matching degree... Enabled only when all of the following conditions are met:

[0056] Track loss time ;

[0057] average velocity of trajectory 1.5 × target average size / frame rate;

[0058] The target was not marked as stationary (historical velocity variance) In other cases, and , .

[0059] S54 performs adaptive parameter optimization, dynamically adjusting online and updating every 30 frames. :

[0060]

[0061] λ=1 was determined through cross-validation. Using the MOT20 and VisDrone datasets, the optimal value of λ was selected within the range [0.5, 1.5] using a grid search method, with the tracking success rate and frame rate (FPS) as optimization objectives. The average target size was the average width of all detection boxes in the current frame, and the average target velocity was the average of the displacement velocities of all trajectories.

[0062] In a further preferred embodiment, step S6 involves a secondary association matching of the target and trajectory that have mismatched in the cascaded matching, specifically including the following steps:

[0063] S61, obtain the target and trajectory of mismatch in the cascade matching stage.

[0064] S62 uses the CMC algorithm to update the position of the detection box in the association matching algorithm.

[0065] S63, set the IoU threshold T=0.5 for secondary matching. When the IoU value between the detection box and the trajectory is greater than 0.5, trigger bidirectional matching verification.

[0066] For each trajectory, calculate its IoU value with all bounding boxes. When the IoU value of a bounding box with this trajectory is greater than a threshold T, the bounding box is considered a candidate matching box for that trajectory.

[0067] For each detection box, calculate its IoU value with all trajectories. When the IoU value of a trajectory with the detection box is greater than a threshold T, the trajectory is considered a candidate matching trajectory for that detection box.

[0068] A detection box is considered to be successfully matched with a trajectory only if it is a candidate matching box for a trajectory in forward matching and the trajectory is also a candidate matching trajectory for the same detection box in backward matching; otherwise, the match is unsuccessful.

[0069] In a further preferred embodiment, step S7 uses a noise-adaptive Kalman filter to update the matching trajectory state, deletes trajectories that still mismatch after secondary association matching and whose loss exceeds a threshold, and outputs the updated target trajectory prediction box and ID; specifically including the following steps:

[0070] S71, a noise adaptation mechanism is added to the Kalman filter algorithm by introducing a trajectory influence factor to adjust the observation noise covariance. This process is defined as follows:

[0071]

[0072] in, As the trajectory influencing factor, It is to test the confidence level. This is the preset minimum noise covariance.

[0073] To control computational complexity, this formula uses interval division. The value is dynamically smoothed, and the value selection method is defined as follows:

[0074]

[0075] in, This represents the number of consecutive confirmations of the trajectory.

[0076] Verified through numerous experiments , , , The system can achieve an optimal balance between target tracking accuracy and computational efficiency. Specifically:

[0077] The minimum number of consecutive confirmation frames required for the initial stabilization of the target trajectory;

[0078] The threshold for a highly reliable trajectory;

[0079] and The results were obtained by fitting historical data using a gradient optimization algorithm.

[0080] S72 updates the state of the matching trajectories obtained from cascaded matching and secondary association matching.

[0081] S73 sets the threshold for the number of frames lost in a trajectory to 30. When a secondary associated matching trajectory loses more than 30 consecutive frames, the trajectory is deleted.

[0082] S74 updates the appearance model of each tracker to adapt to changes in the target's appearance.

[0083] S75 outputs the trajectory of each tracked target, including the target's ID, bounding box, and confidence level.

[0084] This invention constructs a highly efficient and accurate UAV target tracking system through lightweight network design, dynamic feature extraction strategy, adaptive cost function, and noise-robust filtering algorithm. It surpasses existing technologies in computational efficiency, long-term tracking stability, and scene adaptability, providing reliable technical support for key areas such as security patrol and military monitoring. Compared with existing technologies, this invention has the following beneficial effects:

[0085] 1. To address the computational redundancy caused by full-target feature extraction in traditional tracking algorithms, this invention proposes a confidence-based feature extraction decision strategy. Its key feature is the introduction of a dual decision mechanism: first, IoU matching is performed, followed by a confidence rate change determination to filter targets for feature extraction. For targets with stable motion or minimal appearance changes, features from the previous frame are directly reused. This strategy not only reduces computational overhead but also avoids target loss due to brief occlusion. Compared to the BOTSORT fixed extraction strategy, this invention significantly improves system real-time performance while maintaining tracking accuracy, making it particularly suitable for high-speed UAV flight scenarios.

[0086] 2. To address the issues of insufficient feature discrimination and high computational complexity in traditional feature re-identification networks under occluded scenarios, this invention proposes a lightweight feature re-identification network, Rep-OSNet. Its key feature is the introduction of a heavily parameterized SAG structure. During training, a multi-branch convolutional structure is employed, while during inference, this multi-branch structure is fused into a single-path 3×3 convolution, reducing the number of model parameters, improving inference speed, and maintaining robustness in feature discrimination. Compared to the traditional OSNet network, Rep-OSNet significantly optimizes feature extraction efficiency in drone scenarios, solving the bottleneck problem of limited resources on edge devices.

[0087] 3. To address the long-term occlusion tracking failure problem caused by the fixed weights of traditional cost functions, this invention designs a cost function that dynamically fuses motion and appearance features. Its key feature is the adaptive adjustment of motion and appearance weights based on trajectory loss time, velocity variance, and orientation consistency. An appearance weight coefficient formula is introduced; as the target loss time increases or motion uncertainty increases, the IoU matching weight is gradually reduced while the appearance feature weight is increased. Simultaneously, the orientation matching term is only activated for high-speed targets with short-term loss, avoiding directional noise interference from low-speed targets. This design improves the algorithm's recovery success rate in occluded scenarios, demonstrating a significant advantage over DeepSORT's fixed-weight strategy.

[0088] 4. To address the poor noise adaptability of Kalman filtering in complex scenes, this invention improves its state update mechanism. Its key feature is the proposal of a noise adaptive model, which dynamically adjusts the observation noise covariance by combining trajectory influence factors with detection confidence. This ensures that even with low confidence in a single frame, a long-term stably tracked target can maintain tracking continuity through historical performance. Compared to StrongSORT's NSA Kalman algorithm, this method improves trajectory integrity in blurred and partially occluded scenes, and controls computational complexity through interval partitioning, avoiding exponential resource consumption. Attached Figure Description

[0089] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0090] Figure 1 This is a structural diagram of a UAV target tracking method in an example of the present invention;

[0091] Figure 2 This is a flowchart of a UAV target tracking method in an example of the present invention;

[0092] Figure 3 This is a flowchart of the feature extraction and determination process in an example of the present invention;

[0093] Figure 4 This is a diagram of the Rep-OSNet network structure in an example of the present invention;

[0094] Figure 5 This is a diagram of the reparameterized SAG structure in an example of the present invention; Detailed Implementation

[0095] To make the objectives, technical solutions, and advantages of this invention clearer, the specific implementation process of this invention will be described in more detail below with reference to specific embodiments and accompanying drawings, so as to facilitate those skilled in the art to more accurately understand this invention and apply it to various specific fields.

[0096] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0097] See Figure 1 This invention discloses a method for tracking targets using unmanned aerial vehicles (UAVs), the specific process of which is as follows: Figure 2 As shown. The steps are:

[0098] Step S1 involves obtaining the bounding box and confidence score of the drone target in the current frame using a drone target detector; specifically, this includes the following steps:

[0099] S11, construct a UAV target detection network, construct a feature extraction backbone network, construct a multi-scale feature fusion network and a detection head.

[0100] S12, using the collected dataset, select The loss function and Adam optimizer are used to train the UAV target detection network. (The above...) The loss function is a weighted sum of the localization loss, classification loss, and confidence loss, and its definition is as follows:

[0101]

[0102] In the above formula, Indicates location loss. Represents classification loss, This represents the confidence loss.

[0103] S13, use the NMS algorithm to process the redundant detection boxes output by the network and output the detection results.

[0104] Step S2 involves using a noise-adaptive Kalman filter algorithm to predict the position and motion state of the target in the current frame; specifically, it includes the following steps:

[0105] S21, Based on the target's motion model, predict the target's current state according to the state estimate from the previous moment;

[0106] S22, calculate the correlation metric between the predicted state and the currently detected target position. By setting a threshold, compare the correlation metric with the threshold to determine whether the trajectory is in a confirmed or unconfirmed state. If the correlation metric is less than the threshold, it is in a confirmed state; otherwise, it is in an unconfirmed state.

[0107] Step S3, propose a feature extraction and determination strategy based on confidence, such as... Figure 3 As shown, the trajectory and detection box are matched using IoU, and the confidence change rate of the target is calculated. Targets with an IoU matching degree below a threshold and a confidence change rate above a threshold are marked. Targets with an IoU matching degree above a threshold or a confidence change rate below a threshold are reused from the features of the previous frame. Specifically, the steps include:

[0108] S31, Calculate the IoU between the trajectory prediction bounding box and the detection bounding box in the current frame, mark targets with an IoU matching degree lower than the threshold, and define the judgment condition as follows:

[0109]

[0110] in, Let i be the predicted bounding box of the i-th trajectory. To detect bounding boxes, extensive experiments have verified that when the IoU matching threshold c=0.7, it can effectively distinguish between sudden changes in target motion state and normal jitter.

[0111] If this condition is met, the target motion state is considered stable, the appearance features of the previous frame are directly reused, the feature extraction of the current frame is skipped, and the target with an IoU matching degree lower than the threshold is marked.

[0112] S32, for targets with an IoU matching degree below the threshold, further calculate their confidence change rate, specifically defined as follows:

[0113]

[0114] in, Detect the confidence level for the current frame. The confidence level is the same as the previous frame's confidence level; the confidence level change rate threshold is determined based on the statistical distribution (90th percentile) of significant changes in the target's appearance. ;

[0115] If satisfied If the condition is not met, the target's appearance is considered to have changed significantly, and it is marked; otherwise, the target's appearance is considered not to have changed significantly, and there is no need to re-extract features.

[0116] Step S4: Construct a lightweight feature re-identification network, Rep-OSNet, as follows: Figure 4 As shown, targets marked in step S3 above with an IoU matching degree lower than the threshold and a confidence change rate higher than the threshold are obtained, and appearance features are extracted, specifically including the following steps:

[0117] S41 introduces the reparameterized structure SAG into the OSNET network to construct Rep-OSNet.

[0118] S42 and SAG have multi-branch convolutional structures during the training phase, such as... Figure 5 As shown;

[0119] Branch 1 generates spatial attention weights using 3×3 convolution, average pooling, multilayer perceptron (MLP), and sigmoid.

[0120] Branch 2 is a 3×3 depthwise separable convolution;

[0121] The output is the weight of branch 1 × the feature of branch 2 + the original input.

[0122] During the inference phase, the multi-branch structure is equivalently transformed into a single 3×3 convolution, reducing computational complexity and resource consumption. The parameter fusion formula is as follows:

[0123]

[0124] in, The weight matrix for the 3×3 depthwise separable convolution in branch 2. The weight parameters of the MLP in branch 1, This is the attention weight matrix after the Sigmoid activation function, used to spatially weight the features of branch 2.

[0125] Step S5 involves designing an adaptive cost function that integrates motion direction and appearance similarity to perform cascaded matching between the detected target and the confirmed trajectory; specifically, it includes the following steps:

[0126] S51, normalize the calculation for multiple matching items, calculate the IoU between the trajectory prediction box and the detection box, and define the normalized IoU distance as:

[0127]

[0128] in, Let i be the predicted bounding box of the i-th trajectory. To detect bounding boxes;

[0129] Extract the appearance feature vector of the detected target, calculate its cosine similarity with the trajectory history feature database, and define the appearance distance as:

[0130]

[0131] in, To detect the appearance features of the target, The k-th feature in the trajectory history feature database;

[0132] Based on the historical trajectory movement direction and the relative position of the detection box, the direction difference angle is calculated and normalized as follows:

[0133]

[0134] in, To predict the direction angle for the trajectory, The azimuth angle of the detection box relative to the position of the previous frame on the trajectory.

[0135] S52, sets a dynamic weight adaptive mechanism based on trajectory loss time. Frames and maximum allowable loss time Calculate the appearance feature weighting coefficient α:

[0136]

[0137] in, Let i be the time when trajectory i was lost. To adapt the time loss threshold, it is dynamically adjusted according to the scenario. Let be the velocity variance of trajectory i in the most recent 5 frames; γ=0.05 is the velocity variance adjustment factor.

[0138] By fusing motion and appearance matching terms, the final matching cost function is generated.

[0139]

[0140] in, , And all distance items take values ​​in the range [0,1].

[0141] S53, Set the conditions for enabling direction matching. To reduce computational overhead, the direction matching degree... Enabled only when all of the following conditions are met:

[0142] Track loss time ;

[0143] average velocity of trajectory 1.5 × target average size / frame rate;

[0144] The target was not marked as stationary (historical velocity variance) In other cases, and , .

[0145] in, and 1.5 × average target size / frame rate, derived through optimization using the target kinematics model.

[0146] S54 performs adaptive parameter optimization, dynamically adjusting online and updating every 30 frames. :

[0147]

[0148] λ=1 was determined through cross-validation. Using the MOT20 and VisDrone datasets, the optimal value of λ was selected within the range [0.5, 1.5] using a grid search method, with the tracking success rate and frame rate (FPS) as optimization objectives. The average target size was the average width of all detection boxes in the current frame, and the average target velocity was the average of the displacement velocities of all trajectories.

[0149] Step S6 involves performing a secondary association matching on the targets and trajectories that have mismatched in the cascaded matching; specifically, it includes the following steps:

[0150] S61, obtain the target and trajectory of mismatch in the cascade matching stage.

[0151] S62, uses the CMC algorithm to update the position information of the detection boxes in the association matching algorithm, and updates the position of the detection boxes. The CMC algorithm is applied for compensation, and this process is defined as follows:

[0152]

[0153] In the above formula, This represents the estimated value of the target feature obtained after compensation. This indicates the original detection box position. Represents the motion compensation matrix. This represents the motion compensation term, which is used to correct the position of the detection box.

[0154] S63, set the IoU threshold T=0.5 for secondary matching. When the IoU value between the detection box and the trajectory is greater than 0.5, trigger bidirectional matching verification.

[0155] For each trajectory, calculate its IoU value with all bounding boxes. When the IoU value of a bounding box with this trajectory is greater than a threshold T, the bounding box is considered a candidate matching box for that trajectory.

[0156] For each detection box, calculate its IoU value with all trajectories. When the IoU value of a trajectory with the detection box is greater than a threshold T, the trajectory is considered a candidate matching trajectory for that detection box.

[0157] A detection box is considered to be successfully matched with a trajectory only when it is a candidate matching box for a trajectory in forward matching and the trajectory is also a candidate matching trajectory for the detection box in reverse matching; otherwise, the match is unsuccessful.

[0158] Step S7 involves updating the matched trajectory status using a noise-adaptive Kalman filter, deleting trajectories that still mismatch after secondary association matching and whose loss time exceeds a threshold, and outputting the updated target trajectory prediction box and its ID. Specifically, this includes the following steps:

[0159] S71, a noise adaptation mechanism is added to the Kalman filter algorithm by introducing a trajectory influence factor to adjust the observation noise covariance. This process is defined as follows:

[0160]

[0161] in, As the trajectory influencing factor, It is to test the confidence level. This is the preset minimum noise covariance.

[0162] To control computational complexity, this formula uses interval division. The value is dynamically smoothed, and the value selection method is defined as follows:

[0163]

[0164] in, This represents the number of consecutive confirmations of the trajectory.

[0165] Extensive experimental verification revealed that... , , , The system can achieve an optimal balance between target tracking accuracy and computational efficiency. Specifically:

[0166] The minimum number of consecutive confirmation frames required for the initial stabilization of the target trajectory;

[0167] The threshold for a highly reliable trajectory;

[0168] and The results were obtained by fitting historical data using a gradient optimization algorithm.

[0169] S72 updates the state of the matching trajectories obtained from cascaded matching and secondary association matching.

[0170] S73 sets the threshold for the number of frames lost in a trajectory to 30. When a secondary associated matching trajectory loses more than 30 consecutive frames, the trajectory is deleted.

[0171] The threshold of 30 frames for trajectory loss takes into account the maximum effective recall time of the target re-identification algorithm and the optimal value is determined by the ROC curve.

[0172] S74 updates the appearance model of each tracker to adapt to changes in the target's appearance.

[0173] S75 outputs the trajectory of each tracked target, including the target's ID, bounding box, and confidence level.

[0174] Obviously, the above steps of the present invention are merely examples to clearly illustrate the technical solution of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A method for tracking targets using an unmanned aerial vehicle (UAV), characterized in that, include: S1, UAV target detection: Obtain the bounding box and confidence score of the UAV target in the current frame through the UAV target detector; S2, State Prediction: The position and motion state of the target in the current frame are predicted using the noise adaptive Kalman filter algorithm; S3, Feature Extraction Decision: A confidence-based feature extraction decision strategy is proposed. The trajectory and detection box are matched by IoU, and the confidence change rate of the target is calculated. Targets with IoU matching degree below the threshold and confidence change rate above the threshold are marked. Targets with IoU matching degree above the threshold or confidence change rate below the threshold are reused from the previous frame. S4, Extract appearance features: By constructing a lightweight feature re-identification network Rep-OSNet, appearance features are extracted for targets in the current frame whose IoU matching degree is lower than the threshold and whose confidence change rate is higher than the threshold. S5, Cascaded Matching: By designing an adaptive cost function that integrates motion direction and appearance similarity, the detected target and the confirmed trajectory are cascaded matched. S6, Secondary Association Matching: Perform secondary association matching on targets and trajectories that have mismatched in cascaded matching; S7, State Update: Update the matching trajectory state using noise adaptive Kalman filter, delete trajectories that still mismatch after secondary association matching and whose loss time is greater than the threshold; output the updated target trajectory prediction box and ID.

2. The UAV target tracking method as described in claim 1, characterized in that: The confidence-based feature extraction and determination strategy includes performing IoU matching between the trajectory and the detection box, calculating the confidence change rate, and directly using the feature information of the previous frame for targets with an IoU matching degree higher than the threshold or a confidence change rate lower than the threshold. For targets with an IoU matching degree lower than the threshold and a confidence change rate higher than the threshold, they are marked.

3. The UAV target tracking method as described in claim 1, characterized in that: The lightweight feature re-identification network Rep-OSNet includes the introduction of a reparameterized structure. The Rep-OSNet network sets up a reparameterized structure SAG, which has a multi-branch convolutional structure during the training phase, including an input layer, an average pooling layer, an MLP, an activation function Sigmoid, and an output layer. During the inference phase, it is fused into a single-path 3×3 convolution, reducing computational complexity and memory usage.

4. The UAV target tracking method as described in claim 1, characterized in that, In the cascaded matching process, a cost function that dynamically fuses motion and appearance features is used. The matching weights are adaptively adjusted based on trajectory loss duration, motion direction consistency, and target velocity stability. This process specifically includes: S4-1, Calculation of IoU distance, appearance distance, and normalized orientation difference angle: Define the normalized IoU distance as: in, Let i be the predicted bounding box of the i-th trajectory. To detect bounding boxes; Define the appearance distance as: in, To detect the appearance features of the target, The k-th feature in the trajectory history feature database; Calculate the direction difference angle and normalize it as follows: in, To predict the direction angle for the trajectory, The azimuth angle of the detection box relative to the position of the previous frame on the trajectory; S4-2, setting up a dynamic weight adaptive mechanism, including: Calculate the weighting coefficient for appearance features: in, Let i be the time when trajectory i was lost. To adapt the time loss threshold, it is dynamically adjusted according to the scenario. Let γ be the velocity variance of trajectory i in the most recent 5 frames, and γ = 0.05 be the velocity variance adjustment factor. S4-3, Set the conditions for enabling direction matching, only when the trajectory is lost. Average velocity of the trajectory 1.5 × target average size / frame rate and Enabled at time ; S4-4 performs adaptive parameter optimization, updating every 30 frames. : in, λ∈[0.8,1.2] is determined through offline training, the average target size is the average width of all detection boxes in the current frame, and the average target velocity is the mean of all trajectory displacement velocities; S4-5, Dynamic Fusion of Motion and Appearance Matching Items, The final matching cost is defined as follows: in, , And all distance items take values ​​in the range [0,1].

5. The UAV target tracking method as described in claim 1, characterized in that: During the state update process, a noise adaptive mechanism is added to the Kalman filter algorithm, introducing a trajectory influence factor to construct a noise adaptive Kalman filter that balances the confidence level of target detection with historical tracking performance. This process is defined as follows: in, As the trajectory influencing factor, It is to test the confidence level. The minimum noise covariance is preset. To control computational complexity, this formula uses a piecewise function, and the calculation method is as follows: in, This represents the number of consecutive confirmations of the trajectory.

Citation Information

Patent Citations

  • Unmanned aerial vehicle multi-target tracking method based on ByteTrack

    CN116630376A

  • Method and device for target tracking, and storage medium

    EP4310781A1

Cited By

  • Time-sharing detection tracking double-loop method, device and system, and storage medium

    CN122134660A

  • A scale adaptive target tracking method in unmanned airport scene

    CN122657519A