Unmanned aerial vehicle cluster target long-time robust tracking method in low-altitude airspace complex environment
By combining the YOLO detection model, lightweight graph convolutional network, and bidirectional long short-term memory network, and utilizing environmental perception and improved matching metrics, the problem of long-term tracking drift and trajectory breakage of UAV swarm targets in complex low-altitude airspace environments was solved, achieving stable and accurate long-term tracking results.
Patent Information
- Application Number
- CN202511235238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-30
AI Technical Summary
Traditional methods struggle to achieve long-term stable tracking of UAV swarm targets in complex low-altitude airspace environments, especially under conditions of occlusion, scale changes, and complex dynamic interactions, resulting in tracking drift and trajectory breakage issues.
A pre-trained YOLO target detection model is used in conjunction with a lightweight graph convolutional network and a bidirectional long short-term memory network. The detection results are dynamically adjusted through an environment perception mechanism, and data association is performed using a spatial-temporal attention mechanism and an improved UAV_IoU matching index to achieve long-term robust tracking of UAV swarm targets.
It improves the detection quality and matching accuracy of UAV swarm targets, enhances the ability to distinguish small targets, and achieves stable, continuous, and long-term tracking in complex environments.
Smart Images

Figure CN120949800A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target detection and trajectory prediction technology, specifically relating to a long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments. Background Technology
[0002] With the widespread use of low-altitude flying targets (such as drones and small aircraft), their application scenarios in both civilian and military fields are becoming increasingly complex. Low-altitude flying targets are characterized by low flight altitude, high maneuverability, and small scale, and often face complex environmental interference (such as tree obstruction, building reflection, and weather changes). In long-term tracking scenarios, targets may experience challenges such as prolonged occlusion, drastic scale changes, non-rigid deformation, and sudden changes in viewpoint. Traditional tracking methods based on Kalman filtering or kernel correlation filtering, due to their reliance on fixed state transition models or fixed learning rate update strategies, are difficult to adapt to the need for re-detection after the target has disappeared for a long time, and their ability to model dynamic interaction relationships (such as formation flying and obstacle avoidance behavior) is insufficient, leading to frequent problems such as tracking drift and trajectory breakage.
[0003] While existing short-term tracking algorithms can handle fast-moving targets, they are susceptible to sudden changes in target appearance or occlusion in long-term scenarios, and lack the ability to deeply mine temporal dependencies. Furthermore, traditional methods do not fully consider the dynamic interactions between multiple targets (such as cooperative obstacle avoidance and path adjustment in swarm flight), making it difficult to achieve stable and continuous long-term tracking in complex airspace environments. Therefore, there is an urgent need for a long-term tracking method that integrates temporal feature modeling and target relationship learning to overcome the accuracy and robustness bottlenecks of existing technologies in long-term tracking of targets flying in low-altitude airspace. Summary of the Invention
[0004] Therefore, the purpose of this invention is to provide a long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments, in order to solve the problems existing in the prior art.
[0005] The technical solution provided by this invention is: a long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments, which involves real-time acquisition of flight videos of the UAV swarm to be tracked and real-time processing of the current frame and the next frame to achieve tracking of the UAV swarm. The processing of the current frame and the next frame includes the following steps:
[0006] S1: Use the pre-trained YOLO target detection model to perform UAV target detection on the current frame and obtain the target detection result. The target detection result includes 4-dimensional normalized position information, 1-dimensional target category information and 1-dimensional target detection box confidence score information. The target category corresponding to the UAV target is 0.
[0007] S2: Based on the environmental state information of the current frame, dynamically adjust the detection results of the target in the current frame to obtain the final target detection result that meets the target detection box confidence threshold corresponding to the environmental state of the current frame;
[0008] S3: The final target detection result is input in parallel into the pre-trained lightweight graph convolutional network model and the bidirectional long short-term memory network module to obtain the spatial interaction features and temporal dynamic features of the target's predicted position in the next frame, respectively. Then, the spatial interaction features and temporal dynamic features are fused through a spatial-temporal attention mechanism to obtain the target's predicted position in the next frame.
[0009] S4: Based on the confidence score of the target detection box, perform a two-stage data association between the predicted location of the target and the actual detection result of the target.
[0010] Preferably, in S2, the current frame environment state is obtained by the environment perception module. The specific acquisition method is as follows: First, extract features that can represent visibility and weather conditions from the current frame. Then, input the features into a lightweight machine learning model to obtain the classification label of the environment state. Finally, based on the label, dynamically adjust the confidence threshold of the target detection box and use it to adaptively filter the detection results to obtain the final target detection result.
[0011] Further preferably, in S3, the lightweight graph convolutional network model is a two-layer graph convolutional network structure built based on the node feature matrix and the adjacency matrix, wherein the first layer of the graph convolutional network uses the node feature matrix H nx6 Using the adjacency matrix A as input, the node feature matrix is constructed as follows: each detected target is considered as a node in the graph, and the feature vector of each node is a 6-dimensional feature vector composed of the target detection results. n is the total number of UAV targets in the current frame. The adjacency matrix A is constructed based on the Euclidean distance between nodes. If the distance between node i and node j is less than a preset threshold, then element A in the adjacency matrix A is... ij =1, indicating that there is a connection between node i and node j; otherwise, A ij =0 indicates that there is no direct connection between node i and node j;
[0012] The lightweight graph convolutional network model is as follows:
[0013] In the first layer of the graph convolutional network, for each node i, k nodes are sampled from its neighbors, and the feature information of its neighbors is summarized by the mean aggregation method to obtain the aggregated neighbor representation. Then, the aggregated neighbor representation is concatenated with the node's own features, and nonlinear updates are performed through linear transformation and ReLU activation function to obtain the node representation output by the first layer of the graph convolutional network, i.e., the local spatial interaction features.
[0014] The second-layer graph convolutional network takes the local spatial interaction features output by the first-layer graph convolutional network and the original adjacency matrix A as input, and repeatedly performs neighbor sampling, feature aggregation and node update operations to further capture high-order neighborhood information, and finally outputs the high-order spatial interaction features of each target.
[0015] In a further preferred embodiment, in S3, the bidirectional long short-term memory network is used to model the long-term temporal dynamics of the target, taking the historical position of the target for T consecutive frames as input and outputting the temporal dynamic features of the target's predicted position.
[0016] Specifically, the input to the bidirectional long short-term memory network is a time-series tensor constructed from the historical positions of each target across T consecutive frames, where the time-series of the historical positions corresponding to the i-th target is X. i ∈R Tx6 It contains 6 normalized features, namely: (x t y t ), (w t h t ), class t and ID t , where (x t y t (w) represents the center coordinates of the target bounding box. t h t ) represents the width and height of the detection box, class t For the target category, id t A unique identifier for the target, used to ensure consistency across frame sequences;
[0017] The bidirectional long short-term memory network includes two layers of Bi-LSTM. Each Bi-LSTM consists of a forward LSTM and a backward LSTM. The forward LSTM processes the sequence in chronological order, and the backward LSTM processes the sequence in reverse chronological order. The outputs of the two are concatenated in the channel dimension to obtain a hidden state representation that integrates bidirectional temporal dependencies.
[0018] The two-layer Bi-LSTM is as follows:
[0019] First layer Bi-LSTM: Input is the historical location sequence X∈R for each target. n×T×6 Each frame outputs 64-dimensional temporal features. The formula is as follows:
[0020]
[0021] Among them, W res ∈R 64×6, is the linear projection matrix; ReLU is the activation function; Bi_LSTM represents a bidirectional long short-term memory network;
[0022] Second layer Bi-LSTM: The output of the first layer Bi-LSTM As input to the second layer of Bi-LSTM, maintaining a 64-dimensional output per frame, the output tensor in, The formula is as follows:
[0023]
[0024] Next, time-average pooling will be used to... Aggregated into a single vector, where the final temporal dynamic features of the i-th target are... The formula is as follows:
[0025]
[0026] Then, the final temporal dynamic feature H is constructed using the set of temporal dynamic features of all targets. LSTM ∈R n×64 .
[0027] Further optimization, in S3, the method for fusing the spatial interaction features and temporal dynamic features using a spatial-temporal attention mechanism to obtain the target prediction location is as follows:
[0028] S301: For each target in the current frame, its spatial interaction features and temporal dynamic features are concatenated along the channel dimension to form the joint feature vector of the target;
[0029] S302: The joint feature vector is input into a lightweight multilayer perceptron to automatically learn the importance weight of each target in the fusion process. The lightweight multilayer perceptron consists of two fully connected layers and a ReLU nonlinear activation function, and outputs the importance score of each target in the current frame. Then, the softmax function is used to normalize the scores to obtain the attention weight vector for all targets: α = [α1, α2, ..., α...]. n ]∈R n ;
[0030] S303: The joint feature vector of each target is weighted and fused using the attention weight vector to generate the final fused feature;
[0031] S304: The final fused feature is used as input to a fully connected prediction head and mapped to the target position coordinate prediction value of the next frame. The fully connected prediction head is composed of multiple fully connected layers, and the mean square error is calculated with the real coordinates of the next frame as the loss. Finally, the predicted position of the target in the current frame in the next frame is output.
[0032] Further optimization: In S4, the first-stage data association designs a matching cost function based on UAV_IoU and introduces a direction consistency auxiliary mechanism for detection results and predicted positions with a confidence level higher than 0.7. The successfully matched positions are added to the trajectory management sequence, and the unmatched positions enter the second-stage data association. The second-stage data association is used to process the unmatched predicted positions in the first-stage data association and match them with detection boxes with a confidence level between 0.3 and 0.7. After that, the successfully matched positions are included in the trajectory sequence.
[0033] Further optimization, in S4, the UAV_IoU cost function and the direction consistency auxiliary mechanism are as follows:
[0034] UAV_IoU introduces scale-aware weights on top of EIoU. The formula for UAV_IoU is as follows:
[0035]
[0036] In the formula, IoU represents the intersection-union ratio between the actual object detection box and the predicted tracking box; c represents the squared Euclidean distance between the center point of the actual object detection box and the center point of the predicted tracking box; 2 This represents the square of the diagonal length of the smallest closed bounding box; This represents the width difference penalty divided by the square of the width of the smallest closed frame; The height difference penalty term is divided by the square of the height of the smallest closed box; scale_weight represents the scale-aware weight.
[0037] The direction consistency auxiliary mechanism is as follows: Assume that the historical motion direction of the i-th trajectory is a unit vector. The direction vectors of the detection box in the current frame and the trajectory prediction position in the previous frame are: Calculate the cosine similarity between the two as an indicator of directional consistency: And map it to matching weights:
[0038]
[0039] Then, the matching weights are combined with UAV_IoU to form a new matching cost function.
[0040] Furthermore, the long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments further includes a step of managing the tracking trajectory to update successfully matched trajectories and discard unmatched trajectories. The specific management strategy is as follows:
[0041] 1) If a certain trajectory continuously exceeds N lost If a frame fails to match, it is marked as "missing" and deleted, where N lost This is the default value;
[0042] 2) If an unmatched target detection box appears in the image and its detection box confidence is higher than the preset confidence threshold, it is initialized as a new trajectory and assigned a unique target number.
[0043] 3) Mark the successfully matched trajectory as a tracking status.
[0044] This invention provides a robust long-term tracking method for UAV swarm targets in complex low-altitude airspace environments. By introducing an environmental perception mechanism, it can dynamically adjust detection results based on scene features such as illumination, occlusion, and density, improving the detection quality of weak targets. Simultaneously, it innovatively employs a lightweight graph convolutional neural network to model the spatial interaction relationships between multiple targets and combines it with a bidirectional long short-term memory network to learn the temporal evolution characteristics of trajectories, achieving collaborative modeling between targets and making predictions more accurate and responses more sensitive. In terms of feature fusion, a spatial-temporal fusion attention mechanism is adopted, which enhances the dynamic representation capability of target behavior by adaptively learning the fusion weights between spatial and temporal features, effectively overcoming the performance bottleneck of simple splicing or static weighting in heterogeneous feature fusion in traditional methods. In the data association stage, an improved UAV_IoU matching index is used, introducing scale-aware weights to enhance the discrimination capability of small targets, and combining directional consistency auxiliary weights to improve the matching accuracy of low-confidence targets. Attached Figure Description
[0045] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:
[0046] Figure 1 The flowchart shows the long-term robust tracking method for UAV swarm targets in complex low-altitude airspace provided by this invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] In order to solve the problems existing in the current technology, such as Figure 1 As shown, this invention provides a long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments. It involves real-time acquisition of flight videos of the UAV swarm to be tracked and real-time processing of the current and next frames to achieve tracking of the UAV swarm. The processing of the current and next frames includes the following steps:
[0049] S1: Use the pre-trained YOLO target detection model to perform UAV target detection on the current frame and obtain the target detection result. The target detection result includes 4-dimensional normalized position information, 1-dimensional target category information and 1-dimensional target detection box confidence score information. The target category corresponding to the UAV target is 0.
[0050] S2: Based on the environmental state information of the current frame, dynamically adjust the detection results of the target in the current frame to obtain the final target detection result that meets the target detection box confidence threshold corresponding to the environmental state of the current frame;
[0051] The current frame environment state is obtained by the environment awareness module, and the specific method is as follows:
[0052] By analyzing the visual features of the current frame (such as overall brightness, contrast, sharpness, and noise level), the current environmental state is inferred. Specifically, firstly, features such as average brightness, gradient magnitude, and frequency domain sharpness index are extracted from the current frame as representative variables of visibility and weather conditions. Then, these features are input into a lightweight machine learning model to obtain classification labels for the environmental state (such as "good", "average", or "bad"). Finally, based on the labels, the confidence threshold of the target detection box is dynamically adjusted to adaptively filter the detection results, thereby improving the robustness of detection in complex environments.
[0053] S3: Input the final target detection result in parallel into the pre-trained lightweight graph convolutional network model GraphSAGE and the bidirectional long short-term memory network (Bi-LSTM) module to obtain the spatial interaction features H of the target's predicted position in the next frame. GNN and temporal dynamic features HLSTM Then, the spatial interaction features and temporal dynamic features are fused through a spatial-temporal attention mechanism to obtain the target prediction position in the next frame;
[0054] The lightweight graph convolutional network model GraphSAGE is a two-layer graph convolutional network structure built based on node feature matrices and adjacency matrices. The first layer of the graph convolutional network uses the node feature matrix H... nx6 Using the adjacency matrix A as input, the node feature matrix is constructed as follows: Each detected target is considered as a node in the graph, and the feature vector of each node is a 6-dimensional feature vector composed of the target detection results (including normalized position information (4-dimensional), target category information (1-dimensional), and confidence score (1-dimensional), forming a 6-dimensional feature vector), where n is the total number of UAV targets in the current frame, and the adjacency matrix A ∈ {0, 1}. nxn Based on the Euclidean distance between nodes: if the distance between node i and node j is less than a preset threshold, then element A in the adjacency matrix A is... ij =1, indicating that there is a connection between node i and node j; otherwise, A ij =0 indicates that there is no direct connection between node i and node j;
[0055] The lightweight graph convolutional network model is as follows:
[0056] In the first layer of the graph convolutional network, for each node i, k nodes are sampled from its neighbors, and the feature information of its neighbors is summarized by the mean aggregator to obtain the aggregated neighbor representation. Then, the aggregated neighbor representation is concatenated with the node's own features, and nonlinear updates are performed through linear transformation and ReLU activation function to obtain the node representation output by the first layer, i.e., the local spatial interaction features.
[0057] The second-layer graph convolutional network takes the local spatial interaction features output by the first-layer graph convolutional network and the original adjacency matrix A as input, and repeatedly performs neighbor sampling, feature aggregation and node update operations to further capture high-order neighborhood information, and finally outputs the high-order spatial interaction features of each target.
[0058] The bidirectional long short-term memory network (Bi-LSTM) is used to model the long-term temporal dynamics of the target. During training, the historical position of the target in T consecutive frames is taken as input, and the temporal dynamic features of the target's predicted position in the next frame are output.
[0059] Specifically, the input to the Bidirectional Long Short-Term Memory (Bi-LSTM) network is a time-series tensor constructed from the historical positions of each target across T consecutive frames, where the time-series of the historical positions corresponding to the i-th target is X. i∈R Tx6 It contains 6 normalized features, namely: (x t ,y t ), (w t ,h t ), class t and ID t , where (x t ,y t (w) represents the center coordinates of the target bounding box. t ,h t ) represents the width and height of the detection box, classt represents the target category, and idt represents the unique identifier of the target, used to ensure consistency across frame sequences;
[0060] The Bidirectional Long Short-Term Memory Network (Bi-LSTM) consists of two Bi-LSTM layers. Each Bi-LSTM layer is composed of a forward LSTM and a backward LSTM. The forward LSTM processes the sequence in chronological order t = 1 → T to capture the "influence of the past on the present". The backward LSTM processes the sequence in reverse chronological order t = T → 1 to capture the "contextual cues of the future state on the present". The outputs of the two are concatenated in the channel dimension to obtain a hidden state representation that integrates bidirectional temporal dependencies.
[0061] The two-layer Bi-LSTM is as follows:
[0062] First layer Bi-LSTM: Input is the historical location sequence X∈R for each target. n×T×6 Each frame outputs 64-dimensional temporal features. Since the input feature dimension is low (6-dimensional), a linear mapping is introduced after the LSTM output to expand the input dimension to 64-dimensional, and a residual connection is performed with the LSTM output to obtain:
[0063]
[0064] Among them, W res ∈R 64×6 , is the linear projection matrix; the ReLU activation function is used to enhance the nonlinear expressive power; the input tensor X is dimension-matched in the channel dimension through linear transformation to achieve the addition operation with the Bi-LSTM output;
[0065] Second layer Bi-LSTM: The output of the first layer Bi-LSTM As input to the second layer of Bi-LSTM, maintaining a 64-dimensional output per frame, the output tensor Similarly, by introducing a residual structure to enhance the stability of temporal features and the information transmission capability, we obtain:
[0066]
[0067] Next, temporal average pooling is used to extract the dynamic features of the entire time series. Aggregated into a single vector, where the final temporal dynamic features of the i-th target are... The formula is as follows:
[0068]
[0069] This operation averages along the time dimension, effectively compressing the time dimension while preserving long-term dependency information;
[0070] The final temporal dynamic feature H is constructed by using the set of temporal dynamic features of all targets. LSTM ∈R n×64 , where H LSTM Each row in the table corresponds to the temporal modeling result of a target, which serves as the input for the subsequent spatial-temporal attention fusion module;
[0071] The method for fusing spatial interaction features and temporal dynamic features using a spatial-temporal attention mechanism to obtain the predicted position of the target (in the next frame) is as follows:
[0072] S301: For each target in the current frame, its spatial interaction features and temporal dynamic features are concatenated along the channel dimension to form the joint feature vector of the target;
[0073] S302: The joint feature vector is input into a lightweight multilayer perceptron to automatically learn the importance weight of each target in the fusion process. The lightweight multilayer perceptron consists of two fully connected layers and a ReLU nonlinear activation function, and outputs the importance score of each target in the current frame. Then, the softmax function is used to normalize the scores to obtain the attention weight vector for all targets: α = [α1, α2, ..., α...]. n ]∈R n This is used to reflect the weight balance ratio of each target when spatial and temporal features are fused;
[0074] S303: The joint feature vector of each target is weighted and fused using the attention weight vector to generate the final fused feature;
[0075] S304: The final fused feature is used as the input of a fully connected prediction head and mapped to the target position coordinate prediction value of the next frame. The fully connected prediction head is composed of multiple fully connected layers, and the mean square error is calculated with the real coordinates of the next frame as the loss. Finally, the predicted position of the target in the current frame in the next frame is output.
[0076] S4: Based on the target detection box confidence score, perform a two-stage data association between the predicted position of the target (in the next frame) and the actual detection result of the target (in the next frame);
[0077] In the first stage of data association, for detection results with a confidence level higher than 0.7 and predicted locations, a matching cost function based on UAV_IoU is designed and a direction consistency auxiliary mechanism is introduced. Successfully matched locations are added to the trajectory management sequence, while unmatched locations enter the second stage of data association. The UAV_IoU cost function and the direction consistency auxiliary mechanism are as follows:
[0078] UAV_IoU introduces scale-aware weights on top of EIoU. Used for weighted compensation of small targets, it retains meaningful comparisons even when IOU is extremely small. The formula for UAV_IoU is as follows:
[0079]
[0080] In the formula, IoU represents the intersection-union ratio between the actual object detection box and the predicted tracking box; c represents the squared Euclidean distance between the center point of the actual object detection box and the center point of the predicted tracking box; 2 This represents the square of the diagonal length of the smallest closed bounding box; This represents the width difference penalty divided by the square of the width of the smallest closed frame; The height difference penalty term is divided by the square of the height of the smallest closed box; scale_weight represents the scale-aware weight.
[0081] To further improve the matching reliability of low-confidence targets, a direction consistency auxiliary mechanism is introduced, assuming that the historical motion direction of the i-th trajectory is a unit vector. The direction vectors of the detection box in the current frame and the trajectory prediction position in the previous frame are: Calculate the cosine similarity between the two as an indicator of directional consistency: Map it to matching weights:
[0082]
[0083] Then, the matching weights are combined with UAV_IoU to form a new matching cost function;
[0084] The second-stage data association process handles the unmatched predicted locations from the first-stage data association and matches them with detection boxes with confidence levels between 0.3 and 0.7. Then, the successfully matched locations are incorporated into the trajectory sequence.
[0085] To maintain the continuity of trajectory status and improve system stability, as an improvement to the technical solution, a step of managing the tracking trajectory is also included to update successfully matched trajectories and discard unmatched trajectories. The specific management strategy is as follows:
[0086] 1) If a certain trajectory continuously exceeds N lost = 50 frames failed to match, marked as "missing" and deleted;
[0087] 2) For an unmatched target detection box in the image with a confidence level higher than 0.5, initialize it as a new trajectory and assign it a unique target number;
[0088] 3) Mark the successfully matched trajectory as a tracking status.
[0089] This method for long-term robust tracking of UAV swarm targets in complex low-altitude airspace environments introduces an environmental perception mechanism to dynamically adjust detection results based on scene features such as illumination, occlusion, and density, thereby improving the detection quality of weak targets. Simultaneously, it innovatively employs a lightweight graph convolutional neural network to model the spatial interaction relationships between multiple targets and combines this with a bidirectional long short-term memory network to learn the temporal evolution characteristics of trajectories, achieving collaborative modeling among targets and making predictions more accurate and responses more sensitive. In terms of feature fusion, a spatial-temporal fusion attention mechanism is adopted, which adaptively learns the fusion weights between spatial and temporal features, enhancing the dynamic representation capability of target behavior and effectively overcoming the performance bottleneck of simple concatenation or static weighting in heterogeneous feature fusion in traditional methods. In the data association stage, an improved UAV_IoU matching index is used, introducing scale-aware weights to enhance the discrimination capability of small targets, and combining directional consistency auxiliary weights to improve the matching accuracy of low-confidence targets.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A robust long-term tracking method for UAV swarm targets in complex low-altitude airspace environments, characterized in that, The system acquires flight video of the drone swarm to be tracked in real time and processes the current frame and the next frame in real time to achieve drone swarm tracking. The processing of the current frame and the next frame includes the following steps: S1: Use the pre-trained YOLO target detection model to perform UAV target detection on the current frame and obtain the target detection result. The target detection result includes 4-dimensional normalized position information, 1-dimensional target category information and 1-dimensional target detection box confidence score information. The target category corresponding to the UAV target is 0. S2: Based on the environmental state information of the current frame, dynamically adjust the detection results of the target in the current frame to obtain the final target detection result that meets the target detection box confidence threshold corresponding to the environmental state of the current frame; S3: The final target detection result is input in parallel into the pre-trained lightweight graph convolutional network model and the bidirectional long short-term memory network module to obtain the spatial interaction features and temporal dynamic features of the target's predicted position in the next frame, respectively. Then, the spatial interaction features and temporal dynamic features are fused through a spatial-temporal attention mechanism to obtain the target's predicted position in the next frame. S4: Based on the confidence score of the target detection box, perform a two-stage data association between the predicted location of the target and the actual detection result of the target.
2. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: In S2, the current frame environment state is obtained by the environment perception module. The specific acquisition method is as follows: First, extract features that can represent visibility and weather conditions from the current frame. Then, input the features into a lightweight machine learning model to obtain the classification label of the environment state. Finally, based on the label, dynamically adjust the confidence threshold of the target detection box and use it to adaptively filter the detection results to obtain the final target detection result.
3. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: In S3, the lightweight graph convolutional network model is a two-layer graph convolutional network structure built based on the node feature matrix and the adjacency matrix. The first layer of the graph convolutional network uses the node feature matrix H... nx6 Using the adjacency matrix A as input, the node feature matrix is constructed as follows: each detected target is considered a node in the graph, and the feature vector of each node is a 6-dimensional feature vector composed of the target detection results. n is the total number of UAV targets in the current frame. The adjacency matrix A is constructed based on the Euclidean distance between nodes. If the distance between node i and node j is less than a preset threshold, then the element Aij in the adjacency matrix A is 1, indicating that there is a connection between node i and node j; otherwise, Aij is not connected. i j = 0 indicates that there is no direct connection between node i and node j; The lightweight graph convolutional network model is as follows: In the first layer of the graph convolutional network, for each node i, k nodes are sampled from its neighbors, and the feature information of its neighbors is summarized by the mean aggregation method to obtain the aggregated neighbor representation. Then, the aggregated neighbor representation is concatenated with the node's own features, and nonlinear updates are performed through linear transformation and ReLU activation function to obtain the node representation output by the first layer of the graph convolutional network, i.e., the local spatial interaction features. The second-layer graph convolutional network takes the local spatial interaction features output by the first-layer graph convolutional network and the original adjacency matrix A as input, and repeatedly performs neighbor sampling, feature aggregation and node update operations to further capture high-order neighborhood information, and finally outputs the high-order spatial interaction features of each target.
4. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: In S3, the bidirectional long short-term memory network is used to model the long-term temporal dynamics of the target, taking the historical position of the target for T consecutive frames as input and outputting the temporal dynamic features of the target's predicted position. Specifically, the input to the bidirectional long short-term memory network is a time-series tensor constructed from the historical positions of each target across T consecutive frames, where the time-series of the historical positions corresponding to the i-th target is X. i ∈R Tx6 It contains 6 normalized features, namely: (x t y t ), (w t ,h t ), class t and ID t , where (x t y t (w) represents the center coordinates of the target bounding box. t h t ) represents the width and height of the detection box, class t For the target category, id t A unique identifier for the target, used to ensure consistency across frame sequences; The bidirectional long short-term memory network includes two layers of Bi-LSTM. Each Bi-LSTM consists of a forward LSTM and a backward LSTM. The forward LSTM processes the sequence in chronological order, and the backward LSTM processes the sequence in reverse chronological order. The outputs of the two are concatenated in the channel dimension to obtain a hidden state representation that integrates bidirectional temporal dependencies. The two-layer Bi-LSTM is as follows: First layer Bi-LSTM: Input is the historical location sequence X∈R for each target. n×T×6 Each frame outputs 64-dimensional temporal features. The formula is as follows: Among them, W res ∈R 64×6 , is the linear projection matrix; ReLU is the activation function; Bi_LSTM represents a bidirectional long short-term memory network; Second layer Bi-LSTM: The output of the first layer Bi-LSTM As input to the second layer of Bi-LSTM, maintaining a 64-dimensional output per frame, the output tensor in, The formula is as follows: Next, time-average pooling will be used to... Aggregated into a single vector, where the final temporal dynamic features of the i-th target are... The formula is as follows: Then, the final temporal dynamic feature H is constructed using the set of temporal dynamic features of all targets. LSTM ∈R n×64 .
5. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: In S3, the spatial interaction features and temporal dynamic features are fused using a spatial-temporal attention mechanism to obtain the target prediction location as follows: S301: For each target in the current frame, its spatial interaction features and temporal dynamic features are concatenated along the channel dimension to form the joint feature vector of the target; S302: The joint feature vector is input into a lightweight multilayer perceptron to automatically learn the importance weight of each target in the fusion process. The lightweight multilayer perceptron consists of two fully connected layers and a ReLU nonlinear activation function, and outputs the importance score of each target in the current frame. Then, the softmax function is used to normalize the scores to obtain the attention weight vector for all targets: α = [α1, α2, ..., α...]. n ]∈R n ; S303: The joint feature vector of each target is weighted and fused using the attention weight vector to generate the final fused feature; S304: The final fused feature is used as input to a fully connected prediction head and mapped to the target position coordinate prediction value of the next frame. The fully connected prediction head is composed of multiple fully connected layers, and the mean square error is calculated with the real coordinates of the next frame as the loss. Finally, the predicted position of the target in the current frame in the next frame is output.
6. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: In S4, the first stage of data association designs a matching cost function based on UAV_IoU and introduces a direction consistency auxiliary mechanism for detection results and predicted positions with a confidence level higher than 0.
7. Successfully matched positions are added to the trajectory management sequence, while unmatched positions enter the second stage of data association. The second stage of data association is used to process the unmatched predicted positions in the first stage of data association and match them with detection boxes with a confidence level between 0.3 and 0.
7. After that, the successfully matched positions are added to the trajectory sequence.
7. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 6, characterized in that: In S4, the UAV_IoU cost function and the direction consistency auxiliary mechanism are as follows: UAV_IoU introduces scale-aware weights on top of EIoU. The formula for UAV_IoU is as follows: In the formula, IoU represents the intersection-union ratio between the actual object detection box and the predicted tracking box; c represents the squared Euclidean distance between the center point of the actual object detection box and the center point of the predicted tracking box; 2 This represents the square of the diagonal length of the smallest closed bounding box; This represents the width difference penalty divided by the square of the width of the smallest closed frame; The height difference penalty term is divided by the square of the height of the smallest closed box; scale_weight represents the scale-aware weight. The direction consistency auxiliary mechanism is as follows: Assume that the historical motion direction of the i-th trajectory is a unit vector. The direction vectors of the detection box in the current frame and the trajectory prediction position in the previous frame are: Calculate the cosine similarity between the two as an indicator of directional consistency: And map it to matching weights: Then, the matching weights are combined with UAV_IoU to form a new matching cost function.
8. The long-term robust tracking method for UAV swarm targets in complex low-altitude airspace environments according to claim 1, characterized in that: It also includes steps for managing tracking trajectories to update successfully matched trajectories and discard unmatched trajectories. The specific management strategy is as follows: 1) If a certain trajectory continuously exceeds N lost If a frame fails to match, it is marked as "missing" and deleted. Where N lost This is the default value; 2) If an unmatched target detection box appears in the image and its detection box confidence is higher than the preset confidence threshold, it is initialized as a new trajectory and assigned a unique target number. 3) Mark the successfully matched trajectory as a tracking status.
Citation Information
Patent Citations
Basketball game goal event prediction method based on graph convolution network and long-short-term memory network
CN111488815A
End-to-end multi-target broiler behavior recognition method fusing space-time attention mechanism
CN117333948A
Target tracking method based on space-time joint attention
CN118096826A
Traffic flow prediction method based on multimode dynamic memory graph convolutional network
CN119107798A
Improved YOLOv8 network model-based flat peach fruit identification method
CN120299030A
Cited By
Multi-target tracking method and system under fast motion view angle of unmanned aerial vehicle
CN121236119A