Unmanned aerial vehicle-to-ground multi-target real-time tracking method and device

CN122090092BActive Publication Date: 2026-09-18BEIJING ZHONGCEZHIHUI SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610241405.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-28
Publication Date
2026-09-18
Estimated Expiration
2046-02-28

AI Technical Summary

Technical Problem

[0006]有鉴于此,本发明提供一种无人机对地多目标实时跟踪方法及装置,以解决或缓解现有技术中存在的技术问题,至少提供一种有益的选择

Benefits of technology

一、本发明通过在目标检测网络中引入基于Transformer的预测头结构,并结合多尺度特征表达,有效增强了对小尺度目标的特征建模能力,克服了无人机高空视角下目标尺寸小、细节信息不足而导致检测性能下降的问题,适用于复杂无人机航拍场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090092B_ABST
    Figure CN122090092B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV), comprising the following steps: S1, image acquisition: periodically acquiring ground image data through the UAV's onboard camera and transmitting it to the onboard processing unit; S2, image preprocessing: performing distortion correction on the acquired images to obtain preprocessed image frames; S3, target detection: processing the image frames using an improved TPH-YOLOv5 detection network, outputting target category, confidence level, and bounding box coordinates; the network uses CSPDarknet as the backbone, introduces a Transformer Encoder with two-dimensional relative position bias in the Neck layer, and uses a P3 / P4 / P5 multi-scale prediction head in the Head. This invention effectively enhances the feature modeling capability for small-scale targets by introducing a Transformer-based prediction head structure into the target detection network and combining it with a multi-scale feature expression mechanism. It overcomes the problem of decreased detection performance caused by small target size and insufficient detail information in the UAV's high-altitude view, and is suitable for complex UAV aerial photography scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and target tracking technology, and in particular to a method and apparatus for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV). Background Technology

[0002] In recent years, unmanned aerial vehicle (UAV) technology has experienced rapid development and widespread application in fields such as environmental monitoring, aerial photography, and public security. This has also brought new challenges to multi-target tracking technology from the UAV's perspective. Therefore, the ability to perform real-time and accurate multi-target tracking of UAVs has become a critical technical challenge that urgently needs to be addressed in key areas such as security and defense.

[0003] Despite the progress made in existing visual target tracking technologies, they still face many limitations when applied to drone scenarios. First, drone targets are inherently small, often accompanied by high-speed movement, environmental interference, and drastic scale changes during flight. Second, the flight environment of drones (such as urban skies and complex ground backgrounds) is highly variable, easily affected by drastic changes in lighting conditions, background blur, and temporary occlusion by buildings or trees, placing extremely high demands on the robustness of tracking algorithms.

[0004] Current mainstream visual tracking algorithms struggle to overcome these challenges simultaneously. For example, MOSSE and KFC filtering methods, based on features and classifiers, achieve high-speed tracking through efficient computation in the Fourier domain. Some target tracking algorithms based on Siamese networks (SiamFC, SiamMask, etc.) employ correlation filtering or linear correlation operations, which limit their expressive power when dealing with the nonlinear relationship between the target and the background. While some advanced Transformer-based algorithms enhance global feature extraction, their training heavily relies on general target datasets, resulting in poor adaptability to specific small targets (such as drones) or small-scale specialized datasets. Furthermore, they may neglect crucial fine-grained positional information for tracking due to over-reliance on long-range dependencies.

[0005] Furthermore, in the field of multi-target tracking, the widely adopted "detection-based tracking" (TBD) paradigm (such as SORT or DeepSORT and their variants) is highly dependent on the accuracy of the detector. When a UAV encounters prolonged occlusion, drastic changes in lighting, or rapid changes in attitude, its appearance features are easily rendered ineffective. At the same time, relying solely on motion models such as Kalman filtering for state estimation makes it difficult to accurately predict when the target undergoes abrupt changes in motion, which can easily lead to tracking drift, identity (ID) switching, and even the inability to re-lock on the target after it is lost. To address this, a method and device for real-time tracking of multiple ground targets by unmanned aerial vehicles (UAVs) are proposed. Summary of the Invention

[0006] In view of this, the present invention provides a method and apparatus for real-time tracking of multiple ground targets by unmanned aerial vehicles (UAVs) to solve or alleviate the technical problems existing in the prior art, and at least provides a beneficial alternative.

[0007] The technical solution of this invention is implemented as follows: A method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV), comprising the following steps: S1. Image Acquisition: The UAV's onboard camera periodically acquires ground image data and transmits it to the onboard processing unit. S2. Image preprocessing: Distortion correction is performed on the acquired image to obtain preprocessed image frames; S3. Target Detection: An improved TPH-YOLOv5 detection network is used to process image frames and output the target category, confidence score and bounding box coordinates. The network uses CSPDarknet as the backbone, introduces a Transformer Encoder with two-dimensional relative position bias in the Neck layer, and uses a P3 / P4 / P5 multi-scale prediction head in the Head. S4. State Prediction: Predict the target state based on historical trajectory using Kalman filtering, and construct a state vector. ; S5. Feature Extraction: An improved OSNet network trained with three types of loss is used to extract the target appearance feature vector and calculate the similarity to construct the appearance cost. S6. Inter-frame matching: Based on the intersection-union gating constraint, the correlation matrix is ​​constructed by fusing motion and appearance costs, and the optimal matching is solved by the Hungarian algorithm. S7. Trajectory Update and Identity Restoration: Update the trajectory of successfully matched targets. For unmatched targets that are occluded or have sudden changes in motion, restore the original identity through high-confidence re-identification. S8. Output Results: Output the target tracking trajectory and its corresponding identification.

[0008] More preferably, the Transformer Encoder input of the TPH-YOLOv5 detection network is the P3, P4, and P5 multi-scale fusion features output by the Backbone, and each prediction head is connected in parallel with the Transformer prediction branch and the convolutional prediction branch.

[0009] More preferably, the improved OSNet network includes a Primary OSNet backbone network and an independent high-resolution feature branch. This branch consists of a Stem module and an HRConv module, which does not perform depth downsampling. It is fused with the third-stage features of the OSBlock x2 backbone network through lateral connections, and generates an appearance feature vector through a global multi-scale pooling module.

[0010] Further preferably, the three types of losses include identity classification loss (Softmax Loss), metric learning loss (Triplet Loss), and view class loss (View Class Loss). The view class loss supervises the network to learn the shooting view class, so that the extracted identity features are view-independent.

[0011] More preferably, the Kalman filter employs a constant velocity model, and the state transition matrix is:

[0012] The motion cost is calculated based on the Mahalanobis distance between the predicted bounding box and the detected bounding box. The time interval between adjacent frames.

[0013] More preferably, the association cost matrix is ​​constructed based on the intersection-union ratio to form the initial cost. , For the prediction box, The detection box is set and adjusted based on appearance similarity and motion deviation; the intersection-union ratio (IU) gating constraint threshold is set to 0.3, and items below this threshold are marked as unmatched.

[0014] Further preferably, the occlusion or motion mutation criterion includes a prediction deviation exceeding a preset threshold and / or a number of consecutive unmatched frames exceeding a preset number of frames; the high-confidence re-identification condition includes a feature cosine similarity exceeding a first preset threshold, and a similarity difference with the suboptimal matching result exceeding a second preset threshold.

[0015] More preferably, the onboard processing unit is an embedded GPU platform that accelerates network inference through the TensorRT framework and may employ quantization or pruning optimization to ensure that the inference frame rate meets the real-time tracking requirements.

[0016] A further preferred embodiment, without affecting the implementation of the multi-target real-time tracking function on the UAV, also includes a coordinate transformation step: Combining UAV attitude information, camera intrinsic and extrinsic parameters, and SLAM depth data, through formulas

[0017] fusion depth value, The weighting coefficients are 0 to 1, which convert image coordinates into geographic coordinates.

[0018] The present invention also provides a real-time ground target tracking device for unmanned aerial vehicles, comprising: Image acquisition module, used to periodically acquire ground image data; The airborne processing unit is configured to execute the UAV-based real-time multi-target tracking method, and to complete target detection, state prediction, feature extraction, association matching, trajectory update and identity recovery. The pose and depth information acquisition module is used to acquire UAV attitude, camera parameters, and SLAM depth data; The coordinate transformation module is used to convert image coordinates into geographic coordinates; The data feedback module is used to transmit the target trajectory, identity identifier, and geographic coordinates back to the ground terminal; The ground terminal includes a communication receiving module, a result calculation module, and a display module, which are used to receive and visualize the tracking results.

[0019] The embodiments of the present invention have the following advantages due to the adoption of the above technical solutions: I. This invention introduces a Transformer-based prediction head structure into the target detection network and combines it with multi-scale feature representation, which effectively enhances the feature modeling capability for small-scale targets. It overcomes the problem of decreased detection performance caused by the small target size and insufficient detail information in the high-altitude view of UAVs, and is suitable for complex UAV aerial photography scenarios.

[0020] Second, by introducing an attention mechanism and a global feature modeling approach, this invention enables the detection network to adaptively focus on the salient areas of the target, suppressing the influence of complex ground backgrounds, noise interference, and non-target areas, thereby improving the accuracy and stability of target detection in complex environments.

[0021] Third, this invention combines the motion prediction of Kalman filtering with the global optimal matching strategy of the Hungarian algorithm, and integrates the appearance features extracted by deep learning to achieve a continuous and stable correlation between the detection results and historical trajectories. This effectively reduces the target loss and identity switching problems caused by occlusion, viewpoint changes or short-term detection omissions, and improves the overall stability of the multi-target tracking process.

[0022] Fourth, by introducing high-resolution feature branches, multi-source feature fusion architecture, and view-independent feature learning mechanism, this invention enhances the ability to identify targets under different viewpoints and scales, thereby improving the accuracy of target tracking and identity recovery.

[0023] Fifth, while maintaining high detection and tracking accuracy, this invention enables the overall algorithm to perform real-time inference on an embedded computing platform through model structure optimization, lightweight design, and pruning strategies, thus meeting the requirements of real-time performance and limited computing resources in the flight state of UAVs.

[0024] VI. This invention achieves collaborative operation between the UAV and ground-based systems by combining target detection and tracking results from the UAV with optional coordinate calculation and transmission. The ground-based system can perform map annotation and dynamic updates based on the transmitted target identity and location information, providing reliable data support for subsequent target analysis, situational awareness, and mission decision-making.

[0025] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating the overall system flow of the present invention. Figure 2 This is a schematic diagram of the algorithm flow of the present invention; Figure 3 This is a diagram of the network structure of the present invention; Figure 4 Two diagrams illustrating the network structure of this invention. Detailed Implementation

[0028] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0029] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0030] This invention provides a method and apparatus for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV), which mainly includes two parts: the UAV terminal and the ground terminal.

[0031] It consists of two parts: the UAV end and the ground end, forming a closed-loop link of "data acquisition-processing-tracking-positioning-data transmission-display". The system architecture and data flow are as follows: Figure 1 As shown.

[0032] (I) Core Modules of the Unmanned Aerial Vehicle Image acquisition module: It adopts an airborne camera to periodically acquire real-time image data of the ground, covering multiple types of target scenes such as ships, vehicles, buildings, and personnel, and provides raw data for subsequent processing; Airborne processing unit: NVIDIA Jetson Orin embedded computing equipment is selected, and TPH-YOLOv5 detection network, improved OSNet re-identification network, Kalman filter algorithm and Hungarian matching algorithm are deployed to undertake core computing tasks such as image preprocessing, target detection, state prediction, feature extraction and trajectory update to ensure real-time processing requirements. Pose and depth information acquisition module: used to acquire UAV attitude information and target scene depth data, providing basic parameter support for coordinate transformation; Coordinate transformation module: Based on UAV attitude information, camera intrinsic and extrinsic parameters and depth data, it converts image pixel coordinates into geographic coordinates.

[0033] Data feedback module: Using a wireless communication module, the target identification, tracking trajectory and geographic coordinates are packaged and transmitted back to the ground.

[0034] (II) Ground-based core module Communication receiving module: Establishes communication with the UAV data return module, receives the returned packaged data and completes the parsing.

[0035] The results processing module processes the parsed data, extracts the target ID, geographic coordinates, and trajectory time series information, and completes the format standardization.

[0036] Display module: Equipped with map visualization software, it marks the target location on the electronic map in real time based on geographic coordinates, and dynamically updates the trajectory according to the target ID, realizing simultaneous monitoring and tracing of multiple targets.

[0037] Combination Figure 1 and Figure 2 The specific implementation steps of the real-time tracking method of the present invention are as follows: S1. Start the UAV and establish a communication connection with the ground equipment, and initialize the model parameters and algorithm thresholds of the airborne processing unit. After the UAV takes off, the airborne camera periodically collects ground image data and transmits it to the NVIDIA Jetson Orin airborne processing unit in real time. If no new image frame is detected, the system enters a waiting state; if a new frame is detected, the subsequent processing flow is triggered.

[0038] S2. After receiving the image data, the airborne processing unit performs the following preprocessing operations in sequence to provide high-quality input for target detection: Distortion correction: Based on camera intrinsic parameters, optical distortion is eliminated through corresponding algorithms to ensure the accuracy of target position.

[0039] Brightness equalization: For scenes with variable lighting at high altitudes, adjust the image brightness to avoid excessive brightness or darkness that could lead to loss of target features and improve image quality.

[0040] ROI extraction: Based on the drone's flight altitude and field of view, the effective area of ​​the image is automatically delineated, and invalid pixels at the edges are excluded, reducing the amount of computation and improving processing efficiency.

[0041] S3. In this stage, an improved TPH-YOLOv5 detection network is used to perform target recognition on the preprocessed image frames. The network structure is as follows: Figure 3 As shown, the specific implementation is as follows: 1. Network architecture deployment The TPH-YOLOv5 network consists of three parts: Backbone, Neck, and Head, and is specifically improved to meet the needs of small target detection for UAVs. Backbone: Adopting the CSPDarknet structure, it extracts multi-level features from low-level texture to high-level semantics of images through residual connections and cross-stage local fusion, laying the foundation for multi-scale object detection.

[0042] Neck: Introduces the Transformer Encoder structure to enhance the network's ability to capture long-distance dependencies, improve the global feature representation of small targets, and solve the problem of poor detection performance of traditional algorithms for small targets at high altitudes.

[0043] Head: Employs a three-scale prediction head (P3, P4, P5) and integrates a Transformer prediction module to enhance attention to targets at different scales, enabling multi-scale detection output and obtaining target category, confidence level, and bounding box coordinates.

[0044] 2. Network Training and Testing Execution Dataset Construction: Collect multi-scene imagery data from drones, including various types of targets such as ships, vehicles, buildings, and people. Manually annotate bounding box coordinates and category labels to construct a training dataset.

[0045] Model training: The TPH-YOLOv5 network was trained based on the constructed dataset to optimize network parameters and improve the model's ability to detect targets in drone scenes.

[0046] Detection execution: The preprocessed image frame is input into the trained and deployed TPH-YOLOv5 model, which outputs the target category, confidence score, and bounding box coordinates of the current frame, forming a set of detection boxes. If no legitimate target is found in the detection results (confidence score is below a preset threshold), the system returns to the image acquisition stage and waits for the next frame; if a legitimate target is found, the system enters the tracking stage.

[0047] S4. Continuous target tracking is achieved through the synergistic effect of Kalman filtering, the Hungarian algorithm, and the improved OSNet re-identification network. The specific implementation is as follows: 1. Kalman Filter State Prediction The Kalman filter algorithm is used to predict the target state in historical trajectories, achieving a smooth trajectory estimation. Define the target state vector: ,in, The coordinates of the target center are , For the velocity component.

[0048] State transition model: The state transition matrix

[0049] in, The time interval between adjacent frames. This is process noise; Observation model: The observation matrix

[0050] in, To observe noise.

[0051] Prediction phase: Calculate the prior estimate of the target state in the current frame.

[0052] with prior covariance ,in This is the posterior estimate for the previous frame. The posterior covariance of the previous frame. Let be the process noise covariance.

[0053] Through the above prediction process, the set of prediction boxes for the current frame is obtained, and the motion cost is calculated based on the spatial relationship between the prediction results and the detection boxes.

[0054] S5. Cropping the target image region corresponding to the detection box set, and inputting it into the improved OSNet re-identification network to extract high-dimensional appearance feature vectors. The network structure is as follows: Figure 4 As shown: Network Improvement Design: In addition to the Primary OSNet Backbone, a high-resolution feature branch consisting of the Stem module and the HRConv module is designed independently. This branch does not perform depth downsampling and is directly injected into the back-end tracking module to avoid the loss of spatial location information of small targets.

[0055] Features from the Primary OSNet Backbone intermediate layer (OSBlock x2 third stage) are introduced through lateral connections and combined with the output features of the independent Feature Fusion branch, converging into the global multi-scale pooling module. This integrates pooling results from different receptive fields, adapting to scenarios with drastic changes in target scale.

[0056] During the training phase, Softmax Loss (identity classification), Triplet Loss (metric distance), and View Class Loss (viewpoint classification loss) are jointly optimized. The supervised network learns the shooting viewpoint and extracts viewpoint-independent identity features to combat the interference of frequent drone viewpoint switching.

[0057] The improved OSNet network is trained to extract high-dimensional feature vectors of the target. The cosine similarity between the feature vectors of the target detected in the current frame and the target in the historical trajectory database is calculated to construct the appearance cost matrix.

[0058] S6. The Hungarian algorithm is used to fuse the motion cost matrix and the appearance cost matrix to generate the final association cost matrix, and the optimal matching relationship between the detection box and the prediction box is solved: Construct the cost matrix: The initial cost matrix is ​​constructed based on , where For the first One prediction box, For the first One detection box, This indicates the intersection, union, and comparison of the two.

[0059] Optimization goal:

[0060] The constraints are:

[0061] The globally optimal matching result is obtained by solving the problem using the Hungarian algorithm. This achieves a one-to-one correspondence between detection boxes and prediction boxes.

[0062] S7. Based on the matching results of the Hungarian algorithm, perform trajectory updates and identity management: For a successfully matched target, update the Kalman filter state and perform posterior estimation. and subsequent covariance ,in For Kalman gain, (As an identity matrix), correct the target trajectory and update the historical information of the trajectory.

[0063] For unmatched targets, if occlusion or sudden motion changes are detected that cause motion prediction deviation to exceed a preset threshold, re-identification matching is performed based on the appearance cost matrix. When the high confidence condition is met, the original identity of the target is restored instead of a new identity is assigned.

[0064] For unmatched detection boxes, if they are determined to be new targets, a new identity is assigned, the Kalman filter state and trajectory information are initialized, and the boxes are added to the historical trajectory library.

[0065] Unmatched targets enter a loss count state. When the number of lost frames exceeds a preset threshold, the target trajectory is deleted.

[0066] S8. For targets that have been successfully matched or whose identities have been restored, the image coordinates are converted into geographic coordinates (latitude, longitude, or UTM coordinates) by combining the UAV attitude information, camera intrinsic and extrinsic parameters, and depth data output by the SLAM module, generating a 3D localization result. This result is then fused with the Kalman prediction result using the depth information output by the SLAM module. The fusion formula is as follows:

[0067] in, To predict depth for Kalman spectroscopy, For SLAM module depth, The weighting coefficients (ranging from 0 to 1) are used to calculate geographic coordinates, improving spatial positioning accuracy and geographic registration stability.

[0068] The UAV packagees the target identification, tracking trajectory, and geographic coordinates into a data package and sends it to the ground terminal via a wireless communication module. The ground terminal's communication receiving module receives the data, processes it through a result calculation module, and then the display module marks the target location on a map in real time based on latitude, longitude, or UTM coordinates. It also dynamically updates the target trajectory based on the tracking ID, providing a visual basis for subsequent target analysis and mission decision-making.

[0069] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV), characterized in that: Includes the following steps: S1. Image Acquisition: The UAV's onboard camera periodically acquires ground image data and transmits it to the onboard processing unit. S2. Image preprocessing: Distortion correction is performed on the acquired image to obtain preprocessed image frames; S3. Target Detection: An improved TPH-YOLOv5 detection network is used to process image frames and output the target category, confidence score and bounding box coordinates. The network uses CSPDarknet as the backbone, introduces a Transformer Encoder with two-dimensional relative position bias in the Neck layer, and uses a P3 / P4 / P5 multi-scale prediction head in the Head. S4. State Prediction: Predict the target state based on historical trajectory using Kalman filtering, and construct a state vector. And calculate the motion cost based on the spatial relationship between the prediction results and the detection box; S5. Feature Extraction: An improved OSNet network trained with three types of loss is used to extract the target appearance feature vector and calculate the similarity to construct the appearance cost. The improved OSNet network includes a Primary OSNet backbone network and an independent high-resolution feature branch. This branch consists of a Stem module and an HRConv module, which do not perform depth downsampling. It is fused with the third-stage features of the OSBlock x2 backbone network through lateral connections, and then generates an appearance feature vector through a global multi-scale pooling module. S6. Inter-frame matching: Based on the intersection-union gating constraint, the motion cost and the appearance cost are fused to construct an association matrix, and the optimal matching is solved by the Hungarian algorithm. S7. Trajectory Update and Identity Restoration: Update the trajectory of successfully matched targets. For unmatched targets that are occluded or have sudden changes in motion, restore the original identity through high-confidence re-identification. S8. Output Results: Output the target tracking trajectory and its corresponding identification.

2. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The TPH-YOLOv5 detection network's Transformer Encoder input is the P3, P4, and P5 multi-scale fusion features output by the Backbone, and each prediction head is connected in parallel with the Transformer prediction branch and the convolutional prediction branch.

3. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The three types of losses include identity classification loss (Softmax Loss), metric learning loss (Triplet Loss), and view class loss (View Class Loss). The view class loss supervises the network to learn the shooting view categories, making the extracted identity features view-independent.

4. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The Kalman filter employs a constant velocity model, and the state transition matrix is: ; The motion cost is calculated based on the Mahalanobis distance between the predicted bounding box and the detected bounding box. The time interval between adjacent frames.

5. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The correlation matrix is ​​constructed based on the intersection-union ratio (OCR) to determine the initial cost. , For the prediction box, The detection box is set and adjusted based on appearance similarity and motion deviation; the intersection-union ratio (IU) gating constraint threshold is set to 0.3, and items below this threshold are marked as unmatched.

6. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The occlusion or motion mutation criteria include prediction deviation exceeding a preset threshold and / or the number of consecutive unmatched frames exceeding a preset number of frames; the high-confidence re-identification conditions include feature cosine similarity exceeding a first preset threshold and the similarity difference with the suboptimal matching result exceeding a second preset threshold.

7. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: The onboard processing unit is an embedded GPU platform that accelerates network inference through the TensorRT framework and can optionally employ quantization or pruning optimization to ensure that the inference frame rate meets the real-time tracking requirements.

8. The method for real-time tracking of multiple ground targets by an unmanned aerial vehicle (UAV) according to claim 1, characterized in that: Without affecting the real-time multi-target tracking function of the UAV, a coordinate transformation step is also included: Combining UAV attitude information, camera intrinsic and extrinsic parameters, and SLAM depth data, through formulas ; fusion depth value, The weighting coefficients are 0 to 1, which convert image coordinates into geographic coordinates. in, For the first The fusion depth value at any given moment. For the first Kalman prediction depth at time step For the first The depth of the SLAM module at any given time.

9. A real-time ground target tracking device for unmanned aerial vehicles (UAVs), characterized in that: include: Image acquisition module, used to periodically acquire ground image data; The airborne processing unit is configured to perform the method described in any one of claims 1 to 8, and to complete target detection, state prediction, feature extraction, association matching, trajectory update and identity recovery; The pose and depth information acquisition module is used to acquire UAV attitude, camera parameters and SLAM depth data; The coordinate transformation module is used to convert image coordinates into geographic coordinates; The data feedback module is used to transmit the target trajectory, identity identifier, and geographic coordinates back to the ground terminal; The ground terminal includes a communication receiving module, a result calculation module, and a display module, which are used to receive and visualize the tracking results.

Citation Information

Patent Citations

  • Water surface target detection algorithm based on improved YOLO V5

    CN116912694A

  • Multi-scale pedestrian re-identification method based on human body part information assistance

    CN120412012A