Cross-camera target detection and tracking system and method based on unmanned aerial vehicle
By associating multi-camera video data with GIS maps, combined with appearance features and spatiotemporal constraints, the problem of unstable cross-camera target association in traditional systems is solved, high-precision target detection and tracking is achieved, and global scene analysis and rapid response are supported.
Patent Information
- Application Number
- CN202511115593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional target detection and tracking systems rely on visual data from fixed cameras or single drones, making it difficult to achieve high-precision external parameter calibration in dynamic and complex scenes. The cross-camera target association is unstable, and there is a lack of geographic information integration, making it impossible to achieve global scene analysis and rapid response.
Videos are collected by multiple cameras, pixel coordinates are extracted and associated with points of the same name in the GIS map, and cross-camera target trajectories are associated by combining appearance features and spatiotemporal constraints, and dynamically visualized on the GIS platform.
It achieves high-precision cross-camera target detection and tracking, reduces the ID switching rate, improves the target geolocation accuracy, and supports global scene analysis and rapid response.
Smart Images

Figure CN120599239A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image association models, and in particular relates to a system and method for detecting and tracking targets across cameras of unmanned aerial vehicles. Background Art
[0002] With the rapid development of drone technology, drones are increasingly being used in security monitoring, traffic management, disaster relief, and other fields. Traditional target detection and tracking systems rely heavily on visual data from fixed cameras or single drones. This presents the following issues: Camera extrinsic calibration relies on manual calibration plates. Traditional methods, such as checkerboard calibration, struggle to achieve high-precision extrinsic calibration in dynamic and complex scenes, resulting in large errors in target geolocation. Cross-camera target association is unstable, and reliance on a single appearance feature results in a high ID switching rate, making it impossible to effectively handle perspective differences, occlusions, and lighting changes. The lack of geographic information integration and 3D visualization support for target trajectories makes global scene analysis and rapid response difficult. Summary of the Invention
[0003] The purpose of the present invention is to provide a system and method for detecting and tracking targets across cameras of a drone, which can achieve efficient and accurate tracking of the targets by associating the targets captured by multiple perspective cameras including the drone camera.
[0004] To solve the above technical problems, the present invention is achieved through the following technical solutions: The present invention provides a method for detecting and tracking targets across cameras of unmanned aerial vehicles, comprising: Videos are collected by multiple cameras including drone cameras, and the pixel coordinates of each target in the images from the perspectives of the multiple cameras are obtained by extracting the videos, wherein the images collected and extracted by the drone cameras are drone images; Extract the same-name points in the UAV image and the GIS map, and obtain the geometric transformation relationship between the pixel coordinates in each frame of the UAV image and the GIS longitude and latitude; Continuously detect the target in each frame of the video captured by each camera, and track the target detected by a single camera; Extract the appearance features of the target detected by each camera, and associate the target trajectories across cameras based on spatiotemporal constraints to obtain the associated running trajectory of each target; The associated running trajectory of each target is mapped to the GIS platform for dynamic visualization.
[0005] The present invention also discloses a target detection and tracking system based on a UAV cross-camera, comprising: A drone-assisted calibration module is used to capture video using multiple cameras, including a drone camera, and extract the video to obtain pixel coordinates of each target in images from multiple camera perspectives, wherein the image captured and extracted by the drone camera is a drone image; Extract the same-name points in the UAV image and the GIS map, and obtain the geometric transformation relationship between the pixel coordinates in each frame of the UAV image and the GIS longitude and latitude; The target detection and tracking module is used to continuously detect the target in each frame of the video captured by each camera, and track the target detected by a single camera; The cross-camera association module is used to extract the appearance features of the targets detected by each camera, and associate the target trajectories across cameras in combination with spatiotemporal constraints to obtain the associated running trajectory of each target; The GIS visualization module is used to map the associated running trajectory of each target to the GIS platform for dynamic visualization display.
[0006] The extensible knowledge base module is used to store camera parameters, target feature templates, and historical trajectory data of the target, and perform incremental updates and automatically eliminate custom data.
[0007] The present invention uses a cross-camera association module to match the appearance features of multi-view targets detected by the target detection and tracking module. Combined with spatiotemporal constraints, the cross-camera target trajectory is associated to accurately derive the target's trajectory across camera perspectives. This enables dynamic visualization on the GIS platform.
[0008] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic diagram of the process and modules of an embodiment of a system and method for detecting and tracking targets across cameras of a drone according to the present invention; Figure 2 A schematic diagram of a multi-camera target associated with an embodiment of the present invention; Figure 3 This is a system architecture block diagram of an embodiment of a cross-camera target detection and tracking system based on a drone according to the present invention; Figure 4A schematic diagram of a multi-view detection module in an object detection and tracking module according to an embodiment of the present invention; Figure 5 A schematic diagram of a spatial alignment unit in an object detection and tracking module according to an embodiment of the present invention; Figure 6 Schematic diagram of an external parameter calibration unit in the target detection and tracking module of the present invention in one embodiment; Figure 7 A schematic diagram of the process of the DeepSORT tracking algorithm according to one embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of the improved ResNet50 network according to one embodiment of the present invention; Figure 9 A schematic diagram of an embodiment of the GIS trajectory visualization interface of the present invention; In the accompanying drawings, the components represented by the reference numerals are as follows: 1- UAV assisted calibration module, 2- target detection and tracking module, 3- cross-camera association module, 4- GIS visualization module, 5- scalable knowledge base module. DETAILED DESCRIPTION
[0011] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0012] It should be noted that the terms "first," "second," and the like in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application as detailed in the appended claims.
[0013] See also Figures 1 to 3 As shown, the present invention provides a UAV-based cross-camera target detection and tracking system, which is divided into a UAV auxiliary calibration module 1, a target detection and tracking module 2, a cross-camera association module 3, a GIS visualization module 4 and an extensible knowledge base module 5 from the functional module perspective.
[0014] During operation, the system first executes step S1, where the drone-assisted calibration module 1 captures video from multiple cameras, including the drone camera, and extracts the pixel coordinates of each target in the images from the multiple camera perspectives. The images captured by the drone camera are called drone images, while those captured by the ground camera are called ground camera images. Next, step S2 is executed to extract the same-named points in the drone image and the GIS map, obtaining the geometric transformation relationship between the pixel coordinates in each frame of the drone image and the GIS latitude and longitude.
[0015] The hardware includes one drone equipped with a 4K resolution camera (3840×2160 pixels), flying at an altitude of 91 meters; seven ground-based fixed cameras (1080p resolution, 32.5mm focal length, principal point coordinates (960,540)); and GIS map data (accuracy ≤ 0.1 meter). The drone and ground-based cameras capture video synchronously, with a timestamp alignment error of ≤ 1ms, covering a multi-view monitoring area around the building.
[0016] See also Figure 4 As shown in the figure, during the multi-view detection unit operation in the UAV-assisted calibration module 1, the UAV and ground cameras synchronously capture video, and the pixel coordinates (u, v) of the target are extracted using the improved YOLOv5-SE model. The improved YOLOv5 model is used to extract the bounding box and center coordinates of the target in each frame. The improved YOLOv5 model integrates the SE attention mechanism in the Backbone layer and generates channel attention weights through global average pooling and fully connected layers. The DeepSORT algorithm is then used to achieve single-camera multi-target tracking.
[0017] During drone image preprocessing, Mosaic technology is used to preprocess ground camera images to improve the detection accuracy of small targets (such as pedestrians and vehicles). The ground camera extracts the pixel coordinates of at least 10 targets in each frame to ensure the diversity of the calibration data.
[0018] See also Figure 5 As shown, during the spatial alignment process in the UAV-assisted calibration module 1, the SIFT algorithm is used to extract at least four pairs of points with the same name from the UAV image and the GIS map. The RANSAC algorithm is then used to remove outliers, ensuring a registration error of ≤0.1 pixel. The homography matrix H is calculated to convert the UAV image coordinates into GIS latitude and longitude. q and Q represent the pixel locations of the target in the UAV image and the map image, respectively. H describes the geometric transformation from the UAV perspective to the geographic perspective. Using the transformation matrix H, the pixel coordinates (ud, vd) of the target in the UAV image are converted to the corresponding pixel coordinates (um, vm) of the target in the map image.
[0019]
[0020] See also Figure 6 As shown in the figure, during the operation of the external parameter optimization unit in the UAV-assisted calibration module 1, the image captured by the ground camera is the ground camera image. Combining the pixel coordinates of the ground camera with the GIS geographic coordinates, a nonlinear least squares optimization model is constructed to correct the translation vector (that is, the spatial position of the camera) and the rotation matrix (that is, the spatial angle of the camera) of the ground camera to obtain the external parameters of the ground camera. In a specific application, the pixel coordinates of the ground camera and the GIS geographic coordinates are combined to construct a nonlinear least squares optimization model, and the objective function is:
[0021] in, are the real geographical coordinates of the target, The geographic coordinates are estimated based on the camera's internal and external parameters. After optimization, the accuracy of the external parameters achieved an average geographic error of 0.38 meters across the seven cameras, and a reprojection error of 2.87 pixels. See the experimental results in Table 1 for details.
[0022] Table 1 Experimental results
[0023] Applicability Verification: In both single-camera and multi-camera scenarios, the calibration error meets the design requirements (geographic error ≤ 0.4 meters, reprojection error ≤ 3 pixels).
[0024] The cross-camera object tracking system was deployed using an NVIDIA RTX 3090 GPU, Ubuntu 18.04, and PyTorch 1.10.1. The system consisted of seven 1080p cameras and one drone, covering the area surrounding the building. The drone provided real-time bird's-eye view data. The system also included object detection and tracking, cross-camera correlation, and GIS visualization modules.
[0025] Please continue reading Figure 1 and Figure 7 As shown, the target detection and tracking module 2 in this system is used to execute step S3 to continuously detect the target in each frame of the video captured by each camera, and track the target detected by a single camera. Specifically, the improved YOLOv5-SE model is used to detect targets in real time, with a detection speed of ≥24 frames / second and mAP ≥80%. The DeepSORT algorithm combines Kalman filtering and Hungarian matching to reduce the number of ID switches to 30 per thousand frames. Figure 7 , is the flow chart of DeepSORT tracking algorithm, including motion prediction, Hungarian matching and state update.
[0026] Please continue to refer to Figure 1 、 Figure 2 and Figure 8 As shown, the cross-camera association module 3 in this system is used to execute step S4 to extract the appearance features of the targets detected by each camera, and combine spatio-temporal constraints to perform cross-camera target trajectory association, obtaining the associated running trajectories of each target. The target appearance features are extracted based on the improved ResNet50 network (L2 normalization, feature dimension 2048), combined with spatio-temporal constraints (time window ≤ 5 seconds, speed direction consistency threshold ≥ 0.8). For details, see Figure 8 , which records the structure diagram of the improved ResNet50 network, including the SE attention module and the L2 normalization layer.
[0027] Feature matching formula: ··········Formula (1) Among them, dj represents the jth detected target, and yi and Si respectively represent the predicted position and predicted covariance matrix of the ith trajectory.
[0028] ·······Formula (2) rj represents the feature vector of the jth detection, represents all the feature vectors of the ith trajectory.
[0029] In formula (1), and in formula (2), The weighted combination of the two is used to calculate the final matching degree:In the above process, an improved ResNet50 network (including L2 normalization and Batch Normalization) is used to jointly optimize the feature space through triplet loss and center loss; by integrating spatiotemporal constraints (time window ≤ 5 seconds, speed direction consistency ≥ 0.8, cosine distance ≤ 0.3), the Rank-1 recognition rate on the PRID2011 dataset reaches 77.5%.
[0031] Unmatched targets are processed according to the set rules and may be deleted or continue to be tracked.
[0032] Please continue reading Figure 1 and Figure 9 As shown, the GIS visualization module 4 in this system is used to execute step S5 to map the associated running trajectory of each target to the GIS platform for dynamic visualization. The target trajectory is mapped to the geographic coordinate system, supporting 72-hour historical backtracking and three-dimensional dynamic display. Figure 9 , Figure 9 Schematic diagram of the GIS trajectory visualization interface, showing the 3D trajectory overlay and dynamic timeline filtering functions.
[0033] When an emergency event is triggered, the system can request the edge server (timeout 3 seconds) and the cloud (timeout 5 seconds) in parallel to obtain the fastest response result.
[0034] The EPFL dataset was used for verification. The experimental comparison results of this scheme compared with the traditional SORT (Simple Online and Realtime Tracking) scheme are shown in Table 2.
[0035] Table 2 EPFL dataset performance
[0036] Real-time verification: The system processes 1080p video streams with a delay of ≤50ms, meeting industrial-grade real-time requirements.
[0037] The expandable knowledge base module 5 is used to execute step S6 to store camera parameters, target feature templates, and historical target trajectory data, and to perform incremental updates, automatically removing custom data. The camera parameters, target feature templates, and historical target trajectory data are stored in the knowledge base. The initial knowledge base data in this solution contains 1000 sets of target feature templates (with the top 500 most frequent templates retained). The incremental update test in this solution simulates the addition of 200 sets of target features (including scenarios with occlusion and lighting changes), with a daily update cycle.
[0038] During the update of feature templates in the knowledge base, the newly added target features are extracted using the improved ResNet50 and then the cosine similarity is calculated with the existing templates in the knowledge base. If the similarity is less than 0.3, the template is considered new and stored in the knowledge base; otherwise, it is merged into the similar template group.
[0039] During the retrieval and decision optimization process, the cross-camera association module calls the updated knowledge base and optimizes the target feature similarity threshold from 0.3 to 0.25 to improve association accuracy.
[0040] The experimental results of the performance improvement after incremental update of the extensible knowledge base are detailed in Table 3.
[0041] Table 3 Experimental results of performance improvement after update
[0042] Conclusion: Incremental updates of the knowledge base significantly improve the system's adaptability to complex scenarios.
[0043] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, systems, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and the part for the module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be performed substantially in parallel, and they can sometimes also be performed in the opposite order, depending on the function involved.
[0044] It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by hardware that performs the corresponding function or action, such as a circuit or ASIC (Application Specific Integrated Circuit), or can be implemented by a combination of hardware and software, such as firmware.
[0045] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0046] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for detecting and tracking targets across cameras of unmanned aerial vehicles, characterized in that: include, Videos are collected by multiple cameras including drone cameras, and the pixel coordinates of each target in the images from the perspectives of the multiple cameras are obtained by extracting the videos, wherein the images collected and extracted by the drone cameras are drone images; Extract the same-name points in the UAV image and the GIS map, and obtain the geometric transformation relationship between the pixel coordinates in each frame of the UAV image and the GIS longitude and latitude; Continuously detect the target in each frame of the video captured by each camera, and track the target detected by a single camera; Extract the appearance features of the target detected by each camera, and associate the target trajectories across cameras based on spatiotemporal constraints to obtain the associated running trajectory of each target; The associated running trajectory of each target is mapped to the GIS platform for dynamic visualization.
2. The method for detecting and tracking targets across cameras of a drone according to claim 1, wherein: Also includes, The image captured and extracted by the ground camera is the ground camera image. Combining the pixel coordinates of the ground camera with the GIS geographic coordinates, a nonlinear least squares optimization model is constructed to correct the translation vector and rotation matrix of the ground camera to obtain the external parameters of the ground camera.
3. The method for detecting and tracking targets across cameras of a UAV according to claim 1, wherein: The steps of continuously detecting and obtaining the target in each frame of the video captured by each camera, and tracking the target captured and detected by a single camera, include: An improved YOLOv5 model is used to extract the bounding box and center coordinates of the target in each frame. The improved YOLOv5 model integrates the SE attention mechanism in the Backbone layer and generates channel attention weights through global average pooling and fully connected layers. Single camera multi-target tracking is achieved through the DeepSORT algorithm.
4. The method for detecting and tracking targets across cameras of a UAV according to claim 1, wherein: The step of extracting the appearance features of the target detected by each camera, and correlating the target trajectories across cameras in combination with spatiotemporal constraints to obtain the associated running trajectory of each target, include, Extract the appearance features of the target detected by each camera; The target detected by each camera is matched and associated across cameras by jointly optimizing the feature space through triplet loss and center loss, and the associated trajectory of each target is obtained. The spatiotemporal constraints include a time window of ≤5 seconds and a speed direction consistency threshold of ≥0.
8.
5. The method for detecting and tracking targets across cameras of a UAV according to claim 4 is characterized in that: The step of extracting the appearance features of the target detected by each camera includes: The feature extraction network is an improved ResNet50, which includes an L2 normalization layer and a Batch Normalization layer. The feature dimension is 2048, which is consistent with the original output dimension of the improved ResNet50. The target feature similarity is matched by cosine distance ≤ 0.
3.
6. The method for detecting and tracking targets across cameras of a UAV according to claim 4, wherein: The step of performing cross-camera matching and association of the targets detected by each camera by jointly optimizing the feature space through triplet loss and center loss to obtain the associated running trajectory of each target, include, Construct the feature matching formula between the targets detected by the camera: Where dj represents the jth detected target, yi and Si represent the predicted position and predicted covariance matrix of the i-th track respectively; Among them, rj represents the feature vector of the jth detection, represents all feature vectors of the i-th trajectory; the weighted combination of the two is used to calculate the final matching degree: .
7. The method for detecting and tracking targets across cameras of a UAV according to claim 1 or 4, wherein: The step of extracting the appearance features of the target detected by each camera, and correlating the target trajectories across cameras in combination with spatiotemporal constraints to obtain the associated running trajectory of each target, include, If there are targets that fail to be associated among the targets detected by the camera, the cloud-based LLM reasoning is triggered to generate answers and knowledge update packages, which are then merged with the edge RAG results to output the inferred running trajectory of the target that failed to be associated.
8. The method for detecting and tracking targets across cameras of a UAV according to claim 1, wherein: Also includes, Store camera parameters, target feature templates, and target historical trajectory data, perform incremental updates, and automatically remove custom data; For the target captured by the camera, it is first compared with the stored feature template and historical trajectory data to obtain the associated trajectory of the target.
9. The method for detecting and tracking targets across cameras of a UAV according to claim 1, wherein: Also includes, In the process of detecting the target captured by the camera, continuously detect whether there is an urgent keyword; If so, a target association matching request is sent to the edge server and the cloud at the same time, and the fastest response result is used as the post-association running trajectory of the target.
10. A target detection and tracking system based on a drone with multiple cameras, characterized in that: include, A drone-assisted calibration module is used to capture video using multiple cameras, including a drone camera, and extract the video to obtain pixel coordinates of each target in images from multiple camera perspectives, wherein the image captured and extracted by the drone camera is a drone image; Extract the same-name points in the UAV image and the GIS map, and obtain the geometric transformation relationship between the pixel coordinates in each frame of the UAV image and the GIS longitude and latitude; The target detection and tracking module is used to continuously detect the target in each frame of the video captured by each camera, and track the target detected by a single camera; The cross-camera association module is used to extract the appearance features of the targets detected by each camera, and associate the target trajectories across cameras in combination with spatiotemporal constraints to obtain the associated running trajectory of each target; GIS visualization module, used to map the associated running trajectory of each target to the GIS platform for dynamic visualization display; The extensible knowledge base module is used to store camera parameters, target feature templates, and historical trajectory data of the target, and perform incremental updates and automatically eliminate custom data.
Citation Information
Patent Citations
Unmanned aerial vehicle motion target tracking and positioning method under geographic information space-time constraint
CN105352509A
Multi-view target track generation method and device and electronic equipment
CN112686178A
Multi-target tracking positioning and motion state estimation method based on unmanned aerial vehicle
CN113269098A
Multi-information matching unmanned system cross-camera multi-target tracking method
CN116363694A
Pedestrian geographical trajectory extraction method and device
CN116612493A