An unmanned aerial vehicle single-target real-time tracking method, system and edge intelligent device

CN119091328BActive Publication Date: 2026-10-09BEIJING HUAXIAXING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411185842.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-10-09
Estimated Expiration
2044-08-27

AI Technical Summary

Technical Problem

[0007]本发明的主要目的在于提供一种无人机单目标实时追踪方法、系统、边缘智能设备及计算机可读存储介质,旨在解决现有技术中使用无人机进行目标的监控和追踪时,无人机监控识别速度较慢、追踪不稳定、数据传输鲁棒性较低的问题

Benefits of technology

[0041] In this invention, the drone video stream is preprocessed to obtain a preprocessed image. An optimized YOLOv8n-OBB model is used to load a thread, which then detects the preprocessed image to obtain target information. The target information is then subjected to dimensionality upscaling to obtain the target's three-dimensional world coordinates. Based on the three-dimensional world coordinates and the target's orientation, the optimal shooting position for the drone is determined. This optimal shooting position is sent to the drone, which moves to it and acquires the target video stream. Two-dimensional coordinates are calculated using an optimized YOLOv8n-DET model and converted to three-dimensional coordinates. A second optimal shooting position is determined based on these three-dimensional coordinates and sent to the drone to control its movement and tracking. This invention improves the speed and tracking stability of drone monitoring and identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091328B_ABST
    Figure CN119091328B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned plane single target real-time tracking method, system and edge intelligent equipment, the method includes: unmanned plane video stream is preprocessed, and the image after preprocessing is obtained, using optimization yolov8n-obb model loading thread, using thread to detect the image after preprocessing, obtain target information;Target information is converted to dimension, and the three-dimensional world coordinate information of the target in reality is obtained, and the best shooting position of unmanned plane is determined according to the orientation of three-dimensional world coordinate information and target information;The best shooting position is sent to unmanned plane, and unmanned plane moves to the best shooting position, obtains the target video stream of unmanned plane, calculates two-dimensional coordinates by optimization yolov8n-det model, converts two-dimensional coordinates into three-dimensional coordinates, determines the second best shooting position according to three-dimensional coordinates, and the second best shooting position is sent to unmanned plane to control unmanned plane to move and track.The application improves the speed and tracking stability of unmanned plane monitoring identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control technology, and in particular to a method, system, edge intelligent device, and computer-readable storage medium for real-time tracking of a single target on a UAV. Background Technology

[0002] In the field of smart cities, the technology for monitoring and tracking targets mainly relies on cameras, drone patrols, etc. Images and videos obtained by cameras monitoring a certain area are usually received and stored on local servers. Then, AI (Artificial Intelligence) deep learning target recognition models are used to detect and locate targets on the server or locally, thereby achieving target tracking.

[0003] However, existing technologies suffer from problems such as slow monitoring and identification speed, unstable tracking, and low robustness of data transmission. The slow monitoring and identification speed manifests itself in the following ways: after collecting data from the monitored area, current monitoring technologies transmit it to the backend server for target detection. During this data transmission process, data packets may be lost, preventing identification and requiring subsequent retransmission. After target detection, the backend transmits signals to instruct the camera to adjust its pose. The speed at which the target is detected and a response is needed to improve.

[0004] The instability in tracking manifests in the following ways: at a certain moment, the data collected from the surveillance camera's perspective may have low confidence or fail to detect the target at all after being detected by the background object detection model, leading to the conclusion that the target did not appear. In such cases, existing deep learning models may exhibit fatal errors in processing video frames captured from the camera's perspective, resulting in errors in target object recognition.

[0005] Low robustness in data transmission manifests in the fact that the monitoring process must be robust to different poses of the target in the lens (such as when the target first enters the monitoring area, or when the lens is tracking the target), as well as changes in the target's movement just before the lens adjusts. Low-probability target poses may act as anomalies, reducing the performance of target recognition; sudden changes in target direction or acceleration captured by the lens may cause the lens to lag behind the target's movement. Therefore, techniques such as data augmentation and post-detection position prediction are needed to improve robustness.

[0006] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0007] The main objective of this invention is to provide a method, system, edge intelligent device, and computer-readable storage medium for real-time tracking of a single target using a drone, aiming to solve the problems of slow drone monitoring and identification speed, unstable tracking, and low data transmission robustness when using drones for target monitoring and tracking in the prior art.

[0008] To achieve the above objectives, the present invention provides a real-time tracking method for a single target from an unmanned aerial vehicle (UAV), the method comprising the following steps:

[0009] The video stream sent by the drone is acquired, the video stream is decoded to obtain processable video frames, and the processable video frames are preprocessed to obtain a preprocessed image.

[0010] The first preset model and the second preset model are optimized. The optimized first preset model is used to load a thread, and the thread is used to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target.

[0011] The two-dimensional pixel coordinate information is transformed to obtain the three-dimensional world coordinate information of the target in reality. The first optimal shooting position of the drone is determined based on the three-dimensional world coordinate information and the orientation.

[0012] Send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target;

[0013] The system acquires the target video stream obtained by the drone at the optimal shooting position in real time, inputs the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, converts the target real-time coordinate data into target three-dimensional coordinate data, determines the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and sends the second optimal shooting position to the drone to control the drone to perform real-time tracking.

[0014] Optionally, the aforementioned real-time single-target tracking method for unmanned aerial vehicles (UAVs), wherein acquiring the video stream sent by the UAV and decoding the video stream to obtain processable video frames specifically includes:

[0015] Once the drone's camera captures a video stream, the video stream transmitted by the drone via the RTSP or RTP protocol is acquired.

[0016] The video stream is read and decoded using an ffmpeg converter to obtain processable video frames.

[0017] Optionally, the real-time single-target tracking method for UAVs, wherein preprocessing the processable video frames to obtain a preprocessed image specifically includes:

[0018] The OpenCV framework is preloaded, and the OpenCV framework is used to perform geometric transformations on the processable video frame to obtain the transformed video frame.

[0019] The transformed video frame is divided into equal spatial proportions to obtain a preprocessed image, which is a preset number of equal regions.

[0020] Optionally, in the UAV single-target real-time tracking method, the first preset model is the yolov8n-obb model, and the second preset model is the yolov8n-det model;

[0021] The optimization of the first preset model and the second preset model specifically includes:

[0022] The yolov8n-obb model and the yolov8n-det model are respectively subjected to lightweight processing to obtain the lightweight yolov8n-obb model and the lightweight yolov8n-det model;

[0023] The lightweight yolov8n-obb model and the lightweight yolov8n-det model are accelerated using the TensorRT plugin and model acceleration algorithm, respectively, to obtain the optimized yolov8n-obb model and the optimized yolov8n-det model.

[0024] Optionally, the aforementioned real-time single-target tracking method for unmanned aerial vehicles (UAVs) includes, in which the optimized first preset model loading thread is used to detect the preprocessed image to obtain the target's two-dimensional pixel coordinate information and the target's orientation, specifically including:

[0025] The optimized yolov8n-obb model is used to load a preset number of threads, and the preset number of threads are used to perform target detection on the preset number of equal regions respectively.

[0026] Once a target is detected, its two-dimensional pixel coordinates and orientation are obtained.

[0027] Optionally, in the aforementioned real-time single-target tracking method for unmanned aerial vehicles, the step of performing a dimensionality-up transformation on the two-dimensional pixel coordinate information to obtain the target's three-dimensional world coordinate information specifically includes:

[0028] Obtain the drone camera parameters used to capture the video stream, including the camera's factory settings, rotation matrix, and translation matrix;

[0029] Based on the drone camera parameters and the two-dimensional pixel coordinate information, the three-dimensional world coordinate information of the target in reality is obtained by reverse calculation using the formula from the three-dimensional world coordinate system to the two-dimensional pixel coordinate system.

[0030] Optionally, the real-time single-target tracking method for UAVs, wherein determining the first optimal shooting position of the UAV based on the three-dimensional world coordinate information and the orientation specifically includes:

[0031] The optimal shooting position of the drone is calculated based on the three-dimensional world coordinate information;

[0032] The distance that the target may move in the direction is predicted, and the optimal shooting position is adjusted according to the distance to obtain the first optimal shooting position of the drone.

[0033] Furthermore, to achieve the above objectives, the present invention also provides a real-time tracking system for a single target of an unmanned aerial vehicle (UAV), wherein the real-time tracking system for a single target of an UAV includes:

[0034] The video preprocessing module is used to acquire the video stream sent by the drone, decode the video stream to obtain processable video frames, and preprocess the processable video frames to obtain a preprocessed image.

[0035] The target information acquisition module is used to optimize the first preset model and the second preset model, load the thread using the optimized first preset model, and use the thread to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target.

[0036] The optimal shooting position determination module is used to perform dimensionality transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality, and determine the first optimal shooting position of the drone based on the three-dimensional world coordinate information and the orientation.

[0037] The drone movement module is used to send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target.

[0038] The drone tracking module is used to acquire the target video stream obtained by the drone at the optimal shooting position in real time, input the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, convert the target real-time coordinate data into target three-dimensional coordinate data, determine the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and send the second optimal shooting position to the drone to control the drone to perform real-time tracking.

[0039] Furthermore, to achieve the above objectives, the present invention also provides an edge intelligent device, wherein the edge intelligent device includes: a memory, a processor, and a drone single-target real-time tracking program stored in the memory and executable on the processor, wherein when the drone single-target real-time tracking program is executed by the processor, it implements the steps of the drone single-target real-time tracking method as described above.

[0040] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a real-time tracking program for a single UAV target, and when the real-time tracking program for a single UAV target is executed by a processor, it implements the steps of the real-time tracking method for a single UAV target as described above.

[0041] In this invention, the drone video stream is preprocessed to obtain a preprocessed image. An optimized YOLOv8n-OBB model is used to load a thread, which then detects the preprocessed image to obtain target information. The target information is then subjected to dimensionality upscaling to obtain the target's three-dimensional world coordinates. Based on the three-dimensional world coordinates and the target's orientation, the optimal shooting position for the drone is determined. This optimal shooting position is sent to the drone, which moves to it and acquires the target video stream. Two-dimensional coordinates are calculated using an optimized YOLOv8n-DET model and converted to three-dimensional coordinates. A second optimal shooting position is determined based on these three-dimensional coordinates and sent to the drone to control its movement and tracking. This invention improves the speed and tracking stability of drone monitoring and identification. Attached Figure Description

[0042] Figure 1 This is a flowchart of the real-time single-target tracking method for UAVs according to the present invention;

[0043] Figure 2 This is a schematic diagram of the real-time single-target tracking method for UAVs of the present invention;

[0044] Figure 3 This is a schematic diagram of the UAV image transmission and video stream decoding in the UAV single target real-time tracking method of the present invention;

[0045] Figure 4 This is a comparison chart of the execution speed of the yolov8n model before and after acceleration in the real-time single-target tracking method for UAVs of the present invention;

[0046] Figure 5 This is a diagram showing the effect of parallel detection using the optimized yolov8n-obb model in the real-time single-target tracking method of the UAV in this invention;

[0047] Figure 6This is a schematic diagram illustrating the transformation from the world coordinate system to the pixel coordinate system in the real-time single-target tracking method for UAVs of the present invention;

[0048] Figure 7 This is a schematic diagram of a preferred embodiment of the UAV single-target real-time tracking system of the present invention;

[0049] Figure 8 This is a schematic diagram of the operating environment of a preferred embodiment of the edge intelligent device of the present invention. Detailed Implementation

[0050] This application provides a method and related equipment for real-time tracking of a single target using an unmanned aerial vehicle (UAV). To make the purpose, technical solution, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.

[0051] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0052] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0053] The preferred embodiment of the UAV single-target real-time tracking method of the present invention, such as... Figure 1 and Figure 2 As shown, the real-time single-target tracking method for UAVs includes the following steps:

[0054] Step S10: Acquire the video stream sent by the drone, decode the video stream to obtain processable video frames, preprocess the processable video frames to obtain preprocessed images.

[0055] The process of acquiring the video stream sent by the drone and decoding the video stream to obtain processable video frames specifically includes:

[0056] Once the drone's camera captures a video stream, the video stream transmitted by the drone via the RTSP or RTP protocol is acquired.

[0057] The video stream is read and decoded using an ffmpeg converter to obtain processable video frames.

[0058] In this embodiment, the drone integrates a camera and transmits the video stream captured by the camera via RTSP (Real-Time Streaming Protocol) or RTP (Real-Time Transport Protocol). The received video stream is then read and decoded on an edge computing device (not limited to edge computing devices, but can also be other small-scale intelligent computing devices, etc.) using either an ffmpeg converter (ffmpeg is a powerful media file conversion tool, commonly used for transcoding, with a wide range of optional commands, allowing free selection of encoder, video duration, frame rate, resolution, pixel format, sampling format, bitrate, cropping options, number of audio channels, etc.) or the OpenCV framework (OpenCV is a cross-platform computer vision and machine learning software library released under the Apache 2.0 license, which can run on Linux, Windows, Android, and Mac OS operating systems. It is lightweight and efficient—consisting of a series of C functions and a small number of C++ classes) to obtain processable video frames.

[0059] like Figure 3 As shown, the process of decoding the video stream includes: receiving compressed data (the video stream to be decoded), which is usually transmitted in H.264 or H.265 compression format; buffering the compressed data in memory for decoding processing to ensure a smooth decoding process; the decoder parses the data packets in the compressed video stream, which include frame headers and frame data, and the frame data contains encoded information and compressed image data; performing frame separation processing on the parsed frame data to obtain independent frames (such as I-frames, P-frames, and B-frames); the decoder uses motion estimation to recover the differences between frames, performs interpolation based on the previous and subsequent frame data, and restores the quantized coefficients to the original DCT (Discrete Cosine Transform) coefficients, and then uses IDCT (Inverse Discrete Cosine Transform) to convert the frequency domain data back to pixel domain data; converting the pixel domain data to color space from YUV color space to RGB color space for display and processing; finally, decoding is completed to obtain a processable video frame.

[0060] Further, the preprocessing of the processable video frames to obtain the preprocessed image specifically includes:

[0061] The OpenCV framework is preloaded, and the OpenCV framework is used to perform geometric transformations on the processable video frame to obtain the transformed video frame.

[0062] The transformed video frame is divided into equal spatial proportions to obtain a preprocessed image, which is a preset number of equal regions.

[0063] It is understood that the OpenCV framework is preloaded, and the OpenCV framework can be used to preprocess the processable video frames. The result of the preprocessing is that the image is divided into a preset number of equal regions. In this embodiment, preferably, the preset number is 9, that is, the processable video frame is divided into 9 small regions proportionally by the OpenCV framework.

[0064] Additionally, it should be noted that when monitoring with a drone, if the drone's altitude is below a preset threshold, you can choose whether to allow the drone to slowly spin to prevent the target from being cut off by the boundary of a small area (under the premise of a high-altitude perspective, the probability of being cut off by the boundary of a small area is relatively small).

[0065] Step S20: Optimize the first preset model and the second preset model, load the thread using the optimized first preset model, and use the thread to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target.

[0066] It is understood that the first preset model is the yolov8n-obb model, and the second preset model is the yolov8n-det model. The network structure of the yolov8n-obb model (directional boundary recognition model) and the yolov8n-det model (object detection model) mainly consists of three modules: Backbone, Neck, and Head. The Backbone module references the CSPDarkNet-53 network, the Neck module uses PAN-FPN, a high-efficiency and fast dual-stream FPN, and the Head module adopts the YOLOX Decoupled Head network structure.

[0067] Specifically, the yolov8n-obb and yolov8n-det models are optimized by first lightweighting them to obtain lightweight yolov8n-obb and lightweight yolov8n-det models. For example, pruning can be performed on these models. Pruning is a common method for lightweighting models, which reduces the number of parameters and computational complexity by deleting some unimportant connections and neurons.

[0068] Furthermore, the lightweight yolov8n-obb and lightweight yolov8n-det models are accelerated using TensorRT plugins and model acceleration algorithms to obtain optimized yolov8n-obb and yolov8n-det models, respectively. TensorRT is a high-performance deep learning inference optimizer and runtime acceleration library that provides low-latency, high-throughput deployment inference for deep learning applications. Model acceleration algorithms mainly include methods such as reducing model computation, quantization operations, structural reparameterization, using dedicated hardware, data parallelism, pipeline parallel algorithms, and improved sampling algorithms. Figure 4 As shown, Figure 4 This section compares the inference speed of the model on CPU and GPU before and after acceleration using TensorRT. Here, yolov8n.pt represents the original model, yolov8n.onnx represents the intermediate transformation model, and yolov8n.engine represents the TensorRT-accelerated model. It is evident that lightweight acceleration using TensorRT can significantly improve recognition efficiency. Furthermore, the directional boundary recognition model and object detection model can be further optimized by expanding the dataset.

[0069] Furthermore, the step of using the optimized first preset model loading thread to detect the preprocessed image and obtain the target's two-dimensional pixel coordinate information and orientation specifically includes:

[0070] The optimized yolov8n-obb model is used to load a preset number of threads, and the preset number of threads are used to perform target detection on the preset number of equal regions respectively.

[0071] Once a target is detected, its two-dimensional pixel coordinates and orientation are obtained.

[0072] In this embodiment, the optimized yolov8n-obb model loads a preset number of threads. The number of threads must correspond to the number of equal regions mentioned above. Each thread is used to detect one region. Taking nine equal regions as an example, the nine threads detect these nine equal regions in parallel, and the target to be tracked (e.g., ...) is found from these nine equal regions. Figure 5 As shown, Figure 5 The image shows the results of parallel detection using the optimized yolov8n-obb model (where the small vehicle in region 5 is the detected target). The system obtains the 2D pixel coordinates and orientation of the target. As can be seen, after detection, the system instantly calculates the precise pixel position of each target in the original image, ensuring efficient and accurate target localization.

[0073] Step S30: Perform dimensionality transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality, and determine the first optimal shooting position of the UAV based on the three-dimensional world coordinate information and the orientation.

[0074] The step of performing a dimensionality-up transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality specifically includes:

[0075] Obtain the drone camera parameters used to capture the video stream, including the camera's factory settings, rotation matrix, and translation matrix;

[0076] Based on the drone camera parameters and the two-dimensional pixel coordinate information, the three-dimensional world coordinate information of the target in reality is obtained by reverse calculation using the formula from the three-dimensional world coordinate system to the two-dimensional pixel coordinate system.

[0077] In this embodiment, the acquired target information is used as the input of the coordinate dimension transformation algorithm. Based on the algorithm of converting the two-dimensional pixel coordinate system to the three-dimensional world coordinate system, the coordinates of the target information in the real world are obtained.

[0078] Specifically, the camera's intrinsic parameters (camera parameters, factory-defined) and extrinsic parameters (camera pose: rotation matrix, translation matrix) are obtained. Then, the target's world coordinate system position is deduced using the formula for converting from the 3D world coordinate system to the 2D pixel coordinate system. When the target object is on the ground, Zw is the Z-axis value of the target's world coordinates, corresponding to a height of 0. Therefore, the only remaining value is (Xw, Yw, 0), where Xw is the X-axis value of the target's world coordinates, and Yw is the Y-axis value. This coordinate dimensionality transformation can actually be seen as a transformation between two-dimensional spaces.

[0079] like Figure 6As shown, the transformation process from the world coordinate system to the pixel coordinate system is as follows: The world coordinates (Xw, Yw, Zw) are transformed into camera coordinates (Xc, Yc, Zc) through rigid body transformation based on the rotation moment R (3*3 matrix) and translation matrix T (3*1 matrix) in the camera extrinsic parameters. Here, Xc, Yc, and Zc represent the X-axis, Y-axis, and Z-axis values ​​of the camera coordinates, respectively. It should be noted that... Figure 6 The camera extrinsic matrix is ​​padded into a square matrix to facilitate subsequent calculations and obtaining its inverse matrix.

[0080] The camera coordinates are projected using a perspective projection matrix (f represents the camera's focal length) to obtain two-dimensional image coordinates (x, y). Finally, the two-dimensional image coordinates are transformed again based on the camera's intrinsic parameters to obtain pixel coordinates (u, v), where x and y represent the X-axis and Y-axis values ​​of the two-dimensional image coordinates, respectively, and u and v represent the X-axis and Y-axis values ​​of the pixel coordinates, respectively. It should be noted that in... Figure 6 In the image sensor, the camera intrinsic parameter matrix is ​​also completed into a square matrix. fx and fy represent the actual physical size of each pixel in the camera in the x and y directions, respectively. fx and fy are calculated by dividing the focal length f by the physical size of a single pixel on the image sensor. In other words, fx and fy actually reflect the actual length unit corresponding to each pixel of the image sensor.

[0081] The process of deducing the target's world coordinate system position using the formula from the three-dimensional world coordinate system to the two-dimensional pixel coordinate system can be expressed as follows:

[0082]

[0083] Among them, Z c =(Zw+β / α, β=[R -1 T](row=2, column=0), u0 and v0 represent the origin coordinates in the pixel coordinate system, Zw is the Z-axis value in the target world coordinates, and β is the matrix [R -1 The value corresponding to row 2 and column 0 in [T], where α is the matrix... The values ​​corresponding to row 2 and column 0 in the matrix, where row represents the row of the matrix and column represents the column of the matrix.

[0084] Furthermore, determining the first optimal shooting position of the drone based on the three-dimensional world coordinate information and the orientation specifically includes:

[0085] The optimal shooting position of the drone is calculated based on the three-dimensional world coordinate information;

[0086] Calculate the distance the target may move toward the drone based on the orientation information, and adjust the preferred shooting position based on the distance to obtain the first optimal shooting position of the drone.

[0087] It is understandable that, based on the acquired 3D world coordinate information (Xw, Yw, Zw), the target detection conf value (i.e., confidence threshold, an important parameter in target detection that determines the model's confidence score requirement for the detected object; an object will only be detected if its confidence score is higher than the set conf value) is calculated at the shooting position with the highest conf value, which is then used as the better shooting position for the UAV (for changing target positions, there are their own optimal shooting positions, and the relative positions of the two remain unchanged, as has been proven by a large amount of experimental data).

[0088] Considering that the drone needs a certain amount of time to move from high-altitude monitoring to low-altitude tracking, the target may move a certain distance toward the drone based on the orientation information. Then, the target orientation detected by the optimized yolov8n-obb model (the direction perpendicular to the forward direction of the line connecting the pixel coordinates of the target object) is used to calculate the target may move a certain distance toward the drone, that is, to predict the distance that the target is likely to move in this direction, so as to make slight adjustments to the optimal detection position to be transmitted to the drone.

[0089] Step S40: Send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target.

[0090] Specifically, the optimal shooting position is sent to the drone. After receiving the optimal shooting position, the drone calculates the movement signal of this shooting position. Based on the movement signal, the drone moves to the better shooting position and adjusts the corresponding camera pose.

[0091] Step S50: Acquire the target video stream obtained by the UAV at the optimal shooting position in real time, input the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, convert the target real-time coordinate data into target three-dimensional coordinate data, determine the second optimal shooting position of the UAV based on the target three-dimensional coordinate data, and send the second optimal shooting position to the UAV to control the UAV to perform real-time tracking.

[0092] Understandably, when the drone moves to low altitude, the core of the target recognition algorithm is switched to the optimized yolov8n-det model. After the drone reaches the designated optimal position, it begins to detect and track the target in real time. The optimized yolov8n-det model has an extremely fast recognition speed. The backend of the edge computing device continuously performs coordinate transformations through the optimized yolov8n-det model and sends the transformed coordinates to the drone to control the drone to move and track, and repeats the above process.

[0093] In summary, this invention utilizes a drone carrying a high-definition camera and an edge intelligent computing device. The edge intelligent computing device integrates a high-efficiency OpenCV image preprocessing script, a lightweight and accelerated YOLOv8n-obb.engine and YOLOv8n.engine model, and an algorithm for converting 2D pixel coordinates to 3D world coordinates. These components ensure efficient execution of computational tasks and a comprehensive improvement in intelligent analysis capabilities. The drone performs a high-altitude, center-view observation within the monitored area, with a fixed angle of 90°, and slowly rotates to optimize target acquisition. During this process, the edge intelligent computing device receives video frames transmitted from the camera in real time. Each frame undergoes preliminary processing using the OpenCV framework, including geometric transformations and spatial partitioning to divide it into nine equal smaller images, thus preparing it for subsequent model detection.

[0094] Building upon this foundation, this invention utilizes nine parallel threads to simultaneously detect these nine small images. This slow rotation strategy aims to reduce false detections caused by target factor image segmentation, thereby improving the accuracy and comprehensiveness of detection. After a target is detected, a coordinate system transformation algorithm is used to obtain the world coordinate position and orientation of the target object. Then, the UAV moves to the optimal tracking and shooting position at this point, activates the target detection model algorithm, and begins tracking. Before each movement, the target's orientation is used to predict the possible location the target might reach after the UAV moves to the optimal position. This invention effectively solves the problems of low accuracy in target identification and inefficient real-time computation in existing urban surveillance systems. Based on the built-in algorithms and models of edge intelligent computing devices, mounted on UAVs, different target detection models are used at different times to achieve accurate identification and real-time tracking.

[0095] Furthermore, such as Figure 7 As shown, based on the above-described real-time single-target tracking method for UAVs, the present invention also provides a real-time single-target tracking system for UAVs, wherein the real-time single-target tracking system for UAVs includes:

[0096] Video preprocessing module 41 is used to acquire the video stream sent by the drone, decode the video stream to obtain processable video frames, and preprocess the processable video frames to obtain preprocessed images.

[0097] The target information acquisition module 42 is used to optimize the first preset model and the second preset model, load the thread using the optimized first preset model, and use the thread to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target.

[0098] The optimal shooting position determination module 43 is used to perform dimensionality transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality, and determine the first optimal shooting position of the UAV based on the three-dimensional world coordinate information and the orientation.

[0099] The drone movement module 44 is used to send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target.

[0100] The drone tracking module 45 is used to acquire the target video stream obtained by the drone at the optimal shooting position in real time, input the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, convert the target real-time coordinate data into target three-dimensional coordinate data, determine the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and send the second optimal shooting position to the drone to control the drone to perform real-time tracking.

[0101] Furthermore, such as Figure 8 As shown, based on the above-mentioned UAV single-target real-time tracking method and system, the present invention also provides an edge intelligent device, which includes a processor 10, a memory 20 and a network receiver 30. Figure 8 Only some components of the edge smart device are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0102] In some embodiments, the memory 20 may be an internal storage unit of the edge intelligent device, such as a hard drive or memory of the edge intelligent device. In other embodiments, the memory 20 may be an external storage device of the edge intelligent device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the edge intelligent device. Further, the memory 20 may include both internal and external storage units of the edge intelligent device. The memory 20 is used to store application software and various types of data installed on the edge intelligent device, such as the program code installed on the edge intelligent device. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a UAV single-target real-time tracking program 40, which can be executed by the processor 10 to implement the UAV single-target real-time tracking method of this application.

[0103] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the UAV single-target real-time tracking method.

[0104] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the edge smart device and to display a visual user interface. The components 10-30 of the edge smart device communicate with each other via a system bus.

[0105] In one embodiment, when the processor 10 executes the UAV single-target real-time tracking program 40 in the memory 20, the following steps are performed:

[0106] The video stream sent by the drone is acquired, the video stream is decoded to obtain processable video frames, and the processable video frames are preprocessed to obtain a preprocessed image.

[0107] The first preset model and the second preset model are optimized. The optimized first preset model is used to load a thread, and the thread is used to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target.

[0108] The two-dimensional pixel coordinate information is transformed to obtain the three-dimensional world coordinate information of the target in reality. The first optimal shooting position of the drone is determined based on the three-dimensional world coordinate information and the orientation.

[0109] Send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target;

[0110] The system acquires the target video stream obtained by the drone at the optimal shooting position in real time, inputs the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, converts the target real-time coordinate data into target three-dimensional coordinate data, determines the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and sends the second optimal shooting position to the drone to control the drone to perform real-time tracking.

[0111] The step of acquiring the video stream sent by the drone and decoding the video stream to obtain processable video frames specifically includes:

[0112] Once the drone's camera captures a video stream, the video stream transmitted by the drone via the RTSP or RTP protocol is acquired.

[0113] The video stream is read and decoded using an ffmpeg converter to obtain processable video frames.

[0114] Specifically, the step of preprocessing the processable video frames to obtain the preprocessed image includes:

[0115] The OpenCV framework is preloaded, and the OpenCV framework is used to perform geometric transformations on the processable video frame to obtain the transformed video frame.

[0116] The transformed video frame is divided into equal spatial proportions to obtain a preprocessed image, which is a preset number of equal regions.

[0117] Wherein, the first preset model is the yolov8n-obb model, and the second preset model is the yolov8n-det model;

[0118] The optimization of the first preset model and the second preset model specifically includes:

[0119] The yolov8n-obb model and the yolov8n-det model are respectively subjected to lightweight processing to obtain the lightweight yolov8n-obb model and the lightweight yolov8n-det model;

[0120] The lightweight yolov8n-obb model and the lightweight yolov8n-det model are accelerated using the TensorRT plugin and model acceleration algorithm, respectively, to obtain the optimized yolov8n-obb model and the optimized yolov8n-det model.

[0121] The step of using the optimized first preset model loading thread to detect the preprocessed image and obtain the target's two-dimensional pixel coordinate information and orientation specifically includes:

[0122] The optimized yolov8n-obb model is used to load a preset number of threads, and the preset number of threads are used to perform target detection on the preset number of equal regions respectively.

[0123] Once a target is detected, its two-dimensional pixel coordinates and orientation are obtained.

[0124] Specifically, the step of performing a dimensionality-up transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality includes:

[0125] Obtain the drone camera parameters used to capture the video stream, including the camera's factory settings, rotation matrix, and translation matrix;

[0126] Based on the drone camera parameters and the two-dimensional pixel coordinate information, the three-dimensional world coordinate information of the target in reality is obtained by reverse calculation using the formula from the three-dimensional world coordinate system to the two-dimensional pixel coordinate system.

[0127] Specifically, determining the first optimal shooting position of the drone based on the three-dimensional world coordinate information and the orientation includes:

[0128] The optimal shooting position of the drone is calculated based on the three-dimensional world coordinate information;

[0129] The distance that the target may move in the direction is predicted, and the optimal shooting position is adjusted according to the distance to obtain the first optimal shooting position of the drone.

[0130] In summary, this invention provides a method and related equipment for real-time single-target tracking by a drone. The method includes: preprocessing the drone video stream to obtain a preprocessed image; loading a thread using an optimized YOLOv8n-OBB model; using the thread to detect the preprocessed image to obtain target information; performing a dimensionality upscaling transformation on the target information to obtain the target's three-dimensional world coordinates; determining the optimal shooting position of the drone based on the three-dimensional world coordinates and the orientation of the target information; sending the optimal shooting position to the drone; the drone moving to the optimal shooting position; acquiring the target video stream; calculating two-dimensional coordinates using an optimized YOLOv8n-DET model; converting the two-dimensional coordinates to three-dimensional coordinates; determining a second optimal shooting position based on the three-dimensional coordinates; and sending the second optimal shooting position to the drone to control its movement and tracking. This invention effectively solves the problems of low accuracy in target identification and inefficient real-time calculation in existing urban surveillance systems. Based on the built-in algorithms and models of edge intelligent computing devices, mounted on a drone, different target detection models are used at different times to achieve accurate identification and real-time tracking.

[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or edge intelligent device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or edge intelligent device. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or edge intelligent device that includes that element.

[0132] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0133] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for real-time tracking of a single target using an unmanned aerial vehicle (UAV), characterized in that, The aforementioned real-time single-target tracking method for unmanned aerial vehicles includes: The video stream sent by the drone is acquired, the video stream is decoded to obtain processable video frames, and the processable video frames are preprocessed to obtain a preprocessed image. The first preset model and the second preset model are optimized. The optimized first preset model is used to load a thread, and the thread is used to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target. The two-dimensional pixel coordinate information is transformed to obtain the three-dimensional world coordinate information of the target in reality. The first optimal shooting position of the drone is determined based on the three-dimensional world coordinate information and the orientation. Send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target; The system acquires the target video stream obtained by the drone at the optimal shooting position in real time, inputs the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, converts the target real-time coordinate data into target three-dimensional coordinate data, determines the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and sends the second optimal shooting position to the drone to control the drone to perform real-time tracking. The step of preprocessing the processable video frames to obtain a preprocessed image specifically includes: The OpenCV framework is preloaded, and the OpenCV framework is used to perform geometric transformations on the processable video frame to obtain the transformed video frame. The transformed video frame is divided into equal spatial proportions to obtain a preprocessed image, wherein the preprocessed image consists of a preset number of equal regions; The first preset model is the yolov8n-obb model, and the second preset model is the yolov8n-det model; The optimization of the first preset model and the second preset model specifically includes: The yolov8n-obb model and the yolov8n-det model are respectively subjected to lightweight processing to obtain the lightweight yolov8n-obb model and the lightweight yolov8n-det model; The lightweight yolov8n-obb model and the lightweight yolov8n-det model were accelerated using the TensorRT plugin and model acceleration algorithm, respectively, to obtain the optimized yolov8n-obb model and the optimized yolov8n-det model. The step of determining the first optimal shooting position of the drone based on the three-dimensional world coordinate information and the orientation specifically includes: The optimal shooting position of the drone is calculated based on the three-dimensional world coordinate information; Predict the distance the target may move in the direction, and adjust the preferred shooting position based on the distance to obtain the first optimal shooting position of the drone; When monitoring with a drone, if the drone's altitude is below a preset threshold, the drone will be controlled to slowly rotate to prevent the target from being cut off by the boundary of a small area.

2. The real-time tracking method for a single target using a drone according to claim 1, characterized in that, The process of acquiring the video stream sent by the drone and decoding the video stream to obtain processable video frames specifically includes: Once the drone's camera captures a video stream, the video stream transmitted by the drone via the RTSP or RTP protocol is acquired. The video stream is read and decoded using an ffmpeg converter to obtain processable video frames.

3. The real-time tracking method for a single target using a drone according to claim 1, characterized in that, The process of loading the optimized first preset model thread and using the thread to detect the preprocessed image to obtain the target's two-dimensional pixel coordinate information and orientation specifically includes: The optimized yolov8n-obb model is used to load a preset number of threads, and the preset number of threads are used to perform target detection on the preset number of equal regions respectively. Once a target is detected, its two-dimensional pixel coordinates and orientation are obtained.

4. The real-time tracking method for a single target using a drone according to claim 1, characterized in that, The step of performing a dimensionality-up transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality specifically includes: Obtain the drone camera parameters used to capture the video stream, including the camera's factory settings, rotation matrix, and translation matrix; Based on the drone camera parameters and the two-dimensional pixel coordinate information, the three-dimensional world coordinate information of the target in reality is obtained by reverse calculation using the formula from the three-dimensional world coordinate system to the two-dimensional pixel coordinate system.

5. A real-time single-target tracking system for unmanned aerial vehicles (UAVs), characterized in that, The real-time single-target tracking system for unmanned aerial vehicles (UAVs) is applied to the real-time single-target tracking method for UAVs according to any one of claims 1-4, and the real-time single-target tracking system for UAVs comprises: The video preprocessing module is used to acquire the video stream sent by the drone, decode the video stream to obtain processable video frames, and preprocess the processable video frames to obtain a preprocessed image. The target information acquisition module is used to optimize the first preset model and the second preset model, load the thread using the optimized first preset model, and use the thread to detect the preprocessed image to obtain the two-dimensional pixel coordinate information of the target and the orientation of the target. The optimal shooting position determination module is used to perform dimensionality transformation on the two-dimensional pixel coordinate information to obtain the three-dimensional world coordinate information of the target in reality, and determine the first optimal shooting position of the drone based on the three-dimensional world coordinate information and the orientation. The drone movement module is used to send the first optimal shooting position to the drone and control the drone to move to the first optimal shooting position to perform low-altitude real-time tracking of the target. The drone tracking module is used to acquire the target video stream obtained by the drone at the optimal shooting position in real time, input the target video stream into the optimized second preset model for calculation to obtain the target real-time coordinate data, convert the target real-time coordinate data into target three-dimensional coordinate data, determine the second optimal shooting position of the drone based on the target three-dimensional coordinate data, and send the second optimal shooting position to the drone to control the drone to perform real-time tracking.

6. An edge intelligent device, characterized in that, The edge intelligent device includes: a memory, a processor, and a drone single-target real-time tracking program stored in the memory and executable on the processor. When the drone single-target real-time tracking program is executed by the processor, it implements the steps of the drone single-target real-time tracking method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a real-time tracking program for a single drone target, which, when executed by a processor, implements the steps of the real-time tracking method for a single drone target as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Unmanned aerial vehicle monitoring and tracking method and device, electronic equipment and storage medium

    CN116486290A