Real-time target detection tracking and prediction system based on edge device
By integrating YOLOv5 and the improved DeepSort algorithm on edge devices, combined with an asynchronous pipeline scheduling module, real-time target detection, tracking, and prediction are achieved. This solves the problems of insufficient real-time performance and poor tracking stability on edge devices, and provides multimodal output and advanced business integration.
Patent Information
- Application Number
- CN202511900336.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies lack an integrated system architecture on edge devices that can deeply integrate high-performance detection and tracking, real-time prediction capabilities, resource-aware scheduling, and multimodal output interfaces. This results in insufficient real-time performance, poor tracking stability, and a lack of motion prediction capabilities in resource-constrained environments.
The YOLOv5 model is optimized using a multi-threaded asynchronous mechanism and the TensorRT framework. Combined with the improved DeepSort algorithm and Kalman filter, the asynchronous pipeline scheduling module achieves deep fusion of target detection, tracking and prediction, and supports multimodal video streams and structured data output.
It enables real-time processing of high-definition video streams on edge devices, ensuring end-to-end latency within 200 milliseconds, improving the ability to maintain identity in situations with target occlusion and complex scenarios, and providing diverse video and data outputs, while supporting advanced business logic integration.
Smart Images

Figure CN121708053A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and edge computing, specifically a real-time target detection, tracking and prediction system based on edge devices. Background Technology
[0002] With the deep integration of artificial intelligence and Internet of Things (IoT) technologies, vision-based intelligent perception systems are being widely applied in key areas such as intelligent security, autonomous driving, and industrial inspection. These applications typically have stringent requirements for real-time performance, environmental complexity, and deployment costs, driving the migration of visual processing tasks from the cloud to the network edge, where computation is performed directly at the edge, such as cameras and embedded devices. Under this trend, single-stage object detection algorithms, represented by the YOLO (You Only Look Once) series, have become the mainstream choice for real-time object detection on edge devices due to their good balance between accuracy and speed. Meanwhile, multi-object tracking algorithms, represented by DeepSort, can maintain the target's identity (ID) in continuous video frames by fusing Kalman filter prediction and deep appearance feature matching, providing a temporal continuity basis for behavior analysis. Furthermore, inference optimization frameworks, represented by TensorRT, can significantly improve the execution efficiency of neural network models on edge GPUs through operations such as layer fusion and precision quantization, making it one of the key technologies for solving the bottleneck of edge computing power.
[0003] However, existing technologies still suffer from a significant architectural flaw in the practical deployment and system integration of edge devices. Most existing solutions simply "stack" or "stitch together" the aforementioned advanced algorithms, that is, independently optimizing the detection model, independently running the tracking module, and then considering the output results separately. This loosely coupled approach fails to design from the perspective of overall system efficiency and functional integrity, often leading to a dilemma of "performance and functionality being mutually exclusive" in resource-constrained edge environments. For example, a system that may have high-precision detection and tracking capabilities may suffer from accumulated processing latency and loss of real-time performance under high frame rate input due to the lack of efficient pipeline scheduling and load management mechanisms; or, a system that guarantees low latency may only output raw video or simple logs, unable to simultaneously provide structured data for machine decision-making and enhanced video streams for human monitoring, greatly limiting its application value in closed-loop automation scenarios. In short, existing technologies lack an integrated system architecture that can deeply integrate high-performance detection and tracking, real-time prediction capabilities, resource-aware scheduling, and multimodal output interfaces, which is the key obstacle restricting edge vision systems from reaching higher levels of intelligence and practicality. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a real-time target detection, tracking, and prediction system based on edge devices, which solves the problems of insufficient real-time performance, poor tracking stability, and lack of motion prediction capabilities caused by loose coupling of algorithm modules in the prior art.
[0005] A real-time target detection, tracking, and prediction system based on edge devices includes:
[0006] The input acquisition module is used to acquire video streams and uses a multi-threaded asynchronous mechanism to acquire and preprocess video frames;
[0007] The inference processing module detects preprocessed video frames based on the optimized YOLOv5 target detection model and extracts the appearance feature vectors of the targets.
[0008] The tracking and prediction module, based on the improved DeepSort tracking algorithm, uses the target's appearance feature vector and detection box information to associate the target ID, and uses a Kalman filter to predict the target's future position and perform trajectory smoothing.
[0009] The asynchronous pipeline scheduling module is used to build a multi-threaded parallel processing framework. It uses a lock-free queue or a circular buffer to transfer data between the input acquisition module, inference processing module, tracking and prediction module, video rendering and streaming module, and data interface service module to achieve asynchronous execution of each module.
[0010] The video rendering and streaming module is used to render video frames that have been overlaid with detection boxes, target categories, target IDs, confidence scores and predicted trajectory information, and output them in real time through streaming media protocols;
[0011] The data interface service module is used to provide structured detection, tracking, prediction, and system status data to external systems in real time via HTTP or WebSocket interfaces.
[0012] Preferably, the input acquisition module includes a dynamic frame skipping processing unit, which is used to dynamically adjust the processing frequency of video frames according to the real-time load and inference time of the edge device.
[0013] Preferably, the inference processing module uses the TensorRT framework to optimize the YOLOv5 detection model and appearance feature extraction network. The optimization methods include at least one of layer fusion, convolution operator optimization, and FP16 half-precision inference.
[0014] Preferably, the tracking prediction module improves the DeepSort algorithm by using a GPU to accelerate the extraction and matching process of appearance feature vectors; adjusting the state noise and observation noise parameters of the Kalman filter; and using an exponential moving average method to smooth the predicted trajectory.
[0015] Preferably, the asynchronous pipeline scheduling module divides the system tasks into acquisition threads, preprocessing threads, inference threads, tracing threads, rendering threads, storage threads, and streaming threads, and each thread communicates through a data queue based on a "producer-consumer" model.
[0016] Preferably, the video rendering and streaming module supports at least one of the streaming media protocols RTSP, HTTP-FLV and WebRTC for video stream output, and supports local storage of processed video data while streaming.
[0017] Preferably, the structured data output by the data interface service module includes at least one of the following: target ID, category, confidence level, bounding box coordinates, historical trajectory point list, velocity vector, predicted coordinates for several future frames, real-time frame rate, current total number of targets, and device resource utilization rate.
[0018] Preferably, when the tracking and prediction module associates target IDs, it simultaneously calculates the IoU distance between the predicted bounding box and the detection bounding box and the cosine distance between the appearance feature vectors, and performs comprehensive matching based on the Hungarian algorithm.
[0019] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the system described above.
[0020] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the system described above.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] By deeply integrating target detection, tracking, and prediction algorithms and systematically optimizing for edge computing environments, significant improvements have been achieved in performance, functionality, and ease of use. Firstly, regarding real-time performance and stability, the TensorRT framework was introduced to optimize the YOLOv5 detection model and ReID feature network through layer fusion and FP16 half-precision inference, significantly reducing model computation latency. Simultaneously, by combining a dynamic frame skipping strategy and a multi-threaded asynchronous pipeline architecture, the system can intelligently schedule resources based on device load, effectively avoiding frame processing backlog. This enables stable processing of high-definition video streams exceeding 25 frames per second on typical edge devices such as Jetson, ensuring end-to-end latency is controlled within 200 milliseconds, significantly overcoming real-time bottlenecks and stuttering issues caused by limited computing power in edge scenarios.
[0023] By fusing IoU motion information with appearance feature matching based on cosine distance and utilizing GPU to accelerate matching calculations, the ID preservation capability in complex scenarios such as target occlusion, intersection, and lighting changes is significantly improved, reducing identity jumps. Furthermore, trajectory prediction based on Kalman filter combined with exponential moving average smoothing can stably output the target's motion position for several future frames, enabling the system not only to describe the target's historical trajectory but also to predict short-term future behavior, providing direct support for applications requiring forward-looking judgment, such as security early warning and autonomous driving obstacle avoidance.
[0024] With its built-in streaming media engine, the system can generate real-time visualized video streams overlaid with detection boxes, IDs, categories, and predicted trajectories, supporting multiple protocols such as RTSP, HTTP-FLV, and WebRTC to meet diverse viewing needs, from professional monitoring to web browsing. On the other hand, through an independent HTTP / WebSocket service interface, it provides precise data, including target coordinates, velocity, predicted location, and system status, in structured JSON format, enabling external systems to easily obtain information and trigger advanced business logic. This dual output mechanism of "video stream + data interface" bridges the gap from underlying visual perception to upper-level intelligent decision-making, greatly enhancing the system's practicality and integrability. Attached Figure Description
[0025] Figure 1 This is a diagram showing the overall architecture of the system of the present invention;
[0026] Figure 2 This is a schematic diagram of the target detection algorithm of the present invention;
[0027] Figure 3 This is a flowchart of the target tracking algorithm of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Example 1
[0030] Reference Figure 1-3 This is the first embodiment of the present invention, which provides a specific implementation of a real-time target detection, tracking and prediction system based on edge devices.
[0031] S1: System Architecture and Module Functions
[0032] This system mainly includes an input acquisition module, an inference processing module, a tracking and prediction module, an asynchronous pipeline scheduling module, a video rendering and streaming module, and a data interface service module.
[0033] Input Acquisition Module: This module acquires video streams from a webcam (RTSP stream) or a local USB camera. It employs a producer-consumer multi-threaded model: one acquisition thread (producer) continuously reads raw video frames from the camera's hardware interface; a preprocessing thread (consumer) retrieves frames from a thread-safe, lock-free queue and performs standardization processes such as scaling and color space conversion (e.g., BGR to RGB). To address the computing power limitations of edge devices (e.g., Jetson TX2), this module integrates a dynamic frame skipping unit. Its working principle is to monitor the device's CPU utilization, GPU utilization, memory usage, and the single-frame processing time of the inference module in real time. Preset load thresholds (e.g., CPU utilization ≥ 80%) and latency thresholds (e.g., inference time > 33ms, corresponding to a 30fps frame interval). When any indicator exceeds the threshold, the system automatically skips the subsequent 1 to N frames (N is dynamically calculated based on the degree of threshold exceedance), processing only key frames. This sacrifices a small amount of frame continuity for overall system real-time stability, preventing frame backlog from causing a continuous increase in visual latency.
[0034] Inference Processing Module: Receives preprocessed video frames (image matrix). The core of this module is object detection based on the YOLOv5s (lightweight version) model. To achieve real-time performance on edge devices, we optimized the standard YOLOv5 model using TensorRT.
[0035] Model conversion and optimization: TensorRT is used to convert the PyTorch format YOLOv5 model into a highly optimized engine file. The optimization process includes:
[0036] Layer fusion: Combines consecutive convolutional (Conv), batch normalization (BatchNorm), and activation function (SiLU) layers in the network into a single computational layer, reducing kernel startup and memory read / write overhead.
[0037] Accuracy calibration: FP16 half-precision inference is adopted. Calibration is performed during model conversion. While ensuring that the mAP accuracy loss on the CO dataset is less than 2%, the model weights and activation values are converted from FP32 to FP16, which halves the model memory usage and improves the inference speed by about 2-3 times.
[0038] Dynamic batch processing: Supports dynamically adjusting the batch size (1, 2, 4) based on the number of frames in the input queue to fully utilize the parallel computing capabilities of the GPU.
[0039] Detection and Feature Extraction: The optimized YOLOv5 engine performs forward inference on the input frames, outputting the bounding box, category (e.g., "person", "car"), and confidence score for each target. Simultaneously, for each detected target region (a sub-image cropped from the bounding box), this module calls a lightweight ReID (Re-identification) feature extraction network (based on MobileNetV2) optimized with TensorRT. This network encodes the target sub-image, generating a 512-dimensional appearance feature vector. This vector represents the target's color, texture, and other appearance information, and exhibits the characteristic of "high similarity for similar targets and low similarity for dissimilar targets," providing crucial information for subsequent tracking.
[0040] Tracking and Prediction Module: Receives detection results (bounding boxes, categories, confidence scores) and appearance feature vectors from the inference module. This module implements multi-target tracking and trajectory prediction based on an improved DeepSort algorithm.
[0041] Target association (matching):
[0042] For each target with an established tracking trajectory, a Kalman filter is used to predict its bounding box position in the current frame based on its motion state (position, velocity) in the previous frame.
[0043] Calculate two distance metrics for data association:
[0044] Motion similarity ( Distance): Calculate the predicted bounding box With the current frame detection box Crossover ratio (CLOUD) The formula is:
[0045]
[0046] The larger the value (the closer to 1), the higher the spatial overlap, and the greater the probability of a match.
[0047] Appearance similarity (cosine distance): Calculates the historical appearance feature vector of the tracked target. Compared with the appearance feature vector of the detection box in the current frame The cosine similarity between them is calculated using the following formula:
[0048]
[0049] The closer the value is to 1, the more similar the appearance.
[0050] Comprehensive matching: The Hungarian algorithm is used to... The weighted sum of the distance and cosine distance is used as the cost matrix to solve for the globally optimal match, and the detection box is assigned to the existing tracking ID. In this embodiment, the matching threshold of the standard DeepSort is specifically optimized, and GPU acceleration is used to calculate the cosine distance matrix between large-scale feature vectors to cope with dense target scenes.
[0051] Trajectory prediction and smoothing:
[0052] State Prediction and Update: The Kalman filter follows a "prediction-update" cycle. The prediction phase is based on the state estimate from the previous time step. And the motion model (assuming uniform linear motion) predicts the current state. Covariance The update phase combines the current observations (coordinates of successfully matched detection boxes) with the data. ), calculate Kalman gain And correct the predicted values to obtain a more accurate state estimate. Covariance Based on this, its position in the next 1-3 frames can be predicted (future position prediction result).
[0053] Trajectory smoothing: To avoid jitter in the predicted trajectory due to observation noise, the position sequence (including historical trajectories and future predicted points) output by the Kalman filter is smoothed. Perform exponential moving average (EMA) smoothing. The smoothing formula is:
[0054]
[0055] in, These are the smoothed trajectory points of the current frame. For the original predicted trajectory points of the current frame, The smooth trajectory points of the previous frame, The smoothing coefficient is set between 0.1 and 0.3 in this embodiment (usually 0.2). The smoothed trajectory is more stable, which is beneficial for visualization and subsequent analysis.
[0056] The asynchronous pipeline scheduling module, acting as the system's central nervous system, coordinates all processing flows. It decomposes tasks into multiple independent execution threads: acquisition thread, preprocessing thread, inference thread, tracing thread, rendering thread, storage thread, and streaming thread. Data is transferred between threads using lock-free queues or circular buffers. For example, the preprocessing thread places processed frames into the "inference queue," the inference thread retrieves frames from it for processing, and then places the results into the "tracing queue." This "producer-consumer" model ensures that temporary blocking at any stage (such as slow inference of a frame) does not cause the entire pipeline to stall, greatly improving system throughput and robustness. The core of the lock-free queue is the use of CAS (Compare And Swap) atomic operations to achieve thread-safe data reading and writing, avoiding the thread suspension and wake-up overhead caused by traditional mutexes.
[0057] The video rendering and streaming module obtains results with target ID, category, bounding box, and smoothed trajectory information from the tracking and prediction module. The rendering thread uses a GPU-accelerated graphics library (such as OpenCV's CUDA module) to draw semi-transparent colored bounding boxes, category labels, target IDs, and predicted trajectory lines on the raw video frames in real time. The rendered frames are then fed into the streaming thread. This module integrates the FFmpeg library, supporting the encoding of video streams and real-time pushing via various protocols.
[0058] RTSP: Based on RTP / UDP transmission, with low latency (100~300ms), suitable for professional security platforms.
[0059] HTTP-FLV: Transmitted via HTTP protocol, with moderate latency (300~1000ms), and supports playback in web browsers without plugins.
[0060] WebRTC: Uses UDP transmission, supports point-to-point communication, has extremely low latency (50~200ms), and is suitable for interactive remote monitoring.
[0061] Meanwhile, the storage thread saves the rendered video frames to the local disk in H.264 format and generates CSV or JSON log files containing information such as timestamps, target IDs, and coordinates for later review and analysis.
[0062] Data Interface Service Module: This module runs an independent HTTP / WebSocket service (e.g., based on the Flask framework) in parallel with the core processing pipeline. It collects real-time status data from various system modules, encapsulates it in JSON format, and provides it externally through an API interface, including at least the following two types of information:
[0063] a: Real-time target information array: Each array element corresponds to a tracked target object, containing fields such as target ID (id), class name (class), detection confidence, bounding box coordinates (bbox, usually represented in the format [x, y, width, height] or [top left x, top left y, bottom right x, bottom right y]), motion velocity vector (e.g., [vx, vy]), and predicted future position coordinates based on Kalman filter (e.g., [px, py]).
[0064] b: System status information object: contains real-time system operating status indicators, such as the frame rate (fps), the total number of objects in the current screen (object_count), the CPU utilization (cpu_usage), the GPU utilization (gpu_usage), and other fields.
[0065] External applications (such as central management platforms and mobile alarm apps) can obtain this data in real time through simple GET requests or WebSocket subscriptions, which can be used for electronic map annotation, crowd density analysis, automatic alarm triggering, or robot obstacle avoidance decisions.
[0066] S2: System Workflow and Beneficial Effects
[0067] After system startup, the threads of each module run in parallel. After video stream acquisition and preprocessing, optimized YOLOv5 quickly detects and extracts features, while the improved DeepSort algorithm stably associates target IDs and predicts future trajectories. An asynchronous scheduling mechanism ensures a smooth and unblocked process throughout. Users can watch real-time video streams overlaid with rich information or obtain precise structured data via API.
[0068] The embodiments of the present invention are given for the purposes of illustration and description. Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Any changes, modifications, substitutions and variations made by those skilled in the art to the above embodiments within the scope of the present invention should be included within the protection scope of the present invention.
Claims
1. A real-time target detection, tracking, and prediction system based on edge devices, characterized in that, include: The input acquisition module is used to acquire video streams and uses a multi-threaded asynchronous mechanism to acquire and preprocess video frames; The inference processing module detects preprocessed video frames based on the optimized YOLOv5 target detection model and extracts the appearance feature vectors of the targets. The tracking and prediction module, based on the improved DeepSort tracking algorithm, uses the target's appearance feature vector and detection box information to associate the target ID, and uses a Kalman filter to predict the target's future position and perform trajectory smoothing. The asynchronous pipeline scheduling module is used to build a multi-threaded parallel processing framework. It uses a lock-free queue or a circular buffer to transfer data between the input acquisition module, inference processing module, tracking and prediction module, video rendering and streaming module, and data interface service module to achieve asynchronous execution of each module. The video rendering and streaming module is used to render video frames that have been overlaid with detection boxes, target categories, target IDs, confidence scores and predicted trajectory information, and output them in real time through streaming media protocols; The data interface service module is used to provide structured detection, tracking, prediction, and system status data to external systems in real time via HTTP or WebSocket interfaces.
2. The system according to claim 1, characterized in that, The input acquisition module includes a dynamic frame skipping processing unit, which is used to dynamically adjust the processing frequency of video frames according to the real-time load and inference time of the edge device.
3. The system according to claim 1, characterized in that, The inference processing module uses the TensorRT framework to optimize the YOLOv5 detection model and appearance feature extraction network. The optimization methods include at least one of layer fusion, convolution operator optimization, and FP16 half-precision inference.
4. The system according to claim 1, characterized in that, The tracking prediction module improves the DeepSort algorithm by using GPU to accelerate the extraction and matching process of appearance feature vectors; adjusting the state noise and observation noise parameters of the Kalman filter; and using the exponential moving average method to smooth the predicted trajectory.
5. The system according to claim 1, characterized in that, The asynchronous pipeline scheduling module divides system tasks into acquisition threads, preprocessing threads, inference threads, tracing threads, rendering threads, storage threads, and streaming threads. Each thread communicates through a data queue based on a "producer-consumer" model.
6. The system according to claim 1, characterized in that, The video rendering and streaming module supports at least one of the following streaming media protocols for video stream output: RTSP, HTTP-FLV, and WebRTC, and supports local storage of processed video data while streaming.
7. The system according to claim 1, characterized in that, The structured data output by the data interface service module includes at least one of the following: target ID, category, confidence level, bounding box coordinates, historical trajectory point list, velocity vector, predicted coordinates for several future frames, real-time frame rate, current total number of targets, and device resource utilization rate.
8. The system according to claim 1, characterized in that, When the tracking and prediction module associates target IDs, it simultaneously calculates the IoU distance between the predicted bounding box and the detection bounding box and the cosine distance between the appearance feature vectors, and performs comprehensive matching based on the Hungarian algorithm.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the system as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the system as described in any one of claims 1-8.