Real-time target detection tracking method and system for multi-dimensional cooperative remote monitoring
Through the multi-dimensional collaborative remote monitoring method, the improved YOLO model and adaptive motion modeling algorithm are used, combined with the target triggering mechanism and the hierarchical communication protocol, the system bottleneck of real-time object detection in the Internet of Things is solved, efficient target tracking and data value mining are achieved, and the real-time and robustness of the system are improved.
Patent Information
- Application Number
- CN202510447063.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The prior art has problems in the field of the Internet of Things, such as limited system scenarios, high response delay, large network bandwidth pressure, insufficient terminal computing power, limited imaging quality, limited data value mining and lack of remote interactive interfaces, which are difficult to meet real-time requirements and dynamic tracking and control capabilities.
The real-time object detection and tracking method of multi-dimensional collaborative remote monitoring is adopted, and the environment image data is acquired for multi-level preprocessing, and the improved YOLO object detection model is used for object detection and association matching. Combined with the multi-objective association tracking algorithm and target triggering mechanism of adaptive motion modeling, the motion trajectory event picture is constructed, and the end-cloud and end-end instruction transmission is realized through MQTT and RTSP protocols.
It improves tracking robustness in nonlinear motion scenarios, reduces position prediction errors, improves target tracking accuracy and data value mining capabilities, and achieves stable maintenance and real-time tracking of the target in the center of the picture.
Smart Images

Figure CN120355740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet of Things artificial intelligence, and particularly to a real-time target detection and tracking method and system for multi-dimensional collaborative remote monitoring. Background Art
[0002] With the rapid development of Internet of Things, artificial intelligence and edge computing technologies, target detection technology has been widely applied in multiple fields, but there are still significant defects in the current technical system. Traditional solutions rely too much on the cloud for data processing, resulting in limited system scenarios, high response latency, and large network bandwidth pressure, making it difficult to meet real-time requirements. At the same time, existing technologies are mostly limited to single detection functions and fail to improve the capabilities in aspects such as dynamic tracking control, remote monitoring interaction, and data value mining, resulting in fragmented system functions and limited application scenarios. Specifically, monitoring devices generally have technical bottlenecks such as imaging quality being limited by a fixed perspective, terminal computing power being insufficient to support real-time processing requirements, limited mining of effective data value, and lack of remote interaction interfaces. Summary of the Invention
[0003] To solve the above technical problems, the object of the present invention is to provide a real-time target detection and tracking method and system for multi-dimensional collaborative remote monitoring, which can effectively improve the tracking robustness in non-linear motion scenarios, reduce the position prediction error, and ultimately improve the accuracy of target tracking.
[0004] The first technical solution adopted by the present invention is: a real-time target detection and tracking method for multi-dimensional collaborative remote monitoring, comprising the following steps:
[0005] Obtain environmental image data and perform multi-level preprocessing on the image data to obtain preprocessed environmental image data;
[0006] Perform target detection processing and association matching on the preprocessed environmental image data based on an improved YOLO target detection model to obtain a motion trajectory of a numbered target;
[0007] Dynamically cache and analyze the motion trajectory of the numbered target through a target trigger mechanism to construct a motion trajectory event picture;
[0008] Judge the number of targets in the motion trajectory event picture and perform target tracking according to the judgment result to achieve real-time target detection and tracking.
[0009] Further, the step of obtaining environmental image data and performing multi-level preprocessing on the image data to obtain preprocessed environmental image data specifically includes:
[0010] Construct a thread pool architecture based on priority scheduling, and the thread pool architecture includes high-level parallel threads and low-level threads;
[0011] Obtain environmental image data through high - level parallel threads, and write the environmental image data into a buffer queue through low - level threads to obtain buffered environmental image data;
[0012] Perform pre - processing operations of size scaling and color correction on the buffered environmental image data in sequence, and transfer and store it in a pre - processing data queue to obtain pre - processed environmental image data.
[0013] Further, the step of performing object detection processing and association matching on the pre - processed environmental image data based on the improved YOLO object detection model to obtain the motion trajectory of numbered objects specifically includes:
[0014] Based on the YOLO object detection model, perform INT8 quantization and add several detection head output terminals to construct an improved YOLO object detection model;
[0015] Based on the backbone network module of the improved YOLO object detection model, extract features from the pre - processed environmental image data to obtain environmental image feature data;
[0016] Based on the neck network module of the improved YOLO object detection model, perform multi - scale fusion on the environmental image feature data to obtain fused environmental image feature data;
[0017] Based on the detection head module of the improved YOLO object detection model, perform object detection on the fused environmental image feature data to obtain object detection results, and the object detection results include category, location, and confidence information;
[0018] Perform association matching on the object detection results based on a multi - object association tracking algorithm based on adaptive motion modeling to obtain the motion trajectory of numbered objects.
[0019] Further, the step of performing association matching on the object detection results based on a multi - object association tracking algorithm based on adaptive motion modeling to obtain the motion trajectory of numbered objects specifically includes:
[0020] Initialize the trajectory, establish a trajectory sequence, and the trajectory sequence includes active trajectories, unactivated trajectories, and lost - pursuit trajectories;
[0021] According to the confidence information of the object detection results, determine high - score object detection frames and low - score object detection frames;
[0022] Based on an adaptive extended Kalman filter, construct a prediction model for non - linear motion trajectories, predict the active trajectories, and obtain an adaptive extended Kalman filter prediction frame;
[0023] The Hungarian algorithm is used to match and associate the predicted frames of the adaptive extended Kalman filter with the high-score target detection boxes. If the association fails, the high-score target detection boxes are loaded into the unactivated trajectories. If the association is successful, the high-score target detection boxes are loaded into the activated trajectories;
[0024] The Hungarian algorithm is used to associate and match the low-score target detection boxes with the unactivated trajectories. If the association fails, the low-score target detection boxes are loaded into the lost-tracking trajectories. If the association is successful, the low-score target detection boxes are loaded into the activated trajectories;
[0025] The Hungarian algorithm is used to match and associate the unactivated trajectories with the lost-tracking trajectories. If the association fails, the confidence is attenuated. If the attenuated confidence value is greater than the preset threshold, the current lost-tracking trajectory is retained. If the attenuated confidence value is less than the preset threshold, the current lost-tracking trajectory is deleted. If the association is successful, the unactivated trajectories are loaded into the activated trajectories;
[0026] Anomaly trajectory detection is performed on the final activated trajectories, and the motion trajectories of the numbered targets are output.
[0027] Furthermore, the step of dynamically caching and analyzing the motion trajectories of the numbered targets through the target trigger mechanism to construct the motion trajectory event picture specifically includes:
[0028] Effective target detection is performed on the motion trajectories of the numbered targets. If there are effective targets, the current timestamp is recorded and stored in the video recording cache queue;
[0029] The effective frame sequence in the video recording cache queue is compressed into an H.264 standard bitstream by the FFmpeg tool, and an MP4 video file is encapsulated and generated. The effective frame sequence represents the pre-recorded frames buffered before the recording start time and the subsequent real-time frames;
[0030] An independent log file and a statistical report are generated according to the MP4 video file;
[0031] Incremental learning processing and analysis are performed on the independent log file and the statistical report, and combined retrieval is performed to construct the motion trajectory event picture.
[0032] Furthermore, the step of judging the number of targets in the motion trajectory event picture and performing target tracking according to the judgment result to achieve real-time target detection and tracking specifically includes:
[0033] The number of targets in the motion trajectory event picture is judged to obtain a judgment result;
[0034] If the number of targets in the judgment result is greater than the preset threshold, the motion trajectory event picture is corrected to obtain a corrected motion trajectory event picture;
[0035] Perform target tracking based on the corrected moving trajectory event screen;
[0036] If the number of targets in the judgment result is less than the preset threshold, perform target tracking on the moving trajectory event screen to achieve real-time target detection and tracking.
[0037] Further, the step of, if the number of targets in the judgment result is greater than the preset threshold, correcting the moving trajectory event screen to obtain the corrected moving trajectory event screen specifically includes:
[0038] If the number of targets in the judgment result is greater than the preset threshold, set all targets to form a data point set, and there is a hypothetical data point with the smallest total distance value to the center coordinates of other data points;
[0039] Search for the hypothetical data point and allow excluding some data points that are too far away from the group;
[0040] Iteratively search according to the target function with the smallest total distance value to obtain the corrected center coordinates of several targets;
[0041] Calculate the distance between the corrected center coordinates of several targets and the center of the screen until the preset termination condition is met, update the total distance value, determine the remaining data points as targets, and obtain the corrected moving trajectory event screen.
[0042] Further, the step of, if the number of targets in the judgment result is less than the preset threshold, performing target tracking on the moving trajectory event screen to achieve real-time target detection and tracking specifically includes:
[0043] If the number of targets in the judgment result is less than the preset threshold, calculate the distance deviation value between the center coordinates of the first target and the center of the screen;
[0044] Perform median filtering and cubic spline interpolation on the distance deviation value to obtain the interpolated and filtered distance deviation value;
[0045] Based on the PID fuzzy controller, perform fuzzification, fuzzy inference, and defuzzification on the distance deviation value between the interpolated and filtered target deviation and the center of the screen, and output a PWM control signal;
[0046] Drive the pan-tilt motor to move closer to the target according to the PWM control signal to achieve real-time target detection and tracking.
[0047] Further, it also includes:
[0048] Implement instruction transmission between the end-cloud and end-end through the Message Queuing Telemetry Transport protocol;
[0049] By sending content in JSON format, after entering the preset interface instructions and specific parameters in the command item, the IP address of the device can be obtained, the reset and four-way movement of the pan-tilt can be directly controlled, and the switching of the device operation mode can be achieved;
[0050] Build a real-time video transmission channel based on the real-time streaming protocol.
[0051] The second technical solution adopted by the present invention is: a multi-dimensional collaborative remote monitoring real-time target detection and tracking system, including:
[0052] The first module is used to obtain environmental image data and perform multi-level image data preprocessing to obtain preprocessed environmental image data;
[0053] The second module is used to perform target detection processing and correlation matching on the preprocessed environmental image data based on an improved YOLO target detection model to obtain the movement trajectories of numbered targets;
[0054] The third module is used to dynamically cache and analyze the movement trajectories of numbered targets through a target trigger mechanism to construct movement trajectory event pictures;
[0055] The fourth module is used to judge the number of targets in the movement trajectory event pictures and perform target tracking according to the judgment results to achieve real-time target detection and tracking.
[0056] The beneficial effects of the method and system of the present invention are: by obtaining environmental image data and performing multi-level image data preprocessing, and then performing target detection processing and correlation matching on the preprocessed environmental image data based on an improved YOLO target detection model, relying on CPU-NPU multi-threaded pipeline to process data, without relying on the cloud, quickly and stably keeping the target in the center of the picture, effectively improving the tracking robustness in non-linear motion scenarios, reducing the position prediction error, increasing the success rate of identity recovery after occlusion or missed detection, reducing the trajectory drift rate caused by false detection or sudden noise, and then dynamically caching and analyzing the movement trajectories of numbered targets through a target trigger mechanism, and finally judging the number of targets in the movement trajectory event pictures and performing target tracking according to the judgment results to improve the accuracy of target tracking. Description of the Drawings
[0057] Figure 1 is the step flow chart of a multi-dimensional collaborative remote monitoring real-time target detection and tracking method of the present invention;
[0058] Figure 2 is the structural block diagram of a multi-dimensional collaborative remote monitoring real-time target detection and tracking system of the present invention;
[0059] Figure 3It is a schematic diagram of the system framework for multi-dimensional collaborative remote monitoring provided by a specific embodiment of the present invention;
[0060] Figure 4 It is a schematic diagram of the real-time target detection and tracking method provided by a specific embodiment of the present invention;
[0061] Figure 5 It is a schematic diagram of the input-output structure of the target detection model provided by a specific embodiment of the present invention;
[0062] Figure 6 It is a schematic flowchart of the multi-target association tracking algorithm provided by a specific embodiment of the present invention;
[0063] Figure 7 It is a schematic flowchart of the data mining provided by a specific embodiment of the present invention;
[0064] Figure 8 It is a schematic diagram of the control flow provided by a specific embodiment of the present invention. Specific Embodiments
[0065] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0066] Refer to Figure 1 , the present invention provides a real-time target detection and tracking method for multi-dimensional collaborative remote monitoring, and the method includes the following steps:
[0067] S100. Obtain environmental image data and perform multi-level preprocessing on the image data to obtain the preprocessed environmental image data;
[0068] S110. Construct a thread pool architecture based on priority scheduling, and the thread pool architecture includes high-level parallel threads and low-level threads;
[0069] First of all, it should be noted that as Figure 3As shown in the figure, it includes an image acquisition module for acquiring image data; an edge computing module for performing image processing, model inference, image encoding, and video synthesis; a motion control module for calculating motion trajectories and outputting PWM control signals; a pan-tilt for carrying the image acquisition module and driving the adjustment of the camera's viewing angle through its own rotation; an external storage module for saving valid files, including videos, log files, statistical reports, etc.; and a remote communication module for realizing remote data communication with a server or other devices. Among them, the image acquisition module is connected to the edge computing module; the remote communication module is respectively connected to the edge computing module and the motion control module to form two bidirectional data transmission channels; the motion control module is connected to the edge computing module and receives the control instructions sent by the edge computing module. The pan-tilt is connected to the motion control module and receives the control instructions of the motion control module to achieve tracking of the target object. The edge computing module is connected to the external storage module. The image acquisition module transmits the acquired image data to the edge computing module. The image acquisition module is installed on the pan-tilt, and the viewing angle of the camera is adjusted by the rotation of the pan-tilt to center the target in the image.
[0070] S120. Obtain environmental image data through high-level parallel threads and write the environmental image data into a buffer queue through low-level threads to obtain buffered environmental image data.
[0071] Specifically, the image acquisition module is fixedly installed at the top of the system and circularly acquires environmental image data through a camera. The images form a continuous video stream in time series. To optimize resource utilization, the system constructs a thread pool architecture based on priority scheduling: among which 6 high-level parallel threads are used for ADC image acquisition, and 2 low-level threads are used to write the acquired raw image data into the buffer queue Raw_Queue.
[0072] S130. Perform preprocessing operations of size scaling and color correction on the buffered environmental image data in sequence and transfer them to a preprocessing data queue to obtain preprocessed environmental image data.
[0073] Specifically, after the image data is extracted from the buffer queue Raw_Queue, it first needs to go through multiple levels of preprocessing and then be transferred to the preprocessing data queue Pre_Queue. The preprocessing operations include size scaling (uniformly adjusting the input resolution to 640×640×3) and color correction (converting the BGR format to the RGB format and then optimizing the brightness through histogram equalization). Subsequently, the edge computing module inputs the data in the preprocessing data queue Pre_Queue into an improved YOLO object detection model.
[0074] S200. Perform object detection processing and correlation matching on the preprocessed environmental image data based on the improved YOLO object detection model to obtain the motion trajectories of numbered objects.
[0075] Specifically, based on the YOLO object detection model, perform INT8 quantization and add several detection head output terminals to construct an improved YOLO object detection model; based on the backbone network module of the improved YOLO object detection model, extract features from the preprocessed environmental image data to obtain environmental image feature data; based on the neck network module of the improved YOLO object detection model, perform multi-scale fusion on the environmental image feature data to obtain the fused environmental image feature data; based on the detection head module of the improved YOLO object detection model, perform object detection on the fused environmental image feature data to obtain the object detection results, where the object detection results include category, location, and confidence information; perform correlation matching on the object detection results based on the multi-object association tracking algorithm of adaptive motion modeling to obtain the motion trajectories of numbered objects.
[0076] In this embodiment, first obtain the preprocessed image tensor data; then, based on the edge computing requirements, improve the YOLO model to obtain the improved YOLO model. The improvement contents are as follows:
[0077] 1) Model structure adjustment: Remove the post-processing module of the original model, expand the output heads from 6 to 9, and add 3 new output heads to respectively count the total confidence of small, medium, and large objects for threshold filtering of fast post-processing. The input and output structure of the model is as Figure 5 shown.
[0078] 2) Model quantization and conversion: Perform INT8 quantization (FP16→INT8) on the pre-trained model to compress the model volume and improve the inference speed; then convert the model format to adapt to the edge computing module to ensure hardware compatibility and improve the computing efficiency.
[0079] Further, input the image data into the improved YOLO model for inference, extract features through the backbone network composed of several modules such as C3k2 and C2PSA to obtain three-level features of 80×80×C1, 40×40×C2, and 20×20×C3; perform multi-scale fusion on the above three-level feature data through the neck network improved based on the PANet structure to obtain three-level feature data of 80×80×C4, 40×40×C5, and 20×20×C6 after fusion; input the fused features into the lightweight classification detection head through the head network to output the object detection results.
[0080] Among them, considering the performance limitations of edge devices, the edge computing module adopts a strategy of multi-threaded concurrent cooperation between the CPU and the NPU. Specifically, the CPU is mainly responsible for preprocessing and postprocessing, and the NPU is responsible for most of the inference calculations. The overall throughput is improved through a multi-threaded + pipeline design.
[0081] Finally, after the model inference result, the output detection head result of the model undergoes post-processing steps (coordinate transformation, confidence filtering, non-maximum suppression, classification and focusing loss), and the results of object detection are extracted: class, position, and confidence information: [classes, boxes(x, y, h, w), scores].
[0082] Furthermore, initialize the trajectory and establish a trajectory sequence, where the trajectory sequence includes active trajectories, inactive trajectories, and lost pursuit trajectories; determine high-confidence object detection boxes and low-confidence object detection boxes according to the confidence information of the object detection results; construct a prediction model for non-linear motion trajectories based on the adaptive extended Kalman filter to predict the active trajectories and obtain the adaptive extended Kalman filter prediction frames; match and associate the adaptive extended Kalman filter prediction frames with the high-confidence object detection boxes through the Hungarian algorithm. If the association fails, load the high-confidence object detection boxes into the inactive trajectories. If the association is successful, load the high-confidence object detection boxes into the active trajectories; match and associate the low-confidence object detection boxes with the inactive trajectories through the Hungarian algorithm. If the association fails, load the low-confidence object detection boxes into the lost pursuit trajectories. If the association is successful, load the low-confidence object detection boxes into the active trajectories; match and associate the inactive trajectories with the lost pursuit trajectories through the Hungarian algorithm. If the association fails, perform confidence attenuation. If the attenuated confidence value is greater than the preset threshold, retain the current lost pursuit trajectory. If the attenuated confidence value is less than the preset threshold, delete the current lost pursuit trajectory. If the association is successful, load the inactive trajectories into the active trajectories; perform abnormal trajectory detection on the final active trajectories and output the motion trajectories of numbered objects.
[0083] In this embodiment, after obtaining the object detection results, they are input into a multi-object association tracking algorithm for association matching of the detection results to generate the numbers of each object and the motion trajectories between consecutive images. In a dynamic tracking scenario, the object will undergo non-linear motions such as accelerating or sudden direction changes. When the camera's perspective is adjusted, the motion trajectories of the object in the image coordinate system will also exhibit significant non-linear characteristics. The embodiment of the present invention proposes a multi-object association tracking algorithm based on adaptive motion modeling, as Figure 6 shown, and the specific content is as follows:
[0084] 1) Initialize the trajectory and establish a trajectory sequence: active trajectories, inactive trajectories, and lost pursuit trajectories;
[0085] 2) Input the results of object detection and divide them into high-confidence boxes D according to the confidencehigh and the low-score box D low ;
[0086] 3) Use the Adaptive Extended Kalman Filter (AEKF) to construct a prediction model for the non-linear motion trajectory, predict the next frame of the activation trajectory, and dynamically adjust the process noise covariance matrix Q and the observation noise covariance matrix R;
[0087] 4) First association matching: Associate the current high-score box D high with the AEKF prediction frame of the previous frame through the Hungarian algorithm. If there is no match, the current frame will be loaded into the unactivated trajectory; otherwise, the current frame will be loaded into the activated trajectory and the trajectory will be updated;
[0088] 5) Second association matching: Associate the current low-score box D low with the unactivated trajectory through the Hungarian algorithm. If there is no match, the current frame will be loaded into the lost tracking trajectory; otherwise, the current frame will be loaded into the activated trajectory and the trajectory will be updated;
[0089] 6) The nth association matching: Associate the unactivated trajectory and the lost tracking trajectory through the Hungarian algorithm. If there is no match, the confidence of the current frame will be attenuated (scores k = β × scores k-1 , where β is the attenuation coefficient). If the confidence is greater than the preset threshold LOSS_THRESHOLD, the frame will still be retained in the lost tracking trajectory; otherwise, the current frame and the accumulated lost tracking trajectory will be deleted; if the match is successful, the current frame will be loaded into the activated trajectory and the trajectory will be updated;
[0090] 7) When the association matching is successful, the result will be subjected to abnormal trajectory detection. When the acceleration change amount of MAX_A2ERR consecutive frames exceeds the preset threshold, it is determined as an abnormal trajectory and the trajectory will be reset. If there is no abnormality, the target number will be generated and the trajectory will be output.
[0091] This algorithm effectively improves the tracking robustness in non-linear motion scenarios, reduces the position prediction error, increases the success rate of identity recovery after occlusion or missed detection, and reduces the trajectory drift rate caused by false detection or sudden noise.
[0092] S300. Dynamically cache and analyze the motion trajectory of the numbered target through the target trigger mechanism to construct the motion trajectory event picture;
[0093] Specifically, perform effective target detection on the motion trajectory with numbered targets. If there are effective targets, record the current timestamp and store it in the video recording cache queue. Compress the sequence of valid frames in the video recording cache queue into an H.264 standard bitstream through the FFmpeg tool, and encapsulate and generate an MP4 video file. The sequence of valid frames represents the pre-recorded frames buffered before the recording start time and subsequent real-time frames. Generate an independent log file and a statistical report based on the MP4 video file. Perform incremental learning processing and analysis on the independent log file and the statistical report, and perform combined retrieval to construct the motion trajectory event picture.
[0094] In this embodiment, when an effective target is first detected, the following steps are executed:
[0095] 1) Record the current timestamp. If the system is not in the recording state, start the video recording process, mark both the pre-recorded frames buffered before the recording start time and subsequent real-time frames as valid frames, and store them in the video recording cache queue Record_Queue.
[0096] 2) When the target persists and the stop condition is not triggered, maintain the above operations.
[0097] 3) When any stop condition is met (including the continuous disappearance duration of the target exceeding the preset threshold (such as 5 seconds), or the cumulative duration of the cache reaching the upper limit (such as 10 minutes)), compress the sequence of valid frames in the cache into an H.264 standard bitstream through the FFmpeg tool, encapsulate and generate an MP4 video file, and store it in the external storage module. At the same time, clear Record_Queue and release other cache resources.
[0098] Furthermore, each MP4 file is associated with an independent log file and a statistical report. Among them, the log file contains basic attribute information (creation time of the video file and the log file, storage path, checksum), and detailed data of each detection during the video recording process (detection timestamp, target classification label, quantity statistics, and coordinate area). The statistical report is automatically generated in JSON format, including: creation time; frequency statistics: the cumulative detection times of each target category during the video period; peak statistics: the maximum simultaneous presence quantity of the same type of target in a single frame; duration: the time interval from the first appearance to the last disappearance of the target.
[0099] Furthermore, the stored data can support the following extended applications: constructing an incremental learning dataset: driving model retraining by screening low-confidence detection samples in the log to achieve iterative optimization of the model; supporting behavior analysis of specific targets and providing a structured data source for the construction of an industry knowledge graph; constructing a database and performing combined conditional retrieval: implementing compound queries based on time range, target type, and quantity threshold through a preset script.
[0100] S400 determines the number of targets in the moving trajectory event screen, and performs target tracking based on the judgment result to achieve real-time target detection and tracking.
[0101] Specifically, determine the number of targets in the moving trajectory event screen to obtain a judgment result; if the number of targets in the judgment result is greater than a preset threshold, correct the moving trajectory event screen to obtain a corrected moving trajectory event screen; perform target tracking based on the corrected moving trajectory event screen; if the number of targets in the judgment result is less than the preset threshold, perform target tracking on the moving trajectory event screen to achieve real-time target detection and tracking.
[0102] Among them, if the number of targets in the judgment result is greater than the preset threshold, set all targets to form a data point set, and there is a hypothetical data point with the smallest total distance value to the central coordinates of other data points; search for the hypothetical data point and allow excluding some data points that are too far away from the group; perform iterative search according to the target function with the smallest total distance value to obtain the corrected central coordinates of several targets; calculate the distance between the corrected central coordinates of several targets and the center of the screen until the preset termination condition is met, update the total distance value, and determine the remaining data points as targets to obtain a corrected moving trajectory event screen.
[0103] If the number of targets in the judgment result is less than the preset threshold, calculate the distance deviation value between the central coordinate of the first target and the center of the screen; perform median filtering and cubic spline interpolation on the distance deviation value to obtain an interpolated and filtered distance deviation value; perform fuzzification, fuzzy inference, and defuzzification on the distance deviation value between the interpolated and filtered target deviation and the center of the screen based on a PID fuzzy controller, and output a PWM control signal; drive the pan-tilt motor to move towards the target according to the PWM control signal to achieve real-time target detection and tracking.
[0104] First, if there are multiple targets in the screen, first calculate the corrected centers of multiple targets. The calculation process of the corrected center is as follows:
[0105] 1) Basic assumptions, symbol definitions, and target functions;
[0106] Since the number of targets collected and detected in the screen is limited, that is, the total number of targets has an upper limit n;
[0107] Let the central coordinates of each target be a point set: P (t) ={P1(x1, y1), …, P n (x n , y n )}, P i (x i , y i ) represents the original central coordinate of the target point, and its set is P (i), i = 1, 2, ..., n;
[0108] Assume there exists a point such that the total distance S to other target center points is minimized, and it is allowed to
[0109] search for point C * During the process, it is allowed to exclude at most L (t) outliers that are too far away from P max among the targets.
[0110] Objective function:
[0111]
[0112] It represents the minimum total distance S from the remaining point set * to C after excluding at most L max points. to C. min .
[0113] 2) Initialization;
[0114] Set the current point set P (t) ;
[0115] Initial center
[0116] Initial total distance
[0117] The cumulative number of excluded points L = t;
[0118] The number of iterations k.
[0119] 3) Set the termination condition;
[0120] Stop during the iteration process when the following conditions are met:
[0121] The number of excluded points reaches the limit: L ≥ L max ;
[0122] The total distance converges: |S (k) - S (k-1) | < ε, where ε is a very small term.
[0123] 4) Update the iteration;
[0124] For the k-th iteration (k > 1):
[0125] Calculate the distance from each point to C (k) :
[0126] Select the point with the farthest distance as the candidate outlier:
[0127] Exclude this point: L = L + 1, update the point set:
[0128] Based on the updated point set, calculate the center:
[0129] Total distance update:
[0130] Furthermore, if the number of targets in the picture is small, then track the target numbered 1, and use the center coordinates of this target as the overall center coordinates.
[0131] Furthermore, calculate the center coordinates [X k Y k T of the target and the deviation value from the center point [X half Y half T of the picture:
[0132]
[0133] Among them, ΔX k is the pixel deviation in the horizontal direction between the target center of the k-th frame and the image center captured by the camera, and ΔY k is the pixel deviation in the vertical direction between the target center of the k-th frame and the image center captured by the camera.
[0134] After receiving the deviation data sent by the edge computing module, the control module loads the new data in a circular queue structure of finite size, and makes the data more continuous and smooth through median filtering and cubic spline interpolation. The motion control module uses a fuzzy incremental PID control algorithm to generate a PWM signal, thereby driving the pan-tilt head.
[0135] The control parameters K x = [K p K i K d are updated in the fuzzy controller, including three parts: fuzzification, fuzzy inference (the fuzzy rule base is represented by M later), and defuzzification. Specifically as follows:
[0136] Fuzzification: The target deviation e(k) and the deviation change rate Δe(k) are mapped into the fuzzy sets {Negative Big NB, Negative Medium NM, Zero ZO, Positive Medium PM, Positive Big PB} through the triangular membership function. Triangular membership function:
[0137]
[0138] Among them, K x is the collective name of K p 、K i 、K d , represents Kx Membership degree in the rule base, c is the center point of the membership function, and w is the half-width of the scaled base.
[0139] Fuzzy inference: Establish a fuzzy rule base M with 25 rules, and the rule form is:
[0140] R i : IF e(k) = A AND Δe(k) = B THEN ΔK p = C, ΔK i = D, ΔK d = E
[0141] Among them, A, B, C, D, and E represent an item in the fuzzy rule base M, and the specific content is as follows:
[0142]
[0143] Defuzzification:
[0144]
[0145] Among them, ΔK x is the collective term of ΔK p , ΔK i , ΔK d and represents the increment of K x after defuzzification; represents the defuzzified value defined by the τ-th rule in the fuzzy rule base of K x (corresponding to the constants represented by PB, PM, ZO, NM, and NB).
[0146] Parameter update:
[0147]
[0148] Furthermore, the incremental PID formula: Δu(k) = K p Δe(k) + K i e(k) + K d Δ 2 e(k).
[0149] Among them, Δu(k) is the increment of the output PWM signal, and e(k) is the pixel deviation of the k-th frame.
[0150] Furthermore, the PID output result is limited by an upper limit and a lower limit, and after passing through a low-pass filter, the PWM signal is output to the pan-tilt motor.
[0151] Furthermore, when the corrected center position of the target is near the center of the screen, the pan-tilt does not respond and a more stable effect can be achieved.
[0152] Furthermore, when the gimbal adjusts the horizontal angle and pitch angle of the camera, the camera gradually approaches the target, and finally the target is always in the center of the picture. Through the above operation, a smooth and fast real-time tracking effect can be achieved.
[0153] Finally, the embodiment of the present invention also proposes a communication method for remote streaming and remote operation and an interface implementation scheme thereof, which specifically includes:
[0154] 1) On the one hand, the remote communication module can access the Internet through the commonly used WiFi technology, and use its convenient wireless connection characteristics to meet the needs of mobile devices or other inconvenient wiring scenarios; on the other hand, it can also be accessed through the RJ45 interface, relying on the stability and high speed of the wired network to ensure the reliability of data transmission, thereby achieving stable connection in complex network environments, whether in urban environments with more signal interference or in remote areas with limited network coverage, it can ensure that the device can be connected to the Internet normally.
[0155] 2) The MQTT protocol (Message Queuing Telemetry Transport) is lightweight and low-overhead. The system deploys an MQTT client service that establishes a connection with the MQTT server (Broker) to achieve end-to-cloud and end-to-end command transmission (such as IP acquisition, PTZ control, mode switching, etc.), and supports QoS1 level message reliability to ensure that important commands can be delivered accurately when the network fluctuates, avoiding device operation errors or functional abnormalities due to command loss.
[0156] Furthermore, after the device subscribes to the topic and starts the MQTT service, the user sends JSON format content to the topic. It can be used for multiple devices, and the device ID item can be used to specify the device number for targeted adjustments; after entering the preset interface command and specific parameters in the command item, the device IP address can be obtained, the PTZ reset and four-way movement can be directly controlled, the device operation mode can be switched (restart or shut down, sleep, automatic patrol, manual control) and other functions can be controlled.
[0157] 3) Optimized RTSP (Real-Time Stream Protocol) video streaming design to build a highly robust real-time video transmission channel. Specifically: Deploy an RTSP server on the device side, supporting the transmission of high-definition video streams at 1080P@30fps under the H.264 encoding standard. Through an adaptive bitrate adjustment mechanism, it dynamically matches the network bandwidth (adjustment range: 500Kbps - 8Mbps), can maintain an end-to-end transmission delay of <300ms, and allows access through a web server and a streaming media server. In addition, the SRTP protocol (Secure Real-time Transport Protocol) is used to encrypt the video stream transmission, combined with a two-way authentication mechanism for device identity certificates to prevent unauthorized access.
[0158] In summary, as Figure 4 shown, in the embodiment of the present invention, the video stream is first obtained in real time through the image acquisition module; after preprocessing, the video frames are input into the improved object detection model for inference, and the inference results are post-processed to obtain the detection results. The improved multi-object association tracking algorithm is used to generate object numbers and motion trajectories; after the system mines the data value of the valid frames, it encodes and stores them in the external storage module; calculates the position deviation, and the motion control module performs filtering interpolation and fuzzy incremental PID operations on the deviation data, and outputs a control signal to drive the pan-tilt to track in real time. The remote communication module adopts the MQTT-RTSP (Message Queuing Telemetry Transport Protocol - Real-Time Stream Protocol) hierarchical protocol architecture, thus forming a multi-dimensional collaborative remote monitoring system. The present invention has the following significance: It constructs a physical tracking closed-loop system at the edge side, without relying on cloud processing, and has high robustness; pays attention to data value mining and can be used for various purposes such as model iteration and specific target behavior research; the hierarchical communication architecture has a convenient remote human-computer interaction interface.
[0159] Therefore, compared with the prior art, the embodiment of the present invention has the following improvement points:
[0160] 1) Constructed a physical tracking closed-loop system at the edge side. Through the improved object detection model, the improved multi-object association tracking algorithm, and the adaptive pan-tilt tracking technology, relying on the CPU-NPU multi-threaded pipeline to process data, without relying on the cloud, quickly and stably keeps the target in the center of the screen.
[0161] 2) Pay attention to the mining of data value. The target trigger mechanism automatically records videos, and automatically counts various valid data and generates logs and statistical reports. The saved data is highly effective and can be developed subsequently according to the needs of different scenarios, such as subsequent model iteration, behavior research of specified targets, data sorting and query, etc.
[0162] 3) Hierarchical communication architecture. By integrating the MQTT protocol and the RTSP protocol, it has the functions of remote control instruction interface and real-time image encrypted remote transmission, providing a convenient interface method for the server and users to call.
[0163] Referring to Figure 2 , a real-time target detection and tracking system for multi-dimensional collaborative remote monitoring, including:
[0164] The first module 201 is used to obtain environmental image data and perform multi-level image data preprocessing to obtain preprocessed environmental image data;
[0165] The second module 202 is used to perform target detection processing and correlation matching on the preprocessed environmental image data based on an improved YOLO target detection model to obtain the motion trajectory of the numbered target;
[0166] The third module 203 is used to dynamically cache and analyze the motion trajectory of the numbered target through a target trigger mechanism to construct a motion trajectory event picture;
[0167] The fourth module 204 is used to judge the number of targets in the motion trajectory event picture and perform target tracking according to the judgment result to achieve real-time target detection and tracking.
[0168] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0169] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A real-time target detection and tracking method for multi-dimensional collaborative remote monitoring, characterized in that, Including the following steps: Obtain environmental image data and perform multi-level preprocessing of the image data to obtain preprocessed environmental image data; Perform object detection processing and association matching on the preprocessed environmental image data based on an improved YOLO object detection model to obtain a motion trajectory of numbered objects; Dynamically cache and analyze the motion trajectory of numbered objects through an object trigger mechanism to construct a motion trajectory event screen; Judge the number of objects in the motion trajectory event screen and perform object tracking according to the judgment result to achieve real-time object detection and tracking.
2. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 1, wherein The step of obtaining environmental image data and performing multi-level preprocessing of the image data to obtain preprocessed environmental image data specifically includes: Construct a thread pool architecture based on priority scheduling, and the thread pool architecture includes high-level parallel threads and low-level threads; Obtain environmental image data through high-level parallel threads and write the environmental image data into a buffer queue through low-level threads to obtain cached environmental image data; Perform preprocessing operations of size scaling and color correction on the cached environmental image data in sequence and transfer it to a preprocessing data queue to obtain preprocessed environmental image data.
3. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 2, characterized in that, The step of performing object detection processing and association matching on the preprocessed environmental image data based on an improved YOLO object detection model to obtain a motion trajectory of numbered objects specifically includes: Based on the YOLO object detection model, perform INT8 quantization and add several detection head output ends to construct an improved YOLO object detection model; Extract features from the preprocessed environmental image data based on the backbone network module of the improved YOLO object detection model to obtain environmental image feature data; Perform multi-scale fusion on the environmental image feature data based on the neck network module of the improved YOLO object detection model to obtain fused environmental image feature data; Perform object detection on the fused environmental image feature data based on the detection head module of the improved YOLO object detection model to obtain object detection results, and the object detection results include category, position, and confidence information; Perform association matching on the object detection results based on a multi-object association tracking algorithm based on adaptive motion modeling to obtain a motion trajectory of numbered objects.
4. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 3, characterized in that, The step of performing association matching on the object detection results based on a multi-object association tracking algorithm based on adaptive motion modeling to obtain a motion trajectory of numbered objects specifically includes: Initialize the trajectory and establish a trajectory sequence, and the trajectory sequence includes active trajectories, unactivated trajectories, and lost pursuit trajectories; Determine high-score object detection frames and low-score object detection frames according to the confidence information of the object detection results; Construct a prediction model for non-linear motion trajectories based on adaptive extended Kalman filtering to predict the active trajectories and obtain an adaptive extended Kalman filtering prediction frame; Perform matching association on the adaptive extended Kalman filtering prediction frame and the high-score object detection frame through the Hungarian algorithm. If the association fails, load the high-score object detection frame into the unactivated trajectory. If the association is successful, load the high-score object detection frame into the active trajectory; The low-score target detection boxes and unactivated trajectories are associated and matched through the Hungarian algorithm. If the association fails, the low-score target detection boxes are loaded into the lost tracking trajectories. If the association is successful, the low-score target detection boxes are loaded into the activated trajectories; The unactivated trajectories and the lost tracking trajectories are matched and associated through the Hungarian algorithm. If the association fails, the confidence is attenuated. If the attenuated confidence value is greater than the preset threshold, the current lost tracking trajectory is retained. If the attenuated confidence value is less than the preset threshold, the current lost tracking trajectory is deleted. If the association is successful, the unactivated trajectories are loaded into the activated trajectories; An abnormal trajectory detection is performed on the final activated trajectories, and the motion trajectories of the numbered targets are output.
5. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 4, characterized in that, The step of dynamically caching and analyzing the motion trajectories of the numbered targets through the target trigger mechanism and constructing the motion trajectory event picture specifically includes: An effective target detection is performed on the motion trajectories of the numbered targets. If there are effective targets, the current timestamp is recorded and stored in the video recording cache queue; The effective frame sequences in the video recording cache queue are compressed into an H.264 standard bitstream through the FFmpeg tool, and an MP4 video file is encapsulated and generated. The effective frame sequences represent the pre-recorded frames buffered before the recording start time and the subsequent real-time frames; An independent log file and a statistical report are generated based on the MP4 video file; An incremental learning process and analysis are performed on the independent log file and the statistical report, and a combined retrieval is performed to construct the motion trajectory event picture.
6. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 5, characterized in that, The step of judging the number of targets in the motion trajectory event picture and performing target tracking according to the judgment result to achieve real-time target detection and tracking specifically includes: The number of targets in the motion trajectory event picture is judged to obtain a judgment result; If the number of targets in the judgment result is greater than the preset threshold, the motion trajectory event picture is corrected to obtain a corrected motion trajectory event picture; Target tracking is performed based on the corrected motion trajectory event picture; If the number of targets in the judgment result is less than the preset threshold, target tracking is performed on the motion trajectory event picture to achieve real-time target detection and tracking.
7. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 6, characterized in that, The step of, if the number of targets in the judgment result is greater than the preset threshold, correcting the motion trajectory event picture to obtain a corrected motion trajectory event picture specifically includes: If the number of targets in the judgment result is greater than the preset threshold, all the targets are set to form a data point set, and there is a hypothetical data point with the smallest total distance value to the central coordinates of other data points; Search for the hypothetical data point, and allow excluding some data points that are too far away from the outliers; Iterative search is performed according to the target function with the smallest total distance value to obtain the corrected central coordinates of several targets; Calculate the distances between the corrected central coordinates of several targets and the center of the picture until the preset termination condition is met, update the total distance value, determine the remaining data points as targets, and obtain the corrected motion trajectory event picture.
8. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 7, characterized in that The step of, if the number of targets in the judgment result is less than the preset threshold, performing target tracking on the motion trajectory event picture to achieve real-time target detection and tracking specifically includes: If the number of targets in the judgment result is less than the preset threshold, calculate the distance deviation value between the center coordinates of the first target and the center of the screen; Perform median filtering and cubic spline interpolation on the distance deviation value to obtain the distance deviation value after interpolation filtering; Based on the PID fuzzy controller, perform fuzzification, fuzzy inference, and defuzzification on the distance deviation value between the interpolated and filtered target deviation and the center of the screen, and output a PWM control signal; Drive the pan-tilt motor to move towards the target according to the PWM control signal to achieve real-time target detection and tracking.
9. The real-time target detection and tracking method for multi-dimensional collaborative remote monitoring according to claim 8, wherein, It also includes: Realize the instruction transmission between the end-cloud and end-end through the Message Queuing Telemetry Transport protocol; By sending the content in JSON format, after entering the preset interface instruction and specific parameters in the command item, the IP address of the device, the reset and four-way movement of the pan-tilt can be directly controlled, and the switching of the device operation mode can be obtained; Build a real-time video transmission channel based on the Real-Time Streaming Protocol.
10. A real-time target detection and tracking system for multi-dimensional collaborative remote monitoring, characterized in that, It includes the following modules: The first module is used to obtain environmental image data and perform multi-level preprocessing of the image data to obtain the preprocessed environmental image data; The second module is used to perform target detection processing and correlation matching on the preprocessed environmental image data based on the improved YOLO target detection model to obtain the motion trajectory of the numbered target; The third module is used to dynamically cache and analyze the motion trajectory of the numbered target through the target trigger mechanism to construct a motion trajectory event screen; The fourth module is used to judge the number of targets in the motion trajectory event screen and perform target tracking according to the judgment result to achieve real-time target detection and tracking.
Citation Information
Patent Citations
Multi-target tracking method and system
CN117495915A
Pig multi-target tracking method
CN118230358A
Ship detection and tracking method and device
US20240371012A1