Video detection method, device and equipment for edge algorithm all-in-one machine and medium

Through global collaborative scheduling and dynamic parameter management, the problems of resource consumption and detection delay in edge algorithm all-in-one machines are solved, adaptive matching and efficient resource allocation of video detection are realized, and video analysis performance in edge computing scenarios is improved.

CN120388313APending Publication Date: 2025-07-29BEIJING NORTH STAR DIGITAL REMOTE SENSING TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510276832.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The traditional video detection method for edge algorithm all-in-one machines cannot dynamically adapt to fluctuations in the number of devices and detection requirements, resulting in increased resource consumption, data loss or delay, and lack of global collaborative management.

Method used

The global collaborative scheduling mechanism is adopted to obtain task parameters by polling the scheduling task, create a detection queue, dynamic management and detection analysis of video frames based on the global frame fetching timetable, and model training and deployment are combined with PyTorch and RKNN tool chains to achieve adaptive matching and resource optimization.

Benefits of technology

Improve resource utilization and real-time detection, reduce redundant calculations and data frame drops, and provide highly robust video analysis solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388313A_ABST
    Figure CN120388313A_ABST
Patent Text Reader

Abstract

The invention relates to a video detection method and device for an edge algorithm all-in-one machine, equipment and a medium, and the method comprises the steps: obtaining task parameters corresponding to a polling scheduling task for each polling scheduling task; starting a plurality of detection algorithm instances corresponding to each algorithm set, and creating a detection queue for each detection algorithm instance; based on a global framing timetable corresponding to the plurality of polling scheduling tasks, performing framing on video streams collected by a plurality of video input devices corresponding to the plurality of to-be-detected video channel sets to obtain a plurality of video frames, and placing the plurality of video frames in respective corresponding detection queues; and for each detection algorithm instance, continuously extracting the video frame from the corresponding detection queue, and carrying out detection analysis on the extracted video frame. According to the invention, the adaptive matching of the frame taking operation and the detection requirement is realized, and the resource utilization rate and the detection real-time performance are improved under the limited computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video frame extraction, and in particular to a video detection method, device, equipment and medium for edge algorithm all-in-one machines. Background Art

[0002] As one of the core applications in the field of computer vision, video object detection technology has been widely used in scenarios such as security monitoring, intelligent transportation, and industrial quality inspection.

[0003] Traditional video detection methods for edge algorithm all-in-one machines usually adopt a fixed task scheduling mechanism, and process multiple video input devices through single-threaded or static resource allocation. However, with the increase in the number of video channels and algorithm complexity in edge computing scenarios, traditional methods allocate computing power resources through fixed channel numbers or static binding methods, unable to dynamically adapt to fluctuations in the number of devices and detection requirements, easily resulting in redundant calculations, exacerbating the resource consumption of edge devices, and lacking global collaborative management of the frame extraction time of devices, leading to frame extraction conflicts or overclocked access of devices, resulting in data frame loss or delays.

[0004] Therefore, there is an urgent need for a video frame extraction and detection solution for multi-tasks and multi-devices, which can achieve adaptive matching between frame extraction operations and detection requirements through a global collaborative scheduling mechanism, so as to improve resource utilization and detection real-time performance under limited computing power. Summary of the Invention

[0005] In order to achieve adaptive matching between frame extraction operations and detection requirements, and improve resource utilization and detection real-time performance under limited computing power, this application provides a video detection method, device, equipment and medium for edge algorithm all-in-one machines.

[0006] In a first aspect, this application provides a video detection method for an edge algorithm all-in-one machine, including:

[0007] For each polling scheduling task, obtain the task parameters corresponding to the polling scheduling task, where the task parameters include a set of video channels to be detected, a set of algorithms, and a frame extraction interval. The set of video channels to be detected is a set of video input devices that need to perform target detection, the set of algorithms is a set of corresponding detection algorithm instances, and the frame extraction interval is a parameter representing the time interval between two adjacent video stream frame extractions;

[0008] Start multiple detection algorithm instances corresponding to each of the sets of algorithms, and create a detection queue for each of the detection algorithm instances;

[0009] Based on the global frame fetching schedule corresponding to multiple polling scheduling tasks, fetch frames from the video streams collected by multiple video input devices corresponding to each of the multiple video channel sets to be detected, obtain multiple video frames, and place the multiple video frames into their respective corresponding detection queues; the global frame fetching schedule includes the corresponding relationships between the task IDs of each of the polling scheduling tasks, the algorithm IDs of the detection algorithm instances, the device IDs of the video input devices, the previous frame fetching time, and the current frame fetching interval;

[0010] For each of the detection algorithm instances, continuously take out the video frames from the corresponding detection queue and perform detection and analysis on the taken-out video frames.

[0011] The beneficial effects of this application are as follows: Through global collaborative scheduling and dynamic parameter management, continuously perform detection and analysis on video frames, achieving efficient allocation of computing power resources, multi-task collaborative optimization, and guarantee of detection real-time performance, and significantly improving the comprehensive performance of the video analysis system in the edge computing scenario. Upgrading the traditional static scheduling to an adaptive dynamic closed-loop control provides a highly robust solution for scenarios such as large-scale video surveillance and industrial quality inspection.

[0012] Furthermore, after performing the detection and analysis on the taken-out video frames, it further includes:

[0013] If a preset alarm event is detected in the taken-out video frame, based on a preset confidence threshold and detection range, filter the detection data corresponding to the taken-out video frame, store the filtered detection data, and synchronously push it to the message queue.

[0014] The beneficial effects of adopting the above further solution are as follows: By setting double filtering (confidence + detection range), invalid alarms can be reduced, and the signal-to-noise ratio of the system is improved. By only processing key alarm data, storage and network bandwidth consumption are reduced. RabbitMQ can ensure that alarm information is transmitted to the downstream system in a timely manner, achieving real-time response. The detection range and confidence threshold can be dynamically adjusted through the management platform to adapt to different scenario requirements.

[0015] Furthermore, the task parameters further include the maximum number of channels. The fetching of frames from the video streams collected by multiple video input devices corresponding to each of the multiple video channel sets to be detected based on the global frame fetching schedule corresponding to multiple polling scheduling tasks includes:

[0016] Judge whether the total number of the multiple video input devices that need to perform frame fetching exceeds the maximum number of channels;

[0017] If the total quantity does not exceed the maximum number of channels, based on the global frame acquisition schedule, obtain the frame acquisition time requirements corresponding to each of the multiple video input devices; if the current time reaches at least one of the frame acquisition time requirements, then perform frame acquisition on the video streams collected by the corresponding at least one video input device.

[0018] The beneficial effects of adopting the above further solution are as follows: By comparing the total number of devices with the maximum number of channels in real time, automatically switch the processing mode to adapt to the computing power limitations of edge devices. Based on the global frame acquisition schedule, strictly restrict the frame acquisition intervals of devices to avoid resource contention or data redundancy caused by high-frequency frame acquisition. In the single-batch mode, all devices are processed in parallel, reducing waiting time and improving the response speed of high-priority tasks.

[0019] Further, if the total quantity exceeds the maximum number of channels, the method includes:

[0020] Perform rolling batch processing on the multiple video input devices to obtain the correspondence between different processing batches and the multiple video input devices;

[0021] For each current processing batch, based on the global frame acquisition schedule and the correspondence, obtain the frame acquisition time requirements corresponding to each of the multiple video input devices in the current processing batch; if the current time reaches at least one of the frame acquisition time requirements, then perform frame acquisition on the video streams collected by the corresponding at least one video input device in the current batch.

[0022] The beneficial effects of adopting the above further solution are as follows: When the total number of video input devices that need to have frames acquired exceeds the maximum number of channels, rolling batch processing can dynamically allocate computing power resources to ensure that all video input devices are reasonably scheduled under limited hardware capabilities.

[0023] Further, the task parameter further includes a channel switching interval. After performing frame acquisition on the video streams collected by the corresponding at least one video input device in the current batch, it further includes:

[0024] If the processing duration of the current processing batch meets the channel switching interval, then perform batch switching, take the next processing batch as the new current processing batch, and repeat the step of, for each current processing batch, based on the global frame acquisition schedule and the correspondence, obtaining the frame acquisition time requirements corresponding to each of the multiple video input devices in the current processing batch.

[0025] The beneficial effects of adopting the above further solution are as follows: When performing batch switching, it can switch to the next batch in a preset order, mark the next batch as the new current processing batch, and reset the timing start point. Repeat the frame acquisition time check and operation for the new batch to form a closed-loop scheduling.

[0026] Furthermore, the task parameters further include a silence period, which is the minimum time interval between two adjacent triggerings of alarm events. Obtaining the frame acquisition time requirements corresponding to each of the multiple video input devices in the current processing batch includes:

[0027] For each video input device, if a preset alarm event is detected in the video frame taken out last time by this video input device, the frame acquisition time requirement corresponding to this video input device is the sum of the last frame acquisition time and the silence period;

[0028] For each video input device, if a preset alarm event is not detected in the video frame taken out last time by this video input device, the frame acquisition time requirement corresponding to this video input device is the sum of the last frame acquisition time and the frame acquisition interval.

[0029] The beneficial effects of adopting the above further solution are as follows: According to the detection result of the previous frame acquisition of the video input device (i.e., whether an alarm is triggered), the current frame acquisition frequency is dynamically adjusted, and a balance between resource optimization and alarm response efficiency can be achieved. When the detection result of the previous frame acquisition is the alarm state, the frame acquisition interval can be extended through the silence period, reducing the processing of meaningless frames, reducing redundant calculations, reducing the consumption of computing resources, avoiding repeated alarm storms, and improving the accuracy of operation and maintenance response; when the detection result of the previous frame acquisition is the normal state, the conventional frame acquisition interval is maintained to ensure the continuity of monitoring.

[0030] Furthermore, before starting the multiple detection algorithm instances corresponding to each of the algorithm sets, it further includes:

[0031] Based on the Pytorch framework, the YOLOv10 model is trained and optimized in a training server;

[0032] The trained and optimized YOLOv10 model is exported in ONNX format;

[0033] Based on the RKNN Toolkit, the ONNX format YOLOv10 model is converted into an inference model adapted to the edge algorithm all-in-one machine, and the inference model is used as the detection algorithm instance.

[0034] The beneficial effects of adopting the above further solution are as follows: Through the complete process of PyTorch training → ONNX intermediate conversion → RKNN edge adaptation, a standardized and reusable technical path is provided for algorithm deployment in edge computing scenarios. Rapidly iterating the model structure based on the PyTorch ecosystem reflects the flexibility of training. Shielding framework differences through ONNX and supporting multi-hardware backend deployment reflect cross-platform compatibility. Utilizing the quantization and hardware acceleration capabilities of the RKNN toolchain to meet real-time requirements reflects edge efficiency.

[0035] In a second aspect, the present application provides a video detection device for an edge algorithm all-in-one machine, including:

[0036] A task parameter acquisition module, configured to, for each polling scheduling task, acquire the task parameters corresponding to the polling scheduling task, where the task parameters include a set of video channels to be detected, a set of algorithms, and a frame fetching interval. The set of video channels to be detected is a set of video input devices that need to perform target detection, the set of algorithms is a set of corresponding detection algorithm instances, and the frame fetching interval is a parameter representing the time interval between two adjacent video stream frame fetches;

[0037] A queue creation module, configured to start multiple detection algorithm instances corresponding to each of the sets of algorithms, and create a detection queue for each of the detection algorithm instances;

[0038] A frame fetching module, configured to, based on a global frame fetching schedule corresponding to multiple polling scheduling tasks, fetch video streams collected by multiple video input devices corresponding to each of the sets of video channels to be detected to obtain multiple video frames, and place the multiple video frames into the corresponding detection queues respectively; the global frame fetching schedule includes the corresponding relationship between the task ID of each polling scheduling task, the algorithm ID of the detection algorithm instance, the device ID of the video input device, the previous frame fetching time, and the current frame fetching interval;

[0039] A detection and analysis module, configured to, for each of the detection algorithm instances, continuously take out the video frames from the corresponding detection queue and perform detection and analysis on the taken-out video frames.

[0040] In a third aspect, the present application provides an electronic device, including a processor and a memory, and the processor is coupled to the memory;

[0041] The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of the first aspects.

[0042] Fourthly, the present application provides a computer-readable storage medium, including a computer program or instruction. When the computer program or instruction runs on a computer, the computer is enabled to execute the method according to any one of the first aspect. Description of the Drawings

[0043] Figure 1 It is a schematic flowchart of a video detection method for an edge algorithm all-in-one machine according to an embodiment of the present application;

[0044] Figure 2 It is a schematic diagram showing that according to an embodiment of the present application, video frames are placed into corresponding detection queues according to triple frame-taking information;

[0045] Figure 3 It is a schematic diagram showing a global frame-taking schedule according to an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram showing a frame-taking process according to an embodiment of the present application;

[0047] Figure 5 It is a structural block diagram of a video detection device for an edge algorithm all-in-one machine according to an embodiment of the present application;

[0048] Figure 6 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments

[0049] The present application will be further described in detail below with reference to the accompanying drawings.

[0050] An embodiment of the present application provides a video detection method for an edge algorithm all-in-one machine. This method can be executed by a device, and this device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a desktop computer, etc., but is not limited thereto.

[0051] As Figure 1 shown, a video detection method for an edge algorithm all-in-one machine, with an electronic device as the execution subject, the main process of the method is described as follows (Steps S101 to S104):

[0052] Step S101: For each polling scheduling task, obtain the task parameters corresponding to the polling scheduling task. The task parameters include a set of video channels to be detected, a set of algorithms, and a frame-taking interval. The set of video channels to be detected is a set of video input devices that need to perform target detection. The set of algorithms is a set of corresponding detection algorithm instances. The frame-taking interval is a parameter characterizing the time interval between two adjacent video stream frame-takings.

[0053] In this embodiment, the video input device can be a camera, and the task parameters mainly include:

[0054] Set of video channels to be detected {D}: A set of video input devices that need to perform target detection;

[0055] Set of algorithms {A}: A set of instances of detection algorithms that need to be executed;

[0056] Maximum number of channels NC: The maximum number of video channels for decoding and frame extraction at the same time;

[0057] Frame extraction interval FI: The time interval between two adjacent frame extraction processes;

[0058] Silence period SI: To prevent the algorithm from generating alarm messages too densely for the same event, a sleep time needs to be set for the algorithm after detecting a target. During this sleep time, no new alarm events will be triggered, that is, the minimum time interval between two consecutive alarm events;

[0059] Channel switching interval CSI: When the number of video channels to be detected (i.e., the total number of the multiple video input devices that need to perform frame extraction) exceeds the set maximum number of channels NC, detection needs to be performed in batches, that is, the time interval between two adjacent batches;

[0060] Detection period DP: Which days of the week to perform detection;

[0061] Detection time DT: Which time period of the day to perform detection;

[0062] Algorithm confidence threshold CT: A set threshold to filter out the detection results with a confidence level lower than this threshold;

[0063] Detection range DR: By drawing one or more polygons within the video or image frame range, the detection range is specified, and the detection results outside this range will not trigger an alarm.

[0064] Step S102: Start multiple detection algorithm instances corresponding to each of the sets of algorithms, and create a detection queue for each of the detection algorithm instances.

[0065] As Figure 2 shown, in this embodiment, the number of detection algorithm instances that start and create detection queues can be more than the total number of detection algorithm instances corresponding to all sets of algorithms, that is, it is allowed that the detection algorithm instances that start and create detection queues do not perform frame extraction processing.

[0066] Create an independent detection queue for each instance of the detection algorithm, and uniformly schedule the frame acquisition operations of multiple tasks through the subsequent global frame acquisition schedule. The detection algorithm instance only needs to obtain frame data from the queue without being aware of the device's frame acquisition logic, improving the modularity and scalability of the system.

[0067] In this embodiment, before step S102, it further includes: based on the Pytorch framework, training and tuning the YOLOv10 model on a training server; exporting the trained and tuned YOLOv10 model in ONNX format; based on RKNN Toolkit, converting the ONNX format YOLOv10 model into an inference model adapted to the edge algorithm all-in-one machine, and using the inference model as the detection algorithm instance.

[0068] Complete the customized training of the object detection model based on a training server (high-performance GPU cluster) to ensure that the model accuracy and generalization ability meet the actual scenario requirements (such as personnel intrusion detection, vehicle recognition, etc.).

[0069] When training and tuning the model, the dataset and augmentation strategy can be prepared first. The dataset is labeled actual scenario data (COCO / VOC format), and the dataset can be divided into a training set, a validation set, and a test set. The augmentation strategy can include dynamic Mosaic splicing, random HSV adjustment, or CutMix.

[0070] Hyperparameter tuning can include: learning rate, using the Cosine annealing strategy (initial value 3e-4, minimum 1e-6); loss function, SIoU loss (replacing CIoU to improve the stability of bounding box regression).

[0071] When performing distributed training and validation, the hardware configuration can be 8×A100 GPUs, with mixed-precision training (FP16 + gradient scaling); the validation metrics can be mAP@0.5 (the main accuracy metric), FPS (single-GPU inference speed, initially evaluating the feasibility of edge deployment).

[0072] By performing model format conversion (PyTorch→ONNX), cross-frame compatibility can be achieved, providing an intermediate representation for subsequent edge hardware adaptation. Then, using RKNN Toolkit, convert the ONNX model into a dedicated inference format for Rockchip NPU (such as RK3588 / RK3566) to maximize the hardware acceleration performance. In the RKNN environment, that is, in the edge algorithm all-in-one machine (the hardware used in this embodiment can be RK3588), load the detection algorithm instance and perform subsequent inference calculations.

[0073] The complete process of training through PyTorch → intermediate conversion to ONNX → edge adaptation of RKNN provides a standardized and reusable technical path for algorithm deployment in edge computing scenarios. Rapidly iterating the model structure based on the PyTorch ecosystem reflects the flexibility of training. Shielding framework differences through ONNX and supporting multi-hardware backend deployment reflect cross-platform compatibility. Utilizing the quantization and hardware acceleration capabilities of the RKNN toolchain to meet real-time requirements reflects edge efficiency.

[0074] Step S103: Based on the global frame acquisition schedule corresponding to multiple polling scheduling tasks, frame the video streams collected by multiple video input devices corresponding to each of the multiple sets of video channels to be detected, obtain multiple video frames, and place the multiple video frames into their respective corresponding detection queues; the global frame acquisition schedule includes the corresponding relationships between the task IDs of each polling scheduling task, the algorithm ID of the detection algorithm instance, the device ID of the video input device, the previous frame acquisition time, and the current frame acquisition interval.

[0075] As Figure 3 shown, according to the task ID, algorithm ID, and device ID in the task parameters, triple frame acquisition information can be formed, and this triple frame acquisition information is used as the primary key to create a data record in the frame acquisition schedule for each unique "task - algorithm - device" group, forming a global frame acquisition schedule. When starting a polling scheduling task, a data record corresponding to this polling scheduling task can be added to the global frame acquisition schedule. Conversely, when terminating a polling scheduling task, the data record corresponding to this polling scheduling task in the global frame acquisition schedule can be deleted.

[0076] By defining the frame acquisition interval (FI) in the task parameters and combining it with the global schedule for management. Different algorithm instances (such as pedestrian detection and vehicle recognition) can be configured with independent frame acquisition intervals according to different application scenarios, which can meet different task priority requirements. Exemplarily, in high-frequency alarm scenarios, the frame acquisition interval can be shortened for devices that trigger alarms to improve detection real-time performance; in low-frequency static scenarios, the frame acquisition interval can be automatically extended to reduce redundant calculations and lower device power consumption.

[0077] The global frame acquisition schedule forcibly restricts the minimum frame acquisition interval (FI) of the device and combines timestamp management, reducing the possibility of data frame loss or bandwidth congestion caused by overclocking frame acquisition of the video input device, and ensuring that key video input devices are processed within a reasonable time window.

[0078] Step S104: For each detection algorithm instance, continuously take out the video frames from the corresponding detection queue and perform detection and analysis on the taken-out video frames.

[0079] In this embodiment, through global collaborative scheduling and dynamic parameter management, continuous detection and analysis of video frames are carried out, realizing efficient allocation of computing power resources, multi-task collaborative optimization, and ensuring real-time detection, significantly improving the comprehensive performance of the video analysis system in the edge computing scenario. Upgrading the traditional static scheduling to an adaptive dynamic closed-loop control provides a highly robust solution for scenarios such as large-scale video surveillance and industrial quality inspection.

[0080] In this embodiment, after step S104, it further includes: if a preset alarm event is detected in the retrieved video frame, based on a preset confidence threshold and detection range, filtering the detection data corresponding to the retrieved video frame, storing the filtered detection data, and synchronously pushing it to the message queue.

[0081] Alarm event data (detection data) includes the following fields:

[0082] Device ID: Identifies the video input device (such as the camera number).

[0083] Algorithm ID: Identifies the detection algorithm instance that triggers the alarm (such as "Vehicle Recognition Algorithm - 001").

[0084] Label category: Alarm target category (such as "pedestrian", "vehicle").

[0085] Confidence: The confidence score of the model for the detection result (range 0 - 1).

[0086] Detection box information: The position of the target in the image, in the format of rectangular box parameters (x, y, width, height), where: (x, y): The upper left corner coordinates of the detection box (normalized values relative to the width and height of the image, range 0 - 1), width, height: The width and height of the detection box (normalized values).

[0087] When the detection algorithm instance retrieves a video frame from the detection queue and completes the analysis, if a preset alarm event (such as a pedestrian breaking in, a vehicle illegally parked, etc.) is detected, the following filtering logic can be executed:

[0088] Confidence threshold filtering: If the confidence of the detection data is lower than the algorithm confidence threshold CT (such as 0.7), it can be determined as a false alarm with low confidence and can be directly discarded, that is, the retention condition can be expressed as: confidence ≥ CT

[0089] Detection range filtering: The detection range DR is a preset area of interest, which can be a polygon area in the image (such as ROI, Region of Interest); judge the intersection of the detection box and DR, and calculate the overlapping area or intersection status of the detection box R(x, y, w, h) and the polygon DR; if there is no intersection between the detection box and DR, it can be determined that the target is in a non - area of interest, and this alarm can be discarded.

[0090] The filtered alarm data can be written into a database (such as MySQL, MongoDB) or a time - series database (such as InfluxDB), and the stored fields can include device ID, algorithm ID, timestamp, label category, confidence level, detection box coordinates, etc.

[0091] For message queue push, RabbitMQ can be used as the message queue, which supports high concurrency, reliable transmission, and message persistence.

[0092] Specifically, the synchronous push to the message queue can include: encapsulating the filtered alarm data into JSON format; publishing it to a specified Exchange (such as alarm_exchange) and binding a routing key (such as alarm.camera.{device_id}); the backend service (such as the alarm notification system) consumes the data by subscribing to the queue to achieve real - time response.

[0093] By setting double filtering (confidence level + detection range), invalid alarms can be reduced, and the signal - to - noise ratio of the system is improved. By only processing key alarm data, the storage and network bandwidth consumption are reduced. RabbitMQ can ensure that alarm information is promptly transmitted to downstream systems (such as SMS notifications, large - screen displays), achieving real - time response. The detection range DR and the confidence level threshold can be dynamically adjusted through the management platform to adapt to different scenario requirements (such as construction site safety, traffic monitoring).

[0094] In this embodiment, step S103 can specifically include the following processing: judge whether the total number of the multiple video input devices that need to capture frames exceeds the maximum channel number; if the total number does not exceed the maximum channel number, based on the global frame - capture schedule, obtain the frame - capture time requirements corresponding to each of the multiple video input devices; if the current time reaches at least one of the frame - capture time requirements, capture the video stream collected by the corresponding at least one video input device.

[0095] Count the total number of all video input devices that need to capture frames, and compare it with the preset maximum channel number (NC). If the total number ≤ NC, there is no need for batching, and the single - batch processing flow can be directly entered.

[0096] Obtain the frame capture time requirements of all video input devices from the global frame capture schedule table, filter out the video input devices whose frame capture time has reached the current time, and perform frame capture processing.

[0097] Automatically switch between single-batch or rolling-batch modes by comparing the total number of devices with NC in real time to adapt to the computing power limitations of edge devices. Based on the global frame capture schedule table, strictly restrict the frame capture interval of devices to avoid resource contention or data redundancy caused by high-frequency frame capture. In the single-batch mode, all devices are processed in parallel, reducing waiting time and improving the response speed of high-priority tasks.

[0098] In this embodiment, if the total quantity exceeds the maximum number of channels, the method includes: performing rolling-batch processing on multiple video input devices to obtain the correspondence between different processing batches and the multiple video input devices; for each current processing batch, based on the global frame capture schedule table and the correspondence, obtain the respective frame capture time requirements of the multiple video input devices in the current processing batch; if the current time reaches at least one of the frame capture time requirements, perform frame capture on the video stream collected by at least one of the video input devices corresponding to the current batch.

[0099] When the total number of video input devices that need to capture frames exceeds the maximum number of channels (NC), rolling-batch processing can dynamically allocate computing power resources to ensure that all video input devices are reasonably scheduled under limited hardware capabilities.

[0100] Divide the video input devices into multiple batches according to a rolling window, and each batch contains at most NC devices. Exemplarily, NC = 3, and the total number of video input devices = 5, including d1, d2, d3, d4, and d5. The correspondence between the processing batches and the video input devices can be expressed as:

[0101] Batch 1: {d1, d2, d3}, Batch 2: {d4, d5, d1}, Batch 3: {d2, d3, d4}, Batch 4: {d5, d1, d2},..., and so on in a cycle.

[0102] The rolling window can cover all video input devices, avoiding resource skew caused by fixed grouping. By combining the global frame capture schedule table, the video input devices in the current processing batch are triggered to capture frames only when the time requirements are met, avoiding invalid operations. Each batch processes at most NC video input devices, preventing hardware overload, and can make full use of the computing power resources of the edge algorithm all-in-one machine to detect and analyze more video data with limited resources.

[0103] In this embodiment, the task parameter further includes a channel switching interval. After taking frames from the video streams collected by at least one of the video input devices corresponding to the current batch, the following steps are further included: If the processing duration of the current processing batch meets the channel switching interval, batch switching is performed, the next processing batch is taken as the new current processing batch, and the step of, for each current processing batch, obtaining the frame-taking time requirements corresponding to each of the multiple video input devices in the current processing batch based on the global frame-taking schedule and the corresponding relationship is repeated.

[0104] Perform frame-taking operations on the devices within the current batch (check time requirements based on the global frame-taking schedule). Exemplarily: The first batch of processing devices {d1, d2, d3}, where d1 meets the frame-taking time requirements, performs frame-taking and updates its next frame-taking time.

[0105] The starting point of timing is the time counted from the moment (tstart) when the current processing batch starts to be processed. The switching condition is that batch switching is triggered when and only when the elapsed time Δt of the current processing batch ≥ CSI. Exemplarily: If CSI = 5s, then regardless of whether the current processing batch has completed the processing of all video input devices, it can be forced to switch to the next batch after 5 seconds.

[0106] When batch switching, it can be switched to the next batch in a preset order (such as circular right shift), mark the next batch as the new current processing batch, and reset the timing starting point tstart. Repeat the frame-taking time check and operation for the new batch to form a closed-loop scheduling.

[0107] In this embodiment, the obtaining of the frame-taking time requirements corresponding to each of the multiple video input devices in the current processing batch may specifically include the following processing:

[0108] For each video input device, if a preset alarm event is detected in the video frame taken out by the video input device last time, the frame-taking time requirement corresponding to the video input device is the sum of the last frame-taking time and the silence period;

[0109] For each video input device, if a preset alarm event is not detected in the video frame taken out by the video input device last time, the frame-taking time requirement corresponding to the video input device is the sum of the last frame-taking time and the frame-taking interval.

[0110] Such as Figure 3 and Figure 4As shown in the figure, after the polling scheduling task is started, the data in the global frame acquisition schedule can be traversed in units of a fixed time interval (such as 1 second). If the last frame acquisition time corresponding to the current polling scheduling task + the current frame acquisition interval ≤ the current time, the last frame acquisition time is set to the current time, and at the same time, the frame acquisition operation is executed. According to the algorithm ID, the obtained video frame and the "task - algorithm - device" triple frame acquisition information are placed into the detection queue.

[0111] Figure 3 Among them, the preset frame acquisition intervals corresponding to T1 - A1 - D1 and T1 - A1 - D2 are 5s, and the silent periods corresponding to T1 - A1 - D1 and T1 - A1 - D2 are 7s; the preset frame acquisition intervals corresponding to T2 - A1 - D1 and T2 - A1 - D3 are 7s, and the silent periods corresponding to T2 - A1 - D1 and T2 - A1 - D3 are 12s. Figure 4 Among them, the preset frame acquisition interval corresponding to the polling scheduling task 1 is 5s, the silent period corresponding to the polling scheduling task 1 is 12s, the preset frame acquisition interval corresponding to the polling scheduling task 2 is 7s, and the silent period corresponding to the polling scheduling task 1 is 15s.

[0112] After the detection queue completes the detection, if a target alarm event is detected, the detection interval in the global frame acquisition schedule is updated to the silent period according to the "task - algorithm - device" triple frame acquisition information; if no target alarm event is detected, the detection interval in the frame acquisition schedule is updated to the preset frame acquisition interval.

[0113] According to the detection result of the previous frame acquisition of the video input device (i.e., whether an alarm is triggered), the current frame acquisition frequency can be dynamically adjusted, which can achieve a balance between resource optimization and alarm response efficiency. When the detection result of the previous frame acquisition is in the alarm state, the frame acquisition interval can be extended through the silent period, reducing the processing of meaningless frames, reducing redundant calculations, reducing the consumption of computing resources, avoiding repeated alarm storms, and improving the accuracy of operation and maintenance response; when the detection result of the previous frame acquisition is in the normal state, the conventional frame acquisition interval is maintained to ensure the continuity of monitoring.

[0114] Based on the same technical concept, the present application also provides a video detection device for an edge algorithm all - in - one machine, as Figure 5 shown. The video detection device 200 for the edge algorithm all - in - one machine mainly includes:

[0115] A task parameter acquisition module 201, which is used to obtain the task parameters corresponding to each polling scheduling task. The task parameters include a set of video channels to be detected, a set of algorithms, and a frame acquisition interval. The set of video channels to be detected is a set of video input devices that need to perform target detection, the set of algorithms is a set of corresponding detection algorithm instances, and the frame acquisition interval is a parameter representing the time interval between two adjacent video stream frame acquisitions.

[0116] Create a queue module 202 for starting multiple detection algorithm instances corresponding to each of the algorithm sets and creating a detection queue for each of the detection algorithm instances;

[0117] A frame fetching module 203 for fetching video streams collected by multiple video input devices corresponding to each of the multiple sets of video channels to be detected based on a global frame fetching schedule corresponding to multiple polling scheduling tasks, obtaining multiple video frames, and placing the multiple video frames into the corresponding detection queues respectively; the global frame fetching schedule includes the corresponding relationships between the task IDs of the polling scheduling tasks, the algorithm IDs of the detection algorithm instances, the device IDs of the video input devices, the previous frame fetching time, and the current frame fetching interval;

[0118] A detection and analysis module 204 for continuously taking out the video frames from the corresponding detection queue for each of the detection algorithm instances and performing detection and analysis on the taken out video frames.

[0119] Optionally, after the detection and analysis module 204, it further includes:

[0120] A filtering and pushing module for filtering the detection data corresponding to the taken out video frames based on a preset confidence threshold and detection range when a preset alarm event is detected in the taken out video frames, storing the filtered detection data, and synchronously pushing it to a message queue.

[0121] Optionally, the task parameters further include a maximum number of channels, and the frame fetching module 203 includes:

[0122] A judgment sub-module for judging whether the total number of the multiple video input devices that need to fetch frames exceeds the maximum number of channels;

[0123] A single-round frame fetching sub-module for obtaining the frame fetching time requirements corresponding to the multiple video input devices based on the global frame fetching schedule when the total number does not exceed the maximum number of channels; if the current time reaches at least one of the frame fetching time requirements, then fetch the video streams collected by the corresponding at least one video input device.

[0124] Optionally, the frame fetching module 203 further includes:

[0125] A rolling batching sub-module for performing rolling batching processing on the multiple video input devices when the total number exceeds the maximum number of channels to obtain the corresponding relationships between different processing batches and the multiple video input devices;

[0126] The batch frame extraction sub-module is used to, for each current processing batch, based on the global frame extraction schedule and the corresponding relationship, obtain the frame extraction time requirements corresponding to each of the multiple video input devices in the current processing batch; if the current time reaches at least one of the frame extraction time requirements, frame extraction is performed on the video streams collected by at least one of the video input devices corresponding to the current batch.

[0127] Optionally, the task parameters further include a channel switching interval. After the batch frame extraction sub-module, it further includes:

[0128] The batch switching sub-module is used to perform batch switching when the processing duration of the current processing batch meets the channel switching interval, take the next processing batch as the new current processing batch, and repeat the step of, for each current processing batch, based on the global frame extraction schedule and the corresponding relationship, obtaining the frame extraction time requirements corresponding to each of the multiple video input devices in the current processing batch.

[0129] Optionally, the task parameters further include a silence period, where the silence period is the minimum time interval between two adjacent trigger alarm events. The batch frame extraction sub-module includes:

[0130] The first time requirement sub-module is used to, for each video input device, if a preset alarm event is detected in the video frame taken out by the video input device last time, the frame extraction time requirement corresponding to the video input device is the sum of the last frame extraction time and the silence period;

[0131] The second time requirement sub-module is used to, for each video input device, if a preset alarm event is not detected in the video frame taken out by the video input device last time, the frame extraction time requirement corresponding to the video input device is the sum of the last frame extraction time and the frame extraction interval.

[0132] Optionally, before the queue creation module 202, it further includes:

[0133] The training and tuning module is used to perform training and tuning processing on the YOLOv10 model in the training server based on the Pytorch framework;

[0134] The first format conversion module is used to export the YOLOv10 model after training and tuning processing to the ONNX format;

[0135] The second format conversion module is used to convert the YOLOv10 model in the ONNX format into an inference model adapted to the edge algorithm all-in-one machine based on the RKNN Toolkit, and use the inference model as the detection algorithm instance.

[0136] In one example, the modules in any of the above devices may be one or more integrated circuits configured to implement the above methods. For example: one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0137] For another example, when the modules in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0138] In this application, names may be given to various objects such as various messages / information / devices / network elements / systems / devices / actions / operations / processes / concepts, etc. It can be understood that these specific names do not constitute a limitation on the relevant objects, and the given names may change with factors such as the scenario, context, or usage habits. The understanding of the technical meaning of the technical terms in this application should be mainly determined from the functions and technical effects reflected / executed in the technical solution.

[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0140] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0141] Based on the same technical concept, this application also provides an electronic device, as Figure 6 shown, the electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.

[0142] Among them, the processor 301 is used to control the overall operation of the electronic device 300 to complete all or part of the steps in the above video detection method for the edge algorithm all-in-one machine; the memory 302 is used to store various types of data to support the operation of the electronic device 300. These data can include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a hard disk, or an optical disc, or one or more of them.

[0143] The I / O interface 303 provides an interface between the processor 301 and other interface modules. These other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 304 is used to test the wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 304 can include: a Wi-Fi component, a Bluetooth component, an NFC component.

[0144] The communication bus 305 can include a path for transmitting information between the above components. The communication bus 305 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 305 can be divided into an address bus, a data bus, a control bus, etc.

[0145] The electronic device 300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the video detection method for the edge algorithm all-in-one machine given in the above embodiments.

[0146] The electronic device 300 can include, but is not limited to, mobile terminals such as digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc., and can also be a server, etc.

[0147] Based on the same technical concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above video detection method for the edge algorithm all-in-one machine are implemented.

[0148] The computer-readable storage medium can include various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0149] The term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.

[0150] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0151] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0152] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A video detection method for an edge algorithm all-in-one machine, characterized in that, Including: For each polling scheduling task, obtain the task parameters corresponding to the polling scheduling task. The task parameters include a set of video channels to be detected, a set of algorithms, and a frame fetching interval. The set of video channels to be detected is a set of video input devices that need to perform target detection. The set of algorithms is a set of corresponding detection algorithm instances. The frame fetching interval is a parameter representing the time interval between two adjacent video stream frame fetches; Start multiple detection algorithm instances corresponding to each of the sets of algorithms, and create a detection queue for each of the detection algorithm instances; Based on the global frame fetching schedule corresponding to multiple polling scheduling tasks, fetch video frames from the video streams collected by multiple video input devices corresponding to each of the sets of video channels to be detected, obtain multiple video frames, and place the multiple video frames into their respective corresponding detection queues; The global frame fetching schedule includes the corresponding relationships between the task IDs of each of the polling scheduling tasks, the algorithm IDs of the detection algorithm instances, the device IDs of the video input devices, the previous frame fetching time, and the current frame fetching interval; For each of the detection algorithm instances, continuously fetch the video frames from the corresponding detection queue, and perform detection and analysis on the fetched video frames.

2. The video detection method for an edge algorithm all-in-one machine according to claim 1, wherein After performing the detection and analysis on the fetched video frames, it further includes: If a preset alarm event is detected in the fetched video frame, based on a preset confidence threshold and detection range, filter the detection data corresponding to the fetched video frame, store the filtered detection data, and synchronously push it to a message queue.

3. The video detection method for an edge algorithm all-in-one machine according to claim 1, characterized in that, The task parameters further include a maximum number of channels. The fetching of video frames from the video streams collected by multiple video input devices corresponding to each of the sets of video channels to be detected based on the global frame fetching schedule corresponding to multiple polling scheduling tasks includes: Determine whether the total number of the multiple video input devices that need to perform frame fetching exceeds the maximum number of channels; If the total number does not exceed the maximum number of channels, based on the global frame fetching schedule, obtain the frame fetching time requirements corresponding to each of the multiple video input devices; if the current time reaches at least one of the frame fetching time requirements, fetch the video stream of the corresponding at least one video input device.

4. The video detection method for an edge algorithm all-in-one machine according to claim 3, characterized in that, If the total number exceeds the maximum number of channels, the method includes: Perform rolling batch processing on the multiple video input devices to obtain the corresponding relationships between different processing batches and the multiple video input devices; For each current processing batch, based on the global frame fetching schedule and the corresponding relationships, obtain the frame fetching time requirements corresponding to each of the multiple video input devices in the current processing batch; if the current time reaches at least one of the frame fetching time requirements, fetch the video stream of the corresponding at least one video input device in the current batch.

5. The video detection method for an edge algorithm all-in-one machine according to claim 4, characterized in that, The task parameters further include a channel switching interval. After fetching the video stream of the corresponding at least one video input device in the current batch, it further includes: If the processing duration of the current processing batch meets the channel switching interval, batch switching is performed. The next processing batch is used as the new current processing batch, and the step of, for each current processing batch, obtaining the frame fetching time requirements corresponding to each of the multiple video input devices in the current processing batch based on the global frame fetching schedule table and the corresponding relationship is repeated.

6. The video detection method for an edge algorithm all-in-one machine according to claim 4, wherein The task parameters further include a silence period, which is the minimum time interval between two consecutive triggering of alarm events. Obtaining the frame fetching time requirements corresponding to each of the multiple video input devices in the current processing batch includes: For each of the video input devices, if a preset alarm event is detected in the video frame fetched last time by the video input device, the frame fetching time requirement corresponding to the video input device is the sum of the last frame fetching time and the silence period; For each of the video input devices, if a preset alarm event is not detected in the video frame fetched last time by the video input device, the frame fetching time requirement corresponding to the video input device is the sum of the last frame fetching time and the frame fetching interval.

7. A video detection method for an edge algorithm all-in-one machine according to claim 1, characterized in that, Before starting the multiple detection algorithm instances corresponding to each of the algorithm sets, it further includes: Based on the Pytorch framework, perform training and tuning processing of the YOLOv10 model on the training server; Export the YOLOv10 model after training and tuning processing to the ONNX format; Based on the RKNN Toolkit, convert the YOLOv10 model in the ONNX format into an inference model adapted to the edge algorithm all-in-one machine, and use the inference model as the detection algorithm instance.

8. A video detection device for an edge algorithm all-in-one machine, characterized in that, It includes: A task parameter acquisition module, configured to, for each polling scheduling task, acquire the task parameters corresponding to the polling scheduling task. The task parameters include a set of video channels to be detected, an algorithm set, and a frame fetching interval. The set of video channels to be detected is a set of video input devices that need to perform target detection. The algorithm set is a set of corresponding detection algorithm instances. The frame fetching interval is a parameter characterizing the time interval between two consecutive video stream frame fetches; A queue creation module, configured to start the multiple detection algorithm instances corresponding to each of the algorithm sets and create a detection queue for each of the detection algorithm instances; A frame fetching module, configured to, based on the global frame fetching schedule table corresponding to multiple polling scheduling tasks, fetch video streams collected by multiple video input devices corresponding to each of the sets of video channels to be detected to obtain multiple video frames, and place the multiple video frames into the corresponding detection queues respectively; The global frame fetching schedule table includes the corresponding relationship between the task ID of each polling scheduling task, the algorithm ID of the detection algorithm instance, the device ID of the video input device, the last frame fetching time, and the current frame fetching interval; A detection and analysis module, configured to, for each of the detection algorithm instances, continuously fetch the video frames from the corresponding detection queue and perform detection and analysis on the fetched video frames.

9. An electronic device, characterized in that, Comprising a processor and a memory, the processor being coupled to the memory; The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Comprising a computer program or instructions, when the computer program or instructions are run on a computer, causing the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • NPU high-concurrency reasoning method and system based on single-pipeline multi-channel video stream batch processing

    CN121212238A

  • Npu high concurrency inference method and system based on single-pipeline multi-path video stream batch processing

    CN121212238B