Target detection method and system based on edge calculation
Through a multi-deep learning model and a multi-frame parallel inference architecture, combined with a detection and tracking algorithm, the problems of insufficient hardware utilization and single object detection in edge computing are solved, efficient and real-time object detection and tracking are achieved, and duplicate alarm rate is reduced.
Patent Information
- Application Number
- CN202510178076.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, a single model in edge computing processes video streams leads to insufficient hardware usage, difficult to meet high performance and multitasking requirements, and the inability to effectively detect target locations and track targets, resulting in the problems of duplicate alarms and low detection rates.
Using a multi-deep learning model and a multi-frame parallel inference architecture, video streams are processed in parallel through multiple inference threads, target detection and tracking is achieved by combining detection and tracking algorithms, video decoding and encoding is used by GStreamer and MPP hardware accelerator, and RTSP server is deployed on local devices.
It improves the usage rate of NPU and GPU, realizes low-power, low-cost real-time video stream processing, and can complete object detection and tracking at the edge, reducing the repeated alarm rate and improving the detection rate.
Smart Images

Figure CN120259933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing and computer vision technology, and in particular to a target detection method and system based on edge computing. Background Art
[0002] Most current technologies rely on a single model to process video streams, which is difficult to meet the requirements of high performance and multi-tasking. In commercial development and edge deployment, this single model approach to processing video streams has the following problems: first, it may lead to insufficient hardware utilization of the NPU and GPU; second, if additional target detection objects need to be added, the entire dataset must be re-annotated, which is not only costly, but also not conducive to dataset maintenance and has poor flexibility.
[0003] In the prior art, a Chinese invention patent document with publication number CN118279785A and publication date July 2, 2024 is proposed to solve the above-mentioned technical problems. The technical solution disclosed in the patent document is as follows: a multi-channel video stream target recognition method based on an edge computing device, which uses multi-threading to access multiple video streams, calls multiple processes in parallel to accelerate target recognition for multiple target detection models, and performs a highly flexible customized comprehensive judgment on the obtained recognition results according to user needs, realizes real-time analysis of video streams, and provides corresponding alarm analysis data and real-time recognition of video streams.
[0004] The above technical solution will have the following problems during actual use: the target recognition method is single and only supports binary target detection. It does not have the ability to detect the position of the target in the image and the relative position relationship of different targets. As a result, this method is only applicable to simple logical judgments of "yes" or "no", and cannot determine the specific location of a person or object, nor can it make further judgments based on the correlation of the location. In addition, binary recognition cannot track the detected object, making it difficult to balance the detection rate and alarm frequency, which means that increasing the detection rate will lead to repeated alarms, and reducing repeated alarms will reduce the detection rate. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a target detection method and system based on edge computing, which can determine the specific location of the detection object and track the detection object.
[0006] The present invention is achieved by adopting the following technical solutions: A target detection method based on edge computing includes the following steps: Step S1. Obtain a video stream, decode the video stream into a picture in BGR format, and obtain the original picture; Step S2. Preprocessing the original image to obtain an input image; Step S3. Input multiple frames of input pictures into different inference threads respectively. Each inference thread includes multiple independently running deep learning models; the input pictures of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to achieve parallel inference of multiple deep learning models and multiple frames of input pictures; meanwhile, save the corresponding original pictures. Step S4. Perform non-maximum suppression (NMS) processing on each deep learning model after inference. Step S5. Wait until all deep learning models have completed NMS processing, then merge the inference results of multiple deep learning models to generate the final detection result, and match the detection result with the corresponding original picture. Step S6. Run the detection and tracking algorithm to achieve continuous tracking of the detected object, complete the active safety identification and screening, and draw corresponding prompt boxes on the original picture. Step S7. Encode the original picture with the prompt box into a video stream.
[0007] The preprocessing of the original picture specifically refers to: performing scaling and padding processing on the decoded original picture.
[0008] Before matching the detection result with the corresponding original picture, it also includes converting the detection result into a standardized format.
[0009] Completing the active safety identification and screening specifically refers to: detecting whether there are predefined violation behaviors or abnormal states. If so, record the relevant information and push the information to the specified monitoring platform or data center.
[0010] Step S3 specifically includes the following steps: Step S 31 . Initialize inference threads that are independent of each other and have the same number as the number of simultaneous inference frames; each inference thread loads multiple specified deep learning models respectively. Step S 32 . Check whether there is an idle inference thread. If there is, send a new frame of input picture and its metadata into multiple deep learning models in this inference thread at the same time, and the multiple deep learning models perform parallel inference.
[0011] When it is judged in Step S 32 that there is no idle inference thread, discard the current frame of input picture or temporarily store the current frame of input picture in a queue with a limited size, and wait to be processed when the next inference thread becomes idle.
[0012] It also includes visualizing the detection result, including the following steps: Convert the labels output by the deep learning model into corresponding Chinese descriptions through a predefined mapping table. Draw a detection box on the original image according to the detection result.
[0013] Step S6 specifically includes the following steps: Step S 61 . Initialize the tracking list and filtering conditions according to the target name set in the configuration file; Step S 62 . After obtaining the detection result of a frame of input image, traverse all the detection results in the tracking list to determine whether there is an object to be tracked; Step S 63 . If there is a tracked object in the detection result, determine whether the object meets the tracking conditions; Step S 64 . If it meets the conditions, query the tracking list to check whether the newly detected object already exists; if it exists, update the detection box position, last detection time, violation type, and violation duration of the object; if the violation time exceeds the set threshold and the alarm flag is false, trigger an alarm and set the alarm flag to true; if the object does not exist in the tracking list, add it to the list and record the violation type and relevant timestamp; Step S 65 . After traversing all the detection results, delete the objects in the tracking list whose last update time exceeds the set threshold; Step S 66 . Draw corresponding prompt boxes on the original image according to the valid objects in the tracking list and their violation situations.
[0014] During tracking, by calculating the IOU similarity between the detection box in the new frame of target detection result and all the detection boxes in the tracking list, it is determined whether the objects in the input images of different time frames are the same object.
[0015] It further includes step S8, where the encoded video stream is pushed to the RTSP server, and the RTSP server is deployed on the local device.
[0016] Multiple deep learning models in each inference thread are pre-loaded into the hardware accelerator, and the prompt box or detection box is drawn through RGA hardware acceleration.
[0017] Pull the RTSP video stream through GStreamer and decode the video stream into a BGR-format image through the MPP hardware decoder.
[0018] An object detection system based on edge computing includes a hardware layer, a middleware driver layer, a software library, and a business layer; the hardware layer includes a video decoder, a video encoder, a 2D image engine, a multi-core NPU, and a multi-core CPU, which are used for basic hardware acceleration and processing; the business layer includes an inference algorithm, a post-processing algorithm, and a communication middleware; the inference algorithm includes several inference threads, and each inference thread includes multiple independently running deep learning models, which are used to realize the real-time processing of video streams and the multi-frame parallel inference of multiple deep learning models; the post-processing algorithm includes a detection and tracking algorithm, which is used to realize the continuous tracking of detected objects, and complete the active safety identification and screening to judge whether there are predefined violation behaviors or abnormal states.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention innovatively proposes a multi-frame parallel inference architecture for multiple deep learning models, parallelizes multiple inference threads, and each inference thread simultaneously processes multiple deep learning models in parallel, improving the utilization rate of NPU and GPU, and greatly improving the processing efficiency and real-time performance of the system. The same frame of input image is sent to multiple deep learning models for parallel inference, and each deep learning model is relatively independent, greatly reducing the difficulty of deep learning model development and training and the later optimization cost. The thread pool management mechanism is used to realize the dynamic distribution and processing of image frames. Through the asynchronous waiting mechanism and fine-grained parallel strategy, the advantages of the CPU+NPU heterogeneous computing platform can be fully utilized.
[0020] Furthermore, the object detection in the present invention is completely developed based on video frames. Through technologies such as multi-process, multi-thread, and queue buffer, the AI processing is seamlessly integrated into the video stream processing pipeline, enabling the AI inference results to be displayed in the video stream in real time and meeting the requirements of real-time video display.
[0021] Furthermore, through the detection and tracking algorithm, the consistency of the object can be maintained in continuous video frames, and tracking can be maintained even when the object is briefly occluded or the detection fails. And this tracking algorithm can also be deployed on edge computing hardware with low performance, without limitations in processing time and accuracy, can complete active safety identification and screening, can effectively manage and identify unsafe actors, and can meet the requirements of high detection rate and low false alarm rate at the same time.
[0022] Even further, the above detection method can realize the complete processing of data at the edge end, and can solve the problems existing in the prior art: some real-time video streams need to be first transmitted back to the cloud for analysis and then transmitted back to the edge end for display. This process will occupy a large amount of bandwidth and processing power, and essentially does not realize the complete processing of data at the edge end, and does not realize the decoupling of the cloud and the edge end. The edge end only serves as a filter for video content in the whole method, and the core video analysis is still processed in the cloud.
[0023] In summary, through the above detection method, real-time video stream processing and object detection with low power consumption and low cost can be achieved.
[0024] 2. The present invention determines whether the input pictures in different time frames are the same object by calculating the IOU similarity between the detection boxes in the newly calculated object detection results of a frame and all the detection boxes in the tracking list, with low calculation overhead, fundamentally solving the contradiction between the detection rate and repeated alarms. This method is particularly suitable for edge computing platforms with limited computing power, providing higher performance and lower computing resource occupancy rate, taking into account high practicality and low cost.
[0025] Furthermore, the present invention maintains a dynamic tracking list to record the position, timestamp, and violation status of the detection box of the object; adopts a time threshold mechanism to avoid repeated alarms while ensuring a high detection rate.
[0026] 3. The present invention is based on a hardware-accelerated video processing pipeline of GStreamer + MPP. Seamlessly integrates AI inference into video stream processing to achieve real-time overlay display of detection results. Achieves efficient video encoding, decoding, and streaming through RGA hardware acceleration and mediaMTX.
[0027] 4. In the present invention, using the original pictures in the decoded BGR format facilitates subsequent image processing and model input.
[0028] 5. The present invention ensures the quality of the input pictures and meets the size requirements of the deep learning model by performing scaling and padding processing on the decoded original pictures.
[0029] 6. The present invention saves the original pictures during inference for subsequent drawing of detection boxes and video streaming.
[0030] 7. The present invention matches the detection results with the corresponding original pictures to ensure the synchronization of the detection results and the image frames, providing a basis for subsequent object tracking and visualization.
[0031] 8. In the present invention, the RTSP server is deployed on a local device, which can reduce network latency and simplify the system architecture of the method for implementing object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The following will further elaborate on the present invention in conjunction with the drawings in the specification and specific embodiments, where: Figure 1 is the flow schematic of the present invention Figure 1 ; Figure 2 is the flow schematic of the present invention Figure 2 ; Figure 3 Schematic diagram of multi-model inference in the present invention; Figure 4 Schematic diagram of the tracking process in the present invention; Figure 5 Schematic diagram of the operation framework structure for implementing the object detection method in the present invention. Specific implementation manners
[0033] Embodiment 1 As a basic implementation manner of the present invention, the present invention includes an object detection method based on edge computing, comprising the following steps: Step S1. Obtain a video stream, decode the video stream into a BGR-format picture to obtain an original picture.
[0034] Step S2. Preprocess the original picture to obtain an input picture.
[0035] Step S3. Respectively input multiple frames of input pictures into different inference threads, and each inference thread includes multiple independently running deep learning models; the input pictures of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to implement parallel inference of multiple deep learning models and multiple frames of input pictures; meanwhile, save the corresponding original pictures.
[0036] Step S4. Perform nms processing respectively after each deep learning model finishes inference.
[0037] Step S5. Wait until all deep learning models have completed nms processing, then merge the inference results of multiple deep learning models to generate a final detection result, and match the detection result with the corresponding original picture.
[0038] Step S6. Run a detection and tracking algorithm to realize continuous tracking of the detected object, complete active safety identification and screening, and draw corresponding prompt boxes on the original picture.
[0039] Step S7. Encode the original picture with prompt boxes into a video stream.
[0040] Embodiment 2 As a preferred implementation manner of the present invention, the present invention includes an object detection method based on edge computing, comprising the following steps: Step S1. Obtain a video stream, decode the video stream into a BGR-format picture to obtain an original picture; Step S2. Preprocess the original picture to obtain an input picture. Specifically, preprocessing the original picture refers to: performing scaling and padding processing on the decoded original picture.
[0041] Step S3. Input multiple frames of input pictures into different inference threads respectively. Each inference thread includes multiple independently running deep learning models; the input pictures of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to achieve parallel inference of multiple deep learning models and multiple frames of input pictures; meanwhile, save the corresponding original pictures. Step S4. After each deep learning model finishes inference, perform nms processing respectively. Step S5. Wait until all deep learning models have completed nms processing, then merge the inference results of multiple deep learning models to generate the final detection result. After converting the detection result into a standardized format, match the detection result with the corresponding original picture.
[0042] Step S6. Run the detection and tracking algorithm to achieve continuous tracking of the detected object, complete the active safety identification and screening, and draw the corresponding prompt box on the original picture. If a predefined violation behavior or abnormal state is detected, record the relevant information and push the information to the specified monitoring platform or data center.
[0043] Step S7. Encode the original picture with the prompt box into a video stream.
[0044] Embodiment 3 As another preferred embodiment of the present invention, the present invention includes an object detection method based on edge computing, which includes the following steps: Step S1. Obtain the video stream, decode the video stream into pictures in BGR format to obtain the original pictures.
[0045] Step S2. Preprocess the original pictures to obtain the input pictures.
[0046] Step S3. Input multiple frames of input pictures into different inference threads respectively. Each inference thread includes multiple independently running deep learning models; the input pictures of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to achieve parallel inference of multiple deep learning models and multiple frames of input pictures; meanwhile, save the corresponding original pictures. Specifically, it includes the following steps: Step S 31 . Initialize inference threads that are independent of each other and have the same number as the number of simultaneous inference frames. Each inference thread loads multiple specified deep learning models respectively.
[0047] Step S 32 . Check whether there is an idle inference thread. If so, send a new frame of input picture and its metadata into multiple deep learning models in this inference thread at the same time, and the multiple deep learning models perform parallel inference. When step S 32Determine that there is no idle inference thread, abandon the current frame input image or temporarily store the current frame input image in a queue with a limited size, and wait to be processed when the next inference thread becomes idle.
[0048] Step S4. After each deep learning model completes inference, perform NMS processing respectively.
[0049] Step S5. Wait until all deep learning models have completed NMS processing, then merge the inference results of multiple deep learning models to generate the final detection result, and match the detection result with the corresponding original image.
[0050] Step S6. Run the detection and tracking algorithm to achieve continuous tracking of the detected object, complete the active safety identification and screening, and draw corresponding prompt boxes on the original image.
[0051] Step S7. Encode the original image with the prompt box into a video stream.
[0052] Embodiment 4 As another preferred embodiment of the present invention, the present invention includes an object detection method based on edge computing, comprising the following steps: Step S1. Obtain a video stream, decode the video stream into a BGR-format image to obtain the original image.
[0053] Step S2. Preprocess the original image to obtain the input image.
[0054] Step S3. Input multiple frames of input images into different inference threads respectively. Each inference thread includes multiple independently running deep learning models; the input images of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to achieve parallel inference of multiple deep learning models and multiple frames of input images; at the same time, save the corresponding original images.
[0055] Step S4. After each deep learning model completes inference, perform NMS processing respectively.
[0056] Step S5. Wait until all deep learning models have completed NMS processing, then merge the inference results of multiple deep learning models to generate the final detection result, and match the detection result with the corresponding original image.
[0057] Step S6. Run the detection and tracking algorithm to achieve continuous tracking of the detected object, complete the active safety identification and screening, and draw corresponding prompt boxes on the original image. Specifically, it includes the following steps: Step S 61 . Initialize the tracking list and screening conditions according to the target names set in the configuration file.
[0058] Step S62 After obtaining the detection results of a frame of input image, traverse all the detection results in the tracking list to determine whether there are objects to be tracked.
[0059] Step S 63 . If there is a tracked object in the detection results, determine whether the object meets the tracking conditions.
[0060] Step S 64 . If it meets the conditions, query the tracking list to check whether the newly detected object already exists; if it exists, update the position of the detection box, the last detection time, the violation type, and the violation duration of the object; if the violation time exceeds the set threshold and the alarm flag is false, trigger an alarm and set the alarm flag to true; if the object does not exist in the tracking list, add it to the list and record the violation type and the relevant timestamp.
[0061] Step S 65 . After traversing all the detection results, delete the objects in the tracking list whose last update time exceeds the set threshold.
[0062] Step S 66 . According to the valid objects in the tracking list and their violation situations, draw corresponding prompt boxes on the original image.
[0063] Step S7. Encode the original image with the prompt box into a video stream.
[0064] Embodiment 5 As another preferred embodiment of the present invention, referring to the accompanying drawings of the specification Figure 5, the present invention includes an edge-computing-based object detection system, which from bottom to top is successively a hardware layer, a middleware driver layer, a software library, and a service layer. Among them, the hardware layer includes a video decoder, a video encoder, a 2D image engine, a multi-core NPU (Neural Network Processing Unit), and a multi-core CPU. These components are responsible for basic hardware acceleration and processing. The middleware driver layer includes a multimedia processing platform MPP (Media Processing Platform), a 2D raster graphics acceleration unit RGA (Raster Graphic Acceleration), and the neural network acceleration rknn of Rockchip. The software library includes a multimedia stream processing framework GStreamer, a real-time media server MediaMTX, a computer vision library OpenCV, and a C++ library boost. The service layer includes an inference algorithm, a post-processing algorithm, a communication middleware lcm, a communication middleware MQTT, and a communication middleware json. Among them, the inference algorithm is used to implement parallel processing of models such as object detection and pose detection, as well as real-time processing of video streams. The post-processing algorithm is used for specific algorithm services developed based on the active safety video analysis scenario, such as electronic fence, area intrusion, labor protection wear, etc. The communication middleware lcm is Lightweight Communications and Marshalling, a lightweight communication library. The communication middleware MQTT is the Message Queuing Telemetry Transport protocol. The communication middleware json is used to transmit and parse configuration files and alarm and other information in json format.
[0065] Through the above system, a complete stack from hardware to software and then to communication is realized. The system utilizes hardware acceleration (such as NPU and dedicated video codecs), provides powerful processing capabilities through the middleware driver layer and the software library, and executes advanced tasks through the inference algorithm and the post-processing algorithm. Finally, it uses standardized communication protocols to interact with other systems. Application scenarios that can achieve real-time processing of a large amount of video data while performing complex computing tasks (such as computer vision or machine learning), especially the active safety video analysis scenario, can be realized.
[0066] Specifically, the inference algorithm takes video frames as the core, pulls the RTSP video stream through Gstreamer + opencv, processes it in the way of multi-frame inference and parallel inference of multiple models (such as model a, b, c, etc.), then performs drawing through RGA hardware acceleration, and finally pushes it to other network devices through the mediaMTX RTSP server. The post-processing algorithm takes time as the core, maintains a set of unsafe behavior objects, records information such as the number, type, and generation time of each object, and the system manages the object set according to time and finally pushes the unsafe behavior alarm information. This design realizes a complete closed-loop of real-time video analysis and behavior warning.
[0067] The general high-performance edge computing box EC13211 adopts a multi-core heterogeneous architecture, with dedicated cores for specific purposes, taking into account functions such as video processing, AI inference, real-time control, TSN, and 5G communication.
[0068] Embodiment 6 As another preferred embodiment of the present invention, referring to the attached Figure 1 and the attached Figure 2 , the present invention includes an object detection method based on edge computing, comprising the following steps: Step S1. Obtain a video stream, decode the video stream into a BGR-format picture to obtain the original picture. Specifically, it includes the following steps: Step S 11 . At startup, first load the configuration file. This configuration file is usually in JSON or YAML format and contains all necessary parameter settings, such as the RTSP address of the camera, the model path, the push stream address, etc. Then, the system initializes the modules related to GStreamer push-pull stream, including setting the pipeline, configuring the codec, etc. At the same time, it also initializes the modules related to RKNN (Rockchip Neural Network Accelerator), such as loading the YOLOv5 model, setting the inference parameters, etc. In addition, the system also initializes other necessary hardware modules, such as the GPU (if available), MPP (Media Process Platform), etc.
[0069] Step S 12 . Use GStreamer to create a pipeline to pull the RTSP real-time video stream of the camera. This process includes setting the RTSP source, configuring network buffering, handling network jitter, etc.
[0070] Step S 13. Use the MPP (Media Process Platform) hardware decoder to decode the H.264 video stream into a BGR format image. MPP is a hardware acceleration module on Rockchip chips, which can efficiently perform video decoding and significantly reduce CPU usage and power consumption compared to software decoding. The decoded BGR format image is convenient for subsequent image processing and model input.
[0071] Step S2. Preprocess the original image to obtain the input image. Specifically, scale and pad the decoded original image to match the input requirements of the deep learning model. Specifically, the long side of the original image will be scaled to the input size of the AI model in RKNN format (such as 640x640 or 416x416), and at the same time, without stretching the image, the short side will be filled with black or gray to maintain the original aspect ratio of the image. This step ensures the quality of the input image and meets the size requirements of the deep learning model.
[0072] Step S3. Since multiple complex deep learning models need to perform inference and post-processing, the time consumption usually cannot meet the requirement of real-time video stream pushing at 25 frames per second. To solve this problem, the present invention innovatively adopts a scheme of combining multi-model and multi-frame simultaneous inference, which greatly improves the processing efficiency and real-time performance of the system. Specifically, refer to the attached Figure 3 description. Input multiple frames of input images into different inference threads respectively. Each inference thread includes multiple independently running deep learning models. The same-frame input images are simultaneously input into different deep learning models of the corresponding inference thread. For example, multiple models such as object detection, face recognition, and pose estimation can be run simultaneously; realize parallel inference of multiple deep learning models and multiple frames of input images. At the same time, save the corresponding original images for subsequent drawing of detection frames and video streaming. More specifically, it includes the following steps: Step S 31 . According to the number of simultaneous inference frames set in the configuration file (such as 4 or 8), the system will initialize the corresponding number of inference threads. Each inference thread is independent and can process different image frames in parallel. In each thread, the system will load multiple specified deep learning models, such as YOLOv5 for object detection, ResNet50 for image classification, OpenPose for pose estimation, etc. These models will be pre-loaded into hardware accelerators such as GPUs or RKNNs for fast switching and inference. At the same time, the system will allocate necessary memory buffers for each thread to store intermediate results and final outputs.
[0073] Step S 32. After a frame of image is preprocessed, the system checks for idle inference threads. This checking process is implemented by a thread pool manager. If there are idle threads, the new frame of image and its metadata (such as timestamp, frame number, etc.) are sent to the thread for inference. If all threads are busy, the system will handle this situation according to predefined policies. The common approach is to discard the frame of image to ensure the real-time performance of the system. However, it is also possible to choose to temporarily store it in a queue with a limited size and wait for the next thread to become idle for processing. This mechanism ensures that the system can operate stably at high frame rates without crashing or suffering severe delays due to insufficient processing power.
[0074] After receiving the image, the inference thread immediately starts the parallel processing flow. It sends the image into the inference thread pools of multiple models simultaneously. Here, a fine-grained parallel strategy is used, that is, each model runs in an independent sub-thread. The use of a thread pool can effectively manage thread resources and avoid the overhead caused by frequent creation and destruction of threads. The inference thread will wait for all models to complete inference and post-processing. This process adopts an asynchronous waiting mechanism, which can be implemented through condition variables or semaphores. This parallel processing method makes full use of the advantages of modern multi-core CPUs and CPU+NPU heterogeneous computing platforms, significantly improving the overall inference efficiency.
[0075] Step S4. After each deep learning model completes inference, NMS processing is performed respectively. Specifically, for object detection models such as YOLOv5, after a single model completes inference, the system immediately performs non-maximum suppression (NMS) processing on its multi-scale inference results. NMS is a post-processing technique used to eliminate overlapping detection boxes and retain the best detection results. After processing, the system obtains refined detection box information, including bounding box coordinates, confidence scores, and class labels. These information will be temporarily stored and wait to be integrated with the results of other models.
[0076] Step S5. Wait until all deep learning models have completed NMS processing, then merge the inference results of multiple deep learning models to generate the final detection result. This merging process involves complex decision-making logic and requires comprehensive consideration of the outputs of each model. For example, the results of object detection may need to be combined with the results of pose estimation to judge the behavior of a person; the results of image classification may be used to verify or supplement the results of object detection. The system will use predefined rules or lightweight machine learning models to perform this multi-modal information fusion to generate the final detection result.
[0077] These detection results usually include the detected target categories (labels), bounding box coordinates, and confidence scores. The system converts these results into a standardized format and matches them with the original image corresponding to the current detection result. This matching process ensures the synchronization of the detection results with the image frames, providing a basis for subsequent object tracking and visualization.
[0078] After result fusion, the system performs visualization processing. First, the English label information is replaced with Chinese fonts to improve readability. This requires a predefined mapping table to convert the labels output by the model into corresponding Chinese descriptions. Then, the system draws detection boxes, key points, or other visual cues on the original image according to the fused results. This process is completed using an efficient image processing library (such as OpenCV or RGA) to ensure that it does not become a performance bottleneck.
[0079] After processing is completed, the main thread retrieves the processing results and the image with visualization information from the inference thread. This process usually involves data transfer between threads and needs to be carefully designed to avoid unnecessary data copying. After the data transfer is completed, the main thread sets the status of the inference thread to idle and updates the status information of the thread pool. This enables the thread to be immediately allocated to process new image frames, thus ensuring the continuous and efficient operation of the system. Through this innovative multi-model and multi-frame parallel inference scheme, the system can achieve near-real-time video analysis performance with limited hardware resources.
[0080] Step S6. Run the detection and tracking algorithm to achieve continuous tracking of the detected objects, complete the active safety identification and screening, and draw corresponding prompt boxes on the original picture.
[0081] Conventional object detection algorithms can only identify targets, but repeated detection is a relatively prominent problem. In practical applications, for video-based behavior or object detection, for example, in hard hat identification, if a person without a hard hat continuously appears in the picture, when to push an alarm, and how to avoid repeated alarms for the same person are all problems that need to be solved. Therefore, it is necessary to track personnel or other target detection objects to identify the same object in multiple frames. To reduce false alarms and avoid repeated alarms, it is necessary to track the detected objects, that is, to find the same object in frames at different times, so as to achieve the tracking of all targets.
[0082] Through this tracking algorithm, the consistency of the target can be maintained in consecutive video frames, and tracking can be maintained even when the target is briefly occluded or the detection fails. Further, the system will draw corresponding prompt boxes on the original picture to display the tracked target and its ID. If predefined violations or abnormal states (such as entering a restricted area, abnormal aggregation, etc.) are detected, the system will record relevant information, including the timestamp, location, violation type, etc. This information will be pushed to the specified monitoring platform or data center through HTTP POST requests or the MQTT protocol.
[0083] Refer to the attached Figure 4 , and the detection and tracking algorithm specifically includes the following steps: Step S 61 . Initialize the tracking list and filtering conditions according to the target name set in the configuration file.
[0084] Step S 62 . After obtaining the detection results of a frame of input picture, traverse all the detection results in the tracking list to determine whether there is an object to be tracked.
[0085] Step S 63 . If there is a tracked object in the detection results, determine whether the object meets the tracking conditions.
[0086] Step S 64 . If it meets the conditions, query the tracking list to check whether the newly detected object already exists; if it exists, update the position of the detection box, the last detection time, the violation type, and the violation duration of the object; if the violation time exceeds the set threshold and the alarm flag is false, trigger an alarm and set the alarm flag to true; if the object does not exist in the tracking list, add it to the list and record the violation type and the relevant timestamp.
[0087] Step S 65 . After traversing all the detection results, delete the objects in the tracking list whose last update time exceeds the set threshold.
[0088] Step S 66 . Draw corresponding prompt boxes on the original picture according to the valid objects in the tracking list and their violation situations.
[0089] The active safety identification and screening based on the target detection results include two basic types: the presence or absence of target A, and the presence or absence of target B in target A. These two basic types can be combined in series to form a judgment logic, such as "the human detection box does not contain the safety helmet detection box", "the human detection box contains the head detection box", "the human detection box does not contain the safety helmet detection box and contains the head detection box", "the human detection box does not contain the work clothes detection box", etc. Through this method, the accuracy of target recognition can be improved.
[0090] Suppose the rule content is defined as follows: A Represents the set of detection results of the target (e.g., human detection box). B Represents the set of detection results of the target (e.g., safety helmet detection box, head detection box, work clothes detection box, etc.). Then different screening logics can be expressed by the following formula: 1. The human detection box does not contain the safety helmet detection box: Here, B Represents the safety helmet detection box.
[0091] 2. The human detection box does not contain the head detection box: Here, represents the head detection box.
[0092] 3. The human detection box does not contain the safety helmet detection box and does not contain the head detection box: Here, B 1 Represents the safety helmet detection box, B 2 Represents the head detection box.
[0093] 4. The human detection box does not contain the work clothes detection box: Here, B Represents the work clothes detection box.
[0094] The lookup method used is to compare the IOU. Its principle is to calculate the overlap degree between two detection boxes A and B. If the overlap degree is greater than the set threshold, it is determined that on the time scale, the two detection boxes describe the same actual object. Suppose the pixel area of detection box A is S1 and the pixel area of detection box B is S2, then the formula for calculating the overlap degree C is C = (S1 ∩ S2) / (S1 ∪ S2). Set the threshold β. When C > β, it is judged that two detection boxes A and B at different times represent the same actual object.
[0095] In the tracking method of object detection, the most difficult part is how to determine whether different bounding boxes in two consecutive frames represent the same person or object. Theoretically, taking a video stream of 25 frames per second as an example, the time interval between two frames is 0.04 seconds. In such a short time, the object usually has only a small movement. Therefore, by calculating the IOU similarity between the bounding box in the object detection result of the new frame and all the detection boxes in the object list, and selecting the bounding box with the maximum similarity value, it can be determined that they represent the same actual object.
[0096] Step S7. Use GStreamer to encode the original picture with a prompt box into an H.264 video stream. During this process, the system will set various attributes of the video according to the parameters specified in the configuration file, including: Video bitrate: It can be set according to the network bandwidth and picture quality requirements, such as ranging from 2Mbps to 8Mbps.
[0097] Resolution: It can maintain the original resolution or be scaled according to requirements, such as 1080p or 720p.
[0098] Frame rate: It is usually set to 25fps or 30fps to balance smoothness and bandwidth occupancy.
[0099] Key frame interval: It is generally set to 1 - 3 seconds, that is, an I frame is inserted every 25 - 90 frames.
[0100] Step S8. The encoded video stream will be pushed to an RTSP server based on mediaMTX. Although the RTSP server can be deployed on other devices, in this solution, it is chosen to be deployed on a local device to reduce network latency and simplify the system architecture. The target address for pushing the stream is set to the local machine IP (such as 192.168.1.100) or localhost, and the port is customized to 8554. In this way, other clients can access this video stream through the RTSP protocol (such as rtsp: / / 192.168.1.100:8554 / stream) to achieve the functions of remote monitoring and video distribution.
[0101] In summary, after reading the present invention document, various other corresponding transformation schemes made by those of ordinary skill in the art without creative mental labor according to the technical solutions and technical concepts of the present invention all fall within the scope protected by the present invention.
Claims
1. A target detection method based on edge computing, characterized in that: It includes the following steps: Step S1. Obtain a video stream, decode the video stream into a picture in BGR format to obtain an original picture; Step S2. Preprocess the original picture to obtain an input picture; Step S3. Input multiple frames of input pictures into different inference threads respectively. Each inference thread includes multiple independently running deep learning models; the input pictures of the same frame are simultaneously input into different deep learning models of the corresponding inference thread to implement parallel inference of multiple deep learning models and multiple frames of input pictures; at the same time, save the corresponding original pictures; Step S4. Perform nms processing on each deep learning model after it completes inference; Step S5. Wait until all deep learning models have completed nms processing, then merge the inference results of multiple deep learning models to generate a final detection result, and match the detection result with the corresponding original picture; Step S6. Run a detection and tracking algorithm to achieve continuous tracking of the detected object, complete active safety identification and screening, and draw corresponding prompt boxes on the original picture; Step S7. Encode the original picture with prompt boxes into a video stream.
2. The object detection method based on edge computing according to claim 1, wherein: The preprocessing of the original picture specifically refers to: performing scaling and padding processing on the decoded original picture.
3. The object detection method based on edge computing according to claim 1, characterized in that: Before matching the detection result with the corresponding original picture, it also includes converting the detection result into a standardized format.
4. A target detection method based on edge computing according to claim 1, characterized in that: Completing active safety identification and screening specifically refers to: detecting whether there are predefined violation behaviors or abnormal states. If so, record the relevant information and push the information to a specified monitoring platform or data center.
5. The object detection method based on edge computing according to claim 1, wherein: Step S3 specifically includes the following steps: Step S 31 . Initialize inference threads that are independently set and have the same number as the simultaneous inference detection quantity; each inference thread loads multiple specified deep learning models respectively; Step S 32 . Check if there is an idle inference thread. If so, send the new input image frame and its metadata to multiple deep learning models in this inference thread simultaneously, and the multiple deep learning models perform parallel inference.
6. The object detection method based on edge computing according to claim 5, characterized in that: When in step S 32 it is determined that there is no idle inference thread, the current frame input image is discarded or temporarily stored in a queue with a limited size, and processed when the next inference thread becomes idle.
7. A target detection method based on edge computing according to claim 1, characterized in that: It also includes visualizing the detection result, including the following steps: Convert the labels output by the deep learning model into corresponding Chinese descriptions through a predefined mapping table; Draw detection boxes on the original picture according to the detection result.
8. The object detection method based on edge computing according to claim 1, wherein: The said Step S6 specifically includes the following steps: Step S 61 . Initialize the tracking list and filtering conditions according to the target name set in the configuration file; Step S 62 . After obtaining the detection results of a frame of input image, traverse all the detection results in the tracking list to determine whether there are objects to be tracked; Step S 63 . If there is a tracked object in the detection result, determine whether the object meets the tracking conditions; Step S 64 . If the conditions are met, query the tracking list to check whether the newly detected object already exists; if it exists, update the detection box position, the last detection time, the violation type, and the violation duration of the object; if the violation time exceeds the set threshold and the alarm flag is false, trigger an alarm and set the alarm flag to true; if the object does not exist in the tracking list, add it to the list and record the violation type and the relevant timestamp; Step S 65 . After traversing all detection results, delete the objects in the tracking list whose last update time exceeds the set threshold; Step S 66 . Draw corresponding hint boxes on the original picture according to the valid objects in the tracking list and their violations.
9. The object detection method based on edge computing according to claim 8, characterized in that: During tracking, determine whether the input pictures in different time frames are the same object by calculating the IOU similarity between the detection boxes in the detection results of the new frame target and all detection boxes in the tracking list.
10. The object detection method based on edge computing according to claim 1, wherein: It also includes Step S8, where the encoded video stream is pushed to an RTSP server, and the RTSP server is deployed on a local device.
11. A target detection method based on edge computing according to claim 1, characterized in that: Multiple deep learning models in each inference thread are pre-loaded into a hardware accelerator, and the prompt boxes or detection boxes are drawn through RGA hardware acceleration.
12. A target detection method based on edge computing according to claim 1, characterized in that: Pull an RTSP video stream through GStreamer and decode the video stream into a picture in BGR format through an MPP hardware decoder.
13. An object detection system based on edge computing, characterized in that: It includes a hardware layer, a middleware driver layer, a software library, and a business layer; the hardware layer includes a video decoder, a video encoder, a 2D image engine, a multi-core NPU, and a multi-core CPU, which are used for basic hardware acceleration and processing; the business layer includes an inference algorithm, a post-processing algorithm, and a communication middleware; the inference algorithm includes a number of inference threads, and each inference thread includes multiple independently running deep learning models, which are used to achieve real-time processing of video streams and multi-frame parallel inference of multiple deep learning models; the post-processing algorithm includes a detection and tracking algorithm, which is used to achieve continuous tracking of detected objects, and complete active safety identification and screening to determine whether there are predefined violation behaviors or abnormal states.
Citation Information
Patent Citations
Multi-channel video stream target identification method based on edge computing device
CN118279785A
Cited By
Video data processing method and device, electronic equipment and storage medium
CN121486645A