Edge computing device post-processing optimization method based on target detection model

By building a confidence threshold dictionary and multi-dimensional array on edge computing devices, combining tensor broadcasting and post-processing of CUDA kernel functions to optimize the object detection model, the problem of class sensitivity is not regulated and computational efficiency is solved, and efficient and flexible object detection is achieved.

CN120580468APending Publication Date: 2025-09-02KERRY ARTIFICIAL INTELLIGENCE CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510455005.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing target detection model has technical defects in the edge computing device that category sensitivity cannot be independently regulated and computational efficiency is inefficient, resulting in an increase in target missed detection or false detection rate in specific scenarios, and the CPU-GPU data interaction delay affects performance.

Method used

By constructing a confidence threshold dictionary and multidimensional array, combining tensor broadcasting, CUDA kernel function and Numba JIT compilation, the post-processing method of the object detection model is optimized to achieve independent confidence regulation and efficient calculation of each detection category.

Benefits of technology

It improves the computing efficiency of edge computing devices, supports near-real-time processing of large-scale detection results, and reduces the error detection rate, improving computing performance and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580468A_ABST
    Figure CN120580468A_ABST
Patent Text Reader

Abstract

The invention provides an edge computing device post-processing optimization method based on a target detection model, and relates to the technical field of computer vision, and the method comprises the steps: carrying out the initialization operation of a target detection model, so as to determine an initialization parameter; performing target detection on the to-be-detected data through the target detection model to obtain original output data; and carrying out post-processing optimization on the original output data by considering the application condition of the edge computing equipment to obtain final output data. According to the method, high efficiency and flexibility of post-processing of the target detection model are realized, a large batch of detection results can be processed on edge computing equipment at a nearly real-time speed, and independent confidence regulation and control on each detection category are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a post-processing optimization method for edge computing devices based on a target detection model. Background Art

[0002] Edge computing device post-processing enables further data analysis, optimization, and decision-making on edge devices close to the data source after data collection and initial processing. This reduces reliance on the cloud or data center, improves response speed, reduces network bandwidth consumption, and enhances privacy protection. During edge computing device post-processing, machine learning models (such as object detection models) or other algorithms are used to analyze locally collected data in real time to extract valuable information or patterns.

[0003] Object detection models, based on deep learning technology, identify and locate objects in images or videos, mark the location of each object of interest, and label its category. However, current object detection systems based on object detection models have the following technical drawbacks when deployed on edge computing devices:

[0004] (1) Confidence threshold unification problem: The existing edge computing device post-processing method adopts a global unified confidence threshold, which cannot achieve independent sensitivity control of different detection categories, resulting in missed target detection or increased false detection rate in specific scenarios.

[0005] (2) Computational efficiency bottleneck: When the traditional cyclic post-processing method is executed on the CPU side, it has a serious impact on the performance of edge computing devices due to the efficiency limitations of the Python interpreter and the CPU-GPU data interaction delay. In complex systems, it may fail to meet real-time requirements.

[0006] In response to the problems of the prior art, the present invention provides an edge computing device post-processing optimization method based on a target detection model. Summary of the Invention

[0007] In response to the problems of the current existing technology, the purpose of the present invention is to solve the technical defects of the existing target detection model post-processing method, such as the category sensitivity cannot be independently adjusted and the edge device computing efficiency is low.

[0008] The present invention provides an edge computing device post-processing optimization method based on a target detection model, the method comprising the following steps:

[0009] Initialize the target detection model to determine the initialization parameters;

[0010] Performing target detection on the data to be detected by the target detection model to obtain original output data;

[0011] Considering the applicability of the edge computing device, the original output data is post-processed and optimized to obtain the final output data.

[0012] According to one embodiment of the present invention, the initialization parameters include: a confidence threshold dictionary, which is used to characterize the confidence threshold corresponding to each detection category.

[0013] According to one embodiment of the present invention, the initialization parameters further include: inference parameters, which are used to determine the confidence, IOU threshold, device selection, floating point format and data enhancement measures corresponding to the target detection model.

[0014] According to one embodiment of the present invention, the original output data is obtained by the following steps:

[0015] Inputting the data to be detected into the target detection model;

[0016] The target detection model is used to perform frame-by-frame target detection processing on the data to be detected to obtain the original output data, wherein the original output data includes: a detection box result, a detection category result, and a confidence result.

[0017] According to one embodiment of the present invention, the final output data is obtained by the following steps:

[0018] When the edge computing device is applicable only to a GPU, performing post-processing optimization on the original output data using a second optimization algorithm to obtain the final output data;

[0019] When the edge computing device is compatible with a GPU and a CPU, the original output data is post-processed and optimized using a first optimization algorithm or a third optimization algorithm to obtain the final output data.

[0020] According to one embodiment of the present invention, the first optimization algorithm comprises the following steps:

[0021] Constructing a confidence threshold matrix based on the confidence threshold dictionary;

[0022] Based on the original output data, generating a detection confidence matrix matching the dimension of the confidence threshold matrix through a tensor broadcast mechanism;

[0023] The confidence threshold matrix and the detection confidence matrix are compared, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

[0024] According to one embodiment of the present invention, the second optimization algorithm comprises the following steps:

[0025] Convert the confidence threshold dictionary into a multidimensional array and record it as a confidence threshold array;

[0026] Convert the raw output data into a multidimensional array and record it as a detection confidence array;

[0027] Based on the confidence threshold array and the detection confidence array, the filter function is compiled into a kernel function, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

[0028] According to one embodiment of the present invention, the third optimization algorithm comprises the following steps:

[0029] Convert the confidence threshold dictionary into a multidimensional array and record it as a confidence threshold array;

[0030] Convert the raw output data into a multidimensional array and record it as a detection confidence array;

[0031] Based on the confidence threshold array and the detection confidence array, the filter function is dynamically compiled into machine code, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

[0032] According to another aspect of the present invention, a storage medium is provided, which contains a series of instructions for executing the method steps described in any one of the above.

[0033] According to another aspect of the present invention, there is also provided an edge computing device post-processing optimization system based on a target detection model, which executes the method described in any one of the above items, and the system includes:

[0034] An initialization module is used to initialize the target detection model to determine the initialization parameters;

[0035] A target detection module, which is used to perform target detection on the data to be detected using the target detection model to obtain original output data;

[0036] The post-processing optimization module is used to consider the applicability of the edge computing device and perform post-processing optimization on the original output data to obtain the final output data.

[0037] This invention provides an edge computing device post-processing optimization method based on a target detection model. Compared with the existing technology, it has the following advantages:

[0038] 1) The present invention solves two major technical problems existing in the post-processing methods of existing target detection models: the category sensitivity cannot be independently adjusted and the computing efficiency of edge devices is low.

[0039] 2) The present invention proposes a category-adaptive mask matrix construction method that is compatible with CPU and GPU scenarios and improves filtering efficiency.

[0040] 3) The present invention proposes a heterogeneous computing optimization method based on just-in-time compilation, which can be applied to GPU scenarios and scenarios with a large number of detection categories. It can process a large number of detection boxes in parallel and improve computing efficiency.

[0041] 4) The present invention dynamically compiles the filter function into machine code, reducing the computational overhead of calling the interpreter and improving computational efficiency.

[0042] 5) The present invention achieves high efficiency and flexibility in post-processing of target detection models, can process large batches of detection results at near real-time speeds on edge computing devices, and supports independent confidence control for each detection category.

[0043] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0045] Figure 1 A flowchart of a post-processing optimization method for edge computing devices based on a target detection model according to an embodiment of the present invention is shown;

[0046] Figure 2 A flowchart of a post-processing optimization method for an edge computing device based on a target detection model according to another embodiment of the present invention is shown;

[0047] Figure 3 A flowchart of the steps of a first optimization algorithm according to one embodiment of the present invention is shown;

[0048] Figure 4 A flowchart of the steps of a second optimization algorithm according to one embodiment of the present invention is shown;

[0049] Figure 5 A flowchart of the steps of a third optimization algorithm according to an embodiment of the present invention is shown.

[0050] In the accompanying drawings, the same reference numerals are used for the same parts. In addition, the accompanying drawings are not drawn according to the actual scale. DETAILED DESCRIPTION

[0051] To make the objectives, technical solutions and advantages of the present invention more clear, embodiments of the present invention are described in further detail below with reference to the accompanying drawings.

[0052] The prior art (CN117292352B) discloses an obstacle recognition and avoidance method and a trolley system for open-world target detection. First, a target detector with open-world target detection capability is pre-trained, and the pre-trained Omnidata model is used to extract the geometric features of objects in the RGB image. For each RGB image, an image based on depth information and an image based on normal information are generated respectively. The Bounding Box of a known class is used as a supervisory signal to guide the depth image and the normal image to generate category-independent pseudo boxes. Then, for all pseudo boxes, the open vocabulary target detection model Detic is used to complete the classification of known categories and unknown new categories by providing a detection vocabulary, thereby completing the training of the target detector with open-world target detection capability. After completing the preset of the running path Route of the moving device, the obstacle position labels Obstaclei and Bboxi are obtained, and different operations are performed according to the category of the object label classi to complete obstacle recognition and avoidance. The system has good applicability and scalability, and is more in line with actual needs.

[0053] Prior art (US20210370993A1) discloses a real-time pixel-level railway track component detection system based on computer vision. Using improved single-stage instance segmentation models such as YOLACT-Res2Net-50 and YOLACT-Res2Net-101, the system adopts a hybrid activation function and deep learning algorithm for real-time component detection. It can operate under different lighting conditions and integrate with existing detection equipment. High precision and real-time processing are achieved, and track components are detected with 61.5bbox mAP and 59.1maskmAP, enabling fast and accurate inspection, reducing the need for extensive training and expensive equipment, and enhancing track safety by timely identifying defects.

[0054] Prior art (US20220317328A1) discloses a method, system, and computer program for detecting buried pipelines and leaks, comprising programming a drone to fly over a geographic area and capturing geophysical data using a geophysical device in the drone during the flight. Furthermore, the method includes capturing images with a camera in the drone during the flight. A machine learning model is used to identify the locations of buried pipelines and leaks based on the captured geophysical data and the captured images. Furthermore, the method includes presenting the locations of the identified buried pipelines and leaks in a map of the geographic area.

[0055] However, the above-mentioned existing technologies are unable to solve the technical defects of the existing target detection model post-processing method, such as low detection accuracy (the unification of confidence thresholds leads to missed target detection or increased false detection rate in specific scenarios) and computing efficiency bottlenecks (data interaction delays affect the performance of edge computing devices).

[0056] In response to the problems of the existing technology, the present invention provides an edge computing device post-processing optimization method based on the target detection model, which achieves high efficiency and flexibility in target detection model post-processing, can process large quantities of detection results at a near real-time speed on the edge computing device, and supports independent confidence control for each detection category.

[0057] Figure 1 A flowchart of the steps of a post-processing optimization method for an edge computing device based on a target detection model according to an embodiment of the present invention is shown.

[0058] like Figure 1 As shown, in step S1, the target detection model is initialized to determine the initialization parameters.

[0059] In one embodiment, in step S1, the initialization parameters include: a confidence threshold dictionary, which is used to represent the confidence threshold corresponding to each detection category.

[0060] In actual applications, in the Python language code environment, use the target detection model loading tool provided by Ultralytics to initialize the target detection model and read the confidence threshold for each detection category from the configuration file.

[0061] Specifically, Ultralytics is an open-source Python library for object detection and computer vision tasks, focusing on implementing the YOLO (You Only Look Once) family of algorithms. It provides a simple and easy-to-use interface for training, evaluating, and deploying YOLO models.

[0062] It should be noted that corresponding confidence thresholds can be set according to different detection categories, and the present invention does not limit the detection category and confidence threshold.

[0063] In one embodiment, in step S1, the initialization parameters further include: inference parameters, which are used to determine the confidence, IOU threshold, device selection, floating point format and data enhancement measures corresponding to the target detection model.

[0064] In actual applications, in the Python language code environment, the inference parameters of the target detection model are set through the configuration file, including confidence, IOU threshold, device selection (CPU or GPU), whether to use half-precision floating point numbers, whether to use data enhancement, etc.

[0065] Specifically, in the initialization operation, the confidence of the target detection model is set. If the value is not provided in the configuration file, the default is 0.0001; the IOU threshold of non-maximum suppression (NMS) is set. If the value is not provided in the configuration file, the default is 0.5; the floating point format is set. If the value is not provided in the configuration file, half-precision floating point numbers are used by default; the device selection is set, that is, the device used to specify the target detection model to run (such as GPU or CPU); the data enhancement measures are set, that is, whether to enable data enhancement. If the value is not provided in the configuration file, data enhancement is not used by default.

[0066] like Figure 1 As shown, in step S2, target detection is performed on the data to be detected through the target detection model to obtain original output data.

[0067] In one embodiment, in step S2, the original output data is obtained by the following steps: inputting the data to be detected into the target detection model; using the data to be detected to perform frame-by-frame target detection processing on the data to be detected to obtain the original output data, wherein the original output data includes: detection box results, detection category results and confidence results.

[0068] In actual applications, in the Python language code environment, the input data to be detected (such as a video stream) is processed frame by frame, and the target detection model is used to perform target detection to obtain the detection results (such as the original output data), including the detection box, detection category and confidence.

[0069] Furthermore, target detection models include but are not limited to: YOLO, EfficientDet, RetinaNet, DETRv2, CenterNet++, FCOS, Swin Transformer, DINO, ViTAE, and BEiT.

[0070] It should be noted that, in practical applications, the target detection model to be used can be determined according to specific needs, and the present invention does not limit the type of target detection model.

[0071] like Figure 1 As shown, in step S3, considering the applicability of the edge computing device, the original output data is post-processed and optimized to obtain the final output data.

[0072] In one embodiment, in step S3, the final output data is obtained by the following steps: when the applicability of the edge computing device is only applicable to the GPU, the original output data is post-processed and optimized by the second optimization algorithm to obtain the final output data.

[0073] In one embodiment, in step S3, the final output data is obtained through the following steps: when the edge computing device is compatible with GPU and CPU, the original output data is post-processed and optimized through the first optimization algorithm or the third optimization algorithm to obtain the final output data.

[0074] Figure 2 A flowchart of the steps of a post-processing optimization method for an edge computing device based on a target detection model according to another embodiment of the present invention is shown.

[0075] like Figure 2 As shown, in step S1, the target detection model is initialized to determine the initialization parameters.

[0076] like Figure 2 As shown, in step S2, the data to be detected is input into the target detection model, and the target detection is performed by the target detection model to obtain the original output data.

[0077] like Figure 2 As shown, in step S3, it is determined whether it is only applicable to GPU. If the judgment result is yes, the second optimization algorithm is used to post-process and optimize the original output data to obtain the final output result; if the judgment result is no, the first optimization algorithm or the third optimization algorithm is used to post-process and optimize the original output data to obtain the final output result.

[0078] In another embodiment, in step S3, the original output data is post-processed and optimized to obtain the final output data, taking into account the data volume of the data to be detected and the applicability of the edge computing device.

[0079] Specifically, in step S3, when the amount of data to be detected is greater than the preset data amount, and the applicability of the edge computing device is only applicable to GPU, the original output data is post-processed and optimized through the second optimization algorithm to obtain the final output data.

[0080] Specifically, in step S3, when the amount of data to be detected is less than the preset data amount, and the edge computing device is compatible with GPU and CPU, the original output data is post-processed and optimized through the first optimization algorithm or the third optimization algorithm to obtain the final output data.

[0081] It should be noted that the amount of detection data can be characterized by the size of the storage space occupied by the data, or by the number of detection categories. The present invention does not impose any restriction on the value of the preset data amount.

[0082] like Figure 2As shown, in step S4, performance indicators are calculated to evaluate the effect of post-processing optimization, wherein the performance indicators include: average post-processing time and frames per second.

[0083] In one embodiment, in step S4, the average post-processing time is calculated by the following expression:

[0084]

[0085] Where: avgpost represents the average post-processing time; Time total Indicates post-processing optimization time; frame count Indicates the frame number.

[0086] In one embodiment, in step S4, the number of frames per second is calculated using the following expression:

[0087]

[0088] Where: fps means frames per second; Time pipeline Represents the target detection process and post-processing optimization time.

[0089] In the Python language code environment, this paper uses the cuda.jit decorator of the Numba library to perform JIT (runtime) compilation of post-processing functions. A dedicated cache is established in the GPU video memory to store the compiled machine code. A double-buffering mechanism is designed to process video streams, ensuring time overlap between computation and data transmission. Memory pooling technology is used to reduce CUDA memory allocation overhead. CUDA kernel functions are used to process a large number of detection boxes in parallel on the GPU, improving computational efficiency.

[0090] Figure 3 A flow chart showing the steps of a first optimization algorithm according to an embodiment of the present invention is shown.

[0091] like Figure 3 As shown, in step S31, a confidence threshold matrix is ​​constructed based on the confidence threshold dictionary.

[0092] Specifically, in step S31, a confidence threshold matrix with a dimension of [C×N] is constructed based on the confidence threshold dictionary, where C is the total number of detection categories and N is the maximum number of detection frames in a single frame.

[0093] like Figure 3 As shown, in step S32, based on the original output data, a detection confidence matrix matching the dimension of the confidence threshold matrix is ​​generated through a tensor broadcast mechanism.

[0094] Specifically, in step S32, based on the raw output data, a dimensionally matched detection confidence matrix is ​​generated using PyTorch's tensor broadcast mechanism in a Python language code environment. PyTorch is an open source machine learning framework used for building, training, and deploying deep learning models.

[0095] like Figure 3 As shown, in step S33, the confidence threshold matrix and the detection confidence matrix are compared, and according to the detection category results, the detection box results with confidence results greater than the confidence threshold are filtered to obtain the final output data.

[0096] Specifically, in step S33, when the GPU is available, the vectorized confidence threshold matrix and the detection confidence matrix are placed in the GPU, and the CUDA default kernel function is used to perform matrix element-level comparison operations. The parallel processing of multiple category masks is achieved through bit operations to obtain the final output data.

[0097] The first optimization algorithm provided by this invention uses the principle of vectorized filtering and utilizes PyTorch's tensor operations to convert the confidence results and confidence thresholds into tensors. This is then filtered using vectorized operations and can be accelerated using CUDA on a GPU, enabling efficient filtering of detection results.

[0098] Figure 4 A flow chart showing the steps of a second optimization algorithm according to an embodiment of the present invention is shown.

[0099] like Figure 4 As shown, in step S41, the confidence threshold dictionary is converted into a multi-dimensional array and recorded as a confidence threshold array.

[0100] Specifically, in step S41, based on the confidence threshold dictionary, it is converted into a NumPy array in the Python language code environment to obtain a confidence threshold array.

[0101] like Figure 4 As shown, in step S42, the original output data is converted into a multi-dimensional array and recorded as a detection confidence array.

[0102] Specifically, in step S42, based on the original output data, it is converted into a NumPy array in the Python language code environment to obtain a detection confidence array.

[0103] like Figure 4 As shown, in step S43, based on the confidence threshold array and the detection confidence array, the filter function is compiled into a kernel function, and according to the detection category results, the detection box results with confidence results greater than the confidence threshold are filtered to obtain the final output data.

[0104] Specifically, in step S43, in the Python language code environment, the confidence result and the confidence threshold are converted into NumPy arrays on the CPU, and Numba is used to accelerate the filtering operation to obtain the final output data.

[0105] The second optimization algorithm provided by the present invention utilizes Numba CUDA to accelerate filtering, uses Numba's cuda.jit decorator to compile the filter function into a CUDA kernel function, and processes a large number of detection frames in parallel on the GPU to improve the efficiency of the filtering operation.

[0106] Figure 5 A flowchart of the steps of a third optimization algorithm according to an embodiment of the present invention is shown.

[0107] like Figure 5 As shown, in step S51, the confidence threshold dictionary is converted into a multi-dimensional array and recorded as a confidence threshold array.

[0108] Specifically, in step S51, based on the confidence threshold dictionary, it is converted into a NumPy array in the Python language code environment to obtain a confidence threshold array.

[0109] like Figure 5 As shown, in step S52, the original output data is converted into a multi-dimensional array and recorded as a detection confidence array.

[0110] Specifically, in step S52, based on the original output data, it is converted into a NumPy array in the Python language code environment to obtain a detection confidence array.

[0111] like Figure 5 As shown, in step S53, based on the confidence threshold array and the detection confidence array, the filter function is dynamically compiled into machine code, and according to the detection category results, the detection box results with confidence results greater than the confidence threshold are filtered to obtain the final output data.

[0112] Specifically, in step S53, in the Python language code environment, Numba's JIT is used to dynamically compile the filter function into machine code, reducing the computational overhead of calling the Python interpreter and obtaining the final output data.

[0113] The third optimization algorithm provided by the present invention uses Numba dynamic compilation to accelerate filtering, converts the confidence results and confidence thresholds into NumPy arrays, and uses Numba's JIT runtime compilation to accelerate the filtering operation and improve computing efficiency.

[0114] In order to demonstrate the performance of the post-processing optimization algorithm of the present invention, performance tests were conducted on Jetson series devices (e.g., Jetson Xavier NX) to verify the efficiency and accuracy of different filtering methods.

[0115] In one embodiment, a YOLO v11m model with 14 detection categories was used for testing, and results were obtained on a Jetson Xavier NX device. The results showed that the average post-processing time using a traditional For loop was 6.92ms, the average post-processing time using the first optimization algorithm was 3.00ms, the average post-processing time using the second optimization algorithm was 7.66ms, and the average post-processing time using the third optimization algorithm was 3.79ms.

[0116] It can be found that since the current test model has a total of 14 detection types, when the number of detection boxes output in each frame is not large, the performance of using the first and third optimization algorithms is significantly improved (compared to the traditional For loop, the processing speed is increased by about 1 times). However, the performance improvement of using the second optimization algorithm is not obvious at this time. This is because when the amount of data is small, the delay in migrating data to the GPU is much longer than the GPU calculation time.

[0117] In one embodiment, a pre-trained object detection 11m model was used with the COCO dataset, which has 80 categories, to obtain results on a Jetson Xavier NX device. The results showed that the average post-processing time using a traditional For loop was 14.68ms, the average post-processing time using the first optimization algorithm was 5.21ms (approximately 1 / 3 of the time required for a traditional For loop), the average post-processing time using the second optimization algorithm was 10.51ms (approximately 2 / 3 of the time required for a traditional For loop), and the average post-processing time using the third optimization algorithm was 5.73ms (approximately 1 / 3 of the time required for a traditional For loop).

[0118] It can be seen that when the number of detection categories reaches 80, the number of detection frames per frame increases significantly. From the results, it can be seen that the three optimization algorithms proposed in this invention can significantly improve the post-processing speed. When the data increases further, the performance improvement of the second optimization algorithm will be more obvious.

[0119] Through the above tests, the present invention can achieve the following when running the YOLOv11m model on a Jetson Xavier NX (16GB) device: the category-level detection sensitivity is independently adjustable, and the false detection rate is reduced by about 30% after parameter adjustment (COCO dataset test); the post-processing time is reduced from 6.92ms / frame of the traditional method to 3.79ms / frame, and the YOLO model inference performance is improved from 14.69FPS to 17.07FPS, and the computing performance is improved by 16.2%.

[0120] The target detection model-based edge computing device post-processing optimization method provided by the present invention may also be used in conjunction with a computer-readable storage medium having a computer program stored thereon. The computer program is executed to execute the target detection model-based edge computing device post-processing optimization method. The computer program is capable of executing computer instructions, which include computer program code. The computer program code may be in source code form, object code form, executable file, or some intermediate form.

[0121] Computer-readable storage media may include: any entity or device that can carry computer program code, recording media, USB flash drives, mobile hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0122] It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunications signals.

[0123] According to another aspect of the present invention, a post-processing optimization system for edge computing devices based on a target detection model is provided, which executes a post-processing optimization method for edge computing devices based on a target detection model. The system includes: an initialization module, a target detection module, and a post-processing optimization module.

[0124] In one embodiment, the initialization module is used to initialize the target detection model to determine the initialization parameters; the target detection module is used to perform target detection on the data to be detected through the target detection model to obtain the original output data; the post-processing optimization module is used to consider the applicability of the edge computing device, perform post-processing optimization on the original output data, and obtain the final output data.

[0125] In summary, the present invention provides an edge computing device post-processing optimization method based on a target detection model, which has the following advantages over the existing technology:

[0126] 1) The present invention solves two major technical problems existing in the post-processing methods of existing target detection models: the category sensitivity cannot be independently adjusted and the computing efficiency of edge devices is low.

[0127] 2) The present invention proposes a category-adaptive mask matrix construction method that is compatible with CPU and GPU scenarios and improves filtering efficiency.

[0128] 3) The present invention proposes a heterogeneous computing optimization method based on just-in-time compilation, which can be applied to GPU scenarios and scenarios with a large number of detection categories. It can process a large number of detection boxes in parallel and improve computing efficiency.

[0129] 4) The present invention dynamically compiles the filter function into machine code, reducing the computational overhead of calling the interpreter and improving computational efficiency.

[0130] 5) The present invention achieves high efficiency and flexibility in post-processing of target detection models, can process large batches of detection results at near real-time speeds on edge computing devices, and supports independent confidence control for each detection category.

[0131] It should be understood that the embodiments disclosed herein are not limited to the specific structures, processing steps, or materials disclosed herein, but should extend to equivalent substitutions of these features understood by those skilled in the relevant art. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting.

[0132] In the description of the present invention, unless otherwise specified, "plurality" means two or more; terms such as "upper," "lower," "left," "right," "inner," "outer," "front end," "rear end," "head," and "tail" indicate positions or relationships based on those shown in the accompanying drawings. These terms are intended solely to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, terms such as "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0133] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "connected" and "connection" should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integral connection; mechanical connection, electrical connection; direct connection, or indirect connection through an intermediary. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0134] Certain terms are used throughout this application document to indicate specific system components. As will be appreciated by those skilled in the art, different names may be used to indicate the same component, and thus this application document is not intended to distinguish between components that are only different in name but not in function. In this application document, the terms "comprise," "include," and "have" are used in an open format and should therefore be interpreted as meaning "including, but not limited to...". In addition, the terms "substantially," "substantially," or "approximately" that may be used herein refer to industry-accepted tolerances for the corresponding terms. The term "coupling," as used herein, includes direct coupling and indirect coupling via another component, element, circuit, or module, wherein for indirect coupling, the intervening component, element, circuit, or module does not change the information of the signal but can adjust its current level, voltage level, and / or power level. Inferred coupling (e.g., one element is coupled to another element by inference) includes direct and indirect coupling between two elements in the same manner as "coupling."

[0135] References in this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, appearances of the phrases "one embodiment" or "an embodiment" in various places throughout this specification do not necessarily refer to the same embodiment.

[0136] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

[0137] Although the embodiments disclosed herein are as described above, the contents described herein are merely embodiments for facilitating understanding of the present invention and are not intended to limit the present invention. Any person skilled in the art of the present invention may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed herein. However, the scope of patent protection of the present invention shall still be subject to the scope defined by the appended claims.

Claims

1. A post-processing optimization method for edge computing devices based on a target detection model, characterized in that: The method comprises the following steps: Initialize the target detection model to determine the initialization parameters; Performing target detection on the data to be detected by the target detection model to obtain original output data; Considering the applicability of the edge computing device, the original output data is post-processed and optimized to obtain the final output data.

2. The method according to claim 1, wherein The initialization parameters include: a confidence threshold dictionary, which is used to characterize the confidence threshold corresponding to each detection category.

3. The method according to claim 2, wherein The initialization parameters also include: inference parameters, which are used to determine the confidence, IOU threshold, device selection, floating point format and data enhancement measures corresponding to the target detection model.

4. The method according to claim 3, wherein The original output data is obtained by the following steps: Inputting the data to be detected into the target detection model; The target detection model is used to perform frame-by-frame target detection processing on the data to be detected to obtain the original output data, wherein the original output data includes: a detection box result, a detection category result, and a confidence result.

5. The method according to claim 4, wherein The final output data is obtained by the following steps: When the edge computing device is applicable only to a GPU, performing post-processing optimization on the original output data using a second optimization algorithm to obtain the final output data; When the edge computing device is compatible with a GPU and a CPU, the original output data is post-processed and optimized using a first optimization algorithm or a third optimization algorithm to obtain the final output data.

6. The method according to claim 5, wherein The first optimization algorithm comprises the following steps: Constructing a confidence threshold matrix based on the confidence threshold dictionary; Based on the original output data, generating a detection confidence matrix matching the dimension of the confidence threshold matrix through a tensor broadcast mechanism; The confidence threshold matrix and the detection confidence matrix are compared, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

7. The method according to claim 5, wherein The second optimization algorithm comprises the following steps: Convert the confidence threshold dictionary into a multidimensional array and record it as a confidence threshold array; Convert the raw output data into a multidimensional array and record it as a detection confidence array; Based on the confidence threshold array and the detection confidence array, the filter function is compiled into a kernel function, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

8. The method according to claim 5, wherein The third optimization algorithm comprises the following steps: Convert the confidence threshold dictionary into a multidimensional array and record it as a confidence threshold array; Converting the raw output data into a multidimensional array and recording the result as a detection confidence array; Based on the confidence threshold array and the detection confidence array, the filter function is dynamically compiled into machine code, and according to the detection category result, the detection box results whose confidence results are greater than the confidence threshold are screened to obtain the final output data.

9. A storage medium, characterized in that: It contains a series of instructions for executing the method steps according to any one of claims 1 to 8.

10. A post-processing optimization system for edge computing devices based on a target detection model, characterized in that: The method according to any one of claims 1 to 8 is performed, wherein the system comprises: An initialization module is used to initialize the target detection model to determine the initialization parameters; A target detection module, which is used to perform target detection on the data to be detected using the target detection model to obtain original output data; The post-processing optimization module is used to consider the applicability of the edge computing device and perform post-processing optimization on the original output data to obtain the final output data.

Citation Information

Patent Citations

  • Obstacle recognition and avoidance method and car system for open world target detection

    CN117292352B

  • Computer vision based real-time pixel-level railroad track components detection system

    US20210370993A1

  • Detection of buried pipelines and spills

    US20220317328A1