Target detection method, terminal device and storage medium based on heterogeneous platform
By adopting a heterogeneous platform in the terminal device and working together with the first processor and the second processor, the problem of performance degradation in a single processor during multi-object detection is solved, and efficient object detection and low-power operation is achieved.
Patent Information
- Application Number
- CN202080000866.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-08-20
AI Technical Summary
The prior art causes the problem of degradation of processor performance when detecting multiple targets using a single processor.
By employing a heterogeneous platform in the terminal device, using a first processor and a second processor to work together, the first processor is responsible for preprocessing of video stream images and execution of object detection models, and the second processor is responsible for feature point extraction and complex calculations to reduce the operating pressure of a single processor.
It is achieved by reducing the processor operating pressure through heterogeneous processing when the target number is large, and reducing power consumption when the target number is small, which improves the overall target detection efficiency and system performance.
Smart Images

Figure CN114072854B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a target detection method based on a heterogeneous platform, a terminal device, and a computer-readable storage medium. Background Art
[0002] In order to improve the practical application of AI (Artificial Intelligence) technology, many processors are embedded with NPU (Neural-network Processing Unit) specially designed for neural network computing. Different from the traditional X86 (The X86 architecture, the computer language instruction set executed by the microprocessor) architecture or ARM (a 32-bit reduced instruction set architecture) architecture CPU (Central Processing Unit) or GPU (Graphics Processing Unit) based on parallel core computing, NPU is specially designed for convolutional networks and can handle CNN (Convolutional Neural Networks) calculations with high computational complexity.
[0003] Therefore, how to achieve target detection based on multiple processing in heterogeneous platforms has become an urgent problem to be solved. Summary of the invention
[0004] The present application provides a target detection method, terminal device and storage medium based on a heterogeneous platform, which can solve the technical problem of reduced processor performance when detecting a large number of targets using a single processor.
[0005] In one aspect, a terminal device is provided, comprising: a first processor and a second processor, wherein:
[0006] The first processor includes a general calculator and a first memory storing a first target detection model, wherein the general calculator is configured to: receive a video stream image, and pre-process the N+Kth frame image in the video stream image, and when it is determined that the number of targets in the Nth frame image in the video stream image is greater than or equal to the target threshold, send the pre-processed N+Kth frame image to the second processor through the first memory, wherein N and K are positive integers;
[0007] The second processor includes a second memory storing a second target detection model, wherein the second processor is configured to: receive the preprocessed N+Kth frame image sent by the first memory, extract feature points of the N+Kth frame image based on the second target detection model, and store the obtained feature points of the N+Kth frame image in the second memory;
[0008] The universal calculator is further configured to: read feature points of the N+Kth frame image from the second memory through the first memory, and determine a target frame in the N+Kth frame image according to the feature points;
[0009] The general calculator is also configured to: when it is determined that the number of targets in the Nth frame image in the video stream image is less than the target threshold, perform feature point extraction on the preprocessed N+Kth frame image based on the first target detection model stored in the first memory to obtain feature points of the N+Kth frame image, and determine the target box in the N+Kth frame image based on the feature points.
[0010] On the one hand, a target detection method based on a heterogeneous platform is provided, wherein the heterogeneous platform includes a first processor and a second processor, and the target detection method includes:
[0011] Acquire a video stream image to be detected, and the first processor preprocesses the N+Kth frame image in the video stream image to be detected;
[0012] Determine whether the number of targets in the Nth frame image in the video stream image is greater than or equal to a target threshold;
[0013] If yes, the pre-processed N+Kth frame image is sent to the second processor; wherein the second processor extracts feature points from the N+Kth frame image;
[0014] The first processor obtains the feature points of the N+Kth frame image obtained by the operation processing of the second processor;
[0015] If not, the first processor extracts feature points from the preprocessed N+Kth frame image to obtain feature points of the N+Kth frame image;
[0016] A target frame in the N+Kth frame image is determined according to the feature points of the N+Kth frame image.
[0017] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned target detection method based on a heterogeneous platform is implemented.
[0018] In another aspect, another terminal device is provided, comprising:
[0019] a first processor and a second processor;
[0020] A memory communicatively connected to the first processor and the second processor; wherein,
[0021] The memory stores instructions that can be executed by the first processor and the second processor, and the instructions are executed by the first processor and the second processor so that the first processor and the second processor can perform the above-mentioned target detection method based on a heterogeneous platform.
[0022] The technical solution of the embodiment of the present application is to receive the video stream image through the first processor, and when performing target detection on the N+K frame image in the video stream image, the number of targets in the N frame image can be first determined. When the number of targets in the N frame image is greater than or equal to the target threshold, the first processor can send the preprocessed N+K frame image to the second processor, so that the second processor cooperates with the first processor to complete the target detection of the N+K frame image. The use of the heterogeneous mode can reduce the operating pressure of the first processor; when the number of targets in the N frame image is less than the target threshold, the first processor can perform target detection on the N+K frame image alone, thereby reducing power consumption.
[0023] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0025] Figure 1 It is a structural block diagram of a terminal device according to an embodiment of the present application.
[0026] Figure 2 It is a structural block diagram of a terminal device according to another embodiment of the present application.
[0027] Figure 3 is a flow chart of a target detection method based on a heterogeneous platform according to an embodiment of the present application;
[0028] Figure 4 is a flow chart of a target detection method based on a heterogeneous platform according to another embodiment of the present application;
[0029] Figure 5 is an example diagram of a target detection method based on a heterogeneous platform according to an embodiment of the present application;
[0030] Figure 6is an example diagram of a processing and scheduling process of a dynamic task deployment strategy based on target detection according to an embodiment of the present application;
[0031] Figure 7 It is a structural diagram of a terminal device according to another embodiment of the present application. DETAILED DESCRIPTION
[0032] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0033] It should be noted that deep learning has been a hot topic in recent years. In the application of machine vision, the use of deep convolutional networks (CNNs) instead of hand-designed operators can extract higher-dimensional features to enhance the application of classic machine vision. Whether it is face recognition, pedestrian target detection or gesture recognition, it can be extracted from them. However, complex neural networks consume huge computing resources and have high processing time when computing applications. In practical applications, they rely on expensive server-level GPUs for computing. However, server-level computing is often affected by the network environment, is relatively unstable, and has network delays. In recent years, in order to improve the practical application of AI technology, many processors have embedded NPU neural network processors designed for neural network computing. Unlike traditional x86 or ARM architecture CPUs or GPUs based on parallel core computing, NPUs are specially designed for convolutional networks and can cope with CNN computing with high computational complexity. Therefore, how to achieve target detection based on multiple processing in heterogeneous platforms has become an urgent problem to be solved.
[0034] The present application proposes a target detection method, terminal device and computer-readable storage medium based on a heterogeneous platform, that is, debugging and optimization are performed by adopting a dynamic adjustment method. When the number of detection targets is large, multiple processors are used for collaborative processing, which can reduce the operating pressure of a single processor. When the number of people is small, a single processor is used to work independently, which can reduce power consumption.
[0035] It can be understood that a heterogeneous platform mainly refers to a computing unit with different types of instruction sets and architectures. It can be composed of a central processing unit CPU, a graphics processing unit GPU, a DSP (Digital Signal Processor), a neural network processor NPU and other processors. It should be noted that the heterogeneous platform described in the embodiment of the present application may include but is not limited to a first processor and a second processor. Among them, the first processor is a central processing unit CPU, which can act alone, and the second processor is a dedicated processor that can perform specific operations. As an example, the second processor may be a graphics processing unit GPU, a neural network processor NPU, etc. Preferably, the second processor may be a neural network processor NPU.
[0036] Specifically, the target detection method based on a heterogeneous platform, a terminal device, and a computer-readable storage medium according to an embodiment of the present application are described below with reference to the accompanying drawings.
[0037] Figure 1 is a structural block diagram of a terminal device according to an embodiment of the present application. Figure 1 As shown, the terminal device 100 may include: a first processor 10 and a second processor 20. Figure 1 As shown, the first processor 10 may include a general purpose calculator 11 and a first memory 12 storing a first target detection model; the second processor 20 may include a second memory 21 storing a second target detection model. The first processor 10 and the second processor 20 may exchange data and transfer scheduling instructions through their respective memories to complete the collaborative work of processors of different architectures.
[0038] In an embodiment of the present application, the general calculator 11 is configured to: receive a video stream image, and preprocess the N+Kth frame image in the video stream image, and when it is determined that the number of targets in the Nth frame image in the video stream image is greater than or equal to the target threshold, send the preprocessed N+Kth frame image to the second processor 20 through the first memory 12, where N and K are positive integers.
[0039] The second processor 20 is configured to: receive the preprocessed N+Kth frame image sent by the first memory 12, extract feature points of the N+Kth frame image based on the second target detection model, and store the obtained feature points of the N+Kth frame image in the second memory 21.
[0040] In the embodiment of the present application, the general calculator 11 is further configured to: read the feature points of the N+Kth frame image from the second memory 21 through the first memory 12, and determine the target box in the N+Kth frame image according to the feature points.
[0041] In an embodiment of the present application, the general calculator is also configured to: when it is determined that the number of targets in the Nth frame image in the video stream image is less than the target threshold, perform feature point extraction on the preprocessed N+Kth frame image based on the first target detection model stored in the first memory 12 to obtain feature points of the N+Kth frame image, and determine the target box in the N+Kth frame image based on the feature points.
[0042] That is to say, the general calculator 11 on the first processor 10 can receive the video stream image to be processed, and pre-process the N+Kth frame image in the video stream image, wherein the pre-processing may include but is not limited to image size adjustment and grayscale processing, etc. For example, the size of the N+Kth frame image can be adjusted to the size of the input image required by the first target detection model. If the input image of the first target detection model needs to meet the size of 832*832, the size of the N+Kth frame image can be adjusted to 832*832 first, and the N+Kth frame image can be grayscale processed. Afterwards, the general calculator 11 can determine whether the number of targets in the Nth frame image in the video stream image is greater than or equal to the target threshold. If so, the pre-processed N+Kth frame image can be sent to the second processor 20 through the first memory 12.
[0043] The second processor 20 may receive the preprocessed N+Kth frame image sent by the first memory 12 through the second memory 21, and extract feature points of the N+Kth frame image based on the second target detection model stored in the second processor 20, and store the obtained feature points of the N+Kth frame image in the second memory 21 to wait for the first processor 10 to read. After completing the feature point extraction of the N+Kth frame image, the second processor 20 may send a budget completion instruction to the first processor 10. After receiving the budget completion instruction sent by the second processor 20, the general calculator 11 in the first processor 10 sends a read command to the second processor 20 to transfer the feature points of the N+Kth frame image from the second memory 21 of the second processor 20 back to the first memory 12 of the first processor 10. The general calculator 11 may read the feature points of the N+Kth frame image from the second memory 21 through the first memory 12, and then determine the target box in the N+Kth frame image according to the feature points.
[0044] In the embodiment of the present application, when the first processor 10 determines that the number of targets in the Nth frame image in the video stream image is less than the target threshold, the first processor 10 can extract feature points of the preprocessed N+Kth frame image based on the first target detection model stored in the first memory 12 to obtain the feature points of the N+Kth frame image, and determine the target frame in the N+Kth frame image according to the feature points. That is, if the number of targets in the first N frames of the image detected exceeds the target threshold, the heterogeneous mode (i.e., the first processor and the second processor co-process) is used to perform target detection and recognition on the subsequent K frames in the video stream, and when the number of targets in the current N frames is small, the first processor can be used to perform target detection and recognition on the subsequent K frames in the video stream alone. In the embodiment of the present application, the above target threshold can be determined according to the processing performance of the first processor. As an example, in order to ensure the operating efficiency of the first processor, after testing, the target threshold is selected to be 5. The value of the above K can be determined according to actual needs, wherein the value range of K can be [1, MN], that is, the value range of K can be greater than or equal to 1 and less than or equal to MN, wherein M is the total number of frames of the video stream image.
[0045] In some embodiments of the present application, Figure 2 As shown, the terminal device 100 may also include: an image collector 30. The image collector 30 is configured to collect video stream images of the target scene, and send the collected video stream images to the first memory 12 of the first processor 10 for storage, so that the first processor 10 obtains the video stream images sent by the image collector 30 from the first memory 12, and then performs target detection on the video stream images.
[0046] It should be noted that when the first target detection model and the second target detection model process the image, multiple detection frames may appear near the same target, and the probabilities of the multiple detection frames representing the detection targets are different. The probability can be understood as the proportion of the area covering the target. The larger the probability, the larger the proportion of the area covering the target, that is, the larger the proportion of the area of the detection frame used to cover the target. The smaller the probability, the smaller the proportion of the area covering the target, that is, the smaller the proportion of the area of the detection frame used to cover the target. For adjacent or close targets, in order to avoid the redundancy of repeated detection, an NMS (Non-Maximum Suppression) algorithm may be added when determining the pedestrian target frame in the image. Optionally, in some embodiments of the present application, the general calculator 11 is further configured to: obtain multiple detection frames in the N+K frame image according to the feature points of the N+K frame image, wherein the multiple detections are used to represent different probabilities of detecting the target, and the probability is used to represent the area ratio of the covered target; select the detection frame with the highest probability of detecting the target from the multiple detection frames and determine it as the standard frame; calculate the overlap in area between the non-standard frames and the standard frames in the multiple detection frames, wherein the non-standard frames are the detection frames other than the standard frames in the multiple detection frames; delete the non-standard frames whose overlap in area with the standard frames exceeds a preset threshold, and retain the non-standard frames whose overlap in area with the standard frames does not exceed the preset threshold; determine the standard frames and the retained non-standard frames as the target frames of all targets in the N+K frame image.
[0047] That is, the general calculator 11 can select the detection frame with the largest area ratio to cover the target from multiple detection frames and determine it as the standard frame. For example, the number of detection frames obtained according to the feature points can be 5, such as detection frame 1, detection frame 2, detection frame 3, detection frame 4 and detection frame 5. The area ratio of detection frame 1 covering target 1 is 90%; the area ratio of detection frame 2 covering target 2 is 80%; the area ratio of detection frame 3 covering target 3 is 60%; the area ratio of detection frame 4 covering target 4 is 70%; and the area ratio of detection frame 5 covering target 5 is 50%. At this time, the detection frame with the largest area ratio can be selected from the 5 detection frames. If detection frame 1 is the standard frame, the standard frame can be considered as the target frame of a target in the image to be detected. Afterwards, the general calculator 11 calculates the overlap in area between the non-standard frame and the standard frame in multiple detection frames. If the overlap in area between the non-standard frame and the standard frame exceeds a preset threshold, it means that the compared non-standard frame is a redundant detection frame, and the redundant detection frame is deleted. If the overlap in area between the non-standard frame and the standard frame does not exceed the preset threshold, it means that the detection frame belongs to another target, and the detection frame needs to be retained. In this way, redundant detection frames can be deleted while ensuring that detection frames of similar objects are not deleted by mistake.
[0048] It can be understood that the multiple detection frames may belong to the same target, or there may be multiple detection frames caused by multiple targets being too close to each other. For adjacent or similar targets, in order to avoid the redundancy of repeated detection, in an embodiment of the present application, the overlap in area between non-standard frames and standard frames in multiple detection frames can be calculated, wherein the non-standard frame is a detection frame other than the standard frame in multiple detection frames. For example, taking the example given above as an example, after determining that detection frame 1 is a standard frame, the overlap in area between detection frame 2, detection frame 3, detection frame 4 and detection frame 5 and the standard frame (i.e., detection frame 1) can be calculated. If the overlap is relatively high, the detection frame compared with the standard frame can be directly confirmed as corresponding to the same target as the standard frame. That is, for different targets that are too close or adjacent, since they are too adjacent, in order to avoid the redundancy of repeated detection, the non-standard frame whose area overlap with the standard frame exceeds the preset threshold can be deleted, and the non-standard frame whose area overlap with the standard frame does not exceed the preset threshold can be retained. The retained non-standard frame can be used as the comparison content of another target. After executing all iterative detections, the target frames of all targets in the image to be detected can be determined.
[0049] In some embodiments of the present application, the general calculator 11 is further configured to: determine the target number of the targets in the N+Kth frame image according to the target frame in the N+Kth frame image. In other words, the general calculator 11 can identify the target number of all targets in the N+Kth frame image according to the target frame in the N+Kth frame image.
[0050] It should be noted that the present application determines whether the first processor alone performs target detection or the first processor and the second processor jointly perform target detection based on the number of targets in the Nth frame image. The number of targets in the Nth frame image can be determined by the first processor alone. Specifically, in some embodiments of the present application, the general calculator 11 is also configured to: pre-process the Nth frame image in the video stream image; extract feature points of the pre-processed Nth frame image based on the first target detection model stored in the first memory 12 to obtain feature points of the Nth frame image; determine the target frame in the Nth frame image based on the feature points of the Nth frame image; determine the number of targets in the Nth frame image based on the target frame in the Nth frame image.
[0051] It can be understood that the number of targets in the above-mentioned N-th frame image can also be determined by the second processor in collaboration with the first processor. Specifically, in some embodiments of the present application, the general calculator 11 is also configured to: pre-process the N-th frame image in the video stream image, and send the target detection model and the pre-processed N-th frame image to the second memory 21 through the first memory 12. The second processor 20 is also configured to: read the pre-processed N-th frame image sent by the first memory 12 from the second memory 21, and extract feature points of the N-th frame image based on the second target detection model to obtain feature points for the N-th frame image, and store the obtained feature points of the N-th frame image in the second memory 21. The general calculator 11 is also configured to: read the feature points of the N-th frame image from the second memory 21 through the first memory 12, and determine the target frame in the N-th frame image according to the feature points of the N-th frame image, and determine the number of targets in the N-th frame image according to the target frame in the N-th frame image.
[0052] It is worth noting that when the first processor and the second processor are just awakened to perform target detection on the video stream image, the model parameters of the target detection model stored in the first memory can be moved to the second memory of the second processor as the second target detection model for storage. At this time, the operation sequence to be performed by the second processor needs to be initialized to inform the second processor of the operation to be performed. The instructions of this part (wherein, the instructions need to be sent from the first processor to the second processor when the program is running) are stored on the small microprocessor core built into the second processor. Based on the instructions, the second processor can use the target detection model stored on the second processor to extract feature points of the input image, thereby assisting the first processor to complete the target detection of the image.
[0053] For the video stream detection of continuous frames, the initialization work of the second processor only needs to be performed once. The purpose of the initialization is to enable the model parameters of the target detection model stored in the first processor to be moved to the second processor to form a second target detection model and store it, so that the second processor can extract feature points of the image based on the second target detection model to assist the first processor in completing the detection of the target in the image. When the model parameters of the target detection model stored in the first processor and the instructions containing the operation order are initialized to the second processor, the work of the first processor is released. When the first processor sends the packaged data to the second processor, the second processor independently completes the reasoning calculation of the entire neural network. The result of the calculation will be temporarily stored in the memory space of the second processor (i.e., the second storage area) and wait for the first processor to read it. After the first processor receives the instruction from the second processor to complete the budget, it sends a read command to the second processor and transmits the result from the memory of the second processor back to the memory of the first processor to continue the subsequent post-processing work, such as determining the target frame of the target in the image based on the feature points in the image, and then based on the target frame, the number of all targets contained in the image can be determined.
[0054] It should be noted that, in some embodiments of the present application, the above-mentioned first target detection model and the second target detection model can be Yolov3 (You Only Look Onece, a target detection algorithm) algorithm, whose basic idea is: divide the input image into S*S grids, if the center coordinates of a certain target fall in a certain grid, then the object is predicted by the grid, and each grid will predict the corresponding bounding box to select the detected target.
[0055] In order to simplify the calculation steps and reduce the calculation cost, in some embodiments of the present application, the above-mentioned first target detection model and the second target detection module may be Tiny-Yolov3. Compared with Yolov3, the Tiny version compresses the network a lot, does not use the Res layer, and only uses two Yolov output layers of different scales. The overall idea can refer to Yolov3. In an embodiment of the present application, the size of the input image of the Tiny-Yolov3 model is 832*832, and the model training comes from the training results of the public VOC data set. The model of this implementation is relatively lightweight after conversion, and it is only less than 10MB, which is suitable for deployment on terminal devices. This implementation is based on the Tensorflow framework, and the model is characterized by lightweight. During deployment, it can well correspond to the detection requirements of terminal devices after quantization through tools.
[0056] In order to implement the above embodiments, the present application also proposes a target detection method based on a heterogeneous platform. It should be noted that the target detection method based on a heterogeneous platform of the embodiment of the present application can be applied to the terminal device of the embodiment of the present application. It should be noted that in the embodiment of the present application, the terminal device may include an image acquisition module for acquiring video images of the target scene, and a heterogeneous platform based on multi-core heterogeneity. Among them, the heterogeneous platform includes a first processor and a second processor. As an example, the first processor may be a central processing unit CPU, and the second processor may be a central processing unit CPU, a graphics processing unit GPU, a neural network processor NPU, etc. Preferably, the second processor may be a neural network processor NPU. It should be noted that the target detection method of the embodiment of the present application is described from the side of the first processor.
[0057] like Figure 3 As shown, the target detection method based on the heterogeneous platform may include:
[0058] Step 301: Obtain a video stream image to be detected, and a first processor preprocesses the N+Kth frame image in the video stream image to be detected.
[0059] In an embodiment of the present application, the above-mentioned video stream image to be detected can be a pre-photographed video stream image, or a video stream image collected in real time by an image collector. For example, assuming that the target detection method based on a heterogeneous platform in an embodiment of the present application is applied to a terminal device of a digital smart billboard, the terminal device has an image collector, and the image collector can be used to collect a video stream image of the target scene in a video acquisition mode, and the real-time collected video stream image is used as the video stream image to be detected, so that the video stream image to be detected can be obtained.
[0060] After obtaining the video stream image to be detected, target detection can be performed on the video stream image. When performing target detection on the N+Kth frame image in the video stream image, the first processor can first preprocess the N+Kth frame image so that the preprocessed N+Kth frame image can meet the input requirements of the target detection model. For example, the preprocessing may include but is not limited to image size adjustment and grayscale processing, for example, the size of the N+Kth frame image can be adjusted to the size of the input image required by the first target detection model. For example, if the input image of the first target detection model needs to meet the size of 832*832, the size of the N+Kth frame image can be first adjusted to 832*832, and the N+Kth frame image can be grayscale processed to obtain the preprocessed N+Kth frame image.
[0061] Step 302, determine whether the number of targets in the Nth frame of the video stream image is greater than or equal to the target threshold. If yes, execute step 303; if no, execute step 305.
[0062] After preprocessing the N+Kth frame image, before using the target detection model to perform target detection on the N+Kth frame image, the total number of all targets in the Nth frame image can be determined first, and then it can be determined whether the total number of all targets in the Nth frame image is greater than or equal to the target threshold. If so, execute step 303; if not, execute step 305.
[0063] Step 303: Send the pre-processed N+Kth frame image to the second processor; wherein the second processor extracts feature points from the N+Kth frame image.
[0064] It should be noted that when the terminal device uses the target detection method of the embodiment of the present application to detect the target in the video stream image, it is necessary to initialize the second processor first, that is, to move the model parameters of the target detection model stored in the first processor to the second processor to store it as a second target detection model, so that the second processor can extract feature points from the image based on the second target detection model to assist the first processor in completing the detection of the target in the image. Among them, for the video stream detection of continuous frames, the second processor initialization work only needs to be performed once.
[0065] In the embodiment of the present application, when the first processor determines that the number of targets in the Nth frame image is greater than or equal to the target threshold, that is, when the number of targets in the captured scene is large, the preprocessed N+Kth frame image may be sent to the second processor. After receiving the preprocessed N+Kth frame image sent by the first processor, the second processor may extract feature points of the N+Kth frame image using the second target detection model on the second processor, and store the extracted feature points to wait for the first processor to read them.
[0066] Step 304: The first processor obtains feature points of the N+Kth frame image obtained by operation processing by the second processor.
[0067] In an embodiment of the present application, when the first processor receives an instruction from the second processor to complete the budget, it can send a read command to the second processor to transfer the feature points of the N+Kth frame image from the memory of the second processor back to the memory of the first processor to continue the subsequent post-processing work, such as determining the target box of the target in the image based on the feature points in the image, and then determining the number of all targets contained in the image based on the target box.
[0068] Step 305: The first processor extracts feature points from the preprocessed N+Kth frame image to obtain feature points of the N+Kth frame image.
[0069] In an embodiment of the present application, when the first processor determines that the number of targets in the Nth frame image is less than the target threshold, the first processor can use the first target detection model stored in the first processor itself to extract feature points from the Nth frame image. In other words, when the number of targets in the captured scene is small, the first processor can extract feature points from the N+Kth frame image alone.
[0070] Step 306: Determine a target frame in the N+Kth frame image according to the feature points of the N+Kth frame image.
[0071] That is to say, after obtaining the feature points of the N+Kth frame image, the first processor may determine the target frame in the N+Kth frame image according to the feature points of the N+Kth frame image.
[0072] The technical solution of the embodiment of the present application is to receive the video stream image through the first processor, and when performing target detection on the N+K frame image in the video stream image, the number of targets in the N frame image can be first determined. When the number of targets in the N frame image is greater than or equal to the target threshold, the first processor can send the preprocessed N+K frame image to the second processor, so that the second processor cooperates with the first processor to complete the target detection of the N+K frame image. The use of the heterogeneous mode can reduce the operating pressure of the first processor; when the number of targets in the N frame image is less than the target threshold, the first processor can perform target detection on the N+K frame image alone, thereby reducing power consumption.
[0073] It should be noted that after the neural network processing, multiple detection frames will appear near the same object. These multiple detection frames can represent different probabilities of being the detection targets, where the probability can be understood as the proportion of the area covering the target. The larger the probability, the larger the proportion of the area covering the target, that is, the larger the proportion of the area of the detection frame used to cover the target; the smaller the probability, the smaller the proportion of the area covering the target, that is, the smaller the proportion of the area of the detection frame used to cover the target. For adjacent or close targets, in order to avoid the redundancy of repeated detection, the NMS algorithm can be added when determining the pedestrian target frame in the image. Optionally, in some embodiments of the present application, such as Figure 4 As shown, the specific implementation process of determining the target frame in the N+Kth frame image according to the feature points of the N+Kth frame image can be as follows:
[0074] Step 401, obtaining multiple detection frames in the N+Kth frame image according to the feature points of the N+Kth frame image, wherein the multiple detection frames are used to represent different probabilities of detecting the target, and the probability is used to represent the area ratio of the covered target.
[0075] In this embodiment, the probability can be understood as the proportion of the area covering the pedestrian target. For example, the larger the probability, the larger the proportion of the area covering the target, that is, the larger the proportion of the area of the detection frame used to cover the target; the smaller the probability, the smaller the proportion of the area covering the target, that is, the smaller the proportion of the area of the detection frame used to cover the target.
[0076] Step 402: From multiple detection frames, a detection frame with the highest probability of detecting a target is selected as a standard frame.
[0077] That is to say, the first processor can select the detection frame with the largest area ratio to cover the target from multiple detection frames and determine it as the standard frame. For example, the number of detection frames obtained by the first processor based on the feature points can be 5, such as detection frame 1, detection frame 2, detection frame 3, detection frame 4 and detection frame 5, and the area ratio of detection frame 1 covering target 1 is 90%; the area ratio of detection frame 2 covering target 2 is 80%; the area ratio of detection frame 3 covering target 3 is 60%; the area ratio of detection frame 4 covering target 4 is 70%; and the area ratio of detection frame 5 covering target 5 is 50%. At this time, the detection frame with the largest area ratio can be selected from the 5 detection frames. If detection frame 1 is the standard frame, the standard frame can be considered as the target frame of a target in the image to be detected.
[0078] Step 403: Calculate the overlap between the non-standard frame and the standard frame in terms of area among the multiple detection frames, wherein the non-standard frame is the detection frame other than the standard frame among the multiple detection frames.
[0079] Step 404 , the non-standard frames whose area overlaps with the standard frame exceeding a preset threshold are deleted, and the non-standard frames whose area overlaps with the standard frame not exceeding the preset threshold are retained.
[0080] That is to say, if the overlap degree of another detection frame with the standard frame in area exceeds the preset threshold, it means that the compared detection frame is a redundant detection frame, and the redundant detection frame is deleted; if the overlap degree of another detection frame with the standard frame in area does not exceed the preset threshold, it means that the detection frame belongs to another object, and the compared detection frame needs to be retained. In this way, redundant detection frames can be deleted while ensuring that detection frames of similar objects are not deleted by mistake.
[0081] Step 405 , determining the standard frame and the reserved non-standard frame as the target frame of all objects in the N+Kth frame image.
[0082] It can be understood that the multiple detection frames may belong to the same target, or there may be multiple detection frames caused by multiple targets being too close to each other. For adjacent or similar targets, in order to avoid the redundancy of repeated detection, in this embodiment, the overlap in area between the non-standard frames and the standard frames in multiple detection frames can be calculated, wherein the non-standard frames are detection frames other than the standard frames in multiple detection frames. For example, taking the example given in step 402 above as an example, after determining that detection frame 1 is a standard frame, the overlap in area between detection frame 2, detection frame 3, detection frame 4 and detection frame 5 and the standard frame (i.e., detection frame 1) can be calculated. If the overlap is relatively high, the detection frame compared with the standard frame can be directly confirmed as corresponding to the same target as the standard frame. That is, for different targets that are too close or adjacent, since they are too adjacent, in order to avoid the redundancy of repeated detection, the non-standard frame whose area overlap with the standard frame exceeds the preset threshold can be deleted, and the non-standard frame whose area overlap with the standard frame does not exceed the preset threshold can be retained. The retained non-standard frame can be used as the comparison content of another target. After executing all iterative detections, the target frames of all targets in the image to be detected can be determined.
[0083] The target detection method of the embodiment of the present application can be applied to the scene of detecting the number of targets, for example, the number of pedestrians walking on the street can be detected. In some embodiments of the present application, the number of targets in the N+K frame image can be determined based on the target frame in the N+K frame image. Among them, a target frame represents a target, so the number of target frames in the N+K frame image can be counted, and the number of targets contained in the N+K frame image can be determined based on the number.
[0084] It should be noted that the present application determines whether the first processor performs target detection alone or the first processor and the second processor perform target detection together based on the number of targets in the Nth frame image. The number of targets in the Nth frame image can be determined by the first processor alone. Specifically, in some embodiments of the present application, the first processor can preprocess the Nth frame image in the video stream image, and extract feature points from the preprocessed Nth frame image based on the first target detection model to obtain feature points of the Nth frame image, and then determine the target frame in the Nth frame image based on the feature points of the Nth frame image, and determine the number of targets in the Nth frame image based on the target frame in the Nth frame image. Among them, N can be 1, that is, when receiving the video stream image, the first processor can preprocess, extract feature points and determine the number of targets in the first frame image of the video stream image alone, and then, based on the number of targets in the first frame image, it can be determined whether the next one or more frames of images are processed by the first processor alone, or whether the second processor needs to assist the first processor to complete target detection.
[0085] It can be understood that the number of targets in the above-mentioned N-th frame image can also be determined by the second processor in collaboration with the first processor. Specifically, in some embodiments of the present application, the first processor can preprocess the N-th frame image in the video stream image, and send the preprocessed N-th frame image to the second processor, wherein the second processor extracts feature points of the N-th frame image based on the second target detection model. Afterwards, the first processor obtains the feature points of the N-th frame image obtained by the operation and processing of the second processor, and determines the target frame in the N-th frame image based on the feature points of the N-th frame image, and determines the number of targets in the N-th frame image based on the target frame in the N-th frame image.
[0086] In summary, for heterogeneous computing, there is time loss in data transmission between different processors. When the task volume is large, this part of the loss accounts for a lower proportion of the overall task processing time; when the task volume is small, due to the reduction in the amount of calculation, the transmission loss between processors accounts for a larger proportion, and this part causes redundancy in processing time. In order to reduce this partial redundancy and further improve operational efficiency. The present application carries out dynamic task deployment based on target detection, that is, the benchmark strategy designed is as follows: when the task volume is small (that is, the number of detection targets in the current scene is less than the target threshold, such as 5), a single ARM architecture can be used for direct processing to achieve target detection; when the task volume is large (that is, the number of detection targets in the current scene is greater than or equal to the target threshold), the neural network processor NPU co-processing method can be used for heterogeneous processing to achieve target detection. For example, Figure 5 As shown in the figure, when processing the video stream, the processor regards it as a sequence of continuous frames. In a 30fps video, the time interval between each frame is only 33ms. In the time scene, the number of people in the scene often changes less within this time interval. When the number of people is greater than or equal to 5, the entire system considers it as a high task load. At this time, the processing scheduling deploys the work on the neural network processor NPU, and the execution time is to actually deploy and schedule the execution of the N+Kth frame. That is, when the number of targets >= 5 is detected in the Nth frame, the NPU is deployed and executed in the N+Kth frame. The coprocessor scheduling method is used to move the complex neural network model to the neural network processor NPU side. By processing the same task in a heterogeneous way, the CPU memory usage can be reduced and the overall response speed of the algorithm can be improved. When the number of targets detected in a certain frame is less than 5, the processing of the N+Kth frame is returned to the ARM side for processing, avoiding the redundancy of data transmission between heterogeneous processors and reducing power consumption.
[0087] In some embodiments of the present application, the processing and scheduling process of the dynamic task deployment strategy based on target quantity detection can be as follows: Figure 6As shown, starting from the detection result of the Nth frame, the detection of the N+1th frame is judged first. If the detection result of the Nth frame (the detection result is the target number, such as the number of pedestrians) is greater than or equal to 5, the current task is considered to be a higher processing cycle, and the calculation is performed in NPU co-processing mode. The pre-processed data is imported into the NPU and performs high-speed and large-task calculations; when the detection result of the Nth frame is less than 5, the detection is performed on the current ARM to reduce the data handling loss in this scenario.
[0088] The heterogeneous platform of the embodiment of the present application can be implemented using Rk3399 Pro. For example, the heterogeneous acceleration platform can use a 6-core ARM chip for the main budget, configure the NPU in the 3Tops calculation for co-processing operations, and use a low-voltage version of DDR3 memory with a capacity of 3G. On this basis, ARM completes image preprocessing, post-processing and other detection operations, and the NPU performs neural network reasoning calculations. It consists of an RK3399 Pro development board, a display, and a camera. The network camera is responsible for real-time data collection, and the display displays the results through HDMI to verify the feasibility of the overall design. The system used uses fedora 28linux.
[0089] In the experiment, the reasoning part of the neural network was compared before and after acceleration. The actual performance of the algorithm is shown below. The camera uses Logitech C670i, with a resolution of 1080p and a frame rate of 30fps. In the VOC dataset, the average accuracy is 88.62%. The model can detect pedestrian targets standing, facing away from or sideways to the camera. When performing road detection, this feature is not easy to miss people passing by, and can better count the number of pedestrians passing by at the observation point. The experimental results show that when using NPU for co-processing, the processing speed is increased from 8.7fps to 27fps, and the performance on the terminal side has been greatly improved. At the same time, in the implementation of NPU, in order to reduce the amount of calculation, the original 32-bit float floating-point data is converted to 8-bit int type for processing during NPU calculation, which reduces the overall amount of calculation. Quantization processing is performed, and the number of bits of processed data is reduced, which reduces the memory usage, reduces the load of the entire system, improves the overall performance of the detection system, and reduces the memory usage by 50%.
[0090] In order to implement the above embodiment, the present application also proposes another terminal device.
[0091] Figure 7 is a structural block diagram of a terminal device according to another embodiment of the present application. Figure 7As shown, the terminal device 700 may include: a first processor 710, a second processor 720 and a memory 730. The memory 730 is in communication connection with the first processor 710 and the second processor 720. The memory 730 stores instructions that can be executed by the first processor 710 and the second processor 720, and the instructions are executed by the first processor 710 and the second processor 720 so that the first processor 710 and the second processor 720 can perform the target detection method based on a heterogeneous platform described in any of the above embodiments of the present application.
[0092] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the target detection method based on a heterogeneous platform described in any of the above embodiments of the present application is implemented.
[0093] In the description of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0094] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0095] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0096] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.
[0097] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0098] A person skilled in the art may understand that all or part of the steps in the above-mentioned embodiment method may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0099] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0100] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A terminal device, characterized in that: include: A first processor and a second processor, wherein The first processor includes a general calculator and a first memory storing a first target detection model, wherein the general calculator is configured to: receive a video stream image, and pre-process the N+Kth frame image in the video stream image, and when it is determined that the number of targets in the Nth frame image in the video stream image is greater than or equal to the target threshold, send the pre-processed N+Kth frame image to the second processor through the first memory, wherein N and K are positive integers; The second processor includes a second memory storing a second target detection model, wherein the second processor is configured to: receive the preprocessed N+Kth frame image sent by the first memory, extract feature points of the N+Kth frame image based on the second target detection model, and store the obtained feature points of the N+Kth frame image in the second memory; The universal calculator is further configured to: read feature points of the N+Kth frame image from the second memory through the first memory, and determine a target frame in the N+Kth frame image according to the feature points; The universal calculator is further configured to: when it is determined that the number of targets in the Nth frame image in the video stream image is less than the target threshold, extract feature points from the preprocessed N+Kth frame image based on the first target detection model stored in the first memory to obtain feature points of the N+Kth frame image, and determine the target box in the N+Kth frame image according to the feature points; The general purpose calculator is further configured to: According to the feature points, a plurality of detection frames in the N+Kth frame image are obtained, wherein the plurality of detection frames are used to represent different probabilities of detecting a target, and the probability is used to represent the area ratio covering the target; From the multiple detection frames, a detection frame with the highest probability of detecting the target is selected as the standard frame; Calculating the overlap between the non-standard frame and the standard frame in terms of area among the multiple detection frames, wherein the non-standard frame is a detection frame among the multiple detection frames except the standard frame; Deleting non-standard frames whose area overlap with the standard frame exceeds a preset threshold, and retaining non-standard frames whose area overlap with the standard frame does not exceed the preset threshold; The standard frame and the retained non-standard frame are determined as target frames of all objects in the N+Kth frame image.
2. The terminal device according to claim 1, characterized in that: The general purpose calculator is further configured to: According to the target frame in the N+Kth frame image, the target quantity of the targets in the N+Kth frame image is determined.
3. The terminal device according to claim 2, characterized in that: The general purpose calculator is further configured to: Preprocessing the Nth frame image in the video stream image; Extracting feature points from the preprocessed N-th frame image based on the first target detection model stored in the first memory to obtain feature points of the N-th frame image; Determine a target frame in the Nth frame image according to feature points of the Nth frame image; According to the target frame in the Nth frame image, the target quantity of the target in the Nth frame image is determined.
4. The terminal device according to claim 2, characterized in that: The universal calculator is further configured to: preprocess the Nth frame image in the video stream image, and send the target detection model and the preprocessed Nth frame image to the second memory through the first memory; The second processor is further configured to: read the preprocessed N-th frame image sent by the first memory from the second memory, extract feature points of the N-th frame image based on the second object detection model to obtain feature points for the N-th frame image, and store the obtained feature points of the N-th frame image in the second memory; The universal calculator is also configured to: read the feature points of the N-th frame image from the second memory through the first memory, determine the target frame in the N-th frame image according to the feature points of the N-th frame image, and determine the target number of targets in the N-th frame image according to the target frame in the N-th frame image.
5. The terminal device according to claim 2, characterized in that: The first processor is a central processing unit, and the second processor is a dedicated processor.
6. A target detection method based on a heterogeneous platform, characterized in that: The heterogeneous platform includes a first processor and a second processor, and the target detection method includes: Acquire a video stream image to be detected, and the first processor preprocesses the N+Kth frame image in the video stream image to be detected; Determine whether the number of targets in the Nth frame image in the video stream image is greater than or equal to a target threshold; If yes, the pre-processed N+Kth frame image is sent to the second processor; wherein the second processor extracts feature points from the N+Kth frame image; The first processor obtains the feature points of the N+Kth frame image obtained by the operation processing of the second processor; If not, the first processor extracts feature points from the preprocessed N+Kth frame image to obtain feature points of the N+Kth frame image; Determine a target frame in the N+Kth frame image according to the feature points of the N+Kth frame image; Determining a target frame in the N+Kth frame image according to the feature points of the N+Kth frame image includes: Acquire a plurality of detection frames in the N+Kth frame image according to the feature points of the N+Kth frame image, wherein the plurality of detection frames are used to represent different probabilities of detecting a target, and the probability is used to represent a proportion of an area covering the target; From the multiple detection frames, a detection frame with the highest probability of detecting the target is selected as the standard frame; Calculating the overlap between the non-standard frame and the standard frame in terms of area among the multiple detection frames, wherein the non-standard frame is a detection frame among the multiple detection frames except the standard frame; Deleting non-standard frames whose area overlap with the standard frame exceeds a preset threshold, and retaining non-standard frames whose area overlap with the standard frame does not exceed the preset threshold; The standard frame and the retained non-standard frame are determined as target frames of all objects in the N+Kth frame image.
7. The target detection method based on a heterogeneous platform according to claim 6, characterized in that: Also includes: According to the target frame in the N+Kth frame image, the target quantity of the targets in the N+Kth frame image is determined.
8. The target detection method based on a heterogeneous platform according to claim 7, characterized in that: Also includes: Preprocessing the Nth frame image in the video stream image; Extracting feature points from the preprocessed N-th frame image based on the first target detection model to obtain feature points of the N-th frame image; Determine a target frame in the Nth frame image according to feature points of the Nth frame image; According to the target frame in the Nth frame image, the target quantity of the target in the Nth frame image is determined.
9. The target detection method based on a heterogeneous platform according to claim 7, characterized in that: Also includes: Preprocessing the Nth frame image in the video stream image; Sending the preprocessed N-th frame image to the second processor; wherein the second processor extracts feature points from the N-th frame image based on the second target detection model; Acquire feature points of the Nth frame image obtained by operation processing by the second processor; Determine a target frame in the Nth frame image according to feature points of the Nth frame image; According to the target frame in the Nth frame image, the target quantity of the target in the Nth frame image is determined.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method based on a heterogeneous platform as described in any one of claims 6 to 9 is implemented.
11. A terminal device, characterized in that: include: a first processor and a second processor; A memory communicatively connected to the first processor and the second processor; wherein, The memory stores instructions that can be executed by the first processor and the second processor, and the instructions are executed by the first processor and the second processor so that the first processor and the second processor can perform the target detection method based on a heterogeneous platform as described in any one of claims 6 to 9.
Citation Information
Patent Citations
Image processing method and device, and storage medium
CN111079669A
Feature extraction method and system for point-to-read image, terminal equipment and storage medium
CN111079771A