A detection frame processing method and related device for target detection
By dynamically adjusting the crossover threshold and non-maximum value suppression processing, the detection box deviation problem when the image of the same scale scale in object detection is solved, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202210211664.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-03-04
AI Technical Summary
During the object detection process, when images of the same scale are repeatedly detected multiple times, multiple identical detection boxes will appear at the same position, resulting in a decrease in the deviation and accuracy of the detection results.
The crossover threshold is dynamically adjusted by the number of initial detection boxes, and a non-maximum suppression process is performed based on the dynamically adjusted crossover threshold and face coordinate information to eliminate the repeated redundant detection boxes.
It effectively avoids detection deviations caused by the same detection box when images of the same zoom scale are repeatedly detected, and improves the accuracy of target detection.
Smart Images

Figure CN114581983B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target recognition, and in particular to a detection frame processing method, a detection frame processing device, a server and a computer-readable storage medium for target detection. Background Art
[0002] In the face recognition system, when detecting the target data set, it is necessary to calibrate the real bounding box of the target person as clearly as possible. Generally speaking, when a face image is input into the target detection model, a large number of target boxes will be parsed and output. According to the detection branch of the model on the feature pyramid, anchors (detection boxes) of different scales will be generated for each pixel position of the original image. The specific number of target boxes is determined by the number of anchors, among which many repeated boxes of different scales will be located on the same target. The standard nms (Non-Maximum Suppression) algorithm, that is, the non-maximum suppression algorithm, or various deformed nms algorithms are usually used to remove and filter out repeated boxes and obtain the real target box.
[0003] In the related art, in order to effectively improve the detection accuracy of the target detection process, it is currently generally adopted to scale the original image into images of different resolutions, that is, multi-scale targets, and input them into the detection model in sequence to improve the detection accuracy. However, there is a situation where if the detection source is a video stream, and the content of multiple video frames within a period of time is basically unchanged, and is repeatedly sent to the detection model multiple times, it is approximately equivalent to the model input being multiple identical detection images of the same scale, and then using the vote-nms algorithm to merge and filter the detection frame, which will be different from the single input of the original image to obtain a different detection result. That is, when the image of the same scale is repeatedly detected multiple times, multiple identical frames will appear in the same position, which will be added to the detection frame collection, and the voting algorithm will be uniformly performed with other detection frames, and finally different detection frames will be output, resulting in deviations in the detection results and reducing the accuracy of target detection.
[0004] Therefore, how to improve the accuracy of the target detection frame is a key issue that technicians in this field are concerned about. Summary of the invention
[0005] The purpose of the present application is to provide a detection frame processing method, a detection frame processing device, a server and a computer-readable storage medium for target detection, so as to improve the accuracy of the detection frame during the target detection process.
[0006] In order to solve the above technical problems, the present application provides a detection frame processing method for target detection, comprising:
[0007] Inputting the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and a confidence score corresponding to each of the initial detection frames;
[0008] Taking the initial detection frame whose confidence score is greater than the confidence threshold as the detection frame;
[0009] Determine an intersection-over-union ratio threshold based on the number of the plurality of detection frames;
[0010] Based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame, non-maximum suppression processing is performed on the multiple detection frames to obtain a target detection frame.
[0011] Optionally, determining an intersection-over-union ratio threshold based on the number of the plurality of detection frames includes:
[0012] Determine whether the number of the plurality of detection frames is greater than a number threshold;
[0013] If yes, then setting the first threshold to the intersection-over-union ratio threshold;
[0014] If not, the second threshold is set to the intersection-over-union ratio threshold; wherein the first threshold is greater than the second threshold.
[0015] Optionally, performing non-maximum suppression processing on the plurality of detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame includes:
[0016] Calculate the intersection-and-union ratio of the detection frame with the highest confidence score with other detection frames to obtain the intersection-and-union ratio values corresponding to multiple detection frames;
[0017] Dividing the multiple detection frames based on the IoU values corresponding to the multiple detection frames to obtain a detection frame with the highest confidence score, a plurality of adjacent detection frames whose IoU values are greater than the IoU threshold, and a plurality of remaining detection frames whose IoU values are less than or equal to the IoU threshold;
[0018] Eliminating independent detection frames from the multiple adjacent detection frames based on the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames to obtain multiple adjacent detection frames after elimination;
[0019] Based on the weight of each detection frame, non-maximum suppression processing is performed on the detection frame with the highest confidence score and the multiple adjacent detection frames to obtain the target detection frame.
[0020] Optionally, the multiple detection frames are divided based on IoU values corresponding to the multiple detection frames to obtain a detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold, including:
[0021] The detection frames other than the detection frame with the highest confidence score whose IoU values are greater than the IoU threshold are used as the adjacent detection frames, and the remaining detection frames are used as the remaining detection frames;
[0022] If the number of adjacent detection frames is 1, send an execution command.
[0023] If the number of the adjacent detection frames is greater than 1, and the intersection-and-union ratio value of each of the adjacent detection frames is 1, delete the adjacent detection frames until the number is 1;
[0024] When the number of the adjacent detection frames is 1 and the number of the remaining detection frames is 0, the adjacent detection frames are copied as the remaining detection frames.
[0025] Optionally, based on the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames, independent detection frames in the multiple adjacent detection frames are removed to obtain the multiple adjacent detection frames after removal, including:
[0026] Determine the difference between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the plurality of adjacent detection frames;
[0027] Adjacent detection frames whose difference is greater than the difference threshold are eliminated to obtain remaining adjacent detection frames.
[0028] Optionally, also include:
[0029] Determine the degree of overlap between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the plurality of adjacent detection frames;
[0030] The weights of the adjacent detection frames whose overlap is greater than the overlap threshold are downgraded to obtain a new weight of the detection frame.
[0031] Optionally, performing non-maximum suppression processing on the detection frame with the highest confidence score and the multiple adjacent detection frames based on the weight of each detection frame to obtain the target detection frame includes:
[0032] Based on the intersection-and-union ratio value, uncertainty and variance between adjacent detection frames, each adjacent detection frame is weighted to obtain the weight of each adjacent detection frame;
[0033] Non-maximum suppression processing is performed based on the weights of the adjacent detection frames to obtain the target detection frame.
[0034] The present application also provides a detection frame processing device for target detection, comprising:
[0035] A video stream detection module is used to input the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and a confidence score corresponding to each of the initial detection frames;
[0036] A detection frame filtering module, used to take the initial detection frame whose confidence score is greater than the confidence threshold as the detection frame;
[0037] An IoU calculation module, used to determine an IoU threshold based on the number of the plurality of detection frames;
[0038] The non-maximum suppression processing module is used to perform non-maximum suppression processing on the multiple detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame.
[0039] The present application also provides a server, comprising:
[0040] Memory for storing computer programs;
[0041] A processor is used to implement the steps of the detection frame processing method as described above when executing the computer program.
[0042] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the detection frame processing method as described above are implemented.
[0043] The present application provides a method for processing a detection frame for target detection, comprising: inputting an original face video stream into a target detection model in units of frames to obtain a plurality of initial detection frames and a confidence score corresponding to each of the initial detection frames; taking an initial detection frame whose confidence score is greater than a confidence threshold as a detection frame; determining an intersection-over-union threshold based on the number of the plurality of detection frames; and performing non-maximum suppression processing on the plurality of detection frames based on the intersection-over-union threshold and face coordinate information of each detection frame to obtain a target detection frame.
[0044] The intersection-over-union (IoU) threshold is dynamically adjusted according to the number of initial detection frames, and then non-maximum suppression processing is performed based on the dynamically adjusted IoU threshold and face coordinate information to exclude repeated and redundant detection frames, avoid detection deviations caused by the same detection frame when images of the same zoom scale are repeatedly detected, and improve the accuracy of target detection.
[0045] The present application also provides a detection frame processing device, a server and a computer-readable storage medium for target detection, which have the above beneficial effects and are not elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0047] Figure 1 A flowchart of a detection frame processing method for target detection provided by an embodiment of the present application;
[0048] Figure 2 A schematic diagram of the structure of a detection frame processing device for target detection provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The core of this application is to provide a detection frame processing method, a detection frame processing device, a server and a computer-readable storage medium for target detection to improve the accuracy of the detection frame during the target detection process.
[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] In the related art, in order to effectively improve the detection accuracy of the target detection process, it is currently generally adopted to scale the original image into images of different resolutions, that is, multi-scale targets, and input them into the detection model in sequence to improve the detection accuracy. However, there is a situation where if the detection source is a video stream, and the content of multiple video frames within a period of time is basically unchanged, and is repeatedly sent to the detection model multiple times, it is approximately equivalent to the model input being multiple identical detection images of the same scale, and then using the vote-nms algorithm to merge and filter the detection frame, which will be different from the single input of the original image to obtain a different detection result. That is, when the image of the same scale is repeatedly detected multiple times, multiple identical frames will appear in the same position, which will be added to the detection frame collection, and the voting algorithm will be uniformly performed with other detection frames, and finally different detection frames will be output, resulting in deviations in the detection results and reducing the accuracy of target detection.
[0052] Therefore, the present application provides a detection frame processing method for target detection, which dynamically adjusts the intersection-and-union ratio threshold by the number of initial detection frames, and then performs non-maximum suppression processing based on the dynamically adjusted intersection-and-union ratio threshold and face coordinate information, so as to exclude repeated and redundant detection frames, avoid detection deviations caused by the same detection frame when images with the same zoom scale are repeatedly detected multiple times, and improve the accuracy of target detection.
[0053] The following describes a detection frame processing method for target detection provided by the present application through an embodiment.
[0054] Please refer to Figure 1 , Figure 1 A flowchart of a detection frame processing method for target detection provided in an embodiment of the present application.
[0055] In this embodiment, the method may include:
[0056] S101, inputting the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and a confidence score corresponding to each initial detection frame;
[0057] This step aims to input the original face video stream into the target detection model in units of frames, and obtain multiple initial detection frames and the confidence scores corresponding to each initial detection frame. That is, the original face video stream is intercepted frame by frame and input into the target detection model in sequence, and the RPN (Region Proposal Network) is used to generate anchors of different scales and the corresponding anchor confidence scores. If the content of multiple video frames does not change substantially over a period of time, it is approximately equivalent to the model input being multiple detection images with the same content and the same scaling.
[0058] S102, taking the initial detection frame whose confidence score is greater than the confidence threshold as the detection frame;
[0059] Based on S101, this step aims to use the initial detection frame with a confidence score greater than the confidence threshold as the detection frame. That is, the detection frames are screened and the detection frames that cannot meet the composite requirements are eliminated. Further, this step can be to sort the anchors from high to low according to the confidence score, and screen out the anchors with confidence scores higher than the threshold and further merge them, that is, to execute the subsequent operation steps.
[0060] S103, determining an intersection-over-union ratio threshold based on the number of the multiple detection frames;
[0061] Based on S102, this step aims to determine the intersection-over-union threshold based on the number of multiple detection frames. That is, the intersection-over-union threshold in this embodiment, that is, the nms threshold, is not fixed and can be dynamically adjusted based on the number of detection frames to improve the detection accuracy.
[0062] Furthermore, this step may include:
[0063] Step 1, determine whether the number of multiple detection boxes is greater than the number threshold;
[0064] Step 2: If yes, set the first threshold as the intersection-over-union ratio threshold;
[0065] Step 3: If not, set the second threshold as the intersection-and-union ratio threshold; wherein the first threshold is greater than the second threshold.
[0066] It can be seen that this step mainly explains how to determine the intersection-and-union ratio threshold based on the number of detection frames. A dynamically adjusted intersection-and-union ratio threshold (nms threshold) is implemented. The nms threshold of the existing nms algorithm usually adopts a fixed constant. If the threshold is too low, when the detection objects in the image are dense, missed detection is likely to occur. If the threshold is too high, when the detection objects are sparse, it is easy to cause redundant detection frames. Therefore, when setting the threshold, the present invention will dynamically select different nms thresholds according to the number of detection frames, thereby improving the detection accuracy.
[0067] S104, performing non-maximum suppression processing on multiple detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame.
[0068] On the basis of S103, this step aims to perform non-maximum suppression processing on multiple detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame.
[0069] Among them, based on whether the facial coordinate information corresponding to each detection frame is relatively independent, it is judged whether the relationship between the adjacent frame and the frame with the highest confidence score is redundant or an independent detection frame. If the frame with the highest confidence score is compared with its adjacent frame, if the facial features and contours depicted by the facial coordinate information of the two are relatively independent, then even if the IOU of the two detection frames is relatively high, the adjacent frame will be eliminated from the subsequent merging frame process. If the facial coordinate information corresponding to the two adjacent frames has a high degree of overlap, the weight of the adjacent frame is reduced, and it is sent to the subsequent process to participate in the detection frame merging. In other words, based on the facial coordinate information, independent or redundant detection frames in the detection frame can be effectively eliminated, thereby improving the accuracy of detection frame processing.
[0070] The facial coordinate information may be coordinate point information of facial features, or may be more detailed coordinate point information.
[0071] Furthermore, this step may include:
[0072] Step 1: Calculate the intersection-and-union ratio of the detection frame with the highest confidence score with other detection frames to obtain the intersection-and-union ratio values corresponding to multiple detection frames;
[0073] Step 2: divide the multiple detection frames based on the IoU values corresponding to the multiple detection frames to obtain the detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold;
[0074] Step 3, based on the detection frame with the highest confidence score and the face coordinate information of multiple adjacent detection frames, the independent detection frames in the multiple adjacent detection frames are eliminated to obtain multiple adjacent detection frames after elimination;
[0075] Step 4: Based on the weight of each detection frame, non-maximum suppression is performed on the detection frame with the highest confidence score and multiple adjacent detection frames to obtain the target detection frame.
[0076] It can be seen that this optional solution mainly explains how to perform maximum value suppression processing. Among them, it mainly uses the face coordinate information as auxiliary information to determine whether two adjacent detection frames correspond to the same detection object, and determines whether the adjacent frames circle the same target detection object based on whether the face coordinate information is relatively independent.
[0077] Furthermore, step 2 in the previous optional solution may include:
[0078] Step 201, the detection frames except the detection frame with the highest confidence score whose IoU values are greater than the IoU threshold are regarded as adjacent detection frames, and the remaining detection frames are regarded as remaining detection frames;
[0079] Step 202: If the number of adjacent detection boxes is 1, send an execution command.
[0080] Step 203, if the number of adjacent detection frames is greater than 1, and the intersection-and-union ratio value of each adjacent detection frame is 1, delete adjacent detection frames until the number is 1;
[0081] Step 204: when the number of adjacent detection frames is 1 and the number of remaining detection frames is 0, copy the adjacent detection frames as remaining detection frames.
[0082] It can be seen that this optional solution mainly explains how to divide the detection frames. Through this optional solution, multiple detection frames can be divided into the detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold. In addition, the number of adjacent detection frames and remaining detection frames is kept at least 1.
[0083] Furthermore, step 3 in the previous optional solution may include:
[0084] Step 301, determining the difference between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of multiple adjacent detection frames;
[0085] Step 302: remove adjacent detection frames whose difference is greater than a difference threshold to obtain remaining adjacent detection frames.
[0086] It can be seen that this optional solution mainly explains how to remove the detection frame based on the face coordinate information, avoid processing independent detection frames, reduce the impurities in the detection frame processing, and improve the accuracy of target detection.
[0087] Furthermore, it may also include:
[0088] Step 303, determining the degree of overlap between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of multiple adjacent detection frames;
[0089] Step 304 , down-weighting the weights of adjacent detection frames whose overlap is greater than the overlap threshold to obtain a new weight for the detection frame.
[0090] Based on the previous optional solution, this optional solution mainly explains how to reduce the weights of detection frames that are highly overlapped.
[0091] Further, step 4 in the previous optional solution may include:
[0092] Step 401, weighting each adjacent detection frame based on the intersection-over-union ratio value, uncertainty and variance between each adjacent detection frame to obtain the weight of each adjacent detection frame;
[0093] Step 402: Perform non-maximum suppression processing based on the weights of each adjacent detection frame to obtain a target detection frame.
[0094] In summary, this embodiment dynamically adjusts the IoU threshold by the number of initial detection frames, and then performs non-maximum suppression processing based on the dynamically adjusted IoU threshold and face coordinate information to exclude repeated and redundant detection frames, avoid detection deviations caused by the same detection frame when images of the same zoom scale are repeatedly detected multiple times, and improve the accuracy of target detection.
[0095] The following further illustrates a detection frame processing method for target detection provided by the present application through another specific embodiment.
[0096] In this embodiment, the nms threshold is first adjusted dynamically, and the threshold of the number of anchors is preset. When the number of anchors detected in the image exceeds the threshold, it means that the image targets are dense, and a high nms threshold is selected to reduce the number of adjacent anchors that need to be filtered out and reduce missed detections. For example, when the number of detected anchors is lower than the threshold, a low nms threshold is selected to filter out as many duplicate detection frames as possible, and the detection frame with the highest confidence is retained at the same position.
[0097] Then, based on whether the landmarks (coordinate information) of the facial features corresponding to each target detection frame are relatively independent, it is determined whether the adjacent frame and the frame with the highest confidence score are redundant or independent detection frames. If the frame with the highest confidence score is compared with its adjacent frame, if the features and contours depicted by the landmarks of the two are relatively independent, then the adjacent frame is considered an independent frame, and the detection object is not the same as that of the frame with the highest confidence score. If the landmark coordinates corresponding to the two adjacent frames have a high degree of overlap, then the adjacent frame is considered a redundant frame, and the weight of the adjacent frame is reduced, and it is sent to the subsequent process to participate in the detection frame merging.
[0098] It can be seen that the problem of inconsistent detection results obtained by multiple detections and single detections of the same image after vote-nms filtering is solved. When merging detection frames, first make a judgment. If there are multiple adjacent anchors, and the iou of the adjacent anchor and the selected anchor with the highest confidence score is 1, that is, the coordinates of the two completely overlap, then the redundant completely overlapping adjacent anchors are deleted, and the number of adjacent anchors is reset to one, and they will no longer participate in the subsequent detection frame merging.
[0099] According to the above claims, the method mainly includes the following steps:
[0100] Step 1: Cut the original face video stream frame by frame and input it into the target detection model in sequence. Use the RPN network to generate anchors of different scales and the corresponding anchor confidence scores. If the content of multiple video frames has basically not changed over a period of time, it is approximately equivalent to the model input being multiple detection images with the same content and the same scale.
[0101] Step 2: Sort the anchors according to the confidence scores from high to low, and select the anchors with confidence scores higher than the threshold for further merging.
[0102] Step 3: Get the number of all anchors and preset the anchor threshold. If the total number of anchors exceeds the anchor threshold, it means that the target is densely distributed, and the nms threshold is increased to a high threshold A. On the contrary, when the target is sparsely distributed, the nms threshold is lowered to a low threshold B.
[0103] Step 4: Find the anchor with the highest confidence score, calculate the iou (intersection-over-union) of the anchor and other anchors respectively, filter out the adjacent anchors whose iou exceeds the preset nms threshold, and record the idx (index) of these anchors.
[0104] Step 5: Filter out the adjacent anchors with relatively high IOU in step 4 and the remaining anchors and divide them into 2 groups.
[0105] Step 6: If there is only one adjacent anchor, the adjacent anchors and the remaining anchors are grouped and directly perform the subsequent voting weight allocation step.
[0106] Step 7: If there are multiple adjacent anchors, and the iou of the adjacent anchor and the selected anchor with the highest confidence score is 1, that is, the coordinates of the two completely overlap, then delete the redundant adjacent anchors that completely overlap, and reset the number of adjacent anchors to one.
[0107] Step 8: Under the conditions of step 6 or 7, if the number of anchors whose iou is lower than the threshold is zero, copy the adjacent anchors as the remaining anchors and participate in the subsequent weight allocation step together.
[0108] Step 9: Remove independent frames based on facial landmark information, and compare the landmarks corresponding to the anchor with the highest confidence score and all its adjacent anchors to see if they are relatively independent. If the landmarks of the frame with the highest confidence score and the adjacent anchors include but are not limited to: the facial features of the two faces are far apart, the contours of the two faces are very different (one face is the foreground and the other face is the background), and the angles of the two faces are very different (one is the front face and the other is the side face), the frame with the highest confidence score and the adjacent anchors are separated, and the adjacent anchors are extracted as the detection frame for the next round of overall iteration, and no longer participate in the vote-nms detection frame merging.
[0109] In step 10, if the landmarks corresponding to the two detection frames have a high overlap, which is the same as the situation in the previous step, then the adjacent anchor is likely to be a redundant frame. A lower weight is assigned to the anchor and it is sent to the next step to participate in the detection frame merging calculation.
[0110] Step 11: The box with the highest confidence score and all the adjacent anchors retained in the previous step participate in vote-nms. Assign higher weights to anchors that are closer (higher IOU) and have lower uncertainty. Calculate the variance of each adjacent anchor, assign lower weights to adjacent anchors with high variance and anchors with smaller IOU with the selected anchor, and use them as the final detection set.
[0111] Step 12: Based on the detection set in the previous step, all coordinates are weighted averaged to calculate new coordinates as the first final target box, and the anchors and their confidence scores of these detection sets are cleared.
[0112] Step 13: After one cycle, find the anchor with the highest confidence score in the remaining detection frames, repeat steps 3-12 to form the target frame of the next face, until there are almost no overlapping frames, completing the process of merging all anchors.
[0113] It can be seen that this embodiment dynamically adjusts the IoU threshold by the number of initial detection frames, and then performs non-maximum suppression processing based on the dynamically adjusted IoU threshold and face coordinate information to exclude repeated and redundant detection frames, avoid detection deviations caused by the same detection frame when images of the same zoom scale are repeatedly detected multiple times, and improve the accuracy of target detection.
[0114] The following is an introduction to the detection frame processing device for target detection provided in an embodiment of the present application. The detection frame processing device for target detection described below and the detection frame processing method for target detection described above can be referenced to each other.
[0115] Please refer to Figure 2 , Figure 2 A schematic diagram of the structure of a detection frame processing device for target detection provided in an embodiment of the present application.
[0116] In this embodiment, the device may include:
[0117] The video stream detection module 100 is used to input the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and the confidence score corresponding to each initial detection frame;
[0118] A detection frame filtering module 200, configured to use an initial detection frame having a confidence score greater than a confidence threshold as a detection frame;
[0119] An IoU calculation module 300, configured to determine an IoU threshold based on the number of detection frames;
[0120] The non-maximum suppression processing module 400 is used to perform non-maximum suppression processing on multiple detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame.
[0121] The present application also provides a server, including:
[0122] Memory for storing computer programs;
[0123] A processor is used to implement the steps of the detection frame processing method as described in the above embodiment when executing the computer program.
[0124] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the detection frame processing method described in the above embodiment are implemented.
[0125] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0126] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0127] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0128] The above is a detailed introduction to a detection frame processing method, a detection frame processing device, a server and a computer-readable storage medium for target detection provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A detection frame processing method for target detection, characterized in that: include: Inputting the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and a confidence score corresponding to each of the initial detection frames; Taking the initial detection frame whose confidence score is greater than the confidence threshold as the detection frame; Determine an intersection-over-union ratio threshold based on the number of the plurality of detection frames; Based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame, a non-maximum suppression process is performed on the multiple detection frames to obtain a target detection frame; Among them, the non-maximum suppression processing is performed on the multiple detection frames based on the IoU threshold and the face coordinate information of each detection frame to obtain the target detection frame, including: performing IoU calculation on the detection frame with the highest confidence score and other detection frames to obtain IoU values corresponding to the multiple detection frames; dividing the multiple detection frames based on the IoU values corresponding to the multiple detection frames to obtain the detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold; based on the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames, the independent detection frames in the multiple adjacent detection frames are eliminated to obtain the eliminated multiple adjacent detection frames; based on the weight of each detection frame, the detection frame with the highest confidence score and the multiple adjacent detection frames are non-maximum suppression processed to obtain the target detection frame; The method divides the multiple detection frames based on the IoU values corresponding to the multiple detection frames to obtain the detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold, including: taking the detection frames other than the detection frame with the highest confidence score whose IoU values are greater than the IoU threshold as the adjacent detection frames, and taking the remaining detection frames as the remaining detection frames; if the number of the adjacent detection frames is 1, sending an execution command; if the number of the adjacent detection frames is greater than 1, and the IoU value of each of the adjacent detection frames is 1, deleting the adjacent detection frames until the number is 1; when the number of the adjacent detection frames is 1, and the number of the remaining detection frames is 0, copying the adjacent detection frames as the remaining detection frames; the execution command is a command for grouping the adjacent detection frames and the remaining detection frames to directly execute the subsequent voting weight allocation step; Based on the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames, the independent detection frames in the multiple adjacent detection frames are eliminated to obtain the multiple adjacent detection frames after elimination, including: determining the difference between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames; eliminating the adjacent detection frames whose difference is greater than the difference threshold to obtain the remaining adjacent detection frames; The difference thresholds include a distance threshold between the facial features of two faces, a size difference threshold between the contours of two faces, and a difference threshold between the angles of two faces. The detection frame processing method also includes: determining the degree of overlap between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames; and downgrading the weights of the adjacent detection frames whose overlap is greater than the overlap threshold to obtain a new weight for the detection frame.
2. The detection frame processing method according to claim 1, characterized in that: Determining an intersection-over-union ratio threshold based on the number of the plurality of detection frames includes: Determine whether the number of the plurality of detection frames is greater than a number threshold; If yes, then setting the first threshold to the intersection-over-union ratio threshold; If not, the second threshold is set to the intersection-over-union ratio threshold; wherein the first threshold is greater than the second threshold.
3. A detection frame processing device for target detection, characterized in that: include: A video stream detection module is used to input the original face video stream into the target detection model in units of frames to obtain multiple initial detection frames and a confidence score corresponding to each of the initial detection frames; A detection frame filtering module, used to take the initial detection frame whose confidence score is greater than the confidence threshold as the detection frame; An IoU calculation module, used to determine an IoU threshold based on the number of the plurality of detection frames; A non-maximum suppression processing module, used for performing non-maximum suppression processing on the plurality of detection frames based on the intersection-over-union ratio threshold and the face coordinate information of each detection frame to obtain a target detection frame; Among them, the non-maximum suppression processing module is specifically used to perform IoU calculation on the detection frame with the highest confidence score and other detection frames to obtain IoU values corresponding to multiple detection frames; divide the multiple detection frames based on the IoU values corresponding to the multiple detection frames to obtain the detection frame with the highest confidence score, multiple adjacent detection frames whose IoU values are greater than the IoU threshold, and multiple remaining detection frames whose IoU values are less than or equal to the IoU threshold; based on the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames, remove the independent detection frames in the multiple adjacent detection frames to obtain multiple adjacent detection frames after removal; perform non-maximum suppression processing on the detection frame with the highest confidence score and the multiple adjacent detection frames based on the weight of each detection frame to obtain the target detection frame; The non-maximum suppression processing module is specifically used to use the detection frames other than the detection frame with the highest confidence score whose intersection-over-union ratio values are greater than the intersection-over-union ratio threshold as the adjacent detection frames, and use the remaining detection frames as the remaining detection frames; if the number of the adjacent detection frames is 1, send an execution command; if the number of the adjacent detection frames is greater than 1, and the intersection-over-union ratio value of each of the adjacent detection frames is 1, delete the adjacent detection frames until the number is 1; when the number of the adjacent detection frames is 1, and the number of the remaining detection frames is 0, copy the adjacent detection frames as the remaining detection frames; the execution command is a command for grouping the adjacent detection frames and the remaining detection frames to directly execute the subsequent voting weight allocation step; The non-maximum suppression processing module is specifically used to determine the difference between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames; remove the adjacent detection frames whose difference is greater than the difference threshold to obtain the remaining adjacent detection frames; The difference thresholds include a distance threshold between the facial features of two faces, a size difference threshold between the contours of two faces, and a difference threshold between the angles of two faces. The detection frame processing device is also used to determine the degree of overlap between the face coordinate information of the detection frame with the highest confidence score and the face coordinate information of the multiple adjacent detection frames; and to downgrade the weights of the adjacent detection frames whose overlap is greater than the overlap threshold to obtain a new weight for the detection frame.
4. A server, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the detection frame processing method as described in any one of claims 1 to 2 when executing the computer program.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the detection frame processing method according to any one of claims 1 to 2 are implemented.
Citation Information
Patent Citations
Moving target object detection method and device and storage medium
CN112347810A
Target detection method and device, computer equipment and computer readable storage medium
CN112749590A