Target detection result post-processing method and device, electronic equipment, and storage medium
Patent Information
- Application Number
- CN202110432170.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-21
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-04-21
AI Technical Summary
但是,此方法对检测框的尺寸没有约束,仍然会保留不少尺寸明显不当的误检框
[0024]与现有技术相比,本发明实施例的技术方案具有有益效果。
Smart Images

Figure CN115222952B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a post-processing method, apparatus, electronic device, and storage medium for target detection results. Background Technology
[0002] Currently, the most widely used post-processing method in object detection is Non-maximum Suppression (NMS). NMS can be used to eliminate redundant detection boxes by ranking the confidence scores of the boxes and calculating the Intersection Over Union (IoU) between two boxes of the same category, thus removing those with low confidence scores and high overlap with other boxes of the same category. However, this method does not impose constraints on the size of the detection boxes, and still retains many false positives with obviously inappropriate sizes. Summary of the Invention
[0003] The technical problem solved by the embodiments of the present invention is to eliminate false detection boxes in the target detection results.
[0004] To address the aforementioned technical problems, this invention provides a post-processing method for target detection results, comprising: acquiring target detection results of an image to be processed, the target detection results including detection boxes of target objects in the image to be processed; extracting the maximum connected region of the target object within the detection box; acquiring data information of the maximum connected region; determining whether the data information is within the valid data range, if so, retaining the detection box, otherwise eliminating the detection box; wherein, the target detection results are obtained based on target detection of the image to be processed using a preset detection network model, and the valid data range is determined based on the data information of the maximum connected region corresponding to the target object in the validation set of the detection network model.
[0005] Optionally, the target detection result includes the category corresponding to the target object, and the detection box includes the detection box corresponding to the category that needs to be post-processed.
[0006] Optionally, the post-processing method includes: obtaining the false detection rate of the category; determining whether the false detection rate is greater than or equal to the false detection threshold; if so, determining that the category needs to be post-processed.
[0007] Optionally, the post-processing method includes: selecting test images that do not include the category corresponding to the target object to form a test set; using a detection network model to perform target detection on each test image in the test set to obtain test detection results corresponding to each test image; obtaining test false detection boxes that include the category corresponding to the target object in the test detection results; and determining the false detection rate as the ratio of the number of test images containing the test false detection boxes to the total number of all test images in the test set.
[0008] Optionally, the false detection threshold includes 0.5%.
[0009] Optionally, extracting the maximum connected region of the target object within the detection box includes: obtaining a first mask based on the HSV color mode and a second mask based on the grayscale mode for each pixel in the detection box; performing logical operations on the first and second masks to obtain the final mask for each pixel; and determining the maximum connected region based on the final mask; wherein the values of the first mask, the second mask, and the final mask each include either a first value or a second value.
[0010] Optionally, the largest connected region is determined based on the region where the pixel with the final mask has the second value.
[0011] Optionally, the logical operation includes logical AND and NOT operations; performing logical operations on the first mask and the second mask to obtain the final mask of each pixel includes: comparing the first mask and the second mask; setting the final mask of the corresponding pixel to a second value when both the first mask and the second mask are first values, and setting the final mask of the corresponding pixel to a first value when at least one of the first mask and the second mask is a second value.
[0012] Optionally, the post-processing method includes: obtaining the color parameter range of the target object in HSV color mode and the color parameter of each pixel in HSV color mode; comparing the color parameter of each pixel with the color parameter range respectively; setting the first mask of pixels whose color parameters are within the color parameter range to a first value, and setting the first mask of pixels whose color parameters are outside the color parameter range to a second value.
[0013] Optionally, the post-processing method includes: obtaining the average gray value of each pixel in grayscale mode; comparing the gray value of each pixel with the average gray value; taking a first value for the second mask of pixels whose gray value is greater than or equal to the average gray value, and taking a second value for the second mask of pixels whose gray value is less than the average gray value.
[0014] Optionally, the post-processing method includes: selecting verification images containing the target object to form a verification set, wherein the verification images include the bounding boxes of the target object; using a detection network model to perform target detection on each verification image in the verification set to obtain verification detection results corresponding to each verification image, wherein the verification detection results include the verification boxes of the target object; extracting the maximum connected regions of the target object in the bounding boxes and verification boxes respectively; obtaining the annotation data information and verification data information of the maximum connected regions corresponding to the bounding boxes and verification boxes respectively; obtaining the union of the annotation data information and the verification data information, and using the union as the effective data range.
[0015] Optionally, the data information includes size data, which includes at least two of the following: the area of the largest connected region and the length of its longest side, the length of its shortest side, and the aspect ratio of its smallest bounding rectangle.
[0016] Optionally, the target detection result includes the shape information of the target object, the data information includes the proportion data, and the effective data range includes the effective proportion range corresponding to the proportion data; determining whether the data information is within the effective data range includes: identifying the shape category corresponding to the shape information, and determining whether the proportion data is within the effective proportion range based on the shape category.
[0017] Optionally, the shape category includes ring shape, the ratio data includes the effective pixel ratio of the largest connected region, and the effective ratio range includes the effective pixel ratio range.
[0018] Optionally, the shape category includes rectangle or approximate rectangle, the proportion data includes the area percentage of the largest connected region, and the effective proportion range includes the effective area percentage range.
[0019] Optionally, the post-processing method includes: when the data information is determined to be within the valid data range, obtaining the sample probability corresponding to each size data, where the sample probability is the probability of the corresponding size data appearing in the validation set; determining whether the sample probability is less than or equal to a probability threshold, and if so, determining that the sample probability is a small sample probability and that the size data corresponding to the small sample probability is a low-probability size data; determining whether other size data besides the low-probability size data are within a first reference size range, and if so, retaining the detection box, and if not, eliminating the detection box; wherein, the first reference size range is the range in the validation set where other size data are below the probability threshold.
[0020] Optionally, the target detection result includes the confidence level of the detection box; the post-processing method includes: when it is determined that the data information is within the valid data range, obtaining the confidence level, determining whether the confidence level is less than or equal to the confidence level threshold, and for detection boxes with a confidence level less than or equal to the confidence level threshold, determining whether at least one size data corresponding to it is within a second reference size range. If so, the detection box is retained; if not, the detection box is eliminated. The second reference size range is the range of size data corresponding to the verification boxes with a confidence level less than or equal to the confidence level threshold in the verification set.
[0021] This invention also provides a post-processing device for target detection results, comprising: a first acquisition module for acquiring target detection results of an image to be processed, the target detection results including detection boxes of target objects in the image to be processed; an extraction module for extracting the maximum connected region of the target object within the detection box; a second acquisition module for acquiring data information of the maximum connected region; a judgment module for judging whether the data information is within the valid data range; and a processing module for determining whether to retain the detection box when the data information is within the valid data range and to eliminate the detection box when the data information is outside the valid data range; wherein the target detection result is obtained based on target detection of the image to be processed using a preset detection network model, and the valid data range is determined based on the data information of the maximum connected region corresponding to the target object in the validation set of the detection network model.
[0022] This invention also provides an electronic device, including: a processor; a memory storing a computer program that can run on the processor; wherein, when the computer program is executed by the processor, it implements the post-processing method for target detection results provided in this invention.
[0023] This invention also provides a computer-readable storage medium storing a computer program, which, when executed, implements the post-processing method for target detection results provided in this invention.
[0024] Compared with the prior art, the technical solutions of the embodiments of the present invention have beneficial effects.
[0025] For example, determining whether to eliminate a detection box based on whether the relevant data information of the largest connected region of the target object in the detection box is within the valid data range can not only effectively eliminate false detection boxes, but also avoid sacrificing recall.
[0026] For example, the data information may include the size of the largest connected region, thereby eliminating false detection boxes with obviously inappropriate sizes based on the size of the largest connected region.
[0027] For example, for low-probability size data, other size data besides the low-probability size data can be added for secondary judgment to improve the accuracy of false detection box processing.
[0028] For example, for detection boxes with low confidence, at least one of their corresponding size data can be re-evaluated to improve the accuracy of false detection box processing.
[0029] For example, for some specially shaped target objects, the corresponding detection box can be eliminated based on its shape information to improve the accuracy of false detection box processing. Attached Figure Description
[0030] Figure 1 This is a flowchart of the post-processing method for target detection results in an embodiment of the present invention;
[0031] Figure 2 This is a schematic diagram of the target detection results of the image to be processed in an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram of the post-processing device for target detection results in an embodiment of the present invention. Detailed Implementation
[0033] In existing technologies, false detection boxes in target detection results cannot be effectively eliminated.
[0034] To address the aforementioned technical problems, this invention provides a post-processing method for target detection results, comprising: acquiring target detection results of an image to be processed, the target detection results including detection boxes of target objects in the image to be processed; extracting the maximum connected region of the target object within the detection box; acquiring data information of the maximum connected region; determining whether the data information is within the valid data range, if so, retaining the detection box, otherwise eliminating the detection box; wherein, the target detection results are obtained based on target detection of the image to be processed using a preset detection network model, and the valid data range is determined based on the data information of the maximum connected region corresponding to the target object in the validation set of the detection network model.
[0035] Compared with existing technologies, the technical solutions of the embodiments of the present invention have beneficial effects. For example, determining whether to eliminate a detection box based on whether the relevant data information of the largest connected region of the target object in the detection box is within the valid data range can not only effectively eliminate false detection boxes, but also avoid loss of recall rate.
[0036] To make the objectives, features, and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It is to be understood that the specific embodiments described below are for illustrative purposes only and are not intended to limit the present invention. Furthermore, for ease of description, only the parts relevant to the present invention are shown in the accompanying drawings, not all of the structures.
[0037] Figure 1 This is a flowchart of the post-processing method for target detection results in an embodiment of the present invention.
[0038] Reference Figure 1 The present invention provides a post-processing method for target detection results, comprising:
[0039] S1, Obtain the target detection result of the image to be processed, which includes the detection box of the target object in the image to be processed;
[0040] S2, extract the maximum connected region of the target object within the detection box;
[0041] S3, obtain data information of the largest connected region;
[0042] S4. Determine whether the data information is within the valid data range. If yes, retain the detection box; otherwise, remove the detection box.
[0043] In this embodiment of the invention, the target detection result is obtained by performing target detection on the image to be processed using a preset detection network model.
[0044] In practice, the pre-defined detection network model can be implemented using any known conventional techniques in the field.
[0045] In practice, the target detection results include the category of the target object and its corresponding detection box.
[0046] Figure 2 This is a schematic diagram of the target detection results of the image to be processed in an embodiment of the present invention.
[0047] Reference Figure 2 The image to be processed 100 has three target detection results, namely the first target detection result 101, the second target detection result 102 and the third target detection result 103.
[0048] Specifically, the first target detection result 101 includes metal (i.e., the category of the target object) and its corresponding detection box. The second target detection result 102 includes glass (i.e., the category of the target object) and its corresponding detection box. The third target detection result 103 includes an umbrella (i.e., the category of the target object) and its corresponding detection box.
[0049] In some embodiments, it can be determined whether the corresponding detection box of the target object needs to be post-processed based on the category of the target object. That is, the detection boxes that need to be post-processed can be filtered out by determining whether the category of the target object is a category that needs to be post-processed.
[0050] For example, by determining whether metal, glass, and umbrellas are categories that require post-processing, one can filter out which categories(s) of metal, glass, and umbrellas correspond to detection boxes that require post-processing.
[0051] Specifically, the false detection rate of the target object's corresponding category can be obtained, and it can be determined whether the false detection rate is greater than or equal to the false detection threshold. If so, it is determined that the category of the target object needs to be post-processed.
[0052] In some embodiments, the post-processing method for the target detection result may further include:
[0053] S11, obtain the false detection rate of the target object's corresponding category;
[0054] S12, determine whether the false detection rate is greater than or equal to the false detection threshold. If yes, determine that the category corresponding to the target object needs to be post-processed; otherwise, determine that the category corresponding to the target object does not need to be post-processed.
[0055] In some embodiments, the false detection rate of the target object's corresponding category in step S11 may include:
[0056] S111, Select test images that do not include the category corresponding to the target object to form a test set;
[0057] S112, using a preset detection network model to perform target detection on each test image in the test set to obtain test detection results corresponding to each test image;
[0058] S113, Obtain test false detection boxes that include the category corresponding to the target object in the test detection results;
[0059] S114, the ratio of the number of test images containing false detection boxes to the total number of all test images in the test set is determined as the false detection rate of the target object's corresponding category.
[0060] Typically, when using a detection network model to perform object detection on an image, a certain probability of false detections will occur. For example, when performing object detection on an image that does not include the category corresponding to the target object (i.e., the test image), detection results may include the category corresponding to the target object. Specifically, this detection result includes the category corresponding to the target object and its test false detection bounding box.
[0061] For example, when performing object detection on a test set consisting of test images that do not include umbrellas, some test images may show false detection boxes that include umbrellas.
[0062] In practice, the number of test images containing false detection boxes corresponding to the target object category can be counted, and the ratio of this number to the total number of test images in the test set can be used as the false detection rate of the target object category.
[0063] For example, the number of test images containing false positive boxes for umbrellas can be counted, and the ratio of this number to the total number of test images in the test set can be used as the false positive rate for umbrellas.
[0064] When the false detection rate of the target object's corresponding category is greater than or equal to the false detection threshold, the target object's corresponding category can be determined as a category that needs post-processing, and the detection boxes corresponding to this category need to be post-processed.
[0065] For example, when the false detection rate of an umbrella is greater than or equal to the false detection threshold, it can be determined that the detection box corresponding to the umbrella needs to be post-processed.
[0066] In some embodiments, the false detection threshold includes 5‰.
[0067] In some embodiments, step S2, extracting the maximum connected region of the target object within the detection frame, may include:
[0068] S21, obtain the first mask based on the HSV color mode and the second mask based on the grayscale mode for each pixel in the detection box (i.e., the detection box corresponding to the category of the target object. For example, the detection box of an umbrella).
[0069] S22, Perform logical operations on the first mask and the second mask to obtain the final mask for each pixel;
[0070] S23, determine the maximum connected region based on the final mask.
[0071] In practice, the values of the first mask, the second mask, and the final mask include either the first value or the second value.
[0072] Specifically, the first mask can take either a first value or a second value. The second mask can also take either a first value or a second value. The final mask can also take either a first value or a second value.
[0073] In some embodiments, the first value may include 1, and the second value may include 0.
[0074] In other embodiments, the first value may include 255, and the second value may include 0.
[0075] In some embodiments, step S21, obtaining the first mask for each pixel in the detection frame based on the HSV color mode, may include:
[0076] S211, Obtain the color parameter range of the target object in HSV color mode, and the color parameters of each pixel in the detection box in HSV color mode;
[0077] S212, compare the color parameter of each pixel with the color parameter range respectively;
[0078] S213, set the first mask of pixels whose color parameters are within the range of color parameters to a first value, and set the first mask of pixels whose color parameters are outside the range of color parameters to a second value.
[0079] In some embodiments, the image to be processed may be in Red, Green, Blue (RGB) color mode. In this case, the image to be processed can first be converted from RGB color mode to Hue, Saturation, Value (HSV) color mode.
[0080] In practice, the conversion of the image to be processed from RGB color mode to HSV color mode is achieved using conventional techniques in this field.
[0081] In practice, obtaining the color parameter range of the target object in HSV color mode, as well as the color parameters of each pixel in the detection frame in HSV color mode, can be achieved using conventional techniques in this field.
[0082] In a specific implementation, the color parameter of each pixel is compared with the color parameter range; for pixels whose color parameters are within the color parameter range, their first mask is set to a first value (e.g., 255); for pixels whose color parameters are outside the color parameter range, their first mask is set to a second value (e.g., 0).
[0083] In some embodiments, step S21, obtaining the second mask based on the grayscale mode for each pixel in the detection frame, may include:
[0084] S214, Obtain the average grayscale value of each pixel in the detection box in grayscale mode;
[0085] S215, compare the gray value of each pixel with the average gray value;
[0086] S216, the second mask of the pixels whose gray value is greater than or equal to the average gray value is taken as the first value, and the second mask of the pixels whose gray value is less than the average gray value is taken as the second value.
[0087] In practice, the image to be processed is first converted to grayscale mode, and the grayscale value of each pixel in the detection frame is obtained in grayscale mode. The average grayscale value is obtained by summing the grayscale values of each pixel and dividing by the total number of pixels in the detection frame. Then, the grayscale value of each pixel is compared with the average grayscale value. For pixels with a grayscale value greater than or equal to the average grayscale value, the second mask is set to the first value (e.g., 255). For pixels with a grayscale value less than the average grayscale value, the second mask is set to the second value (e.g., 0).
[0088] In practice, the first value of the first mask and the first value of the second mask are the same, and the second value of the first mask and the second value of the second mask are also the same. For example, the first value of both the first mask and the first value of both the first mask are 255, and the second value of both the first mask and the second value of both the second mask are 0.
[0089] In some embodiments, the logical operation in step S22 may include a logical AND-NOT operation.
[0090] Accordingly, step S22, which involves performing logical operations on the first mask and the second mask to obtain the final mask for each pixel, may include:
[0091] S221, compare the first mask and the second mask;
[0092] S222, when both the first mask and the second mask are first values, the final mask of the corresponding pixel is set to the second value; when at least one of the first mask and the second mask is the second value, the final mask of the corresponding pixel is set to the first value.
[0093] In a specific implementation, the first mask and the second mask of each pixel are compared; when both the first mask and the second mask are first values (e.g., 255), the final mask of the pixel is set to a second value (e.g., 0); when at least one of the first mask and the second mask is a second value (e.g., 0), the final mask of the pixel is set to a first value (e.g., 255).
[0094] Accordingly, step S23, which involves determining the maximum connected region based on the final mask, may include:
[0095] S231, determine the maximum connected region of the target object in the detection box based on the region where the pixel with the final mask is the second value (e.g., 0).
[0096] Specifically, the region where the pixel with the final mask is the second value (e.g., 0) is located is determined as the region where the target object is located, and the maximum connected region of the target object is the maximum connected region of the target object in the detection box.
[0097] In practice, obtaining the maximum connected region of the target object's location is achieved using conventional techniques in this field.
[0098] In some embodiments, the data information in step S3 may include size data.
[0099] Specifically, the size data may include the area of the largest connected region of the target object in the detection frame and at least two of the following: the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region.
[0100] In practice, the area of the largest connected region and the length of its longest side, the length of its shortest side, and the aspect ratio of its smallest bounding rectangle are calculated using conventional techniques in this field.
[0101] Accordingly, step S4, determining whether the data information is within the valid data range, includes:
[0102] S41, determine whether all the size data are within the corresponding valid data range. If yes, retain the detection box; otherwise, eliminate the detection box.
[0103] In practice, the effective data range is determined based on the data information of the largest connected region corresponding to the target object in the validation set of the detection network model.
[0104] Specifically, the maximum connected region in the validation set includes the maximum connected region of the target object in its bounding box and the maximum connected region of the target object in its validation box.
[0105] In some embodiments, the post-processing method for the target detection result may further include:
[0106] S51, Select verification images including the target object to form a verification set;
[0107] S52, using a detection network model to perform target detection on each verification image in the verification set to obtain verification detection results corresponding to each verification image, the verification detection results including the verification bounding box of the target object;
[0108] S53, extract the maximum connected region of the target object in the annotation box and the verification box respectively;
[0109] S54, respectively obtain the annotation data information and verification data information of the largest connected region corresponding to the annotation box and the verification box;
[0110] S55, obtain the union of labeled data information and verification data information, and use this union as the valid data range.
[0111] In practice, the maximum connected regions of the target object in the bounding box and the verification box are extracted separately. This includes extracting the maximum connected regions of the bounding box and the verification box corresponding to the target object for each verification image in the verification set.
[0112] For example, for a verification set consisting of verification images including umbrellas, the bounding boxes and verification boxes for umbrellas in each verification image in the verification set should be obtained, and the maximum connected region of the umbrella in the bounding boxes and verification boxes should be extracted respectively.
[0113] In practice, extracting the maximum connected region of the target object in the annotation box and the verification box respectively includes extracting the maximum connected region of the target object in the annotation box and extracting the maximum connected region of the target object in the verification box.
[0114] In specific implementation, extracting the maximum connected region of the target object in the annotation box and extracting the maximum connected region of the target object in the verification box can both be achieved using the technical means described in the embodiments of the present invention for extracting the maximum connected region of the target object in the detection box.
[0115] In some embodiments, the annotation data information of the largest connected region corresponding to the annotation box may include the size data of the annotation box.
[0116] Specifically, the size data of the annotation box may include the area of the largest connected region of the target object in the annotation box, and at least two of the following: the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region.
[0117] In some embodiments, the verification data information for the largest connected region corresponding to the verification box may include the size data of the verification box.
[0118] Specifically, the size data of the verification box may include the area of the largest connected region of the target object in the verification box, and at least two of the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region.
[0119] In practice, for any one of the following factors in the detection box size data, such as the area of the largest connected region of the target object in the detection box and the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region, the union of the size data of all the labeled boxes of the target object in the validation set and the size data of all the validation boxes shall be used as the corresponding valid data range.
[0120] For example, to determine the area of the largest connected region of the umbrella in the detection frame, we can obtain all the labeled boxes and all the verification boxes related to the umbrella in each verification image of the verification set. The sum of the areas of the largest connected regions corresponding to all the labeled boxes related to the umbrella is recorded as the first area information, and the sum of the areas of the largest connected regions corresponding to all the verification boxes related to the umbrella is recorded as the second area information. The union of the first area information and the second area information is taken as the effective area range.
[0121] Similarly, for any one of the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle that may be included in the size data of the detection frame, the above-mentioned technical means can also be used to obtain the corresponding effective data range, including the effective long side length range, the effective short side length range, and the effective aspect ratio range.
[0122] In the specific implementation of step S41, it is determined whether all the size data are within the corresponding valid data range.
[0123] Specifically, the size data described in S41 may include the area of the largest connected region of the target object in the detection frame and at least two of the long side length, short side length, and aspect ratio of the smallest bounding rectangle of the largest connected region.
[0124] For example, the size data may include the area of the largest connected region of the target object within the detection frame and the length of the long side of the smallest bounding rectangle of that largest connected region.
[0125] Accordingly, step S41, determining whether all the size data falls within the corresponding valid data range, may include:
[0126] S411, determine whether the area of the largest connected region of the target object in the detection frame is within the effective area range, and determine whether the length of the longest side of the smallest bounding rectangle of the largest connected region is within the effective length range.
[0127] For example, the size data may also include the area of the largest connected region of the target object in the detection frame, as well as all of the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region.
[0128] Accordingly, step S41, determining whether all the size data falls within the corresponding valid data range, may include:
[0129] S412, determine whether the area of the largest connected region of the target object in the detection frame is within the effective area range, and determine whether the long side length of the smallest bounding rectangle of the largest connected region is within the effective long side length range, and determine whether the short side length of the smallest bounding rectangle of the largest connected region is within the effective short side length range, and determine whether the aspect ratio of the smallest bounding rectangle of the largest connected region is within the effective aspect ratio range.
[0130] In some embodiments, for target objects with special shapes, it can also be determined whether the corresponding detection box should be eliminated based on their shape information.
[0131] Specifically, the target detection result in step S1 may further include the shape information of the target object. Correspondingly, the data information in step S3 may further include proportional data related to the shape information of the target object; the effective data range in step S4 may further include the effective proportional range corresponding to the proportional data.
[0132] Furthermore, step S4, determining whether the data information is within the valid data range, may also include:
[0133] S42, determine whether the ratio data is within the valid ratio range.
[0134] Specifically, step S42, determining whether the ratio data is within the valid ratio range, may include:
[0135] S421, Identify the shape category corresponding to the shape information.
[0136] S422, based on shape category, determine whether the proportion data is within the valid proportion range.
[0137] In practice, shape information includes the shape category corresponding to the target object.
[0138] In some embodiments, the shape category may include a ring. Accordingly, the scale data may include the effective pixel ratio of the largest connected region of the target object within the detection frame, and the effective scale range may include an effective pixel ratio range.
[0139] Specifically, the effective pixel ratio refers to the ratio of the number of pixels in the area where the target object is located within the detection box to the number of pixels in the area where the final mask is located.
[0140] In specific implementation, the effective pixel ratio range can be achieved using the technical means (steps S51 to S55) described in the embodiments of the present invention for obtaining the effective data range.
[0141] In other embodiments, the shape category may include a rectangle or an approximate rectangle. Accordingly, the scale data may include the area percentage of the largest connected region of the target object within the detection frame, and the effective scale range may include an effective area percentage range.
[0142] Specifically, the area ratio of the largest connected region of the target object within the detection frame refers to the ratio of the area of the largest connected region within the detection frame to the area of its smallest bounding rectangle.
[0143] In specific implementation, the effective area ratio range can be achieved using the technical means (steps S51 to S55) described in the embodiments of the present invention for obtaining the effective data range.
[0144] In some embodiments, the post-processing method for the target detection result may further include:
[0145] S61, when it is determined that the data information is within the valid data range, obtain the sample probability corresponding to each size data, and the sample probability is the probability of the corresponding size data appearing in the validation set;
[0146] S62, determine whether the sample probability is less than or equal to the probability threshold. If yes, determine the sample probability as a small sample probability and the size data corresponding to the small sample probability as a small probability size data and proceed to step S63. If no, retain the corresponding detection box.
[0147] S63, determine whether other size data besides low-probability size data are within the range of the first reference size. If yes, retain the detection box; otherwise, eliminate the detection box.
[0148] In the specific implementation of step S61, for any one of the area of the largest connected region that may be included in the size data and the length of the long side, the length of the short side, and the aspect ratio of its smallest bounding rectangle, there is a corresponding sample probability.
[0149] Specifically, the sample probability is the probability that data of the corresponding size will appear in the validation set.
[0150] For example, the size data includes the area of the largest connected region corresponding to the umbrella in the detection frame, specifically 100 square centimeters. Then, the sample probability of the area of the largest connected region corresponding to the umbrella is the probability that the area of the largest connected region corresponding to the umbrella in the validation set is 100 square centimeters.
[0151] In the specific implementation of step S62, the probability threshold can be used to determine whether the sample probability is a small sample probability, and the size data corresponding to the small sample probability can be determined as the small probability size data.
[0152] In some embodiments, the probability threshold can be 1%.
[0153] In practice, when the sample probability is less than or equal to the probability threshold, the sample probability can be determined as a small sample probability; when the sample probability is greater than the probability threshold, the sample probability cannot be determined as a small sample probability; when none of the sample probabilities can be determined as a small sample probability, the corresponding detection box can be directly retained; when at least one of the sample probabilities is determined as a small sample probability, proceed to step S63.
[0154] Specifically, the size data may include the area of the largest connected region of the target object within the detection box, and at least two of the following: the length of the long side, the length of the short side, and the aspect ratio of the smallest bounding rectangle of the largest connected region. For any one of these, a corresponding sample probability is obtained, and it is determined whether the corresponding sample probability is less than or equal to a corresponding probability threshold. Sample probabilities less than or equal to the probability threshold are defined as small sample probabilities, and the size data corresponding to the small sample probabilities are defined as small probability size data. Sample probabilities greater than the probability threshold cannot be defined as small sample probabilities. When all sample probabilities cannot be defined as small sample probabilities, the corresponding detection box can be directly retained.
[0155] For example, the size data may include the area of the largest connected region of the target object within the detection box and the length of the longest side of the smallest bounding rectangle of that largest connected region. When neither the sample probability corresponding to the area nor the sample probability corresponding to the longest side can be determined to be a small sample probability, the corresponding detection box can be directly retained. When at least one of the sample probability corresponding to the area and the sample probability corresponding to the longest side can be determined to be a small sample probability, proceed to step S63.
[0156] In the specific implementation of step S63, it is determined whether other size data besides low-probability size data are within the range of the first reference size. If so, the detection box is retained; otherwise, the detection box is eliminated.
[0157] For example, the size data includes the area of the largest connected region of the target object within the detection frame and the length of the longest side of the smallest bounding rectangle of that largest connected region. Only the area of the largest connected region of the target object within the detection frame is determined as a low-probability size data point. In this case, it is then determined whether the length of the longest side of the smallest bounding rectangle falls within the corresponding first reference size range.
[0158] In practice, the first reference size range is the range of other size data under the probability threshold in the verification set.
[0159] For example, the area of the largest connected region of the target object in the validation set is greater than 100 square centimeters, and the length of the longest side of the minimum bounding rectangle of the largest connected region of the target object in the validation set is greater than 15 centimeters. Meanwhile, for a small number of samples in the validation set (e.g., samples with a probability less than or equal to 1%), the area of the corresponding largest connected region is greater than 150 square centimeters, and the length of the longest side of the corresponding minimum bounding rectangle is greater than 25 centimeters. When the area of the largest connected region of the target object in the detection box (e.g., 160 square centimeters) is determined to be a low-probability size because it is greater than 150 square centimeters and falls within the range of a small number of samples, it can be determined whether the length of the longest side of the minimum bounding rectangle of the largest connected region of the target object in the detection box is within the first reference size range, i.e., whether the length of the longest side of the minimum bounding rectangle is greater than 25 centimeters. If yes, the corresponding detection box is retained; otherwise, the corresponding detection box is eliminated.
[0160] In the above example, a value greater than 150 square centimeters corresponds to the first reference size range corresponding to the area of the largest connected region of the target object in the detection frame, and a value greater than 25 centimeters corresponds to the first reference size range corresponding to the long side length of the smallest bounding rectangle of the area of the largest connected region of the target object in the detection frame.
[0161] In some embodiments, the target detection result may also include the confidence level of the detection box.
[0162] Accordingly, the post-processing method for the target detection result may also include:
[0163] S71, when it is determined that the data information is within the valid data range, obtain the confidence level;
[0164] S72, determine whether the confidence level is less than or equal to the confidence threshold.
[0165] S73, for a detection box with a confidence level less than or equal to the confidence level threshold, determine whether its corresponding at least one size data is within the range of the second reference size. If yes, retain the detection box; otherwise, eliminate the detection box.
[0166] The second reference size range is the range of size data corresponding to the verification boxes in the verification set whose confidence is less than or equal to the confidence threshold.
[0167] For example, in the validation set, the confidence scores of the validation bounding boxes for umbrellas are all greater than 0.3, and the longest side of the minimum bounding rectangle of the largest connected region of each validation bounding box is greater than 15 cm. Meanwhile, for a small number of samples in the validation set, the confidence scores of their validation bounding boxes are all less than 0.4, and the longest side of the minimum bounding rectangle of the largest connected region of each corresponding validation bounding box is greater than 25 cm. Therefore, when the confidence score of a detection box is less than 0.4, we can determine whether the longest side of the minimum bounding rectangle corresponding to the detection box is greater than 25 cm. If it is, the corresponding detection box is retained; otherwise, it is eliminated.
[0168] In the example above, greater than 25 cm is the second reference size range for the long side length of the minimum bounding rectangle corresponding to the detection frame.
[0169] In a specific implementation, the at least one size data mentioned in step S73 may also include the area of the largest connected region of the target object in the detection frame and the short side length and aspect ratio of its smallest bounding rectangle.
[0170] In this embodiment of the invention, the target detection result of the image to be processed may include multiple target objects and their corresponding categories. For target objects of different categories, their corresponding detection boxes can be post-processed separately.
[0171] For target objects of the same category, multiple corresponding detection boxes can be included in the same image to be processed, and each detection box can be post-processed separately.
[0172] Figure 3 This is a schematic diagram of the post-processing device for target detection results in an embodiment of the present invention.
[0173] Reference Figure 3This invention also provides a post-processing device 200 for target detection results. The post-processing device 200 includes a first acquisition module 201, an extraction module 202, a second acquisition module 203, a judgment module 204, and a processing module 205.
[0174] Specifically, the first acquisition module 201 is used to acquire the target detection result of the image to be processed, which includes the detection box of the target object in the image to be processed; the extraction module 202 is used to extract the maximum connected region of the target object in the detection box; the second acquisition module 203 is used to acquire the data information of the maximum connected region; the judgment module 204 is used to judge whether the data information is within the valid data range; the processing module 205 is used to determine whether to retain the detection box when the data information is within the valid data range, and to eliminate the detection box when the data information is outside the valid data range; wherein, the target detection result is obtained by performing target detection on the image to be processed using a preset detection network model, and the valid data range is determined based on the data information of the maximum connected region corresponding to the target object in the validation set of the detection network model.
[0175] In specific implementation, the first acquisition module 201, the extraction module 202, the second acquisition module 203, the judgment module 204, and the processing module 205 can be implemented based on the technical solution of the post-processing method for target detection results disclosed in the embodiments of the present invention.
[0176] This invention also provides an electronic device.
[0177] The electronic device includes a processor and a memory. The memory stores a computer program that can run on the processor; when executed by the processor, the computer program implements the post-processing method for target detection results disclosed in embodiments of the present invention.
[0178] This invention also provides a computer-readable storage medium.
[0179] The computer-readable storage medium stores a computer program that, when executed, implements the post-processing method for target detection results disclosed in the embodiments of the present invention.
[0180] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the invention, even when only a single embodiment is described with respect to a particular feature. The feature examples provided in this disclosure are intended to be illustrative and not limiting, unless otherwise stated. In practice, one or more technical features of the dependent claims may be combined with the technical features of the independent claims as needed and where technically feasible, and the technical features from the respective independent claims may be combined in any suitable manner rather than solely by the specific combinations listed in the claims.
[0181] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A post-processing method for target detection results, characterized in that, include: Obtain the target detection result of the image to be processed, wherein the target detection result includes the detection box of the target object in the image to be processed; Extract the maximum connected region of the target object within the detection frame; Obtain the data information of the largest connected region; Determine whether the data information is within the valid data range. If yes, retain the detection box; otherwise, eliminate the detection box. The data information includes size data, which includes at least two of the following: the area of the largest connected region and the length of its longest side, the length of its shortest side, and the aspect ratio of its smallest bounding rectangle. When it is determined that the data information is within the valid data range, the sample probability corresponding to each size data is obtained, where the sample probability is the probability of the corresponding size data appearing in the verification set; it is determined whether the sample probability is less than or equal to a probability threshold. If so, the sample probability is determined to be a small sample probability, and the size data corresponding to the small sample probability is determined to be a low probability size data; it is determined whether other size data besides the low probability size data are within a first reference size range. If so, the detection box is retained; otherwise, the detection box is eliminated; wherein, the first reference size range is the range in which other size data in the verification set are located below the probability threshold. The target detection result is obtained by performing target detection on the image to be processed using a preset detection network model, and the effective data range is determined based on the data information of the largest connected region corresponding to the target object in the validation set of the detection network model.
2. The post-processing method according to claim 1, characterized in that, The target detection result includes the category corresponding to the target object, and the detection box includes the detection box corresponding to the category that needs to be post-processed.
3. The post-processing method according to claim 2, characterized in that, include: Obtain the false detection rate for the category; Determine whether the false detection rate is greater than or equal to the false detection threshold. If so, determine that the category needs post-processing.
4. The post-processing method according to claim 3, characterized in that, include: A test set is constructed by selecting test images that do not include the category corresponding to the target object; The detection network model is used to perform target detection on each test image in the test set to obtain test detection results corresponding to each test image. The test detection results include false detection boxes corresponding to the category of the target object. The false detection rate is determined by the ratio of the number of test images containing the false detection box to the total number of all test images in the test set.
5. The post-processing method according to claim 3, characterized in that, The false detection threshold includes 5‰.
6. The post-processing method according to claim 1, characterized in that, The step of extracting the maximum connected region of the target object in the detection frame includes: Obtain the first mask for each pixel in the detection frame based on the HSV color mode, and the second mask based on the grayscale mode; The final mask for each pixel is obtained by performing logical operations on the first mask and the second mask. The maximum connected region is determined based on the final mask; The values of the first mask, the second mask, and the final mask each include either a first value or a second value.
7. The post-processing method according to claim 6, characterized in that, The maximum connected region is determined based on the region where the pixel whose final mask is the second value is located.
8. The post-processing method according to claim 7, characterized in that, The logical operation includes logical AND and NOT operations; the logical operation on the first mask and the second mask to obtain the final mask for each pixel includes: Compare the first mask and the second mask; When both the first mask and the second mask are the first value, the final mask of the corresponding pixel is set to the second value; when at least one of the first mask and the second mask is the second value, the final mask of the corresponding pixel is set to the first value.
9. The post-processing method according to claim 8, characterized in that, include: Obtain the color parameter range of the target object in the HSV color mode, and the color parameters of each pixel in the HSV color mode; The color parameter of each pixel is compared with the color parameter range; The first mask for pixels whose color parameters are within the range of the color parameters takes the first value, and the first mask for pixels whose color parameters are outside the range of the color parameters takes the second value.
10. The post-processing method according to claim 8, characterized in that, include: Obtain the average grayscale value of each pixel in the grayscale mode; The grayscale value of each pixel is compared with the average grayscale value. The second mask for pixels whose grayscale value is greater than or equal to the average grayscale value is taken from the first value, and the second mask for pixels whose grayscale value is less than the average grayscale value is taken from the second value.
11. The post-processing method according to claim 1, characterized in that, include: The verification set is formed by selecting verification images that include the target object, wherein the verification images include the bounding box of the target object; The detection network model is used to perform target detection on each verification image in the verification set to obtain a verification detection result corresponding to each verification image. The verification detection result includes the verification bounding box of the target object. Extract the maximum connected region of the target object in the annotation box and the verification box, respectively; Obtain the annotation data information and verification data information of the largest connected region corresponding to the annotation box and the verification box, respectively; Obtain the union of the labeled data information and the verification data information, and use the union as the effective data range.
12. The post-processing method according to claim 1, characterized in that, The target detection result includes the shape information of the target object, the data information includes proportional data, and the effective data range includes the effective proportional range corresponding to the proportional data. The determination of whether the data information is within the valid data range includes: Identify the shape category corresponding to the shape information. Based on the shape category, determine whether the ratio data is within the effective ratio range.
13. The post-processing method according to claim 12, characterized in that, The shape category includes ring shape, the ratio data includes the effective pixel ratio of the largest connected region, and the effective ratio range includes the effective pixel ratio range.
14. The post-processing method according to claim 12, characterized in that, The shape category includes rectangles or approximate rectangles, the ratio data includes the area percentage of the largest connected region, and the effective ratio range includes the effective area percentage range.
15. The post-processing method according to any one of claims 12 to 14, characterized in that, The target detection result includes the confidence level of the detection box; the post-processing method includes: When it is determined that the data information is within the range of valid data, the confidence level is obtained. Determine whether the confidence level is less than or equal to the confidence threshold. For a detection box with a confidence level less than or equal to the confidence threshold, determine whether its corresponding at least one size data is within the range of the second reference size. If yes, retain the detection box; otherwise, eliminate the detection box. Wherein, the second reference size range is the range of size data corresponding to the verification boxes in the verification set whose confidence level is less than or equal to the confidence threshold.
16. A post-processing device for target detection results, characterized in that, include: The first acquisition module is used to acquire the target detection result of the image to be processed, wherein the target detection result includes the detection box of the target object in the image to be processed; An extraction module is used to extract the maximum connected region of the target object within the detection frame; The second acquisition module is used to acquire data information of the maximum connected region; The judgment module is used to determine whether the data information is within the valid data range; the data information includes size data, which includes at least two of the following: the area of the largest connected region and the length of its longest side, the length of its shortest side, and the aspect ratio of its smallest bounding rectangle; The processing module is configured to determine whether to retain the detection box when the data information is within the valid data range, and to eliminate the detection box when the data information is outside the valid data range; When it is determined that the data information is within the valid data range, the sample probability corresponding to each size data is obtained, where the sample probability is the probability of the corresponding size data appearing in the verification set; it is determined whether the sample probability is less than or equal to a probability threshold. If so, the sample probability is determined to be a small sample probability, and the size data corresponding to the small sample probability is determined to be a low probability size data; it is determined whether other size data besides the low probability size data are within a first reference size range. If so, the detection box is retained; otherwise, the detection box is eliminated; wherein, the first reference size range is the range in which other size data in the verification set are located below the probability threshold. The target detection result is obtained by performing target detection on the image to be processed using a preset detection network model, and the effective data range is determined based on the data information of the largest connected region corresponding to the target object in the validation set of the detection network model.
17. An electronic device, characterized in that, include: processor; The memory stores computer programs that can run on the processor; The computer program, when executed by the processor, implements the method as described in any one of claims 1 to 15.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed, it implements the method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Fine-grained vehicle type recognition method based on weak surveillance localization and subclass similarity measurement
CN109359684A
Instrument panel pointer reading prediction method and device, computer equipment and storage medium
CN112115896A
Dust detection method and device based on movement detection, medium and terminal equipment
CN112365486A