A target detection method, device, and equipment

By dividing images into blocks and resizing those with low target density for further detection, the method addresses target scale mismatch, improving detection precision in target detection networks.

CN114626477BActive Publication Date: 2025-07-15AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210283691.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-07-15
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

The existing deep learning-based object detection algorithm has the problem of target scale mismatch, resulting in inaccurate detection results.

Method used

The image to be detected is input to the target detection network, and the first target detection result is obtained and the area of interest with a small density is filtered through coordinate axes projection statistics. After scale transformation, the network detection is input again, and the detection results are fused twice to improve the accuracy.

Benefits of technology

By matching the detection scale of the target detection network and the actual image scale, the accuracy of target detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626477B_ABST
    Figure CN114626477B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an object detection method, device, and equipment. The object detection method includes: inputting an image to be detected into an object detection network to obtain a first object detection result, where the first object detection result includes a first rectangular detection frame for the target object. The image to be detected is segmented to obtain a plurality of segmented images, and the first object detection result is subjected to coordinate axis projection statistics to screen out the segmented images with less target object density. These segmented images are considered as areas with poor detection effects, and then the scale transformation is performed by scaling these segmented images and sent into the object detection network for re-detection to obtain a second object detection result of the target segmented images. The first object detection result and the second object detection result are fused to determine the final object detection result of the image to be detected. Thereby, the detection scale of the object detection network is made to match the actual scale of the image to be detected as much as possible, achieving the purpose of enhancing the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an object detection method, apparatus, and device. Background Art

[0002] Object detection technology is one of the basic problems in the field of computer vision. The purpose is to identify target objects of specific categories in an image and calibrate the positions of target objects of each specific category in the image. This technology has important application value and research value in fields such as pedestrian recognition, license plate recognition, driverless, and defect detection. With the development of deep learning technology, the combination of object detection technology and convolutional neural networks has become closer, and the accuracy and speed of object detection algorithms have been greatly improved compared with before.

[0003] Object detection algorithms based on deep learning often require a large amount of training data that matches the actual situation, and also require a large amount of computing power resources and time to train network parameters in order to achieve better detection effects in actual applications. However, currently, object detection networks usually have the problem of mismatched object scales, resulting in inaccurate object detection results. Summary of the Invention

[0004] In view of this, embodiments of this application provide an object detection method, apparatus, and device to solve the technical problem of inaccurate object detection results.

[0005] To solve the above problems, the technical solutions provided by the embodiments of this application are as follows:

[0006] An object detection method, the method includes:

[0007] Input the image to be detected into an object detection network to obtain a first object detection result of the image to be detected, where the first object detection result includes at least one first rectangular detection box for a target object;

[0008] Divide the image to be detected into blocks to obtain a plurality of block images;

[0009] By performing coordinate axis projection statistics on the first object detection result, obtain an interested region in the image to be detected where the density of the target object is less than a first threshold;

[0010] Determine the proportion of the interested region in each block image, and determine the block images with a proportion greater than a second threshold as target block images;

[0011] After performing scale transformation on the target block images, input them into the object detection network to obtain a second object detection result of the target block images, where the second object detection result includes at least one second rectangular detection box for the target object;

[0012] Determine the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target segmented image.

[0013] In a possible implementation manner, the step of segmenting the image to be detected to obtain a plurality of segmented images includes:

[0014] Determine the size levels of each of the first rectangular detection frames, and calculate the proportion of the quantity of each size level;

[0015] Determine the number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction according to the proportion of the quantity of each size level and the size of the image to be detected;

[0016] Segment the image to be detected according to the number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction to obtain a plurality of segmented images.

[0017] In a possible implementation manner, the step of determining the number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction according to the proportion of the quantity of each size level and the size of the image to be detected includes:

[0018] Calculate 1 divided by the minimum value among the proportions of the quantity of each size level to obtain a first value;

[0019] Calculate the length value of the image to be detected divided by a target sum value, where the target sum value is the sum of the length value and the width value of the image to be detected, to obtain a second value;

[0020] Calculate the width value of the image to be detected divided by the target sum value to obtain a third value;

[0021] Calculate the integer obtained by multiplying the first value by the second value to obtain the number of divisions of the image to be detected in the horizontal axis direction;

[0022] Calculate the integer obtained by multiplying the first value by the third value to obtain the number of divisions of the image to be detected in the vertical axis direction.

[0023] In a possible implementation manner, the step of obtaining the region of interest where the density of the target object in the image to be detected is less than a first threshold by performing coordinate axis projection statistics on the first target detection result includes:

[0024] Project the first rectangular detection frame onto the horizontal coordinate axis to obtain the projection quantity on the horizontal coordinate axis and the horizontal coordinate axis ranges respectively corresponding to the projection quantities on each horizontal coordinate axis;

[0025] Project the first rectangular detection frame onto the vertical coordinate axis to obtain the number of projections on the vertical coordinate axis and the range of the vertical coordinate axis corresponding to each of the numbers of projections on the vertical coordinate axis;

[0026] Determine the target interval as the range of the horizontal coordinate axis where the number of projections on the horizontal coordinate axis is less than or equal to a third threshold and the range of the vertical coordinate axis where the number of projections on the vertical coordinate axis is less than or equal to a fourth threshold;

[0027] Determine the region corresponding to the target interval in the image to be detected as the region of interest where the density of the target object in the image to be detected is less than a first threshold.

[0028] In a possible implementation, the third threshold is determined according to the count median of the number of projections on the horizontal coordinate axis, and the fourth threshold is determined according to the count median of the number of projections on the vertical coordinate axis.

[0029] In a possible implementation, the method further includes:

[0030] Filter out the second rectangular detection frames within the preset edge region of the target sub-image.

[0031] In a possible implementation, determining the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target sub-image includes:

[0032] Use the non-maximum suppression algorithm to fuse the first target detection result of the image to be detected and the second target detection result of the target sub-image to determine the final target detection result of the image to be detected.

[0033] A target detection device, the device includes:

[0034] A first detection unit, configured to input an image to be detected into a target detection network to obtain a first target detection result of the image to be detected, where the first target detection result includes at least one first rectangular detection frame for a target object;

[0035] A partitioning unit, configured to partition the image to be detected to obtain a plurality of sub-images;

[0036] A projection unit, configured to obtain the region of interest where the density of the target object in the image to be detected is less than a first threshold by performing coordinate axis projection statistics on the first target detection result;

[0037] A first determination unit, configured to determine the proportion of the region of interest in each of the sub-block images, and determine the sub-block images with the proportion greater than a second threshold as target sub-block images;

[0038] A second detection unit, configured to input the target sub-block images into the target detection network to obtain a second target detection result of the target sub-block images, where the second target detection result includes at least one second rectangular detection frame for the target object;

[0039] A second determination unit, configured to determine a final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target sub-block images.

[0040] A target detection device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the target detection method as described above is implemented.

[0041] A computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute the target detection method as described above.

[0042] Thus, the embodiments of the present application have the following beneficial effects:

[0043] In the embodiments of the present application, the image to be detected is input into the target detection network to obtain a first target detection result, where the first target detection result includes a first rectangular detection frame for the target object. The image to be detected is partitioned into multiple sub-block images, and the first target detection result is subjected to coordinate axis projection statistics to screen out the sub-block images with a relatively small density of the target object. These sub-block images are considered as regions with poor detection effects, and then scale transformation is performed by scaling these sub-block images and they are input into the target detection network for re-detection to obtain a second target detection result of the target sub-block images. The first target detection result and the second target detection result are fused to determine the final target detection result of the image to be detected. Thereby, the detection scale of the target detection network is made to match the actual scale of the image to be detected as much as possible, achieving the purpose of enhancing the detection accuracy. Description of the Drawings

[0044] Figure 1 A schematic diagram of a scenario example provided by an embodiment of the present application;

[0045] Figure 2 A flowchart of a target detection method provided by an embodiment of the present application;

[0046] Figure 3 A schematic diagram of a target detection result provided by an embodiment of the present application;

[0047] Figure 4 It is a flowchart of the specific steps of the object detection method provided by the embodiment of the present application;

[0048] Figure 5 It is another flowchart of the specific steps of the object detection method provided by the embodiment of the present application;

[0049] Figure 6 It is yet another flowchart of the specific steps of the object detection method provided by the embodiment of the present application;

[0050] Figure 7 It is a schematic diagram of the projection quantity of the first rectangular detection frame provided by the embodiment of the present application;

[0051] Figure 8 It is a schematic diagram of the object segmented image provided by the embodiment of the present application;

[0052] Figure 9 It is a structural diagram of an object detection device provided by the embodiment of the present application. Specific embodiments

[0053] To make the above objects, features, and advantages of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] To facilitate the understanding and explanation of the technical solutions provided by the embodiments of the present application, the background technology of the embodiments of the present application will be described below.

[0055] Object detection algorithms based on deep learning often require a large amount of training data that matches the actual situation, and also require a large amount of computing power resources and time to train network parameters in order to achieve better detection effects in actual applications. However, in the actual test process of the current object detection network, there is usually a problem of mismatched object scales, resulting in inaccurate object detection results. There are mainly two reasons for the problem of mismatched object scales:

[0056] On the one hand, there is a mismatch in object scales between the training set and the actual test data. For example, if the training data in the training set are all large objects, but the actual test data are all small objects, the detection effect will be significantly reduced.

[0057] Second aspect, the training strategy of the object detection network affects the actual results. Taking the YOLOv3 network as an example, although the YOLOv3 network fuses features sampled 8 times, 16 times, and 32 times from the original image, in the COCO2014 training dataset it uses, the number of objects accounting for 5% of the image size reaches 80.74% of the total number of objects. After high-magnification downsampling, the information of these objects that only account for 5% of the size is surely lost and becomes invalid training data. From the experimental results of the YOLOv3 network, the mAP (Mean Average Precision) of the detection results for small objects is 18.3, the mAP of the detection results for medium objects is 35.4, and the mAP of the detection results for large objects is 41.9. So actually, the network is most successful in training large objects and most failed in training small objects.

[0058] Based on this, the embodiments of the present application provide an object detection method, device, and equipment. To facilitate understanding of the object detection method provided by the embodiments of the present application, the following will be described in combination with Figure 1 the following scene example. Among them, Figure 1 is a schematic diagram of a scene example provided by the embodiments of the present application. This method can be applied to the terminal device 101.

[0059] In practical applications, the terminal device 101 acquires the image to be detected, inputs the image to be detected into the object detection network, and obtains the first object detection result. The first object detection result includes the first rectangular detection frame for the target object. The image to be detected is segmented to obtain multiple segmented images, and the first object detection result is statistically projected along the coordinate axes to screen out the segmented images with a relatively small density of target objects. These segmented images are considered as areas with poor detection effects, and then the scale transformation is performed by scaling these segmented images and they are sent into the object detection network for re-detection to obtain the second object detection result of the target segmented image. The first object detection result and the second object detection result are fused to determine the final object detection result of the image to be detected.

[0060] Those skilled in the art can understand that Figure 1 the framework schematic diagram shown is only an example in which the embodiments of the present application can be implemented. The applicable scope of the embodiments of the present application is not limited by any aspect of this framework.

[0061] Based on the above description, the object detection method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0062] Figure 2 is a flowchart of an object detection method provided by the embodiments of the present application. As Figure 1 shown, this method includes S10 - S60:

[0063] S10. Input the image to be detected into the object detection network to obtain the first object detection result of the image to be detected. The first object detection result includes at least one first rectangular detection box for the target object.

[0064] Among them, the first rectangular detection box is a rectangular box for locating the target.

[0065] The object detection network can be any object detection network that outputs detection boxes. Object detection networks are mainly divided into two categories. One category is the second-order detection network, such as Fast RCNN (Fast region-based convolutional neural network), Faster RCNN (Faster region-based convolutional neural network), Mask RCNN (Mask region-based convolutional neural network), etc. The characteristics of this type of method are good detection accuracy but slow speed. The other category is the first-order detection network, such as the YOLO (You only look once) series of networks, retinaNet, SSD (Single Shot MultiBox Detector), etc.

[0066] In a specific example, before inputting the image to be detected into the object detection network, it is necessary to scale (i.e., zoom in or out) the image to be detected to the scale required by the input object detection network.

[0067] In a specific example, Figure 3 is a schematic diagram of the first object detection result provided by the embodiment of the present application. As Figure 3 shown, the first object detection result 001 of the image to be detected includes the first rectangular detection boxes 002, 003, 004, and 005 for the target object.

[0068] S20. Divide the image to be detected into multiple sub-images.

[0069] Due to target mismatch, some regions in the image to be detected have good detection effects, while some regions have poor detection effects. In order to perform secondary detection on the regions with poor detection effects in the image to be detected, the image to be detected can be divided into blocks in a preset manner, so as to further determine the target sub-images that need secondary detection from the sub-images later.

[0070] In a possible implementation manner, the image to be detected can be divided into blocks according to the size of the image to be detected. ThenFigure 4 This is the specific step flowchart of step S20 provided by the embodiment of the present application. As Figure 4 shown, S20 includes S201 - S203:

[0071] S201. Determine the size levels of each first rectangular detection frame, and calculate the proportion of the number of each size level.

[0072] In a specific example, the definition of the size level of the first rectangular detection frame is as follows: the first rectangular detection frame with a size of 96 * 96 or more is a large target detection frame, the first rectangular detection frame with a size between 96 * 96 - 32 * 32 is a medium target detection frame, and the first rectangular detection frame with a size of 32 * 32 or less is a small target detection frame.

[0073] It should be noted that the size level and its standard can be defined according to actual needs, and the embodiment of the present application does not limit this.

[0074] In a specific example, according to the histogram of the number of target detection frames of each size level in the first target detection result, calculate the proportion of the number of each size level.

[0075] S202. According to the proportion of the number of each size level and the size of the image to be detected, determine the number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction.

[0076] In order to divide the image to be detected into image blocks of appropriate size, scientific division is carried out according to the proportion of the number of each size level combined with the size (length and width dimensions) of the image to be detected. The number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction is the number of divisions on the length or width of the image to be detected.

[0077] Figure 5 This is the specific step flowchart of step S202 provided by the embodiment of the present application. As Figure 5 shown, S202 includes S2021 - S2025:

[0078] S2021. Calculate 1 divided by the minimum value in the proportion of the number of each size level to obtain a first value.

[0079] S2022. Calculate the length value of the image to be detected divided by the target sum value, where the target sum value is the sum of the length value and the width value of the image to be detected, to obtain a second value.

[0080] S2023. Calculate the width value of the image to be detected divided by the target sum value to obtain a third value.

[0081] S2024. Calculate the integer obtained by multiplying the first value by the second value to obtain the number of divisions of the image to be detected in the horizontal axis direction.

[0082] S2025. Calculate the integer part after multiplying the first value by the third value to obtain the number of divisions of the image to be detected in the vertical axis direction.

[0083] In a specific example, the calculation formula for the number of divisions in the horizontal axis direction is:

[0084]

[0085] The calculation formula for the number of divisions in the vertical axis direction is:

[0086]

[0087] Where m is the number of divisions in the horizontal axis direction, n is the number of divisions in the vertical axis direction, i is the length of the image to be detected, j is the width of the image to be detected, a, b, and c are the proportions of the number of large, medium, and small target detection frames in the first rectangular detection frame respectively, and [] is the integer part symbol.

[0088] It should be noted that the embodiments of the present application do not limit the execution order of S2021, S2022, and S2023. In one possible implementation, S2021, S2022, and S2023 can be executed simultaneously, or S2022 and S2023 can be executed first, and then S2021 can be executed. It is also possible to execute S2023 first, then S2022, and finally S2021. It is also possible to execute S2022 first, then S2023, and finally S2021, etc. The embodiments of the present application do not limit the execution order of S2024 and S2025. In one possible implementation, S2024 and S2025 can be executed simultaneously, or S2024 can be executed first, and then S2025 can be executed.

[0089] S203. Divide the image to be detected according to the number of divisions in the horizontal axis direction and the vertical axis direction of the image to be detected to obtain a plurality of divided images.

[0090] In a specific example, the image to be detected is divided into m*n divided images.

[0091] S30. By performing coordinate axis projection statistics on the first target detection result, obtain the region of interest in the image to be detected where the density of the target object is less than the first threshold.

[0092] In order to screen out the regions with poor detection effects, the first target detection result, that is, the first rectangular detection frame, is subjected to coordinate axis projection statistics. The projection statistical result can reflect the density distribution of the target objects in different ranges of the coordinate axis, and then the region with low target density can be determined, which is the region of interest.

[0093] Figure 6It is a flowchart of the specific steps of S30 provided by the embodiments of this application. As Figure 6 shown, the method S30 includes S301 - S304:

[0094] S301. Project the first rectangular detection frame onto the horizontal coordinate axis to obtain the projection quantity on the horizontal coordinate axis and the horizontal coordinate axis ranges respectively corresponding to the projection quantities on each horizontal coordinate axis.

[0095] If the projections of N first rectangular detection frames overlap on the horizontal coordinate axis, the range count of the overlapping part on the horizontal coordinate axis needs to be increased by N, and the range count of the horizontal coordinate axis range without the projection of the first rectangular detection frame is 0.

[0096] S302. Project the first rectangular detection frame onto the vertical coordinate axis to obtain the projection quantity on the vertical coordinate axis and the vertical coordinate axis ranges respectively corresponding to the projection quantities on each vertical coordinate axis.

[0097] Similarly, if the projections of M first rectangular detection frames overlap on the vertical coordinate axis, the range count of the overlapping part on the vertical coordinate axis needs to be increased by M, and the range count of the vertical coordinate axis range without the projection of the first rectangular detection frame is 0.

[0098] S303. Determine the horizontal coordinate axis ranges with the projection quantity on the horizontal coordinate axis less than or equal to the third threshold and the vertical coordinate axis ranges with the projection quantity on the vertical coordinate axis less than or equal to the fourth threshold as the target intervals.

[0099] In a possible implementation, the third threshold is determined according to the median of the count of the projection quantity on the horizontal coordinate axis, and the fourth threshold is determined according to the median of the count of the projection quantity on the vertical coordinate axis.

[0100] Among them, when determining the median, the 0 count is excluded.

[0101] S304. Determine the region corresponding to the target interval in the image to be detected as the region of interest where the density of the target object in the image to be detected is less than the first threshold.

[0102] It should be noted that the embodiments of this application do not limit the execution order of S301 and S302. In a possible implementation, S301 and S302 can be executed simultaneously, or S302 can be executed first and then S301.

[0103] Through Figure 7 the optional examples shown to illustrate S301 - S304.

[0104] Figure 7 It is a schematic diagram of the projection quantity of the first rectangular detection frame provided by the embodiments of this application. As Figure 7 shown, project Figure 3The detection results in the coordinate axis projection are statistically analyzed, and the first rectangular detection frame 002-005 is projected onto the horizontal coordinate axis. It is obtained that the number of projections in the range x2-x3 on the horizontal coordinate axis is 2, and the numbers of projections in the ranges x3-x4, x5-x6, and x7-x8 are 1 respectively.

[0105] The first rectangular detection frame 002-005 is projected onto the vertical coordinate axis, and the number of projections on the vertical coordinate axis for the range y2-y3 is 1, the number of projections on the range y3-y4 is 3, and the number of projections on the range y5-y6 is 1.

[0106] The median of the projection counts 1, 1, 1, 2 on the horizontal coordinate axis is 1, and the third threshold is determined to be 1. The median of the projection counts 1, 1, 3 on the vertical coordinate axis is 1, and the fourth threshold is determined to be 1.

[0107] The horizontal coordinate axis range where the number of projections on the horizontal coordinate axis is less than or equal to the third threshold 1 is: x1-x2, x3-x9. The vertical coordinate axis range where the number of projections on the vertical coordinate axis is less than or equal to the fourth threshold 1 is: y1-y3, y4-y7. The target intervals are determined to be x1-x2, x3-x9 and y1-y3, y4-y7.

[0108] The area corresponding to the target interval in the image to be detected 001, that is, the area consisting of the horizontal coordinate axis range: x1-x2, x3-x9, and the vertical coordinate axis range: y1-y3, y4-y7, is determined as the area of interest in the image to be detected where the density of the target object is less than the first threshold, such as Figure 7 Fills the area with a medium diagonal line.

[0109] S40: Determine the proportion of the region of interest in each block image, and determine the block image whose proportion is greater than a second threshold as the target block image.

[0110] The region of interest is an area where the target object density is less than the first threshold, that is, an area with low target density. If the region of interest in a certain block image accounts for a large proportion, that is, the target density in most areas of the block image is low, it is very likely that the detection effect is poor due to the mismatch of the target scale. In this case, the block image is determined as a target block image, that is, the target block image is considered to be an area with poor detection effect.

[0111] Among them, the second threshold can be set according to needs, and this application does not limit this.

[0112] In a specific example, Figure 8 A schematic diagram of a target block image provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, among the four block images 801, 802, 803, and 804, 802 and 804, in which the proportion of the region of interest is greater than 40%, are used as target block images.

[0113] S50. After performing scale transformation on the target segmented image, input it into the target detection network to obtain the second target detection result of the target segmented image. The second target detection result includes at least one second rectangular detection frame for the target object.

[0114] Perform scale transformation on the poorly detected area to make its scale match the detection scale of the target detection network as much as possible.

[0115] S60. Determine the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target segmented image.

[0116] In a possible implementation, S60 includes: using the non-maximum suppression algorithm to fuse the first target detection result of the image to be detected and the second target detection result of the target segmented image to determine the final target detection result of the image to be detected.

[0117] In a possible implementation, the target detection method further includes:

[0118] Filter out the second rectangular detection frames within the preset edge area of the target segmented image.

[0119] In a specific example, a dog in the image to be detected is taken as a target, and only the head of a dog is divided into the preset edge area of a target segmented image. At this time, the target corresponding to the second rectangular detection frame is the head of a dog, but in fact it is only a small part of one target. Therefore, such second rectangular detection frames need to be filtered out.

[0120] In the embodiment of the present application, the image to be detected is input into the target detection network to obtain the first target detection result, and the first target detection result includes the first rectangular detection frame for the target object. The image to be detected is segmented to obtain multiple segmented images, and the first target detection result is subjected to coordinate axis projection statistics to screen out the segmented images with less density of the target object. These segmented images are considered as areas with poor detection effects, and then scale transformation is performed by scaling these segmented images and sent into the target detection network for re-detection to obtain the second target detection result of the target segmented image. The first target detection result and the second target detection result are fused to determine the final target detection result of the image to be detected. Thus, the detection scale of the target detection network is made to match the actual scale of the image to be detected as much as possible, achieving the purpose of enhancing the detection accuracy.

[0121] Figure 9 The structural diagram of a target detection device provided by the embodiment of the present application is as Figure 9 shown. The device includes:

[0122] The first detection unit 901 is configured to input the image to be detected into the target detection network to obtain a first target detection result of the image to be detected, where the first target detection result includes at least one first rectangular detection box for the target object;

[0123] The chunking unit 902 is configured to chunk the image to be detected to obtain a plurality of chunked images;

[0124] The projection unit 903 is configured to obtain a region of interest in the image to be detected where the density of the target object is less than the first threshold by performing coordinate axis projection statistics on the first target detection result;

[0125] The first determination unit 904 is configured to determine the proportion of the region of interest in each chunked image, and determine the chunked image with a proportion greater than the second threshold as the target chunked image;

[0126] The second detection unit 905 is configured to input the target chunked image into the target detection network to obtain a second target detection result of the target chunked image, where the second target detection result includes at least one second rectangular detection box for the target object;

[0127] The second determination unit 906 is configured to determine the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target chunked image.

[0128] In a possible implementation, the chunking unit includes:

[0129] A calculation subunit is configured to determine the size levels of the respective first rectangular detection boxes and calculate the proportion of the number of each size level;

[0130] A first determination subunit is configured to determine the number of divisions of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction according to the proportion of the number of each size level and the size of the image to be detected;

[0131] A chunking subunit is configured to chunk the image to be detected according to the number of divisions of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction to obtain a plurality of chunked images.

[0132] In a possible implementation, the first determination subunit is specifically configured to:

[0133] Calculate 1 divided by the minimum value among the proportions of the number of each size level to obtain a first value;

[0134] Calculate the length value of the image to be detected divided by the target sum value, where the target sum value is the sum of the length value and the width value of the image to be detected, to obtain a second value;

[0135] Calculate the width value of the image to be detected divided by the target sum value to obtain a third value;

[0136] Round the product of the first value and the second value to obtain the number of divisions of the image to be detected in the horizontal axis direction;

[0137] Round the product of the first value and the third value to obtain the number of divisions of the image to be detected in the vertical axis direction.

[0138] In a possible implementation, the projection unit includes:

[0139] The first projection subunit is configured to project the first rectangular detection frame onto the horizontal axis to obtain the number of projections on the horizontal axis and the horizontal axis ranges respectively corresponding to the numbers of projections on each horizontal axis;

[0140] The second projection subunit is configured to project the first rectangular detection frame onto the vertical axis to obtain the number of projections on the vertical axis and the vertical axis ranges respectively corresponding to the numbers of projections on each vertical axis;

[0141] The second determination subunit is configured to determine the horizontal axis ranges where the number of projections on the horizontal axis is less than or equal to the third threshold and the vertical axis ranges where the number of projections on the vertical axis is less than or equal to the fourth threshold as the target intervals;

[0142] The third determination subunit is configured to determine the region corresponding to the target interval in the image to be detected as the region of interest where the density of the target object in the image to be detected is less than the first threshold.

[0143] In a possible implementation, the third threshold is determined according to the median count of the number of projections on the horizontal axis, and the fourth threshold is determined according to the median count of the number of projections on the vertical axis.

[0144] In a possible implementation, the apparatus further includes:

[0145] The filtering unit is configured to filter out the second rectangular detection frames within the preset edge region of the target block image.

[0146] In a possible implementation, the second determination unit is specifically configured to:

[0147] Use the non-maximum suppression algorithm to fuse the first target detection result of the image to be detected and the second target detection result of the target block image to determine the final target detection result of the image to be detected.

[0148] In the embodiment of the present application, the image to be detected is input into the target detection network to obtain the first target detection result, and the first target detection result includes the first rectangular detection frame for the target object. The image to be detected is segmented to obtain a plurality of segmented images, and the first target detection result is subjected to coordinate axis projection statistics to screen out the segmented images with a relatively small density of target objects. These segmented images are considered as areas with poor detection effects, and then the scale transformation is performed by scaling these segmented images and they are sent into the target detection network for re-detection to obtain the second target detection result of the target segmented images. The first target detection result and the second target detection result are fused to determine the final target detection result of the image to be detected. Thereby, the detection scale of the target detection network is made to match the actual scale of the image to be detected as much as possible, achieving the purpose of enhancing the detection accuracy.

[0149] It should be noted that the embodiments in this specification are described in a progressive manner. The key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method section.

[0150] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0151] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0152] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.

[0153] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target detection method, characterized in that, The method includes: Inputting the image to be detected into the target detection network to obtain a first target detection result of the image to be detected, where the first target detection result includes at least one first rectangular detection frame for the target object; Dividing the image to be detected into blocks to obtain a plurality of block images; Obtaining a region of interest in the image to be detected where the density of the target object is less than a first threshold by performing coordinate axis projection statistics on the first target detection result; Determining the proportion of the region of interest in each block image, and determining the block images with a proportion greater than a second threshold as target block images; Inputting the target block images after scale transformation into the target detection network to obtain a second target detection result of the target block images, where the second target detection result includes at least one second rectangular detection frame for the target object; Determining a final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target block images; The dividing the image to be detected into blocks to obtain a plurality of block images includes: determining the size levels of each of the first rectangular detection frames, and calculating the proportion of the number of each size level; determining the number of divisions of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction according to the proportion of the number of each size level and the size of the image to be detected; dividing the image to be detected according to the number of divisions of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction to obtain a plurality of block images; Among them, the calculation formula for the number of divisions in the horizontal coordinate axis direction is: The calculation formula for the number of divisions in the vertical coordinate axis direction is: Among them, m is the number of divisions in the horizontal coordinate axis direction, n is the number of divisions in the vertical coordinate axis direction, i is the length of the image to be measured, j is the width of the image to be measured, a, b, and c are the proportions of the number of large, medium, and small target detection frames in the first rectangular detection frame respectively, and [] is the rounding symbol.

2. The method according to claim 1, characterized in that, The determining the number of divisions of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction according to the proportion of the number of each size level and the size of the image to be detected includes: Calculating 1 divided by the minimum value among the proportions of the number of each size level to obtain a first value; Calculating the length value of the image to be detected divided by the target sum value, where the target sum value is the sum of the length value and the width value of the image to be detected, to obtain a second value; Calculating the width value of the image to be detected divided by the target sum value to obtain a third value; Calculating the first value multiplied by the second value and then taking the integer to obtain the number of divisions of the image to be detected in the horizontal coordinate axis direction; Calculating the first value multiplied by the third value and then taking the integer to obtain the number of divisions of the image to be detected in the vertical coordinate axis direction.

3. The method according to claim 1, wherein The obtaining a region of interest in the image to be detected where the density of the target object is less than a first threshold by performing coordinate axis projection statistics on the first target detection result includes: Project the first rectangular detection frame onto the horizontal coordinate axis to obtain the projection quantity on the horizontal coordinate axis and the horizontal coordinate axis ranges corresponding to the projection quantities on the respective horizontal coordinate axes; Project the first rectangular detection frame onto the vertical coordinate axis to obtain the projection quantity on the vertical coordinate axis and the vertical coordinate axis ranges corresponding to the projection quantities on the respective vertical coordinate axes; Determine the target intervals as the horizontal coordinate axis ranges where the projection quantity on the horizontal coordinate axis is less than or equal to a third threshold and the vertical coordinate axis ranges where the projection quantity on the vertical coordinate axis is less than or equal to a fourth threshold; Determine the region corresponding to the target interval in the image to be detected as the region of interest in the image to be detected where the density of the target object is less than a first threshold.

4. The method according to claim 3, characterized in that, The third threshold is determined according to the count median of the projection quantity on the horizontal coordinate axis, and the fourth threshold is determined according to the count median of the projection quantity on the vertical coordinate axis.

5. The method according to claim 1, wherein The method further includes: Filter out the second rectangular detection frames within the preset edge region of the target segmented image.

6. The method according to claim 1 or 5, characterized in that, The determining the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target segmented image includes: Using the non-maximum suppression algorithm to fuse the first target detection result of the image to be detected and the second target detection result of the target segmented image to determine the final target detection result of the image to be detected.

7. A target detection device, characterized in that, The device includes: A first detection unit configured to input an image to be detected into a target detection network to obtain a first target detection result of the image to be detected, where the first target detection result includes at least one first rectangular detection frame for a target object; A segmentation unit configured to segment the image to be detected to obtain a plurality of segmented images; A projection unit configured to obtain the region of interest in the image to be detected where the density of the target object is less than a first threshold by performing coordinate axis projection statistics on the first target detection result; A first determination unit configured to determine the proportion of the region of interest in each of the segmented images, and determine the segmented images with the proportion greater than a second threshold as target segmented images; A second detection unit configured to input the target segmented image into the target detection network to obtain a second target detection result of the target segmented image, where the second target detection result includes at least one second rectangular detection frame for the target object; A second determination unit configured to determine the final target detection result of the image to be detected according to the first target detection result of the image to be detected and the second target detection result of the target segmented image; The segmentation unit includes: A calculation sub-unit configured to determine the size levels of the respective first rectangular detection frames and calculate the proportion of the quantity of each size level; A first determination sub-unit configured to determine the division quantity of the image to be detected in the horizontal coordinate axis direction and the vertical coordinate axis direction according to the proportion of the quantity of each size level and the size of the image to be detected; A block sub-unit, configured to block the image to be detected according to the number of divisions of the image to be detected in the horizontal axis direction and the vertical axis direction, so as to obtain a plurality of block images; Wherein, the calculation formula for the number of divisions in the horizontal axis direction is: The calculation formula for the number of divisions in the vertical axis direction is: Wherein, m is the number of divisions in the horizontal axis direction, n is the number of divisions in the vertical axis direction, i is the length of the image to be detected, j is the width of the image to be detected, a, b, and c are respectively the proportions of the number of large, medium, and small target detection frames in the first rectangular detection frame, and [] is the rounding symbol.

8. A target detection device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the target detection method according to any one of claims 1-6 is implemented.

9. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on the terminal device, the terminal device is caused to execute the target detection method according to any one of claims 1-6.