Object detection device and method
The object detection device and method address the challenge of high-definition image processing by using a density estimation model to adjust threshold values and sizes, ensuring accurate and efficient object detection in high-definition images and videos.
Patent Information
- Application Number
- PCT/JP2024/000022
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-10
AI Technical Summary
Existing object detection methods struggle with high-definition images and videos due to limitations in input size, leading to increased computational load or inaccurate detection when using threshold-based region extraction, as they either oversize or undersize the regions for object detection, affecting accuracy.
An object detection device and method that utilizes a density estimation model to determine optimal regions for object detection by adjusting threshold values and sizes based on probability density, merging overlapping regions, and classifying object sizes to ensure accurate detection.
Enables accurate object detection from high-definition images and videos by automatically adjusting threshold values and sizes, reducing computational load and maintaining detection accuracy by optimizing region selection.
Smart Images

Figure JP2024000022_10072025_PF_FP_ABST
Abstract
Description
Object detection device and method
[0001] The disclosed technology relates to an object detection device and an object detection method.
[0002] An object detection device is a device that estimates the class of an object, such as a person or a vehicle, contained in an input image, the coordinate information of a rectangular bounding box that surrounds the object area in the image, and the reliability of the detection result. In recent years, several object detection models using deep learning have been proposed. As object detection models based on deep learning, YOLO (You Only Look Once) and RetinaNet, which infer bounding boxes and object classes simultaneously, have been proposed (Non-Patent Documents 1 and 2). In addition, R-CNN, which separately detects object candidate areas and classifies objects, and an improved version of it, Faster R-CNN, have been proposed (Non-Patent Documents 3 and 4).
[0003] Furthermore, when performing object detection from high-resolution images or videos such as 4K (3840 x 2160 pixels) or 8K (7680 x 4320 pixels), the input image size to the object detection model is limited to a maximum of 1536 x 1536 pixels in YOLO v5, for example. Therefore, high-resolution images cannot be input to the object detection model at their original size. Therefore, a method has been proposed in which the input high-resolution image is divided to match the input image size of the object detection model, the results of object detection from each divided image are aggregated, and the result of object detection from the entire image is output. Several methods have been proposed for this purpose, depending on the image division method. For example, a method has been proposed in which a group of images equally divided to match the input image size of the object detection model and an entire image reduced in size are input to the object detection model, respectively. In this method, the coordinate information of the obtained bounding box is scaled, and the detection results of each divided image and the reduced image are combined to output the final result (Non-Patent Document 5).
[0004] Also proposed are techniques that use probability density estimation or cluster detection to estimate the distribution of regions in an image where an object exists, extract only regions where the object is predicted to exist, and apply an object detection model (Non-Patent Documents 6 and 7). For example, in the technique of Non-Patent Document 6, a high-resolution image serving as an input image is first reduced, and a density map that estimates the distribution of regions in the image where the object exists is output. Next, based on the output density map, a rectangular region in the image where object detection is to be applied is extracted. At this time, a single threshold value for the probability density of an object where the object is deemed to exist is set in advance, and a portion with a probability density equal to or greater than the threshold value is extracted as a rectangular region. Finally, a portion corresponding to the extracted rectangular region is extracted from the input image, and an object detection model is applied to obtain a detection result.
[0005] J. Redmon et al., "You Only Look Once: Unified, Real-Time Object Detection," 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779-788.T. -Y. Lin et al., "Focal Loss for Dense Object Detection," 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2999-3007.R. Girshick et al., "Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation," 2014 IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 580-587.S. Ren et al., "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137-1149, 1 June 2017.H. Uzawa et al., "High-definition object detection technology based on AI inference scheme and its implementation", IEICE Electronics Express, 2021, Volume 18, Issue 22, Pages 20210323.C. Li et al., "Density Map Guided Object Detection in Aerial Images," 2020 IEEE / CVF CVPRW, 2020, pp. 737-746.F.Yang et al., "Clustered Object Detection in Aerial Images," 2019 IEEE / CVF ICCV, 2019, pp. 8310-8319.
[0006] In a method such as the method described in Non-Patent Document 5, which equally divides a high-resolution image and performs object detection, it is not possible to selectively apply object detection to each divided image. Therefore, as the number of divided images increases with the increase in resolution of the input image, the amount of calculation required for object detection increases. On the other hand, in methods based on cluster detection or probability density estimation such as the methods described in Non-Patent Documents 6 and 7, the area in the image to which object detection is applied is narrowed down, which may enable a significant reduction in the amount of calculation depending on the image.
[0007] However, when performing probability density estimation, a high-density distribution with low variance is generally obtained from areas of an image where small objects are densely packed, while a low-density distribution with high variance is obtained from areas where large objects are packed. When extracting rectangular regions using a single threshold, as in conventional methods, setting a high threshold to match small objects may result in the extraction of rectangular regions corresponding to large objects, or in the extraction of small rectangular regions that only partially crop out the large objects. On the other hand, setting a low threshold to match large objects may result in the rectangular regions being larger than the input size of the object detection model. This requires the cropped image to be reduced when input to the object detection model, potentially resulting in small objects being distorted and undetectable. Thus, when extracting regions above a single threshold from the results of probability density estimation as regions to be input to the object detection model, it may be impossible to extract regions of optimal size, potentially resulting in a deterioration in object detection accuracy.
[0008] The disclosed technology has been made in consideration of the above points, and aims to cut out an area from a target image of an optimal size according to the scene of the target image as an area to which an object detection model is applied.
[0009] A first aspect of the present disclosure is an object detection device including: an estimation unit that estimates a density map of a target image that is the subject of object detection using a density estimation model that has been generated in advance by machine learning, so as to estimate a density map that indicates the probability density of an object being present at each position in the image; an extraction unit that extracts a region that is the subject of object detection based on the probability density in the density map estimated by the estimation unit; a detection unit that detects an object by applying an object detection model that has been generated in advance by machine learning to detect an object from an image to a region of the target image that corresponds to the region extracted by the extraction unit; and a processing unit that performs processing to set the size of the region extracted by the extraction unit according to the density map or the object detection result by the detection unit.
[0010] A second aspect of the present disclosure is an object detection method, in which an estimation unit estimates a density map of a target image that is the subject of object detection using a density estimation model generated in advance by machine learning so as to estimate a density map that indicates the probability density of an object being present at each position in the image, an extraction unit extracts a region that is the subject of object detection based on the probability density in the density map estimated by the estimation unit, a detection unit detects an object by applying an object detection model generated in advance by machine learning to detect an object from an image to a region of the target image that corresponds to the region extracted by the extraction unit, and a processing unit performs processing to set the size of the region to be extracted by the extraction unit according to the density map or the object detection result by the detection unit.
[0011] According to the disclosed technology, it is possible to extract an area from a target image with an optimal size according to the scene of the target image as an area to which an object detection model is applied.
[0012] FIG. 1 is a block diagram showing the hardware configuration of an object detection device. FIG. 2 is a block diagram showing the functional configuration of an object detection device according to a first embodiment. FIG. 3 is a diagram showing pairs of thresholds and rectangular area sizes. FIG. 4 is a diagram for explaining extraction of rectangular areas. FIG. 5 is a diagram for explaining merging of rectangular areas. FIG. 6 is a diagram showing an overview of threshold update. FIG. 7 is a diagram showing an overview of threshold update. FIG. 8 is a diagram for explaining adjustment of the size of rectangular areas. A flowchart showing an example of object detection processing according to the first embodiment. A flowchart showing an example of threshold adjustment processing. A flowchart showing an example of size adjustment processing. A block diagram showing the functional configuration of an object detection device according to a second embodiment. A diagram for explaining extraction of rectangular areas based on local maxima. A diagram for explaining search for local maxima. A flowchart showing an example of object detection processing according to the second embodiment.
[0013] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that the same or equivalent components and parts in each drawing are given the same reference numerals. Also, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.
[0014] 1 is a block diagram showing the hardware configuration of an object detection device 10 according to a first embodiment. As shown in Fig. 1, the object detection device 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage 14, an input unit 15, a display unit 16, and a communication I / F (Interface) 17. Each component is connected to each other via a bus 19 so as to be able to communicate with each other.
[0015] The CPU 11 is a central processing unit that executes various programs and controls each component. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 controls the above components and performs various arithmetic processing in accordance with the program stored in the ROM 12 or the storage 14. In this embodiment, the ROM 12 or the storage 14 stores an object detection program for executing the object detection processing described below.
[0016] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores programs or data as a working area. The storage 14 is configured with storage devices such as a hard disk drive (HDD) or a solid state drive (SSD), and stores various programs including an operating system and various data.
[0017] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs. The display unit 16 is, for example, a liquid crystal display, and displays various information. The display unit 16 may also function as the input unit 15 by employing a touch panel system. The communication I / F 17 is an interface for communicating with other devices. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.
[0018] Next, the functional configuration of the object detection device 10 according to the first embodiment will be described. Fig. 2 is a block diagram showing an example of the functional configuration of the object detection device 10. As shown in Fig. 2, the object detection device 10 includes, as its functional configuration, an estimation unit 22, an extraction unit 24, a detection unit 26, and a processing unit 28. Furthermore, a density estimation model 30 and an object detection model 32 are stored in a predetermined storage area of the object detection device 10. Each functional configuration is realized when the CPU 11 reads out an object detection program stored in the ROM 12 or the storage 14, expands it in the RAM 13, and executes it.
[0019] The density estimation model 30 is a machine learning model generated in advance by machine learning so as to estimate a density map indicating the probability density of an object at each position in an image. Specifically, the density estimation model 30 is generated by using an image and ground truth data of the density map for that image as training data and updating the parameters of the density estimation model 30 so that the density map output from the density estimation model 30 approaches the ground truth data.
[0020] The estimation unit 22 estimates a density map of a target image that is a target for object detection, using the density estimation model 30. Specifically, the estimation unit 22 sequentially acquires each frame of a video input to the object detection device 10 as a target image. The estimation unit 22 inputs the acquired target image to the density estimation model 30 and acquires a density map output from the density estimation model 30. The estimation unit 22 passes the estimated density map to the extraction unit 24.
[0021] The extraction unit 24 extracts a target region for object detection based on the probability density in the density map passed from the estimation unit 22, using pairs of multi-stage thresholds and rectangular region sizes supplied from the processing unit 28, which will be described later. The threshold is a value for identifying a portion of the density map where the probability density is equal to or greater than the threshold, and the size of the rectangular region is a size determined so as to correspond one-to-one with the threshold of each stage. In this embodiment, the region is a rectangle, and the length of one side of the rectangular region (number of pixels) is the size of the rectangular region. Also, as shown in FIG. 3, the number of pairs is N (an arbitrary natural number), and the threshold is T i , threshold T i The length of one side of the rectangular area corresponding to i where i = 1, 2, ..., N, and T 1 >T 2 > ... > T N , and S 1 <S 2 <...<S N Let us assume that:
[0022] Specifically, as shown in FIG. 4, the extraction unit 24 extracts the density map estimated by the estimation unit 22 when the probability density is greater than or equal to a threshold T i The area of size S is centered on the area where the above pixels are continuous. iThe extraction unit 24 extracts the rectangular area as a target area for object detection. i A rectangular area is extracted for (i=1, 2, . . . , N).
[0023] In addition, when the extraction unit 24 extracts multiple rectangular areas from the target image, it merges rectangular areas that have an overlapping degree equal to or greater than a predetermined value and for which the minimum size of a detectable object satisfies a predetermined condition relative to the size of the rectangular area after merging.
[0024] Specifically, as shown in the upper diagram of FIG. i and a threshold T j Assume that rectangular area j extracted based on the rectangular area i has an overlapping area (shaded area in FIG. 5). The extraction unit 24 calculates, for example, IoU (Intersection over Union) as the overlapping degree. Note that IoU = area of overlapping area between rectangular area i and rectangular area j / area of union of rectangular area i and rectangular area j. The extraction unit 24 calculates the overlapping degree when the overlapping degree is equal to or greater than a predetermined value T iou In the above cases, the rectangular area that contains rectangular area i and rectangular area j is set as the rectangular area after merging, as shown in the lower diagram of Figure 5. By merging rectangular areas in this way, it is possible to reduce the areas where object detection is performed redundantly when applying object detection model 32, and reduce the amount of calculation required for object detection.
[0025] However, the extraction unit 24 does not merge rectangular areas when the following formula (1) is satisfied. min i and a min j is the threshold T i and T j the minimum size of the object detected from the image corresponding to the rectangular area extracted based on ij is the area of the rectangular region after merging, A min is the minimum size of an object that can be detected by the object detection model 32. This makes it possible to merge rectangular regions while preventing objects that are smaller than the detectable size from being missed in object detection.
[0026]
[0027] The extraction unit 24 assigns an ID, which is identification information, to each of the extracted rectangular areas, and passes rectangular information including the ID, the position of the rectangular area, and the size (vertical and horizontal lengths) of the rectangular area to the extraction unit 24 and the detection unit 26. The position of the rectangular area may be the coordinates of a predetermined point of the rectangular area, such as the center coordinates of the rectangular area or the coordinates of the upper left corner.
[0028] The object detection model 32 is a machine learning model that has been generated in advance by machine learning to detect objects from images. The object detection model 32 outputs metadata including the class of an object contained in an input image, a bounding box indicating the area of the object, and the confidence level of the detection.
[0029] The detection unit 26 detects an object by applying the object detection model 32 to a region of the target image corresponding to the rectangular region extracted by the extraction unit 24. Specifically, the detection unit 26 identifies a region on the target image corresponding to the rectangular region extracted by the extraction unit 24, based on the rectangle information passed from the extraction unit 24. The detection unit 26 cuts out an image of the identified region from the target image, enlarges or reduces the cut-out image to match the input size of the object detection model 32, and inputs the image to the object detection model 32. The detection unit 26 outputs the metadata output by the object detection model 32 as a detection result, and passes the detection result to the processing unit 28.
[0030] The processing unit 28 performs processing to set the size of the rectangular region to be extracted by the extraction unit 24, in accordance with the density map estimated by the estimation unit 22 or the object detection result by the detection unit 26. Specifically, the ... above-mentioned multi-stage threshold T i and the size of the rectangular area S i The initial value may be set arbitrarily. When the target image is the first frame of the video, the processing unit 28 sets the initial value T i and S i If the target image is a frame two or later, the processing unit 28 calculates a threshold T i and the size of the rectangular area S iThen, the processing unit 28 adjusts the adjusted T i and S i and the pair are supplied to the extraction unit 24.
[0031] Threshold T i The adjustment of the probability density in the density map is described in more detail below. i The area of the part where the above pixels are consecutive is the threshold T i Paired with size S i The area corresponding to (S i ×S i ), the threshold T i In addition, the processing unit 28 reduces the probability density in the density map by a predetermined value or a predetermined rate. i The area of the part where the above pixels are continuous is size S i The area corresponding to (S i ×S i ), the threshold T i is increased by a predetermined value or a predetermined rate.
[0032] For example, the processing unit 28 detects whether the probability density is greater than or equal to a threshold value T i The average area a of the part specified as the above i Calculate the average area a i and a constant α (0<α<1), the threshold T i is updated using the following equation (2).
[0033]
[0034] In FIG. i ×S i >a i When the threshold T i The outline of updating S is shown below. The lower the threshold value, the larger the area of the part where the probability density is equal to or greater than the threshold value, and the higher the threshold value, the smaller the area of the part where the probability density is equal to or greater than the threshold value. i ×S i >a i In this case, the threshold T i By lowering the threshold, the threshold is updated so as to expand the area where the probability density is above the threshold.
[0035] In FIG.i ×S i ≦a i When the threshold T i The processing unit 28 updates S i ×S i ≦a i In this case, the threshold T i As a result, the processing unit 28 detects an area with a high probability density where smaller objects are distributed by increasing the threshold T i and size S i The threshold is adjusted so that the pair is extracted using the threshold T i+1 A pair of a threshold value or a lower threshold value and a size associated with that threshold value is a threshold T i is the threshold for extracting the area where objects larger than the target object are distributed.
[0036] By adjusting the threshold value, the adjacent threshold value T k and threshold T k+1 But, |T k -T k+1 If |<ε (ε is an arbitrary positive number), the processing unit 28 integrates these thresholds. For example, T i By lowering i T i+1 Approaching |T i -T i+1 If |<ε, or T i By increasing i T i-1 Approaching |T i-1 -T i |<ε. Specifically, the processing unit 28 integrates the thresholds T k and size S k+1 and are considered as a new pair, and the threshold T k+1 and size S k will be abolished.
[0037] Next, the adjustment of the size of the rectangular region will be described in more detail. The processing unit 28 adjusts the size S iThe processing unit 28 classifies the size of the object detected from the area into a plurality of categories of different sizes. When the number of objects included in the category of large object sizes is greater than the number of objects included in the category corresponding to the average object size, the processing unit 28 classifies the size S i In addition, when the number of objects included in a category with small object sizes is greater than the number of objects included in a category with an average object size, the processing unit 28 enlarges the size S i is reduced by a predetermined percentage.
[0038] For example, the processing unit 28 classifies the size of the bounding box of an object detected in a certain frame into three categories, S (Small), M (Medium), and L (Large), based on the lower limit of the object size detectable by the object detection model 32. The processing unit 28 may classify the size based on the length of one side of the bounding box of the object, or may classify the size based on the area of the bounding box.
[0039] The processing unit 28 also determines the size S i Among all the objects detected from the image cut out from the target image based on the rectangular area of the size S, the object of the most numerous category is the M category, which is an example of a category where the size of the object corresponds to the average. i This is because if the S category is the most common, there is a possibility that a smaller object that has not been detected by the object detection model 32 exists in the cut-out image. Also, if the L category is the most common, the object is adjusted to the size S. i This is because the size is larger than the size S and there is a possibility that the entire image cannot be detected as a single object from the cut-out image. i (i=1, 2, . . . , N). At this time, the processing unit 28 i If the image cut out by size S does not exist, i If the most common category of objects detected from the image cut out in is M categories, then size S i No adjustments are made.
[0040] The case where the area of the bounding box is used as the classification criterion will be described in more detail. The processing unit 28 determines the upper limit A of the area of the bounding box of an object that can be detected by the object detection model 32. max and lower limit A min The processing unit 28 also obtains the upper limit A of the area. max and lower limit A min Based on the values of M For example, A M = (A max +A min ) / 2.
[0041] Furthermore, the processing unit 28 classifies the detected objects into S, M, and L using predetermined criteria based on the detection results of the detection unit 26 in each frame. For example, Bbox But, A min ≦a Bbox If X, then S category, X≦a Bbox If <2X, M category, 2X≦a Bbox ≦A max In this case, it may be classified as L category. max -A min ) / 3.
[0042] The processing unit 28 calculates the average area of the bounding boxes of the objects belonging to each category as a S , a M , and a L Let the maximum area of the detected object be a max , the minimum area is a min Using these values, the size S is calculated using the following formula (3): i Update.
[0043]
[0044] However, as described above, when the M category contains the most objects, the size of the rectangular region is not adjusted. Figure 8 shows an example of adjusting the size of the rectangular region using the above formula (3). As shown in Figure 8(a), when the L category contains the most objects, the size of the rectangular region is increased to capture the entire image of the object and make it easier to detect. On the other hand, as shown in Figure 8(b), when the S category contains the most objects, the size of the rectangular region is reduced to prevent small objects from being overlooked. This makes it possible to appropriately adjust the size of the rectangular region within a range in which the detected object does not exceed the upper or lower limit of the size of a detectable object.
[0045] Next, the operation of the object detection device 10 according to the first embodiment will be described. Fig. 9 is a flowchart showing the flow of the object detection process performed by the object detection device 10. The object detection process is performed by the CPU 11 reading an object detection program from the ROM 12 or the storage 14, expanding it into the RAM 13, and executing it. Note that the object detection process is an example of an object detection method disclosed herein.
[0046] In step S10, the CPU 11, functioning as the estimation unit 22, acquires, as a target image, a frame of a video input to the object detection device 10. Next, in step S12, the CPU 11, functioning as the estimation unit 22, uses the density estimation model 30 to estimate a density map of the target image.
[0047] Next, in step S14, the CPU 11, as the extraction unit 24, extracts a rectangular area to be subjected to object detection based on the probability density in the density map estimated in step S12, using the multi-stage threshold T i and size S i Here, the multi-stage T i and S i When the target image is the first frame of the video, the pair i and S i In addition, when the target image is a frame from the second frame onward, the T i and S i It is a pair with.
[0048] Next, in step S16, the CPU 11, functioning as the extraction unit 24, merges rectangular regions in the target image that have an overlapping degree equal to or greater than a predetermined value and whose minimum detectable object size satisfies a predetermined condition relative to the size of the rectangular region after merging. The CPU 11, functioning as the extraction unit 24, assigns an ID, which is identification information, to each extracted rectangular region, and passes rectangular information including the ID, the position of the rectangular region, and the size of the rectangular region to the extraction unit 24 and the detection unit 26.
[0049] Next, in step S18, the CPU 11, functioning as the detection unit 26, identifies an area on the target image that corresponds to the rectangular area extracted by the extraction unit 24, based on the rectangle information passed from the extraction unit 24. The CPU 11, functioning as the detection unit 26, cuts out an image of the identified area from the target image, enlarges or reduces the cut-out image to match the input size of the object detection model 32, and inputs the image to the object detection model 32. The CPU 11, functioning as the detection unit 26, then acquires metadata that is the object detection result output by the object detection model 32.
[0050] Next, in step S20, the CPU 11, functioning as the detection unit 26, outputs the object detection result and transfers the detection result to the processing unit 28. Next, in step S22, the CPU 11, functioning as the processing unit 28, determines whether the target image is the last frame of the video. If it is the last frame, the object detection process ends, and if there are subsequent frames, the process proceeds to step S30.
[0051] In step S30, a threshold adjustment process is executed, which will now be described with reference to FIG.
[0052] In step S32, the CPU 11 functions as the processing unit 28 and sets the variable i to 1. Next, in step S34, the CPU 11 functions as the processing unit 28 and causes the extraction unit 24 to extract the probability density in the density map of the target image that is equal to or greater than the threshold T i The average area a of the part specified as the above i Then, the CPU 11, as the processing unit 28, calculates S i ×S i ≦a iIt is determined whether or not i ×S i ≦a i In this case, the process proceeds to step S36. i ×S i >a i In this case, the process proceeds to step S38.
[0053] In step S36, the CPU 11 controls the processing unit 28 to calculate the threshold value T i On the other hand, in step S38, the CPU 11 controls the processing unit 28 to increase the threshold value T i Next, in step S40, the CPU 11 controls the processing unit 28 to reduce the adjacent threshold T k and threshold T k+1 Regarding |T k -T k+1 Determine whether |<ε. |T k -T k+1 If |<ε, the process proceeds to step S42, and |T k -T k+1 If |≧ε, the process proceeds to step S44.
[0054] In step S42, the CPU 11 controls the processing unit 28 to calculate the threshold value T k and size S k+1 and are considered as a new pair, and the threshold T k+1 and size S k By abolishing the threshold T k and threshold T k+1 Next, in step S44, the CPU 11, functioning as the processing unit 28, determines whether the variable i has reached N, which is the number of pairs of thresholds and rectangular area sizes. If i<N, the process proceeds to step S46, and the CPU 11, functioning as the processing unit 28, increments i by 1 and returns to step S34. On the other hand, if i=N, the threshold adjustment process ends, and the process returns to the object detection process ( FIG. 9 ).
[0055] Next, in step S50, a size adjustment process is executed. The size adjustment process will now be described with reference to FIG.
[0056] In step S52, the CPU 11, functioning as the processing unit 28, classifies the sizes of the bounding boxes of objects detected from the target image into three categories, S, M, and L, based on the lower limit of the object size detectable by the object detection model 32. Next, in step S54, the CPU 11, functioning as the processing unit 28, sets a variable i to 1.
[0057] Next, in step S56, the CPU 11, as the processing unit 28, calculates the size S i It is determined whether the rectangular area extracted by the size S exists. i If the rectangular area extracted in step S5 exists, the process proceeds to step S58, and if not, the process proceeds to step S66.
[0058] In step S58, the CPU 11, as the processing unit 28, calculates the size S i It is determined whether the most common category of all objects detected from the image cut out from the target image based on the rectangular area of size S is the L category. If it is the L category, the process proceeds to step S60, and the CPU 11 causes the processing unit 28 to calculate the size S. i On the other hand, if the most popular category is not category L, the process proceeds to step S62.
[0059] In step S62, the CPU 11 controls the processing unit 28 to calculate the size S i It is determined whether the most common category of all objects detected from the image cut out from the target image based on the rectangular area of size S. If it is the S category, the process proceeds to step S64, and the CPU 11 causes the processing unit 28 to calculate the size S. i On the other hand, if the most common category is not the S category, that is, if the most common category is the M category, the process proceeds to step S66.
[0060] In step S66, the CPU 11, functioning as the processing unit 28, determines whether the variable i has reached N, which is the number of pairs of a threshold value and a rectangular area size. If i<N, the process proceeds to step S68, where the CPU 11, functioning as the processing unit 28, increments i by 1 and returns to step S56. On the other hand, if i=N, the size adjustment process ends and the process returns to the object detection process ( FIG. 9 ). Then, the process returns to step S10, where the next frame of the video is acquired as the target image.
[0061] In the object detection process described above, either the threshold adjustment process in step S30 or the size adjustment process in step S50 may be performed first, or only one of them may be performed.
[0062] As described above, the object detection device according to the first embodiment estimates a density map of a target image to be subjected to object detection using a density estimation model previously generated by machine learning, so as to estimate a density map indicating the probability density of an object's presence at each position in the image. The object detection device also extracts a region to be subjected to object detection based on the probability density in the estimated density map. The object detection device then detects the object by applying an object detection model previously generated by machine learning for detecting objects from images to a region of the target image corresponding to the extracted region. Furthermore, the object detection device performs processing to set the size of the region to be extracted based on the region extraction result or object detection result. Specifically, the object detection device adjusts at least one of the threshold for identifying a portion where the probability density is equal to or greater than a threshold and the size of the region paired with the threshold, based on the extraction result or object detection result for a frame prior to the target image in the video.
[0063] As a result, the object detection device according to the first embodiment can extract an area from the input image of an optimal size according to the scene in the target image as an area to which the object detection model is applied. This makes it possible to detect objects from high-definition video while suppressing degradation of object detection accuracy due to image reduction or cutting off of objects caused by an inappropriate size of the extracted area.
[0064] In general, when there are multiple pairs of thresholds and associated rectangular area sizes, it is difficult to manually adjust these to improve object detection accuracy. On the other hand, the object detection device according to the first embodiment can automatically perform these adjustments, which has the effect of reducing the effort required for adjustment.
[0065] In the first embodiment, the size of the bounding box of a detected object is classified into three categories: S, M, and L. This is because three is the minimum number that can accommodate cases where the object size is too large, appropriate, and too small, respectively, but the number of categories may be four or more.
[0066] Second Embodiment Next, a second embodiment will be described. In the object detection device according to the second embodiment, components similar to those of the object detection device 10 according to the first embodiment are denoted by the same reference numerals, and detailed descriptions thereof will be omitted. Furthermore, the hardware configuration of the object detection device according to the second embodiment is similar to the hardware configuration of the object detection device 10 according to the first embodiment shown in FIG. 1 , and therefore description thereof will be omitted.
[0067] The functional configuration of the object detection device according to the second embodiment will be described. As shown in Fig. 12, the object detection device 210 includes, as its functional configuration, an estimation unit 22, an extraction unit 224, a detection unit 26, and a processing unit 228. A density estimation model 30 and an object detection model 32 are stored in a predetermined storage area of the object detection device 210. Each functional configuration is realized when the CPU 11 reads out an object detection program stored in the ROM 12 or storage 14, expands it in the RAM 13, and executes it.
[0068] The processing unit 228 searches for a pixel where the probability density has a maximum value in the density map estimated by the estimation unit 22. The processing unit 228 also assumes a normal distribution centered on the pixel where the probability density has a maximum value, and calculates a variance that serves as a threshold value based on the maximum value.
[0069] Specifically, the processing unit 228 searches the entire density map for pixels where the probability density is a maximum value, as shown in the upper diagram of Fig. 13. For example, the processing unit 228 may search for the maximum value using a sliding window, as shown in Fig. 14. The example in Fig. 14 uses a sliding window of 3 x 3 pixels. Each square in Fig. 14 represents a pixel in a portion of the density map, and the number in each square roughly represents the probability density of each pixel in the density map.
[0070] The processing unit 228 compares the central pixel of the sliding window with the eight surrounding pixels in size, and if the probability density of the central pixel is greater than the probability densities of any of the other pixels, it determines that probability density as a maximum value and the coordinates of that pixel as maximum value coordinates. i (i=1, 2, ..., K), the maximum coordinates are (x i , y i ) The method for searching for the maximum value is not limited to the above example, and any method may be applied.
[0071] Furthermore, the processing unit 228 assumes that the objects are distributed in a normal distribution pattern centered on the maximum value coordinates, as shown in the middle diagram of Fig. 13. Specifically, the processing unit 228 assumes that the objects are distributed in a normal distribution pattern centered on the maximum value coordinates (x i , y i ) is a normal distribution centered on f i (x, y). In general, normal distribution f i (x, y) has a covariance of 0 in the x-axis and y-axis directions and a variance of σ i Assuming that, it is given by the following equation (4).
[0072]
[0073] The processing unit 228 is i (x i , y i ) = M i Therefore, f i Variance of (x, y) σ i is calculated by the following formula (5).
[0074]
[0075] The processing unit 228 finds the maximum value M i and the maximum coordinate (x i , y i ) and the calculated variance σ i and is passed to the extraction unit 224.
[0076] The extraction unit 224 extracts a portion where pixels having a probability density equal to or greater than a threshold are consecutive around the local maximum coordinates searched for by the processing unit 228. Specifically, the extraction unit 224 extracts a portion that includes the range of variance calculated as the threshold in the normal distribution assumed by the processing unit 228 as a region to be subjected to object detection.
[0077] Specifically, the extraction unit 224 extracts the variance σ i and a constant T, the probability density is f i (x i ±Tσ i , yi±Tσ i ) (optional compound symbol) or more, that is, M i exp(-T 2 ) or more, the part of the continuous pixels is defined as the maximum coordinate (x i , y i ) is searched. The constant T is a positive real number that represents the interval of the variance to be extracted, and is set to an arbitrary value externally. Then, the extraction unit 224 extracts a rectangular area surrounding the searched portion, as shown in the lower diagram of FIG.
[0078] The extraction unit 224 also merges overlapping rectangular areas using the same method as the extraction unit 24 in the first embodiment. The extraction unit 224 passes rectangle information about the extracted rectangular areas to the detection unit 26.
[0079] Next, the operation of the object detection device 210 according to the second embodiment will be described. Fig. 15 is a flowchart showing the flow of the object detection process performed by the object detection device 210. The object detection process is performed by the CPU 11 reading out an object detection program from the ROM 12 or storage 14, expanding it into the RAM 13, and executing it. Note that in the object detection process of the second embodiment, processes that are similar to those in the object detection process of the first embodiment are assigned the same step numbers, and detailed descriptions thereof will be omitted.
[0080] After steps S10 and S12, in step S200, the CPU 11, as the processing unit 228, searches for pixels whose probability density has a maximum value from the entire density map estimated in step S12. The CPU 11, as the processing unit 228, calculates the K maximum values found by the search as M i (i=1, 2, ..., K), the maximum coordinates are (x i , y i )
[0081] Next, in step S202, the CPU 11, as the processing unit 228, calculates a normal distribution f i Assuming (x, y), the maximum value M i From f i Variance of (x, y) σ i The CPU 11, as the processing unit 228, calculates the searched maximum value M i and the maximum coordinate (x i , y i ) and the calculated variance σ i and is passed to the extraction unit 224.
[0082] Next, in step S204, the CPU 11, functioning as the extraction unit 224, extracts the variance σ i and a constant T, the probability density is M i exp(-T 2 ) or more, the part of the continuous pixels is defined as the maximum coordinate (x i , y i Then, the CPU 11 functions as the extraction unit 224 to extract a rectangular area surrounding the searched portion, and passes rectangular information about the extracted rectangular area to the detection unit 26.
[0083] Next, after steps S16 to S20, in the next step S222, the CPU 11 as the processing unit 228 determines whether or not the target image is the last frame of the video. If it is the last frame, the object detection process ends, and if there are subsequent frames, the process returns to step S10.
[0084] As described above, the object detection device according to the second embodiment searches for the local maximum coordinates in a density map where the probability density reaches a local maximum, and assumes a normal distribution centered on the local maximum coordinates. The object detection device then calculates the variance of the assumed normal distribution based on the local maximum, and extracts a rectangular region that encloses the portion where the probability density is equal to or greater than a threshold determined based on the variance. This makes it possible to extract a rectangular region of a size corresponding to the density map, and, as in the first embodiment, it is possible to cut out from the input image an area of an optimal size corresponding to the scene in the target image as an area to which the object detection model is applied.
[0085] Furthermore, in the second embodiment, by using the coordinates of the local maximum values found, it is possible to extract a rectangular area that matches the size of the object using only constants that specify the range of variance, without using a threshold value and the size of the rectangular area as in the first embodiment, which significantly reduces the burden of setting and tuning the constants.
[0086] In the above embodiments, the extraction unit extracts a rectangular region, but regions of other shapes may be extracted. In this case, when extracting the corresponding region from the target image or when inputting the extracted image to the object detection model, it is necessary to process the image into a shape and size that can be input to the object detection model.
[0087] In addition, the object detection process executed by the CPU in each of the above embodiments by reading the software (program) may be executed by various processors other than the CPU. Examples of processors in this case include PLDs (Programmable Logic Devices) whose circuit configuration can be changed after manufacture, such as FPGAs (Field-Programmable Gate Arrays), and dedicated electrical circuits, which are processors having a circuit configuration designed specifically for executing specific processes, such as ASICs (Application Specific Integrated Circuits). The object detection process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, etc.). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor elements.
[0088] In addition, in each of the above embodiments, the object detection processing program is described as being pre-stored (installed) in the ROM 12 or the storage 14, but this is not limiting. The program may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The program may also be downloaded from an external device via a network.
[0089] The following additional notes are provided regarding the above-described embodiments.
[0090] (Supplementary Item 1) An object detection device including: an estimation unit that estimates a density map of a target image that is the subject of object detection using a density estimation model that has been generated in advance by machine learning, so as to estimate a density map that indicates the probability density of an object being present at each position in the image; an extraction unit that extracts a region that is the subject of object detection based on the probability density in the density map estimated by the estimation unit; a detection unit that detects an object by applying an object detection model that has been generated in advance by machine learning to detect an object from an image to a region of the target image that corresponds to the region extracted by the extraction unit; and a processing unit that performs processing to set the size of the region extracted by the extraction unit, depending on the density map or the object detection result by the detection unit.
[0091] (Supplementary Item 2) When each frame of an input video is the target image, the processing unit adjusts at least one of the threshold for identifying a portion where the probability density is equal to or greater than a threshold and the size of the region paired with the threshold, based on the density map for a frame in the video that precedes the target image or the object detection result by the detection unit.
[0092] (Supplementary Item 3) The object detection device according to Supplementary Item 2, wherein the processing unit lowers the threshold by a predetermined value or a predetermined percentage when the area of a portion of consecutive pixels whose probability density is equal to or greater than the threshold is smaller than an area corresponding to the size of the region paired with the threshold, and raises the threshold by a predetermined value or a predetermined percentage when the area of the portion is larger than an area corresponding to the size of the region paired with the threshold.
[0093] (Supplementary Item 4) The object detection device according to Supplementary Item 2, wherein the processing unit classifies the sizes of objects detected from a first size region into multiple categories of different sizes in the object detection result by the detection unit for the previous frame, and performs at least one of a process of enlarging the first size by a predetermined percentage if the number of objects included in the category with a larger object size is greater than the number of objects included in the category with a smaller object size is greater than the number of objects included in the category with a smaller object size is greater than the number of objects included in the category with a smaller object size.
[0094] (Supplementary Item 5) The object detection device described in Supplementary Item 1, wherein the processing unit searches for a pixel in the density map estimated by the estimation unit where the probability density has a maximum value, assumes a normal distribution centered on the pixel where the probability density has a maximum value, and calculates the variance of the normal distribution based on the maximum value, and the extraction unit extracts a portion around the pixel searched for by the processing unit that includes the range of the variance in the normal distribution as a region to be subjected to object detection.
[0095] (Supplementary Item 6) The object detection device according to any one of Supplementary Items 1 to 5, wherein when the extraction unit extracts multiple regions, it merges regions that have an overlapping degree equal to or greater than a predetermined value and for which the minimum size of a detectable object satisfies a predetermined condition relative to the size of the region after merging.
[0096] (Supplementary Item 7) An object detection method, comprising: an estimation unit estimating a density map of a target image that is the target of object detection using a density estimation model generated in advance by machine learning so as to estimate a density map indicating the probability density of an object being present at each position in the image; an extraction unit extracting a region that is the target of object detection based on the probability density in the density map estimated by the estimation unit; a detection unit detecting an object by applying an object detection model generated in advance by machine learning for detecting an object from an image to a region of the target image that corresponds to the region extracted by the extraction unit; and a processing unit performing processing to set the size of the region to be extracted by the extraction unit according to the density map or a result of object detection by the detection unit.
[0097] (Supplementary Item 8) An object detection program that causes a computer to function as: an estimation unit that estimates a density map of a target image that is the subject of object detection using a density estimation model that has been generated in advance by machine learning, so as to estimate a density map that indicates the probability density of an object being present at each position in the image; an extraction unit that extracts a region that is the subject of object detection based on the probability density in the density map estimated by the estimation unit; a detection unit that detects an object by applying an object detection model that has been generated in advance by machine learning to detect an object from an image to a region of the target image that corresponds to the region extracted by the extraction unit; and a processing unit that performs processing to set the size of the region extracted by the extraction unit, depending on the density map or the object detection result by the detection unit.
[0098] (Supplementary Item 9) An object detection device including: a memory; and at least one processor connected to the memory, wherein the processor is configured to: estimate a density map of a target image that is the subject of object detection using a density estimation model generated in advance by machine learning so as to estimate a density map that indicates the probability density of an object being present at each position in the image; extract a region that is the subject of object detection based on the probability density in the estimated density map; detect an object by applying an object detection model generated in advance by machine learning to detect an object from an image to a region of the target image that corresponds to the extracted region; and perform processing to set the size of the region to be extracted according to the density map or the object detection result.
[0099] (Supplementary Item 10) A non-transitory recording medium storing a program executable by a computer to perform object detection processing, the object detection processing including: estimating a density map of a target image that is the target of object detection using a density estimation model generated in advance by machine learning so as to estimate a density map that indicates the probability density of an object being present at each position of the image; extracting a region that is the target of object detection based on the probability density in the estimated density map; detecting an object by applying an object detection model generated in advance by machine learning to a region of the target image that corresponds to the extracted region to detect an object from an image; and performing processing to set the size of the region to be extracted according to the density map or the object detection result.
[0100] REFERENCE SIGNS LIST 10, 210 Object detection device 11 CPU 12 ROM 13 RAM 14 Storage 15 Input unit 16 Display unit 17 Communication I / F 19 Bus 22 Estimation unit 24, 224 Extraction unit 26 Detection unit 28, 228 Processing unit 30 Density estimation model 32 Object detection model
Claims
1. An object detection device, comprising: - an estimation unit that estimates a density map of a target image to be subjected to object detection by using a density estimation model generated in advance by machine learning so as to estimate a probability density indicating the presence of an object at each position of the image; - an extraction unit that extracts a region to be subjected to object detection based on the probability density in the density map estimated by the estimation unit; - a detection unit that detects an object by applying an object detection model generated in advance by machine learning to the region of the target image corresponding to the region extracted by the extraction unit; and - a processing unit that performs processing for setting the size of the region to be extracted by the extraction unit according to the density map or the object detection result by the detection unit.
2. The object detection device according to claim 1, wherein when each frame of the input video is used as the target image, the processing unit adjusts at least one of the threshold for specifying a portion where the probability density is equal to or greater than the threshold and the size of the region paired with the threshold based on the density map of a frame preceding the target image in the video or the object detection result by the detection unit.
3. The object detection device according to claim 1, wherein the processing unit searches for a pixel having a maximum probability density in the density map estimated by the estimation unit, assumes a normal distribution centered on the pixel having the maximum probability density, calculates a variance of the normal distribution based on the maximum value, and the extraction unit extracts, as a region to be subjected to object detection, a portion including the range of the variance in the normal distribution around the pixel searched by the processing unit.
4. An object detection method, comprising: estimating a density map indicating a probability density of an object existing at each position of an image by using a density estimation model generated in advance by machine learning; extracting a region to be an object detection target based on the probability density in the density map estimated by the estimation unit; detecting an object by applying an object detection model generated in advance by machine learning to a region of the target image corresponding to the region extracted by the extraction unit to detect the object from the image; and performing processing for setting a size of a region to be extracted by the extraction unit according to the density map or an object detection result by the detection unit.
Citation Information
Patent Citations
Image detection device and control program and image detection method
JP2014089626A
Image processing apparatus, and image processing method
JP2019032773A
Object detection device, object detection method and object detection program
JP2021149666A
Information processing device, information processing method, and program
JP2022057352A