Object detection device
The dual detection system in the object detection device addresses the challenge of detecting small objects in high-resolution images by combining difference image analysis and deep learning-based object detection in divided images, thereby improving detection accuracy and reducing upgrade costs.
Patent Information
- Application Number
- JP2021081250
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-05-12
AI Technical Summary
Current object detection devices struggle with detecting small objects in high-resolution images, leading to decreased detection accuracy and the need for significant upgrades to process high-resolution image information.
The object detection device employs a dual detection system, where the first detection means uses difference images and moving object detection to identify objects based on size and movement, and the second detection means uses deep learning to detect objects in divided images, improving detection accuracy.
This approach enhances the detection accuracy of small objects in high-resolution images, reducing the processing burden and costs associated with upgrading existing devices to handle high-resolution inputs.
Smart Images

Figure 0007699464000001 
Figure 0007699464000002 
Figure 0007699464000003
Abstract
Description
Technical Field
[0001] The present invention relates to an object detection device that detects an object based on image information.
Background Art
[0002] As a technology for detecting an object, a technology for detecting an object using deep learning (a machine learning method using a multi-layer neural network) (an object detection technology using AI) has been studied. For example, many end-to-end methods such as SSD (Single Shot MultiBox Detector) and YOLO (You Only Look Once) that simultaneously detect the position and category of an object based on image information indicating a captured image captured by an imaging means such as a CCD camera have been proposed. These methods are based on multi-task learning that simultaneously performs learning by a multi-layer neural network for detecting the position of an object and learning by a multi-layer neural network for discriminating the category of an object. The object detection technology by SSD is disclosed in, for example, Non-Patent Document 1, and the object detection technology by YOLO is disclosed in, for example, Non-Patent Document 2.
[0003] In recent years, the performance of imaging means has improved, and the number of pixels (resolution) of image information indicating a captured image has a tendency to increase. For example, currently, many surveillance cameras with a resolution of 2K (2 million pixels) or less are used, but in the future, those corresponding to 4K (8 million pixels) or 8K (33 million pixels) may become popular. If an imaging means with a large number of pixels (high resolution) can be used, high-definition image information can be obtained, and the detection accuracy of an object is improved. On the other hand, the current object detection device is configured to process image information with a resolution of less than 2K (for example, several hundred × several hundred pixels). Therefore, when processing image information with a resolution corresponding to 4K, 8K, or higher by the current object detection device, there are many detection omissions, and the detection accuracy of the object may decrease. Considerable labor and cost are required to configure the current object detection device to be capable of processing high-resolution image information. Therefore, the inventor has developed and filed a technology that can improve the detection accuracy of an object while using a current object detection device by dividing a captured image into a plurality of divided images and processing the divided image information indicating each divided image.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, when using deep learning to detect an object existing in a monitoring area based on a captured image of a wide monitoring area captured by an imaging means arranged at a distance, the image of the object included in the captured image may become very small. For example, when detecting whether a person exists in a monitoring area such as a riverbed downstream of a dam, in the captured image of the monitoring area, the image of the person existing in the monitoring area is very small. As described above, when the image of an object included in the captured image is small, the object cannot be detected even by using the technique of dividing the captured image and processing the divided image information indicating each divided image as described above. The inventor of the present invention has variously studied a technique for detecting an object when the image of the object included in the captured image is small. As a result, it has been found that the object can be detected by paying attention to the size and moving distance (moving speed) of the object image even when the object image is small. The present invention has been devised in view of such points, and an object thereof is to provide a technique capable of improving the detection accuracy of an object included in a captured image.
Means for Solving the Problems
[0006] The object detection device of the present invention includes an imaging means and 、 a first detection means and a second detection means; and is provided with. The imaging means sequentially outputs image information indicating a captured image obtained by capturing a monitoring area. As the imaging means, known imaging means such as a CCD camera can be used. The first detection means detects a moving object (for example, a person who is a detection target) existing in the monitoring area based on the image information output from the imaging means. The first detection means includes a difference image creation means, a moving object detection means, and a first object detection means. The difference image creation means sequentially creates difference image information indicating a difference image of the captured images indicated by two pieces of image information output from the imaging means at different times. The difference image shows an image of a moving object obtained by removing a stationary background image from two captured images. As a method for creating the difference image information, various known methods can be used. The interval between the two pieces of image information is set to an integer multiple (including "1") of the interval at which the image information is output from the imaging means, for example. The moving object detection means detects the moving speed and moving trajectory of the moving object based on the difference image information. The moving speed and moving trajectory of the moving object can be detected by using various known methods. The first object detection means detects the moving object as an object to be detected when the moving speed and moving trajectory of the moving object detected by the moving object detection means satisfy the set conditions set corresponding to the object to be detected. As the set conditions, conditions specific to the object to be detected are used. For example, when the object to be detected is a person, conditions such as "the average moving speed within the first set period is within the range of a lower limit value (e.g., the speed of walking slowly) and an upper limit value (e.g., the speed of running fast)" and "the moving trajectory within the second set period forms a trajectory of a continuous predetermined shape" are set. As the first set period and the second set period, a period capable of detecting the unique movement of the person to be detected is set. The second detection means; Detect an object (e.g., a person, etc.) to be detected included in the captured image indicated by the image information output from the imaging means. The second detection means includes a second object detection means, an image segmentation means, and a detection result synthesis means. The image segmentation means divides the captured image indicated by the image information output from the imaging means into a plurality of divided images. As a method of dividing the captured image into a plurality of divided images (the number of divided images, the size of the divided images, the number of divisions, etc.), an appropriate method can be used. For example, a method of dividing at equal intervals in the vertical and horizontal directions, or a method of dividing at equal intervals in the vertical and horizontal directions with the same number of divisions is used. The second object detection means; Based on each divided image information, detect the objects included in each divided image. As the second object detection means, known image processing means such as SSD and YOLO that simultaneously detect the position and category of an object based on image information are used. The detection result synthesis means synthesizes the object detection results for each divided image by the second object detection means and outputs them as the object detection result of the captured image (the object detection result of the captured image based on the object detection results for each divided image). As a method of synthesizing the object detection results for each divided image, for example, a method of converting the position information of the object in the divided image into the position information of the object in the captured image is used. In the present invention, even when the size of an object included in a captured image is small, the object can be detected. Thereby, the detection accuracy of the object included in the captured image can be improved. Further, since object detection by the first detection means and object detection by the second detection means can be performed, the detection accuracy of the object can be further enhanced. In different forms of the present invention, the image segmentation means divides a captured image indicated by image information output from the imaging means into at least a first number of first divided images and into a second number of second divided images. The first number and the second number are set such that boundary portions of the first divided images and boundary portions of the second divided images do not overlap in parallel. In this form, it is allowed that boundary portions of the first divided images and boundary portions of the second divided images intersect. Thereby, for example, an object existing across a boundary portion of one of the divided images that cannot be detected in object detection for one of the divided images can be detected by object detection for the other divided image. As the first number and the second number, appropriate numbers can be set. The types of the divided images are not limited to two types of the first number of first divided images and the second number of second divided images. In this form, a decrease in the detection accuracy of an object at a boundary portion of one of the first divided image and the second divided image can be compensated for by a detection result of an object for the other divided image. The mode of using the object detection result by the first detection means and the object detection result by the second detection means can be set as appropriate. In different forms of the present invention, first, object detection by the second detection means is executed, and when the object cannot be detected by the second detection means, object detection by the first detection means is configured to be executed. In this form, the processing burdens of the first detection means and the second detection means can be reduced. In different forms of the present invention, it is configured to execute object detection by the first detection means and object detection by the second detection means in parallel. In this form, an object included in the captured image can be detected in a short time.
Effect of the Invention
[0007] The present invention can improve the detection accuracy of an object included in a captured image.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0009] Hereinafter, embodiments of the object detection apparatus of the present invention will be described with reference to the drawings. FIG. 1 shows a block diagram of an object detection apparatus 10 according to an embodiment. The object detection apparatus 10 according to an embodiment includes a processing means 20, an imaging means 50, a storage means 60, an input means 70, an output means 80, etc., which are connected by a wired communication line, a wireless communication line, or the like. The imaging means 50 is constituted by, for example, a digital camera using a CCD or a CMOS. The imaging means 50 outputs image information indicating the captured image at set time intervals (frame rate). Note that the imaging means 50 is arranged so that the monitoring area is included in the captured image. The imaging means 50 corresponds to the "imaging means" of the present invention, and the image information output from the imaging means 50 corresponds to the "image information indicating the captured imaging image" of the present invention. The storage means 60 is constituted by a ROM, a RAM, or the like, and stores a program for executing the processing of the processing means 20 and various data. The input means 70 is constituted by a keyboard, a touch panel, or the like, and inputs various information. The output means 80 is constituted by a display means constituted by a liquid crystal display device, an organic EL display device, or the like, a printing means, or the like, and outputs various information. When a display means capable of inputting information by touching a display portion displayed on the display screen is used as the display means, the input means 70 is constituted by a touch sensor. The imaging means 50, the memory means 60, the input means 70, the output means 80, etc. may be arranged at a location separate from the processing means 20.
[0010] The processing means 20 is constituted by a CPU or the like. The processing means 20 has a first detection means 30 and a second detection means 40. Based on the image information output from the imaging means 50, the second detection means 40 uses deep learning or the like to detect an object included in the captured image indicated by the image information. Preferably, the second detection means is set to detect an object existing in the monitoring area included in the captured image. Since the second detection means can detect the category and position of an object, it can also detect a specific object (the object to be detected). When the size of an object included in the captured image indicated by the image information is small and the second detection means 40 cannot detect the object, the first detection means 30 detects the object included in the captured image. The first detection means 30 detects that the object included in the image information is the object to be detected when the moving speed and moving trajectory of an object included in the captured image (preferably, an object existing in the monitoring area included in the captured image) satisfy the set conditions set corresponding to the object to be detected.
[0011] First, the second detection means 40 will be described. The second detection means 40 has a second object detection means 41, an image segmentation means 42, and a detection result synthesis means 43.
[0012] The image segmentation means 42 divides the captured image (referred to as the "original image") indicated by the image information output from the imaging means 50 into a plurality of divided images. The image information also includes the image information output from the imaging means 50 and stored in the memory means 60. The method of dividing the captured image by the image segmentation means 42 will be described later.
[0013] The second object detection means 41 detects an object and its position included in the divided image based on the divided image information indicating the divided image divided by the image division means 42. Note that the second object detection means 41 can also detect an object and its position included in the captured image based on the image information indicating the captured image. As the second object detection means 41, various known object detection means that use deep learning to detect the category and position of an object included in the captured image or the divided image based on the image information indicating the captured image or the divided image information indicating the divided image can be used. For example, object detection means that detect the category and position of an object using the SSD or YOLO method can be used. For example, as shown in FIG. 2, SSD is based on a multi-layer CNN (Convolutional Neural Network) and is composed of a layer that estimates the candidate region of the object's existence and a layer that discriminates the object within the candidate region of the existence region. In the layer that estimates the candidate region of the object's existence, the image information is divided into rectangular regions (default boxes) of a plurality of predetermined sizes, and the candidate region of the object's existence (bounding box) is estimated while considering the deviation of the rectangular regions. In the layer that discriminates the object within the candidate region of the existence region, a separately learned CNN is used to discriminate the object within the candidate region of the existence region.
[0014] The detection result synthesis means 43 synthesizes the object detection results for each divided image by the second object detection means 41 and outputs them as the object detection results for the captured image. For example, the object detection results for each divided image are synthesized in a state where the position information of the object in each divided image is converted into the position information in the captured image, and are output as the object detection results for the captured image (the "original image") (in this case, the object detection results for the captured image based on the object detection results for each divided image). Note that the detection result synthesis means 43 can synthesize the object detection results for each divided image by the second object detection means 41 and the object detection result for the captured image ("original image"), and output it as the object detection result for the captured image ("original image") (in this case, the object detection result for the captured image based on the object shadow detection result for each divided image and the object detection result for the captured image). At this time, if the synthesized object detection results include objects with approximately the same position, for example, the one with a higher score indicating objectness, which is used during object detection, is selected. Alternatively, both can be output. Also, the detection result synthesis means 43 can output the object detection result for the captured image as the object detection result for the captured image (in this case, "the object detection result for the captured image based on the object detection result for the captured image").
[0015] Here, the object detection accuracy when using the current image processing means as the object detection means and performing object detection processing on the captured image, and when performing object detection processing on the divided images obtained by dividing the captured image, will be described with reference to FIG. 3. (M1) shows an image including two people existing in the distance. (M2) shows an image obtained by reducing the image (M1) to a size corresponding to the captured image (X) used for object detection processing. Note that by using the zoom function of the imaging means 50, the size of the image (M2) in the captured image (X) changes. (M3) shows an image obtained by reducing the image (M1) to a size corresponding to the divided images obtained by dividing the captured image (X) (in FIG. 3, a 4 - divided image obtained by dividing equally in both the vertical and horizontal directions). (N1) shows the image (processing target image) of the region corresponding to the image (M2) in the captured image (X). (N2) shows the image (processing target image) of the region corresponding to the image (M3) in the divided image. When performing object detection processing on the processing target image (N1) of the captured image (X) using the current image processing means, the object image (M2) could not be detected. On the other hand, when performing object detection processing on the processing target image (N2) of the divided image, the object image (M3) could be detected.
[0016] As described above, by performing object detection processing by the second detection means 40 (second object detection means 41) on the divided image obtained by dividing the captured image, it becomes possible to detect an object in a small image included in the captured image that could not be detected by the current image processing means. That is, the object detection accuracy can be improved. In the experiment, when using the current image processing means, the minimum detection size of the object in the captured image was about [10×30 pixels], but when using the second detection means 40, the minimum detection size of the object in the 4-divided image was about [7×15 pixels].
[0017] Next, the operation of the second detection means 40 will be described. A first embodiment of the operation of the second detection means 40 will be described with reference to FIG. 4. In the first embodiment, the image dividing means 42 divides the captured image (X) into four (two equally spaced in the horizontal direction × two equally spaced in the vertical direction) divided images (a) to (d) by four dividing lines (one horizontal dividing line and one vertical dividing line). Note that the positions of the respective divided images (a) to (d) in the captured image (X) (for example, the corner positions of the respective divided images (a) to (d) on the coordinates of the captured image (X)) are stored in the storage means 60. In the first embodiment, the captured image (X) is divided into 2 squared (2 2 )(two equally spaced in each of the vertical and horizontal directions) divided images (a) to (d). That is, the aspect ratios of the divided images (a) to (d) are equal to (including "substantially equal to") the aspect ratio of the captured image (X). Therefore, when performing object detection processing on each of the divided images (a) to (d) by the second object detection means 41, there is no distortion due to image scaling, and the object detection performance is not affected.
[0018] The second object detection means 41 executes object detection processing on each of the four divided images (a) to (d) and outputs an object detection result for each of the four divided images (a) to (d). Also, the second object detection means 41 executes object detection processing on the captured image (X) and outputs an object detection result for the captured image (X). The detection result combining means 43 combines the object detection results for each of the four divided images (a) to (d) and the object detection result for the captured image (X), and outputs it as the object detection result of the captured image (X) (the object detection result of the captured image (X) based on the object detection results for each of the divided images (a) to (d) and the object detection result for the captured image (X)). For example, the object detection results for each of the four divided images (a) to (d) are combined with the object detection result for the captured image (X) in a state where the position information of the objects included in the object detection results for each of the four divided images (a) to (d) is converted into the position information in the captured image (X).
[0019] In the first embodiment, by executing object detection processing on the four four-divided images (a) to (d) obtained by dividing the captured image (X) into four, it is possible to detect an object that cannot be detected by the object detection processing for the captured image (X). Thereby, the detection accuracy of the object can be improved. When the captured image (X) is divided into four four-divided images (a) to (d), there is a possibility that an object existing in the boundary portion of each of the four-divided images (a) to (d), for example, an object straddling the boundary portion of the four-divided images (a) to (d) cannot be detected. For example, as shown in FIG. 5, an object (P) straddling the four-divided images (a) and (b) may not be detected by the object detection processing for the four-divided images (a) and (b). In the first embodiment, by performing object detection processing on the captured image (X), object detection processing for the captured image (X) is output. Then, the object detection result for the captured image (X) and the object detection results for each of the four divided images (a) to (d) are combined. As a result, objects existing in the boundary portions of each of the four divided images (a) to (d) that cannot be detected by the object detection processing for each of the four divided images (a) to (d) can be detected by the object detection processing for the captured image (X).
[0020] A second embodiment of the operation of the second detection means 40 will be described with reference to FIG. 5. In the second embodiment, the image dividing means 42 divides the captured image (X) into four divided images (a) to (d) by four dividing lines and also divides it into nine divided images (A) to (I) (three horizontal dividing lines and three vertical dividing lines, a total of nine, equally spaced horizontally by three and equally spaced vertically by three) by nine dividing lines. Note that the positions of each of the four divided images (a) to (d) and each of the nine divided images (A) to (I) in the captured image (X) (for example, the positions of the corner portions of each of the divided images (a) to (d), (A) to (I) on the coordinates of the captured image (X)) are stored in the storage means 60. In the second embodiment, the captured image (X) is divided into 2 squared (2 2 )(two equally spaced in each of the vertical and horizontal directions) four divided images (a) to (d) and 3 squared (3 2 )(three equally spaced in each of the vertical and horizontal directions) nine divided images (A) to (I). That is, the aspect ratios of the four divided images (a) to (d) and the nine divided images (A) to (I) are equal to (including "substantially equal to") the aspect ratio of the captured image. Therefore, when the second object detection means 41 performs object detection processing on the four divided images (a) to (d) and the nine divided images (A) to (I), there is no distortion due to image scale change and no impact on object detection performance.
[0021] The second object detection means 41 executes object detection processing on each of the four divided images (a) to (d) and each of the nine divided images (A) to (I), and outputs object detection results for each of the four divided images (a) to (d) and each of the nine divided images (A) to (I). Also, the second object detection means 41 executes object detection processing on the captured image (X), and outputs an object detection result for the captured image (X). The detection result combining means 43 combines the object detection results for each of the four divided images (a) to (d) and each of the nine divided images (A) to (I) with the object detection result for the captured image (X), and outputs the object detection result for the captured image (X) (the object detection result for the captured image (X) based on the object detection results for each of the divided images (a) to (d), (A) to (I) and the object detection result for the captured image (X)). For example, the object detection results for each of the four divided images (a) to (d) and each of the nine divided images (A) to (I) are combined with the object detection result for the captured image (X) in a state where the position information of the objects included in the object detection results for each of the four divided images (a) to (d) and each of the nine divided images (A) to (I) is converted into the position information in the captured image (X). The combining process of the object detection results by the detection result combining means 43 can use the same method as the combining process in the first embodiment.
[0022] In the second embodiment, by executing object detection processing on the four four-divided images (a) to (d) obtained by dividing the captured image (X) into four parts and the nine nine-divided images (A) to (I) obtained by dividing the captured image (X) into nine parts, it is possible to detect objects that cannot be detected by the object detection processing for the captured image (X). Thereby, the detection accuracy of the object can be improved. Also, the captured image (X) is divided into four parts into four four-divided images (a) to (d) which is 2 to the power of 2 (2 2 ) which is an even number, and divided into nine parts into nine nine-divided images (A) to (I) which is 3 to the power of 3 (3 2 ) which is an odd number. Thereby, the boundary portions (vertical boundary lines, horizontal boundary lines) of the four-divided images (a) to (d) and the boundary portions (vertical boundary lines, horizontal boundary lines) of the nine-divided images (A) to (I) intersect, but do not overlap parallel to each other. Therefore, the decrease in object detection accuracy at the boundary portions of each of the four divided images (a) to (d) (for example, an object existing across the boundary portion cannot be detected) can be compensated for by the object detection results of each of the nine divided images (A) to (I). For example, as shown in FIG. 5, an object (P) existing across the four divided images (a) and (b) may not be detected by the object detection process for the four divided images (a) and (b), but can be detected by the object detection process for the nine divided image (B). Similarly, the decrease in object detection accuracy at the boundary portions of each of the nine divided images (A) to (I) can be compensated for by the object detection results of each of the four divided images (a) to (d). Furthermore, it can also be compensated for by the object detection process for the captured image (X). Therefore, the detection accuracy of the object can be further improved.
[0023] A third embodiment of the operation of the second detection means 40 will be described with reference to FIG. 6. In the third embodiment, the image dividing means 42 divides the captured image (X) into four four-divided images (a) to (d) by four dividing lines, nine nine-divided images (A) to (I) by nine dividing lines, and 16 (4 equally spaced in the horizontal direction × 4 equally spaced in the vertical direction) 16-divided images (1) to (16) by 16 dividing lines (4 horizontal dividing lines and 4 vertical dividing lines). Note that the positions of each of the four divided images (a) to (d), each of the nine divided images (A) to (I), and each of the 16 divided images (1) to (16) in the captured image (X) (for example, the positions of the corner portions of each of the divided images (a) to (d), (A) to (I), (1) to (16) on the coordinates of the captured image (X)) are stored in the storage means 60. In the third embodiment, the captured image (X) is composed of 2 squared (2 2 )(2 equally spaced in each of the vertical and horizontal directions) four-divided images (a) to (d), 3 squared (3 2 )(3 equally spaced in each of the vertical and horizontal directions) nine-divided images (A) to (I), and 4 squared (4 2)(It is) divided into 16 divided images (1) to (16) with 4 images at equal intervals in both the vertical and horizontal directions. That is, the aspect ratios of the 4-divided images (a) to (c), the 9-divided images (A) to (I), and the 16-divided images (1) to (16) are equal to (including "substantially equal to") the aspect ratio of the captured image. For this reason, when the second object detection means 41 executes object detection processing on the 4-divided images (a) to (d), the 9-divided images (A) to (I), and the 16-divided images (1) to (16), there is no distortion due to image scaling, and there is no impact on object detection performance.
[0024] The second object detection means 41 executes object detection processing on each of the 4-divided images (a) to (d), each of the 9-divided images (A) to (I), and each of the 16-divided images (1) to (16), and outputs object detection results for each of the 4-divided images (a) to (d), each of the 9-divided images (A) to (I), and each of the 16-divided images (1) to (16). Also, the second object detection means 41 executes object detection processing on the captured image (X) and outputs an object detection result for the captured image (X). The detection result combining means 43 combines the object detection results for each of the 4-divided images (a) to (d), the object detection results for each of the 9-divided images (A) to (I), and the object detection results for each of the 16-divided images (1) to (16) with the object detection result for the captured image (X), and outputs it as the object detection result for the captured image (X) (the object detection result for the captured image (X) based on the object detection results for each of the divided images (a) to (d), (A) to (I), (1) to (14) and the object detection result for the captured image (X)). The combining process of the object detection results by the detection result combining means 43 can use the same method as the combining process in the first embodiment and the second embodiment.
[0025] In the third embodiment, by executing object detection processing on the 4 four-divided images (a) to (d) obtained by dividing the captured image (X) into four, the 9 nine-divided images (A) to (I) obtained by dividing it into nine, and the 16 sixteen-divided images (1) to (16) obtained by dividing it into sixteen, it is possible to detect objects that cannot be detected in the object detection processing for the captured image (X). This makes it possible to improve the detection accuracy of the object. In addition, the captured image (X) is divided into four parts by 2 to the power of 2 (2 2 ) and into 16 parts by 4 to the power of 4 (4 2 ), and is divided into nine parts by 3 to the power of 3 (3 2 ), which is odd. As a result, a part of the boundary between the four-divided images (a) to (d) and the 16-divided images (1) to (16) overlaps in parallel, but the boundary between the four-divided images (a) to (d) and the boundary between the 16-divided images (1) to (16) do not overlap in parallel with the boundary between the nine-divided images (A) to (I). Therefore, the decrease in the object detection accuracy at the boundary of each of the nine-divided images (A) to (I) (for example, an object existing across the boundary cannot be detected) can be compensated for by the object detection results of each of the four-divided images (a) to (d) and each of the 16-divided images (1) to (16). Similarly, the decrease in the object detection accuracy at the boundary of each of the four-divided images (a) to (d) and each of the 16-divided images (1) to (16) can be compensated for by the object detection results of each of the nine-divided images (A) to (I). Furthermore, it can also be compensated for by the object detection result of the captured image (X). Therefore, the detection accuracy of the object can be further improved.
[0026] Next, the first detection means 30 will be described. The first detection means 30 includes a differential image creation means 31, a moving body detection means 32, and a first object detection means 33.
[0027] The differential image creation means 31 sequentially creates differential image information indicating a differential image of the captured image represented by two pieces of image information output from the imaging means 50 at different times. The differential image is an image obtained by removing a common image (background image) from each of the two captured images. That is, the differential image shows an image of the moving body. As a method for creating the differential image information, various known methods can be used. For example, the background subtraction method of OpenCV can be used.
[0028] An example of a method for creating differential image information will be described with reference to FIGS. 7 and 8. Note that FIG. 7 shows a captured image (X[t−1]) represented by image information output from the imaging means 50 at time [t−1], which is one timing before time [t]. The captured image (X[t−1]) includes a background image (still image) (M[t−1]) and images of moving bodies (Y1[t−1]) and (Y2[t−1]). FIG. 8 shows a captured image X[t]) represented by image information output from the imaging means 50 at time [t]. The captured image (X[t]) includes a background image (still image) (M[t]) and images of moving bodies (Y1[t]) and (Y2[t]). The positions of the moving bodies (Y1[t]) and (Y2[t]) included in the captured image (X[t]) are different from the positions of the moving bodies (Y1[t−1]) and (Y2[t−1]) included in the captured image (X[t−1]). That is, the moving bodies (Y1) and (Y2) have moved between time [t−1] and time [t].
[0029] The differential image is created, for example, by comparing the state (such as brightness) of each pixel in the captured image (X[t]) with the state (such as brightness) of the corresponding pixel in the captured image (X[t−1]) and extracting pixels that are determined to be different. Specifically, a differential image is created that includes the images of the moving bodies (Y1[t]) and (Y2[t]) included in the captured image (X[t]) shown in FIG. 8 and the background image at the positions corresponding to the moving bodies (Y1[t−1]) and (Y2[t−1]).
[0030] The moving body detection means 32 detects the positions of the moving bodies (Y1[t]) and (Y2[t]) (included in the captured image (X[t])) at time [t] based on the differential image information created by the differential image creation means 31. Also, the positions of the moving bodies (Y1[t−1]) and (Y2[t−1]) (included in the captured image (X[t−1])) at time [t−1] are detected. Note that (Y1s[t]), (Y2s[t]), (Y1s[t-1]), and (Y2s[t-1]) respectively indicate the sizes of the moving objects (Y1[t]), (Y2[t]), (Y1[t-1]), and (Y2[t-1]). The size of the moving object is determined, for example, by the number of pixels. That is, it is determined by the number of pixels (number of vertical pixels × number of horizontal pixels) in the rectangular or square region circumscribing the image of the moving object. Then, the distance (moving distance) and direction (moving direction) between the positions of the moving objects (Y1[t-1]) and (Y2[t-1]) at time point [t-1] and the positions of the moving objects (Y1[t]) and (Y2[t]) at time point [t] are detected. As methods for detecting the distance and direction, various known methods can be used. For example, the "Optical Flow" method can be used. In FIG. 8, the distance and direction between the position of the moving object (Y1[t-1]) and the position of (Y1[t]) are indicated by the movement vector [Y1(Wt)], and the distance and direction between the position of the moving object (Y2[t-1]) and the position of (Y2[t]) are indicated by the movement vector [Y2(Wt)]. Furthermore, based on the movement vectors at each time point for each moving object, the average moving speed of each within the first set period and the movement trajectory of each moving object within the second set period are detected. The first set period and the second set period are appropriately set corresponding to the object to be detected. The movement vector may be a vector indicating movement on a two-dimensional plane (for example, a two-dimensional plane including the left-right direction and the up-down direction of the captured image), but preferably, a vector indicating movement in a three-dimensional space (for example, a three-dimensional space including the left-right direction, the up-down direction, and the front-back direction of the captured image) is used. Moving body Preferably, the moving object detection means 32 is configured to detect a moving object existing within the monitoring area included in the captured image. The monitoring area is defined, for example, by setting the positions of the boundary portions of the monitoring area in the captured image.
[0031] The first object detection means 33 detects whether the moving objects (Y1) and (Y2) are the objects to be detected (hereinafter referred to as "target objects"). As a method for detecting whether the moving bodies (Y1) and (Y2) are target objects, various methods can be used. In this embodiment, it is detected whether the moving bodies (Y1) and (Y2) are target objects based on whether the moving speeds and moving trajectories of the moving bodies (Y1) and (Y2) satisfy the set conditions set corresponding to the target object. As set conditions regarding the moving speed and moving trajectory for detecting that the moving body is a target object, set conditions are uniquely used for the target object. As the moving speed and moving trajectory, the moving speed and moving trajectory on a two-dimensional plane (for example, a two-dimensional space including the left-right direction and the up-down direction of the captured image) may be used, but preferably, the moving speed and moving trajectory in a three-dimensional space (for example, a three-dimensional space including the left-right direction, the up-down direction, and the front-back direction of the captured image) are used. For example, when the target object is a person, the following set conditions are used. (1) The average moving speed within the first set period is within the range of the lower limit value and the upper limit value. As the lower limit value, for example, it is set to the speed at which a person walks slowly, and as the upper limit value, for example, it is set to the speed at which a person runs fast. As the first set period, an appropriate period for determining the average moving speed of a person is set. The average moving speed of a person is different from the average moving speeds of moving bodies such as birds and vehicles. (2) The moving trajectory within the second set period forms a trajectory of a continuous predetermined shape. The moving trajectory of a person is different from the moving trajectories of moving bodies such as birds and vehicles. As the second set period, an appropriate period for determining the moving trajectory of a person is set. The set conditions can be set by collecting data related to the movement of a person and setting set conditions specific to the movement of a person based on the collected data. Alternatively, it can also be set by learning using deep learning.
[0032] As described above, by detecting the movement of the moving object and detecting the object to be detected based on the movement of the moving object, it has become possible to detect an object in a small image included in the captured image that cannot be detected by the second detection means 40 (the second object detection means 41). That is, the detection accuracy of the object can be improved. In the experiment, the minimum detection size of the object in the captured image (the minimum size of (Y1s[t]), (Y2s[t]), (Y1s[t - 1]), and (Y2s[t - 1]) shown in FIG. 8) was approximately [4×7 pixels], which is smaller than the minimum detection size of the object in the 4-divided image [7×15 pixels] when the second detection means 40 was used.
[0033] An example of combining the object detection result by the first detection means 30 and the object detection result by the second detection means 40 is shown in FIG. 9. FIG. 9 shows a captured image (X) of the riverbed downstream of the dam, which is the monitoring area, captured by the imaging means 50. FIG. 9 shows a state where moving objects (1) to (4) exist in the monitoring area. The moving objects (1) to (4) are images of a person who is the target object. In the captured image (X), the images of the moving objects (1) and (2) are large, and the images of the moving objects (3) and (4) are small. Since the images of the moving objects (1) and (2) are large, the moving objects (1) and (2) can be detected by the object detection process (AI detection) using deep learning by the second detection means 40. Since the images of the moving objects (3) and (4) are small, the moving objects (3) and (4) cannot be detected by the object detection process (AI detection) by the second detection means 40, but the moving objects (3) and (4) can be detected by the object detection process (detection of movement) based on the movement of the moving object by the first detection means 30.
[0034] As described above, in the object detection process by the second detection means 40 (second object detection means 41), various objects can be detected with higher accuracy compared to the object detection process by the first detection means 30 (first object detection means 33). On the other hand, in the object detection process by the first detection means 30 (first object detection means 33), objects with a small image size can be detected compared to the object detection process by the second detection means 40 (second object detection means 41). Therefore, by combining the object detection process by the first detection means 30 (first object detection means 33) and the object detection process by the second detection means 40 (second object detection means 41), the detection accuracy of the object can be improved.
[0035] In the first embodiment, first, the object detection process by the second detection means 40 (second object detection means 41) is executed. And when the object to be detected cannot be detected by the object detection process by the second detection means 40, the object detection process by the first detection means 30 (first object detection means 33) is executed. In the first embodiment, the processing burdens of the first detection means 30 (first object detection means 33) and the second detection means 40 (second object detection means 41) can be reduced. In the first embodiment, if the object could not be detected by the object detection process by the second detection means 40 but the object to be detected is detected by the object detection process by the first detection means 30, it can also be configured to further perform a notification to prompt the staff to confirm the object. Alternatively, it can be configured to use the zoom function of the imaging means 50 to enlarge the captured image and execute the object detection process by the second detection means 40 (second object detection means 41) based on the image information indicating the enlarged captured image.
[0036] In the second embodiment, the object detection process by the first detection means 30 (first object detection means 33) and the object detection process by the second detection means 40 (second object detection means 41) are executed in parallel (simultaneously). In the second embodiment, an object can be detected in a short time. In the second embodiment, when an object is detected by the object detection process of the first detection means 30, further, similarly to the first embodiment, it is configured to perform notification for prompting an attendant to confirm the object, or to use the zoom function of the imaging means 50 to enlarge the captured image, and based on the image information indicating the enlarged captured image, it is also possible to configure to execute the object detection process by the second detection means 40.
[0037] Also, when the image of the object included in the captured image is large, it can be detected by the object detection process by the second detection means 40 (the second object detection means 41). On the other hand, when the image of the object included in the captured image is small, it cannot be detected by the object detection process by the second detection means 40 (the second object detection means 41), but it can be detected by the object detection process by the first detection means 30 (the first object detection means 33). Therefore, when the image of the object included in the captured image is small, it can also be configured to detect the object included in the captured image by executing the object detection process only by the first detection means 30 (the first object detection means 33). That is, the present invention can also be configured only by the first detection means 30 (the first object detection means 33).
[0038] In the above embodiments, as the divided images, a divided image group including a four-divided image obtained by dividing the captured image into 2 to the power of 2 (2 2 ) pieces (divided into two equal parts in the vertical and horizontal directions), a four-divided image obtained by dividing the captured image into 2 to the power of 2 (2 2 ) pieces, and a nine-divided image obtained by dividing the captured image into 3 to the power of 2 (3 2 ) pieces (divided into three equal parts in the vertical and horizontal directions), a divided image group including a four-divided image obtained by dividing the captured image into 2 to the power of 2 (2 2 ) pieces, a nine-divided image obtained by dividing the captured image into 3 to the power of 2 (3 2 ) pieces, and a sixteen-divided image obtained by dividing the captured image into 4 to the power of 2 (4 2 ) pieces (divided into four equal parts in the vertical and horizontal directions) are used. However, the divided images or combinations of divided images constituting the divided image group are not limited to this. The object detection results for each divided image and the object detection results for the captured image are combined and output as the object detection results for the captured image. However, it is also possible to configure it to combine the object detection results for the divided images and output them as the object detection results for the captured image. One group of divided images is used. However, it is also possible to use a plurality of groups of divided images, perform object detection processing on the divided images that make up one selected group of divided images, and if no object is included in the object detection results, select a different group of divided images and configure it to perform object detection processing on the divided images that make up the selected group of divided images. The repetition of object detection processing for different groups of divided images can be terminated at an appropriate timing. For example, it can be terminated when the number of groups of divided images for which object detection processing has been performed reaches a set value or when a set time has elapsed since the start of the object detection processing. When the object detection processing by the second object detection means 41 is to be performed for the purpose of detecting the presence or absence of an object (the presence of at least one object), it can be configured as follows. The second object detection means 41 performs object detection processing on the captured image, and when no object is included in the object detection results for the captured image, it performs object detection processing on each divided image. The detection result combining means 43 outputs the object detection results for the captured image as the object detection results for the captured image when an object is included in the object detection results for the captured image, and when no object is included in the object detection results for the captured image, it combines the object detection results for each divided image and outputs them as the object detection results for the captured image. Note that when no object is included in the object detection results for each divided image, it is also possible to configure it to perform object detection processing on different numbers of each divided image, combine the object detection results for different numbers of each divided image, and output them as the object detection results for the captured image. The repetition of object detection processing for different numbers of each divided image can be terminated at the same timing as, for example, the timing for terminating the repetition of object detection processing for different groups of divided images described above. In this case, the number of times of object detection processing by the second object detection means 41 can be reduced.
[0039] The present invention can also be configured as follows. “(Aspect 1) An object detection device according to any one of claims 2 to 4, wherein the image segmentation means divides the captured image represented by the image information output from the imaging means into at least a first number of first divided images and a second number of second divided images, and the first number and the second number are set such that a boundary portion of the first divided image and a boundary portion of the second divided image do not overlap in parallel, characterized in that it is an object detection device.” In this aspect, it is allowed that a boundary portion of the first divided image and a boundary portion of the second divided image intersect. Thereby, for example, an object existing across a boundary portion of one divided image that cannot be detected by object detection for one divided image can be detected by object detection for the other divided image. As the first number and the second number, appropriate numbers can be set. The types of divided images are not limited to two types, namely, the first number of first divided images and the second number of second divided images. In this aspect, a decrease in detection accuracy of an object at a boundary portion of one of the first divided image and the second divided image can be compensated for by a detection result of an object for the other divided image. Also, “(Aspect 2) An object detection device according to any one of claims 2 to 4 and Aspect 1, wherein the image segmentation means divides the captured image represented by the image information output from the imaging means into at least a first odd number of squared first divided images and a first even number of squared second divided images, characterized in that it is an object detection device.” Preferably, the captured image is divided at equal intervals with the same number of divisions (odd or even) in the vertical and horizontal directions. The types of divided images are not limited to two types, namely, the first odd number of squared first divided images and the first even number of squared second divided images. In this aspect, as the first divided image and the second divided image, divided images having an aspect ratio substantially the same as the aspect ratio of the captured image can be used. Therefore, even when object detection processing is performed on the divided images using the image processing means used in a normal object detection device, there is no distortion due to image scaling change, and the object detection performance is not affected. Also, "(Aspect 3) An object detection device according to any one of claims 2 to 4 and aspects 1 and 2, wherein the image dividing means can divide the captured image indicated by the image information output from the imaging means into a plurality of divided image groups including at least one type of divided image and having different total numbers of divided images, and the detection result combining means detects an object that is the detection target included in each divided image based on the divided image information indicating each divided image constituting one divided image group, and when the object that is the detection target is included in any of the detection results for each divided image, combines the object detection results for each divided image and outputs them as the object detection result for the captured image indicated by the image information output from the imaging means, and when the object that is the detection target is not included in the object detection results for each divided image, performs the same processing on different divided image groups. An object detection device characterized by that. The repetition of the object detection processing for different divided image groups can be terminated at an appropriate timing. For example, it can be terminated when the number of divided image groups for which the object detection processing has been executed reaches a set value or when a set time has elapsed since the start of the object detection processing. This aspect can preferably be used when detecting that at least one object is included in the captured image. In this aspect, since the object detection processing can be terminated when it is detected that an object exists in the captured image, the processing load on the second detection means can be reduced. Also, "(Aspect 4) An object detection device according to any one of claims 2 to 4 and aspects 1 and 2, The second object detection means uses deep learning to detect the object to be detected included in the captured image indicated by the image information output from the imaging means. The detection result combining means combines the object shadow detection result for each divided image and the object detection result for the captured image, and outputs the result as the object detection result for the captured image. An object detection apparatus can be configured as described above. As a method for combining the object detection result for each divided image and the object detection result for the captured image, for example, a method of converting the position information of the object in the divided image into the position information in the captured image can be used. Also, when the same object with the same category and position is included in a plurality of object detection results, for example, a method of selecting the object with a higher score indicating the likelihood of a human body, which is used in the object detection process, can be used. In this aspect, the detection accuracy of the object can be further improved. Also, “(Aspect 5) An object detection apparatus according to any one of claims 2 to 3 and Aspects 1 to 4, The second object detection means uses deep learning to detect the object to be detected included in the captured image. When the object to be detected is not included in the object detection result for the captured image, the object to be detected included in each divided image is detected. The detection result combining means outputs the object detection result for the captured image as the object detection result for the captured image when the object to be detected is included in the object detection result for the captured image. When the object to be detected is not included in the object detection result for the captured image, the object detection results for each divided image are combined and output as the object detection result for the captured image. An object detection apparatus can be configured as described above. This aspect can preferably be used when detecting that at least one object exists in the captured image. This aspect can reduce the processing burden on the second detection means.
[0040] The present invention is not limited to the configurations described in the embodiments, and various modifications, additions, and deletions are possible. The difference image generation means, the moving body detection means, the first object detection means, the second object detection means, the image segmentation means, and the detection result synthesis means are not limited to the configurations described in the embodiments. The first detection means is not limited to the configuration described in the embodiment. The second detection means is not limited to the configuration described in the embodiment. Each configuration described in the embodiment can be used alone or in combination of a plurality of appropriately selected ones.
Explanation of Reference Numerals
[0041] 10 Object detection device 20 Processing means 30 First detection means 31 Difference image creation means 32 Moving body detection means 33 First object detection means 40 Second detection means 41 Second object detection means 42 Image segmentation means 43 Detection result synthesis means 50 Imaging means 60 Storage means 70 Input means 80 Output means
Claims
1. An object detection device comprising an imaging means, a first detection means, and a second detection means, wherein the first detection means includes a differential image creation means, a moving object detection means, and a first object detection means, the imaging means outputs image information indicating an imaged image, the differential image creation means creates differential image information indicating a differential image of the imaged images indicated by two pieces of image information output from the imaging means at different times, the moving object detection means detects the moving speed and moving trajectory of a moving object based on the differential image information created by the differential image creation means, the first object detection means detects the detected moving object as the object to be detected when the moving speed and moving trajectory of the moving object detected by the moving object detection means satisfy the set conditions set corresponding to the object to be detected, the second detection means includes a second object detection means, an image segmentation means, and a detection result synthesis means, the image segmentation means divides the imaged image indicated by the image information output from the imaging means into a plurality of divided images and outputs divided image information indicating each divided image, the second object detection means detects the object to be detected included in each divided image indicated by each divided image information output from the image segmentation means, and the detection result synthesis means synthesizes the object detection results for each divided image by the second object detection means and outputs the result as the object detection result for the imaged image indicated by the image information output from the imaging means.
2. The object detection device according to claim 1, wherein the image segmentation means divides the imaged image indicated by the image information output from the imaging means into at least a first number of first divided images and a second number of second divided images, and the first number and the second number are set such that the boundary portions of the first divided images and the boundary portions of the second divided images do not overlap in parallel.
3. The object detection device according to claim 1 or 2, wherein the first detection means is configured to operate when the object to be detected is not detected by the second detection means.
4. The object detection device according to claim 1 or 2, wherein the first detection means and the second detection means are configured to operate in parallel.
Citation Information
Patent Citations
Image recognition device
JP2006155167A
Fire monitoring system
JP2018101416A
Object detection device, object detection method and program
JP2019021001A
Program, information processing apparatus, and image processing system
JP2019057849A