Image processing apparatus, image processing method, and program

The image processing apparatus addresses the issue of wasted regions and decreased detection rates in object detection by generating an input image with a second image that is an enlarged or reduced partial region of the first image, thereby improving detection rates and expanding the detection area.

JP7692287B2Active Publication Date: 2025-06-13KOITO ELECTRIC IND LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2021089992
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-28
Publication Date
2025-06-13
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

In object detection processing using deep learning, images that do not match a predetermined size require resizing, leading to wasted regions and decreased detection rates due to reduced image sizes.

Method used

An image processing apparatus that includes a preprocessing unit generating an input image of a predetermined size by combining a first image and a second image, where the second image is an enlarged or reduced image of a partial region of the first image, thereby expanding the detection area and improving object detection rates.

Benefits of technology

The proposed solution enhances object detection rates while expanding the detection area by effectively utilizing regions that would otherwise be wasted, thereby improving the efficiency of object detection processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692287000001
    Figure 0007692287000001
  • Figure 0007692287000002
    Figure 0007692287000002
  • Figure 0007692287000003
    Figure 0007692287000003
Patent Text Reader

Abstract

To provide an image processing device, image processing method, and program, which allow for improving the object detection rate while expanding a detection area.SOLUTION: An image processing device 100 is provided, comprising a preprocessing unit 12 and a detection unit 13. The preprocessing unit generates an input image of a predetermined size, including a first image and a second image representing a magnified or reduced image of a partial area of the first image. The detection unit detects a detection target object from the input image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program for detecting an object from an input image.

Background Art

[0002] A technique for detecting an object from an image using a machine learning model is widely known. For example, in Patent Document 1, using a pre-trained model, candidate regions for each of a plurality of parts of a target object are detected, and based on the detected candidate regions, a highly reliable part with relatively high reliability and a non-highly reliable part with relatively low reliability are selected from the plurality of parts, and the non-highly reliable parts are rearranged based on the highly reliable parts, and a technique for detecting the positions of a plurality of parts of the target object is disclosed.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In object detection processing of an image using deep learning, which is a type of machine learning, in order to facilitate the detection of an object, the input image is set to a predetermined size (for example, a square). Therefore, for an input image that is not of the predetermined size, it is necessary to adjust the sizes of the long side and the short side of the image to the predetermined size. For example, when the image size of the input image does not match the predetermined size, the image size is often converted (mainly reduced) while maintaining the image ratio. In this case, regions (blank portions) that do not match the aspect ratio are filled in black, etc., resulting in wasted regions that are not substantially used for the detection process. Also, when the input image is reduced to the predetermined size, there is a problem that the detection rate of the object in the image decreases.

[0005] In view of the above circumstances, an object of the present invention is to provide an image processing apparatus, an image processing method, and a program that can improve the detection rate of an object while expanding a detection area.

Means for Solving the Problems

[0006] An image processing apparatus according to an aspect of the present invention includes a preprocessing unit and an object detection unit. The preprocessing unit generates an input image of a predetermined size including a first image and a second image that is an enlarged image or a reduced image of a partial region of the first image. The detection unit detects an object to be detected from the input image.

[0007] In the above image processing apparatus, since it includes a preprocessing unit that generates an input image of a predetermined size including a first image and a second image that is an enlarged image or a reduced image of a partial region of the first image, it is possible to improve the detection rate of an object in the detection unit while expanding the detection area.

[0008] The preprocessing unit may have an image adjustment unit that reduces the original image of the first image so as to fit within the predetermined size. Thereby, it is possible to perform object detection using a deep learning model.

[0009] The preprocessing unit may further have an image synthesis unit that synthesizes the second image with a blank portion of the input image obtained when the first image is reduced so as to fit within the predetermined size.

[0010] The preprocessing unit may further have an image extraction unit that extracts an enlarged image of a preset region of the first image as the second image.

[0011] The image extraction unit may be configured to convert the resolution of the second image to the resolution of the original image.

[0012] The image extraction unit may be configured to extract, as the second image, enlarged images of a plurality of different regions of the first image, respectively.

[0013] The detection unit may be configured to detect the type and position of the object in the input image.

[0014] The detection unit may include a deep learning device.

[0015] An object detection method according to an aspect of the present invention generates an input image of a predetermined size including a first image and a second image that is an enlarged image of a partial region of the first image, and detects an object to be detected from the input image.

[0016] A program according to an aspect of the present invention includes steps of generating an input image of a predetermined size including a first image and a second image that is an enlarged or reduced image of a partial region of the first image, and detecting an object to be detected from the input image and causes a computer to execute them.

Advantages of the Invention

[0017] According to the present invention, it is possible to improve the object detection rate while expanding the detection region.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0020] [First Embodiment] FIG. 1 is a block diagram showing the configuration of an image processing apparatus 100 according to an embodiment of the present invention. The image processing apparatus 100 of the present embodiment is configured to detect an object to be detected from an image of a camera 1 that photographs a road such as an intersection, determine its type and position, and transmit the determination result to a control center 2 that manages traffic lights installed at the intersection or on the road, or a traffic light control unit 3 that controls the traffic lights.

[0021] In the present embodiment, the objects to be detected include, for example, vehicles traveling on the road, as well as pedestrians on the sidewalk or crosswalk. Vehicles include automobiles, motorcycles, light vehicles, etc. In addition to this, the detection target can be arbitrarily set according to the detection purpose.

[0022] The camera 1 is a video camera that photographs the shooting area at a predetermined frame rate (for example, 30 frames / second). The installation location of the camera 1 is not particularly limited and may be a traffic light, a utility pole near the shooting area, or the rooftop of a building.

[0023] [Image Processing Apparatus] The image processing apparatus 100 is typically composed of a computer including a CPU (Central Processing Unit), a memory, and the like. The image processing apparatus 100 includes a control unit 10, a storage unit 21, and a communication unit 22. Part of the CPU processing may be processed by a GPU.

[0024] The control unit 10 has an acquisition unit 11, a preprocessing unit 12, and a detection unit 13 as functional blocks of the CPU.

[0025] The acquisition unit 11 acquires image data of a shooting area from the camera 1 and stores it in the storage unit 21. The preprocessing unit 12 adjusts the size of the image acquired by the acquisition unit 11 to an image size that can be processed by the detection unit 13 as described later. The detection unit 13 detects an object (object) set in advance as a detection target from the input image whose size has been adjusted by the preprocessing unit 12. In the present embodiment, the detection unit 13 includes a deep learning device, and while using the learning model stored in the storage unit 21, determines the presence or absence of the detection target in the input image, and when the detection target exists, detects its type (such as a vehicle, a pedestrian, etc.) and position for each individual detection target.

[0026] The storage unit 21 is composed of a storage device such as a semiconductor memory or a hard disk drive. The storage unit 21 stores a program and arithmetic parameters for causing the control unit 10 to execute various functions described later, and further stores a learning model referred to during object detection processing in the detection unit 13. The storage unit 21 is not limited to being built into the image processing apparatus 100, and may be a storage device separate from the image processing apparatus 100, or may be a cloud server or the like that can be connected to the control device 10 via a network.

[0027] The communication unit 22 is composed of a communication module that transmits and receives information to and from the traffic control center 2 or the signal control unit 3. The communication method is not particularly limited, and may be wired or wireless.

[0028] Note that the control center 2 is a traffic control center that manages traffic lights belonging to the control area. The traffic light control unit 3 is, for example, a control device that comprehensively controls a plurality of traffic lights within an intersection. When the traffic lights to be controlled belong to the control area, the traffic light control unit 3 controls each traffic light based on an instruction from the control center 2. When the traffic lights to be controlled belong to a non-control area, the traffic light control unit 3 mainly controls each traffic light.

[0029] In the object detection process of an image using deep learning, in order to facilitate the detection of an object, as shown in FIG. 2(A), the input image is set to a predetermined size S. Therefore, for an input image that does not have the predetermined size S, it is necessary to adjust the sizes of the long side and the short side of the image to the predetermined size S. For example, when the image size of the input image does not match the predetermined size S, as shown in FIG. 2(B), the image size is often converted (mainly reduced) while maintaining the image ratio. However, regions (blank portions) that do not match the vertical and horizontal ratios are filled with black, etc., resulting in regions that are not substantially used for the detection process.

[0030] In addition, when there are a largely captured portion and a small captured portion in the image, the small captured portion becomes smaller than the size at which deep learning can easily detect it due to the above reduction process, resulting in a problem that the detection rate of the detection target decreases.

[0031] Therefore, in this embodiment, the image processing apparatus 100 is configured to be able to utilize deep learning without waste by utilizing the regions that are not originally used in the input image. Specifically, as shown in FIG. 3, the preprocessing unit 12 generates a first image G1 constituting the input image and a second image G2 that is an enlarged image of a partial region G1a thereof, and embeds the second image G2 in the unused region of the input image to improve the detection rate of the object. Hereinafter, the details of the preprocessing unit 12 will be described.

[0032] [Details of the Preprocessing Unit] As shown in FIG. 3, the preprocessing unit 12 generates an input image G of a predetermined size S including a first image G1 and a second image G2 which is an enlarged image of a partial region G1a of the first image G1. The first image G1 corresponds to the captured image of the camera 1 acquired by the acquisition unit 11. The region G1a extracted as the second image G2 is a pixel region preset by the user in the first image G1. The predetermined size S is the size of the image input to the detection unit 13 constituting the deep learning device, and is typically square.

[0033] In the present embodiment, the preprocessing unit 12 includes an image adjustment unit 121, an image extraction unit 122, and an image composition unit 123.

[0034] The image adjustment unit 121 is configured to reduce the original image (the captured image of the camera 1) of the first image G1 so as to fit within the predetermined size S. For example, when the captured image of the camera 1 is 1080 pixels in height and 1920 pixels in width (aspect ratio 9:16), without changing the image ratio, as shown in FIG. 3, it is processed into the first image G1 whose size is reduced so as to fit within the predetermined size S. The first image G1 is typically an image with reduced resolution by thinning out the pixel data of the captured image of the camera 1.

[0035] The image extraction unit 122 extracts an enlarged image of a preset region G1a of the first image G1 as the second image G2. The region G1a can be set arbitrarily and can be set arbitrarily according to the imaging target, the wide angle or magnification of the camera 1, etc. The number of regions extracted as the second image G2 is not limited to one and may be two or more. Also, the second image G2 preferably has a higher resolution than the resolution of the first image G1. For example, the resolution of the second image G2 is converted to the resolution of the original image of the first image G1.

[0036] The image synthesis unit 123 is configured to synthesize the second image G2 in the margin of the input image G obtained when the first image G1 is reduced to fit within a predetermined size S. The above margin corresponds to the unused area in FIG. 2(B) and corresponds to the part that is not originally used as a detection area in the detection unit 13. In the present embodiment, while effectively utilizing this margin, the margin is filled with an enlarged image (second image G2) of an area where it is difficult to detect an object at the image size of the first image G1. As a result, object detection can be performed at the two image sizes of the first image G1 and the second image G2, so that the detection rate of the object can be improved.

[0037] [Image processing method] Subsequently, a typical operation of the image processing apparatus 100 configured as described above will be described. FIG. 4 is a flowchart showing an example of the processing procedure executed in the control device 10.

[0038] The acquisition unit 11 acquires the captured image of the camera 1 (step 101). FIG. 5 shows an example of the captured image G0 of the camera 1. Here, as the captured image G0, a wide-angle camera image with a part of the area of a road intersection as the subject is shown. In this captured image G0, a plurality of vehicles V waiting for a signal at the intersection and a plurality of crosswalks C are captured.

[0039] Subsequently, the preprocessing unit 12 (image adjustment unit 121) adjusts the captured image G0 to a predetermined size S (see FIG. 2(A)) necessary for inputting it to the detection unit 13 (step 102). In the present embodiment, while maintaining the aspect ratio of the captured image G0 (for example, 9:16), as shown in FIG. 6, a first image G1 obtained by reducing the captured image G0 to fit within a predetermined size S is generated.

[0040] Subsequently, the preprocessing unit 12 (image extraction unit 122) extracts an enlarged image of a preset area G1a (see FIG. 6) of the first image G1 as the second image G2 (step 103). The position and size of the area G1a can be arbitrarily set. For example, when the purpose is to detect the number of vehicles V waiting for a signal, in the first image G1, the area G1a is set so that the rear part of the vehicle queue waiting for the signal can be extracted as the second image G2.

[0041] In particular, when a wide-angle lens is used for camera 1, for example, the farther the area from camera 1, the smaller the object will be photographed. If a detection target exists in that area, when the size adjustment process (step 102) of the captured image G0 is executed in the previous process, the detection target (for example, vehicle V1 in FIG. 5) existing in that area will become even smaller, and there is a possibility that appropriate detection processing cannot be performed in the object detection process by the detection unit 13. Therefore, in the present embodiment, by setting or designating such an area as an extraction area as the second image G2, the above-mentioned adverse effects associated with the reduction process of the captured image G0 are eliminated, and the detection rate of the object photographed at a distant position is improved.

[0042] Subsequently, the preprocessing unit (image synthesis unit 123) executes a process of synthesizing the enlarged image of the extracted area G1a as the second image G2 into the unused area of the input image (step 104).

[0043] The image synthesis unit 123 executes a process of converting the second image G2 to the resolution of the captured image G0, which is the original image of the first image G1. Thereby, since the second image G2 can be generated with higher precision, it becomes possible to appropriately perform the object detection process in the detection unit 13.

[0044] The image synthesis unit 123 further stores the coordinate information of the second image G2 synthesized into the unused area of the input image in the storage unit 21 in association with the coordinate information in the area G1a of the first image G1. Thereby, it becomes possible to specify the coordinate position of the object detected by the second image G2 in association with the coordinate position on the first image G1.

[0045] Subsequently, the detection unit 13 executes a detection process for an object to be detected based on the input image including the first image G1 and the second image G2 (step 105). The object detection process detects the presence or absence of an object in the input image (the first image G1 and the second image G2), and if the object exists, its type and position are detected respectively. For the object detection process, for example, a deep learning model using a convolutional neural network (CNN) is used.

[0046] Subsequently, the detection unit 13 executes a process of converting the detection position of the detected object, particularly the coordinate position of the object detected in the second image G2, into the coordinate system of the first image G1 (step 106). Thereby, for example, when an object (e.g., vehicle V1) that was not detected in the first image G1 is detected in the second image G2, the position of the detected object V1 can be easily grasped as the position on the first image G1.

[0047] Subsequently, the control unit 10 outputs the object detection result in the detection unit 13 to the control center 2 or the traffic signal control unit 3 (step 107). The control center 2 or the traffic signal control unit 3 switches the display data of the number of seconds for controlling the stages of the traffic signal corresponding to the imaging area of the camera 1 based on the output of the image processing apparatus 100. For example, when the number of vehicles V1 stopped waiting for a signal is equal to or more than a predetermined number, the number of seconds of the red display of the traffic signal is shortened, or when the number of stopped vehicles in the right turn lane is equal to or more than a predetermined number, the display time of the right turn arrow signal is lengthened, etc., and traffic signal control is executed. Thereby, the occurrence of congestion at the intersection can be suppressed.

[0048] By repeatedly executing the above processes each time an image is acquired from the camera 1, the traffic situation in the shooting area can be grasped in real time, so that appropriate traffic signal control can be realized.

[0049] Note that the object detection information output from the image processing apparatus 100 may be displayed as an image on a display device or the like. Further, based on the object detection information output from the image processing apparatus 100, driving support information that can be presented to vehicles passing through the intersection may be generated. Examples of the driving support information include various alert information for notifying the presence of pedestrians to a right-turning vehicle or the like passing through the crosswalk when a pedestrian is detected on the crosswalk C.

[0050] As described above, according to the present embodiment, since the recognition process of the image in the detection unit 13 can be executed by making full use of the image sizes processed by the deep learning, the detection area can be expanded, and thereby the object detection process of the captured image can be efficiently performed.

[0051] Further, since a plurality of images with different resolutions are synthesized into one input image, the object detection result can be obtained without executing the deep learning a plurality of times. In addition, since an image that cannot be detected by one image can be detected by the other image, the object detection rate can be improved.

[0052] Furthermore, according to the present embodiment, since the enlarged image of a partial region G1a of the first image G1 is extracted and synthesized as the second image G2, the extraction region can be freely set according to the detection purpose, the type of the imaging object, and the like.

[0053] <Modification Example> In the above description, the second image G2 is an enlarged image of a partial region G1a of the first image G1, but the present invention is not limited thereto, and the second image G2 may be a reduced image of a partial region G1a of the first image G1.

[0054] A case where the second image G2 is a reduced image of a partial region G1a of the first image G1 will be described. When the detection unit 13 determines the presence or absence of a detection target in the input image, an object detection process is performed using the learning model stored in the storage unit 21. Here, when the image size of the object to be detected that is learned by the learning model is larger than the image size of the object to be detected displayed in the first image G1, there is a problem that the detection rate of the detection target decreases because it becomes larger than the size at which deep learning can easily detect.

[0055] Therefore, the image processing apparatus 100 is configured to be able to improve the detection rate of the detection target by utilizing the region that is not originally used in the input image. Specifically, the preprocessing unit 12 generates a first image G1 that constitutes the input image and a second image G2 that is a reduced image of a partial region G1a thereof, and embeds the second image in the region that is not used in the input image. As a result, since it is reduced to a size at which deep learning can easily detect, it is possible to improve the detection rate of the object. Regarding the size of the second image G2, it is preferably approximately the same as the size of the object to be detected during learning by the learning model.

[0056] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments and can of course be variously modified.

[0057] For example, in the above embodiment, a partial region G1a of the first image G1 is extracted as the second image G2, but the number of regions to be extracted is not limited to one, and may be two or more as shown in FIG. 8, for example. In this case, in the same figure, the enlarged image of the region G1a is used as the second image G2a, and the enlarged image of the region G1b is used as the second image G2b, and they are respectively synthesized into the input image. The sizes and magnification ratios of the regions G1a and G1b can be arbitrarily set.

[0058] Also, in the above embodiment, the process of detecting an object (mainly a vehicle) from a captured image of a road intersection has been described as an example, but of course, the present invention is not limited to this, and the present invention is applicable to various monitoring systems in traffic facilities and commercial facilities.

Explanation of Symbols

[0059] 1…Camera 10…Control device 11…Acquisition unit 12…Pre - processing unit 13…Detection unit 100…Image processing device 121…Image adjustment unit 122…Image extraction unit 123…Image synthesis unit G0…Captured image (original image) G1…First image G2…Second image

Claims

1. A preprocessing unit that generates the input image such that a first image and a second image, which is an enlarged or reduced image of a partial region of the first image, are included in one input image of a predetermined size, a detection unit that detects an object to be detected from the input image An image processing apparatus comprising:

2. The image processing apparatus according to claim 1, wherein the preprocessing unit has an image adjustment unit that reduces the original image of the first image so as to fit within the predetermined size An image processing apparatus.

3. The image processing apparatus according to claim 2, wherein the second image has a resolution different from that of the first image An image processing apparatus.

4. The image processing apparatus according to claim 2 or 3, wherein the preprocessing unit further has an image synthesis unit that synthesizes the second image in a blank portion of the input image obtained when the first image is reduced so as to fit within the predetermined size An image processing apparatus.

5. The image processing apparatus according to any one of claims 2 to 4, wherein the preprocessing unit further has an image extraction unit that extracts an enlarged image of a preset region of the first image as the second image An image processing apparatus.

6. The image processing apparatus according to claim 5, wherein the image extraction unit converts the resolution of the second image to the resolution of the original image An image processing apparatus.

7. The image processing apparatus according to claim 5 or 6, wherein the image extraction unit extracts enlarged images of a plurality of different regions of the first image as the second image respectively An image processing apparatus.

8. The image processing apparatus according to any one of claims 1 to 7, wherein the detection unit detects the type and position of the object in the input image An image processing apparatus.

9. The image processing apparatus according to any one of claims 1 to 8, wherein the detection unit includes a deep learning device An image processing apparatus.

10. Generate an input image such that a first image and a second image, which is an enlarged image of a partial region of the first image, are included in one input image of a predetermined size, Detect an object to be detected from the input image An image processing method.

11. A step of generating an input image such that a first image and a second image, which is an enlarged image of a partial region of the first image, are included in one input image of a predetermined size, A program for causing a computer to execute a step of detecting an object to be detected from the input image.

Citation Information

Patent Citations

  • Camera used both for silver halide photographing and electronic image pickup

    JP2002162681A

  • Image processing device, imaging device, and image processing program

    JP2012114549A

  • Image processing apparatus, image processing method, and program

    JP2017016593A

  • Image processing apparatus, image processing method, program, and printing system

    JP2019134324A

  • Detection method, detection program, and detection device

    JP2020071615A