Image analyzer, image analysis system, image analysis method, and computer program

The image analysis device addresses inaccuracies in vehicle detection by integrating detection areas based on vehicle compartment information, rotation correction, and error suppression, enhancing detection accuracy in parking lot surveillance.

JP2025136631APending Publication Date: 2025-09-19CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024035340
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing vehicle detection methods using bounding boxes in parking lot surveillance face issues with erroneous merging or non-detection of vehicle areas due to variations in vehicle size and tilt, leading to inaccurate vehicle detection.

Method used

An image analysis device that integrates vehicle detection areas by calculating detection target areas based on vehicle compartment information, using a combination of manual and automatic setting, rotation correction, and error suppression techniques to enhance accuracy.

Benefits of technology

The device effectively suppresses erroneous detection or non-detection of vehicle areas, ensuring accurate integration and detection of vehicles in parking lots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025136631000001_ABST
    Figure 2025136631000001_ABST
Patent Text Reader

Abstract

To provide an image analyzer capable of suppressing erroneous detection or non-detection in a vehicle detection area.SOLUTION: An image analyzer includes: acquisition means that acquires an image at a parking lot; setting means that acquires, on the basis of the image, vehicle space information of the parking lot; calculation means that calculates, on the vehicle space information, a detection target area for vehicle detection; detection means that outputs, on the basis of the image, a vehicle detection area, which is an area including a vehicle; and integration means that integrates, on the basis of the detection target area, a plurality of vehicle detection areas.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image analysis device, an image analysis system, an image analysis method, a computer program, and the like. [Background technology]

[0002] Conventionally, there has been a need to detect whether or not a vehicle (e.g., an automobile) is parked in each parking area of ​​a parking lot in order to know the congestion level and vacant parking status of the parking lot. One method for detecting the presence or absence of a vehicle in each parking area is to install a vehicle detection device such as an infrared sensor to detect objects.

[0003] However, these vehicle detection devices generally have a limited detection range, and there is a risk that a huge number of vehicle detection devices will be installed in a large parking lot. On the other hand, another method for detecting the presence or absence of a vehicle in each parking area is to photograph the parking area with a camera and apply image processing to the photographed image to detect whether a vehicle is parked there.

[0004] This method has the advantage that by capturing images of multiple parking areas, it is possible to detect the presence or absence of vehicles in multiple parking areas using an image captured by a single camera. For example, Patent Document 1 describes a technology that determines the presence or absence of a vehicle for each pixel in an image and compares it with a pre-set vehicle compartment area to detect the presence or absence of a vehicle compartment.

[0005] However, the technology described in Patent Document 1 does not use a bounding box, but instead determines whether each pixel in the image is a vehicle, which results in a very high computational cost.On the other hand, for example, Non-Patent Document 1 discloses a method for detecting vehicles at high speed using a bounding box. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-206462

[0007] [Non-Patent Document 1] J. Redmon, A. Farhadi, “YOLO 9000:Better Faster Stronger”, Computer Vision and Pattern Recognition (CVPR) 2016. Summary of the Invention [Problem to be solved by the invention]

[0008] However, methods using bounding boxes have the problem that, depending on the size and tilt of the vehicle interior, multiple detected vehicle areas may be mistakenly merged into one, or conversely, a falsely detected area that should not exist between two vehicles may be output.

[0009] The present invention has been made in consideration of these problems, and one of its objects is to provide an image analysis device that can suppress erroneous detection or non-detection of a vehicle detection area. [Means for solving the problem]

[0010] In the image analysis device, an acquisition means for acquiring an image of the parking lot; a setting means for acquiring parking space information of the parking lot based on the image; a calculation means for calculating a detection target area for vehicle detection based on the vehicle compartment information; a detection means for outputting a vehicle detection area, which is an area including a vehicle, based on the image; an integration means for integrating the vehicle detection areas based on the detection target areas; It has. [Effects of the Invention]

[0011] According to the present invention, it is possible to provide an image analysis device that can suppress erroneous detection or non-detection of a vehicle detection area. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram showing an example of the configuration of an image analysis system 101 according to a first embodiment of the present invention. [Figure 2] 1 is a block diagram illustrating an example of the hardware configuration of an imaging device according to a first embodiment. [Figure 3] 1 is a functional block diagram showing an example of the functional configuration of an imaging device according to a first embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a server according to the first embodiment. [Figure 5] FIG. 2 is a functional block diagram showing an example of the functional configuration of a server according to the first embodiment. [Figure 6] FIG. 2 is a diagram showing an example of a setting screen of the server 130 according to the first embodiment. [Figure 7] FIG. 2 is a diagram showing an example of a parking lot information setting screen according to the first embodiment. [Figure 8] FIG. 8 is a diagram showing an example of the coordinates of the vehicle interior area set on the screen of FIG. 7. [Figure 9] FIG. 2 is a functional block diagram showing an example of the functional configuration of an analysis unit 505 according to the first embodiment. [Figure 10] 9 is a diagram showing an example of a detection area detected by a detection unit 901. FIG. [Figure 11] FIG. 2 is a diagram showing an example of a target region according to the first embodiment. [Figure 12] FIG. 4 is a diagram showing an example of allocation scores between a target region and a detection region according to the first embodiment. [Figure 13] 10 is a flowchart illustrating a flow of calculating a vehicle area executed by an area integration unit 903 in the server 130 according to the first embodiment. [Figure 14] 10 is a flowchart illustrating a flow of calculating a target region, which is executed by a target region calculation unit 902 according to the second embodiment. [Figure 15] FIG. 11 is a functional block diagram showing an example of the functional configuration of a server according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.

[0014] <Embodiment 1> 1 is a diagram showing an example of the configuration of an image analysis system 101 according to a first embodiment of the present invention. In the following, an example of a parking situation determination system will be described as an image analysis system, but this embodiment is not limited to this and can be applied to any system that analyzes images and outputs predetermined information.

[0015] The image analysis system 101 is composed of imaging devices 110a to 110d, a network 120, a server 130, etc. That is, the image analysis system of this embodiment is composed of at least one imaging device and an image analysis device, and analyzes the parking state in a parking lot having spaces where vehicles can be parked. Note that, hereinafter, the imaging devices 110a to 110d are collectively referred to as "imaging devices 110."

[0016] Although the following describes an example of a parking lot having a parking space where vehicles can be parked, the vehicle may be a mobile object such as a drone or a robot, the parking space may be a section for accommodating the mobile object, and the parking lot may be a place for accommodating the mobile object. This embodiment can also be applied to facilities for such mobile objects.

[0017] The imaging device 110 includes, for example, a network camera for acquiring images of the parking lot and transmitting them to an image analysis device. In this embodiment, the imaging device 110 is assumed to have a built-in computing device capable of processing the images, but is not limited to this. For example, the images may be processed by an external computer such as a PC (personal computer) connected to the imaging device 110, and the imaging device 110 may be a system formed by combining a camera and an external computer.

[0018] The server 130 is a computer such as a PC, functions as an image analysis device, and has the function of accepting operation inputs from a user and outputting information to the user (for example, displaying information).

[0019] The imaging device 110 and the server 130 are communicably connected via a network 120. The network 120 includes a plurality of routers, switches, cables, etc. that comply with a communication standard such as Ethernet (registered trademark).

[0020] The network 120 may be any network that enables communication between the image capture device 110 and the server 130, and may be constructed with any scale, configuration, and in accordance with any communication standard.

[0021] For example, the network 120 may be the Internet, a wired LAN (Local Area Network), a wireless LAN, a WAN (Wide Area Network), or the like.

[0022] The network 120 may be configured to enable communication using a communication protocol that complies with the ONVIF (Open Network Video Interface Forum) standard, for example. However, the network 120 is not limited to this, and may be configured to enable communication using a proprietary communication protocol or other communication protocols.

[0023] 2 is a block diagram showing an example of the hardware configuration of the imaging device according to embodiment 1. The imaging device 110 includes, for example, an imaging unit 201, an image processing unit 202, an arithmetic processing unit 203, a distribution unit 204, and the like.

[0024] The imaging unit 201 has a lens unit for focusing light and an imaging element such as a CMOS image sensor that converts the focused light into an analog signal. The lens unit has a zoom function for adjusting the angle of view and an aperture function for adjusting the amount of light.

[0025] The image sensor also has a gain function that adjusts sensitivity when converting light into an analog signal. These functions are adjusted based on setting values ​​supplied from the image processing unit 202. The analog signal acquired by the image capturing unit 201 is converted into a digital signal by an analog-to-digital conversion circuit and transferred to the image processing unit 202 as an image signal.

[0026] The image processing unit 202 includes an image processing engine and its peripheral devices, etc. The peripheral devices include, for example, a RAM (Random Access Memory) and drivers for each I / F (Interface).

[0027] The image processing unit 202 generates image data by performing image processing such as development, filtering, sensor correction, and noise removal on the image signal acquired from the imaging unit 201. The image processing unit 202 can also transmit setting values ​​to the lens unit and the imaging element to perform exposure adjustment so that an image with proper exposure can be acquired. The image data generated in the image processing unit 202 is transferred to the arithmetic processing unit 203.

[0028] The arithmetic processing unit 203 is composed of one or more computers such as a CPU or MPU, memories such as RAM or ROM, drivers for each I / F, etc. CPU is an abbreviation for Central Processing Unit, MPU is an abbreviation for Micro Processing Unit, and ROM is an abbreviation for Read Only Memory.

[0029] The distribution unit 204 includes a network distribution engine and peripheral devices such as a RAM and an ETH PHY module. The ETH PHY module is a module that executes processing of the physical (PHY) layer of Ethernet.

[0030] The distribution unit 204 converts the image data and processing result data acquired from the arithmetic processing unit 203 into a format that can be distributed over the network 120 , and outputs the converted data to the network 120 .

[0031] Fig. 3 is a functional block diagram showing an example of the functional configuration of the imaging device according to embodiment 1. Note that some of the functional blocks shown in Fig. 3 are realized by causing a CPU or the like serving as a computer included in the imaging device to execute a computer program stored in a memory serving as a storage medium.

[0032] However, some or all of these functions may be implemented by hardware, which may be a dedicated circuit (ASIC) or a processor (reconfigurable processor, DSP).

[0033] Furthermore, the functional blocks shown in Fig. 3 do not have to be built into the same housing, but may be configured as separate devices connected to each other via signal paths. The above explanation regarding Fig. 3 also applies to Figs. 5, 9, and 15.

[0034] The imaging device 110 includes, as its functions, an imaging control unit 301, a signal processing unit 302, a storage unit 303, a control unit 304, an analysis unit 305, a network communication unit 306, and the like.

[0035] The imaging control unit 301 executes control to capture an image of the surrounding environment via the imaging unit 201. The signal processing unit 302 performs predetermined processing on the image captured by the imaging control unit 301 to generate captured image data. Note that, hereinafter, this captured image data may be simply referred to as a "captured image" or a "captured image."

[0036] The signal processing unit 302 encodes, for example, an image captured by the imaging control unit 301. The signal processing unit 302 encodes a still image using an encoding method such as JPEG (Joint Photographic Experts Group).

[0037] Furthermore, the signal processing unit 302 encodes the moving image using an encoding method such as H.264 / MPEG-4AVC (hereinafter referred to as "H.264") or HEVC (High Efficiency Video Coding).

[0038] In addition, the signal processing unit 302 may encode the image using an encoding method selected by the user from a plurality of pre-set encoding methods, for example, via an operation unit (not shown) of the imaging device 110.

[0039] The storage unit 303 stores temporary data for various processes. The control unit 304 includes the aforementioned arithmetic processing unit 203, executes computer programs stored in memory, and controls the signal processing unit 302, storage unit 303, analysis unit 305, and network communication unit 306. The analysis unit 305 performs image analysis processing on the captured image. The network communication unit 306 communicates with the server 130 via the network 120.

[0040] Fig. 4 is a diagram showing an example of the hardware configuration of a server according to embodiment 1. For example, as shown in Fig. 4, the server 130 is configured with a processor 401 such as a CPU as a computer, memories such as a RAM 402 and a ROM 403, storage devices such as an HDD 404, and a communication I / F 405. The server 130 executes various functions by the processor 401 executing programs stored in the memory and storage device.

[0041] 5 is a functional block diagram showing an example of the functional configuration of the server according to embodiment 1. The server 130 includes, as its functional configuration, for example, a network communication unit 501, a control unit 502, a display unit 503, an operation unit 504, an analysis unit 505, a storage unit 506, and a setting processing unit 507.

[0042] The network communication unit 501 is connected to, for example, the network 120, and executes communication with an external device such as the image capturing device 110 via the network 120. Note that the network communication unit 501 may be configured to communicate directly with the image capturing device 110 without going through the network 120 or another device.

[0043] The control unit 502 executes a computer program stored in memory to control the network communication unit 501, the display unit 503, the operation unit 504, the analysis unit 505, the storage unit 506, the setting processing unit 507, etc. The display unit 503 presents (displays) information to the user via, for example, an LCD display.

[0044] In this embodiment, information is presented to the user by displaying the results of rendering by the browser on a display, but information may also be presented by methods other than a screen display, such as sound or vibration. The operation unit 504 includes a mouse, keyboard, etc. for receiving operations from the user, and the user operates these to input user operations into the browser. The operation unit 504 may also include, for example, a touch panel, a microphone, etc.

[0045] The analysis unit 505 performs vehicle detection processing on the captured image using a vehicle compartment area, which will be described later. In this embodiment, the vehicle compartment refers to a parking space for each vehicle, separated by, for example, a white line, where the vehicle is parked. The storage unit 506 stores data such as the vehicle compartment area, detection area, vehicle area, and trained vehicle detector, which will be described later. The setting processing unit 507 performs setting processing, which will be described later.

[0046] 6 is a diagram showing an example of a setting screen of the server 130 according to the first embodiment, and shows an example of the configuration of the setting screen on which settings related to vehicle detection are made by the setting processing unit 507 of the server 130. The setting screen 600 is made up of a parking lot setting button 601 and the like.

[0047] Pressing the parking lot setting button 601 transitions to a screen for setting the parking lot. The setting screen 600 and the like are displayed on the display unit 503 (display) of the server 130. Furthermore, each setting information set by the user via the operation unit 504 is saved in the saving unit 506 of the server 130.

[0048] When the parking lot setting button 601 is pressed, the screen transitions to a vehicle compartment setting screen 700 shown in Fig. 7. Fig. 7 is a diagram showing an example of a parking lot information setting screen according to the first embodiment, and shows an example of the display of the vehicle compartment setting screen 700.

[0049] The cabin area, which is configured by the coordinates of the four vertices of the corners of the cabin, is set on the cabin setting screen 700. The cabin area is set by the user manually specifying the coordinates of the four vertices of the corners of each cabin while looking at the captured image 701 displayed on the screen.

[0050] The set vehicle compartment area is stored in the storage unit 506 of the server 130. Vehicle compartment 720 is an example of a vehicle compartment area on the captured image 701. The vehicle compartment area of ​​vehicle compartment 720 is set by the user specifying vertex coordinates 721, 722, 723, and 724 of the vehicle compartment. In this way, the setting processing unit 507 functions as a setting means that executes a setting step for acquiring vehicle compartment information of a parking lot based on an image.

[0051] Fig. 8 is a diagram showing an example of the coordinates of the vehicle interior area set on the screen of Fig. 7. A vehicle interior area 800 corresponding to the vehicle interior 720 is made up of four set vertex coordinates 721, 722, 723, and 734. In this embodiment, the origin of the imaging device coordinate system is set at the top left of the vehicle interior setting screen 700 of Fig. 7, and the horizontal x coordinate and vertical y coordinate of the vehicle interior setting screen 700 of Fig. 7 are used. The unit of the vehicle interior coordinates is pixels.

[0052] 9 is a functional block diagram showing an example of the functional configuration of the analysis unit 505 according to embodiment 1. A detection unit 901 acquires a captured image 900 captured by the imaging device 110, detects a vehicle shown in the captured image 900, and outputs a detection area consisting of the center x coordinate, center y coordinate, width, and height of the area where the vehicle exists, as well as the reliability of the detection area.

[0053] In this embodiment, the captured image 900 is acquired directly from the imaging device 110, but the captured image 900 may be acquired by downloading the captured image 900 recorded in a recording device such as a server (not shown). That is, the analysis unit 505 functions as an acquisition means that executes an acquisition step of acquiring an image of the parking lot from the imaging device 110 or a server.

[0054] The detection unit 901 outputs the above-mentioned detection area using a vehicle detector trained by machine learning and stored in the storage unit 506, and also outputs the reliability of the detection area. Note that the above-mentioned vehicle detector may be any object detection neural network that can simultaneously estimate the detection area and its reliability, and for example, the technology described in Literature 1 may be used. Note that the reliability of the detection area is output as a real number ranging from 0 to 1, with 0 being the lowest reliability and 1 being the highest reliability.

[0055] FIG. 10 is a diagram showing an example of a detection area detected by the detection unit 901, and shows an example of a detection area 1020 detected by the detection unit 901 when a vehicle 1010 is captured in the captured image 701.

[0056] The detection unit 901 detects a vehicle 1010 in the captured image 701 and outputs a center x coordinate 1031, a center y coordinate 1032, a width 1033, a height 1034, and its reliability 1035 as a detection area 1021. The detection unit 901 stores the output detection area 1020 in the storage unit 506 of the server 130.

[0057] The target area calculation unit 902 in Figure 9 uses the information on the vehicle interior area 800 set by the setting processing unit 507 and stored in the storage unit 506 of the server 130 to calculate target areas that are candidates for the target area corresponding to each vehicle interior, and stores the target areas in the storage unit 506 of the server 130. Here, the target area calculation unit 902 functions as a calculation means, and executes a calculation step of calculating a detection target area for vehicle detection based on the vehicle interior information.

[0058] The target area consists of the center x coordinate, center y coordinate, width, and height. The area that forms a circumscribing rectangle of the vehicle interior area is the range of the target area corresponding to that vehicle interior. The center x coordinate of the target area can be calculated by averaging the minimum and maximum values ​​of the x coordinates of the vertices of the vehicle interior area.

[0059] Similarly, the y-coordinate of the center of the target area is calculated by averaging the minimum and maximum y-coordinates of the vertices of the vehicle interior area. The width of the target area can be calculated by the difference between the maximum and minimum x-coordinates of the vertices of the vehicle interior area. Similarly, the height of the target area can be calculated by the maximum and minimum y-coordinates of the vertices of the vehicle interior area.

[0060] 11 is a diagram showing an example of a target area according to the first embodiment, and shows an example of a target area 1100 that is output when the vehicle interior area 800 is obtained. The vertex x coordinate of the vehicle interior area 800 has the maximum value at the x coordinate of the vertex coordinate 723, and the vertex x coordinate has the minimum value at the x coordinate of the vertex coordinate 721.

[0061] As a result, the center x coordinate of target area 1100 corresponding to vehicle interior area 800 is the value shown in center x coordinate 1111, and the width is the value shown in width 1113. Similarly, the center y coordinate of target area 1100 is center y coordinate 1112, and the height is height 1114.

[0062] The area integration unit 903 outputs the center x coordinate, center y coordinate, width, height, and reliability of the vehicle area where the final vehicle is located, based on the detection area detected by the detection unit 901 and the target area calculated by the target area calculation unit 902.

[0063] The area integration unit 903 calculates the allocation score (described later) between one target area and all detection areas, and stores the detection area with the highest allocation score in the storage unit 506 in the server 130 as the vehicle interior area corresponding to that target area.

[0064] Furthermore, if the assignment score is equal to or greater than a predetermined assignment threshold, the detected region is considered to have been assigned to the target region and is deleted. Similarly, the assigned target region is also deleted as having been assigned. This process is performed for all target regions.

[0065] In addition, the detection areas that remain after all target areas have been deleted are treated as detection areas for vehicles that exist in areas other than the passenger compartment area, and the detection area frames are integrated using a method such as NMS (Non-Maximum Suppression).

[0066] The allocation score used to allocate a detection region to a target region is calculated by multiplying the similarity between the target region and the detection region by the reliability of the detection region. The similarity between the target region and the detection region is calculated using, for example, IoU (Intersection-over-Union). However, as long as the similarity between the two regions can be calculated, the score may be calculated using a similarity index such as the DICE coefficient or the Simpson coefficient.

[0067] Fig. 12 is a diagram showing an example of allocation scores between target regions and detection regions according to embodiment 1. In Fig. 12, two vehicles 1201a and 1201b are shown in a captured image 1200, a target region 1202 is acquired by the target region calculation unit 902, and detection regions 1203a, 1203b, and 1203c are acquired by the detection unit 901.

[0068] The values ​​constituting the target region 1202 correspond to the values ​​shown in table 1204. The values ​​constituting the detection regions 1203a, 1203b, and 1203c correspond to rows 1205a, 1205b, and 1205c of table 1205, respectively.

[0069] In this example, the allocation threshold is set to, for example, 0.4. Calculating the allocation scores between the target region 1202 and the detection regions 1203a, 1203b, and 1203c results in allocation scores 1206a, 1206b, and 1206c, respectively.

[0070] Since the allocation threshold is 0.4, the detection areas corresponding to the target area 1202 are 1203a and 1203b, and of these, the detection area 1203a with the highest allocation score is assigned to the target area 1202. Furthermore, since the detection area 1203c is below the allocation threshold, it is not deleted, and the NMS performs frame integration processing.

[0071] As a result, the region integration unit 903 outputs a detection region 1203a as the vehicle region corresponding to the vehicle 1201a, and outputs a detection region 1203c as the vehicle region corresponding to the vehicle 1201b.

[0072] Next, we will explain an example of the flow of processing executed by the image analysis system 101. Each of the following processes is realized by the processor 401 of the server 130 executing a program stored in the RAM 402. However, this is just one example, and some or all of the processes described below may be realized not only by the server 130 but also by the imaging device 110 or dedicated hardware.

[0073] Fig. 13 is a flowchart illustrating a flow for calculating a vehicle area executed by the area integration unit 903 in the server 130 according to the first embodiment. Note that the operation of each step in the flowchart in Fig. 13 is performed sequentially by a CPU or the like serving as a computer in the server 130 executing a computer program stored in a memory.

[0074] Using Figure 13, we will explain an image analysis method in which the area integration unit 903 integrates the detection areas output from the detection unit 901 using the target area output from the target area calculation unit 902, and calculates the vehicle area.

[0075] First, in step S1301, the target area output by the target area calculation unit 902 and stored in the storage unit 506 of the server 130 is acquired. Next, in step S1302, the detection area output by the detection unit 901 and stored in the storage unit 506 of the server 130 is acquired. Here, step S1302 functions as a detection step (detection means) that outputs a vehicle detection area, which is an area that includes a vehicle, based on the image.

[0076] Next, in step S1303, it is determined whether there are any unallocated target areas that have not yet been processed in step S1309. If there are no unallocated target areas, the process proceeds to step S1310. If it is determined that there are unallocated target areas, one of the existing unallocated target areas is selected, and the process proceeds to step S1304.

[0077] In the following explanation, the unassigned target area selected at this time will be referred to as target area A. Also, an initial value of the allocation score for the candidate detection area that will be assigned to target area A and will be a candidate for the detection area is set. The initial value of this allocation score is set to 0. This allocation score will be referred to as the candidate allocation score hereinafter.

[0078] Next, in step S1304, it is determined whether there is a detection area for which an allocation score has not been calculated for target area A. If the determination in step S1304 is No, the process proceeds to step S1309.

[0079] If the answer in step S1304 is Yes, one detection area is selected from among the detection areas for which an allocation score has not been calculated for target area A, and the process proceeds to step S1305. The detection area selected at this time will be referred to as detection area B in the following description.

[0080] Next, in step S1305, an allocation score is calculated from the target area A and the detection area B. The allocation score at this time is called allocation score C.

[0081] Next, in step S1306, it is determined whether the allocation score C is equal to or greater than the allocation threshold to determine whether the detection area B corresponds to the target area A. If the determination in step S1306 is Yes, the process proceeds to step S1307. If the determination in step S1306 is No, the process returns to step S1304.

[0082] In step S1307, it is determined whether the allocation score C is greater than the candidate allocation score in order to select the most appropriate detection area for allocation from among the detection areas corresponding to target area A. If the determination in step S1307 is Yes, the process proceeds to step S1308. If the determination in step S1307 is No, the detection area B is deleted and the process returns to step S1304.

[0083] In step S1308, the allocation candidate detection area that is a candidate for allocation to the target area A is updated. The allocation candidate detection area is updated with the detection area B, and the candidate allocation score is updated with the allocation score C. Then, the process proceeds to step S1304.

[0084] If the determination in step S1304 is No, in step S1309, the allocation candidate detection area is allocated to the target area A. That is, the allocation candidate detection area is stored in the storage unit 506 of the server 130 as a vehicle area corresponding to the target area A, and the allocated target area A is deleted. Then, the process proceeds to step S1303.

[0085] If the determination in step S1303 is No, in step S1310, the region integration unit 903 performs integration processing such as integrating frames of the remaining detected regions that exist in regions other than the target region.

[0086] As a result, even if there are multiple vehicle detection areas (vehicle detection frames) for one vehicle compartment, they will ultimately be integrated into one vehicle detection area (vehicle detection frame) for one vehicle compartment. Here, step S1310 functions as an integration step (integration means) that integrates multiple vehicle detection areas based on the detection target area.

[0087] <Variation 1> In the first embodiment, the user manually inputs the vehicle compartment area while viewing the vehicle compartment setting screen, but the vehicle compartment area may be automatically set without the need for the user to input corresponding points. That is, the setting processing unit 507 may be configured to automatically acquire vehicle compartment information of the parking lot based on the image.

[0088] Specifically, for example, a known Hough transform or the like is used to detect white lines that separate the vehicle interior area in the captured image 701, and multiple points are calculated using edges that are intersections of the white lines as feature points. Then, from the multiple points, points that are likely to be vertices may be selected using a known algorithm such as RANSAC (RANdom Sample Consensus), thereby automatically defining the vehicle interior area.

[0089] <Variation 2> In the first embodiment, four corner vertices of the vehicle interior are specified when the vehicle interior area is set. However, the number of vertices is not limited to four, and three vertices or more than four vertices may be set.

[0090] That is, in the setting processing unit 507, the target area calculation unit 902 may calculate the circumscribing rectangle of an area surrounded by more than four vertices set by the user for one vehicle compartment as the target area corresponding to that vehicle compartment.

[0091] <Embodiment 2> In the first embodiment, the captured image is input without processing to the detection unit 901. However, due to the inclination of the vehicle interior in the overhead image obtained by the imaging device, two target areas may overlap, and this overlap may result in an incorrect detection area being assigned to the target area.

[0092] Therefore, in embodiment 2, incorrect allocation is reduced by rotating the captured image. The system configuration of embodiment 2 is the same as embodiment 1 shown in Fig. 1, and the following description will focus on the differences from embodiment 1.

[0093] In the second embodiment, in addition to the contents of the first embodiment, the storage unit 506 of the server 130 stores a correction angle that minimizes overlap between the target regions.

[0094] That is, the target area calculation unit 902 calculates the rotation angle of the captured image that minimizes overlap between target areas and the target areas after rotation that become candidates for the detection area corresponding to each vehicle compartment from the vehicle interior area 800 set by the setting processing unit 507. The target area after rotation is composed of the center x coordinate, center y coordinate, width, and height.

[0095] Rotation is a coordinate transformation that rotates counterclockwise with the center of the captured image as the origin in the image capture device coordinate system. If the center coordinates of the image are (Cx, Cy) and the rotation angle is θ, the point (x1, y1) on the captured image and the point (x2, x2) on the rotated image can be expressed as the following equation 1 using the transformation matrix Rθ.

[0096]

number

[0097] The area that forms a circumscribing rectangle of the cabin area after rotation is the target area corresponding to the cabin. The center x-coordinate of the target area can be calculated by averaging the minimum and maximum x-coordinates of the vertices of the cabin area after rotation.

[0098] Similarly, the y-coordinate of the center of the target area is calculated as the average of the minimum and maximum y-coordinates of the vertices of the rotated vehicle interior area. The width of the target area can be calculated as the difference between the maximum and minimum x-coordinates of the vertices of the rotated vehicle interior area. Similarly, the height of the target area can be calculated as the maximum and minimum y-coordinates of the vertices of the rotated vehicle interior area.

[0099] To determine the correction angle that minimizes overlap between target regions and the target regions at that angle, the similarity between target regions at a certain rotation angle is calculated. The similarity is calculated using, for example, IoU. The similarity is output as a real number between 0 and 1, with 0 being the minimum and 1 being the maximum.

[0100] However, it is sufficient if the similarity between two regions can be calculated, and similarity indices such as the DICE coefficient and Simpson coefficient can be used. The similarity is calculated between all overlapping target regions, and the average similarity is determined by calculating the arithmetic mean of all similarities.

[0101] The rotation angle is determined in increments of 5°, for example, within a range from 0° to 90°, and the average similarity is calculated for each angle. The rotation angle at which the average similarity is smallest is stored as the correction angle in storage unit 506 of server 130. That is, the rotation angle of the image is determined based on the similarity between the multiple detection target regions calculated for each rotation angle.

[0102] The target region at this time is stored in the storage unit 506 of the server 130. However, the range of rotation angles and the increments of rotation angles are not limited to those described above. Any range of rotation angles and increments of angles can be used to calculate the rotation angle that minimizes the average similarity.

[0103] The detection unit 901 detects an area where a vehicle exists based on an image captured by the imaging device 110, which is rotated by a correction angle output by the target area calculation unit 902 and stored in the storage unit 506 of the server 130. Then, the detection unit 901 outputs a detection area consisting of the center x coordinate, center y coordinate, width, and height of the rotated area where the vehicle exists, and the reliability of the detection area, as in the first embodiment.

[0104] The region integration unit 903 finally outputs a vehicle region where a vehicle exists, based on the detection region detected by the detection unit 901 and the target region calculated by the target region calculation unit 902. The vehicle region is composed of a center x coordinate, a center y coordinate, a width, a height, and a reliability.

[0105] After the vehicle area is determined as described above, the coordinates of this vehicle area are inversely transformed back to the original coordinate system by rotating them by the opposite angle to the correction angle. That is, the inverse transformation is performed by finding the inverse matrix of the transformation matrix Rθ in Equation 1, and using this inverse matrix, transforming the center coordinates of the vehicle area. The vehicle area after the inverse transformation is stored in the storage unit 506 of the server 130.

[0106] 14 is a flowchart illustrating the flow of calculating the target area, which is executed by the target area calculation unit 902 according to embodiment 2. Note that the operation of each step in the flowchart in FIG. 14 is performed sequentially by a CPU or the like serving as a computer in the server 130 executing a computer program stored in memory.

[0107] First, in step S1401, the rotation angle for rotating the image is initialized. The initial value of the rotation angle is set to 0. Also, the initial value of the average similarity when the overlap between the target regions is smallest is set. The initial value of this average similarity is set to 1. Hereinafter, this average similarity will be referred to as the candidate average similarity.

[0108] Next, in step S1402, it is determined whether the rotation angle is less than 90 degrees. If it is less than 90 degrees, the process proceeds to step S1403. If it is 90 degrees or more, the process proceeds to step S1409.

[0109] In step S1403, the vehicle interior area set by the setting processing unit 507 and stored in the storage unit 506 is acquired. Next, in step S1404, it is determined whether there is a vehicle interior area for which the target area has not been calculated.

[0110] If the determination in step S1404 is No, the process proceeds to step S1406. If the determination in step S1404 is Yes, one of the existing vehicle interior areas is selected, and the process proceeds to step S1405. The vehicle interior area selected at this time is called vehicle interior area A.

[0111] In step S1405, the target area corresponding to the vehicle interior area A is calculated. The coordinates of each vertex of the target area are rotated using a transformation matrix of the current rotation angle. After that, the circumscribing rectangle of the area enclosed by the four vertices after rotation is calculated as the target area of ​​the vehicle interior area. Here, step S1405 functions as a rotation step (rotation means) that rotates the image. Then, the process returns to step S1404.

[0112] In step S1406, the average similarity is calculated for all combinations of overlapping target regions from among all target regions corresponding to all vehicle interior regions calculated in step S1405. Hereinafter, this average similarity will be referred to as average similarity B.

[0113] Next, in step S1407, it is determined whether the average similarity B is smaller than the candidate average similarity. If the determination in step S1407 is No, the process proceeds to step S1402. If the determination in step S1407 is Yes, the process proceeds to step S1408.

[0114] In step S1408, the correction angle is updated to the current rotation angle. The rotation angle is then incremented by 5°. In this way, in steps S1406 to S1408, the image rotation angle is determined based on the similarity between the multiple detection target regions calculated for each rotation angle. Then, the process proceeds to step S1402.

[0115] If the determination in step S1402 is No, the correction angle and the target area at that time are saved in the saving unit 506 of the server 130 in step S1409.

[0116] <Variation 3> In the second embodiment, the entire captured image is rotated, but a partial image extracted from the captured image may also be rotated. Furthermore, if multiple partial images are extracted from the captured image, the rotation angle may be determined for each of the partial images. That is, the rotation angle of the detection target region may be determined based on the similarity between overlapping detection target regions calculated for each rotation angle.

[0117] <Embodiment 3> In the first and second embodiments, the error between the position and size of the detection area output from the detection unit and the actual position and size in the captured image is not taken into consideration, but there may be cases where the detection area cannot be assigned to the target area due to the influence of the error.

[0118] Therefore, a method for suppressing the influence of errors will be described in embodiment 3. The system configuration of embodiment 3 is the same as embodiment 1 shown in Fig. 1, and the following description will focus on the differences from embodiment 1.

[0119] 15 is a functional block diagram showing an example of the functional configuration of a server according to embodiment 3. In addition to the configuration of embodiment 1, the server 130 includes a statistical processing unit 1501. A control unit 502 controls the statistical processing unit 1501 to execute the processing of embodiment 1 in addition to the processing of embodiment 1. A storage unit 506 stores error information, evaluation information, minimum allowable size, and maximum allowable size, which will be described later, in addition to the information of embodiment 1.

[0120] The statistical processing unit 1501 uses the trained vehicle detector stored in the storage unit 506 to output error information, minimum allowable size, and maximum allowable size from the evaluation information stored in the storage unit 506, and stores them in the storage unit 506. The error information consists of an x-coordinate deviation, a y-coordinate deviation, a width deviation, and a height deviation.

[0121] The minimum and maximum allowable sizes are composed of width and height errors. The evaluation information is prepared in advance by the user and is composed of one or more evaluation images showing a vehicle and an evaluation vehicle area that accurately represents the area of ​​the vehicle corresponding to each image.

[0122] The error distribution for each center x coordinate, center y coordinate, width, and height of the detection area and evaluation vehicle area obtained when all evaluation images are input into the vehicle detector is calculated. The quartile deviation is calculated from each obtained error distribution and saved as error information.

[0123] The first quartile of the width error distribution is set as the minimum allowable width error, and the third quartile is set as the maximum allowable width error, and these are stored in storage unit 506. Similarly, the first quartile of the height error distribution is set as the minimum allowable height error, and the third quartile is set as the maximum allowable height error, and these are stored in storage unit 506.

[0124] The target area calculation unit 902 uses the vehicle interior area set by the setting processing unit 507 and stored in the storage unit 506 of the server 130 to output target areas that are candidates for the target area corresponding to each vehicle interior, and stores them in the storage unit 506 of the server 130. The target area is composed of a center x coordinate, a center y coordinate, a width, and a height.

[0125] The center coordinates of the circumscribing rectangle of the vehicle interior area become the center x-coordinate and center y-coordinate of the target area. The width of the circumscribing rectangle of the vehicle interior area plus the x-coordinate deviation and width deviation of the error information becomes the width of the target area. Similarly, the height of the circumscribing rectangle of the vehicle interior area plus the y-coordinate deviation and height deviation of the error information becomes the width of the target area.

[0126] The region integration unit 903 outputs the final vehicle region where the vehicle is located from the detection region detected by the detection unit 901 and the target region calculated by the target region calculation unit 902. The vehicle region is composed of the center x coordinate, center y coordinate, width, height, and reliability. In this way, in this embodiment, the detection target region is calculated based on the statistical error information consisting of the quartile deviation of the error distribution of the detection results of the detection means and the vehicle interior information.

[0127] In the first embodiment, a predetermined value is used as the allocation score threshold, but in the third embodiment, in order to suppress the influence of errors in the detection unit 901, an allocation score threshold is determined for each target region.

[0128] That is, first, an area is calculated by subtracting the minimum allowable width and height errors from the width and height of the target area. This area is called the minimum allowable area. Similarly, an area is calculated by adding the maximum allowable width and height errors to the width and height of the area. This area is called the maximum allowable area.

[0129] Furthermore, the similarity between the minimum allowable area and the maximum allowable area for the target area is calculated. This similarity is calculated using, for example, IoU. The smaller of the similarities between the minimum allowable area and the maximum allowable area is used as the allocation threshold for this target area. In this way, it is possible to determine whether to allocate a vehicle detection area to the detection target area using the allocation threshold calculated based on statistical error information.

[0130] The present invention has been described above in detail based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible based on the spirit of the present invention, and these are not excluded from the scope of the present invention.

[0131] For example, in the above-described first to third embodiments, an example has been described in which the captured image captured by the imaging device 110 is analyzed by the server 130, but the captured image captured by the imaging device 110 may be configured to be analyzed by the imaging device 110.

[0132] The present invention also includes those that realize the functions of the above-described embodiments using at least one processor or circuit such as a CPU, etc. Also, it is possible to use multiple processors to perform distributed processing.

[0133] In order to realize part or all of the control in the above embodiments, a computer program that realizes the functions of the above embodiments may be supplied to an image analysis device or the like via a network or various storage media. Then, a computer (or a CPU, MPU, or the like) in the image analysis device or the like may read and execute the program. In this case, the program and the storage medium storing the program constitute the present invention. The present invention also includes the following combinations.

[0134] (Configuration 1) An image analysis device having an acquisition means for acquiring an image of a parking lot, a setting means for acquiring vehicle compartment information of the parking lot based on the image, a calculation means for calculating a detection target area for vehicle detection based on the vehicle compartment information, a detection means for outputting a vehicle detection area that is an area containing a vehicle based on the image, and an integration means for integrating multiple vehicle detection areas based on the detection target area.

[0135] (Configuration 2) The image analysis device according to Configuration 1, characterized in that the calculation means has a rotation means for rotating the image, and determines the rotation angle of the image based on the similarity between the plurality of detection target regions calculated for each rotation angle.

[0136] (Configuration 3) The image analysis device according to configuration 1 or 2, characterized in that the calculation means has a rotation means for rotating the image, and determines a rotation angle of the detection target region based on the similarity between the overlapping detection target regions calculated for each rotation angle.

[0137] (Configuration 4) The calculation means calculates the detection target area based on statistical error information consisting of quartile deviation of an error distribution of the detection result of the detection means and the vehicle interior information; 4. The image analysis device according to any one of configurations 1 to 3, characterized in that:

[0138] (Configuration 5) The integration means determines whether to allocate the vehicle detection area to the detection target area using an allocation threshold calculated based on the statistical error information; 5. The image analysis device according to configuration 4,

[0139] (Configuration 6) An image analysis system that is composed of at least one imaging device and an image analysis device and analyzes the parking condition in a parking lot having spaces where vehicles can be parked, wherein the imaging device acquires an image of the parking lot and transmits it to the image analysis device, and the image analysis device has a setting means that acquires space information for the parking lot based on the image, a calculation means that calculates a detection target area for vehicle detection based on the space information, a detection means that outputs a vehicle detection area that is an area that includes a vehicle based on the image, and an integration means that integrates multiple vehicle detection areas based on the detection target areas.

[0140] (Method) An image analysis method comprising an acquisition step of acquiring an image of a parking lot, a setting step of acquiring vehicle compartment information of the parking lot based on the image, a calculation step of calculating a detection target area for vehicle detection based on the vehicle compartment information, a detection step of outputting a vehicle detection area that is an area that includes a vehicle based on the image, and an integration step of integrating multiple vehicle detection areas based on the detection target area.

[0141] (Program) A computer program for controlling each means of the image analysis device according to any one of configurations 1 to 5 or the image analysis system according to configuration 6 by a computer. [Explanation of symbols]

[0142] 201: Imaging unit 505: Analysis Department 507: Setting processing section 901: Detection unit 902: Target area calculation unit 903: Area Integration Department

Claims

1. an acquisition means for acquiring an image of the parking lot; a setting means for acquiring parking space information of the parking lot based on the image; a calculation means for calculating a detection target area for vehicle detection based on the vehicle compartment information; a detection means for outputting a vehicle detection area, which is an area including a vehicle, based on the image; an integration means for integrating the vehicle detection areas based on the detection target areas; An image analysis device having the above.

2. the calculation means has a rotation means for rotating the image, 2. The image analysis device according to claim 1, wherein the rotation angle of the image is determined based on the similarity between the plurality of detection target regions calculated for each rotation angle.

3. the calculation means has a rotation means for rotating the image, 2. The image analysis device according to claim 1, wherein the rotation angle of the detection target region is determined based on the similarity between the overlapping detection target regions calculated for each rotation angle.

4. the calculation means calculates the detection target area based on statistical error information consisting of quartile deviation of an error distribution of the detection result of the detection means and the vehicle interior information; 2. The image analysis device according to claim 1,

5. the integrating means determines whether to allocate the vehicle detection area to the detection target area using an allocation threshold calculated based on the statistical error information; 5. The image analysis device according to claim 4,

6. An image analysis system that analyzes a parking state in a parking lot having a vehicle space where a vehicle can be parked, the image analysis system comprising at least one imaging device and an image analysis device, the imaging device acquires an image of the parking lot and transmits it to the image analysis device; The image analysis device a setting means for acquiring parking space information of the parking lot based on the image; a calculation means for calculating a detection target area for vehicle detection based on the vehicle compartment information; a detection means for outputting a vehicle detection area, which is an area including a vehicle, based on the image; an integration means for integrating the vehicle detection areas based on the detection target areas; An image analysis system comprising:

7. an acquisition step of acquiring an image of the parking lot; a setting step of acquiring vehicle space information of the parking lot based on the image; a calculation step of calculating a detection target area for vehicle detection based on the vehicle interior information; a detection step of outputting a vehicle detection area, which is an area including a vehicle, based on the image; an integration step of integrating the plurality of vehicle detection areas based on the detection target area; An image analysis method comprising:

8. A computer program for controlling each means of the image analysis device according to any one of claims 1 to 5 or the image analysis system according to claim 6 by a computer.

Citation Information

Patent Citations

  • Method for measuring parking lot occupancy state from digital camera image

    JP2013206462A