Image processing device, imaging apparatus, image processing method, program, and storage medium

The image processing device addresses accuracy issues in obstacle detection by calculating three-dimensional position information relative to a reference plane, ensuring reliable obstacle detection without calibration values, thus improving precision in ADAS and autonomous driving systems.

JP2025169111APending Publication Date: 2025-11-12CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024074143
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing obstacle detection methods in ADAS and autonomous driving systems face accuracy issues due to the use of calibration values that can change over time or with in-vehicle environment variations, leading to reduced precision in obstacle detection.

Method used

An image processing device that calculates three-dimensional position information without requiring a calibration value by using height and depth information to determine obstacles relative to a reference plane, employing methods like SfM, triangulation, and reference plane height calculation to establish obstacle presence.

Benefits of technology

Enables accurate obstacle detection using scale-invariant three-dimensional position information, reducing errors and enhancing the reliability of obstacle determination in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169111000001_ABST
    Figure 2025169111000001_ABST
Patent Text Reader

Abstract

To provide an image processing device, an imaging apparatus, a mobile object, an image processing method and a program which can determine an obstacle without requiring a calibration value by performing calculation as three-dimensional position information with undefined scale.SOLUTION: In a mobile object comprising: an imaging apparatus including an imaging element and an image processing device; a distance acquisition device; a vehicle information acquisition device; an external world recognition device; a control device; and an alarm, the image processing device 111 comprises: a position information acquisition unit 310 which detects a feature point in an image acquired by the imaging element and acquires height information showing height of the feature point on the image and depth information showing distance to the feature point; a reference plane height calculation unit 320 which calculates height of a predetermined reference plane as reference plane height on the basis of the height information; and an obstacle determination unit 340 which determines whether the feature point shows an object to be an obstacle of movement of the mobile object by using the reference plane height.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device that detects an obstacle from an image. [Background technology]

[0002] In recent years, ADAS (Advanced Driver-Assistance Systems) and autonomous driving have been attracting attention. To realize these, it is necessary to detect areas where vehicles cannot pass (hereinafter referred to as obstacles), such as fallen objects on the road and depressions such as grooves.

[0003] Patent Document 1 discloses a method for detecting obstacles using images. In Patent Document 1, an obstacle is determined to exist if the three-dimensional position estimated using SfM (Structure from Motion) is higher than a predetermined value. The values ​​acquired using SfM are values ​​with an indefinite scale. Scale is the degree of size of one unit of a certain index. In other words, the value acquired using SfM does not have a fixed metric. Therefore, a calibration value, such as the position information of the imaging device that captured the image, is used to convert the value to a real scale. However, Patent Document 1 has the risk of reducing the accuracy of the calibration value, resulting in a reduction in the accuracy of obstacle detection. This is because the calibration value may change over time or depending on the in-vehicle environment. For example, position information, such as the installation height of the imaging device, changes depending on the load and its distribution.

[0004] In view of the above, Patent Document 2 discloses a method for correcting calibration values ​​by estimating the installation height of an imaging device. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent Publication No. 2018-156222 [Patent Document 2] Patent Publication No. 2023-39777 [Non-patent literature]

[0006] [Non-Patent Document 1] R. Ranftl,et al.Towards robust monocular depth estimation:Mixing datasets for zero-shot cross-dataset transfer.IEEE TPAMI,44(3):1623-1637,2020. Summary of the Invention [Problem to be solved by the invention]

[0007] In the method disclosed in Patent Document 2, values ​​are corrected, so there is a risk that errors that cannot be corrected will remain.

[0008] Therefore, an object of the present invention is to provide an image processing device that can perform an obstacle determination without requiring a calibration value by performing calculations using three-dimensional position information with an indefinite scale. [Means for solving the problem]

[0009] An image processing device according to one aspect of the present invention is an image processing device installed on a moving body, and includes an image acquisition means for acquiring an image, an information acquisition means for detecting feature points within the image and acquiring height information indicating the height of the feature points on the image and depth information indicating the distance to the feature points, a reference plane height calculation means for calculating the height of a predetermined reference plane as the reference plane height based on the height information, and a determination means for using the reference plane height to determine whether the feature points indicate an object that will obstruct the movement of the moving body.

[0010] Other objects and features of the present invention will be described in the following embodiments. [Effects of the Invention]

[0011] According to the present invention, it is possible to provide an image processing device that can determine whether or not an object is an obstacle based on three-dimensional position information with an indefinite scale. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is an explanatory diagram of an imaging device including an image processing device according to a first embodiment; [Figure 2] FIG. 1 is an explanatory diagram of a light beam received by an image sensor according to a first embodiment; [Figure 3] FIG. 1 is an explanatory diagram of an image processing apparatus according to a first embodiment; [Figure 4] FIG. 1 is an explanatory diagram of a location information acquisition method according to a first embodiment; [Figure 5] 1 is a diagram illustrating features of the first embodiment; [Figure 6] FIG. 1 is an explanatory diagram of triangulation according to a first embodiment; [Figure 7] FIG. 1 is an explanatory diagram of a reference surface height estimation method according to a first embodiment; [Figure 8] FIG. 1 is an explanatory diagram of an obstacle determination method according to a first embodiment; [Figure 9] FIG. 1 is an explanatory diagram of block matching according to a first embodiment; [Figure 10] FIG. 10 is an explanatory diagram of an imaging device according to a second embodiment; [Figure 11] FIG. 10 is an explanatory diagram of a moving body according to a second embodiment; [Figure 12] FIG. 10 is an explanatory diagram of a block matching method according to a second embodiment. [Figure 13] FIG. 10 is an explanatory diagram of an imaging device according to a modified example of the second embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0013] The present invention will be described in detail using embodiments and drawings. The present invention is not limited to the contents described in each embodiment. In addition, each embodiment may be combined appropriately.

[0014] First Embodiment (Device configuration) FIG. 1 is a diagram schematically showing the configuration of an imaging device according to an embodiment of the present invention.

[0015] In FIG. 1A, a moving object 100 includes an imaging device 110, a distance acquisition device 120, a vehicle information acquisition device 130, an external environment recognition device 140, a control device 150, and an alarm device 160.

[0016] The mobile object 100 is an object that moves using a power source and is capable of moving by itself, such as an automobile, a ship, an aircraft, a drone, an industrial robot, etc. In the following description, the mobile object 100 will be described as an automobile.

[0017] In FIG. 1B, the imaging device 110 includes an image processing device 111 and an imaging unit 112.

[0018] The imaging unit 112 includes an imaging element 112-1 and an optical system 112-2. The image processing device 111 can be configured using a logic circuit. Alternatively, the image processing device 111 may be configured with a central processing unit (CPU) and a memory that stores a processing program.

[0019] The optical system 112-2 is a photographing lens of the imaging device 110, and has the function of forming an image of a subject on the imaging element 112-1 (on the imaging element). The optical system 112-2 is composed of a plurality of lens groups (not shown) and an aperture (not shown), and has an exit pupil 123 at a predetermined distance from the imaging element 112-1. Note that in this specification, the z-axis is parallel to the optical axis 130 of the optical system 112-2. Furthermore, the x-axis and y-axis are perpendicular to each other and to the optical axis.

[0020] The image sensor 112-1 is composed of a CMOS (complementary metal-oxide semiconductor) or a CCD (charge-coupled device). The subject image formed on the image sensor 112-1 via the optical system 112-2 is photoelectrically converted by the image sensor 112-1 to generate an image signal based on the subject image.

[0021] For example, the imaging device 110 is installed at a predetermined position inside the vehicle close to the front (or rear) windshield of the moving body 100, and the imaging unit 112 captures an image of the forward field of view (or rear field of view) of the moving body 100.

[0022] The distance acquisition device 120 is a sensor that acquires distance information around the vehicle, and may be configured with, for example, a millimeter wave radar or a LiDAR (Light Detection and Ranging).

[0023] FIG. 2 is a flowchart showing the operation of the moving body 100 of this embodiment.

[0024] S200 is executed by the image processing device 111. The image processing device 111 first acquires image information from the imaging device 110, and acquires distance information from the distance calculation device 120. Then, it acquires information about an obstacle from this image information and the distance information acquired from the distance calculation device 120. Details will be described later.

[0025] S210 is executed by the vehicle information acquisition device 130. The vehicle information acquisition device 130 acquires one or more pieces of vehicle information (information about a moving body) from among a moving speed, a roll angle, a pitch angle, and the like.

[0026] S220 is executed by the external environment recognition device 140. The external environment recognition device 140 recognizes the degree of danger in the external environment from obstacle information acquired by the imaging device 110 and information on the moving object acquired by the vehicle information acquisition device 130. For example, it recognizes whether there is danger if the moving object moves in the current direction and speed. More specifically, it may recognize danger when the moving object is moving and there is an obstacle close to the moving object's direction of travel. If the degree of danger is recognized to be low, the processing ends.

[0027] S230 is executed by the control device 150. When the external environment recognition device 140 recognizes that the danger level is high, the control device 150 controls the moving body to avoid or reduce the danger. For example, the control device 150 may brake the moving body 100 or change the direction of travel.

[0028] S240 is executed by the warning device 160. When the external environment recognition device 140 recognizes that the danger level is high, the warning device 160 issues a warning to passengers or people around. For example, the warning device 160 performs processing to generate a warning sound or the like, or to display warning information on a display screen of a car navigation device or the like, a head-up display, or the like. Alternatively, the warning device 160 warns the driver of the moving object 100 by vibrating a seat belt or the steering wheel, for example.

[0029] (Explanation of image processing device) The image processing device 111 will be described below. The image processing device 111 acquires three-dimensional position information, and calculates a determination result as to whether or not that position is an obstacle and a distance value thereof. Note that the three-dimensional position information acquired at this time may have an indefinite scale.

[0030] Fig. 3(A) is a diagram schematically showing the configuration of an image processing device 111 according to an embodiment of the present invention. In Fig. 3(A), the image processing device 111 includes a position information acquisition unit 310, a reference surface height calculation unit 320, a threshold determination unit 330, an obstacle determination unit 340, a region integration unit 350, and a distance information acquisition unit 360. Fig. 3(B) is a flowchart showing the operation of the image processing device 111 according to this embodiment. When image processing according to this embodiment starts, the process proceeds to step S310.

[0031] In step S310, photography is performed using the imaging device 110, an image is generated and acquired, and the acquired image is stored in the main body memory (not shown). That is, image acquisition is performed.

[0032] Furthermore, the image acquired in step S310 may be subjected to processing to correct imbalances in light intensity, which are primarily caused by vignetting in the optical system 112-2. Specifically, the light intensity balance can be corrected by correcting the luminance value of the image so that it becomes approximately constant regardless of the angle of view, based on the results of the image capture device 110 previously capturing an image of a surface light source with a constant luminance. Furthermore, the acquired image may be subjected to filtering processing using a band-pass filter, a low-pass filter, or the like, to reduce the effects of optical shot noise generated by the image capture element 112-1, for example. Alternatively, the image may be downsized to reduce calculation costs. Alternatively, the image may be made higher-resolution using a known method to determine obstacles at higher resolution.

[0033] Step S320 is performed by the position information acquisition unit 310. In step S320, three-dimensional position information is calculated from the acquired image. Any known method may be used, but here a method using SfM will be described with reference to Fig. 4. Fig. 4 is a flowchart detailing the processing flow of step S320.

[0034] In step S321, feature points are matched using a known method from the images acquired in step S310. Note that at least two images (image It1 and image It2) must be acquired here. These images are taken at different times t1 and t2.

[0035] Feature point matching will be explained in detail using Figures 5(A) and (B). First, feature points of images It1 and It2 are calculated. Any known method can be used, but here the Harris corner detection algorithm is used. Feature points 501 calculated for image It1 are shown in Figure 5(A), and feature points 502 calculated for image It2 are shown in Figure 5(B). Here, the respective feature points are indicated by stars. Next, correspondence is established between feature points 501 and 502. Any known method can be used, but here the KLT (Kanade-Lucas-Tomasi) feature tracking algorithm is used.

[0036] The algorithm used to calculate feature points and feature quantities is not limited to the methods described here, and may be, for example, FAST (Features from Accelerated Segment Test), BRIEF (Binary Robust Independent Elementary Features), ORB (Oriented FAST and Rotated BRIEF), etc.

[0037] Furthermore, matching can be performed after removing feature points at specific pixel positions. For example, it is highly likely that valid feature points cannot be calculated in areas such as the hood, so removing such points or limiting the acquisition of such points will enable subsequent calculation of 3D position information with higher accuracy.

[0038] In step S322, a rotation matrix R representing the amount of rotational movement of the camera between frames and a translation vector T representing the amount of translational movement are estimated using the correspondence result acquired in step S321.

[0039] First, find the fundamental matrix F. When x1 and x2 are corresponding points, the relationship in Equation 1 holds. Note that x1 and x2 are three-dimensional vectors that express the coordinates of corresponding points in the image coordinate system in the simultaneous coordinate system. For example, with the 5-point algorithm, the fundamental matrix F can be found if the image positions of at least five pairs of corresponding points are obtained. Also, with the 8-point algorithm, these can be found if the image positions of at least eight pairs of corresponding points are obtained. Furthermore, if there are more corresponding points than necessary, a least-squares solution can be used. Alternatively, the RANSAC (Random Sample Consensus) method can be used to remove outliers and use the results.

[0040]

number

[0041] Next, the fundamental matrix E of the camera is calculated using Equation 2. K1 and K2 are internal matrices of the camera that indicate the values ​​of parameters such as the focal length of the camera and the center position of the two-dimensional coordinates. The internal matrices K1 and K2 of the camera may be set to default values.

[0042]

number

[0043] Finally, we calculate the rotation matrix R and the translation vector T. The fundamental matrix E can be decomposed into the rotation matrix R and the translation vector T as shown in Equation 3. E=T×R (Equation 3)

[0044] Although the translation vector T obtained here has some uncertainty regarding the constant factor, the process may proceed directly to the next step S323. When scaling is performed, the amount of camera movement may be obtained from various measuring devices, specifically, an IMU or GNSS, or in the case of an in-vehicle camera, vehicle speed information or map information, and then scaling may be performed.

[0045] Note that, among the correspondences used in the above calculations, feature points calculated from subjects that are not stationary relative to the world coordinate system to which the imaging device belongs may be excluded from the processing. The above camera movement estimation calculates various parameters assuming the subject is a stationary object, which can become a source of error if the subject is a moving object. Therefore, excluding feature points calculated from moving objects can improve the accuracy of calculating various parameters. Moving objects are determined by classifying the subject using image recognition technology or by comparing the relative value of the time-series change in the acquired distance information with the movement of the imaging device.

[0046] In step S322, the rotation matrix R and translation vector T acquired in step S321 are used to calculate the three-dimensional position information of the matching point according to the principle of triangulation. The principle of triangulation will be explained using FIG. 5. FIG. 6 shows the three-dimensional space coordinate X of the matching point, the three-dimensional space coordinate C1 centered on one camera, and the three-dimensional space coordinate C2 centered on another camera. The angle θ1 at both ends of the triangle XC1C2 can be calculated from the coordinates of the feature points on the image and the rotation matrix R. Once the angles θ1 and θ2 are known, the positions of the vertices can be measured. Note that here, the relative positions may be calculated when the distance between C1 and C2 is normalized to 1.

[0047] Furthermore, bundle adjustment, a well-known method, may be used to calculate the amount of rotational and translational movement of the camera and the positional relationship between the subject and the camera. The relationships between the camera fundamental matrix and corresponding points, including internal camera parameters such as focal length, can be analytically calculated together using the nonlinear least squares method to improve consistency.

[0048] Furthermore, the reliability of the three-dimensional position information calculated by the optical system 112-2 may be calculated, and if the reliability is low, the feature point may be eliminated. In this case, the three-dimensional position information that is presumed to have been calculated incorrectly can be eliminated, allowing subsequent obstacle determination to be performed with high accuracy.

[0049] Although the above describes an example in which SfM is used, three-dimensional position information may be calculated from an acquired image using a model that has been trained in advance by machine learning or the like.

[0050] For example, depth estimation (estimation of depth information) of a single image can be performed using a convolutional neural network as described in Non-Patent Document 1. From there, conversion into three-dimensional position information can be performed using information from the optical system 112-2 of the imaging device 110.

[0051] In many models, the depth estimation result has several times the uncertainty, but the process may proceed to the next step S323 as described above, or scaling may be performed.

[0052] Step S330 is performed by the reference plane height calculation unit 320. In step S330, the height of the reference plane is calculated from the three-dimensional position information acquired by the position information acquisition unit 310. The reference plane may be set to a road surface, a floor inside a building, or the like.

[0053] 7 is a detailed flowchart of the process flow of step S330. The height of the reference plane may be calculated using any known method, but the following describes the process based on the flow in FIG.

[0054] In step S331, the image is divided into two or more regions. For example, the image may be divided horizontally at equal intervals without being divided vertically. Alternatively, the image may be divided so that the angle of light incident on the image sensor is equal to the intervals based on information from the optical system 112-2.

[0055] In step S332, feature points estimated to belong to the reference plane are extracted from each region divided in the region division process S331, and the extracted point group is used as the reference plane. For example, if the reference plane is the road surface, the position with the lowest height is likely to be the road surface, so a predetermined number of points may be extracted from those with the lowest height estimated in step S320. In this case, a predetermined condition is set to extract a predetermined number of points from those with the lowest estimated height, and those that match are extracted. Alternatively, a predetermined number of points may be extracted from pixel positions in the image that correspond to the reference plane, such as the lower region on the image.

[0056] In step S333, the height of the reference plane is calculated from the three-dimensional position information of the feature points extracted in step S332. For example, the average, median, or mode of the heights of the extracted feature points may be calculated as the height of the reference plane. Alternatively, the feature points may be fitted to a specific function such as a plane, and the resulting shape may be used as the reference plane. In this case, accuracy is improved because roll angles and pitch angles can be taken into account when the imaging device is not level with the reference plane.

[0057] In the process based on the flow in FIG. 7, points to be used for extracting the reference plane are selected from each divided region, so that it is possible to prevent a decrease in accuracy due to pixel position.

[0058] Step S340 is performed by threshold value determination unit 330. Step S340 determines the threshold value based on the obstacle to be detected, i.e., the ratio of height to depth of the detection target. For example, if an object of Bm or more at Am distance is to be set as an obstacle, B / A may be set as the threshold. Alternatively, if a depression of Dm or more at Cm distance is to be set as an obstacle, D / C may be set as the threshold. Furthermore, multiple threshold values ​​may be set.

[0059] Step S350 is performed by obstacle determination unit 340. Step S350 determines whether or not each feature point is an obstacle based on the three-dimensional position information calculated in step S320, the reference plane height calculated in step S330, and the threshold value calculated in step S340. Figure 8 is a flowchart showing the detailed processing flow of step S350.

[0060] In step S351, the position is corrected in accordance with the height of the reference plane calculated in step S330. The object to be corrected is the three-dimensional position information calculated in step S320 or the threshold value calculated in step S340.

[0061] When correcting three-dimensional position information, coordinate conversion or change of coordinate values ​​is performed so that the height of the reference plane becomes 0.

[0062] When correcting or changing the threshold, add the value obtained by dividing the height of the reference plane by the depth of the point to the threshold.

[0063] In step S352, the three-dimensional position information is compared with a threshold value to determine whether or not there is an obstacle. If an obstacle is determined to exist when the ratio of height to depth is greater than the threshold value, a fallen object or a wall can be determined to be an obstacle. Also, if an obstacle is determined to exist when the ratio of height to depth is less than the threshold value, a depression in the reference surface can be determined to be an obstacle. If there are multiple threshold values, it is possible to separately determine whether an obstacle is determined to exist when the value is greater than the threshold value, when the value is smaller than the threshold value, or when the value is outside the threshold range. "Outside the threshold range" refers to outside the range between the upper and lower limits. More specifically, an obstacle is determined to exist when the value is greater than the upper limit value or smaller than the lower limit value.

[0064] Step S360 is performed by the region integration unit 350. In step S360, among the feature points determined to be obstacles, regions that are close and have a similar depth or regions within a predetermined range are determined to be the same object, and classification is performed. For example, a morphological operation may be performed using an arbitrary range of depth values, with pixel values ​​near the feature points determined to be obstacles set to 1 and other values ​​set to 0. In this case, if an obstacle close in depth is present nearby, the pixel value of the entire pixel value of that region becomes 1, and this region can be determined to be the same object. This has the effect of not determining that movement between the multiple feature points is possible, for example, when multiple feature points are detected from obstacles of the same object.

[0065] In step S370, distance acquisition device 120 performs measurements to acquire distance information around the vehicle. The distance values ​​at this time are preferably distance values ​​with a known scale. By acquiring the distance values ​​of the area determined to be an obstacle, the distance to the obstacle can be calculated.

[0066] More preferably, the obstacle area may be expanded to the pixels representing the reference plane by using the distance to the obstacle and information from the optical system 112-2.

[0067] The image processing device of this embodiment can determine an obstacle even with scale-invariant 3D position information by utilizing the ratio of height to depth. In addition, because the distance can be measured after detecting an obstacle, measurement can be performed even if the angular resolution of the distance is low.

[0068] More preferably, a calibration unit is provided and the error of the distance measuring device can be reduced by estimating the coefficients.

[0069] The calibration unit calibrates the distance values ​​of each obstacle region calculated in step S360 by referring to the distance values ​​calculated in step S370. For example, coefficients a and b that minimize the squared error may be calculated using Equation 4. The values ​​converted using coefficients a and b are the real-scale distance values ​​of each obstacle region. D1=a×D2+b (Equation 4) D1: Distance value calculated in step S370 D2: Distance value of each obstacle area calculated in step S360

[0070] If the defocus amount is calculated as a distance value in step S370, the distance value of each obstacle area calculated in step S360 may be converted into a defocus amount using Equation 5.

[0071] <Second embodiment> The second embodiment of the present invention will be described in detail below with reference to the drawings. Note that the components described in this embodiment are merely examples, and the scope of the present invention is not limited to the components described in this embodiment.

[0072] (Device configuration) Fig. 9 is a diagram showing a schematic configuration of an image processing device according to an embodiment of the present invention. In Fig. 9, the same parts as those in Fig. 1 are assigned the same numbers as in Fig. 1, and the description thereof will be omitted. This method of omitting the description will also be used in the embodiments described below.

[0073] 9A, a moving object 900 includes an imaging device 910, a vehicle information acquisition device 130, an external environment recognition device 140, a control device 150, and an alarm device 160.

[0074] The mobile object 900 is an object that moves by a power source, and is, for example, an automobile, a ship, an aircraft, a drone, an industrial robot, etc. In the following description, the mobile object 900 is an automobile.

[0075] In FIG. 9B, the imaging device 910 includes an image processing device 911 and an imaging unit 912 .

[0076] The imaging unit 912 includes an imaging element 912-1 and an optical system 912-2. The image processing device 911 can be configured using a logic circuit. Alternatively, the image processing device 911 may be configured with a central processing unit (CPU) and a memory that stores a processing program.

[0077] The optical system 912-2 is a photographing lens of the imaging device 910, and has the function of forming an image of a subject on the imaging element 912-1 (on the imaging element). The optical system 912-2 is composed of a plurality of lens groups (not shown) and an aperture (not shown), and has an exit pupil 923 at a predetermined distance from the imaging element 912-1. Note that in this specification, the z-axis is parallel to the optical axis 930 of the optical system 912-2. Furthermore, the x-axis and y-axis are perpendicular to each other and to the optical axis.

[0078] The image sensor 912-1 is composed of a CMOS (complementary metal-oxide semiconductor) or a CCD (charge-coupled device). The subject image formed on the image sensor 912-1 via the optical system 912-2 is photoelectrically converted by the image sensor 912-1 to generate an image signal based on the subject image.

[0079] 9C is an xy cross-sectional view of the image sensor 912-1. The image sensor 912-1 is configured by arranging multiple pixel groups 914 in a 2-row x 2-column array. The pixel group 914 is configured by arranging green pixels 914G1 and 914G2 in the diagonal direction, and a red pixel 914R and a blue pixel 914B in the other two pixels.

[0080] FIG. 9(D) is a schematic diagram showing the I-I' cross section of the pixel group 914. Each pixel is composed of a light receiving layer 917 and a light guide layer 916. The light receiving layer 917 has two photoelectric conversion units (a first photoelectric conversion unit 915-1 and a second photoelectric conversion unit 915-2) arranged therein for photoelectrically converting received light. In other words, the pixel has multiple photoelectric conversion units. The light guide layer 916 has arranged therein a microlens 918 for efficiently guiding the light beam incident on the pixel to the photoelectric conversion unit, a color filter (not shown) that transmits light of a predetermined wavelength band, wiring (not shown) for image reading and pixel driving, and the like. Each pixel also has wiring (not shown) through which each pixel can send an image signal (output signal) to the image processing device 911. 1C and 1D show examples of a photoelectric conversion unit that is divided into two in one pupil division direction (x-axis direction), but depending on the specifications, an image sensor having a photoelectric conversion unit that is divided into two pupil division directions (x-axis direction and y-axis direction) may be used. The pupil division direction and the number of divisions are arbitrary.

[0081] 10 shows the exit pupil 912-3 of the optical system 912-2 as viewed from the intersection (central image height) of the optical axis 913 and the image sensor 912-1. A first light beam that has passed through a first pupil region 1010, which is a different region of the exit pupil 912-3, and a second light beam that has passed through a second pupil region 1020 are incident on the photoelectric conversion unit 915-1 and the photoelectric conversion unit 915-2, respectively. The photoelectric conversion unit 915-1 and the photoelectric conversion unit 915-2 in each pixel photoelectrically convert the incident light beams to generate image signals corresponding to image A (first image) and image B (second image), respectively. The generated image signals are transmitted to the image processing device 911.

[0082] 10 shows the center of gravity of the first pupil region 1010 (first center of gravity position 1011) and the center of gravity of the second pupil region 1020 (second center of gravity position 1021). In this embodiment, the first center of gravity position 1011 is decentered (moved) from the center of the exit pupil 912-3 along a first axis 1000. On the other hand, the second center of gravity position 1021 is decentered (moved) along the first axis 1000 in the opposite direction to the first center of gravity position 1011. The direction connecting the first center of gravity position 1011 and the second center of gravity position 1021 is called the pupil division direction. The distance between the centers of gravity of the first center of gravity position 1011 and the second center of gravity position 1021 is the base length 1030.

[0083] (Explanation of image processing device) The image processing device 911 will be described below.

[0084] Fig. 11(A) is a diagram schematically showing the configuration of an image processing device 111 according to an embodiment of the present invention. In Fig. 11(A), the image processing device 911 includes a position information acquisition unit 310, a reference surface height calculation unit 320, a threshold determination unit 330, an obstacle determination unit 340, a region integration unit 350, a distance information acquisition unit 1160, and a calibration unit 370. Fig. 3(B) is a flowchart showing the operation of the image processing device 911 according to this embodiment. When image processing according to this embodiment starts, the process proceeds to step S1110.

[0085] In step S1110, an image is captured using the imaging device 910, an image is generated and acquired, and the acquired image is stored in the main memory (not shown). Note that the image acquired may be either image A or image B, or an image obtained by combining the two.

[0086] Furthermore, the image acquired in step S1110 may be subjected to a process for correcting the imbalance in light intensity, which occurs mainly due to vignetting of the optical system 912-2. Specifically, the light intensity balance can be corrected by correcting the luminance value of the image so that it becomes approximately constant regardless of the angle of view, based on the results of the image capture device 910 previously capturing an image of a surface light source with a constant luminance. Furthermore, in order to reduce the influence of optical shot noise generated by the image sensor 912-1, for example, the acquired image may be subjected to filtering using a band-pass filter or a low-pass filter. Alternatively, the image may be reduced in size to reduce calculation costs. Alternatively, the image may be made higher-resolution using a known method in order to determine obstacles at a higher resolution.

[0087] Step S1170 is performed by distance information acquisition unit 1160. Distance values ​​at a real scale are acquired in step S1170. A method for acquiring distance values ​​from images A and B captured by imaging device 910 will be described below.

[0088] First, the block matching method for calculating the parallax from the A image and the B image will be described with reference to FIG.

[0089] FIG. 12(A) shows image A 1210A, and FIG. 12(B) shows image B 1210B. In step S370, first, a partial region including feature point 1220 and its neighboring pixels is extracted from image A 1210A and set as base image 1211. Next, an area having the same area (image size) as base image 1211 is extracted from image B 1210B and set as reference image 1212. Next, the position from which reference image 1212 is extracted is moved on image B 1210B, and the correlation value between reference image 1212 and base image 1211 for each amount of movement (each position) is calculated. This calculates a correlation value consisting of a correlation value data string corresponding to each amount of movement. Finally, the amount of movement that results in the highest correlation is calculated from the correlation value data string, and this amount of movement is calculated as the parallax.

[0090] The correlation value may be calculated using any known method as long as it can evaluate the degree of correlation between the base image 1211 and the reference image 1212. For example, the sum of squared differences (SSD), the sum of absolute differences (SAD), or normalized cross-correlation (NCC) may be used.

[0091] Furthermore, a more detailed parallax may be calculated by performing sub-pixel estimation, which is a known method.

[0092] Next, we will explain how to convert the parallax into the distance (defocus amount) from the image sensor 912-1 to the image formation point by the imaging optical system 912-2. Hereinafter, the coefficient for converting the parallax amount into the defocus amount will be called the BL value. When the BL value is BL, the defocus amount is ΔL, and the parallax amount is d, the parallax amount d can be converted into the defocus amount ΔL using Equation 5. ΔL=BL×d (Formula 5)

[0093] Finally, we will explain how to convert the defocus amount into distance. To convert the defocus amount into the subject distance, we can use the lens formula in geometric optics shown in Equation 6. 1 / A+1 / B=1 / f (Equation 6) A: Distance from the object surface to the imaging optical system 112-2 B: Distance from the principal point of the imaging optical system 112-2 to the image plane f: focal length of imaging optical system 112-2

[0094] In Equation 5, the focal length is a known value. The value of B can be calculated using the defocus amount. Therefore, the distance A to the object surface, i.e., the distance, can be calculated using the focal length and the defocus amount.

[0095] If the imaging optical system is the same, the defocus amount and the distance have a unique relationship, and therefore the defocus amount may be calculated as a distance value.

[0096] <Modification> The imaging device 913 may have the configuration shown in Fig. 13. The imaging section 1320 includes two imaging elements 1321 and 1322 and two optical systems 1323 and 1324. The optical systems 1323 and 1324 are photographing lenses of the imaging device 1300, and have the function of forming an image of a subject on the imaging element 1321 or 1322. The optical systems 1323 and 1324 are composed of a plurality of lens groups (not shown), an aperture (not shown), etc., and have exit pupils 1325 and 1326 at positions separated by a predetermined distance from the imaging element 1321 or 1322. In this case, the optical axes of the optical systems 1323 and 1324 are 1341 and 1342, respectively.

[0097] By calibrating parameters such as the positional relationship of the optical systems in advance, the distance can be calculated accurately in step S1170. In addition, the distance can be calculated accurately by correcting lens distortion in each optical system.

[0098] In FIG. 13, there are two optical systems that capture images A and B, each with a parallax according to distance, but it may also be configured with a stereo camera that is made up of three or more optical systems and corresponding image sensors.

[0099] In the image processing device of this embodiment, the distance acquisition device is not required, and therefore the device can be made smaller.

[0100] The present invention encompasses not only a distance measuring device but also a computer program. The computer program of this embodiment causes a computer to execute predetermined processes to calculate distance or parallax. The program of this embodiment is installed in a computer of a distance measuring device or an imaging device such as a digital camera equipped with the same. The installed program is executed by the computer to realize the above functions, enabling high-speed and high-precision calculation of parallax.

[0101] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]

[0102] 100 Mobile 110 Imaging device 120 Distance acquisition device 130 Vehicle information acquisition device 140 External world recognition device 150 control device 160 Alarm device

Claims

1. An image processing device installed in a moving body, image acquisition means for acquiring an image; an information acquisition means for detecting a feature point in the image and acquiring height information indicating the height of the feature point on the image and depth information indicating the distance to the feature point; a reference surface height calculation means for calculating the height of a predetermined reference surface as a reference surface height based on the height information; a determination means for determining whether the feature point indicates an object that will be an obstacle to the movement of the moving object, using the reference plane height; 1. An image processing device comprising:

2. the determining means determines whether the object is an obstacle to the movement of the moving body based on a ratio between the height information and the depth information, the reference plane height, and a threshold value.

2. The image processing device according to claim 1, wherein:

3. a threshold determining means; the threshold value determining means determines the threshold value based on a ratio between a height of a detection target and a depth to the detection target.

3. The image processing device according to claim 2.

4. an extraction means for extracting the feature points that meet predetermined conditions from the image and generating an extracted point group; a calculation means for calculating the predetermined reference plane height from the height information of each of the extracted points constituting the extracted point group, 3. The image processing device according to claim 2.

5. the extraction means vertically or horizontally divides the image into two or more regions and extracts the feature points from each of the regions; 5. The image processing device according to claim 4.

6. The extraction means extracts the feature points whose height information does not satisfy a predetermined value as the predetermined condition.

5. The image processing device according to claim 4.

7. the extraction means extracts the feature points located at pixel positions corresponding to the reference plane in the image as the predetermined condition; 5. The image processing device according to claim 4.

8. the calculation means calculates, as the reference plane height, one of an average value, a median value, and a mode value of the height information of each of the feature points constituting the extracted point group.

5. The image processing device according to claim 4.

9. the calculation means fits the extracted point group to a predetermined function and calculates the function as the reference surface height; 5. The image processing device according to claim 4.

10. the determining means determines that the object is an obstacle to movement of the moving object when a ratio of the height information to the depth information is greater than the threshold value.

3. The image processing device according to claim 2.

11. the determining means determines that the object is an obstacle to movement of the moving body when a ratio of the height information to the depth information is smaller than the threshold value and the height information is lower than the reference plane height.

3. The image processing device according to claim 2.

12. the determining means determines that the object is an obstacle to movement of the moving object when a ratio of the height information to the depth information is outside a range between an upper limit value and a lower limit value of the threshold value.

3. The image processing device according to claim 2.

13. the determining means changes a value indicating the height information based on the reference surface height.

3. The image processing device according to claim 2.

14. The determination means changes the threshold value based on the reference surface height.

3. The image processing device according to claim 2.

15. the information acquisition means limits acquisition of the height information and the depth information located at a predetermined pixel in the image.

3. The image processing device according to claim 2.

16. the information acquisition means calculates reliability from at least one of the height information and the depth information and information of an optical system that captured the image, and when the reliability is lower than a predetermined value, restricts acquisition of the height information and the depth information.

3. The image processing device according to claim 2.

17. A region integration means is provided, the region integration means integrates the two or more feature points determined to be objects that obstruct the movement of the moving body as the same object when the depth information of the two or more feature points is within a predetermined range; 16. The image processing device according to claim 1,

18. having an imaging means; 16. An imaging device comprising the image processing device according to claim 1.

19. the imaging means includes an optical system and an imaging element; the optical system forms an image of a subject on the imaging element; the imaging element includes a plurality of first photoelectric conversion units for generating a first image and a plurality of second photoelectric conversion units for generating a second image; 19. The imaging device according to claim 18.

20. The imaging means a first imaging element; a first optical system that forms an image of a subject on the first image sensor; Equipped with acquiring a first image with the first image sensor; 19. The imaging device according to claim 18.

21. The imaging means a first imaging element; a first optical system that forms an image of a subject on the first image sensor; a second imaging element; a second optical system that forms an image of a subject on the second image sensor; Equipped with acquiring a first image with the first image sensor and acquiring a second image with the second image sensor; 19. The imaging device according to claim 18.

22. A moving body that can move by itself and that is equipped with the imaging device according to claim 18, A control device is provided. The control device controls the moving body based on the determination by the determination means. A moving object characterized by:

23. An image processing method for a CPU installed in a moving body to determine whether an object is an obstacle to the movement of the moving body, comprising: image acquisition means for acquiring an image; an information acquisition means for detecting a feature point in the image and acquiring height information indicating the height of the feature point on the image and depth information indicating the distance to the feature point; a reference surface height calculation means for calculating the height of a predetermined reference surface as a reference surface height based on the height information; a determination means for determining whether the feature point indicates an object that will be an obstacle to the movement of the moving object, using the reference plane height; the determining means determines whether the object is an obstacle to the movement of the moving body based on a ratio between the height information and the depth information, the reference plane height, and a threshold value. An image processing method comprising:

24. A program for causing a computer to execute each step of the image processing method according to claim 23.

Citation Information

Patent Citations

  • Obstacle detection device

    JP2018156222A

  • Obstacle detection device, obstacle detection method and obstacle detection program

    JP2023039777A