Information processing device for acquiring distance and method for controlling the same

By integrating split-pupil imaging and SfM methods, the device corrects for environmental-induced errors in distance measurement, enhancing accuracy and reliability in distance estimation.

JP2025149761APending Publication Date: 2025-10-08CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024050588
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-10-08

AI Technical Summary

Technical Problem

Existing distance measurement devices, particularly those using a split-pupil image plane phase difference method, suffer from measurement errors due to changes in the surrounding environment, such as temperature and shock, which affect the baseline length and optical properties, leading to inaccurate distance calculations.

Method used

The device employs a combination of first and second distance information acquisition methods, utilizing a split-pupil imaging plane phase difference method for initial distance measurement and Structure from Motion (SfM) for error reduction, generating correction values by comparing the ratio of image-side defocus amounts to correct for temporal measurement errors.

Benefits of technology

This approach significantly reduces distance measurement errors caused by environmental changes, ensuring accurate and reliable distance calculations over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025149761000001_ABST
    Figure 2025149761000001_ABST
Patent Text Reader

Abstract

To solve such a problem that a temporal range-finding error occurs because an imaging apparatus is affected by an ambient environment such as heat and an impact.SOLUTION: First distance information including an error about a distance between an imaging part and a subject is acquired through an optical system, and second distance information whose error is smaller than that of the first distance information is acquired. A first correction value for correcting a first defocus amount is generated on the basis of a ratio between the first defocus amount corresponding to deviation in an optical axis direction between a sensor surface and an image forming surface used for the acquisition of the first distance information and a second defocus amount when obtaining the second distance information, and the distance between the imaging part and the subject is calculated by using the first defocus amount corrected by the correction value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a device having a distance measurement function that is used in digital cameras, digital video cameras, in-vehicle sensor devices, robot vision sensor devices, etc., and a control method for the same. [Background technology]

[0002] Imaging devices have been proposed that are equipped with a distance measurement function that calculates the parallax of a subject based on images captured from different viewpoints and can obtain distance information such as the distance to the subject or the defocus state from the calculated parallax. These include stereo cameras consisting of at least two cameras, and split-pupil image plane phase difference ranging cameras that are composed of a single camera and can obtain parallax images by receiving each light beam that passes through different pupil regions of the optical system. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2014-52335 Summary of the Invention [Problem to be solved by the invention]

[0004] In imaging devices with such distance measurement functions, the value of the baseline length changes due to influences from the surrounding environment such as heat and shock, resulting in distance measurement errors over time. Patent Document 1 discloses a method in which the expansion coefficient with respect to temperature of a jig that maintains the camera spacing, which is the baseline length of a stereo camera, is acquired and stored in advance, and the amount of change in the length of the jig, i.e., the amount of change in the baseline length, is calculated and corrected.

[0005] However, the technique in Patent Document 1 can only be used to correct the baseline length in the case of a stereo camera, where the baseline length is defined by the physical distance between the cameras, and cannot be used in the case of a rangefinder camera that uses a split-pupil image plane phase difference method. This is because in the case of a split-pupil image plane phase difference rangefinder camera, the baseline length is not defined by the physical spatial length but by the distance between the pupil regions through which each light beam passes.

[0006] In view of the above-mentioned problems, an object of the present invention is to provide an information processing device capable of acquiring distance information from a captured image, which is capable of reducing distance measurement errors that occur over time. [Means for solving the problem]

[0007] The information processing device according to the present invention has the following configuration: first acquisition means for acquiring, via an optical system, first distance information including an error regarding the distance between an imaging unit and a subject, second acquisition means for acquiring second distance information having a smaller error than the first distance information, generation means for generating a first correction value for correcting the first defocus amount based on the ratio between a first defocus amount corresponding to a deviation in the optical axis direction between a sensor plane and an imaging plane used to acquire the first distance information and a second defocus amount used when the second distance information was acquired, and calculation means for calculating the distance between the imaging unit and the subject using the first defocus amount corrected by the correction value. [Effects of the Invention]

[0008] According to the present invention, in a distance measurement device capable of acquiring distance information from a captured image, it is possible to reduce distance measurement errors that occur over time due to the influence of the surrounding environment, such as heat and shock. [Brief explanation of the drawings]

[0009] [Figure 1] Example of distance acquisition device configuration [Figure 2] Details of the camera device, image sensor, and unit pixel [Figure 3] Relationship between unit pixel and exit pupil [Figure 4] Schematic diagram of the image sensor and imaging optical system [Figure 5] Relationship between parallax and defocus amount [Figure 6] Diagram explaining how to obtain distance information using the SfM method [Figure 7] FIG. 10 is a diagram illustrating generation of correction information using a defocus amount. [Figure 8] Flowchart explaining the processing flow [Figure 9] Relationship between parallax and defocus amount [Figure 10] FIG. 10 is a diagram illustrating generation of correction information using a defocus amount. [Figure 11] Flowchart explaining the processing flow [Figure 12] Example of a vehicle equipped with a distance acquisition device DETAILED DESCRIPTION OF THE INVENTION

[0010] The present invention will be described in detail using embodiments and drawings. The present invention is not limited to the contents described in each example. In addition, each embodiment may be combined as appropriate. In the description with reference to the drawings, even if the figure numbers are different, the same reference numerals will be used to indicate the same parts, and duplicate explanations will be avoided.

[0011] Example 1 1 shows an example of the configuration of a distance acquisition device 100. The distance acquisition device 100 will be described as an information processing device that acquires distances. It includes a camera device 101, a stereo distance measurement calculation device 102, a feature point distance measurement calculation device 103, and a correction value calculation device 104. The camera device 101 can also be configured as an external device of the distance acquisition device 100.

[0012] In the first distance information acquisition process shown in Fig. 8, the stereo distance measurement calculation device 102 calculates image signals acquired by the camera device 101 to acquire the distance from the camera device to the subject (first distance information). In the second distance information acquisition process, the feature point distance measurement calculation device 103 calculates image signals acquired by the camera device 101 to acquire the distance from the imaging unit of the camera device to the subject (second distance information). A correction value calculation device 104 calculates a correction value based on the acquired first distance information and second distance information, and outputs a distance measurement value with distance measurement errors corrected over time based on the calculated correction value. Hereinafter, the distance from the camera device to the subject will also be referred to as the subject distance.

[0013] A camera device 101, which is a distance measuring camera using a split-pupil imaging plane phase difference method, will be described with reference to Fig. 2. Fig. 2(a) shows the configuration of the camera device 101. The camera device 101 includes an optical system 201, an image sensor 202, an image storage memory 203, and a signal transmission unit 204.

[0014] An image signal is obtained by photoelectrically converting an image of a subject formed on an image sensor 202 via an optical system 201. The acquired image signal is stored in an image storage memory 203 and transmitted to the outside of the camera device 101 by a signal transmission unit 204. In this embodiment, the z axis is set parallel to an optical axis 205 of the imaging optical system 201, and the x and y axes are set perpendicular to each other and perpendicular to the optical axis.

[0015] The image sensor 202 is an image sensor that can acquire a set of images from different viewpoints (hereinafter referred to as parallax images). The image sensor 202 is shown in detail in FIG. 2(b). FIG. 2(b) is an xy cross-sectional view of the image sensor 202.

[0016] The image sensor 202 is configured by arranging a plurality of unit pixels 210 in each of the x and y directions. The light receiving layer of the unit pixel 210 has two photoelectric conversion units, a first photoelectric conversion unit 211 and a second photoelectric conversion unit 212. FIG. 2(c) schematically shows the I-I' cross section of the unit pixel 210. The unit pixel 210 is configured from a light guide layer 221 and a light receiving layer 222. The light guide layer 221 is provided with a microlens 223 for efficiently guiding a light beam incident on the unit pixel to the photoelectric conversion unit, a color filter (not shown) that transmits light of a predetermined wavelength band, and wiring (not shown) for image reading and pixel driving. The light receiving layer 222 is provided with two photoelectric conversion units, a first photoelectric conversion unit 211 and a second photoelectric conversion unit 212, for photoelectrically converting received light. An imaging element having such a unit pixel structure can acquire a parallax image consisting of a pair of a first image and a second image taken from different viewpoints using a configuration consisting of one optical system and one imaging element.

[0017] The principle of such a pupil division imaging plane phase difference method will be explained with reference to FIGS.

[0018] FIG. 3( a) is a diagram showing the relationship between a unit pixel near the central image height as a representative example of the unit pixels 210 in the image sensor 202 and its exit pupil 301 in the optical system. The microlens 223 in the unit pixel 210 is arranged so that the exit pupil 301 and the light receiving layer 222 are optically conjugate. As a result, a light beam that passes through a first pupil region 311 on the exit pupil 301 is incident on the first photoelectric conversion unit 211. Similarly, a light beam that passes through a second pupil region 312 is incident on the second photoelectric conversion unit 212. As shown in FIG. 3( b), even if the unit pixel is located at the peripheral image height, the chief ray is tilted and is obliquely incident on the microlens and the light receiving layer, but the corresponding relationships between the first and second pupil regions, light beams, and photoelectric conversion units do not change. The image sensor 202 is configured with a plurality of unit pixels 210 arranged on the same plane, and a signal photoelectrically converted by a first photoelectric conversion unit 211 of each unit pixel is read out to generate a first image of a first viewpoint. Similarly, a signal photoelectrically converted by a second photoelectric conversion unit 212 of each unit pixel is read out to generate a second image of a second viewpoint. In this way, a parallax image can be obtained, which is a plurality of signals having parallax that are output by receiving light beams that have passed through different pupil regions of a monocular optical system at a single image sensor.

[0019] The amount of parallax between the first image signal and the second image signal corresponds to the amount of defocus, which is the amount of deviation from the focal point on the image sensor. The relationship between the amount of parallax and the amount of defocus will be explained using FIG. 4. FIGS. 4(a), (b), and (c) are schematic diagrams showing the image sensor 202 and the imaging optical system 201. In the diagram, 401 denotes a first light beam passing through the first pupil region 311, and 402 denotes a second light beam passing through the second pupil region 312.

[0020] FIG. 4(a) shows the in-focus state, where a first light beam 401 and a second light beam 402 emitted from an object 400 located at a focus position 403 converge on the image sensor 202. At this time, the relative positional shift between the first image signal formed by the first light beam 401 and the second image signal formed by the second light beam 402, i.e., the parallax, is zero. FIG. 4(b) shows the state in which the object 400 is located farther away from the focus position 403 and is defocused in the negative direction of the z-axis on the image side. At this time, the relative positional shift in the x-axis between the first image signal formed by the first light beam and the second image signal formed by the second light beam is not zero but is a negative value. FIG. 4(c) shows the state in which the object 400 is located closer to the focus position 403 and is defocused in the positive direction of the z-axis on the image side. At this time, the relative positional shift between the first image signal formed by the first light beam and the second image signal formed by the second light beam is not zero but is a positive value.

[0021] As shown in Figures 4(a), (b), and (c), a parallax proportional to the defocus amount occurs between the first and second light beams received on the image sensor, and the sign of the parallax is reversed depending on whether the defocus amount is positive or negative. Therefore, the amount of parallax between the first image signal and the second image signal can be detected, and the detected amount of parallax can be converted into the amount of defocus using a predetermined conversion coefficient.

[0022] A known method is used to detect the amount of parallax. For example, the correlation value between the first image signal and the second image signal is calculated using SSD (Sum of Squared Difference) or the like, and the amount of parallax can be detected with sub-pixel accuracy by function approximation of the minimum cost. If the amount of parallax is r and the amount of defocus is d, then the conversion coefficient k is used to calculate (Equation 1) d=kr The parallax can be converted into the defocus amount by using the following equation. Here, the conversion coefficient k can be obtained in advance by calibration, such as by measuring the parallax amount r and defocus amount d for a known distance. If the detected defocus amount d is the focal length f and the subject distance z, then the following can be obtained using the imaging formula: (Equation 2) 1 / z=1 / f-1 / d The distance value of the image pickup surface phase coarse ranging method, which is the first distance information, is detected by converting the distance to the subject from the above relationship.

[0023] Here, we will explain again the distance measurement error in a distance measuring camera using the pupil division image plane phase difference method. Figure 5 shows the relationship between the amount of parallax and the amount of defocus. Since the parallax between each light beam is defined by the interval between the centroid rays of each light beam, Figure 5(a) shows the centroid ray 501 of the first light beam and the centroid ray 502 of the second light beam in a certain defocus state. The figure shows the defocus amount d, which is the distance in the optical axis direction from the image sensor (sensor plane) 202 to the focal point (imaging plane) 500, and the parallax r, which is the distance from the centroid ray 501 of the first light beam to the centroid ray 502 of the second light beam on the image sensor. The figure also shows the base length w, which is the distance from the centroid ray 501 of the first light beam to the centroid ray 502 of the second light beam on the exit pupil 301, and the exit pupil distance p, which is the distance from the image sensor to the exit pupil. These are determined by the geometric relationship as follows: (Equation 3) (p+d):w=d:r Here, the defocus amount d is generally a displacement amount on the order of micromillimeters (um), while the exit pupil distance p is on the order of millimeters (mm), so p + d ≒ p, (Equation 4) d=kr (where k=p / w) As shown in FIG. 5B, the parallax r and the defocus amount d have a linear relationship 511 with a slope k expressed by the exit pupil distance p and the base line length w.

[0024] At this time, due to changes in the temperature of the surrounding environment or external impacts, the positions of the lenses forming the optical system change, as well as changes in optical properties such as refractive indices, that is, over time, the exit pupil distance p and the baseline length w, which is the distance between each chief ray, change from the calibrated state. As a result, the value of the tilt k changes to form a linear relationship 512, resulting in a ranging error. That is, when the parallax amount r1 is detected, in the calibrated state, the defocus amount d1 is detected according to the linear relationship 511, but due to the error caused by changes over time, it is detected as the defocus amount d2 in the case of the linear relationship 512. When this defocus amount d2 is converted into the subject distance using the previous formula 2, a distance value different from the calibrated state, that is, a ranging error occurs. Since the change in the conversion coefficient k due to changes over time is caused by complex factors such as the aforementioned changes in lens position and optical properties, it is difficult to prepare, as disclosed in Patent Document 1, the previous characteristic changes as a conversion coefficient table. Therefore, as shown in this embodiment, correction information is generated by comparing as the ratio of the image-side defocus amounts.

[0025] The second distance information acquisition process acquires the second distance information by calculating the image signal acquired by the camera device 101 in the feature point ranging arithmetic unit 103. The image signal S1 acquired by the camera device 101 at time t1 and the image signal S₂ acquired at time t2 are read from the image storage memory 203 through the signal transmission unit 204 and calculated in the feature point ranging arithmetic unit 103. The relationship between the times is t1 < t2, and t1 is in a time series prior to t2. Here, the image signals S1 and S₂ are image signals generated by a light beam that has passed through the entire region of the exit pupil obtained by adding the signals read from the aforementioned first photoelectric conversion unit 211 and the second photoelectric conversion unit 212.

[0026] In the second distance information acquisition process, the second distance information is acquired using the well-known SfM (Structure from Motion) method. Specifically, feature points (e.g., SIFT feature points) are calculated in each image using a well-known method, and the calculated feature points are associated with each other using a well-known method to calculate an optical flow and calculate the distance to the subject. This process will be explained using FIGS. 6(a), 6(b), and 6(c). The feature points are calculated using the well-known Harris corner detection algorithm for the acquired image signals S1 and S2. The feature points 601 calculated for the image signal S1 at time t1 are indicated by stars in FIG. 6(a), and the feature points 602 calculated for the image signal S2 at time t2 are indicated by stars in FIG. 6(b). FIG. 6(c) shows an optical flow 603 calculated by associating the feature points between the calculated image signals S1 and S2 using the well-known Kanade-Lucas-Tomasi (KLT) feature tracking algorithm. The algorithms used to calculate feature points, feature amounts, and optical flow are not limited to the methods described here. It is also preferable to use FAST (Features from Accelerated Segment Test), BRIEF (Binary Robust Independent Elementary Features), ORB (Oriented FAST and Rotated BRIEF), etc.

[0027] The calculated optical flow 603 is used to calculate the distance to the subject using a known method. The camera fundamental matrix F is obtained using the 8-point algorithm to satisfy the epipolar constraint using the feature point 601 at time t1, the feature point 602 at time t2, and the optical flow 603 representing their correspondence. It is also preferable to use the RANSAC (Random Sample Consensus) method to efficiently remove outliers and perform calculations using a stable method. The camera fundamental matrix F is decomposed into the camera fundamental matrix E using a known method, and the camera extrinsic parameters, rotational movement R(ωx,ωy,ωz) and translational movement T(tx,ty,tz), are calculated from the camera fundamental matrix E. The calculated camera extrinsic parameters are relative displacements of the camera movement from time t1 to time t2, and the scaling is indefinite. Therefore, the translational movement T(tx,ty,tz) in particular is a normalized relative value. This is scaled to obtain the actual translational movement T(tx,ty,tz). Specifically, the amount of translational movement T, that is, the amount of movement tz parallel to the optical axis of the camera, is obtained from the difference between the distance information for image S1 at time t1 and the distance information for image S2 at time t2, both of which are obtained using the image plane phase difference ranging method, which is the first distance information acquisition process described above. Other components are also scaled from the obtained actual movement tz, which has also been scaled, to obtain the actual translational movement T(tx, ty, tz) from time t1 to t2. Then, distance information z is detected from Equation 5 and Equation 6, which are well-known relationships.

[0028]

number

[0029]

number

[0030] Here, the coordinates in the image coordinate system of the object whose distance is to be calculated are (u, v), the optical flow of the object whose distance is to be calculated is (Δu, Δv), and the distance to the object is z. The camera movements between images used to calculate the optical flow are the rotational movement (ωx, ωy, ωz), the translational movement (tx, ty, tz), and the focal length of the camera f.

[0031] The feature point ranging calculation device 103 converts the distance from the camera at each coordinate on the image of the feature point 602 of the image signal S2 at time t2 thus acquired into the image-side defocus amount using equation 2, thereby acquiring second distance information.

[0032] The second distance information thus obtained is calculated from the optical flow that associates images obtained at a time interval much shorter than the time interval at which time-dependent changes occur, and therefore does not contain any error over time due to changes in the surrounding environment, such as temperature changes in the surrounding environment or external impacts.

[0033] Here, the method for scaling the camera movement amount is not limited to this method, and it is also suitable to perform scaling by obtaining the camera movement amount from various measuring devices, specifically an IMU (Inertial Measurement Unit) or a GNSS (Global Navigation Satellite System), or in the case of an in-vehicle camera, vehicle speed information or map information.

[0034] It is also suitable to use bundle adjustment, a well-known method, to calculate the camera movement amount and the positional relationship between the subject and the camera. The relationships between variables such as the camera fundamental matrix and optical flow, including internal camera parameters such as focal length, can be analytically calculated using the nonlinear least squares method to achieve good consistency.

[0035] Of the feature points used to calculate the camera movement amount, it is also preferable to exclude from the processing feature points calculated from a subject that is not a stationary object with respect to the world coordinate system to which the imaging device belongs. Known methods for estimating camera movement amount calculate various parameters assuming the subject is a stationary object, which can become a source of error if the subject is a moving object. Therefore, excluding feature points calculated from moving objects can improve the accuracy of calculating various parameters. Moving objects are determined by classifying the subject using image recognition technology or by comparing the relative value of the time-series change in the acquired distance information with the movement amount of the imaging device.

[0036] Figure 7 shows how the correction value calculation device 104 compares the first distance information acquired in the first distance information acquisition process with the second distance information acquired in the second distance information acquisition process as a ratio of the image-side defocus amount to generate a correction value.

[0037] FIG. 7(a) shows the defocus amount D1 when first distance information is acquired by the first distance information acquisition process at time t2. FIG. 7(b) shows the defocus amount D2 when second distance information is acquired by the second distance information acquisition process at time t2. Because distance information is acquired at the pixel position corresponding to the feature point 602 in FIG. 7(b), the data is a sparse group of data corresponding to the coordinates of the feature point 602 on the image. FIG. 7(c) shows the image-side defocus amount along I-I' in the figure. The discontinuous line segment (i) in FIG. 7(c) represents the defocus amount D1, and the point data (ii) in FIG. 7(c) represents the defocus amount D2. The defocus amount D1 includes errors due to changes in the camera device 101 over time and corresponds to the linear relationship 512 in FIG. 5(b). On the other hand, the defocus amount D2 corresponds to the linear relationship 511 in FIG. 5(b), which is not affected by changes in the camera device 101 over time. Therefore, when the ratio between the two is calculated, the value corresponds to the change in the slope coefficient k affected by changes over time. Linear relationship 511: d1 = k1·r1, and linear relationship 512: d2 = k2·r1, resulting in d1 / d2 = k1 / k2 (the change in the slope coefficient k). Figure 7(d) shows the change = (i) / (ii), obtained by dividing (i), the defocus amount D1 in Figure 7(c), by (ii), the defocus amount D2. The ratio data 702 of the change = (i) / (ii) can actually be acquired at the data acquisition coordinates 701 shown in Figure 7(d), which are the coordinates of the feature point 602. In contrast, the change affected by changes over time occurs across the entire angle of view. The change in slope between angles of view, i.e., between pixels, is smoothly connected, even considering that the optical characteristics change continuously. Therefore, the angle of view can be interpolated by fitting the ratio data 702 acquired at each data acquisition coordinates 701 using polynomial approximation. The amount of change 703 thus obtained from polynomial approximation is shown by the dashed line in FIG. 7(d). For the sake of explanation, the amount of change due to polynomial approximation is expressed one-dimensionally along the line segment I-I', but the actual amount of change on the image side is two-dimensional data on the xy plane. Therefore, using the acquired difference data that is discrete with respect to the angle of view, surface fitting is performed using polynomial approximation on the xy plane to estimate the amount of change. The approximate surface data that is the calculated amount of change is taken as the correction value kc. Using the correction value kc, k', which is the conversion coefficient k corrected as described above, is calculated as (Equation 7) k'=k / kc Then, the correction value calculation device 104 calculates the relationship between the corrected parallax r and the defocus amount d′ as follows: (Equation 8) d′=k′r By using the corrected defocus amount d' in this way to obtain the subject distance z using Equation 2, it is possible to calculate a distance value that is less affected by distance measurement errors due to changes over time.

[0038] FIG. 8 shows a calculation flow according to this embodiment. FIG. 8(a) shows the overall flow. The following flowcharts are assumed to be realized by the CPU executing a control program. A first distance information acquisition process S810 is performed based on parallax images acquired by the camera device 101. As shown in FIG. 8(b), the first distance information acquisition process S810 performs light intensity correction and noise reduction in preprocessing S811. Then, parallax is detected using the aforementioned method in parallax amount detection process S812, and the defocus amount d and subject distance z are calculated in distance conversion process S813 as described above. In the overall flow shown in FIG. 8(a), a second distance information acquisition process S820 is performed based on images acquired by the camera device 101 at times t1 and t2. As shown in FIG. 8(c), the second distance information acquisition process S820 performs the aforementioned feature point detection and association in optical flow calculation process S811 to calculate optical flow. In SfM processing S822, the subject distance at each feature point is calculated from the optical flow and camera movement amount, and second distance information is calculated. In the overall flow of FIG. 8(a), a correction value is calculated in correction information calculation processing S330 based on the calculated first distance information and second distance information. As shown in FIG. 8(d), in change amount calculation processing S831, the ratio of the defocus amount is calculated as the change amount, as described above. In correction information calculation processing S832, surface fitting is performed, and a correction value k' is calculated by correcting the conversion coefficient k used during calibration. In the overall flow of FIG. 8(a), correction processing as shown in Equation 8 is performed based on the correction value k' calculated in correction processing S840.

[0039] In this way, the first distance information obtained by the image plane phase difference ranging method and the second distance information obtained by the SfM method are compared as a ratio of the image-side defocus amount, a correction value is calculated, and the defocus amount is corrected. By obtaining the subject distance using the corrected defocus amount, it is possible to provide a distance acquisition device that reduces ranging errors due to changes over time.

[0040] The method used by the feature point ranging calculation device 103 is not limited to the SfM method described in this embodiment. It is also preferable to use LiDAR (Light Detection and Ranging) or millimeter-wave radar, which acquires distance information based on the reflection intensity of irradiated electromagnetic waves, as a ranging method that is not affected by ranging errors due to changes over time in the camera device 101. In this case, it is also preferable to detect the amount of movement of the camera device 101 using an IMU (Inertial Measurement Unit) or the like, and switch to and select the LiDAR or millimeter-wave radar as the sensor for obtaining the second distance information, since a small amount of movement results in a large error in the optical flow. Furthermore, the feature point ranging calculation device 103 can also calculate the subject distance using an image obtained from a camera device with a configuration other than the camera device 101.

[0041] Example 2 This section explains a correction method for when errors due to changes over time exist not only in the tilt component but also in the offset component in the relationship between parallax and defocus amount. Figure 9 shows the relationship between parallax r and defocus amount d when an offset component error also exists. Compared to the linear relationship 901 at the time of calibration, changes in conversion coefficient k due to changes over time cause a tilt component ranging error, resulting in a linear relationship 902 in which an offset component ranging error b also occurs. Offset component b corresponds to a change in the focal point on the image sensor due to changes in lens position and optical characteristics caused by changes over time, i.e., the amount of field curvature. This section explains a method for calculating the correction value for offset component b. As described above, first distance information and second distance information are acquired, and the two are compared as the difference in image-side defocus amount to calculate the correction value. This is shown in Figure 10.

[0042] FIG. 10(a) shows the defocus amount D1 when the first distance information is acquired by the first distance information acquisition process at time t2. FIG. 10(b) shows the defocus amount D2 when the second distance information is acquired by the second distance information acquisition process at time t2. As described above, these are shown in FIG. 10(c) as the image-side defocus amount along the line I-I' in the figure. The discontinuous line segment (i) in FIG. 10(c) represents the defocus amount D1, and the point data (ii) in FIG. 7(c) represents the defocus amount D2. The defocus amount D1 includes an error due to changes in the camera device 101 over time, as shown by the linear relationship 902 in FIG. 9. On the other hand, the defocus amount D2 is not affected by changes in the camera device 101 over time, as shown by the linear relationship 901 in FIG. 9. Therefore, when the difference between the two is calculated, the resulting value corresponds to the offset error amount b, which is affected by changes over time. However, depending on the subject distance, i.e., the value of the defocus amount, there may also be an error in the slope component, so this does not necessarily match b. This will be discussed later. FIG. 10(d) shows the difference value (i)-(ii) obtained by subtracting the defocus amount D2 (ii) from the defocus amount D1 (i) in FIG. 10(c). The difference data 1002 of the difference value (i)-(ii) can actually be acquired on the discrete data acquisition coordinates 1001. The amount of change 1003 calculated by fitting this with a polynomial approximation is shown by the broken line in FIG. 10(d). Similarly, surface fitting with a polynomial approximation on the xy plane is performed to estimate the amount of in-plane change. The approximate surface data, which is the calculated amount of change, is set as the correction value bc. Using the correction value bc, (Equation 9) d´´=kr-bc By doing so, the parallax r with the offset component error reduced is converted into the defocus amount d''.

[0043] Here, in the relationship between the parallax r and the defocus amount d shown in Fig. 9, there is also an error in the tilt component in the linear relationship 902. In this embodiment, the error in the tilt component and the error in the offset component described in the first embodiment are corrected alternately, and both errors are reduced by performing an iterative process.

[0044] The processing flow is shown in FIG. 11. In the first distance information acquisition process S1110 and the first distance information acquisition process S1120, first distance information and second distance information are acquired in the same manner as described above. In the correction information generation process S1130, a correction value kc for the tilt component is calculated in the same manner as described above, and in the correction process S1140, the tilt component is corrected in the same manner as described above. In the judgment selection process S1150, it is determined whether the correction of the ranging error is sufficient. This determination is made based on whether a preset upper limit on the number of corrections has been reached or whether a preset upper limit on the processing time has been reached. If the set upper limit has not been reached, the offset component error b is selected as the correction target, and the process returns to the first distance information acquisition process S1110. Then, in the correction information generation process S1130, a correction value bc for the offset component is calculated in the same manner as described above, and in the correction process S1140, the offset component is corrected in the same manner as described above. If the upper limit has not been reached in the determination and selection process S1150, the correction target is again selected as the tilt component error k, and the process returns to the first distance information acquisition process S1110. This correction is repeated multiple times until the upper limit is reached in the determination and selection process S1150, thereby reducing both error components through iteration processing. This makes it possible to provide a distance acquisition device that reduces distance measurement errors that change over time even when an offset component error occurs.

[0045] In the judgment and selection process S1150, it is also preferable to judge whether the correction is sufficient not based on the upper limit of the number of times or time, but based on whether the calculated distance measurement value for a calibration target whose distance is known is smaller than a preset threshold value of distance error.In addition, it is also preferable to end the correction process if the difference between the defocus amount obtained in this correction and the defocus amount obtained in the previous correction is equal to or less than a predetermined value (almost converged).

[0046] In the iteration process, starting with the tilt component k as the first correction target is preferable because the effect of the tilt component error is large when the value of k is large, when the distance value of the data acquisition coordinate 1001 deviates from the focus, or when the offset component error b is small. Since the tilt k is large when the baseline length is long, if the designed camera has a baseline length that is longer than normal, starting with the correction of the tilt k will result in faster convergence.

[0047] Furthermore, in the iteration process, starting the first correction target with the offset component b is preferable because when the value of k is small, when the distance value of the data acquisition coordinate 1001 is near the focus point, and when the error b of the offset component is large, the effect of the error in the offset component is large. In camera design where focus changes due to temperature are likely to occur, starting with the correction of the offset component will result in even faster convergence.

[0048] (Basic configuration, overall configuration) Fig. 12(a) is a schematic diagram of the device configuration according to the embodiment mounted on a vehicle 1200, which is a moving body, and Fig. 12(b) is a block diagram thereof. The moving body is not limited to automobiles, but may be various public transportation means such as trains and airplanes, or various robots such as small mobility vehicles and AGVs (Automatic Guided Vehicles).

[0049] 12, a vehicle 1200 includes an imaging device 1210, a millimeter-wave radar device 1220, a LiDAR device 1230 (Light Detection and Ranging: LiDAR), a vehicle information measuring instrument 1240, a route generation ECU 1250 (Electronic Control Unit: ECU), and a vehicle control ECU 1260. As an alternative to the route generation ECU 1250 and the vehicle control ECU 1260, they may be configured with a central processing unit (CPU) and a memory that stores a processing program. The imaging device 1210 is, for example, the camera device shown in FIG. 2.

[0050] The imaging device 1210 that acquires the first distance information captures an image of the surrounding environment including the road on which the vehicle 1200 is traveling, generates image information representing the captured image, and distance image information having information representing the distance to the subject for each pixel, and outputs the information to the route generation ECU 1250. The imaging device 1210 is placed near the top edge of the windshield of the vehicle 1200 as shown in Fig. 12, and captures an image of an area in a predetermined angular range (hereinafter referred to as the imaging angle of view) facing forward of the vehicle 1200.

[0051] The information representing the distance to the subject may be information that can be converted into the distance from the image capture device 1210 to the subject within the imaging angle of view, and may be information that can be converted using a predetermined lookup table or a predetermined conversion coefficient and conversion formula. For example, the distance value may be assigned to a predetermined integer value and output to the route generation ECU 1230. Alternatively, information that can be converted into an optically conjugate distance value (the distance from the image capture element to the conjugate point (so-called defocus amount) or the distance from the optical system to the conjugate point (the distance from the image-side principal point to the conjugate point)) that can be converted into the distance to the subject may be output to the route generation ECU 1230.

[0052] As the vehicle information measuring instrument 1240, the vehicle 1200 is equipped with a traveling speed measuring instrument 1241, a steering angle measuring instrument 1242, and an angular velocity measuring instrument 1243. The traveling speed measuring instrument 1241 is a measuring instrument that detects the traveling speed of the vehicle 1200. The steering angle measuring instrument 1242 is a measuring instrument that detects the steering angle of the vehicle 1200. The angular velocity measuring instrument 1243 is a measuring instrument that detects the angular velocity of the vehicle 1200 in the turning direction.

[0053] The route generation ECU 1250 is configured using logic circuits. Based on subject distance information obtained from image information from the imaging device 1210, the route generation ECU 1250 generates target route information relating to at least one of a target travel trajectory and a target travel speed of the vehicle 1200, and sequentially outputs the generated information to the vehicle control ECU 1260. The route generation ECU 1250 can further utilize measurement signals from the vehicle information measuring instrument 1240, distance information from the radar device 1220, and distance information from the LiDAR device 1230. The route generation ECU 1250 also controls the movement of the vehicle by outputting a control value for controlling the vehicle so that the distance between the vehicle and an obstacle that obstructs the movement is equal to or greater than a predetermined distance. In addition, if the vehicle 1200 is equipped with an HMI (Human Machine Interface) 1270 that displays an image or notifies the driver 1201 by voice, the target route information generated by the route generation ECU 1250 can be notified to the driver 1201 via the HMI 240.

[0054] By applying the distance correction according to this embodiment to the imaging device 1210, the accuracy of the distance information output is improved, the accuracy of the target route information output from the route generation ECU 1250 is improved, and safer vehicle driving control is achieved.

[0055] The second distance information according to this embodiment can be acquired by the SfM method using the imaging device 1201, the radar device 1220, the LiDAR device 1230, or the vehicle information measuring device 1240.

[0056] The present invention can also be realized by executing the following process. That is, software (programs) that realize the functions of the above-described embodiments are supplied to a system or device via a network or various storage media, and the computer (or CPU, MPU, etc.) of the system or device reads and executes the programs. The programs may also be provided by recording them on a computer-readable recording medium. [Explanation of symbols]

[0057] 100 distance acquisition device 101 Camera equipment 102 Stereo distance measurement calculation device 103 Feature point distance measurement calculation device 104 Correction value calculation device

Claims

1. a first acquisition means for acquiring first distance information including an error regarding the distance between the imaging unit and the subject via an optical system; a second acquiring means for acquiring second distance information having an error smaller than that of the first distance information; a generating means for generating a first correction value for correcting the first defocus amount based on a ratio between a first defocus amount corresponding to a deviation in the optical axis direction between a sensor surface and an image plane used to acquire the first distance information and a second defocus amount when the second distance information is acquired; a calculation unit that calculates a distance between the imaging unit and the subject using the first defocus amount corrected by the correction value.

2. the generating means further generates a second correction value for correcting the first defocus amount based on a difference between the first defocus amount and the second defocus amount; 2. The information processing apparatus according to claim 1, wherein the calculation means calculates the distance between the imaging unit and the subject using the first defocus amount corrected using the first correction value and the second correction value.

3. The information processing device described in claim 1, characterized in that the first acquisition means acquires first distance information based on multiple signals having parallax output by receiving each light beam that has passed through different pupil regions of a monocular optical system with a single image sensor.

4. 2. The information processing apparatus according to claim 1, wherein the second acquisition means uses an SfM (Structure from Motion) technique.

5. 2. The information processing apparatus according to claim 1, wherein the second obtaining means uses a technique for obtaining distance information based on the reflection intensity of an irradiated electromagnetic wave.

6. The information processing apparatus according to claim 1 , wherein the second acquisition means selects a method to be used based on the amount of movement of the imaging unit.

7. the calculation means corrects the first defocus amount a plurality of times using the first correction value and the second correction value; 3. The information processing apparatus according to claim 2, wherein the first correction value is applied before the second correction value.

8. the calculation means corrects the first defocus amount a plurality of times using the first correction value and the second correction value; 3. The information processing apparatus according to claim 2, wherein the second correction value is applied before the first correction value.

9. An information processing device mounted on a moving body, a first acquisition means for acquiring first distance information including an error regarding the distance between the imaging unit and the subject via an optical system; a second acquiring means for acquiring second distance information having an error smaller than that of the first distance information; a generating means for generating a first correction value for correcting the first defocus amount based on a ratio between a first defocus amount corresponding to a deviation in the optical axis direction between a sensor surface and an image plane used to acquire the first distance information and a second defocus amount when the second distance information is acquired; a calculation means for calculating a distance between the imaging unit and the subject using the first defocus amount corrected by the correction value; and control means for outputting a control value for controlling movement of the moving body based on the distance between the imaging unit and the subject calculated by the calculation means.

10. A mobile device equipped with the information processing device according to claim 9.

11. a first acquisition step of acquiring first distance information including an error regarding the distance between the imaging unit and the subject via an optical system; a second acquiring step of acquiring second distance information having an error smaller than that of the first distance information; a generating step of generating a first correction value for correcting the first defocus amount based on a ratio between a first defocus amount corresponding to a deviation in the optical axis direction between a sensor surface and an image plane used to acquire the first distance information and a second defocus amount used when the second distance information is acquired; and calculating a distance between the imaging unit and the subject using the first defocus amount corrected by the correction value.

12. A program that causes a computer to function as each of the means of the information processing device according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processor and stereo camera device

    JP2014052335A