Image processing device and image processing method
The image processing device maintains image resolution and reduces computational cost by parallelizing images from stereo cameras with fisheye lenses, addressing memory and accuracy challenges in stereo matching for vehicle collision prevention and surveillance.
Patent Information
- Application Number
- PCT/JP2025/011592
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-15
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-23
AI Technical Summary
Stereo cameras with fisheye lenses face challenges in achieving both memory efficiency and accurate parallax calculation due to image distortion and resolution loss during rectification, which is crucial for applications like vehicle collision prevention and surveillance.
An image processing device that parallelizes images captured by multiple cameras to maintain resolution and reduce computational cost, using a rectification process that aligns vertical coordinates and adjusts pixel resolution based on the baseline direction, ensuring accurate parallax calculation without degrading image quality.
The solution enables both memory savings and accurate distance measurement by maintaining image resolution, enhancing the performance of stereo cameras in applications such as vehicle control and surveillance.
Smart Images

Figure JP2025011592_23102025_PF_FP_ABST
Abstract
Description
Image processing device and image processing method
[0001] The present invention relates to an image processing device and an image processing method.
[0002] A stereo camera captures images of a three-dimensional object using multiple cameras, for example, two cameras. The stereo camera's image processing device calculates the parallax of the three-dimensional object in the images captured by each camera and measures the depth distance of the object using the principle of triangulation, thereby recognizing the three-dimensional object in three dimensions. Therefore, stereo cameras are used in a wide range of applications, such as vehicle collision prevention functions, robot sensors and surveillance cameras, and security. For example, stereo cameras are installed in automobiles and other vehicles to detect the positions of surrounding automobiles, pedestrians, bicycles, motorcycles, and other three-dimensional objects, and then perform automatic vehicle control, such as automatic emergency braking, based on the information.
[0003] In order to accommodate a variety of applications, stereo cameras are required to have a wider angle of view and be able to measure distances to more distant three-dimensional objects. When measuring distances to particularly wide-angle three-dimensional objects, cameras with extremely wide-angle lenses are often used. Fisheye lenses introduce significant distortion into the image, making it possible to capture subjects with a wide angle in a single image.
[0004] In stereo cameras, the stereo matching process, which calculates the disparity (relative position) between the same images captured by the left and right cameras, is the most computationally expensive process. Since real-time performance is required, particularly in automotive applications, it is necessary to reduce the computational cost as much as possible. Therefore, before performing the stereo matching process, the left and right images are rectified in the stereo camera. By rectifying the images, the search range for the same image can be limited to the same parallel line, significantly reducing the computational cost. Generally, the rectified images are converted into perspective projection images (undistorted images). However, a problem with perspective projection images of wide-angle images captured with a fisheye lens is that the image size is large.
[0005] Patent document 1 describes a mapper that "applies a mapping to the first wide-angle image to generate a corrected image having a corrected projection," and that "provides a vertical mapping function that corresponds to a non-linear vertical mapping from the perspective projection to a corrected vertical projection of the corrected projection, and a horizontal mapping function that corresponds to a mapping from the first projection to the perspective projection followed by a non-linear horizontal mapping from the perspective projection to a corrected horizontal projection of the corrected projection."
[0006] Special Publication No. 2022-506104
[0007] As described above, a nonlinear mapping function is applied to the images captured by the left and right image capture devices instead of perspective projection to generate a more compressed rectified image than in the past. However, the technology described in Patent Document 1 does not describe a specific compression function and is limited to performing some kind of compression on the perspective projection. While the greater the compression, the more the image can be compressed, compressing it to the point where the image resolution deteriorates leads to a deterioration in the accuracy of the parallax calculation. Furthermore, even if the image is enlarged, it is only enlarged by interpolation processing, and therefore cannot essentially exceed the resolution of the original image. In summary, compression that degrades the resolution of the original image also degrades the parallax accuracy, and enlargement simply increases memory capacity unnecessarily.
[0008] In order to achieve both parallax accuracy and image size, the challenge is to generate the most memory-efficient rectified image while maintaining the resolution of the original image.
[0009] The present invention has been made in consideration of the above points, and its purpose is to provide an image processing device that can achieve both memory savings and distance measurement accuracy by maintaining resolution when converting images captured by multiple cameras into parallelized images.
[0010] The image processing device of the present invention that solves the above problem includes: an image acquisition unit that acquires a first captured image and a second captured image, each captured by a plurality of imaging units; a parallelization processing unit that parallelizes the first captured image and the second captured image to generate a first parallelized image and a second parallelized image, such that vertical coordinates of identical images captured in the first captured image and the second captured image coincide at least at any point in the first captured image and the second captured image, with a vertical axis being a direction perpendicular to a baseline direction defined by the plurality of imaging units and to the optical axis direction of the plurality of imaging units; a parallax image generation unit that generates parallax based on the plurality of parallelized images; and an object recognition unit that recognizes an object based on the parallax, wherein the parallelization processing unit parallelizes the first captured image and the second captured image so that the resolution per unit pixel in the baseline direction exceeds that of the first captured image and the second captured image within a predetermined angle of view range in a plane including the vertical axis.
[0011] According to the present invention having the above configuration, by converting images captured by multiple cameras into parallelized images, it is possible to achieve both memory saving and distance measurement accuracy by maintaining resolution. Further features related to the present invention will become apparent from the description of this specification and the accompanying drawings. Furthermore, problems, configurations, and effects other than those described above will become apparent from the description of the following embodiments.
[0012] 1 is a functional block diagram illustrating the configuration of a stereo camera image processing device according to the first embodiment. FIG. 2 is a flowchart illustrating stereo calibration of the stereo camera image processing device according to the first embodiment. FIG. 3 is a schematic diagram of the relationship between light rays and image coordinates in a fisheye camera image according to the first embodiment. FIG. 4 is a three-dimensional schematic diagram of triangulation in a stereo camera according to the first embodiment. FIG. 5 is a schematic diagram of the relationship between light rays and image coordinates in a rectified image according to the first embodiment. FIG. 6 is a two-dimensional schematic diagram of triangulation in a stereo camera according to the first embodiment. FIG. 7 is an example of resolution-maintaining rectified image generation according to the first embodiment. FIG. 8 is a block diagram illustrating the configuration of a stereo camera image processing device according to the first embodiment. FIG. 9 is an example of a vertical angle of view range of interest according to the second embodiment. FIG. 10 is an example of resolution-maintaining rectified image generation according to the second embodiment. FIG. 11 is an example of a horizontal and vertical angle of view range of interest according to the second embodiment. FIG. 12 is an example of resolution-maintaining rectified image generation according to the second embodiment. FIG. 13 is an example of a resolution-maintaining rectified image generation according to the second embodiment. FIG. 14 is a schematic diagram of the relationship between light rays and image coordinates in a rectified image according to the third embodiment. FIG. 15 is a two-dimensional schematic diagram of triangulation in a stereo camera according to the third embodiment.
[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted. The present invention is applicable to, for example, a computing device for vehicle control capable of communicating with an on-board ECU (Electronic Control Unit) for an Advanced Driver Assistance System (ADAS) or Autonomous Driving (AD).
[0014] First Embodiment [Configuration Example of Image Processing Device 10] First, the configuration of the image processing device 10 according to the first embodiment of the present invention will be described. Below, a configuration example in which the image processing device 10 is mounted on a vehicle will be described. Note that the application of the present invention is not limited to this, and the image processing device 10 may also be mounted on a robot sensor, a surveillance camera, or the like. Furthermore, in the following embodiment, an example in which the image processing device of the present invention is used in a stereo camera having two imaging units will be described, but the configuration of the present invention is not limited to a stereo camera, and can also be applied to a multi-camera system having three or more cameras.
[0015] The image processing device 10 according to this embodiment is provided in an in-vehicle stereo camera 100. In this embodiment, the stereo camera 100 includes a first imaging unit (left imaging unit) 12 and a second imaging unit (right imaging unit) 13 as multiple imaging units 11 (see FIG. 8 (1)). The left imaging unit 12 and the right imaging unit 13 are each monocular cameras equipped with a lens and an image sensor, and capture images of three-dimensional objects in front of the vehicle. The lenses are fisheye lenses, and the captured images are circular fisheye images. Hereinafter, when there is no need to distinguish between the left imaging unit 12 and the right imaging unit 13, the left imaging unit 12 and the right imaging unit 13 will be collectively referred to as the "left and right imaging units 11."
[0016] FIG. 1 is a functional block diagram showing an example of the configuration of an image processing device 10 according to this embodiment. The image processing device 10 is configured as a computer system including an arithmetic unit, a storage device, an input / output device, etc. The image processing device 10 realizes the following functions by having the arithmetic unit read and execute a software program stored in the storage device. As shown in FIG. 1 , the image processing device 10 includes an image acquisition unit 101, an image pre-processing unit 103, a rectification processing unit 104, an image interpolation unit 105, a parallax image generation unit 106, a three-dimensional object recognition unit 107, a vehicle control unit 108, and a rectification image conversion table storage unit 109.
[0017] The image acquisition unit 101, image pre-processing unit 103, parallelization processing unit 104, image interpolation unit 105, parallax image generation unit 106, three-dimensional object recognition unit 107, and vehicle control unit 108 are connected in this order, and the parallelization image conversion table storage unit 109 is connected to the parallelization processing unit 104.
[0018] The image acquisition unit 101 acquires an image P1 captured by the left imaging unit 12 and an image P2 captured by the right imaging unit 13 (image acquisition step). The images P1 and P2 are first and second captured images obtained by capturing optical images formed by optical systems having the image height characteristics of the lenses of the left imaging unit 12 and the right imaging unit 13. The image pre-processing unit 103 performs pre-processing on the images P1 and P2 input from the left and right imaging units 11 by adjusting the luminance gain of the images P1 and P2 and performing shading correction to correct uneven luminance in the images P1 and P2. The image pre-processing unit 103 outputs the images P1 and P2 after shading correction to the parallelization processing unit 104.
[0019] The rectification image conversion table storage unit 109 stores a rectification image conversion table used in the processing performed by the rectification processing unit 104. The rectification image conversion tables for the left and right imaging units 11 are generated and stored during advance stereo calibration. The rectification image conversion table storage unit 109 stores conversion information for the radial distortion of the lenses mounted on the first imaging unit and the second imaging unit 13 (conversion information storage unit).
[0020] The parallelization processing unit 104 parallelizes the multiple captured images based on the conversion information stored in the parallelization image conversion table storage unit 109 so that the parallelization magnifications of the vertical and horizontal axes of the parallelized images are determined by the optical distortion of the lens according to the radial distortion of the lens. Specifically, the parallelization processing unit 104 performs image conversion processing (parallelization processing) on the images P1 and P2 after luminance gradation conversion so that the coordinates of the vertical axis of the image coordinate system of the same point in the left and right images are always equal (the left and right epipolar lines are both parallel to the horizontal axis of the image coordinate system). In other words, parallelization is performed so that the difference in the vertical coordinates of the image coordinate system of the same image captured in the images P1 and P2 is substantially constant over substantially the entire image.
[0021] In this parallelization process, the direction perpendicular to the baseline direction defined by the left and right imaging units 11 and the optical axis direction of the left and right imaging units 11 is used as the vertical axis, and images P1 (first image) and P2 (second image) are parallelized by image processing to generate a first parallelized image and a second parallelized image so that the vertical coordinates of the same image captured in images P1 (first image) and P2 (second image) coincide at least at any point in images P1 and P2 (parallelization processing step).
[0022] The rectification processing unit 104 rectifies the images P1 and P2 to generate a first rectified image and a second rectified image so that the resolution per unit pixel in the baseline direction is greater within a predetermined range of angles of view in a plane including the vertical axis. Specifically, the rectification processing unit 104 performs rectification using a rectification function that represents an image height per angle of view in the horizontal direction after rectification, and that is designed so that the resolution at any angle of view within a predetermined range of vertical and horizontal angles of view in the rectified images is equal to or greater than the resolution along the image central axis in either the first imaging unit (left imaging unit) 12 or the second imaging unit (right imaging unit) 13 and is equal to or less than the resolution along the widest-angle epipolar line, where the number of pixels per unit angle of view is the resolution.
[0023] The images P1 and P2 after the rectification process (first rectified image and second rectified image) are output from the rectification processing unit 104 to the image interpolation unit 105. In the rectification process, coordinate conversion of each image coordinate in the image conversion may be performed sequentially, but recalculation is inefficient. Therefore, an image conversion table is prepared that holds the image coordinates on the original image for each pixel of the image after the rectification process, and the rectified image is generated based on this. The rectification image conversion table is generated during prior calibration of the stereo camera.
[0024] The image interpolation unit 105 performs demosaicing on the images P1 and P2 after the rectification process. Demosaicing is, for example, a process of converting a Bayer image into a color image or a grayscale image. The image interpolation unit 105 also outputs the images P1 and P2 after the demosaicing process to the parallax image generation unit 106.
[0025] The parallax image generation unit 106 generates parallax by searching, for an image appearing at any vertical coordinate in the reference image, for an image appearing at a vertical coordinate in the reference image that is a certain difference from the vertical coordinate in the reference image, along a substantially horizontal direction from the vertical coordinate in the reference image (parallax generation step). The parallax image generation unit 106 is an example of a parallax calculation unit, and calculates parallax to generate a parallax image using the image captured by the first imaging unit (image P1) and the image captured by the second imaging unit (image P2), which have been subjected to image correction processing by the image interpolation unit 105. Here, the parallax image generation unit 106 performs stereo matching processing on the images P1 and P2 input from the image interpolation unit 105 to calculate parallax. The images P1 and P2 input from the image interpolation unit 105 are images P1 and P2 after image preprocessing by the image preprocessing unit 103 and various processing by the parallelization processing unit 104 and the image interpolation unit 105.
[0026] Stereo matching is a process of searching for a region of the same area in image P1 that corresponds to a specific range centered on a pixel of interest in image P2, and locating the location where the similarity between the regions is highest. Examples of matching methods include sum of absolute difference (SAD) and sum of squared difference (SSD). In this embodiment, the parallax image generation unit 106 determines the horizontal difference between corresponding pixels in image P2 and image P1 as the parallax of the pixel, calculates the parallax for all pixels in image P2, and generates a parallax image. The parallax image generation unit 106 outputs the generated parallax image to a three-dimensional object recognition unit 107 (object recognition unit).
[0027] The three-dimensional object recognition unit 107 detects three-dimensional objects (objects) using parallax images, identifies the type of the three-dimensional object (e.g., pedestrian, bicycle, car, motorcycle, light spot, etc.) through image recognition, and recognizes an image of a predetermined area including the identified three-dimensional object as an image of the three-dimensional object (object recognition step). The three-dimensional object recognition unit 107 also measures the distance to the recognized three-dimensional object. The three-dimensional object recognition unit 107 outputs the image of the recognized three-dimensional object, the type of the three-dimensional object, and information regarding the measured distance to the vehicle control unit 108.
[0028] The vehicle control unit 108 performs vehicle control such as acceleration / deceleration and steering of the vehicle for emergency braking and following a preceding vehicle, based on the image of the three-dimensional object recognized by the three-dimensional object recognition unit 107, information regarding the type of the three-dimensional object, distance, and the vehicle's traveling speed, etc.
[0029] [Stereo Calibration Processing] Next, a description will be given of stereo calibration in the image processing device 10. Fig. 2 is a flowchart showing the stereo calibration processing in the image processing device 10 according to this embodiment.
[0030] First, a calibration chart is photographed (step S21). Next, the photographed calibration chart is used to determine the internal and external parameters of the camera (step S22). Here, the internal parameters include the origin of the image plane, the focal lengths in the vertical and horizontal directions of the image, and radial distortion parameters that represent the radial distortion of the image point on the subject image relative to the origin. The radial distortion parameters are the coefficients of each term in a Taylor expansion of the image height characteristic function, which follows the projection equation of the fisheye camera, centered on the angle of incidence, and k 2n+1 It is expressed as:
[0031] The extrinsic parameters are the rotation matrix and translation matrix that transform the camera coordinate system of one camera into that of the other. A representative method for calculating these parameters is "Zhengyou Zhang. A flexible new technique for camera calibration. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 22(11):1330-1334, 2000," which is also implemented in OpenCV, but there are no limitations on the method as long as the intrinsic and extrinsic parameters can be accurately estimated.
[0032] Next, a resolution-maintaining rectified image conversion table that maintains the resolution of the sensor image is generated from the determined internal and external parameters (step S23). The generated rectified image conversion table is stored in the rectified image conversion table storage unit 109. Next, the generation of this resolution-maintaining rectified image conversion table will be described in detail.
[0033] [Generation of Resolution-Preserving Parallel Image Conversion Table] Figure 3 shows the relationship between incident light rays and their incident positions on the image coordinate system. As shown in Figure 3 (1), the camera coordinate system is (X, Y, Z), with Z being the optical axis direction of the camera. Furthermore, when the virtual image plane of the light rays incident on the camera for the subject P is a sphere with a radius equal to the focal length f, the sphere and the light point of the light ray are denoted as p. When this sphere is expanded to form a typical fisheye image, the point corresponding to p is p', as shown in Figure 3 (2).
[0034] If the fisheye lens mounted on the camera is designed as an orthogonal projection system, the image height characteristic, which is the distance between the formed image and the image center, is expressed as follows, where r is the distance from the origin (image center): The angle θ is the angle between the light ray incident on the fisheye lens and the optical axis of the fisheye lens.
[0035] In an actual lens, since variations from the design are included, it is expressed using the radial distortion parameter, which is an internal parameter described above, but for simplicity, the description will be given assuming a lens with Equation (2) as the image height function. The coordinates of p' are expressed using the image height function f sin θ according to θ described above, and φ in the figure.
[0036] Figure 4 shows a stereo camera with two fisheye cameras arranged side by side. The coordinate system after the stereo camera is parallelized is (X r ,Y r ,Z r ), (X l ,Y l ,Z l ), and the coordinate system before rectification is (X' r ,Y' r ,Z' r ), (X' l ,Y' l ,Z' l) In addition, when the ray of light incident on the camera for the subject P is a virtual image plane of a sphere with a radius of the focal length f, the sphere and the light point of the ray are defined as p l ,p r To achieve parallelization, it is necessary to transform the coordinate system (the optical axis after transformation is called the parallelized optical axis) so that the optical axis of the camera is oriented perpendicular to the line connecting the centers of the left and right cameras (called the baseline). There is some freedom in the direction of the parallelized optical axis, but it is set in the direction defined as the front relative to the position where the housing is attached.
[0037] In addition to setting the coordinate system, how to define the image coordinates is also important for rectification. l ,E r is the epipolar line. When the sphere is expanded onto the image plane, this epipolar line must be parallel to the horizontal axis of the image coordinate system.
[0038] Figure 5 shows one way of setting the image coordinate system to realize such parallelization. The direction of the light beam to the object P is defined as the pitch angle φ′ and the distance between the object P and the origin 0 on the left and right, as shown in Figure 5 (1). l ,0 r The horizontal axis is expressed as f sin θ' and the vertical axis is f sin φ', and as shown in Figure 5 (2), the horizontal axis of the image coordinate system is f sin θ' and the vertical axis is f sin φ', respectively. There is a degree of freedom in how the horizontal and vertical axes of the image coordinate system are chosen, and the horizontal axis can be x(θ'), a function of any θ', and the vertical axis can be y(φ'), a function of any φ'. Here, the function x(θ') is called the rectification function.
[0039] FIG. 6 shows the subject p in FIG. l and the left and right origin 0 l ,0 r The plane containing Sinθ' is shown. l , Sinθ′ r From there, By calculating the distance D from the base line length B, the three-dimensional position of the object P can be measured.
[0040] As mentioned above, there is a degree of freedom in how the image coordinate system is set after rectification, but to accurately calculate the distance D in equation (3), it is necessary to perform stereo matching with high precision. If the horizontal resolution deteriorates during conversion to a rectified image, the accuracy of the matching position will also deteriorate, so particular care must be taken when designing this rectification function.
[0041] FIG. 7 shows the process of converting an image into a parallelized image in this embodiment, with FIG. 7(1) being the image before parallelization, FIG. 7(2) being an intermediate image between the pre-parallelization and post-parallelization states, and FIG. 7(3) being the image after parallelization. Image 91 in FIG. 7(1) is a diagram showing an image in the coordinate system of a fisheye lens with a field angle of 180° before parallelization, along with the epipolar line. The lens does not necessarily have to have a field angle of 180°, and any field angle may be used, but for ease of understanding, it is assumed to be 180°. P n , P s are the two ends of the image central axis (epipolar line E0), and are the points that correspond to the so-called North and South Poles in latitude and longitude coordinates.
[0042] The parallelization processing unit 104 parallelizes the image 91 to generate a parallelized image so that the resolution per unit pixel in the baseline direction is higher than that of the pre-parallelization image 91 within a predetermined range of angles of view in a plane including the vertical axis. Here, parallelization is performed using a parallelization function that represents the image height per angle of view in the horizontal direction after parallelization, and is designed so that the resolution at an arbitrary angle of view is equal to or higher than the resolution along the image central axis in the first imaging unit 12 and the second imaging unit 13 and is equal to or lower than the resolution along the widest-angle epipolar line.
[0043] In Figures 7(1) to 7(3), the epipolar line of image 91 before rectification is E, the epipolar line of intermediate image 92 is E', and the epipolar line of image 93 after rectification is E''. First, the image is unfolded so that all epipolar lines E of the image in the fisheye lens coordinate system shown in image 91 before rectification in Figure 7(1) become horizontal while maintaining the distortion, as shown in intermediate image 92 in Figure 7(2). In other words, the image is unfolded so that all epipolar lines become parallel straight lines while maintaining the resolution (length). Then, as shown in image 91 after rectification in Figure 7(3), in this embodiment, the image is rectified so that it matches the epipolar line E'3 with the widest vertical angle of view that appears on the screen in intermediate image 92 in Figure 7(2).
[0044] As can be seen from the epipolar line E' of the intermediate image 92, in the intermediate image 92, the larger the vertical coordinate of the epipolar line, the higher the resolution of the epipolar line E'3 relative to the epipolar line E'0, for example. In this case, if the image is rectified using the rectification function fsin θ', this is equivalent to rectifying the image to the shortest epipolar line E'0, and therefore the epipolar line E''3 of the rectified image 93 will be compressed compared to the epipolar line E'3 of the intermediate image 92, resulting in a degradation of resolution compared to the original image 91 before rectification.
[0045] In contrast to this, in this embodiment, rectification is performed using a rectification function that represents the image height per horizontal angle of view after rectification. Specifically, as in the rectified image 93 in FIG. 7(3), in order to make the resolution constant over the entire screen, the image is rectified using the epipolar line E'3, which has the widest vertical angle of view on the screen, as the reference line. In other words, the rectification function is The image is rectified as follows. As a result, the resolution at any angle of view within the image is equal to or greater than the resolution along the image central axis and equal to or less than the resolution along the widest-angle epipolar line. Therefore, it is possible to rectify the image without degrading the resolution of any epipolar line E. Further enlargement would waste image memory capacity, so this can be said to indicate an upper limit for enlargement. However, if the epipolar line E'3 of the intermediate image 92 is larger than the angle of view of the camera lens, the epipolar line E'3 of the intermediate image 92 will not be included in the intermediate image 92. However, even in this case, the effect of the present invention, that is, rectification without degrading resolution, can be obtained.
[0046] In this embodiment, the image height characteristic of the fisheye lens is r=f sin θ, but in reality, the image height characteristic of the lens includes variations from the design. Furthermore, the projection method is not limited to orthogonal projection f sin θ, and may be, for example, a general equidistant projection fθ, solid angle projection 2f sin(θ / 2), or equisolid angle projection 2f tan(θ / 2), or the effects of the present invention can be obtained with other projections. Therefore, the image height characteristic may be generalized as r=r(θ). Furthermore, the epipolar line when generalized is expressed as follows in the fisheye coordinate system of the image 91 before parallelization, with φ as a constant: Here, x(θ′) obtained by horizontally expanding the epipolar line of the vertical angle of view φ′ as in the intermediate image 92 can be expressed as follows: This can be written as: and can be calculated using the relationship cosθ=cosθ′cosφ′.
[0047] [Example of Hardware Configuration of Image Processing Device 10] Next, a description will be given of the hardware configuration of the image processing device 10 according to this embodiment. Fig. 8 is a block diagram showing an example of the hardware configuration of the image processing device 10 according to this embodiment.
[0048] 8, the image processing device 10 includes a CPU (Central Processing Unit) 81, a ROM (Read Only Memory) 82, a RAM (Random Access Memory) 83, a storage device 84, an input / output interface 85, and a bus 86. The bus 86 is a signal path that electrically connects the components and allows input and output of information data between the components.
[0049] The CPU 81 controls the operation of each unit in the image processing device 10. For example, the CPU 81 controls the processing in each of the image acquisition unit 101, the image pre-processing unit 103, and the image interpolation unit 105. The CPU 81 also controls the processing in each of the parallelization processing unit 104, the image interpolation unit 105, the parallax image generation unit 106, the three-dimensional object recognition unit 107, and the vehicle control unit 108. Note that a GPU (Graphics Processing Unit) may be used instead of the CPU 81, or the CPU 81 and the GPU (Graphics Processing Unit) may be used together.
[0050] The ROM 82 is configured as a storage medium such as a nonvolatile memory, and stores programs and data executed and referenced by the CPU 81. The RAM 83 is configured as a storage medium such as a volatile memory, and temporarily stores information (data) required for each process performed by the CPU 81.
[0051] The storage device 84 is configured as a computer-readable, non-transitory recording medium storing a software program executed by the CPU 81, and is configured as a storage device such as an HDD (Hard Disk Drive). The storage device 84 stores programs for the CPU 81 to control various parts, an OS (Operating System), controller programs, and other data. Note that some of the programs and data stored in the storage device 84 may be stored in the ROM 82. The computer-readable, non-transitory recording medium storing the program executed by the CPU 81 is not limited to an HDD, and may be, for example, a solid state drive (SSD), a compact disc (CD)-ROM, a digital versatile disc (DVD)-ROM, or other recording medium.
[0052] The input / output interface 85 transmits and receives information data to and from the outside of the image processing device 10 under the control of the CPU 81. The image processing device 10 embodies the various components described in FIG. 1 through cooperation between these hardware components and software programs.
[0053] [Effect] As described above, in the image processing device 10 according to this embodiment, a rectification image conversion table is generated that performs rectification so as to maintain the resolution, and is stored in the rectification image conversion table storage unit 109. This rectification image conversion table makes it possible to generate a rectified image that achieves both high ranging accuracy and memory saving.
[0054] It should be noted that the present invention is not limited to the above-described embodiments, and it goes without saying that various other applications and modifications are possible without departing from the gist of the present invention as set forth in the claims.
[0055] For example, the above-described embodiments have described in detail and specifically the configuration of an image processing device in order to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0056] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected.
[0057] Second Embodiment In the first embodiment, the reference line is selected by using the epipolar line with the widest vertical angle of view that appears on the screen to perform collimation, with the aim of maintaining a constant resolution across the entire screen. However, the range of vertical angles of view for which accurate distance measurement is required may be limited. In other words, image collimation may be performed so as to maintain a constant resolution across a predetermined range of vertical angles of view, rather than across the entire image.
[0058] 9 shows the positional relationship between a bump and an on-board camera. For example, a speed breaker (bump) 120 that encourages drivers to reduce their speed is installed on a road surface 121. If the vehicle V attempts to pass over this bump at a high speed, a large impact will be applied to the vehicle V, which may affect the passenger comfort, driving operation, and the running of the vehicle V. Therefore, it is necessary to accurately measure the distance to the bump 120 and use the information to warn the passengers and to control the braking of the vehicle V.
[0059] When the height of the stereo camera 100 from the ground is h and the distance in the optical axis direction from the stereo camera 100 to the bump 120 is the distance measurement distance L1, the vertical angle of view α is tan -1 For example, if the host vehicle V is a normal vehicle such as a passenger car, when h=1.6 m and L=5 m, α=approximately 20°.
[0060] FIG. 10 shows the process of converting to a parallelized image in this embodiment, where FIG. 10(1) is the image before parallelization, FIG. 10(2) is the intermediate image between before and after parallelization, and FIG. 10(3) is the image after parallelization.
[0061] In the example shown in FIG. 10, an epipolar line E passing through the vertical angle of view α α By performing parallelization in this manner, the resolution is not degraded in the field of view where the bump 120, which is the detection target, is to be accurately detected.
[0062] The rectification processing unit 104 rectifies the image 140 to generate a rectified image 142 so that the resolution per unit pixel in the baseline direction is higher than that of the pre-rectification image 140 in a predetermined angle of view range r(α) in a plane including the vertical axis, but is not higher than that of the pre-rectification image 140 in a range outside the predetermined angle of view range r(α). That is, in a region (within the angle of view range r(α)) perpendicular to the direction in which the multiple imaging units are arranged, the first rectified image and the second rectified image have a larger number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image, and in a region (outside the angle of view range r(α)) perpendicular to the direction in which the multiple imaging units are arranged that is larger than the vertical angle of view α, the first rectified image and the second rectified image have a smaller number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image. In Figures 10(1) to 10(3), the epipolar line of the image 140 before rectification is Eα, the epipolar line of the intermediate image 141 is E'α, and the epipolar line of the image 142 after rectification is E''α.
[0063] The image is expanded so that all epipolar lines E of the image in the fisheye lens coordinate system shown in image 140 before rectification in FIG. 10(1) become horizontal while maintaining the distortion, as shown in intermediate image 141 in FIG. 10(2). In other words, the image is expanded so that all epipolar lines become parallel straight lines while maintaining the resolution (length). Then, as shown in image 142 after rectification in FIG. 10(3), the epipolar line E'' passing through the vertical angle of view α αThe image is parallelized using the reference line of the parallelization function as a reference line. As a result, the resolution per unit pixel in the baseline direction exceeds that of the image 140 before parallelization within a predetermined range of angle of view r(α) on a plane including the vertical axis, and the image can be parallelized without degrading the resolution within the field of view in which the bump 120, which is the detection target, is to be accurately detected.
[0064] The parallelization processing unit 104 may also set the vertical angle of view depending on the type of object. The vertical angle of view condition can be set by dividing an area by the horizontal angle of view and imposing different conditions on each area. Fig. 11 shows a bump 120 across the width of a road surface 121 and a pedestrian 122 about to cross the road surface 121. Fig. 12 shows an image 140 in the coordinate system of the fisheye lens before parallelization, plotted against an epipolar line E. α , E β Assuming that the bump 120 and the pedestrian 122 are located closest to the host vehicle V, the vertical angle of view will be within the angle of view range r(α) if they are within the range of the horizontal angle of view γ1, and the vertical angle of view will be within the angle of view range r(β) if they are within the range of the horizontal angle of view γ2.
[0065] At this time, as shown in FIG. 12, for each region of the horizontal field of view, an epipolar line E corresponding to the vertical field of view is drawn. α and E β In other words, within the range of the horizontal angle of view γ1, the epipolar line E'' passing through the vertical angle of view α is α The image is rectified using the reference line of the rectification function, and the epipolar line E'' passing through the vertical angle of view β within the range of the horizontal angle of view γ2 β The image can be rectified using the horizontal angle of view γ1 as the reference line of the rectification function. The vertical angles of view α and β can also be used depending on the speed of the host vehicle V. In other words, the vertical angles of view α and β can be changed depending on the speed of the moving object. FIG. 12 shows an image 140 before rectification, and at a horizontal angle of view γ1, the epipolar line E α At the horizontal angle of view γ2, the epipolar line E βIn other words, the parallelization processing unit 104 can parallelize the captured image based on a parallelization function that outputs the image height in the direction in which the multiple image capturing units are arranged, which differs depending on the angle of view ranges of the multiple image capturing units.
[0066] In addition, an upper limit to the vertical angle of view can also occur when the camera's angle of view is less than 180°. The upper limit epipolar line may not be a simple epipolar line for lenses with a polar angle of view other than 180°.
[0067] Figure 13 shows the collimation in the case of a wide-angle lens with an angle of view of less than 180°. Image 183 in Figure 13(1) shows the image in the coordinate system of the wide-angle lens before collimation along with the epipolar line, Figure 13(2) shows an intermediate image between before and after collimation, and Figure 13(3) shows the image after collimation.
[0068] If the maximum vertical angle of view of the camera is α, then the epipolar line E α is determined, and rectification is performed in the range outside the maximum vertical angle of view α so that the resolution per unit pixel in the baseline direction does not exceed the maximum vertical angle of view α. At this time, the size of the final image varies depending on the maximum horizontal angle of view of the camera. For example, in the case of a wide-angle lens whose maximum horizontal angle of view is smaller than that of image 180, as in image 183 before rectification in Figure 13(1), the horizontal size will also be smaller, as in intermediate image 181 in Figure 13(2) and image 182 after rectification in Figure 13(3).
[0069] Furthermore, if the epipolar line of the vertical angle of view is followed, it is not possible to ensure the resolution of the area outside the diagonal corner of the camera's imaging range. Therefore, as shown in Figure 14, it is also possible to determine the maximum epipolar line for each horizontal angle of view γ1, γ2 and perform parallelization so that the resolution does not exceed that value.
[0070] Furthermore, up until now, an appropriate epipolar line has been selected, and parallelization has been performed according to a parallelization function that follows the epipolar line. However, from the viewpoint of maintaining resolution, it is not necessary to use a parallelization function that follows the epipolar line. As long as the resolution, defined by the number of pixels per unit angle of view, is at least greater than the epipolar line E0 in FIG. 7 and smaller than the epipolar line E3 at any horizontal angle of view, the image resolution will not be degraded across the entire image, and there will be no unnecessary enlargement. This is because, assuming that the parallelization function that follows the epipolar line E0 is x0(θ'), the parallelization function that follows the epipolar line E3 is x3(θ'), and the parallelization function is x(θ'), then at any θ', For example, using the formulas (2) and (4) used in the first embodiment, and Then, the desired x(θ') can be obtained.
[0071] Under the above-mentioned resolution conditions, it is also possible to choose a reference line instead of the epipolar line. Figure 15 shows an example of a suitable reference line L.
[0072] Third Embodiment Most of the configuration and processing in this embodiment is the same as in the first embodiment. The difference is that the relationship between the incident light ray and the incident position on the image coordinate system after collimation is as shown in Figure 16 instead of Figure 5. Specifically, the angle from the X axis when the light ray is projected onto the ZX plane is defined as θ'.
[0073] 17 shows the projection of the subject P onto the ZX plane in FIG. 6. In this case, D, which is calculated in the same manner as in Equation (3), always represents the depth, not the distance from the baseline. In this embodiment, the depth is directly calculated from the parallax, making this a configuration that is well suited as a projection method for a front-sensing sensor.
[0074] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the above-described embodiments, and various design modifications can be made without departing from the spirit of the present invention as defined in the claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0075] 10...Image processing device, 11...Imaging unit (plurality of imaging units), 12...First imaging unit (left imaging unit), 13...Second imaging unit (right imaging unit), 100...Stereo camera, 101...Image acquisition unit, 103...Image pre-processing unit, 104...Parallelization processing unit, 105...Image interpolation unit, 106...Parallax image generation unit, 107...Three-dimensional object recognition unit (object recognition unit), 108...Vehicle control unit, 109...Parallelization image conversion table storage unit (conversion information storage unit)
Claims
1. An image processing device comprising: an image acquisition unit that acquires first and second captured images captured by a plurality of imaging units; a parallelization processing unit that parallelizes the first and second captured images to generate first and second parallelized images, so that vertical coordinates of identical images captured in the first and second captured images coincide at least at any point in the first and second captured images, with a vertical axis being a direction perpendicular to a baseline direction defined by the plurality of imaging units and to the optical axis directions of the plurality of imaging units; a parallax image generation unit that generates parallax based on the plurality of parallelized images; and an object recognition unit that recognizes an object based on the parallax, wherein the parallelization processing unit parallelizes the first and second captured images so that the resolution per unit pixel in the baseline direction exceeds that of the first and second captured images within a predetermined angle of view in a plane including the vertical axis.
2. An image processing device according to claim 1, characterized in that, when a predetermined vertical angle of view α is set in a direction perpendicular to the direction in which the multiple imaging units are arranged, in an area equal to or smaller than the vertical angle of view α in the direction perpendicular to the direction in which the multiple imaging units are arranged, the first parallelized image and the second parallelized image have a larger number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image, and in an area larger than the vertical angle of view α in the direction perpendicular to the direction in which the multiple imaging units are arranged, the first parallelized image and the second parallelized image have a smaller number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image.
3. An image processing device according to claim 1, wherein the parallelization processing unit performs the parallelization using a parallelization function expressed by the following equation (4): (where f is the focal length, and θ′ is the yaw angle in a plane including the subject and the direction in which the multiple imaging units are arranged) 4. An image processing device according to claim 1, characterized in that the parallelization processing unit parallelizes the captured image based on a parallelization function that outputs an image height in the direction in which the multiple imaging units are arranged, which differs depending on the angle of view range of the multiple imaging units.
5. An image processing device according to claim 2, wherein, when the height of the plurality of imaging units from the ground is h and the corresponding distance measurement distance is L1, the vertical angle of view α is tan -1 (h / L1) or more.
6. An image processing device according to claim 1, wherein the epipolar lines of the plurality of imaging units are parallel to the direction in which the plurality of imaging units are arranged in the first rectified image and the second rectified image.
7. An image processing device according to claim 1, further comprising a conversion information storage unit storing conversion information for radial distortion of lenses mounted on said plurality of imaging units, and wherein said parallelization processing unit parallelizes said first captured image and said second captured image based on said conversion information so that the parallelization magnifications of the vertical and horizontal axes of said first parallelized image and said second parallelized image are determined by the optical distortion of said lens in accordance with the radial distortion of said lens.
8. An image processing device according to claim 1, wherein the parallelization processing unit unfolds the first captured image and the second captured image so that the epipolar lines of the first captured image and the second captured image become parallel straight lines while maintaining their lengths, selects an epipolar line that makes the resolution of the entire captured image constant as a reference line, and parallelizes the first captured image and the second captured image to match the reference line.
9. An image processing device according to claim 8, wherein the parallelization processing unit selects the epipolar line in accordance with a predetermined vertical angle of view appearing in the captured image.
10. An image processing device according to claim 9, wherein the parallelization processing unit selects the epipolar line whose vertical angle of view is the widest.
11. An image processing device according to claim 9, wherein the parallelization processing unit sets the vertical angle of view according to the type of the object.
12. An image processing device according to claim 9, wherein the parallelization processing unit changes the vertical angle of view in accordance with the speed of a moving object on which the image processing device is mounted.
13. An image processing device according to claim 9, wherein the parallelization processing unit sets the vertical angle of view for each region of the horizontal angle of view of the captured image.
14. An image processing method comprising: an image acquisition step of acquiring first and second captured images respectively captured by a plurality of imaging units; a parallelization processing step of parallelizing the first and second captured images to generate first and second parallelized images so that vertical coordinates of identical images captured in the first and second captured images coincide at least at any point in the first and second captured images, with a vertical axis being a direction perpendicular to a baseline direction defined by the plurality of imaging units and the optical axis direction of the plurality of imaging units; a parallax generation step of generating parallax based on the plurality of parallelized images; and an object recognition step of recognizing an object based on the parallax, wherein in the parallelization processing step, the first and second captured images are parallelized to generate the first and second parallelized images so that the resolution per unit pixel in the baseline direction exceeds that of the first and second captured images within a predetermined angle of view in a plane including the vertical axis.
Citation Information
Patent Citations
Image processing device, image processing method, and vehicle
WO2017212928A1
Image processing device, image processing method, and program
WO2019003910A1
Disparity estimation from a wide angle image
WO2020094575A1