Image processing device and image processing method
The image processing device addresses the challenge of maintaining resolution and reducing memory usage in stereo cameras by converting images into parallelized formats, achieving efficient memory use and accurate distance measurement.
Patent Information
- Application Number
- JP2024065069
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-10-27
AI Technical Summary
Existing stereo camera systems face challenges in achieving both memory efficiency and distance measurement accuracy while maintaining image resolution, particularly when using fisheye lenses that introduce significant distortion and require computationally expensive stereo matching processes.
The image processing device employs a parallelization processing unit that converts images captured by multiple cameras into parallelized images, ensuring the resolution per unit pixel in the baseline direction exceeds a predetermined range, thereby maintaining image resolution and reducing memory usage.
This approach allows for both memory savings and accurate distance measurement by generating rectified images that preserve resolution, enhancing the performance of stereo cameras in applications like vehicle collision prevention and surveillance.
Smart Images

Figure 2025162008000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device and an image processing method. [Background technology]
[0002] A stereo camera captures images of a three-dimensional object using multiple cameras, for example, two cameras. The stereo camera's image processing device calculates the parallax of the object in the images captured by each camera and measures the depth distance of the object using the principle of triangulation, thereby recognizing the object in three dimensions. Therefore, stereo cameras are used in a wide range of applications, including vehicle collision prevention functions, robot sensors and surveillance cameras, and security. For example, stereo cameras are installed in automobiles and can detect the positions of surrounding automobiles, pedestrians, bicycles, motorcycles, and other three-dimensional objects, and then use this information to automatically control the vehicle, such as by applying automatic emergency braking.
[0003] In order to accommodate a variety of application cases, stereo cameras are required to have a wider angle of view and be able to measure distances to more distant three-dimensional objects. When measuring distances to particularly wide-angle three-dimensional objects, extremely wide-angle cameras with fisheye lenses are often used. Fisheye lenses introduce significant distortion into the image, making it possible to capture subjects with a wide angle in a single image.
[0004] In stereo cameras, the stereo matching process, which calculates the disparity (relative position) between the same images captured by the left and right cameras, is the most computationally expensive process. Since real-time performance is required, particularly in automotive applications, it is necessary to reduce the computational cost as much as possible. Therefore, before performing the stereo matching process, the left and right images are rectified in the stereo camera. By rectifying the images, the search range for the same image can be limited to the same parallel line, significantly reducing the computational cost. Generally, the rectified images are converted into perspective projection images (undistorted images). However, a problem with perspective projection images of wide-angle images captured with a fisheye lens is that the image size is large.
[0005] Patent document 1 describes a mapper that "applies a mapping to the first wide-angle image to generate a corrected image having a corrected projection," and that "provides a vertical mapping function that corresponds to a non-linear vertical mapping from the perspective projection to a corrected vertical projection of the corrected projection, and a horizontal mapping function that corresponds to a mapping from the first projection to the perspective projection followed by a non-linear horizontal mapping from the perspective projection to a corrected horizontal projection of the corrected projection." [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Special Publication No. 2022-506104 Summary of the Invention [Problem to be solved by the invention]
[0007] As described above, a nonlinear mapping function is applied to the images captured by the left and right image capture devices instead of perspective projection to generate a more compressed rectified image than conventional methods. However, the technology described in Patent Document 1 does not specify a specific compression function and is limited to performing some kind of compression on the perspective projection. While the greater the compression, the greater the image compression, compressing the image to the point where the image resolution is degraded will lead to a deterioration in the accuracy of the parallax calculation. Furthermore, even if the image is enlarged, it is only enlarged by interpolation processing, and therefore cannot essentially exceed the resolution of the original image. In summary, compression that degrades the resolution of the original image will also degrade the parallax accuracy, and enlargement will only increase memory capacity unnecessarily.
[0008] In order to achieve both parallax accuracy and image size, the challenge is to generate the most memory-efficient rectified image while maintaining the resolution of the original image.
[0009] The present invention has been made in consideration of the above points, and its purpose is to provide an image processing device that can achieve both memory savings and distance measurement accuracy by maintaining resolution when converting images captured by multiple cameras into parallelized images. [Means for solving the problem]
[0010] The image processing device of the present invention that solves the above problems comprises: an image acquisition unit that acquires a first captured image and a second captured image captured by the plurality of imaging units; a parallelization processing unit that parallelizes the first captured image and the second captured image to generate a first parallelized image and a second parallelized image, with a direction perpendicular to a baseline direction defined by the plurality of image capturing units and to the optical axis directions of the plurality of image capturing units as a vertical axis, so that vertical coordinates of identical images captured in the first captured image and the second captured image coincide with each other at least at any point in the first captured image and the second captured image; a parallax image generator that generates parallax based on the plurality of parallelized images; an object recognition unit that recognizes an object based on the parallax, the parallelization processing unit parallelizes the first captured image and the second captured image so that a resolution per unit pixel in the base line direction exceeds a resolution per unit pixel in a predetermined range of an angle of view in a plane including the vertical axis, with respect to the first captured image and the second captured image. It is characterized by: [Effects of the Invention]
[0011] According to the present invention having the above configuration, by converting images captured by multiple cameras into parallelized images, it is possible to achieve both memory saving and distance measurement accuracy by maintaining resolution. Further features related to the present invention will become apparent from the description of this specification and the accompanying drawings. Furthermore, problems, configurations, and effects other than those described above will become apparent from the description of the following embodiments. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a functional block diagram illustrating the configuration of a stereo camera image processing device according to a first embodiment. [Figure 2] 4 is a flowchart illustrating stereo calibration of the stereo camera image processing device according to the first embodiment. [Figure 3] 5 is a schematic diagram illustrating the relationship between light rays and image coordinates in a fisheye camera image according to the first embodiment. [Figure 4] FIG. 2 is a three-dimensional schematic diagram of triangulation in the stereo camera according to the first embodiment. [Figure 5] FIG. 3 is a schematic diagram illustrating the relationship between light rays and image coordinates in a parallelized image according to the first embodiment. [Figure 6] FIG. 2 is a two-dimensional schematic diagram of triangulation in the stereo camera according to the first embodiment. [Figure 7] 10 shows an example of generation of a resolution-maintaining parallelized image according to the first embodiment. [Figure 8] FIG. 1 is a block diagram illustrating the configuration of a stereo camera image processing device according to a first embodiment. [Figure 9] 10 shows an example of a vertical angle of view range of interest according to the second embodiment. [Figure 10] 10 shows an example of generation of a resolution-preserving parallelized image according to the second embodiment. [Figure 11] 10 shows an example of a horizontal and vertical angle of view range of interest according to the second embodiment. [Figure 12] 10 shows an example of generation of a resolution-preserving parallelized image according to the second embodiment. [Figure 13] 10 shows an example of generation of a resolution-preserving parallelized image according to the second embodiment. [Figure 14] 10 shows an example of generation of a resolution-preserving parallelized image according to the second embodiment. [Figure 15] 10 shows an example of generation of a resolution-preserving parallelized image according to the second embodiment. [Figure 16] FIG. 11 is a schematic diagram illustrating the relationship between light rays and image coordinates in a parallelized image according to the third embodiment. [Figure 17] FIG. 11 is a two-dimensional schematic diagram of triangulation in a stereo camera according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted. The present invention is applicable to, for example, a computing device for vehicle control capable of communicating with an on-board ECU (Electronic Control Unit) for an Advanced Driver Assistance System (ADAS) or Autonomous Driving (AD).
[0014] First Embodiment [Configuration example of image processing device 10] First, the configuration of an image processing device 10 according to a first embodiment of the present invention will be described. Below, an example configuration will be described in which the image processing device 10 is mounted on a vehicle. Note that the application of the present invention is not limited to this, and the image processing device 10 may also be mounted on a robot sensor, a surveillance camera, or the like. Furthermore, in the following embodiment, an example will be described in which the image processing device of the present invention is used in a stereo camera having two imaging units, but the configuration of the present invention is not limited to a stereo camera, and can also be applied to a multi-camera system having three or more cameras.
[0015] The image processing device 10 according to this embodiment is provided in an in-vehicle stereo camera 100. In this embodiment, the stereo camera 100 includes a first imaging unit (left imaging unit) 12 and a second imaging unit (right imaging unit) 13 as multiple imaging units 11 (see FIG. 8(1)). The left imaging unit 12 and the right imaging unit 13 are each a monocular camera equipped with a lens and an image sensor, and capture images of three-dimensional objects in front of the vehicle. The lens is a fisheye lens, and the captured image is a circular fisheye image. Hereinafter, when there is no need to distinguish between the left imaging unit 12 and the right imaging unit 13, the left imaging unit 12 and the right imaging unit 13 will be collectively referred to as the "left and right imaging units 11."
[0016] FIG. 1 is a functional block diagram showing an example of the configuration of an image processing device 10 according to this embodiment. The image processing device 10 is configured by a computer system including an arithmetic unit, a storage device, an input / output device, etc. The image processing device 10 realizes the following functions by having the arithmetic unit read and execute a software program stored in the storage device. As shown in FIG. 1, the image processing device 10 includes an image acquisition unit 101, an image preprocessing unit 103, a rectification processing unit 104, an image interpolation unit 105, a parallax image generation unit 106, a three-dimensional object recognition unit 107, a vehicle control unit 108, and a rectification image conversion table storage unit 109.
[0017] The image acquisition unit 101, image pre-processing unit 103, parallelization processing unit 104, image interpolation unit 105, parallax image generation unit 106, three-dimensional object recognition unit 107, and vehicle control unit 108 are connected in this order, and the parallelization image conversion table storage unit 109 is connected to the parallelization processing unit 104.
[0018] The image acquisition unit 101 acquires an image P1 captured by the left imaging unit 12 and an image P2 captured by the right imaging unit 13 (image acquisition step). The images P1 and P2 are a first captured image and a second captured image obtained by capturing optical images formed by optical systems having the image height characteristics of the lenses of the left imaging unit 12 and the right imaging unit 13. The image pre-processing unit 103 performs pre-processing on the images P1 and P2 input from the left and right imaging units 11 by adjusting the luminance gain of the images P1 and P2 and performing shading correction to correct uneven luminance of the images P1 and P2. The image pre-processing unit 103 outputs the images P1 and P2 after shading correction to the parallelization processing unit 104.
[0019] The rectification image conversion table storage unit 109 stores a rectification image conversion table used in the processing performed by the rectification processing unit 104. The rectification image conversion tables for the left and right imaging units 11 are generated and stored during advance stereo calibration. The rectification image conversion table storage unit 109 stores conversion information for radial distortion of the lenses mounted on the first imaging unit and the second imaging unit 13 (conversion information storage unit).
[0020] The parallelization processing unit 104 parallelizes the multiple captured images based on the conversion information stored in the parallelization image conversion table storage unit 109 so that the parallelization magnifications of the vertical and horizontal axes of the parallelized images are determined by the optical distortion of the lens according to the radial distortion of the lens. Specifically, the parallelization processing unit 104 performs image conversion processing (parallelization processing) on the images P1 and P2 after luminance gradation conversion so that the coordinates of the vertical axis of the image coordinate system of the same point in the left and right images are always the same (the left and right epipolar lines are both parallel to the horizontal axis of the image coordinate system). In other words, parallelization is performed so that the difference in the vertical coordinates of the image coordinate system of the same image captured in the images P1 and P2 is approximately constant over almost the entire image.
[0021] In this parallelization process, images P1 and P2 are parallelized by image processing so that the vertical coordinates of the same images captured in images P1 (first image) and P2 (second image) coincide at least at any point in images P1 and P2, with the direction perpendicular to the baseline direction defined by the left and right imaging units 11 and the optical axis direction of the left and right imaging units 11 as the vertical axis, to generate a first parallelized image and a second parallelized image (parallelization processing step).
[0022] The rectification processing unit 104 rectifies the images P1 and P2 to generate a first rectified image and a second rectified image so that the resolution per unit pixel in the baseline direction is greater within a predetermined range of angles of view in a plane including the vertical axis. Specifically, the rectification processing unit 104 performs rectification using a rectification function that represents the image height per angle of view in the horizontal direction after rectification, and is designed so that the resolution at any angle of view within a predetermined range of vertical and horizontal angles of view in the rectified images is equal to or greater than the resolution along the image central axis in either the first imaging unit (left imaging unit) 12 or the second imaging unit (right imaging unit) 13 and is equal to or less than the resolution along the widest-angle epipolar line, where the resolution is the number of pixels per unit angle of view.
[0023] The images P1 and P2 after the rectification process (first rectified image and second rectified image) are output from the rectification processing unit 104 to the image interpolation unit 105. In the rectification process, coordinate conversion of each image coordinate in the image conversion may be performed sequentially, but since recalculation is inefficient, an image conversion table is prepared that holds the image coordinates on the original image for each pixel of the image after the rectification process, and the rectified image is generated based on this. The rectification image conversion table is generated during prior calibration of the stereo camera.
[0024] The image interpolation unit 105 performs demosaicing on the images P1 and P2 after the rectification process. Demosaicing is, for example, a process of converting a Bayer image into a color image or a grayscale image. The image interpolation unit 105 also outputs the images P1 and P2 after the demosaicing process to the parallax image generation unit 106.
[0025] The parallax image generation unit 106 uses one of the first rectified image and the second rectified image as a base image and the other as a reference image, and generates parallax by searching the reference image for an image that appears at a vertical coordinate in the base image that is a certain difference from the vertical coordinate of the image that appears at the vertical coordinate of the reference image along a substantially horizontal direction from the vertical coordinate in the reference image (parallax generation step). The parallax image generation unit 106 is an example of a parallax calculation unit, and calculates parallax to generate a parallax image using the image captured by the first imaging unit (image P1) and the image captured by the second imaging unit (image P2) that have been subjected to image correction processing by the image interpolation unit 105. Here, the parallax image generation unit 106 performs stereo matching processing on the images P1 and P2 input from the image interpolation unit 105 to calculate parallax. The images P1 and P2 input from the image interpolation unit 105 are images P1 and P2 that have undergone image preprocessing by the image preprocessing unit 103 and various processes by the rectification processing unit 104 and the image interpolation unit 105.
[0026] Stereo matching is a process of searching for a region of the same area in image P1 that corresponds to a specific range centered on a pixel of interest in image P2, and locating a location where the similarity between the regions is highest. Examples of matching methods include Sum of Absolute Difference (SAD) and Sum of Squared Difference (SSD). In this embodiment, the parallax image generation unit 106 determines the horizontal difference between corresponding pixels in image P2 and image P1 as the parallax of the pixel, and calculates the parallax for all pixels in image P2 to generate a parallax image. The parallax image generation unit 106 outputs the generated parallax image to a three-dimensional object recognition unit 107 (object recognition unit).
[0027] The three-dimensional object recognition unit 107 detects a three-dimensional object (object) using the parallax image, identifies the type of the three-dimensional object (for example, a pedestrian, bicycle, car, motorcycle, light spot, etc.) by image recognition, and recognizes an image of a predetermined area including the identified three-dimensional object as an image of the three-dimensional object (object recognition step).The three-dimensional object recognition unit 107 also measures the distance to the recognized three-dimensional object.The three-dimensional object recognition unit 107 outputs the image of the recognized three-dimensional object, the type of the three-dimensional object, and information regarding the measured distance to the vehicle control unit 108.
[0028] The vehicle control unit 108 performs vehicle control such as acceleration / deceleration and steering of the vehicle for emergency braking and following a preceding vehicle, based on the image of the three-dimensional object recognized by the three-dimensional object recognition unit 107, information regarding the type of the three-dimensional object, distance, and the traveling speed of the vehicle.
[0029] [Stereo calibration processing] Next, a description will be given of stereo calibration in the image processing device 10. Fig. 2 is a flowchart showing the stereo calibration process of the image processing device 10 according to this embodiment.
[0030] First, a calibration chart is photographed (step S21). Next, the photographed calibration chart is used to determine the internal and external parameters of the camera (step S22). Here, the internal parameters include the origin of the image plane, the focal lengths in the vertical and horizontal directions of the image, and radial distortion parameters that represent the radial distortion of the image point on the subject image relative to the origin. The radial distortion parameters are the coefficients of each term in the Taylor expansion of the image height characteristic function, which follows the projection equation of the fisheye camera, centered on the angle of incidence, and are given by TIFF2025162008000002.tif14158k 2n+1 It is expressed as:
[0031] The extrinsic parameters are the rotation matrix and translation matrix that transform the camera coordinate system of one camera into that of the other. A representative method for calculating these parameters is "Zhengyou Zhang. A flexible new technique for camera calibration. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 22(11):1330-1334, 2000," which is also implemented in OpenCV, but there are no limitations on the method as long as the intrinsic and extrinsic parameters can be accurately estimated.
[0032] Next, a resolution-maintaining rectified image conversion table that maintains the resolution of the sensor image is generated from the determined internal and external parameters (step S23). The generated rectified image conversion table is stored in rectified image conversion table storage unit 109. Next, the generation of this resolution-maintaining rectified image conversion table will be described in detail.
[0033] [Resolution-preserving rectification image conversion table generation] Figure 3 shows the relationship between incident light rays and their incident positions on the image coordinate system. As shown in Figure 3(1), the camera coordinate system is (X, Y, Z), with Z being the direction of the camera's optical axis. Furthermore, when the virtual image plane of the light rays incident on the camera for the subject P is a sphere with a radius equal to the focal length f, the sphere and the light point of the light ray are designated as p. When this sphere is expanded to form a typical fisheye image, the point corresponding to p is designated as p', as shown in Figure 3(2).
[0034] If the fisheye lens mounted on the camera is designed as an orthogonal projection system, the image height characteristic, which is the distance between the formed image and the image center, is expressed as follows, where r is the distance from the origin (image center): It is expressed as TIFF2025162008000003.tif5156. The angle θ is the angle between the light ray incident on the fisheye lens and the optical axis of the fisheye lens.
[0035] In an actual lens, variations from the design are included, and therefore it is expressed using the radial distortion parameter, which is an internal parameter described above, but for simplicity, the explanation will be given for a lens whose image height function is Equation (2). The coordinates of p' are expressed using the image height function fsinθ according to θ described above, and φ in the figure.
[0036] Figure 4 shows a stereo camera with two fisheye cameras arranged side by side. The coordinate system after the stereo camera is parallelized is (X r ,Y r ,Z r ), (X l ,Y l ,Z l ), and the coordinate system before rectification is (X' r ,Y' r ,Z' r ), (X' l ,Y' l ,Z' l ) In addition, when the ray of light incident on the camera for the subject P is a virtual image plane of a sphere with a radius of the focal length f, the sphere and the light point of the ray are defined as p l ,p rTo achieve parallelization, it is necessary to transform the coordinate system so that the optical axis of the camera is oriented perpendicular to the line connecting the centers of the left and right cameras (called the baseline) (the optical axis after transformation is called the parallelized optical axis). There is some freedom in the direction of the parallelized optical axis, but it is set in the direction defined as the front relative to the position where the housing is attached.
[0037] In addition to setting the coordinate system, how to define the image coordinates is also important for rectification. l ,E r is the epipolar line. When the sphere is expanded onto the image plane, this epipolar line must be parallel to the horizontal axis of the image coordinate system.
[0038] Figure 5 shows one way of setting the image coordinate system to realize such parallelization. The direction of the light beam to the object P is defined as the pitch angle φ′ and the distance between the object P and the origin 0 on the left and right, as shown in Figure 5(1). l ,0 r The horizontal axis is expressed as f sin θ' and the vertical axis is f sin φ', and as shown in Figure 5(2), the horizontal axis of the image coordinate system is f sin θ' and the vertical axis is f sin φ', respectively. There is a degree of freedom in how the horizontal and vertical axes of the image coordinate system are chosen, and the horizontal axis can be x(θ'), a function of any θ', and the vertical axis can be y(φ'), a function of any φ'. Here, the function x(θ') is called the rectification function.
[0039] Figure 6 shows the subject p in Figure 4. l and the left and right origin 0 l ,0 r The plane containing Sinθ' is shown. l , Sinθ′ r From there, By calculating the distance D from the base line length B as TIFF2025162008000004.tif10158, the three-dimensional position of the subject P can be measured.
[0040] As mentioned above, there is a degree of freedom in how the image coordinate system is set after rectification, but to accurately calculate the distance D in equation (3), it is necessary to perform stereo matching with high precision. If the horizontal resolution deteriorates during conversion to a rectified image, the accuracy of the matching position will also deteriorate, so particular care must be taken when designing this rectification function.
[0041] FIG. 7 shows the process of converting into a parallelized image in this embodiment, where FIG. 7(1) is the image before parallelization, FIG. 7(2) is an intermediate image between before and after parallelization, and FIG. 7(3) is the image after parallelization. Image 91 in FIG. 7(1) is a diagram showing an image in the coordinate system of a fisheye lens with a field angle of 180° before parallelization, along with the epipolar line. The lens does not necessarily have to have a field angle of 180°, and any field angle can be used, but for ease of understanding, it is set to 180°. P n , P s are the two ends of the image central axis (epipolar line E0), and are the points that correspond to the so-called North and South Poles in latitude and longitude coordinates.
[0042] The rectification processing unit 104 rectifies the image 91 to generate a rectified image so that the resolution per unit pixel in the baseline direction is higher than that of the pre-rectification image 91 within a predetermined range of angles of view in a plane including the vertical axis. Here, rectification is performed using a rectification function that represents the image height per angle of view in the horizontal direction after rectification, and is designed so that the resolution at any angle of view is equal to or higher than the resolution along the image central axis in the first imaging unit 12 and the second imaging unit 13 and is equal to or lower than the resolution along the widest-angle epipolar line.
[0043] In Figures 7(1) to 7(3), the epipolar line of image 91 before rectification is E, the epipolar line of intermediate image 92 is E', and the epipolar line of image 93 after rectification is E''. First, the image is unfolded so that all epipolar lines E of the image in the fisheye lens coordinate system shown in image 91 before rectification in Figure 7(1) become horizontal while maintaining the distortion, as shown in intermediate image 92 in Figure 7(2). In other words, the image is unfolded so that all epipolar lines become parallel straight lines while maintaining the resolution (length). Then, as shown in image 91 after rectification in Figure 7(3), in this embodiment, the image is rectified so that it matches the epipolar line E'3 with the widest vertical angle of view that appears on the screen in intermediate image 92 in Figure 7(2).
[0044] As can be seen from the epipolar line E' of the intermediate image 92, in the intermediate image 92, the larger the vertical coordinate of the epipolar line, the higher the resolution of the epipolar line E'3 relative to the epipolar line E'0, for example. In this case, if the image is rectified using the rectification function fsinθ', this is the same as rectifying the image to match the shortest epipolar line E'0, and therefore the epipolar line E''3 of the rectified image 93 will be compressed compared to the epipolar line E'3 of the intermediate image 92, resulting in a degradation of resolution compared to the original image 91 before rectification.
[0045] In contrast, in this embodiment, rectification is performed using a rectification function that represents the image height per horizontal angle of view after rectification. Specifically, as in the rectified image 93 in FIG. 7(3), in order to make the resolution constant across the entire screen, the image is rectified using the epipolar line E'3, which has the widest vertical angle of view on the screen, as the reference line. In other words, the rectification function is The image is rectified as TIFF2025162008000005.tif11158. This ensures that the resolution at any angle of view within the image is equal to or greater than the resolution along the image's central axis and equal to or less than the resolution along the widest-angle epipolar line. Therefore, it is possible to rectify the image without degrading the resolution of any epipolar line E. Further enlargement would waste image memory capacity, so this can be considered an upper limit for enlargement. However, if the epipolar line E'3 of the intermediate image 92 is larger than the angle of view of the camera lens, the epipolar line E'3 of the intermediate image 92 will not be included in the intermediate image 92. However, even in this case, the effect of the present invention, which allows rectification without degrading resolution, can be achieved.
[0046] In this embodiment, the image height characteristic of the fisheye lens is r=f sin θ, but in reality, the image height characteristic of the lens includes variations from the design. Furthermore, the projection method is not limited to orthogonal projection f sin θ, and may be, for example, a general equidistant projection fθ, solid angle projection 2f sin(θ / 2), or equisolid angle projection 2f tan(θ / 2), or the effects of the present invention can be obtained with other projections. Therefore, the image height characteristic may be generalized as r=r(θ). Furthermore, the epipolar line when generalized is expressed as follows in the fisheye coordinate system of the image 91 before parallelization, with φ as a constant: TIFF2025162008000006.tif12156. Here, x(θ′) obtained by horizontally expanding the epipolar line of the vertical angle of view φ′ as in intermediate image 92 is given by TIFF2025162008000007.tif14156, which means cosθ=cosθ′cosφ′ It can be calculated using the following relationship:
[0047] [Example of hardware configuration of image processing device 10] Next, a description will be given of the hardware configuration of the image processing device 10 according to this embodiment. Fig. 8 is a block diagram showing an example of the hardware configuration of the image processing device 10 according to this embodiment.
[0048] 8, the image processing device 10 includes a CPU (Central Processing Unit) 81, a ROM (Read Only Memory) 82, a RAM (Random Access Memory) 83, a storage device 84, an input / output interface 85, and a bus 86. The bus 86 is a signal path that electrically connects the components together and allows input and output of information data between the components.
[0049] The CPU 81 controls the operation of each unit in the image processing device 10. For example, the CPU 81 controls the processing in each of the image acquisition unit 101, the image pre-processing unit 103, and the image interpolation unit 105. The CPU 81 also controls the processing in each of the parallelization processing unit 104, the image interpolation unit 105, the parallax image generation unit 106, the three-dimensional object recognition unit 107, and the vehicle control unit 108. Note that a GPU (Graphics Processing Unit) may be used instead of the CPU 81, or the CPU 81 and a GPU (Graphics Processing Unit) may be used together.
[0050] The ROM 82 is configured as a storage medium such as a nonvolatile memory, and stores programs and data executed and referenced by the CPU 81. The RAM 83 is configured as a storage medium such as a volatile memory, and temporarily stores information (data) required for each process performed by the CPU 81.
[0051] The storage device 84 is configured as a computer-readable non-transitory recording medium storing a software program executed by the CPU 81, and is configured as a storage device such as an HDD (Hard Disk Drive). The storage device 84 stores programs for the CPU 81 to control various parts, an OS (Operating System), controller programs, and other data. Note that some of the programs and data stored in the storage device 84 may be stored in the ROM 82. The computer-readable non-transitory recording medium storing the program executed by the CPU 81 is not limited to an HDD, and may be, for example, a recording medium such as an SSD (Solid State Drive), a CD (Compact Disc)-ROM, or a DVD (Digital Versatile Disc)-ROM.
[0052] The input / output interface 85 transmits and receives information data to and from the outside of the image processing device 10 under the control of the CPU 81. The image processing device 10 embodies the various components described in FIG. 1 through cooperation between these hardware configurations and software programs.
[0053] [effect] As described above, in the image processing device 10 according to this embodiment, a rectification image conversion table is generated that performs rectification so as to maintain the resolution, and is stored in the rectification image conversion table storage unit 109. This rectification image conversion table makes it possible to generate a rectified image that achieves both high ranging accuracy and memory saving.
[0054] It should be noted that the present invention is not limited to the above-described embodiments, and various other applications and modifications are possible without departing from the gist of the present invention as set forth in the claims.
[0055] For example, the above-described embodiments have described in detail and specifically the configuration of an image processing device in order to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0056] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected.
[0057] Second Embodiment In the first embodiment, the reference line is selected by using the epipolar line with the widest vertical angle of view that appears on the screen to perform collimation, with the aim of maintaining a constant resolution across the entire screen. However, the range of vertical angles of view for which accurate distance measurement is required may be limited. In other words, image collimation may be performed so as to maintain a constant resolution across a specified range of vertical angles of view, rather than across the entire image.
[0058] Figure 9 shows the relative positions of the bump and the on-board camera. For example, a speed breaker (bump) 120 that encourages drivers to slow down is installed on the road surface 121, and if the vehicle V attempts to pass over this at a high speed, a large impact will be applied to the vehicle V, which may affect the ride comfort of the occupants, driving operation, and the running of the vehicle V. Therefore, it is required to measure the distance to the bump 120 precisely and use the information to warn the occupants and to control the braking of the vehicle V.
[0059] When the height of the stereo camera 100 from the ground is h and the distance in the optical axis direction from the stereo camera 100 to the bump 120 is the measurement distance L1, the vertical angle of view α is tan -1For example, if the host vehicle V is a normal vehicle such as a passenger car, then when h=1.6 m and L=5 m, α=approximately 20°.
[0060] FIG. 10 shows the process of converting to a parallelized image in this embodiment, where FIG. 10(1) is the image before parallelization, FIG. 10(2) is the intermediate image between before and after parallelization, and FIG. 10(3) is the image after parallelization.
[0061] In the example shown in Figure 10, the epipolar line E passing through the vertical angle of view α α By performing parallelization in this manner, there is no degradation in the resolution within the field of view where it is desired to accurately detect the bump 120, which is the detection target.
[0062] The rectification processing unit 104 rectifies the image 140 to generate a rectified image 142 so that the resolution per unit pixel in the baseline direction is higher than that of the pre-rectification image 140 in a predetermined angle of view range r(α) in a plane including the vertical axis, but is not higher than that of the pre-rectification image 140 in a range outside the predetermined angle of view range r(α). That is, in a region (within the angle of view range r(α)) perpendicular to the direction in which the multiple imaging units are arranged, the first rectified image and the second rectified image have a larger number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image, and in a region (outside the angle of view range r(α)) perpendicular to the direction in which the multiple imaging units are arranged that is larger than the vertical angle of view α, the first rectified image and the second rectified image have a smaller number of pixels per unit angle of view in at least the direction in which the multiple imaging units are arranged than the first captured image and the second captured image. In Figures 10(1) to 10(3), the epipolar line of the image 140 before rectification is Eα, the epipolar line of the intermediate image 141 is E'α, and the epipolar line of the image 142 after rectification is E''α.
[0063] The image is expanded so that all epipolar lines E of the image in the fisheye lens coordinate system shown in the image 140 before rectification in FIG. 10(1) become horizontal while maintaining the distortion, as shown in the intermediate image 141 in FIG. 10(2). In other words, the image is expanded so that all epipolar lines become parallel straight lines while maintaining the resolution (length). Then, as shown in the image 142 after rectification in FIG. 10(3), the epipolar line E'' passing through the vertical angle of view α α The image is collimated using the reference line of the collimation function as the reference line of the collimation function. As a result, the resolution per unit pixel in the baseline direction exceeds that of the image 140 before collimation within a predetermined range of angle of view r(α) on a plane including the vertical axis, and the image can be collimated without degrading the resolution within the field of view where the bump 120, which is the detection target, is to be accurately detected.
[0064] The parallelization processing unit 104 may also set the vertical angle of view depending on the type of object. The vertical angle of view condition can be set by dividing an area by the horizontal angle of view and imposing different conditions on each area. Fig. 11 shows a bump 120 across the width of a road surface 121 and a pedestrian 122 about to cross the road surface 121. Fig. 12 shows an image 140 in the coordinate system of the fisheye lens before parallelization, plotted along an epipolar line E. α , E β Assuming that the bump 120 and the pedestrian 122 are located closest to the host vehicle V, the vertical angle of view will be within the angle of view range r(α) if they are within the range of the horizontal angle of view γ1, and the vertical angle of view will be within the angle of view range r(β) if they are within the range of the horizontal angle of view γ2.
[0065] At this time, as shown in Figure 12, for each horizontal field of view area, an epipolar line E corresponding to the vertical field of view is drawn. α and E β It is also possible to select the parallelization functions and connect them to form a new parallelization function. In other words, within the range of the horizontal angle of view γ1, the epipolar line E'' passing through the vertical angle of view α is α The image is rectified using the reference line of the rectification function. In the range of horizontal angle of view γ2, the epipolar line E'' passing through the vertical angle of view β is βThe image can be rectified using the horizontal angle of view γ1 as the reference line of the rectification function. The vertical angles of view α and β can also be used depending on the speed of the host vehicle V. In other words, the vertical angles of view α and β can be changed depending on the speed of the moving object. FIG. 12 shows an image 140 before rectification, and at a horizontal angle of view γ1, the epipolar line E α At the horizontal angle of view γ2, the epipolar line E β In other words, the parallelization processing unit 104 can parallelize the captured image based on a parallelization function that outputs the image height in the direction in which the multiple image capturing units are arranged, which differs depending on the angle of view ranges of the multiple image capturing units.
[0066] In addition, an upper limit to the vertical angle of view can also occur when the camera's angle of view is less than 180°. The upper limit epipolar line may not be a simple epipolar line for lenses with a polar angle of view other than 180°.
[0067] FIG. 13 is a diagram showing collimation in the case of a wide-angle lens with an angle of view of less than 180°. Image 183 in Figure 13(1) is a diagram showing an image in the coordinate system of a wide-angle lens before collimation along with the epipolar line, Figure 13(2) is an intermediate image between before and after collimation, and Figure 13(3) is the image after collimation.
[0068] If the maximum vertical angle of view of the camera is α, then the epipolar line E is α is determined, and rectification is performed in the range outside the maximum vertical angle of view α so that the resolution per unit pixel in the baseline direction does not exceed the maximum vertical angle of view α. At this time, the size of the final image varies depending on the maximum horizontal angle of view of the camera. For example, in the case of a wide-angle lens whose maximum horizontal angle of view is smaller than that of image 180, as in image 183 before rectification in Figure 13(1), the horizontal size will also be smaller, as in intermediate image 181 in Figure 13(2) and image 182 after rectification in Figure 13(3).
[0069] Furthermore, if the epipolar line of the vertical angle of view is followed, it is not possible to ensure the resolution of areas beyond the diagonal corners of the camera's imaging range. Therefore, as shown in Figure 14, it is also possible to determine the maximum epipolar line for each horizontal angle of view γ1, γ2 and perform parallelization so that the resolution does not exceed that value.
[0070] Furthermore, up until now, an appropriate epipolar line has been selected, and rectification has been performed according to a rectification function that follows the epipolar line. However, from the viewpoint of maintaining resolution, it is not necessary to use a rectification function that follows the epipolar line. As long as the resolution, defined by the number of pixels per unit angle of view, is at least larger than the epipolar line E0 in FIG. 7 and smaller than the epipolar line E3 at any horizontal angle of view, the image resolution will not be degraded across the entire image, and there will be no unnecessary enlargement. This is because, if the rectification function that follows the epipolar line E0 is x0(θ'), the rectification function that follows the epipolar line E3 is x3(θ'), and the rectification function is x(θ'), then at any θ', It is sufficient if x(θ′) is such that TIFF2025162008000008.tif5157. For example, using the formulas (2) and (4) used in the first embodiment, TIFF2025162008000009.tif15157, TIFF2025162008000010.tif5159, the desired x(θ') can be obtained.
[0071] Under the above resolution constraints, it is possible to choose a reference line instead of the epipolar line. Figure 15 shows an example of a suitable reference line L.
[0072] Third Embodiment Most of the configuration and processing in this embodiment is the same as in the first embodiment. The difference is that the relationship between the incident light ray after collimation and the incident position on the image coordinate system is as shown in Figure 16 instead of Figure 5. Specifically, the angle from the X axis when the light ray is projected onto the ZX plane is defined as θ'.
[0073] 17 shows the projection of the subject P onto the ZX plane in FIG. 6. In this case, D, which is calculated in the same manner as in Equation (3), always represents the depth, not the distance from the baseline. In this embodiment, the depth is directly calculated from the parallax, making this a configuration that is well suited as a projection method for a front-sensing sensor.
[0074] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the above-described embodiments, and various design modifications can be made without departing from the spirit of the present invention as defined in the claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations. [Explanation of symbols]
[0075] 10...image processing device, 11...imaging unit (plurality of imaging units), 12...first imaging unit (left imaging unit), 13...second imaging unit (right imaging unit), 100...stereo camera, 101...image acquisition unit, 103...image pre-processing unit, 104...parallax image processing unit, 105...image interpolation unit, 106...parallax image generation unit, 107...three-dimensional object recognition unit (object recognition unit), 108...vehicle control unit, 109...parallelization image conversion table storage unit (conversion information storage unit)
Claims
1. an image acquisition unit that acquires a first captured image and a second captured image captured by the plurality of imaging units; a parallelization processing unit that parallelizes the first captured image and the second captured image to generate a first parallelized image and a second parallelized image, with a direction perpendicular to a baseline direction defined by the plurality of image capturing units and to the optical axis directions of the plurality of image capturing units as a vertical axis, so that vertical coordinates of identical images captured in the first captured image and the second captured image coincide with each other at least at any point in the first captured image and the second captured image; a parallax image generator that generates parallax based on the plurality of parallelized images; an object recognition unit that recognizes an object based on the parallax, the parallelization processing unit parallelizes the first captured image and the second captured image so that a resolution per unit pixel in the base line direction exceeds a resolution per unit pixel in a predetermined range of an angle of view in a plane including the vertical axis, with respect to the first captured image and the second captured image.
1. An image processing device comprising:
2. 2. The image processing device according to claim 1, When a predetermined vertical angle of view α is set with respect to a direction perpendicular to the direction in which the plurality of imaging units are arranged, in a region equal to or smaller than the vertical angle of view α in a direction perpendicular to the direction in which the plurality of imaging units are arranged, the first parallelized image and the second parallelized image have a larger number of pixels per unit angle of view in at least the direction in which the plurality of imaging units are arranged than the first captured image and the second captured image, In a region where the vertical angle of view α is larger than the vertical angle of view in a direction perpendicular to the direction in which the plurality of imaging units are arranged, the first parallelized image and the second parallelized image have a smaller number of pixels per unit angle of view in at least the direction in which the plurality of imaging units are arranged than the first captured image and the second captured image.
1. An image processing device comprising:
3. 2. The image processing device according to claim 1, The parallelization processing unit The parallelization is performed using a parallelization function shown in the following equation (4):
1. An image processing device comprising: (where f is the focal length, and θ′ is the yaw angle in a plane including the subject and the direction in which the multiple imaging units are arranged)
4. 2. The image processing device according to claim 1, the parallelization processing unit parallelizes the captured image based on a parallelization function that outputs an image height in an arrangement direction of the plurality of image capturing units, the image height varying depending on an angle of view range of the plurality of image capturing units.
1. An image processing device comprising:
5. 3. The image processing device according to claim 2, When the height of the plurality of imaging units from the ground is h and the corresponding distance measurement distance is L1, The vertical angle of view α is expressed as tan -1 (h / L1) or more, 1. An image processing device comprising:
6. 2. The image processing device according to claim 1, epipolar lines of the plurality of imaging units are parallel to a direction in which the plurality of imaging units are arranged in the first rectified image and the second rectified image; 1. An image processing device comprising:
7. 2. The image processing device according to claim 1, a conversion information storage unit storing conversion information for radial distortion of lenses mounted on the plurality of imaging units; the parallelization processing unit parallelizes the first captured image and the second captured image based on the conversion information so that parallelization magnifications of the vertical axis and the horizontal axis of the first parallelized image and the second parallelized image are determined by optical distortion of the lens according to radial distortion of the lens.
1. An image processing device comprising:
8. 2. The image processing device according to claim 1, The parallelization processing unit unfolding the first captured image and the second captured image so that epipolar lines of the first captured image and the second captured image become parallel straight lines while maintaining their lengths; selecting an epipolar line that makes the resolution of the captured image constant throughout the entire image as a reference line; Parallelizing the first captured image and the second captured image in accordance with the reference line.
1. An image processing device comprising:
9. 9. The image processing device according to claim 8, The parallelization processing unit An image processing device comprising: an image processing unit for selecting the epipolar line in accordance with a predetermined vertical angle of view of the captured image;
10. 10. The image processing device according to claim 9, The parallelization processing unit An image processing device comprising: selecting an epipolar line having the widest vertical angle of view.
11. 10. The image processing device according to claim 9, The parallelization processing unit an image processing device that sets the vertical angle of view depending on the type of the object;
12. 10. The image processing device according to claim 9, The image processing device, wherein the parallelization processing unit changes the vertical angle of view in accordance with the speed of a moving object on which the image processing device is mounted.
13. 10. The image processing device according to claim 9, The parallelization processing unit 10. An image processing device, comprising: a first area of a captured image having a horizontal angle of view; a second area of a captured image having a vertical angle of view;
14. an image acquiring step of acquiring a first captured image and a second captured image captured by the plurality of imaging units; a parallelization processing step of parallelizing the first captured image and the second captured image to generate a first parallelized image and a second parallelized image, with a direction perpendicular to a baseline direction defined by the plurality of image capturing units and to the optical axis directions of the plurality of image capturing units as a vertical axis, so that vertical coordinates of identical images captured in the first captured image and the second captured image coincide with each other at least at any point in the first captured image and the second captured image; a parallax generating step of generating parallax based on the plurality of rectified images; an object recognition step of recognizing an object based on the parallax, In the rectification processing step, the first captured image and the second captured image are rectified to generate the first rectified image and the second rectified image so that the resolution per unit pixel in the base line direction is greater than that of the first captured image and the second captured image within a predetermined range of angle of view in a plane including the vertical axis. An image processing method comprising:
Citation Information
Patent Citations
Disparity estimation from wide-angle images
JP2022506104A