Image processing apparatus, image processing method, and program
The image processing device addresses the resolution decrease at joints in bird's-eye-view images by generating patch groups and setting high-resolution magnification for low-resolution camera data, resulting in improved image quality with consistent resolution across the composite image.
Patent Information
- Application Number
- JP2023181864
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2025-05-08
AI Technical Summary
When combining high-resolution and low-resolution images from multiple cameras to generate bird's-eye-view images, the resolution decreases at the joints or overlapping areas, affecting the subjective quality of the image.
An image processing device that synthesizes image data from multiple cameras, generates patch groups from low-resolution camera data, and sets high-resolution magnification for each patch to match the target resolution based on the high-resolution camera, thereby reducing resolution differences at joints.
The solution effectively reduces the decrease in resolution at the joints of bird's-eye-view images, improving the overall image quality by ensuring consistent resolution across the composite image.
Smart Images

Figure 2025071580000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device and an image processing method. [Background technology]
[0002] There is known a vehicle display device that stitches together images captured by multiple cameras mounted on the vehicle to generate and display an image (bird's-eye view image) of the surroundings of the vehicle as seen from a virtual viewpoint above the vehicle. For example, in Patent Document 1, an image from an imaging device is projected onto a curved, bowl-shaped surface, and an image of the curved surface as seen from a virtual viewpoint is generated. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 3286306 specification Summary of the Invention [Problem to be solved by the invention]
[0004] However, when combining images from a relatively high-resolution imaging device and a relatively low-resolution imaging device at the seams or overlapping areas between imaging devices in an overhead image, the resolution decreases at the seams. For example, as shown in Figure 1, when imaging devices are installed on a long vehicle such as a truck, the distance between the left or right camera and the rear camera is large. In this positional relationship, at the seams or overlapping areas, the image from the left or right camera will be low-resolution due to the distance, while the image from the rear camera will be high-resolution due to the relatively close distance. If such a resolution difference occurs during composition, the subjective quality of the image quality will be impaired. If a blending process is applied to prevent this, a low-resolution image will be blended with a high-resolution image, causing a decrease in resolution in the blended area.
[0005] An object of the present invention is to provide an image processing device that synthesizes images acquired from multiple imaging devices to generate an overhead image, and that can reduce the reduction in resolution when synthesizing images from an imaging device with a relatively high resolution and an imaging device with a relatively low resolution at the joints or overlapping areas between the imaging devices. [Means for solving the problem]
[0006] In order to achieve the above-mentioned object, the image processing device of the present invention is an image processing device that is equipped with at least two or more imaging devices, combines image data acquired from the imaging devices, and generates a composite image, and has: a first imaging device that provides relatively high resolution in the joints or overlapping areas of the composite image; a second imaging device that provides relatively low resolution in the joints or overlapping areas; a generation means that generates a group of patches from the image data acquired from the second imaging device (Figure 4 - generation unit 12); and a magnification setting means that sets a high-resolution magnification for each patch of the generated group of patches (Figure 4 - magnification setting unit 13), and is characterized in that the high-resolution magnification is set to a value corresponding to the target resolution set based on the first imaging device. [Effects of the Invention]
[0007] According to the present invention, in an image processing device that synthesizes images acquired from multiple imaging devices to generate an overhead image, it is possible to reduce the reduction in resolution when synthesizing images from an imaging device with a relatively high resolution and an imaging device with a relatively low resolution at the joints or overlapping areas between the imaging devices. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 10 is a diagram showing the arrangement of imaging devices and joints. [Figure 2] FIG. 2 is a diagram illustrating geometric conditions of an imaging device and a subject. [Figure 3] FIG. 1 is a diagram illustrating a hardware configuration of an image processing apparatus. [Figure 4] FIG. 1 is a diagram illustrating a functional configuration of an image processing apparatus according to a first embodiment. [Figure 5] 4 is a flowchart showing processing performed by the image processing device according to the first embodiment. [Figure 6] 10 is a flowchart of a process for setting a high resolution magnification in the first embodiment. [Figure 7] 10 is a flowchart of a process for setting a high resolution magnification in the second embodiment. [Figure 8] 10 is a second flowchart of the process of setting a high resolution magnification in the second embodiment. [Figure 9] FIG. 10 is a diagram illustrating a functional configuration of an image processing apparatus according to a third embodiment. [Figure 10] 10 is a flowchart showing processing by an image processing device according to a third embodiment. [Figure 11] 10 is a flowchart of a process for setting a high resolution magnification in the third embodiment. [Figure 12] FIG. 10 is a diagram illustrating an interpolation process. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present invention, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the present invention. Note that the same components will be described with the same reference numerals.
[0010] <About resolution> In an embodiment of the present invention, the resolution is defined as the length of an object captured in real space at one pixel in image data. Resolution is calculated for each pixel of image data in the imaging device. If one pixel is a pixel of interest, the resolution at the pixel of interest is calculated from the distance between the imaging device and the object captured by the pixel of interest. The resolution calculation method is described with reference to FIG. 2. FIG. 2 shows the geometric conditions between the imaging device and the object at the pixel of interest. Here, the focal length of the imaging device is f [mm], the horizontal width of the image sensor is SW [mm], the total number of pixels in the horizontal direction of the sensor is x [pixel], and the angle between the projection line onto the object captured by the pixel of interest and the optical axis of the lens is θ [rad]. Furthermore, d [mm] represents the distance from the imaging device to the object captured by the pixel of interest in the optical axis direction. In this case, the pitch w [mm / pixel] is SW / x. Furthermore, the coordinates of the pixel of interest in the image data are defined as (u0, v0). In the pinhole camera model, Equation (1) holds due to the similarity of triangles.
[0011]
number
[0012] Here, S is the resolution [mm / pixel] and is the length of the object captured by pixel of interest 1 [pixel]. From equation (1), equation (2) can be obtained.
[0013]
number
[0014] The focal length f and pitch w can be obtained by obtaining design values or performing internal calibration in advance. In other words, the resolution can be determined by determining the distance d.
[0015] [First embodiment] In the first embodiment, a method for setting the resolution increasing magnification of a low-resolution camera in accordance with one of the resolutions of a relatively high-resolution camera will be described.
[0016] <Hardware configuration of image processing device> An example of the configuration of an image processing device in this embodiment will be described with reference to Fig. 3. The image processing device 100 in this embodiment includes a CPU 101, a RAM 102, a ROM 103, a secondary storage device 104, an input interface 105, and an output interface 106. The components of the image processing device 100 are connected to one another by a system bus 107. The image processing device 100 is also connected to an external storage device 108 via the input interface 105, and to the external storage device 108 and a display device 109 via the output interface 106.
[0017] The CPU 101 is a processor that uses the RAM 102 as a work memory, executes programs stored in the ROM 103, and performs overall control of each component of the image processing device 100 via a system bus 107. This allows various processes, which will be described later, to be executed.
[0018] The secondary storage device 104 is a storage device that stores various data handled by the image processing device 100, and in this embodiment, a HDD is used. The CPU 101 can write data to the secondary storage device 104 and read data stored in the secondary storage device 104 via the system bus 107. Note that the secondary storage device 104 can be various storage devices other than a HDD, such as an optical disk drive or flash memory.
[0019] The input interface 105 is a serial bus interface such as USB or IEEE1394, and data, commands, and the like are input from an external device to the image processing device 100 via this input interface 105. The image processing device 100 acquires data from an external storage device 108 (e.g., a storage medium such as a hard disk, memory card, CF card, SD card, or USB memory) via this input interface 105. Note that input devices such as a mouse and buttons (not shown) can also be connected to the input interface 105. The output interface 106, like the input interface 105, has a serial bus interface such as USB or IEEE1394. Alternatively, a video output terminal such as DVI or HDMI can also be used. Data and the like are output from the image processing device 100 to an external device via this output interface 106. The image processing device 100 displays image data by outputting processed image data and the like to a display device 109 (e.g., various image display devices such as a liquid crystal display) via this output interface 106. Note that the image processing device 100 includes other components than those described above, but these are not the focus of the present invention and therefore will not be described here.
[0020] <Functional configuration of image processing device> 4 is a block diagram showing the functional configuration of an image processing device, which is made up of an input unit 11, a generation unit 12, a magnification setting unit 13, and an output unit 14.
[0021] The input unit 11 reads image data captured by multiple imaging devices installed in the vehicle via the input interface 105 or from the secondary storage device 104. It also acquires imaging device information and three-dimensional information. The image processing device 100 may be connected to multiple imaging devices and configured as an image processing system including the image processing device 100 and multiple imaging devices. With this configuration, it is possible to generate overhead images in real time. Furthermore, when multiple imaging devices capture video, the image processing device 100 can perform processing using frames captured by the multiple imaging devices at approximately the same time. The input unit 11 outputs image data, imaging device information, and three-dimensional information to the generation unit 12 and the magnification setting unit 13.
[0022] The generation unit 12 generates a set of patches (patch group) that are units for setting a high-resolution magnification. The generation unit 12 outputs the patch group to the magnification setting unit 13.
[0023] The magnification setting unit 13 sets a high-resolution magnification indicating how much the resolution should be increased by super-resolution processing for each patch in the patch group. The magnification setting unit outputs the high-resolution magnification to the output unit 14.
[0024] The output unit 14 outputs the resolution-enhancing magnification.
[0025] <Processing of image processing device> 5 is a flowchart showing the processing executed by the image processing device. Hereinafter, each step (process) will be represented by adding an S before the reference number.
[0026] In step S101, the input unit 11 acquires imaging device information. FIG. 1 shows a hardware configuration representing the arrangement of multiple imaging devices in this embodiment. FIG. 1 shows the imaging devices, which are configured as a front camera 21, a left camera 22, a right camera 23, a rear camera 24, and a vehicle 25. When compositing images such as an overhead image, the position indicated by the dashed line in FIG. 1 becomes a seam. Because the distance from the seam to the left camera 22 is greater than the distance from the seam to the rear camera 24, the left camera 22 has low resolution near the seam or in the overlapping area. As a result, a difference in resolution occurs at the seam. This also applies to the relationship between the right camera 23 and the rear camera 24. Note that the low resolution near the seam is not necessarily determined by the distance to the seam; the resolution at the seam or overlapping area also varies depending on the sensor size and lens characteristics of the imaging devices. The positional relationship between imaging devices with relatively low resolution and imaging devices with relatively high resolution is acquired as imaging device information.
[0027] The positional relationship includes the position of the high-resolution camera in the low-resolution camera coordinate system (translation matrix T) and a rotation matrix R that represents the rotation from the low-resolution camera coordinate system to the high-resolution camera coordinate system. The translation matrix T and rotation matrix R are called extrinsic parameters and can be obtained in advance using a known external calibration method. Here, the camera coordinate system is one in which the positive direction of the z-axis is the optical axis direction of the imaging device, the positive direction of the y-axis is the downward direction of the imaging device, and the positive direction of the x-axis is the direction perpendicular to the y-axis and z-axis. Note that the camera coordinate system may also use other axes.
[0028] In the present embodiment, the following description will focus on the processing of image data from the left camera 22 and rear camera 24. Hereinafter, the left camera 22 will be referred to as the low-resolution camera, and the rear camera 24 will be referred to as the high-resolution camera. Note that cameras other than the left camera 22 and rear camera 24 may be used as long as there is a low-resolution and high-resolution relationship near the seam. For example, the right camera 23 may be the low-resolution camera, and the rear camera 24 may be the high-resolution camera, and the same processing may be applied.
[0029] The imaging device information may also include the focal length or principal point position, the number of pixels in the vertical and horizontal directions, distortion parameters indicating distortion of the image data, F-number, shutter speed, white balance, angle of view, etc.
[0030] In S102, the input unit 11 reads image data from the low-resolution camera and image data from the high-resolution camera. In this embodiment, the acquired image data will be described as having three RGB channels.
[0031] In S103, the input unit 11 reads a three-dimensional model. This three-dimensional model indicates the relative positional relationship between the low-resolution camera and the subject, and the relative positional relationship between the high-resolution camera and the subject. In this embodiment, a distance map (depth map) between the low-resolution camera and the subject, and a distance map between the high-resolution camera and the subject are acquired as three-dimensional models. Methods for generating distance maps are well known, and any method can be used. For example, a stereo matching method can be used to generate a distance map of the subject. Alternatively, a distance map of the subject can be generated using some kind of tracker or the like, and a distance map can be generated based on this three-dimensional model. Alternatively, a distance map can be acquired by measuring the distance from each camera to the corresponding subject in advance using a range sensor or the like.
[0032] In S104, the generation unit 12 generates a group of patches from the image data of the low-resolution camera. Specifically, the upper left pixel of the image data of the low-resolution camera is set as the pixel of interest. A square patch centered on this pixel of interest is generated. The group of patches is generated by raster scanning, in which the generated square patch is shifted pixel by pixel. Note that the patch does not have to be square and may be a shape such as a rectangle or a circle. Furthermore, the pixel of interest does not have to be the center. The pixel of interest may be located at a position other than the center of the patch, or may be outside the patch. Furthermore, if the pixel of interest is located near the edge of the image data, there may be no image data even if a square centered on the pixel of interest is created. In this case, padding can be used to enable processing. Furthermore, instead of setting the pixel of interest from the upper left of the image data, the pixel of interest may be set at a position half the width of the square to the lower right from the upper left pixel. Furthermore, the pixel of interest does not have to be the upper left of the image data, and may be set only within the range of the image data to be used after synthesis. For example, only pixels used in the overhead image may be set. Furthermore, the raster scan may be performed by shifting multiple pixels instead of by shifting one pixel at a time.
[0033] In S105, the magnification setting unit 13 selects one of the patches generated by the generation unit 12 as a target patch, and the process proceeds to S106. If the processing of all patches has been completed, the process proceeds to S107.
[0034] In S106, the magnification setting unit 13 sets a high-resolution magnification for the set target patch. Figure 6 is a flowchart for explaining the process of S106 in detail.
[0035] In S1001, the magnification setting unit 13 obtains the resolution Slow at the pixel of interest using equation (2). At this time, the focal length f, pitch w, principal point position (Cx, Cy), and distance d(u0, v0) read from the input unit 11 are used.
[0036] In S1002, the magnification setting unit 14 sets a target resolution. For example, the target resolution Starget is set to the lowest resolution among the resolutions of image data from a high-resolution camera used to generate the overhead video. Note that depending on the angle of view, distant objects can be captured. Even if the resolution is relatively high at the seams or overlapping areas, the resolution will be low in the distance. In this case, the target resolution Starget will be small. To prevent this, a range of image data for which the target resolution is set may be set in advance. For example, the range of pixels used as the overhead image may be stored in advance as a mask image in a storage unit such as the secondary storage device 104, and this may be read and used as the lowest resolution within the range of that area.
[0037] In S1003, if the resolution Slow at the pixel of interest is equal to or greater than the target resolution Starget, the magnification setting unit 14 sets the resolution-enhancing magnification to 1, and the process proceeds to S105. If Slow is less than Starget, the process proceeds to S1004.
[0038] In S1004, the magnification setting unit 14 sets a high-resolution magnification, which is calculated by Starget / Slow, and then the process proceeds to S105.
[0039] In S107, the output unit 14 outputs the resolution increase magnification set for each patch.
[0040] The image processing device 100 may perform super-resolution using this high-resolution magnification. Super-resolution methods are well known, and any method can be used. For example, super-resolution using deep learning technology or, in overlapping areas, a reconstruction-type super-resolution method using image data from a high-resolution camera may be applied. Furthermore, for images in which the high-resolution magnification is 1 in S1003, super-resolution processing may not be applied.
[0041] Furthermore, the magnification setting unit 14 calculates the super-resolution magnification by Starget / Slow, but in many cases super-resolution only supports magnifications that are integers or multiples of 2. In such cases, a compatible magnification may be obtained by rounding up, rounding down, or performing a remainder operation.
[0042] In addition, in this embodiment, the camera is installed on a truck as shown in Fig. 1, but the imaging device may be installed on other vehicles such as a passenger car. Also, the imaging device does not have to be a vehicle as long as multiple imaging devices are installed.
[0043] Furthermore, although the above description is given with respect to a four-camera configuration consisting of a front camera, a left camera, a right camera, and a rear camera, the configuration is not necessarily limited to four cameras, and any configuration with two or more cameras may be used.
[0044] Furthermore, in this embodiment, RGB image data has been described as the target, but image data in other color spaces or other wavelength regions, or image data with a number of channels other than 3ch, may also be used.
[0045] Furthermore, the imaging device information does not necessarily have to be the same between imaging devices, and the focal length, sensor size, number of pixels, etc. may differ between imaging devices.
[0046] Although the pinhole camera model is used in this embodiment, other models may be used. For example, when a central projection camera model is used, the resolution S is calculated by the following equation (3).
[0047]
number
[0048] When an equidistant projection fisheye lens model is used, the resolution S is calculated using equation (4).
[0049]
number
[0050] In this embodiment, the resolution is determined by the length of the object in real space represented by one pixel in the image data, but other indices may be used. The spatial frequency components of the image data may be calculated to evaluate the sharpness and determine the resolution. In this case, for example, the spatial frequency components of a patch centered on a target pixel are acquired, and the spatial frequency components are used as the target resolution to set the resolution enhancement magnification of the target patch.
[0051] As described above, according to this embodiment, when combining image data from a high-resolution camera and a low-resolution camera, the image data from the low-resolution camera can be super-resolved to match the resolution of the high-resolution camera, and a drop in resolution near the seams of overhead images can be reduced.
[0052] In the first embodiment, a method was described in which one of the resolutions of a relatively high-resolution camera is set as a target resolution (Starget) and the resolution magnification of a low-resolution camera is set. However, the resolution near the joints of the high-resolution cameras is not necessarily uniform. In such cases, even if the image data of the low-resolution camera is super-resolved to match a single target resolution, there is a possibility that a resolution step will remain. In the method of the second embodiment, in order to further reduce the resolution step, the target resolution (Starget) is made variable and a high-resolution magnification of the low-resolution camera image data is set.
[0053] The hardware configuration and functional configuration of the image processing device in this embodiment are the same as those in the first embodiment, and therefore will not be described. The following mainly describes the parts that are different from the first embodiment. The same components will be described with the same reference numerals.
[0054] <Processing of image processing device> The flowchart showing the processing executed by the image processing device is the same as that shown in FIG.
[0055] In S106, the magnification setting unit 13 sets a high-resolution magnification for the set target patch. Fig. 7 is a flowchart for explaining the process of S106 in this embodiment in detail.
[0056] In S2001, the magnification setting unit 13 acquires the distance as the positional relationship between the three-dimensional position of the object captured in the pixel of interest and the high-resolution camera. The three-dimensional position (x0, y0, z0) of the object captured in the pixel of interest is calculated using equation (5).
[0057]
number
[0058] Here, (u0, v0) are the coordinates of the pixel of interest in the image data of the low-resolution camera, f is the focal length of the low-resolution camera, and (Cx, Cy) are the principal point position. Also, d(u0, v0) is the distance in the optical axis direction (z-axis direction) from the viewpoint of the low-resolution camera to the object captured at the pixel of interest, as shown in the distance map. Then, the Euclidean distance d between the three-dimensional position (x0, y0, z0) of the object captured at the pixel of interest and the position of the high-resolution camera (xcam, ycam, zcam) in the low-resolution camera coordinate system obtained from the input unit 11 is calculated.
[0059] In S2002, the magnification setting unit 13 obtains and sets the target resolution by substituting the distance d into equation (2).
[0060] In S106, the magnification setting unit may use the resolution of the high-resolution camera at the closest position among the subjects captured by the high-resolution cameras, instead of the resolution according to the distance to the high-resolution camera as shown in Fig. 7. A flowchart showing an example of this is shown in Fig. 8.
[0061] In S3001, the magnification setting unit 13 first calculates the three-dimensional position (x0, y0, z0) of the object captured at the pixel of interest using equation (5). Next, among the objects captured by the high-resolution camera, the nearest pixel in the image data of the high-resolution camera that is nearest to the three-dimensional position (x0, y0, z0) of the object captured at the pixel of interest is found. A pixel of the high-resolution camera is defined as (uihigh, vihigh). i is a value indicating the ith pixel in the image data of the high-resolution camera. By substituting the focal length of the high-resolution camera as fhigh and the principal point position as (Cxhigh, Cyhigh) into equation (5), equation (6) is obtained.
[0062]
number
[0063] Here, (xihigh, yihigh, zihigh) is the position of the object captured at the i-th pixel in the image data of the high-resolution camera in the high-resolution camera coordinate system. To calculate the distance, the coordinate system is converted to the low-resolution camera coordinate system. The conversion formula is expressed by Equation (7).
[0064]
number
[0065] In equation (7), (xilow, yilow, zilow) represent the three-dimensional position of the object captured by (uihigh, vihigh) in the low-resolution camera coordinates. R represents the rotation matrix from the low-resolution camera coordinates to the high-resolution camera coordinates. T represents the translation matrix from the low-resolution camera coordinates to the high-resolution camera coordinates.
[0066] For all pixels of the high-resolution camera, the processing shown in equations (6) and (7) is applied to find the number i of the pixel that is closest to (x0, y0, z0). The pixel in the image data of the high-resolution camera indicated by number i is the nearest pixel.
[0067] In S3002, the resolution of the i-th pixel of the high-resolution camera image data is calculated using equation (2), and this is set as the target resolution.
[0068] As described above, according to this embodiment, the target resolution can be set according to the three-dimensional distance between the high-resolution camera and the object captured by the pixel of interest. Also, according to this embodiment, among the objects captured by the high-resolution camera, the object that is closest to the three-dimensional position of the object captured by the pixel of interest is found, and the resolution of the pixel of the high-resolution camera capturing that object can be set as the target resolution. This makes it possible to set the target resolution for each patch according to the resolution of the high-resolution camera, and if a difference in resolution remains, the difference can be further reduced. Embodiment 3
[0069] In the first and second embodiments, a method for setting the resolution magnification of a low-resolution camera was described. In the third embodiment, in order to reduce processing costs, the area in which super-resolution processing of the low-resolution camera is performed is set only in the overlapping area between the low-resolution camera and the high-resolution camera. A method for performing blending processing to reduce the resolution difference in the overlapping area of overhead images is well known. However, when blending processing is performed in the overlapping area of overhead images, even though the images were captured by the high-resolution camera, the image data from the low-resolution camera is blended, resulting in a lower resolution after synthesis than the resolution of the high-resolution camera. Therefore, in the third embodiment, this reduction in resolution is reduced by super-resolving the image data from the low-resolution camera in the overlapping area.
[0070] The hardware configuration of the image processing device in this embodiment is the same as that of the first and second embodiments, and therefore will not be described here. The following mainly describes the differences between this embodiment and the first or second embodiment. The same components will be denoted by the same reference numerals.
[0071] <Functional configuration of image processing device> 9 is a block diagram showing the functional configuration of an image processing device, which is made up of an input unit 11, a generation unit 12, a magnification setting unit 13, an output unit 14, and an area setting unit 21.
[0072] The input unit 11 outputs image data, imaging device information, and three-dimensional information to the generation unit 12, the magnification setting unit 13, and the region setting unit 21.
[0073] The region setting unit 21 sets a region for setting a high resolution magnification, and outputs region information to the generation unit 12.
[0074] <Processing of image processing device> The process executed by the image processing device will be described below with reference to the flowchart shown in FIG.
[0075] In S201, the region setting unit 21 sets region information. For example, the region information is a mask image in which the overlap region between the low-resolution camera and the high-resolution camera is 1. The overlap region information can be obtained, for example, by projecting each pixel of the image data of the low-resolution camera using a three-dimensional model and external parameters and determining whether it is within the range of the angle of view of the high-resolution camera. First, a pixel of interest in the image data of the low-resolution camera is determined, and the position (xilow, yilow, zilow) of the subject captured at the pixel of interest in the camera coordinates of the low-resolution camera is calculated using equation (8), where i is a value indicating the ith pixel in the low-resolution camera image data.
[0076]
number
[0077] Here, (xilow, yilow, zilow) are the position of the subject in the low-resolution camera coordinate system, and (uilow, vilow) are the coordinates of the pixel of interest in the image data of the low-resolution camera. Here, the focal length of the low-resolution camera is defined as flow, and the principal point position is defined as (Cxlow, Cylow). Also, dlow(uilow, vilow) represents the distance in the optical axis direction (z-axis direction) from the viewpoint of the low-resolution camera to the subject captured at the pixel of interest, as shown in the distance map.
[0078] Next, the position of the object captured at the pixel of interest, calculated from equation (8), in the low-resolution camera coordinate system is converted into the high-resolution camera coordinate system according to equation (9).
[0079]
number
[0080] In equation (8), (xihigh, yihigh, zihigh) represent the position of the object captured at the pixel of interest (uilow, vilow) in the high-resolution camera coordinate system. R represents the rotation matrix from the low-resolution camera coordinate system to the high-resolution camera coordinate system, and T represents the translation matrix from the low-resolution camera coordinate system to the high-resolution camera coordinate system. The rotation matrix R and the translation matrix T (external parameters) are acquired from the input unit 11.
[0081] Next, the position of the pixel of interest in the high-resolution camera coordinates obtained by equation (9) is converted into a position in the image coordinates of the high-resolution camera by equation (10).
[0082]
number
[0083] Here, the focal length of the high-resolution camera is fhigh, and the principal point position is (Cxhigh, Cyhigh). (uihigh, vihigh) are the coordinates obtained by converting the pixel of interest (uilow, vilow) of the low-resolution camera into the image coordinates of the high-resolution camera.
[0084] All pixels of the low-resolution camera are converted using equation (10). If (uihigh, vihigh) is within the range of the image data, it is determined to be an overlapping area. In other words, equation (11) holds.
[0085]
number
[0086] Here, (umax, vmax) indicate the total number of vertical and horizontal pixels of the image data, respectively. A mask image is obtained by applying the processes shown in equations (8) to (11) to all pixels of the low-resolution camera. Note that the overlapping area may be obtained by reading an area previously stored in the secondary storage device 104. After obtaining the mask image, the process proceeds to S202.
[0087] In S202, the generation unit 12 generates a set of patches from image data of the low-resolution camera. Raster scanning is performed within the area set by the area setting unit 21 where the mask image is at the position of 1, and a set of patches is generated.
[0088] In S106, the magnification setting unit 14 sets a high-resolution magnification for the set target patch. Fig. 11 is a flowchart for explaining the process of S106 in detail.
[0089] In S4001, the magnification setting unit 14 sets a target resolution for performing resolution interpolation. If super-resolution is applied only to the overlapping region, a resolution step will occur at the boundary between the region to which super-resolution is applied and the region to which super-resolution is not applied. FIG. 12 is a diagram showing an example of interpolation. FIG. 12 shows the relationship between position and resolution when the vicinity of the seam is cut perpendicular to the seam. To prevent this step, the target resolution is set to gradually match the resolution of the low-resolution camera, as shown by the thick line in FIG. 12. This is achieved by calculating a resolution that interpolates the resolution Starget_high of the high-resolution camera closest to the pixel of interest and the resolution Starget_low of the boundary between the low-resolution camera closest to the pixel of interest and the super-resolution region.
[0090] First, the three-dimensional position (x0, y0, z0) of the object captured at the pixel of interest is calculated using equation (5). Next, the resolution of the boundary between the low-resolution camera and the super-resolution region closest to the pixel of interest is calculated. This is done by obtaining the pixel at the boundary between 1 and 0 from the mask image, and for each of these pixels, calculating the three-dimensional position of the captured object using equation (5), and then obtaining the nearest pixel. Here, the distance between the three-dimensional position of the object captured at the nearest pixel and (x0, y0, z0) is defined as dlow. Then, the resolution Starget_low is calculated using equation (6).
[0091] Next, the resolution Starget_high of the high-resolution camera closest to the pixel of interest is calculated. This is calculated using the same process as in S3001. Here, the distance between the three-dimensional position of the subject captured by the nearest pixel and (x0, y0, z0) is defined as dhigh.
[0092] Then, interpolation is performed using the distance shown in equation (11) to set the target resolution.
[0093]
number
[0094] In this embodiment, the area setting unit 21 sets an overlap area between the image data from the low-resolution camera and the image data from the super-resolution camera, but it does not necessarily have to be an overlap area. For example, an area may be set that is several pixels wider or narrower than the overlap area. Also, an area several pixels wide near the seam may be set. By using object recognition technology, an important object existing near the seam between the cameras may be recognized, and only the range of that object may be set.
[0095] Although the example in which linear interpolation based on distance is performed is shown as the interpolation method, other interpolation methods may be used, such as interpolation according to a multidimensional function.
[0096] As described above, according to this embodiment, by setting the area where super-resolution is performed, it is possible to reduce the processing cost of the super-resolution process. [Explanation of symbols]
[0097] 11 Input section 12 Generation part 13 Magnification setting section 14 Output section 100 Image processing device 101 CPU 102 RAM 103 ROM 104 Secondary storage device 105 Input Interface 106 Output Interface 107 System Bus 108 External storage device
Claims
1. An image processing device that is equipped with at least two or more imaging devices, combines image data acquired from the imaging devices, and generates a composite image, a first imaging device that provides a relatively high resolution in a seam or overlapping area of the composite image; a second imaging device that has a relatively low resolution in the seam or overlap area; a generation means for generating a patch set from image data acquired from the second imaging device; a magnification setting means for setting a high resolution magnification for each patch in the generated patch group; 2. The image processing apparatus according to claim 1, wherein the resolution increasing magnification is set to a value corresponding to a target resolution set based on the first image capturing device.
2. 2. The image processing apparatus according to claim 1, wherein the target resolution is one resolution determined from image data of the first image capture device.
3. The image processing apparatus according to claim 1 , wherein the target resolution is set for each patch in the set of patches.
4. the target resolution is a resolution at which a difference between the resolution of each patch in the patch group and the resolution of the first image capture device is reduced in the seam or overlap region; 4. The image processing apparatus according to claim 3, wherein the resolution increase magnification is a magnification by which the resolution of each patch in the patch group approaches the target resolution.
5. 5. The image processing device according to claim 4, wherein the magnification setting means sets the target resolution of each patch based on the distance between the three-dimensional position of the subject captured by each patch of the patch group and the position of the first imaging device.
6. The image processing device described in claim 4, characterized in that the magnification setting means acquires, for each patch in the patch group, a nearest pixel in the image data of the first imaging device based on the three-dimensional position of the subject captured by the patch, and sets the high-resolution magnification for each patch in the patch group using the resolution of the nearest pixel as the target resolution.
7. 4. The image processing apparatus according to claim 2, wherein the target resolution is calculated based on spatial frequency components of the image data.
8. 4. The image processing apparatus according to claim 2, wherein the target resolution is calculated using at least one of information about the image capturing apparatus, the lens model, the focal length, the sensor size, and the number of pixels.
9. 9. The image processing device according to claim 1, further comprising: an area setting unit that sets an area in the image acquired from the second imaging device, in which the high-resolution magnification is to be set.
10. 10. The image processing apparatus according to claim 9, wherein the area set by the area setting means is the overlap area.
11. The image processing device described in claim 9 or 10, characterized in that the magnification setting means obtains a first resolution, which is the resolution of the nearest pixel in the image data of the first imaging device, based on the three-dimensional position of the subject captured by the patch, and obtains a second resolution, which is the resolution of the nearest pixel among the pixels at the boundary between the area and an area other than the area in the image data of the second imaging device, based on the three-dimensional position of the subject captured by the patch, and the target resolution in the area is set to a resolution that interpolates the first resolution and the second resolution.
12. The image processing device according to claim 1 , wherein super-resolution is performed on the image data of the second imaging device in accordance with the resolution-high magnification.
13. An image processing device that is equipped with at least two or more imaging devices, combines image data acquired from the imaging devices, and generates a composite image, a first imaging device that provides a relatively high resolution in a seam or overlapping area of the composite image; a second imaging device that has a relatively low resolution in the seam or overlap area; a generation means for generating a patch set from image data acquired from the second imaging device; a magnification setting means for setting a high resolution magnification for each patch in the generated patch group; The image processing method according to claim 1, wherein the resolution increasing magnification is set to a value corresponding to a target resolution set based on the first image pickup device.
14. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 13.
Citation Information
Patent Citations
Image generation device, image generation method
JP3286306B2