Method and device for restoring image obtained from array camera
Patent Information
- Application Number
- KR1020210025724
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-25
- Publication Date
- 2026-09-09
- Estimated Expiration
- 2041-02-25
Smart Images

Figure 112021023085702-PAT00039_ABST
Abstract
Description
Technology Field
[0001] The following embodiments relate to a technique for restoring images acquired through an array camera. Background Technology
[0003] Due to advancements in optical and image processing technologies, imaging devices are being utilized in a wide range of fields, including multimedia content, security, and recognition. These devices are mounted on mobile devices, cameras, vehicles, and computers to capture images, recognize objects, or acquire data for controlling the equipment. The volume of an imaging device can be determined by the size of the lens, the focal length of the lens, and the size of the sensor; to reduce volume, multi-lenses composed of small lens elements may be used. means of solving the problem
[0005] An image restoration method according to one embodiment may include: acquiring a plurality of images through each lens element included in an array camera; estimating global parameters for the images; generating first processed images by transforming the viewpoint of the images using the estimated global parameters; estimating local parameters for each pixel of the first processed images; generating second processed images by transforming the first processed images using the estimated local parameters; and synthesizing the second processed images to generate a composite image of a target viewpoint.
[0006] The step of estimating the global parameters may include the step of estimating the global parameters using a neural network model that takes the acquired images as input.
[0007] The step of estimating the global parameters may include the step of estimating the global parameters based on the depth value of the scene captured by the array camera.
[0008] The step of estimating the local parameters above may include the step of obtaining an offset value for a pixel position for each pixel of the first processed image using a neural network model that inputs the first processed image.
[0009] The step of generating the second processed images may include generating the second processed images by performing image transformation pixel by pixel of the first processed image based on the offset value.
[0010] An image restoration device according to one embodiment includes a processor; and a memory for storing instructions executed by the processor, wherein the processor receives images corresponding to a plurality of viewpoints, estimates global parameters for the images, generates first processed images by transforming the viewpoint of the images using the estimated global parameters, estimates local parameters for each pixel of the first processed images, generates second processed images by transforming the first processed images using the estimated local parameters, and generates a composite image of a target viewpoint by synthesizing the second processed images.
[0011] A mobile device according to one embodiment may include: an imaging device that acquires images corresponding to a plurality of viewpoints; and a processor that generates first processed images by estimating global parameters for the images and converting the viewpoint of the images using the estimated global parameters, generates second processed images by estimating local parameters for each pixel of the first processed images and converting the first processed images using the estimated local parameters, and generates a composite image of a target viewpoint by synthesizing the second processed images. Brief explanation of the drawing
[0013] FIG. 1 is a diagram illustrating the general process of image restoration according to one embodiment. FIG. 2 is a drawing illustrating a mobile device including an array camera according to one embodiment. FIG. 3 is a flowchart illustrating an image restoration method according to one embodiment. FIG. 4 is a diagram illustrating the process of generating an image of a target point in time according to one embodiment. FIG. 5 is a diagram illustrating the process of performing image warping according to one embodiment. FIG. 6 is a diagram illustrating the positional relationship of sensing elements included in an array camera according to one embodiment. FIG. 7 is a diagram illustrating the process of generating a composite image according to one embodiment. FIG. 8 is a diagram illustrating the process of generating a composite image according to another embodiment. FIG. 9 is a block diagram illustrating the configuration of an image restoration device according to one embodiment. FIG. 10 is a block diagram illustrating the configuration of an electronic device according to one embodiment. Specific details for implementing the invention
[0014] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, actual implementations are not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or substitutions included in the technical concept described by the embodiments.
[0015] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.
[0016] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or joined to that other component, or that there may be other components in between.
[0017] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0018] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.
[0019] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.
[0021] FIG. 1 is a diagram illustrating the general process of image restoration according to one embodiment.
[0022] Referring to FIG. 1, an image restoration device (e.g., the image restoration device (900) of FIG. 9) is a device that restores an image based on information sensed about a scene. The image restoration device may include an imaging device (110) corresponding to an array camera, and the imaging device (110) may include a lens array in which a plurality of lens elements are arranged, and an image sensor that senses light passing through each lens element of the lens array. The lens array may be a multi-aperture lens, such as a CEV (compound eye vision) lens, for example. The image sensor may include a sensing array (112) in which a plurality of sensing elements are arranged along a plane.
[0023] The quality of the image acquired by the imaging device (110) can be determined by the number of sensing elements included in the image sensor and the amount of light incident on the sensing elements. The resolution of the acquired image is determined by the number of sensing elements included in the sensing array (112), and the sensitivity of the image can be determined by the amount of light incident on the sensing elements. The amount of light incident on the sensing elements can be determined based on the size of the sensing elements, and as the size increases, the amount of incident light increases and the dynamic range of the sensing array (112) increases, making it possible to capture high-quality images. In addition, as the number of sensing elements included in the sensing array (112) increases, the imaging device (110) can capture high-resolution images, and as the size of the sensing elements increases, the imaging device (110) can operate advantageously for capturing high-sensitivity images in low light conditions. The sensing element is a device that senses light passing through a lens array, and may be an image sensing element composed of, for example, a CMOS (complementary metal oxide semiconductor), a CCD (Charge-Coupled Device), or a photodiode.
[0024] The volume of the imaging device (110) can be determined by the focal length of the lens elements included in the lens array. Since the lens elements and the sensing elements must be spaced apart by the focal length of the lens elements in order for the sensing elements of the sensing array (112) to collect light refracted by the lens elements, the volume of the imaging device (110) can be determined by the spacing between the lens elements and the sensing elements. The focal length of the lens elements is determined by the field of view of the imaging device (110) and the size of the lens elements. For example, when the field of view is fixed, the focal length increases in proportion to the size of the lens elements, and in order to capture an image within a constant field of view range, the size of the lens elements must increase as the size of the sensing array (112) increases. In order to increase the sensitivity of the image while maintaining the field of view and the resolution of the image, the size of each sensing element must be increased while maintaining the number of sensing elements included in the sensing array (112), so the size of the sensing array (112) must be increased. At this time, in order to maintain the field of view, as the size of the sensing array (112) increases, the size of the lens element increases, and the focal length of the lens element becomes longer, so the volume of the imaging device (110) increases.
[0025] For a smaller size imaging device (110), the lens array may include lens elements corresponding to multiple viewpoints. The lens elements may be arranged along the plane of the lens array. The sensing elements of the sensing array (112) may be divided into sensing regions corresponding to each of the lens elements. The plane of the lens array and the plane of the sensing array (112) may be arranged parallel to each other, spaced apart by the focal length of the lens elements included in the lens array. The lens array may be a micro-lens array (MLA). As the size of each lens element included in the lens array decreases, the focal length of the lens element decreases, and as the focal length decreases, the thickness of the imaging device (110) may decrease. Thus, a thin camera may be implemented through a lens array containing small-sized lens elements. In a thin camera structure using a micro-lens array, the image restoration device can restore high-quality and high-resolution images by more accurately aligning each image through the image restoration method described below.
[0026] Each lens element included in the lens array can cover a sensing area of the sensing array (112) corresponding to its lens size. For example, the sensing area (113) covered by the lens element (111) in the sensing array (112) can be determined according to the lens size of the corresponding lens element (111). The sensing area (113) corresponds to an area on the sensing array (112) that light within a certain field of view reaches after passing through the lens element (111). Each sensing element of the sensing array (112) can sense the intensity value of the light that has passed through the corresponding lens element as sensing information.
[0027] The imaging device (110) may be divided into a plurality of sensing units. Each sensing unit may be distinguished by a lens element unit constituting a lens array. For example, each sensing unit may include a lens element (111) and sensing elements of a sensing area (113) covered by the lens element (111). The position where each lens element is placed in the imaging device (110) may correspond to a viewpoint. A viewpoint may represent a point where a subject is observed and / or photographed. The imaging device (110) may acquire low-resolution images (or viewpoint images) (120) corresponding to each viewpoint of the lens elements based on light received through each lens element placed at any viewpoint.
[0028] A process is required to acquire low-resolution images (120) with different viewpoints through a lens array and to generate a high-resolution composite image (190) having a target viewpoint from the acquired low-resolution images (120). The image restoration device described below can restore a high-resolution composite image (190) by rearranging and combining the low-resolution images (120) captured at each lens element. In one embodiment, the image restoration device can restore a high-resolution composite image (190) based on a reference image (121) among the acquired low-resolution images (120). The reference image (121) can be captured by a lens element (115) and a corresponding sensing area (116) corresponding to the target viewpoint. In the embodiment of FIG. 1, it is assumed that the reference image (121) is the image located in the center among the plurality of images (120), but the scope of the embodiment is not limited thereto, and an image at a location other than the center (e.g., image (122)) may be used as the reference image. The image restoration device may obtain first processed images by performing image alignment (or image warping) on other low-resolution images based on the reference image (121), obtain second processed images by refining the offset for each pixel of the processed images, and then obtain a high-resolution composite image (190) by synthesizing the second processed images.The image restoration device can perform image alignment more accurately by using a projection matrix determined based on the depth value of the scene to perform image alignment between images (120) without the process of performing camera calibration using camera intrinsic parameters and camera extrinsic parameters of each lens element, and can obtain a higher quality composite image (190) by refining the local offset value of each image.
[0030] FIG. 2 is a drawing illustrating a mobile device including an array camera according to one embodiment.
[0031] Referring to FIG. 2, an array camera (220) can be placed on a mobile device (210), such as a smartphone, to capture multiple images. The array camera (220) may be mounted on various devices, such as a DSLR camera, a vehicle, a drone, a surveillance camera such as a CCTV, a webcam camera, a VR (virtual reality) camera, or an AR (augmented reality) camera, in addition to the mobile device (210). The array camera (220) may be implemented in a thin or curved structure and used as a camera for object recognition.
[0032] The array camera (220) includes a lens array in which a plurality of lens elements are arranged, and may be positioned on the front or rear of the mobile device (210). FIG. 2 illustrates an embodiment in which the array camera (220) is positioned on the rear. The lens elements may be arranged adjacent to each other and may be arranged on the same plane. The array camera (220) can acquire low-resolution images with different viewpoints through the lens elements, and an image restoration device included in the mobile device (210) can process the low-resolution images acquired through the array camera (220) to restore a high-resolution composite image.
[0034] FIG. 3 is a flowchart illustrating an image restoration method according to one embodiment. The image restoration method may be performed by an image restoration device described herein (e.g., the image restoration device (900) of FIG. 9). The image restoration device can restore an image acquired through an array camera to generate a high-resolution image.
[0035] Referring to FIG. 3, in step (310), the image restoration device can acquire multiple images through each lens element included in the array camera. The lens elements of the array camera may be spaced apart from each other at equal distances on the same plane. The images acquired through the lens elements may be viewpoint images corresponding to different viewpoints of the lens elements.
[0036] In step (320), the image restoration device can estimate global parameters for the images acquired in step (310). Global parameters are parameters applied to the acquired images and are parameters for performing image alignment (or image warping) that transforms the image viewpoint to a target viewpoint. Global parameters may include, for example, rotation parameters, translation parameters, and scale parameters. Rotation parameters represent the degree of rotation between one viewpoint and a target viewpoint, translation parameters represent the degree of translation between one viewpoint and a target viewpoint, and scale parameters represent the scale difference between one viewpoint and a target viewpoint.
[0037] Image alignment can be performed by applying a projection matrix to each image, and the matrix elements included in the projection matrix may correspond to global parameters. An image restoration device can estimate global parameters using a neural network model that takes the acquired images as input. The neural network model is a neural network trained to output global parameters based on input data. Here, the input data may include data in which the corresponding images are concatenated, or data in which feature maps extracted from the images are concatenated. The feature map may represent feature data and / or feature vectors extracted from images sensed through individual lens elements for scene capture. The image restoration device can estimate eight matrix elements included in the projection matrix, and in this case, can estimate global parameters based on the depth value of the scene captured by the array camera.
[0038] According to another embodiment, the image restoration device may acquire a plurality of images in step (310) and perform a process of converting each of the acquired images into high-resolution images. The image restoration device may estimate global parameters using a neural network model that takes the high-resolution images as input. Here, the neural network model may be a learned neural network that outputs global parameters for performing image alignment using data combined from high-resolution images or data combined from feature maps extracted from high-resolution images as input data.
[0039] In step (330), the image restoration device can generate first processed images by transforming the viewpoints of the images using the global parameters estimated in step (320). The image restoration device can warp the images into first processed images having a target viewpoint using the global parameters. The image restoration device can generate first processed images in which each viewpoint of the images is transformed to a target viewpoint by determining a projection matrix in which the global parameters derived from the neural network model are matrix parameters, and applying the determined projection matrix to each image. Some of the global parameters may be commonly used to transform the viewpoints of the images.
[0040] In step (340), the image restoration device can estimate local parameters for each pixel of the first processed images. Local parameters are parameters applied to each pixel of the first processed images and are parameters for correcting the disparity error of each pixel. Local parameters can be obtained through a learned neural network model that takes the first processed images as input. The image restoration device can obtain an offset value for the pixel position for each pixel of the first processed images as a local parameter using the said neural network model. Here, the offset value may include errors caused by misalignment, such as errors regarding the object's distance value or errors that occurred during the image alignment process.
[0041] In step (350), the image restoration device can generate second processed images by transforming the first processed images using local parameters estimated in step (340). The image restoration device can generate second processed images by performing image transformation on each pixel of the first processed image based on an offset value for a pixel position obtained for each pixel of the first processed image. The offset value corresponds to a parallax error of the pixel, and the image restoration device can generate second processed images by correcting the parallax error.
[0042] In step (360), the image restoration device can synthesize the second processed images to generate a synthesized image of a target viewpoint. The image restoration device can combine the pixels of the second processed images to generate a synthesized image with a higher resolution than the second processed images. The image restoration device can generate a synthesized image rearranged from the second processed images to a target viewpoint, which is a single viewpoint, by performing pixel shuffling on the second processed images. Pixel shuffling may include a process of synthesizing the second processed images by rearranging pixels representing the same and / or similar points in the second processed images of multiple viewpoints to be adjacent to each other. The synthesized image is an image in which the pixels of the second processed images are registered, and may have a resolution equal to or higher than the resolution of the second processed images.
[0044] FIG. 4 is a diagram illustrating the process of generating an image of a target point in time according to one embodiment.
[0045] Referring to FIG. 4, input images (410) are acquired through an imaging device such as an array camera. The input images (410) may correspond to multiple viewpoint images that are of lower resolution than the image (430) of the target viewpoint. Viewpoint images corresponding to the viewpoint of each camera may be acquired by each camera included in the array camera.
[0046] The process of acquiring the image (430) at the target viewpoint may include a global transformation process that performs image warping for viewpoint switching of the input image (410) and a local offset refinement process that corrects the offset of the pixel position value for each pixel of the individual image. An image restoration device (e.g., the image restoration device (900) of FIG. 9) can restore a high-resolution image (430) at the target viewpoint from low-resolution input images (410) by using an image processing model (420) that performs image warping for image alignment between input images (410) and performs offset refinement for each pixel.
[0047] The image restoration device uses a neural network model (422) for obtaining global parameters to obtain global parameters Calculate and the calculated global parameters Viewpoint images can be converted using [this]. The image restoration device estimates global parameters suitable for the structure of the imaging device using a neural network model (422) without a separate calibration process. The neural network model (422) may be a neural network trained to output global parameters from information of an input image (410). The neural network can perform image restoration, such as image registration, by mapping input data and output data that have a non-linear relationship to each other based on deep learning. Deep learning is a machine learning technique for solving image registration problems from big data sets, and can map input data and output data to each other through supervised or unsupervised learning. The neural network may include an input layer, a plurality of hidden layers, and an output layer. Data input through the input layer is propagated through the plurality of hidden layers, and output data can be output from the output layer. However, data may be directly input into a hidden layer instead of an input layer, or output data may be output from a hidden layer instead of an output layer. The neural network can be trained, for example, through a back propagation technique. The neural network model (422) may be a convolutional neural network (CNN) implemented, for example, by a combination of a convolution layer and a fully connected layer. The image restoration device can extract feature data by performing convolution filtering on the data input to the convolution layer. The feature data is data in which the features of the image are abstracted, and may, for example, represent the result of a convolution operation according to the kernel of the convolution layer.However, the structure of the neural network model (422) is not limited to this and can be implemented in various combinations.
[0048] In one embodiment, the image restoration device uses a neural network model (422) to transform a matrix for transforming a single viewpoint as in Equation 1 below. (424) can be obtained. An input image (410) is input to the neural network model (422), and a matrix such as Equation 1 is obtained from the neural network model (422). (424) constituting ~ 8 global parameters of can be obtained.
[0049]
[0050] Acquired global parameters ~ Some of these can be commonly used in image warping processes that transform the viewpoint of another input image (or viewpoint image). For example, , Assuming that two global parameters are used commonly, the image restoration device obtains for each viewpoint image , For each of them, a representative value such as the average value can be calculated, and the calculated representative value can be used to transform the viewpoints of the images at each viewpoint. Here, the average value is merely an example, and various types of values (e.g., maximum or minimum values, etc.) may also be used as representative values. The remaining global parameters ( For ), global parameters obtained for each viewpoint image can be used. By using global parameters commonly, the number of global parameters that need to be trained can be reduced.
[0051] procession In (424), parameter Z represents the depth value of the scene appearing in the input image (410). Assuming that the lens elements in the imaging device are located on the same plane as each other and the sensing elements are located on the same plane as each other, the depth value of the scene (or object) in the viewpoint images obtained from the imaging device can be considered to be the same in all viewpoint images. If the depth value between the viewpoint images is assumed to be the same, parameter Z can be assumed to be the same between the input images (410). Additionally, on the plane where the lens elements and the sensing elements are each placed, the lens elements and the sensing elements can be assumed to be placed at equal intervals in the x and y directions, respectively. By considering these placement characteristics of the lens elements and the sensing elements, the number of global parameters requiring learning can be reduced.
[0052] The image restoration device is a matrix based on global parameters Image warping can be performed by applying (424) to the input image (410) (425) to transform the viewpoint of the input image (410). Through image warping, first processed images can be obtained in which the viewpoint of each of the input images (410) is transformed to the same viewpoint as the target viewpoint.
[0053] After performing image warping, the image restoration device may calculate offset values for each local location within the first processed image using a neural network model (426). The offset values include errors (e.g., parallax errors) resulting from image warping. The neural network model (426) may be, for example, a neural network trained to calculate feature values extracted by passing through several convolution layers with the first processed image as input as offset values for each pixel location. Based on the offset values, the image restoration device [applies] local parameters to each pixel of the first processed image (428) can be estimated. The image restoration device can estimate local parameters on the first processed image. By applying (428) (429), a second processed image can be generated in which the offset value is corrected for each pixel of the first processed image. The offset value may include a value regarding the position offset in the x-axis direction and the position offset in the y-axis direction for each pixel position of the first processed image. The image restoration device can generate a second processed image corresponding to the image (430) at the target viewpoint by correcting the position of each pixel of the first processed image based on the offset value. The image restoration device can generate second processed images by performing the same process for other first processed images, and can generate a synthesized image of a single target viewpoint (or reference viewpoint) by synthesizing the second processed images. The image restoration device can generate a synthesized image with a higher resolution than the second processed images by combining the pixels of the second processed images through pixel shuffling, etc.
[0055] FIG. 5 is a diagram illustrating the process of performing image warping according to one embodiment.
[0056] The image restoration device can perform image warping using a trained neural network model without a camera calibration process to align low-resolution viewpoint images acquired through an array camera. Viewpoint transformation of the viewpoint images can be performed as follows through an image transformation model between two independent and different cameras.
[0057] Referring to FIG. 5, the relationship between 2D positions (p1, p2) within the viewpoint images from two different cameras for a single 3D point p0 is illustrated. In the first viewpoint image (510) captured by the first camera, the coordinates of position p1 of point p0 are (x c1 , y c1) and the coordinates of the position p2 of point p0 in the second viewpoint image (520) captured by the second camera are (x c2 , y c2 ) and the coordinates of 3D point p0 expressed relative to the coordinate system of the first camera are (X c1 , Y c1 , Z c1 ) and the coordinates of 3D point p0 expressed relative to the second camera's coordinate system are (X c2 , Y c2 , Z c2 If we say ), the relationship between each coordinate can be expressed as in the following mathematical formulas 2 to 4.
[0058]
[0059]
[0060]
[0061] Equation 2 represents the projection of 3D to 2D, and Equation 4 represents the projection of 2D to 3D. Equation 3 represents the application of 3D homography. Transformation between 3D points can be expressed as 3D homography represented by 16 independent parameters, and 3D homography is equal to the product of matrices based on camera intrinsic parameters and camera extrinsic parameters as shown in Equation 5 below.
[0062]
[0063] Here, R and t are camera extrinsic parameters representing rotation and translation, respectively. K represents the camera intrinsic parameter, and T represents the transpose.
[0064] Since the scene depth values in the viewpoint images are the same, Z c1 =Z c2Assuming =Z, as shown in the following mathematical equation 6, x c2 , y c2 The transformation expression for can be represented based on a 2D homography dependent on Z, as shown in the following mathematical equation 7.
[0065]
[0066]
[0067] Based on Equation 7, the number of global parameters required for image warping of viewpoint images is ~ It is the value obtained by multiplying the number of 8 by the number of cameras. According to Equation 7, for 3D points with the same depth value, image warping (image coordinate transformation) between two cameras can be performed by applying a matrix consisting of 8 independent global parameters to the coordinate values of the pixels in the viewpoint image.
[0069] FIG. 6 is a diagram illustrating the positional relationship of sensing elements included in an array camera according to one embodiment.
[0070] Assuming that the surface on which lens elements are placed and the surface on which sensing elements are placed in an array camera are each coplanar, it can be assumed that the depth values of the scenes (or objects) within the viewpoint images captured by the array camera are identical across the viewpoint images. Additionally, since the positions where lens elements and sensing elements are placed in the array camera are known in advance, the translation information between cameras is not independent of each other. Considering these constraints, image transformation can be expressed with a smaller number of global parameters. For example, FIG. 6 illustrates the positional relationship of sensing elements (or lens elements) (600) of a 5x5 array camera that are spaced at equal intervals by a horizontal distance of d and a vertical distance of d. The imaging plane on which the viewpoint images are captured exists on the same plane, and the distance between adjacent sensing elements differs by a horizontal distance of d and a vertical distance of d. If we assume that the position of the central sensing element (610) is the reference position (0, 0), then the position of the sensing element (612) is (-2d, -2d) and the position of the sensing element (614) can be defined as (2d, d).
[0071] Assuming that the interval is predetermined as d based on the reference sensing element (610) in the array camera and there is no movement in the z direction, the translation component occurs proportionally to the interval for each camera. In image conversion between two cameras, the coordinate translation component caused by translation occurs inversely proportional to the depth value and proportionally to the magnitude of movement in the x and y directions. When the index of each camera included in the array camera is expressed as (i, j), image warping can be expressed as shown in the following mathematical formula 8, taking into account the camera placement information.
[0072]
[0073] Here, , corresponds to image coordinate shift components caused by the camera moving by an interval of d in the x and y directions, respectively, and can be commonly used in image warping of different viewpoint images. , By using it commonly, the number of global parameters required to perform image warping for a 5x5 array camera is 8( ~ (Number of) x 25 (Number of individual cameras included in a 5x5 array camera) = 200, from which 6 ( Number of) x 25 (Number of individual cameras included in a 5x5 array camera) + 2 ( , The number of items) can be reduced to 152.
[0075] FIG. 7 is a diagram illustrating the process of generating a composite image according to one embodiment.
[0076] Referring to FIG. 7, the image restoration device can generate first processed images by converting the viewpoint to a target viewpoint through an image alignment process (720) of a low-resolution input image (710) (e.g., 25 viewpoint images (C1~C25) having height HX and width W), and generate second processed images by correcting the offset value for the pixel position for each of the first processed images. H is the number of pixels arranged along the height of the input image (710), and W is the number of pixels arranged along the width of the input image (710), and each may be a natural number greater than or equal to 1.
[0077] The image restoration device can generate a high-resolution composite image (740) (e.g., a composite image having a height of 5H X a width of 5W) by performing a high-resolution image processing process (730) that restores a high-resolution image by merging the second processed images. The image restoration device can perform a high-resolution image processing process (730) that includes pixel concatenation and pixel shuffling processes on the second processed images.
[0078] FIG. 8 is a diagram illustrating the process of generating a composite image according to another embodiment.
[0079] According to another embodiment, the image restoration device may first convert each of the low-resolution input images (810) (e.g., 25 viewpoint images (C1–C25) having height HX width W) into high-resolution images (820). Subsequently, the image restoration device may perform image warping in an image alignment process (830) to convert each of the high-resolution images into images having a target viewpoint. By performing image warping on the high-resolution images, the image restoration device may have a higher accuracy of image warping than in the embodiment of FIG. 7. The image restoration device may generate a high-resolution composite image (850) (e.g., a composite image having height 5H X width 5W) by performing an image synthesis process (840) that integrates the images having the target viewpoint through the concatenation of pixels.
[0081] FIG. 9 is a block diagram illustrating the configuration of an image restoration device according to one embodiment.
[0082] Referring to FIG. 9, the image restoration device (900) may include an imaging device (910), a processor (920), and a memory (930). According to another embodiment, the imaging device (910) may be located outside the image restoration device (900), or the image restoration device (900) may be implemented by being integrated with the imaging device (910).
[0083] The imaging device (910) can acquire images corresponding to multiple viewpoints. The imaging device (910) can correspond to an array camera that acquires multiple images through a multi-lens array including lens elements placed at different locations. The imaging device (910) can capture a multi-lens image including multiple viewpoint images corresponding to multiple viewpoints, and the processor (920) can generate input data from the multi-lens image.
[0084] The memory (930) can temporarily or permanently store data required for the execution of the image restoration method. For example, the memory (930) can store instructions executed by the processor (920), images acquired by the imaging device (910), various parameters (such as global parameters and local parameters), a neural network model for estimating parameters for image restoration, and synthetic images.
[0085] The processor (920) controls the overall operation of the image restoration device (900) and can execute functions and instructions to be executed within the image restoration device (900). The processor (920) receives images corresponding to multiple viewpoints from the imaging device (910) and can estimate global parameters for the images using a learned neural network model that takes the acquired images as input. The processor (920) can generate first processed images by defining a projection matrix based on the estimated global parameters and applying the projection matrix to each image to transform the viewpoint of the images. The processor (920) can obtain an offset value for a pixel position for each pixel of the first processed image as a local parameter using a learned neural network model that takes the first processed image as input. The processor (920) can generate second processed images by correcting the pixel-by-pixel offset value of the first processed image and can generate a composite image of the target viewpoint by synthesizing the second processed images. The processor (920) can generate a composite image with a higher resolution than the second processed images by combining the pixels of the second processed images by performing pixel shuffling. For reference, the operation of the processor (920) is not limited thereto, and the processor (920) may perform one or more of the operations described above in FIGS. 1 to 8 simultaneously or sequentially.
[0087] FIG. 10 is a block diagram illustrating the configuration of an electronic device according to one embodiment.
[0088] Referring to FIG. 10, the electronic device (1000) is a device that generates a high-resolution synthetic image by performing the image restoration method described above, and can perform the function of the image restoration device (900) described in FIG. 9. The electronic device (1000) may be a mobile device such as, for example, an image processing device, a smartphone, a wearable device, a tablet computer, a netbook, a PDA (personal digital assistant), an HMD (head mounted display), and a camera device. The electronic device (1000) may also be implemented as a vision camera device for vehicles, drones, and CCTVs, a webcam camera device for video calls, a 360-degree shooting camera device, a VR camera device, or an AR camera device.
[0089] The electronic device (1000) may include a processor (1010), memory, an imaging device (1030), a storage device (1040), an input device (1050), an output device (1060), and a communication device (1070). Each component of the electronic device (1000) may communicate with each other via a communication bus (1080).
[0090] The processor (1010) controls the overall operation of the electronic device (1000) and executes functions and instructions to be executed within the electronic device (1000). The processor (1010) can perform one or more of the operations described above through FIGS. 1 to 9.
[0091] The memory (1020) stores information necessary for the processor (1010) to perform an image restoration method. For example, the memory (1020) may store instructions to be executed by the processor (1010) and may store relevant information while software or a program is running on the electronic device (1000). The memory (1020) may include RAM, DRAM, SRAM, or other forms of non-volatile memory known in the art.
[0092] The imaging device (1030) includes an array camera and can acquire images corresponding to each of a plurality of lens elements. The electronic device (1000) can generate a high-resolution composite image by performing an image restoration process based on the acquired images.
[0093] The storage device (1040) includes a computer-readable storage medium or a computer-readable storage device (1040) and can store raw images and enhanced images. For example, the storage device (1040) may include storage, a magnetic hard disk, an optical disk, a flash memory, an electrically programmable memory (EPROM), etc.
[0094] The input device (1050) can receive input from a user through tactile, video, audio, or touch input. For example, the input device (1050) may include a keyboard, mouse, touchscreen, microphone, or any other device capable of detecting input from a user and transmitting the detected input to an electronic device (1000).
[0095] The output device (1060) can provide the output of the electronic device (1000) to the user through a visual, auditory, or tactile channel. The output device (1060) may include a display, a touch screen, a speaker, a vibration generating device, or any other device capable of providing the output to the user. The communication device (1070) can communicate with an external device through a wired network or a wireless network.
[0097] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0098] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.
[0099] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program instructions, data files, data structures, etc., either individually or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0100] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0101] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0102] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 An image restoration method for restoring an image obtained through an array camera, comprising: a step of obtaining a plurality of images through each lens element included in the array camera; a step of estimating a global parameter for the images; a step of generating first processed images by transforming the viewpoint of the images using the estimated global parameter; a step of estimating a local parameter for each pixel of the first processed images; a step of generating second processed images by transforming the first processed images using the estimated local parameter; and a step of synthesizing the second processed images to generate a synthesized image of a target viewpoint, wherein the step of estimating the global parameter includes a step of estimating the global parameter based on the depth value of the scene captured by the array camera. Claim 2 An image restoration method according to claim 1, wherein the step of estimating the global parameters comprises the step of estimating the global parameters using a neural network model that takes the acquired images as input. Claim 3 In paragraph 2, the step of estimating the global parameters comprises the step of estimating matrix elements included in the projection matrix, in an image restoration method. Claim 4 delete Claim 5 An image restoration method according to claim 1, wherein the step of generating a composite image at the target point in time comprises the step of combining pixels of the second processed images to generate the composite image having a higher resolution than the second processed images. Claim 6 In claim 5, the step of generating a synthetic image at the target point in time comprises the step of generating the synthetic image from the second processed images using pixel shuffling. Claim 7 An image restoration method according to claim 1, further comprising the step of converting each of the acquired images into high-resolution images, and the step of estimating the global parameters comprising the step of estimating the global parameters using a neural network model that takes the high-resolution images as input. Claim 8 An image restoration method according to claim 1, wherein the step of estimating the local parameters comprises the step of obtaining an offset value for a pixel position for each pixel of the first processed image using a neural network model that inputs the first processed image. Claim 9 In claim 8, the step of generating the second processed images comprises generating the second processed images by performing image transformation pixel by pixel of the first processed image based on the offset value. Claim 10 An image restoration method according to claim 1, wherein the step of generating the first processed images comprises the step of warping the images into the first processed images having the target viewpoint using the global parameters. Claim 11 An image restoration method according to claim 1, wherein the lens elements of the array camera are spaced apart from each other at equal distances on the same plane. Claim 12 An image restoration method according to claim 1, wherein the images obtained through the lens elements are viewpoint images corresponding to different viewpoints. Claim 13 A computer-readable recording medium storing one or more computer programs comprising instructions for performing any one of the methods of paragraphs 1 through 3 and paragraphs 5 through 12. Claim 14 An image restoration device comprising: a processor; and a memory for storing instructions executed by the processor, wherein the processor receives images corresponding to a plurality of viewpoints, estimates global parameters for the images, generates first processed images by transforming the viewpoint of the images using the estimated global parameters, estimates local parameters for each pixel of the first processed images, generates second processed images by transforming the first processed images using the estimated local parameters, and generates a composite image of a target viewpoint by synthesizing the second processed images, and wherein the processor estimates the global parameters based on the depth values of the images. Claim 15 In claim 14, the image restoration device, wherein the processor estimates the global parameters using a neural network model that takes the received images as input. Claim 16 In claim 14, the above processor is an image restoration device that obtains an offset value for a pixel position for each pixel of the first processed image using a neural network model that inputs the first processed image. Claim 17 In claim 14, the processor is an image restoration device that combines pixels of the second processed images to generate the composite image, which has a higher resolution than the second processed images. Claim 18 In claim 14, the image restoration device, wherein the processor converts each of the received images into high-resolution images and estimates the global parameters using a neural network model that takes the high-resolution images as input. Claim 19 A mobile device comprising: an imaging device for acquiring images corresponding to multiple viewpoints; and a processor for estimating global parameters for the images and generating first processed images by transforming the viewpoint of the images using the estimated global parameters, estimating local parameters for each pixel of the first processed images and generating second processed images by transforming the first processed images using the estimated local parameters, and synthesizing the second processed images to generate a composite image of a target viewpoint, wherein the processor estimates the global parameters based on depth values of the images.
Citation Information
Patent Citations
Image-processing device and image-processing method
WO2014132754A1