Image restoration apparatus and method
The video restoration method addresses the challenge of aligning and combining multi-lens videos by generating warping video information for each disparity and using a video restoration model, resulting in a high-resolution output video with improved quality and resolution.
Patent Information
- Application Number
- JP2020155697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-30
- Filing Date
- 2020-09-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-09-16
AI Technical Summary
Existing technologies face challenges in restoring a high-quality single video from a multi-lens video, particularly in aligning and combining input videos captured by multiple lenses with different positions and disparities.
A video restoration method that involves acquiring multiple input video information, generating warping video information for each disparity, and using a video restoration model to combine and process the input and warping video information, resulting in a high-resolution output video.
The method effectively generates a high-resolution output video by aligning and combining input videos from multiple lenses, improving video quality and resolution without requiring accurate depth detection for each pixel.
Smart Images

Figure 0007687790000038 
Figure 0007687790000039 
Figure 0007687790000040
Abstract
Description
Technical Field
[0001] Provided are a method and an apparatus for restoring a multi-lens video, which is a technology for restoring a multi-lens video, and restoring a video based on a plurality of input videos captured by a plurality of image sensors or captured by an image sensor including a multi-lens array.
Background Art
[0002] With the development of optical technology and video processing technology, imaging devices are being utilized in a wide range of fields such as multimedia content, security, and recognition. For example, imaging devices are mounted on mobile devices, cameras, vehicles, computers, etc., and can capture videos, recognize objects, and acquire data for controlling devices. The volume of an imaging device is determined by factors such as the size of the lens, the focal length of the lens, and the size of the sensor, and multi-lenses composed of small lenses are used to reduce the volume.
Summary of the Invention
Problems to be Solved by the Invention
[0003] A video restoration apparatus according to an embodiment is to restore a single video from a multi-lens video.
Means for Solving the Problems
[0004] A video restoration method according to an embodiment includes steps of acquiring a plurality of input video information, generating a plurality of warping video information for each of a plurality of disparities based on the plurality of input video information, and generating an output video by using a video restoration model based on the plurality of input video information and the plurality of warping video information.
[0005] The plurality of input video information may include a plurality of input videos captured through lenses arranged at different positions from each other.
[0006] The step of generating the plurality of warping video information includes generating a plurality of warping videos for each of the plurality of disparities as the warping video information by warping each of the plurality of input videos to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities.
[0007] The step of generating the warping video can include generating a warping video by warping all pixels in a first input video among the plurality of input videos to a pixel coordinate system corresponding to the target video using a single depth corresponding to a first disparity among the plurality of disparities.
[0008] Disparity is set for the input video with reference to the target video, and the depth corresponding to the disparity can be based on the disparity and the distance between the detection units that captured the target video and the input video.
[0009] The step of generating the output video can include generating the output video by providing, as an input, data obtained by combining the plurality of input videos and the plurality of warping videos corresponding to each of the plurality of disparities to the video restoration model.
[0010] The plurality of input video information can include a plurality of input feature maps extracted from the plurality of input videos using a feature extraction model.
[0011] The step of generating the plurality of warping video information can include generating a plurality of warping feature maps for each of the plurality of disparities as the plurality of warping video information by warping each of the plurality of input feature maps to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities.
[0012] The step of generating the output video can include the step of generating the output video by providing, as an input, data obtained by combining the plurality of input feature maps and the plurality of warping feature maps corresponding to the respective ones of the plurality of disparities to the video restoration model.
[0013] The video restoration model can be a neural network including at least one convolutional layer configured to apply convolutional filtering to input data.
[0014] The plurality of disparities are equal to or less than a maximum disparity and equal to or greater than a minimum disparity, the maximum disparity being based on the focal length of the detection units, the distance between the detection units, and the minimum shooting distance of the detection units, and the detection units can be configured to capture input images corresponding to the plurality of input image information.
[0015] The plurality of disparities can be a finite number.
[0016] The step of generating the output video can include the step of generating the output video without depth detection to a target point corresponding to an individual pixel.
[0017] The step of generating the plurality of warping video information includes the step of generating warping video information by applying a coordinate mapping function to an input video corresponding to the input video information, and the coordinate mapping function can be predetermined with respect to a detection unit configured to capture the input image and a target detection unit configured to capture a target image.
[0018] The resolution of the output video may be higher than the resolution of each of the plurality of input video information.
[0019] The plurality of input video information includes multi-lens video captured by an image sensor including a multi-lens array, and the multi-lens video can include a plurality of input videos.
[0020] The plurality of input video information can include a plurality of input videos individually captured by a plurality of image sensors.
[0021] A video restoration device according to an embodiment includes an image sensor that acquires a plurality of input video information, and based on each of the plurality of input video information, generates a plurality of warping video information for each of a plurality of disparities, and generates an output video using a video restoration model based on the plurality of input video information and the plurality of warping video information.
[0022] The video restoration device can include a lens array including a plurality of lenses, and a plurality of detection elements that detect light that has passed through the lens array. The plurality of detection elements include detection regions individually corresponding to the plurality of lenses, a detection array configured to acquire a plurality of input information, and a processor that generates a plurality of warping information for each of a plurality of disparities based on each of the plurality of input information, and generates an output video using a video restoration model based on the plurality of input information and the plurality of warping information.
[0023] The resolution of the output video may be higher than the resolution corresponding to the plurality of input information.
[0024] The processor can generate the plurality of warping information for each of the plurality of disparities by warping each of the plurality of input information to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities.
[0025] The processor can generate warping information by warping all pixels corresponding to the input information among the plurality of input information to pixel coordinate systems corresponding to the target video using a single depth corresponding to a first disparity among the plurality of disparities.
[0026] The processor can generate the output video by providing, as an input to the video restoration model, data obtained by combining the warping information corresponding to each of the plurality of input information and the plurality of disparities.
[0027] The processor can extract a plurality of input feature maps as the plurality of input information from a plurality of input videos using a feature extraction model.
[0028] The processor can generate, as the plurality of warping information, a plurality of warping feature maps for each of the plurality of disparities by warping each of the plurality of input feature maps to a pixel coordinate system corresponding to the target video using a depth corresponding to each of the plurality of disparities.
[0029] The processor can generate the output video by providing, as an input to the video restoration model, data obtained by combining the plurality of input feature maps and the warping feature maps corresponding to each of the plurality of disparities.
Advantages of the Invention
[0030] The video restoration device according to one embodiment can generate a warping video for each disparity and restore a high-resolution video from the warping video using a neural network.
Brief Description of the Drawings
[0031]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0032] The embodiments described below can be variously modified. The scope of the patent application is not limited or restricted by such embodiments. The same reference numerals presented in each drawing indicate the same members.
[0033] The specific structural or functional descriptions disclosed in this specification are merely exemplified for the purpose of explaining the embodiments, and the embodiments can be implemented in various different forms, and the present invention is not limited to the embodiments described in this specification.
[0034] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "comprising" or "having" indicate the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0035] Unless otherwise defined, all terms used herein, including technical and scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Commonly used pre-defined terms should be interpreted as having a meaning consistent with their meaning in the context of the related art, and should not be interpreted in an idealized or overly formal sense unless clearly defined herein.
[0036] Also, when describing with reference to the drawings, the same components are given the same reference numerals regardless of the reference signs in the drawings, and redundant descriptions thereof are omitted. When it is determined that a specific description of a known technology related to the description of an embodiment unnecessarily obscures the gist of the present invention, the detailed description thereof is omitted.
[0037] FIG. 1 illustrates a schematic process of video restoration according to an embodiment.
[0038] The quality of the image captured and restored by the image sensor 110 according to an embodiment is determined according to the number of detection elements included in the detection array 112 and the amount of light incident on the detection elements. For example, the resolution of the image is determined according to the number of detection elements included in the detection array 112, and the sensitivity of the image is determined according to the amount of light incident on the detection elements. The amount of light incident on the detection elements is determined based on the size of the detection elements. The larger the size, the more light is incident, and the dynamic range of the detection array 112 increases. Therefore, by increasing the number of detection elements included in the detection array 112, the image sensor 110 can capture a high-resolution image, and by increasing the size of the detection elements, the image sensor 110 can advantageously capture a high-sensitivity image at low illuminance.
[0039] The volume of the image sensor 110 is determined by the focal length of the lens element 111. More specifically, the volume of the image sensor 110 is determined by the distance between the lens element 111 and the detection array 112. However, in order to collect the light refracted by the lens element 111, the lens element 111 and the detection array 112 must be arranged at a distance equal to the focal length of the lens element 111.
[0040] The focal length of the lens element 111 is determined by the viewing angle of the image sensor 110 and the size of the lens element 111. For example, when the viewing angle is fixed, the focal length increases in proportion to the size of the lens element 111. In order to capture an image within a certain viewing angle range, the size of the lens element 111 must increase as the size of the detection array 112 increases. As described above, in order to increase the sensitivity of an image while maintaining the viewing angle and the resolution of the image, the volume of the image sensor 110 is increased. For example, in order to increase the sensitivity of an image while maintaining the resolution of the image, the size of each detection element included in the detection array 112 must be increased while maintaining the number of detection elements included in the detection array 112, so the size of the detection array 112 is increased. Here, in order to maintain the viewing angle, as the size of the detection array 112 increases, the size of the lens element 111 increases, and the focal length of the lens element 111 becomes longer, so the volume of the image sensor 110 is increased.
[0041] Referring to FIG. 1, the image sensor 110 includes a lens array and a detection array 112. The lens array includes lens elements, and the detection array 112 includes detection elements. The lens elements are arranged along the plane of the lens array, and the detection elements are arranged along the plane of the detection array 112. The detection elements of the detection array 112 are divided into detection regions corresponding to each of the lens elements. The plane of the lens array is parallel to the plane of the detection array 112 and is separated by the focal length of the lens element 111 included in the lens array. The lens array can be shown as a micro multi lens array (MMLA) or a multi lens array.
[0042] According to one embodiment, the smaller the size of each lens element included in the lens array, in other words, the larger the number of lenses included in the same area on the lens array, the smaller the focal length of the lens element 111 and the thinner the image sensor 110. Therefore, a thin camera is realized. In this case, the image sensor 110 can restore a high-resolution output video 190 by rearranging and combining the low-resolution input videos 120 captured by each lens element 111.
[0043] Each individual lens element 111 of the lens array can cover a certain detection area 113 of the detection array 112 corresponding to the size of its own lens. The detection area 113 covered by the lens element 111 in the detection array 112 is determined according to the size of the lens of the corresponding lens element 111. The detection area 113 indicates the area on the detection array 112 where light rays within a certain field of view angle range reach after passing through the corresponding lens element 111. The size of the detection area 113 can be expressed as the distance from the center of the detection area 113 to the outermost point or the length of the diagonal, and the size of the lens corresponds to the diameter of the lens.
[0044] Each detection element of the detection array 112 generates detection information based on the light rays that have passed through the lenses of the lens array. For example, the detection element can detect the intensity value of the light received through the lens element 111 as detection information. The image sensor 110 can determine the intensity information corresponding to the original signal regarding the points included in the field of view of the image sensor 110 based on the detection information output by the detection array 112, and can restore the captured image based on the determined intensity information. For example, the detection array 112 may be an image detection module composed of CMOS (complementary metal oxide semiconductor) or CCD (Charge-Coupled Device), etc.
[0045] In addition, the detection element includes a color filter for detecting a desired color and generates the intensity value of the color corresponding to the specific color as detection information. Each of the plurality of detection elements constituting the detection array 112 may be arranged to detect a color different from that of the adjacent detection elements adjacent in space.
[0046] When the diversity of the detection information is sufficiently ensured and a full rank relationship is formed between the original signal information corresponding to the points included in the field of view of the image sensor 110 and the detection information, a captured image corresponding to the maximum resolution of the detection array 112 is derived. The diversity of the detection information is ensured based on the parameters of the image sensor 110, such as the number of lenses included in the lens array and the number of detection elements included in the detection array 112.
[0047] Furthermore, the detection region 113 covered by the individual lens element 111 includes non-integer detection elements. The multi-lens array structure according to one embodiment is realized as a fractional alignment structure. For example, when the lens elements included in the lens array have the same lens size, the number of lens elements included in the lens array and the number of detection elements included in the detection array 112 may have a relatively prime relationship with each other. The ratio P / L between the number L of lens elements of the lens array and the number P of detection elements corresponding to one axis of the detection array 112 is determined as a real number. Each of the lens elements can cover the same number of detection elements as the pixel offset corresponding to P / L.
[0048] With the fractional alignment structure as described above, the image sensor 110 may have the optical center axes (OCA, optical center axis) of the respective lens elements 111 arranged slightly differently from each other with respect to the detection array 112. In other words, the lens elements 111 may be arranged eccentrically with respect to the detection elements. Accordingly, each lens element 111 of the lens array receives different light field information from each other. For reference, a light field (LF, light field) may be emitted from any target point and represents a field indicating the direction and intensity of light rays reflected from any point on a subject. Light field information represents information in which a plurality of light fields are combined. Since the direction of the chief ray of each lens element 111 can also vary, each detection region 113 receives different light field information from each other, so that a plurality of slightly different input information (for example, input video information) may be obtained in a plurality of detection regions. Through the plurality of slightly different input information, the image sensor 110 can optically acquire more detection information.
[0049] The above-described image sensor 110 may be divided into a plurality of sensing units. Each of the plurality of sensing units is distinguished into lens units that constitute a multi-lens array. For example, each sensing unit includes a lens and detection elements in a detection area 113 covered by the corresponding lens. According to one embodiment, the image sensor 110 generates individual input videos from the detection information acquired for each detection area 113 corresponding to each lens. In other words, each of the plurality of sensing units can individually acquire an input video. As described above, since the plurality of sensing units acquire different light field information, the input videos captured by each sensing unit may capture slightly different scenes. The image sensor 110 may include N lenses and may be distinguished into N sensing units. Since the N sensing units individually capture input videos, the image sensor 110 can acquire N input videos 120. Here, N may be an integer of 2 or more. In FIG. 1, the multi-lens array may include N = 5 × 5 = 25 lenses, and the image sensor captures 25 low-resolution input videos 120. As a different example, the multi-lens video may be composed of N = 6 × 6 = 36 input videos. For reference, an example in which one image sensor 110 includes a plurality of sensing units has been described, but the sensing units are not limited thereto. The sensing unit may indicate an independent image detection module (for example, a camera sensor), and in this case, each sensing unit may be arranged at a position different from other sensing units.
[0050] Hereinafter, the image sensor 110 generates a plurality of low-resolution input videos 120 from various detection information obtained as described above, and among the plurality of low-resolution input videos 120, a high-resolution output video 190 can be restored based on the target video 121. For reference, in FIG. 1, the target video 121 is determined to be the central video among the plurality of input videos 120, but it is not limited thereto, and other input videos that are not in the center may be used as the target video. Further, the image sensor 110 can also use the video of another separate additional image sensor as the target video. The additional image sensor may be a camera sensor capable of capturing a higher-resolution video than the image sensor 110.
[0051] FIG. 2 is a flowchart for explaining a video restoration method according to an embodiment. FIG. 3 illustrates video restoration using a video restoration model according to an embodiment.
[0052] First, in step S210, the video restoration device acquires a plurality of input video information. The input video information may be the input video itself, but is not limited thereto, and may be an input feature map extracted from the input video using a feature extraction model. For reference, an example where the input video information is the input video itself will be described with reference to FIGS. 4 to 8 below, and an example where it is an input feature map will be described with reference to FIGS. 9 and 10 below.
[0053] According to an embodiment, the video restoration device captures a plurality of input videos via the image sensor 310 shown in FIG. 3. For example, an image sensor 310 including a multi-lens array in the video restoration device captures a multi-lens video 320 including a plurality of input videos. In the multi-lens video 320, each input video is captured by an individual detection unit constituting the image sensor 310. The first input video to the Nth input video are respectively individually the first detection unit C 1 ~the Nth detection unit C NIt is captured by [the relevant means]. As a different example, in the video restoration device, each of the plurality of image sensors 310 may capture the input video. Here, each of the detection units may be an independent image sensor 310.
[0054] And in step S220, the video restoration device generates a plurality of warping information (for example, warped video information 330) corresponding to a plurality of disparities from each of the plurality of input information (for example, input video information). Disparity is the difference in position for the same target point in any two videos, and for example, it indicates the difference in pixel coordinates. According to one embodiment, for each input video, the disparity with the target video is set to an arbitrary value, and the virtual distance from the image sensor 310 to the target point is determined by the set disparity. The video restoration device generates the warping video information 330 using the distance determined by the set disparity. The warping video information 330 is the warped video itself in which the input video is converted to the pixel coordinate system of the target video, but is not limited thereto, and may be a warped feature map in which the input feature map extracted from the input video is converted to the pixel coordinate system of the target detection unit that captured the target video. The virtual depth determined by the above-described disparity and the warping using the virtual depth will be described with reference to FIG. 4 below. For reference, in this specification, the depth value indicates the distance to the target point.
[0055] For example, as shown in FIG. 3, for each input video included in the multi-lens video 320, the video restoration device has a warping video information 330 corresponding to the minimum disparity d min to the warping video information 330 corresponding to the maximum disparity d max is generated based on the camera calibration parameter 319. The minimum disparity d minIf it is 0, the warping video is the input video itself. The camera calibration parameter 319 will be described with reference to FIG. 7 below. When the number of disparities is D, the video restoration device generates D pieces of warping video information 330 for each of the N input videos, so a total of N×D pieces of warping video information 330 are generated. Here, D may be an integer of 1 or more.
[0056] Next, in step S230, the video restoration device generates an output video 390 using a video restoration model 340 based on a plurality of input video information and a plurality of warping video information 330. According to one embodiment, the video restoration model 340 may be a model trained to output an output video 390 from the input video information. The video restoration model 340 may be, for example, a neural network as a machine learning structure. The neural network maps input data and output data that are in a non-linear relationship to each other based on deep learning to perform video restoration by video registration (image registration), etc. Deep learning is a machine learning method for solving the video registration problem from a large dataset, and maps input data and output data to each other through supervised or unsupervised learning. The neural network includes an input layer 341, a plurality of hidden layers 342, and an output layer 343. The data input through the input layer 341 propagates through the plurality of hidden layers 342 and is output from the output layer 343. However, instead of the input layer 341 and the output layer 343, data may be directly input to the hidden layer 342 or directly output from the hidden layer 342. The neural network may be trained, for example, through back propagation.
[0057] The above-described video restoration model 340 can be implemented in a convolutional neural network. A convolutional neural network is a neural network including convolutional layers, and the hidden layer 342 of the convolutional neural network includes convolutional layers. For example, a convolutional neural network includes a convolutional layer in which nodes are connected via kernels. The convolutional neural network may be a network pre-trained to output a high-resolution output video from a plurality of input video information and a plurality of warping video information based on training data. The output video may be, for example, a video in which pixels matching the target video in the input video and the warping video are registered (aligned), and the resolution of the output video is higher than the resolution corresponding to the plurality of input information (e.g., the input video). For reference, the video restoration device can extract feature data by performing convolutional filtering on the data input to the convolutional layer. The feature data can be data in which the features of the video are abstracted, and can indicate, for example, the result value of the convolutional operation by the kernel of the convolutional layer. The video restoration device may perform a convolutional operation using the element values of the kernel on the pixels at any position and the surrounding pixels in the video. The video restoration device calculates the convolutional operation value for each pixel while sweeping the kernel over the pixels of the video. An exemplary implementation of the convolutional neural network of the video restoration model 340 will be described in detail with reference to FIG. 8 below.
[0058] For example, the video restoration device can provide the N input video information obtained in step S210 and the N×D warping video information 330 generated in step S220 to the video restoration model 340. As described above, the video restoration model 340 can include a convolutional layer that applies convolutional filtering to the input data. Therefore, the video restoration device can apply convolutional filtering to the N input video information and the N×D warping video information 330 using the video restoration model 340, and as a result, generate a high-resolution output video 390.
[0059] FIG. 4 illustrates the generation of a warping video for input to a video restoration model according to an embodiment.
[0060] A video restoration apparatus according to an embodiment generates a plurality of warping information (e.g., warping videos) by warping each of a plurality of input information (e.g., input videos) to pixel coordinate systems corresponding to a target video 430 (target image) using depths corresponding to respective disparities of the plurality of disparities. For example, FIG. 4 illustrates a warping video in which the i-th input video 420 among N input videos is warped to a pixel coordinate system corresponding to the target video 430.
[0061] For reference, in this specification, the world coordinate system can represent a coordinate system based on an arbitrary point on the world as a three-dimensional coordinate system. The camera coordinate system can be represented as a three-dimensional coordinate system based on a camera. For example, the principal point of the detection unit is used as the origin, the optical axis direction of the detection unit is the z-axis, the vertical direction of the detection unit is the y-axis, and the horizontal direction of the detection unit is the x-axis. The pixel coordinate system can also be referred to as an image coordinate system and can represent the two-dimensional coordinates of pixels within a video.
[0062] For example, assume that the world coordinates of an arbitrary target point 490 separated from the image sensor are X, Y, and Z. Assume that the pixel coordinates detected by the i-th detection unit 411C among N detection units for the target point 490 are u and v. Assume that the pixel coordinates detected by the target detection unit 412C i are u' and v'. However, it is difficult to accurately determine the distance to the target point 490 only from the pixel values detected by each detection unit. A video restoration apparatus according to an embodiment can warp an input video to the pixel coordinate system of the target video 430 using a distance value corresponding to a hypothesized disparity, assuming that the input video has an arbitrary disparity with respect to the target video 430. T
[0063] First, the video restoration device normalizes the pixel coordinates for the individual pixels of the i-th input video 420 TIFF0007687790000001.tif13134 to calculate the normalized coordinates TIFF0007687790000002.tif8134 of the i-th input video 420 as shown in the following formula (1).
[0064]
Equation
[0065] Then, the video restoration device uses the depth TIFF0007687790000006.tif9134 corresponding to TIFF0007687790000007.tif13134 to calculate the three-dimensional camera coordinates i TIFF0007687790000008.tif9134 for the i-th detection unit 411C as shown in the following formula (2).
[0066]
Number
[0067] As described above, it is difficult to accurately estimate the depth value to the target point 490 indicated by the corresponding pixel only from the pixel value of the input video. However, the video restoration device according to an embodiment can perform coordinate conversion according to the above-described formula (2) using the depth values corresponding to some disparities within the limited range of disparities. Here, the range of disparities may be limited to [d min , d max , and the depth value may also be TIFF0007687790000015.tif14134 limited. Z min may be, for example, 10 cm as the minimum shooting distance of the image sensor. For example, assume that the i-th input video 420 in FIG. 4 has a disparity of d = 1 with respect to the target video 430, and the depth value corresponding to d = 1 (for example, z 1 ) can be used. In the above-described formula (2), the depth TIFF0007687790000016.tif13134 is z 1Here, the image decompressor can set all pixels of the i-th input image 420 to have the same disparity with respect to the target image 430, and convert the coordinates of all pixels using the same depth value (e.g., the depth value corresponding to d=1). Similarly, the image decompressor can set all pixels of the i-th input image 420 to have the same disparity with respect to the target image 430 using d=2, 3, 4, .., d=1, ... max In other words, the image decompressor can use the depth value z corresponding to the disparity d=2. 2 The 3D camera coordinates are transformed using the depth value z corresponding to the disparity d=3. 3 The 3D camera coordinates are transformed using the depth value z corresponding to the disparity d=4. 4 3D camera coordinates transformed using d=d max The depth value z corresponding to the disparity min The three-dimensional camera coordinate values converted using the above formula can be individually obtained. For reference, in FIG. 4, the disparity having an integer value is taken as an example for explanation, but the present invention is not limited thereto.
[0068] The image restoration device converts the 3D camera coordinates of the i-th input image 420, which has been converted using disparity according to the above-mentioned equation (2), into a target detection unit 412C. T 3D camera coordinates for This can be converted to TIFF0007687790000017.tif13134 using the following formula (3).
[0069]
number
[0070] The video restoration device can normalize the three-dimensional camera coordinate T with respect to the target detection unit 412 C TIFF0007687790000021.tif13134 as shown in the following formula (4).
[0071]
Equation
[0072] Finally, the video restoration device is the target detection unit 412C T Coordinates normalized with respect to Pixel coordinates in the pixel coordinate system corresponding to the target video 430 from TIFF0007687790000026.tif13134 TIFF0007687790000027.tif14134 can be calculated as shown in the following formula (5).
[0073]
Equation
[0074] According to the above formulas (1) to (5), the video restoration device can i obtain the pixel coordinates of the i-th detection unit 411C TIFF0007687790000033.tif13134 as the pixel coordinates of the target detection unit By converting to TIFF0007687790000034.tif13134, the i-th input video 420 can be warped into the pixel coordinate system corresponding to the target video 430. A series of operations according to the above-described mathematical formulas (1) to (5) are shown as warping operations. For the sake of convenience of explanation, the warping operation has been described in time series, but it is not limited thereto, and an operation in which the operations according to the above-described mathematical formulas (1) to (5) are combined (for example, a unified matrix operation, etc.) may be used.
[0075] The video restoration device according to an embodiment can generate one warped video by warping all the pixels of the corresponding input video (for example, input video) into the pixel coordinate system corresponding to the target video 430 using a single depth corresponding to one of a plurality of disparities. For example, when the disparity d has a value of "j", the depth value z corresponding to the disparity d = j is used, and all the pixels of the j-th warped video generated from the i-th input video 420 are warped using the same depth value z. Here, j may be an integer of 1 or more and d or less, but is not limited thereto, and may be a real number of 0 or more and d or less. For reference, the maximum disparity d is determined as in the following mathematical formula (6). j Using it, all the pixels of the j-th warped video generated from the i-th input video 420 are warped using the same depth value z. j Here, j may be an integer of 1 or more and d max or less, but is not limited thereto, and may be a real number of 0 or more and d max or less. For reference, the maximum disparity d max is determined as follows in the following mathematical formula (6).
[0076]
Equation
[0077] The depth corresponding to one of the plurality of disparities is determined based on the corresponding disparity set for the input video with respect to the target video 430 and the interval b between the detection units C i , C T where the depth of all target points 490 shown in the external scene is the same as z j , all pixels of the j-th warping video corresponding to one of the plurality of disparities can be accurately aligned with respect to the target video 430. However, since the depth of the actual subject is diverse, only some pixels in the input video are aligned with the target video 430.
[0078] For example, as shown in FIG. 4, the video restoration device generates warping videos corresponding to a plurality of disparities from the i-th input video 420. The plurality of warping videos include the first warping video 421 generated using the depth z 1 corresponding to d = 1, the second warping video 422 generated using the depth z 2 corresponding to d = 2, the third warping video 423 generated using the depth z 3 corresponding to d = 3, the fourth warping video 424 generated using the depth z 4 corresponding to d = 4, or the warping video 425 generated using the depth z max corresponding to d = d min . For the sake of convenience of explanation, a part of the input video and each warping video are shown one-dimensionally, but it is not limited thereto, and each video may be two-dimensional.
[0079] Any target point 490 is detected from a target pixel 439 in a target video 430 and from an input pixel 429 in an input video. When the disparity between the input video and the target video 430 is set to d = 1, the video restoration device can generate a first warped video 421 by warping the input video so that a pixel in the input video at a position separated from the target pixel 439 by the above-described disparity (e.g., d = 1) from the target pixel 439 is aligned with the target pixel 439 of the target video 430. The second warped video 422 may be a video in which the input video is warped so that a pixel separated from the target pixel 439 by a disparity of d = 2 from the target pixel 439 is aligned with the target pixel 439. In the remaining warped videos 423 to 425, they are warped from the input video so that pixels separated from the target pixel 439 by the set disparity are aligned with the target pixel 439. As shown in FIG. 4, in the first warped video 421, the second warped video 422, and the warped video 425, the input pixel 429 is aligned at a position different from the target pixel 439. However, in the third warped video 423 and the fourth warped video 424, the input pixel 429 may be aligned with the target pixel 439 within one pixel error. Pixel alignment between the warped video and the target video will be described with reference to FIG. 5 below.
[0080] FIG. 5 is a diagram for explaining the matching between the pixels of the warped video and the pixels of the target video according to an embodiment.
[0081] According to one embodiment, in each of a plurality of warped images warped from an input image 520 using a plurality of disparities, at least one of the pixels included in the corresponding warped image may show an error of one pixel or less with a corresponding target pixel in the target image 530. As a result, even if accurate depth estimation for a target point is omitted, the image restoration apparatus can match at least one pixel in at least one of the plurality of warped images to the target point by generating a warped image using a depth corresponding to a preset disparity. For example, in FIG. 5, the first pixel 501 of the first warped image 521 warped from the input image 520 may be matched with the pixel 531 of the target image 530. Also, the second pixel 502 of the second warped image 522 may be matched with the pixel 532 of the target image 530.
[0082] In FIG. 5, for the sake of convenience of explanation, an example in which an arbitrary pixel in the warped image is matched with the target image 530 has been described, but the present invention is not limited thereto. Any region in the input image may include the same optical information as the corresponding region in the target image, and among the warped images warped from the corresponding input image, some of the corresponding regions may be matched with the corresponding regions in the target image.
[0083] FIG. 6 is a diagram for explaining generation of an output image through alignment of warped images according to one embodiment.
[0084] According to one embodiment, the video restoration device generates warping videos 631 to 635 from a plurality of input videos 620. For example, the video restoration device generates a first warping video 631 using depth values corresponding to arbitrary disparities from a first input video 621. The second warping video 632 may be a video warped from a second input video 622, the third warping video 633 may be a video warped from a third input video 623, the fourth warping video 634 may be a video warped from a fourth input video 624, and the fifth warping video 635 may be a video warped from a fifth input video 625. In the first to fifth input videos 621 to 625, a first pixel 601 is matched to a target video. The target video is selected as one of the input videos, but is not limited thereto. In the second warping video 632, a second pixel 602, in the third warping video 633, a third pixel 603, and in the fourth warping video 6349, a fourth pixel 604 may be respectively matched to the target video. There are pixels that match the target video in the remaining warping videos, which are omitted for simplicity of explanation.
[0085] The video restoration device provides a plurality of input videos 620 and warping videos 631 to 635 to a video restoration model 640. The video restoration model 640 may include a convolutional neural network including convolutional layers as described above, and is trained to output a high-resolution output video 690 from input video information and warping video information. For example, the video restoration device can generate a high-resolution output video 690 by aligning (registering: position alignment) pixels that match the target video with various video information using the video restoration model 640.
[0086] FIG. 7 is a diagram for explaining a camera calibration process according to one embodiment.
[0087] According to one embodiment, the video restoration device stores in advance information for generating warping video information.
[0088] For example, in step S710, the video restoration device performs camera calibration. Although a plurality of detection units included in the image sensor are designed to be in an aligned state 701, the actually manufactured image sensor may show a misaligned state 702. The video restoration device performs camera calibration using a checkerboard. The video restoration device detects the principal point with respect to the x-axis and y-axis in the detection unit as an internal camera parameter through camera calibration TIFF0007687790000036.tif13134, and the focal length with respect to the x-axis and y-axis in the detection unit TIFF0007687790000037.tif13134 is calculated. Also, the video restoration device calculates the rotation information R with respect to the world coordinate system of the detection unit and the translation information T with respect to the world coordinate system of the detection unit as external parameters through camera calibration i , the translation information T with respect to the world coordinate system of the detection unit i is calculated.
[0089] Then, in step S720, the video restoration device generates and stores depth information for each disparity. For example, the video restoration device calculates a depth value corresponding to a given disparity between input videos detected by two detection units based on the arrangement relationship between the detection units (for example, the angle between each optical axis, the distance between the detection units, etc.). As described above, the disparity is composed of a finite number within a limited range. For example, the disparity may be composed of integer disparities, but is not limited thereto.
[0090] According to an embodiment, the video restoration device can calculate in advance (for example, before step S210) a coordinate mapping function applied as a warping operation using internal camera parameters and external parameters. The coordinate mapping function is a function that converts the coordinates of each pixel of the input video into a pixel coordinate system corresponding to the target video using the internal camera parameters, external parameters, and depth corresponding to the given disparity described above. For example, it represents a function in which a series of operations according to formulas (1) to (5) are integrated. The video restoration device can calculate and store the coordinate mapping function in advance for each individual disparity and for each detection unit.
[0091] In step S220 shown in FIG. 2 described above, the video restoration device loads the coordinate mapping function calculated in advance for any one of the plurality of input videos with respect to the other detection unit and the target detection unit that captured the corresponding input video in order to generate warping video information. The video restoration device can generate warping video information by applying the coordinate mapping function calculated and stored in advance to the input video, and can quickly generate the warping video information provided to the video restoration model in order to generate a high-resolution output video while minimizing the amount of calculation.
[0092] However, the coordinate mapping function does not necessarily have to be calculated and stored in advance as described above. The video restoration device may store the internal camera parameters and external parameters instead of the coordinate mapping function calculated in advance. The video restoration device can load the internally stored camera parameters and external parameters, calculate the coordinate mapping function, and generate warping video information for the input video using the calculated coordinate mapping function.
[0093] FIG. 8 is a diagram showing the structure of a video restoration model according to an embodiment.
[0094] The video restoration device according to an embodiment can generate an output video by providing data obtained by concatenating a plurality of input information (e.g., input video) and warping information (e.g., warping video) to the input of a video restoration model.
[0095] For example, the video restoration device generates concatenated data 841 by concatenating a plurality of warping video information 829 generated from the input video information 820 and the input video information 820 as described above. For example, the video restoration device can concatenate D warping videos generated for each of the N input videos together with the N input videos acquired from N detection units. As shown in FIG. 8, since the concatenated data 841 is obtained by concatenating the input video information and the warping video information, it includes (D + 1) × N videos. The resolution of each video may be H × W, where H represents the number of pixels corresponding to the height of the video and W represents the number of pixels corresponding to the width of the video. The concatenation operation may be included as part of the operations of the video restoration model.
[0096] The video restoration device extracts feature data from the concatenated data 841 via a convolutional layer 842. The video restoration device may perform a shuffle 843 so that pixel values indicating the same location in the extracted plurality of feature data are adjacent to each other. The video restoration device can generate a high-resolution output video from the feature data via residual blocks 844 and 845. A residual block refers to a block that outputs feature data extracted from the data input to the corresponding block and residual data (residual data) between the data input to the corresponding block. Since the resolution of the output video is (A × H) × (A × W), it is higher than the resolution of H × W of each of the plurality of input videos.
[0097] For reference, the subject is from the image sensor at [z min , z maxIf within the distance between them, each region of the target video contains information similar to the region at the same position of at least one of the (D + 1) × N reconstructed videos included in the combined data 841 described above (see FIGS. 5 and 6). Therefore, by providing the combined data 841 to the video restoration model 340, the video restoration device can use the information of the regions containing information similar to the target video in each input video and warping video, thus improving the performance of video restoration. Even if the depth information of the target point indicated by the individual pixels of the input video is not given, the video restoration device can generate a relatively high-resolution output video. Also, even if there is no alignment between the input videos, the video restoration device can restore the video as long as it knows only the camera parameter information.
[0098] In FIGS. 1 to 8, an example of directly warping the input video has been mainly described, but it is not limited thereto. Hereinafter, with reference to FIG. 9, an example of warping the feature data extracted from the input video will be described.
[0099] FIG. 9 is a diagram for explaining a video restoration process using a video warping model and a video restoration model according to an embodiment.
[0100] A video restoration device according to an embodiment can also use a video warping model 950 together with the video restoration model. The video warping model 950 includes a feature extraction model 951 and a warping operation 952. The video warping model 950 is a model trained to extract feature maps from the input video 920 respectively and warp the extracted feature maps. The parameters (for example, connection weights) of the feature extraction model 951 can be varied by training, but the warping operation 952 can be constant as the operations according to the above-described formulas (1) to (6).
[0101] For example, the video restoration device extracts a plurality of input feature maps from a plurality of input videos as a plurality of input video information using the feature extraction model 951. The feature extraction model 951 may include, for example, one or more convolutional layers, and the input feature maps are the result values of convolutional filtering. The video restoration device can generate a warping feature map as warping video information by warping each of the plurality of input feature maps to the pixel coordinate system corresponding to the target video using the depth corresponding to each of the plurality of disparities. The feature map warped to the pixel coordinate system of the target detection unit using the depth corresponding to a specific disparity in the input feature map can be shown as the warping feature map. Since the warping operation 952 applied to the input feature map is the same as the warping operation 952 applied to the input video 920 by the above-described mathematical formulas (1) to (5), a detailed description thereof is omitted.
[0102] As a reference, when an input video captured in a Bayer pattern is directly warped to the pixel coordinate system of the target detection unit, the Bayer pattern may be lost in the warped video warped from the corresponding input video. Color information may be lost in the warped video while the color information of each channel is mixed by warping. The video restoration device according to an embodiment extracts an input feature map from the input video before color information is lost by the warping operation 952, so that the color information is stored in the input feature map. The video restoration device calculates a warping feature map by applying the warping operation 952 to the input feature map extracted from the state in which the color information is stored. Therefore, the video restoration device can generate a high-resolution output video 990 in which color information is stored by providing data obtained by combining the plurality of input feature maps and the warping feature maps to the input of the video restoration model. As described above, the video restoration device can minimize the loss of color information.
[0103] An exemplary detailed structure of the video warping model will be described with reference to FIG. 10 below.
[0104] FIG. 10 is a diagram showing the detailed structure of a video warping model according to an embodiment.
[0105] According to an embodiment, the video restoration device generates an input feature map and a warping feature map from a plurality of input videos using a video warping model 950. For example, the video restoration device extracts an input feature map from each of the plurality of input videos using a feature extraction model. The feature extraction model includes one or more convolutional layers 1051 as described above. The feature extraction model may also include a residual block 1052. For example, in FIG. 10, the feature extraction model includes one convolutional layer and M feature residual blocks. M is an integer of 1 or more. The video restoration device can extract an input feature map as a result value obtained by applying convolutional filtering to the individual input video 1020.
[0106] Then, the video restoration device applies a warping operation to the extracted input feature map. As described above, the video restoration device warps the input feature map corresponding to each detection unit to the pixel coordinate system of the target detection unit using the calibration information 1019 (for example, internal parameters and external parameters, etc.) of the image sensor 1010 and the depth corresponding to a plurality of disparities. For example, the video restoration device generates D warping feature maps for one input feature map by performing a warping operation on the depth corresponding to D disparities for each input feature map. The video restoration device generates combined data 1053 by combining a plurality of input feature maps and warping feature maps. The combined data 1053 includes information regarding N input feature maps and N×D warping feature maps.
[0107] The image restoration device can generate an output image 1090 with a high resolution (for example, a resolution increased by A times compared to the resolution of individual input images) by providing the combined data 1053 as an input to the image restoration model 340. For example, as shown in FIG. 10, the image restoration model 340 includes one convolutional layer 1042 and a plurality of residual blocks 1044, 1045. Among the plurality of residual blocks 1044, 1045, the residual block 1044 to which the combined data 1053 is input can receive the data to which shuffling 1043 is applied so that pixel values indicating the same point in the combined data 1053 are adjacent to each other.
[0108] The above-described video warping model 950 and image restoration model 340 may be trained simultaneously and / or sequentially during training. Since the warping operation that induces loss of color information is included in the video warping model 950, the video warping model 950 learns parameters that minimize color loss through training. The video warping model 950 and the image restoration model 340 may be trained via backpropagation. For example, the video warping model 950 and the image restoration model 340 are trained so that a high-resolution training output (for example, a ground truth image of a high-resolution correct value image) is output from a low-resolution training input (for example, a plurality of low-resolution images). The video warping model 950 and the image restoration model 340 during training can be respectively shown as a temporary video warping model 950 and a temporary image restoration model 340. From any training input, the temporary video warping model 950 and the temporary image restoration model 340 generate a temporary output, and the parameters (for example, connection weight values between nodes) of the temporary video warping model 950 and the temporary image restoration model 340 can be adjusted so that the loss between the temporary output and the correct value image is minimized.
[0109] FIG. 11 is a block diagram showing the configuration of an image restoration device according to an embodiment.
[0110] The video restoration device 1100 according to one embodiment includes an image sensor 1110, a processor 1120, and a memory 1130.
[0111] The image sensor 1110 acquires a plurality of input video information. According to one embodiment, the image sensor 1110 acquires a plurality of input videos captured through lenses arranged at different positions as the plurality of input video information. For example, the image sensor 1110 includes a detection unit that acquires each of the plurality of input video information. To acquire N pieces of input video information, the image sensor 1110 includes N detection units. However, it is not limited to the case where N detection units are included in a single image sensor 1110, and each of the N image sensors 1110 may include a detection unit.
[0112] The processor 1120 can generate a plurality of warped image information corresponding to a plurality of disparities from each of the plurality of input video information, and generate an output video using a video restoration model based on the plurality of input video information and the plurality of warped image information. The processor 1120 can skip the depth detection to the target point corresponding to an individual pixel and generate an output video without performing the depth detection operation.
[0113] However, the operation of the processor 1120 is not limited thereto, and the processor 1120 may perform at least one of the operations described above with reference to FIGS. 1 to 10 simultaneously or sequentially.
[0114] The memory 1130 can store temporarily or permanently the data required for the execution of the video restoration method. For example, the memory 1130 stores input video information, warped image information, and output video. Further, the memory 1130 may store a video warping model and its parameters, and a video restoration model and its parameters. The parameters of each model may already be trained.
[0115] FIG. 12 is a block diagram showing a computing device according to an embodiment.
[0116] Referring to FIG. 12, the computing device 1200 is a device that generates high-resolution video using the video restoration method described above. In one embodiment, the computing device 1200 corresponds to the device 1100 described with reference to FIG. 11. The computing device 1200 may be, for example, a video processing device, a smartphone, a wearable device, a tablet computer, a netbook, a laptop, a desktop, a PDA (personal digital assistant), or an HMD (head mounted display). Also, the computing device 1200 may be realized as a non-dedicated camera device for vehicles, drones, and CCTVs. As a different example, the computing device 1200 may also be realized as a webcam camera device for video calls, a 360-degree shooting VR camera device, a VR and AR camera device.
[0117] Referring to FIG. 12, the computing device 1200 includes a processor 1210, a storage device 1220, a camera 1230, an input device 1240, an output device 1250, and a network interface 1260. The processor 1210, the storage device 1220, the camera 1230, the input device 1240, the output device 1250, and the network interface 1260 communicate via a communication bus 1270.
[0118] The processor 1210 executes functions and instructions for execution within the computing device 1200. For example, the processor 1210 processes instructions stored in the storage device 1220. The processor 1210 can perform one or more operations described above with reference to FIGS. 1-11.
[0119] The storage device 1220 stores information or data necessary for the execution of the processor. The storage device 1220 includes a computer-readable storage medium or a computer-readable storage device. The storage device 1220 stores instructions for execution by the processor 1210 and stores related information while software or an application is being executed by the computing device 1200.
[0120] The camera 1230 captures a plurality of input videos. Also, although the video has been mainly described as a still image above, the camera 1230 is not limited thereto and may capture a video composed of one or more image frames. For example, the camera 1230 may generate frame videos corresponding to each of a plurality of lenses. In this case, the computing device 1200 can generate a high-resolution output video for each frame from the plurality of input videos corresponding to the individual frames using the video warping model and the video restoration model described above.
[0121] The input device 1240 receives input from the user by tactile, video, audio, or touch input. The input device 1240 includes a keyboard, a mouse, a touch screen, a microphone, or any other device that can detect input from the user and transmit the detected input.
[0122] The output device 1250 provides the output of the computing device 1200 to the user via a visual, auditory, or tactile channel. The output device 1250 may include, for example, a display, a touch screen, a speaker, a vibration generating device, or any other device capable of providing output to the user. The network interface 1260 communicates with external devices via a wired or wireless network. According to one embodiment, the output device 1250 can provide the user with at least one of visual information, auditory information, and haptic information, such as the result of processing data. For example, the computing device 1200 can visualize the generated high-resolution output video via a display.
[0123] The above-described device is implemented by hardware components, software components, or a combination of hardware components and software components. For example, the devices and components described in this embodiment are implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an ALU (arithmetic logic unit), a digital signal processor, a microcomputer, an FPA (field programmable array), a PLU (programmable logic unit), a microprocessor, or different devices that execute and respond to an instruction. The processing device executes an operating system (OS) and one or more software applications executed on the operating system. Further, the processing device accesses, stores, manipulates, processes, and generates data in response to the execution of the software. For the sake of convenience of understanding, the processing device may be described as being one, but those skilled in the art will understand that the processing device includes a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing device includes a plurality of processors or one processor and one controller. Also, other processing configurations, such as a parallel processor, are possible.
[0124] Software includes a computer program, code, instructions, or a combination of one or more of them, and can configure a processing device to operate as desired or command the processing device independently or in combination. Software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or signal wave being transmitted, in order to be interpreted by the processing device or to provide instructions or data to the processing device. Software can be distributed on a computer system connected to a network and stored or executed in a distributed manner. Software and data can be stored in one or more computer-readable recording media.
[0125] The method according to this embodiment is embodied in the form of program instructions implemented via various computer means and recorded on a computer-readable recording medium. The recording medium includes program instructions, data files, data structures, etc. alone or in combination. The recording medium and the program instructions may be specially designed and configured for the purpose of the present invention, or may be known and usable to those skilled in the art of computer software technology. Examples of computer-readable recording media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROMs, RAMs, flash memories, etc. Examples of program instructions include not only machine language code generated by a compiler, but also high-level language code executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to execute the operations shown in the present invention, and vice versa.
[0126] As described above, the embodiments have been described with reference to the limited drawings. However, those skilled in the art can apply various technical modifications and variations based on the above description. For example, the described technology may be executed in an order different from the described method, and / or the components such as the described system, structure, device, circuit, etc. may be combined or assembled in a form different from the described method, or appropriate results can be achieved even if they are replaced or substituted by other components or equivalents.
[0127] Therefore, the scope of the present invention is not defined by being limited to the disclosed embodiments, but is defined by the claims and equivalents thereof.
Claims
1. In a video restoration method, a step of obtaining a plurality of input video information of a plurality of input videos; a step of generating a plurality of warping video information for each of a plurality of disparities based on the plurality of input video information; a step of generating an output video by using a video restoration model based on the plurality of input video information and the plurality of warping video information, wherein the plurality of disparities are preset, and the plurality of warping video information is generated by using the distances to target points of the plurality of input videos determined based on the preset plurality of disparities, the video restoration method.
2. The plurality of input videos of the plurality of input video information are captured through lenses arranged at different positions from each other, the video restoration method according to claim 1.
3. The step of generating the plurality of warping video information includes the step of warping each of the plurality of input videos to a pixel coordinate system corresponding to a target video by using a depth corresponding to each of the plurality of disparities, thereby generating a plurality of warping videos for each of the plurality of disparities as the warping video information, the video restoration method according to claim 2.
4. The step of generating the warping video includes the step of warping all pixels in a first input video among the plurality of input videos to a pixel coordinate system corresponding to the target video by using a single depth corresponding to a first disparity among the plurality of disparities, thereby generating a warping video, the video restoration method according to claim 3.
5. The disparity is set for the input video with reference to the target video, the depth corresponding to the disparity is based on the disparity and the distance between the target video and the detection units that captured the input video, the video restoration method according to claim 3.
6. The step of generating the output video includes the step of providing, as an input, data obtained by combining the plurality of input videos and the plurality of warping videos corresponding to each of the plurality of disparities to the video restoration model, thereby generating the output video, the video restoration method according to claim 3.
7. The video restoration method according to any one of claims 1 to 6, wherein the plurality of input video information includes a plurality of input feature maps extracted from the plurality of input videos using a feature extraction model.
8. The step of generating the plurality of warping video information includes generating, as the plurality of warping video information, a plurality of warping feature maps for each of the plurality of disparities by warping each of the plurality of input feature maps to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities. The video restoration method according to claim 7.
9. The step of generating the output video includes generating the output video by providing, as an input to the video restoration model, data obtained by combining the plurality of input feature maps and the plurality of warping feature maps corresponding to each of the plurality of disparities. The video restoration method according to claim 8.
10. The video restoration model is a neural network including at least one convolutional layer configured to apply convolutional filtering to input data. The video restoration method according to any one of claims 1 to 9.
11. The plurality of disparities are less than or equal to a maximum disparity and greater than or equal to a minimum disparity, The maximum disparity is based on the focal length of the detection unit, the distance between the detection units, and the minimum shooting distance of the detection unit, The detection unit is configured to capture an input image corresponding to the plurality of input video information. The video restoration method according to any one of claims 1 to 10.
12. The plurality of disparities are a finite number. The video restoration method according to claim 11.
13. The step of generating the output video includes generating the output video without detecting the depth to the target point corresponding to an individual pixel. The video restoration method according to any one of claims 1 to 12.
14. The step of generating the plurality of warping video information includes generating warping video information by applying a coordinate mapping function to an input video corresponding to the input video information. The coordinate mapping function is the video restoration method according to any one of claims 1 to 13, which is predetermined for a detection unit configured to capture an input image and a target detection unit configured to capture a target image.
15. The resolution of the output video is higher than the resolution of each of the plurality of input video information, and the video restoration method according to any one of claims 1 to 14.
16. The plurality of input video information includes a multi-lens video captured by an image sensor including a multi-lens array, The multi-lens video includes the plurality of input videos, and the video restoration method according to any one of claims 1 to 15.
17. The plurality of input videos of the plurality of input video information are individually captured by a plurality of image sensors, and the video restoration method according to any one of claims 1 to 16.
18. A computer-readable recording medium storing one or more computer programs including instruction words for performing the method according to any one of claims 1 to 17.
19. A video restoration device, An image sensor that acquires a plurality of input video information of a plurality of input videos, Based on each of the plurality of input video information, a plurality of warping video information is generated for each of the plurality of disparities, and an output video is generated using a video restoration model based on the plurality of input video information and the plurality of warping video information. A processor, Including, The plurality of disparities are preset, and the processor is configured to generate the plurality of warping video information by using the distances to the target points of the plurality of input videos determined based on the preset plurality of disparities. Video restoration device.
20. A lens array including a plurality of lenses, Including a plurality of detection elements for detecting light that has passed through the lens array, the plurality of detection elements include detection regions corresponding individually to the plurality of lenses, and a detection array configured to acquire a plurality of input information of a plurality of input videos. Based on each of the plurality of input information, a plurality of warping information is generated for each of the plurality of disparities, and an output video is generated using a video restoration model based on the plurality of input information and the plurality of warping information. A processor, Including, The plurality of disparities are preset, and the processor is configured to generate the plurality of warping information by using the distances to the target points of the plurality of input videos determined based on the preset plurality of disparities, a video restoration device.
21. The video restoration device according to claim 20, wherein the resolution of the output video is higher than the resolution of each of the plurality of input information.
22. The video restoration device according to claim 20 or 21, wherein the processor generates the plurality of warping information for each of the plurality of disparities by warping each of the plurality of input information to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities.
23. The video restoration device according to claim 22, wherein the processor generates warping information by warping all pixels corresponding to the input information to a pixel coordinate system corresponding to the target video using a single depth corresponding to a first disparity among the plurality of disparities.
24. The video restoration device according to claim 22, wherein the processor generates the output video by providing, as an input, data obtained by combining the plurality of input information and the warping information corresponding to each of the plurality of disparities to the video restoration model.
25. The video restoration device according to any one of claims 20 to 24, wherein the processor extracts a plurality of input feature maps as the plurality of input information from the plurality of input videos using a feature extraction model.
26. The video restoration device according to claim 25, wherein the processor generates a plurality of warping feature maps as the plurality of warping information for each of the plurality of disparities by warping each of the plurality of input feature maps to a pixel coordinate system corresponding to a target video using a depth corresponding to each of the plurality of disparities.
27. The video restoration device according to claim 26, wherein the processor generates the output video by providing, as an input, data obtained by combining the plurality of input feature maps and the warping feature maps corresponding to each of the plurality of disparities to the video restoration model.
Citation Information
Patent Citations
Image super-resolution method and device based on optical-field collection device
CN108074218A
Depth estimation apparatus, reconfigured image generation device, depth estimation method, reconfigured image generation method and program
JP2013178684A
Image processing device and method, recording medium, and program
JP2016178678A
Method and apparatus for restoring image
JP2020042801A