Depth of Field Image Refocusing
By generating a depth mask and using a fuzzy core for refocusing processing, the problem of low image refocusing efficiency in the prior art is solved, and efficient multi-plane refocusing and shallow depth of field images are realized, which is suitable for the implementation of learnable pipelines.
Patent Information
- Application Number
- CN201980074801.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-03-21
AI Technical Summary
The prior art is difficult to achieve efficient image refocusing through digital image processing, especially in the simulation and synthesis of shallow depth of field images.
Multi-plane refocusing of the input image is achieved by generating a depth mask and performing refocusing using a fuzzy kernel. This method uses a differentiable function to evaluate the relationship between the image region and the depth plane, generates a masked image and performs a stacking process to generate a refocused image.
A differentiable refocusing step is implemented so that the refocusing process can be implemented as a layer within the learnable pipeline, supporting the simulation and synthesis of shallow depth of field images.
Smart Images

Figure CN113016003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to performing image refocusing. Some aspects relate to training an image processing model that performs refocusing of an input image. Background Art
[0002] The depth of field (DOF) of an image refers to the depth range within which all objects are sharp (in focus). For example, shallow DOF images can be used to separate an object from its background, and are often used in portraits and macro photography.
[0003] The depth of field of an image depends in part on the size of the aperture of the device used to capture the image. Therefore, some image processing devices with narrow apertures (such as smartphones) are generally unable to optically render shallow depth of field images, and instead capture essentially full-focus images.
[0004] However, through appropriate digital image processing techniques, it is possible to simulate or synthesize a shallow depth of field effect by refocusing a digital image.Some methods of computationally refocusing a digital image exploit the depth information of the image. Figure 1 An example of how to computationally refocus an input digital image is shown.
[0005] Figure 1 An all-focus image 101 is shown. A map containing depth information for image 101 is shown at 103. In this example, the map is a disparity map, but in other examples it may be a depth map. A disparity map encodes spatial displacements between blocks of a pair of (usually stereo) images representing a common scene. When the pair of images are captured by two spatially separated cameras, a point in the scene is projected to different block locations within each image. The coordinate shift between these different block locations is the disparity. Therefore, each block in the disparity map encodes a spatial displacement between a portion of the scene represented by a pair of images. The blocks of the disparity map encode the coordinate shift between a portion of the scene represented by a juxtaposed block in one of a pair of images and the portion of the scene represented by the other of the pair of images. The disparity values of the blocks of the disparity map represent the coordinate shift in the corrected image domain.
[0006] A block is a block of one or more pixels of an image. Parts of a scene that are deeper from the image plane have relatively smaller block displacements between a pair of images than parts of the scene that are shallower from the image plane. Therefore, the values of a disparity map are usually inversely proportional to the depth values of the corresponding depth map.
[0007] Based on the input image 101 and the image 103, the refocused images 105 and 107 can be generated by computationally refocusing the image 101. In the image 105, the background is blurred by refocusing and the foreground remains in focus; in the image 107, the foreground is blurred by refocusing and the background remains in focus. Summary of the invention
[0008] The Summary introduces some concepts that are further described in the Detailed Description. The Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0009] According to one aspect, an image processing device is provided, comprising a processor, the processor being used to generate a refocused image based on an input image and a map indicating depth information of the input image through the following steps: for each plane of a plurality of planes associated with corresponding depths within the image: generating a depth mask, the value of the depth mask indicating whether an area of the input image is within a specified range of the plane, wherein whether an area is within the specified range of the plane is evaluated by evaluating a differentiable function of a range between the area of the input image and the plane determined according to the map; generating a mask image based on the input image and the generated depth mask; refocusing the mask image using a blur kernel to generate a refocused partial image; generating the refocused image based on the plurality of refocused partial images.
[0010] This makes it possible to perform a differentiable refocusing step, allowing the refocusing step to be implemented as a layer within a learnable pipeline.
[0011] Multiple planes can define a depth range, and the input image can have a depth of field that spans the depth range of the multiple planes. This can enable shallow depth of field images to be refocused.
[0012] The map indicating the depth information of the image may be a disparity map, and each of the plurality of planes may be a plane at a corresponding disparity value; or (ii) the map indicating the depth information of the image may be a depth map, and each of the plurality of planes may be a plane at a corresponding depth. This allows for the flexibility of using either a depth map or a disparity map in the refocusing stage.
[0013] The function can smoothly vary between a lower bound and an upper bound defining a range of function values, and when the image region for which the function is being evaluated is within the specified range of the plane, the function has a value within the first half of its range; when the image region for which the function is being evaluated is not within the specified range of the plane, the function has a value within the second half of its range. This enables the function value returned from the calculation function to indicate whether the region of the image is within the specified range of the plane.
[0014] The differentiable function can be evaluated to assess whether The condition of , where D(x) is the value of the disparity map or depth map of the image area at the block x of the input image, d′ is the disparity or depth of the plane, a is the size of the aperture, and when the condition is met, it is determined that the image area is within the specified range of the plane; when the condition is not met, it is determined that the image area is not within the specified range of the plane. This provides a quantitative condition to determine whether the image area is within the specified range of the plane.
[0015] The depth mask can specify a value for each block of one or more pixels of the input image. This allows each block of the image to be evaluated to determine whether the block is within a specified range of the plane.
[0016] The blur kernel is a radial kernel whose radius is a function of the aperture size and the depth or parallax difference between the plane and the focal plane. This makes the blur dependent on the depth of the plane.
[0017] The resulting mask image can contain only those areas of the input image that are determined to be within the specified range of the plane. This allows the partial images for each plane to be layered on top of each other.
[0018] According to a second aspect, a method for training an image processing model is provided, the method comprising:
[0019] receiving a plurality of image tuples, each image tuple comprising a pair of images representing a common scene and a refocused training image of the scene;
[0020] For each image tuple:
[0021] (i) processing the pair of images using the image processing model to estimate a map indicative of depth information of the scene;
[0022] (ii) for each plane of a plurality of planes associated with a corresponding depth within the image:
[0023] generating a depth mask, wherein a value of the depth mask indicates whether a region of the input image is within a specified range of the plane, wherein whether the region is within the specified range of the plane is evaluated by evaluating a differentiable function of a range between the region of the input image and the plane determined according to the estimation map;
[0024] generating a mask image based on the input image and the generated depth mask;
[0025] refocusing the mask image using a blur kernel to generate a refocused partial image;
[0026] (iii) generating a refocused image based on the plurality of refocused partial images;
[0027] (iv) estimating the difference between the generated refocused image and the refocused training image;
[0028] (v) adjusting the image processing model based on the estimated difference.
[0029] This provides a way to train image processing models to generate refocused images that can be learned through refocusing.
[0030] The generated map may be of the same resolution as the pair of images, and the input image may be one of the pair of images. This allows the refocusing stage to be performed sequentially with the depth estimation stage, and a full-resolution refocused image may be generated from the refocusing stage without upsampling in that stage.
[0031] The generated map may have a reduced resolution compared to the pair of images, and the input image may be formed as a reduced resolution image of one of the pair of images. This enables the steps of the refocusing phase to be performed on fewer image patches.
[0032] The step (iii) of generating the refocused image may be performed using the resolution reduction map, and may further comprise the steps of: generating a refocused image with reduced resolution from the plurality of refocused partial images; upsampling the refocused image with reduced resolution to generate the refocused image. This enables the generation of a refocused image at full resolution.
[0033] According to a third aspect, a method for adjusting an image processing model is provided, the method comprising:
[0034] receiving a plurality of image tuples, each image tuple comprising a pair of images representing a common scene and a refocused training image of the scene;
[0035] For each image tuple:
[0036] Processing the pair of images using an image processing model includes:
[0037] (i) processing the pair of images using a computational neural network to extract features from the pair of images;
[0038] (ii) comparing the features for a plurality of disparity planes to generate a cost volume;
[0039] (iii) generating a mask for each of a plurality of planes associated with a corresponding depth within the image based on a corresponding upsampled strip of the cost volume;
[0040] Perform the refocusing phase, including:
[0041] (iv) for each plane in the plurality of planes:
[0042] generating a mask image based on the input image and the generated mask;
[0043] refocusing the mask image using a blur kernel to generate a refocused partial image;
[0044] (v) generating a refocused image based on the plurality of refocused partial images;
[0045] (vi) estimating a difference between the generated refocused image and the refocused training image;
[0046] (vii) adjusting the image processing model based on the estimated difference.
[0047] This enables refocusing to be performed directly from the cost volume and can reduce the number of steps in the learnable pipeline.
[0048] An image processing model may be provided, which is adjusted by the method of this article to generate a refocused image from a pair of input images representing a common scene. This may provide an image processing model that has been adjusted or trained to generate a refocused image from the input images.
[0049] An image processing device may be provided, comprising a processor and a memory storing instructions in a non-transient form executable by the processor to implement an image processing model adapted by any of the methods herein to generate a refocused image from a pair of input images representing a common scene. This may enable the adapted or trained image processing model to be implemented within the device to perform refocusing.
[0050] The above features may be combined as appropriate, as will be apparent to the skilled person, and may be combined with any aspects of the examples described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Examples will now be described with reference to the accompanying drawings, in which:
[0052] Figure 1 An example of computationally generating a refocused image from an input digital image is shown;
[0053] Figure 2 shows an overview of the stages for generating a refocused image from an input image and a map indicative of depth information of the input image;
[0054] Figure 3 is an example of an image processing device for generating a refocused image from input images and a map indicating depth information of these images;
[0055] Figure 4 is a flowchart of a method for generating a refocused image according to an input image and a map indicating depth information of the input image provided by an embodiment of the present invention;
[0056] Figure 5 is a graph that approximates the Heaviside step function using a differentiable function;
[0057] Figure 6 is a schematic diagram of an exemplary StereoNet pipeline for estimating a graph from a pair of stereo images;
[0058] Figure 7 is a schematic diagram of an exemplary learnable pipeline for generating a refocused image from a pair of stereo images;
[0059] Figure 8 Shows the use of Figure 7 An exemplary predicted image obtained by the pipeline of
[0060] Fig. 9 is a schematic diagram of another exemplary learnable pipeline for generating a refocused image from a pair of stereo images;
[0061] Fig.10 is a flow chart of a method for adjusting an image processing model;
[0062] Fig.11 is a schematic diagram of another exemplary learnable pipeline for generating a refocused image from a pair of stereo images;
[0063] Fig.12 is used Fig.11 Flowchart of a method for pipeline adjustment of an image processing model in FIG.
[0064] The accompanying drawings show various examples. It will be appreciated by the skilled person that the element boundaries (e.g., boxes, box groups or other shapes) shown in the accompanying drawings represent an example of boundaries. In some examples, an element may be designed as multiple elements, or multiple elements may be designed as one element. Where appropriate, common reference numerals are used in all the accompanying drawings to represent similar features. DETAILED DESCRIPTION
[0065] The following description is presented by way of example to enable a person skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to a person skilled in the art. The embodiments are described by way of example only.
[0066] Figure 2 is a schematic diagram of a pipeline provided by the present invention for performing computational image refocusing.
[0067] The computational pipeline is divided into two general stages: depth estimation 201 and refocusing 203. The depth estimation stage 201 can be performed by a depth estimation module, and the refocusing stage can be performed by a refocusing module. The depth estimation stage processes an input image 205 representing a common scene to generate a map 207 containing depth information of the scene. In this example, the map is a disparity map, but in other examples, it can be a depth map. The map (depth map or disparity map) is an image that encodes the depth information of the scene represented by the input image 205. The pair of images can be stereo images.
[0068] One of the input images and the graph 207 are input into the refocusing stage 203. Other parameters input into the refocusing stage 203 of the pipeline are the virtual aperture size and the virtual focal plane used to refocus the final image. The values of the virtual aperture size and focal plane are configurable.
[0069] A method of performing the refocusing stage 203 which may be implemented by the refocusing module is now described.
[0070] First, a set of planes is defined. Depending on whether the graph is a depth map or a disparity map, these planes can be depth planes or disparity planes. A depth plane is a plane located at a specific depth within the acquired scene (i.e., the scene represented by image 205). A disparity plane is a plane located at a specific disparity value within the acquired scene, i.e., a plane associated with a specific disparity value. The (depth or disparity) planes are parallel to the image plane and the focal plane. These planes are also parallel to each other. These planes span the depth or disparity range within the image. The inter-plane spacing of the planes can be equal to 1 / a, where a is the aperture size. The span of the planes can be such that the depth range covered by the planes is less than or equal to the depth of field of the input image; that is, the input image can have a depth of field that spans the depth range of the planes.
[0071] The refocusing phase is then performed by iteratively executing a sequence of processing steps for a set of planes. This sequence of steps is now described.
[0072] 1. For every plane d∈[d min , d max ,], generate a depth mask, denoted as M d The depth mask contains values indicating whether an area of the input image is within the specified range of the plane d. The mask contains an array of values corresponding to the corresponding blocks of the input image, where each block is a block of one or more pixels. The portion of the image within a block x of the input image is considered to be within the specified range of the plane d if the following conditions are met:
[0073]
[0074] Where D(x) is the value of image 207 of block x, d′ is the depth value or disparity value of plane d, and a is the aperture size.
[0075] Whether the conditions of equation (1) are satisfied can be assessed by evaluating the Heaviside step function, which is written as:
[0076]
[0077] Where ξ is an arbitrary quantity. Therefore, in one example, the refocusing module can calculate the value of the depth mask by evaluating equation (2) as follows:
[0078]
[0079] Among them, M d (x) is the value of the depth mask for plane d and block x. In this example, a value of '0' indicates that the area of the image is not within the specified depth of plane d, and a value of '1' indicates that the area of the image is within the specified range of plane d.
[0080] 2. After calculating the depth mask M d Afterwards, the refocusing module calculates a mask image for the plane d based on the input image and the depth mask calculated for the plane d. The mask image only contains those parts of the input image that are determined to be within a specified range of the depth plane d.
[0081] Mathematically, the refocusing module can calculate the mask image of plane d as follows, denoted as I d :
[0082]
[0083] Where I is the input image, is the entry-wise Hadamard product operator.
[0084] 3. The refocusing module calculates the blur kernel K used to blur the mask image r In this example, the blur kernel is a radial kernel with a radius (denoted as r). The value of the radius depends on the distance of the plane d from the focal plane and the aperture size a. The radius can be calculated according to the following equation:
[0085] r=a|d′-d f | (5)
[0086] Among them, d f is the depth value or disparity value of the focal plane.
[0087] 4. The refocusing module then applies the blur kernel to the mask image I d , to generate the refocused partial image, expressed as Also apply the blur kernel to the depth mask M d, to generate a refocused or blurred depth mask, expressed as
[0088] The refocused partial image and the refocused depth mask may be generated by convolving the blur kernel with the partial image and the depth mask, respectively. That is, the refocusing module may generate the refocused partial image and the refocused depth mask by performing the calculation:
[0089]
[0090] Among them, K r is the blur kernel with radius r and * denotes convolution.
[0091] 5. The refocusing module then combines the refocused partial images of plane d with the refocused partial images of planes already calculated in the previous iteration of the refocusing phase. The refocused partial images may be combined from back to front in terms of depth, i.e., these refocused partial images are stacked, wherein the refocused partial images from the planes associated with the shallower depth values are stacked on top of the refocused partial images from the planes associated with the deeper depth values. The refocused partial images may be combined by accumulating block values of the partial images in a buffer. That is, the refocused partial images may be integrated to form a refocused image.
[0092] Mathematically, the refocusing module can calculate the stack image I in each iteration of the refocusing phase as follows f :
[0093]
[0094] I f The value of can be initialized with 0.
[0095] The refocused depth mask of plane d may similarly be combined with the refocused depth masks of planes computed in previous iterations of the refocusing stage, ie the refocused depth masks may be integrated.
[0096] Mathematically, the refocusing module can calculate the stacked refocused depth mask M at each iteration of the refocusing phase as follows f :
[0097]
[0098] M f The value of can be initialized with 0.
[0099] 6. After performing the sequence of steps 1 to 5 for each plane in the set, the depth mask M is obtained by using the stacked refocused depth mask M. f Standardized stacked image I fTo generate the refocused image. That is, the refocusing module calculates the refocused image as follows:
[0100]
[0101] Where O is the final refocused image (such as Figure 2 209), is the item-wise Hadamard division operator.
[0102] Although the above algorithm enables the calculation of a refocused image from an input image, it has the disadvantage that it cannot be used as a layer within a learnable pipeline. In other words, it is not possible to learn with the refocusing algorithm. It has been learned that the reason why it is not possible to learn with the refocusing algorithm is that, since the algorithm is not differentiable, the gradient information cannot be back-propagated through the various stages of the algorithm to adjust the parameters used in the algorithm, and the algorithm is not differentiable because the Heaviside step function used to generate the depth mask in step 1 is not differentiable because it is not a continuous function.
[0103] This paper describes a method for performing image refocusing that can be learned by suitable machine learning techniques.
[0104] Figure 3 An image processing device 300 is shown that includes a depth estimation module 302 and a refocusing module 304. Modules 302 and 304 may be implemented as part of a processor of the device 300. The depth estimation module 302 receives as its input digital images (generally indicated at 306) and generates a map containing depth information of these digital images from these digital images. The refocusing module 304 receives: (i) a generated map D from the depth estimation module 302, which contains depth information of the digital image I; (ii) the digital image I for which the map has been generated; and (iii) a focal plane d f and an indication of the pore size a. d f The values of and a can be user-set, i.e., configurable. They can be called a virtual focal plane and a virtual aperture. In the example below, the image D can be a disparity map or a depth map. The image is an image that encodes the depth information of the image I. The disparity map contains disparity values for blocks of one or more pixels of the input image, and the depth map contains depth values for blocks of one or more pixels of the input image. Due to the inverse relationship between disparity and depth, the disparity map can usually be calculated based on the depth map, and vice versa.
[0105] The refocusing module 304 is used to digitally refocus the received input image I using the received image D. The operation of the refocusing module is combined with Figure 4 Flowchart description of the .
[0106] The refocusing module 304 defines a plurality of planes d. If the map is a disparity map, the planes may be disparity planes (i.e., planes at specific disparity values in the scene represented by the input image); or if the map is a depth map, the planes may be depth planes (i.e., planes at specific depth values in the scene represented by the input image). Due to the relationship between depth and disparity, the (disparity or depth) planes may be associated with depth within the image.
[0107] These planes are parallel to the image plane and the focal plane. These planes are also parallel to each other. These planes span the depth or disparity range within the image, ranging from the minimum depth level (denoted as d min ) to the plane associated with the maximum depth level (denoted as d max ) associated with a plurality of planes. For the sake of clarity, it should be noted that 'minimum depth level' refers to the minimum depth level of a plurality of planes, and 'maximum depth level' refers to the maximum depth level of a plurality of planes, and does not imply any limitation or constraint on the values of the maximum depth and the minimum depth of the image. The inter-plane spacing of the planes may be equal to 1 / a, where a is the aperture size. The depth span of the planes may be less than or equal to the depth of field of the input image I; that is, the input image may have a depth of field that spans the depth range of the planes. In some examples, the input image may be an all-focus image, but is not necessarily an all-focus image. In other examples, the depth span of the planes may be independent of the aperture size a, but may be set empirically.
[0108] Steps 402 to 408 are performed for each of the plurality of planes. In other words, the refocusing module 304 iteratively performs the sequence of processing steps (steps 402 to 408) for the plurality of planes. min or plane d max At the beginning, steps 402 to 408 are performed for each plane in turn. That is, the refocusing module can min Steps 402 to 408 are performed for each successively deeper depth plane, or may be performed from plane d max Steps 402 to 408 are initially performed for each successively shallower depth plane.
[0109] In step 402, the refocusing module 304 generates a depth mask of plane d, denoted as M d . The depth mask contains values indicating whether a region of the input image I is within a specified range of the plane d. The depth mask contains an array of values corresponding to corresponding blocks of the image I, where each block is a block of one or more pixels. The depth mask may contain a value corresponding to each pixel of the image I. A region of the image I within a block x of the image is considered to be within a specified range of the plane d if the condition specified in equation (1) above (repeated below for convenience) is satisfied:
[0110]
[0111] where D(x) is the value of the map for block x, d′ is the depth (if D is a depth map) or disparity (if D is a disparity map) of plane d, and a is the aperture size.
[0112] The refocusing module 304 evaluates whether the region is within a specified range of the depth plane by evaluating a differentiable function. The differentiable function is a function of the range between the region of the input image and the plane d determined according to the graph D(x). In other words, the refocusing module 304 evaluates the differentiable function of the block x of the image I to determine whether the region of the image within the block is within the specified range of the plane d; that is, the value of the evaluation function indicates whether the region is within the specified range of the plane d.
[0113] Therefore, the refocusing module 304 evaluates a differentiable function of the block x to evaluate whether the block satisfies the condition of equation (1).
[0114] A differentiable function is differentiable over the domain in which it is defined (i.e., continuous in its first derivative). A function can vary smoothly between a lower bound and an upper bound, which define the range of function values. When evaluated for an image region within a specified range of a plane d, a differentiable function preferentially returns a value within a first subrange of the function value; when evaluated for an image region not within a specified range of a plane d, a differentiable function preferentially returns a value within a second non-overlapping subrange of the function value. This enables the value of the evaluated function to indicate whether the image region is within the specified range of a plane d.
[0115] For example, the function can vary from a substantially constant first value to a substantially constant second value, such as by varying smoothly from a first asymptotic value to a second asymptotic value. The function can be centered around a point that has a range equal to the specified range (i.e., ) transitions from a first value to a second value. The function value transitions from a first subrange of values to a second subrange of values at a point where the range equals the specified range.
[0116] The differentiable function can be a smooth approximation of the Heaviside step function. It is known that an advantageous differentiable approximation of the Heaviside function of equation (2) is the tanh function:
[0117]
[0118] Figure 5 An exemplary tanh function y=tanh(x) (denoted by 502) that approximates the Heaviside step function H(x) (denoted by 504) is shown.
[0119] Therefore, in one example, the refocusing module 304 may calculate the depth mask M by evaluating equation (10): d The value of
[0120]
[0121] Where α is a constant that controls the sharpness of the approximation transition. The value of α is configurable. It can be set empirically. Typically, larger values of α can be used to more closely approximate the Heaviside function. For example, the value of α can be greater than 500 or greater than 1000.
[0122] It should be appreciated that in other examples, the differentiable function may not be a tanh function, but may take a different form.
[0123] In step 404, the refocusing module 304 performs a refocusing operation based on the input image I and the depth mask M. d Generate a mask image. The mask image or sub-image is represented as I d , where the superscript 'd' indicates that a mask image is generated for plane d. The mask image only contains those parts of the input image that are determined to be within a specified range of plane d, and can be calculated according to equation (4) above.
[0124] The refocusing module 304 may also calculate a blur kernel as described above in connection with step 3.
[0125] In step 406, the refocusing module 304 uses the blur kernel to refocus the mask image I d , to calculate the refocused partial image, expressed as As described above in conjunction with step 4 and equation (6).
[0126] The refocus module 304 also computes a blurred depth mask As described above in conjunction with step 4 and equation (6).
[0127] In step 408, the refocusing module 304 refocuses the partial image superimposed on the partially reconstructed refocused image I generated from the previous iteration of the refocusing phase f In this example, the refocused partial images are combined from back to front in terms of depth, i.e., they are stacked such that the refocused partial images associated with the shallower depth planes are stacked on top of the refocused partial images associated with the deeper depth planes. For each iteration of the method, the refocused partial images can be combined by accumulating block values of the partial images in a buffer. That is, the refocused partial images can be integrated to form a refocused image.
[0128] Mathematically, the refocusing module 304 may create a stack of images I at each iteration of the refocusing phase as follows: f :
[0129]
[0130] That is, the refocusing module 304 combines the refocused partial image of plane d with the refocused partial images of planes that have been calculated in the previous iteration. f and The multiplication reduces intensity leakage by preventing block values in areas not within the specified range of plane d from being blurred by the refocused partial image update.
[0131] The refocused depth mask of plane d may similarly be combined with the refocused depth masks of planes calculated in previous iterations of the refocusing stage, ie the refocused depth masks may also be integrated.
[0132] Mathematically, the refocusing module 304 may calculate the stacked refocused depth mask M at each iteration of the refocusing phase as follows: f :
[0133]
[0134] For multiple planes d∈[d min , d max After all steps 402-408 are executed, the refocusing module 304 has formed a completed stacked image I by stacking the refocused partial images. f , and a stacked refocused depth mask has been formed. Thus, the refocusing module generates a refocused image, denoted as O. The refocused image is obtained by using the completed stacked refocused depth mask M f The completed stacked image I f That is, the refocusing module 304 calculates the refocused image as follows:
[0135]
[0136] Therefore, the refocusing module 304 generates a refocused image based on the plurality of refocused partial images.
[0137] The above method uses a differentiable function to generate a depth mask for each plane, wherein the depth mask indicates the area in the input image that is within a specified range of the plane. Figure 4 The entire algorithm described is differentiable. This has the advantage of enabling the method to be used as part of a pipeline for generating a refocused image from an input image, which can be learned by a range of learning techniques. Some examples of how the method can be implemented within a learnable refocusing pipeline are now described.
[0138] Although combined Figure 4The methods described are suitable for implementation using a range of learning techniques, but for illustrative purposes only, examples referencing the "StereoNet" pipeline are provided.
[0139] Figure 6 An example of a StereoNet pipeline is shown. The StereoNet pipeline 600 is used to estimate a map D based on a pair of images representing a common scene. The pair of images can be stereo images. The map D can be a disparity map or a depth map.
[0140] The pair of stereo images is shown at 602, including a 'left image' and a 'right image'. In other examples, the images 602 are not 'left' and 'right' images, but some other pair of images representing a common scene. The pair of images is input into a feature extraction network 604, which extracts features from these images. These features can be extracted at a reduced resolution compared to the resolution of the input image 602.
[0141] The extracted features are then compared for multiple disparity planes. To do this, features extracted from one image in a pair of images (the left image in this example) are transferred to each of the multiple disparity planes and then compared to features extracted from the other image in the pair of images (the right image in this example). The comparison is performed by calculating the difference between the transferred features of one of the images and the non-transferred features of the other image. This can be done by calculating the difference (e.g., value) between juxtaposed blocks of the transferred features and the non-transferred features. The calculated difference for each disparity plane forms a corresponding strip of the cost volume 606.
[0142] The cost volume 606 passes through a softmin function 608, and from this a disparity map 610 is calculated. The disparity map 610 is a low-resolution image, i.e., the resolution is reduced compared to the input image 602. In this example, the resolution of the disparity map is 1 / 8 of the resolution of the input image. The low-resolution map is then upsampled using an upsampling network 612 to generate a map 614, which in this example is a disparity map. In other examples, the disparity map can be processed to generate a corresponding depth map. The map 614 is an image of the same resolution as the image 602.
[0143] Figure 7 An exemplary architecture is shown for implementing a differentiable method for computing a refocused image from an input image to form a learnable refocusing pipeline.
[0144] Figure 7 The refocusing pipeline 700 in FIG. 7 includes a depth estimation stage 600 and a refocusing stage 702. The depth estimation stage 600 may be performed by Figure 3 The depth estimation module 302 is implemented, and the refocusing stage can be implemented by the refocusing module 304.
[0145] To train or tune the refocusing pipeline 700, multiple image tuples may be used. Each image tuple may include a pair of images representing a common scene and a training image of the scene. The pair of images may be a stereo image formed by a 'left image' and a 'right image'. The pair of images is shown at 602. The training image may be a ground truth image. The training image may be a shallow DOF image. For completeness, it should be noted that the training images are not shown in FIG. Figure 7 Shown in.
[0146] For an image tuple, a pair of images representing a common scene in the tuple is input into the depth estimation stage 600. The depth estimation stage 600 estimates a graph 614 based on the pair of images, as described above in conjunction with Figure 6 described.
[0147] The calculated map 614 is then input into the refocusing stage 702. One of the images 602 is also input into the refocusing stage 702. The refocusing stage 702 then combines the above Figure 4 The described approach calculates the refocused image 704 based on the estimated map 614 and the input image.
[0148] The difference between the calculated refocused image 704 and the training image of the tuple is then calculated and fed back to the depth estimation stage 600. The difference between the estimated refocused image and the training image can be fed back as residual data. The residual data is fed back to the depth estimation stage 600 to adjust the parameters used by this stage to estimate the map from a pair of images.
[0149] Figure 8 Shown is an example of a predicted refocused image 802 that the inventors have obtained from a ground truth image 804. An estimation map 806 is also shown.
[0150] Fig. 9 Another exemplary architecture is shown for implementing a differentiable method for computing a refocused image from an input image to form a learnable refocusing pipeline.
[0151] Fig. 9 The refocusing pipeline 900 in FIG. 1 includes a depth estimation stage 902 and a refocusing stage 904. The depth estimation stage 902 may be performed by Figure 3 The refocusing stage 904 can be implemented by the depth estimation module 302 .
[0152] To train or tune the refocusing pipeline 900, multiple image tuples may be used. Each image tuple may include a pair of images representing a common scene and a training image of the scene. The pair of images may be a stereo image formed by a 'left image' and a 'right image'. The pair of images is shown at 906. The training images may be ground truth shallow DOF images.
[0153] For each image tuple, a pair of images of the tuple representing a common scene is input into the depth estimation stage 900. The depth estimation stage 900 estimates a map 908 from the pair of images.
[0154] The pair of images is input into a feature extraction network 910, which extracts features from the images to estimate map 908. The features may be extracted at a reduced resolution compared to the resolution of the input image 906.
[0155] The extracted features are then compared for multiple disparity planes. To do this, features extracted from one image in a pair of images (the left image in this example) are transferred to each of the multiple disparity planes and then compared to features extracted from the other image in the pair of images (the right image in this example). The comparison is performed by calculating the difference between the transferred features of one of the images and the non-transferred features of the other image. This can be done by calculating the difference (e.g., value) between juxtaposed blocks of the transferred features and the non-transferred features. The calculated difference for each disparity plane forms a corresponding strip of the cost volume 912.
[0156] The cost volume 912 is passed through a softmin function 914 and a map 908 is calculated therefrom. The map 908 is a low resolution image, i.e., the resolution is reduced compared to the pair of images 906. In this example, the resolution of the map is 1 / 8 of the resolution of the input images. The output of the softmin function 914 is a disparity map, so the map 908 can be a disparity map, or can be further processed to generate a depth map.
[0157] The calculated map 908 is then input into the refocusing stage 904. One of the pair of images 906 (the left image in this example) is downsampled and also input into the refocusing stage 904. The input image may be downsampled to have the same resolution as the map 908. Fig. 9 Downsampling is not shown.
[0158] Then, the refocusing stage 904 calculates a refocused image 916 with reduced resolution based on the estimated map 908 and an input image (the input image is a downsampled version of one of the pair of images 906). The refocusing stage 904 calculates a refocused image 916 with reduced resolution based on the estimated map 908 and the input image (the input image is a downsampled version of one of the pair of images 906). Figure 4 The method computes a reduced resolution refocused image 916 based on these inputs.
[0159] The reduced-resolution refocused image 916 is then upsampled by an upsampling network 918 to generate a refocused image 920. The refocused image 920 is a full-resolution image, i.e., it has the same resolution as the image 906. The upsampling may also be performed by the refocusing module 304.
[0160] The difference between the refocused image 920 and the training image of the tuple is then calculated and fed back to the depth estimation stage 902. The difference between the estimated refocused image 920 and the training image can be fed back as residual data. The residual data is fed back to the depth estimation stage 902 to adjust the parameters used by the stage to estimate the map from a pair of images.
[0161] Now combine Fig.10 The flowchart of provides an overview of how pipelines 700 and 900 can be trained or tuned to generate a refocused image from an input pair of images.
[0162] In step 1002, a plurality of image tuples are received, for example by the depth estimation module 302. Each image tuple includes a pair of images representing a common scene (eg, image 602 or 906), and a refocused training image that may be a ground truth image or the like.
[0163] Steps 1004 to 1018 are performed for each image tuple.
[0164] In step 1004, a pair of images of the tuple is processed using an image processing model to estimate the image D. Step 1004 may be performed by the depth estimation module 302. The image processing model may correspond to Figure 7 The depth estimation stage 600 or Fig. 9 The depth estimation stage 902 is shown. In other words, the performance of the image processing model may include performing the steps of the depth estimation stage 600 or 902. The map may be a disparity map or a depth map. The resolution of the map may be the same as the resolution of the pair of images in the tuple (e.g., map 614), or may have a reduced resolution compared to the pair of images (e.g., map 908).
[0165] Steps 1006 to 1014 may be performed by the refocusing module 304. These steps may correspond to Figure 7 The refocusing stage 702 or Fig. 9 The refocusing stage 904 is shown. These steps are performed using the image D generated in step 1004 and the input image I. The input image can be one of a pair of images in a tuple (e.g., as in pipeline 700), or can be an image formed by one of a pair of images in a tuple (e.g., as in pipeline 900, where the input image of the refocusing stage 902 is formed by downsampling one of the stereo images 906).
[0166] Steps 1006 to 1012 are performed for each of the plurality of planes associated with a corresponding depth within the input image. Each plane may be a depth plane (if the map generated in step 1004 is a depth map) or a disparity plane (if the map generated in step 1004 is a disparity map).
[0167] In step 1006, a depth mask is generated The value of the depth mask indicates whether a region of the input image is within a specified range of the plane. Whether a region is within a specified range of the plane is evaluated by evaluating a differentiable function of the range between the region of the input image and the plane, the plane being determined from the estimation map and combined with the above Figure 4 described.
[0168] In step 1008, according to the depth mask I d And the input image I generates a mask image.
[0169] In step 1010, the blur kernel K is used r Refocus the mask image to generate refocused partial images
[0170] In step 1012, the refocused partial image is reconstructed in the partially refocused image I f Upper stack.
[0171] Steps 1006 to 1012 correspond to Figure 4 Therefore, detailed descriptions of these steps will not be repeated here.
[0172] After performing steps 1006 to 1012 for each plane, a refocused image O is generated in step 1014 .
[0173] The refocused image can be directly obtained by using a stacked blur depth mask M f For the final stacked image I f Normalization is performed to generate, for example, a refocused image 704 as in pipeline 700. When the final stacked image I f This approach can be used when the resolution of is equal to that of a pair of images in a tuple (e.g., as in pipeline 700).
[0174] Alternatively, the refocused image can be generated as follows: first use a stacked blurred depth mask M f For the final stacked image I f Normalization is performed to generate a refocused image with reduced resolution, which is then upsampled to generate a refocused image O. This is the case, for example, in pipeline 900, where the final stacked image I f The image 916 is of reduced resolution compared to the image 906 and is used to generate a reduced resolution refocused image 916. The image 916 is then upsampled to generate a refocused image 920.
[0175] In step 1016, the difference between the refocused image and the training image of the tuple is estimated. The estimated difference is then fed back to the depth estimation stage and used to adjust the image processing model (step 1018).
[0176] The above examples illustrate how a differentiable function can be used to generate a refocused image from an input image to produce a fully learnable pipeline for performing refocusing.
[0177] Fig.11 Another exemplary architecture is shown for computing a refocused image from an input image to form a learnable refocusing pipeline. The pipeline 1100 includes a depth processing stage 1102 and a refocusing stage 1104 implemented using an image processing model. The depth processing stage 1102 can be performed by the depth estimation module 302, and the refocusing stage 1104 can be performed by the refocusing module 304.
[0178] The operations of pipeline 1100 are combined Fig.12 Flowchart description of the .
[0179] In step 1202, a plurality of image tuples are received, for example by the depth estimation module 302. Each image tuple includes a pair of images representing a common scene, and a refocused training image, which may be a ground truth image or the like. Each pair of images may be a stereo image including a 'left image' and a 'right image'. An example of a pair of stereo images is shown at 1106. The training images may be all-focus images, but not necessarily all-focus images.
[0180] Steps 1204 to 1218 are performed for each image tuple.
[0181] Steps 1204 to 1208 form part of an image processing model for processing the pair of images 1106 as part of the depth processing stage 1102. In other words, the depth processing stage 1102 is performed by processing the pair of images 1106 according to steps 1204 to 1208 using the image processing model.
[0182] In step 1204, a pair of images 1106 are input into a feature extraction network 1108. The feature extraction network 1108 is a computational neural network that extracts features from the images. These features may be extracted at a reduced resolution compared to the resolution of the input images 1106.
[0183] In step 1206, the extracted features are compared for a plurality of disparity planes. To this end, features extracted from one image of a pair of images (in this example, the left image) are transferred to each of a plurality of disparity planes and then compared to features extracted from the other image of the pair of images (in this example, the right image). The comparison is performed by computing the difference between the transferred features of one of the images and the non-transferred features of the other image. This may be done by computing the difference (e.g., value) between corresponding blocks of the transferred features and the non-transferred features. The computed differences for each disparity plane form a corresponding strip of the cost volume 1110. The strips of the cost volume are in Fig.11 Each strip of the cost volume represents the (reduced resolution) disparity information of the input pair of images 1106. Therefore, each strip of the cost volume is associated with a corresponding disparity plane.
[0184] In step 1208, a mask M is generated for each of the plurality of planes d associated with a corresponding depth within one of the pair of images 1106 by upsampling the corresponding strip of the cost volume 1110. d In other words, each strip of the cost volume is upsampled to form the corresponding mask M d , the mask is associated with plane d. Upsampling is shown at 1112.
[0185] The plane associated with the mask may be a disparity plane or a depth plane. The mask associated with the corresponding disparity plane may be generated directly by upsampling the corresponding strip of the cost volume. To generate the mask associated with the depth plane, the cost volume is upsampled, and the resulting mask (associated with the disparity plane) is further processed to generate a mask associated with the corresponding depth plane. The mask created by upsampling the corresponding strip of the cost volume is shown at 1114.
[0186] Then, the calculated mask M d is used in the refocusing stage 1104 of the pipeline. Therefore, it is different from using the depth mask M generated from the estimated depth / disparity map in the refocusing stage. d Compared to the above method, in this method, a mask M created by upsampling a strip of the cost volume is used in the refocusing stage. d In other words, the refocusing stage is performed using the cost volume instead of the estimated depth / disparity map D. This has the benefit of reducing the number of processing steps in the pipeline compared to the pipeline for estimating depth / disparity maps.
[0187] The refocusing stage of pipeline 1104 is performed by executing steps 1210 to 1214. These steps are performed using the mask 1114 generated in step 1208 and an input image I. In this example, the input image is one of the pair of images in the tuple, namely one of the pair of images 1106.
[0188] Steps 1210-1212 are performed for each of the plurality of planes associated with the generated mask. As described above, these planes may be disparity planes or depth planes.
[0189] In step 1210, based on the input image and the mask M generated for the plane d Generate a mask image of plane d. The mask image is denoted as I d The mask image may be generated by multiplying the mask with the input image I. The mask image is shown at 1116 in image 11.
[0190] In step 1212, the blur kernel k is used. r Refocusing mask image I d , to generate the refocused partial images Mask M d Also uses blur kernel refocusing to generate blur mask The refocused partial image and blur mask are shown at 1118 .
[0191] Then, the refocused partial images are Superimposed on the partially reconstructed refocused image I f superior.
[0192] After performing steps 1210 and 1212 for each plane, a refocused image O is generated in step 1214 in the manner described above. The refocused image is shown at 1120 .
[0193] In step 1216, the difference between the refocused image 1120 and the training image of the tuple is estimated. The estimated difference is then fed back to the depth estimation stage and used to adjust the image processing model (step 1218).
[0194] This article describes differentiable methods for refocusing input images that enable the formation of pipelines for refocusing input images that can be learned through machine learning techniques. In some examples of this article, the pipeline generates a depth map / disparity map between the depth estimation stage and the refocusing stage, so that the estimated depth map can be improved by suitable deep learning techniques. The methods and pipelines described in this article can be made computationally efficient and easy to handle, so that they can be deployed in various computing architectures (such as smartphone architectures). In addition, the method of refocusing images described in this article produces a bokeh driven by the depth information of the image, which can provide higher quality results than other types of blur (such as Gaussian blur). The technology described in this article also enables the refocused image to be presented at any equivalent aperture size or focal plane, for example by appropriately setting the parameters in equation (1).
[0195] Although this article has described an example of implementing the refocusing stage within a pipeline that uses StereoNet to perform depth estimation, it should be understood that this is for illustrative purposes only, and the refocusing method described herein is not dependent on the specific type of depth estimation being used. It can be seen that the refocusing stage described herein can be implemented with a variety of deep learning techniques to form a learnable pipeline and is not limited to implementation with a StereoNet pipeline.
[0196] The image processing device described herein can be used to perform any method described herein. Generally, any of the above methods, techniques or components can be implemented in software, hardware or any combination thereof. The terms "module" and "unit" can generally be used to represent software, hardware or any combination thereof. In the case of software implementation, a module or unit represents a program code that performs a specified task when executed in a processor. Any algorithm and method described herein can be executed by one or more processors, and the one or more processors execute the code that enables one or more processors to execute the algorithm / method. The code can be stored in a non-transient computer-readable storage medium, such as a random-access memory (RAM), a read-only memory (ROM), a CD, a hard disk storage and other storage devices. The code can be any suitable executable code, such as a computer program code or a computer-readable instruction.
[0197] Applicants hereby disclose separately each individual feature described herein and any combination of two or more such features. With the common knowledge of those skilled in the art, such features or combinations can be implemented as a whole according to this specification, without considering whether such features or combinations of features can solve any problem disclosed herein. In view of the above description, it will be obvious to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. An image processing device, comprising a processor, characterized in that: The processor is configured to generate a refocused image according to an input image and a map indicating depth information of the input image by the following steps: For each plane of a plurality of planes associated with a corresponding depth within the image: Generate a depth mask, the value of the depth mask indicating whether the area of the input image is within a specified range of the plane, wherein the map indicating the depth information of the image is a disparity map or a depth map, and evaluate whether the differentiable function of the range between the area of the input image and the plane determined according to the map is satisfied by evaluating The conditions, among which, is a block in the input image The value of the disparity map or depth map of the image area at is the disparity or depth of the plane, is the size of the aperture, when the condition is met, it is determined that the image area is within the specified range of the plane; when the condition is not met, it is determined that the image area is not within the specified range of the plane, the function varies smoothly between a lower bound and an upper bound defining a range of function values, when the image area being evaluated for the function is within the specified range of the plane, the function has a value within the first half of its range; when the image area being evaluated for the function is not within the specified range of the plane, the function has a value within the second half of its range; generating a mask image based on the input image and the generated depth mask; refocusing the mask image using a blur kernel to generate a refocused partial image; The refocused image is generated from the plurality of refocused partial images.
2. The device according to claim 1, characterized in that The plurality of planes define a depth range, and the input image has a depth of field that spans the depth range of the plurality of planes.
3. The device according to claim 1 or 2, characterized in that: (i) when the map indicating the depth information of the image is a disparity map, each of the multiple planes is a plane at a corresponding disparity value; or (ii) when the map indicating the depth information of the image is a depth map, each of the multiple planes is a plane at a corresponding depth.
4. The device according to any one of claims 1 to 3, characterized in that The depth mask specifies a value for each block of one or more pixels of the input image.
5. The device according to any one of claims 1 to 4, characterized in that The blur kernel is a radial kernel whose radius is a function of the aperture size and the depth or disparity difference between the plane and the focal plane.
6. The device according to any one of claims 1 to 5, characterized in that The generated mask image contains only those regions of the input image that are determined to be within the specified range of the plane.
7. A method for training an image processing model, characterized in that: The method comprises: receiving a plurality of image tuples, each image tuple comprising a pair of images representing a common scene and a refocused training image of the scene; For each image tuple: (i) processing the pair of images using the image processing model to estimate a map indicative of depth information of the scene; (ii) for each plane of a plurality of planes associated with a corresponding depth within the image: Generate a depth mask, the value of the depth mask indicating whether the area of the input image is within a specified range of the plane, wherein the map indicating the depth information of the image is a disparity map or a depth map, and evaluate whether the differentiable function of the range between the area of the input image and the plane determined according to the map is satisfied by evaluating The conditions, among which, is a block in the input image The value of the disparity map or depth map of the image area at is the disparity or depth of the plane, is the size of the aperture, when the condition is met, it is determined that the image area is within the specified range of the plane; when the condition is not met, it is determined that the image area is not within the specified range of the plane, the function varies smoothly between a lower bound and an upper bound defining a range of function values, when the image area being evaluated for the function is within the specified range of the plane, the function has a value within the first half of its range; when the image area being evaluated for the function is not within the specified range of the plane, the function has a value within the second half of its range; generating a mask image based on the input image and the generated depth mask; refocusing the mask image using a blur kernel to generate a refocused partial image; (iii) generating a refocused image based on the plurality of refocused partial images; (iv) estimating a difference between the generated refocused image and the refocused training image; (v) adjusting the image processing model based on the estimated difference.
8. The method according to claim 7, characterized in that The partial image has the same resolution as the pair of images, and the input image is one of the pair of images.
9. The method according to claim 7, characterized in that: The partial image has a reduced resolution compared to the pair of images, and the input image is formed as a reduced resolution image of one of the pair of images.
10. The method according to claim 9, characterized in that The step (iii) of generating the refocused image is performed using the resolution reduction map and further comprises the following steps: generating a refocused image with reduced resolution from the plurality of refocused partial images; The reduced-resolution refocused image is upsampled to generate the refocused image.
11. A method for adjusting an image processing model, characterized in that: The method comprises: receiving a plurality of image tuples, each image tuple comprising a pair of images representing a common scene and a refocused training image of the scene; For each image tuple: Processing the pair of images using an image processing model includes: (i) processing the pair of images using a computational neural network to extract features from the pair of images; (ii) comparing the features with respect to a plurality of disparity planes to generate a cost volume; (iii) generating a mask for each of a plurality of planes associated with a corresponding depth within the image based on a corresponding upsampled strip of the cost volume; Perform the refocusing phase, including: (iv) for each plane of the plurality of planes: Generate a depth mask, the value of the depth mask indicating whether the area of the input image is within a specified range of the plane, wherein the map indicating the depth information of the image is a disparity map or a depth map, and evaluate whether the differentiable function of the range between the area of the input image and the plane determined according to the map is satisfied by evaluating The conditions, among which, is a block in the input image The value of the disparity map or depth map of the image area at is the disparity or depth of the plane, is the size of the aperture, when the condition is met, it is determined that the image area is within the specified range of the plane; when the condition is not met, it is determined that the image area is not within the specified range of the plane, the function varies smoothly between a lower bound and an upper bound defining a range of function values, when the image area being evaluated for the function is within the specified range of the plane, the function has a value within the first half of its range; when the image area being evaluated for the function is not within the specified range of the plane, the function has a value within the second half of its range; generating a mask image based on the input image and the generated mask; refocusing the mask image using a blur kernel to generate a refocused partial image; (v) generating a refocused image based on the plurality of refocused partial images; (vi) estimating a difference between the generated refocused image and the refocused training image; (vii) adjusting the image processing model based on the estimated difference.
12. An image processing device, characterized in that: Comprising a processor and a memory storing instructions in a non-transitory form, the instructions being executable by the processor to implement an image processing model trained by the method according to any one of claims 7 to 10, or adjusted by the method according to claim 11, to generate a refocused image from a pair of input images representing a common scene.
Citation Information
Patent Citations
Depth-based application of image effects
US20170061635A1