Layer merging and post-twist error correction in multi-plane images
The described method optimizes MPI layer merging and error correction in multi-camera scenarios to improve image quality and compatibility, addressing inconsistencies and resolution constraints in MPI rendering.
Patent Information
- Application Number
- CN202380083074.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-15
AI Technical Summary
The existing multi-planar image processing technology is prone to the problems of visual artifacts and insufficient resolution during the generation and decoding process, especially in multi-phase airport scenes, resulting in poor quality of reconstruction images.
Optimize multi-planar image data through layer merging and distortion error correction methods, including layer weight array interpolation, partition optimization, and distortion error minimization, ensuring high-quality reconstruction of images at new view locations.
It effectively reduces the number of multi-planar image layers, improves the spatial resolution and quality of the image, reduces distortion errors, and ensures real-time decoding capabilities on terminal devices.
Smart Images

Figure CN120322792A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Priority Application No. 63 / 429,875, filed on December 2, 2022, the content of which is incorporated herein by reference. Background Art
[0003] Multi - plane image (MPI) scene representations are commonly used for three - dimensional (3D) scene reconstruction. This representation stores the front - parallel planes of a scene at discrete sampled fixed depth ranges from a reference coordinate system. The information stored in each plane contains texture (represented by R, G, B values) and transparency (represented by the α (A) channel). This structure is advantageous for synthesizing images from new viewpoints of the scene because it can avoid occlusion - related problems commonly encountered in RGB images. Common MPI - based applications require reliable storage and communication of multi - plane images on a bitstream. Importantly, the reconstructed images on the bitstream decoder side have high quality to ensure a satisfactory user experience. However, imperfect MPI generation sources may result in rendered images containing visual artifacts. Additionally, the resolution constraints of standard video codecs may lead to further quality degradation on the decoder side. Pipelines must be designed to process the generated MPI data before feeding the generated MPI data into the bitstream to ensure that the reconstructed multi - plane images are visually appealing. Summary of the Invention
[0004] Two methods are described herein to address the multiple requirements of an MPI processing pipeline. The first method performs layer merging for both single - camera and multi - camera scenes, which optimally reduces the number of MPI layers to meet the spatial resolution constrained by the real - time decoding video capabilities in the end - device. Due to warping and MPI generation errors, the MPI data at multiple camera positions may not be completely consistent. The second method addresses these MPI to mitigate this error inconsistency by modeling and formulating an optimization problem. The algorithm is useful in cases where MPI data from multiple camera positions are fed through a parallel bitstream for 3D scene reconstruction on the decoder side.
[0005] In a first aspect, a method for merging layers of a multi-planar image (MPI) includes steps (1) to (5). Step (1) includes determining a layer weight array by: for each of the D1 image layers of the MPI, each located at a respective one of D1 layer depths, determining a respective one of D1 total weights as the element-wise sum of a weighted factor array for that image layer. Step (2) includes interpolating the layer weight array to produce an interpolated layer weight signal. Step (3) includes dividing the interpolated layer weight signal into D2 segments to produce D2 optimized layer depth intervals, where D2 is less than D1. Step (4) includes, for each optimized layer depth interval, determining (i) the layer depths within the interval among the D1 layer depths that are within the optimized layer depth interval; and (ii) the layer weights within the interval among the D1 total weights that are each associated with a respective one of the layer depths within the interval. Step (5) includes, for each of the D2 optimized layer depth intervals, determining the corresponding output layer depth as the average of the layer depths within the interval, each weighted by the respective layer weight within the interval.
[0006] In a second aspect, a method for reducing warping errors in a multi-planar image (MPI) dataset is disclosed. The dataset includes N reference images of a scene, each reference image having D1 image layers, and each of the N reference images has been captured at a respective one of N camera positions. The method includes steps (1), (2), (3).
[0007] Step (1) includes determining a plurality of warped images of quantity N·N b ·D1 by: for each of the N camera positions, for each of the N b adjacent camera positions other than that camera position among the N camera positions, determining a respective sub-plurality of warped images by: for each of the D1 image layers of the MPI dataset: determining a normalized reference image as the element-wise quotient of (i) the reference image captured from the adjacent camera position among the N reference images and (ii) the disparity map of the reference image; and warping the normalized reference image from the adjacent camera position to that camera position, where N b is less than or equal to (N - 1).
[0008] Step (2) includes, for each of the N reference images, determining N b weighted alpha channel arrays, each array equal to the contribution of a respective one of the N b adjacent camera positions to the alpha channel of the reference image.
[0009] Step (3) includes, for each pixel location among the plurality of pixel locations of the MPI dataset, determining a coefficient array that minimizes the total warping error contributed by each of the N reference images, the total warping error being a function of: (i) N b weighted alpha channel arrays, (ii) D1 weighting factors, each weighting factor being evaluated at a corresponding pixel location among D1 image layers of one of the N reference images and derived from the alpha channel of that one reference image, (iii) a plurality of warped images, and (iv) the texture channels of each of the N reference images, all evaluated at the pixel location; and for each of the D1 image layers of each of the N reference images, updating the value of the texture channel by multiplying the value of the texture channel of the image layer at the pixel location by the elements of the coefficient array corresponding to the image layer of the reference image. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a diagram illustrating an MPI 3D scene representation.
[0011] Figure 2 is an example graph showing weighting factors as a function of depth for example multi-plane images.
[0012] Figure 3 is a graphical illustration of how to perform layer merging in a multi-plane image using partition intervals in one embodiment.
[0013] Figure 4 is a flowchart showing a method for merging layers of a multi-layer image in one embodiment.
[0014] Figure 5 depicts an image showing artifacts generated when cleaning is not applied before and after layer merging.
[0015] Figure 6 is a flowchart showing a full layer merging processing pipeline in a single camera scene in one embodiment.
[0016] Figure 7 is a flowchart illustrating a full layer merging processing pipeline in a multi-camera scene.
[0017] Figure 8 includes (i) a reference image of the scene captured from a reference position and (ii) a synthetic image of the scene captured from an adjacent position and warped to the reference position.
[0018] Figure 9 is a schematic diagram of a layer merger for a single camera in one embodiment.
[0019] Figure 10 is illustrative of what can be achieved in one embodiment byFigure 9 Flowchart of a single-camera MPI layer merging method implemented by a layer merger
[0020] Figure 11 Schematic diagram of a layer merger for multiple cameras in an embodiment
[0021] Figure 12 Illustrates that in an embodiment, it can be achieved through Figure 11 Flowchart of a multi-camera MPI layer merging method implemented by a layer merger
[0022] Figure 13 Schematic diagram of a distortion error mitigator in an embodiment
[0023] Figure 14 Illustrates that in an embodiment, it can be implemented by Figure 11 Flowchart of a distortion error mitigation method implemented by a distortion error mitigator Detailed implementation
[0024] 1. Overview of multi-plane imaging
[0025] In this section, the framework and related processes of MPI are first reviewed. Then, this MPI scenario will be formulated as an optimization problem. Although symbols are introduced inline in the context of this document, Appendix A contains references to all symbols mentioned in this document.
[0026] 1.1 MPI representation
[0027] Figure 1 Schematic diagram of a multi-plane image 100. The MPI image 100 has D layers of images, which can be RGB A images. The farthest layer is layer 0, and layer D - 1 is the layer closest to the reference camera position. The RGB value of the p-th layer at the camera position s (e.g., the viewpoint) is denoted as Its dimension is H×W×3. The pixel value at (x, y) for color channel c is denoted as The α value of the p-th layer is Its dimension is H×W. The pixel value at (x, y) is denoted as The depth distance between the p-th layer and the reference camera position is d p . The image from the original reference view (without moving the camera) is denoted as R, with texture pixel values R(x, y, c). It should be noted that in MPI, the distance between two adjacent layers has a fixed equal interval. Given a set of multi-view images or single-view images according to the selected algorithm, the algorithm will output a multi-plane image that contains p = 0,..., D - 1}.
[0028] 1.2 MPI rendering
[0029] Given a set of multi - plane images p = 0, …, D - 1}, which are used to render images (for virtual views) at a new camera position. For new view rendering, there are two main processes: warping and compositing.
[0030] 1.2.1 Warping
[0031] Each layer needs to be warped from the current viewpoint position (ν s ) to the new viewpoint position (Vt).
[0032]
[0033] The warping function is expressed as
[0034]
[0035] where v s =(u s , v s ), v t =(u t , v t ). Ks and Kt are the intrinsic camera models of the reference view and the target view respectively. R and t are the extrinsic camera models of rotation and translation, n is the normal vector [0 0 1] T , a is the distance from the plane parallel to the front of the source camera at a depth σd p .
[0036] 1.2.2 Compositing
[0037] A new view can be rendered from those warped textures and α - images via the following formula.
[0038]
[0039] The disparity map in the source view can be calculated as:
[0040]
[0041] A view can also be rendered at the original reference camera position. In this case, the warping process can be skipped, and the weighting factor can be calculated as
[0042]
[0043] For ease of discussion, the conversion from the α - channel to the weighting factor is expressed as the following function.
[0044]
[0045] The rendering process is further represented by the following function:
[0046]
[0047] It should be noted that if the weighting factor is known, the first step can be skipped.
[0048] 2. Multi-plane Image Layer Merging
[0049] In this section, the layer merging problem is formulated and various methods for solving this problem are considered. The motivation for this problem comes from the fact that the initially generated multi-plane image requires a storage capacity larger than what the current state-of-the-art in the market can store. The generated MPI can consist of D depth layers, where a typical value of D is 32, and the typical resolution of each depth image is 1080p. For these specifications, reconstructing the depth information MPI would require storing an image with a resolution of 16k, which is much higher than what most current market products support. The task of this module is to perform layer merging according to the device's support capabilities to reduce the number of layers. It is important to do so in a way that preserves the quality of the MPI as much as possible.
[0050] There are two versions of the problem considered: (1) single-camera transmission scenario and (2) multi-camera transmission scenario. In the single-camera transmission scenario, only one MPI is fed through the bitstream. The goal in this case is to optimize the merging of the layers of the original MPI such that the quality of the MPI is preserved after local distortion. In the multi-camera transmission scenario, multiple MPIs captured at different camera positions are encoded into a compressed bitstream. The information in these MPIs is jointly used to generate a new global view between the original camera positions.
[0051] 2.1 Problem Formulation
[0052] 2.1.1 Single-Camera Problem Formulation
[0053] Let p = 0, …, D - 1} be the original multi-plane images provided by the MPI generation source. Let D’ < D denote the number of output layers that need to be had after applying layer merging. Let p′ = 0, …, D’ - 1} denote the multi-plane images obtained through layer merging, and let t be any new view to which it is to be distorted. Then the goal of layer merging in the single-camera problem formulation is to solve the following minimization problem.
[0054]
[0055] In short, the problem statement is to merge the MPI layers in such a way that when the merged MPI is warped to any new view position, the resulting rendered image should be similar to the image rendered by the original MPI warped to that position. If there are an infinite number of new poses t to be tested, directly solving this optimization problem is computationally infeasible. Two solutions are provided for this problem.
[0056] 2.1.2 Naive Solution to the Single-Camera Problem
[0057] The first solution proposed for this problem works by simply combining adjacent layers consecutively. Since depth information is not considered when trying to solve the problem, this solution is called the "naive approach".
[0058] Let D’ < D be a number such that mod(D, D’) = 0. Also, let Then, the proposed solution is given by:
[0059]
[0060] This solution works effectively by uniformly interpolating the original sampled depths. In addition to not considering depth information when merging consecutive layers, this method also has the drawback that the number of layers to be merged must be a multiple of D. However, in many cases, it may be desirable to perform fractional downsampling of the number of layers. To address these two drawbacks, a second solution to the problem is provided in the next subsection.
[0061]
[0062] 2.1.2 Interpolation Solution to the Single-Camera Problem
[0063] To address the drawbacks of the naive approach for layer merging, a method is designed as follows: This method uses constrained optimization to select the best depth positions where the merged layers will appear and which layers should be merged to construct the output layers at these depth positions. This method uses interpolation to merge the layers together. Therefore, this solution is called the "interpolation method".
[0064] Each of the original MPI layers contains different amounts of information in its weighting factors, and the amount of information present in each layer typically varies significantly. Figure 2 An example of this is shown in the graph of the interpolated layer weight signal q (s) (z) defined in equations (11) and (12). The total weighting factor information within each of the original 32 layers of the multi-plane image is given by the open circles. At Figure 2In the example, layers 15 to 20 of the MPI contain a large amount of cumulative weights, while layers 25 to 32 contain very few cumulative weights. Intuitively, it is more important to preserve the quality of the layer that contains more information in its weighting factors than the layer that contains little information, because the layer with more weighted information contributes more information to the final rendered image. This observation inspires the following solution.
[0065] The cumulative weights contributing to the weighting factor at depth p can be obtained according to the following formula, where a (s) (p) is the "layer weight" of layer p.
[0066]
[0067] By performing linear interpolation to fill samples between points in the layer weight a (s) the layer weight a (s) is converted into a continuous piecewise linear signal q (s) , as follows. For any point z falling between two depths d p-1 and d p , q (s) can be represented by Equation (12). Although Equation (12) uses linear interpolation, other types of interpolation can also be used without departing from the scope of this article.
[0068]
[0069] Then, the goal becomes to optimally partition q (s) into D' segments, which will be used to generate the MPI for output layer merging. For this purpose, a constrained formula of the Lloyd-Max quantizer is used. Let t0,..., t D’ represent the partition boundaries of q (s) , and r0,..., r D’-1 represent the optimized reconstruction levels of the quantized version of q (s) (which is denoted as q ′(s) ). The original formula of the Lloyd-Max quantizer is solved by minimizing the following cost function:
[0070]
[0071] This cost function indicates that if one wishes to use D' reconstruction levels to quantize q (s) , then the undersampled regions of q (s) should be penalized with a large total weight. By taking partial derivatives with respect to {t p′} and {r p′}, the necessary conditions for the optimal solution can be obtained:
[0072]
[0073] A solution that meets these conditions can be iteratively implemented by alternately updating the reconstruction level {r p′} and the partitioning level {t p′}.
[0074] The problem with this solution is that it cannot guarantee that each partitioning interval contains the depth layers of the original MPI within it. That is, according to this formula, some partitioning intervals may be left empty, which is wasteful. Therefore, this problem is modified to a constrained form to ensure that each partitioning interval contains at least one of the original MPI depths within it. Observing that the distance between the originally generated camera depths is constant, given by the constant a constraint is imposed to ensure that each partitioning interval contains at least the information of one layer from the original MPI within it. These constraints can be incorporated into the Lloyd - Max objective function in the form of hard constraints as follows:
[0075]
[0076] where
[0077]
[0078] χ + represents the positive characteristic function, which reaches infinity if any constraint is violated and equals 0 otherwise. Specifically,
[0079]
[0080] Since the new term in the cost function does not depend on the reconstruction level, the optimality conditions for the reconstruction level remain unchanged. On the other hand, if any partitioning interval becomes too small, the cost function will reach infinity. Therefore, if modifying the partitioning level causes any partitioning interval to become too small, it will be optimal to keep that partitioning level unchanged. If modifying the partitioning level does not violate the constraints, the second term in the cost function becomes 0, and it is optimal to update the partitioning interval according to the previous optimality conditions. In summary, for the constrained Lloyd - Max problem, the following optimality conditions can be obtained:
[0081]
[0082] where The function to implement this iterative process is provided below.
[0083]
[0084] So far, how to determine the partitioning intervals for layer merging has been discussed, but how to actually generate the MPI information for layer merging has not been addressed. Now an explanation is given. For any partitioning interval [tp′-1 ,t p′ ) includes the following MPI layer information: j = p min (p′), …, p max (p′)}. Here p min (p′) and p max (p′) are such that d pmin(p′) is greater than or equal to the lower limit t p′-1 of d p and the minimum value of d pmax(p′) is less than the upper limit t p′ of d p and the maximum value of d. Update this information according to the following formula:
[0085]
[0086] In Figure 3 a diagram of this process is presented, which illustrates how to perform layer merging using partition intervals. The squares represent the partition levels. The positions of the samples of the layer weights a (s) between two partition levels are used to determine the positions of the new sampling depths {d′ p′}. In this example, the original MPI contains 8 layers. The total weights a (s) of each layer are given by the 8 dark gray dots in the figure. Interpolation is performed on these 8 points to obtain a piecewise linear continuous approximation of the cumulative depth information in this MPI, given by the curve q (s) . After obtaining q (s) , constrained Lloyd-Max optimization is performed. The output depths {d′ p′} and the partition levels {t p′} are represented by light gray circles and squares respectively. It can be observed that between every two consecutive partition levels, there is always a value of the layer weight a (s) , meaning that no partition level is empty. The MPI depth information within each of these partition levels is used to perform layer merging. Thus, each partition level outputs one layer in the output MPI.
[0087] Figure 4 is a flowchart showing method 400 for merging layers of a multi-layer image. Method 400 employs the function single_camera_interpolation described below.
[0088]
[0089] 2.1.3 Full single-camera layer merging processing pipeline
[0090] Regardless of which layer merging method is used, directly applying layer merging without any pre - processing or post - processing will result in significant artifacts in the rendered multi - plane image. Examples of such artifacts can be seen when comparing the naive and interpolation methods with Figure 5 the reference image 510. Figure 5 shows the artifacts caused when cleaning is not applied before and after layer merging. Figure 5 includes the reference image 510, the layer - merged image 520 rendered using the naive layer merging, and the rendered image 530 rendered using interpolation. Therefore, MPI cleaning is performed before and after layer merging to address these issues. Figure 6 is a flowchart 600 showing the full layer - merging processing pipeline in a single - camera scenario.
[0091] 2.1.4 Multi - camera Problem Formulation
[0092] As mentioned before, it may be desirable to transmit multiple MPIs and use them to jointly construct a new global view of the scene. The problem setup for this case is formulated in this section.
[0093] Let be the original set of multi - plane images provided by the MPI generation source at N camera positions. Let D’ < D denote the number of output depths required after applying layer merging to each MPI. Let
[0094] denote the multi - plane image obtained through layer merging. Let t be an arbitrary new view to which the images are to be warped. Let T t be the set of all neighboring cameras at position t. The rendered image and the α - channel of the camera s’ at the new view position are given by and respectively in the case of the original MPI, and by and respectively in the case of the layer - merged MPI. Assume that a weighted average of the rendered images of all cameras at the new view position is taken for both the original and layer - merged MPI cases to obtain the final RGB. Specifically, the output rendered RGB images generated by the original and layer - merged MPI at the new view position are given by the following equations:
[0095]
[0096]
[0097] where α (s→t) and α ′(s→t) weight the contribution of the image captured by camera s to the new view t according to the distance between camera s and the new view t. Specifically,
[0098]
[0099] where f represents the camera focal length, and p s and p t represent the poses of camera s and the new view t. Also in equations (21) and (22), d min is the distance d p , where p = (D - 1), and d' min is the distance d' p , where p = (D' - 1). These equations are based on Equation 8 of Mildenhall et al. [2], and for the sake of brevity, the symbol B (s→t) is used to represent the following weighting factor:
[0100]
[0101] Then the goal of layer merging in the multi-camera problem formulation is to solve the following minimization problem.
[0102]
[0103] In short, this problem states the desire to merge the MPI layers in such a way that when they are merged to generate a rendered image at the new view, the resulting rendered image should be similar to the image generated by the original MPI. Given an infinite number of new poses t that will need to be tested, directly solving this optimization problem is computationally infeasible. Therefore, we use the insights of the problem to provide two solutions to this problem. Since these solutions are a summary of the single-camera algorithms presented in the previous two sections, only one section will be used for their formulation and discussion.
[0104] 2.1.5 Multi-Camera Layer Merging Schemes: Naive and Interpolation
[0105] The key to the solutions proposed in the multi-camera case is to ensure depth consistency across cameras when performing layer merging. Specifically, for all cameras s = 0,..., N - 1, the output merged MPI layers are positioned at the same depth spacing {d' p′}.
[0106] The naive method can be easily extended from the single-camera to the multi-camera scenario. When each camera initially has the same depth {d p}. Applying the following equation from the single-camera naive method,
[0107]
[0108] will result in the same output depth {d' p′} for each camera. Therefore, the multi-camera solution of the naive method is given by the following equation:
[0109]
[0110]
[0111] In contrast, directly applying the single-camera interpolation method to each of multiple cameras will not produce the same set of depths {d′ p} for each camera. This is because the choice of the output depth provided for a single camera depends on the cumulative weight information in the weighting factors for each layer, which can be different for each camera. Specifically, q (s) can vary according to s. To address this issue, a global cumulative weighting signal q is selected, which will be used to obtain the set of output depths {d′ p} for all camera positions according to Equation (27).
[0112]
[0113] q represents the cumulative weight over all cameras. This means that when the constrained Lloyd-Max optimizer decides how to create the uniform depth partition, the weighting factors at each depth (and for each camera) will be considered. An overview of the algorithm is provided in the function multi_camera_intumbution, where the key step change is replacing q (s) with q in step - 5 below.
[0114]
[0115]
[0116] 2.1.6 Full multi-camera layer merging processing pipeline
[0117] Similar to the single-camera scenario, in the multi-camera scenario, directly applying layer merging without any preprocessing or postprocessing will result in significant artifacts in the rendered multi-plane image. Therefore, a similar processing pipeline is applied in the multi-camera scenario, where the multi-plane image is cleaned before and after layer merging. Figure 7 is a block diagram illustrating the process.
[0118] 4. Multi-plane image multi-camera post-warp clean
[0119] In this section, the multi-camera post-warping cleaning problem is formulated as a regression problem, and a solution to this problem is provided. The motivation for this problem comes from the fact that the MPI generation source may be defective. This results in inconsistent MPI data being generated, including inaccuracies in the warping function. When synthesizing new views from multiple different cameras using Equation (20), the resulting inaccuracies degrade the quality of the final rendered image. Figure 8 An example is shown where the synthesized new view image 820 captured from the second position and warped to the reference position appears blurry compared to the reference image 810 captured from a reference position adjacent to the second position.
[0120] The MPI generation source provides a grid of MPI cameras that can be used to generate new views. If the MPI information at a particular camera position is warped to one of its neighboring cameras in the grid, artifacts similar to those seen at other new view positions are observed. Fortunately, each neighbor in the grid contains a reference image. Therefore, by mapping the MPI at the current camera position to the adjacent camera positions, the error between the synthesized image and the reference image can be calculated. Our problem formulation proposes propagating this error information back to the original camera position so that the MPI information at that position can be corrected. By correcting the warping problems at multiple local new view positions, the model may reduce the warping errors at other adjacent new view positions.
[0121] 3.1 Problem Formulation
[0122] Let \(t\) be the new view position that contains the reference image \(R\) (t) Suppose an image at position \(t\) is synthesized using Equation (20) with MPI information from neighboring cameras in the MPI grid Let \(T\) t be the set of all neighboring cameras of position \(t\). Then, the error between the reference image and the synthesized image is given by:
[0123]
[0124] where the second expression holds because at each pixel position and It is observed that each term in the double summation in this equation represents the error contributed by each layer of the MPI of a single neighboring camera to the overall deviation at the new view of the pixel \((x,y)\). This observation motivates two views of the problem:
[0125] 1. Viewpoint 1 : The goal is to minimize the squared error received from its neighbors at the new view position for all new views that contain the reference image.
[0126] 2. Viewpoint 2: The goal is to minimize the sum of the squares of the errors contributed by each camera pair to each adjacent new view containing the reference image.
[0127] Both viewpoints aim to solve the same problem, but the presentation of Viewpoint 2 will be focused on.
[0128] Let denote the warping from position t to position s. Let The following equation represents the error contributed by camera s to camera t:
[0129] B (t→s) (x,y)(R (t→s) (x,y)-I (s) (x,y)). (29)
[0130] Let T s be the set of all adjacent cameras to camera s. Then, the total error contributed by camera s to all its neighbors can be obtained by summing over t:
[0131]
[0132] Let N denote the total number of cameras in the grid generated by the MPI source. Then, the total error contributed by all cameras in the grid to all their neighbors can be obtained by summing the previous equation over s:
[0133]
[0134] Observe that R (t→s) and I (s) are rendering forms. Equation (31) can be decomposed into the sum of products of the texture at each depth position and a weighting factor:
[0135]
[0136] The summations can be combined for p to obtain:
[0137]
[0138] This provides a useful form for formulating the linear regression problem in the next section. The aim is to design a solution that makes this equation close to zero.
[0139] 3.1 Problem solution via optimization
[0140] Our solution assumes that only the texture channels of the MPI of any camera are allowed to be modified linearly to improve the quality of the synthesized new views (however, this assumption can be generalized to non-linear forms of regression). Thus, by setting Denote the coefficients for modifying the texture information of the \(p\)-th depth of camera \(s\) at pixel \((x,y)\). The goal is to solve the following minimization problem:
[0141]
[0142] It should be noted that the terms inside the square of the minimization problem can be reorganized as follows:
[0143]
[0144] In the last two expressions of Equation (35), \(\odot\) represents the Hadamard product, and 1N and 1D are vectors of length \(N\) and \(D\) containing entries of 1. \(F\), \(G\), and \(\beta\) are matrices, where and Each of the matrices \(F\), \(G\), and \(\beta\) has dimension \(N\times D\), where \(N\) is the number of cameras and \(D\) is the number of depth planes. This allows the previous optimization problem to be converted into a simple element-wise least squares problem given by:
[0145]
[0146] The solution is a per-pixel solution, meaning the matrix will be used to update the given pixel position \((x,y)\) over all \(N\) cameras and \(D\) depths. Specifically, \(\beta\) is applied to the depth of each camera at the fixed pixel position \((x,y)\) through point-wise multiplication. Therefore, the matrix \(\beta\) must be solved for each pixel position of the image. It should be noted that \(\beta\) is applied uniformly to each texture color channel, which means that, in an embodiment, only one \(\beta\) is solved to update all three texture color channels. The solution to the problem can be found by taking the gradient of the expression in the minimization problem, setting it to 0, and solving for the optimal coefficients:
[0147] \(G\odot(F - G\odot\beta)=0\)
[0148] \(G\odot F = G\odot G\odot\beta\)
[0149]
[0150] where represents the Hadamard division operator.
[0151]
[0152] Figure 9Schematic diagram of an image layer combiner 900, which performs the function single_camera_interpolation (single-camera interpolation) introduced in the description of method 400. The image layer combiner 900 includes a processor 986 and a memory 904. The memory 904 stores a multi-plane image 910, software 920, an intermediate output 930, and an in-range output 940.
[0153] The memory 904 may include one or more types of data storage devices. The memory 904 may be temporary and / or non-temporary, and may include volatile memory (e.g., SRAM, DRAM, computational RAM, other volatile memory, or any combination thereof) and non-volatile memory (e.g., flash memory, ROM, magnetic media, optical media, other non-volatile memory, or any combination thereof), either one or both. A part or all of the memory 904 may be integrated into the processor 986.
[0154] The multi-plane image 910 includes D1 image layers 912(0, 1, 2, ..., (D1 - 1)), and each image layer is located at a corresponding one of the D1 layer depths 914(0, 1, 2, ..., (D1 - 1)), where D1 is a positive integer greater than 1. D1 is an example of D in the function α_to_weight (α to weight) and subsequent functions. The multi-plane image 910 is Figure 1 an example of the multi-plane image 100. The memory 904 stores non-transitory computer-readable instructions as software 920. When executed by the processor 986, the software 920 causes the processor 986 to implement the functions of the image layer combiner 900 as described herein. The software 920 may be firmware or include firmware.
[0155] The software 920 includes a transparency-to-weight converter 921, a weight array generator 922, an interpolator 923, a partitioner 925, a depth selector 926, and a weighted averager 929. The intermediate output 930 includes a weighted factor array 931, a layer weight array 932, an interpolated layer weight signal 933, an optimized layer depth interval 935, an in-range layer depth 936, and an in-range layer weight 937. The layer weight array 932 includes D1 elements.
[0156] The in-range output 940 includes in-range layer depths 941(0, 1, 2, ..., (D2 - 1)), and may also include at least one of in-range texture channels 942(0, 1, 2, ..., (D2 - 1)) and in-range α channels 943(0, 1, 2, ..., (D2 - 1)). D2 represents the number of output depths, is a positive integer less than D1, and may be stored in the memory 904. D2 is an example of D' in the function single_camera_naive (single-camera naive) and subsequent functions.
[0157] Examples of the weighting factor array 931, the layer weight array 932, and the interpolated layer weight signal 933 are, respectively: W of the function α_to_weight p (x, y), the layer weight a of Equation (11) (s) (p), and q of Equation (12) (s) (z). The partition interval is an example of the optimized layer depth interval 935, where p′ = 1, …, (D'-1), as in the function constrained_Lloyd_Max (constrained Lloyd_Max).
[0158] An example of the in - interval layer depth 936 was discussed before Equation (19). The layer depth 936 is, for example, the depth d of the value of the index j between j = pmin(p') and j = pmax(p'). j . For a given optimized layer depth interval 935, the in - interval layer depth 936 is the layer depth 914 within this optimized layer depth interval 935. Similarly, the in - interval layer weight 937 is the value of the layer weight array 932 corresponding to one of the in - interval layer depths 936.
[0159] The expressions of Equations (19a), (19b), and (19c) are examples of the in - interval layer depth 941, the in - interval texture channel 942, and the in - interval α channel 943, respectively.
[0160] Figure 10 is a flowchart showing a method 1000 for merging layers of a multi - plane image (MPI). The method 1000 can be implemented within one or more aspects of the image layer combiner 900. In an embodiment, the method 1000 is implemented by a processor 986 executing computer - readable instructions of software 920. The method 1000 includes steps 1010, 1020, 1030, 1040, 1050, and 1060. The method 1000 may also include step 1010. In an implementation, there is at least one of the following: (a) steps 1010, 1020, and 1030 respectively correspond to steps 1 - 3 of the function single_camera_interpolation; (b) steps 1040 and 1050 correspond to step 4 of the function single_camera_interpolation, and (c) step 1060 corresponds to step 5 of the function single_camera_interpolation.
[0161] Step 1020 includes determining a layer weight array by, for each of D1 image layers of the MPI each located at a respective one of D1 layer depths, determining a respective one of D1 total weights as the element-wise sum of the weighted factor array of that image layer. In an example of step 1020, the weight array generator 922 determines the layer weight array 932.
[0162] When method 1000 includes step 1010, step 1010 is before step 1020. Step 1010 includes, for each of the D1 image layers, determining the weighted factor array of that image layer based on (i) the alpha channel of that image layer, and (ii) the alpha channels of the image layers among the D1 image layers whose layer depth is less than that of that layer. In an example of step 1010, the weight converter 921 executes equation (7) to determine the weighted factor array 931.
[0163] Step 1010 may include step 1012, which includes, for each pixel coordinate of a layer, determining a respective per-pixel weighted factor based on (i) the value of the alpha channel at that pixel coordinate and (ii) the value of the alpha channel at that pixel coordinate. In an example of step 1012, the weight converter 921 determines a respective pixel weighted factor W p (x, y) according to the function alpha_to_weight described in section 1.2.2 for each pixel coordinate (x, y).
[0164] Step 1030 includes interpolating the layer weight signal to produce an interpolated layer weight signal. The interpolation may be linear interpolation. In an example of step 1030, the interpolator 923 interpolates the layer weight array 932 to produce the interpolated layer weight signal 933.
[0165] Step 1040 includes dividing the interpolated layer weight signal into D2 segments to obtain D2 optimized layer depth intervals. D2 is less than D1. In an example of step 1040, the partitioner 925 divides the interpolated layer weight signal 933 into D2 segments to produce the optimized layer depth intervals 935(0, 1, 2,..., (D2 - 1)). Step 1040 may include minimizing an objective function, such as J of equation (13), which includes an increasing function of the interpolated layer weight signal. The function may be at least one of monotonically non-decreasing and monotonically increasing (also known as strictly increasing).
[0166] Step 1040 may include at least one of steps 1041 - 1044. In an embodiment, steps 1041 - 1044 respectively correspond to steps 1 - 4 of the function constrained_Lloyd_Max described in section 2.1.2. The partitioner 925 may execute each of steps 1041 - 1044.
[0167] Apply steps 1042 - 1044 in the case where each pair of adjacent layer depths in D1 layer depths are separated by the same initial distance. This same initial distance can be generated by step 1041, which includes initializing D2 layer depth intervals such that each of the D2 layer depth intervals has an equal length.
[0168] Step 1042 includes determining D2 weighted depths by: for each of the D2 layer depth intervals each having a respective one of a plurality of lower boundaries, determining the respective one of the D2 weighted depths to be the average of the layer depth values within the layer depth interval weighted by the value of the interpolation layer weight signal within the layer depth interval. In step - 2 of constrained_Lloyd_Max, each value of r p′ is the corresponding weighted depth.
[0169] Step 1043 includes updating the D2 layer depth intervals by: for each of the D2 layer depth intervals, if the initial distance is less than the difference between (i) the average R of the associated weighted depth and the subsequent weighted depth of the layer depth interval and (ii) the lower boundary of the layer depth interval, changing the value of the lower boundary to be equal to the average R.
[0170] Step 1044 includes obtaining D2 optimized layer depth intervals by iteratively performing the step of determining the D2 weighted depths and the step of updating the D2 layer depth intervals until the plurality of lower boundaries converge within a predetermined tolerance.
[0171] Step 1050 includes, for each of the D2 optimized layer depth intervals, determining (i) the in - interval layer depths among the D1 layer depths that are within the optimized layer depth interval; and (ii) the in - interval layer weights among the D1 total weights, each in - interval layer weight being associated with a respective one of the in - interval layer depths. In the example of step 1050, depth selector 926 determines, for each optimized layer depth interval 935, (i) the in - interval layer depths 936 among the layer depths 914 that are within layer depth interval 935; and (ii) the in - interval layer weights 937 of layer weight array 932, each associated with a respective in - interval layer depth 936.
[0172] Step 1060 includes step 1061 and may further include at least one of steps 1062 and 1063. Step 1061 includes, for each of the D2 optimized layer depth intervals, determining a corresponding output layer depth as an average of the several intra-interval layer depths weighted by a corresponding one of the several intra-interval layer weights. In an example of step 1061, the weighted averager 929 determines a corresponding intra-interval layer depth 941(p') from the layer depth 936 and the total weight 937 for each optimized layer depth interval 935(p'). In steps 1061 - 1063, p' ranges from zero to (D2 - 1).
[0173] When each of the several intra-interval layer depths is the layer depth of a corresponding one of the several intra-interval image layers in the D1 image layers, step 1060 may include steps 1062 and 1063. Each intra-interval image layer has a corresponding one of the several intra-interval texture channels and a corresponding one of the several intra-interval α channels.
[0174] Step 1062 includes, for each of the D2 optimized layer depth intervals, determining a corresponding texture channel as an average of the several intra-interval texture channels weighted by a corresponding one of the several intra-interval layer weights. An example of a texture channel is that of Equation (19a). In an example of step 1062, the weighted averager 929 determines a corresponding texture channel 942(p') for each optimized layer depth interval 935(p') based on the layer depth 936, the total weight 937, and the texture data of the multi-plane image 910 corresponding to the layer depth 936. In an embodiment, the texture data is C for j values between p min (p') and p max (p'), as in Equation (19b). j (s)
[0175] Step 1063 includes, for each of the D2 optimized layer depth intervals, determining a corresponding α channel as an average of the several intra-interval α channels weighted by a corresponding one of the several intra-interval layer weights. An example of an α channel is that of Equation (19b). In an example of step 1063, the weighted averager 929 determines a corresponding α channel 943(p') for each optimized layer depth interval 935(p') based on the layer depth 936, the total weight 937, and the α channel of the multi-plane image 910 corresponding to the layer depth 936. In an embodiment, each α channel is the corresponding one for j values between p min (p') and p max (p'), as in Equation (19c).
[0176] Figure 11 It is a schematic diagram of the image layer combiner 1100 that executes the function multi_camera_interpolation (multi-camera interpolation) described in Section 2.1.5. The image layer combiner 1100 is Figure 9 an example of the image layer combiner 900 and includes a memory 1104 and software 1120, which are examples of the memory 904 and software 920 respectively. The image layer combiner 1100 determines the output layer depth based on multiple multi-plane images rather than just a multi-plane image.
[0177] Thus, the memory 904 of the image layer combiner 1100 stores N multi-plane images 910(1, 2...., N), and the software 1120 includes an adder 1124. The weight converter 921 generates N sets of weighted factor arrays 1131, each set including D1 weighted factor arrays 931, one array for each layer of the MPI 910. Each weighted factor array is an example of the weighted factor array 931. From each set of weighted factor arrays 1131(n), the weight array generator 922 generates a corresponding set of layer weight arrays 1132(n), where n is a positive integer less than or equal to N. Each layer weight array 1132 is an example of the layer weight array 932.
[0178] The interpolator 923 generates a corresponding one of N interpolated layer weight signals 1133 from each of the N layer weight arrays 1132. Each signal 1133(n) is an example of the signal 933 corresponding to the multi-plane image 910(n). The adder 1124 sums the signals 1133 to produce a global interpolated layer weight signal 1134, which is input to the partitioner 925.
[0179] The image layer combiner 1100 includes an intermediate output 1130, which is an example of the intermediate output 930, which includes the layer weight array 932, the global interpolated layer weight signal 1134, and may include the weighted factor array 931. The intermediate output 1130 also optimizes the layer depth interval 1135, the in-interval layer depth 1136, and the in-interval layer weights 1137(1-N), which are examples of the layer depth interval 935, the in-interval layer depth 936, and the in-interval layer weights 937 respectively. The memory 1104 also stores the in-interval output 1140, which includes (D2×N) in-interval texture channels 1141, (D2×N) in-interval α channels 1142, and D2 in-interval depths 1143, which are examples of the texture channels 942, the α channels 943, and the depth 941 respectively.
[0180] Figure 12is a flowchart showing method 1200 for merging layers of multiple multi - plane images (MPIs) (e.g., N MPIs from N different camera poses of a scene). Method 1200 can be implemented within one or more aspects of image layer combiner 1100. In an embodiment, method 1200 is implemented by a processor 986 executing computer - readable instructions of software 1120. At least one step of method 1200 corresponds to a step of function multi_camera_interpolation introduced after equation (28).
[0181] Method 1200 includes steps 1220, 1230, 1240, 1050, and 1261. Method 1200 may also include at least one of steps 1210, 1262, and 1263. Step 1050 is introduced in the description of method 1000. Steps 1261 - 1263 are corresponding examples of steps 1061 - 1063.
[0182] Step 1210 is an example of step 1010 and, for each of the D1 image layers of each of the N reference images, determines a corresponding one of the D1 weighting factors based on (i) the alpha channel of the image layer and (ii) the alpha channels of the image layers among the D1 image layers whose layer depth is less than that of the layer. In the example of step 1210, weight converter 921 executes equation (7) to determine the set of N arrays of weighting factors 1131.
[0183] Step 1220 includes determining N interpolated - layer weight signals by performing the following steps for each of the N MPIs: (i) step 1020 of method 1000 for determining an array of layer weights, and (ii) step 1030 of method 1000 for linearly interpolating the array of layer weights. In the example of step 1220, weight - array generator 922 generates a corresponding one of the layer - weight arrays 1132(1 - N) from each set of weighting - factor arrays 1131, and interpolator 923 interpolates each of the N layer - weight arrays 1132 to produce a corresponding interpolated - layer weight signal 1133.
[0184] Step 1220 may also include performing step 1010 of method 1000 for each of the N MPIs. In such an embodiment, transparency weight converter 921 generates the set of weighting - factor arrays 1131(1 - N).
[0185] Step 1230 includes obtaining a globally interpolated layer - weight signal as the sum of each of the N interpolated - layer weight signals. In the example of step 1230, adder 1124 performs step - 4 of function multi_camera_intension to generate a globally interpolated layer - weight signal 1134 from the interpolated - layer weight signals 1133.
[0186] Step 1240 includes dividing the global interpolation layer weight signal into D2 segments to produce D2 globally optimized layer depth intervals. D2 is less than D1. Step 1240 is an example of step 1040 and thus may include at least one of steps 1041 - 1044. In an example of step 1240, the partitioner 925 divides the global interpolation layer weight signal 1134 into D2 segments to produce the optimized layer depth intervals 1135(0, 1, 2, ..., (D2 - 1)).
[0187] Step 1050 is introduced in the description of method 1000. In an example of step 1050, the depth selector 926 determines for each optimized layer depth interval 1135 (i) the in - interval layer depth 1136 of the layer depth 914 within the layer depth interval 1135; and (ii) the in - interval layer weights 1137(1 - N) of the layer weight array 1132, with each in - interval layer weight associated with the corresponding in - interval layer depth 1136. For each value of the positive index n < N, each total in - interval weight 1137(n) may include more than one of the D1 weights of the layer weight array 1132(n), depending on the number of layer depths 1136 within each optimized layer depth interval 1135.
[0188] Step 1261 includes, for each of the D2 globally optimized layer depth intervals, determining the corresponding output layer depth as the average of the several in - interval layer depths weighted by the corresponding one of the several in - interval layer weights. In an example of step 1261, the weighted averager 929 determines the corresponding in - interval layer depth 1136(p') for each optimized layer depth interval 1135(p'). In steps 1261 - 1263, the index p' ranges from zero to (D2 - 1).
[0189] Step 1262 includes, for each of the N MPIs, determining the corresponding texture channel by performing step 1062. In an example of step 1262, the weighted averager 929 determines the corresponding texture channel 1142(n, {0, 1, 2,...(D2 - 1)}) for each MPI 910(n).
[0190] Step 1263 includes, for each of the N MPIs, determining the corresponding α channel by performing step 1063. In an example of step 1263, the weighted averager 929 determines the corresponding α channel 1143(n, {0, 1, 2,...(D2 - 1)}) for each MPI 910(n).
[0191] Figure 13Schematic diagram of the distortion error mitigator 1300, which executes the function N_camera_post_warp_clean (N camera post-warp clean) introduced after Equation (38). The distortion error mitigator 1300 includes a processor 1386 and a memory 1304. The memory 1304 stores N multi-plane images 910, software 1320, and an output 1330. The memory 1304 may include one or more types of data storage devices, such as those listed for the memory 904. Part or all of the memory 1304 may be integrated into the processor 1386. Each multi-plane image 910 has D1 image layers 912.
[0192] The software 1320 includes a normalizer 1322, an image warper 1323, an alpha channel array generator 1324, a warping error minimizer 1326, and a texture channel updater 1328. The output 1330 includes a normalized reference image 1332, a warped image 1333, a weighted alpha channel array 1334, a coefficient array 1336, and texture channel values 1338, which are generated by the normalizer 1322, the image warper 1323, the alpha channel array generator 1324, the warping error minimizer 1326, and the texture channel updater 1328, respectively.
[0193] The software 1320 may also include a transparency-to-weight converter 921, introduced as part of the image layer combiner 900. As in the image layer combiner 1100, the transparency-to-weight converter 921 generates a set of N weighted factor arrays 1131.
[0194] An example of the normalized reference image 1332 is which is introduced after Equation (29) and after step -1 of the function N_camera_post_warp_clean. An example of the warped image 1333 is which can be obtained by applying Equation (1) to the texture channel and the transparency channel of the warped reference image to warp the image from position t to position s. The weighted alpha channel array 1334 includes N arrays. An example of the weighted alpha channel array 1334 is B (t→s) which is the B in Equation (24) in the case where the camera positions s and t are switched (t→s) .
[0195] The sizes of both the coefficient array 1336 and the texture channel values 1338 can be N by D1. An example of the coefficient array is β, introduced in Equation (36) and described in subsequent equations and the function N_camera_post_warp_clean. An example of the texture channel values 1338 is that of Equation (35) Where the coordinates (x, y) are pixel positions. In an embodiment, the coefficient array 1336 has N rows and D1 columns such that the element in the j-th row and k-th column of the coefficient array 1336 corresponds to the k-th image layer of the j-th reference image, where j and k are positive integers less than or equal to N and D1, respectively.
[0196] Figure 14 FIG. 14 is a flowchart illustrating a method 1400 for reducing warping errors in an MPI dataset. The MPI dataset includes N reference images of a scene, each reference image having D1 image layers, and each of the N reference images has been captured at a respective one of N camera positions. Method 1400 may be implemented within one or more aspects of the warping error reducer 1300. In an embodiment, method 1400 is implemented by a processor 1386 executing computer-readable instructions of software 1320. Method 1400 includes steps 1410, 1440, 1460, and in an embodiment, also includes step 1210. Step 1410 includes steps 1420 and 1430, and step 1460 includes steps 1470 and 1480. Method 1400 may also include step 1210 introduced in the description of method 1200.
[0197] Step 1410 may be performed by implementing three nested loops: an outer loop, an intermediate loop within the outer loop, and an inner loop within the intermediate loop. The outer loop has N iterations, iterating once for each camera position of the N reference images. The intermediate loop associated with step 1420 has (N - 1) iterations, iterating once for each of the N - 1 reference images except the one for which the outer loop is iterating. The inner loop associated with step 1430 has D1 iterations, each iteration corresponding to a respective image layer of the reference image.
[0198] Step 1410 includes determining a plurality of warped images, numbering N·N b ·D1, by performing step 1420 for each of the N camera positions. N b is a positive integer less than or equal to (N - 1). For example, step 1420 includes determining a respective sub - plurality of warped images by performing step 1430 for each of the D1 image layers in the MPI dataset for N b adjacent camera positions among the N camera positions other than that camera position. In an embodiment, N b equals 8, such that the N b adjacent camera positions are the four side - adjacent and four diagonal - adjacent camera positions relative to the camera position.
[0199] Step 1430 includes, for each of the D1 image layers in the MPI dataset, performing steps 1432 and 1434. Step 1432 includes determining the normalized reference image as the element-wise quotient of (i) the reference images captured from adjacent camera positions among the N reference images and (ii) the disparity map of the reference images. An example of the disparity map is the In an example of step 1432, the normalizer 1322 implements step (1a) of N_camera_post_warp_clean to determine the normalized reference image 1332(1, 2, …, T), where T is equal to N(N - 1)·D1.
[0200] Step 1434 includes warping the normalized reference image from adjacent camera positions to the camera position. In an example of step 1434, the image warper 1323 implements step (1b) of N_camera_post_warp_clean to generate the warped images 1333(1, 2,..., T) from the T normalized images 1332.
[0201] Step 1440 includes, for each of the N reference images, determining N b weighted α-channel arrays. Each of the N b weighted α-channel arrays is equal to the contribution of the α-channel of that reference image to the corresponding one of the N b adjacent camera positions. In an example of step 1440, the α-channel array generator 1324 executes step 2 of N_camera_post_warp_clean to determine (N - 1) weighted α-channel arrays 1334.
[0202] Step 1460 includes, for each of the plurality of pixel positions in the MPI dataset, performing steps 1470 and 1480. Step 1470 includes determining a coefficient array that minimizes the total warping error contributed by each of the N reference images. The total warping error is a function of: (i) the N b weighted α-channel arrays, (ii) D1 weighting factors, each of which is evaluated at a pixel position corresponding to a respective one of the D1 image layers of one of the N reference images and is derived from the α-channel of that one reference image, (iii) the plurality of warped images, and (iv) the texture channels of each of the N reference images, all evaluated at the pixel position. In an example of step 1470, the warping error minimizer 1326 implements step 3 of the function N_camera_post_warp_clean to determine, for each pixel position of the multi-plane image 910, a corresponding coefficient array 1336 from the set of weighting factor arrays 1131, the warped images 1333, and the weighted α-channel arrays 1334.
[0203] Step 1480 includes, for each of the D1 image layers of each of the N reference images, updating the value of the texture channel at a pixel location by multiplying the value of the texture channel of the image layer at the pixel location by an element of a coefficient array corresponding to the image layer of the reference image. In an example of step 1480, the texture channel updater 1328 implements step 4 of the function N_camera_post_warp_clean to update the value of the texture channel of each image layer {0, 1, 2, ..., (D1-1)} of each image 910 to the corresponding texture channel value 1338(1, 2, ..., N).
[0204] Changes may be made to the above methods and systems without departing from the scope of this embodiment. Accordingly, it should be noted that the content included in the above description or shown in the drawings should be construed as illustrative rather than restrictive. In this document, unless otherwise specified, the phrase "in an embodiment" is equivalent to the phrase "in some embodiments" and does not refer to all embodiments. The following claims are intended to cover all general and specific features described herein, as well as all statements of the scope of the method and system, which may be said to fall between them in terms of language.
[0205] Aspects of the present invention can be recognized from the following exemplary embodiments (EEE) listed:
[0206] EEE1. A method for merging layers of a multi-plane image (MPI), comprising:
[0207] Determining a layer weight array {{a (s)}} by: for each of the D1 image layers of the MPI, determining a corresponding one of the D1 total weights as the element-wise sum of a weighted factor array {{W p (x,y)}}, each of the image layers being at a corresponding one of the D1 layer depths;
[0208] Interpolating the layer weight array to produce an interpolated layer weight signal {{q (s) (z)}};
[0209] Dividing the interpolated layer weight signal into D2 segments to produce D2 optimized layer depth intervals {{with edges t p ’}}, where D2 is less than D1;
[0210] For each optimized layer depth interval, determine (i) several in-layer depths within the optimized layer depth interval among the D1 layer depths; and (ii) several in-layer weights among the D1 total weights, each in-layer weight being associated with a corresponding one of the several in-layer depths; and
[0211] For each of the D2 optimized layer depth intervals, determine a corresponding output layer depth {{d′}} as an average of the several in-layer depths weighted by the corresponding one of the several in-layer weights respectively.
[0212] EEE2. The method according to EEE1, further comprising, before determining the layer weight array:
[0213] For each of the D1 image layers, determine a weighted factor array {{W p (x,y)}} according to (i) the alpha channel of the image layer and (ii) the alpha channels of the image layers among the D1 image layers whose layer depths are less than that of this layer.
[0214] EEE3. The method according to EEE2, wherein determining the weighted factor array includes: for each pixel coordinate of the layer, determining a corresponding per-pixel weighted factor according to (i) the value of the alpha channel at the pixel coordinate and (ii) the value of the alpha channel at the pixel coordinate.
[0215] EEE4. The method according to any one of the foregoing EEEs, wherein the partitioning includes minimizing an objective function, the objective function including a monotonically non-decreasing function of the interpolated layer weight signal.
[0216] EEE5. The method according to any one of the foregoing EEEs, wherein each pair of adjacent layer depths among the D1 layer depths are separated by the same initial distance, and the partitioning includes:
[0217] Determine D2 weighted depths {{r p ′}} by: for each of the D2 layer depth intervals each having a corresponding one of a plurality of lower bounds {{t p′}}, determine a corresponding one of the D2 weighted depths as an average of the layer depth values within the layer depth interval weighted by the value of the interpolated layer weight signal within the layer depth interval;
[0218] Update the D2 layer depth intervals by: for each layer depth interval among the D2 layer depth intervals, for the layer depth interval, if the initial distance is less than the difference between (i) the average R of the associated weighted depth and the subsequent weighted depth of the layer depth interval and (ii) the lower bound {{t p′}} of the layer depth interval, change the value of the lower bound to be equal to the average R; and
[0219] Obtain D2 optimized layer depth intervals by iteratively performing the steps of determining the D2 weighted depths and updating the D2 layer depth intervals until the multiple lower bounds converge within a predetermined tolerance.
[0220] EEE6. The method according to EEE5 further includes, before any determination instance, initializing the D2 layer depth intervals such that each of the D2 layer depth intervals has an equal length.
[0221] EEE7. The method according to any one of the foregoing EEEs, wherein each in-layer depth in the several intervals is the layer depth of the corresponding one of the in-layer images in the several intervals of the D1 image layers, and each in-layer image has a corresponding one of the several in-layer texture channels between {{pMin and pMax}} }, and the method further includes:
[0222] For each of the D2 optimized layer depth intervals, determine the corresponding texture channel to be the average value of the several in-layer texture channels weighted by the corresponding one of the several in-layer weights.
[0223] EEE8. The method according to any one of EEE1 to EEE6, wherein each in-layer depth in the several intervals is the layer depth of the corresponding one of the in-layer images in the several intervals of the D1 image layers, and each in-layer image has a corresponding one of the several in-layer α channels between {{pMin and pMax}} }}, and the method further includes:
[0224] For each of the D2 optimized layer depth intervals, determine the corresponding α channel to be the average value of the several in-layer α channels weighted by the corresponding one of the several in-layer weights.
[0225] EEE9. A method for merging the layers of N MPIs from N different camera poses of a scene, including:
[0226] By, for each of the N MPIs, performing (i) the step of determining the layer weight array {{a (s)}} of EEE1 and (ii) the step of interpolating the layer weight array of EEE1, determine N interpolated layer weight signals {{q (s) (z)}};
[0227] Obtain a global interpolation layer weight signal {{q(z)}} as the sum of each of the N interpolation layer weight signals;
[0228] Divide the global interpolation layer weight signal into D2 segments to generate D2 global optimization layer depth intervals {{with margin t p′}}, where D2 is less than D1;
[0229] For each of the D2 global optimization layer depth intervals, determine (i) several in-layer depths among the D1 layer depths that are within the optimization layer depth interval; and (ii) several in-layer weights among the D1 total weights, with each in-layer weight associated with a corresponding one of the several in-layer depths; and
[0230] For each of the D2 global optimization layer depth intervals, determine the corresponding output layer depth {{d’}} as the average of the several in-layer depths weighted by a corresponding one of the several in-layer weights.
[0231] EEE 10. The method according to EEE9 further includes, for each MPI among the N MPIs, by performing the method of EEE7, determining a corresponding texture channel for each of the D2 optimization layer depth intervals.
[0232] EEE11. The method according to EEE9 further includes, for each MPI among the N MPIs, by performing the method of EEE8, determining a corresponding α channel for each of the D2 optimization layer depth intervals.
[0233] EEE 12. A method for reducing distortion errors in a multi-plane image (MPI) dataset, the multi-plane image dataset including N reference images of a scene, each reference image having D1 image layers, and each of the N reference images having been captured at a corresponding one of N camera positions, the method including:
[0234] Determine a plurality of distorted images of quantity N·N b ·D1 by the following operations where N b ≤(N - 1): For each camera position among the N camera positions,
[0235] For each of the N adjacent camera positions among the N camera positions {{t∈T s}} other than the camera position {{s}}, determine the corresponding sub-plurality of distorted images by the following b way
[0236]
[0237] For each image layer in the D1 image layers of the MPI dataset: Determine a normalized reference image For (i) the reference images captured from adjacent camera positions among the N reference images and (ii) the disparity maps of the reference images Of the element-by-element quotient; and
[0238] Warp the normalized reference image from the adjacent camera position to this camera position;
[0239] For each of the N reference images, determine N b Weighted alpha channel arrays {{B (t→s)}}, each array equal to the contribution of the corresponding one of the N b Adjacent camera positions to the alpha channel of the reference image;
[0240] For each pixel position in the plurality of pixel positions of the MPI dataset,
[0241] Determine a coefficient array that minimizes the total warping error contributed by each of the N reference images, the total warping error being a function of: (i) the N b Weighted alpha channel arrays {{B (t→s)}}, (ii) D1 weighting factors Each weighting factor is evaluated at the pixel position corresponding to a respective one of the D1 image layers of one of the N reference images and is derived from the alpha channel of that one reference image, (iii) the plurality of warped images And (iv) the texture channels of each of the N reference images Are all evaluated at the pixel position; and
[0242] For each image layer of the D1 image layers of each of the N reference images, update the value of the texture channel by multiplying the value of the texture channel of the image layer at the pixel position by the element of the coefficient array corresponding to the image layer of the reference image.
[0243] EEE13. The method according to EEE12, further comprising:
[0244] For each image layer of the D1 image layers of each of the N reference images, determine the corresponding one of the D1 weighting factors {{W p (x,y)}} according to (i) the alpha channel of the image layer and (ii) the alpha channels of the image layers whose layer depths in the D1 image layers are less than the layer depth of this layer.
[0245] EEE14. The method according to EEE12 or EEE13, wherein, in the step of determining the coefficient array, the coefficient array has N rows and D1 columns, such that the element in the j-th row and k-th column of the coefficient array corresponds to the k-th image layer of the j-th reference image, where j and k are positive integers less than or equal to N and D1, respectively.
[0246] EEE15. An image layer combiner, comprising:
[0247] a processor; and
[0248] a memory storing a multi-plane image (MPI) and machine-readable instructions that, when executed by the processor, cause the processor to perform the method according to any one of EEE1 to EEE8.
[0249] EEE16. A warping error mitigator, comprising:
[0250] a processor; and
[0251] a memory storing (i) a multi-plane image (MPI) data set including N reference images of a scene, each reference image having D1 image layers, each of the N reference images having been captured at a respective one of N camera positions, and (ii) machine-readable instructions that, when executed by the processor, cause the processor to perform the method according to any one of EEE12 to EEE14.
[0252] Appendix A. Notes
[0253]
Claims
1. A method for merging layers of a multi-plane image (MPI), comprising: Determine the layer weight array {{a (s)}} by the following operations: For each of the D1 image layers of the MPI, determine the corresponding one of the D1 total weights as the element-wise sum of the weighted factor array {{W p (x,y)}}, where each of the image layers is located at a corresponding one of the D1 layer depths; Interpolate the layer weight array to generate an interpolated layer weight signal {{q (s) (z)}}; Divide the interpolation layer weight signal into D2 segments to generate D2 optimized layer depth intervals {{with edge t p’}}, where D2 is less than D1; For each optimized layer depth interval, determining (i) several in-layer depths within the optimized layer depth interval among the D1 layer depths; and (ii) several in-layer weights among the D1 total weights, each in-layer weight being associated with a corresponding one of the several in-layer depths; and For each of the D2 optimized layer depth intervals, determining a corresponding output layer depth {{d′}} as the average of the several in-layer depths weighted by the corresponding one of the several in-layer weights.
2. The method according to claim 1, further comprising, before determining the layer weight array: For each of the D1 image layers, determine a weighted factor array {{W p (x,y)}} based on (i) the alpha channel of the image layer and (ii) the alpha channels of the image layers among the D1 image layers whose layer depths are less than the layer depth of this layer.
3. The method according to claim 2, wherein Determining the weighted factor array includes: for each pixel coordinate of the layer, determining a corresponding per-pixel weighted factor according to (i) the value of the α channel at the pixel coordinate and (ii) the value of the α channel at the pixel coordinate.
4. The method according to any one of the preceding claims, wherein, The partitioning includes minimizing an objective function, which includes a monotonically non-decreasing function of the interpolated layer weight signal.
5. The method according to any one of the preceding claims, wherein, Each pair of adjacent layer depths among the D1 layer depths is separated by the same initial distance, and the partitioning includes: Determine the D2 weighted depths {{r p′}} by the following operations: For each of the D2 layer depth intervals each having a respective one of a plurality of lower bounds {{t p′}}, determine the respective one of the D2 weighted depths to be the average of the layer depth values within the layer depth interval weighted by the value of the interpolation layer weight signal within the layer depth interval; Update the D2 layer depth intervals by the following operations: For each such layer depth interval among the D2 layer depth intervals, change the value of the lower boundary to be equal to the average value R, for which the initial distance is less than the difference between (i) the average value R of the associated weighted depth and the subsequent weighted depth of the layer depth interval and (ii) the lower boundary {{t p′}} of the layer depth interval; and Obtaining D2 optimized layer depth intervals by iteratively performing the step of determining the D2 weighted depths and the step of updating the D2 layer depth intervals until the plurality of lower boundaries converge within a predetermined tolerance.
6. The method according to claim 5, further comprising, before any determination instance, initializing the D2 layer depth intervals such that each of the D2 layer depth intervals has an equal length.
7. The method according to any one of the preceding claims, wherein, Each of the depths of the inner layers of the several intervals is the depth of the corresponding one of the inner layers of the images in the several intervals of the D1 image layers, and each inner layer of the images in the interval has a corresponding one of the several texture channels in the interval {{between pMin and pMax }}, and the method further includes: For each of the D2 optimized layer depth intervals, determine the corresponding texture channels which is the average of the texture channels within the several intervals weighted by the corresponding one of the layer weights within the several intervals.
8. The method according to any one of claims 1-6, wherein Each of the inner-layer depths of the several intervals is the layer depth of the corresponding one of the in-interval image layers in the D1 image layers, and each in-interval image layer has a corresponding one of the several in-interval alpha channels {{A between pMin and pMax p}}, and the method further includes: For each of the D2 optimized layer depth intervals, determine the corresponding alpha channel which is the average of the alpha channels within the several intervals weighted by the corresponding one of the layer weights within the several intervals.
9. A method for merging layers of N MPIs from N different camera poses of a scene, comprising: For each of the N MPIs, performing (i) the step of determining a layer weight array {a (s)} according to claim 1} and (ii) the step of interpolating the layer weight array according to claim 1 to determine N interpolated layer weight signals {q (s) (z)}; Obtaining a global interpolated layer weight signal {{q(z)}} as the sum of each of the N interpolated layer weight signals; Dividing the global interpolation layer weight signal into D2 segments to generate D2 global optimization layer depth intervals {{with edge t p′}}, where D2 is less than D1; For each of the D2 global optimized layer depth intervals, determining (i) several in-layer depths within the optimized layer depth interval among the D1 layer depths; and (ii) several in-layer weights among the D1 total weights, each in-layer weight being associated with a corresponding one of the several in-layer depths; and For each of the D2 global optimized layer depth intervals, determining a corresponding output layer depth {{d’}} as the average of the several in-layer depths weighted by the corresponding one of the several in-layer weights.
10. The method according to claim 9, further comprising, for each of the N MPIs, determining a corresponding texture channel for each of the D2 optimized layer depth intervals by performing the method according to claim 7.
11. The method according to claim 9, further comprising, for each of the N MPIs, determining a corresponding α channel for each of the D2 optimized layer depth intervals by performing the method according to claim 8.
12. A method for reducing warping errors in a multi-plane image (MPI) dataset, the multi-plane image dataset including N reference images of a scene, each reference image having D1 image layers, each of the N reference images having been captured at a respective one of N camera positions, the method comprising: Determine the quantity of N·N through the following operations b ·Multiple distorted images of D1 where N b ≤(N - 1): For each camera position {{s}} among the N camera positions, For the N camera positions {{t ∈ T s}}, except for the camera position {{s}}, the N b adjacent camera positions, determine the corresponding sub-plurality of distorted images in the following manner For each of the D1 image layers in the MPI dataset: Determine a normalized reference image as the element-wise quotient of (i) the reference images captured from adjacent camera positions among the N reference images and (ii) the disparity map of the reference images ; and warping the normalized reference image from the adjacent camera position to this camera position; For each of the N reference images, determine N b weighted alpha channel arrays {{B (t→s)}}, each array being equal to the contribution of the corresponding one of the N b adjacent camera positions to the alpha channel of the reference image; for each pixel position in a plurality of pixel positions of the MPI dataset, Determine a coefficient array {β} that minimizes the total distortion error contributed by each of N reference images, where the total distortion error is a function of: (i) N b weighted α-channel arrays {B (t→s)}, (ii) D1 weighting factors each of which is evaluated at a pixel location corresponding to a respective one of D1 image layers of one of the N reference images and is derived from the α-channel of that one reference image, (iii) a plurality of distorted images and (iv) the texture channels of each of the N reference images each of which is evaluated at the pixel location; and for each image layer of the D1 image layers of each of the N reference images, updating the value of the texture channel of the image layer at the pixel position by multiplying the value of the texture channel of the image layer at the pixel position by an element of a coefficient array corresponding to the image layer of the reference image.
13. The method according to claim 12, further comprising: For each of the D1 image layers of each of the N reference images, determine the corresponding one of the D1 weighting factors {{W p (x,y)}} according to (i) the alpha channel of the image layer and (ii) the alpha channels of the image layers among the D1 image layers whose layer depths are less than the layer depth of this layer.
14. The method according to claim 12 or 13, wherein, in the step of determining the coefficient array, the coefficient array has N rows and D1 columns such that the element in the j-th row and k-th column of the coefficient array corresponds to the k-th image layer of the j-th reference image, where j and k are positive integers less than or equal to N and D1, respectively.
15. An image layer combiner, comprising: a processor; and a memory that stores a multi-plane image (MPI) and machine-readable instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-8.
16. A warping error mitigator, comprising: a processor; and a memory that stores (i) a multi-plane image (MPI) dataset including N reference images of a scene, each reference image having D1 image layers, each of the N reference images having been captured at a respective one of N camera positions, and (ii) machine-readable instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 12-14.