Integrated multiview multiplane image representation
The integrated multiview MPI framework addresses occlusion and artifact issues in MPI generation by using occlusion-guided residuals and inter-layer texture fillers, resulting in accurate and efficient 3D scene representation.
Patent Information
- Application Number
- PCT/US2025/011734
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2025-01-15
- Publication Date
- 2025-07-24
AI Technical Summary
Existing multiplane image (MPI) generation methods face challenges in accurately representing 3D scenes with continuous depth variations, leading to occlusion artifacts and inefficient data transmission requirements, especially when using single-view or multi-view approaches.
An integrated multiview MPI framework that utilizes multiple camera views to generate an MPI representation by constructing occlusion maps, applying occlusion-guided residuals (OGR) to fill occluded areas, and using an inter-layer texture filler to reduce artifacts, while incorporating modules to correct camera pose and disparity map inaccuracies.
The framework produces a high-quality MPI representation that is robust against occlusion and inter-layer artifacts, achieving data-efficient volumetric rendering with improved accuracy and reduced data transmission needs.
Smart Images

Figure IMGF000011_0001 
Figure IMGF000011_0002 
Figure IMGF000011_0003
Abstract
Description
Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP INTEGRATED MULTIVIEW MULTIPLANE IMAGE REPRESENTATION 1. Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from EP application 24159664.2, filed 26 February 2024 and U.S. Provisional Application Ser. No. 63 / 621,204, filed on 16 January 2024, which is incorporated by reference herein in its entirety. 2. Field of the Disclosure
[0002] Various example embodiments relate generally to multiplane images (MPIs) and, more specifically but not exclusively, to generation and rendition of MPIs. 3. Background
[0003] Multiplane images embody a relatively new approach to storing volumetric content. Multiplane imaging can be used to render both still images and video and represents a three- dimensional (3D) scene within a view frustum using, e.g., 8, 16, or 32 planes of texture and transparency (alpha) information per camera. Example applications of MPIs include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS
[0004] Disclosed herein are various embodiments of a method and system for generating an MPI representation that integrates information from multiple views of a scene. In some examples, the method generates such MPI representation based on a set of inputs including the multiple views, the corresponding camera poses, and the depth information (e.g., in the form of a disparity map). The method initializes the MPI based on the source view by decomposing a source view’s pixel into MPI layers and estimating the surface opacities. Based on the surface- opacity estimates, the method operates to construct an occlusion map that indicates the respective occluded regions within each MPI layer. This occlusion map is propagated to other target view positions through homography warping, thereby enabling retrieval of relevant pixels from the other camera views to fill in the occluded areas. Some examples of the method additionally introduce an inter-layer texture filler, which has the learned RGB textures on intermediate disparity between the MPI layers. This texture filler can beneficially be used to mitigate boundary artifacts between the layers in the rendered scene, e.g., when dealing with scenes characterized by continuous disparity. By combining pixels from the source image,Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP occlusion data from the multiple views, and the learned inter-layer texture filler, the method operates to generate RGB textures for the MPI representation that are beneficially robust against occlusion and inter-layer boundary artifacts.
[0005] While the alpha channel plays an important role in shaping the RGB layers of the MPI representation, the reconstructed RGB textures from a previous training iteration also contribute to accurately shaping the surface opacity of the scene. Thus, through training iterations, both the scene opacity (A) and texture (RGB) channels are jointly optimized leading to a more-accurate MPI representation.
[0006] Some embodiments also incorporate modules configured to address potential inaccuracies originating from the input data, such as disparity maps and camera poses. In some examples, a Camera Pose Adjust Net is used to rectify potential errors arising from uncertainties in the camera pose estimations. Furthermore, a low-pass filtering operation is used within the disparity fidelity-based Occlusion Guided Residuals (OGR) combination procedure to reduce potential influence of local inaccuracies in the disparity maps.
[0007] Experimental results indicate that the integrated multiview MPI obtained using the disclosed method and system excels in terms of the source view fidelity and view-interpolation capabilities. Furthermore, in at least some examples, the integrated multiview MPI demonstrates performance comparable to that of a conventional MPI having twice the number of MPI layers, showcasing the integrated multiview MPI’s data-efficient volumetric representation capability.
[0008] According to an example embodiment, a method of generating an MPI representation of a scene comprises: generating a first MPI representation of the scene based on a first image of the scene; constructing an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtaining, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and generating a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
[0009] In some examples, the obtaining comprises, for each image of the one or more second images of the scene: warping the occlusion map to the respective camera pose of the image; determining a first occlusion guided residuals (OGR) set by copying pixel values of the image toDolby Laboratories Licensing Corporation C17491EP0 / D23150EP a region defined by the warped occlusion map; computing a depth fidelity map based on a deviation between (a) a depth representation of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a depth representation map of the image, and constructing a second OGR set by modifying the first set of pixel values by applying depth- based pixel removal based on the computed deviation. For example, the warping may be performed as shown in Figure 14 by a warping block of an OGR generator that warps the occlusion map to various target camera positions and the first OGR set (e.g. the initial OGR set 1410) may be determined by pulling every pixel from the target image onto the regions specified by the warped occlusion map. The depth fidelity map may be determined for example as described with reference to Figure 17, in which disparity fidelity weights are computed according to Eq. 13, in which the numerator term defines an absolute difference between a scalar value of disparity for a given layer of the MPI representation and the disparity map of the target view. Depth-based pixel removal may be performed, for example, as described herein with reference to Figures 18A-18B, in which pixels of each layer are removed based on their disparity conformity based on the disparity fidelity mask. As described with reference to the process of Figure 17, operations of the process include multiplying the disparity fidelity mask (map) and the initial OGR set, allowing only the pixels with matching disparities in each layer to remain while removing the rest.
[0010] In some examples, the method further comprises warping the depth fidelity map for each of the one or more second images to the camera pose of the first image to obtain a warped depth fidelity map; wherein the selectively filling includes: for a layer of the first MPI representation, comparing the warped fidelity maps for the one or more second images of the scene, and selecting, from the respective sets of the pixel values, one or more subsets for which the warped fidelity map value indicates substantial matching of depth representation between the first image and a respective one of the one or more second images; and filling an occluded portion of the layer using the one or more subsets. For example, the warping of the fidelity map may be performed as described herein with reference to Figures 18A and 18B, where a block 1730 is configured to warp the disparity fidelity mask (map) back to the source camera view for additional uses including disparity-fidelity-based OGR combination. The comparing of the warped fidelity maps for the various second images of the scene and selection of pixels based on warped fidelity map values may be performed as described with reference to Figures 20A-20B, in a process described as disparity-fidelity-based combination, in which the collected OGR sets from multiple target views are combined by comparing disparity fidelity values across differentDolby Laboratories Licensing Corporation C17491EP0 / D23150EP views and choosing, for each pixel location, the OGR from the view with the highest value. It is noted that by performing this comparison and selection at each pixel location, this results in a subset of pixels for which the warped fidelity map value indicates substantial matching of depth representation (in this case disparity) between the source image and a given second image.
[0011] According to another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method.
[0012] According to yet another example embodiment, an apparatus for generating an MPI representation of a scene comprises: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a first MPI representation of the scene based on a first image of the scene; construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtain, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
[0013] According to yet another example embodiment, an apparatus for generating an MPI representation of a scene comprises: a plurality of neural networks; an MPI compositor configured to generate a first MPI representation of the scene based on a first set of outputs generated by the plurality of neural networks in response to a first image of the scene; and an occlusion guided residuals (OGR) generator configured to construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation, the occlusion map being constructed based on a second set of outputs generated by the plurality of neural networks, wherein the plurality of neural networks is configured to obtain, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and wherein the MPI compositor is further configured to generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:
[0015] FIG. 1 depicts an example process for a video / image delivery pipeline.
[0016] FIG. 2 pictorially illustrates a 3D-scene representation with a multiplane image according to an embodiment.
[0017] FIG. 3 pictorially illustrates a multiview-based approach to generating the multiplane image of FIG. 2 according to an embodiment.
[0018] FIG. 4 is a flowchart illustrating a method of generating an integrated multiview multiplane image of FIG. 2 according to various embodiments.
[0019] FIG. 5 is a block diagram illustrating a feature-encoder block used in the method of FIG. 4 according to an embodiment.
[0020] FIG. 6 is a block diagram illustrating a disparity-information-processor block used in the method of FIG. 4 according to an embodiment.
[0021] FIG. 7 is a block diagram illustrating an alpha-generator block used in the method of FIG. 4 according to an embodiment.
[0022] FIG. 8 is a block diagram illustrating an inter-layer texture filler generator block used in the method of FIG. 4 according to an embodiment.
[0023] FIG. 9 is a block diagram illustrating a camera pose adjust net used in the method of FIG. 4 according to an embodiment.
[0024] FIG. 10 pictorially illustrates certain operations of the method of FIG. 4 according to an embodiment.
[0025] FIG. 11 is a flowchart illustrating operations performed in an occlusion guided residuals (OGR) generator used in the method of FIG. 4 according to an embodiment.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0026] FIG. 12 pictorially illustrates formulation of an occlusion map with the OGR generator of FIG. 11 according to an example.
[0027] FIG. 13 pictorially illustrates an example occlusion map constructed with the OGR Generator of FIG. 11 based on different camera poses according to an embodiment.
[0028] FIG. 14 is a block diagram illustrating certain operations directed to OGR extraction that are performed by the OGR Generator of FIG. 11 according to an embodiment.
[0029] FIG. 15 shows an example pseudo code for the operations illustrated in FIG. 14 according to an embodiment.
[0030] FIG. 16 is a block diagram illustrating how the disparity distance between the MPI layers is modified based on the camera rotations using a subset of operations performed in the OGR Generator of FIG. 11 according to an embodiment.
[0031] FIG. 17 is a block diagram illustrating a process of generating the disparity fidelity mask implemented in the OGR Generator of FIG. 11 according to an embodiment.
[0032] FIGS. 18A-18B pictorially illustrate an example of effective removal of pixels based on disparity conformity performed by the OGR Generator of FIG. 11 according to an embodiment.
[0033] FIGS. 19A-19B show an example pseudo code for the process illustrated in FIG. 17 according to an embodiment.
[0034] FIGS. 20A-20B illustrate certain operations directed to disparity fidelity-based combination performed in the OGR Generator of FIG. 11 according to an embodiment.
[0035] FIG. 21 graphically illustrates an example effect on the view selection map of a low pass operation performed in the OGR Generator of FIG. 11 according to an embodiment.
[0036] FIG. 22 shows a table listing various inputs received by the MPI Compositor block from other blocks of the method of FIG. 4 according to an embodiment.
[0037] FIG. 23 is a block diagram illustrating a training procedure according to an embodiment.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0038] FIG. 24 shows a table that lists inputs and outputs of the trainable blocks that are being trained with the training procedure of FIG. 23 according to an embodiment.
[0039] FIG. 25 is a flowchart illustrating various operations of the training procedure of FIG. 23 according to an embodiment.
[0040] FIG. 26 pictorially illustrates an alpha shape mask used in the training procedure of FIG. 23 according to one example.
[0041] FIG. 27 pictorially illustrates the process of computing one the loss functions used in the training procedure of FIG. 23 according to one example.
[0042] FIG. 28 pictorially illustrates effects of weights in the formulation of render fidelity loss on the rendered image quality according to one example.
[0043] FIGS. 29A-29B pictorially illustrate the rendered results corresponding to a loss function configuration not employing the occlusion loss according to one example.
[0044] FIGS. 30A-30C show scatter plots corresponding to FIGS. 29A-29B for three color channels (R, G, B) according to an example.
[0045] FIG. 31 shows a histogram corresponding to the occlusion map of FIG. 29B according to an example.
[0046] FIGS. 32A-32B graphically illustrate the effect of the occlusion loss component of the loss function on the rendered occlusion map according to an example.
[0047] FIG. 33 is a block diagram illustrating a computing device used to implement various embodiments. DETAILED DESCRIPTION Example Video / Image Delivery Pipeline
[0048] FIG. 1 depicts an example process of a video / image delivery pipeline (100), showing various stages from video / image capture to video / image-content display according to an embodiment. A sequence of video / image frames (102) may be captured or generated using an image-generation block (105). The frames (102) may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and / orDolby Laboratories Licensing Corporation C17491EP0 / D23150EP image data (107). Alternatively, the frames (102) may be captured on film by a film camera. Then, the film may be translated into a digital format to provide the video / image data (107). In some examples, the image-generation block (105) includes generating an MPI image or video.
[0049] In a production phase (110), the data (107) may be edited to provide a video / image production stream (112). The data of the video / image production stream (112) may be provided to a processor (or one or more processors, such as a central processing unit, CPU) at a post- production block (115) for post-production editing. The post-production editing of the block (115) may include, e.g., adjusting or modifying colors or brightness in particular areas of an image to enhance the image quality or achieve a particular appearance for the image in accordance with the video creator’s creative intent. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, removal of artifacts, etc.) may be performed at the block (115) to yield a “final” version (117) of the production for distribution. In some examples, operations performed at the block (115) include enhancing texture and / or alpha channels in multiplane images / video. During the post-production editing (115), video and / or images may be viewed on a reference display (125).
[0050] Following the post-production (115), the data of the final version (117) may be delivered to a coding block (120) for being further delivered downstream to decoding and playback devices, such as television sets, set-top boxes, movie theaters, and the like. In some embodiments, the coding block (120) may include audio and video encoders, such as those defined by the ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate a coded bitstream (122). In a receiver, the coded bitstream (122) is decoded by a decoding unit (130) to generate a corresponding decoded signal (132) representing a copy or a close approximation of the signal (117). The receiver may be attached to a target display (140) that may have somewhat or completely different characteristics than the reference display (125). In such cases, a display management (DM) block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Depending on the embodiment, the decoding unit (130) and display management block (135) may include individual processors or may be based on a single integrated processing unit.
[0051] A codec used in the coding block (120) and / or the decoding block (130) enables video / image data processing and compression / decompression. The compression is used in the coding block (120) to make the corresponding file(s) or stream(s) smaller. The decoding processDolby Laboratories Licensing Corporation C17491EP0 / D23150EP carried out by the decoding block (130) typically includes decompressing the received video / image data file(s) or streams(s) into a form usable for playback and / or further editing. Example coding / decoding operations that can be used in the coding block (120) and the decoding unit (130) according to various embodiments are described in more details below. Multiplane Imaging
[0052] A multiplane image comprises multiple image planes, with each of the image planes being a “snapshot” of the 3D scene at a certain depth with respect to the camera position. Information stored in each plane includes the texture information (e.g., represented by the R, G, B values) and transparency information (e.g., represented by the alpha (A) values). Herein, the acronyms R, G, B stand for red, green, and blue, respectively. In some examples, the three texture components can be (Y, Cb, Cr), or (I, Ct, Cp), or another functionally similar set of values. There are different ways in which a multiplane image can be generated. For example, two or more input images from two or more cameras located at different known viewpoints can be co-processed to generate a corresponding multiplane image. Alternatively, a multiplane image can be generated using a source image captured by a single camera.
[0053] FIG. 2 pictorially illustrates a 3D scene representation using a multiplane image (200) according to an embodiment. The multiplane image (200) has Nlplanes or layers (P0, P1, …, P(Nl−1)), where Nl is an integer greater than one. Typically, the planes (layers) are indexed such that the most remote layer, from the source view (s) (which may also be referred to as the reference camera position (RCP)), is labeled as the (Nl−1)-th layer. The index is decremented by one for each next layer located closer to the RCP. The plane (layer) that is the closest to the RCP is the layer (P0). Each of the planes (P0, P1, …, P(Nl−1)) is orthogonal to a base plane (202) which is parallel to the XZ-coordinate plane. The RCP is at a vertical height h0 above the base plane (202). The XYZ triad shown in FIG. 2 indicates the general orientation of the multiplane image (200) and the planes (P0, P1, …, P(Nl−1)) with respect to the X, Y, and Z dimensions of the 3D scene. In various examples, the number Nlcan be 32, 16, 8, or any other suitable integer greater than one.
[0054] Let us denote the three channel RGB plane of the ^^^layer at camera position ^ as ^^^. Similarly, let us denote the one-channel α plane of the ^^^layer at camera position ^ as ^^^. Each^^^ MPI layer has a four-channel RGBA plane denoted as (^^ ^^, ^^). The whole MPI (200) canthus be collectively represented as:Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP MPI(s) = ^(^^, ^^)^^^^^^ ^ ^^^ (1)The MPI dimension is ^^ × 4 × ℎ × ^. The ℎ and ^ refer to the resolutions of height and widthof the original reference view (without moving camera), denoted as ^^. The depth distance between the ^^^layer to the reference camera position is ^^^. Note that the distance between two neighboring layers does not need to be a fixed equal interval for different pairs of layers. In some examples, the distances can be adaptive distances, e.g., based on different contents and selected to provide optimal novel view rendering. It is straightforward to extend this still MPI image representation to a video representation, provided that the camera position s is kept static overtime. This video representation is given by Eq. (2): ^^^^where t denotes time.
[0055] As already indicated above, a multiplane image, such as the multiplane image (200), can be generated from a single source image ^ or from two or more source images. Such generation may be performed, e.g., during the production phase (110). The corresponding MPI generation algorithm(s) may typically output the multiplane image (200) containing XYZ- resolved pixel values.
[0056] By processing the multiplane image (200) represented by ^, an MPI-rendering algorithm can generate a viewable image corresponding to the RCP or to a new virtual camera position that is different from the RCP. An example MPI-rendering algorithm (often referred to as the “MPI viewer”) that can be used for this purpose may include the steps of warping and compositing. Other suitable MPI viewers may also be used. The rendered multiplane image (200) can be viewed, e.g., on the reference display (125).
[0057] During the warping step of the MPI-rendering algorithm, each MPI^, ^ )needs to be warped from the source view ^ to a novel target view ^. This can be done by applying homography warping !(⋅) which establishes a correspondence between the sourcepixel coordinates (#^, $^) and the target pixel coordinates (#^, $^). The correspondence for the^^^MPI layer is given as: %1 where )^and )^are the intrinsic camera parameters at the source (^) and target (^) positions, respectively. The functions ^ and , are the extrinsic camera parameters describing rotation andDolby Laboratories Licensing Corporation C17491EP0 / D23150EP translation between two camera positions. The n is the normal vector
[0001] T, and ^^^is the distance to a plane that is fronto-parallel to the source camera. The warp amount is different for different layers due to the effect of layer depth (^^^) in Eq. (3).
[0058] We express each MPI layer that have been warped from view ^ to view ^ as (→4) ( →4) the compositing step of the MPI-rendering algorithm, we can render ausing these warped MPI layers, e.g., using processing operations corresponding to the following equations: ^where the weights 6^expressed as: ^^^The weight the visibility weight of the ^^^color channel, where values for each layer at each pixel location are determined by the ray presence up until the current layer times the surface opacity of the current layer as expressed by Eq. (5).
[0059] Depth information can be represented in multiple forms. The term ‘depth representation’ is used herein to include any form of representing of depth. One form of depth representation is the distance between an observer’s (or camera’s) position to the specific point in the 3D scene. The ^-th layer’s depth (^^^) indicated in FIG. 2, for instance, use this form of the ‘depth’ representation.
[0060] Another form of depth representation is referred to as the ‘disparity’. The ‘disparity’ represents the horizontal shift between the corresponding points of the left and right images of a stereo pair. Some depth databases provide stereo image pairs along with their disparity maps to aid tasks like depth estimation, 3D scene reconstruction, or other depth-aware computer vision applications. Various depth estimation models trained on these datasets learn to output a disparity map from the given images. Some examples of the proposed framework are based on this disparity representation, where the depth of each MPI layer is specified using a disparity vector (<^), and the depth information of multiple views is provided in the form of disparity maps. Herein, the disparity map refers to a one-channel 2D plane having the same width and height as the corresponding image. Each pixel of this 2D plane stores the disparity value of the corresponding image pixel. In some examples, these values are normalized to fall within the [0,1] range, where higher disparity values (i.e., values that are closer to one) signify closerDolby Laboratories Licensing Corporation C17491EP0 / D23150EP distances with respect to the observer. While the example algorithms described herein operate primarily on disparities, it is noted that the methods described herein can equally be applied to other forms of depth representation, for example depth representations referring directly to a distance between an observer or camera position to a point in the 3D scene. Disparity-based representations of depth can be mapped to the distance-based form of depth described in the previous paragraph by means of a defined mapping (see equations 6-9).
[0061] We note that, although the rendering algorithm described later herein can mainly operate in the disparity domain, it needs to convert the MPI’s disparity vector (<^) to the depth vector (@^) when applying the warping operation expressed by Eq. (3). For this conversion, thealgorithm needs to “know” the depth range ([^ABCD , ^ECD]) of the scene. For typical contents fordepth related tasks, this depth range is usually provided along with the camera parameters. When not provided, it can be obtained using a suitable alternative method, such as COLMAP. Assuming we have these depth ranges available, we first convert the depth ranges to a corresponding disparity representation as: FABCD = 1 / ^ABCD (6)FECD = 1 / ^ECD (7)Then, we apply min-max normalization on the disparity vector (<^) to obtain the rescaled disparity vector (<H^). If we denote the ^-th element of the disparity vector <^as F^^and the ^-th element of the rescaled disparity vector <H^as FI^^, then: FINote that FI^^, F^^, FABCD, and FECDare scalar values. The min(<^) and max(<^) are also scalar values, with each representing the minimum and maximum values, respectively, from the vector<^. Through Eq. (8), we are rescaling the disparities to cover the ranges of ([FECD , FABCD]). Bytaking the reciprocal of each FI^^, we obtain the depth value ^^^as follows: ^^^^ = JX10The resulting depth vector (@^) properly covers the depth range [^ABCD , ^ECD] of the scene.
[0062] In various examples, the multiplane image (200) can be generated from multiple views or from a single view. In one example single-view-based approach, the source view and its depth information can be used to generate MPI layers with adaptive depth distances. The adaptive depth can be used to optimize the allocation of layers per scene, which facilitates efficient layer utilization from a data transmission standpoint. However, this method may haveDolby Laboratories Licensing Corporation C17491EP0 / D23150EP limitations in terms of the range of pose spans due to the restricted information from a single view. For example, there might be information loss in occluded areas and significant distortions when transitioning to adjacent views.
[0063] One example multi-view-based approach combines multiple neighboring MPIs when rendering a novel view. This approach can be used to handle a broader pose span and address occlusion issues to some extent by selecting different sets of multiple MPIs for each pose. However, example problems associated with this approach include imprecise object separations and holes in cumulative alpha. Consequently, rendering a view under this multi-view-based approach often necessitates sending multiple MPIs, resulting in higher data transmission requirements.
[0064] FIG. 3 pictorially illustrates a multiview-based approach to generating the multiplane image (200) according to an embodiment. Under this multiview-based approach, the multiplane image (200) is generated by integrating information from a plurality of different views (3021- 302N) of the same 3D scene. As a nonlimiting example, FIG. 3 illustrates a configuration in which N=5. In other examples, other values of N can similarly be used. In at least some examples, the multiview-based approach pictorially illustrated in FIG. 3 can beneficially be used to overcome at least some of the above-indicated limitations of other approaches to generating MPIs.
[0065] Example features of the multiview-based approach pictorially illustrated in FIG. 3 include one or more of the following: • The disclosed MPI generation framework can identify and collect occluded RGB pixels from other views. For example, we describe a way to utilize alpha layers to form occlusion maps. These occlusion maps are transmitted to other target views to retrieve relevant pixels that share matching depths, effectively filling in the occluded areas. We denote these retrieved pixels as Occlusion Guided Residuals (OGR). • The disclosed MPI framework can iteratively learn accurate surface opacity (A) of the scene through reconstructed RGB textures and through supervision derived from its novel view rendering capability. Through iterations of training, both the scene opacity (A) and textures (RGB) are jointly optimized leading to an accurate MPI representation. • A scene with continuous depth poses challenges to the MPI representation with a discrete set of depths. Accordingly, we introduce an inter-layer texture filler, whichDolby Laboratories Licensing Corporation C17491EP0 / D23150EP includes a learned RGB texture for intermediate depths between the discrete MPI layers. In at least some examples, the inter-layer texture filler can beneficially be used for reducing inter-layer artifacts, especially for scenes with continuous depths. • In some examples, we integrated different modules to mitigate the adverse effects of potential inaccuracies in the input data, such as in the disparity maps and camera poses. In one example, we use a Camera Pose Adjust Net to correct potential errors in pose estimations. Additionally, we designed an algorithm within the OGR combination procedure to reduce potential influence of local inaccuracies in disparity maps. These and other pertinent features of the disclosed MPI framework are described in more detail below. Overall Framework of Integrated Multiview MPI
[0066] In this section, the overall framework of integrated multiview MPI is presented. Considered at a high level, an example method of generating an integrated multiview MPI uses information collected from the plurality of different views (3021-302N) and integrates that information into a single MPI (200). The resulting MPI (200) includes RGB texture pixels collected from the plurality of different views (3021-302N) and alpha layers featuring an accurate surface geometry from multiview supervision. These characteristics enable the corresponding rendering algorithm to present a high-quality rendered scene from a variously selected camera angle that is substantially free of occlusion artifacts.
[0067] FIG. 4 is a flowchart illustrating a method (400) of generating an integrated multiview MPI (200) according to various embodiments. The method (400) corresponds to the inference stage. Several blocks (410, 420, 430, 460, 470) of the method (400) include processing modules with trainable network parameters. Example training methods that can be used to train those modules of the blocks (410, 420, 430, 460, 470) are described below, e.g., in reference to FIGS. 5-22. In the below-described examples, multiview images are provided with corresponding disparity maps and camera poses. However, as noted above, disparity is just one possible depth representation, and the methods below can also be implemented for other types of depth representation. For example, the images may be provided with a map of depth values representing a distance between points of the scene and the observer or camera. In this case, the steps in the below-described methods relating to disparity (i.e. disparity maps, vectors and / or values) and implemented by the Feature Encoder, the Disparity Information Processor, and theDolby Laboratories Licensing Corporation C17491EP0 / D23150EP OGR Generator, would instead be performed for the provided depth values, and the conversion from the disparity domain to determine depth values in the warping step, as described earlier, would not be required. Processing steps applied to disparity values should be understood to be equally applicable to any form of depth representation. Similarly, any components described as disparity-related are equally applicable to other forms of depth representation.
[0068] Inputs to the method (400) include multiview images, the corresponding disparity maps, and camera poses. The views include a source view (402), denoted as ^, which corresponds to the reference camera position that we construct the MPI representation on. In some examples, the center position, such as that corresponding to the view (3023), among the available plurality of views (3021-302N) as the source view. The target views (404), denoted as^Y, represent the rest of the views (3021-302N) other than the source view ^. Here, Z ∈[0, ^Y − 2], where ^Y is the total number of views available. Note that the upper bound of the Zis given as ^Y − 2 because the index starts from 0 and we are excluding the source view from thecount.
[0069] The outputs of the method (400) include a first output (496), MPI ^(^^^, ^^^)^^^^^^^^ ,where ^^refers to the number of layers used in the MPI representation, and a second output (498), a disparity vector, that represents the disparity at which each MPI layer is at. Note that the outputs (496, 498) are typically sent as a bitstream to the corresponding edge device side and are sufficient for performing volumetric representation operations thereat.
[0070] Processing blocks of the method (400) include a Feature Encoder (410), an Inter-layer Texture Filler Generator (420), an alpha generator (430), an OGR Generator (440), an MPI Compositor (450), a Disparity Information Processor (460), and a Camera Pose Adjust Net (470). The respective inputs and outputs to each of these blocks are as follows: • The Feature Encoder (410) receives the source view’s (402) image (^^) and disparity map (]^) as inputs. The Feature Encoder (410) outputs feature maps (^^) that capture hierarchical features of the given RGB and disparity information. • The Disparity Information Processor (460) receives the source view’s (402) image (^^) and disparity map (]^). The Disparity Information Processor (460) outputs (i) a disparity vector (498) (<^), of dimension ^^ × 1, that indicates the disparity of each MPI layer and(ii) an Alpha Shape Mask (_^), of dimension ^^ × 1 × ℎ × ^, that contains informationof the Alpha Shape on each MPI layer.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP • The Alpha Generator (430) receives the feature maps (^^) from the Feature Encoder (410) and the Alpha Shape Mask (_^) from the Disparity Information Processor (460). The Alpha Generator (430) outputs the alpha layers (^^) of dimension ^^ × 1 × ℎ × ^.• The Inter-layer Texture Filler Generator (420) receives the feature maps (^^) from the Feature Encoder (410). The Inter-layer Texture Filler Generator (420) outputs the inter- layer texture filler of dimension ^^ × 3 × ℎ × ^. This inter-layer texture filler (^^)represents the RGB on intermediate disparity between the MPI layers and serves to minimize possible artifacts on a scene with continuous disparity. • The Camera Pose Adjust Net (470) receives camera poses (a) of the views (402, 404) and outputs the corresponding adjusted camera poses (ab). • The Occlusion Guided Residual (OGR) Generator (440) receives the alpha layers (^^) from the Alpha Generator (430) which serve to construct the occlusion map of the current scene. The OGR Generator (440) also receives the disparity vector (<^) and the adjusted camera poses (ab) from the Disparity Information Processor (460) and the Camera Pose Adjust Net (470), respectively, and utilizes those parameters to apply warp and warp-back operations on the occlusion map. The OGR Generator (440) also receives multiple target view’s (404) images (^^c) and their corresponding disparity maps (]^c) and uses them to pull appropriate RGB pixels with matching disparities that could fill-in the occlusion map. These RGB pixels are referred to as the Occlusion Guided Residual (OGR). The OGR Generator (440) operates to combine the OGRs from multiple views to output a combined OGR (^def^) of dimension ^^ × 3 × ℎ × ^. The OGR Generator (440) also collects aDisparity Fidelity Mask (DFM) from the multiple views (404) and forms a combined DFM (_g ]h) of dimension ^^ × 1 × ℎ × ^ that contains information on the extent to which thesecombined OGR pixels maintain disparity accuracy for each layer. As noted elsewhere herein, disparity is just one possible depth representation that can be used with the present techniques. Where other forms of depth representation, a fidelity map is constructed for that depth representation. The term ‘depth fidelity map’ may be used herein to refer to a fidelity map constructed for any type of depth representation. • The MPI Compositor (450) receives the Inter-layer texture filler (^^) from Inter-layer Texture Filler Generator (420). The MPI Compositor (450) also receives the combined OGR (^def^) and combined DFM (_g ]h) from the OGR generator (440). The MPICompositor (450) also receives the source view’s (402) image (^^). The ^^, ^^and ^def^Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP are combined to from RGB layers (^^). The MPI Compositor (450) also receives the alpha layers (^^) from the Alpha Generator (430) and concatenates them to form the final MPI (496) (^^, ^^) of dimension ^^ × 4 × ℎ × ^.Example implementations of the processing blocks (410-470) of the method (400) are described in more detail below in reference to FIGS. 5-22. Example Implementations of Processing Blocks (410-470)
[0071] FIG. 5 is a block diagram illustrating the feature encoder (410) according to an embodiment. In the example shown, the feature encoder (410) is implemented using an encoder component of a U-net architecture. Various examples of the U-net architecture are described in more detail, e.g., in (i) Han, Yuxuan, Ruicheng Wang, and Jiaolong Yang, “Single-view view synthesis in the wild with learned adaptive multiplane images,” ACM SIGGRAPH 2022 Conference Proceedings, 2022; (ii) Godard, Clément, et al., “Digging into self-supervised monocular depth estimation,” Proceedings of the IEEE / CVF international conference on computer vision, 2019; and (iii) Li, Jiaxin, et al., “Mine: Towards continuous depth mpi with nerf for novel view synthesis,” Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021, each of which is incorporated herein by reference in its entirety. In other examples, other suitable architectures may also be used to implement the feature encoder (410).
[0072] In the example shown, the feature encoder (410) employs a down-sampling branch (502) of the U-net topology including convolutional neural networks (CNNs) arranged as indicated in FIG. 5. The CNNs of the branch (502) are trained to learn mapping from the source view’s (402) RGB (^^) and disparity information (]^) to the multi-resolution feature maps ^^. The multi-resolution information provided by the feature maps ^^is beneficial for a wide range of tasks as it captures both large-scale and low-scale features, thereby enabling the feature encoder (410) to comprehend semantic details present in the input scene.
[0073] A proper design of a loss function is important for training the weights of the CNNs of the feature encoder (410). In particular, it is important to optimally align the feature space with the intended goal of reconstructing good alpha layers (^^) and the inter-layer texture filler (^^) for the output MPI representation (496). A detailed explanation of the training process, including the choice and application of the loss function, is presented below in the section entitled “Networks Training.”Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0074] FIG. 6 is a block diagram illustrating the disparity information processor (460) according to an embodiment. It is noted that while this component is called the disparity information processor in the context of the present example implementation, in which a disparity maps are received with the images, it can also be configured to process other types of depth representations (e.g. maps of depth values defining a distance between a camera or observer and a point in a scene) where other depth representations are available, with equivalent effect. The disparity information processor (460) receives the source view’s (406) image (^^) and disparity map (]^) and outputs the adaptive and scene-specific disparity vector (498) (<^) and the corresponding Alpha Shape Mask (_^). These adaptive disparity distances optimize the layer usage for each scene and enable robust volumetric representation with a fewer number of layers. The resulting efficiency is particularly beneficial from the data transmission perspective.
[0075] In the example shown, the disparity information processor (460) includes a disparity vector prediction stage (602) and a mask generation stage (604) constructed using certain features the disparity processing architectures described in the above-cited publication of Han, Yuxuan, Ruicheng Wang, and Jiaolong Yang, “Single-view view synthesis in the wild with learned adaptive multiplane images,” ACM SIGGRAPH 2022 Conference Proceedings, 2022. In other examples, uniform disparity distances or alternative methods for determining disparity distances may also be used to implement the disparity information processor (460).
[0076] In the disparity vector prediction stage (602), the source view image (^^) and the disparity map (]^) are first subjected to input processing (610) where they are downsampled to a quarter resolution and replicated ^^times. The input volume is concatenated in the inputprocessing (610) with the initial disparity vector (<i) volume of dimension ^^ × 1 × (ℎ × ^) / 4,wherein every pixel of each plane contains its respective initial disparity values <i(j), j ∈[0, ^^ − 1]. In some examples, the initial disparity vector (<i) is set based on a uniformdisparity distance. The output of the input processing (610) provides a volume (612) ofdimension ^^ × 5 × (ℎ × ^) / 4. Each layer within the volume (612) comprises five channels,including three channels for RGB information, one for the disparity map, and one for current layer’s disparity scalar value. The volume (612) is processed by a context encoder (620), composed of multiple ResNet layers, which encode this information into a lower-dimensional feature representation (622). The feature representation (622) is subjected to a self-attention operation (630) to exchange the feature level geometry and appearance information among different layers. The resulting adjusted feature (632) is processed by a linear layer (640) toDolby Laboratories Licensing Corporation C17491EP0 / D23150EPfurther reduce the dimension to ^^ × 1, thereby providing the adjusted disparity vector (498)(<^).
[0077] The mask generation stage (604) uses input processing (650) to concatenate the volume constructed from the adjusted disparity vector (498) (<^) with the image (^^) and disparity map (]^) that are replicated ^^times, this time at full resolution. This concatenationoperation results in a volume (652) of dimension ^^ × 5 × (ℎ × ^), wherein each layer includesfive channels: three for the RGB information, one the disparity map, and one for the current layer’s disparity scalar value. The volume (652) goes through a series (660) of 2D convolutionblocks to produce the Alpha Shape Mask (_^) of dimension ^^ × 1 × ℎ × ^. This maskencapsulates information about the estimated Alpha Shape for each layer in the MPI and is subsequently employed in the process of constructing the alpha channel within the Alpha Generator block (430).
[0078] FIG. 7 is a block diagram illustrating the alpha generator (430) according to an embodiment. In the example shown, the alpha generator (430) is implemented using a decoder component (720) of the above-mentioned U-net architecture. The Alpha Generator (430) takes as inputs the multi-resolution feature maps (^^) received from the Feature Encoder (410) and the Alpha Shape Mask (_^) received from the Disparity Information Processor (460) and processes these inputs to generate the source alpha layers (^^).
[0079] In addition to the decoder component (720), the alpha generator (430) has an initial feature processing stage (710). In the stage (710), each resolution component of the feature maps (^^) is replicated ^^times, effectively expanding the final representation to ^^layers. TheAlpha Shape Mask (_^), which is of dimension ^^ × 1 × ℎ × ^, serves as an auxiliaryinformation input to the stage (710). The Alpha Shape Mask (_^) is concatenated in the stage (710) with each resolution component of the feature maps after subsampling them to a corresponding resolution. This operation effectively adds an extra feature plane to each resolution component, as indicated in FIG. 7.
[0080] An example overall decoding procedure implemented in alpha generator (430) isas follows. Let us denote the lowest resolution feature map volume as ^^^. This volume first goes through an expansive path, which is an upcovolution operation indicated in FIG. 7 in accordance with the shown legend. This operation involves the nearest-neighbor upsampling bya factor of 2 followed by a 3 × 3 convolution operation. The upsampled feature is thenDolby Laboratories Licensing Corporation C17491EP0 / D23150EP concatenated with the features from the higher resolution, denoted as as ^^^. The concatenated feature is then processed through the corresponding convolutional layers, thereby effectively integrating the information to a smaller number of channels. The expansive process is repeated until a desired plane resolution is achieved. Before the final output ^^, the decoding result is subjected to a sigmoid function to ensure control over the output range. Details of the training process for the Alpha Generator (430) are described in the section entitled “Networks Training.”.
[0081] FIG. 8 is a block diagram illustrating the Inter-layer Texture Filler Generator (420) according to an embodiment. In the example shown, the Inter-layer Texture Filler Generator (420) has an arrangement of modules that is qualitatively similar to that of the alpha generator (430) shown in FIG. 7. More specifically, the Inter-layer Texture Filler Generator (420) is also implemented using a decoder component (820) of the above-mentioned U-net architecture. The Inter-layer Texture Filler Generator (420) takes as inputs the multi-resolution feature maps (^^) received from the Feature Encoder (410) outputs the Inter-layer texture filler (^^). The latter represents the RGB textures on intermediate disparity between the MPI layers and helps to reduce possible boundary artifacts between the layers in the rendered scene, particularly when applying a large pose span on a scene with continuous disparity.
[0082] In addition to the decoder component (820), the Inter-layer Texture Filler Generator (420) has an initial feature processing stage (810). The stage (810) is generally similar to the stage (710), with a noteworthy difference being that the stage (810) does not use explicit auxiliary information, such as the Alpha Shape Mask (_^) in the stage (710). This is possible because the weight supervision for the Inter-layer Texture Filler Generator (420) is derived from gradients passing through the MPI compositor (450), where the mask guidance for ^^is inherently embedded in the formulation. Details of the training process for the Inter-layer Texture Filler Generator (420)) are described in the section entitled “Networks Training.”.
[0083] FIG. 9 is a block diagram illustrating the Camera Pose Adjust Net (470) according to an embodiment. A first operation in the Camera Pose Adjust Net (470) is directed at extracting rotation matrices (^) and translation vectors (,) from the received camera poses (a) of the multiple views (402, 404). The extracted rotation matrices (^) and translation vectors (,) are then subjected to respective adjustments in first and second branches (902, 904) of the Camera Pose Adjust Net (470).Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0084] In some examples, the Camera Pose Adjust Net (470) take the COLMAP camera pose as initialization and is jointly trained with the MPI representation based on the render fidelity on other views. The corresponding loss function is described in detail in in the section entitled “Networks Training.”. Here, we note that, while this pose adjustment can be employed in the rendering stage, for example, using the camera positions as reference points and generating a path around it, this is not obligatory. Any suitable path can serve as a target novel path. The described adjustment and alignment processes primarily pertain to the MPI generation stage, where pixels from multiple views are integrated into a single MPI (200).
[0085] Regarding the rotation adjustment, a first converter module (910) of the first branch(902) operates to convert the 4 × 4 rotation matrices into the corresponding 4 × 1 quaternionrepresentation. In some cases, directly manipulating the rotation matrices can be problematic as orthogonality and determinant constraints need to be maintained for validity but can be perturbed during the adjustment. On the other hand, quaternions provide efficiency by employing fewer parameters compared to the full rotation matrices, leading to a more streamlined network. Moreover, quaternions inherently enforce the unit quaternion constraint, guaranteeing both valid rotations and numerical stability during the network adjustments. Thus, the first branch (902) is configured to use the quaternions for the task of adjusting rotations. The quaternions go through two fully connected layers (FCs) (912, 914) to project them into a feature space. Then, a set (916) of weight banks processes the projections, thereby enabling each cameras’ rotation to be adjusted with its dedicated weights. These weights are purposefully designed to finely refine the alignment for each camera pose. Note that first converter module (910) has a skip connection (918) to guide the network to build upon the initial rotation estimation as a reference point while learning the residuals. A second converter module (920) of the first branch (902) operates to convert the adjusted quaternions back into the corresponding rotation matrices (^H).
[0086] The second branch (904) of the Camera Pose Adjust Net (470) performs the translation adjustment in a similar manner. More specifically, the translation vectors (,) go through two FC layers (922, 924) and then through a set (926) of weight banks, thereby enabling each camera’s translation to be adjusted with its dedicated weights. A skip connection (928) is used to guide the network to build upon the initial translation estimation as a reference point while learning the residuals. Finally, the adjusted translation vectors (,l) generated with the second branch (904) are then combined with the adjusted rotation matrices (^H) generated with the first branch (902) to form the adjusted camera poses (ab).Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0087] The effectiveness of the above-described camera pose adjustment has been confirmed on various scene examples. For example, with the camera pose adjustment described above, the hallucination artifact and the doubled feature artifact caused by slight errors in the target camera position can beneficially be substantially avoided.
[0088] FIG. 10 pictorially illustrates certain operations of the method (400) according to an embodiment. One challenge in generating an MPI representation is to find a way to reconstruct RGB textures that are occluded by the objects located in the foreground layers of the scene. A first image (1002) shown in FIG. 10 is an example image corresponding to a certain source camera position. The first image (1002) has inherently limited information on the textures behind the present foreground object, i.e., the guitar man. However, views from other cameras, possibly from different positions and / or angles, may contain some information on those textures. Through homography warping, the method (400) can be configured to align those textures as if seen from the source camera position. Accordingly, the method (400) operates to collect, aggregate, and combine those textures from the other views (402, 404) as pictorially indicated by a middle collage (1004). The method (400) then operates to reconstruct (1005) the occluded textures for the background layers (1006) corresponding to the first image (1002) based on the pertinent information contained in the middle collage (1004).
[0089] FIG. 11 is a flowchart illustrating operations performed in the OGR Generator (440) according to an embodiment. For illustrations purposes, some operations are grouped into three respective sets of operations denoted as (1101, 1102, 1103). The first set (1101) is directed to Occlusion Map Construction. The second set (1102) is directed to OGR extraction. The third set (1103) is directed to Disparity Fidelity-based Combination. Each of the sets (1101, 1102, 1103) is described in more detail below.
[0090] Referring to the first set (1101) directed to Occlusion Map Construction, the Alpha (^^) contains the opacity information of each layer and thus can be used to construct the Occlusion Map (e^), which contains information indicating which regions are occluded within each layer due to the opacity of the preceding layers. Both ^^and e^are of dimension^ × 1 × ℎ × ^ ^^ . Herein, e is defined as follows:e^^ = (1 − m^^) ^^^ (10)where^= :^^ P1 − ^:Q (11)Dolby Laboratories Licensing Corporation C17491EP0 / D23150EPHerein, the m^ represents ^^ the transmittance of a ray up to the ^-th layer. Thus, the term (1 − m^ )in equation (10) represents the ray obstruction up to the ^-th layer. Eq. (10) shows how the likelihood of occlusion for each ^-th layer is determined by identifying regions with both highray obstructions (1 − m^ ^^ ) and high alpha (^^) values. These regions can be interpreted as areaswhere a ray is obstructed by previous layers but encounters an opaque surface, which aligns with the definition of occlusion.
[0091] FIG. 12 pictorially illustrates formulation of the occlusion map (e^) using operations of the first set (1101) according to an example. Therein, a first panel (1202) represents the rayobstruction term (1 − m^) for different layers i=0, …, 15. A second panel (1204) represents thesurface opacity (^^) for the different layers. A third panel (1208) represents the occlusion map (e^) obtained through a multiplication operation (1206) in accordance with Eq. (10).
[0092] A fourth panel (1210) represents the different layers of the corresponding RGBA MPI (200) obtained using a prior-art method described in the above-cited publication of Han, Yuxuan, Ruicheng Wang, and Jiaolong Yang, “Single-view view synthesis in the wild with learned adaptive multiplane images,” ACM SIGGRAPH 2022, Conference Proceedings. In this case, the prior-art method is a single view-based method and, as such, it has limitations in filling-in the occluded areas solely based on the source image’s pixels. Thus, this prior-art method composites the RGB layers of the MPI (200) using the source image’s pixels on the regionsindicated by m^ and predicts the rest of the (1 − m^) area using an in-painting network. Sincethe RGBA representation shows alpha-weighted RGB pixels, these in-painted pixels appear onregions specified by (1 − m^)^^, which correspond to the occlusion map (e^) of the third panel(1208). Upon inspecting the RGB textures in the fourth panel (1210) within the regions specified in the third panel (1208), it becomes evident that some textures appear blurred. This blurring is a pictorial manifestation of certain limitations of the in-painting network employed in the prior-art method. In contrast, operations of the OGR Generator (440) beneficially provide significant improvements over these limitations, as described in more detail below.
[0093] FIG. 13 pictorially illustrates an example occlusion map (e^) constructed with the first set (1101) of operations performed in the OGR Generator (440) based on different camera poses according to an embodiment. When rendering the occlusion map e^(1306) back at its source camera view position (s), as depicted in a first panel (1302), we anticipate the map (1306) to be a sparse map because the occluded regions are not expected to be revealed much from the source camera perspective. Thus, the source view can be reconstructed with high fidelity solelyDolby Laboratories Licensing Corporation C17491EP0 / D23150EP based on the pixels from the source view (^^). However, when we warp to a new camera position and then render, as indicated in a second panel (1304), we observe that the occluded regions are now revealed as indicated by an occlusion map (1308). The OGR Generator (440) is configured to use this information to extract the relevant pixels from other target views and then integrate them back into the source camera view.
[0094] FIG. 14 is a block diagram illustrating certain steps of the second set (1102) of operations performed in the OGR Generator (440) and directed to OGR extraction according to an embodiment. After forming an occlusion map (e^) (1404) using operations of an occlusion map construction block (1402), a warping block (1406) performs warping to various target camera positions using the adjusted camera poses (ab) and the disparity vectors (<^). In some examples, the warping operations of the warping block (1406) are implemented according to the general procedures described above in the section entitled “Multiplane Imaging” (e.g., see Eq.(3) and the corresponding description). The resulting warped occlusion map e(^→^c), Z ∈[0, ^Y − 2], holds information about where the occluded region of each layer is placed on eachone of the target views. Accordingly, the OGR Generator (440) operates to form an initial OGR set (1410) ^^ecf^by pulling every pixel from the target image (^^c) onto the regions specified by the warped occlusion map e(^→^c)via a multiplication operation (1408).
[0095] FIG. 15 shows an example pseudo code (1500) for the operations illustrated in FIG. 14 according to an embodiment.
[0096] Referring back to FIG. 14, each layer of the initialOGR set (1410) includes respective elements that do not match the layer’s disparity. For example, this effect manifests itself as a presence of the “guitar man” in later layers with lower disparity (or greater depth). Thus, disparity-based supervision is applied to eliminate pixels in each layer that do not have matching disparity.
[0097] To properly implement the disparity-based supervision according to an embodiment, two pieces of information are needed. A first of the two pieces is the disparity vector that contains the disparity information of each MPI layer. A second of the two pieces is the disparity map (]^c) of the target view ^Yto discern and eliminate irrelevant pixels from each layer. Regarding the disparity vector, the OGR Generator (440) is configured to use the adjusted disparity vector <(^→^c)which takes account of the disparity change coming from the warping to a new camera position.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0098] FIG. 16 is a block diagram illustrating how the disparity distance between the MPI layers is modified based on the camera rotations in the second set (1102) of operations performed in the OGR Generator (440) according to an embodiment. In some examples, the adjustment is computed as follows: (^→^c) ^ ^^Herein, the quantity F^is a original disparity vector (<^). The quantity F(^→^^c)is a scalar referring to the ^-th element of the adjusted disparity vector(< ^ ^ 1] is the index of the MPI layers. The angles st and su areyaw angle and the pitch angle, respectively.
[0099] FIG. 17 is a block diagram illustrating a process (1700) of generating the disparity fidelity mask implemented in the second set (1102) of operations performed in the OGR Generator (440) according to an embodiment. In at least some examples, the process (1700) facilitates the disparity-based pixel removal. The process (1700) includes generating (1710) a set (1712) of disparity fidelity weights (6^]ch ) using the disparity map of the target view (]^c) and the disparity vectorThe set (1712) ofhis a volume of^^ × 1 × ℎ × ^, where we have ^^ layers of 2D planes, each having the dimension of ℎ × ^. Ifwe denote the ^-th layer plane ofit is expressed
[0100] The numerator termin Eq. (13) is a 2D plane with dimension ℎ ×This is because it takes absolute difference between a scalar valuec)and a 2D plane ]^cwhich itself has dimension ℎ × ^. It defines a plane in which each pixel's value represents thedisparity deviation of the corresponding pixel location relative to the disparity of the current ^-th layer.
[0101] The denominator term^^^ t^^ y^^ − ^is a scalar value which take the of absolute differences between the disparity map(]^c) and the disparity vector (< c ) over all possible (#, $, ^) within their specified ranges.Here, the ^ ∈ [0, ^^ − 1] refers to the index of the layers, $ ∈ [0, ℎ − 1] refers to the verticalcoordinate of a plane, and # ∈ [0, ^ − 1] refers to the horizontal coordinate of a plane. ThisDolby Laboratories Licensing Corporation C17491EP0 / D23150EP denominator term serves as a normalization factor, ensuring that the value of the numerator falls within the range of [0,1].
[0102] The final 6^]ch,^is computed using Eq. (13) by subtracting this normalized disparity deviation plane from one, effectively inverting its meaning where a higher value now signifies lower disparity deviation (or higher disparity fidelity). In summary, each plane (6^]ch,^) contains an absolute difference map with values normalized to [0,1]. The regions with higher values (closer to one) indicate highly matching disparity to the current ^-th layer's disparity.
[0103] The example set (1712) shown in FIG. 17 visualizes the 6^]ch with 16 layers. For the top left earlier layers with greater disparity (closer depth) of the shown set (1712), one can see that the closer object, such as the guitar man, is indicated brighter (higher value), while the far away background wall regions are indicated darker. For the bottom right later layers with less disparity (greater depth) of the shown set (1712), one can see the opposite tendency, where far away background wall regions are indicated brighter whereas the close-by guitar man is indicated darker.
[0104] While the Disparity Fidelity Weights (6^]ch ) contains information that can be used for disparity-based pixel removal, it cannot be used as is because it presents continuous variations. For instance, in the shown set (1712), one can see that the intermediate layers are overall showing bright maps that give ambiguity the process of distinguishing the regions for the decision to remove the pixels or not. Thus, to highlight only the areas with a high degree of disparity fidelity, the process (1700) is configured to apply soft thresholding (1720)ch using an exponential weighting function. An example output the soft thresholding (1720) is visualized by a set (1722) illustrating the disparity fidelity map]h,^corresponding to the shown set (1712). In one formulated as follows: _Byu (C0)where we apply layer-adaptive parameter ^^, ^ ∈ [0, ^^].
[0105] The adaptive thresholding in accordance with Eq. (14) gives room to control threshold sensitivity adaptively on different layers. In general, the early and intermediate layers need higher value of ^^to transfer OGRs only on the regions with significantly higher disparity fidelity. Otherwise, unnecessary residuals may be transferred to these early and intermediateDolby Laboratories Licensing Corporation C17491EP0 / D23150EP layers, showing up as an artifact on the rendered scene. The later layers can have a lower value of ^^. Imposing the higher ^^value on the later layers may cause too few OGRs being transferred, thereby reducing their disocclusion capability.
[0106] In some examples, we set ^^adaptively to take larger values on earlier layers and gradually decreasing on later layers, e.g., as follows: ^^In some examples, ^xCy = = for 16 layers. Othervalues may be used for other (than 16) numbers of layers. The different threshold sensitivities corresponding to different ^^values are illustrated in FIG. 17 in a graph inset (1724).
[0107] The set (1722) visualizes the resulting _^]ch for 16 layers. Therein, one can see that only the regions with a relatively high degree of disparity fidelity are indicated bright, and the rest of the regions are indicated dark. Operations of the process (1700) include multiplying the disparity fidelity mask (_^]c ^h ) and the initial ^ecf^, allowing only the pixels with matching disparities in each layer to remain while removing the rest.
[0108] FIGS. 18A-18B pictorially an example of effective removal of pixels in each layer based on their disparity conformity performed by the OGR Generator (440) according to an embodiment. More specifically, FIG. 18A illustrates the ^^ecf^before the pixel removal. FIG. 18B illustrates the final ^^ecf^after the pixel removal in a block (1730) of the process (1700) performed using the disparity fidelity mask (1722) _^]ch . The resulting MPI layers now contain the properly separated residuals for foreground and background, corresponding to the guitar man and the brick wall, respectively. Operations of the block (1730) also include warping the final OGR (^^ecf^) back to the source camera view. The block (1730) is also configured to warp backthe disparity fidelity mask]halongside the finalf^for various additional uses, e.g., including disparity fidelity-based OGR combination and inter-layer texture filler (^^) masking.
[0109] FIGS. 19A-19B show an example pseudo code (1900) for the process (1700) according to an embodiment.
[0110] FIGS. 20A-20B illustrate certain operations of the third set (1103) of operations performed in the OGR Generator (440) and directed to disparity fidelity-based combination according to an embodiment. More specifically, FIG. 20A is a block diagram illustrating a viewDolby Laboratories Licensing Corporation C17491EP0 / D23150EP selection procedure (2002). FIG. 20B graphically illustrates an example view selection map (2004).
[0111] In various examples, the collected ^(^efc→^)^ from multiple target views are combined using the disparity fidelity as pertinent criteria. More specifically, for each pixel location of the layer’s occluded region, defined by e^, the OGR Generator (440) operates to compare the (^ →^)across all views and choose the OGR from the view with the highest value. As recap, a higher _(^]hc→^)value indicates closer proximity to the MPI layer’s disparity. The ^-th layer of the combined OGR (^def^,^) can be expressed as follows: ^^where ^
[0112] In Eq. (17), the function ^(⋅) serves as a target view index selector that takes thepixel position (#, $) and layer index (^) as inputs and compares the values c→^)h,^ (#, $)all available view Z ∈ [0, ^Y − 2]. Then, the function ^(⋅) outputs the view index that gives themaximum value. Based on these selected view indices for each layer and pixel, the ^def^is constructed by extracting the pixel values from the respective view’s ^(^efc→^)^ . The view selection procedure (2002) of FIG. 20A is the procedure for an ^-th layer of an example MPI. The differently shades rectangles (2014, 2016, 2018) represent OGR pixels from different target views denoted as ^^, ^^, and ^^, respectively. Indicated in parentheses are pixels corresponding to disparity fidelity values _(^hc,^. disparity fidelity value is the highest, its OGR is selected (2012) and gets pasted to the final combined OGR. In the example of OGR selection map (2004) shown in FIG. 20B, OGRs from different views are chosen at different spatial locations to construct the final combined OGR for each layer.
[0113] Given that the combination process (1103) relies on the disparity information, the accuracy of the provided disparity maps significantly affects the quality of the combined OGR. For example, local inaccuracies on disparity maps can lead to scatter artifacts, due to which some objects may appear scattered and / or deformed. To minimize a potential impact of such local inaccuracies in the disparity map, the OGR Generator (440) is configured to apply an adjustment to the ^(⋅) function of Eq. (17), e.g., as follows:Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP ^X(#, $, ^) = ^^^^^# (^c→^)Y (#, $)), Z ∈ [0, ^Y − 2] (18)where(^c→^) = ^ ^ ^ ⋅ (^c→^)^^^^ + +and^:^ ^^ ^ ^ ^^
[0114] In some examples, the parameter ^ in Eq. (20) is set to ^ = 3 to form a low passkernel which is a zero-mean 2D Gaussian with the standard deviation of 3. The parameters Land M in Eq. (19) are set to ^ = 6 and ^ = 6 to have the kernel window sampled out to twostandard deviations for both horizontal and vertical directions. In accordance with Eq. (19), the OGR Generator (440) operates to apply the low pass kernel on each layer of the disparity fidelity mask ensure stable fidelity evaluation, even in the presence of local fluctuationsthe disparity maps. Note that the OGR Generator (440) is not applying the low pass operation on the OGR pixels themselves, but rather on the disparity fidelity mask), which is used a combination criterion in Eqs. (18) and (16). As a result, blurs on the pixels are typically avoided while stable combination is ensured.
[0115] FIG. 21 graphically illustrates an example effect on the view selection map of a low pass operation (2104) corresponding to Eqs. (18)-(20). The example illustrated in FIG. 21 represents content characterized by a noisy disparity map. Accordingly, an unfiltered view selection map (2102) exhibits significant local variations in the view selections. When the corresponding view selection is used to construct ^def^in accordance with Eq. (16), it leads to noticeable scatter artifacts. After the low pass operation (2104), a resulting filtered view selection map (2106) has the OGRs grouped in larger, more coherent clusters. The combined OGR using the view selection corresponding to the map (2106) can be constructed similar to the previously described ^ef^,^ ef^,^In this case, the view selection scheme is changed from ^(⋅) to ^X(⋅) (also see Eq. (16)). This change tends to cause the combined OGRs to have stable structures and be substantially free of scatter artifacts.
[0116] Alongside the ^def^generated based on Eq. (21), the OGR Generator (440) alsooperates to generate a combined disparity fidelity mask (_g ]h) as an additional output from theDolby Laboratories Licensing Corporation C17491EP0 / D23150EP thirst set (1103) of operations. Similar to how the OGRs are combined based on equation (21),_g ]h can be formed as follows:(^^H(^,q,0)→^)The resulting _g ]h is ainformation on thedextent to which ^ef^maintains disparity for each layer and is used in the MPI compositor (450).
[0117] FIG. 22 shows a table listing various inputs received by the MPI Compositor (450) from other blocks of the method (400) according to an embodiment. The processing of the received Alpha information in the MPI Compositor (450) is relatively straightforward, with the input ^^being directly concatenated to the final MPI representation (496). The processing of the received RGB information in the MPI Compositor (450) includes fusing three kinds of plane / volume holding RGB information pieces, which include: • Source Image (^^): Pixels from source camera view. These pixels can be used to fill- in the non-occluded regions of each layer. • Combined OGR (^def^): Pixels collected from multiple target camera views on regions that are occluded from the source camera view. These pixels can be used to fill-in the occluded regions of each layer. • Inter-layer texture filler (^^): Pixels that are learnt from a network to complete the remaining non-zero alpha regions within each layer that are not covered by the source image pixels or OGRs. These pixels represent the RGB textures at intermediate disparity between MPI layers, which may not be explicitly captured in either the source view pixel or the OGRs but become visible from certain novel camera poses. Including these pixels helps to reduce inter-layer artifacts in scenes with continuous disparity.
[0118] In some examples, the composite formulation for constructing MPI RGB layers for the output (496) can be as follows: ^^ ^ ^ ^^ = m^ ^ + (1 − m^) ^^,^ (23)where ^^refers to the RGB textures that are occluded in the reference view. m^^is a composite weight primarily influenced by ^^as defined in Eq. (11). An example approach to constructing ^^is to train an inpainting network. However, the inpainting network has certain qualityDolby Laboratories Licensing Corporation C17491EP0 / D23150EP limitations and may introduce artifacts related to temporal inconsistency when generating MPIs for multiple frames.
[0119] In some other examples, the MPI Compositor (450) is configured to use: (i) ^def^collected from multiple views and _g ]h, conveying the layer disparity fidelity of the pixels in^def^; and (ii) ^^representing the learned RGB textures capable of filling in surfaces that are not covered by source and target view pixels. In such examples, the MPI Compositor (450) can be configured to formulate the MPI composition as follows: ^^ ^ ^^Here, we formulate the occluded textures ^^, based on Eq. (23) as the convex combination of ^def^and ^^, with greater weights assigned to ^def^pixels in the regions characterized by high_g ]h values. This is rational as these ^def^ pixels from other views possess closely matchingdisparities to the current layer, thereby making them into more-suitable candidates for texture placement in those regions. The regions not covered by either ^def^or ^^are expressed as(1 − m^^ )P1 − _g ]h,^Q, which serves as a mask for ^^. The Inter-layer Texture Filler Generator(420) receives weight supervision from gradients that propagate through this mask, thereby allowing ^^to adapt and fill these regions with relevant textures accordingly.
[0120] The effectiveness of the approach exemplified by Eq. (24) has been verified by computer simulations. For example, cases with continuous disparity pose challenges to MPI representation with discrete set of disparities. In such cases, direct application of the approach exemplified by Eq. (23) typically causes inter-layer artifacts coming from limitations of representing a continuous scene with quantized disparity level. In contrast, application of the approach exemplified by Eq. (24) enables the MPI Compositor (450) to effectively learn how to fill intermediate disparity textures without significant observable artifacts.
[0121] In some examples, to form the final MPI representation (^^, ^^) (496), the MPICompositor (450) operates to concatenate the RGB layers (^^) generated in accordance with Eq. (24) and the alpha layers (^^). Networks Training
[0122] FIG. 23 is a block diagram illustrating a training procedure (2300) according to an embodiment. The training procedure (2300) is configured to provide network parameters of the trainable networks of the several blocks (410, 420, 430, 460, 470) of the method (400). In FIG.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP 23, the blocks (410, 420, 430, 460, 470) that are being trained with the training procedure (2300) are collectively denoted as a trainable block (2310). Arrows (2302, 2304, 2306, 2308) indicate operations related to loss functions and weight updates used in the training procedure (2300). Arrows (2301, 2303) indicate external inputs provided for the training procedure (2300). Arrows (2305, 2307, 2309) indicate components that are used to build the MPI representation.
[0123] FIG. 24 shows a table that lists inputs and outputs of each of the blocks (410, 420, 430, 460, 470) of the trainable block (2310) that are being trained with the training procedure (2300) according to an embodiment.
[0124] FIG. 25 is a flowchart (2500) illustrating various operations of the training procedure (2300) according to an embodiment. Several processing blocks (2506, 2508, 2518, 2520) of the flowchart (2500) include operations related to loss functions and weight updates. The flowchart (2500) is described below with continued reference to FIGS. 4 and 23-25.
[0125] In a block (2502) of the training procedure (2300), appropriate inputs are obtained. In various examples, the inputs include the source view’s (402) image (^^) and disparity map (]^). In a block (2504) of the training procedure (2300), these inputs are applied to the Feature Encoder (410), Disparity Information Processor (460), Inter-layer Texture Filler Generator (420), and Alpha Generator (430) of the trainable block (2310) to get the corresponding outputs including the source alpha (^^), disparity vectors (<^), inter-layer texture filler (^^), and alpha shape mask (_^). Operations of the block (2502) also include obtaining the input camera poses (a). Operations of the block (2504) also include applying the input camera poses (a) to the Camera Pose Adjust Net (470) of the trainable block (2310) to get the corresponding adjusted camera poses (ab).
[0126] From the set of outputs generated in the block (2504), one can compute two kinds of losses: the occlusion loss (^¢££^¤^^¢A) and the disparity loss (^J^^uCD^^t), which are computed in the block (2506) and the block (2508), respectively. We briefly describe these losses here and will present more details below in reference to Eqs.-(34). The occlusion loss ^¢££^¤^^¢Ais computed using ^^, which provides supervision on ^^to reduce or minimize ray obstructions.We compute ^J^^uCD^^t using <^, ]^, and _^, which provides constraints on the adaptivedisparity distances and disparity-based constraints on the alpha shape.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0127] Another important loss used in the training procedure (2300) is the render fidelity loss (^DBAJBD), which provides supervision on the MPI representation through a comparison with the ground truth target views. To compute this loss, we first construct the MPI representation which involves collecting pertinent OGRs from the OGR generator (440). In some implementations of this computation, we first collect the target view images (^,¥) and disparity maps (],¥) in a block (2510) of the training procedure (2300). From the OGR Generator (440), we use ^^, <^, ab, ^,¥,and ],¥ to generate the combined OGR (^def^) and DFM (_g ]h) in a block (2512) of the trainingprocedure (2300). In a block (2514) of the training procedure (2300), the MPI Compositor (450)operates generate the final MPI representation ((^^, ^^)) using ^^, ^d ^ef^, ^^, g^ , and _ ]hgenerated in the previous blocks (2504, 2512). In a block (2516) of the training procedure (2300), we warp this MPI to various target views (,¥). In the block (2518) of the training procedure (2300), we assess the similarity of the rendered novel view and the ground truth ^,¥using a range of image quality models, collectively forming ^DBAJBD. We use the gradients of these loss functions to update the weights of the network modules in the block (2520) of the training procedure (2300). The process is repeated for a specified number of iterations (2522).
[0128] The training procedure (2300) uses loss functions, which are denoted as ^J^^uCD^^t, ^DBAJBD, and ^¢££^¤^^¢A. In some examples, these loss functions are consolidated to form the combined loss function denoted as ^^¢^C^: ^^¢^C^ = ¦J^^uCD^^t ^J^^uCD^^t + ¦DBAJBD^DBAJBD + ¦¢££^¤^^¢A^¢££^¤^^¢A (25)In some examples, where we setThecombined loss function ^^¢^C^is used to propagate gradients and update the weights of all network modules of the of the trainable block (2310). Below, we provide a more-detailed description of the loss function ^J^^uCD^^t, ^DBAJBD, and ^¢££^¤^^¢Aaccording to an example embodiment.
[0129] The Disparity Loss (^J^^uCD^^t) refers to a set of disparity-based constraints on thealpha shape (_^) and disparity vector (<^). While ^^and <^are also driven by other sets of supervisions, such as the render fidelity, the disparity loss provides additional control over <^to be in a correct disparity order and over the alpha shape (_^) of each layer to be disparity- consistent with <^.^J^^uCD^^t = ¦¢DJBD^¢DJBD + ¦xC^^^xC^^ (26)whereDolby Laboratories Licensing Corporation C17491EP0 / D23150EP ^= ^∑^^^^ ^ ^^^^ ^^# (0, F^^^ − F^ ) (27)and^= ^ ^^^^ ^^^ ^^^ ⋅ ^ − F^In somepenalize any switched orders among the disparity vector <^, ensuring that <^is guided to maintain a sorted order. The ^xC^^of Eq. (28) provides disparity-based constraints on the alphashape mask (_^). The _^ is a volume of dimension ^^ × 1 × ℎ × ^, where each layer containsthe approximate alpha shape at the current disparity.
[0130] FIG. 26 pictorially illustrates the alpha shape mask _^(2600) used in Eq. (28) according to one example. Each layer of the mask (2600) contains a plane with values in the range [0,1], where the regions with higher values (closer to one) have better matching disparity to the current layer. These regions are likely to have high alpha values. However, the mask _^is only a rough estimate of a possible alpha locations based on the disparity information and does not necessarily have an accurate scene opacity information. For example, the last few layers(e.g., ^ = 13, 14, and 15) in the example of FIG. 26 have low values for regions that werecovered by the foreground guitar man. In reality, there should also be opaque background wall surfaces in those regions. Thus, the Alpha Generator (430) is configured to take this _^as a shape guidance information and is further configured to use supervision from other constituent loss functions of Eq. (25) to form the final alpha layers (^^).
[0131] FIG. 27 pictorially illustrates the process of computing the loss function ^xC^^using Eq. (28) according to one example. A second panel (2704) in FIG. 27 shows an example of the absolute difference map between the disparity map (]^) and the layer 0’s disparity (F^^). The darker regions in the panel (2704)) represent regions with matching disparity. As ^xC^^is driven to be minimized, it effectively guides the shape of the mask (_^) to match these low- valued regions. With a properly matching mask, e.g., as the one depicted in a first panel (2702)of FIG. 27, a multiplication operation (2703) _ ^ ^^,^ ⋅ |] − F^| will yield a sparse map, as themap depicted a third panel (2706) of FIG. 27, which helps to minimize ^xC^^.
[0132] The Render Fidelity Loss (^ ) is to provide supervision to the ensuring that and rendered target view (©(^→^c)) closely aligns with the ground truth target view (©^c). For evaluating the fidelity between two images, we employ the L1 loss (^^), SSIM loss (^^^^x), and LPIPS loss (^^u^u^), which are combined to form ^DBAJBDas follows:Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP ^DBAJBD = ¦^^^ + ¦^^^x^^^^x + ¦^u^u^^^u^u^ (29)In some examples, we set ¦^ = 0.5, ¦^^^x = 0.7, and ¦^u^u^ = 1. Note that each of these losses,i.e., ^^, ^^^^x, and ^^u^u^,a different respective aspect of fidelity, which gives a different training direction. The ^^also supervises the overall fidelity between the compared images. Some implementations of the SSIM loss (^^^^x) and LPIPS loss (^^u^u^) may benefit from certain features disclosed in: (i) Wang, Zhou, et al. “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing 13.4 (2004): 600-612, and (ii) Zhang, Richard, et al. “The unreasonable effectiveness of deep features as a perceptual metric,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, both of which are incorporated herein by reference in their entirety.
[0133] FIG. 28 pictorially illustrates the effects of weights used in Eq. (29) for the formulation of the render fidelity loss on the rendered image quality according to one example. For example, the absence of ^^(i.e., setting ¦^to ¦^=0) can significantly affect the overall result, potentially leading to incorrect images or significantly increased convergence times. However, when ^^is given a higher relative weightit tends to introduce blurriness into the image. Therefore, in some examples, the weightis set to ¦^=0.5 to have less blurriness in the results. This weight value of 0.5 also initially served as a baseline for all other losses. Then, as depicted in FIG. 28, we conducted ablation studies on different ratios of losses. A first panel (2802) in FIG. 28 shows the ground truth image at the view ^^. Panels (2804, 2806) in FIG. 28 show the warped and rendered novel viewgenerated from the MPIs of source view ^, trained with different respective loss function ratios indicated under those panels.
[0134] As indicated by the image in the second panel (2804) of FIG. 28, increased weights on ^^u^u^emphasize the overall perceptual fidelity and, thus, shape the MPI to give relatively sharp and natural looking rendered view. However, as it gives less emphasis to the fidelity of other fine details, it has some problems with the capture of the correct angular fidelity, as indicated by the different face angles between the images of the second and third panels (2804, 2806).
[0135] Giving more weight to ^^^^x, on the other hand, shapes the MPI to give substantially correct structural angles on the rendered views. However, it blurs the image significantly, thereby reducing the overall naturalness of the image. Through experimentation and computersimulation, we determined that the relative weights of ¦^ = 0.5, ¦^^^x = 0.7, and ¦^u^u^ = 1Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP provide nearly optimal results for this particular example. To arrive at these weights, we increased the contribution of the ^^^^xfrom that corresponding to the second panel (2804) case to achieve a nearly optimal balance between the angular correction and sharpness.
[0136] The ^DBAJBDtypically affects the weight of all network modules that contribute to theconstruction of the MPI(^^, ^^). It is also worth noting that ^^ and ^^ affect each other throughiterations of training. This interplay occurs because ^^not only serves as the surface opacity information for the final MPI, but also as occlusion information, guiding the selection of OGR pixels from multiple views, which constitutes ^^. Conversely, the dis-occluded RGB textures from the previous training iteration accurately shape the surface opacity (^^), leading to the more accurate revelation of concealed textures. Thus, the training procedure can be viewed as a joint-optimization process, where scene opacity (^^) and the RGB textures (^^) are interdependent on each other.
[0137] FIGS. 29A-29B pictorially illustrate the rendered results corresponding to a loss function configuration not employing the occlusion loss ^¢££^¤^^¢Aaccording to one example. More specifically, FIG. 29A depicts the mean absolute difference (MAD) map between the rendered view and the reference view. FIG. 29B depicts the rendered occlusion map. Multi- view-based training facilitates MPI representations that yield high-quality rendering on novel views. However, this approach can inadvertently introduce opacities optimized for rendering on expansive poses, potentially compromising the faithful restoration of the source view. This effect is demonstrated in FIGS. 29A, 29B.
[0138] In the example shown, the MAD map between the source reference view and the rendered view is as:−where ^^̄ is a rendered view=^^^and 6^ = ^ ∏^^^ ^^ ^^ ∙ :^^ P1 − ^:Q (32)Note that Eqs. (31) and (32) are analogous to Eqs. (6) and (7), with the target camera pose beingset as the source (^ = ^). Brighter locations indicate regions with a higher error.
[0139] The rendered as:e°^ ^ ^^ = ^^^ e^6^ (33)Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP where e^follows the definition given in Eq. (10). The e°^measures the amount of ray obstruction present at the source view. The brighter locations indicate regions with a higher level of obstructions. As illustrated by FIGS. 29A-29B, the MAD map and the rendered occlusion map show a high degree high correlation, which indicates that the rendering error mainly comes from unnecessary ray obstruction.
[0140] FIGS. 30A-30C show scatter plots for the MAD and e°^of FIGS. 29A-29B for three color channels (R, G, B) according to an example.
[0141] FIG. 31 shows the histogram of the occlusion map e°^illustrated in FIG. 29B according to an example. As can be seen in FIG. 31, most of the values of e°^are around 0, indicating the overall sparsity expected for the ray obstruction map at the source view. On these low values of e°^, despite having many pixels, the error (MAD) distributions are highly compacted to lower values. In contrast, as the e°^value gets larger, the error distribution becomes wider, and the overall error magnitude tends to get bigger.
[0142] This result aligns well with the visual tendencies observed in FIGS. 29A-29B where regions with higher ray obstruction in e°^mainly contribute to the rendering error, MAD, on the source view. This alignment suggests that reducing the ray obstruction at the source view can improve the rendering results. Thus, the ^¢££^¤^^¢Aconstraint is used to sparsify e°^, which is formulated as follows:Ideally, the occluded regions are not revealed from the source view. Therefore, through having ^¢££^¤^^¢A, we provide supervision for ^^to reduce the extent of unnecessary ray obstructions at the source view.
[0143] FIGS. 32A-32B graphically illustrate the effect of ^¢££^¤^^¢Aon the rendered occlusion map (e°^) according to one example. More specifically, FIG. 32A shows the rendered occlusion map (e°^) from the alpha (^^) trained without the occlusion loss ^¢££^¤^^¢A. FIG. 32B shows the rendered occlusion map (e°^) from the alpha (^^) trained with the occlusion loss ^¢££^¤^^¢A. Comparison of FIGS. 32A-32B reveals how training with the ^¢££^¤^^¢Aterm included has effectively shaped ^^to remove unnecessary ray obstruction at the source view. This also manifests itself as improvements in the source view render fidelity results, as indicated in the table below, which presents PSNR and SSIM values computed on 30 frames of anDolby Laboratories Licensing Corporation C17491EP0 / D23150EP example (“barn”) content. We see a significant improvement on the PSNR and SSIM of the MPIs trained with the ^¢££^¤^^¢A. We find that this ^¢££^¤^^¢Aplays a significant role in improving the render fidelity on the source view and maintaining the global brightness of the scene on novel camera poses. PSNR [dB] SSIM Without ^¢££^¤^^¢A28.22 0.9557 With ^¢££^¤^^¢A44.40 0.9985 On Temporal Stability
[0144] Unlike the case of static images, when creating Multi-Plane Images (MPIs) or any other volumetric representation for video, it is important to pay attention to temporal stability. Representations that are individually tuned for each frame may lack coherence when played sequentially, and if we introduce a rendering path with different viewpoints from the 3D scene, it could produce video artifacts, such as flickering. In this sub-section, we present factors that affect the temporal stability and improvements that can address (e.g., mitigate) such factors.
[0145] Regarding the temporal consistency of the depth (or disparity) map, we note that the method (400) uses the depth information of the scene that can be given in terms of the depth map or disparity map. Consequently, the temporal stability of these maps plays an important role for constructing a representation that is temporally coherent.
[0146] Video contents that provide matching depth information that are from the actual depth sensors readings will face less of these temporal consistency issues. However, in a typical scenario, a starting point is a set of image / videos without the depth information, and one can use various tools to estimate the depth or disparity from that starting point.
[0147] There exist various single view-based depth estimation methods that use the latest deep learning models, such as transformer or diffusion models. However, those methods rely on the information from a single view and have limitations in giving coherent results for multi- views or across multiple frames. For example, for some prior art methods, we found issues regarding these temporal and multi-view consistencies and encountered artifacts, such as flickering or inaccurate dis-occluded textures.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0148] There also exist methods that take multiple view images to construct depth information that is consistent across the multiple views. However, many of such methods do not consider the depth consistency issues across many frames. Thus, to ensure the stability of the method (400) across a wide range of content, it is important to interface the method (400) with suitable depth estimation methods that consider consistency across multiple views and frames.
[0149] For example, during the training stage, the training procedure (2300) may be configured to update the network module weights to enhance their ability to generate accurate MPI representations for the data they are trained on. While training separate weights for individual frames may give optimal rendering results for each static frame, it can lead to undesirable temporal artifacts, such as flickering, as it has not considered coherency across different frames. Also, under this approach, it may take a longer time to train weights for each and every one of the frames separately.
[0150] To address these issues, the training procedure (2300) can be modified to expose the network modules to multiple frames, thereby enabling the weights to learn consistent patterns that span multiple frames for mapping the RGBD inputs to an output MPI representation. For example, in one experiment, we covered a 1-second segment (30 frames) of the scene using a single set of weights. For every training iteration, we randomly select one frame from the pool of 30 frames and update the weights based on supervisions from four randomly selected target views out of the 15 considered views. This process is repeated for a specified number of iterations (2522). In one example, this specified number is 2000. It is also worth noting that the duration which a single set of weights can effectively cover may vary depending on the scene characteristics, such as motion complexity or scene change occurrences.
[0151] Assuming that we have sufficient GPU memory to cover multiple frames concurrently during a single training iteration, one feasible approach is to add another loss term, ^^Bxu¢DC^, to the loss function ^^¢^C^of Eq. (25). In one example, the loss-function component ^^Bxu¢DC^is a loss term constructed to promote inter-frame consistency of MPI layers’ disparities formulated as:^Bxu¢DC^ ^^ ^^^where^²³ ED^^ − ^²³ ED^^Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0152] Here, we have a new notation F^^ED,^^, with an additional term ¶^ in the subscript. This term refers to the disparity of the ^-th MPI layer at ¶^-th frame. The ^EDrefers to the number of frames dealt concurrently during a single training iteration. The ±^values in Eq. (36) measure the variance of the ^-th layer’s disparity across different frames. The ^^Bxu¢DC^in Eq. (35) calculates the average of these ±^values with respect to the number of layers. Thus, byadding ^^Bxu¢DC^to the final loss to ^^¢^C^, we are effectively minimizing the variations of the disparity vectors on different frames. In other words, the MPI layers’ disparities are guided to be consistent throughout the frame progression. Example Hardware
[0153] FIG. 33 is a block diagram illustrating a computing device (3300) used to implement the method (400) and / or the training procedure (2300) according to various embodiments. The computing device (3300) comprises input / output (I / O) devices (3310), a processing engine (3320), and a memory (3330). The I / O devices (3310) may be used to enable the device (3300) to receive various input signals (3302) and to output various output signals (3304). For example, the I / O devices (3310) may be operatively connected to send and receive signals via a communication channel.
[0154] The memory (3330) may have buffers to receive data. Once the data are received, the memory (3330) may provide parts of the data to the processing engine (3320) for processing therein. The processing engine (3320) includes a processor (3322) and a memory (3324). The memory (3324) may store therein program code, which when executed by the processor (3322) enables the processing engine (3320) to perform various data processing operations, including but not limited to at least some operations associated with the method (400) and / or the training procedure (2300).
[0155] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-33, provided is a method of generating an MPI representation of a scene comprising: generating a first MPI representation of the scene based on a first image of the scene; constructing an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtaining respective sets of pixel values corresponding to the occlusion map from one or more second images of the scene, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and generating aDolby Laboratories Licensing Corporation C17491EP0 / D23150EP second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
[0156] In some embodiments of the above method, the method further comprises obtaining disparity information corresponding to the first image and the one or more second images, wherein the selectively filling includes: for a layer of the first MPI representation, selecting, from the respective sets of the pixel values, one or more subsets for which the disparity information indicates substantial matching of disparity between the first image and a respective one of the one or more second images; and filling an occluded portion of the layer using the one or more subsets.
[0157] In some embodiments of any of the above methods, the selecting includes applying spatial lowpass filtering to select the one or more subsets from a plurality of candidate subsets corresponding to different ones of the second images.
[0158] In some embodiments of any of the above methods, the method further comprises generating an inter-layer texture filler for the first MPI representation, wherein the generating of the second MPI representation is further based on the inter-layer texture filler.
[0159] In some embodiments of any of the above methods, the obtaining comprises, for each image of the one or more second images of the scene: warping the occlusion map to the respective camera pose of the image, determining a first set of pixel values by copying pixel values of the image to a region defined by the warped occlusion map, computing a fidelity map based on a deviation between (a) a depth representation of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a depth representation map of the image, and constructing a second set of pixel values by modifying the first set of pixel values by applying depth-based pixel removal based on the computed deviation. In some embodiments of any of the above methods, when processing depth information based on disparity, the obtaining comprises: warping the occlusion map to the respective camera pose of a selected second image; forming a first OGR set using pixel values of the selected second image from a region defined by the warped occlusion map; constructing a disparity fidelity map of the selected second image based on disparity information corresponding to the first image and the selected second image; and constructing a second OGR set by modifying the first OGR set based on the disparity fidelity map.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0160] In some embodiments of any of the above methods, the obtaining further comprises: warping the second OGR set back to the camera pose corresponding to the first image; and warping the disparity fidelity map of the selected second image to the camera pose corresponding to the first image.
[0161] In some embodiments of any of the above methods, the generating of the second MPI representation comprises: generating a combined OGR set based on a plurality of the warped second OGR sets corresponding to a plurality of second images; and generating a combined disparity fidelity map based on the warped disparity fidelity maps of the plurality of second images.
[0162] In some embodiments of any of the above methods, the generating of the second MPI representation is based on the combined OGR set and the combined disparity fidelity map.
[0163] In some embodiments of any of the above methods, the modifying comprises removing from the first OGR set pixel values for which a corresponding disparity fidelity value is smaller than a threshold value.
[0164] In some embodiments of any of the above methods, the threshold value is layer- dependent.
[0165] In some embodiments of any of the above methods, training a plurality of neural network modules on the first image and the one or more second images at least to: perform the generating of the first MPI representation; and provide guidance to the constructing and the obtaining.
[0166] In some embodiments of any of the above methods, the training comprises using one or more loss functions selected from the group consisting of: an occlusion loss; a disparity loss; and a render fidelity loss.
[0167] In some embodiments of any of the above methods, the training comprises using a loss function including a weighted sum of an occlusion loss, a disparity loss, and a render fidelity loss.
[0168] In some embodiments of any of the above methods, the occlusion loss is computed using alpha values corresponding to the first image of the scene to provide supervision on the alpha values configured to reduce or minimize ray obstruction.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0169] In some embodiments of any of the above methods, the disparity loss is computed using disparity vectors corresponding to the first image of the scene, a disparity map corresponding to the first image of the scene, and an alpha shape mask corresponding to the first image of the scene and is configured to provide constraints on adaptive disparity distances and further provide disparity-based constraints on an alpha shape.
[0170] In some embodiments of any of the above methods, the render fidelity loss is configured to provide supervision on the second MPI representation through a comparison with ground truth of target views.
[0171] In some embodiments of any of the above methods, the training comprises selecting weight values for the weighted sum based on an optimization criterion.
[0172] In some embodiments of any of the above methods, the loss function is configured to jointly optimize alpha and texture channels of the second MPI representation.
[0173] In some embodiments of any of the above methods, the training includes training the plurality of neural network modules on a video sequence including: a sequence of first images, each corresponding to a different respective time in the video sequence; and a sequence of corresponding sets of second images, each of the sets being associated with a respective one of the first images.
[0174] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.
[0175] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-33, provided is an apparatus for generating an MPI representation of a scene, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a first MPI representation of the scene based on a first image of the scene; construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtain respective sets of pixel values corresponding to the occlusion map from one or more second images of the scene, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding toDolby Laboratories Licensing Corporation C17491EP0 / D23150EP the first image; and generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
[0176] In some embodiments of the above apparatus, the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to generate an inter-layer texture filler for the first MPI representation, wherein the generating of the second MPI representation is further based on the inter-layer texture filler.
[0177] In some embodiments of any of the above apparatus, to obtain the respective sets of pixel values corresponding to the occlusion map, the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to: warp the occlusion map to the respective camera pose of a selected second image; form a first OGR set using pixel values of the selected second image from a region defined by the warped occlusion map; construct a disparity fidelity map of the selected second image based on disparity information corresponding to the first image and the selected second image; and construct a second OGR set by modifying the first OGR set based on the disparity fidelity map.
[0178] In some embodiments of any of the above apparatus, the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to train a plurality of neural network modules on the first image and the one or more second images at least to: perform the generating of the first MPI representation; and provide guidance to the constructing and the obtaining.
[0179] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-33, provided is an apparatus for generating an MPI representation of a scene, the apparatus comprising: a plurality of neural networks; an MPI compositor configured to generate a first MPI representation of the scene based on a first set of outputs generated by the plurality of neural networks in response to a first image of the scene; and an occlusion guided residuals (OGR) generator configured to construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation, the occlusion map being constructed based on a second set of outputs generated by the plurality of neural networks, wherein the plurality of neural networks is configured to obtain, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second imagesDolby Laboratories Licensing Corporation C17491EP0 / D23150EP corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and wherein the MPI compositor is further configured to generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
[0180] In some embodiments of the above apparatus, the plurality of neural networks includes a first neural network configured to generate feature maps in response to the first image of the scene; and wherein the plurality of neural networks is configured to generate the first and second sets of outputs based on the feature maps.
[0181] In some embodiments of any of the above apparatus, the first neural network includes a down-sampling branch of a U-net architecture.
[0182] In some embodiments of any of the above apparatus, the plurality of neural networks includes a second neural network configured to generate an alpha shape mask and a disparity vector in response to the first image of the scene; and wherein the plurality of neural networks is configured to generate the first and second sets of outputs based on the alpha shape mask.
[0183] In some embodiments of any of the above apparatus, the second neural network includes: a disparity vector prediction module configured to generate the disparity vector in response to the first image of the scene; and a mask generation module configured to generate the alpha shape mask in response to the first image of the scene and the disparity vector.
[0184] In some embodiments of any of the above apparatus, the plurality of neural networks includes a third neural network configured to generate alpha values corresponding to the first image of the scene based on the feature maps and the alpha shape mask; and wherein each of the first and second sets of outputs includes the alpha values.
[0185] In some embodiments of any of the above apparatus, the third neural network includes a decoder component of the U-net architecture.
[0186] In some embodiments of any of the above apparatus, the plurality of neural networks includes a fourth neural network configured to generate an interlayer texture filler corresponding to the first image of the scene based on the feature maps; and wherein the first set of outputs includes the interlayer texture filler.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0187] In some embodiments of any of the above apparatus, the fourth neural network includes another decoder component of the U-net architecture.
[0188] In some embodiments of any of the above apparatus, the plurality of neural networks includes a fifth neural network configured to generate one or more adjusted camera poses in response to camera pose information corresponding to the one or more second images of the scene; and wherein the second set of outputs includes the one or more adjusted camera poses.
[0189] In some embodiments of any of the above apparatus, the fifth neural network includes first and second parallel branches configured to perform a translation adjustment of a pose and a rotation adjustment of the pose, respectively.
[0190] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0191] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0192] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0193] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
[0194] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.
[0195] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0196] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0197] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.
[0198] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0199] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0200] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”
[0201] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.
[0202] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP
[0203] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
[0204] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.
[0205] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0206] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such asDolby Laboratories Licensing Corporation C17491EP0 / D23150EP a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0207] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0208] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP CLAIMS What is claimed is:
1. A method of generating a multiplane image (MPI) representation of a scene, the method comprising: generating a first MPI representation of the scene based on a first image of the scene; constructing an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtaining, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and generating a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
2. A method according to claim 1, wherein the obtaining comprises, for each image of the one or more second images of the scene: warping the occlusion map to the respective camera pose of the image; determining a first occlusion guided residuals (OGR) set by copying pixel values of the image to a region defined by the warped occlusion map; computing a depth fidelity map based on a deviation between (a) a depth representation of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a depth representation map of the image, and constructing a second OGR set by modifying the first set of pixel values by applying depth-based pixel removal based on the computed deviation.
3. The method of claim 2, further comprising warping the depth fidelity map for each of the one or more second images to the camera pose of the first image to obtain a warped depth fidelity map; wherein the selectively filling includes: for a layer of the first MPI representation, comparing the warped fidelity maps for the one or more second images of the scene, and selecting, from the respective sets of the pixel values, one or more subsets for which the warped fidelity map value indicates substantial matching ofDolby Laboratories Licensing Corporation C17491EP0 / D23150EP depth representation between the first image and a respective one of the one or more second images; and filling an occluded portion of the layer using the one or more subsets.
4. The method of claim 3, further comprising obtaining disparity information corresponding to the first image and the one or more second images, wherein the depth representation of each layer of the first MPI representation comprises a disparity value corresponding to each layer of the first MPI representation, wherein, for each of the second images, the depth fidelity map is a disparity fidelity map computed based on a deviation between (a) a disparity value of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a disparity map of the image.
5. The method of claim 3, wherein the selecting includes applying spatial lowpass filtering to select the one or more subsets from a plurality of candidate subsets corresponding to different ones of the second images.
6. The method of any preceding claim, further comprising generating an inter-layer texture filler for the first MPI representation, wherein the generating of the second MPI representation is further based on the inter- layer texture filler.
7. The method of claim 4 or any claim dependent thereon, wherein the obtaining further comprises: warping the second OGR set back to the camera pose corresponding to the first image; and wherein the generating of the second MPI representation comprises: generating a combined OGR set based on a plurality of the warped second OGR sets corresponding to a plurality of second images; and generating a combined disparity fidelity map based on the warped disparity fidelity maps of the plurality of second images; and wherein the generating of the second MPI representation is based on the combined OGR set and the combined disparity fidelity map.
8. The method of claim 7,Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP wherein the modifying comprises removing from the first OGR set pixel values for which a corresponding disparity fidelity value is smaller than a threshold value; and wherein the threshold value is layer-dependent.
8. The method of any of claims 1 to 7, further comprising training a plurality of neural network modules on the first image and the one or more second images at least to: perform the generating of the first MPI representation; and provide guidance to the constructing and the obtaining.
9. The method of claim 8, wherein the training comprises using one or more loss functions selected from the group consisting of: an occlusion loss; a disparity loss; and a render fidelity loss.
10. The method of claim 9, wherein the occlusion loss is computed using alpha values corresponding to the first image of the scene to provide supervision on the alpha values configured to reduce or minimize ray obstruction.
11. The method of claim 9 or 10, wherein the disparity loss is computed using disparity vectors corresponding to the first image of the scene, the disparity vectors indicating a disparity of each layer of the first MPI representation, a disparity map corresponding to the first image of the scene, and an alpha shape mask corresponding to the first image of the scene and is configured to provide constraints on adaptive disparity distances and further provide disparity- based constraints on an alpha shape.
12. The method of any of claims 9 to 11, wherein the render fidelity loss is configured to provide supervision on the second MPI representation through a comparison with ground truth of target views.
13. The method of any of claims 9 to 12, wherein the training comprises: using a loss function including a weighted sum of an occlusion loss, a disparity loss, and a render fidelity loss; andDolby Laboratories Licensing Corporation C17491EP0 / D23150EP selecting weight values for the weighted sum based on an optimization criterion; and wherein the loss function is configured to jointly optimize alpha and texture channels of the second MPI representation.
14. The method of any of claims 8 to 13, wherein the training includes training the plurality of neural network modules on a video sequence including: a sequence of first images, each corresponding to a different respective time in the video sequence; and a sequence of corresponding sets of second images, each of the sets being associated with a respective one of the first images.
15. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any of claims claim 1 to 14.
16. An apparatus for generating an MPI representation of a scene, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a first MPI representation of the scene based on a first image of the scene; construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation; obtain respective sets of pixel values corresponding to the occlusion map from one or more second images of the scene, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
17. The apparatus of claim 16, wherein the obtaining comprises, for each image of the one or more second images: warping the occlusion map to the respective camera pose of the image;Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP determining a first occlusion guided residuals (OGR) set by copying pixel values of the image to a region defined by the warped occlusion map; computing a depth fidelity map based on a deviation between (a) a depth representation of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a depth representation map of the image, and constructing a second OGR set by modifying the first set of pixel values by applying depth-based pixel removal based on the computed deviation.
18. The apparatus of claim 17, wherein the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to: warp the depth fidelity map for each of the one or more second images to the camera pose of the first image to obtain a warped depth fidelity map; wherein the selectively filling includes: for a layer of the first MPI representation, comparing the warped fidelity maps for the one or more second images of the scene, and selecting, from the respective sets of the pixel values, one or more subsets for which the warped fidelity map value indicates substantial matching of the depth representation between the first image and a respective one of the one or more second images; and filling an occluded portion of the layer using the one or more subsets.
19. The apparatus of claim 18, wherein the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to obtain disparity information corresponding to the first image and the one or more second images; wherein the depth representation of each layer of the first MPI representation comprises a disparity value corresponding to each layer of the first MPI representation, wherein, for each of the second images, the depth fidelity map is a disparity fidelity map computed based on a deviation between (a) a disparity value of each layer of the first MPI representation corresponding to the camera pose of the image and (b) a disparity map of the image.
20. The apparatus of any one of claims 16 to 19, wherein the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to generate an inter-layer texture filler for the first MPI representation, wherein the generating of the second MPI representation is further based on the inter-layer texture filler.Dolby Laboratories Licensing Corporation C17491EP0 / D23150EP 21. The apparatus of any of claims 16 to 20, wherein the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to train a plurality of neural network modules on the first image and the one or more second images at least to: perform the generating of the first MPI representation; and provide guidance to the constructing and the obtaining.
22. An apparatus for generating an MPI representation of a scene, the apparatus comprising: a plurality of neural networks; an MPI compositor configured to generate a first MPI representation of the scene based on a first set of outputs generated by the plurality of neural networks in response to a first image of the scene; and an occlusion guided residuals (OGR) generator configured to construct an occlusion map that indicates respective occluded portions for different layers of the first MPI representation, the occlusion map being constructed based on a second set of outputs generated by the plurality of neural networks; wherein the plurality of neural networks is configured to obtain, from one or more second images of the scene, respective sets of pixel values corresponding to the occlusion map, each of the second images corresponding to a respective camera pose that is different from a camera pose corresponding to the first image; and wherein the MPI compositor is further configured to generate a second MPI representation of the scene by selectively filling parts of the respective occluded portions of the different layers of the first MPI representation using the respective sets of the pixel values.
21. The apparatus of claim 20, wherein the plurality of neural networks includes a first neural network configured to generate feature maps in response to the first image of the scene; and wherein the plurality of neural networks is configured to generate the first and second sets of outputs based on the feature maps.
22. The apparatus of claim 20 or 21, wherein the first neural network includes a down- sampling branch of a U-net architecture.
23. The apparatus of any one of claims 20-22, wherein the plurality of neural networks includes a second neural network configured to generate an alpha shape mask and a disparityDolby Laboratories Licensing Corporation C17491EP0 / D23150EP vector in response to the first image of the scene the disparity vectors indicating a disparity of each layer of the first MPI representation; and wherein the plurality of neural networks is configured to generate the first and second sets of outputs based on the alpha shape mask.
24. The apparatus of any one of claims 20-23, wherein the second neural network includes: a disparity vector prediction module configured to generate the disparity vector in response to the first image of the scene; and a mask generation module configured to generate the alpha shape mask in response to the first image of the scene and the disparity vector.
25. The apparatus of any one of claims 20-24, wherein the plurality of neural networks includes a third neural network configured to generate alpha values corresponding to the first image of the scene based on the feature maps and the alpha shape mask; and wherein each of the first and second sets of outputs includes the alpha values.
26. The apparatus of claim 25, wherein the third neural network includes a decoder component of the U-net architecture.
27. The apparatus of any one of claims 20-26, wherein the plurality of neural networks includes a fourth neural network configured to generate an interlayer texture filler corresponding to the first image of the scene based on the feature maps; and wherein the first set of outputs includes the interlayer texture filler.
28. The apparatus of claim 27, wherein the fourth neural network includes another decoder component of the U-net architecture.
29. The apparatus of any one of claims 20-28, the plurality of neural networks includes a fifth neural network configured to generate one or more adjusted camera poses in response to camera pose information corresponding to the one or more second images of the scene; and wherein the second set of outputs includes the one or more adjusted camera poses.
30. The apparatus of claim 29, wherein the fifth neural network includes first and second parallel branches configured to perform a translation adjustment of a pose and a rotation adjustment of the pose, respectively.
Citation Information
Patent Citations
Free-viewpoint photorealistic view synthesis from casually captured video
US20200226816A1