Layered surface lightfields

Layered surface lightfields address the challenge of delivering high-quality 3D video content by efficiently encoding and rendering volumetric scenes with reduced bandwidth, achieving improved visual fidelity and real-time processing.

WO2025184103A1PCT designated stage Publication Date: 2025-09-04DOLBY LABORATORIES LICENSING CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017222
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-15
Filing Date
2025-02-25
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing video encoding technologies struggle to deliver high-quality 3D video content efficiently, leading to inaccurate representations and visible artifacts due to limited bandwidth and device capabilities.

Method used

The use of layered surface lightfields (LSLFs) to represent volumetric scenes, which combine main and additional surface layers with depth, transparency, and view-dependent visual information, allowing for efficient encoding and rendering of immersive videos with reduced bandwidth usage.

Benefits of technology

LSLFs enable high-quality rendering of volumetric videos with reduced bandwidth, supporting real-time or near-real-time processing and higher fidelity images/videos, overcoming limitations of traditional methods like ray marching rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017222_04092025_PF_FP_ABST
    Figure US2025017222_04092025_PF_FP_ABST
Patent Text Reader

Abstract

An input lightfield comprising a spatial distribution of visual and geometric information in a volumetric scene in a three-dimensional (3D) space is received. A spatial sequence of two-dimensional (2D) surface layers is selected to be included in a layered surface lightfield (LSLF) set. The input lightfield is collapsed into the spatial sequence of 2D surface layers of the LSLF set. Each 2D surface layer in the spatial sequence of 2D surface layers includes a depth data portion, a transparency data portion, a color data portion, etc. The LSLF set is encoded into a coded bitstream. The coded bitstream causes a recipient device of the coded bitstream to generate a display image of the volumetric scene for rendering on an image display.
Need to check novelty before this filing date? Find Prior Art

Description

LAYERED SURFACE LIGHTFIELDSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 559,371, filed February 29, 2024 and European Patent Application No. 24176142.8, filed May 15, 2024, each of which is incorporated by reference herein in its entirety.TECHNOLOGY

[0002] The present invention relates generally to lightfield coding and more particularly to layered surface lightfields.BACKGROUND OF THE INVENTION

[0003] Video techniques are being developed to support transmitting and rendering volumetric video content based on available bandwidths supported by contemporary computing and network infrastructure. For example, video encoders and decoders may be extended or developed to support encoding and decoding three-dimensional (3D) video content for rendering with computing devices incorporating non-MPEG codecs.

[0004] An end user computing device may be installed or configured with video codecs of relatively limited capabilities. If 3D video content is not encoded and delivered with relatively high quality and limited bandwidth usage, even if rendered, decoded 3D video content may include incorrect or inaccurate interpretation or representation of the original 3D video content, and may produce visible artifacts in shapes, colors and luminance values in rendered images.

[0005] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.BRIEF DESCRIPTION OF DRAWINGS

[0006] The present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0007] FIG. 1A illustrates example two-dimensional (2D) intersection surfaces for lightfield parametrization; FIG. IB illustrates example surface layer portions occluded in a target view but visible from neighboring views; FIG. 1C illustrates an example layered surface lightfield (LSLF) set; FIG. ID and FIG. IE illustrate an example scene represented by three surface layers of LSLF sets customized to two different views;

[0008] FIG. 2A illustrates an example surface layer representation with overlayering; FIG. 2B through FIG. 2D illustrate example LSLF sets with target specialization;

[0009] FIG. 3A illustrates an example process flow for LSLF data generation and encoding operations; FIG. 3B and FIG. 3C illustrate example process flows for LSLF rendering operations;

[0010] FIG. 4A and FIG. 4B illustrate example process flows; and

[0011] FIG. 5 illustrates an example hardware platform on which a computer or a computing device as described herein may be implemented.DETAILED DESCRIPTION OF THE INVENTION

[0012] Example embodiments, which relate to layered surface lightfields, are described herein. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present invention.

[0013] Example embodiments are described herein according to the following outline:1. GENERAL OVERVIEW2. SURFACE RADIANCE FIELD AND LAYERED SURFACE LIGHTFIELD3. VIEW SYNTHESIS AND PLANAR SURFACE LIGHTFIELD4. NON-PLANAR SURFACE LIGHTFIELD5. EXAMPLE LSLF SETS6. LSLF DATA GENERATION AND ENCODING7. OVERLAYERING8. TARGET SPECIALIZATION9. TARGET CUSTOMIZATION10. LSLF FEATURES AND OPERATIONS11 . LSLF DATA DECODING AND RENDERING12. EXAMPLE PROCESS FLOWS13. IMPLEMENTATION MECHANISMS - HARDWARE OVERVIEW14. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS1 . GENERAL OVERVIEW

[0014] This overview presents a basic description of some aspects of an example embodiment of the present invention. It should be noted that this overview is not an extensive or exhaustive summary of aspects of the example embodiment. Moreover, it should be noted that this overview is not intended to be understood as identifying any particularly significant aspects or elements of the example embodiment, nor as delineating any scope of the example embodiment in particular, nor the invention in general. This overview merely presents some concepts that relate to the example embodiment in a condensed and simplified format, and should be understood as merely a conceptual prelude to a more detailed description of example embodiments that follows below.

[0015] Techniques as described herein can be implemented or performed to project or convert input or source lightfields of relatively large data sizes to layered surface lightfields. The resultant layered surface lightfields may be used to enable encoding and transmitting volumetric or immersive videos / images with relatively high quality at relatively low bitrates.

[0016] A lightfield as described herein refers to a description of, or a data set describing or specifying, some or all light in a given volume or three-dimensional (3D) scene. A full lightfield specifies all light for all points in the volume or scene. In practice, it is relatively difficult if not impossible to capture a full lightfield of a volume or 3D scene.

[0017] Instead of using a full lightfield representation, a volumetric scene may be represented by capturing or representing light rays that pass through a two-dimensional (2D)plane between a viewer and the volume or scene. Light information of the lightfield can then be defined by, or reduced to, positions on the (2D) plane and directions, intensities, colors, luminances, or chrominances of the light rays passing through the positions on the 2D plane.

[0018] Surface lightfields under techniques as described herein can help overcome constraints and limitations of 2D planes. Under these techniques, surfaces (or surface layers) of arbitrary shapes can be used in capturing or generating a layered surface lightfield (LSLF) set in place of a full lightfield.

[0019] Some or all surface layers in an LSLF set representing a lightfield depicting a volumetric or immersive scene may spatially conform or adjacent to, or curvilinear with, real or actual surfaces of physical or visual objects depicted in the scene. Hence, the viewer is enabled to closely and fully explore the volumetric scene or physical or visual objects therein - with relatively high spatial resolution - without risk of crossing these surfaces on which light information and other attendant visual or non-visual information such as depth and transparency is represented or captured.

[0020] Surface or surface layers included or defined in an LSLF set as described herein represent corresponding surface or 2D lightfields in terms of where and with what spatial directions light rays pass through each of the surfaces or locations thereon. Light rays passing through such a surface (lightfield) may be direction- or view-dependent light including but not limited to specular reflections. The direction- or view-dependent light may be only or more visible from certain viewpoints (or user view directions) and may exhibit different intensities and colors to different view directions.

[0021] Under techniques as described herein, in some operational scenarios, direction- or view-dependent light or visual information in surface lightfields can be represented or recorded with spherical harmonic (SH) coefficients specified for a set of SH basis functions up to a maximum SH basis function order. For the purpose of illustration only, as an example, SH basis functions and coefficients may be used to represent or capture view- or direction-dependent light information. It should be noted that in other operational scenarios, non-SH basis functions may be used as represent or capture view- or direction-dependent light information instead of or in addition to SH basis functions.

[0022] Direction- or view-dependent visual information in a volumetric scene may be spatially local, such as small reflective surface portions on a reflective surface of an airplane or a car.

[0023] Overlayering (techniques) as described herein may be applied to generate asequence of multiple additional layered surface lightfields (or additionally surface layer portions) on top of each other in addition to a main surface light field (or main surface layer portion). These additional layered surface lightfields correspond to or at different or separate depths from the same main surface light field. Each of the main or additional layered surface light fields (or portions) can carry its respective depth data, transparency / opacity values, direction- or view-dependent visual or color information. Some of the direction- or viewdependent visual or color information in the main or additional surface layer portions may be represented with SH coefficients for each of color components or channels of a selected (e.g., RGB, YUV, YCbCr, IPT, etc.) color space.

[0024] In comparison with using only a main layered surface lightfield, the combination of main and additional layered surface lightfields as described herein can represent a lightfield with relatively high precision exceeding or comparable to that represented by a single surface lightfield with relatively large numbers of spherical harmonic coefficients / orders. To reduce or prevent data redundancy and excessive bandwidth usage, the overlayering techniques may (e.g., only, etc.) be utilized on certain parts of certain surfaces needing relatively high levels of view dependency visual information.

[0025] As a result, a relatively low maximum SH basis function order can be used in an LSLF set as described herein to carry direction- or view-dependent visual information, thereby reducing or capping an overall data size of the LSLF set to represent a volumetric scene.

[0026] In some operational scenarios, multiple LSLF sets with multiple different target views and / or with multiple different spatial projections / mappings may (e.g., actually, possibly, etc.) be generated by an upstream device for a given time point of a volumetric scene.

[0027] User viewing direction data relating to real time or near real time view directions of a viewer may be monitored, tracked or collected by a downstream device and provided to the upstream device.

[0028] From the user viewing direction data, the upstream device can determine, predict or estimate a specific view direction (e.g., a novel direction not coinciding with any of the target views, a to-be-synthesized or -reconstructed view direction, etc.).

[0029] Based at least in part on the specific view direction determined from the user view direction data, the upstream device can customarily select or generate a specific LSLF set with a specific target view (among the multiple different target views) and a specific spatialprojection / mapping (among the multiple different spatial projections / mappings) to enable the downstream device to generate or reconstruct a relatively high quality and / or high resolution display image for rendering on an image display operating in conjunction with the downstream device.

[0030] Hence, under techniques as described herein, visual quality of volumetric or immersive videos / images as rendered by end user devices can be significantly improved without excessive bandwidth usages or bitrates of volumetric video content.

[0031] As used herein, volumetric or immersive videos / images may, but are not necessarily limited to only, relate to any of: audiovisual programs, movies, video programs, TV broadcasts, computer games, augmented reality (AR) content, virtual reality (VR) content, mixed reality (MR) content, remote presence content, automobile entertainment content, etc.

[0032] As used herein, “upstream device” may refer to one or more computing devices that prepare and stream (e.g., immersive, volumetric, etc.) video content to one or more downstream client or end-user computing devices such as video decoders in order to render at least a portion of the video content on one or more image displays.

[0033] Example upstream and / or downstream devices may include, but are not necessarily limited to, any of: cloud-based video streaming servers located remotely from video streaming client(s), local video streaming servers connected with video streaming client(s) over local wired or wireless networks, VR devices, AR devices, automobile entertainment devices, digital media devices, digital media receivers, set-top boxes, gaming machines (e.g., an Xbox), general purpose personal computers, tablets, dedicated digital media receivers such as the Apple TV or the Roku box, etc.

[0034] Example embodiments described herein relate to encoding 3D visual content. An input lightfield comprising a spatial distribution of visual and geometric information in a volumetric scene in a three-dimensional (3D) space is received. A spatial sequence of two- dimensional (2D) surface layers is selected to be included in a layered surface lightfield (LSLF) set. The input lightfield is converted, mapped, collapsed or reduced into the spatial sequence of 2D surface layers of the LSLF set. Each 2D surface layer in the spatial sequence of 2D surface layers includes a depth data portion, a transparency data portion, a color data portion, etc. The LSLF set is encoded into a coded bitstream. The coded bitstream causes a recipient device of the coded bitstream to generate a display image of the volumetric scene for rendering on an image display.

[0035] Example embodiments described herein relate to decoding 3D visual content. A layered lightfield (LSLF) set is decoded from a coded bitstream. The LSLF set includes a spatial sequence of two-dimensional (2D) surface layers with respective depth data portions, transparency data portions, color data portions, etc. The LSLF set is used to generate a display image. The display image is rendered on an image display.

[0036] In some example embodiments, mechanisms as described herein form a part of a media processing system, including but not limited to any of: a wearable device, a handheld computing device, game machine, television, laptop computer, netbook computer, tablet computer, desktop computer, computer workstation, computer kiosk, or various other kinds of computing devices and media processing units.

[0037] Various modifications to the preferred embodiments and the generic principles and features described herein will be readily apparent to those skilled in the art. Thus, the disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.2. LAYERED SURFACE LIGHTFIELD

[0038] Layered surface lightfield may be used to implement operations, methods or algorithms for rendering volumetric or immersive video, for example represented in layered surface lightfield (LSLF) data sets.

[0039] The LSLF data sets provide a relatively high-quality representation of the immersive or volumetric video. The LSLF rendering operations, methods or algorithms may be implemented or performed based on the LSLF data sets, in place of or in addition to other rendering methods / algorithms such as ray marching rendering, NeRF, Plenoxels, DNR, MPI, etc. Example LSLF data sets and formats are described in U.S. Provisional Patent Application No. / , (Attorney Docket No. 60175-0558; D23155USP1), titled “SURFACE LIGHTFIELD COMPRESSION,” by Vijay Sundaram, Peng Yin, Guan-Ming Su, Taoran Lu, Vijay Kamarshi, filed on equal day, the contents of which are incorporated herein by reference in its entirety.

[0040] The LSLF data sets carrying a high quality representation of the immersive or volumetric video in accordance with specific LSLF formats may be processed, compressed, transmitted, decoded, reconstructed and rendered in real time or near real time. Some or all of these image / video processing and rendering operations can be implemented or performed using a wide variety of available codecs and / or GPU hardware including but not limited to those already deployed in the field or end user devices.

[0041] Depth data may be used as prior(s) (e.g., knowns, beliefs, initial conditions, etc.) along with other inputs, knowns, priors (e.g., alpha map, spherical harmonic (SH) coefficients, etc.) if available in the LSLF-based image / video processing or rendering operations, methods or algorithms. These operations, methods or algorithms can be performed to render volumetric or immersive video or views - e.g., a scene or corresponding video data covering a scene level time interval such as 10-15 milliseconds or longer, etc. - with relatively high visual quality and / or fidelity in real time or near real time,

[0042] In comparison with other approaches such as ray marching based rendering, the LSLF based operations, methods or algorithms under techniques as described herein can relatively efficiently process and render real time or near real time immersive or volumetric video, produce higher fidelity higher visual quality images / videos, and support a wide variety of target frame rates and / or target image resolutions.3. VIEW SYNTHESIS AND PLANAR SURFACE LIGHTFIELD

[0043] To view rendered or display images relating to volumetric or immersive video, a viewer may choose different views - which may be referred to as “novel views,” “synthesized views,” or “reconstructed views” - from camera views used to capture visual data of a 3D volume or scene.

[0044] Image / view synthesis or reconstruction operations refer to those performed for the purpose of creating or generating display or rendered images corresponding to novel views of a 3D volume or scene from arbitrary viewpoints or non-camera viewpoints. These image / view synthesis or reconstruction operations (or methods / algorithms) can be implemented or performed based at least in part on a lightfield representation of the 3D volume or scene such as five-dimensional (5D) plenoptic functions that are able to represent (nearly) all the light or light rays passing through the 3D volume or scene.

[0045] A complete or full (e.g., volume-based, etc.) lightfield represented in a 5D plenoptic function may provide a description of all light in the volume or scene, or a collection of all the light rays flowing from all points in all directions in the volume or scene: three dimensions for specifying (3D) location / position and two dimensions for specifying (2D) orientations / directions of light rays at any given location / position. Given this complete or full lightfield, it is sufficient to simply select rays that pass through a target or desired viewpoint within a target or desired field of view (e.g., of an image display, of a viewport, etc.) to create a corresponding image of the scene.

[0046] However, it may be impractical to capture a complete or full lightfield, even atrelatively low sampling densities, as that would need full panoramic captures over the entire volume or scene. In practice, lightfields may be captured with camera arrays with densely or sparsely populated cameras or camera elements spatially arranged in grids or constellations instead of an entire 3D volume.

[0047] Indeed, a relatively large portion of all visual data contained in a complete or full lightfield may be (e.g., highly, etc.) redundant, and much of the visual data may not be needed or used depending on the target or desired view and the viewed region of the volume or scene. Hence, an alternative formulation / representation of the lightfield that uses - or is parametrized with - only four dimensions (4D) with two-plane parameterization may be adopted in immersive or volumetric video applications.

[0048] The 4D parameterization can be based on the following observation: if a plane - such as representing that of a viewport or field of view - is placed between a viewer and the scene, any ray that the viewer sees passes through the plane. Therefore, it is sufficient to capture these intersecting rays to the plane for the purpose of rendering an image of the scene to the viewer. These (intersecting) rays can be defined or specified by their respective positions on the plane (2D) and their respective direction (2D).

[0049] This 4D parametrization approach can be implemented or used to directly map the plane to 2D camera arrays with camera (spatial) densities and resolutions corresponding to the sampling densities of the lightfield. On the other hand, the viewer may not be immersed in the scene in the same way as the complete or full (5D) light field, as the viewer and the scene under the 4D parametrization approach are to be located on opposite sides of the plane.4. NON-PLANAR SURFACE LIGHTFIELD

[0050] If a 3D volume or scene is behind a 2D (intersection) surface in reference to a viewer, any view of the volume or scene (e.g., only, etc.) by the viewer is composed of rays that pass through the 2D surface. Visual data (e.g., intensity, color, directional dependence, etc.) of some or all of these intersecting rays of some or all orientations / directions at some or all spatial locations / positions of the 2D surface may be referred to - or captured / represented - as a surface lightfield.

[0051] Surface lightfields as described herein may use 2D intersection surfaces that can have arbitrary shapes (as opposed to flat planes only). These arbitrary surfaces may or may not correspond to actual surfaces of real physical objects. Hence, a 2D intersection surface as described herein - which may be used to intersect light rays or define / specify a surface lightfield - may or may not be planar or a plane. In various operational scenarios, a 2Dintersection surface as described herein may be any 2D manifold or surface. The 2D intersection surface may or may not coincide with a (e.g., outer, exterior, etc.) surface of a physical object.

[0052] FIG. 1 A illustrates example 2D intersection surfaces for lightfield parametrization. These 2D surfaces may include, but are not necessarily limited to only, any, some or all of: such as a planar surface, a curved surface that does not correspond to any actual surface of a physical object, a surface that corresponds to (e.g., adjacent to, coincide with, etc.) an actual surface of a physical object, and so on.

[0053] As shown in FIG. 1 A, light rays or radiant light coming off (e.g., reflected from, emitted by, etc.) one or more physical objects can be parameterized into a surface lightfield using a 2D intersection surface formed with one or more actual (e.g., outer, exterior, etc.) surfaces of the physical objects or positions. The surface light field provides visual data of the light rays intersecting respective positions or locations of the 2D surface as well as respective orientations / directions of the light rays in space at these positions / locations of the 2D surface.

[0054] It should be noted that, while the actual surfaces of the physical objects forming the 2D surface may be a part of the (e.g., 3D, etc.) physical objects, the intersecting surface may be defined or specified with a 2D space coinciding, slightly (e.g., within a relatively small absolute or relative distance range, etc.) away from, or with a certain spatial separation from, a fixed surface shape of the physical objects.

[0055] Regardless of how the 2D intersecting surface is selected, there is no extra or 3rd degree of freedom available beyond surface dimensions spanning over the 2D intersecting surface. As a result, the surface lightfield can still be represented or expressed as a 4D parametrized (or 4D plenoptic) function, regardless of whether the 2D surface used to specify the surface lightfield is or includes the actual surfaces of the physical objects.

[0056] Surface lightfields can be defined, specified or generated to represent visual or light information of a (viewed) 3D volume or scene with one or more (parametrization) surfaces separating the viewed volume or scene from a spatial region (or viewing volume) in which a viewer is located. The viewer may be spatially or virtually represented in the spatial region (or viewing volume) in a depicted 3D space that also includes the viewed volume or scene. Visual depictions such as image(s) or video(s) of the viewed volume or scene may be presented, displayed or rendered to the viewer represented in the viewing volume.

[0057] The one or more (parametrization) surfaces used to define, specify or generate the (corresponding) surface lightfields may be determined, generated or selected based at least inpart on an estimate of a scene geometry of the viewed volume or scene. The estimate of the scene geometry of the viewed volume or scene can be used to constrain the placement of the one or more surfaces - e.g., separating the viewed volume or scene from the viewing volume in the 3D space - used in surface lightfield parameterization. However, this estimate need not be particularly accurate, so long as the selected or constrained (parametrization) surfaces can contain the scene geometry within its defined volume. In other words, the (parametrization) surfaces can be selected to separate the viewing volume and the viewed volume in the 3D space so the corresponding surface lightfields can be defined or specified on the (parametrization) surfaces with sufficient visual or light information and hence can function properly for constructing visual images of the viewed volume or scene.

[0058] In some operational scenarios, one or more depth estimation algorithms, methods, and / or operations and / or depth data acquired by depth cameras may be implemented or performed to produce relatively highly accurate results of geometry information of the viewed volume or scene. These accurate results of the geometry information of the viewed volume or scene can in turn be used to provide a relatively sufficiently bounded surface estimate of the viewed volume or scene. The surface estimate may then be used to constrain, select or define the one or more (parametrization) surfaces used to define or specify the corresponding surface lightfields.

[0059] As illustrated in FIG. 1 A, some scene geometry estimates or resultant (parametrization) surfaces do not need to correspond to physical surfaces or actual surfaces of physical or visual objects. Additionally, optionally or alternatively, some scene geometry estimates or resultant (parametrization) surfaces may correspond to or follow along relatively closely or even coincide with physical surfaces or actual surfaces of physical or visual objects.

[0060] In some operational scenarios, there may be advantages for a parametrization surface to follow along relatively closely true or actual surfaces of physical or visual objects in the scene. One of the benefits is that such a parametrization surface allows the viewer to get closer (or to be at closer viewing positions) to the objects while keeping the viewer from crossing the surface boundary or the parametrization surface that separates the viewer from the viewed volume or scene. An additional benefit of the parametrization surface staying relatively close to the real surfaces of physical or visual objects is to simplify the lightfield representation. For example, if a physical or visual object to be viewed by the viewer is mostly Lambertian (without view-dependent effects), and if the scene geometry manifold orthe parametrization surface hugs surface(s) of the object, then there is no direction-based or direction-dependent color variation in light rays.

[0061] Surface lightfields are a powerful way of representing scenes. As lightfields, they are capable of representing all the complexity in a scene. As such, surface lightfields lend themselves very well to form the foundation on which a number of view synthesis methods can be constructed.5. EXAMPLE LSLF SETS

[0062] As noted, layered surface lightfield (LSLF) as described herein may be used to provide a layered representation of a scene targeting relatively high quality Novel View Synthesis (NVS).

[0063] In some operational scenarios, multiple cameras with overlapping fields-of-view can be used to capture a volumetric representation of the scene. Image and / or depth data generated with the multiple cameras may be used to generate an LSLF representation (e.g., one or more LSLF data sets, etc.) that provides a relatively compact way of representing this scene or volume.

[0064] An LSLF set - e.g., for a given time point / instance in a plurality of time points / instances covering a time duration of a video constructed from a plurality of LSLF sets respectively for the plurality of time points / instances, etc. - represents a scene with multiple surface layers such as illustrated in FIG. 1C, FIG. ID and FIG. IE.

[0065] Intensity and color mappings - which may be view dependent - and occlusion information along with depth information can be encoded into or decoded from a surface layer of the LSLF set. Spherical harmonic (SH) basis functions and their corresponding SH coefficients may be used for the surface layer or points therein to describe view dependent variations of intensity / color at each point.

[0066] Occlusion information - which may be specified or represented with alpha (map, channel or parameter) values - may be used to identify which parts of the scene is occluded or not visible from a target view represented by the LSLF set. Some or all of the occluded parts of the scene not visible from the target view may become visible from neighboring (e.g., camera, etc.) views, as illustrated in FIG. IB.

[0067] FIG. 1C illustrates an example LSLF set that includes three (or multiple) surface layers. As illustrated, each point in layer 1 or layer 2 in the multiple surface layers of the LSLF set is visible to at least one camera among cameras or camera elements used to generate original image data and / or depth data for constructing the multiple surface layers ofthe LSLF are located. An occluded layer in the multiple surface layers of the LSLF set may not be visible from any target (or camera) view supported by the cameras or camera elements - used to generate the original image data and / or depth data for the LSLF set. Hence, the occluded layer may not be present or included in the LSLF set.

[0068] FIG. ID illustrates three example texture images corresponding to three (or multiple) surface layers of an LSLF set. Each texture image in the three texture images may correspond to a respective surface layer in the multiple surface layers of the LSLF set. These texture images may be used to capture intensity / color (e.g., RGB, YCbCr, SH coefficients per color channel, etc.) mapping or information for the multiple surface layers.

[0069] With the information / data specified or included in the LSLF set or the multiple surface layers therein, a translated and reoriented (e.g., virtual camera, novel, etc.) view of the scene can be synthesized. For example, as illustrated in FIG. ID, the back wall is occluded from target view(s) of the camera or camera elements by the guitar player’s body. However, as visual or image information about the back wall is included or specified in the LSLF set, any synthesized view that “looks behind” the guitar player’s body may be reconstructed or synthesized correctly.

[0070] The quality of these reconstructed or synthesized views from the LSLF set depends on (e.g., a total number of, specific locations of, etc.) the reference or target views captured and processed with the cameras or camera elements as well as depends on the total number of surface layers created for the LSLF set. The greater the total number of the multiple surface layers in the LSLF set is, the higher quality the reconstructed or synthesized view images are. Also, the greater a portion, volume or spatial region of the scene is included in the LSLF set, the larger a view area or a set of synthesized or novel views can be reconstructed or synthesized from the LSLF set. However, a relatively large number of the multiple surface layers in the LSLF set and a relatively large portion of the scene represented in the LSLF set may come at the cost of increased transmission bandwidth. Techniques as described herein can be implemented to provide multiple ways of reducing this bandwidth cost.

[0071] Each surface layer of the multiple surface layers in the LSLF set is composed of a collection of points with implied connectivity. Each point in the collection of points in the surface layer may carry a relatively precise depth value, opacity information (alpha channel / parameter value - possibly view dependent), and (e.g., view dependent, etc.) color information (e.g., a set of SH coefficients for a set of SH basis functions per colorchannel / component of a color space).

[0072] There are some differences between a point in an LSLF set and a pixel in a 2D image. A 2D image pixel typically has a certain number of bits (e.g., 8 bits, etc.) allocated for each of color component (e.g., R, G, B, etc.) values in a color space (e.g., an RGB color space, etc.). An optional alpha value may also be included.

[0073] In comparison, a point in the LSLF set may include or may be specified with a combination of: a depth (value), SH coefficients, and an alpha (parameter / channel) value.

[0074] In some operational scenarios, the depth (value) for the point may be represented as a floating point number (or fixed point for a relatively limited range), which may be a half float number having sixteen (16) bits of depth.

[0075] The total number of the SH coefficients for the point may be NA2 per color channel, where N represents the total number (N) of orders of SH basis functions used. In some operational scenarios, the total number of orders of SH basis functions is three (N = 3): namely, the Oth, first and second orders. Hence, the total number of the SH basis functions used as well as the total number of the corresponding SH coefficients is nine (NA2 = 9). Each of these SH coefficients may be twelve (12) bits with optional sub-sampling (e.g., in chroma channels in an YCbCr color space, etc.). The SH basis functions with their respective SH coefficients can be used to define, specify or represent view dependent color values (e.g., RGB values, luma values, chroma values, etc.).

[0076] The alpha (parameter / channel) value for the point may be represented as a fixed point of twelve (12) bits.

[0077] For each two dimensional grid of points of a target view in 3D space in which the viewed volume or scene resides, an LSLF set as described herein or one or more surface layers therein may include multiple points (with different sets of depth, alpha and / or SH coefficients). For example, these points may include a first point carrying foreground information and a second point carrying background information.

[0078] LSLF sets - e.g., a sequence of time consecutive LSLF sets, etc. - may be used by surface light field algorithms / methods / operations to generate immersive or volumetric videos or images. These algorithms / methods / operations may use surface depths in the scene as input or priors for the purpose of generating the videos or images. There may be a wide variety of ways to convey surface depth information, including but not necessarily limited to only, any, some or all of: signed distance function (SDF), meshes, depth maps etc. In some operational scenarios, depth maps may be used by surface light field algorithms / methods / operations asdescribed herein. These depth maps may be relatively easy or efficient to process, compress / encode, transmit and decode in real time or near real time with available codecs or GPU hardware such as those supporting H.265 codecs operations.6. LSLF DATA GENERATION AND ENCODING

[0079] FIG. 3A illustrates an example process flow for LSLF data generation and encoding operations. The process flow may be implemented with one or more computing devices including but not limited to volumetric or immersive video encoding devices equipped or deployed with available video or image codecs.

[0080] Block 352 comprises receiving input light field data such as represented by or sampled into or with a 3D point cloud. The input 3D point cloud includes a spatial distribution of points located at a plurality of spatial locations in a represented 3D space, a viewed volume, a scene, etc. These points may be specified with different vision related attributes such as colors, intensities or luminance values, chrominance values, reflectances, etc., some or all of which may or may not be direction- or view-dependent. It should be noted that, in various operational scenarios, other form of volumetric or immersive video or image data (e.g., 4D or 5D plenoptic functions, multi-view images, voxels, 3D object rendering models, 3D grids or meshes with visual attributes, etc.) may be used - in addition to or in place of point clouds - as input by a system as described herein to generate LSLF data.

[0081] Block 354 comprises generating a specific LSLF surface geometry representation for an LSLF set to be generated from the input light field or point cloud. Example surface geometry operations for generating the specific LSLF surface geometry representation may include, but are not necessarily limited to only, any, some or all of: overlayering 358, target specialization 360, target customization 362, and so on. The specific LSLF surface geometry representation may specify or identify one or more surface layers to be included in the LSLF set, target view(s) to be represented by the LSLF set, specific projection operation(s) to be performed to generate depth, alpha map and SH coefficients, etc.

[0082] Block 354 further comprises generating the LSLF set from the input light field or point cloud based at least in part on the specific LSLF surface geometry representation. Projection operations (e.g., perspective, orthographic, radial, isometric, fisheye, etc.) and sampling operations (e.g., interpolation, extrapolation, etc.) may be used to project, sample and / or collapse the input light field or point cloud into the surface layers or points therein and generate respective visual attributes such as depth data, alpha (transparence, opacity, occlusion and / or disocclusion) data, SH coefficients (for colors, intensities, luminancesand / or chrominances), etc., for some or all points represented in each surface layer included in the LSLF set.

[0083] Block 356 comprising compressing or encoding LSLF data (e.g., depth, alpha, SH coefficients, etc.) in the LSLF set into one or more (volumetric or immersive video / image) coded bitstreams, sub-streams, or video files.

[0084] In some operational scenarios, some or all of the operations of FIG. 3A may be performed in real time or near real time. Additionally, optionally or alternatively, some or all of these operations may be performed offline. Some or all of these operations can be repeatedly or iteratively performed to process a sequence of (e.g., time sequential, consecutive, etc.) input light fields or point clouds into a corresponding sequence of (e.g., time sequential, consecutive, etc.) of LSLF sets covering a sequence of (e.g., time sequential, consecutive, etc.) time points or instances or steps.7. OVERLAYERING

[0085] A surface lightfield may use a single surface layer as a basic representation, with view-dependent visual appearance or intensity / color data stored or represented by SH coefficients at points throughout the single surface layer.

[0086] In some operational scenarios, view-dependent visual appearances of a scene can be relatively (e.g., highly, etc.) variable. For example, the scene may include physical or visual objects around (e.g., local, etc.) spatial areas or surfaces reflecting specular light or (relatively highly) direction- or view-dependent light rays. At any given single surface layer in or around these (e.g., local, etc.) spatial areas or surfaces, a series of SH basis functions used to represent the direction- or view-dependent light rays may converge slowly. Relatively high orders of SH basis functions may become significant in the series of SH basis functions to represent accurately the view-dependent visual appearances of the scene with the single surface layer. However, as the total number of SH coefficients grows quadratically with the total number of orders used in the SH basis functions, a relatively large amount of SH coefficient data may have to be used with the basic representation of the surface lightfield that uses only a single surface layer.

[0087] An LSLF set as described herein may be implemented to enable an alternative representation for surface lightfields that involves (e.g., local, etc.) spatial regions or surfaces with relatively high direction or view variations in visual appearances. This alternative representation can support using multiple (e.g., local, etc.) surface layers around a single surface or around the (e.g., local, etc.) spatial regions or surfaces. The multiple (e.g., local,etc.) surface layers in the LSLF for the (e.g., local, etc.) spatial regions or surfaces with the relatively high direction or view variations in visual appearances may include a main (e.g., local, etc.) layer and one or more additional (e.g., local, etc.) layers. The additional layers can still follow along a specific surface or portions thereof represented by the main layer, but with (e.g., relatively small inter-layer, etc.) spatial separations from the main layer.

[0088] Each of these multiple surface layers in the alternative representation can be used to hold or maintain same types of visual data / information as a single surface layer in the basic representation.

[0089] FIG. 2A illustrates an example surface layer representation 202 with overlayering in an LSLF as described herein. As illustrated, the surface layer representation 202, which may be referred to as a main surface layer, includes three spatial regions or portions 202- 1 , 202-2 and 202-3.

[0090] In some operational scenarios, the main surface layer may be a surface corresponding to, coincide with, or adjacent to, an actual surface of physical or visual object(s) in a 3D volume or scene. The physical or visual object(s) may be located on the upper space of the main surface layer emitting or reflecting light (or light rays) that may or may not be directional- or view- dependent.

[0091] For the purpose of illustration, visual appearance of the scene as viewed from the main surface layer may be relatively accurately captured with a single surface layer such as the main layer portions 202-1 and 202-3. In other words, SH coefficients - up to a specific maximum order of SH basis functions - for each point of some or all points included in the main layer portions 202-1 and 202-3 relatively accurately describe direction- or viewdependent light rays passing through that point.

[0092] On the other hand, visual appearance of the scene as viewed from the main surface layer may not be relatively accurately captured with a single surface layer such as the main layer portion 202-2. In other words, SH coefficients - up to a specific maximum order of SH basis functions - for each point (e.g., a, b, c, etc.) of some or all points included in the main layer portion 202-2 may not accurately describe direction- or view-dependent light rays passing through that point.

[0093] One or more additional layer 204-1 through 204-5 may be added in the LSLF set to complement the main layer (portion) 202-2. These additional layers may generally follow along curvilinear contours with respect to the main layer (portion) 202-2. In some operational scenarios, these additional layers are not intersecting and located with relatively small spatialseparations - except at edges such as those joining with the main layers 202-1 and 202-3 - from one another and from the main layer (portion) 202-2. In some operational scenarios, the additional layers do not join with the main layers 202-1 and 202-3.

[0094] Each point included in a surface layer of the multiple surface layers 202-1 and 204-1 through 204-5 may include depth data for that point, alpha map data for that point, and SH coefficients for that point (up to the maximum order of the SH basis functions used to represent visual data or light in the LSLF set, etc.). The use of multiple layers as described herein provides more capacity for representing view-dependent visual appearance or data (than other approaches), without increasing the maximum order of the SH basis functions. Furthermore, the use of multiple layers as described herein can be done selectively, for example only for surface portions or patches (e.g., with relatively highly direction- or viewdependent visual appearance, etc.) that would benefit. This multi-layer technique used selectively for relatively high direction- or view-dependent visual appearance may be referred to as “overlayering”.

[0095] Consider a cone 206 of light rays that are coming from a given surface point such as point a of FIG. 2 A. At the start, the cone 206 only intersects one point - the origin (or point a) at the surface, providing a single sample (or data point) of the visual appearance. At the second layer or additional layer 204-1, the cone 206 intersects a small section of that layer, providing a few more samples. The further out it goes, the cone 206 will intersect an increasingly larger surface patch, with the number of samples increasing with the area (or quadratically). The larger the number of samples, the greater the variation - higher spatial frequency - of the visual appearance can be represented. Collectively, all the samples from the multiple layers 202-2 and 204-1 through 204-5 - such as those in the cone 206 - may be used to relatively accurately represent direction- or view-dependent appearance of light rays from the single point a on the main surface (layer).

[0096] A tradeoff of this overlayering approach is that samples or data on the additional layers such as 204-1 through 204-5 are shared between or among - or contributed by light rays from - multiple points such as points b and c on the main layer (portion) 202-2. While the base or main layer (portion) such as 202-2 carries data exclusive to each its surface point, all other layers are shared or contributed by multiple points on the base or main layer (portion) 202-2, with the outer (or more outward) additional layers shared between or among more points of the base or main layer (portion) 202-2 than the inner (or more inward) additional layers.

[0097] However, each “shared” point or its corresponding samples on the additional layers may be itself direction- or view-dependent and has its respective capacity to - e.g., with a separate set of SH coefficients for up to the maximum order of the SH basis functions, semi-independently, etc. - represent each surface point (on the main surface (layer) 202-2) that shares that (“shared”) point. This capacity of representation by the shared point depends on the total number of SH coefficients carried or the total number of orders of the SH basis functions. In some operational scenarios, multi-layer alpha blending operations may be performed to combine and render visual data carried by the main layer (portion) 202-2 and the additional layers 204-1 through 204-5. The multi-layer alpha blending operations may be based at least in part on transparency / opaqueness values as indicated by alpha map values set for these different layers. Each of the shared points or samples on the additional layers 204-1 through 204-5 may be used to represent a portion of corresponding main surface points’ visual appearance.

[0098] Under the overlayering techniques as described herein, visual data or appearance relating to a surface lightfield or a portion thereof can be distributed across multiple layers, for example locally or globally. A relatively efficient mechanism may be implemented to reduce redundancies that may exist in the visual data, for example by solving an optimization problem, by removing opaque layers, by removing layer portions not to be used in view synthesis or reconstruction, etc. Example optimization problems as described herein may include, but are not necessarily limited to only, those specified or formulated or optimized / minimized with an image quality and / or bitrate utilization based loss / objective / error functions. Some or all of the total number of additional layers and their positions (depth or spatial separation), the total number of SH coefficients carried by each layer (for intensity / color and / or depth data and / or alpha data), etc., may be all adjustable or optimizable operational parameters for the optimization problem used to reduce data redundancy.

[0099] A specific combination of optimized parameters (or parameter values) that provides the most compact representation may be generated or selected among many different combinations of candidate parameters (or parameter values) as a solution to the optimization problem. In some operational scenarios, different combinations of optimized parameters (or parameter values) may be generated for different surface light field patches and their corresponding visual appearances.8. TARGET SPECIALIZATION

[0100] An LSLF (geometry and / or texture / luma / chroma data and / or alpha data) representation provides compactness in representing light field(s) of a 3D volume or scene by collapsing volumetric lightfield(s) of the scene into multiple 2-dimensional surface layers. As noted, in some operational scenarios, some or all surface layers in an LSLF set may represent, correspond to or relatively closely approximate actual surfaces of physical or visual objects depicted in the scene.

[0101] Some or all of these surfaces or surface layers may be generated using 3D reconstruction algorithms / methods / operations (implemented in a combination of hardware or software), for example using point clouds as input.

[0102] Qualities of LSLF data and LSLF based view synthesis may depend on how points from the point cloud are grouped (and interpolated / extrapolated) into these surface layers. It is possible that some of the surfaces or surface layers - as generated from these 3D reconstruction algorithms / methods / operations - may not correspond to represent, correspond to or relatively closely approximate actual surfaces of physical or visual objects giving rise to the point clouds.

[0103] Target specialization as described herein may be implemented or performed to select a specific mapping from among a plurality of candidate mappings. Each of these candidate mappings may be capable of mapping from a scene geometry to a target manifold or a corresponding 2D surface. Each of some or all of the candidate mappings may be a (spatial) projection. Example candidate or selected mappings as described herein may include, but are not necessarily limited to only, any of: a perspective projection, an orthographic projection, a radial projection, an isometric projection, a fisheye projection, and so on.

[0104] The selection of a specific mapping used to map the scene to a specific surface layer in an LSLF set may be based at least in part on a range of locations from which view synthesis is to be performed. The range of locations can be identified or specified based on manual input or specification, for example from a media content creator. The specific mapping may be selected from candidate mappings to allow relatively high quality video or display images of the scene to be reconstructed, synthesized or resynthesized from any, some or all of these locations in the range of locations. As a result, image synthesis quality can be relatively high at these locations, while bandwidth usage for transmitting the LSLF data can be minimized.

[0105] FIG. 2B and FIG. 2C illustrate two example LSLF representations / sets of the same 3D volume or scene. These LSLF representations / sets may be selected or generated with target constraints or specialization based on (e.g., expected, manually or programmatically specified, etc.) ranges of locations from which view syntheses or reconstructions are to be made.

[0106] As illustrated in FIG. 2B, in response to determining that a first range of locations is to be used for view synthesis and reconstruction, a first mapping 208-1 that represents a projection on a cylindrical surface centered at axis 210-1 (corresponding to a first target view for a first LSLF set) may be selected from a plurality of candidate mappings.

[0107] In comparison, as illustrated in FIG. 2C, in response to determining that a second range of locations is to be used for view synthesis and reconstruction, a second mapping 208- 2 that represents a projection from a second reference position 210-2 (corresponding to a second target view for a second LSLF set) may be selected from the plurality of candidate mappings.

[0108] As illustrated in FIG. 2B and FIG. 2C, the first LSLF set generated with the first mapping may include more surface layers than the second LSLF set.

[0109] The first LSLF set can be used to provide or support relatively high quality visual data for a novel or to-be-synthesized view from or near first locations in the first range of locations 208-1. In operational scenarios in which a volumetric or immersive video application is expected to reconstruct or synthesize views from the first locations, the first LSLF set can be generated to help better reduce bandwidth usage than the second LSLF set.

[0110] The second LSLF set can be used to provide or support relatively high quality visual data for a novel or to-be-synthesized view from or near second locations in the second range of locations 208-2. In operational scenarios in which a volumetric or immersive video application is expected to reconstruct or synthesize views from the second locations, the second LSLF set can be generated to help better reduce bandwidth usage than the first LSLF set.9. TARGET CUSTOMIZATION

[0111] In some operational scenarios, a system as described herein - an upstream device that is tasked to generate LSLF sets (e.g., used to generate a sequence of immersive or volumetric video images, etc.) - may receive or have access to view synthesis details.

[0112] Example view synthesis details may include, but are not necessarily limited to only, a viewer’s target pose or view; a spatial location / orientation of an image display used torender immersive videos or images; field-of-view constraints or ranges or dimensions relating to the image display or the viewer; the viewer’s viewing direction and / or location and / or a viewing volume in which the viewer may have degrees of freedom to move - as represented in a 3D space that includes both the viewed volume or scene and the viewing volume; and so on.

[0113] Some or all of these view synthesis details may be generated through real time monitoring of the viewer’s gazes and / or poses, for example, by an attendant tracking device operating in conjunction with the system or by a downstream device that uses the LSLF sets to render immersive videos / images.

[0114] Additionally, optionally or alternatively, some or all of the view synthesis details may be pre-programed, for example based at least in part on an immersive content creator’s manual input.

[0115] In some operational scenarios, the upstream device can (e.g., further, after LSLF target specialization, etc.) customize the LSLF representation in the LSLF sets for the scene based on some or all of the view synthesis details such as the viewer’s (e.g., actual, predicted, estimated, etc.) target pose and field-of-view constraints of the viewer and / or the image display.

[0116] FIG. ID and FIG. IE illustrate two example LSLF sets that may have been target customized to viewer’s two different view directions (or orientations) and / or positions, respectively. More specifically, a first LSLF set (e.g., three first texture images in three first LSLF surface layers of the first LSLF set, etc.) as illustrated in FIG. ID may be generated or customized to a first view from the left side, whereas a second LSLF set (e.g., three second texture images in three second LSLF surface layers of the second LSLF set, etc.) as illustrated in FIG. IE may be generated or customized to a second view from the right side.

[0117] The LSLF target customization (operations) performed by the upstream device may be used to (e.g., significantly, relatively efficiently, etc.) compress - or reduce the size / amount of - the LSLF data to be encoded into or carried by a coded bitstream. As an example, only a relatively small subset of point cloud (video / image) data may be collapsed into or represented by LSLF surface layers in the LSLF sets. Hence, the coded bitstream can be directly or indirectly transmitted or delivered by the upstream device to downstream recipient devices with reduced transmission bitrate or bandwidth usage.

[0118] The LSLF target customization as described herein can be used to generate LSLF data and resultant videos / images customized or constrained to some or all of the viewsynthesis details such as the viewing or image rendering device position, orientation and field-of-view limitations.

[0119] In some operational scenarios, the viewing or image rendering device positions and / or orientations may be - continuously or at a relatively high sampling rate in real time or near real time - monitored and reported by a downstream device operating with the viewing or image rendering device to a server such as the upstream device that is to generate the customized LSLF data.

[0120] In response to receiving the reported (e.g., estimated, predicted, etc.) viewing or image rendering device position and / or orientation for a given time point or instance, the upstream device can generate and respond with a right or specifically customized set of LSLF data for the given time point or instance in accordance with the viewing or image rendering position and / or orientation.

[0121] FIG. 2D illustrates example LSLF target customization with respect to a toy of a hexagon shape. Two target views are shown: View A (on the left side of FIG. 2D) and View B (on the right side of FIG. 2D).

[0122] Without any (target view) constraints or LSLF target customization, an LSLF set for each of the two views (View A and View B) would carry or include three (3) LSLF surface layers, respectively.

[0123] In comparison, under the (target view) constraints or LSLF target customization as described herein, as illustrated on the left side of FIG. 2D, in response to determining that the viewer’s first viewing position and / or orientation (or direction) at a first time point / instance corresponds to or is relatively close to View A, the system can generate a first LSLF set specifically customized to View A. Layer 1 may be dropped from the first LSLF set. Even if the viewer’s first viewing position and / or orientation (or direction) does not coincide with View A, a downstream device receiving the first LSLF set - for example, by way of a coded bitstream - can perform first view synthesis or reconstruction operations such as first image warping operations with the first LSLF set to generate, synthesize or reconstruct a first image for a first novel view represented by the viewer’s first viewing position and / or orientation (or direction).

[0124] Likewise, as illustrated on the right side of FIG. 2D, in response to determining that the viewer’s second viewing position and / or orientation (or direction) at a second time point / instance corresponds to or is relatively close to View B, the system can generate a second LSLF set specifically customized to View B. Layer 1 may be dropped from thesecond LSLF set. Even if the viewer’s second viewing position and / or orientation (or direction) does not coincide with View B, the downstream device receiving the second LSLF set - for example, by way of the coded bitstream - can second perform view synthesis or reconstruction operations such as second image warping operations with the second LSLF set to generate, synthesize or reconstruct a second image for a second novel view represented by the viewer’s second viewing position and / or orientation (or direction).

[0125] Hence, LSLF sets generated by the upstream device with knowledge of target views at the downstream device can exclude or drop layers or portions thereof that are not to be used by the downstream device in the latter’s image synthesis or reconstruction operations. The LSLF data generation operations can be performed by the upstream device dynamically and / or adaptively based on the target views estimated or reported with view synthesis details collected or reported by the downstream device to ensure the downstream device to receive sufficient LSLF layers or data for synthesizing or reconstructing images correctly for views such as any orientation between Views A and B.10. LSLF FEATURES AND OPERATIONS

[0126] In some embodiments, an LSLF set as described herein may be generated in accordance with an LSLF format that collapses a volumetric scene into a sequence of two- dimensional data elements (surface layers) that carry LSLF data relating to depth, transparency and view dependent intensity / color information or values associated with surfaces - corresponding or giving rise to the 2D data elements or surface layers in the LSLF set - in the scene. The surface layers may or may not hug or coincide with real surfaces of visual or physical objects depicted in the scene. Also, in some operational scenarios, each of some or all of the (e.g., real, etc.) surfaces in the scene or surface portions thereof may be associated with one or more (local or global; main; additional) surface layers at relative or different depths - which may or may not be of constant depths - to aid in fully describing the surface (e.g., direction- or view-dependent visual appearances or intensity / color information, etc.).

[0127] In some embodiments: the total number of (local or global; main; additional) surface layers associated with a (e.g., virtual or real, etc.) surface in the scene can be increased to increase expressive capacity of the LSLF set in representing that surface.

[0128] In some embodiments, the LSLF surface layers in the LSLF set are set up to honor or comply with a specific occlusion ordering, which may be used to determine a specific order in applying or performing alpha blending operations with respect to these differentLSLF surface layers. For example, an LSLF image Tenderer implemented by a downstream recipient device of the LSLF set - for example, by way of a coded bitstream - may relatively easily or efficiently implement “back-to-front” blending operations of view dependent colors, across the successive LSLF surface layers based on the specific occlusion ordering. This provides a rendering mechanism to relatively accurately and efficiently represent transparency (of different surface layers and their corresponding visual appearances) in the scene, as well as to help increase the expressive capacity of the LSLF set or data.

[0129] In some embodiments, an LSLF set as described herein may be built for a particular target view neighborhood. The LSLF data can be expressed as or generated with a target view (spatial) mapping from a scene or its corresponding geometry to the target view manifold or surface. This target view mapping can be a perspective projection, an orthographic projection, a radial projection, an isometric projection, a fisheye projection, and so on, to closely match specific type(s) of to-be-synthesized views desired or intended. This target view mapping can be specifically selected among all the candidate projections to allow or support relatively high (e.g., the highest, etc.) fidelity reconstruction of the synthetic or novel views without increasing a total size or amount of the LSLF data in the LSLF set.

[0130] In some embodiments, the scene may be described by multiple (e.g., target view customized, etc.) LSLF sets generated from or customized for multiple target views (with different combinations of view positions / locations and / or orientations / directions). An adaptive target view switching mechanism may be implemented by an upstream device and / or a downstream (synthesizing / decoding / rendering) device, for example switching between or among the different LSLF sets customized for different target view based on real time or near real time target view details. Each of some or all of the target- view customized LSLF sets may be simplified by removing redundant occluded surface layers or surface layer portions.11. LSLF DATA DECODING AND RENDERING

[0131] Many different approaches may be implemented by a downstream recipient / decoder device to use LSLF sets as described herein - which may be received by way of decoding a coded bitstream generated by an upstream encoder device - to generate display images corresponding to novel or reconstructed views and render these display images on one or more image displays operating with the downstream device.

[0132] FIG. 3B illustrates an example process flow for LSLF rendering operations. The process flow may be implemented with one or more computing devices including but notlimited to end user devices equipped or deployed with available video or image codecs. Some or all of the LSLF rendering operations such as illustrated in FIG. 3B may be implemented or performed based at least on 3D graphics, rasterizing textured meshes, etc.

[0133] The upstream device can encode some or all LSLF data (e.g., depth maps, SH coefficients, alpha data, etc.) one or more coded bitstream such as H.265 bitstreams or compressed H.265 video files. The LSLF data may be encoded by the upstream device using video coding syntaxes or syntax elements supported by available (e.g., hardware, firmware, a combination of hardware and software, etc.) H.265 decoders or codecs deployed with some or all of a wide variety of downstream devices or end user devices. Hence, the LSLF data may be stored by the upstream device in compressed H.265 video files or bitstreams. Some or all of the compressed LSLF data in the video files or bitstreams may be decoded and converted by the downstream device into textured meshes for rendering.

[0134] Block 302 comprises decoding one or more bitstreams or video files to recover or retrieve LSLF data in an LSLF set. The LSLF data may be coded or decoded as one or more surface (or parameter) layers in the LSLF set; the one or more surface layers may, for example, correspond to the one or more bitstreams or video files, respectively.

[0135] Each of the surface layers in the LSLF set may include a specific depth map 310, a specific alpha map, specific SH coefficients 308 forming a specific texture image 312 for the surface layer.

[0136] Block 304 comprises converting the depth maps 310 in the LSLF set into 3D meshes 314. By way of illustration but not limitation, the depth maps 310 include 3 surface layers representing three surfaces of a cube, a sphere and a pyramid, respectively. The 3D meshes 314 converted from the depth maps 310 have two aspects: vertex connectivity (e.g., forming or defining mesh topology, etc) and vertex (or grid) positions. The vertex connectivity - which defines, for a given spatial resolution, a fixed mesh or surface topology such as one or more of the three surfaces of the cube, the sphere and the pyramid - is implied by or expressed with a 2D grid of pixels (e.g., a quad for each set of neighboring 4 pixels in a 2D grid, etc.) of one or more of the 3D meshes 314.

[0137] Vertex positions forming a 3D mesh 314 or a corresponding 2D grid - e.g., representing a surface layer in the LSLF set or a visual object therein - may be relatively efficiently determined or estimated from a depth map by inverting or reversing a projection that was used by the upstream device to create the depth map. In an example, vertex coordinates of a pixel in a 3D mesh 314 or a corresponding 2D grid of pixels as describedherein may be represented in an XYZ (Cartesian) coordinate system based at least in part on pixel coordinates and depth of that pixel.

[0138] In some operational scenarios, spatial locations or coordinates of the pixels in the 3D meshes 314 or the corresponding 2D grids can also be used to define specific spatial coordinates used for texture (e.g., RGB, luma / chroma, YUV or YCbCr, etc.) sampling operations. These texture sampling operations uses the spatial locations or coordinates of the pixels to sample SH coefficient data to derive texture (e.g., pixel, RGB, YUV, YCbCr, etc.) values 316 for these pixels, for example through interpolation operations.

[0139] Block 306 comprises rendering the 3D scene (or viewed volume) as previously represented in the LSLF set and now represented as a (layered) collection of textured meshes, for example using a (e.g., standard, proprietary, etc.) rasterization-based programmable graphics rendering pipeline. The rendering pipeline may transform and project the 3D meshes 314 - corresponding to the surface layers of the LSLF set - and the texture values 316 of the pixels or vertexes (forming the 3D meshes 314) from a target view represented in the LSLF set into a specific (e.g., novel, synthetic, desired, target, viewer’s, etc.) view for rendering or visualization purposes.

[0140] The rendering pipeline may include or invoke one or more shaders such as fragment shaders to sample (e.g., 9, etc.) SH coefficients for each color channel or component using (e.g., regular, non-regular, etc.) texture sampling. The shaders may also compute a view dependent SH multiplier, based on a specific angle between the target view and the specific (e.g., novel, etc.) view. A dot (or inner) product between the SH (spherical harmonic) coefficients and the view dependent multiplier may be applied or performed to determines (e.g., R, G, B, Y, Cb, Cr, etc.} color or luma / chroma values at the specific (e.g., novel, etc.) view for each surface layer or each corresponding 3D mesh. The textured mesh layers (or texture values of pixels in the 3D meshes) may be (e.g., alpha, etc.) blended from back to front with depth testing and alpha blending enabled.

[0141] FIG. 3C illustrates an example algorithm / method to project a first pixel and second pixel into a specific (target or novel) view in rendering operations. The shaders can be programmed to automatically cull any unnecessary spatial surface portions or faces, for example when there is no depth defined for a pixel and therefore representing an invalid vertex / face.12. EXAMPLE PROCESS FLOWS

[0142] FIG. 4A illustrates an example process flow according to an embodiment. In someembodiments, one or more computing devices or components (e.g., one or more video codecs, an encoding device / module, a transcoding device / module, a coding device / module, a volumetric or immersive video server system, etc.) may perform this process flow. In block 402, an image processing system receives an input lightfield comprising a spatial distribution of visual and geometric information in a volumetric scene in a three-dimensional (3D) space.

[0143] In block 404, the system selects a spatial sequence of two-dimensional (2D) surface layers to be included in a layered surface lightfield (LSLF) set.

[0144] In block 406, the system collapses the input lightfield into the spatial sequence of 2D surface layers of the LSLF set. Each 2D surface layer in the spatial sequence of 2D surface layers includes a depth data portion, a transparency data portion, a color data portion, etc.

[0145] In block 408, the system encodes the LSLF set into a coded bitstream. The coded bitstream causes a recipient device of the coded bitstream to generate a display image of the volumetric scene for rendering on an image display.

[0146] In an embodiment, the surface layer represents one of: a surface adjacent to a real surface of a visual object depicted in the volumetric scene, or a surface not curvilinear to any real surface of any visual object depicted in the volumetric scene.

[0147] In an embodiment, the surface layer includes an overlayering portion formed by a main surface at a main depth and a set of one or more additional surfaces at one or more different depths from the main depth; each surface in the set of one or more additional surfaces carries additional depth, transparency and view-dependent color data to be used along with main depth, transparency and color data carried in the main surface in blending operations that provide more accurate view-dependent colors than colors carried in the main surface alone.

[0148] In an embodiment, a total number of surfaces in the set of additional surfaces is set based at least in part on a target expressive capacity of the LSLF set representing the main surface.

[0149] In an embodiment, the surface layer includes a non-overlayered portion not covered by any set of additional surfaces; the overlayered portion includes visual information of greater view-dependency for the original input lightfield than the non-overlayered portion.

[0150] In an embodiment, the transparency data portion includes alpha channel values.

[0151] In an embodiment, the display image is generated through blending the spatial sequence of surface layers in the LSLF set from back to front based on respective depth andtransparency data carried by the surface layers in the spatial sequence of surface layers to produce a relatively accurate visual representation of the volumetric scene.

[0152] In an embodiment, the LSLF set is specifically built for a particular spatial neighborhood in reference to a specific target view using a specific spatial mapping from a 3D geometry of the volumetric scene to 2D geometries of the surface layers in the LSLF set; the specific spatial mapping represents one of: a perspective mapping, an orthographic mapping, a radial projection, an isometric projection, a fisheye projection, or another spatial projection.

[0153] In an embodiment, one or both of the specific target view and the specific spatial mapping are selected based at least in part on one or more synthesized views to be reconstructed from the LSLF set with relatively high fidelity by one or more recipient devices of the coded bitstream.

[0154] In an embodiment, at least one of the specific target view and the specific spatial mapping is selected based at least in part on constraining a total amount of data to be carried with the LSLF set in the coded bitstream to a specific data size limit.

[0155] In an embodiment, multiple LSLF sets corresponding multiple target views are generated from the same input lightfield; wherein synthesized view data communicated from the downstream device is used to determine a specific target view used to select the LSLF set from among the multiple LSLF sets.

[0156] In an embodiment, one or more surface layer portions occluded to the specific target view are removed from the set of surface layers in the LSLF sets from being encoded into the coded bitstream.

[0157] FIG. 4B illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., one or more video codecs, a coding device / module, a transcoding device / module, a decoding device / module, a volumetric or immersive video client system, etc.) may perform this process flow. In block 452, an image processing system decodes a layered lightfield (LSLF) set from a coded bitstream. The LSLF set includes a spatial sequence of two-dimensional (2D) surface layers with respective depth data portions, transparency data portions, color data portions, etc.

[0158] In block 454, the system uses the LSLF set to generate a display image.

[0159] In block 456, the system renders the display image on an image display.

[0160] In an embodiment, the spatial sequence of 2D surface layers is converted into a set of three-dimensional (3D) meshes; wherein the display image is generated by performingblending operations with the set of 3D meshes.

[0161] In an embodiment, a computing device such as a display device, a mobile device, a set-top box, a multimedia device, etc., is configured to perform any of the foregoing methods. In an embodiment, an apparatus comprises a processor and is configured to perform any of the foregoing methods. In an embodiment, a non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of any of the foregoing methods.

[0162] In an embodiment, a computing device comprising one or more processors and one or more storage media storing a set of instructions which, when executed by the one or more processors, cause performance of any of the foregoing methods.

[0163] Note that, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.13. IMPLEMENTATION MECHANISMS - HARDWARE OVERVIEW

[0164] Embodiments of the present invention may be implemented with a computer system, systems configured in electronic circuitry and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or apparatus that includes one or more of such systems, devices or components. The computer and / or IC may perform, control, or execute instructions relating to the adaptive perceptual quantization of images with enhanced dynamic range, such as those described herein. The computer and / or IC may compute any of a variety of parameters or values that relate to the adaptive perceptual quantization processes described herein. The image and video embodiments may be implemented in hardware, software, firmware and various combinations thereof.

[0165] Certain implementations of the inventio comprise computer processors which execute software instructions which cause the processors to perform a method of the disclosure. For example, one or more processors in a display, an encoder, a set top box, a transcoder or the like may implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processors. Embodiments of the invention may also be provided in the form of a program product. The program product may comprise any non-transitory medium which carries a set of computer-readable signals comprising instructions which, when executed by adata processor, cause the data processor to execute a method of an embodiment of the invention. Program products according to embodiments of the invention may be in any of a wide variety of forms. The program product may comprise, for example, physical media such as magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAM, or the like. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0166] Where a component (e.g. a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including a reference to a "means") should be interpreted as including as equivalents of that component any component which performs the function of the described component (e.g., that is functionally equivalent), including components which are not structurally equivalent to the disclosed structure which performs the function in the illustrated example embodiments of the invention.

[0167] According to one embodiment, the techniques described herein are implemented by one or more special -purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices or any other device that incorporates hard-wired and / or program logic to implement the techniques.

[0168] For example, FIG. 5 is a block diagram that illustrates a computer system 500 upon which an embodiment of the invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general purpose microprocessor.

[0169] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing informationand instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions.

[0170] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions.

[0171] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane.

[0172] Computer system 500 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques as described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0173] The term “storage media” as used herein refers to any non-transitory media that store data and / or instructions that cause a machine to operation in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape,or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge.

[0174] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications.

[0175] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infra-red transmitter to convert the data to an infra-red signal. An infra-red detector can receive the data carried in the infra-red signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.

[0176] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0177] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by anInternet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet” 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are example forms of transmission media.

[0178] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.

[0179] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution.14. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS

[0180] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is claimed embodiments of the invention, and is intended by the applicants to be claimed embodiments of the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.Enumerated Exemplary Embodiments

[0181] The invention may be embodied in any of the forms described herein, including, but not limited to the following Enumerated Example Embodiments (EEEs) which describe structure, features, and functionality of some portions of embodiments of the present invention.

[0182] EEEL A method, comprising: receiving an input lightfield comprising a spatial distribution of visual and geometricinformation in a volumetric scene in a three-dimensional (3D) space; selecting a spatial sequence of surface layers to be included in a layered surface lightfield (LSLF) set, wherein each surface layer in the spatial sequence of surface layers is parametrized using two-dimensional (2D) coordinates; collapsing the input lightfield into the spatial sequence of surface layers of the LSLF set, wherein each surface layer in the spatial sequence of surface layers includes a depth data portion, a transparency data portion and a color data portion; encoding the LSLF set into a coded bitstream, wherein the coded bitstream causes a recipient device of the coded bitstream to generate a display image of the volumetric scene for rendering on an image display.

[0183] EEE2. The method as recited in EEE1 , wherein at least one surface layer in the spatial sequence of surface layers includes an overlayering portion formed by a main surface at a main depth and a set of one or more additional surfaces at one or more different depths from the main depth; wherein each surface in the set of one or more additional surfaces carries additional depth, view-dependent transparency and view-dependent color data to be used along with main depth, transparency and color data carried in the main surface in blending operations that provide more accurate view-dependent colors than colors carried in the main surface alone.

[0184] EEE3. The method as recited in EEE2, wherein a total number of surfaces in the set of additional surfaces is set based at least in part on a target expressive capacity of the LSLF set representing the main surface.

[0185] EEE4. The method as recited in EEE2 or EEE3, wherein the surface layer includes a non-o verlayered portion not covered by any set of additional surfaces; wherein the overlayered portion includes visual information of greater view-dependency for the original input lightfield than the non-o verlayered portion.

[0186] EEE5. The method as recited in any one of EEE1-EEE4, wherein at least one surface layer in the spatial sequence of surface layers represents one of: a surface adjacent to a real surface of a visual object depicted in the volumetric scene, or a surface not curvilinear to any real surface of any visual object depicted in the volumetric scene.

[0187] EEE6. The method as recited in any one of EEE1-EEE5, wherein the transparency data portion includes alpha channel values.

[0188] EEE7. The method as recited in any one of EEE1-EEE6, wherein the display image is generated through blending the spatial sequence of surface layers in the LSLF setfrom back to front based on respective depth and transparency data carried by the surface layers in the spatial sequence of surface layers.

[0189] EEE8. The method as recited in any one of EEE1-EEE7, wherein the LSLF set is specifically built for a particular spatial neighborhood in reference to a specific target view using a specific spatial mapping from a 3D geometry of the volumetric scene to 2D geometries of the surface layers in the LSLF set; wherein the specific spatial mapping represents one of: a perspective mapping, an orthographic mapping, a radial projection, an isometric projection, a fisheye projection, or another spatial projection.

[0190] EEE9. The method as recited in EEE8, wherein one or both of the specific target view and the specific spatial mapping are selected based at least in part on one or more synthesized views to be reconstructed from the LSLF set with relatively high fidelity by one or more recipient devices of the coded bitstream.

[0191] EEE10. The method as recited in EEE8 or EEE9, wherein at least one of the specific target view and the specific spatial mapping is selected based at least in part on constraining a total amount of data to be carried with the LSLF set in the coded bitstream to a specific data size limit.

[0192] EEE11. The method as recited in any one of EEE1 -EEE10, wherein multiple LSLF sets corresponding multiple target views are generated from the same input lightfield; wherein synthesized view data communicated from the downstream device is used to determine a specific target view used to select the LSLF set from among the multiple LSLF sets.

[0193] EEE12. The method as recited in EEE11, wherein one or more surface layer portions occluded to the specific target view are removed from the set of surface layers in the LSLF sets from being encoded into the coded bitstream.

[0194] EEE13. A method, comprising: decoding a layered surface lightfield (LSLF) set from a coded bitstream, wherein the LSLF set includes a spatial sequence of surface layers with respective depth data portions, transparency data portions and color data portions; using the LSLF set to generate a display image; rendering the display image on an image display.

[0195] EEE14. The method as recited in EEE13, wherein at least one surface layer in the spatial sequence of surface layers includes an overlayering portion formed by a main surface at a main depth and a set of one or more additional surfaces at one or more differentdepths from the main depth; wherein each surface in the set of one or more additional surfaces carries additional depth, view-dependent transparency and view-dependent color data to be used along with main depth, transparency and color data carried in the main surface in blending operations that provide more accurate view-dependent colors than colors carried in the main surface alone.

[0196] EEE15. The method as recited in EEE13 or EEE14, wherein the spatial sequence of surface layers is converted into a target three-dimensional (3D) visual data representation; wherein the display image is generated by performing blending operations with the target 3D data representation.

[0197] EEE16. An apparatus performing any one of the methods as recited in EEE1-EEE15.

[0198] EEE17. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any one of the methods as recited in EEE1-EEE15.

[0199] EEE18. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any one of EEE1-EEE15.

Claims

CLAIMS1. A method, comprising: receiving an input lightfield comprising a spatial distribution of visual and geometric information in a volumetric scene in a three-dimensional (3D) space; selecting a spatial sequence of surface layers to be included in a layered surface lightfield (LSLF) set, wherein at least some of the surface layers spatially conform to, or are curvilinear with, real surfaces of physical or visual objects depicted in the scene, and wherein each surface layer in the spatial sequence of surface layers is parametrized using two-dimensional (2D) coordinates; collapsing the input lightfield into the spatial sequence of surface layers of the LSLF set, wherein each surface layer in the spatial sequence of surface layers includes a depth data portion, a transparency data portion and a color data portion; encoding the LSLF set into a coded bitstream, wherein the coded bitstream causes a recipient device of the coded bitstream to generate a display image of the volumetric scene for rendering on an image display.

2. The method as recited in claim 1 , wherein at least one surface layer in the spatial sequence of surface layers includes an overlayering portion formed by a main surface at a main depth and a set of one or more additional surfaces at one or more different depths from the main depth; wherein each surface in the set of one or more additional surfaces carries additional depth, view-dependent transparency and view-dependent color data to be used along with main depth, transparency and color data carried in the main surface in blending operations that provide more accurate view-dependent colors than colors carried in the main surface alone.

3. The method as recited in claim 2, wherein a total number of surfaces in the set of additional surfaces is set based at least in part on a target expressive capacity of theLSLF set representing the main surface.

4. The method as recited in claim 2 or 3, wherein the surface layer includes a nonoverlayered portion not covered by any set of additional surfaces; wherein the overlayered portion includes visual information of greater view-dependency for the original input lightfield than the non-overlayered portion.

5. The method as recited in any one of claims 1-4, wherein at least one surface layer in the spatial sequence of surface layers represents one of: a surface adjacent to a real surface of a visual object depicted in the volumetric scene, or a surface not curvilinear to any real surface of any visual object depicted in the volumetric scene.

6. The method as recited in any one of claims 1-5, wherein the transparency data portion includes alpha channel values.

7. The method as recited in any one of claims 1-6, wherein the display image is generated through blending the spatial sequence of surface layers in the LSLF set from back to front based on respective depth and transparency data carried by the surface layers in the spatial sequence of surface layers.

8. The method as recited in any one of claims 1-7, wherein the LSLF set is specifically built for a particular spatial neighborhood in reference to a specific target view using a specific spatial mapping from a 3D geometry of the volumetric scene to 2D geometries of the surface layers in the LSLF set; wherein the specific spatial mapping represents one of: a perspective mapping, an orthographic mapping, a radial projection, an isometric projection, a fisheye projection, or another spatial projection.

9. The method as recited in claim 8, wherein one or both of the specific target view andthe specific spatial mapping are selected based at least in part on one or more synthesized views to be reconstructed from the LSLF set with relatively high fidelity by one or more recipient devices of the coded bitstream.

10. The method as recited in claim 8 or 9, wherein at least one of the specific target view and the specific spatial mapping is selected based at least in part on constraining a total amount of data to be carried with the LSLF set in the coded bitstream to a specific data size limit.

11. The method as recited in any one of claims 1-10, wherein multiple LSLF sets corresponding to multiple target views are generated from the same input lightfield; wherein synthesized view data communicated from the downstream device is used to determine a specific target view used to select the LSLF set from among the multiple LSLF sets.

12. The method as recited in claim 11, wherein one or more surface layer portions occluded to the specific target view are removed from the set of surface layers in the LSLF sets from being encoded into the coded bitstream.

13. A method, comprising: decoding a layered surface lightfield (LSLF) set from a coded bitstream, wherein the LSLF set includes a spatial sequence of surface layers with respective depth data portions, transparency data portions and color data portions, wherein at least some of the surface layers spatially conform to, or are curvilinear with, real surfaces of physical or visual objects depicted in the scene, and wherein each surface layer in the spatial sequence of surface layers is parametrized using two-dimensional (2D) coordinates; using the LSLF set to generate a display image;rendering the display image on an image display.

14. The method as recited in claim 13, wherein at least one surface layer in the spatial sequence of surface layers includes an overlayering portion formed by a main surface at a main depth and a set of one or more additional surfaces at one or more different depths from the main depth; wherein each surface in the set of one or more additional surfaces carries additional depth, view-dependent transparency and view-dependent color data to be used along with main depth, transparency and color data carried in the main surface in blending operations that provide more accurate view-dependent colors than colors carried in the main surface alone.

15. The method as recited in claim 13 or 14, wherein the spatial sequence of surface layers is converted into a target three-dimensional (3D) visual data representation; wherein the display image is generated by performing blending operations with the target 3D data representation.

16. An apparatus configured to perform the method of any one of claims 1-15.

17. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of the method of any one of claims 1-15.

18. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any one of claims 1-15.

Citation Information

Patent Citations

  • Layered scene decomposition CODEC with layered depth imaging

    US11252392B2