Layered scene decomposition coding system and method

By dividing the light field data into multiple layers and sampling and encoding it, the high bandwidth and latency problems of light field displays are solved, enabling efficient real-time transmission and high-resolution display on light field displays.

CN113748682BActive Publication Date: 2025-10-21AVALON HOLOGRAPHICS INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202080016205.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-22
Filing Date
2020-02-22
Publication Date
2025-10-21
Estimated Expiration
2040-02-22

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the high bandwidth requirements and real-time transmission latency issues of light field displays, resulting in the inability to provide high-resolution multidimensional content in real time on light field displays.

Method used

The light field data is divided into multiple layers, each representing a different part of the scene. The position of the sub-parts is determined based on the geometry of the objects. A second dataset is generated using sampling and encoding methods to reduce the amount of data and achieve efficient transmission and reconstruction.

Benefits of technology

It enables the real-time delivery of high-resolution light field images on light field displays, reduces system transmission latency and increases bandwidth, and is suitable for video streaming and real-time interactive games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113748682B_ABST
    Figure CN113748682B_ABST
Patent Text Reader

Abstract

Systems and methods are provided for CODECs for driving real-time light field displays for multi-dimensional video streaming, interactive gaming, and other light field display applications that apply a layered scene decomposition strategy. As the distance between a given layer and the display surface increases, multi-dimensional scene data is separated into multiple layers of data with increasing depth. The layers of data are sampled using a plenoptic sampling scheme and rendered using hybrid rendering (e.g., perspective and skew rendering) to encode the light field corresponding to each layer of data. The resulting compressed (layered) core representation of the multi-dimensional scene data is produced at a predictable rate, reconstructed and merged in real-time on the light field display by applying a view synthesis protocol that includes edge-adaptive interpolation to reconstruct the pixel array in stages (e.g., columns then rows) from reference elemental images.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority claim

[0002] This application claims priority to U.S. patent application serial number 62 / 809,390, filed on February 22, 2019, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure relates to image (light field) data encoding and decoding, including data compression and decompression systems and methods for providing interactive multi-dimensional content at a light field display. Background Art

[0004] Autostereoscopic, high-angular resolution, wide field of view (FOV), multi-view displays provide users with an improved visual experience. Three-dimensional displays capable of passing the 3D Turing Test (described by Banks et al.) will require light field representations to replace the two-dimensional images projected by standard, existing displays. Realistic light field representations require significant bandwidth to transmit the display data, which can contain at least a billion pixels of data. These bandwidth requirements currently exceed the bandwidth capabilities of previously known technologies; the upcoming consumer video standard is 8K Ultra High Definition (UHD), which provides only 33.1 megapixels of data per display.

[0005] Compressing data for transmission is known in the art. Data may be compressed for various types of transmission, such as, but not limited to, transmitting data over long distances via the Internet or Ethernet networks, or transmitting composite multi-view images created by a graphics processing unit (GPU) and transmitted to a display device. Such data can be used for video streaming, real-time interactive gaming, or any other light field display.

[0006] Several encoder-decoders (CODECs) for compressed light field transmission are previously known in the art. Olsson et al. teach compression techniques in which the entire light field dataset is processed to reduce redundancy and produce a compressed representation. Subcomponents of the light field (i.e., elemental images) are treated as video sequences to exploit redundancy using standard video coding techniques. Vetro et al. teach multi-view specializations of compression standards that exploit redundancy between light field subcomponents to achieve better compression rates, but at the expense of more intensive processing. These techniques may not achieve sufficient compression ratios, and when good compression ratios are achieved, the encoding and decoding processes exceed real-time rates. These methods assume that the entire light field exists on a storage disk or memory before being encoded. Therefore, large light field displays requiring a large number of pixels introduce excessive delay when reading from the storage medium.

[0007] In order to overcome hardware limitations for real-time delivery of multi-dimensional content, various methods and systems are known, however, these methods and systems present their own limitations.

[0008] U.S. Patent No. 9,727,970 discloses a distributed parallel (multi-processor) computing method and apparatus for generating a hologram by dividing 3D image data into data groups, calculating hologram values ​​at different locations to be displayed on a holographic plane from the data groups, and summing the values ​​at each location to generate a holographic display. As the disclosure focuses on generating holographic displays, the strategy employed involves manipulating light at a scale smaller than the light field, in this case characterized by sorting and partitioning the data by color, followed by a color image plane, and then further partitioning the plane image into sub-images.

[0009] US Patent Publication No. 20170142427 describes content-adaptive light field compression based on collapsing multiple elemental images (hogels) into a single hogel. This disclosure describes achieving a guaranteed compression rate, however, the image loss is variable and there is no guarantee that redundancy can be exploited in the combined hogels as disclosed.

[0010] U.S. Patent Publication No. 20160360177 describes a method for full parallax compressed light field synthesis using depth information and relates to the application of a view synthesis method for creating a light field from a set of elemental images that form a subset of a total set of elemental images. The view synthesis technique described herein does not describe or provide methods for handling reconstruction artifacts caused during backwarping.

[0011] US Patent Publication No. 20150201176 describes methods for full-parallax compressed light field 3D imaging systems that subsample elemental images in a light field based on the distances of objects in the captured scene. While these methods describe the possibility of downsampling the light field using simple conditions that can increase encoding speed, in the worst case, there are 3D scenes where no downsampling occurs, and encoding will fall back to transform coding techniques that rely on having the entire light field present before encoding.

[0012] There remains a need for increased data transmission capabilities, improved data encoder-decoders (CODECs), and methods of achieving improved data transmission and CODEC capabilities for delivering multi-dimensional content to light field displays in real time. Summary of the Invention

[0013] The present invention generally relates to 3D image data encoding and decoding for driving light field displays in real time, which overcomes or can be implemented with current hardware limitations.

[0014] The present disclosure aims to provide a CODEC with reduced system transmission latency and high bandwidth rates to provide light field generation in real time and at good resolution on a light field display for applications in video streaming and real-time interactive gaming. Light field or 3D scene data is deconstructed into subsets, which can be called layers (corresponding to layered light fields) or data layers, sampled and rendered to compress the data for transmission, and then decoded to construct and merge light fields corresponding to the data layers in the light field display.

[0015] According to one aspect, there is provided a computer-implemented method comprising:

[0016] receiving a first data set comprising a three-dimensional description of a scene;

[0017] dividing the first data set into a plurality of layers, each layer representing a different portion of the scene at a different position relative to a reference position;

[0018] dividing data corresponding to at least one of the layers into a plurality of sub-portions, wherein positions of particular sub-portions are determined based on the geometry of at least a portion of an object represented within the scene; and

[0019] The plurality of layers and the plurality of sub-portions are encoded to generate a second data set.

[0020] According to another aspect, there is provided a computer-implemented method comprising:

[0021] receiving a first data set comprising a three-dimensional description of a scene, the first data set comprising information about normal directions on surfaces in the scene,

[0022] The normal direction is expressed relative to a reference direction, where

[0023] At least some of the surfaces have non-Lambertian reflectance characteristics;

[0024] dividing the first data set into a plurality of layers, each layer representing a portion of the scene at a position relative to a reference position; and

[0025] The plurality of layers are encoded to generate a second data set, wherein a size of the second data set is smaller than a size of the first data set.

[0026] According to another aspect, a light field image rendering method is provided, comprising the following steps:

[0027] Divide the 3D surface description of the scene into multiple layers, each with an associated light field and sampling scheme;

[0028] further dividing the at least one layer into a plurality of sub-portions, each sub-portion having an associated light field and sampling, wherein a position of a particular sub-portion is determined based on a geometry of at least a portion of an object represented within the scene;

[0029] rendering a first set of pixels for each layer and each sub-portion according to the sampling scheme and corresponding to the sampled light field, including additional pixel information;

[0030] reconstructing the sampled light field for each layer and sub-portion using the first set of pixels; and

[0031] The reconstructed light field images are merged into a single output light field image.

[0032] According to another aspect, there is provided a computer-implemented method comprising:

[0033] receiving a first data set comprising a three-dimensional description of a scene;

[0034] dividing the first data set into a plurality of layers, each layer representing a portion of the scene at a position relative to a reference position;

[0035] For each of the plurality of layers, obtaining one or more polygons representing a corresponding portion of an object in the scene;

[0036] determining a view-independent representation based on the one or more polygons; and

[0037] The view-independent representation is encoded as part of a second data set, wherein a size of the second data set is smaller than a size of the first data set.

[0038] According to another aspect, there is provided a computer-implemented method comprising:

[0039] receiving a first data set comprising a three-dimensional description of a scene;

[0040] dividing the first data set into a plurality of layers, each layer representing a portion of the scene at a position relative to a reference position; and

[0041] The plurality of layers are encoded by performing a sampling operation on the layers to generate a second dataset, including:

[0042] Use the effective resolution function to determine the appropriate sampling rate; and

[0043] Downsample the element image associated with the layer using the appropriate sampling rate,

[0044] The size of the second data set is smaller than the size of the first data set.

[0045] According to another aspect, there is provided a computer-implemented method comprising:

[0046] receiving a first data set comprising a three-dimensional description of a scene, the first data set comprising information regarding transparency of surfaces in the scene;

[0047] dividing the first data set into a plurality of layers, each layer representing a portion of the scene at a position relative to a reference position; and

[0048] The plurality of layers are encoded to generate a second data set, wherein a size of the second data set is smaller than a size of the first data set.

[0049] According to another aspect, a light field image rendering method is provided, comprising the following steps:

[0050] Divide the 3D surface description of the scene into multiple layers, each with an associated light field and sampling scheme;

[0051] further dividing the at least one layer into a plurality of sub-portions, each sub-portion having an associated light field and sampling, wherein a position of a particular sub-portion is determined based on a geometry of at least a portion of an object represented within the scene;

[0052] rendering a first set of pixels for each layer and each sub-portion according to the sampling scheme and corresponding to the sampled light field, including additional pixel information;

[0053] reconstructing the sampled light field for each layer and sub-portion using the first set of pixels; and

[0054] The reconstructed light field images are merged into a single output light field image.

[0055] Implementations may include one or more of the following features.

[0056] In an embodiment of the method, the second data set is transmitted to a remote device for rendering the scene at a display device associated with the remote device.

[0057] In an embodiment of the method, encoding a layer or a sub-portion comprises performing a sampling operation on a corresponding portion of said first data set.

[0058] In an embodiment of the method, the sampling operation is based on a target compression ratio associated with the second data set.

[0059] In an embodiment of the method, encoding the plurality of layers or the plurality of sub-portions includes performing a sampling operation on corresponding portions of the first data set, wherein performing the sampling operation includes:

[0060] Rendering a set of pixels to be encoded using ray tracing; selecting a plurality of element images from a plurality of element images such that the set of pixels is rendered using the selected plurality of element images; and

[0061] The set of pixels is sampled using a sampling operation.

[0062] In an embodiment of the method, the sampling operation comprises selecting a plurality of element images from corresponding portions of the plurality of element images according to a plenoptic sampling scheme.

[0063] In an embodiment of the method, performing the sampling operation includes:

[0064] determining an effective spatial resolution associated with the layer or sub-portion; and

[0065] A plurality of element images are selected from corresponding portions of the plurality of element images according to the determined angular resolution.

[0066] In an embodiment of the method, the angular resolution is determined as a function of the directional resolution associated with the portion of the scene associated with said layer or sub-portion.

[0067] In an embodiment of the method, the angular resolution is determined as a field of view associated with the display device.

[0068] In an embodiment of the method, the three-dimensional description comprises light field data representing images of the plurality of elements.

[0069] In an embodiment of the method, each of the plurality of element images is captured by one or more image acquisition devices.

[0070] In an embodiment of the method, the first data set comprises information on normal directions on surfaces comprised in the scene, said normal directions being expressed relative to a reference direction.

[0071] In an embodiment of the method, the reflective properties of at least some of the surfaces are non-Lambertian.

[0072] In an embodiment of the method, encoding the layer or sub-portion further comprises:

[0073] obtaining, for the layer or sub-portion, one or more polygons representing corresponding portions of objects in the scene;

[0074] determining a view-independent representation based on the one or more polygons; and

[0075] A view-independent representation is encoded in the second data set.

[0076] In an embodiment of the method, the method further comprises:

[0077] receiving the second data set;

[0078] decoding portions of the second data set corresponding to each said layer and each said sub-portion; combining the decoded portions into a representation of a light field image; and

[0079] The light field image is presented on a display device.

[0080] In an embodiment of the method, the method further comprises:

[0081] receiving user input indicating a user's position relative to the light field image; and

[0082] The light field image is updated based on the user input before being presented on the display device.

[0083] In an embodiment of the method, a layer located closer to the display surface achieves a lower compression ratio than a layer of the same width located further away from the display surface.

[0084] In an embodiment of the method, the plurality of layers of the second data set comprises a light field.

[0085] In an embodiment of the method, further comprising merging the light fields to create a final light field.

[0086] In an embodiment of the method, the dividing of the layers includes limiting the depth range of each layer.

[0087] In an embodiment of the method, layers located closer to the display surface are narrower in width than layers located further from the display surface.

[0088] In an embodiment of the method, dividing the first data set into a plurality of layers maintains a uniform compression rate throughout the scene.

[0089] In an embodiment of the method, dividing the first data set into a plurality of layers comprises dividing the light field display into an inner frustum volume set of layers and an outer frustum volume set of layers.

[0090] In an embodiment of the method, the method is used to generate a synthetic light field for a multi-dimensional video stream, a multi-dimensional interactive game, real-time interactive content, or other light field display scenarios.

[0091] In an embodiment of the method, the synthetic light field is generated only in the active viewing zone.

[0092] According to one aspect, there is a computer method for rendering a light field image, comprising:

[0093] Divide the 3D surface description of the scene into multiple layers, each with an associated light field and sampling scheme;

[0094] further dividing the at least one layer into a plurality of sub-portions, each sub-portion having an associated light field and sampling, wherein a position of a particular sub-portion is determined based on a geometry of at least a portion of an object represented within the scene;

[0095] rendering a first set of pixels for each layer and each sub-portion according to the sampling scheme and corresponding to the sampled light field, including additional pixel information;

[0096] reconstructing the sampled light field for each layer and sub-portion using the first set of pixels; and

[0097] The reconstructed light field images are merged into a single output light field image.

[0098] In an embodiment of the method, the first set of pixels and associated additional pixel information are divided into subsets, whereby the reconstruction for each layer of the sampled light field and merged to create some subset of the output light field image is performed using pixels from a single subset in the cache.

[0099] In an embodiment of the method, further comprising reconstructing the sampled light field for each layer by reprojecting pixels in the first group from the cache to create some subset of the output light field image.

[0100] In an embodiment of the method, further comprising performing re-projecting the pixels using a warping process along a single dimension in a first set of pixels followed by a second warping process in a second dimension in the first set of pixels. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] These and other features of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings.

[0102] Figure 1 : is a schematic representation (block diagram) of an implementation of a hierarchical scene decomposition (CODEC) system according to the present disclosure.

[0103] Figure 2 : is a schematic top view of the inner and outer frustum volumes of a light field display.

[0104] Figure 3A : Schematically illustrates the application of edge-adaptive interpolation for pixel reconstruction according to the present disclosure.

[0105] Figure 3B : Illustrate the processing flow for reconstructing the pixel array.

[0106] Figure 4 : Schematically illustrates an element image specified by a sampling scheme within a pixel matrix as part of an image (pixel) reconstruction process according to the present disclosure.

[0107] Figure 5 : Schematically illustrates the column-by-column reconstruction of a pixel matrix as part of an image (pixel) reconstruction process according to the present disclosure.

[0108] Figure 6: illustrates the subsequent row-by-row reconstruction of the pixel matrix as part of the image (pixel) reconstruction process according to the present disclosure.

[0109] Figure 7 : Schematically illustrates an exemplary CODEC system implementation according to the present disclosure.

[0110] Figure 8 : Schematically illustrates an exemplary hierarchical scene decomposition of an image dataset associated with the inner frustum light field of a display (hierarchical scheme of ten layers).

[0111] Figure 9 : Schematically illustrates an exemplary hierarchical scene decomposition of image data relating to the inner frustum and outer frustum light field regions of a display, respectively (two hierarchical schemes of ten layers).

[0112] Figure 10 : illustrates an exemplary CODEC processing flow according to the present disclosure.

[0113] Figure 11 : Illustrate an exemplary process flow for encoding 3D image (scene) data to produce a layered and compressed core coded (light field) representation according to the present disclosure.

[0114] Figure 12 : illustrates an exemplary process flow for decoding a core encoded representation to construct (display) a light field at a display according to the present disclosure.

[0115] Figure 13 : illustrates an exemplary process flow for encoding and decoding residual image data for use with core image data to produce a (display / final) light field at a display according to the present disclosure.

[0116] Figure 14 : illustrates an exemplary CODEC processing flow including layered depth images according to the present disclosure.

[0117] Figure 15 : illustrates an exemplary CODEC processing flow including specular light calculation according to the present disclosure.

[0118] Figure 16 : illustrates an alternative exemplary CODEC processing flow including specular light calculation according to the present disclosure.

[0119] Figure 17 : illustrates an exemplary CODEC processing flow including view-independent rasterization according to the present disclosure.

[0120] Figure 18 : illustrates an exemplary CODEC processing flow including performing a sampling operation using an effective resolution function according to the present disclosure.

[0121] Figure 19 : Illustrate viewer-based construction planes for measuring effective resolution at depth.

[0122] Figure 20 : Graphically illustrate the asymptotic properties of the effective resolution with respect to scene depth.

[0123] Figure 21 : illustrates an exemplary CODEC processing flow including transparency according to the present disclosure. DETAILED DESCRIPTION

[0124] The present invention generally relates to CODEC systems and methods for light field data or multi-dimensional scene data compression and decompression to provide efficient (fast) transmission and reconstruction of the light field at a light field display.

[0125] The various features of the present invention will become apparent from the following detailed description taken in conjunction with the illustrations in the accompanying drawings. The design factors, construction, and use of the hierarchical scene decomposition CODEC disclosed herein are described with reference to various examples representing implementations, which are not intended to limit the scope of the invention as described and claimed herein. Those skilled in the art will appreciate that other variations, examples, and embodiments of the invention not disclosed herein may be practiced in accordance with the teachings of this disclosure without departing from the scope and spirit of the invention.

[0126] definition

[0127] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0128] The word "a" or "an" when used herein with the term "comprising" can mean "one", but is also consistent with the meaning of "one or more", "at least one" and "one or more than one".

[0129] As used herein, the terms "comprising," "having," "including," and "containing," and grammatical variations thereof, are inclusive or open-ended and do not exclude additional, unrecited elements and / or method steps. The term "consisting essentially of," when used herein in connection with a composition, apparatus, article, system, use, or method, indicates that additional elements and / or method steps may be present, but that such addition would not materially affect the functionality of the recited composition, apparatus, article, system, method, or use. The term "consisting of," when used herein in connection with a composition, apparatus, article, system, use, or method, excludes the presence of additional elements and / or method steps. A composition, apparatus, article, system, use, or method described herein as comprising certain elements and / or steps may also consist essentially of those elements and / or steps in certain embodiments, and consist of those elements and / or steps in other embodiments, whether or not such embodiments are specifically mentioned.

[0130] As used herein, the term "about" refers to a variation of approximately + / - 10% from a given value. It should be understood that such variation is always included in any given value provided herein, whether or not specifically mentioned.

[0131] Unless otherwise indicated herein, the recitation of ranges herein is intended to convey the range and individual values ​​falling within the range to the same place value as the number used to express the range.

[0132] The use of any example or exemplary language, such as "for example," "exemplary embodiment," "illustrative embodiment," and "for example," is intended to illustrate or represent aspects, embodiments, variations, elements or features related to the present invention and is not intended to limit the scope of the invention.

[0133] As used herein, the terms "connect" and "connected" refer to any direct or indirect physical association between elements or features of the present disclosure. Thus, these terms can be understood to mean that elements or features are partially or completely contained within, attached, coupled, arranged, joined together, communicate, operatively associated, etc., even if there are other elements or features intervening between the elements or features described as connected.

[0134] As used herein, the term "light field" refers, at a basic level, to a function that describes the amount of light flowing through a point in space in each direction, unobstructed. Thus, a light field represents radiance as a function of the position and direction of light in free space. Light fields can be generated synthetically through various rendering processes, or they can be captured from a light field camera or light field camera array.

[0135] A light field can be most generally described as a mapping between a set of points in 3D space and a corresponding set of directions to one or more sets of energy values. In practice, these energy values ​​will be red, green, blue intensities, or potentially other wavelengths of radiation.

[0136] As used herein, the term "light field display" is a device that reconstructs a light field from a finite number of light field radiation samples input to the device. The radiation samples represent the red, green, and blue (RGB) color components. For reconstruction in a light field display, the light field can also be understood as a mapping from a four-dimensional space to a single RGB color. These four dimensions include the vertical and horizontal dimensions of the display (x, y) and two dimensions that describe the directional components of the light field (u, v). The light field is defined as the function:

[0137] LF: (x, y, u, v) → (r, g, b)

[0138] For a fixed x f ,y f ,LF(x f ,y f ,u,v) represents a two-dimensional (2D) image called an “element image”. The element image is a f ,y f A directional image of the light field at a given location. When multiple element images are connected side by side, the resulting image is called an "overall image." The overall image can be understood as the entire light field required for light field display.

[0139] As used herein, the term "scene description" refers to a geometric description of a three-dimensional scene, which can be a potential source for rendering light field images or videos. Such a geometric description can be represented by, but is not limited to, points, quadrilaterals, and polygons.

[0140] As used herein, the term "display surface" may refer to a set of points and directions defined by the physical spacing of a planar display plane and its individual lightfield hogel elements, as in a traditional 3D display. In the present disclosure, displays as described herein may be formed on curved surfaces, so that the set of points will reside on a curved display surface, or any other desired display surface geometry that can be imagined. In an abstract mathematical sense, a light field can be defined and represented on any geometric surface and does not necessarily correspond to a physical display surface with actual physical energy emission capabilities.

[0141] As used herein, the term "elemental image" refers to a two-dimensional (2D) image LF(x f ,y f ,u,v), for a fixed x f ,y f ,LF(x f ,y f , u, v). The element image is from a fixed xf ,y f Orientation image of the light field at the position.

[0142] As used herein, the term “overall image” refers to multiple element images connected side by side, so the resulting image is called an “overall image.” The overall image can be understood as the entire light field required for light field display.

[0143] As used herein, the term "layer" refers to any two parallel or non-parallel boundaries, with uniform or variable width, parallel or non-parallel to the display surface.

[0144] As used herein, the term "pixel" refers to the light source and light-emitting mechanism used to create a display.

[0145] It is contemplated that any embodiment of the compositions, apparatuses, articles, methods, and uses disclosed herein may be implemented by one skilled in the art as such, or by making such variations or equivalents without departing from the scope and spirit of the invention.

[0146] Layered Scene Decomposition (LSD) CODEC system and method

[0147] A codec according to the present disclosure applies strategies that utilize known sampling, rendering, and view synthesis methods to generate light field displays, adapted for use with the novel hierarchical scene decomposition strategy as disclosed herein, including its derivation, implementation, and application.

[0148] 3D display

[0149] Conventional displays previously known in the art consist of spatial pixels that are substantially uniformly spaced and organized into a two-dimensional array, allowing for ideally uniform sampling. In contrast, three-dimensional displays require both spatial and angular samples. While the spatial sampling of a typical 3D display is consistent, the angular sampling is not necessarily considered uniform in terms of the display's footprint in angular space. For a review of various light field parameterizations for angular light distribution, see U.S. Patent No. 6,549,308.

[0150] The angular samples, also called the directional components of the light field, can be parameterized in various ways, such as the planar parameterization taught by Gortler et al. in "Lumigraph". When the light field function is discretized with respect to position, the light field can be understood as a regularly spaced array of planar parameterized pinhole projectors, as taught by Chai in "Plenoptic Sampling". For a fixed x f ,y f Element image LF(x f ,y f, u, v) represents a two-dimensional image, which can be understood as the image projected by a pinhole projector with arbitrary light parameterization. For a light field display, a continuous elemental image is represented by a finite number of light field radiation samples. For an idealized planar parameterized pinhole projector, this finite number of samples is mapped to the image plane as a regularly spaced array (regular spacing within the plane does not correspond to regular spacing in the corresponding angular direction space).

[0151] In the case of a typical 3D light field display, the set of points and directions would be defined by the physical spacing of the planar display plane and its individual light field hogel elements. However, given that displays can be formed on curved surfaces, the set of points would reside on the curved display surface, or any other desired display surface geometry that can be imagined. In an abstract mathematical sense, a light field can be defined and represented on any geometric surface, not necessarily corresponding to a physical display surface with actual physical energy emission capabilities. The concept of surface light fields in the literature illustrates this situation, as shown by Chen et al.

[0152] The consideration of a plane parameterization is not intended to limit the scope or spirit of the present disclosure, as the directional component of the light field can be parameterized by a variety of other arbitrary parameterizations. For example, lens distortion or other optical effects in a physically embodied pinhole projector can be modeled as distortions of a plane parameterization. Furthermore, display components can be defined by deformation functions, such as taught by Clark et al. in "A transformation method for the reconstruction of functions from nonuniformly spaced samples."

[0153] The deformation function α(u, v) defines the warped plane parameterization of the pinhole projector, producing an arbitrary alternating angular distribution of directional rays in the light field. The angular distribution of rays propagating from the light field pinhole projector is determined by the focal length f of the pinhole projector and the corresponding two-dimensional deformation function α(u, v).

[0154] An autostereoscopic light field display that projects a light field to one or more users is defined as:

[0155] D=(M x , M y , N u , N v ,f,α,D LP )

[0156] Among them (M x , M y ) are the horizontal and vertical dimensions of the display spatial resolution, (N u , N v) are the horizontal and vertical dimensions of the display’s angular resolution components. The display is a set of idealized light field projectors with a spacing D LP , focal length f, and a deformation function α that defines the ray direction distribution of the light field projected by the display.

[0157] Driving light field display D=(M x , M y , N u , N v ,f,α,D LP ) light field LF(x, y, u, v) needs to have M in the x direction x Light field radiation samples, M in the y direction y light field radiation samples, and N in the u and v directions u , and N v light field radiation samples. Although D is defined with a single deformation function α, if there is significant microlens variation in the actual pinhole projector, resulting in significant variations in the angular ray distribution from one microlens to another, then each light field plane parameterized pinhole projector within the idealized light field pinhole projector array may have a unique deformation function α.

[0158] Light field display rendering

[0159] In "Fast computer graphics rendering for full parallax spatial displays", Halle et al. provide a method for rendering objects located within the inner and outer frustum volumes of a display. Figure 2 A light field display representing objects within a volumetric region defined by two separate viewing cones is illustrated, with the inner frustum volume (110) being located behind a display surface (300) (i.e., within the display) and the outer frustum volume (210) being located in front of the display surface (i.e., outside the display). As shown, various objects (schematically shown as prismatic and circular) are located at different depths from the display surface (300).

[0160] Halle et al. teach a dual frustum rendering technique where the inner frustum volume and the outer frustum volume are rendered as two different light fields. O (x, y, u, v) and the outer frustum volume LF P (x, y, u, v) are recombined into a single light field LF(x, y, u, v) through a deep merging process.

[0161] The technique uses a pinhole camera rendering model to generate individual elemental images of a light field. Each elemental image (i.e., each rendered planar parameterized pinhole projector image) requires the use of two cameras: one camera capturing the inner frustum volume and one camera capturing the outer frustum volume. Halle et al. teach the use of a standard orthographic camera and its conjugate pseudo-photographic camera to render pinhole projector images over a sampled region of the light field. For a pinhole camera C, the corresponding conjugate camera is denoted by C * .

[0162] In order to capture elemental images within a light field display using a projector parameterized by a deformation function α, a general pinhole camera based on a reparameterized idealized plane-parameterized pinhole camera is used. As taught by Gortler et al., the rays of a pinhole camera C with focal length f are defined by a parameterization created by two parallel planes. Pinhole camera C captures image I C (u, v), where (u, v) are coordinates in the ray parameterization plane. Universal pinhole camera C α Based on a planar parameterized camera that is deformed using a two-dimensional, continuous, reversible time deformation function, as taught by Clark et al. Using the deformation function α(u, v), the inverse function is γ(u, v). Therefore, C α Image, I Cα =I C (α(u, v)).

[0163] Given a general pinhole camera C α , forming a conjugate generalized camera To complete the double frustum rendering. From the universal pinhole camera M x ×M y The mesh generated view is rendered to render the light field for the light field display.

[0164] Therefore, for a given light field display D = (M x , M y , N u , N v ,f,α,D LP ), the set of all generic pinhole camera pairs that must be rendered to produce the light field LF(x, y, u, v) is defined as:

[0165]

[0166] A set of orthographic cameras (O = {(C α (x, y)|1≤x≤M x , 1≤y≤M y}) captures the light field image corresponding to the inner frustum volume, and a set of conjugate generalized cameras An image corresponding to the outer frustum volume is captured, and as described above, the inner and outer frustum volumes are merged into a single light field.

[0167] Real-time rendering

[0168] It is believed that a usable light field display may require at least 1-10 billion pixels, each representing a different directional ray of light from the light field. Considering a modest interactive frame rate of 30Hz and assuming 24 bits per raw ray pixel, this results in a significant bandwidth requirement of (10 billion pixels) x (24 bits / pixel) x (30 frames / second) = 720 Gbits / second of bandwidth. Due to the eventual demand for higher fidelity displays, this requirement could realistically scale to 100s of Tbits / second as this display technology enters the consumer market and continues to advance in visual fidelity.

[0169] Interactive computer graphics rendering is a computational process that, at least traditionally, requires computing a simulation of a virtual camera imaging a scene. A scene is typically described as a collection of light sources and surfaces or volumes with various materials, colors, and physical optical properties, as well as various viewing camera positions. This rendering computation must be performed quickly enough to produce interactive frame rates (e.g., at least 30 Hz). Rendering fidelity can be adjusted based on how approximated the light transport computations are, which of course reduces computational requirements as more approximations are used. It is for this reason that interactive computer graphics typically have lower visual fidelity than offline rendered graphics where very high-fidelity light transport models are used.

[0170] The need for interactivity implies a specific frame rate with corresponding bandwidth (usually at least 20-30Hz, but often desired to be higher), and also implies reducing latency to support instant graphical response to user input. The combination of high bandwidth and latency requirements creates certain computational challenges.

[0171] In traditional 2D computer graphics, the challenge of low-latency, high-frame-rate graphics has led to the widespread use of specialized hardware designed to accelerate interactive rendering computations, known as graphics processing units (GPUs). These specialized architectures can produce higher visual fidelity at interactive rates than the general-purpose central processing units (CPUs) used in modern computers. While their performance is impressive, these architectures are ultimately optimized for their specific task: rendering a single camera image of a scene at a high frame rate while maximizing visual quality.

[0172] For light field displays, the rendering problem becomes rendering images produced by a virtual light field camera. A light field camera (defined in more detail elsewhere) can be viewed as an array of many traditional 2D camera views. This more general camera model results in computations utilizing significantly different geometries. As a result, the computations do not map well to the framework of existing accelerated computer graphics hardware.

[0173] Traditionally, a rendering computational pipeline process has been defined. This has traditionally been based on rasterization, but ray tracing pipelines have also been standardized (e.g., more recently DirectX Raytracing). In either case, the computing hardware architecture is tailored to the form of these pipelines and their associated required computations, with the ultimate goal of producing a 2D image at video frame rates.

[0174] What is needed is a different pipeline for interactive light field rendering computations that can be implemented with a minimal hardware footprint in order to minimize the cost, size, weight, and power requirements of the rendering architecture. These requirements are driven by the desire to eventually create a consumer product within a comparable price range.

[0175] When considering the pipeline for light field rendering and the large data rates required, a major bottleneck involves the potential memory bandwidth required. In traditional 2D video rendering and processing, buffering entire frames (or even subsequences of frames) in double data rate (DDR) memory is a common operation. It has been observed that the data rate of DDR and its cost-to-capacity relationship make it well suited for these types of applications. However, the light field bandwidth requirements discussed earlier indicate that significant DDR buffering can be very limited in terms of physical footprint and cost.

[0176] The first step in rendering is typically to load a scene description representation from DDR memory or some other storage that is slow relative to the clock rate of the computing hardware. A remarkable aspect of light field rendering is that each light field camera rendering pass can be considered a set of more or less conventional 2D camera rendering passes. Naturally, each of these 2D cameras (hogels) must be rendered twice, for the inner and outer hogel, following the "double frustum rendering" approach suggested by Halle. The number of rays is 2 for each direction represented by the display. An alternative solution, obvious from the prior art, is to define inner and outer far clipping planes, and have rays cast from the outer far clipping plane, through the hogels on the display surface and end at the inner far clipping plane (or vice versa). This results in one ray per pixel.

[0177] In the worst case, each of these 2D camera rendering passes in the array requires loading the entire scene description. Even in the more optimistic case, and especially compared to traditional 2D rendering where the scene is typically accessed at most a small number of times per frame, repeatedly loading the scene description from DDR or other slow memory naturally results in large bandwidth requirements.

[0178] Given the presence of seemingly redundant memory accesses, it's worth considering whether this situation can be addressed through the use of caching strategies. In high-performance computing, when computations are structured such that data is redundantly loaded in a coherent, predictable pattern, caching this data in smaller but faster storage (typically located directly on the die of the chip performing the computation) can significantly alleviate DDR or slow memory bandwidth limitations. In the context of light field rendering, each hogel's elemental image rendering requires the same scene description in the worst case, identifying significant potential redundancy that can be effectively exploited through caching. Modern ray tracing techniques for surface rendering of a single camera view of a scene are able to exploit cache coherence by virtue of the principle that coherent rays on the image plane typically intersect the same geometry (or at least for the primary intersection points). This same coherence is inherently exploited when rasterizing polygons to a single imaging camera, as a single polygon can be cached in hardware after loading, and all pixels it intersects can be computed during hardware-accelerated rasterization. In the context of light fields, it's clear that these same coherence principles can be exploited if ray tracing or rasterization is used to render the single 2D camera view that constitutes the light field.

[0179] It can also be observed that imaging rays from different hogels in a light field camera intersecting the same polygon exhibit additional coherent elements. It is proposed that this coherence can be exploited in a structured form to produce a higher performance interactive rendering system for light field displays. We show how this coherence can be exploited by buffering a structured intermediate rendering form of the output light field. This buffer is referred to herein as the light field surface buffer or surface buffer. Pixels in this buffer can also contain additional pixel information such as color, depth, surface coordinates, normals, material values, transparency values, and other possible scene information.

[0180] It is proposed that a surface buffer can be efficiently rendered using an efficient traditional 2D rendering ray tracing pipeline. The surface buffer is based on the concept of hierarchical scene decomposition and a sampling scheme as presented herein that specifies which pixels will constitute the surface buffer. Based on the analysis presented within this specification, it can be seen that with appropriately chosen hierarchical scene decomposition and sampling schemes, the resulting surface buffer can be determined to contain fewer pixels than required for the rendered output light field image frame, as this can be viewed as a form of data compression scheme.

[0181] Furthermore, an appropriately chosen layered scene decomposition and sampling scheme will result in the surface buffer containing samples of all surface areas in the scene visible to any hogel in the target light field camera's view. Such a surface buffer will contain data to enable reconstruction of the light field associated with each layer and layer sub-portion. Once reconstructed, these light fields can be merged into a single light field image representing the desired rendering output, as described elsewhere in this document.

[0182] It is further proposed that the resulting surface buffer can be partitioned into smaller subsets. This partitioning can occur in such a way that each subset of the surface buffer data can be used individually to reconstruct certain portions of the resulting output light field. A practical implementation involves partitioning the layers and sub-portions based on a ΔEI function, and then selecting a sampling scheme that includes a small number (e.g., 4) of element images per partition, which are then used to reconstruct the unsampled element images within the partition. If this partitioning is chosen appropriately, subsets of the surface buffer can be loaded into a faster cache, from which reconstruction and merging calculations can be performed without resorting to repeated loading from slower system memory.

[0183] So, in summary, an efficient approach to rendering light field video at interactive rates can be described as starting with a 3D description of the scene, rendering a surface buffer, and then rendering the final output frame by reconstructing layers and sub-portions from cached individual partitions of the surface buffer to create corresponding portions of the desired output light field image. When the rendering is structured in this form, as opposed to applying a brute force approach of performing light field rendering as many conventional 2D rendering passes, less slow memory bandwidth is required because the cache memory can be exploited in a structured manner through the surface buffer partitions.

[0184] View-Independent Rasterization

[0185] Maars et al. proposed a generalized multi-view rendering technique using view-independent rasterization. After point generation, we use the point representation to render multiple views in parallel. We perform point rendering by a) streaming the points directly to the pixel shader stage of VIR, or b) storing the points in a separate buffer and dispatching GPU computation threads (e.g., Figure 4 Our simple point rendering kernel reads the world space position of a point, then for each view, applies the corresponding view-projection matrix, snaps the projected position to the nearest neighbor in the view buffer, and performs z-buffering. The atomic function resolves race conditions caused by multiple points projecting onto the same texel.

[0186] The remaining challenges of the technique disclosed by Maars et al. relate to quality and speed. Implementing view-independent rasterization of a layer or subset of a three-dimensional description of a scene may include obtaining one or more polygonal representations based on the geometry of objects in the scene. A view-independent representation is generated based on one or more of these polygons. The generated view-independent representation is encoded as part of a compressed second dataset.

[0187] Figure 17 A computer-implemented method is described, comprising:

[0188] receiving a first data set (420) comprising a three-dimensional description of a scene;

[0189] dividing the first data set into a plurality of subsets, each subset representing a different portion of the scene at a different position relative to a reference position (429);

[0190] For each of the plurality of subsets, obtaining one or more polygons representing a corresponding portion of an object in the scene (430);

[0191] determining a view-independent representation based on the one or more polygons (431); and

[0192] The view-independent representation is encoded as part of a second data set, wherein the size of the second data set is smaller than the size of the first data set (432).

[0193] Data Compression for Light Field Displays

[0194] Piao et al. exploit prior physical properties of light fields to identify redundancy in the data. Based on the observation that elemental images representing adjacent points in space contain significant overlapping information, the redundancy is used to discard elemental images. This avoids performing computationally complex data transformations to identify the information to discard. This approach does not utilize the depth map information associated with each elemental image.

[0195] In "Compression of Full-Parallax Light-Field Displays," Graziosi et al. propose a standard for subsampling elemental images based on a simple pinhole camera cover geometry to reduce light field redundancy. The downsampling technique taught by Graziosi et al. is simpler than the complex basis decomposition of 2D image and video data commonly used in other CODEC schemes. When objects are located deep in the scene, the light field is sampled at a smaller rate. For example, when two separate pinhole cameras provide two different fields of view, there is little difference from one elemental image to the next, and the fields of view of the two pinhole cameras overlap. When the views are subsampled based on geometric (triangle) overlap, the pixels within the views are not compressed. Because these pixels can be large, Graziosi et al. compress the pixels using standard 2D image compression techniques.

[0196] Graziosi et al. teach that the sampling gap (ΔEI) between element images, based on the minimum depth d of the object, can be calculated as follows, where θ represents the field of view of the light field display and P represents the lens spacing of the integral imaging display:

[0197]

[0198] This strategy provides theoretically lossless compression for fronto-parallel planes when there are no image occlusions. As shown in the formula, the sampling gap increases with d, providing improved compression when fewer element images are required. For sufficiently small d, ΔEI can reach 0. Therefore, this downsampling technique does not guarantee compression. In scenes with multiple small objects, where the objects are close to the screen or at screen distances, each element image has at least some pixels with a depth of 0, and this technique does not provide any gain, that is, ΔEI = 0 throughout the entire image.

[0199] Graziosi et al. equate the rendering process to the initial encoding process. Instead of generating all the element images, this approach generates only the number needed to reconstruct the light field while minimizing any information loss. Depth maps are included in the element images selected for encoding, and missing element images are reconstructed using well-established warping techniques associated with depth-image-based rendering (DIBR). The selected element images are further compressed using a method similar to the H.264 / AVC approach, and the images are decompressed before the final DIBR-based decoding stage. While this approach provides improved compression rates with reasonable levels of signal distortion, it does not provide time-based performance results. This encoding and decoding approach does not provide good low-latency performance for high bandwidth rates. Furthermore, the approach is limited to single objects far from the display surface; in scenes with multiple overlapping objects and many objects close to the display surface, compression is forced to revert to H.264 / AVC-style encoding.

[0200] Chai teaches plenoptic sampling theory to determine the amount of angular bandwidth required to represent front-parallel plane objects at a specific scene depth. Zwicker et al. teach that the depth of field of a display is based on angular resolution, with higher resolution resulting in greater depth of field. Thus, objects close to the display surface can be adequately represented with lower angular resolution, while objects farther away require greater angular resolution. Zwicker et al. teach that the maximum display depth of field using an ideal projection lens based on a plane parameterization is:

[0201]

[0202] Among them, P l is the lens spacing, P p is the pixel pitch, and f is the focal length of the lens.u =N v ) in a three-dimensional display, N = P l / P p Therefore, Z DOF =fN.

[0203] To determine the angular resolution required to represent the full spatial resolution of the display, at a given depth d, the formula is rearranged to:

[0204]

[0205] Therefore, each focal distance into the scene adds another pixel of angular resolution required to fully represent the object at a given spatial resolution of the display surface.

[0206] Hierarchical scene decomposition sampling scheme

[0207] The sampling gap taught by Graziosi et al. and the plenoptic sampling theory taught by Zwicker et al. provide complementary light field sampling strategies: Graziosi et al. increase the downsampling of distant objects (ΔEI), while Zwicker et al. increase the downsampling of nearby objects (N res ). However, when downsampling a single light field representing a scene, the combination of these strategies does not guarantee compression. Therefore, the present disclosure divides the multidimensional scene into multiple layers. This division into multiple (data) layers is referred to herein as hierarchical scene decomposition. Where K1 and K2 are natural numbers, we define L = (K1, K2, L O , L P ), dividing the inner and outer frustum volumes of the 3D display. The inner frustum is divided into a set of K1 layers, where Each inner frustum layer consists of a layer at a distance from the display surface and The outer frustum is divided into a set of K2 layers, where Each outer frustum layer is composed of a layer at a distance from the display surface and A pair of boundaries parallel to the display surface is defined, where 1≤i≤K2. In alternative embodiments, the inner and outer frustum volumes may be partitioned by different layering schemes and the boundary pair may or may not be parallel to the display surface.

[0208] Each of the hierarchical scene decomposition layers has an associated light field (also referred to herein as a "light field layer") based on scene restrictions for the planar boundary regions of that layer. Consider a layer with an inner frustum or outer frustum layer The D of the light field display is (M x , My , N u , N v ,f,a,D LP ) of the hierarchical scene decomposition L = (K1, K2, L O , L P ), where 1≤i≤K1, 1≤j≤K2. Inner frustum light field A set of universal pinhole cameras O = {C a (x, y)|1≤x≤M x , 1≤y≤M y The equation is restricted to imaging only the space at a distance d from the light field display surface, where Therefore, for a fixed x, y and C a For the inner frustum of (x, y)∈O, we compute Similarly, the outer frustum light field A set of universal pinhole cameras The equation is restricted to imaging only the space at a distance d from the light field display surface, where Therefore, for a fixed x, y and C a For the outer frustum of (x, y)∈P, we compute

[0209] We can further define a light field set of inner and outer frustum regions relative to the layered scene decomposition L. Assume that the light field display D = (M x , M y , N u , N v ,f,a,D LP ) has a hierarchical scene decomposition L = (K1, K2, L O , L P ). The light field set in the inner frustum region is defined as The light field set in the outer frustum region is defined as

[0210] By definition, a layered scene decomposition generates a light field for each layer. For any layered scene decomposition, the orthographic camera generates an inner frustum volume light field, while the pseudoscopic camera generates an outer frustum volume light field. If the scene captured by these universal pinhole camera pairs consists only of opaque surfaces, then each point of the light field has an associated depth value that indicates the distance from the universal pinhole camera plane to the corresponding imaging point in space. When given a light field or hour, The depth map is formally defined as and The depth map is formally defined as Depth map Dm =∞, where there is no surface intersection corresponding to the rays of the general pinhole camera for related imaging. In their domain, and In other words, the depth map associated with the light field of a hierarchical scene decomposition layer is constrained by the depth boundaries of the layer itself.

[0211] The merge operation reassembles the layered scene decomposition layer set back into the inner and outer frustum volumes, or LF O and LF P . Using the merge operator* m Merge the inner frustum volume light field and the outer frustum volume light field. For example, when given two arbitrary light fields, LF1(x, y, u, v) and LF2(x, y, u, v), where i = argmin j∈{1,2} D m [LF j ](x,y,u,v),* m Defined as:

[0212] LF1(x, y, u, v)* m LF2(x, y, u, v) = LF i (x, y, u, v)

[0213] Therefore, LF O (x, y, u, v) and LF P (x, y, u, v) can be obtained from the set O by merging the light fields associated with the inner and outer frustum layers. LF and P LF For example:

[0214]

[0215]

[0216] The present disclosure provides a hierarchical scene decomposition operation and an inverse operation of merging data to reverse the decomposition. Performing hierarchical scene decomposition using K layers is understood to create a single light field that is K times larger. The value of hierarchical scene decomposition lies in the light fields induced by the layers; these light field layers are more suitable for downsampling than the original total light field or the inner frustum volume light field or the outer frustum volume light field, because the total amount of data required for multiple downsampled hierarchical scene decomposition light field layers with an appropriate sampling scheme is significantly less than the size of the original light field.

[0217] Those skilled in the art will appreciate that there are many types of sampling schemes that can successfully sample a light field. The sampling scheme S provided is not intended to limit or depart from the scope and spirit of the present invention, as other sampling schemes can be employed, such as specifying a separate sampling rate for each element image in a layered scene decomposition layer light field. Relatively simple sampling schemes can provide an efficient CODEC with greater sampling control; therefore, the present disclosure provides a simple sampling scheme to illustrate the present disclosure without limiting or departing from the scope and spirit of the present invention.

[0218] The light field sampling scheme provided in the present disclosure represents a light field encoding method. Given a display D = (M x , M y , N u , N v ,f,a,D LP ) and hierarchical scene decomposition L = (K1, K2, L o , L P ), the present disclosure provides a sampling scheme S associated with L as o or L P Any layer in i Associated M x ×M y Binary matrix M S [l i ] and each layer l i Mapped to a pair R(l i )=(n x , n y )’s mapping function R(l i ). M S [l i ](x m ,y m ) indicates the element image Is it included in the sampling plan? (1) is included, a(0) means Not included. R(l i )=(n x , n y ) represents the light field The element image in n x ×n y The resolution of the sampling.

[0219] The present disclosure also provides a hierarchical scene decomposition light field encoding process using the plenoptic sampling theory. The following description involves the inner frustum volume L of the hierarchical scene decomposition L. o , but the volume of the outer frustum is L P can be coded in a similar manner.

[0220] For each l i ∈L o , corresponding to the light field The depth map is limited to Based on the sampling scheme presented above, the present disclosure uses the following equation to create a sampling scheme S to guide Creation:

[0221]

[0222] In other words, ΔEI guides the M associated with each hierarchical scene decomposition layer. S The distance between "1" entries in the matrix. The following equation sets the image of a single element in the layer Resolution:

[0223]

[0224] This method uses ΔEI and N res The sampling scheme used to drive the sampling rates of the individual hierarchical scene decomposition layers can be considered a hierarchical plenoptic sampling theory sampling scheme (also referred to herein as a "plenoptic sampling scheme"). The plenoptic sampling scheme is based on a display that utilizes the plenoptic sampling theory identity function α(t)=t. This per-layer sampling scheme provides lossless compression for front-parallel planar scene objects, where objects within a layer do not occlude each other.

[0225] The assumption of only fronto-parallel planar scene objects is restrictive and does not represent typical scenes; intra-layer occlusions are inevitable, especially for larger hierarchical scene decomposition layers. To capture and encode the full range of potential scenes without introducing significant perceptible artifacts, the system can utilize information in addition to the disclosed light field plenoptic sampling scheme.

[0226] For example, surfaces are locally approximated by flat surfaces at various tilt angles. In "On the Bandwidth of Plenoptic Functions," Do et al. theorize temporal warping techniques to enable spectral characterization of surfaces in tilted light field displays. This work shows that the necessary reduction in downsampling and the need for accurate characterization of local bandwidth variations are driven by the degree of surface tilt, the depth of objects in the scene, and the positioning of objects at the edges of the FOV. Therefore, if signal distortion from deviations from front-facing parallel geometry is perceptually significant, the residual representation can adaptively send additional or supplementary elemental image data (dynamically changing the static sampling scheme) to compensate for the resulting loss.

[0227] Thus, the present disclosure provides for the identification of "core" or "residual" information used to encode and decode light fields by a CODEC. Given a light field display D and a corresponding hierarchical scene decomposition L with an associated sampling scheme S, the present disclosure considers the encoded, downsampled light field associated with L and S, as well as the number of hierarchical scene decomposition layers and the depth of the layers, as the "core" representation of the light field encoded and decoded by the CODEC. Any additional information that may be needed to be transmitted along with the core (encoded) representation of the light field during the decoding process is considered a "residual" representation of the light field to be processed by the CODEC and used together with the core representation of the light field to produce the final displayed light field.

[0228] Many of the layered scene decompositions and sampling schemes defined in the framework defined above may still exhibit holes due to occlusion after merging and reconstructing the original light field. Observe that object occlusion between objects in different layers does not lead to holes in the reconstruction. However, objects within the same layer that may occlude each other may cause holes, especially for certain sampling schemes.

[0229] Specifically, if the sampling within a particular layer is such that the gap between the sampled element images is large, then it is likely that occluded objects may be underestimated, resulting in holes. One solution to this is to simply sample the element images at a higher rate. However, higher sampling rates result in lower compression rates. Therefore, adding more element images results in the inclusion of a large amount of redundant information. What is needed is a more discriminative method that can include additional information that helps fill the holes, while not leading to redundancy in the overall representation. For example, consider a hierarchical scene decomposition:

[0230] L=(K1,K2,L O , L P )

[0231] For L o or L P Each layer l i , we can define a set of residual layers:

[0232] R(l i )={r(l i )(j)|1≤j≤K i}

[0233] where K i is a natural number, describing the layer l i The number of residual layers required. For each residual layer, such as a hierarchical scene decomposition layer, there is a light field associated with that layer:

[0234]

[0235] In the most general description, these additional layers can be free-form without further restrictions. In practice, additional information that can help deal with occlusions can be represented in these residual layers. One way to achieve this is to have the residual layer have the same sampling scheme as its parent hierarchical scene decomposition layer, but a possible variation might be to sample the residual layer with lower directional resolution to strictly control the compression rate of the LSD plus residual layer combination.

[0236] Specifically, the residual layer can be defined as an additional layer corresponding to the concept of Deep G-Buffers. Therefore:

[0237]

[0238] In this context, residual layers are contrasted with hierarchical scene decomposition layers in that the depth range of each layer is not determined by a predetermined depth partitioning of the hierarchical scene decomposition layer scheme, but is instead based on depth layer characteristics inherent to the geometry in the represented scene.

[0239] Layered Depth Image

[0240] Previously known approaches use view synthesis (DIBR) to recreate light fields from sampled elemental images with depth, where occlusion is a problem. For scenes with objects with sufficient depth complexity, it is possible that information about some scene objects may not be captured when subsampling the elemental images of the light field. Often, surfaces are not captured because they are hidden over a wide range of angles due to surface occlusion. In such cases, the synthesized view will show holes where the surfaces were not captured in the sampled view.

[0241] In the case of view synthesis in the context of hierarchical scene decomposition, problems due to occlusion only arise when the occluding objects or surface fragments are co-located within the same layer.

[0242] In previous work, computer graphics researchers have considered how to represent a scene with sampled views while still capturing occluded information that may be seen from unsampled viewpoints. In computer graphics, a geometry buffer (G-buffer) is the name for an image buffer that stores color, normal, and depth information rendered relative to a specific camera viewpoint. Mara et al. proposed the idea of ​​a deep G-buffer to render layered depth images in the context of global illumination calculations in computer graphics in order to capture information that would otherwise be lost. In this work, normal, color, and depth values ​​are also stored for each layered depth image. The proposed data structure can be used to provide additional geometry information to existing screen-space techniques (using standard G-buffers) in order to improve the quality of lighting calculations based on the use of the additional occlusion information provided by the deep G-buffer.

[0243] The number of layers or subsets is chosen to allow a richer representation of the occluded information, and is typically chosen to be small based on practical constraints and diminishing returns on the increased visual quality achieved. This work also introduces the idea of ​​enforcing a minimum separation distance constraint between layers or subsets, each representing a different part of the scene at a different position relative to a reference position, to avoid having layers representing trivial behind-occluded surfaces that do not contribute to the final image quality.

[0244] It is proposed that a deep geometry buffer may be rendered for each sampled element image, sub-portion in a hierarchical scene decomposition layer or subset. Thus, for each layer or subset, and for each element image or sub-portion in each subset, there will be a set of hierarchical depth image layers or sub-portions. The number of layers is based on some input parameter determined according to the geometry of at least a portion of an object represented within the scene, with the layer depth defined by a minimum separation distance parameter as another input.

[0245] Figure 14 A computer-implemented method is described, comprising:

[0246] receiving a first data set (420) comprising a three-dimensional description of a scene;

[0247] dividing the first data set into a plurality of layers, each layer representing a different portion of the scene at a different position relative to a reference position (429);

[0248] dividing the data corresponding to the at least one subset into a plurality of sub-portions, wherein positions of particular sub-portions are determined based on the geometry of at least a portion of an object represented within the scene (435); and

[0249] The plurality of layers and the plurality of sub-portions are encoded to generate a second data set, wherein a size of the second data set is smaller than a size of the first data set (424).

[0250] Layer-based compression analysis

[0251] Predictable compression rates are needed to create real-time rendering and transmission systems, as well as downsampling standards (which do not indicate achievable compression rates). The following is a compression analysis of the hierarchical scene decomposition coding strategy of the present disclosure.

[0252] As mentioned above, downsampling the light field based solely on the plenoptic sampling theory does not guarantee compression. The present disclosure provides a downsampled light field encoding strategy that allows for low-latency, real-time light field CODEC. In one embodiment, a complementary sampling scheme based on the plenoptic sampling theory is used, using ΔEI and N resTo drive the sampling rate of individual layered scene decomposition layers. Layered scene decomposition represents the entire 3D scene as multiple light fields, expanding the scene representation by a factor of the number of layers. The present disclosure further contemplates that when the layer depth is appropriately chosen, compression rates can be guaranteed when combined with downsampling based on plenoptic sampling theory.

[0253] For a given hierarchical scene decomposition layer l i The corresponding light field The restricted depth range of a layer provides a guaranteed compression rate for the layer's light field. The compression ratio achievable by downsampling a scene completely contained in a single layer can be explained by the following theorem:

[0254] Theorem 1

[0255] Consider an isotropic directional resolution N = N u =N v , hierarchical scene decomposition L and the associated sampling scheme S = (M s , R) of the display D=(M x , M y , N u , N v ,f,a,D LP ). Assume that the hierarchical scene decomposition layer l i The corresponding light field Make d min (l i )<Z DOF (D) and select The distance between "1" entries is set to ΔEI(d min (l i )) and R(l i )=N res (d max (l i )). Relative to the layered scene decomposition layer l i The compression ratio associated with S is

[0256] Proof 1

[0257] Consider a hierarchical scene decomposition layer within the maximum depth of field of the display, where and Where 0<c,d≤Z DOF .therefore, and and Therefore, ΔEI(d min (l i ))=N / c and N res (d max (l i))=N / d.

[0258] Based on this subsampling rate, the system needs to th element images, thus providing 1:(N / c) 2 The compression ratio of the element image is 1:d 2 Therefore, the total compression ratio is 1: (N / c) 2 *1∶d 2 =1∶N 2 (d / c) 2 . Compression factor term Determines the compression ratio.

[0259] There may be an alternative situation where d min (l i )=Z DOF and (d max (l i )) can be extended to any arbitrary depth. We know that ΔEI(Z DOF )=N and N res For all depths d ≥ Z DOF The maximum possible value of N is reached. Based on this subsampling rate, the system needs to th element images, thereby providing a 1:N 2 When representing a front-parallel plane object, the Z DOF Adding additional layers of hierarchical scene decomposition increases redundant representation capabilities. Therefore, when creating the core encoding representation, the maximum depth of field among the layers can be used to best decompose the entire scene.

[0260] Given an expression for the compression of downsampling a hierarchical scene decomposition layer, we can determine how the compression factor varies with the layer parameters. For a fixed width layer, or for some w, d max (l i )-d min (l i )=w, when d max (l i )-d min (l i ) is closest to the display surface, c f The term is minimized. Therefore, the hierarchical scene decomposition layers located closer to the display surface need narrower widths to achieve the same compression ratio as the layers located farther away from the display surface. This compression ratio analysis can be extended from the display surface to the depth Z DOF The space is divided into multiple adjacent facade plane layers.

[0261] Theorem 2

[0262] Consider an isotropic directional resolution N = N u =N v , hierarchical scene decomposition L and the associated sampling scheme S = (M s , R) of the display D=(M x , M y , N u , N v ,f,a,D LP ). Let S LF =M x M Y N u N v , represents the number of image pixels in the light field. The compression ratio of the hierarchical scene decomposition representation can be defined as:

[0263]

[0264] Proof 2

[0265] For a given hierarchical scene decomposition layer downsampled using compression ratio:

[0266]

[0267] To calculate the compression ratio, the size of each layer in its compressed form is calculated and summed, and the total compressed layer size is divided by the light field size. Consider a sum where the size of the compressed set of layers is:

[0268]

[0269] Therefore the compression ratio of the combined layer is:

[0270]

[0271] In a system with variable width of layered scene decomposition, d min (i) and d max (i) represents the front and back boundary depth of layer i respectively. The compression ratio of the layered scene decomposition representation is:

[0272]

[0273] Sum of constant hierarchical scene decomposition layers It decreases monotonically and tends to 1.

[0274] Therefore, hierarchical scene decomposition layers located closer to the display surface achieve lower compression rates than layers of the same width located further away from the display surface. To maximize efficiency, hierarchical scene decomposition layers with narrower widths are placed closer to the display surface, while wider hierarchical scene decomposition layers are placed further away from the display surface; this placement maintains a uniform compression rate across the entire scene.

[0275] The number and size of layers in a hierarchical scene decomposition

[0276] In order to determine the number of layers and the size of the layers required for the layered scene decomposition, a light field display with an identity function of α(t)=t is provided as an example. Consideration of this identity function is not intended to limit the scope or spirit of the present disclosure, as other functions may be utilized. It will be understood by those skilled in the art that although the display D=(M x , M y , N u , N v ,f,a,D LP ) is defined with a single identity function α, but each light field planar parameterized pinhole projector in the planar parameterized pinhole projector array may have a unique identity function α.

[0277] To losslessly represent the frontal planar surface (assuming no occlusion), the front boundary is located at depth Z DOF The single-layer scene decomposition layer represents the Z DOF to infinity. Lossless compression can be defined as a class of data compression algorithms that allows perfect reconstruction of the original data from the compressed data. For the purpose of generating the core representation, hierarchical scene decomposition layers beyond the deepest layer located at the maximum depth of field of the light field display are not considered, as these layers do not provide additional representational power from the perspective of the core representation; this applies to both the inner frustum volume layer set and the outer frustum volume layer set.

[0278] In the region from the display surface to the display's maximum depth of field (for both the inner and outer frustum volume layer sets), the hierarchical scene decomposition layers utilize maximum and minimum distance depths that are integer multiples of the light field display f-number. Hierarchical scene decomposition layers with narrower widths provide better per-layer compression, and thus better overall scene compression. However, the more layers in the decomposition, the greater the processing required for decoding, as more layers must be reconstructed and merged. The present disclosure therefore teaches layer distribution schemes with different layer depths. In one embodiment, hierarchical scene decomposition layers with narrower widths (and the correlation of the light fields represented by the layers) are located closer to the display surface, and the layer width (i.e., the depth difference between the front and back layer boundaries) increases exponentially with increasing distance from the display surface.

[0279] Layer-wise resolution sampling based on asymptotic resolution

[0280] There are two main issues in designing a suitable light field codec using hierarchical scene decomposition. The first is how to decompose the scene into subsets (layers). A natural follow-up design problem is how to downsample the light field associated with each layer. We refer to the downsampling strategy as a "sampling scheme."

[0281] There are many ways to construct a sampling scheme. The method proposed in this disclosure uses ΔEI and N res The full optical sampling scheme is an embodiment. res is an example of an effective resolution function. If there is no occlusion and all objects within the layer are frontal planes, then this scheme will theoretically result in lossless sampling.

[0282] Existing video codecs can be used effectively without being fully lossless; instead, where loss occurs, they are optimized to minimize the perceptual effect of any artifacts that exist. We therefore explore how to use potentially lossy downsampling while minimizing perceptual effects as a viable strategy for designing useful light field codecs.

[0283] The concept of depth of field of light field display proposed by Zwicker et al. res The foundation of the plenoptic sampling standard. This concept is strictly based on representation theory and the total representation capacity of the samples that make up a lightfield image. Total representation capacity is considered outside the context of the viewer viewing the lightfield display; in this sense, it is viewer-independent. In terms of lightfield displays and designing the best experience for viewers, the representation capacity of the lightfield is only relevant to the image perceived by the viewer and how their position relative to the display affects image perception.

[0284] The depth of field concept proposed by Zwicker states that a greater angular resolution of the light field will result in a larger 3D display volume within which the display will display objects at the full spatial resolution of the display. The theory predicts that the resolution of objects beyond the maximum depth of field distance from the display surface will decrease linearly ( Figure 20 ).

[0285] This resolution is viewer-independent. We believe it is overly conservative because it doesn't account for the viewer's proximity to the display. We consider a physical viewer at a certain distance from the display. From a sampling perspective, the angular sampling rate of the viewer's samples as they look at the display is a function of the viewer's distance. It can be shown that, in theory, based on the physical viewer's distance from the display, we can create an equation that estimates the observed resolution as a function of the object's distance from the viewer. We present a specific derivation of such an equation here, but other equations or models of resolution degradation could be used, as well as empirical models of degradation observed through experimental or simulation studies of light-field displays and viewer models.

[0286] The derived model shows that the asymptotic resolution of an object can be calculated. In general, the asymptotic resolution decreases as the viewer distance increases. Therefore, if we can assume a maximum viewer distance, it is reasonable to use the corresponding asymptotic resolution or other related resolution decay functions as the worst-case measure of depth resolution degradation.

[0287] Consider using N as described above res Plenoptic sampling scheme. Assume we have a minimum depth d min (442) and the maximum depth d max (443) layer. Then according to N res Formula d max (443) Determine the required directional resolution for this layer. According to the depth representation theory described in Zwicker, this sampling rate fully represents the scene at its maximum resolution within the range of the given layer. This is Figure 20 Indicated in.

[0288] Let us assume that it is reasonable to define a maximum viewing distance for a light field display in a given practical environment. Based on this maximum viewing distance, we can consider the associated falloff function as a function of the directional resolution chosen for the potential sampling scheme for a given layer.

[0289] like Figure 20 As shown, it is possible to draw a graph smaller than N res (d max )(443) implies a resolution function for the various directional sampling rates, and consider how it works in terms of the min (442) to d max (443) deviates from the ideal value (440) over a range of depths. For the asymptotic function, it can be observed that the deviation becomes larger with increasing depth, but of course does not exceed a maximum value based on the asymptotic value (441). The deviation represents a signal loss; however, it can be quantified and has a bound based on the asymptotic value. Therefore, this suggests a framework within which layer downsampling and any associated losses can be quantified, thereby guiding the design of appropriate sampling schemes for LSD-based light field codecs.

[0290] A three-dimensional description of a scene is divided into a plurality of subsets, and a plurality of layers or subsets are encoded to generate a second dataset that is smaller than the first dataset. Encoding the layers or subsets may include performing a sampling operation on the subsets. An effective resolution function is used to determine a suitable sampling rate. Elementary images associated with the subdivisions are then downsampled using the determined suitable sampling rate.

[0291] Related work has focused on analyzing depth of field and how resolution degrades with depth in displays with multiple light-attenuating layers. This work still analyzes depth of field as a viewer-independent concept similar to Zwicker et al., describing the depth of field of a light-field display as the range of depths within a virtual plane parallel to the display surface that can be reproduced at the display's maximum spatial resolution. However, this framework is viewer-independent and based on effectively orthogonal views of the scene. The information a particular viewer can access from the light field from a given viewpoint is resolved, in terms of the quality of objects at a certain depth in the scene.

[0292] Alpaslan et al. conducted a study to determine how perceived resolution in a lightfield display changes with distance based on variations in optical and spatial angular resolution parameters. The results showed that increasing angular resolution reduced the degradation of perceived resolution with depth in the display. This analysis was based on oscillation patterns measured in units of cycles per spatial period, where spatial refers to the space within the inner and outer frustums of the display, as shown in the equation below.

[0293] p=p o +s*tanφ

[0294] Where p is the minimum feature size, p o is the pixel size, s is the depth into the screen, and φ is the angular distance between the two samples. This formula is based on simple geometric parameters and the assumption that the display's directional light rays are uniformly distributed in angular space. It clearly shows that feature size increases with distance; however, this formulation is independent of the specific viewer and the feature size that the viewer can resolve at depth.

[0295] In contrast, Dodgson analyzed how viewers occupy various viewing zones in front of a 3D display that correspond to the projected density of the display's angular components, but did not directly relate these viewing zones to the apparent viewing quality of objects at depth.

[0296] One approach to dealing with small depth of field, or DoF, involves scaling the content to fit within the target area. This technique does appear to produce good results, but because it involves some optimization of the content, it doesn't seem immediately applicable to real-time datasets in interactive settings. A simpler, fixed-scheme rescaling technique can be applicable in real-time settings, but can introduce unacceptable distortion artifacts, such as cardboarding. Cardboarding can be defined as a common artifact that occurs when visualizing 3D content, the so-called "cardboarding" effect, where objects appear flat due to depth compression.

[0297] In the case where all surfaces are Lambertian, the viewer can be assumed to be a pinhole camera. A more accurate modeling of the real human eye as a finite aperture camera is an approach taken in other 3D display view simulation works. However, for simplicity, a pinhole camera is used, as it can in some sense serve as an upper bound on the quality of the finite aperture case. It is assumed that the canonical image forms the basis of the viewer's image. Therefore, in order to examine the quality of the image, it is necessary to consider the canonical rays. More specifically, it is assumed that the canonical image I c [D, O] can be related to the canonical image through a type of magnification operation. The canonical image is a sampled version of the deformed viewer image. Applying the inverse deformation function to the canonical image, a continuous version of it, will then give the viewer image. The deformation function can also be described as a projection function.

[0298] A 3D light field display can be formally defined based on various design parameters. Without loss of generality, assume that the display is centered at (x, y, z) = (0, 0, 0) in 3D space and points towards the viewer in the positive z direction, with y pointing upwards. The formal definition is as follows:

[0299] Consider the light field display D=(M x , M y , N u , N v ,f,D LP ), where (M x ,M y ) are the horizontal and vertical dimensions of the display spatial resolution, (N u ,N v ) are the horizontal and vertical dimensions of the angular resolution component of the display. Assume that the display is a set of idealized light field projectors with a spacing of D LP , with focal length f. Assume that the light field projector M x ×M y Arrays can be indexed as LFP ij , so that the first coordinate is aligned with the x-axis and the second coordinate is aligned with the y-axis. Therefore, a set of light field projectors is obtained: {LFP ij |1≤i≤M x ,1≤j≤M y For any particular light field projector, for 1≤u≤N u and 1≤u≤N v , can be achieved through LF P ij (u,v) for each individual N u ×N v Pixels are addressed. Based on the focal length f of the display, an angular field of view can be calculated, expressed as θ FOV .

[0300] As is well known, light field displays can represent objects within a volumetric region defined by two separate viewing frustums, including the area in front of and behind the display surface. These two frustums are referred to herein as the inner and outer frustum regions of a given display.

[0301] The viewer O = (X O ,D O ,f O ) is defined as a pinhole camera with a focal length of f O Image the display with the focus at X O And point to direction D O , where D O is a 3D vector. For viewer O, this is called the viewer image and is denoted as I O .

[0302] A particular viewer images a different subset of the possible output ray directions cast by the display's light field projector, depending on their specific position and direction / orientation. These rays can be more precisely defined as:

[0303] Given a display D = (M x , M y , N u , N v ,f,D LP ) and viewer O=(X O ,D O ,f O ). Define a set of rays, one for each light field projector associated with D, connected by X O and the line defined by the center of each light field projector. Let LFP ij Then put a set of lines is defined as the set of canonical rays for the viewer O with respect to the display D. It should be noted that the canonical rays form only a subset of the rays that contribute to the viewer's image. It is easy to observe that the canonical ray set for the viewer is related to the viewer's direction D O and focal length f O Therefore, the set of all possible viewers at a particular location shares the same set of canonical rays relative to the display.

[0304] For any standard ray Here is the corner Yes, it means The spherical coordinates of the associated vector. LF P ij N of (u,v) u ×N v Each of the elements also has a spherical coordinate representation, which can be written as and the space vector representation, expressed as

[0305] A canonical ray for a given display and viewer can be seen sampling intensity values ​​from the light field projected by the display and its light field projector. These intensity values ​​can be observed to form M x ×M y The image is further referred to herein as the canonical image relative to the display D and the viewer O. This image is denoted as I c [D, O](x, y).

[0306] Considering the display's field of view θ FOV , the viewer must be at a minimum distance to be able to see the light from all the light field projector-based pixels on the display. Typically, this distance will be larger for smaller FOVs, and the viewer may be closer for larger FOVs. This distance can be determined trigonometrically as:

[0307]

[0308] Each light field projector represents a continuous and smooth light field. Given a display and a viewer, each canonical ray samples the light field using the intensity within its corresponding light field projector array. Based on the LFP image contained in the light field projector ij The intensity values ​​are reconstructed by performing a resampling operation on the intensity values ​​in . The common model adopted here provides a spot width for each ray implied by the light field projector image. This spot width allows describing the physical reconstruction of the light field by providing a physical angular spread for each projector intensity value.

[0309] To simplify the analysis, the point spread function (PSF) model is ignored to some extent. Instead, it is assumed that the canonical rays are interpolated from a specific LEP using a nearest neighbor interpolation scheme. ij Sample the light field. For some (i, j), consider the light field from the intensity value I c [D, O](i, j) corresponds to the ray of the canonical image. We assume that the ray vector It can be expressed using spherical coordinates as (θ ij ,φ ij ).let

[0310]

[0311] Index (u n , v n ) represents the light field projection pixel with the minimum angular distance to the sampling canonical ray. Therefore, through this nearest neighbor interpolation:

[0312]

[0313] This reconstructed model allows for an initially simpler analysis and understanding of the sampling geometry.

[0314] Based on the display's conceptual depth of field, or DoF, a 3D display's ability to represent spatial resolution degrades when objects move beyond the maximum depth of field. Current lightfield or multi-view displays appear to suffer from small depths of field due to relatively poor angular resolution and sampling density. It's clear from these displays that objects at any significant depth into the screen appear significantly blurry.

[0315] Objects in a 2D display, while lacking the additional perceptual cues of a 3D display, do not become unnaturally blurry at a distance. In a standard 2D display, 3D objects that appear deep in the scene are actually far away from the display surface, degrading in a natural way relative to the display's maximum resolution. That is, as an object gets farther away from the 2D display, its projected area on the 2D display becomes smaller, and so the number of pixels representing that projected area decreases with the size of the area. This corresponds to how farther objects are projected onto a smaller area on the retina (or the imaging plane of the camera), so less detail can be resolved. However, in a 3D display with relatively low angular resolution, distant objects appear blurry and cannot be represented with a resolution proportional to their projected area on the display surface.

[0316] It is proposed to measure the effective spatial resolution of a 3D display at a certain depth based on a comparison with a pseudo-equivalent 2D display. A 3D display is used to simulate a 2D display in some way by considering how the 3D display presents right-parallel planes located in the area of ​​a frustum within the display. The size of the planes increases with depth to fill the entire width of the viewing frustum relative to a given viewer position. Let d p represents the z coordinate of the plane. This plane is referred to herein as plane P C (d p , O).

[0317] Consider the position (x O , z O ) viewer O at . Let the width of the display D be ω. Construct a plane so that for (x O , z O ), no matter what depth the plane is placed at, its size is such that the plane is projected onto every spatial pixel of the display. In other words, what is seen on the plane will occupy the entire space of the display surface, such as Figure 19 shown.

[0318] To calculate the depth d p The width W of the construction plane at , using a formula based on similar triangle geometry. It should be noted that the positive z-axis points from the display towards the viewer, making d pbecomes the negative value in the following formula:

[0319]

[0320] For simplicity, a 1D display has been used for the analysis. For the 1D analysis, the display will be defined as

[0321] D=(M x , N u ,f,D LP )

[0322] Use addressable LFP i (i) A light field projector is used for u within a defined range. The viewer is defined as O = (X O ), where X O are just x and z coordinates. We let LFP i The canonical ray set for viewer O relative to display D is

[0323]

[0324] The resulting canonical image is I C {D, O](x). Standard ray With angle representation θ i .LEP i N of (u) u Each of the elements has an angle representation θ(u) and a space vector representation, expressed as For nearest neighbor interpolation

[0325]

[0326] in:

[0327]

[0328] Effective resolution of the inner frustum

[0329] To answer the question of how the resolution for a given viewer degrades for scene elements with respect to their distance from the display surface, the analysis is restricted to the depth of a frustum located within the display. For simplicity, a one-dimensional display is assumed.

[0330] To quantify the effective resolution at a certain depth, the key question to be answered in this setup is: How do the light field projector rays that contribute to the reconstruction of the incident canonical rays affect the plane P? C (d p, O)? Sampling? Two sampling steps are simulated here: (1) the light field projector ray samples the plane, and (2) the incident canonical ray samples the light field from a subset of the light field projector rays. The problem can be simplified by assuming that the canonical ray sample is constructed using only one element of the light field projector via nearest neighbor interpolation.

[0331] Theorem 1

[0332] Assume that the display D = (M x , N u ,f,D LP ) and viewer O=(X O , D O , f O ). Let z O =ω=M x D LP Therefore, the effective resolution at depth d p It can be estimated as follows:

[0333]

[0334] prove:

[0335] Assume that the plane P(d p , O). The definition of distance is related to how the light of the light field projector is directed to the inner frustum and the plane P(d p , O) is related to sampling. Consider a set of M x Standard light Marked as c i Is with LFP i Intersecting rays. With ray c i The associated intensity is I C [D, O](i).

[0336] Based on the nearest neighbor scheme defined, each ray c i ∈C maps to LFP i The corresponding light in Indices, as defined before. There are two possible cases. In the first case, two adjacent canonical rays c i and c i+1 In its corresponding light field projector, there are nearest neighbors (LFP i , LFP i+1 ) have the same angle. That is, Another way to look at this problem is to map adjacent canonical rays to parallel light field projection rays. In the second possible case, two adjacent canonical rays c i and c i+1are mapped to different rays in their corresponding light field projectors. That is, for integer k ≥ 1, For N<M x and assumes that the viewer is at least d away from the display surface O , this situation will be

[0337] Now we define the distance based on these two cases. For the first case, two adjacent LFP rays are parallel and their distance is defined as D LP In the second case,

[0338] d=D LP +q

[0339]

[0340] This combination of parallel and diverging samples produces a non-uniform sampling pattern. It is suggested that the effective resolution of the display surface depth-viewer triplet is the plane P C (d p , O) divided by the maximum sampling distance. This is because the maximum sampling distance determines the minimum feature size that guarantees sampling.

[0341] That is, the resolution is P X A 2D display will have the same minimum feature size as the particular display-surface-depth-viewer triplet. For the inner frustum, P X will be:

[0342]

[0343] Given the estimation formula of the effective depth resolution, it can be seen that the formula gives a curve that is proportional to the variable d p If the plane is at a very large depth (ie d p →-∞), the asymptotic minimum effective resolution is:

[0344]

[0345] Effective resolution of the outer frustum

[0346] For the outer frustum, d p The value of will be positive based on the current coordinate system. p When q is positive, q becomes negative, so d<D LP .

[0347] It is recommended that the effective resolution of the display surface depth-viewer triplet is the surface size divided by the maximum sampling distance, which is now D LP .

[0348] Disparity encoding / decoding

[0349] The encoded hierarchical scene decomposition representation of the light field resulting from the sampling scheme applied to each layer primarily consists of multiple pixels comprising RGB color and disparity. In general, choosing an appropriate bit width for the disparity (depth) field of a pixel is important, as the width of this field improves the accuracy of operations during reconstruction. However, increasing the number of bits used has a negative impact on the achieved compression rate.

[0350] In the present disclosure, each layer of RGB color and disparity pixels, specified by a given sampling scheme, has a specific disparity range corresponding to each pixel. The present disclosure utilizes this narrow range of disparity within each hierarchical scene decomposition layer to increase the accuracy of depth information. In traditional pixel representations, the disparity range of the entire scene is mapped to a fixed number of values. For example, with 10-bit disparity encoding, only 1024 different depth values ​​are possible. In the hierarchical scene decomposition of the present disclosure, since each layer has known depth boundaries, the same fixed number of values ​​is applied to each hierarchical scene decomposition layer. This is advantageous because it reduces transmission bandwidth by reducing the width of the depth channel while maintaining pixel reconstruction accuracy. For example, when a system implements a disparity width of 8 bits, decomposing the scene into 8 hierarchical scene decomposition layers allows for a total of 2048 different disparity values, with each layer having 256 different possible values ​​based on an 8-bit representation. This is more efficient than mapping the entire range of possible disparity values ​​within the inner or outer frustum to a given number of bits.

[0351] The present disclosure uses the same number of bits, but the bits are interpreted and clearly represent the disparity within each hierarchical scene decomposition layer. Since each hierarchical scene decomposition layer is independent of each other, the depth (bit) encoding of each layer can be different and can be designed to provide a more accurate fixed-point representation. For example, hierarchical scene decomposition layers close to the display surface have smaller depth values ​​and can use a fixed-point format with a small number of integer bits and a large number of decimal places, while hierarchical scene decomposition layers far from the display surface have larger depth values ​​and can use a fixed-point format with a large number of integer bits and a small number of decimal places. The decimal places can be configured on a per-layer basis:

[0352] Minimum number of fixed points = 1 / (2 小数位 )

[0353] Maximum number of fixed points = 2 16-小数位 -Minimum fixed point number

[0354] Disparity is calculated from depth in the light field post-processing stage and encoded using the following formula:

[0355] Scaling factor = (maximum fixed point number - minimum fixed point number) / (near clipping parallax - far clipping parallax)

[0356] Encoded disparity = (disparity - far clipping disparity) * scaling factor + minimum fixed point number

[0357] The disparity is decoded using the following formula:

[0358] Scaling factor = (maximum fixed point number - minimum fixed point number) / (N near clipping parallax - far clipping parallax)

[0359] Uncoded disparity = (coded disparity - minimum fixed point number) / scaling factor + far clipping disparity

[0360] Figure 18 A computer-implemented method is described, comprising:

[0361] receiving a first data set (420) comprising a three-dimensional description of a scene;

[0362] dividing the first data set into a plurality of subsets, each subset representing a different portion of the scene at a different position relative to a reference position (429);

[0363] Encoding the plurality of subsets to generate a second data set, wherein the size of the second data set is smaller than the size of the first data set, wherein encoding the subsets includes performing a sampling operation on the subsets, performing the sampling operation includes (433):

[0364] An appropriate sampling rate is determined using the effective resolution function, and then the elemental images associated with the sub-portions are downsampled using the appropriate sampling rate (434).

[0365] Generalized and illustrative implementation schemes - CODEC implementation and applications

[0366] Overview

[0367] This disclosure defines encoder-decoders for various types of angular pixel parameterizations, such as, but not limited to, planar parameterization, arbitrary display parameterization, combinations of parameterizations, or any other configuration or parameterization type. Generalized and illustrative embodiments of this disclosure provide methods for generating synthetic light fields for multi-dimensional video streams, multi-dimensional interactive games, or other light field display scenarios. A rendering system and process are provided that can drive light field displays with real-time interactive content. Light field displays do not require long-term storage of light fields, but must render and transmit them with low latency to support interactive user experiences.

[0368] Figure 7An overview of a CODEC system is provided for a generalized, illustrative embodiment of the present invention. A game engine or interactive graphics computer (70) sends three-dimensional scene data to a GPU (71). The GPU encodes the data and sends it via a display port (72) to a decoding unit (73) containing a decoding processor (e.g., an FPGA or ASIC). The decoding unit (73) sends the decoded data to a light field display (74).

[0369] Figure 1 Another generalized, exemplary layered scene decomposition CODEC system is illustrated, wherein light field data from a synthetic or video data source (50) is input to an encoder (51). A GPU (43) encodes the inner frustum data, dividing it into multiple layers, and a GPU (53) encodes the outer frustum data, dividing it into another multiple layers. While Figure 1 Separate GPUs (43, 53) dedicated to the inner and outer frustum volume layers are illustrated, but a single GPU can be used to process both the inner and outer frustum volume layers. Each hierarchical scene decomposition layer is sent to a decoder (52), where the multiple inner frustum volume layers (44(1) to 44(*)) and multiple outer frustum volume layers (54(1) to 54(*)) of the light field are decoded and merged into a single inner frustum volume layer (45) and a single outer frustum volume layer (55). According to dual frustum rendering, the inner and outer frustum volumes are then synthesized (merged) into a single, reconstructed light field dataset (56), otherwise referred to herein as the "final light field" or "display light field."

[0370] Figures 10 to 13 An exemplary CODEC process implementation according to the present disclosure is illustrated.

[0371] Figure 10 An exemplary hierarchical scene decomposition CODEC method is illustrated, whereby 3D scene data or light field data in an image description format is loaded into an encoder (400) for encoding, whereupon (sub)sets of data as shown, or, alternatively, the entire data set representing the 3D scene are partitioned (403). Where a subset of 3D scene data is identified for partitioning (402), it will be appreciated that the identification process is a general process step reference intended to simply refer to the ability to partition a data set in a pass or in groups (e.g., encoding inner and outer frustum data layers, as in

[0048] ). Figure 11), as may be required depending on the circumstances. In this regard, the identification of data subsets may imply a pre-encoding processing step or a processing step that also forms part of the encoding sub-processing stage (401). The data subsets may be marked, designated, confirmed, scanned, or even compiled or grouped when partitioned to produce a set of layers (a decomposition of the 3D scene) (403). Following the partitioning of the data subsets (403), each data layer is sampled and rendered in accordance with the present disclosure to produce compressed (image) data (404). Following compression of the data layers, the compressed data is transmitted to a decoder (405) for a decoding sub-process (406), including decompression, decoding, and reassembly steps to (re)construct a set of light fields (407), otherwise referred to herein as "layered light fields", layered light field images, and light field layers. The constructed layered light fields are merged to produce a final light field (408) that displays the 3D scene (409).

[0372] Figure 13 An exemplary parallel codec process for optimizing the delivery of a light field representing a 3D scene in real time (e.g., minimizing artifacts) is shown in FIG. The process includes the following steps: loading 3D scene data into an encoder (700), encoding and compressing a residual coded representation of the final light field (701), transmitting the residual coded representation (702) to a decoder, decoding the residual coded representation and using the residual coded representation and a core coded representation to produce a final light field (703) and displaying the 3D scene on a display (704).

[0373] Figure 11 Shown with Figure 10 An embodiment related to the embodiment shown, wherein two data (sub)sets are identified; an inner frustum layer (502) and an outer frustum layer (503) derived based on 3D scene data (500) are used for partitioning (501), and each data set is divided into layers of different depths, i.e., equivalent to multiple data layers, according to two different layering schemes for each data set (504, 505). Each set (multiple) of data layers (506, 507) representing the inner frustum and outer frustum volumes of a light field display, respectively, is then sampled on a per-layer basis according to a sampling scheme (508, 509); and each sampled layer is rendered to compress the data and produce two sets of compressed (image) data (510, 511) in processing steps (508, 509), respectively. The compressed data sets (510, 511) encoding the light field sets corresponding to the data layer sets (506, 507) are then combined (512) to produce a layered, core-encoded representation (513) (CER) of the final (display) light field.

[0374] Figure 12An embodiment of a CODEC method or process for reconstructing a set of light fields and producing a final light field at a display is illustrated. The set of light fields (layered light fields) is (re)constructed from a core coded representation (513) using a multi-stage view synthesis protocol (600). Protocols (designated VS1-VS8) are applied (601-608) to each of the eight layers of the core coded representation (513), which may be different or the same depending on the characteristics of the light fields of each data layer to be decoded. Each protocol may apply a form of nonlinear interpolation, referred to herein as edge-adaptive interpolation (609), to provide good image resolution and clarity in the set of layered light fields (610) reconstructed from the core coded representation of the fields to ensure image clarity. The layered light fields (610) are merged, in this case illustrating the merging of two sets of light fields (611, 612) corresponding to two data subsets to produce two sets of merged light fields (613, 614). The merged set of light fields (613, 614) may represent, for example, the inner and outer frustum volumes of a final light field and may be merged accordingly (615) to produce said final light field (616) at a display.

[0375] CODEC encoder / encoding

[0376] The encoding according to the present disclosure is designed to support the generation of real-time interactive content (e.g., for games or simulated environments) as well as existing multidimensional datasets captured by a light-field general pinhole camera or camera array.

[0377] For a light field display D, a hierarchical scene decomposition L, and a sampling scheme S, the system encoder produces elemental images associated with the light field corresponding to each hierarchical scene decomposition layer included in the sampling scheme. Each elemental image corresponds to a generic pinhole camera. The elemental images are sampled at the resolution specified by the sampling scheme, and each elemental image includes a depth map.

[0378] Achieving rendering performance to drive real-time interactive content to multi-dimensional displays of significant high resolution and size presents significant challenges that are overcome by applying hybrid or combined rendering approaches to address the shortcomings of relying solely on any one technique as described herein.

[0379] Given an identity function α, a set of generic pinhole cameras specified by the encoding scheme of a given hierarchical scene decomposition layer can be systematically rendered using standard graphics viewport rendering. This rendering approach results in a large number of draw calls, especially for hierarchical scene decomposition layers with sampling schemes that include a large number of underlying element images. Therefore, this rendering approach alone does not provide real-time performance in systems that use hierarchical scene decomposition for realistic autostereoscopic light field displays.

[0380] Rendering techniques using standard graphics draw calls restrict the rendering of a generic pinhole camera plane parameterization (identity function α) to perspective transformations. Hardware-optimized rasterization functions provide the performance required for high-quality real-time rendering on traditional 2D displays. These accelerated hardware functions are based on the plane parameterization. Alternatively, parallel oblique projection can be used to render the generic pinhole camera plane parameterization using the standard rasterization graphics pipeline.

[0381] The present disclosure contemplates the application of rasterization to render a general pinhole camera view by converting groups of triangles into pixels on a display surface. When rendering a large number of views, every triangle must be rasterized in each view; oblique rendering reduces the number of rendering passes required for each hierarchical scene decomposition layer and can accommodate any arbitrary identity function α. ​​The system uses a parallel oblique projection for each angle specified by the identity function α. ​​Once the data is rendered, the system performs a "slice and dice" block transform (see U.S. Patent Nos. 6,549,308 and 7,436,537) to regroup the stored data from its angle-based groupings into elemental image groupings. When there are a large number of angles to render, a "slice and dice" approach alone is inefficient for real-time interactive content that requires many separate oblique rendering draw calls.

[0382] Ray tracing rendering systems can also accommodate arbitrary identity functions α. In ray tracing, specifying arbitrary angles does not require higher performance than accepting plane parameterization. However, for real-time interactive content that requires rendering systems using the latest accelerated GPUs, rasterization offers more reliable performance scalability than ray tracing rendering systems.

[0383] This disclosure provides several hybrid rendering methods for efficiently encoding light fields. In one embodiment, the encoding scheme renders hierarchical scene decomposition layers located closer to the display surface, with more images requiring fewer angular samples, while rendering layers farther from the display surface with fewer images and more angular samples. In a related embodiment, perspective rendering, oblique rendering, and ray tracing are combined to render the hierarchical scene decomposition layers; these rendering techniques can be implemented in a variety of interleaved rendering methods.

[0384] According to a generalized illustrative embodiment of the present disclosure, one or more light fields are encoded by a GPU rendering a two-dimensional pinhole camera array. A rendered representation is created by computing pixels from a sampling scheme applied to each hierarchical scene decomposition layer. A pixel shader executes the encoding algorithm. A typical GPU is optimized to generate a maximum of two to four pinhole camera views per scene in a single transmission frame. The present disclosure requires rendering hundreds or thousands of pinhole camera views simultaneously, thus employing multiple rendering techniques to render the data more efficiently.

[0385] In one optimization approach, a general-purpose pinhole camera located in a hierarchical scene decomposition layer farther from the display surface is rendered using standard graphics pipeline viewport operations, referred to as perspective rendering. A general-purpose pinhole camera located in a hierarchical scene decomposition layer closer to the display surface is rendered using a "slice and dice" block transform. Combining these approaches provides efficient rendering for a hierarchical plenoptic sampling scheme. The present disclosure provides hierarchical scene decomposition layers in which layers farther from the display surface contain a smaller number of higher-resolution element images, while layers closer to the display surface contain a larger number of lower-resolution element images. Using perspective rendering to render the smaller number of element images in layers farther from the display surface is efficient because this approach requires only one draw call per element image. However, at some point, perspective rendering becomes inefficient or inefficient for layers closer to the display surface because these layers contain more element images and require more draw calls. Since the element images in layers closer to the display surface correspond to a relatively small number of angles, oblique rendering can efficiently render these element images while reducing the number of draw calls. In one embodiment, a process is provided for determining where a system should render a hierarchical scene decomposition layer using perspective rendering, oblique rendering, or ray tracing. A threshold algorithm is applied to evaluate each hierarchical scene decomposition layer to compare the number of element images to be rendered (i.e., the number of perspective rendering draw calls) with the size of the element images required for a particular layer depth (i.e., the number of oblique rendering draw calls), and the system implements the rendering method (technique) that requires the fewest number of rendering draw calls.

[0386] In situations where standard graphics calls cannot be used, the system can implement ray tracing instead of perspective or oblique rendering. Thus, in another embodiment, an alternative rendering method uses ray tracing to render layers that are closer to the display surface, or portions of layers that are closer to the display surface.

[0387] In a ray-traced rendering system, each pixel in a layer of a hierarchical scene decomposition is associated with a ray defined by a light field. Each ray is cast and its intersection with the hierarchical scene decomposition is calculated according to standard ray tracing methods. Ray tracing is advantageous when rendering does not conform to the standard planar parameterization expected by standard GPU rendering pipelines, as it can accommodate arbitrary light angles that are challenging for traditional GPU rendering.

[0388] When a hogel projects pixels into space, not every pixel is useful. Consider a display with a top-left hogel projecting a pixel upward and to the left. The only time a viewer can see this pixel is when the viewer is in a position such that the top-left hogel is at the bottom-right boundary of the viewer's field of view. From this position, all other hogels in the display will be viewed from a wider angle than the field of view allows, so all but the top-left hogel will be turned off. This given viewer position is not a useful viewing position; it would be irrelevant if the top-left pixel of the top-left hogel were turned off. This discussion uses the concept of an effective viewport. The effective viewport is the set of all positions in space from which a viewer can view every hogel on the display at some angle within the field of view and therefore receive a pixel from each hogel. This region is where the projected frustums of each hogel intersect.

[0389] The definition of the effective viewport is effectively reduced to where the projection frustums of the four corner hogels intersect. Corners are the most extreme cases, so if a location is within the projection frustums of the four corners, it is also within the effective viewport. This approach also introduces the concept of a maximum viewing distance, a constraint introduced to achieve these savings and efficiencies. Without a maximum viewing distance, the view frustum is a rectangular pyramid with its tip oriented along the negative display normal and its base at an infinite depth from the display (i.e., the standard view frustum). With the introduction of a maximum viewing distance, the base of the rectangular pyramid now has a base at the same distance as the maximum viewing distance. The savings are achieved by not rendering or sending pixels that would not be projected into the effective viewport and would therefore be wasted. The number of pixels required to specify the maximum viewing distance is the hogel fill factor. The hogel fill factor is the ratio between the viewport and the size of the hogel projection at a given depth (i.e., in 2D, if the hogel projection is 1 meter wide and the viewport is 0.5 meters wide, only half as many projected pixels are needed).

[0390] DW represents the display width in meters, MVD is the minimum viewing distance in meters, and FOV is the field of view in degrees. The maximum viewing distance is defined as MVD + y, where y represents the size of the usable range in meters. From a similar geometric shape, angle b is equal to angle a, where angle b is equal to the field of view in degrees. The width of the viewing area, marked as c, is defined by the following equation:

[0391]

[0392] The width of the Hogel projection is defined by the following equation:

[0393]

[0394] The Hogel filling factor in two dimensions is the ratio between c and e, so:

[0395]

[0396] This simplifies to:

[0397]

[0398] If this is applied in 3D, then the Hogel fill factor is applied along (x,y). Therefore, the Hogel fill factor is defined as:

[0399]

[0400] The result of increasing or decreasing the Hogel fill factor is an increase or decrease in the maximum observation depth, respectively.

[0401] Correcting ray traced pixels in sample mode

[0402] The strategy for producing a corrected light field is to rasterize the light field and then apply a per-pixel deformation operation. Where the pixel should go is determined by a characterization routine that involves imaging the display. How the light field is deformed depends on an equation whose form does not change, but the coefficients do. These coefficients are unique to each display. The idea behind the correction (but not literally how it works) is that if a pixel should be at position X, but is measured at X+0.1, the pixel will be deformed to position X-0.1 in anticipation of it being measured at X. The goal is to make the measured position match the expected position.

[0403] This strategy of generating on a uniform grid and then warping to the correct grid can be replaced by using ray tracing to immediately sample the correct grid. Rasterization is a uniform grid operation, while ray tracing is a generalized sampling. This will also help maintain light field integrity. Consider a white pixel in a sea of ​​black. The correction of the first lens system requires horizontal and vertical offsets of +0.5. The result is four gray pixels in a 2x2 grid around 0.5, 0.5. The display lens requires horizontal and vertical offsets of -0.5. The result is a 3x3 grid of illuminated pixels with a bright gray pixel in the center, four dark gray pixels on the four sides, and four dark gray pixels in the four corners. This energy dispersion would not occur if the pixels were sampled correctly in the first place. It seems unlikely that ray tracing would be faster than rasterization, but if the correction is removed and only half of the light field is captured based on the calculated Hogel fill factor, the entire pipeline could be faster.

[0404] Screen-space ray tracing

[0405] An alternative to warping methods for view synthesis is screen-space ray tracing. McGuire et al. proposed applying screen-space ray tracing to multiple depth layers (for robustness). These depth layers are generated by depth peeling. However, depth peeling algorithms are slow, so when using modern GPUs, single-pass methods are preferred, such as those by Mara et al., based on reverse reprojection, multiple viewports, and multiple rasterizers.

[0406] It is possible to combine screen-space ray tracing with hierarchical scene decomposition. A single ray is traced based on a known view. The result is an image that indicates the color of each pixel. For the hierarchical scene decomposition codec process, an encoded form of the light field is created and represented as layers with missing pixels. Screen-space ray tracing can be used to reconstruct these pixels from the pixels present in the encoded representation. For example, the representation can be an element image in the form of a deep G-buffer or a hierarchical element image. McGuire et al. describe a technique for doing this using an acceleration data structure for a hierarchical depth image type representation. This is in contrast to methods that trace rays at the polygon or object level representation, which are also effectively used with data structures that accelerate ray intersections.

[0407] Many real-time rendering techniques operate in screen space for computational efficiency, including techniques for approximating realistic lighting, such as, but not limited to, screen-space ambient occlusion, soft shadows, and camera effects like depth of field. These screen-space techniques are approximations of algorithms that traditionally work by ray tracing 3D geometry. Many of these algorithms use screen-space ray tracing, or more precisely, ray marching as described by Sousa et al. Ray marching is desirable because it eliminates the need to construct additional data structures. Classic ray marching methods, such as the Digital Differential Analyzer (DDA), are prone to oversampling and undersampling unless perspective is taken into account. Most screen-space ray tracing methods use only a single depth layer. Combining this technique with hierarchical scene decomposition allows algorithms to work on a subset of the scene rather than through multiple depth layers, potentially reducing ray hit distances due to the optimized partitioning of the scene into subsets.

[0408] Those skilled in the art will appreciate that there are a variety of rendering methods and combinations of rendering methods that can successfully encode layered scene decomposition element images. Other rendering methods may provide efficiency in different contexts, depending on the underlying computing architecture of the system, the sampling scheme used, and the identity function α of the light field display.

[0409] CODEC decoder / decoding

[0410] The decoding according to the present disclosure is designed to exploit the encoding strategy (sampling and rendering). The core representation is decoded as a set of layered light fields from a downsampled layered scene decomposition to reconstruct the light field LFO and LF P Consider the display D = (M x , M y , N u , N v ,f,α,D LP ) has a hierarchical scene decomposition L = (K1, K2, L O , L P ) and the related sampling scheme S=(M s , R). By downsampling the deconstructed LF from the sampling scheme S O and LF P Light field to reconstruct light field LF O and LF P Decode the element image. Pixels are aligned so that the inner and outer frustum volume layers closer to the display surface are checked first, moving to the inner and outer frustum volume layers further from the display surface until a non-empty pixel is found, and data from the non-empty pixel is transferred to the empty pixel closer to the display surface. In alternative embodiments, a particular implementation may restrict viewing to the inner or outer frustum volume of the light field display, requiring the LF O or LF P One of them is decoded.

[0411] In one embodiment, the decoding process is represented by the following pseudo code:

[0412] Core layered decoding:

[0413] For each l i ∈L o :

[0414]

[0415] or (Before and after comparison)

[0416] Similar procedures to rebuild LF P Each hierarchical scene decomposition layer is reconstructed from a finite number of samples defined by a given sampling scheme S. Each inner frustum volume layer or outer frustum volume layer is merged to reproduce the LF o or LF P .

[0417] ReconLF can be implemented in various forms with different computational and post-CODEC image quality properties. ReconLF can be defined as a function such that, given a light field associated with a layer sampled according to a given sampling scheme S, and the corresponding depth map of the light field, it reconstructs the full light field that has been sampled. The ReconLF input is a function consisting of a given sampling scheme S and the corresponding downsampled depth map. Definition A subset of the data. Depth Image-Based Rendering (DIBR) described by Graziosi et al. can reconstruct the input light field. DIBR can be categorized as a projective rendering method. In contrast to reprojection techniques, ray casting methods, such as screen-space ray casting as taught by Widmer et al., can reconstruct light fields. Ray casting offers greater flexibility than reprojection but increases computational resource requirements.

[0418] In the DIBR method, the element images specified in the sampling scheme S are used as reference "views" to synthesize the missing element images from the light field. As described by Vincent Jantet in "Layered Depth Images for Multi-View Coding" and Graziosi et al., when the system uses DIBR reconstruction, the process usually includes forward warping, merging, and backprojection.

[0419] The application of backprojection techniques avoids the generation of seams and sampling artifacts in synthesized views, such as element images. Backprojection assumes that a depth map or disparity map of the element image is synthesized together with the necessary reference image required to reconstruct the target image; this synthesis typically occurs through a forward warping process. Using the disparity value for each pixel in the target image, the system warps the pixel to the corresponding position in the reference image; typically, this reference image position is not aligned on an integer pixel grid, so values ​​from neighboring pixel values ​​must be interpolated. Implementations of backprojection known in the art use simple linear interpolation. However, linear interpolation can be problematic. If the warped reference image position is on or near an object edge boundary, the interpolation can exhibit noticeable artifacts because information from the edge boundary is included in the interpolation operation. The resulting synthesized image has "dirty" or blurred edges.

[0420] The present disclosure provides a back-projection technique for the interpolation sub-step, producing a high-quality composite image without dirty or blurred edges. The present disclosure introduces edge-adaptive interpolation (EAI), in which the system incorporates depth map information to identify the pixels required for the interpolation operation to calculate the color of the deformed pixels in the reference image. EAI is a nonlinear interpolation procedure that adapts to and preserves edges during the low-pass filtering operation. Consider a display D = (M x , M y , N u , Nv ,f,α,D LP ) with target image I t (x, y), reference image I r (x, y) and depth map D m (It) and D m (Ir). The present disclosure utilizes the depth map D m (I t ) Pinhole camera parameters (f, α, etc.) and the relative position of the display plane parameterized pinhole projector array will be each I t Pixel integer (x, y) is transformed into I r The real number position in (x w ,y w ). In (x w ,y w ) is not located at an integer coordinate position, it must be based on I r Integer sample reconstruction value.

[0421] Linear interpolation methods known in the art reconstruct I from the four nearest integer coordinates located in a 2×2 pixel neighborhood r (x w ,y w ). Alternative reconstruction methods use larger neighborhoods (e.g., 3×3 pixel neighborhoods) that produce similar results with different reconstruction quality (see Marschner et al., “An evaluation of reconstruction filters for volume rendering”). These linear interpolation methods are unaware of the underlying geometric structure of the signal. Smudged or blurred edge images can occur when reconstruction utilizes pixel neighbors belonging to different objects that are separated by edges in the image. Wrong inclusion of colors from other objects can produce ghosting. The present disclosure provides a method for reconstructing the image by using a depth map D m (I r ) to predict the existence of edges created when multiple objects overlap to solve the reconstruction problem by weighting or omitting pixel neighbors.

[0422] Figure 3AA texture (80, 83) is illustrated in which one sampling location, shown as a black dot (86), is back-projected into another image being reconstructed. The sampling location (86) is located near the boundary of a dark object (87) with a white background (88). In the first reconstruction matrix (81), the complete 2x2 pixel neighborhood, each single white pixel represented by a square (89), is reconstructed using known techniques such as linear interpolation to reconstruct the sampling location (86) value. This results in a non-white pixel (82) because the dark object (87) is included in the reconstruction. The second reconstruction matrix (84) uses the EAI technique of the present disclosure to reconstruct the sampling location (86) from three adjacent single white pixels (90). EAI detects object edges and ignores the dark pixels (87), thereby generating a correct white pixel reconstruction (85).

[0423] For the target image I t Any fixed coordinate (x, y) in (x, y) r ,y r ), d t The position depth is defined:

[0424] d t =D m [I r (x r ,y r )]

[0425] Target image coordinates (x r ,y r ) is transformed to the reference image coordinates (x w ,y w ).

[0426] For the close (x w ,y w ) of the point m-sized neighborhood, set N S ={(x i ,y i )|1≤i≤m}. The weight of each neighborhood is defined as:

[0427] w i =f(d t , D m [I r ](x i ,y i )]

[0428] where w i is the depth (x r ,y r ) and (x w ,y w ) is a function of the depth of the neighborhood of . The following equation represents a given threshold t eThe effective w i :

[0429]

[0430] Threshold t e is the feature size parameter. The weight function determines how to reconstruct I r (x r ,y r ):

[0431] I r (x r ,y r )=Recon(w1I r (x1, y1), (w2I r (x2, y2),...(w m I r (x m ,y m ))

[0432] The Recon function can be a simple modified linear interpolation, where w i The weights are combined using a standard weighting procedure and renormalized to keep the total weight equal to 1.

[0433] The present disclosure also provides a performance-optimized decoding method for reconstructing a hierarchical scene decomposition. Consider a display D = (M x , M y , N u , N v ,f,a,D LP ) has a hierarchical scene decomposition L = (K1, K2, L O , L P ) and the associated sampling scheme S=(M s , R). By downsampling the deconstructed LF from the sampling scheme S O and LF P Light Field Reconstruction LF O and LF P As mentioned above, a particular implementation may restrict the viewing of the inner or outer frustum volume of the light field display, thus requiring the decoding of the LF O or LF P one.

[0434] LF OReconstruction can be achieved by decoding the elemental images specified by the sampling scheme S. The ReconLF method for a specific layer does not include inherent constraints on the order in which missing pixels are to be reconstructed from the missing elemental images. The present disclosure aims to reconstruct the missing pixels using a method that maximizes throughput; sufficiently large light fields for effective light field display require significant data throughput to deliver content at interactive frame rates, so improvements in reconstruction data transmission are needed.

[0435] Figure 3B The diagram shows the general process flow for reconstructing a pixel array. Reconstruction begins (30), and the sampling scheme is then implemented (31). Pixels are then synthesized by column (32) and also by row (33) in the array, which can be done in either order. Once all pixels have been synthesized by column and row, pixel reconstruction is complete (34).

[0436] This disclosure introduces a set of basic constraints to improve pixel reconstruction and improve data transmission of content at interactive frame rates. x ×M y A single light field L of the element image i ∈L o , serves as input to ReconLF. The pixels (in other words, the element image) are reconstructed in two basic passes. Each pass operates on a different dimension of the element image array; the system performs the first pass as column decoding and the second pass as row decoding to reconstruct each pixel. Although this disclosure describes a system that employs column decoding followed by row decoding, this is not meant to limit the scope and spirit of the invention, as a system that employs row decoding followed by column decoding may also be utilized.

[0437] In the first pass, the element image specified by the sampling scheme S is used as reference pixels to fill in the missing pixels. Figure 4 The element images in the matrix are displayed as B, or blue pixels (60). The missing pixels (61) are synthesized strictly from reference pixels in the same column. Figure 5 Schematically illustrates the column-wise reconstruction of a pixel matrix as part of an image (pixel) reconstruction process showing column-wise reconstruction (63) of red pixels (62) and blue pixels (60). These newly synthesized column-wise pixels are Figure 5 The blue pixel (60) and the missing pixel (61) are shown next to an R or red pixel (62). The newly reconstructed pixel is written to the buffer and serves as a further pixel reference for the second pass, which reconstructs the pixel reference pixels that are in the same row as the other element images. Figure 6 The subsequent row-wise reconstruction (64) of the pixel matrix is ​​illustrated as part of the image (pixel) reconstruction process along with the column-wise reconstruction (63). These newly synthesized row-wise pixels are shown as G or green pixels (65) next to the blue pixels (60) and red pixels (62).

[0438] In one embodiment, the process for reconstructing the pixel array is represented by the following pseudo-code algorithm:

[0439] Dimensionality decomposition light field reconstruction:

[0440] First pass:

[0441] For L i Each row element image in

[0442] For each missing element in the row image

[0443] For each row in the element image

[0444] Load (cache) pixels from the same row in the reference image

[0445] For each pixel in the missing row

[0446] Reconstruct pixels from reference information and write

[0447] Second pass:

[0448] For L i Each column element image in

[0449] For each missing element in the column image

[0450] For each column in the element image

[0451] Load (cache) reference pixels from the same column

[0452] For each pixel in the missing column

[0453] Reconstruct pixels from reference information and write

[0454] This performance-optimized decoding approach allows row-decoding and column-decoding constraints to limit the effective working dataset required for the reconstruction operation.

[0455] To reconstruct a missing row of an element image, the system only needs the corresponding row of pixels from the reference element image. Similarly, to reconstruct a missing column of an element image, the system only needs the corresponding column of pixels in the reference element image. This method requires a smaller data set, as decoding methods previously known in the art require the entire element image for decoding.

[0456] Even when decoding relatively large element image sizes, a reduced data set may be stored in a buffer while reconstructing the rows and columns of missing element images, thereby providing improved data transfer.

[0457] Once all rendering data has been decoded and each of the multiple inner and outer display volume layers has been reconstructed, the layers are merged into a single inner display volume layer and a single outer display volume layer. The layered scene decomposition layers can be partially decompressed in a staged decompression manner or fully decompressed simultaneously. Algorithmically, the layered scene decomposition layers can be decompressed in a front-to-back or back-to-front process. A final dual-frustum merging process combines the inner and outer display volume layers to create the final light field for the light field display.

[0458] Using computational neural networks with hierarchical scene decomposition

[0459] Martin discussed deep learning in light fields, presenting convolutional neural networks (CNNs). Deep learning research for other light field problems is also ongoing, as it is becoming increasingly popular to train the network end-to-end, meaning that the network learns all aspects of the problem at hand. For example, in view synthesis, this avoids using computer vision techniques such as appearance flow, image inpainting, and depth image-based rendering to model parts of the network. Martin proposed a conceptual framework for a view synthesis pipeline for light field volume rendering. This pipeline can be implemented using the hierarchical scene decomposition method disclosed here.

[0460] Hierarchical scene decomposition enables the decomposition of a multi-dimensional scene into layers, or subsets, and element images, or sub-parts. Machine learning is emerging as a learning-based view synthesis approach. Hierarchical scene decomposition provides a method for downsampling the light field after the decomposition into layers. Previously, this was considered in the context of rendering opaque surfaces using Lambertian shading surfaces. What is needed is a method for downsampling the light field as described previously, but that can be applied to higher order lighting models, including semi-transparent surfaces, such as those based on direct volume rendering. Volume rendering techniques include, but are not limited to, direct volume rendering (DVR), texture-based volume rendering, volumetric lighting, two-pass volume rendering with shadows, or procedural rendering.

[0461] Direct volume rendering is a rendering process that maps a volumetric dataset (e.g., voxel sampling of a scalar field) to a rendered image without intermediate geometry (no isosurfaces). Typically, the scalar field defined by the data is considered a semi-transparent, emissive medium. A transfer function specifies how to map the field to opacity and color, and a raycasting program then accumulates local color and opacity along the path from the camera and through the volume.

[0462] Levoy (1988) first proposed direct volume rendering methods for generating images of 3D volume datasets without explicitly extracting geometric surfaces from the data. Kniss et al. proposed that, although the dataset is interpreted as a continuous function in space, for practical purposes it is represented by a uniform 3D array of samples. In graphics memory, volume data is stored as a stack of 2D texture slices or a single 3D texture object. The term voxel refers to an individual "volume element," analogous to the terms pixel for "picture element" and texel for "texture element." Each voxel corresponds to a location in data space and has one or more data values ​​associated with it. Values ​​at intermediate locations are obtained by interpolating the data from adjacent volume elements. This process, called reconstruction, plays an important role in volume rendering and processing applications.

[0463] The purpose of an optical model is to describe how light interacts with particles within a volume. More complex models account for light scattering effects by taking into account both local illumination and volumetric shadowing. Optical parameters are specified directly from data values, or they are computed by applying one or more transfer functions to the data to classify specific features in the data.

[0464] Martin implements volume rendering using a volume dataset and provides a depth buffer to assign depth values ​​to each individual pixel location. The depth buffer or z-buffer is converted to pixel disparity, and the depth buffer value Z b Z b is converted to normalized coordinates in the range [-1,1], such as Z c Z c =2·Z b Z b -1. The perspective projection is then inverted to give the depth Z of the eye space e ,as follows:

[0465]

[0466] where Z n is the depth of the camera's near plane, Z f is the depth of the far plane in eye space. Wanner et al. proposed Z n Should be set as high as possible to increase depth buffer accuracy, while Z f The effect on accuracy is small. Given the eye depth Z e , which can be converted to a disparity value dr in real units by using a similar triangle:

[0467]

[0468] Where B is the distance between two adjacent cameras in the grid, f is the camera focal length, and Δx is the distance between the principal points of two adjacent cameras. Using similar triangles, the parallax in real units can be converted to the parallax in pixels:

[0469]

[0470] Where dp and dr represent the difference between pixels and real-world units, respectively, Wp is the image width in pixels, and Wr is the image sensor width in real-world units. If the image sensor width in real-world units is unknown, Wr can be calculated from the camera field of view θ and focal length f as:

[0471]

[0472] View synthesis can also be formulated through warping. While warping is a simple way to synthesize new views, it can produce visual artifacts that degrade the visual quality of the warped image. The most common of these artifacts are occlusions, cracks, and ghosting.

[0473] Deocclusion artifacts, or "occlusion holes," occur when a foreground object deforms and the reference view does not contain data for the background pixels that are now in view. Occlusion holes can be fixed by patching the holes with available background information or filling them with actual data captured with additional reference or residual information.

[0474] Deformation cracks occur when a surface is deformed so that two pixels that were adjacent in the reference view are deformed to the new view and are now adjacent for a longer time but separated by a smaller number of pixels. Rounding errors can cause deformation cracks because the newly calculated pixel coordinates must be truncated to integer image coordinates, which can cause adjacent pixels to round differently. Sampling frequency can cause deformation cracks by attempting to deform the surface into an orientation that increases its number of pixels (i.e., a plane tilted with respect to the camera then viewed perpendicularly). The new view will want to display pixels that exceed the sampling frequency of the reference camera, causing cracks in the new image.

[0475] Ghosting can occur during the backwarping interpolation phase. Here, the backprojected pixel neighborhood contains pixels from both background and foreground objects. Pixels in the foreground can bleed color information into the background, causing "halos" or ghosting effects. These typically occur around occlusion holes, where the foreground color bleeds into the background.

[0476] One of the main issues with forward warping is that the warped image may contain deformation cracks, which degrade visual quality. Depth maps generated using forward warping can be easily repaired by merging multiple views warped from different references or applying a crack filter. Filters such as the median filter can effectively remove cracks because the depth map is a very low-frequency image, primarily consisting of subtle gradients or edges around objects. Due to the complexity of object textures, these simple filters are not suitable for color images. One method to eliminate deformation cracks in color images is to use backward warping. In backward warping, the depth image is first forward warped to obtain a depth map for the new view. After filtering the cracked depth map, the filtered depth map is used to warp back to the reference image. To prevent cracks, pixel coordinates are not rounded. Instead, a pixel neighborhood is selected and the correct colors are interpolated using actual pixel weighting. This backward warped color image is now free of cracks. A side effect of the interpolation stage is that ghosting artifacts may be introduced.

[0477] Interactive direct volume rendering is needed to interactively view 4D volumetric data that varies over time, as progressive rendering may not be suitable for the specific use cases proposed by Martin. Example use cases for interactive direct volume rendering include, but are not limited to, rendering static voxel-based data without artifacts during rotation, rendering time-varying voxel-based data such as 4D MRI or ultrasound, CFD, wave, meteorology, visual effects (OpenVDB), and other physics simulations, etc.

[0478] One proposed solution involves using machine learning to learn how to "deform" a volumetric scene view, perhaps constrained to a specific transfer function that conveys how to map the density of different materials to color and then its transparency level. For a fixed transfer function, computational neural networks can be trained very well using moderately sized datasets, allowing for the definition of a decoder that works only with volumetric data and only with a specific transfer function. The potential result is a hardware system or decoding system that, when a desired transfer function is chosen, will slightly change the decoder based on a different training dataset, thus being able to decode the data it has been given.

[0479] The proposed method disclosed herein is applicable to rendering 4D volumetric data. Using current hardware and hardware technology, robust light field rendering of volumetric data is difficult. The proposed method generates a hierarchical scene decomposition of the volumetric data, which is then rendered and decoded using a decoder. The decoded data effectively fills in missing pixels or image elements. Convolutional neural networks can be trained to solve a smaller version of the problem using a system that employs column-by-row decoding, and alternatively, a system that employs row-by-column decoding can be utilized. Martin teaches that to perform fast yet accurate image warping using disparity maps, a form of backward warping with bilinear interpolation is implemented. The estimated disparity map of the central view is used as an estimate for all views. Pixels in the new view that should read data from locations outside the boundaries of the reference view are set to read the nearest boundary pixel in the reference view. Essentially, this stretches the boundaries of the reference view in the new view, rather than creating holes. Since warped pixels rarely fall at integer positions, bilinear interpolation is applied to accumulate information from the four nearest pixels in the reference view. This results in fast warping without holes and good accuracy. Martin further discloses a method for training a neural network to apply this correction function. An improvement on this is to teach a neural network compatible with hierarchical scene decomposition. This can be further extended to apply a convolutional neural network for each layer in the hierarchical scene decomposition, which is trained specifically for that layer, at each depth. One neural network would then be set up to perform row reconstruction, while the other would be set up to perform column reconstruction.

[0480] A light field display simulator can be used to train neural networks based on selected criteria. The light field display simulator provides a high-performance method for exploring parameterizations of simulated virtual 3D light field displays. The method uses a canonical image generation method as part of its computational process to simulate a virtual viewer's view of the simulated light field display. The canonical image method provides a robust, fast, and general method for generating simulated light field displays and the light field content displayed on them.

[0481] Advanced lighting models

[0482] In computer graphics, the color of opaque dielectrics can be modeled using Lambertian reflectance; the color is considered constant with respect to viewing angle. This correlates with the standard color measurement methods used in industry, which are based on the same Lambertian (or near-Lambertian) reflectance.

[0483] The fundamental concept of layered scene decomposition is the ability to partition a scene into sets and subsets, and then reconstruct these partitions to form a light field. This concept is based on the ability to warp pixels to reconstruct the image of missing elements in a layer through deformation. More specifically, the light intensity at a specific point in one image is mapped to a slightly different pixel location in another image based on the geometric shift created when the image moves from left to right in the camera. The deformation method described here can be used to accurately reconstruct missing pixels within a layer under the assumption that the pixels mapped from one image to the next have the same color, as is the case when using the Lambertian illumination model.

[0484] Light fields, especially when constrained to assume a Lambertian illumination model, exhibit significant redundancy. This can be observed as each element image differing only slightly from adjacent images. This redundancy is described in the literature on plenoptic sampling theory. The Lambertian illumination model is sufficient for useful graphics, but not overly realistic. To capture the gloss, haze, and angular flop color of an object, alternative models have been investigated, including but not limited to the specular index of the Phong model, the surface roughness of the Ward model, and the surface roughness of the Cook-Torrance model. Gloss can be defined as a measure of the specular amplitude, and haze can be defined as a parameter that captures the width of the specular lobe.

[0485] In order to exploit the alternating illumination model in the view synthesis aspect of our disclosed layered scene decomposition method, it is proposed that shading can be applied to the reconstructed pixels as a post-process. This post-processing occurs after the deformation process (or other view synthesis reconstruction) has occurred. Surface normal information relative to light positions or points may be known and included in the encoded light field data. This encoded list of light positions allows a decoder to use this normal data when decoding a particular pixel in a layer to compute the specular component. Other parameters may be included in the light field data, for example, properties of whether a surface has a specular component, or the extent of it may be numerically quantified. Material properties may also be included along with the intensity values. This additional data may be sent along with the typical RGB and depth data that is sent in encoded form in each element image or each layer. Material properties may include, but are not limited to, atomic, chemical, mechanical, thermal, and optical properties.

[0486] The concept of storing surface normals in combination with RGB and depth information is called G-Buffer in computer graphics.

[0487] transparency

[0488] In traditional 2D display computer graphics, it is often necessary to simulate the visual effects of non-opaque surfaces. In surface rendering, this usually involves assigning a non-opacity (also known as a synonym for transparency and translucency) measure to surface elements (e.g., polygons, vertices, etc.).

[0489] When a 3D scene composed of such surfaces is rendered from a particular virtual camera view, the transparency metric of each surface element determines the extent to which surfaces occluded by one surface element will optically bleed through the unoccluded, most positive surface. There are various ways to simulate this process, including computing the final color value in the image based on a blend of overlapping surfaces. Of course, the key to such a computation is that each overlapping patch of surface elements has (1) a color and (2) a transparency metric, or alpha, value.

[0490] It is also desirable to be able to represent and render (transmit / decode) scenes consisting of transparent surfaces for light field video purposes. We describe how a scene decomposition-based representation scheme can be enhanced to provide support for transparent surfaces.

[0491] During encoding, it is described how normals or other optical surface properties can be written along with the encoded representation in addition to the color and depth maps. It is also suggested that an alpha (transparency) value α associated with each pixel can be additionally written during encoding.

[0492] This alpha value can then be used during the decoding process to generate a light field image of the scene containing transparent surfaces. The alpha values ​​associated with the sampled pixels generated or selected during encoding can be reprojected along with the depth values ​​during the reconstruction process in which the light field associated with each layer (or scene subset) is reconstructed. This reprojection can be a warping process as described in connection with depth image based rendering (DIBR). In a typical embodiment, the end result is that each pixel in the light field associated with a layer or subset will also have an associated alpha value to represent transparency.

[0493] In order to incorporate this transparency alpha value into the final image, we have to use an operator during the merging process that incorporates some optical blending model. During the layer (or subset) merging process, use * m operator, we can use α as a means of accumulating the color observed along a single light field image pixel over multiple layers (or subsets, etc.).

[0494] It is proposed to naturally combine the layered decoding process based on polygonal surface rendering with the volume rendering process. The basic idea is that in this process, a single layer or a subset is reconstructed and then merged with the adjacent layers (via * m operator). It is suggested that in addition to the merging operator, the composition equation for volume rendering must also be included. Therefore, in this more general hybrid approach, the layer combination operator becomes more general and complex, as it performs a more general function, sometimes acting as a merging operator as before, and sometimes acting as a volume rendering ray accumulation function. This operator is denoted as * c .

[0495] Figure 15 An exemplary hierarchical scene decomposition CODEC method is shown, whereby 3D scene data or light field data in an image description format is loaded into an encoder (400) for encoding, whereupon (sub)sets of data as shown, or the entire data set representing the 3D scene, are partitioned (403). Where a subset of 3D scene data is identified for partitioning (402), it will be appreciated that the identification process is a general process step reference intended to simply refer to the ability to partition a data set in a pass or in groups (e.g., encoding inner and outer frustum data layers, as shown). Figure 11 ), as may be required depending on the circumstances. In this regard, the identification of data subsets may imply a pre-encoding processing step or a processing step that also forms part of the encoding sub-processing stage (401). The data subsets may be marked, specified, confirmed, scanned, or even compiled or grouped when partitioned to produce a set of layers (decomposition of the 3D scene) (403). Following partitioning of the data subsets (403), each data layer is sampled and rendered in accordance with the present disclosure to produce compressed (image) data (404). Following compression of the data layers, the compressed data is transmitted to a decoder (405) for a decoding sub-process, including decompression, decoding, and reassembly steps to (re)construct a set of light fields (407), otherwise referred to herein as "layered light fields", layered light field images, and light field layers. Specular illumination is calculated (411) and the constructed layered light fields are merged to produce a final light field (408) that displays the 3D scene (409).

[0496] Figure 16 A computer-implemented method is described, comprising:

[0497] receiving a first data set (420) comprising a three-dimensional description of a scene;

[0498] The first data set comprises information (421) about the directions of normals on surfaces comprised in the scene; the directions of the normals are expressed relative to a reference direction (422); and

[0499] Optionally, the reflective properties of at least some of the surfaces are non-Lambertian;

[0500] dividing the first data set into a plurality of subsets, each subset representing a different portion of the scene at a different position relative to a reference position (423); and

[0501] The plurality of subsets and the plurality of sub-portions are encoded to generate a second data set, wherein a size of the second data set is smaller than a size of the first data set (424).

[0502] In one embodiment, the method further comprises:

[0503] receiving a second data set (425);

[0504] reconstructing the portion associated with the subset using normal directions on surfaces included in the scene to compute a specular component (426);

[0505] combining the reconstructed parts into a light field (427); and

[0506] The light field image is presented on a display device (428).

[0507] Figure 21 A computer-implemented method is described, comprising:

[0508] receiving a first data set (420) comprising a three-dimensional description of a scene;

[0509] The first data set includes information (429) about the transparency of surfaces included in the scene; and

[0510] dividing the first data set into a plurality of layers, each layer representing a portion of the scene at a position relative to a reference position (423); and

[0511] The plurality of layers are encoded to generate a second data set, wherein a size of the second data set is smaller than a size of the first data set (424).

[0512] In one embodiment, the method further comprises:

[0513] receiving a second data set (425);

[0514] combining the reconstructed parts into a light field (427); and

[0515] The light field image is presented on a display device (428).

[0516] In order to better understand the invention described herein, the following embodiments are described with reference to the accompanying drawings. It should be understood that these embodiments are intended to describe illustrative embodiments of the present invention and are not intended to limit the scope of the present invention in any way.

[0517] Example

[0518] Example 1: Exemplary encoder and encoding method for light field display

[0519] The following illustrative embodiments of the present invention are not intended to limit the scope of the invention as described and claimed herein, as the present invention can successfully implement multiple system parameters. As described above, conventional displays previously known in the art consist of spatial pixels that are substantially evenly spaced and organized into two-dimensional rows, allowing for idealized uniform sampling. In contrast, three-dimensional (3D) displays require both spatial and angular samples. While the spatial sampling of a typical 3D display remains consistent, the angular sampling is not necessarily considered to be consistent in terms of the footprint of the display in angular space.

[0520] In the illustrative embodiment, multiple planar parameterized pinhole projectors provide angular samples, also known as the directional components of the light field. The light field display is designed with a spatial resolution of 640×480 and an angular resolution of 512×512. The multiple planar parameterized pinhole projectors are idealized using the identity function α. ​​The spacing between each of the multiple planar parameterized pinhole projectors is 1 mm, defining a display surface of 640 mm×400 mm. The display has a 120° FOV, corresponding to an approximate focal length f=289 μm.

[0521] This light field display contains 640×480×512×512=80.5 billion RGB pixels. Each RGB pixel requires 8 bits, so one frame of the light field display requires 80.5 billion×8×3=1.93Tb. For a light field display to provide interactive content, driving data at 30 frames per second requires a bandwidth of 1.93Tb×30 frames per second=58.0Tb / s. Current displays known in the art are driven by DisplayPort technology, which offers a maximum bandwidth of 32.4Gb / s. Therefore, such a display would require more than 1024 DisplayPort cables to provide the massive bandwidth required for an interactive light field display, resulting in cost and form factor limitations.

[0522] The illustrative embodiment transmits data from a computer equipped with an accelerated GPU with dual DisplayPort 1.3 cable outputs to a light field display. We consider a conservative maximum throughput of 40Gb / s. The encoded frames must be small enough to be transmitted over the DisplayPort connection to the decoding unit physically closer to the light field display.

[0523] The hierarchical scene decomposition of the illustrative embodiment is designed to allow the required data throughput. Based on the dimensions defined above, the maximum depth of field of the light field display is Z DOF = (289 microns) (512) = 147968 microns = 147.986 mm. Layered scene decomposition Place multiple layered scene decomposition layers within the depth of field of the light field display, ensuring that the distance between the layered scene decomposition layer and the display surface is less than Z DOF . This illustrative embodiment describes a light field display with objects located only within the inner frustum volume of the display. This illustrative embodiment is not intended to limit the scope of the invention, as the invention can successfully implement multiple system parameters, such as a light field display with objects located only within the outer frustum volume of the display, or a light field display with objects located within both the inner and outer frustum volumes of the display; implementations limited to one frustum volume require a smaller number of hierarchical scene decomposition layers, thereby slightly reducing the size of the encoded light field to be generated.

[0524] The illustrative embodiment defines ten scene decomposition layers. Additional hierarchical scene decomposition layers can be added as needed to capture data that might be lost due to occlusion or to improve the overall compression rate. However, additional hierarchical scene decomposition layers require additional computation from the decoder, so the number of hierarchical scene decomposition layers is carefully chosen. The illustrative embodiment specifies ten hierarchical scene decomposition layers based on their front and back boundaries, assuming that the layer division boundaries are parallel to the display surface.

[0525] Each hierarchical scene decomposition layer is located at a defined distance from the display surface, where the distance is specified as a multiple of the focal length f, with a maximum depth of field of 512f. Hierarchical scene decomposition layers with narrower widths are concentrated closer to the display surface, and the layer width (i.e., the depth difference between the front and back layer boundaries) increases exponentially as a power of 2 with increasing distance from the display surface. This embodiment of the invention is not intended to limit the scope of the invention, as other layer configurations may be successfully implemented.

[0526] The following table (Table 1) describes the hierarchical scene decomposition layer configuration of an exemplary embodiment and provides a sampling scheme based on plenoptic sampling theory to create sub-sampled hierarchical scene decomposition layers:

[0527] Table 1

[0528]

[0529]

[0530] In the table above, layer 0 captures the image to be displayed on the display surface, as in conventional two-dimensional displays known in the art. Layer 0 contains 640×480 pixels of fixed depth, so no depth information is required. The total data size is calculated for each pixel, each with an 8-bit RGB value and a depth value (alternative implementations may require larger bit values, such as 16 bits). In the illustrative embodiment, the element image resolution and sampling gap are calculated by the above formula, and the sampling scheme selected reflects the element image resolution and sampling gap constraints.

[0531] As shown in the table above, the total size of the combined hierarchical scene decomposition system is 400.5 Mb. Therefore, to produce data at a rate of 30 frames per second, a bandwidth of 30 × 0.4005 = 12.01 GB / s is required. This encoding, along with the additional information required to represent scene occlusion, is sent over a dual DisplayPort 1.3 cable.

[0532] In an illustrative embodiment, the hierarchical scene decomposition layers are configured by the encoder, effectively implementing an oblique rendering technique to produce layers located closer to the display surface (layers 0 through 5) and a perspective rendering technique to produce layers located further away from the display surface (layers 6 through 9). Each element image corresponds to a rendered view.

[0533] At level 6, the number of separate angles to be rendered (64×64=4096) exceeds the number of views to be rendered (21×16=336); this marks the transition in efficiency between oblique and perspective rendering methods. It should be noted that specific implementation aspects may provide additional overhead that distorts the exact sweet spot. For use with modern graphics acceleration techniques known in the art, perspective rendering can be efficiently implemented using geometry shader instancing. Multiple views are rendered from the same set of input scene geometry without repeatedly accessing the geometry through draw calls or repeatedly accessing memory to retrieve the exact same data.

[0534] Figure 8 An illustrative embodiment is shown having ten hierarchical scene decomposition layers (100-109) in an inner frustum volume (110). The inner frustum volume layers extend from the display surface (300). The layers are defined as described in the table above, for example, the front boundary of inner frustum volume layer 0 (100) is 1f, inner frustum volume layer 1 (101) is 1f, inner frustum volume layer 2 (102) is 2f, inner frustum volume layer 3 (103) is 4f, and so on. Inner frustum volume layers (100-105) 0 to 5, or the layers closest to the display surface (300), are rendered using an oblique rendering technique, and inner frustum volume layers (106-109), the layers 6 to 9 farthest from the display surface, are rendered using a perspective rendering technique.

[0535] Figure 9 An alternative embodiment is illustrated having ten hierarchical scene decomposition layers (100-109) in an inner frustum volume (110) and ten hierarchical scene decomposition layers (200-209) in an outer frustum volume (210). The inner frustum volume layers and the outer frustum volume layers extend from a display surface (300). Although the inner frustum volume layers and the outer frustum volume layers are shown as mirror images of each other, the inner frustum volume and the outer frustum volume can have different numbers of layers, layers of different sizes, or layers of different depths. Inner frustum volume layers 0 to 5 (100-105) and outer frustum volume layers 0 to 5 (200-205) are rendered using an oblique rendering technique, and inner frustum volume layers 6 to 9 (106-109) and outer frustum volume layers 6 to 9 (206-209), which are farther from the display surface (300), are rendered using a perspective rendering technique.

[0536] An alternative embodiment may implement the system using an approach based on ray tracing encoding. Rendering a complete hierarchical scene decomposition layer representation may require increased GPU performance, even using the optimizations described herein, because GPUs are optimized for interactive graphics on traditional two-dimensional displays where accelerated rendering of a single view is required. The computational cost of the ray tracing approach is a direct function of the number of pixels to be rendered by the system. While a hierarchical scene decomposition layer system contains a comparable number of pixels as some two-dimensional single-view systems, the form and arrangement of the pixels are very different due to the layer decomposition and the corresponding sampling scheme. Therefore, there may be implementations where tracing some or all rays is a more efficient implementation.

[0537] Example 2: CODEC decoder and decoding method for light field display

[0538] In an illustrative embodiment of the present invention, the decoder receives 12.01 GB / s of encoded core representation data and any residual representation data from the GPU via dual DisplayPort 1.3 cables. The compressed core representation data is decoded using a custom FPGA, ASIC, or other integrated circuit for efficient decoding (the residual representation data is decoded separately, e.g., Figure 13 For the final light field display, the 12.01GB / s core representation decompresses to 58Tb / s. Note that this core representation does not include the residual representation required to render occlusion. Provides a compression ratio of 4833:1; while this is a high-performance compression ratio, the reconstructed light field data may still exhibit occlusion-based artifacts unless residual representation data is included in the reconstruction.

[0539] for Figure 8 In the illustrative embodiment shown in , the data is decoded by reconstructing separate hierarchical scene decomposition layers and merging the reconstructed layers into the inner frustum volume layer. For alternative embodiments, such as Figure 9 As shown in , the data is decoded by reconstructing separate hierarchical scene decomposition layers and merging the reconstructed layers into inner and outer frustum volume layers.

[0540] A single layer of scene decomposition can be reconstructed from a given data sampling scheme using view synthesis techniques known in the art from the field of image-based rendering. For example, Graziosi et al. specify the use of reference element images to reconstruct the light field in a single pass. The method uses reference element images that are offset from a multi-dimensional reconstructed image. Since the element image data represents a three-dimensional scene point (including RGB color and disparity), the pixels are decoded as a non-linear function (although fixed on the direction vector between the reference element image and the target element image), so decoding the reference element image requires a memory buffer of the same size. This may cause memory storage or bandwidth limitations when decoding larger element images, depending on the decoding hardware.

[0541] For a light field display with an elemental image size of 512×512 pixels and 24-bit color, the decoder requires a buffer capable of storing 512×512 = 262,144 24-bit values ​​(excluding the disparity bits in this example). Current high-performance FPGA devices offer internal block RAM (BRAM) organized as 18 / 20-bit wide memories with 1024 memory locations, which can be used as a 36 / 40-bit wide memory with 512 memory locations. A buffer capable of reading and writing images in the same clock cycle must be large enough to accommodate two reference elemental images, as the nonlinear decoding process causes the write port to use a nondeterministic access pattern. Implementing this buffer in an FPGA device for a 512×512 pixel image requires 1024 BRAM blocks. Depending on the reconstruction algorithm used, multiple buffers may be required in each decoder pipeline. To meet the data rates of high-density light field displays, the system may require over a hundred parallel pipelines, which is several orders of magnitude more pipelines than current FPGA devices. Because each buffer requires a separate read / write port, implementing such a system on current ASIC devices may be impossible.

[0542] The present disclosure circumvents buffer and memory limitations by breaking the pixel reconstruction process into multiple single-dimensional stages. The present disclosure implements one-dimensional reconstruction to fix the direction vector between the reference element image and the target to a correction path. Although the reconstruction remains non-linear, the reference pixel to be transformed to the target position is locked to the same row or column position of the target pixel. Therefore, the decoder buffer only needs to capture one row or column at a time. For the above-mentioned 512×512 pixel element image with 24-bit color, the decoder buffer is organized as a 24-bit wide, 1024 deep memory, requiring two 36 / 40×512BRAMs. Therefore, the present disclosure has reduced the memory footprint by a factor of 512 or multiple orders of magnitude. This allows current FPGA devices to support display pixel fill rates that require more than one hundred decoding pipelines.

[0543] A multi-stage decoding architecture requires two stages to reconstruct the two-dimensional pixel array in a light field display. These two stages are orthogonal to each other and reconstruct rows or columns of the elemental image. The first decoding stage may require a pixel scheduler to ensure that the output pixels are ordered to be compatible with the input pixels of the next stage. Because each decoding stage requires extremely high bandwidth, it may be necessary to reuse some of the output pixels of the previous stage to reduce local storage requirements. In this case, an external buffer can be used to capture all the output pixels of the first stage so that subsequent decoding stages can efficiently access the pixel data, thereby reducing logic resources and memory bandwidth.

[0544] The disclosed multi-stage decoding with an external memory buffer allows the decoding process to shift required memory bandwidth from expensive on-die memory to low-cost memory devices such as double data rate (DDR) memory devices. A high-performance decoded pixel scheduler ensures maximum reuse of reference pixels in this external memory buffer, allowing the system to use narrower or slower memory interfaces.

[0545] The disclosures of all patents, patent applications, publications, and database entries cited in this specification are hereby incorporated by reference in their entirety to the same extent as if each such individual patent, patent application, publication, and database entry was specifically and individually indicated to be incorporated by reference.

[0546] Although the present invention has been described with reference to certain specific embodiments, various modifications thereof will be apparent to those skilled in the art without departing from the spirit and scope of the present invention. All such modifications apparent to those skilled in the art are intended to be included within the scope of the appended claims.

[0547] References

[0548] ALPASLAN, ZAHIR Y., EL-GHOROURY, HUSSEIN S., CAI, JINGBO, "ParametricCharacterization of Perceived Light Field Display Resolution", pp. 1241–1245, 2016.

[0549] BALOGH、TIBOR、 The Holovizio system–New opportunity offered by 3D displays, Proceedings of the TMCE, (May):1–11, 2008.

[0550] BANKS, MARTIN S., DAVID M. HOFFMAN, JOOHWAN KIM AND GORDON WETZSTEIN, "3D Displays", Annual Review of Vision Science, 2016, pp. 397 - 435.

[0551] CHAI, JIN - XIANG, XIN TONG, SHING - CHOW CHAN and HEUNG - YEUNG SHUM. "Plenoptic Sampling"

[0552] CHEN, A., WU M., ZHANG Y., LI N., LU J., GAO S. and YU J.. 2018, "Deep Surface Light Fields", Proc. ACM Comput. Graph. Interact. Tech. 1, 1, Article 14 (July 2018), 17 pages. DOI: https: / / doi.org / 10.1145 / 3203192

[0553] CLARK, JAMES J., MATTHEW R. PALMER and PETER D. LAWRENCE, "A Transformation Method for the Reconstruction of Functions from Nonuniformly Spaced Samples", IEEE Transactions on Acoustics, Speech, and Signal Processing, October 1985, pp1151 - 1165. Vol. ASSP - 33, No. 4.

[0554] DO, MINH N., DAVY MARCHAND - MAILLET and MARTIN VETTERLI, "On the Bandwidth of the Plenoptic Function", IEEE Transactions on Image Processing, pp. 1 - 9.

[0555] DODGSON, N.A. Analysis of the viewing zone of the Cambridge autostereoscopic display,Applied optics,35(10):1705–10,1996。

[0556] DODGSON, N.A. Analysis of the viewing zone of multiview autostereoscopic displays,Electronic Imaging 2002,International Society for Optics and Photonics, pp. 254–265, 2002。

[0557] GORTLER, STEVEN J., RADEK GRZESZCZUK, RICHARD SZELISKI and MICHAEL F. COHEN."The Lumigraph" 43-52。

[0558] GRAZIOSI, D.B., APLASLAN, Z.Y., EL-GHOROURY, H.S., Compression for Full-Parallax Light Field Displays,Proc. SPIE 9011,Stereoscopic Displays and Applications XXV, (MARCH), 90111A, 2014。

[0559] GRAZIOSI, D.B., APLASLAN, Z.Y., EL-GHOROURY, H.S., Depth Assisted Compression of Full Parallax Light Fields,Proc. SPIE 9391, Stereoscopic Displays and Applications XXVI, (FEBRUARY), 93910Y. 2015。

[0560] HALLE, MICHAEL W. and ADAM B. KROPP, "Fast Computer Graphics Rendering for Full Parallax Spatial Displays", Proc. SPIE 3011, Practical Holography XI and Holographic Materials III, (April 10, 1997).

[0561] HALLE, MICHAEL W, Multiple Viewpoint Rendering. In Proceedings of the 25th annual conference on Computer graphics and interactive techniques (SIGGRAPH’98). Association for Computing Machinery, New York, NY, USA, 243–254.

[0562] JANTET, VINCENT, "Layered Depth Images for Multi-View Coding" Multimedia.pp.1-135, Universite Rennes 1, 2012. English.

[0563] LANMAN, D., WETZSTEIN, G., HIRSCH, M. and RASKAR, R., Depth of Field Analysis for Multilayer Automultiscopic Displays, Journal of Physics: Conference Series, 415(1):012036, 2013.

[0564] LEVOY, MARC and PAT HANRAHAN, "Light Field Rendering" SIGGRAPH.pp.1-12.

[0565] MAARS, A., WATSON, B., HEALEY, C.G., Real-Time View Independent Rasterization for Multi-View Rendering. Eurographic Proceedings, The Eurographics Association. 2017.

[0566] MARSCHNER, STEPHEN R. and RICHARD J. LOBB, "An Evaluation of Reconstruction Filters for Volume Rendering" IEEE Visualization Conference 1994.

[0567] MARTIN, S, “View Synthesis in Light Field Volume Rendering using Convolutional Neural Networks”, University of Dublin, August 2018.

[0568] MASIA, B., WETZSTEIN, G., ALIAGA, C., RASKAR, R., GUTIERREZ, D., Display adaptive 3D content remapping. Computers and Graphics (Pergamon), 37(8):983–996, 2013.

[0569] MATSUBARA, R., ALPASLAN, ZAHIR Y., EL-GHOROURY, HUSSEIN S., Light Field Display Simulation for Light Field Quality Assessment. Stereoscopic Displays and Applications XXVI, 7(9391):93910G, 2015.

[0570] PIAO, YAN and XIAOYUAN YAN, "Sub-sampling Elemental Images for Integral Imaging Compression" IEEE.pp.1164-1168.2010.

[0571] VETRO, ANTHONY, THOMAS WIEGAND, and GARY J. SULLIVAN, "Overview of the Stereo and Multiview Video Coding Extensions of the H.264 / MPEG-4 AVC Standard." Proceedings of the IEEE. pp. 626 - 642. April 2011. Vol. 99, No. 4.

[0572] WETZSTEIN, G., HIRSCH, M., Tensor Displays: Compressive Light Field Synthesis using Multilayer Displays with Directional Backlighting. 1920.

[0573] WIDMER, S., D. PAJAK, A. SCHULZ, K. PULLI, J. KAUTZ, M. GOESELE, and D. LUEBKE, An Adaptive Acceleration Structure for Screen-space Ray Tracing. Proceedings of the 7th Conference on High-Performance Graphics, HPG’15, 2015.

[0574] ZWICKER, M., W. MATUSIK, F. DURAND, H. PFISTER, "Antialiasing for Automultiscopic 3D Displays" Eurographics Symposium on Rendering. 2006.

Claims

1. A computer-implemented method comprising: receiving a first data set comprising a three-dimensional description of a scene; dividing the first data set into a plurality of layers, each layer representing a different portion of the scene at a different position relative to a reference position; dividing data corresponding to at least one of the layers into a plurality of sub-portions, wherein positions of particular sub-portions are determined based on a geometry of at least a portion of an object represented within the scene; as well as encoding the plurality of layers and the plurality of sub-portions to generate a second data set, The size of the second data set is smaller than the size of the first data set. 2 . The method of claim 1 , further comprising transmitting the second data set to a remote device for rendering the scene at a display device associated with the remote device.

3. The method according to claim 1 or 2, wherein: Encoding a layer or a sub-portion comprises performing a sampling operation on a corresponding portion of the first data set.

4. The method according to claim 3, wherein: The sampling operation is based on a target compression ratio associated with the second data set.

5. The method according to claim 1 or 2, wherein: Encoding multiple layers and multiple sub-parts involves: Render the set of pixels to be encoded using ray tracing; selecting a plurality of element images from the plurality of element images such that the set of pixels is rendered using the selected plurality of element images; and The set of pixels is sampled using a sampling operation.

6. The method according to claim 5, wherein: The sampling operation includes selecting a plurality of element images from corresponding portions of the plurality of element images according to a plenoptic sampling scheme.

7. The method according to claim 5, wherein: Executing the sampling operation includes: determining an effective spatial resolution associated with the layer or sub-portion; and A plurality of element images are selected from corresponding portions of the plurality of element images according to the determined angular resolution.

8. The method according to claim 7, wherein: The angular resolution is determined as a function of a directional resolution associated with the portion of the scene associated with the layer or sub-portion.

9. The method according to claim 7, wherein: The angular resolution is determined as a field of view associated with the display device.

10. The method according to claim 1 or 2, wherein: The three-dimensional description includes light field data representing a plurality of elemental images.

11. The method according to claim 10, wherein: Each of the plurality of element images is captured by one or more image acquisition devices.

12. The method according to claim 10, wherein: The light field data includes a depth map corresponding to the elemental image.

13. The method according to claim 1 or 2, wherein: The first data set comprises information about normal directions on surfaces comprised in the scene, the normal directions being expressed relative to a reference direction.

14. The method according to claim 13, wherein: The reflective properties of at least some of the surfaces are non-Lambertian.

15. The method according to claim 1 or 2, wherein: Encoding the layer or sub-portion further comprises: obtaining, for the layer or sub-portion, one or more polygons representing corresponding portions of objects in the scene; determining a view-independent representation based on the one or more polygons; and A view-independent representation is encoded in the second data set.

16. The method according to claim 1 or 2, further comprising: receiving the second data set; decoding a portion of the second data set corresponding to each said layer and each said sub-portion; combining the decoded parts into a representation of a light field image; as well as The light field image is presented on a display device.

17. The method according to claim 16, further comprising: receiving user input indicating a position of a user relative to the light field image; as well as The light field image is updated based on the user input before being presented on the display device.

18. The method according to claim 1 or 2, wherein: A layer located closer to the display surface achieves a lower compression ratio than a layer of the same width located further away from the display surface.

19. The method according to claim 1 or 2, wherein: The plurality of layers of the second data set comprises a light field.

20. The method according to claim 19, wherein The light fields are combined to create a final light field.

21. The method according to claim 1 or 2, wherein Dividing the layers includes limiting the depth range of each layer.

22. The method according to claim 1 or 2, wherein: Layers located closer to the display surface are narrower in width than layers located farther from the display surface.

23. The method according to claim 1 or 2, wherein Dividing the first data set into a plurality of layers maintains a uniform compression rate throughout the scene.

24. The method according to claim 1 or 2, wherein: Partitioning the first data set into a plurality of layers includes partitioning the first data set into an inner frustum volume set of layers and an outer frustum volume set of layers.

25. The method according to claim 1 or 2, wherein The method is used to generate a synthetic light field for multi-dimensional video streams, multi-dimensional interactive games, real-time interactive content or other light field display scenes.

26. The method according to claim 25, wherein The synthetic light field is generated only in the effective viewing area.

Citation Information

Patent Citations

  • Methods for Full Parallax Compressed Light Field 3D Imaging Systems

    US20150201176A1

  • Methods for Full Parallax Compressed Light Field Synthesis Utilizing Depth Information

    US20160360177A1

  • Content Adaptive Light Field Compression

    US20170142427A1

  • Unibiased light field models for rendering and holography

    US6549308B1

  • Distributed system for producing holographic stereograms on-demand from various types of source material

    US7436537B2