Light Field Offset Rendering
The light field rendering method employs a light field offset scheme and pixel remapping techniques to address the inefficiencies in existing methods, achieving reduced rendering times and computational requirements, and enabling high-quality light field rendering with fewer cameras.
Patent Information
- Application Number
- JP2024562312
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-18
- Filing Date
- 2023-05-03
- Publication Date
- 2025-05-20
AI Technical Summary
Existing light field rendering methods face challenges in efficiently capturing and rendering high-quality light field images, particularly due to limitations in ray tracing support in content creation tools and high computational power requirements.
A computer-implemented light field rendering method using a light field offset scheme and pixel remapping techniques, which involves defining a display surface, capturing views of a 3D scene with a camera array at an offset distance, and performing pixel remapping to create a light field image rendered on the display surface.
This method reduces rendering times and processing requirements compared to ray tracing methods, enabling efficient high-quality light field rendering suitable for offline applications, while also reducing the number of light field cameras needed to capture a 3D scene.
Smart Images

Figure 2025515588000001_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Patent Application No. 17 / 663,943, filed May 18, 2022, which is incorporated by reference herein in its entirety.
[0002] This disclosure relates to light field rendering methods that use light field offset schemes and pixel remapping techniques, and to the use of image rasterization for rendering light fields. [Background technology]
[0003] Light field displays recreate the experience of looking at the real world through a window by replicating a light field that describes all the light rays associated with a particular scene. Creating an image for a standard display from a three-dimensional (3D) scene is called rendering or light field rendering. One rendering method, ray tracing, is done by modeling the light in the scene. The ray tracing model can then be used in a wide variety of rendering algorithms to generate a digital image from the light model. Another rendering method, called rasterization, takes an image described in a vector graphics format or shape and converts it into a raster image, which is a series of pixels, dots, or lines that, when viewed together, create an image represented through the shapes. Although ray tracing is ideal for light field rendering, many features of content creation tools do not support ray tracing. Thus, rasterization techniques, which sacrifice fast rendering times, are sometimes used as offline renderers to generate high-quality content for light field displays.
[0004] Existing 3D display technologies can be divided into two categories: binocular stereoscopic and autostereoscopic. Binocular stereoscopic displays use special eyewear to facilitate the observation of two slightly different images in the left and right eyes, creating a depth cue. Autostereoscopic displays allow the viewer to see different images with the naked eye depending on where they are looking at the display. This can be compared to a viewer looking through a window, looking from the left side and seeing an entirely different image than the right side. Traditional two-dimensional light field displays, at least the images displayed are not stereoscopic and both eyes see the same image, so they are unable to provide the proper depth cues for the brain to interpret the images seen above them in the same way that they interpret real world objects.
[0005] Multiview3D, a class of autostereoscopic display technology, approximates a light field associated with a scene. Defined as a function describing the amount of light flowing in all directions through a point in space free of occluding objects, the light field contains information about all images that can be seen from all possible combinations of viewing positions and angles relative to the display surface. An actual light field contains an infinitely continuous number of light rays. In practice, to limit the amount of data needed to describe the scene, the rendered light field is discretized by selecting a finite set of "views" that are used to estimate the actual light field at any given position from the 3D display.
[0006] In one example of light field rendering, Do et al. (Do, Luat, and Sveta Zinger. "Quality Improvement Techniques for Free Viewpoint DIBR." Stereoscopic Display and Applications XXI. Vol. 7524. International Society for Optics and Photonics, 2010) describe a rendering algorithm based on depth image warping between two reference views from an existing camera. While experimenting with synthetic data, it is observed that the rendering quality is highly dependent on the scene complexity, and the method is performed using compressed video from surrounding cameras. The overall system quality is governed by the rendering quality and not by the encoding, resulting in long rendering times and high computational power requirements.
[0007] In another example of light field rendering, Li et al. (Li, Hui, et al. "Optimized layered rendering method for real-time interactive holographic display based on ray tracing technology" Optical Engineering 59.10 (2020): 102408) describes an optimized layered rendering method for real-time interactive holographic display based on ray tracing technology to overcome the challenges of light field rendering and realize real-time interaction of three-dimensional scenes. The reconstructed holographic image with real depth cues is demonstrated by experiments, and the optically reconstructed image can be interacted in real time, but the ray tracing-based method requires high computational power.
[0008] In another example, US Patent Application Publication No. 20220012860 to Zhang et al. describes a method and apparatus for synthesizing a six degree of freedom view from a sparse red-green-blue (RGB) depth input. The exemplary apparatus includes at least one memory, instructions therein, and a processor circuit for executing the instructions to reproject a first depth view and a second depth view onto a target camera position to obtain a first reprojected view and a second reprojected view, combine the first reprojected view and the second reprojected view into a mixed view including missing RGB depth information due to at least one of occlusion or non-occlusion, and generate a six degree of freedom synthesized view of the mixed view data, the synthesized view including the missing RGB depth information. The method requires at least two views, reprojecting each view separately, and then blending the views to generate the synthesized view.
[0009] Based on the pixel directionality of light field displays (LFDs), utilizing ray tracing for light field rendering is an organic approach and has been thoroughly studied. It is the fastest method described for light field rendering and can be used for real-time rendering, thereby enabling interactive light field experiences such as video games. However, many existing tools used by content creators do not support ray tracing to the extent required to implement high quality light field rendering. Since integration with existing architectures is crucial for the development and adoption of new technologies, it is desirable to enable designers to move quickly into generating content for LFDs. Thus, there is still a need for high performance light field rendering methods with rasterization that reduce rendering times for applications where quality takes precedence over speed, such as offline rendering. Rendering methods with rasterization that are designed to present images at 30 frames per second or more and can be configured to render light field images in addition to traditional 2D images, are available in commercial real-time rendering software, including but not limited to UnrealEngine®, Maya, Blender®, and Unity.
[0010] This background information is provided for the purpose of making known information believed by the applicant to be of possible relevance to the present invention. It is not necessarily intended, nor should it be construed, that any of the preceding information constitutes prior art against the present invention. Summary of the Invention
[0011] It is an object of the present invention to provide a light field rendering method for rendering a 3D light field using a light field offset scheme and pixel remapping techniques.It is another object of the present invention to provide a method for light field mapping and rendering using image rasterization.
[0012] According to one aspect, a computer-implemented light field rendering method is provided that includes the steps of: defining a display surface for the light field, the light field including an inner frustum volume bounded by the display surface and a far clip plane, and an outer frustum volume bounded by the display surface and a near clip plane; defining an in-draw plane parallel to the display surface and at an integral offset distance therefrom, the in-draw plane comprising a plurality of light field cameras spaced apart by a sample gap; capturing a view of the 3D scene as a source image with each of the plurality of light field cameras; decoding each source image to generate a plurality of hogel cameras on the in-draw plane, each hogel camera providing an elemental image; generating an integral image from the elemental images at the in-draw plane comprising a plurality of pixels; and performing a pixel remapping technique on individual pixels in the integral images to create a light field image rendered on the display surface.
[0013] In one embodiment of the method, the portion of the 3D scene captured at the basin includes image information from an inner frustum volume and an outer frustum volume.
[0014] In another embodiment of the method, the captured 3D scene includes all of the image information within the outer frustum volume.
[0015] In another embodiment of the method, the integral offset distance is calculated from the focal length, the directional resolution of the light field camera, and the offset integer N.
[0016] In another embodiment of the method, the offset integer N>1.
[0017] In another embodiment of the method, the lead-in surface is located at the proximal clip surface.
[0018] In another embodiment, the method further comprises displaying the rendered light field image on a light field display.
[0019] In another embodiment of the method, the surface area of the lead-in surface is greater than the surface area of the display surface.
[0020] In another embodiment of the method, the optical properties of the light field camera are orientation, lens pitch, directional resolution, and field of view.
[0021] In another embodiment, the method further comprises generating a plurality of integral images at a plurality of entrapment planes.
[0022] In another embodiment, the method further includes combining the integral images to create a combined rendered light field image on the display surface.
[0023] In another embodiment of the method, the compositing incorporates transparency data.
[0024] In another embodiment of the method, each light field camera is one of a digital single-lens reflex (DSLR) camera, a pinhole camera, a plenoptic camera, a compact camera, and a mirrorless camera.
[0025] In another embodiment of the method, each light field camera is a computer-generated camera.
[0026] In another embodiment of the method, the pixel remapping technique remaps the pixels to their hogel index (H x ,H y ) from the pull-in surface to the display surface.
[0027] In another embodiment of the method, the plane of intrusion is outside the outer frustum volume.
[0028] In another aspect, a method includes the steps of: capturing a first light field at a drawing plane relative to a light field display surface using a light field camera, the first light field including an array of drawing plane hogels, each hogel having a plurality of pixels; and assigning a hogel index (H ) to each pixel in each drawing plane hogel by applying a pixel remapping technique to select a single pixel from each drawing plane hogel. x ,H y ) and pixel index (P x ,P y ) to indicate its location in the light field display; and, using a compositing function, approximating each pixel to the fascination plane (LF r ) and load each pixel from the light field (LF d ) and generating a light field image on the display surface that includes the remapped pixels.
[0029] In another embodiment of the method, one pixel from the entrapment surface generates one pixel on the display surface.
[0030] In another embodiment of the method, applying the pixel remapping technique comprises: determining a hogel index (H x ,H y ) and change the pixel index (P x ,P y ) remains constant.
[0031] In another embodiment of the method, the pixel remapping technique is a function of the directional resolution and the offset parameter N of the light field display.
[0032] In another embodiment of the method, the pixel remapping technique is according to the formula LF r [H x +(DR x *N)-(N*P x ),H y +(DR y *N)-(N*P y ),P x ,P y ]⇒LF d [H x ,H y ,P x ,P y ] based on
[0033] In another embodiment of the method, the offset integer N>1.
[0034] In another embodiment of the method, the drawing surface is constructed from a sufficient number of hogels to provide a number of pixels to achieve the required directional resolution of the light field display at the viewing surface.
[0035] In another embodiment of the method, the size of the display surface is defined by the directional and spatial resolution of the light field display.
[0036] In another embodiment of the method, the light field camera is a mirror (DSLR) camera, a pinhole camera, a plenoptic camera, a compact camera, or a mirrorless camera. [Brief description of the drawings]
[0037] These and other features of the present invention will become more apparent in the following detailed description, when taken in conjunction with the accompanying drawings.
[0038] [Figure 1A] FIG. 1 is an illustration of one method of rasterization that uses frustum projection to flatten a 3D scene into a 2D image.
[0039] [Figure 1B] FIG. 13 is an example of a front view of a flattened 2D image.
[0040] [Diagram 2] FIG. 1 is an illustration of an example of a ray tracing technique for solving visibility problems.
[0041] [Diagram 3] FIG. 2 is an example of a digital representation of a ground truth image.
[0042] [Figure 4] FIG. 1 is an example diagram showing one example of a method for rendering a light field showing inner and outer frustums.
[0043] [Diagram 5] FIG. 13 is an example of a frustum view in a 4×4 hogel spatial resolution light field.
[0044] [Figure 6A] 4 shows the hogel camera translation in the 3D view.
[0045] [Figure 6B] 13 shows hogel camera translation in grid view.
[0046] [Figure 7A] 1 shows a light field camera and a forward vector of the light field camera at the display surface.
[0047] [Figure 7B] 1 shows a light field camera recessed an offset distance from the viewing surface.
[0048] [Figure 8] 1 shows the relationship between light field pixels at the display surface and light field pixels at an offset distance.
[0049] [Figure 9]1 shows a light field camera array positioned at a display surface, showing captured and non-captured volumes.
[0050] [Figure 10] We present a light field capture method in which the light field camera array is offset with respect to the entrainment plane, allowing the inner frustum volume to be captured.
[0051] [Figure 11] 1 illustrates one embodiment of the disclosed method with an offset parameter of N=1.
[0052] [Figure 12] 1 illustrates one embodiment of the disclosed method with an offset parameter of N=2.
[0053] [Figure 13] 1 shows indexed hogels and pixels in a light field display.
[0054] [Figure 14] FIG. 1 is a flow diagram of a light field rendering method.
[0055] [Figure 15] FIG. 1 is a flow diagram of a light field rendering method including a pixel remapping technique.
[0056] [Figure 16] FIG. 2 is a flow diagram of a light field rendering method process according to the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0057] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0058] The use of the words "a" or "an" when used herein in conjunction with the term "comprising" may mean "one," but is also consistent with the meaning of "one or more," "at least one," and "one or more."
[0059] As used herein, the terms "comprising," "having," "including," and "containing," and grammatical variations thereof, are inclusive or open ended and do not exclude further, unrecited elements and / or method steps. The term "consisting essentially of," when used herein in connection with a composition, device, article, system, use, or method, indicates that additional elements and / or method steps may be present, but that these additions do not materially affect aspects of the recited composition, device, article, system, method, or use functionality. A composition, device, article, system, use, or method described herein as including certain elements and / or steps may also, in certain embodiments, consist essentially of these elements and / or steps, and in other embodiments, consist of these elements and / or steps, regardless of whether these embodiments are specifically mentioned.
[0060] As used herein, the term "about" refers to about a + / - 10% variation from a given value. It should be understood that such a variation is always included in any given value provided herein, whether or not it is specifically referred to.
[0061] The recitation of ranges herein is intended to convey both the range and the individual values falling within the range, unless otherwise indicated herein, values in the same place as the numbers used to denote the range.
[0062] The use of any example or exemplary language, such as "such as," "exemplary embodiment," "illustrative embodiment," and "for example," is intended to illustrate or illustrate aspects, embodiments, variations, elements, or features related to the invention and is not intended to limit the scope of the invention.
[0063] As used herein, the terms "connect" and "connected" refer to any direct or indirect physical association between elements or features of the present disclosure. Thus, these terms may be understood to refer to elements or features that are partially or completely contained within, attached, coupled, positioned, joined together, in communication, operably associated, etc., even if there are other elements or features between the elements or features described as being connected.
[0064] As used herein, the term "pixel" refers to the light source and light emitting mechanism used to create the display.
[0065] As used herein, the term "light field" at a basic level refers to a function that describes the amount of light flowing in all directions through a point in space without occlusions. A light field contains information about all images that can be seen from all possible combinations of viewing positions and angles light flowing in all directions through a point in space without occluding objects for a particular display format. Thus, a light field represents radiance as a function of light position and direction in free space. Light fields can be synthetically generated by various rendering processes or captured from a light field camera or an array of light field cameras. In a broad sense, the term "light field" can be described as an array or subset of hogels.
[0066] As used herein, the term "hogel" is a holographic element, which is a cluster of traditional pixels with directional control. An array of hogels can generate a light field. Whereas pixels describe the spatial resolution of a two-dimensional display, hogels describe the spatial resolution of a three-dimensional display.
[0067] As used herein, the term "light field display" is a device that reconstructs a light field from a finite number of light field radiance samples input to the device. The radiance samples represent red, green, and blue (RGB) color components. For the purposes of reconstruction in a light field display, the light field can also be understood as a mapping from a four-dimensional space to a single RGB color. The four dimensions include the vertical and horizontal dimensions (x,y) of the display, and two dimensions that describe the directional components (u,v) of the light field. The light field is defined as a function, i.e., LF: (x,y,u,v) → (r,g,b) If fixed, x f ,y f ,LF(x f ,y f ,u,v) represent two-dimensional (2D) images called "element images". An element image is a fixed x f ,y f The component images are directional images of the light field from a position. When multiple component images are concatenated side-by-side, the resulting image is called an "integral image". An integral image can be understood as the entire light field required for a light field display.
[0068] As used herein, the term "display surface" refers to a set of points and directions defined by a planar display and the physical spacing of its individual light field hogel elements, as in a conventional 3D display. In an abstract mathematical sense, the light field may be defined and represented on any geometric surface and may not necessarily correspond to a physical display surface with actual physical energy emission capabilities. The inner and outer frustum volumes are light field regions that extend above and below, or behind and in front of, the display surface, respectively. The inner and outer frustum volumes may have different numbers of layers, may have different volumes, may have different depths, and may be rendered using different rendering techniques.
[0069] As used herein, the term "voxel" refers to a single sample or data point on a regularly spaced three-dimensional grid of single data. A voxel is an individual volume element that corresponds to a location in three-dimensional data space and has one or more data values associated with it.
[0070] As used herein, the term "scene description" refers to a geometric description of a three-dimensional scene that may be a potential source from which a light field image or video can be rendered. This geometric description may be represented by, but is not limited to, points, lines, quadrilaterals, textures, parametric surfaces, and polygons.
[0071] As used herein, the term "extra-pixel information" refers to information contained in a scene description, including, but not limited to, color, depth, surface coordinates, normals, material values, transparency values, and other possible scene information.
[0072] As used herein, the term "source image" refers to an image of the light field captured at a single position in the camera array.
[0073] As used herein, the term “element image” refers to LF(x f ,y f ,u,v) fixed position x f ,y f The element image represents a two-dimensional (2D) image of a fixed x f ,y f This is a directional image of the light field from the position.
[0074] As used herein, the term "integral image" refers to multiple elemental images connected side-by-side, and thus the resulting image is called an "integral image." An integral image can be understood as the entire light field required for a light field display. A rendered light field image is a rendering or mapping of the integral image that can be displayed on a display surface.
[0075] It is contemplated that any embodiment of the compositions, devices, articles, methods and uses disclosed herein can be implemented by one of ordinary skill in the art either as is or by making such modifications or equivalents without departing from the scope of the invention.
[0076] Various features of the present invention will become apparent from the following detailed description taken in conjunction with the illustrative drawings. The design factors, construction and use of the light field volume rendering techniques disclosed herein are described with reference to various examples that represent embodiments that are not intended to limit the scope of the invention described and claimed herein. Those skilled in the art to which the present invention pertains may appreciate that there may be other variations, examples and embodiments of the present invention not disclosed herein that can be implemented in accordance with the teachings of the present disclosure without departing from the scope of the present disclosure.
[0077] Described herein is a light field rendering method that uses a light field offset scheme and pixel remapping techniques to capture inner and outer frustum light field data from a 3D scene. In the method, a camera array is positioned at a drawing plane at an integral offset distance from the display surface away from the inner frustum volume, either inside or outside the outer frustum volume, to capture the scene. Capturing the scene with the camera array at the drawing plane allows for the capture of the entire inner frustum volume of the light field as well as some or all of the outer frustum volume. A pixel remapping function is used to transfer the resulting light field back to the display surface to display the light field as if the camera array was originally placed at the display surface. As a result, the light field captured by the camera array includes more of the 3D scene than would be captured if the camera array were placed at the display surface. In particular, the method reduces or eliminates blind volumes in both the inner and outer frustums, allowing for a more complete capture and rendering of the 3D scene. The captured 3D scene can then be presented on a 3D light field display.
[0078] The technique can be implemented in game engines such as Unreal Engine®, for example, and can be used to capture features of light field image data, such as advanced lighting and physically based rendering materials. The light field rendering method described herein also reduces light field rendering time and processing requirements compared to existing offline renderers, such as Octane®, by using rasterization techniques rather than comparatively computationally intensive ray tracing techniques. Overall, using a light field offset method to capture a dual frustum light field with rasterization is a viable technique for use in 3D scene rendering, particularly offline rendering. The placement and number of light field cameras in the entrained camera array can also provide a wide enough area to capture the entire intended frustum of the intended light field volume, thereby significantly reducing the size of the rendered 3D source data while providing a more complete rendering of the 3D scene.
[0079] This disclosure presents and elaborates the concept of using a drawing surface for light field rendering. The drawing surface function described here defines the geometry of the drawing surface, and the sample gap between the cameras in the camera array at the drawing surface specifies the light field camera spacing in the camera array at the drawing surface. The sample gap between the light field cameras can also be understood as the distance between the light field cameras located at the drawing surface, and the intermediate space between the cameras is occupied by a virtual hogel camera that is created with pixel remapping to generate an integral image of the 3D scene at the drawing surface.
[0080] Rendering is the process of converting a 3D scene into a 2D image. Apart from some artistic rendering techniques, the goal of rendering is usually to create photorealistic images that are indistinguishable from reality, a common theme in the field of computer graphics. Before a renderer (the program that performs the rendering) can convert a 3D scene into an image, the geometry must be specified in a format that can be interpreted by the computer. Geometry is represented using a combination of points in 3D space, each of which is defined as a vector with x, y, and z components. Most commonly, geometry is represented using triangles, which are the simplest way to define a plane and are used as the building blocks of all polygons. To generate a 2D image for a light field display from a 3D scene, the geometry in the scene must be "flattened" onto a canvas of a display surface that represents the screen. There are two common methods used to create 2D images from geometry: rasterization and ray tracing. Once a 2D image has been generated that accurately reflects how the eye sees the 3D scene, the appearance of the objects in the scene must also be recreated. The appearance of objects in a scene, such as their color, reflectance, and texture, depends on how light interacts with the materials the objects are made of. These interactions are modeled using mathematical models that are based on the laws of physics.
[0081] By using a drawing plane to capture the 3D scene at a distance offset from the display surface, both the inner and outer frustum volumes of the light field can be captured. The method results in less data to render the 3D scene into a 2D image, resulting in faster rendering, as well as less data to send to the display surface for image display, resulting in an overall smaller amount of data to decode. The light field is composed of an array of hogels, each hogel having a diameter that may be, for example, 0.1 to 25 mm. Capturing images for every hogel in a light field display results in a required range of 100 to 100,000 captured images from a physical camera. Using a light field camera array positioned in the sample gap of the drawing plane to capture the 3D scene at a distance offset from the display surface enables both full light field capture and efficient light field rendering for 3D display.
[0082] FIG. 1A is an illustration of one method of rasterization using frustum projection to flatten a 3D scene into a 2D image. As shown, a cube 12 representing a 3D scene 22 is flattened using frustum projection by tracing multiple rays 20 between each of the vertices 10 of objects in the 3D scene 22 and the observer, or the eye of the observer at the viewing position 16, and then finding the intersections of these rays 20 with a display surface 18 representing a display screen. The intersections on the display surface can then be connected to create a 2D image 14. If the cube 12 is opaque, the output will not accurately reflect how the 3D scene 22 is interpreted by the eye at the viewing position 16, since the front surface of the cube 12 hides all other faces and edges. Known as the "visibility problem," calculating which parts of the 3D scene 22 are visible is one of the fundamental concepts of light field rendering. One of the most common methods to solve visibility problems using rasterization is the Z-buffering algorithm, which is a depth buffer used to render opaque objects in a 3D scene to compensate for polygons that occlude other polygons. The Z-buffering algorithm, also known as the depth-buffering method, is an image-space method that references the pixel being rendered 2D and compares the surface depth at each pixel location on a projection plane, where the z-axis is expressed as depth. If multiple points or vertices 10 in the 3D scene geometry project to the same pixel in the display surface, they are added to a list called the depth buffer. The depth buffer is then sorted by increasing distance from the display surface 18. The first point in the depth buffer is closest to the display surface 18 and is visible through the corresponding pixel of the light field display or screen of the display surface 18.
[0083] Figure 1B shows a front view of a flattened 2D image 14 created by the intersection of each ray traced between the object's vertices at the viewing position from the display surface or screen and the observer's eye. In the projection example shown in Figure 1A, each point or vertex 10 of each object in the 3D scene 22 was projected onto a flat canvas representing the screen at the display surface 18. The process started with a known vertex 10 in the 3D scene 22 and found its location when projected onto the screen at the display surface 18. Performing this series of steps in reverse is the basis of ray tracing.
[0084] FIG. 2 is an illustration of an exemplary ray tracing technique for solving visibility problems. As shown, first, each pixel on the display surface 18 is converted to a point 24 on the display surface 18. Then, as shown, a ray 20 is traced from the eye of the observer at the viewing position 16 through the point 24 into the 3D scene 22. The point where the ray 20 intersects with the scene geometry 26 becomes the point that is seen through the pixel on the display surface 18. Like most algorithms that attempt to simulate the real world, the concept of ray tracing is simple to understand, but is resource and computationally intensive. In particular, ray tracing requires computing the intersection points 24 of many rays 20, each with a different type of geometry, and much research has been done to optimize this computation. Rasterization is a faster solution to solving visibility problems. In other words, to solve visibility, rasterization actually "projects" geometry onto the display surface 18 from a 3D representation onto a 2D representation of said geometry using perspective projection.
[0085] Typical rendering produces realistic 2D images from a 3D scene for viewing on a conventional 2D display. Similarly, at least one realistic 3D image from a 3D scene to create 3D content that can be viewed on a 3D light field display (LFD). Light field rendering produces a light field from a 3D scene that can be loaded and viewed on an LFD. Light field rendering utilizes the same rendering techniques as regular displays, applied differently to support multiple 3D views in a 3D scene. Based on the directionality of pixels in the LFD, utilizing ray tracing for light field rendering can be an organic technique, has been well-studied, and can be used for real-time rendering to enable interactive light field experiences such as video games. However, many existing software tools used by content creators do not support ray tracing to the degree necessary to implement high-quality light field rendering. Since integration with existing architectures is crucial for the development and adoption of new technologies, it is desirable to enable designers to move quickly into generating content for LFDs. Therefore, for applications where quality takes precedence over speed, such as offline renderers, light field rendering with rasterization can be a viable and good solution when ray tracing is not fully supported. One of the most widely used creation tools is Unreal Engine® 4 (UE4), which provides a platform for 3D creation in many technical fields, such as industrial and architectural visualization, games, visual effects, and filmmaking. UE4 supports ray tracing for visual effects such as reflection and refraction, but primary scene rendering is performed using rasterization. UE4's rendering pipeline can be modified to use ray tracing, but this destroys many of the advanced features that designers use to create high-quality content on the UE4 platform.To obtain real-time light field rendering, quality can be sacrificed for speed. For offline rendering where quality takes precedence over speed, an unobtrusive rasterization technique with a high level of interoperability with engine features may be preferred over a faster ray tracing technique with minimal support for advanced engine features. Therefore, a new method for rendering a dual frustum light field using frustum projection as an alternative to modified engines for exporting light fields offline is described herein.
[0086] FIG. 3 shows a ground truth, uncompressed, digital representation of a light field. A digital image for use on a 2D display is simply a grid of pixels, where the height and width of the grid are known as the image resolution. A similar representation is used in an LFD. The image displayed on the LFD, hereafter referred to as the light field, is represented by a hogel grid 40 containing a plurality of hogels 30, with the hogel width 38 (SR x ) and hogel height 36 (SR y ) determine the spatial resolution of the display. Each hogel 30 consists of a grid of directional pixels 28, which generate or emit light rays that convey information about the color, intensity, and direction of the light at that point on the hogel grid 40. The individual hogel widths 32 and hogel heights 34 of each hogel 30 in the hogel grid 40 determine the directional resolution of the LFD, and the image formed by each hogel is known as an "elementary image." Thus, each single hogel 30 represents a single elemental image. In the illustrated example, the spatial resolution of this LFD is 8x8 (hogels in the hogel grid 40) and the directional resolution is 4x4 (directional pixels in each hogel 30).
[0087] FIG. 4 is an illustration of an example of a method for rendering a light field. The illustrated method involves capturing a volume of the scene that can be seen from a surface representing a screen, referred to herein as the display surface 18. The scene volume is defined by two "frustums" or truncated pyramids as shown: an outer frustum volume 42 that is proximal or near the observer or human eye at the viewing position 16, and an inner frustum volume 44 that is distal or away from the observer or viewing position 16. Recall that perspective projection is used to "flatten" the 3D scene into a 2D light field that can be displayed on the viewing surface 18. This process is also known as frustum projection, since the frustum defines the volume of the 3D scene that is visible to the human eye at the viewing position 16. The inner frustum volume 44 extends from the viewing surface 18 into the scene, while the outer frustum volume 42 projects away from it. Objects in the inner frustum volume 44 appear to the eye at the viewing position 16 as being "inside" the display of the display surface 18, while objects in the outer frustum volume 42 appear to be projected from the display of the display surface 18. The distance the frustum moves in and out of the display surface 18 is defined by the far clip plane 48 and the near clip plane 46, respectively, and has practical limits. As the eyes of the observer at the viewing position 16 move in a horizontal direction parallel to the display surface 18, the eyes of the observer at the viewing position 16 perceive a shift in the light field to provide an accurate perspective of the light field from the new position. The shift in perspective, referred to herein as motion parallax, shows the natural relationship between elements in the light field and how that relationship changes as the viewing angle changes. Motion parallax mimics a physical scene and presents the light field as if the observer at the viewing position 16 were present in the physical scene. Enabling the eyes of an observer at viewing position 16 to naturally perceive the light field as if it were present in a physical world scene provides numerous applications for 3D light field displays where two-dimensional representations fall short.
[0088] Light field displays discretize light rays to approximate the light field, and the number of rays used, or directional resolution, and the density with which the rays are packed affect the depth of the display. The density of rays can be measured using the hogel pitch, which is the distance between the centers of two adjacent hogels. The maximum depth beyond which an object appears blurry can be calculated using the directional resolution (NxN), hogel pitch (HP), and field of view (F) as follows: JPEG2025515588000002.jpg17153
[0089] Rendering a light field requires capturing each source image from an individual view and stitching them together to form the entire light field. To do this, the packing algorithm must first take as input a set of LFD configuration and rendering parameters, then iteratively capture all views before generating a data structure containing all the information required to render all views. The light field resolution represents the total number of pixels in the light field and is calculated as the product of the directional resolution and the spatial resolution using the following formula: Light field resolution = directional X * directional Y * spatial X * spatial Y Where: directionalX is the directional resolution along the x-axis of the light field, direction Y is the directional resolution along the y-axis of the light field, Space X is the spatial resolution along the x-axis of the light field, Spatial Y is the spatial resolution along the y-axis of the light field.
[0090] If the light field resolution exceeds the maximum supported frame size, the light field can be split into subframes, with each subframe containing a portion of the total view. All views in a subframe are rendered and then exported to memory. Finally, the subframes are composited using a compositing function to produce the complete light field.
[0091] FIG. 5 is an example diagram of a frustum view of a light field with a spatial resolution of 4×4. For demonstration purposes, a front view of hogel grid 50A, a side view of hogel grid 50B, a top view of hogel grid 50C, and an isometric view of hogel grid 50D are shown, but it should be noted that each view is preferably captured consecutively, not parallel. The combination of all frustum views forms the light field frustum. The common method of rendering a light field using frustum projection at the display surface does not allow for capturing objects in the outer frustum. In particular, when the camera is placed at the display surface, any light field information behind the camera in the outer frustum volume cannot be captured by the camera. Traditionally, capturing light field data at both the inner and outer frustum volumes has been a tedious task. The method according to the present disclosure describes a light field shifting scheme, referred to herein as light field offset, to capture the light field at the drawing surface and then return the light field to the display surface after capturing all the frustum views. The method can be easily implemented by shifting the image capture to an offset distance from the display surface away from the inner frustum volume. This method allows for the capture of a 3D light field scene for the outer frustum volume as well as a more complete 3D light field scene capture for the inner frustum volume. Furthermore, performing this light field offset with the rendering does not require any modification of the engine.
[0092] FIG. 6A illustrates hogel camera translation in a 3D view, and FIG. 6B illustrates hogel camera translation in a grid view. As shown in FIGS. 6A and 6B, the packing algorithm defines a light field camera 52 at the center of the display surface 18 to aid in rendering calculations. Although only a single light field camera 52 is shown for illustrative purposes, it is understood that multiple light field cameras in a camera array can be used to capture the 3D scene in more detail and at multiple positions. For each light field camera view, the location of the associated hogel camera's viewport is defined by an x-direction offset 56 and a y-direction offset 58 from the light field camera 52. The light field camera 52 can be a physical camera or a computer-generated camera, and the hogel camera 54 can be understood as a virtual representation of the camera view with an x, y offset relative to the light field camera 52 on the camera surface. The view information also includes a location in the subframe for writing the rendered elemental image for display on the display surface of the LFD. A hogel camera 54 is defined and configured to represent the viewport of the elemental image at the hogel camera location. Using these offsets, the positions of the multiple hogel cameras 54 shifted in the x and y directions around the display surface 18 can be achieved, thereby rendering all views in a sub-frame to the sub-frame texture at the x,y positions specified in the rendering information. An isometric view of the x-direction offset 56 and y-direction offset 58 (also known as hogel camera translation) from the light field camera 52 to each hogel camera 54 on the display surface 18 is shown in Figure 6A. A front view of the x-direction offset 56 and y-direction offset 58 from the light field camera 52 to an exemplary hogel camera 54 on the display surface 18 for a light field having a spatial resolution of 4x4 is shown in Figure 6B.
[0093] 7A shows a light field camera 52 disposed on the display surface 18 and the light field camera's forward vector 60. When the light field camera 52 is positioned on the display surface 18, only light field information within the inner frustum volume 44 is received by the camera.
[0094] FIG. 7B shows the light field camera 52 retracted an offset distance from the display surface 18 and positioned at the retraction plane 64. Instead of capturing the scene from the display surface 18, the light field camera 52 can be retracted at a specific distance, or offset distance 62, from the display surface 18 to the retraction plane 64 to capture objects within the outer frustum volume 42. Note that the outer frustum volume 42 is shown as a block of space for illustrative purposes only. To apply this shift, the light field camera 52 is translated backward along the trajectory of its forward vector 60. In contrast to the camera translation method described in FIG. 6A / B, the light field camera 52 is retracted the offset distance 62 from the display surface 18 to the retraction plane 64 before one or more hogel cameras 54 are defined and configured. The hogel camera 54 is generated by applying a decoding scheme to the light field camera 52, which decodes a source image with the light field camera and offsets this decoded source image, including the component images, by a translational shift of the light field camera 52 relative to the hogel camera 54 to create a virtual or synthetic image at the location of the hogel camera 54. When the x and y offsets from the light field camera 52 are applied to the hogel camera 54, the hogel camera 54 is translated on the drawing plane 64 behind the display surface, which in this case is also at the near clip plane 46, as shown in FIG. 6A / B. In other words, in this embodiment, the near clip plane 46 is the drawing plane 64, but it should be noted that this does not have to be the case. The drawing plane 64 may also be located anywhere along the trajectory of the forward vector 60 within the outer frustum volume 42 of the light field display at an integral distance from the display surface 18.
[0095] After the light field camera 52 captures a frustum view from the entrainment plane 64 to generate the light field, a pixel remapping function is used to move the light field back from the entrainment plane 64 to the display surface 18. The offset distance D62 for a directional resolution NxN and for all integer values of the shift factor α can be calculated as follows: D=α*N*focal length
[0096] The same pixel remapping can be performed on all light field cameras 52 and all hogel cameras 54 to create an integrated light field image on the display surface 18. It can be shown that each pixel in the display surface can simply be offset an integer number of hogels to shift the light field towards or away from the camera.
[0097] It should be noted that the disclosed rendering method can be used for computer-generated light field images, which in this embodiment include one or more component images, thereby eliminating the decoding of the source image.
[0098] 8 illustrates the relationship between pixels of the light field at display surface 18 and pixels of the light field at attraction surface 64 at offset distance 62 (D). For simplicity, the illustration is two-dimensional (along the xz plane) and focuses on a single hogel 30 at display surface 18, but note that the relationship is the same in the zy plane and for all hogels. The relationship is expressed in terms of the pixel index (P o). A hogel view at the display surface 18 is shown by box 66, and a hogel view at an offset distance 62 (D) at the drawing plane is shown by box 68. Angle θ is defined as FOV / 2 (F / 2), where angle phi, Φ, is the angle between a ray corresponding to any pixel 70 of the hogel 30 perpendicular to the display surface 18 and a centerline 72 of the hogel 30. As shown, the F / 2 (angle theta, θ) of the outer frustum volume represented by box 66 combined with the F / 2 (angle theta, θ) of the inner frustum in box 68 equals the total field of view of the light field display. The ray represented by one pixel 70 is returned from the display surface 18 to the drawing plane 64, which is parallel to the display surface 18 and drawn in by the offset distance 62. Using similar triangular geometry, tangent functions, and the equation relating lens pitch to focal length, it can be shown that if the offset distance is defined by D=α*N*focal length, then each pixel 70 is shifted by an integer multiple of the lens pitch based on the pixel index. The lens pitch is the distance between adjacent hogels, so an offset of N multiplied by the lens pitch simply shifts pixel N hogels. The direction of the shift depends on whether the ray is in the first or second half of the hogel 30.
[0099] The offset distance D62 as a function of the focal length of the display and an integer N is derived by: 1. JPEG2025515588000003.jpg16153 2. JPEG2025515588000004.jpg19153 3. JPEG2025515588000005.jpg13153 Here, the lens pitch is defined as: JPEG2025515588000006.jpg13153Solving Equation 3 for D gives us:D=LP×d Substituting equation 2 for d, we get: JPEG2025515588000007.jpg18153 As a result, D=FL×N
[0100] Pixel indexing, a known step in image processing, describes how an image is indexed to identify every pixel in the image. x ,P y ) is a way for computer software to identify each pixel. Once an image is indexed, computer software can manipulate the image information, which changes the image. In one example, the computer software can render, transmit, store, change color, change transparency, etc. The indexing methodology for 2D images can be applied to light field images by providing indexes not only for pixels but also for hogels. The notation LF[H x ,H y ,P x ,P y ] is the 3D identity of each hogel and pixel in the light field image.
[0101] Dual frustum light field rendering using frustum projection and light field offset is implemented as a UE4 plugin. Light field frustums can be placed in the scene and configured as desired prior to rendering. When rendering is triggered, the light field is captured using the techniques described herein. After the light field is captured from the retraction position using frustum projection, the entire light field is saved in memory. A second light field buffer is then created to perform the remapping. This introduces a very large memory overhead that can be solved using a "scan" method of applying the remapping on the fly as opposed to after the image is fully captured.
[0102] Based on the relationship between the pixels of the light field at display plane 18 and the light field at drawing plane 64, a simple mapping can be applied to the captured light field to shift it towards the camera. The remapping algorithm iterates over each pixel in each hogel in the light field image buffer and calculates its pixel index P o Offset each pixel by "N" hogels based on . To offset a pixel by a full hogel, the spatial index is increased by 1. After calculating the offset, the pixel is stored in the remapped buffer at the determined target location. After all pixels are remapped, the remapped light field is exported.
[0103] One of the most important engine features that the disclosed rendering technique must support is lighting. Gaming engines such as UE4 have three general types of lighting: static lighting, dynamic lighting, and fixed lighting. Each type of lighting is used by designers in different scenarios, and all types should be fully supported. Physically Based Rendering (PBR) uses a program to reproduce the interaction of light with materials such as metal, glass, and mirrors. UE4 uses PBR materials extensively to create realistic scenes, and the disclosed rendering method properly supports PBR. To test this support with the methods and systems described herein, scenes with various PBR materials were rendered and all tested PBR materials were successfully captured. PBR material tests include roughness, metallicity, specularity, and translucency.
[0104] To further develop the disclosed rendering method, the light field camera can be decoded into an array of hogel cameras using multi-view depth image-based rendering. This technique can be described as rendering a synthetic image using multiple views. 3D image warping can further be used to synthesize a virtual view. Depth image-based rendering is a decoding method that converts an image with image information including color and depth into an index-based geometry and then performs a known 3D rendering method. Specifically, depth image-based rendering obtains a 3D scene and its associated depth information, as well as an index for each pixel in the scene. The depth information provides the depth of each pixel, and the index provides the position of the pixel in 3D space. Depth image-based rendering completes the rendering operation by projecting these points onto an image plane, which can be a drawing plane or a display plane, to generate an array of hogel cameras with the same optical properties as the hogels and a geometrically correct elemental image for each hogel camera. After decoding, each hogel has a hogel camera with an elemental image, and the array of hogel cameras collectively includes integral images for the array of hogels to emit. The integral image is now a light field image that can be displayed on a light field display.
[0105] The reprojection of the color+depth image is a decoding method that uses the light field cameras and their forward vector directions. The pixel's depth information and index, specifically the x and y indexes, can be used to calculate the pixel's coordinates in 3D space (x, y, z). This pixel in 3D space can be converted to the depth, x and y indexes of a single hogel and its associated elemental image using the hogel's origin and forward direction. The decoded x and y indexes can be used to write the pixel to an elemental image on an image plane, including but not limited to the drawing plane or the display plane. The reprojection color+depth image decoding method is repeated until all hogels are filled with the pixels written to each hogel's elemental image. Collectively, the array of hogel cameras comprises an integral image. After the array of cameras at the drawing plane is decoded, the integral image at the display plane can be constructed by reindexing the decoded integral image at the drawing plane. The integral image is a light field that can be displayed on a light field display.
[0106] As methods for decoding light field images improve, fewer light field cameras are needed to capture the scene that is decoded to generate the combined light field. One advantage of using fewer light field cameras is that less data is rendered and the data is transmitted at a faster rendering rate, resulting in a higher frames per second (fps) rate. Therefore, it is highly desirable to limit the number of light field cameras needed to capture a 3D description of a scene. In known light field image capture, the light field camera is aligned on the display surface. The light field camera can be a physical camera, such as a pinhole camera, a digital single-lens and mirrorless camera (DSLR), a plenoptic camera, a compact camera, a mirrorless camera, or a computer-generated camera. Each light field camera captures a view from a specific point in space and can generally only capture the volume in front of the light field.
[0107] FIG. 9 is a simplified diagram illustrating a light field capture technique having a light field camera array disposed on a display surface, showing light field volumes captured by the light field cameras and volumes not captured, referred to herein as "blind volumes" or "blind volume regions." As shown, an array of light field cameras 52a, 52b, 52c, 52d, 52e is positioned on the display surface 18, and frustum volumes 74a, 74b, 74c, 74d, 74e are captured by each of the light field cameras 52a, 52b, 52c, 52d, 52e, respectively, within the light field inner frustum volume 44. Although shown as a line of light field cameras 52a, 52b, 52c, 52d, 52e, it will be understood that multiple light field cameras organized in an array along the x,y plane are generally used to capture light fields. For simplicity, a current two-dimensional representation is shown here. Note that the size and shape of the frustum volume captured by each camera depends on the type of light field camera used to capture the light field source image, including its lens and optical properties. As shown herein, known light field image capture techniques do not capture images behind the display surface 18 within the light field outer frustum volume 42.
[0108] Physical cameras when used as light field cameras are further limited by the size and number required to generate a cohesive light field image. A light field is composed of an array of hogels, each hogel having a diameter that may be, for example, 0.1 to 25 mm. Capturing images for every hogel in a light field display results in a required range of 100 to 100,000 captured images from a physical camera. While it is possible to control a scene to capture this many images with multiple physical light field cameras, a controlled scene limits the type of content that can be captured. In particular, it is difficult to create an array of physical light field cameras that is small enough and patterned densely enough to capture a complete light field image of an uncontrolled scene without sparse sampling artifacts, which are areas in an image where a light field camera cannot capture an image. FIG. 9, for example, illustrates a case where there is limited or no overlap between frustum volume 74a of one light field camera 52a and frustum volume 74b of an adjacent light field camera 52b in a camera array, creating a blind volume region 76a where no image can be captured. In a camera array disposed on the display surface 18 as shown, multiple blind volume regions 76a, 76b, 76c, 76d exist in areas where there is no capture of the light field by any camera in the array. Although the figure is shown in cross section, it will be understood that the cameras in the camera array are typically regularly spaced across the 2D display surface 18, and thus the blind volumes span the entire display surface 18. The sparse sampling artifacts present as black areas on the display are a result of these blind volume regions as there is no image information to deliver to the pixels, causing them to appear black or remain off.
[0109] 10 illustrates a light field capture method that allows the entire inner frustum volume 44 to be captured by the light field camera array by offsetting the light field cameras 52a, 52b in the camera array from the display plane 18 to the entrainment plane 64, increasing the overlap of the individual camera frustum volumes 74a, 74b, respectively. Although two light field cameras 52a, 52b are shown for simplicity, it is understood that a camera array according to the present disclosure comprises multiple cameras spaced in a 2D array in the x, y plane of the camera array, which is the entrainment plane 64, as shown here. To increase the overlap of the individual camera frustum volumes 74a, 74b, the entrainment plane 64 is entrained at a vertical integral offset distance 62 from the display plane 18 within the light field outer frustum volume 42. The entrainment plane 64 can be, for example, outside or inside the near clip plane 46. At the entrapment plane 64 is an array of light field cameras positioned such that each light field camera has the same offset distance 62 relative to the display surface 18, and the light field cameras are positioned at sample gaps relative to each other. As previously mentioned, the number of light field cameras used to capture the 3D scene and render the light field can be reduced to the minimum number required based on sample gap calculations. However, it will be understood that camera arrays for producing complex, high quality light field displays will generally comprise three or more cameras.
[0110] The disclosed method of capturing the outer frustum volume 42 includes moving the light field cameras 52a, 52b to the entrainment plane 64 and rendering forward, as described herein. This method is in contrast to the less advantageous method of setting the near and far clip planes 46, 48 to negative values and placing the light field cameras 52a, 52b on the display surface 18. The latter method does not capture the complete outer frustum 42 volume unless all hogels are rendered, thus defeating the purpose of encoding. By entraining the light field cameras 52a, 52b to the entrainment plane 64, the entire scene can be captured in both the inner frustum volume 44 and the outer frustum volume 42. Compositing the captured images to the display surface modifies the images to appear as if they were taken at the display surface, allowing the images to present the inner frustum volume 42 and the outer frustum volume 44. Once composited, the focus of the scene is on the display surface 18, providing an immersive experience to the viewer.
[0111] The arrangement and number of light field cameras 52a, 52b on the entrainment plane 64, as well as the number of these cameras used for any particular image capture, are designed based on the system limitations and processing capabilities. As mentioned above, the number of light field cameras 52 used to capture a 3D scene can be reduced to the minimum number required based on sample gap calculations. The arrangement and number of light field cameras 52 in the camera array can provide a wide enough area to capture the entire intended frustum of the intended light field volume, i.e., the inner frustum volume 44 and / or the outer frustum volume 42. The objective of the disclosed rendering method is to calculate the required number of light field cameras 52 from the entrainment plane 64. In this embodiment, the frustum intended to be captured is the inner frustum volume 44. By entraining the light field cameras 52a, 52b into the entrainment plane 64, the individual camera frustum volumes 74a, 74b of the light field cameras 52a, 52b, respectively, can capture the entire inner frustum volume 44 with the blind volume 76 present only in the outer frustum volume 42. It should be noted that in this embodiment, two light field cameras 52a, 52b are shown, however, any number of light field cameras may be used and the light field cameras may be in any desired 2D orientation, such as a square, rectangular, hexagonal, triangular, or any other reasonable orientation.
[0112] In one example, in a three-dimensional embodiment, four light field camera planes may be on the drawing plane 64 with an integral offset distance of N=1 to capture the inner frustum volume 44 of the light field for the LFD. In another example with an offset distance of N=2, nine light field cameras 52a, 52b planes may be on the drawing plane 64 to capture the outer frustum volume 42 of the light field display. Traditional methods of capturing light field images with physical light field cameras require capturing source images ranging from 100 to 100,000. In comparison, this method can create a light field image of the physical world with four source images (using a 2×2 light field camera array). It should be noted that the 2×2 light field camera array represents only the inner frustum volume 44, which is still considered a valid light field image. Nine source images are required to capture the outer frustum volume 42 (i.e., a minimum of 3×3 light field cameras in the camera array). This reduction in the number of light field cameras required to capture a 3D scene further reduces the amount of data that needs to be transmitted and processed, increasing the feasibility of capturing physical images for light field displays. For comparison, using the disclosed method, which uses a 3x3 light field camera array to obtain 9 source images, reduces the rendered dataset compared to conventional capture methods with a minimum of 100 source images, thereby improving rendering efficiency of the same 3D scene by 91%.
[0113] The dimensions of the drawing plane 64 are derived from the spatial and directional resolution of the light field display and the selected integral offset parameter N. therefore, Retraction surface = [SR x +(DR x *(N-2)),SR y +(DR y *(N-2)),DR x ,DR y ] Where: S.R.x is the spatial resolution of the display on the x-axis, i.e., the number of hogels, S.R. y is the spatial resolution of the display on the y-axis, i.e., the number of hogels, DR x is the lateral resolution of each hogel in the x-axis, i.e., number of pixels, DR y is the directional resolution of each hogel in the y-axis, i.e., number of pixels, N is an integer offset from the display surface, i.e., the number of retracts.
[0114] The drawing plane 64 has an offset distance 62 from the display surface 18 equal to an integer multiple of the focal length multiplied by the directional resolution. Light field displays typically set the inner frustum volume 44 and the outer frustum volume 42 to be equal to the focal length of the light field camera capturing the light field, which is equal to the focal length of the display multiplied by the directional resolution. In this volume, the pixel pitch is equal to the hogel pitch that produces a high quality image, and the drawing plane 64 is on the near clip plane 46. As the inner frustum volume 44 and the outer frustum volume 42 increase to be larger than the focal length of the light field camera multiplied by the directional resolution, the sampling rate limits the image quality at the near clip plane 46 and the far clip plane 48. Light field displays do not need to fit a specific size to prove useful. If the inner frustum volume 44 and the outer frustum volume 42 of the LFD are not based on the focal length and directional resolution, the drawing plane 64 may not be located on the near clip plane 46.
[0115] In this method, the drawing plane 64 has an offset distance 62 from the display surface 18 that is determined by the characteristics of the light field display, which is designed to accommodate and optimize the optical properties of the light field cameras 52a, 52b. therefore, Offset distance = N(f*DR) Where: f is the focal length of the display, DR is the directional resolution of the display surface, N is an integer offset from the display surface.
[0116] Because the offset distance 62 is calculated from the focal lengths and directional resolutions of the light field cameras 52a, 52b, the directional resolution of the light field cameras will be the same as the directional resolution of the display surface 18. All light field cameras 52 in the camera array should have the same optical properties, including but not limited to, orientation lens pitch, directional resolution, and field of view, to allow for the offset distance 62 of the camera array of the drawing plane 64 relative to the camera array of the display surface while preserving all of the inner frustum volume 44 in the field of view of the light field cameras 52a, 52b.
[0117] At any particular distance, a ray traced from a camera pixel in the drawing plane 64 intersects with the display plane 18 such that the distance between pixels or pixel pitch at the display plane 18 is equal to the pixel pitch at the drawing plane 64. Thus, the pixels at the display plane 18 and the pixels at the drawing plane 64 have the same resolution and lens pitch, which is a necessary condition for implementing a pixel remapping function. The term "same resolution" refers to the directional resolution of the pixels at the display plane 18 in the LFD. The pixel remapping function is determined by the hogel index (H x , H y ) and the direction pixel index (P x ,P y) are not changed during remapping, the requirement that the pixels of the display surface 18 have the same directional resolution as the pixels of the drawing surface 64 is required in all circumstances. In addition, the requirement that the pixels of the display surface 18 have the same lens pitch as the pixels of the drawing surface 64 is required in all circumstances because the remapping function is performed by offsetting the drawing surface 64 by a certain distance (offset distance 62) whose pixel pitch matches the lens pitch of each pixel in both the display surface 18 and the drawing surface 64.
[0118] Challenges that arise when capturing and displaying physical light field images have previously prevented the use of captured light field images of the physical world on light field displays. In particular, capturing the inner frustum volume 44 and the outer frustum volume 42 from the real-world light field is very difficult due to sparse sampling artifacts caused by the lack of overlap of the light field camera frustum volumes. The present method and system solves this problem by offsetting the camera array to the entrainment plane 64 and using rendering, decoding, and reindexing methods to regenerate the light field image captured by the camera array and generate a light field image at the display surface 18.
[0119] FIG. 11 shows an embodiment of the disclosed method with an offset parameter of N=1. The drawing plane 64 may be offset from the display surface 18 by an offset distance 62 (N), where N is any integer. The illustrated embodiment shows the drawing plane 64 offset from the display surface 18 by N=1, which is the smallest possible offset. It is understood that for N=0, the drawing plane 64 is on the same plane as the display surface 18. At N=1, the drawing plane 64 is on the near clip plane 46 and the outer frustum volume 42 is determined by the focal length and directional resolution of the display. Otherwise, at N=1, the drawing plane 64 is not on the near clip plane 46. At additional offset distances of N=2, N=3, and N>1, the drawing plane 64 is outside the outer frustum volume 42 and the near clip plane 46.
[0120] Once the multiple light field cameras 52 in the camera array capture the light field at the drawing plane 64, the depth image based rendering completes the rendering operation by projecting the captured light field data onto the image plane. In the rendering method, multiple hogel cameras 54 with the same optical properties as the hogels with geometrically correct elemental images of each hogel camera 54 are generated based on the position of the hogels relative to the light field cameras 52 in the camera array. After decoding, the hogel cameras 54 with associated elemental images are generated, and the multiple elemental images generated by the multiple hogel cameras 54 together with the multiple elemental images generated by the multiple light field cameras 52 collectively comprise an integral image of the light field. The integral image is a light field image that can be displayed on a light field display. As shown, each hogel in the drawing plane is rendered with a 2D (x, y) geometric alignment to provide an elemental image of each hogel on the display plane emitting on the LFD.
[0121] The color+depth image reprojection is a decoding method using the light field camera 52 and its forward vector direction. Using the pixel's depth information and index, specifically the x-index and y-index, the coordinates of the pixel in 3D space (x, y, z) can be calculated. This pixel in 3D space can be converted to the depth, x-index, and y-index of a single hogel and its associated elementary image using the hogel's origin and forward direction. The decoded x-index and y-index can be used to write the pixel to an elementary image on an image plane, including but not limited to the drawing plane 64 or the display surface 18. The pixel remapping function consists of repeating the reprojection color+depth image decoding method until all hogels are filled with pixels written to each hogel's elementary image.
[0122] One solution to the physical limitations of conventional light field capture techniques is an additional advantage of the disclosed light field rendering method for decoding one or more light field cameras 52 by generating hogel cameras 54 in a hogel camera array in the same plane as the light field camera array and synthesizing the hogel camera array to generate a cohesive integral image. After capture, the 3D scene data can be processed to include image information including, but not limited to, color, transparency, and depth information. Pixel indexing, also referred to as pixel remapping, provides the location of each pixel that allows computer software to identify each individual pixel and the image information associated with each pixel. Identifying the pixel location and its image information allows computer software to modify and process the image so that pixel information in the hogel camera at the drawing plane 64 can be remapped to pixel information in the hogel at the display plane 18. The hogel cameras 54 can be generated by decoding methods such as depth image based rendering, color plus depth image reprojection, etc. to provide image information for every pixel in the LFD.
[0123] The decoding method, such as depth image based rendering, color plus depth image reprojection, etc., generates a hogel camera 54 for each single hogel, with each hogel camera having an associated elemental image related to the hogel position in the light field. In other words, the decoding generates elemental images from the source image captured by each light field camera. The decoding method can create small enough hogel cameras 54 that, when patterned densely enough, create integral images while providing a very dense integral image for display on the LFD. In accordance with this disclosure, the input to the decoding in this context is a sparse array of light field cameras 52. What makes the array sparse is the use of sample gaps, which provide the distance or spacing between the light field cameras 52 in the camera array required to obtain a light field image using interstitially generated or synthesized hogel cameras 54 located between the light field cameras 52 on the same x,y plane. In one example, to create an integral light field using a sparse array of light field cameras 52, the multiple light field cameras 52 are decoded as a full array of hogel cameras 54 in the same plane by geometric adjustment of the position of each hogel camera 54. In one example, the decoding process first decodes the sparsely spaced array of light field cameras vertically. After this step, the height of the integral image matches the height of the decoded integral image. The second step is to decode the integral image horizontally. After the second step, the integral image is a decoded integral image with the correct height and width. Utilizing sample gaps, fewer light field cameras 52 need to be decoded to generate the full array of hogel cameras 54 and provide a complete integral image at the camera plane.
[0124] The disclosed rendering method is described for simplicity with respect to a single light field camera 52 and a single hogel camera 54. Of course, it is understood that a light field camera 52 array including multiple light field cameras 52 will generate a hogel camera 54 array including multiple hogel cameras 54. The number of light field cameras 52 required to capture the 3D scene data is determined by a sample gap calculation defined by the encoding and decoding scheme. The sample gap determines the maximum distance between two light field cameras 52 in the camera array that is sufficient to provide light field data between the light field cameras 52 and generate an array of hogel cameras 54. therefore, Sample gap = l*DR Where: l is the lens pitch of the light field camera, DR is the directional resolution of the display surface.
[0125] After decoding, there is one hogel camera 54 for every hogel in the LFD, and the hogel cameras 54 share the same optical properties as those hogels. Each hogel camera 54 is generated to contain an elemental image, and its associated hogel is remapped to emit the elemental image in the LFD. A set of hogel cameras 54, together with the light field camera 52, generate an integral image, which consists of multiple elemental images. Each elemental image is associated with a single hogel, and a collection of hogels is needed to create a light field in the LFD. A hogel can be defined as a light engine that generates light. Each hogel has a field of view, a 2D resolution, an (x, y, z) position in space (origin), and a forward direction. After decoding, the hogel camera pitch, which is the distance between the centers of two adjacent hogel cameras 54 on the entrainment plane 64, is equal to the lens pitch of the light field display. The blind volume 76 exists only in the outer frustum volume 42. At this stage, the hogel cameras 54 are composited with those hogels that require decoding to provide elemental images. All hogel cameras 54 are aligned on a uniform grid, are coplanar, and have the same position and optical properties (including but not limited to orientation, lens pitch, directional resolution, and field of view) as their hogels. Once the hogel cameras 54 are composited with their hogels, the prerequisites for implementing the pixel remapping technique are met.
[0126] FIG. 12 illustrates an embodiment of the present disclosure in which the drawing plane 64 is calculated using an offset parameter of N=2 from the display surface 18. In this 3D embodiment, the frustum volume 74 is captured by nine light field cameras 52 arranged on an x,y grid at the drawing plane 64. Since FIG. 12 is a 2D view of a 3D embodiment, three light field 52 cameras are shown, but it is understood that multiple light field cameras 52 arranged in an array on the drawing plane 64 are required when capturing a light field image. The required number of light field cameras 52 is calculated to ensure sufficient capture of the entire outer frustum volume 42 and inner frustum volume 44 of the light field display. Using an offset parameter of N=2, the blind volume 76 lies outside the near clip plane 46, and the frustum volume 74 captured by the array of light field cameras at the drawing plane 64 is the complete inner frustum volume 44 and outer frustum volume 42 of the LFD. Thus, by setting the offset distance to N=2, both the inner frustum volume 44 and the outer frustum volume 42 can be completely captured without the blind volume 76. It is possible to use N=2 without N=1, so that the entire volume is captured in one layer. As shown in FIG. 11, following the decoding of the light field camera 52, hogel cameras 54 are generated whose hogel camera pitch on the entrapment plane 64 is equal to the lens pitch of the display. The hogel cameras 54 are composited with their hogels, thus fulfilling the prerequisites for implementing the pixel remapping technique. This composition during the decoding stage is performed to generate virtual images at each hogel camera 54 to generate elemental images, each elemental image being associated with a single hogel camera 54.
[0127] FIG. 13 illustrates the indexing of a single hogel 40 in an array of hogels 30 that comprise a pixel 28 in a light field display. In conventional rendering methods, a camera array is placed on the display surface of a light field display to capture an image of a 3D scene. According to the methods of the present disclosure, the light field camera array can be placed on a drawing plane at an offset distance perpendicular to the display surface. Pixel indexing is used to reference the individual pixels 28 and hogels 40 contained within the array of hogels 30 that comprise the light field. Once each pixel of the light field is indexed, computer software can manipulate the image information and thereby manipulate the light field. For example, computer software can render, transmit, store, change the color, change the transparency, etc. of the light field image data captured by the light field camera. The light field has spatial and directional resolution, and each of these display characteristics has an x-component and a y-component. Spatial resolution (SR) is a measure of the spatial resolution of a light field. x 38, S.R. y 36) refers to the number of hogels 30 in the hogel grid 40 used to construct the light field image. x ) 38 denotes the number of hogels 30 along the horizontal direction in the x-axis of the light field display. y ) 36 indicates the number of hogels 30 along the vertical in the y-axis of the light field display. x ,H y ) identifies the location of each hogel 30 within the spatial resolution of the light field display. The indexing provides an initial location for each pixel 28 and hogel 30, allowing computer software to process the image. The light field re-indexing method utilizes this indexing methodology to manipulate the image to present it in a plane parallel to the plane in which the image was captured. The directional resolution (DR x 32, D.R. y34) refers to the number of pixels 28 that comprise each single hogel 30. The directional resolution x (DR x ) 32 indicates the number of pixels 28 along the horizontal direction in the x-axis of the hogel 30. y ) 34 indicates the number of pixels 28 along the vertical in the y-axis of the hogel 30. x ,P y ) identifies the location of each pixel 28 within the directional resolution of that hogel 30.
[0128] Hogel Index (H x ,H y ) and pixel index (P x ,P y ) can be used to index each pixel 28 to indicate its position in the light field display. therefore, LF[H x ,H y ,P x ,P y ] Where: H x is the hogel index along the x-axis, with the leftmost column being 1, H y is the hogel index along the y-axis, with the top row being 1, P x is the pixel index along the x-axis, with the leftmost column being 1, P y is the pixel index along the y-axis, with the top row being 1.
[0129] To further explain pixel and hogel indexing, see Indexed Hogels 80. Identifying a particular pixel, for example pixel 78 within hogel 80, is LF[1,6,2,1], where H x =1, H y =6, P x =2, P y=1, which provides the initial locations of pixels 78 and hogels 80 to allow the computer software to process the image. The pixel remapping technique utilizes the index to manipulate the image to present it in a plane parallel to the plane in which it was captured.
[0130] While pixel indexing is a conventional approach, the pixel remapping technique described herein shifts the perceived location in the light field image while preserving motion parallax. The pixel remapping technique is essentially a hogel remapping technique since hogels are composed of pixels. However, known methods of shifting image planes utilize interpolation or subsampling which reduces image resolution, cannot preserve motion parallax, and increase computational complexity and bandwidth requirements for reading additional pixel data from memory. The present method for shifting light field image planes utilizes pixel remapping techniques to index each pixel into a drawing plane (LF). r ) and loads each pixel from the light field of the display surface (LF d ) light field, the pixel index (P x ,P y ) by the hogel index (H x ,H y ) offset. therefore, LF r [H x +(DR x *N)-(N*P x ),H y +(DR y *N)-(N*P y ),P x ,P y ]⇒LF d [H x ,H y ,P x ,P y ] Where: LF d is the light field at the display surface, LF ris the light field at the basin of attraction, H x is the hogel index along the x-axis, with the leftmost column being 1, H y is the hogel index along the y-axis, with the top row being 1, P x is the pixel index along the x-axis, with the leftmost column being 1, P y is the pixel index along the y-axis, with the top row being 1, DR x is the lateral resolution of each hogel in the x-axis, i.e., number of pixels, DR y is the directional resolution of each hogel in the y-axis, i.e., number of pixels, N is an offset parameter, ie, an integer indicating the number of pulls.
[0131] By applying the present pixel remapping technique, the computer software selects a pixel from the drag surface hogel 30 and maps its indexed location to the following equation LF r [H x +(DR x *N)-(N*P x ),H y +(DR y *N)-(N*P y ),P x ,P y], thus remapping the pixels to different hogels at the display plane. This method is computationally simple for computer software, which only needs to read one pixel from the drawing plane to generate one pixel at the display plane. This minimizes mathematical operations, only integers are used, and few source images are required. In comparison, other methods may require thousands of source images, which increases the amount of data to be captured, rendered, and / or transmitted. Other methods that use few source images may require floating point numbers, i.e., no integer values, increasing the computational complexity and therefore the time and hardware requirements to generate the light field. For example, a known method of moving the image plane is interpolation. To move the image plane by interpolation, each output pixel, e.g., a pixel at the display plane, is read from four source pixels, e.g., a pixel at the drawing plane. In comparison, in the disclosed light field rendering method, each display plane pixel is read from only one drawing plane pixel. Thus, the pixel remapping technique requires only ¼ of the data required for interpolation. It has been found that for commercially available memory devices such as random access memory (RAM), double data rate (DDR) memory, and synchronous dynamic random access memory (SDRAM), the described method can be executed in ¼ the time compared to interpolation methods, resulting in higher frames per second (fps) frame rates. This reduction in time and data allows the light fields to be transmitted and rendered in real time on commercially available systems while maintaining resolution and motion parallax, producing high quality light field images.
[0132] Preferably, the pixel remapping technique assigns pixels to their hogel index (H x ,H y ) from the drawing surface to the display surface. The pixel itself remains the pixel index (P x ,Py ) holds. Thus, the directional resolution of the light field display remains constant while the spatial resolution is changed. The spatial resolution is also changed, but this is due to the hogel index H x ,H y This is because reindexing requires peripheral (additional) hogels.
[0133] In the described method, the display surface hogel receives each pixel from a different drawing surface hogel. The drawing surface must be composed of enough hogels to provide enough pixels to achieve the required directional resolution at the display surface for the light field display. The size of the display surface is dictated by the directional and spatial resolution of the light field display. Display surface=[SR x ,SR y ,DR x ,DR y ] Where: S.R. x is the spatial resolution of the display on the x-axis, i.e., the number of hogels, S.R. y is the spatial resolution of the display on the y-axis, i.e., the number of hogels, DR x is the lateral resolution of each hogel in the x-axis, i.e., number of pixels, DR y is the lateral resolution of each hogel in the y-axis, i.e., number of pixels.
[0134] Following the decoding process, the entrainment plane is constructed from additional hogels, referred to herein as peripheral hogels, that contribute pixels to the display surface hogels. These peripheral hogels ensure that all display surface hogels have the required directional resolution. The number of peripheral hogels depends on the directional resolution of the display surface hogels and the offset parameter N. therefore, Peripheral hogel = [(DR x *N),(DR y *N)]
[0135] Furthermore, the peripheral hogels meet the requirement that all pixels in all hogels on the display surface are populated according to a pixel remapping technique. In order to create enough peripheral hogels to meet the requirements of a light field rendering method for creating high quality light field images, the area of the display surface, and subsequently the number of display surface hogels, must be taken into account when capturing and decoding the light field image at the entrapment surface.
[0136] Compositing is used to combine multiple remapped planes to create a seamless display that presents to both the inner and outer frustum volumes. Compositing is an exact implementation of the light field rendering method that incorporates transparency information when compositing multiple light fields into a single light field image at the display plane. Each remapped light field LF N Pixels from are composited (LF d ,LF N ) and blended with the transparency data to the display surface LF d are stored in a light field and displayed as a single light field image. therefore, Each LF N Regarding: For each hogel, For each pixel, LF d [H x ,H y ,P x ,P y ] =Synthesis(LF d [H x ,H y ,P x ,P y ],LF N [H x +DR x -P x ,H y +DR y -P y ,P x ,P y]) Where: LF d is the light field at the display surface, LF r is the light field at the basin of attraction, H x is the hogel index along the x-axis, with the leftmost column being 1, H y is the hogel index along the y-axis, with the top row being 1, P x is the pixel index along the x-axis, with the leftmost column being 1, P y is the pixel index along the y-axis, with the top row being 1, DR x is the lateral resolution of each hogel in the x-axis, i.e., number of pixels, DR y is the lateral resolution of each hogel in the y-axis, i.e., number of pixels. Here, the synthesis (LF d ,LF N ) returns either a remapped pixel or a blended pixel, depending on the type of transparency data. Each remapped light field is composited onto the display surface in ascending order of the offset parameter, i.e., N,value. If there is no transparency data, then compositing (LF d ,LF N ) writes the remapped pixels to the display surface as follows: Synthesis (LF d ,LF N ) → Return LF N
[0137] After compositing, the remapped pixels are stored on the display surface and are ready to be displayed. Without transparency data, when multiple light fields with different N values are composited onto the light field of a display surface, the last light field composited onto the display surface will be the only visible light field image. Transparency data incorporates various levels of transparency and opacity of each element in the image, such that when multiple light field images are composited together, opaque elements become visible when they are closer to the far clip plane than transparent elements. Transparency data is essential when multiple light fields are composited into a single light field image on a display surface to provide a 3D layered light field image. Commercially available physical light field cameras do not capture transparency data when capturing images, as the camera's sensor captures red, blue, and green color intensities. Source images from physical light field cameras can be processed using methods including, but not limited to, temporal median filtering, interactive foreground extraction, and the like to incorporate transparency data into the image before rendering. If transparency data is present, the compositing (LF d ,LF N ) returns the blend of the light field at the immersion plane and the light field at the display plane. Synthesis (LF d ,LF N ) → Return Blend (LF d ,LF N )
[0138] The two light fields are then combined and blended using the transparency data. Once blended, the pixels are stored on the display surface and are ready to be displayed. In an embodiment with multiple light fields, compositing loads the light fields in series starting with the light field closest to the far clip plane, the light field that originally had the lowest N value and shortest offset distance. Each rendered light field is composited onto the display surface until all light fields are composited onto a single light field at the display surface. Compositing with transparency data ensures that objects maintain their proper depth, transparency, and obscuration when composited into a single light field image, further improving the motion parallax of the light fields. This ensures that objects in the light field closest to the near clip plane are fully visible from all angles within the display's field of view, i.e., no object obscures their foreground, while objects far from the near clip plane are partially or fully covered, with the obscuration of each object depending on the obscuring objects between the object in question and the near clip plane and the transparency of the obscured object. For example, an object placed on the far clip plane may be fully visible if there are no obscuring objects between the object in question and the near clip plane, or if the obscuring objects are partially or fully transparent. Compositing combines multiple light fields from the drawing planes, incorporating transparency data, and storing them in a single light field at the display plane to create a light field image that contains an image of the physical world and presents it to both the inner and outer frustum volumes, while preserving motion parallax.
[0139] FIG. 14 is a flow diagram of an exemplary light field rendering method. Before the method begins, the light field camera array must meet a set of prerequisites. The first step is to set up or decode a planar array of light field cameras 100. The planar array of cameras can be, for example, an array of physical cameras, such as digital single-lens reflex (DSLR) cameras, pinhole cameras, plenoptic cameras, compact cameras, mirrorless cameras, or computer-generated or virtual cameras. Decoding methods such as depth image-based rendering, reprojection of color plus depth images, generate an array of virtual hogel cameras that display images in hogels where physical cameras do not exist or cannot exist due to size limitations. After decoding, the distance between the light field cameras on the drawing surface is equal to the sample gap, which is also the lens pitch of the display. The initial planar array of light field cameras is determined by the final display surface array of cameras to have its shape [SR x ,SR y ,DR x ,DR y ], then, in every hogel, the hogel camera is decoded to generate hogel cameras. x is the spatial resolution of the display on the x-axis, and SR y is the spatial resolution of the display on the y-axis, and DR x is the directional resolution of the display in the x-axis, and DR y is the directional resolution of the display in the y-axis. This sets the parameters of the final light field display after rendering has been performed. Step 102 is to determine whether the light field camera's encircling plane array is [SR x +(DR x *N),SR y +(DR y *N),DR x ,DR y ] is satisfied (102), in which case the offset parameter N indicates the number of retractions. This allows the retraction plane to be adjusted to the omnidirectional resolution (DR) required to create the integral image. x ,DR y) and spatial resolution (SR x ,SR y ) to the viewing surface. If the answer to step 102 is "no", the algorithm adjusts the cameras to match (104) and returns to step 102. If yes, the algorithm proceeds to step 106. Step 106 asks: Do all hogels in the drawing plane have cameras? Every hogel on the drawing plane requires a hogel camera because a camera provides an image to the associated hogel. This step ensures that every hogel on the drawing plane has a camera and therefore each hogel has an image. If no, the algorithm adjusts the cameras to match (108) and returns to step 106. If yes, the algorithm proceeds to step 110. Step 110 asks: Do all hogel cameras have the same position, orientation, lens pitch, directional resolution, and field of view as their associated hogels? Each hogel camera must have the same position as its associated hogel in order for the image the camera provides to be accurate for the hogel's position. The hogel cameras must then have the same optical property orientation, i.e., lens pitch, directional resolution, and field of view, as their associated hogels to ensure that the cameras have images that share the optical properties of that hogel. When the optical properties of the hogel cameras and their respective hogels match, the light rays from pixels in the drawing plane will intersect on the display surface, allowing the method to remap the perceived location of the integral image to the display surface. If no, the algorithm adjusts the cameras to match (112) and returns to step 110. If yes, the algorithm proceeds to step 114 in FIG. 15.
[0140] FIG. 15 is a flow diagram of an exemplary light field rendering method including a pixel remapping technique. Once all preconditions of steps 100-112 of FIG. 14 are met, the algorithm proceeds to begin 114 a pixel remapping technique with hogel[1,1]. The first hogel to be remapped can be any hogel in the display. In this embodiment, hogel[1,1] was used for simplicity. Once a hogel is selected, the pixel remapping technique begins 116 with pixel[1,1]. The first pixel in the hogel to be remapped can be any pixel. In this embodiment, pixel[1,1] was used for simplicity. First, the LF of the entrainment surface light field 118 is r [H x +(DR x *N)-(N*P x ),H y +(DR y *N)-(N*P y ),P x ,P y ], where H x ,H y indicates a particular hogel in the display, and P x ,P y denotes a particular pixel within the hogel. This step remaps each pixel from the drawing plane to the display plane. The boundary of the hogel (H x ,H y ) is the spatial resolution (SR) of the display x ,SR y In this embodiment, the boundary 1≦H x ≦SR x and 1≦H y ≦SR y is chosen for simplicity, but can vary depending on how the hogels are indexed, i.e., if the bounds are 0≦H x ≦SR x -1 and 0≦H y ≦SR y Can be -1. Pixel Bounds (P x ,P y ) is the directional resolution (DR) of the hogelx ,DR y In this embodiment, the boundary 1≦P x ≦DR x and 1≦P y ≦DR y is chosen for simplicity, but can vary depending on how the pixels are indexed, i.e., the bounds are 0 ≤ P x ≦DR x -1 and 0≦P y ≦DR y After the pixel is loaded from the drawing surface, it is added to the LF of the display surface light field 120. d [H x ,H y ,P x ,P y ]. Step 120 stores the remapped pixels on a display surface that allows them to be displayed as a light field image. One important advantage of this method is that the entire operation loads from one address, in this embodiment the drawing surface, and stores to another address, in this embodiment the display surface. Compared to known methods, the disclosed rendering method does not need to perform additional operations on the loaded data. Once loaded on the display surface, the data is an integral image and can be presented as a light field.
[0141] After each pixel has been stored in the display surface light field in step 120, the next step 122 asks: have all pixels in the hogel been remapped? If no, the algorithm repeats the light field reindexing method with the next pixel (124), which begins loading the next pixel in step 118. If yes, the algorithm proceeds to step 126. After all pixels in the hogel have been stored in the display surface light field, the next step checks whether all hogels on the display surface have been filled with remapped pixels (126). If no, the algorithm repeats the remapping technique with the next hogel (128), which begins loading the pixels in the next hogel in step 116. If yes, the algorithm proceeds to step 130. Once all pixels in all hogels have been remapped by loading and storing them in their respective display surface hogels, all conditions of the method are met and the light field display can display the rendered light field image (130).
[0142] FIG. 16 is a flow diagram of a light field rendering method process illustrating a computer-implemented method, the method process comprising: a. defining 140 a light field display having an inner frustum volume bounded by the viewing plane and a far clip plane, and an outer frustum volume bounded by the viewing plane and a near clip plane 140; b. positioning an array of light field cameras in a drawing plane at an offset distance from the display plane, where each light field camera has the same lens pitch, focal length, directional resolution, and field of view, and the offset distance is an integer multiple of the camera focal length multiplied by the directional resolution; c. capturing 144 the 3D scene in the entrainment plane using each light field camera as a set of source images, with one source image for each light field camera; d. decoding 146 the light field camera at the encirclement plane to generate a plurality of hogel cameras, each hogel camera having an associated hogel that includes an elemental image represented by an array of pixels; e. Step 148 of generating an integral image composed of elemental images at the basing plane; f. applying pixel remapping techniques to individual pixels in the integral image to create a light field image rendered on the display surface; Includes.
[0143] All publications, patents, and patent applications mentioned in this specification are indicative of the level of skill of those skilled in the art to which this invention pertains, and are hereby incorporated by reference. The reference to any prior art in this specification is not, and should not be construed as, an acknowledgment or any form of suggestion that such prior art forms part of the common general knowledge.
[0144] The invention thus described, it will be apparent that it may be modified in many ways. Such variations are not to be regarded as a departure from the scope of the invention, and all such modifications as would be apparent to one skilled in the art are intended to be included within the scope of the following claims.
Claims
1. defining a display surface for a light field, the light field including an inner frustum volume bounded by the display surface and a far clip plane, and an outer frustum volume bounded by the display surface and a near clip plane; defining a drawing plane parallel to and at an integral offset distance from the display surface, the drawing plane comprising a plurality of light field cameras spaced apart by a sample gap; capturing a view of the 3D scene as a source image with each of the plurality of light field cameras; decoding each source image to generate a plurality of hogel cameras on the entrapment plane, each hogel camera providing a component image; generating an integral image including a plurality of pixels from the elemental images in the drawing plane; performing a pixel remapping technique on individual pixels in the integral image to create a light field image rendered on the display surface; 1. A computer-implemented light field rendering method comprising:
2. The method of claim 1 , wherein the portion of the 3D scene captured at the drawing plane includes image information from the inner frustum volume and the outer frustum volume.
3. The method of claim 1 or 2, wherein the captured 3D scene includes all of the image information within the outer frustum volume.
4. The method of claim 1 , wherein the integral offset distance is calculated from the focal length, a directional resolution of the light field camera, and an offset integer N.
5. The method of claim 4 , wherein the offset integer N≧1.
6. The method according to claim 1 , wherein the lead-in surface is located on the proximal clip surface.
7. The method of claim 1 , further comprising displaying the rendered light field image on a light field display.
8. The method of claim 1 , wherein the surface area of the lead-in surface is greater than the surface area of the display surface.
9. The method of claim 1 , wherein the optical properties of the light field camera are orientation, lens pitch, directional resolution, and field of view.
10. The method of claim 1 , further comprising generating a plurality of integral images at a plurality of drawing planes.
11. The method of claim 10 , further comprising combining the integral images to create a combined rendered light field image on the display surface.
12. The method of claim 11 , wherein compositing incorporates transparency data.
13. 13. The method of claim 1, wherein each light field camera is one of a digital single lens reflex (DSLR) camera, a pinhole camera, a plenoptic camera, a compact camera, and a mirrorless camera.
14. The method of claim 1 , wherein each light field camera is a computer-generated camera.
15. The pixel remapping technique assigns the pixels their hogel index (H x , H y 15. The method of claim 1 , further comprising changing a surface of the display screen from the leading surface to the display surface.
16. The method of claim 1 , wherein the inlet surface is outside the outer frustum volume.
17. capturing a first light field at a drawing plane relative to a light field display surface using a light field camera, the first light field including an array of drawing plane hogels, each hogel having a plurality of pixels; A hogel index (H x , H y ) and pixel index (P x , P y ) to indicate its position within the light field display; Using the synthesis function, each pixel is r ) from the light field and each pixel is d ) in the light field; generating a light field image on the display surface comprising the remapped pixels; 16. A computer-implemented method for displaying a light field image, comprising:
18. The method of claim 17 , wherein one pixel from the drawing surface generates one pixel at the display surface.
19. Applying the pixel remapping technique involves recalculating the hogel index (H x , H y ) and change the pixel index (P x , P y 19. The method of claim 17 or 18, wherein:
20. 20. The method of claim 17, wherein the pixel remapping technique is a function of a directional resolution of the light field display and an offset parameter N.
21. The pixel remapping technique is based on the formula LF r [H x + (DR x *N) - (N*P x ), H y + (DR y *N) - (N*P y ), P x , P y ]⇒L.F. d [H x , H y , P x , P y 20. The method according to claim 19, wherein the method is based on
22. 21. The method of claim 20, wherein the offset integer N>1.
23. 23. The method of any one of claims 17 to 22, wherein the drawing surface is composed of a sufficient number of hogels to provide a number of pixels to achieve a required directional resolution of a light field display at the viewing surface.
24. 24. The method of claim 17, wherein the size of the display surface is defined by the directional and spatial resolution of the light field display.
25. 25. The method of claim 17, wherein the light field camera is a digital still camera, a pinhole camera, a plenoptic camera, a compact camera, or a mirrorless camera.