Codec for processing scenes with almost unlimited detail

The system addresses the limitations of traditional codecs by providing a flexible scene codec framework for efficient, continuous user experience across multiple clients with varying perspectives and data types, utilizing a plenoptic scene database and machine learning for optimal load balancing.

JP2026004310APending Publication Date: 2026-01-14QUIDIENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025146243
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-05-02
Filing Date
2025-09-03
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing codecs are limited by their tight coupling to specific data types, restricting user experience and struggling with efficient representation and distribution of complex real-world scenes across multiple clients with varying perspectives and data types.

Method used

A system utilizing scene codecs that provide multi-way, just-in-time, only-as-needed scene data through a plenoptic scene database, spatial processing unit, and machine learning for optimal load balancing and continuous user experience.

Benefits of technology

Enables efficient and flexible scene representation and distribution, supporting continuous user experience with minimal data transfer and reduced processing requirements, accommodating diverse client needs and data types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004310000001_ABST
    Figure 2026004310000001_ABST
Patent Text Reader

Abstract

A method and apparatus for a system using a scene codec.SOLUTION: The system is either a provider or a consumer of scene data, including sub-scenes and sub-scene increments, in a multi-way, just-in-time, as-needed basis. An exemplary system using a scene codec comprises a plenoptic scene database containing one or more digital models of a scene, wherein the representation and the organization of the representation are distributable across a plurality of systems such that together the plurality of systems can represent the scene in almost unlimited detail. The system further comprises highly efficient means for processing these representations and the organization of the representations to provide just-in-time, sub-scenes only when needed, and scene increments necessary to ensure a maximally continuous user experience enabled by a minimal amount of newly provided scene information, wherein these highly efficient means include a spatial processing unit.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 665,806, filed May 2, 2018, the entire contents of which are incorporated herein by reference. This application is related to International Patent Application No. PCT / US2017 / 026994, filed April 11, 2017, the entire contents of which are incorporated herein by reference.

[0002] This disclosure relates to scene representation, processing, and acceleration in distributed digital networks. [Background technology]

[0003] Various codecs are well known in the art and are generally devices or programs that compress data to enable high-speed transmission and decompress received data. Typical types of codecs include video (e.g., MPEG, H.264), audio (e.g., MP3, ACC), image (e.g., JPEG, PNG), and data (e.g., PKZIP), with each codec encapsulating a type of data and being tightly coupled to that type of data. While these types of codecs are sufficient for applications that are limited to the type of data, inherent in the tight coupling is a limited end-user experience.

[0004] Codecs are inherently "file-based"; a file is a data representation of some kind of real or synthetic pre-captured sensory experience, and a file (such as a movie, song, or book) necessarily limits the user's experience to the experiential path chosen by the file's creator. Thus, we watch a movie, listen to a song, or read a book within a substantially sequenced experience limited by the creator.

[0005] Technological advances in the marketplace are expanding the types of data and the means for experiencing them. This expansion in data types includes what is often referred to as real-world scene reconstruction, in which sensors such as cameras and distance-measuring devices create scene models of real-world scenes. The inventors of the present invention proposed significant advances in scene reconstruction in International Patent Application No. PCT / 2017 / 026994, entitled "Quotidian Scene Reconstruction Engine," filed April 11, 2017, the entire contents of which are incorporated herein by reference. Improvements in the means for experiencing types of data include higher resolution and better performance 2D and 3D displays, autostereoscopic displays, holographic displays, and extended reality devices and methods, such as virtual reality (VR) and augmented reality (AR) headsets. Other significant technological advances include the proliferation of automata, in which humans are no longer the sole consumers of real-world sensory information, and the proliferation of networks, in which the flow of and access to information is enabling new experiential paradigms.

[0006] Some research is being done to develop new scene-based codecs, where the type of data is reconstruction of real-world scenes and / or computer-generated synthetic scenes. For evaluations of scene codecs, the reader is referred to technical reports such as those published by the Joint ad hoc group for digital representations of light / sound fields for immersive media applications, the entire contents of which are incorporated herein by reference.

[0007] Scene reconstruction and distribution are problematic: reconstruction is difficult in terms of creating representations and organizing those representations that fully describe the complexity of real-world matter and light fields in an efficiently controllable and scalable manner; distribution is difficult in terms of managing active, even live, scene models across a large number of interactive clients, including humans and automata, each potentially requiring any of a virtually unlimited number of scene perspectives, details, and data types.

[0008] Therefore, there is a need to overcome the shortcomings and deficiencies of this technology by providing an efficient and flexible system that can address the many needs and opportunities in the marketplace. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG1N72033 [Non-patent document 2] ISO / IEC JTC1 / SC29 / WG11N16352 Summary of the Invention [Problem to be solved by the invention]

[0010] The following simplified summary is intended to provide a first basic understanding of some aspects of the systems and / or methods described herein. This summary is not an extensive overview of the systems and / or methods described herein. It is not intended to identify all key / critical elements or to delineate the entire scope of such systems and / or methods. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.

[0011] Provided herein are methods and apparatus for supporting systems that use scene codecs, which are either providers or consumers of multi-way, just-in-time, only-as-needed scene data, including subscenes and subscene increments. According to some embodiments, systems that use scene codecs include a plenoptic scene database containing one or more digital models of a scene, and representations and representation configurations are distributable across multiple systems so that collectively the systems can represent scenes with nearly unlimited detail. The systems further include efficient means for processing these representations and representation organization, which provide the just-in-time, only-as-needed subscenes and scene increments necessary to ensure the most continuous user experience possible with a minimum amount of newly provided scene information; these efficient means include a spatial processing unit. [Means for solving the problem]

[0012] Systems according to some embodiments may further comprise application software that performs both executive system functions as well as user interface functions, including any combination of providing a user interface or communicating with an external user interface. The user interface determines explicit and implicit user instructions that are used at least in part to determine user requests for scene data (and other related scene data), and provides the user with either scene data and other scene data responsive to the user's requests.

[0013] The system according to some embodiments may further comprise a scene codec, which may include an encoder and / or decoder, thereby enabling the system to be a scene data provider and / or consumer. The system may optionally interface with or comprise any of the available sensors for sensing real-world actual scene data, any such sensed data being available for reconstruction by the system into entirely new scenes or increments to existing scenes, any one system sensing the data may reconstruct the data into scene information or offload the data to other systems for scene reconstruction, and the other systems performing scene reconstruction may return reconstructed sub-scenes and scene increments to the system that originally sensed it.

[0014] Codecs according to some embodiments support scene models and other types of non-scene data that are either integrated with the scene model or maintained in association with the scene model. Codecs according to some embodiments support multiple system networks. Network connections may be supported to exchange control packets containing user requests, client status, and scene usage data, as well as scene data packets containing requested scene and non-scene data, and optional request identification for client use in performance verification. Support may be provided for one-to-one, one-to-many, and many-to-many system network connections, where again, any system may be able to sense new scene data, reconstruct new scene data, provide scene data, and consume scene data.

[0015] Systems according to some embodiments support the use of machine learning in both reconstruction and delivery of scene data, with key data logging of new types of information providing the basis for machine learning or deterministic algorithms that optimize both individual and networked system performance. For example, the state of all client systems consuming scene data is tracked to ensure that any potential serving systems have valuable prior knowledge of the client's existing scene and non-scene data. User requests, including scene types and scene instances, are categorized and uniquely identified. Individual systems are identified and classified according to their capabilities for scene sensing, scene reconstruction, scene provision, and scene consumption. The extent of scene usage, including usage type, as well as scene consumption path and scene duration, are tracked. The wide variety of categorized and tracked information provides valuable new data for machine learning, and user requests for scene data are intelligently enhanced with look-ahead predictions based on cumulative learning, further ensuring the most continuous user experience possible with a minimal amount of newly provided scene information.

[0016] These and other features and advantages will be better and more completely understood by reference to the following detailed description of exemplary, non-limiting, illustrated embodiments taken in conjunction with the following drawings. [Brief explanation of the drawings]

[0017] [Figure 1A] 1 is a block diagram illustrating a system using a scene codec according to some implementations. [Figure 1B] FIG. 1 is a block diagram illustrating a scene codec including both an encoder and a decoder, according to some embodiments. [Figure 1C] FIG. 2 is a block diagram illustrating a scene codec that includes an encoder but does not include a decoder, according to some embodiments. [Figure 1D]FIG. 2 is a block diagram illustrating a scene codec that includes a decoder but does not include an encoder, according to some embodiments. [Figure 1E] FIG. 1 is a block diagram illustrating a network connecting two or more systems that use a scene codec, according to some implementations. [Figure 1F] 1 is a block diagram illustrating a scene codec including an encoder according to some embodiments. [Figure 1G] FIG. 2 is a block diagram illustrating a scene codec including a decoder according to some embodiments. [Figure 2A] Block diagram of the existing state-of-the-art "Light / Sound Field Conceptual Workflow" as described by the Joint Ad Hoc Group for Digital Representation of Light / Sound Fields for Immersive Media Applications in technical publications ISO / IEC JTC1 / SC29 / WG1N72033, ISO / IEC JTC1 / SC29 / WG11N16352, dated June 2016, issued from Geneva, Switzerland. [Figure 2B] FIG. 1C is a combined block and pictorial diagram of a real-world scene captured by a representative real camera and provided to a system such as the system shown in FIG. 1 in a networked environment as shown in FIG. 1E, according to some embodiments. [Figure 3] 1 is a pictorial diagram of an exemplary network connecting systems that use scene codecs, according to some embodiments. [Figure 4A] 1 is a pictorial diagram of an exemplary real-world scene of unlimited or nearly unlimited detail, such as an interior house scene with a window that views an outdoor scene, according to some embodiments. [Figure 4B] 4B is a pictorial representation of a real-world scene as shown in FIG. 4A , in accordance with some exemplary embodiments, where the representation can be thought of as an abstract model view of the data contained within the plenoptic scene database, as well as other objects such as described and undescribed objects. [Figure 4C] FIG. 1 is a block diagram of some data sets in a plenoptic scene database, in accordance with some example embodiments. [Figure 5] 1 is a flow diagram of a use case involving sharing a larger global scene model with remote clients consuming any of various types of scene model information, according to some embodiments. [Figure 6] 6 is a use case flow diagram similar to FIG. 5, but handling the different case where a client is initially creating a scene model or updating an existing scene model, according to some embodiments. [Figure 7] A use case flow diagram similar to Figures 5 and 6, according to some embodiments, but dealing with the different case where a client system first creates a scene model or updates an existing scene model, and then both the client-side system and the server-side system each reconstruct and distribute the real scene, thus capturing local scene data of the real scene from which sub-scenes and increments to the sub-scenes can be determined and provided. [Figure 8] A composite plenoptic scene model, a synthetically generated image of an ordinary kitchen. [Figure 9] 1 is a geometric diagram illustrating two views of a volume element ("voxel") and a solid angle element ("seil"), according to some embodiments. [Figure 10] FIG. 1 is a geometric diagram showing an overhead plan view of a scene model of a trivial scene, according to some embodiments. [Figure 11] FIG. 2 is a block diagram of a scene database, according to some embodiments. [Figure 12] FIG. 1 is a class diagram illustrating a hierarchy of primitive types used in representing plenoptic fields, according to some embodiments. [Figure 13] A composite plenoptic scene model, a synthetically generated image of a mundane kitchen with two points highlighted. [Figure 14] 14A-14C are diagrams including images from the exterior showing a light cube with incident light entering a point in an open space in the kitchen shown in FIG. 13 according to some embodiments. [Figure 15] 15A-15C are diagrams including six additional views of the light cube shown in FIG. 14 according to some embodiments. [Figure 16] 15 is an image of the light cube shown in FIG. 14 from an internal viewpoint, according to some embodiments. [Figure 17] 15 is an image of the light cube shown in FIG. 14 from an internal viewpoint, according to some embodiments. [Figure 18] 15 is an image of the light cube shown in FIG. 14 from an internal viewpoint, according to some embodiments. [Figure 19] 14 is an image showing the outside of a light cube for light emitted from a point on the surface of the kitchen counter shown in FIG. 13, according to some embodiments. [Figure 20] 1 is an image of a light cube showing the results of BLIF applied to a single incident beam of vertical polarization, according to some embodiments. [Figure 21] FIG. 2 illustrates an octree tree structure according to some embodiments. [Figure 22] FIG. 22 is a geometric diagram illustrating the volume space represented by the nodes of the octree shown in FIG. 21 in accordance with some embodiments. [Figure 23] FIG. 2 illustrates a tree structure of a saeltree, according to some embodiments. [Figure 24] FIG. 24 is a geometric diagram illustrating regions of direction space represented by nodes of the Seyer tree shown in FIG. 23 in accordance with some embodiments. [Figure 25] FIG. 1 is a geometric diagram illustrating a Seyle with the origin at the center of an octree node, according to some embodiments. [Figure 26] FIG. 2 is a geometric diagram illustrating the space represented by the three Seyer trees of a 2D Seyer tree, according to some embodiments. [Figure 27] FIG. 1 is a geometric diagram illustrating two outgoing Seyers of two Seyer trees and the intersection of the two Seyers with two volume octree (VLO) voxels, according to some embodiments. [Figure 28] FIG. 1 is a geometric diagram illustrating two incoming Seyerls of a new Seyerl tree attached to one VLO voxel resulting from two outgoing Seyerls from two Seyerl trees projecting onto a VLO node, in accordance with some embodiments. [Figure 29] FIG. 10 is a geometric diagram illustrating an outgoing Seyel from a new Seyel tree generated for a VLO voxel based on the voxel's incoming Seyel tree and the BLIF associated with the voxel, in accordance with some embodiments. [Figure 30] FIG. 2 is a schematic diagram illustrating the functionality of a spatial processing unit (SPU) according to some embodiments. [Figure 31] FIG. 10 is a schematic diagram illustrating sub-functions of the light field operation functionality of the spatial processing unit, according to some embodiments. [Figure 32] FIG. 10 is a geometric diagram illustrating the numbering of the six faces of the perimeter cube of a Seyer's tree, according to some embodiments. [Figure 33] FIG. 1 is a geometric diagram illustrating the quarter faces of a perimeter cube in which the quarter faces of the perimeter cube face of a Seyel tree are highlighted, according to some embodiments. [Figure 34] FIG. 1 is a geometric diagram showing a side view of a quarter face of a perimeter cube of a Seyel tree, according to some embodiments. [Figure 35] FIG. 1 is a geometric diagram illustrating a 2D side view of a segment of direction space represented by a top-order sphere, according to some embodiments. [Figure 36] FIG. 1 is a geometric diagram illustrating how the projection of a Seyer tree onto a projection plane is represented by its intersection points in their placement on the faces of the Seyer tree's surrounding cube, according to some embodiments. [Figure 37] 1A-1C are geometric diagrams illustrating in 2D the movement of a Seyer tree while maintaining the projection of the Seyer tree on the projection surface, according to some embodiments. [Figure 38] 10A-10C are geometric diagrams illustrating the movement of the projection plane while maintaining the Seyer projection of the Seyer tree, according to some embodiments. [Figure 39] FIG. 1 is a geometric diagram illustrating the geometry of a Seyel across a projection surface, according to some embodiments. [Figure 40] FIG. 10 is a geometric diagram illustrating the relationship between the intersection of the top and bottom and the situation where the top is below the bottom in the projection plane coordinate system, indicating that the projection is not valid (opposite the origin from Seyhel), according to some embodiments. [Figure 41] 1 is a geometric diagram illustrating a subdivision of a Seyer tree into two subtree levels and spatial regions represented by nodes in 2D, according to some embodiments. FIG. [Figure 42] 10A-10C are geometric diagrams illustrating the placement of the intersection of the new seam edge with the projection surface that results in the new top or bottom edge of the sub-sea, according to some embodiments. [Figure 43] FIG. 1 is a geometric diagram illustrating an outgoing Seyel from a Seyel tree that causes the generation of an incoming Seyel in the Seyel tree attached to the VLO node where the outgoing Seyel intersects, according to some embodiments. [Figure 44] 1 is a geometric diagram illustrating a VLO traversal sequence from front to back in 2D within the illustrated range of direction space, according to some embodiments. [Figure 45] 1 is a geometric diagram illustrating in 2D the use of a quadtree as a projection mask in Seymour projection into a scene, according to some embodiments. [Figure 46] FIG. 1 is a geometric diagram illustrating the construction of Seyer's volume space by the intersection of multiple half-spaces, according to some embodiments. [Figure 47] 1 is a geometric diagram illustrating three intersection situations between a Seyer and three VLO nodes, according to some embodiments. [Figure 48]1 is a geometric diagram illustrating a Seyer rotation as part of a Seyer tree rotation, according to some embodiments. [Figure 49] FIG. 1 is a geometric diagram illustrating the geometric configuration of center points between Seyer edges with a projection plane during rotation of a Seyer tree, according to some embodiments. [Figure 50] 1 is a geometric diagram illustrating the geometric operations for performing edge rotations when rotating a Seyer tree, according to some embodiments. FIG. [Figure 51] 10A-10C are schematic diagrams illustrating the calculation of the geometric relationship between the Seyle origin node, the VLO node of the projection plane, or the Seyle and the projection plane when the Seyle is PUSHed, according to some embodiments. [Figure 52] 1 is a schematic diagram illustrating the calculation of the geometric relationship between the Seyle origin node, the VLO node of the projection plane, or the Seyle and the projection plane when the Seyle is pushed, in accordance with some embodiments, where the Seyle origin node and the VLO node of the projection plane are pushed simultaneously. [Figure 53] 1 is a table showing a portion of a spreadsheet tabulating a series of Seyel tree origins, VLOs, and Seyel PUSH series results, according to some embodiments. [Figure 54] This is a continuation of Figure 53. [Figure 55] 55 is a table showing calculation formulas for the spreadsheets of FIGS. 53 and 54. [Figure 56] FIG. 55 is a geometric diagram showing the starting geometric relationships at the beginning of the sequence of PUSH operations tabulated in FIGS. 53 and 54. [Figure 57] FIG. 54 is a geometric diagram showing the geometric relationship between the Seyel and its projection on the projection plane after iteration #1 shown in the spreadsheets of FIGS. 53 and 54 (PUSH of Seyel origin SLT node to child 3), according to some embodiments. [Figure 58]FIG. 54 is a geometric diagram showing the geometric relationship between the Seyel and its projection on the projection plane after iteration #2 shown in the spreadsheets of FIGS. 53 and 54 (PUSH of Seyel origin SLT node to child 2), according to some embodiments. [Figure 59] A geometric diagram showing the geometric relationship between the Seyel and its projection on the projection plane after iteration #3 shown in the spreadsheets of Figures 53 and 54 (PUSH of projection plane VLO node to child 3), according to some embodiments. [Figure 60] A geometric diagram showing the geometric relationship between the Seyel and its projection on the projection plane after iteration #4 shown in the spreadsheets of Figures 53 and 54 (PUSH of projection plane VLO node to child 1), according to some embodiments. [Figure 61] FIG. 54 is a geometric diagram showing the geometric relationship between Seyel and its projection on the projection plane after iteration #5 shown in the spreadsheets of FIGS. 53 and 54 (PUSH of Seyel to Child 1), according to some embodiments. [Figure 62] FIG. 54 is a geometric diagram showing the geometric relationship between Seyel and its projection on the projection plane after iteration #6 shown in the spreadsheets of FIGS. 53 and 54 (PUSH of Seyel to Child 2), according to some embodiments. [Figure 63] FIG. 54 is a geometric diagram illustrating the geometric relationship between the Seyel and its projection on the projection plane after iteration #7 shown in the spreadsheets of FIGS. 53 and 54 (PUSH of projection plane VLO node to child 0), according to some embodiments. [Figure 64] 10 is a table showing a portion of a spreadsheet tabulating the results of a series of Seyer tree origins, VLOs, and a series of Seyer PUSHs when the Seyer tree origin is not at the center of an octree node, according to some embodiments. [Figure 65] This is a continuation of Figure 64. [Figure 66] FIG. 2 is a schematic diagram illustrating application programming interface functionality of a scene codec according to some embodiments. [Figure 67]FIG. 1 is a schematic diagram illustrating the functionality of a query processor function, according to some embodiments. [Figure 68A] 1 is a flowchart of a procedure used to implement a plenoptic projection engine, according to some embodiments. [Figure 68B] 1 is a flowchart of a procedure used to extract sub-scenes from a plenoptic octree for remote transmission, according to some embodiments. [Figure 69] 1 is a flow diagram of a process for extracting sub-scene models from a scene database for the purpose of generating images from multiple viewpoints, according to some embodiments. [Figure 70] 1 is a flow diagram of a process for accumulating plenoptic primitives that provide light contributions to a query string, according to some embodiments. [Figure 71] 1 is a flow diagram of a process for accumulating a media element (“mediel”) and its contributing light field elements (“radiiel”) that provide light contributions to a query ellipses, according to some embodiments. [Figure 72] 1 is an image of a kitchen with a small rectangular area highlighting an analytics portal, according to some embodiments. [Figure 73] 73 is an image of a portion of the kitchen of FIG. 72 enlarged with the rectangular window of FIG. 72 highlighting the analysis portal, according to some embodiments. [Figure 74] 74 is an image of the rectangular area shown in FIG. 73 enlarged to show the analytical elements displayed in the analytical portal of FIG. 73, according to some embodiments. [Figure 75] 1 is a pictorial diagram relating to evidence of effectiveness of an embodiment. [Figure 76] 1 is a pictorial diagram relating to evidence of effectiveness of an embodiment. [Figure 77] 1 is a pictorial diagram relating to evidence of effectiveness of an embodiment. [Figure 78] FIG. 1 is a pictorial diagram showing sub-scene extraction for image generation purposes. DETAILED DESCRIPTION OF THE INVENTION

[0018] In the following description, numerous specific details are set forth, such as specific example components, types of usage scenarios, etc., to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that the present disclosure may be practiced without some of these specific details, or through alternative implementations, some of which are also described herein. In other instances, well-known components or methods have not been shown in detail in order to avoid unnecessarily obscuring the disclosure. Thus, the specific details set forth are for illustrative purposes only. It is contemplated that the specific details may vary and still be within the spirit and scope of the present disclosure.

[0019] A comprehensive solution is realized for providing variably scalable scene representations such as sub-scenes and increments to sub-scenes that both meet complex user requirements using minimal scene data, while "look-ahead" predicting sufficient buffers (extensions to requested scene data) to ensure continuous quality of service. The codec according to an exemplary embodiment handles continuous scene reconstruction commensurate with continuous scene consumption, where multiple entities are providing or consuming scene data at any point in time, and providing scene data includes both reconstructed scene data and unreconstructed data of a newly determined real scene.

[0020] In some exemplary embodiments, scene delivery is less "file-based" (i.e., less focused on a one-to-one unidirectional pipeline of entire scene information) and more "file-segment-based" (i.e., more focused on a many-to-many bidirectional pipeline of just-in-time, only-needed sub-scene and sub-scene increment information). This multidirectional configuration in some exemplary embodiments is self-learning, tracking the provision and consumption of scene data to potentially more scene servers and scene servers. The goal of scene processing in an exemplary embodiment is to determine optimal load balancing and sharing across multiple clients. Scene processing in an exemplary embodiment allows for the amalgamation of all types of data; the scene model is an indexable, extensible, translatable object with connections to virtually all other types of data; the scene then provides a context for various types of data and is itself searchable based on all types of data.

[0021] A scene in exemplary embodiments can be thought of as a region in space and time occupied by a matter field and a light field. An exemplary system according to some embodiments supports scene visualization in freeview, freematter, and freelight, where freeview allows a user to self-navigate a scene; freematter allows a user to objectify, qualify, quantify, enhance, and otherwise translate a scene; and freelight allows a user to recast a scene taking into account the unique spectral outputs of various light sources and even light intensity and polarization considerations, all of which add to the realism of the scene model. The combination of freematter and freelight allows a user to recontextualize a scene into various settings, such as experiencing a city tour of Prague on a winter morning or a summer night.

[0022] While human visualization of scene data is always important, codecs according to some embodiments provide an array of scene data types and capabilities, including metrology, object recognition, scene, and situational awareness. Scene data may include the entire range of data and metadata determinable in the real world, limited only by the range of matter and light field detail contained within the scene model. This range of data must then be formatted for a range of consumers, from humans to AI systems to automata, such as search and rescue automatons crawling or flying over a disaster scene modeled in real time and using advanced object recognition to search for specific objects and people. As such, codecs according to example embodiments are free view, free matter, free writing, and free data.

[0023] Codecs according to some embodiments implement new apparatus and methods for highly efficient sub-scene and scene incrementation, extraction, and insertion, and such technological improvements in efficiency result in significant reductions in computer processing requirements, such as computation time, with associated power requirements. Given the expected growing market requirements for multi-way, just-in-time, only-needed scene reconstruction and delivery, a new type of scene processing unit, including customized computer chips, is needed that embed a new class of instruction sets optimized for new representations and organization of representations of complex, highly detailed real-world scenes.

[0024] Referring to FIG. 1A, a block diagram illustrating key components of a system using a scene codec 1A01 is shown, according to some exemplary embodiments. System 1A01 provides significant technological improvements for the reconstruction, distribution, and processing of scene models, where a real scene is generally understood to be three-dimensional space but may include a fourth dimension, time, such that spatial aspects of the real scene may change over time. The scene model may be either a reconstruction of a real scene, or a computer-generated scene, or a scene augmentation, or any combination thereof. System 1A01 addresses the practical challenge of a global scene model, which is generally understood to represent a larger real-world space, the experience and exploration of which an end user achieves in spatial increments, referred to herein as subscenes. As an example, the global real scene may be a major tourist city such as Prague, where an attempt to explore Prague in the real world would require a large number of subscenes. ,would require days of spatial navigation across sub-scenes containing a significant amount of spatially detailed information.,In particular, for larger real scenes, the combination of scene entry points,,traversal paths, and viewpoints along the traversal paths,produces a virtually infinite amount of information, thus necessitating,intelligent scene modeling and processing, including compression.

[0025] For purposes of efficiency of explanation below, when this disclosure refers to a scene or subscene, it should be understood that this is a scene model or subscene model, and therefore should be understood to exist, as opposed to the real scene or real subscene from which the model is at least partially derived. However, at times, this disclosure may describe a scene as real, i.e., the real world, to discuss the real world without confusion with the modeled world. It should also be understood that the terms viewer and user are used interchangeably without distinction.

[0026] System 1A01 is configured to intelligently provide user access to a virtually infinite number of scenes in a highly efficient real-time or near-real-time manner. A global scene can be thought of as a combination of local scenes, which are less extensive but must be explored in a spatially incremental manner. Local scenes, and therefore global scenes, can have entry points where scene information is initially presented to the user. Scene entry points are essentially subscenes; for example, a scene entry point in a global scene model for "Prague" might be the "Narthex of St. Clement's Cathedral." Again, it will be understood that the data provided by system 1A01 to represent the "Cathedral" subscene is typically substantially less than the entire data for the "Prague" global scene. In some exemplary embodiments, a provided subscene, such as "St. Clement's Cathedral," is determined by the system to be a minimal scene representation sufficient to meet end-use requirements. This determination of sufficiency by the system in some exemplary embodiments has many advantages. Generally, determining sufficiency involves providing subscene model information with varying levels of matter field and / or light field resolution based on at least the desired or expected scene viewing direction. For example, higher resolution information may be provided for nearby objects as opposed to visually distant objects. The term "light field" refers to the flow of light in all directions in all regions of a scene, and the term "matter field" refers to matter occupying regions in a scene. In this disclosure, the terms "light" or "light" refer to electromagnetic waves in frequencies including the visible, infrared, and ultraviolet bands.

[0027] Additionally, according to some exemplary embodiments, system 1A01 intelligently provides spatial buffers to subscenes for purposes such as, for example, providing “look-ahead” scene resolution. In the example “Saint Clement’s Cathedral Narthex” subscene, a minimum resolution might be one in which a viewer stands motionless at the entrance to Saint Clement’s Cathedral and then rotates 360 degrees to look in any direction, e.g., toward or away from the cathedral. This minimum resolution would be sufficient assuming the viewer remains standing within the narthex, but if the viewer approaches and desires to enter the cathedral, the resolution in the direction of the cathedral would eventually fall below a quality of service (QoS) threshold. The system anticipates the viewer’s requested movement and, in response, includes additional non-minimum resolution so that the viewer does not perceive a substantial loss of scene resolution when the viewer moves their free viewpoint. In the current example, this additional non-minimum resolution could include enough resolution to see all of Prague at the QoS threshold, except that this would in turn result in significantly excessive, and in most cases unused, data processing and transmission, which would likely adversely affect an uninterrupted real-time viewer experience. Thus, the concept of a scene buffer is to intelligently determine and provide some additional non-minimum resolution based on all known information, including the viewer's likely traversal path, traversal path viewpoint, and traversal path movement speed.

[0028] System 1A01 exhibits a high degree of contextual awareness regarding both the scene and the users experiencing and requesting access to the scene; in some exemplary embodiments, this contextual awareness is enhanced based on the application of machine learning and / or the accumulation of scene experience logging performed by system 1A01. For a global scene such as Prague experienced by multiple users over time, logging of at least individual users' traversal metrics, including selected entry points, traversal paths, traversal path viewpoints, and traversal travel speeds, provides significant information to system 1A01's machine learning components and helps adjust the size of spatial buffers, thus ensuring a maximal (or substantially maximal) continuous user experience of the scene provided with a minimal (or substantially minimal) amount of provided scene information; this max-min relationship is the focus of system 1A01's scene compression technology in some exemplary embodiments. Another important aspect of scene compression addressed by system 1A01 is scene processing time, which is highly dependent on a novel arrangement of scene model data representing the real-world scene, here generally referred to as a plenoptic scene model and contained in plenoptic scene database 1A07.

[0029] Those familiar with the term “plenoptic” will recognize it as a five-dimensional (5D) representation of a particular point in a scene from which a movement of 4π steradians may be experienced; thus, any point (x, y, z) in the scene may be considered the center of a sphere from which user movement may be experienced in any direction (θ, φ) outward from the center point. Those familiar with light field processing will also understand that plenoptic functions are useful for describing what, at least in the art, is referred to as a light field. As detailed herein, some exemplary embodiments of the present invention implement novel representations of both light fields and matter fields of real scenes such that a user's substantially 5D traversal of a scene model can be efficiently processed in a just-in-time manner to enable a maximally (or substantially maximally) continuous user experience provided with a minimal (or substantially minimal) amount of newly provided scene information.

[0030] The system 1A01 further comprises a spatial processing unit (SPU) 1A09 for substantially processing the plenoptic scene database 1A07 for the purposes of both scene reconstruction and scene distribution. As described herein, reconstruction is generally the process of adding to or building upon a scene database to augment any of various data representations of a scene, including, but not limited to: 1) the spatiotemporal extent, which is the three-dimensional volume of a real scene, for example, ranging from a car hood being inspected for damage to crossing Prague for sightseeing; 2) spatial detail, which includes at least a visual representation of the scene relative to the limits of spatial acuity perceptible to a user experiencing the scene, where visual spatial acuity is generally understood to vary with the human visual system and defines a maximum resolution of detail per approximately 0.5 to 1.0 arcminutes of solid angle that is distinguishable by a human user, such that additional detail is substantially imperceptible to the user unless the user approaches the scene area and thereby changes spatial location to effectively increase the scene area within that solid angle; and 3) the light field dynamic range, including both the light intensity and color gamut, representing the perceived scene, for example, the dynamic range being more for portions of the scene considered to be foreground versus background. 3) a light field dynamic range that can be intelligently modified to provide a large color gamut; 4) a matter field dynamic range that includes both spatial properties (e.g., surface shape) along with light interaction properties that describe the effect of materials in a scene on the transmission, absorption, and reflection of the scene's light field. Sub-scene extraction is then the intelligent and efficient determination by system 1A01 using SPU 1A09 of a minimum dataset of scene information with respect to various dimensions of information representing the scene in plenoptic scene database 1A07, where again, it is paramount to the user's experience that this minimum dataset (sub-scene) provides a substantially continuous experience with sufficient scene resolution (e.g., continuity and / or resolution that meets a predetermined QoS threshold).

[0031] System 1A01, in at least some embodiments, may include a scene solver 1A05 for providing machine learning in one or more of the processes of scene reconstruction and sub-scene distribution, where auxiliary scene information, such as information indicating scene entry points, traversal paths, viewpoints, and effective scene increment pace, may be considered in providing maximum scene compression with minimal or at least acceptable scene loss.

[0032] System 1A01 further comprises a request controller 1A13 for receiving requests indicated through a user interface implemented by application software 1A03. The received requests are translated into control packets 1A17 for communication with another networked system using scene codec 1A11. Thus, system 1A01 is also capable of receiving requests generated by other networked systems 1A01. Received requests are processed by system 1A01 either independently by request controller 1A13 or in combination with both request controller 1A13 and application software 1A03. Control packet 1A17 may include either or both explicit and implicit user requests, where an explicit request represents a conscious decision by the user, such as selecting a particular available entry point into a scene (e.g., St. Clement's Cathedral as the starting point for a Prague tour), and an implicit user request may represent a subconscious decision by the user, such as detecting the user's head orientation with respect to the current scene (e.g., as detected by a camera sensor attached to a holographic display or an inertial sensor provided in a virtual reality (VR) headset). This distinction between explicit and implicit is illustrative, but not limiting, and some user requests are semi-conscious, for example, scene increment pace, which may be indicated by movement of a motion controller in a VR system.

[0033] Scene codec 1A11 is configured to respond to user requests, which may be included in control packets 1A17, and preferably provides just-in-time scene data packets when and if system 1A01 is functioning as a scene provider. Scene codec 1A11 may further be enabled to receive and respond to scene data packets 1A15 when and if system 1A01 is functioning as a scene consumer. For example, system 1A01 may be a provider of scene information that is extracted from plenoptic scene database 1A07 to numerous other systems 1A01 that receive the provided scene information for potential consumption by end users. The scene information included in plenoptic scene database 1A07 need not be strictly limited to visual information and thus may also include, for example, information that is ultimately received by a user viewing some form of image output device in some exemplary embodiments. In some exemplary embodiments, the scene information may be any number of meta-information that is at least partially translated from the matter field and light field of the scene, such as scene metrics (e.g., size of a table) or scene recognition (e.g., placement of light sources), or information that is not a matter field or a light field, but any combination or portion of a matter field and a light field. It should be appreciated that the information may also include related information such as auxiliary information that may be associated with the scene. Exemplary auxiliary information includes, but is not limited to, scene entry points, scene object labels, scene extensions, digital scene signage, etc.

[0034] System 1A01 may be configured for either or both outputting and receiving scene data packets 1A15. Furthermore, the exchange of scene data packets 1A15 between systems such as system 1A01 may not be synchronous or homogeneous, but rather may exhibit minimal responsiveness in order to maximally satisfy user requests, as expressed solely in control packets 1A17 or otherwise through a user or application interface provided by application software 1A03. In particular, with respect to the periodicity of scene data packets 1A15, in contrast to conventional codecs, scene codec 1A11 may operate asynchronously, e.g., providing sub-scene data representing a scene increment having a given scene buffer size both just-in-time and only as needed, or just-in-time but only as expected, where "needed" is often due to explicit user request and "expected" is often due to implicit user request. In particular, with respect to the content construction of the scene data packets 1A15, in contrast to conventional codecs, the scene codec 1A11 can operate to provide heterogeneous scene data packets 1A15, for example, just-in-time packets containing any one or any combination of matter field information, light field information, auxiliary information, or any translation thereof.

[0035] It is also understood that a “user” is not limited to a person but can include any requester, such as another autonomous system 1A01 (see, for example, a terrestrial robot, UAV, computer, or cloud system as depicted in FIG. 3, which will now be described). As those familiar with autonomous systems will appreciate, such autonomous systems may be particularly useful for scene representations that essentially include visual representation information; for example, known images of a scene may be usable by the autonomous system 1A01 performing search and find tasks to compare with visual information captured by the autonomous system 1A01 in real-world scenes that either correspond to or are similar to scenes contained within the plenoptic scene database 1A07. Furthermore, such autonomous systems 1A01 may have preferred applications for non-visual or quasi-visual information, where non-visual information may include scene and scene object measurements and quasi-visual information may include scene lighting attributes. Either the autonomous or human-operated system 1A01, in some exemplary embodiments, may also be configured to collect and provide non-visual representations of the scene for possible spatial or even object collocations in a plenoptic scene database 1A07, or at least for further describing the scene as auxiliary information (see, e.g., FIG. 4C , shown below, illustrating a scene database view). For example, the non-visual representation may include other sensory information, such as somatosensory (touch), olfactory (smell), auditory (hearing), or even gustatory (taste). For purposes of this disclosure, where the focus is on scene representation as visual information, this focus should not be construed as a limitation, but rather as a feature, in which case it is at least understood that the plenoptic scene database can include any sensory information or translation of sensory information, particularly audio / visual data, often requested by human users.

[0036] 1B, a block diagram of a scene codec 1A11 included in one of systems 1A01 is shown, where scene codec 1A11 includes both encoder 1B11a and decoder 1B11b, according to some example embodiments. As described in connection with FIG. 1A, the main function of the encoder is to determine and provide scene data packets 1A15, perhaps over a network, for reception by at least one other system, such as another system 1A01, which is capable of receiving and processing the scene data packets (e.g., by including a decoder such as decoder 1B11b). Encoder 1B11a can receive and respond to control packets 1A17. A use case for system 1A01 including scene codec 1A11 with both encoder 1B11a and decoder 1B11b is described in connection with FIG. 7, which follows below.

[0037] 1C, a block diagram of a scene codec 1A11 including only encoder 1B11a (and therefore no decoder 1B11b as depicted in FIG. 1B) is shown, according to some example embodiments. A use case for a system 1A01 with a scene codec 1A11 including only encoder 1B11a is described below in connection with FIGS. 5 and 6, which will be described next.

[0038] 1D, a block diagram of a scene codec 1A11 including only a decoder 1B11b (and therefore no encoder 1B11a as depicted in FIG. 1B) is shown, according to some example embodiments. A use case for a system 1A01 with a scene codec 1A11 including only a decoder 1B11b is described below in connection with FIGS. 5 and 6, which will be described next.

[0039] 1E, a block diagram of a network 1E01 including a transport layer for connecting two or more systems 1A01 is shown. Network 1E01 may represent any means or communications infrastructure for transmitting information between any two or more computing systems, and in the exemplary embodiment, the computing system of primary focus is system 1A01, but is not limited to such in at least some embodiments. As those familiar with computer networks will appreciate, many variations of networks currently exist, such as personal area networks (PANs), local area networks (LANs), wireless local area networks (WLANs), campus area networks (CANs), metropolitan area networks (MANs), wide area networks (WANs), storage area networks (SANs), passive optical local area networks (POLANs), enterprise private networks (EPNs), and virtual private networks (VPNs), any and all of which may be embodiments of network 1E01 as described herein. As will also be well understood by those familiar with computer networks, the transport layer is generally understood to be a logical division of technology in a layered architecture of protocols in a network stack, and is referred to, for example, as Layer 4 with respect to the Open Systems Interconnection (OSI) communications model. For purposes of this disclosure, the transport layer includes functions that communicate information, such as control packets 1A17 and scene data packets 1A15, exchanged over network 1E01 by any two or more systems 1A01.

[0040] Still referring to FIG. 1E, computing systems such as system 1A01 communicating over network 1E01 are often referred to as residing either on the server side, such as 1E05, or on the client side, such as 1E07. The typical server-client distinction is most often used in reference to web services (server side) being provided to a web browser (client side). It should be understood that within the exemplary embodiment, for example, there is no limitation that application software 1A03 in system 1A01 be implemented using a web browser as opposed to another technology, such as a desktop application or even an embedded application, and as such, the terms server-side and client-side are used herein in their most general sense, such that a server is any system 1A01 that determines and provides scene data packets 1A15, and a client is any system 1A01 that receives and processes scene data packets 1A15. Similarly, a server is any system 1A01 that receives and processes control packets 1A17, and a client is any system 1A01 that determines and provides control packets 1A17. 1A01. System 1A01 may function as either a server or a client, or a single system 1A01 may function as both a server and a client. Accordingly, the block diagrams and descriptions provided herein regarding the network, transport layer, server side, and client side should be considered useful for conveying information rather than as limitations on the exemplary embodiments. FIG. 1E illustrates network 1E01 of systems 1A01 as including one or more systems 1A01 functioning as servers at any given time, and also one or more systems 1A01 functioning as clients at any given time, with the understanding that a given system 1A01 may function alternately or substantially simultaneously as both a server of scene data and a client of scene data.

[0041] 1E also shows optional sensor 1E09 and optional sensor output 1E11, which may be included in some exemplary embodiments. It will be understood that system 1A01 does not require either sensor 1E09 or sensor output 1E11 to perform useful functions, such as, for example, receiving scene data packets 1A15 from other systems 1A01 for use in scene reconstruction or providing scene data packets 1A15 to other systems 1A01 for further scene processing. Alternatively, system 1A01 can include any one or more sensors 1E09, including, but not limited to: 1) imaging sensors for detecting any of the data in the multispectral range, such as ultraviolet, visible, or infrared, filtered for any of the properties of light, such as intensity and polarization; 2) distance or communication sensors that can be used at least in part to determine distance, such as lidar, time-of-flight sensors, ultrasonic, ultra-wideband, microwave, and any other radio frequency-based system; and 3) any of the non-visual sensors that can detect other sensory information, such as somatosensory (touch), olfactory (smell), auditory (hearing), or even gustatory (taste). It is important to understand that the real-world scenes to be represented in plenoptic scene database 1A07 typically include what would commonly be understood to be visual data, although this data is not necessarily limited to what is known as the visible spectrum, and that real-world scenes include a great deal of additional information that can be sensed using any of the sensors available today, as well as additional known or unknown data that can be detected by future sensors. In the spirit of the exemplary embodiments, all such sensors may provide information useful for scene reconstruction, distribution, and processing as described herein, and are therefore sensors 1E07. Similarly, there are many currently known sensor outputs 1E09, including, but not limited to, 2D, 3D, and 4D image display devices, which often include accompanying 1D, 2D, or 3D auditory output devices.The sensor outputs 1E09 include any currently known and future devices for providing any form of sensory information, including visual, auditory, tactile, olfactory, or even gustatory. Any given system 1A01 may include zero or more sensors 1E07 and zero or more sensor outputs 1E09.

[0042] 1F, a block diagram of encoder 1B11a of scene codec 1A11, including at least encoder 1B11a, is shown, according to some example embodiments. An API interface 1F03 of scene codec 1A11 receives and responds to application interface (API) calls 1F01 from an API control host, which may be, for example, application software 1A03. API 1F03 is in communication with various codec components, including packet manager 1F05, encoder 1B11a, and non-plenoptic data control 1F15. API 1F03 receives control signals, such as commands, from a host, such as application software 1A03, provides control signals, such as commands, to various codec components, including 1F05, 1B11a, and 1F15, based at least in part on any of the host control signals, receives control signals, such as component status indications, from various codec components, including 1F05, 1B11a, and 1F15, and controls component status. and providing control signals, such as codec status indications, to a host, such as application software 1A03, based at least in part on any of the indications. The primary purpose of API 1F03 is to provide an external host with a single interaction point for controlling scene codec 1A11, where API 1F03 is, for example, a set of software functions executing on a processing element, and in one embodiment, the processing element for executing API 1F03 is exclusive to scene codec 1A11. Furthermore, API 1F03 can perform functions for controlling ongoing processes commanded by host 1F01, generating multiple signals and communications between API 1F03 and various codec components, including 1F05, 1B11a, and 1F15, with a single host command. At any time during the execution of any of scene codec 1A11's internal processes, API 1F03 determines whether a response, such as a status update, needs to be provided to host 1F01 based at least in part on the interface contract implemented for API 1F03, as will be understood by those familiar with software programming, particularly object-oriented programming.

[0043] Still referring to FIG. 1F, each of the various components 1F05, 1B11a, 1F15 communicate with one another as needed to exchange control signals and data appropriate to any of the internal processes implemented by the scene codec 1A11. During normal operation of the scene codec 1A11, the packet manager 1F05 receives one or more control packets 1A17 for internal processing by the codec 1A11 and provides one or more scene data packets 1A15 based on the internal processing by the codec 1A11. As those familiar with networked systems will appreciate, in one embodiment, the scene codec 1A11 implements a data transfer protocol over what is referred to as a packet-switched network for transmitting data divided into units called packets, each packet including a header describing the packet and a payload, which is the data carried within the packet. As described in connection with FIG. 1E, exemplary embodiments may be implemented over multiple network 1E01 types; for example, multiple systems using the scene codec 1A01 communicate over the packet-switched network 1E01, the Internet. A packet-switched network 1E01 such as the Internet uses a protocol at the transport layer 1E03 such as TCP (Transmission Control Protocol) or UDP (User Datagram Protocol).

[0044] TCP is well known in the art and offers many advantages, such as message acknowledgment, retransmission and timeouts, and proper ordering of transmitted data streams. However, it is typically limited to what is known in the art as unicasting, in which a single server system 1A01 provides data to a single client system 1A01 for each single TCP stream. Using TCP, it is still possible for a single server system 1A01 to set up multiple TCP streams with multiple client systems 1A01, and vice versa, with the understanding that transmitted control packets 1A17 and data packets 1A15 are exchanged exclusively between the two systems forming a single TCP connection. Other data transmission protocols, such as UDP (User Datagram Protocol), are known in the art for supporting what is known in the art as multicasting or broadcasting; unlike unicasting, these protocols allow, for example, multiple client systems 1A01 to receive the same stream of scene data packets 1A15. UDP is limited in that transmitted data is not acknowledged upon receipt by the client and packet transmission order is not maintained. Packet manager 1F05 may be adapted to implement any one of available data transfer protocols based on at least one of the TCP and UDP transport layer protocols for communicating packets 1A17 and 1A15; new protocols may become available in the future, or existing protocols may be further adapted; embodiments should not be unnecessarily limited to any single selection of data transfer protocol or transport layer protocol; rather, the protocol selected to implement a particular configuration of a system using scene codec 1A01 should be selected based on the desired implementation of the many features of a particular embodiment.

[0045] 1F , the packet manager 1F05 analyzes each received control packet 1A17, for example, by processing either the packet's header or payload, to determine various types of packet 1A17 content, including, but not limited to, 1) user requests for plenoptic scene data, 2) user requests for non-plenoptic scene data, 3) scene data usage information, and 4) client state information. The packet manager 1F05 provides any information related to the user request for plenoptic scene data to the encoder 1B11a, which processes the user request, at least in part, using a query processor 1F09 to access the plenoptic scene database 1A07. The query processor 1F09 comprises, at least in part, a subscene extractor 1F11 for efficiently extracting the requested plenoptic scene data, including subscenes or increments to subscenes. The extracted requested plenoptic scene data is then provided to packet manager 1F05 for insertion as a payload into scene packet 1A15 for transmission to requesting (client) system 1A05. In one embodiment, packet manager 1F05 preferably further inserts sufficient information to identify the original user request into scene data packet 1A15 containing the requested plenoptic scene data, such that receiving client system 1A05 receives both an indication of the original user request and the plenoptic scene data provided to fulfill the original request. In operation, any of encoder 1B11a, query processor 1F09, packet manager 1F05, and particularly sub-scene extractor 1F11 may invoke codec SPU 1F13 to efficiently process plenoptic scene database 1A07, which may be configured to implement various technical advantages described herein for efficiently processing representations and organization of representations for plenoptic scene database 1A07.

[0046] The representations in the exemplary embodiments are used to represent real-world scenes as plenoptic scene models, and novel organizations of these representations are used in plenoptic scene database 1A07. The combination of representations and organization used in the exemplary embodiments for processing plenoptic scene database 1A07 provides significant technical advantages, such as the ability to efficiently query plenoptic scene databases potentially representing very large, complex, and detailed real-world scenes, and then quickly and efficiently extract requested sub-scenes or increments to sub-scenes. As those familiar with computer systems will appreciate, scene codec 1A11 can be implemented in many combinations of software and hardware, including, for example, a high-level programming language such as C++ running on a general-purpose CPU, or an embedded programming language running on an FPGA (Field Programmable Gate Array), or a substantially hard-coded instruction set provided on an ASIC (Application Specific Integrated Circuit). Additionally, any of the components and subcomponents of scene codec 1A11 may be implemented in different combinations of software and hardware, and in one embodiment, codec SPU 1F13 is implemented as a substantially hard-coded instruction set such as is provided in an ASIC. Alternatively, in some embodiments, the implementation of codec SPU 1F13 is a separate hardware chip in communication with at least scene codec 1A11, such that codec SPU 1F13 is effectively external to scene codec 1A11.

[0047] As those familiar with computer systems will appreciate, the scene codec 1A11 stores at least part or all of the plenoptic scene database 1A07, or a copied portion of the database 1A07 that is most relevant to the plenoptic scene model. 1A11a。 Scene codec 1A11 may further comprise a memory or some other form of data storage element for maintaining the copied portion, for example, the copied portion may be implemented in what is known in the art as a cache. It is important to note that while plenoptic scene database 1A07 is currently represented as being external to scene codec 1A11, in alternative embodiments of the current scene codec, at least a portion of plenoptic scene database 1A07 is maintained within scene codec 1A11 and also within encoder 1B11a. Therefore, the currently represented block diagram of the scene codec having at least an encoder is illustrative and therefore should not be considered a limitation of the illustrative embodiment, and it is important to understand that many modifications and configurations of the various components and subcomponents of scene codec 1A11 are possible without departing from the spirit of the described embodiment.

[0048] 1F, packet manager 1F05 provides any scene data usage information to encoder 1B11a, which inserts the usage information, or other information calculated at least in part based on the usage information, into plenoptic scene database 1A07. (For more information on usage information, see, inter alia, FIG. 4C for the following plenoptic database data model view and FIGS. 5, 6, and 7 for the following use cases.) As further described, usage information is highly valuable in optimizing the functionality of exemplary embodiments, including at least determining the information range of a sub-scene or scene increment to ideally service a user's request. Packet manager 1F05 also provides any client state information to encoder 1B11a, which maintains client state 1F07 based at least in part on any client state information received from client system 1A01. It is important to understand that the scene codec 1A11 can support multiple client systems 1A01, and a different client state 1F07 is maintained for each supported client system 1A01. As will be further explained in Figures 5, 6, and 7, particularly in connection with the use cases below, the client state 1F07 is at least sufficient to allow the encoder 1B11a to determine the extent of the plenoptic scene database 1A07 information that has already been successfully received by and is available to the client system 1F07.

[0049] Unlike conventional codecs for providing some types of other scene data 1F19 (such as video), the scene codec 1A11 having encoder 1B11a provides either plenoptic scene data 1A07 or other scene data 1F19 to the requesting client system 1A01. Also, unlike conventional codecs, at least the plenoptic scene data 1A07 provided by scene codec 1A11 is of a nature that is not necessarily completely consumed when received and processed by the client system 1A01. For example, when conventional codecs stream video including a series of image frames typically encoded in some format such as MPEG, as the encoded stream of images is decoded by the conventional client system, each next decoded image is essentially presented to the user in real time, after which the decoded image essentially has no further value, or at least no further immediate value when the user is presented with the next decoded image, and so on until the entire stream of images has been received, decoded, and presented.

[0050] In contrast, the current scene codec 1A11 provides at least plenoptic scene data 1A07, such as sub-scenes or scene increments, that are both immediately available to the user of the client system 1A01 while retaining additional substantial future value. As further described in Figures 5, 6, and 7 with respect to at least the following use cases, for example, the codec 1A11 with encoder 1B11a may represent some requested portion of the server's plenoptic scene database 1A07 in scene data packet 1A15. The scene data packet 1A15 transmits the sub-scene or sub-scene increment, which is then received and decoded by the requesting client system 1A01 for two substantially simultaneous purposes, including immediate data provision to the user and insertion into the client's plenoptic scene database 1A07. It should then be further understood that by inserting the received plenoptic sub-scene or sub-scene increment into the client database 1A07, the inserted scene data is then made available for later use in responding to future potential user requests received directly from the client's plenoptic scene database 1A07 without requiring additional scene data from the server's plenoptic scene database 1A07. After inserting the subscene or subscene increment into the client's plenoptic scene database 1A07, the client system 1A01 then provides client state information as feedback to the providing scene codec 1A11, the client state information being provided within a control packet 1A17, and the parsed client state information being used by the encoder 1B11a to update and maintain the corresponding client state 1F07.

[0051] By receiving and maintaining client state 1F07 associated with the stream of scene data packets 1A15 provided to client system 1A01, codec 1A11 with encoder 1B11a can then determine at least the minimum scope of new server plenoptic scene database 1A07 information needed to satisfy the user's next request received from the corresponding client system 1A01. It is also important to understand that in some use cases, client system 1A01 receives plenoptic scene data from two or more server systems 1A01 equipped with scene codec 1A11 with encoder 1B11a. In these use cases, client system 1A01 preferably notifies each server system 1A01 regarding changes in client state information based on scene data packets 1A15 received from all server systems 1A01. In such an arrangement, multiple serving systems 1A01 are used in a load-sharing situation, whereby a user request made from a single client system 1A01 can be conveniently fulfilled using the plenoptic scene database 1A07 of any of the serving systems 1A01, as if all of the serving systems 1A07 were collectively providing a single virtual plenoptic scene database 1A07.

[0052] 1F , packet manager 1F05 provides any of the user requests for non-plenoptic scene data to non-plenoptic data control 1F15, which is in communication with one or more non-plenoptic data encoders 1F17. Non-plenoptic data encoder 1F17 includes any software or hardware component or system that provides other scene data from other scene data databases 1F19 to data control 1F15. It is important to understand that in some embodiments, codec 1A11 with encoder 1B11a does not require access to other scene data such as that contained in other scene data databases 1F19, and therefore does not require access to non-plenoptic data encoder 1F17 or even implementation of non-plenoptic data control 1F15 within codec 1A11. For embodiments of scene codec 1A11 having encoder 1B11a that require, or may anticipate requiring, encoding of some combination of both plenoptic scene data, such as that contained in server plenoptic scene database 1A07, and other scene data, such as that contained in other scene data database 1F19, scene codec 1A11 uses the other scene data, such as that provided by non-plenoptic data encoder 1F15, to determine, at least in part, at least a portion of the payload of any one or more scene data packets 1A15. Exemplary other scene data 1F19 includes any information that is not plenoptic scene database 1A07 information (see particularly FIG. 4C below for a description of plenoptic scene database 1A07 information), including video, audio, graphics, text, or some other digital information, that is determined to be necessary to respond to client system 1A01's request, such as that contained in control packet 1A17.For example, a scene in plenoptic scene database 1A07 may be a house for sale where other scene data contained in database 1F19 includes any related video, audio, graphics, text, or some other digital information, such as product video related to objects in the house, such as appliances.

[0053] It is important to note that plenoptic scene database 1A07 includes the capability to store either conventional video, audio, graphics, text, or some other form of digital information for association with any of the plenoptic scene data (see particularly FIG. 4C below for further details), and thus scene codec 1A11 including encoder 1B11a can provide other data, such as video, audio, graphics, text, or some other form of digital information, retrieved from either server plenoptic scene database 1A07 or other scene data database 1F19. As those familiar with computer systems will understand, it is beneficial to store different forms of data in different forms of databases, and the different forms of databases may then reside on different means of storing and retrieving data, for example, some means being more economical in terms of data storage costs and other means being more economical in terms of retrieval time, and therefore it will be apparent to those skilled in the art that in at least some embodiments it is preferable to substantially separate plenoptic scene data from any other scene data.

[0054] Non-plenoptic data encoder 1F17 comprises any processing element capable of accessing at least other scene data database 1F19 and retrieving at least some other scene data to provide to data control 1F15. In some embodiments of the invention, information relating scene data to other non-scene data is maintained in plenoptic scene database 1A07, whereby non-plenoptic data encoder 1F17 preferably accesses server plenoptic scene database 1A07 to determine what of other scene data 1F19 should be retrieved from other scene data database 1F19 to fulfill a user request. In one embodiment, non-plenoptic data encoder 1F17 comprises any processing element capable of retrieving some other scene data in a first format, translating the first format to a second format, and then providing the translated scene data in the second format to data control 1F15. In at least one embodiment, the first format is, for example, uncompressed video, audio, graphics, text, or some other form of digital information, and the second format is one of a compressed format for representing video, audio, graphics, text, or some other form of digital information. In another embodiment, the first format is, for example, one of a first compressed format for representing video, audio, graphics, text, or some other form of digital information, and the second format is one of a second compressed format for representing video, audio, graphics, text, or some other form of digital information. Also, in at least some embodiments, the non-plenoptic data encoder 1F17 is expected to simply extract other scene data from the database 1F19 to provide to the data control 1F15 without format conversion, and the extracted other scene data may already be in a compressed format or in an uncompressed format.

[0055] Still referring to FIG. 1F, a user request may be for plenoptic scene data and other scene data, for example, based on a rendered view of the plenoptic scene data. In this use case, the initial user request is The packet manager 1F05 extracts the plenoptic scene data from the non-plenoptic data and provides it to encoder 1B11a, which then provides the extracted plenoptic scene data to non-plenoptic data control 1F15. Data control 1F15 then provides the non-plenoptic data to non-plenoptic data encoder 1F17, which can translate the plenoptic scene data, for example, into a requested rendered view of a scene, and this rendered view, which is other scene data, is then provided to data control 1F15 for inclusion in the payload of one or more scene data packets 1A15. Alternatively, as will be made clear particularly in connection with FIG. 1G below, the extracted plenoptic scene data can simply be transmitted to client system 1A01 in one or more scene data packets 1A15, and codec 1A11, including decoder 1B11b on client system 1A01, then renders the requested scene view using the extracted plenoptic scene data received in scene data packets 1A15. As can be seen upon reflection, this flexibility creates a network of communication systems using the scene codec 1A11 with multiple options for most efficiently meeting any given user request.

[0056] 1F, it should be understood that scene codec 1A11 with encoder 1B11a can process an unlimited number of simultaneous streams of scene data, and that codec 1A11 with encoder differs in this respect from conventional codecs with encoders that typically provide a single stream of data to either a single decoder (often referred to as unicast) or multiple decoders (often referred to as multicast or broadcast). In particular, as shown above with respect to FIG. 1E, some exemplary embodiments of the present invention provide a one-to-many relationship between a single serving system 1A01 (including scene codec 1A11 with at least encoder 1B11a) and multiple client systems 1A01 (including scene codec 1A11 with at least decoder 1B11b). Some embodiments of the present invention provide a many-to-one relationship between a single client system 1A01 and multiple serving systems 1A01, and even a many-to-many relationship between multiple server systems 1A01 and multiple client systems 1A01.

[0057] Referring now to FIG. 1G, a block diagram of a scene codec 1A11 including at least a decoder 1B11b is shown. Many of the elements illustrated in FIG. 1G are the same as or similar to those in FIG. 1F and therefore will not be described in greater detail. As previously described, the API control host 1F01 is, for example, application software 1A03 running on or in communication with a system using the scene codec 1A01; for example, the software 1A03 either fully or partially implements a user interface (UI) or communicates with a UI. Ultimately, a user, such as a human or automaton, uses the UI to provide one or more explicit or implicit instructions, and these instructions are used, at least in part, to determine one or more user requests 1G11 for scene data. In a general sense, any system using the scene codec 1A01 to determine user requests 1G11 is referred to herein as a client system 1A01. As explained above, and as further explained in Figures 5, 6, and 7, particularly in connection with the use cases set forth below, it is possible, and even desirable, for client system 1A01 to have sufficient scene data to satisfy a given user request 1G11. However, it should also be understood that the corpus of possible useful scene data is likely to far exceed the capacity of any given client system 1A01; for example, a computing platform for implementing client system 1A01 may be a mobile computing device or a computing element embedded within an automaton such as a drone or robot. Thus, some exemplary embodiments may provide a method for a given client system 1A01, which determines one or more user requests 1G11, to communicate with any number of other systems that use scene codec 1A01. The other systems 1A01 may be accessible over the network 1E01, providing that any one or more of these other systems 1A01 has access to or may contain sufficient scene data to satisfy a given user request 1G11. As will be further explained in connection with this figure, the user request 1G11 may then be communicated over the network to another system 1A01 equipped with a scene codec 1A11 including at least an encoder 1B11a, this other system 1A01 being referred to herein as the server system 1A01, which will ultimately provide the client system 1A01 with one or more scene data packets 1A15 for satisfying the user request 1G11.

[0058] There is no restriction that any given system using scene codec 1A01 should be limited to functioning as only a client system 1A01 or only a server system 1A01; as described particularly in connection with Figures 5, 6, and 7, a given system 1A01 can operate at any time as either a client or a server, or both, where the client has a codec including decoder 1B11b and the server has a codec including encoder 1B11a, such that system 1A01 having codec 1A11 including both decoder 1A11b and encoder 1A11a can function as both client system 1A01 and server system 1A01.

[0059] Still referring to FIG. 1G, in codec 1A11 with decoder 1B11b, API host 1F03 communicates with various codec components, including packet manager 1F05, decoder 1B11b, and non-plenoptic data control 1F15. Packet manager 1F05 receives user request 1G11 in any of many possible forms sufficient to communicate the user's desired scene data to server system 1A01, where scene data is broadly understood to include any of the scene data contained in plenoptic scene database 1A07 and / or any other scene data contained in another scene data database 1G07. The other scene data may include any of video, audio, graphics, text, or some other form of digital information. The other scene data may be contained in and retrieved from plenoptic scene database 1A07, particularly as auxiliary information (see element 4C21 in FIG. 4C below). Throughout this specification, descriptions are provided to illustrate various data types, which are generally referred to as including a plenoptic scene database 1A07, which includes a scene model and auxiliary information, including scene model extensions, translations, indexes, and usage history. The scene model may generally include both a matter field and a light field.

[0060] It should be noted that the various types of scene data and other scene data described in this application can be classified; for example, this classification can take the form of a GUID (globally unique identifier) ​​or UUID (universally unique identifier). Furthermore, the inventive structure described herein for reconstructing a real scene into a scene model for possible association with other scene data is applicable to virtually an infinite number of real-world scenes, and it is also useful to provide a classification for the types of real-world (or computer-generated) scenes that can then be used as scene models. Thus, it is also possible to assign GUIDs or UUIDs to represent various possible types of scene models (e.g., cityscapes, buildings, automobiles, houses, etc.). It may also be possible to use a separate GUID or UUID to uniquely identify a particular instance of a type of scene, such as identifying a type of automobile as "2016 Mustang xyz." It should also be understood that a given user requesting scene information may remain anonymous or may be similarly assigned a GUID or UUID. It is also possible for each system using scene codec 1A01, whether acting as a server and / or client, to be assigned a GUID or UUID. Additionally, user requests 1G11 can be categorized into types of user requests (such as "new subscene request," "subscene increment request," "scene index request," etc.), and both the type of user request and the actual user request can be assigned a GUID or UUID.

[0061] In some embodiments, one or more identifiers, such as a GUID or UUID, are included with a particular user request 1G11 to provide to a packet manager, and the packet manager can then include one or more additional identifiers so that control packets 1A17 issued by scene codec 1A11 with decoder 1B11b contain meaningful user-requested classification data, any of which may be stored in at least one of: 1) a plenoptic scene database 1A07 maintained by server system 1A01 servicing user requests, or an external user-requested database made generally available to one or more systems 1A01, such as server system 1A01 servicing user requests. 1) the request traffic processing agent(s) may be stored in a database such as a client database, and 2) may be used to determine either the routing of user requests 1G11 / control packets 1A17 or the load balancing of scene data provision, and any one or more request traffic processing agents may communicate with any one or more of the client and server systems 1A01 over the network 1E01 to route or reroute control packets 1A17, particularly for the purpose of balancing the load of user requests 1G11 with the availability and network bandwidth of the server system 1A01, all of which would be understood by those familiar with networked systems and network traffic management.

[0062] 1G, in one embodiment, client system 1A01 communicates solely with server system 1A01, where client system 1A01 provides user request 1G11 contained in control packet 1A17, and the server system in return provides scene data packet 1A15 that satisfies user request 1G11. In another embodiment, client system 1A01 is serviced by more than one server system 1A01. As described in connection with FIG. 1F, server system 1A01 preferably includes identifying information in scene data packet 1A15 along with any requested scene data (or other scene data), thereby enabling codec 1A11 on client system 1A01 to track the status of received scene data, including responded-to requests, contained within client system 1A01's plenoptic scene database 1A07 or contained within client system 1A01's other scene data database 1G07. During operation, a given scene data packet 1A15 is received and analyzed by packet manager 1F05, and then any non-plenoptic scene data is provided to non-plenoptic data control, any plenoptic scene data is provided to decoder 1B11b, and any user-requested identification data is provided to decoder 1B11b.

[0063] The non-plenoptic data control 1F15 provides any non-plenoptic scene data to one or more of the non-plenoptic data decoders 1G05 for either decoding and / or inclusion, preferably as auxiliary information, in either another scene data database 1G07 or the plenoptic scene database 1A07 of the client system 1A01 (see, for example, FIG. 4C ). Again, non-plenoptic scene data may include, for example, any of video, audio, graphics, text, or some other form of digital information; decoders for such data are well known in the art and are under constant further development; therefore, it should be understood that any non-plenoptic scene data decoder that is available or to become available can be used as the non-plenoptic data decoder 1G05 according to some embodiments.

[0064] The decoder 1B11b receives the plenoptic scene data and, at least in part, A query processor 1G01 with a scene inserter 1G03 is used to insert plenoptic scene data into the plenoptic scene database 1A07 of the client system 1A01. As previously described with respect to FIG. 1F, the client plenoptic scene database 1A07 may be implemented as any combination of internal or external data memory or storage, for example, decoder 1B11b may be equipped with high-speed internal memory to house a substantial portion of the client plenoptic scene database 1A07 most likely to be needed and requested by a user, or additional portions of the client plenoptic scene database 1A07 may be housed external to decoder 1B11b (but not necessarily external to the system using scene codec 1A01 including decoder 1B11b). Similar to the encoder 1B11a, in operation, any of the decoder 1B11b, the query processor 1G01, and in particular the sub-scene inserter 1G03 may invoke the codec SPU 1F13 to efficiently process the plenoptic scene database 1A07, and the codec SPU 1F13 is intended to implement various technical advantages described in this specification for efficiently processing representations and organization of representations related to the plenoptic scene database 1A07.

[0065] 1G, decoder 1B11b receives either user request identification data for either 1) updating client state 1F05 and 2) notifying API host 1F01 via API 1F03 that the user request has been fulfilled. It is important to note that decoder 1B11b may also update client state 1F05 based on any of its internal operations, and the purpose of client state 1F05 includes at least accurately representing the current state of available client system 1A01's plenoptic scene database 1A07 and other scene data database 1G07. 5, 6, and 7 with respect to use cases, the information in the client state 1F05 is useful to the client system 1A01 for use, at least in part, in efficiently determining whether a given client user request can be fulfilled locally at the client system 1A01 using any of the client system's 1A01's available plenoptic scene database 1A07 and other scene data databases 1G07, or requires additional scene data or other data that must be provided by another server system 1A01, in which case the client user's request is packaged in a control packet 1A17 and transmitted to a particular server system 1A01 or a load balancing component, thereby selecting an appropriate server system 1A01 for fulfilling the user's request. The information in the client state 1F05 is useful to either the load balancing component or ultimately to a particular server system 1A01 for efficiently determining at least the minimum amount of scene data or other scene data sufficient to fulfill the user's request.

[0066] After receiving an indication via API 1F03 that a particular user request has been fulfilled, API-controlled host 1F01, such as application software 1A03, then causes client system 1A01 to provide the requested data to the user, who, again, can be either human or autonomous. It should be understood that there are many possible formats for providing scene data and other scene data, such as a freeview format for use on a display to output video and audio, or an encoded format for use by an automaton that has requested scene object identification information, including local directions to objects and confirmation of the object's appearance. What is important to note is that codec 1A11 with decoder 1B11b is operative to receive and process scene data packet 1A15 to provide user request 1G11 to one or more server systems 1A01, which in turn provides the user request 1G11 to one or more server systems 1A01, and ultimately the user receives the requested data in some form through some user interface means. The codec 1A11 with decoder 1B11b operates to keep track of the current client state 1F05, so that the client system 1A01 can use any of the information in the client state 1F05 to determine whether a given user request is being processed by the client system. It is also important to see that client system 1A01 using codec 1A11 with decoder 1B11b may require scene or other data that can be satisfied, at least in part, locally at server system 1A01 or that must be provided by another server system 1A01. Client system 1A01 using codec 1A11 with decoder 1B11b optionally provides one or many of various possible unique identifiers, including, for example, a classifier, along with any user request 1G11, particularly as encoded in control packet 1A17, and it is further important to see that tracking of various possible unique identifiers by client system 1A01 or at least serving system 1A01 is useful for optimizing the overall performance (e.g., by using machine learning) of any one or more clients 1A01 and any one or more servers 1A01. It is also important to see that, like codec 1A11 with encoder 1B11a, codec 1A11 with decoder 1B11b has access to codec SPU 1F13 to significantly improve at least the execution speed of various extraction and insertion operations, respectively, all of which are described in more detail herein.

[0067] 1G , when codec 1A11 with decoder 1B11b processes scene data packets, any of the processing metrics or information may be provided as usage data along with changes to client state 1F07; usage data differs from user requests at least in that a given user request may be fulfilled by providing a sub-scene (such as a scene model of a house for sale); client system 1A01 then tracks how the user interacts with this provided scene model; for example, the user is a human, and the tracked usage information may refer to rooms within the home scene model accessed by the user, the duration of access, the viewpoint taken in each room, etc. Just as there are in fact an infinite number of possible scene models representing any combination of real-world and computer-generated scene models, it should be understood that there are at least a vast number of usage classifications and some other form of information that can be tracked and that would be at least valuable to the machine learning aspects of the exemplary embodiments; and at least one function of the machine learning described herein is to estimate the best information range of a sub-scene or sub-scene increment when determining how to fulfill a user request.

[0068] For example, if a user requests a tour of a city such as Prague (see FIG. 5 in particular) starting from a particular city location, such as the narthex of St. Clement's Cathedral, the system must determine the information range of the initial narthex sub-scene to provide to the user; a wider range generally allows the user greater initial freedom of scene consumption, while a narrower range generally reduces the transmitted scene data and improves response time. As explained, the scene degrees of freedom include at least free view, free matter, and free lighting; for example, free view includes spatial movement within the sub-scene, such as moving from the narthex to enter St. Clement's Cathedral, or moving from the narthex to walk across the street, turn, and capture a virtual image of St. Clement's Cathedral. Careful consideration reveals that each possible user choice for consuming the provided sub-scenes may require an ever-increasing information range, including the required matter-field and light-field data. In this regard, an exemplary embodiment proposes that by tracking scene usage across multiple users and user incidents, the accumulated usage information can be used by the machine learning components described herein to estimate, for example, the extent of matter or light field information that would be required to enable "x" amount of scene movement by the user, where X can then be related to Y amount of time that scene movement is typically experienced, thereby allowing the system to look ahead and predict when a user is likely to need new scene data based on all known usage and currently tracked user scene movement, and such look ahead can then be automatically used to trigger additional (implicit) user requests 1G11 so that more new scene data is provided by the server system 1A01.

[0069] Although not shown in FIG. 1G, any non-plenoptic or other scene data received in scene data packet 1A15 and processed by codec 1A11 with decoder 1B11b may be provided directly to any suitable sensor output 1E11 (see FIG. 1E), thereby providing the requested data to the requesting user, for example, sensor output 1E11 being a conventional display, a holographic display, or an augmented reality device such as a VR headset or AR glasses; by provided directly, we mean that it is provided from codec 1A11 and not provided from a process that retrieves equivalent scene data from either plenoptic scene database 1A07 or other scene data database 1G07 after it has first been stored by codec 1A11 in its respective database 1A07 or 1G07. Furthermore, any of this non-plenoptic data or other scene data provided to sensor output 1E11 may not be stored as data or may be stored as data in either the client's plenoptic scene database 1A07 or other database 1G17, with the storage operation either prior to provision, substantially simultaneous with provision, or subsequent to provision, and provision to sensor output 1E11 may alternatively be achieved by further processing any of the data contained in database 1A07 or 1G07, thereby retrieving scene data that is equivalent to the scene data received in scene data packet 1A15 for provision to sensor output 1E11.

[0070] Referring now to FIG. 2A, a block diagram is shown as presented on page 13 of the publication entitled "Technical report of the joint ad hoc group for digital representations of light / sound fields for immersive media applications," as provided by the "Joint ad hoc group for digital representations of light / sound fields for immersive media applications," the entire contents of which are incorporated by reference. This publication is directed to the processing of "conceptual light / sound fields," and the diagram shows seven steps in the processing flow. The seven steps include: 1) sensor, 2) sensed data conversion, 3) encoder, 4) decoder, 5) renderer, 6) presentation, and 7) interaction commands. The processing flow is intended to present a real-world or computer-synthesized scene to a user.

[0071] 2B illustrates a combination block and pictorial diagram representing an exemplary use case of some exemplary embodiments for providing a user 2B17 with at least visual information about a real-world scene 2B01 through a sensor output device, such as a 2π-4π freeview display 2B19. In the illustrations herein, the real-world scene 2B01 is sensed using one or more sensors, such as real cameras 2B05-1 and 2B05-2, which can image the real scene 2B01 over some field of view, including, for example, a full 4π spherical steradians (hence, 360×180 degrees). The real cameras, such as 2B05-1 and 2B05-2, as well as other sensors of the real-world scene 2B01, provide captured scene data 2B07 to a system 1A01, which may reside, for example, on a server side 1E05 of a network 1E01. In the figures herein, scene data 2B07 is visual in nature, however, as previously described with respect to FIG. 1E, sensors 1E09 of system 1A01 include, but are not limited to, real cameras such as 2B05-1 and 2B05-2, and further, real cameras such as 2B05-1 and 2B05-2 may be, for example, single-sensor narrow-field-of-view cameras, but are not limited to 4P steradian cameras (often referred to as 360-degree cameras). Also, as previously described, cameras such as 2B05-1 and 2B05-2 may be sensitive across a wide range of frequencies, including, for example, ultraviolet, visible, and infrared. However, in the use case illustrated in this figure for an end user 2B17 to view visual information, the preferred sensors are real cameras or multiple-sensor narrow-field-of-view cameras. The depth and color of a real scene can be perceived across multiple points, all of which will be well understood by those familiar with imaging systems.

[0072] 2B, sensor data 2B07, including example camera images herein, is provided to server-side system 1A01. As discussed above in connection with FIG. 1A and further described below in connection with FIG. 4C, additional extrinsic and intrinsic information regarding sensors, such as real cameras 2B05-1 and 2B05-2, may also be provided to server-side system 1A01, including, for example, sensor placement and orientation, possibly sensor resolution, optical and electronic filters, capture frequency, etc. Using the provided information as well as any captured information, such as images including 2B07, server-side system 1A01 reconstructs real-world scene 2B01, which preferably forms a plenoptic scene model in plenoptic scene database 1A07 under the direction of application software 1A03 in combination with scene solver 1A05 and SPU 1A09. 4A and 4B, shown next, provide further information about both an example real-world scene such as 2B01 (FIG. 4A) and an abstract model view (or plenoptic scene model) of the real-world scene (FIG. 4B). According to some embodiments, the plenoptic scene model describes, at least to some degree of resolution, both the matter field 2B09 and the light field 2B11 of the corresponding real-world scene across various dimensions, including 1) spatial extent, 2) spatial detail, 3) light field dynamic range, and 4) matter field dynamic range.

[0073] 2B, a user 2B17 interacting with client-side system 1A01 requests to view at least a portion of a real-world scene 2B01 represented in a plenoptic scene database 1A07 housed on or accessible to server-side system 1A01. In a preferred embodiment, application software 1A03 executing on client-side system 1A01 presents and controls a user interface to determine at least the user request. The following Figures 5, 6, and 7 provide examples of types of user requests. In an exemplary embodiment, application software 1A03 on client-side system 1A01 interfaces with a scene codec, such as 1B11b, having a decoder to communicate the user request in a control packet 1A17 over network 1E01 (shown in Figure 1E) to a scene codec, such as 1A11b, having an encoder, running on server-side system 1A01. The application software 1A03 of the server-side system 1A01 is preferably in communication with the server-side encoder 1B11a and receives at least explicit user requests, such as a user request to receive scene information spatially starting from a given entry point in the plenoptic scene model (coincident with the spatial entry point in the real-world scene 2B01).

[0074] When processing a request, the server-side system 1A01 preferably determines and extracts from the plenoptic scene database 1A07 the relevant subscene as indicated by the requested scene entry point. The extracted subscene preferably further comprises a subscene spatial buffer. Thus, in this example, the subscene minimally includes image data representing a 2π-4π steradian viewpoint located at the entry point, but then maximally includes additional portions of the database 1A07 sufficient to accommodate any expected path traversal of the scene by the user both with respect to the entry point and for a given minimum time. For example, if the real-world scene is Prague and the entry point is the narthex of St. Clement's Cathedral, the minimal extracted scene would substantially allow the user to perceive a 4π / 2 steradian (half-dome) viewpoint located at the narthex of the cathedral. However, additional information may be added within or to the server-side system 1A01 based on the user request, or the user's typical walking speed and direction based on a given entry point. Based on the auxiliary information available to the server-side system 1A01, application software 1A03 running on the server-side system 1A01 can determine sufficient sub-scene buffers to provide enough additional scene resolution to support 30 seconds of walking from the narthex in any available direction.

[0075] 2B, the determined and extracted sub-scenes are provided in the communication of one or more scene data packets 1A15 (see FIG. 1A) within a scene stream 2B13. Preferably, a minimum number of scene packets 1A15 are communicated so that the user 2B17 perceives acceptable application responsiveness, understanding that transferring the entire plenoptic scene database 1A07 is generally prohibitively large (e.g., due to bandwidth and / or time limitations), particularly as any of the scene dimensions increase, as is the case for at least large international city scenes such as Prague. The techniques used in exemplary embodiments for both organizing the plenoptic scene database and processing the organized database offer substantial technical improvements over other known techniques, such that the user experiences real-time or near-real-time scene entry.

[0076] However, it is possible and affordable for a user to experience some degree of delay in favor of perceiving a continuous experience of an entered scene when the user first enters the scene, a continuous experience directly related to both the size of the entry sub-scene buffer and the provision of supplemental scene increments along the explicitly or implicitly expressed direction of scene traversal. Some exemplary embodiments of the present invention provide a means for balancing the initial entry point resolution and the sub-scene buffer, as well as the periodic or aperiodic event-based rate of sub-scene increments and resolution. Such balancing provides a maximum continuous user experience encoded with a minimum amount of scene information, thus resulting in novel scene compression capable of meeting a predetermined quality of service (QoS) level. Within the asynchronous scene stream 2B13 determined and provided by the exemplary server-side system 1A01, any given transmission of scene data packet 1A15 may include any combination of any form and type of information from the plenoptic scene database 1A07, for example, one scene data packet such as 2B13-a or 2B13-d may include at least a combination of matter field 2B09 and light field 2B11 information (e.g., shown as having both “M” and “L” in 2B13-a and 2B13-d, respectively), while another scene data packet such as 2B13-b may not include at least matter field 2B09 but may include some light field 2B11 (e.g., shown as having only “L” in 2B13-b), while another scene data packet such as 2B13-c may include at least some matter field 2B09 but not some light field 2B11 (e.g., shown as having only “M” in 2B13-c).

[0077] Still referring to FIG. 2B, the scene stream 2B13 is transmitted from the server-side system 1A01 via the transport layer of the network 1E01 and is received and processed by the client-side system 1A01. Preferably, the scene data packets 1A15 are first received on the client-side system 1A01 by a scene codec 1B11b equipped with a decoder. The decoded scene data is then processed under the direction of application software 1A03 accessing the functionality of the SPU 1A09. The decoded and processed scene data is preferably used on the client-side system 1A01 to reconstruct a local plenoptic scene database 2B15 and to provide scene information such as scene entry points in 2π-4π free view. The provided entry points free view allow the user 2B17 to explicitly or implicitly change the viewing point representation with respect to at least the viewing angle (θ, φ) and the spatial viewing position (x, y, z) representing the user's current location within the scene. As user 2B17 explores further within the scene and thus presents an explicit or implicit request to move around within the scene, client-side system 1A01 first determines whether local scene database 2B15 contains sufficient information to provide a sub-scene increment or whether to request that server-side system 1A01 extract additional increments from server-side scene database 1A07. A well-functioning scene reconstruction, distribution, and processing system as described herein balances multiple considerations and intelligently determines the optimal QoS for user 2B17 while providing efficient means for storing and retrieving plenoptic scene model information from plenoptic scene database 1A07.

[0078] 3, a combined block and pictorial diagram of a network 1E01 connecting multiple systems using various forms of scene codec 1A01 representing a variety of possible forms, including, but not limited to, personal mobile devices such as cell phones, display devices such as holographic televisions, cloud computing devices such as servers, local computing devices such as computers, unmanned autonomous vehicles (UAVs) such as land robots and drones, and augmented reality devices such as AR glasses. As previously described, all systems 1A01 include a scene codec 1A11, and the codec 1A11 included in any of systems 1A01 may further include both encoder 1B11a and decoder 1B11b, encoder 1B11a and no decoder 1B11b, or decoder 1B11b and no encoder 1B11a. Thus, any system 1A01 as represented in the figures herein may be both a plenoptic scene data provider and a plenoptic scene data consumer, a provider only, or a consumer only. While a single system 1A01, for example, a computer not connected to the network 1E01, may perform any of the functions described herein as defining a system using the scene codec 1A01, some embodiments may include two or more systems 1A01 operating interactively over the network 1E01, whereby these systems exchange either captured real scene data 2B07 to be reconstructed in the plenoptic scene database 1A07 or exchange either plenoptic scene database 1A07 scene data (see particularly Figures 5, 6, and 7 for the use cases described below).

[0079] 4A, a pictorial representation of an exemplary real-world scene 4A01 with a very high level of detail (e.g., sometimes referred to as unlimited or nearly unlimited detail), such as an interior house scene with a window looking out onto an outdoor scene, is shown. For example, the scene 4A01 includes, but is not limited to, any one or any combination of opaque objects 4A03, microstructured objects 4A05, distant objects 4A07, emissive objects 4A09, highly reflective objects 4A11, featureless objects 4A13, or partially transparent objects 4A15. Also depicted is a user operating a system 1A01 using a scene codec, such as a mobile phone operating either individually or in combination with another (not shown) system 1A01, which provides the user with, for example, a sub-scene image 4A17 along with any number of accompanying interpretations of the real-world scene 4A01, e.g., object measurements, light field measurements, or scene boundary measurements such as a portion of the scene including a window boundary versus an opaque boundary (see particularly the next figure, e.g., FIG. 4B).

[0080] Still referring to FIG. 4A, in general, any real-world scene, such as 4A01, is translated into a plenoptic scene model 1A07 through a process called "scene reconstruction." As explained previously, and as will be explained in more detail later in this disclosure, the SPU 1A09 can also be used to generate scene reconstructions, as well as other databases such as, but not limited to, scene augmentation and scene extraction. The system implements several operations to most efficiently perform both the functions of the subscene 1A07: scene augmentation introduces new or synthetic scene information into a scene model that would not otherwise necessarily or substantially be present in the corresponding real-world scene 4A01; ​​and scene extraction determines and processes portions of a scene model that represent a subscene. Synthetic scene augmentation involves providing higher resolution of reconstructed real-world objects, such as wood or marble floors, so that when a viewer views a real-world object from above a given QoS threshold, the viewer is presented with the real-world reconstruction information as represented in the original plenoptic scene model corresponding to the captured real-world scene. However, as the viewer spatially approaches the object in the provided subscene, eventually exceeding the QoS threshold, the system according to some exemplary embodiments intelligently augments during presentation or includes as augmentation within the provided subscene synthetic information, such as details of the bark or marble floor that were not originally captured (or even present) in the real-world scene. Also, as previously described, the scene solver 1A05 is generally an optional processing element that further applies machine learning techniques to increase the accuracy and precision of any of the aspects of scene reconstruction, enhancement (such as QoS-driven synthesis), extraction, or other forms of scene processing.

[0081] As those familiar with computer systems will appreciate, any combination of system components, including application software 1A03, scene solver 1A05, SPU 1A09, scene codec 1A11, and request controller 1A13, provides functionality and technical improvements that may be implemented in various arrangements of components without departing from the scope and spirit of the exemplary embodiments. For example, one or more of the various novel features of scene solver 1A05 could alternatively be provided in either application software 1A03 or SPU 1A09, such that the presently described summary of functionality describing the various system components should be considered exemplary rather than limiting of the exemplary embodiments, and those skilled in the art of software and computer systems will recognize many possible variations of system components and component functionality without departing from the scope of the exemplary embodiments.

[0082] Still referring to FIG. 4A , the reconstruction of a real-world scene such as 4A01 by system 1A01 involves determining data representations for both the matter field 2B09 and the light field 2B11 of the real-world scene, and these representations, and the organization of these representations, significantly affect the scene reconstruction and, more importantly, the efficiency, including the processing speed, of sub-scene extraction. Exemplary embodiments provide scene representations and organization of scene representations that inherently enable, for example, large-scale global scene models to be made available to users who experience them in real time or near real time. Collectively, FIGS. 4A , 4B , and 4C are directed toward describing these scene representations and organization of scene representations at a level of detail, all of which will be described in further detail later in this disclosure.

[0083] Techniques according to example embodiments described herein may use hierarchical, multi-resolution, and spatially sorted volumetric data structures to describe both the matter field 2B09 and the light field 2B11. This allows for identifying the portions of the scene required for remote viewing based on their positioning, resolution, and visibility as determined by each user's positioning and viewing direction, or as statistically estimated for a group of users. By communicating only the necessary portions, channel bandwidth requirements are minimized. The use of volumetric models also facilitates advanced functionality in virtual worlds, such as collision detection and physics-based simulation (where mass properties are easily calculated). Thus, novel Based on a novel scene reconstruction process of real-world scenes such as 4A01 into a reliable plenoptic scene model representation and organization of the representation, as well as novel processes for sub-scene extraction and user scene interaction monitoring and tracking, exemplary embodiments provide many use-case advantages, some of which are illustrated in FIGS. 5, 6, and 7 below, one of which includes providing a free-viewpoint viewer experience, in which one or more remote viewers can independently change their viewpoints on the transmitted sub-scenes. What is required for the greatest free-viewpoint experience, especially for larger global scene models, is sub-scene provision by system 1A01 to the free-viewpoint viewer both just-in-time and only when needed, or just-in-time and only when anticipated.

[0084] Still referring to FIG. 4A , images and other captured sensor data representing real-world scenes, such as 4A01, include one or more characteristics of light, such as color, intensity, and polarization. By processing this and other real-scene information, system 1A01 determines shape, surface characteristics, material properties, and light interaction information for matter field 2B09 relative to a representation in plenoptic scene database 1A07. Separate determination and characterization of the light field 2B11 of the real scene is used in combination with matter field 2B09 to remove ambiguities in surface characteristics and material properties caused by, among other things, scene illumination (e.g., specular reflection, shadows). The presently described novel processing of real-world scenes into scene models enables effective modeling of transparent materials, highly reflective surfaces, and other challenging situations in everyday scenes. For example, this involves the "discovery" of material and surface properties independent of the actual lighting in the scene; this discovery and accompanying representation and organization of the representation then enables novel sub-scene extraction, including the accurate separation and presentation to a free viewpoint viewer of the matter field 2B09, distinct from the light field 2B11.

[0085] In this way, the free-viewpoint viewing experience achieves another key goal of free lighting: for example, when accessing a scene model corresponding to a real scene such as 4A01, the viewer can request free-viewpoint viewing of the scene, perhaps in “morning sunlight” versus “evening sunlight,” or even “half-moon lighting with available room light.” Preferably, a user interface provided by application software 1A03 allows for the insertion of new lighting sources from a group of template lighting sources, and both the newly specified and available lighting sources can then be modified, for example, to change their emission, reflection, or transmission characteristics. Similarly, the properties and characteristics of matter field 2B09 may also be dynamically changed by the viewer; thus, it is particularly important to note that providing free-matter along with free lighting and free viewpoints, exemplary embodiments enable more precise separation of matter field 2B09 from light field 2B11 of the real scene; a lack of precision in the separation would adversely limit the end-use experience to precisely changing the properties and characteristics of matter field 2B09 and / or light field 2B11. Another benefit of accurate matter fields 2B09 as described herein includes interference and collision detection within objects in the matter field; these and other life simulation features require material properties such as mass, weight, and center of gravity (e.g., in physics-based simulations). Again, as those familiar with object recognition in real-world scenes will well understand, highly accurate matter and light fields offer significant advantages.

[0086] Still referring to FIG. 4A, the matter field 2B09 of the real scene 4A01 includes media, which are finite-volume representations of matter through which light flows or is blocked, and thus have varying degrees of light transmittance that can be characterized as degrees of absorption, reflection, transmission, and scattering. Media are positioned and oriented within the scene space and have associated properties such as material type, temperature, and a two-way light interaction function (BLIF) that relates the incident light field to the outgoing light field caused by the light's interaction with the media. Optically, spatially, and temporally homogeneous co-located media form segments of objects that include surfaces with tactile boundaries, where tactile boundaries are generally understood to be boundaries that a human can sense by touch. Using these and other properties of the matter field 2B09, the various objects shown in Figure 4A are distinguished not only spatially but also, importantly, with respect to their interaction with the light field (2B11), and the various objects again include an opaque object 4A03, a finely structured object 4A05, a distant object 4A07, an emissive object 4A09, a highly reflective object 4A11, a featureless object 4A13, or a partially transparent object 4A15.

[0087] 4B, a pictorial representation of a real-world scene 4A01, such as that shown in FIG. 4A, is shown, where the representation can be considered a view of an abstract scene model 4B01 of data contained within a plenoptic scene database 1A07. The abstract representation of the scene model 4B01 of the real-world scene 4A01 includes an outer scene boundary 4B03 that includes a plenoptic field 4B07 comprising a matter field 2B09 and a light field 2B11 of the scene. The light field 2B11 interacts with any number of objects within the matter field 2B11, as well as other objects, such as, for example, the described object 4B09, the undescribed area 4B11, and so on, as described with respect to FIG. 4A. The real-world scene 4A01 is captured by any one or more of the real sensors, for example, real camera 2B05-1 capturing real image 4B13, while the scene model 4B01 is translated into a real-world data representation, such as an image, using, for example, virtual camera 2B03 providing a real-world representation image 4A17.

[0088] In addition to objects including opaque objects 4A03, microstructured objects 4A05, distant objects 4A07, emissive objects 4A09, highly reflective objects 4A11, featureless objects 4A13, or partially transparent objects 4A15 as shown in Figure 4A, systems according to some embodiments may further tolerate both described objects 4B09 and undescribed regions 4B11, where these general objects and regions contain modifications of the features and properties of the matter field 2B09 as described in Figure 4A. Importantly, the matter field 2B09 is identified by scene reconstruction sufficient to distinguish between multiple types of objects, and any distinct types of objects uniquely located within the model scene may be further processed, for example, by using machine learning to perform object recognition and classification and modifying various characteristics and properties to cause model presentation effects such as visualization changes (generally translation), object augmentation and tagging (see particularly model augmentation 4C23 and model index 4C27 with respect to Figure 4C), and object removal. Along with object removal, object translation (see model translation 4C25 in FIG. 4C below) can be specified to perform any number of geometric translations (such as resizing and rotation) or even object movement based on, for example, object collision or an assigned object path, e.g., an opaque object 4A03 classified through machine learning that then rolls along the floor (opaque outer scene boundary 4B03) and bounces off a wall (opaque outer scene boundary 4B03).

[0089] Still referring to Figure 4B, the characteristics and properties of any object can be modified through additional processing of new real scene sensor data, such as, but not limited to, new camera images, perhaps taken at non-visible light frequencies such as infrared, thus providing at least new BLIF (Bidirectional Light Field Information). The object types as shown in Figures 4A and 4B should be considered as examples rather than limitations of the embodiment, and it will be clear to those familiar with software and databases that the data can be updated and the tagging of the associated data forming the matter field object can be adjusted, which may include at least changes in the naming of the object, such as "featureless" vs. "microstructure." or modifying the object classification threshold, which may be used to classify objects as, for example, "partially transparent" versus "opaque." Other useful modifications of object types, as well as other modifications of matter field 2B09 and light field 2B11 in general, will be apparent to those skilled in the art of scene processing based on this disclosure, most importantly the representation and organization of real-world scenes such as 4A01 as described herein, as well as efficient processing where the combination of these is useful in providing the unique functionality as described herein, again providing at least a free viewpoint, free matter, and free light experience where the viewer uses a scene model for visualization.

[0090] The scene model also includes an exterior scene boundary 4B03 that defines the outermost extent of the represented plenoptic field 4B07. As becomes apparent from careful consideration of a real-world scene, such as the kitchen or outdoor scene (not shown) shown in FIG. 4A, some regions of the plenoptic field near the exterior scene boundary 4B03 act substantially opaque (such as a kitchen wall or counter, or thick fog in an outdoor scene), while other regions near the (imaginary) exterior scene boundary may effectively act as windows and represent arbitrary light field boundaries (such as the sky in an outdoor scene). In a real scene, light may cross back and forth through the space associated with the exterior boundary of the scene model. However, the scene model does not allow such penetration. Rather, a window light field may represent the light field in a real scene (such as a television displaying an image of the moon at night). In the scene model, opaque regions near the outer scene boundary 4B03 represent no substantial transmission of light outside the plenoptic field into the scene (within the real scene), while window regions near the outer scene boundary 4B03 represent substantial transmission of light outside the plenoptic field into the scene (within the real scene). In some embodiments, trees and other outdoor material can be represented as being included in the plenoptic field 4B07, and the outer scene boundary 4B03 is spatially extended to include at least this material. However, as objects in the matter field become even more distant, it is beneficial to terminate the plenoptic field 4B07 at the outer scene boundary 4B03, even though it reaches a distance called the neglect limit, at which features on the object do not change substantially with changes in viewpoint. By using one or more window light elements 4B05, it is possible to represent the light field incident on the real scene along parts of the outer scene boundary 4B03 as if the plenoptic field 4B07 in those parts of the outer scene boundary 4B03 extended infinitely.

[0091] For example, referring to FIG. 4A , the scene boundary 4B03, and therefore the plenoptic field 4B07, can be extended beyond the window into the outdoor area, thereby substantially terminating the scene boundary 4B03 with a representation of media and matter including the countertop, wall, and window, rather than including trees within the scene's matter field. Multiple window light elements 4B05, included along the portion of the outer scene boundary 4B03 that spatially represents the window surface, can then be added to effectively cause the window light field to enter the plenoptic field 4B07. Various light field complexities are possible using the window light elements 4B05, including 2D, 3D, and 4D light fields. Note that, as used herein, "medium" refers to the contents of a volumetric region that contains some matter or no matter at all. A medium can be homogeneous or heterogeneous. Examples of homogeneous media include empty space, air, and water. Examples of heterogeneous media include the surface of a mirror (part air and part silvered glass), the surface of a sheet of glass (part air and part transparent glass), and the contents of a volumetric region containing a pine tree branch (part air and part organic material). Light flows through a medium by phenomena including absorption, reflection, transmission, and scattering. Examples of media with partial transparency include a pine tree branch and a sheet of glass.

[0092] In one exemplary use and advantage of the inventive system, as illustrated in FIG. 4B , in a manner corresponding to a real scene such as that shown in FIG. 4A , scene model 1A07 can be used to estimate the amount of daylight transmitted into a scene (e.g., a kitchen) based on time of day, and estimated time-of-day room temperature or seasonal energy savings opportunities based on various types of window coverings can be calculated based at least in part on the estimated amount of transmitted daylight. Such an example calculation is based at least in part on data representing light field 2B11, including plenoptic field 4B07, whose light field representation and other more fundamental calculations are addressed in more detail herein. While light field 4B07 is treated as a quasi-steady-state light field in which all light propagation is modeled as instantaneous with respect to the scene, suffice it to say that, using the free light principles described herein, a viewer may experience a dynamic state light field through the presentation of a visual scene representation, preferably using application software 1A03.

[0093] 4B , both real camera 2B05-1, capable of capturing images of a view spanning, for example, up to 4π steradians, and virtual camera 2B03, having an exemplary limited viewpoint 4A17 of less than 4π steradians, are shown. Any scene, such as real scene 4A01 having a corresponding scene model 4B01 described in plenoptic scene database 1A07, may include any number of real or virtual cameras, such as 2B05-1 and 2B03, respectively. Any of the cameras, such as 2B05-1 and 2B03, may be designated as fixed with respect to the scene model or movable with respect to the scene model, with movable cameras being associated with, for example, a traversal path. It is important to note that many possible viewpoints, and therefore resulting images, of any real or virtual camera, whether movable or fixed, and whether adjustable within a field of view versus a fixed field of view, can be estimated by processing scene model 4B01, as described with respect to FIG. 4B .

[0094] 4C, a block diagram of major data sets within one embodiment of a plenoptic scene database 1A07 is shown, where database 1A07 contains data representing a real-world scene 4A01, for example, as shown in FIG. 4A. The plenoptic scene database 1A07 data model view of the real-world scene 4A01 typically includes either internal or external data 4C07 describing a sensor 1E07, such as, for example, a real camera 2B05-1 or 2B05-2, or a virtual camera 2B03, where the internal and external data are well known in the art based on the type of sensor, and generally, the internal properties relate to the sensor's data capture and processing capabilities, and the external properties relate to the sensor's physical placement and orientation, typically with respect to a scene local coordinate system or at least any coordinate system that allows for understanding the sensor's spatial placement with respect to a scene including a matter field 2B09 and a light field 2B11.

[0095] The data model view further comprises a scene model 4C09, typically including a plenoptic field 4C11, objects 4C13, segments 4C15, BLIFs 4C17, and features 4C19. The term "plenoptic field" has a range of meanings that are within the current state of the art, and further, the present disclosure provides a novel representation of the plenoptic field 4C11, and this novel representation and organization of the representation underlies, at least in part, many of the technical improvements described herein, including, for example, just-in-time sub-scene extraction that provides a substantially continuous visual free-view experience with sufficient scene resolution (and thus satisfying QoS thresholds) made usable by a minimal dataset (sub-scene). Thus, like other specifically described terms and datasets described herein, the plenoptic field 4C11 of this term and dataset should be understood in light of this specification, and not simply with reference to the current state of the art.

[0096] Still referring to FIG. 4C , the plenoptic field 4C11 includes an organization of representations referred to herein as a plenoptic octree, which holds representations of both the matter field 2B09 and the light field 2B11. A more detailed description of the representations and organization of representations generally with respect to the scene model 4C09, as well as key datasets of the scene model, such as the plenoptic field 4C11, objects 4C13, segments 4C15, BLIF 4C17, and features 4C19, among others, are provided in the remainder of this specification, but generally, the plenoptic octree representation as described herein includes two types of representations for the matter field 2B09 and one type of representation for the light field 2B11. The matter field 2B09 is shown to include both (volumetrically) “medium”-type and “surface”-type matter representations. Medium-type representations describe homogeneous or heterogeneous materials through which light substantially flows (or through which light substantially blocks light). This includes empty space. Light flows through media, including media types, through phenomena including absorption, reflection, transmission, and scattering. The type and degree of light change is contained in property values ​​contained in or referenced by voxels in the plenoptic octree. Surface type representations describe tangible (touchable), (approximately) planar boundaries between matter and empty space (or another medium), which contain media on both sides, which may be the same or different media (where surfaces of different media are referred to as "split surfaces" or "split surfels").

[0097] Surface-type materials, including co-located media that are homogeneous in space and time, form segment representations 4C15, and the co-located segments then form a representation of an object 4C13. The effects of surface-type materials on the light field 2B11 (e.g., reflection, refraction, etc.) are modeled by a two-way light interaction function (BLIF) representation 4C17 associated with the surface-type material. The granularity of the BLIF representation 4C17 extends to associations with at least the segments 4C15 containing the object 4C13, but also to associations with feature representations 4C19, where features are referred to as poses of entities, such as the object 4C13, located within the space described by the scene model 4C09. Examples of features in a scene or image include spots and glitters from a microscopic perspective, or buildings from a macroscopic perspective. The BLIF representation 4C17 is based on the interaction of the light field with the material and relates the transformation of the light field 2B11 incident on the material to the light field 2B11 emanating from the material.

[0098] 4C , the primary dataset of at least one of some exemplary embodiments includes auxiliary information 4C21, such as any of, or any combination of, model extensions 4C23, model translations 4C25, model indexes 4C27, and model usage history 4C29. Model extensions 4C23 include any additional metadata to be associated with a portion of the scene model 4C09 that is not otherwise contained within or that does not otherwise specify a change to the scene model 4C09 (e.g., model translations 4C25 describe some mathematical function or similar for an attribution to the scene model that alters (modifies) a scene model interpretation such as an extracted sub-scene or metric).

[0099] The model augmented representation 4C09 may include, but is not limited to, 1) a virtual scene description including text, graphics, URLs, or other digital information that may be displayed as an augmentation to the viewed sub-scene (similar in concept to augmented reality (AR)), for example, including room features, object pricing, or links to the nearest store for purchasing objects related to the real scene shown in FIG. 4A; 2) sensor information, such as a current temperature reading, associated with a portion of either the matter field 2B09 or the light field 2B11, typically at least one of the matter field 2B09 or the light field 2B11; and 3) sensor information, not based in part on, including any type of data available from currently known or yet unknown future sensors, particularly relating to data associated with somatosensation (touch), olfaction (smell), hearing (hearing), or gustatory (taste); and 3) metrics related to calculations describing either the matter field 2B09 or the light field 2B11, examples including measurements of quantity (such as a dimension of size) or quality (such as a dimension of temperature), preferably where the calculations are based at least in part on either the matter field 2B09 or the light field 2B11. Model extensions 4C23 are directly associated with either the scene models 4C09 or indirectly associated with the scene models 4C09 through either the model translations 4C25 or the model indexes 4C27. Model extensions may be based at least in part on any information contained in the model usage history 4C29, for example, extensions are up-to-date statistics about some logged aspect of scene model usage across multiple users.

[0100] Still referring to FIG. 4C, the model translation 4C25 may include, but is not limited to, 1) a geometric transformation applied to either the matter field 2B09 or light field 2B11 including the scene model 4C09, which maps the spatial location of any matter field or light field element to a new location including spatial shift, rotation, magnification, reduction, etc.; 2) a complex geometric transformation such as a trajectory path to describe the movement of an object within the scene model 4C09, or anything else to either the matter field 2B09 or light field 2B11; 3) a geometric transformation that does not necessarily involve a viewer experiencing, for example, a visual representation of the scene model. 4) a virtual scene path including path point timings for describing the movement and viewpoint of a virtual camera (e.g., 2B03) within the scene model to be guided through a scene without the need for freeview instruction, or for describing a proposed scene model path, such as a set of city tour destinations where, for example, a viewer may reposition from one sub-scene to another, and the sub-scenes may or may not be spatially co-located within the scene model 4C09; and 4) a pre-compilation of any of the data contained within the scene model 4C09 and associated auxiliary information 4C21, including, for example, a JPEG image of St. Clement's Cathedral in Prague. Model translations 4C25 are directly associated with any of the scene models 4C09, or indirectly associated with the scene model 4C09 through either a model extension 4C23 or a model index 4C27. Model translations may be referenced by a model usage history 4C29, e.g., usage of a given translation is logged across multiple users.

[0101] The model index 4C27 includes data useful for presenting to either a human or autonomous user index elements to select any portion of the scene database 1A07, including in particular any of the scene model 4C09 or auxiliary information 4C21, including, but not limited to, 1) a list of index elements including any of text, images, video, audio, or other digital information, where each index element in the list is associated with at least one portion, such as a sub-scene, of the scene database 1A07 to be extracted for the requesting user (human or autonomous), or 2) an encoded list of index elements including either encrypted or unencrypted information useful for selecting a portion of the scene database 1A07 using an executed computer algorithm, for example, a remote computer system using the scene codec 1A01 accesses an encrypted model index of extractable scene information including the type of scene information for algorithmic comparison with a desired type of scene information, and the algorithm then selects scene information for extraction based at least in part on the algorithmic comparison. A given model index 4C27 may include associated permission information for granting or denying access to the index 4C27 (and thus the scene model 4C09 through the index 4C27) by any given user (human or autonomous), the permission information including the given user's authorized type, The transaction may include either a specific given user associated with the transaction, access credential information such as a username and password, and payment or sales transaction relationship information including a link to interact with a remote sales transaction provider such as PayPal.

[0102] The model index 4C27 is directly associated with one of the scene models 4C09. The model index 4C27 may be associated with or trigger the use of model expansion 4C23 (e.g., current sensor readings of a particular type taken throughout a real scene, such as a natural disaster scene, corresponding to the scene model) or model translation 4C25 (e.g., relighting of a scene in a morning, daytime, or evening setting, or automatic entry into the scene in a particular subscene, followed by automatic movement throughout the scene according to a predetermined path). The model index 4C27, or any of its index elements, may be associated with any of the model usage history 4C29, where association is interpreted broadly to include any formulation of the model usage history 4C29, such as a statistical percentage of index element selections by multiple users (human or autonomous) with associated scene model elapsed time usage, which statistical percentage is then used to re-sort the ranking or presentation of index elements in a given model index 4C27.

[0103] Still referring to FIG. 4C, the model usage history 4C29 includes any data known to the system using the scene codec 1A01 that represents the user's requests or instructions, whether the user is human or autonomous. The request or instruction may include, but is not limited to, any of: 1) selection of a model index and model index elements; 2) adjustment of free view, free matter, or free light of the scene model based at least in part on explicit or implicit user instructions; 3) any of the user's explicit or implicit user instructions, including human user tracking body movements, facial expressions, or audible sounds; 4) generalizing scene model propagation information to include elapsed time spent within the subscene, or at least the elapsed time starting from the provision of the subscene before incrementing the subscene (such as a spatial increase to include more of the entire scene) or before requesting to switch to an alternative subscene; or 5) any information logging that records the use of model expansion 4C23 (e.g., use of a URL to access information outside the scene database 1A07) or model translation 4C25 (e.g., use of a pre-specified scene path, such as representing a particular tour of a scene that is a cityscape or a house for sale).

[0104] Further, with respect to FIG. 4C, as those familiar with computer databases will understand, there are many types of database technologies available on the market today or that will become available in the future, and the currently described plenoptic scene database 1A07 may be implemented with any number of these database technologies or combinations of these technologies, each with its own trade-off advantages and each enhanced by technological improvements in at least the representation and organization of the scene model 4C09 as described herein. Those familiar with computer databases will recognize that, although one embodiment of database 1A07 is described as including various datasets 4C07, 4C09, including datasets 4C11, 4C13, 4C15, 4C17, and 4C19, as well as datasets 4C21, including datasets 4C23, 4C25, 4C27, and 4C29, the data described herein as belonging to these datasets may be rearranged into different arrangements of datasets, or some of the described datasets may be divided or combined into other datasets without departing from the true scope and spirit of the exemplary embodiment. It should also be understood that other data not specifically discussed in at least the presentation of FIG. 4C is described herein and may also form its own datasets that are included within plenoptic scene database 1A07 but that do not include those datasets shown in the figures herein. A careful reading of this disclosure will also reveal that not all of the data sets described in the figures herein must be present in database 1A07 for a system using scene codec 1A01 to perform useful or novel functions or otherwise provide any one of the many technical improvements described herein. Thus, plenoptic scene database 1A07 as described with respect to FIG. 4C herein should be considered as an example and not as a limitation of the exemplary embodiment, and many variations are possible without departing from the scope of the embodiments provided herein.

[0105] 5, 6, and 7, a series of three flowcharts are shown illustrating three variations of a general use case for a system using the scene codec 1A01, according to some exemplary embodiments. Each of the three flowcharts shows a series of connected process shapes, either boxes, ovals, or diamonds. It should be understood that each of these process shapes represents a high-level function of a dedicated computer process that includes some of the technical improvements provided in the exemplary embodiments, and all of the shapes and their interconnections then collectively further describe the technical improvements specified herein. The diamond shape represents a function that determines important branching decisions for the system regarding the processing of a user's request within the general use case. In general, each of the shapes can be understood to represent an executable instruction set for processing on a computing device, although these computing devices may be any arrangement of many possible variations, such as a single CPU having multiple cores, each of which performs one or more functions, or multiple CPUs having single or multiple cores, and these multiple CPUs may be distributed over any type of network, as previously described. Also, as previously explained, there are specific computer operations for certain novel techniques described herein that may be further optimized using some form of hardware-specific processing unit, e.g., an FPGA or ASIC or other well-known hardware executing code, often referred to as embedded code. In particular, some exemplary embodiments preferably include a spatial processing unit (SPU) 1A09 (see FIG. 1A), which is a set of operations that run on an embedded system, possibly comprising customized digital circuitry optimized for the key plenoptic scene processing functions described in this disclosure.

[0106] 5, 6, and 7 collectively, the exemplary embodiments provide significant advantages for scene model reconstruction, distribution, and processing; it should be understood that, traditionally, most codecs for transmitting image information, such as movies, do not offer many of the advantages described herein, such as freeview and sub-scene-on-demand, nor the transmission of non-visual scene data (such as scene metric, matter field, or light field), e.g., serving scene data on behalf of human and autonomous users who each access the same model for different forms and aspects of the scene data. There are several systems for transmitting scene models, including virtual reality (VR) systems in particular, but VR systems are typically based on computer modeling that lacks significant realism, such as the real-world scene shown in FIG. 4A. While there are other systems for transmitting reconstructed real-world scenes within a scene model, the system herein provides a unique plenoptic octree representation of the real scene, where the matter field 2B09 and light field 2B11 are separated to a greater degree than in existing systems, and the organization of the matter field 2B09 and light field 2B11 representations, among other things, provides a significant technical improvement over the basic functionality that enables real-time or near-real-time access and consumption of large real-world reconstructed scenes. Thus, the system of the present invention provides a representation of the real scene and an organization of the representations such that the plenoptic octree database 1A07 enables the system to process large global scenes over a distributed network, where the scene may be undergoing intermittent or even continuous reconstruction. ,The fundamental transfer of information is not simply visual data as in ,traditional codecs, or even entire reconstructed models as in ,other state-of-the-art scene model codecs, but a just-in-time ,heterogeneous, asynchronous stream of plenoptic data.

[0107] As will become apparent with reference to Figures 5, 6, and 7 below, there are many ways in which a scene can be processed using two or more systems that use the coded scene 1A01; for example, a first system 1A01 resides on a server and provides scene model information on demand to either a human or autonomous consumer (client), who then receives and processes the scene model information using a second system 1A01 (see, e.g., Figure 5). In another variation, the first system 1A01 is used by a client who wants to capture what might be considered a "local" versus "global" scene, such as their car that was damaged in a recent hailstorm, or their house that was damaged in a storm or is simply ready to be sold. In this type of use case, the client system 1A01 is further adapted to include a scene sensor 1E09 to capture raw data of the local scene (see, in particular, Figure 6). This raw scene data can then be transmitted over the network to a second system 1A01 running on a server that is primarily responsible for reconstructing the local client scene and retransmitting the reconstructed sub-scenes and sub-scene increments back to the client system 1A01.In yet another variation, a first system 1A01 used by a client in a shared scene (such as a disaster site or industrial warehouse) and a second system 1A01 running on a network both share responsibility for reconstructing scene data captured by the client into sub-scenes and scene increments, and these reconstructed sub-scenes and scene increments are then shared between the first and second systems 1A01 via codec functionality, and the shared scene captured locally by the first system 1A01 is compiled (through cooperation between the first and second systems 1A01) into a larger scene model that is still available for sharing with a third system 1A01, which may then be local to the shared scene and also capture raw data, or may be remote from both the first and second systems 1A01 and therefore also remote from the shared scene.

[0108] Referring to Figures 5, 6, and 7 collectively, there are shown several rectangular shapes representing the encoder function of codec 1A11 (505, 515, 517 in Figure 5; 505, 515, 517 in Figure 6; 505, 517 in Figure 7), several shapes representing the decoder function of codec 1A11 (507, 511 in Figure 5; 507, 511 in Figure 6; 507, 511 in Figure 7), and several shapes representing application software 1A03 (501a, 501b, 503, 509, 513a-d in Figure 5; 501a, 501b, 503, 509, 513a-d in Figure 6; 501a, 501b, 503, 509, 513 in Figure 7). As those familiar with computer systems and software architectures will appreciate, deployed embodiments of the various operations and functions depicted in Figures 5, 6, and 7 have many variations, including the use of any one or more processing elements to perform any one or more functions, such as a CPU, GPU, or SPU 1A09, as defined herein, running as an embedded processor in communication with and supporting any of the functions of the codec 1A11 or application software 1A03. Therefore, the following depictions and specifications for Figures 5, 6, and 7 should be considered examples, not limitations, of the exemplary embodiments; the described functions may be further combined or further divided, and these functions in various combinations may be implemented and deployed in many variations without departing from the scope and spirit of the exemplary embodiments. It will also be apparent to those skilled in the art of software systems, networks, conventional compression, scene modeling, and the like, that some other functions have been omitted for clarity but may be obvious based on existing knowledge (such as the functions of the transport layer 1E03, depending on the type of network, if one is used). 7 also shows some connecting lines as being thicker than others, and these thick connecting lines (505-507, 517-507 in FIG. 5; 501b-505, 505-517, 517-507 in FIG. 6; 501b-505, 505-517, 517-507 in FIG. 7) represent the transmission of any combination of information from a plenoptic scene database 1A07, such as stream 2B13 generally described with respect to FIG. 2B, or the transmission of scene data 2B07 (see FIG. 2B) intended for scene reconstruction or annotation as captured by scene sensors 1E09, such as real cameras 2B05-1 and 2B05-2.

[0109] Referring now only to FIG. 5 , a flow diagram of one embodiment is shown that includes sharing a larger global scene model with a remote client, either human or autonomous, consuming any of the various types of scene model information as described herein, including, for example, freeview, freematter, freelighting, metric, traditional 2D, 3D, or 4D visualization, any of the relevant (five) sensor information, any auxiliary information or otherwise related scene model information contained within or associated with plenoptic scene database 1A07 (e.g., database 1A07 includes URL links embedded within the spatial scene that connect to other Internet-accessible content, such as current weather conditions, supporting videos, product information, etc.). In this exemplary embodiment, it may be assumed that the global scene model is not only remote from the consuming client, but also prohibitively large, such that the entire global scene model cannot simply be transferred to the client. For ease of illustration, the flow of FIG. 5 is described with respect to only one of many possible specific use cases: a human user requesting a city tour as a remote client with respect to a global repository of city tours of scene models made available on a server.

[0110] 5, the human client operates a client UI (user interface) 501 executed by application software 1A03 running on a client system 1A01, preferably a mobile device. Also included within or in communication with client system 1A01 are zero or more sensors 1E09 (see FIG. 1E) for sensing at least some data about the human user that can be used, at least in part, to determine any of the user's requests. Exemplary sensors include a mouse, joystick, game controller, webcam, motion detector, etc., and exemplary data preferably includes data that explicitly or implicitly indicates a desired scene movement or viewpoint change, including a direction, path, trajectory, or the like, relative to a tracked current position / viewpoint within an ongoing scene that can be used, at least in part, by system 1A01 to assist in determining a viewpoint change and / or the next scene increment for the current subscene, as will be described in more detail shortly. The client system 1A01 further comprises one or more sensor outputs 1E11 (FIG. 1E) for sending data to a human user, for example, a 2D display, a VR headset, or a holographic display.

[0111] In a first step of this example, a human user accesses client UI 501 to determine a global scene of interest (SOI) 501a. For example, the options are multiple tourist destinations around the world, including major cities around the world. For example, the user selects to take a city tour of Prague. Operation 501a communicates with and provides a scene index from global SOI operation 503. For example, after the user selects to take a virtual city tour of Prague, operation 503 provides an index of multiple possible tours (and thus scene entry points with connected paths, see in particular model translation 4C25 of FIG. 4C ) along with associated auxiliary information such as images, short videos, customer ratings, critic ratings, text, URL links to websites, hotel and restaurant information, and websites. This auxiliary information, along with other scene index information, may be transmitted using well-known conventional techniques and codecs based on the type of data and therefore does not need to be included in plenoptic stream 2B13 (see FIG. 2B ). While operation 503 could be performed on a server-side 1E05 system 1A01 that houses or has access to a plenoptic scene database 1A07 (see FIG. 2B), operation 501a, as well as other processing of client UI 501, is preferably performed on a client-side 1E07 system 1A01, including determining an initial subscene within a global SOI operation 501b. In operation 501b, the user reviews and selects a subscene to enter into the scene model from the scene index provided by operation 503; for example, the user selects a city-cathedral tour starting from the narthex of St. Clement's Cathedral.

[0112] 5, operation 501b communicates the extraction of an initial subscene from global SOI model operation 505 and transmits an indication of a user-selected subscene, e.g., the narthex of St. Clement's Cathedral, to operation 505. Based at least in part on the user-selected subscene and at least in part on the determined subscene buffer size information, operation 505 accesses plenoptic scene database 1A07 to determine a set of at least initial matter field 2B09 and light field 2B11 data to provide to client system 1A01 as a first independent subscene. In one embodiment, operation 505 is solely a function of scene codec 1A11, and operation 505 communicates with operation 501b either directly or through an intermediate system component such as application software 1A03. In another embodiment, application software 1A03 substantially provides more than a communication service between operations 501b and 505, with software 1A03 implementing, for example, the portion of operation 505 primarily responsible for determining buffer sizes, and then causing scene codec 1A11 to extract and then transmit the initial sub-scene by invoking various application program interface (API) calls to scene codec 1A11. In any of these or other possible embodiments possible and understood by those familiar with software systems, processes running as part of scene codec 1A11 may then invoke various API calls to SPU 1A09.

[0113] In yet another embodiment, the scene solver 1A05 is invoked, for example by either the application software 1A03 or the scene codec 1A11, for example when determining a preferred buffer size, and the scene solver 1A05 executes deterministic or non-deterministic (e.g., statistical or probabilistic) algorithms, including machine learning algorithms, to determine or predict the buffer size, preferably based at least in part on auxiliary information 4C21 (see FIG. 4C) contained within the database 1A07, the auxiliary information 4C29 being particularly useful as a basis for performing machine learning based at least in part on data indicating previous buffer sizes and scene motion as logged for previous client sessions of other users accessing either the same or different subscenes. As will be appreciated by those familiar with the art of software systems and architectures, these same deterministic or non-deterministic (e.g., statistical or probabilistic) algorithms, including machine learning algorithms, may also be functions of scene codec 1A11, SPU 1A09, application software 1A03, or even some other components not specifically described herein but which will become apparent based on the description herein; for example, a system using scene codec 1A01 may incorporate another scene-based learning component implemented, for example, using any of the dedicated machine learning hardware currently available, known in the marketplace, or which will become known in the future. Equipped with components.

[0114] At least one technology company, known in the marketplace as NVIDIA, is currently offering a technology referred to as an “AI chip,” which is part of what is referred to as “Infrastructure 3.0” and is implemented on a dedicated GPU that further includes what NVIDIA refers to as “tensor cores.” The disclosure herein provides novel representations and organization of representations of plenoptic scene models, including auxiliary information 4C21, which is not traditionally considered scene data, but rather data such as model usage history 4C29 that directs how the scene model is used in various ways, either by humans or automata. As those familiar with machine learning will appreciate, while the exemplary embodiments provide several novel approaches to implementing scene learning, other approaches and implementations may be apparent, particularly with respect to buffer size determination, and these implementations may be software running on general computing hardware and / or dedicated machine learning hardware; all of these solutions are considered to be within the scope and spirit of the present disclosure.

[0115] Still referring to FIG. 5, a description is given of at least efficient (just-in-time) sub-scene extraction based on some determined or provided scene entry point and some determined or provided scene buffer size or some other form of information that predictably constrains the sub-scenes to be extracted from the entire SOI (e.g., global) scene model, whereby the extracted sub-scenes for the spatial buffer substantially ensure a maximally continuous user experience provided with a minimal amount of provided scene information. After determining plenoptic scene data representing the user's selected sub-scene from database 1A07, the plenoptic scene data is transmitted as an asynchronous just-in-time stream 2B13 of any combination of matter field 2B09 and light field 2B11 data contained within database 1A07, and stream 2B13 is received by the client-side 1E07 system using scene codec 1A01 for processing into sensor output, such as images and corresponding audio, that is provided to user 2B17 through sensor output device 2B19, which can provide 2π to 4π freeview manipulation to a human user, output device 2B19 being a specific example of any sensor output device 1E11 available through client UI 501.

[0116] As a first step in receiving stream 2B13 by the decoder included in codec 1A01, a function is performed to insert the following scene data into client SOI model 507, resulting in a reconstruction or update of the client SOI (i.e., plenoptic scene database 1A07) that mirrors, but is not equivalent to, the global scene model (i.e., plenoptic scene database 1A07) from which the sub-scene was extracted and provided. It is important to note that it is possible, and is considered within the scope of the exemplary embodiments, for the provided stream 2B13, which substantially contains plenoptic scene model data, to be translated into requested user data without first being placed in the client ("local") database 1A07, or even ever being placed in the client database 1A07, the scene translation being performed, for example, via a rendering and presentation step into freeview or other scene data that satisfies the user request. However, what is preferred, and shown herein to provide significant advantages, is that by initially or incrementally restructuring the client database 1A07, and not simply translating the stream 2B13 into requested scene data such as user freeview visualization, it is possible to allow ongoing client-side based scene data provision substantially independent, or at least quasi-independent, from the global scene model, where from time to time the local client scene database 1A07 is updated based on user requests. It is necessary to add or "grow" further, and such growing is referred to as providing a sub-scene increment, which will be explained shortly.

[0117] Still referring to FIG. 5, those familiar with software systems and architectures will understand that operation 507 is preferably implemented as part of the decoder within scene codec 1A11, and that at least a portion of operation 507 can be performed by application software 1A03 implementing client UI 501. For example, application software 1A03 could include, in operation 509, providing instructions and modifications to client UI 501 prior to the actual provision of the requested scene data. In traditional scene processing, operation 509 involves what is commonly referred to as rendering when the requested data is, for example, a free-view visualization. With respect to current rendering techniques, by utilizing a plenoptic scene database 1A07, exemplary embodiments increase free matter and free lighting options, providing even more realistic free views. As will be understood by those familiar with software architectures and based on a careful reading of this disclosure, both insert operation 507 and provide-request-data operation 509 may invoke various application interface (API) calls to scene processing unit (SPU) 1A09. Unlike conventional codecs, the exemplary embodiment provides confirmation 511 that the scene data has been received by the scene data server, which is an especially important feature when considering that future provided scene increments depend on the independent sub-scene that was initially provided, as well as any scene increments that are subsequently provided.

[0118] Still referring to FIG. 5, it is important to note that client SOI database 1A07 may be sufficient to provide any and all of the SOI data required in operation 509. The extent to which an initial, independent subscene is sufficient to satisfy all future data requests is proportional to the size of the initial subscene and inversely proportional to the range of scene data to be requested. As the range of requested or anticipated scene data increases, the burden and cost of transmitting an anticipated initial subscene with sufficient scene buffers eventually becomes prohibitive. For example, if the initial subscene is the narthex of St. Clement's Cathedral, from which the user is only expected to enter the cathedral and stand in the main hall, the size of the subscene may be limited. However, if the user is expected to enter the cathedral or cross the street into another building, the subscene must necessarily increase in size. Thus, in an exemplary embodiment, an initial sub-scene includes an intelligently determined scene buffer that provides for balancing anticipated user demand for up to a certain amount of scene data with the need to minimize transmitted scene data, thereby reducing perceived scene or UI lag, after which the system provides for transmitting further increments of the sub-scene from the global model to satisfy further demands or anticipated further demands based on either explicit or implicit user instructions. Again, this balance is preferably based on machine learning and other deterministic techniques, based at least in part on a history of similar user demands, such that the most continuous user experience is provided with a minimal amount of initially provided, and incrementally provided, scene information.

[0119] As will be apparent to those familiar with various types of prediction systems, as the "look-ahead" (future) time increases, the number of possible scene movement variations increases geometrically, or even exponentially, rather than linearly. For example, if a user is given an initial subscene of the Narthex of St. Clement's Cathedral, a look-ahead time of one minute versus one hour will increase the scene buffer size at least geometrically, such that if the calculated buffer size is X for one minute, the buffer size for one hour Y will be substantially larger than 60*X. In this regard, another important technical advantage of certain embodiments is that both the representation of the plenoptic scene model and the organization of these representations can be shown to significantly reduce the processing time required to extract any initial subscene or scene increment, given any chosen buffer size, with respect to currently known scene processing techniques. Thus, careful consideration of the balance trade-off relationship makes it clear that for the same system response time, a larger initial sub-scene buffer can be supported, with significant time savings in both extracting and processing sub-scenes or scene increments, and smaller sub-scene increment buffers can be supported in favor of more frequent scene increments, with the smaller, more frequent approach reducing the amount of scene data actually transmitted so that user request look-ahead time is reduced.

[0120] Still referring to FIG. 5, the remaining process client request operation 513 and log consumption operation 515 track the user's scene usage and provide intelligent increments of the client SOI model when it is determined or predicted that the initial subscene lacks sufficient scene data to satisfy current or possible future user requests. As the user interacts with the UI 501, for example, receiving updated scene data, these interactions provide an indication of the value and consumption of scene data. Additionally, the client UI 501 preferably allows the user to express indications that can be interpreted as requests for more scene data, such as by moving a mouse, joystick, game controller, or VR headset, which indications are detected using sensors 1E09. Either one or more of these indications of usage or predicted usage enables the system to track user consumption within the client SOI model as operation 513a, and the tracked usage is stored in either the server or client plenoptic scene database 1A07 as model usage history 4C29 (see FIG. 4C).

[0121] When a user display is processed by client UI 501, process client request operation 513 includes operation 513b for determining whether any of the user instructions can be interpreted as a next request for scene data, and then determining whether the next request can be fulfilled solely based on the scene data already contained within the existing local client SOI model. If the next request can be fulfilled solely based on the existing client SOI model, the requested scene data is provided to the user by operation 509. If the next request cannot be fulfilled solely based on the existing client SOI model, operation 513c determines whether the next request increments to an existing subscene (or subscenes) within the client SOI model, or whether the next request is for an entirely new (and therefore independent) subscene. If the request is for a new subscene, operation 513c transfers control to, or otherwise invokes, client UI 501, effectively determining what the new subscene being requested is (e.g., switching to St. Lawrence Cathedral), or even whether the user may be requesting an entirely new global scene (e.g., switching to a tour of Venice). If the request is not for a new subscene, but rather to continue exploring an existing subscene in a manner that requires incremental additions to the current subscene, operation 513d determines a next increment vector for the subscene. The next increment vector represents one of the dimensions of the scene model, such as spatiotemporal extent, spatial detail, light field dynamic range, or matter field dynamic range, and is any information that indicates the minimum range of new scene data required to satisfy the user's request.When determining the vector, operation 513d preferably has access to the user history tracked by log consumption operation 515, and the vector determined to minimally satisfy the user's requirements, along with the usage history (of the current and all other tracked users), can be used to recommend the next scene increment and increment buffer size. These may be combined for use at least in part by the system when determining the buffer size, where again the buffer size extends the scene increment beyond that of the minimum filled vector to include anticipated "look ahead" sub-scene usage.

[0122] Still referring to FIG. 5, as will be understood by those familiar with software systems, other arrangements of operations are possible while performing the preferred step of tracking and logging at least a portion of user instructions and scene model consumption. Additionally, other arrangements of operations are also possible while determining whether the user has explicitly or implicitly requested additional scene data, and if so, whether this additional scene data already exists in the client scene model database 1A07. Still other arrangements of operations are possible in determining the next scene increment and buffer size when the additional scene data does not already exist but is an extension to a subscene already present in the client SOI database 1A07. As such, the functionality provided in the present illustration for processing client requests 513 and logging consumption 515 should be considered examples, not limitations on the embodiment. Furthermore, operations 513, 513a, 513b, 513c, 513d, and 515 can all be implemented to execute concurrently on their own processing elements or sequentially on a single processing element. For example, tracking and logging of user consumption operations 513a and 515 may be performed in parallel with the sequential processing of request tracking operations 513b and 513c. In another consideration, log consumption operation 515 may be performed on both client system 1A01 to update client SOI database 1A07 usage history 4C29 and on server system 1A01 to update global SOI database 1A07 usage history 4C29. Also, some or all of the determined next increment vector for subscene 513d (including buffer size determination) may be performed on either client system 1A01 or server system 1A01.

[0123] It is important to note that user usage of the scene model is tracked and aggregated, client systems initially attempt to fulfill requests for new scene data based solely on client SOI models currently located on or accessible to client system 1A01, and when additional sub-scene increments are required from the global SOI model, a calculation is made to determine the minimum amount of sub-scene increments required to provide the most continuous user experience in terms of both the expected amount of look-ahead usage and the determined quality of service (QoS) level, the determination of the expected amount of look-ahead usage being based at least in part on the tracked usage history.

[0124] Referring now to Figure 6, flowcharts of some exemplary embodiments are shown that build on the description of Figure 5 but address the modified case where a client first creates a scene model or updates an existing scene model rather than accessing an existing model. With respect to Figure 5, operations 507, 509, 511, 513 (including 513a, 513b, 513c, and 513d), 515, and 517 are substantially as described in Figure 5, and therefore only minimal additional description will be provided with respect to this Figure 6. An exemplary use case for Figure 6 is a user working with a mobile device, such as a cell phone, that is a system that uses scene codec 1A01 to create (or update) a scene model of the hood of their car that has been damaged in a hailstorm. Many use cases other than the exemplary use case for Figure 6 are applicable to each of the flowcharts of Figures 5, 6, and 7, such as modeling damage to an automobile. For example, this use case of FIG. 6 may be an insurance policy that requires a user to capture a scene model of either an asset or property, including, for example, a home, or perhaps a model of an asset or property. It is equally applicable if you are a company or an agent such as a real estate agent. Businesses and construction companies can also use the same use case to capture and share with others scene models of important assets or properties, which can be of any scale and have almost unlimited visual detail.

[0125] Still referring to FIG. 6, client UI 501, like FIG. 5, includes both sensor 1E09 and sensor output 1E11; one difference in the use case is that, for the one illustrated in FIG. 6, sensor 1E09 includes one or more sensors for sensing an asset, real estate, or some other form of real scene to be reconstructed into a plenoptic scene model and database 1A07. A typical sensor 1E09 would be one or more real cameras, such as 2B05-1 and 2B05-2 shown in FIG. 2B, but may otherwise be any of a number of sensors and sensor types. Client UI 501 allows a user to instantiate a new client SOI or select an existing client SOI, which could be, for example, their car or even their car hood. In an exemplary embodiment, for example, a plenoptic scene model of either the user's asset, real estate, or some other scene may pre-exist, or the user may specify that they are creating (instantiating) a new scene model. For example, if the user is a rental car company and a renter has just returned a rental car to a scanning platform, client UI 501 might enable the company to scan a barcode from the renter's consent form and then use this information, at least in part, to recall an existing plenoptic scene model of the same vehicle from before the rental began. Thus, the car could be re-scanned, perhaps by a device that is autonomous but still considered a sensor 1E09, and the newly scanned image or other sensed data could be used to update the vehicle's existing scene model.

[0126] Using this approach, as previously noted, the plenoptic scene exists in all four dimensions, including the three spatial dimensions as well as the temporal dimension. Thus, any refinement of an existing plenoptic scene model either permanently alters the plenoptic scene such that the original baseline matter field and light field data is overwritten or otherwise lost, or the refinement can be organized as an additional representation associated with either a specific time, such as April 25, 2019, 10:44 AM EST, or an event name, such as Rental Agreement 970445A. If the matter field and light field are then organized in the time dimension of the plenoptic scene database 1A07, then it is at least possible to 1) create either scene data based on earlier or later times / events for reconstruction and refinement of any real scene, 2) measure or otherwise describe the differences between two points in time within the plenoptic scene database 1A07, and 3) catalog the change history of the plenoptic scene database 1A07 filtered by any of the features of the database 1A07, such as some or all of any part of the scene model including the matter field and light field.

[0127] In the example of a user scanning the hood of their car, for example, to document and measure hail damage, it is anticipated that rather than accessing a remote database of car plenoptic scene models and instantiating a new model without a baseline plenoptic scene, the user first selects an appropriate baseline make and model for their car and then uses this as the basis for scanning their own unique data to rebuild and refine the baseline model. In this case, it is further anticipated that the client UI 501 will provide intelligent user-facing functionality, allowing the user to, for example, adjust the matter field associated with the baseline model, e.g., change the car's color to a custom paint color added by the user, or any similar type of difference between the baseline and the unique real scene. It is further anticipated that any portion of the matter field may be named or tagged, e.g., "car exterior," where this tag is auxiliary information 4C21 that may be considered a model extension 4C23 (see FIG. 4C). As can be seen with careful consideration, by providing a tagged baseline plenoptic scene model, the system provides significant means for creating and refining new custom scene models.

[0128] Some exemplary embodiments further provide multiple tagged plenoptic matter field and light field types and instances along with the baseline plenoptic scene model; for example, an automobile manufacturer might create various plenoptic matter field types representing the various materials used in the construction of one of its automobile makes, which again is represented as the baseline plenoptic scene. In this arrangement, an automobile salesperson could quickly modify the baseline vehicle to select different material (matter field) types to substitute within the baseline, such that translations 4C25 of these makes (see FIG. 4C ) can then be accessed as plenoptic scene models for exploration (such as the general use case of FIG. 5 ). As the careful reader will appreciate, there are virtually countless uses for each of the general use cases presented in FIGS. 5 , 6 , and 7 , and the use cases described specifically with respect to FIGS. 5 , 6 , and 7 should be considered by way of example, and not limitation, not to mention variations of other use cases described, implied, or otherwise apparent herein.

[0129] Still referring to FIG. 6, after establishing the client SOI as database 1A07, the information is transmitted to, for example, a second server-side system 1A01, where operation 503 instantiates / opens a global SOI model corresponding to the client SOI model. As will become apparent, the global model and the client model should be updated to include substantially identical new scene model information as captured by sensor 1E09 of client system 1A01, but there is no requirement that the global model and the client model be substantially identical with respect to scene data in any other way. Preferably, operation 503 is in communication with client UI 501, and after the global SOI model is instantiated or substantially instantiated, client UI 501 prompts the user and enables the user to begin capturing scene data (see 2B07 in FIG. 2B), such as photos or video of the damaged hood of the user's car. Once new scene data, such as images, is captured, the captured data is preferably compressed and transmitted using an appropriate conventional codec, such as a video codec for image data. The compressed new scene data is transmitted to the server-side system 1A01, and operation 505 unpacks the compressed new scene data and then uses it, at least in part, to reconstruct or refine the global SOI model, where reconstruction further references an entirely new scene and refinement further references an existing scene (such as a car plenoptic scene model that already existed but is now being updated). While reconstruction operation 505 is creating a new portion of the global scene model, server-side system 1A01 operation 517 provides the next sub-scene increment (of the new portion) from the global SOI model to be communicated to the client-side system 1A01. After receiving the new sub-scene increment, client-side operation 507 inserts the next scene data (sub-scene increment) into the client SOI model.Careful consideration reveals that the client-side system 1A01 is capturing data while the server-side system 1A01 is doing all of the scene reconstruction, which can be computationally intensive, and the server-side system 1A01 effectively offloads this computationally intensive task from the client-side system 1A01.

[0130] 6, once the client SOI model has been constructed based on the received sub-scene increments provided from the global SOI model (reconstructed based on client sensor data), client system 1A01 can provide requested SOI data operation 509 and provide scene model information to the user through UI 501 in accordance with all of the previous descriptions related to processing client requests 513. Also, as previously described, client system 1A01 preferably tracks user instructions and usage in operation 513a for logging with global SOI database 1A07 through operation 515.

[0131] 7, which builds upon FIGS. 5 and 6 but addresses the modified case where a client first creates a scene model or updates an existing scene model and then captures local scene data for the real scene, both client-side system 1A01 and server-side system 1A01 can each reconstruct the real scene and provide scene increments, as opposed to FIG. 6 where the real scene was reconstructed into a scene model only on server-side system 1A01. With respect to FIG. 5, operations 507, 509, 511, 513 (including 513a, 513b, 513c, and 513d), 515, and 517 are substantially as described in FIG. 5, and therefore, only minimal additional description will be provided with respect to this FIG. 7. With respect to the description of FIG. 6 , operation 501 (including 501a, 501b and sensors 1E09, 1E11) as well as operation 503 are substantially as described in FIG. 6 , and therefore, minimal additional description will be provided with respect to this FIG. 7 . An exemplary use case of FIG. 7 is a user working with a mobile device, such as an industrial tablet, that is a system that uses scene codec 1A01 to create (or update) a scene model of a disaster relief scene (with as many users or autonomous vehicles (not shown) acting as clients 1A01 to capture scene sensor data substantially simultaneously). Many use cases other than the exemplary use case for FIG. 7 are applicable to each of the flowcharts of FIGS. 5 , 6 , and 7 , such as modeling a disaster scene. For example, this use case of FIG. 7 is equally applicable to users capturing a scene model of any shared scene, such as workers in an industrial environment or commuters and pedestrians in an urban environment.

[0132] 7, similar to FIG. 6, the server-side system 1A07 receives the compressed raw data as captured by or at least partially based on the client-side system 1A07, and then reconstructs it into a global SOI model or refines an existing global SOI model in operation 505. The current use case also includes an operation 517 on the server-side system 1A07 for sending the next sub-scene increment from the global SOI model to operation 507 of the client-side system 1A01, which then uses the provided sub-scene increment to update the client SOI model and provide the requested SOI data in operation 509 to the user via UI 501. 6, the client-side system 1A01 also includes an operation 505 for reconstructing the client SOI model, and then an operation 517 for providing the sub-scene increment to the server-side system 1A01, which then includes an insertion operation 507 for reconstructing or refining the global SOI model. Both the client-side system and the server-side system include an operation 511 for validating the received and processed scene data. Preferably, at any given time, any location captured by or under the command of any one or more client-side systems 1A01 is validated, preferably under the direction of, and preferably in shared communication with, application software 1A01 running on either the server-side system 1A01 or the client-side system 1A01. It is important to note that for given real scene data, either or both of the server-side system 1A01 or the client-side system 1A01 can be instructed by their respective application software 1A01 to reconstruct any of the real scene data and then share the reconstructed scene data as scene increments with any of the other systems 1A01, or not to reconstruct any of the real scene data and then receive and process any of the scene increments reconstructed by any of the other systems 1A01.

[0133] The value of this operational arrangement becomes even more apparent in larger use cases with multiple client-side systems 1A01 and even multiple server-side systems 1A01. Those familiar with computer networks and servers will understand that application software 1A01 communicating among multiple systems 1A01 performs scene reconstruction and distribution load balancing. Some of the clients may be users with mobile devices 1A01, while others may be autonomous land-based systems 1A01 or airborne systems 1A01. Each of these different types of clients 1A01 is expected to have different computing and data transmission capabilities. It is also expected that each of these individual clients 1A01 will have different possible real scene sensor 1E09 ranges and needs for plenoptic scene data. The software 1A01's load balancing decisions take into account, at least in part, any one or any combination of the multitude of collected sensor 1E09 data overall, the priority for scene reconstruction, the availability of computing power across all server-side and client-side systems 1A01, the data transmission capabilities across the network 1E01 (see FIG. 1E) between the various systems 1A01, and even the anticipated and on-demand demand for scene data by each of the systems. Similar to the use cases of FIGS. 5 and 6, the use case of FIG. 7 also preferably captures instruction and scene data usage on multiple client-side systems 1A01 and logs this instruction and usage data in any of the appropriate scene databases 1A07 across the multiple systems 1A01; the machine learning (or deterministic) component of the exemplary embodiment can then access this logged scene usage to optimize load balancing, among other benefits and uses already previously described.It is also anticipated that server-side scene reconstruction metrics, such as, but not limited to, the type and amount of raw data received as well as variations in scene reconstruction processing time, will be additionally logged alongside client-side usage, and this additional server-side logging will also be used at least in part by the machine learning (or deterministic) component to determine or provide load balancing needs. Scene Database

[0134] Figure 8 shows a kitchen scene with key attributes associated with mundane (everyday) scenes: transparent media (e.g., glass pitcher and window pane), highly reflective surfaces (e.g., metal pot), microstructured objects (e.g., potted plant and outdoor tree to the right), featureless surfaces (e.g., cabinet door and dishwasher door), and virtually infinite volumetric coverage (e.g., outdoor space visible through a window). The scene in Figure 8 is an example scene that may be included in a scene database 1A07 that the system processes in various use cases using scene codec 1A01. One important aspect of such processing is the volumetric and directional (angular) subdivision of space into addressable containers that serve to store elements of the scene's plenoptic field.

[0135] FIG. 9 shows an exemplary representation of spatial containers with respect to volume and direction. A voxel 901 is a container that bounds a volumetric region of scene space. A solid angle element 903, known by the abbreviation "seil," bounds an angular region of space originating at the vertex of the seil (seil 903 is shown from two different perspectives to help convey its 3D shape). Although seil 903 is shown as a pyramid of finite extent, the seil may extend infinitely outward from its origin. Containers used in embodiments may or may not have the exact shape shown in FIG. 9. For example, non-cubic voxels or seils that do not have a square cross-section are not excluded from use. Further details regarding efficient hierarchical arrangements of voxels and seils are provided below with reference to FIGS. 21-65.

[0136] FIG. 10 shows an overhead plan view of an exemplary scene model 1001 of a mundane scene in an exemplary embodiment different from the embodiment described with reference to FIG. 4B above. The embodiment described here focuses on aspects related to the extraction and insertion of sub-scenes, as opposed to the broader focus of the embodiment of FIG. 4B on overall codec operation. A plenoptic field 1003 is bounded by an outer scene boundary 1005. The plenoptic field 1003 contains plenoptic primitive entities (“plenoptic primitives,” or simply “primitives”) that represent the matter and light fields of the modeled scene. The plenoptic field 1003 is volumetrically indexed by one or more generally hierarchical arrangements of voxels and directionally indexed by one or more generally hierarchical arrangements of spheres. Matter within the plenoptic field is represented as one or more media elements (“medials”), such as 1027, each contained within a voxel. A voxel may be empty, in which case it is said to be "void" or "void-type." Voxels that lie outside the outer scene boundary are void-type. These void voxels, by definition, do not contain plenoptic primitives, although they may point (point to) entities other than plenoptic primitives. Light within the plenoptic field is represented as one or more radiant elements ("radials"), such as 1017, each contained within a radiant element that is located in a medial (in a voxel that contains a medial).

[0137] Light fields in a medium (including those representing only negligible light interactions) include four component light fields: the incident light field, the response light field, the emitted light field, and the window light field. The incident light field represents light carried from other media, including media immediately adjacent to the medium of interest. The response light field represents light emitted from a medium in response to interaction with the incident light. The emitted light field represents light emitted from a medium due to some physical process other than interaction with the incident light (e.g., conversion from another form of energy, such as a light bulb). The window light field represents light entering a medium due to unspecified processes outside the plenoptic light field. An example of this is a window light field representing sunlight injected at the scene boundary outside the plenoptic field when the plenoptic field does not extend far enough to volumetrically represent the sun itself as a radiant source. It is important to note that in some embodiments, the window light field may be composed of multiple window light subfields, which may be considered "window layers," representing, for example, light from the sun in one layer and light from the moon in another layer. Media interact with the incident window light field in the same way that they interact with the incident light field. In the following discussion of BLIF, statements regarding the incident light field apply equally to the window light field (the response light field is determined by both the incident and window light fields).

[0138] In the plenoptic field 1003, the medium 1027, like all mediums, has an associated BLIF. The BLIF typically contains radiometric information and / or BLIF represents the relationship between properties of interest, such as incident and response radii, in a quasi-steady-state light field, or properties that include spectral and / or polarization information. In the context of certain exemplary embodiments, BLIF is useful because it practically represents the interaction of matter with light without resorting to computationally intensive modeling of such interactions at the molecular / atomic level. In a highly generalized BLIF representation, the response-to-incidence ratio of properties of interest can be captured in a sampled / tabular form with an appropriately fine Seymour granularity. When practical, embodiments may use one or more compressed BLIF representations. One such representation is a low-dimensional model that yields response luminance as an analytical function of incident illuminance, parameterized over the direction, spectral bands, and polarization states of the incident and response light. Examples of such low-dimensional models include traditional analytical BRDFs, such as the Blinn-Phong and Torrance-Sparrow microfacet reflectance models. Such compression of BLIF information is well understood by those skilled in the art and is used in some embodiments of the present invention to compress and expand BLIF data. One embodiment may allow for the representation of spatially (volumetrically) varying BLIFs, where one or more BLIF parameters vary over a range of volumetric scene regions.

[0139] The outer scene boundary 1005 is a closed, piecewise continuous, two-dimensional manifold that separates media within the plenoptic field from void voxels outside the plenoptic field. The void voxels lie within the inner boundaries 1007 and 1009. The scene model 1001 does not represent light transport outside the outer scene boundary or within the inner boundaries. Media lying adjacent to void voxels are known as "boundary media." The light field of a boundary media may include an incident light field carried from other media in the plenoptic field, as well as a window light field representing light incident into the plenoptic field due to unspecified phenomena outside the plenoptic field. The window light field at one or more boundary voxels in a scene can generally be thought of as a four-dimensional light field volumetrically positioned on the piecewise continuous manifold defined by the boundary.

[0140] An example of an outer scene boundary is the sky in a mundane outdoor scene. In the plenoptic field of a scene model, there is air media up to some reasonable distance (e.g., the parallax resolution limit), and void voxels beyond that. For example, clear sky or moonlight is represented by a window light field in the air media at the outer scene boundary. Similarly, light from unspecified phenomena inside an inner scene boundary is represented by a window light field in the media bordering the inner scene boundary. An example of an inner scene boundary is the boundary of a volumetric region for which a complete reconstruction has not yet been performed. The 4D window light field of the adjacent boundary media contains all (currently) available light field information for the bounded void region. This may change if a subsequent reconstruction operation succeeds in discovering a model of the matter field, which places the previous window light field within the previous void region as incident light carried from the newly discovered (resolved) media, as described below.

[0141] In addition to the plenoptic field 1003, the scene model 1001 includes other entities. Media 1027 and other nearby non-air media are referred to in various groupings that are useful in the display, manipulation, reconstruction, and other potential operations performed by a system using the scene codec 1A01. One grouping is known as a feature, where plenoptic primitives are grouped together by some pattern of their characteristics of interest, possibly including spatial orientation. 1029 is a shape feature, meaning that the constituent media of a feature are grouped by virtue of their spatial arrangement. In one embodiment, a system using the scene codec 1A01 may consider a feature 1027 to be a protrusion or bump for some purpose. 1021 are BLIF features, meaning that the constituent media of the features are grouped based on their associated BLIF patterns. A system using scene codec 1A01 may consider features 1021 to be contrast boundaries, color boundaries, boundaries between materials, etc.

[0142] A plenoptic segment is a feature subtype defined by similarity (rather than any pattern) in some set of characteristics. Segments 1023 and 1025 are matterfield segments, in this case defined (within some tolerance) by the uniformity of the BLIFs of each segment's medial. Objects such as 1019 are matterfield feature subtypes defined by their recognition by one or more humans as "objects" in natural language and cognition. Example objects include a kitchen table, a glass window, and a tree.

[0143] Camera path 1011 is a feature subtype that represents a 6DOF path traced by a camera observing plenoptic field 1003. Aspects of potentially useful embodiments of camera path include kinematic modeling and spherical linear interpolation (slerp). At locations along camera path 1011, there is a focal plane, such as 1013, at the camera viewpoint where the light field is recorded. The collection of radii incident on the focal plane is typically referred to as an image. Exemplary embodiments do not limit the camera representation to having a planar array of pixels (light-sensing elements). Other arrangements of pixels are similarly representable. Focal plane 1013 records light exiting object 1019. Features can be defined in the matter field, the light field, or a combination of the two. Item 1015 is an exemplary feature of the light field, in this case including radii at focal plane 1013. The pattern of radii in this case defines the feature. In traditional image processing terms, a system using a scene codec can consider 1015 to be a feature that is detected as a 2D pattern within the image pixels.

[0144] FIG. 11 is a block diagram of a scene database in an embodiment different from that described with reference to FIG. 4C above. The embodiment described here focuses on aspects related to the extraction and insertion of subscenes, as opposed to the broader focus of the embodiment of FIG. 4C on overall codec operation. Scene database 1101 includes, among other entities not shown, one or more scene models, a BLIF library, an activity log, and camera calibration. Scene model 1103 includes one or more plenoptic fields, such as 1105, and a collection of features, such as 1107, potentially including features of type segment (e.g., 1109), object (e.g., 1111), and camera path (e.g., 1113). In addition, one or more scene graphs, such as 1115, point to entities within the plenoptic fields. Scene graphs may also point to analytical entities not currently manifested within the plenoptic fields. A scene graph is arranged in a hierarchy of nodes that define spatial and other relationships between referenced entities. Multiple plenoptic fields and / or scene graphs typically exist together in a particular single scene model if the system using the scene codec expects them to register to a common spatiotemporal reference frame at some appropriate point in time. Without this expectation, multiple plenoptic fields and / or scene graphs will typically exist in separate scene models.

[0145] The BLIF library 1119 holds BLIF models (representations). As described above, the scene database may contain BLIFs in various forms, from spectrally polarized outgoing / incoming ratios to efficient low-dimensional parametric models. The BLIF library 1119 includes a material sublibrary 1125, which represents the light interaction and other properties of media that may exist in the matter field. Example entries in the material library 1125 include dielectrics, metals, wood, stone, fog, air, water, and the near-vacuum of outer space. The BLIF library 1119 also includes a roughness sublibrary 1127, which represents the roughness properties of media. Example entries in the roughness library 1127 include various surface microfacet distributions, grit categories for sandpaper, and the distribution of impurities in volumetric scattering media. Media in the plenoptic field may reference entries in the BLIF library or may have "locally" defined BLIFs that are not included in any BLIF library.

[0146] Activity log 1121 holds a log 1129 of sensing (including imaging) activity, a log 1131 of processing activity (including activity related to encoding, decoding, and reconstruction), and other related activity / events. Camera calibration 1123 holds correction parameters and other data related to the calibration of cameras used for imaging, display, or other analysis operations on a scene model.

[0147] Figure 12 shows a class diagram 1200 of the hierarchy of types of primitive entities in a plenoptic field. The root plenoptic primitive 1201 has subtypes, medial 1203 and radial 1205. Medial 1203 represents a medium in the matter field resolved to be contained in a particular voxel. A homogeneous medium 1209 is a medium that is uniform throughout that voxel in one or more properties of interest, within some tolerance. Examples of homogeneous media 1211 include suitably uniform solid glass, air, water, and fog. A heterogeneous medium 1211 is a medium that does not have such uniformity in the properties of interest.

[0148] A surfel 1225 is a heterogeneous medial having two distinct regions of different media separated by a piecewise continuous 2D manifold. The manifold has an average spatial orientation represented by a normal vector and, in an exemplary embodiment, a spatial offset represented by the closest approach point between the manifold and the volumetric center of the voxel containing the surfel. Subtypes of surfels 1225 include simple surfels 1227 and split surfels 1229. Simple surfels 1227 are similar to those described for their supertype, surfel 1225. Examples of simple surfels 1227 include the surface of a wall, the surface of a glass sculpture, and the surface of calm water. For split surfels 1229, on one side of the intra-medial surfel boundary, the medial is additionally split into two subregions separated by another piecewise continuous 2D manifold. An example of a split surfel 1229 is the region of a chessboard surface where a black square and a white square intersect.

[0149] A smoothly varying medium 1211 represents a medium in which one or more properties of interest vary smoothly across the volume range of the medium. Spatially varying BLIF will typically be employed to represent smooth changes in light interaction properties throughout the volume of the smoothly varying medium 1211. Examples of smoothly varying mediums 1219 include surfaces painted with a smooth color gradient, and areas where a thin layer of ground level fog gives way to clearer air above.

[0150] A radiel 1205 represents light in a scene's light field resolved to be contained in a particular ray. A radiel 1205 has subtypes: isotropic radiel 1213 and anisotropic radiel 1215. An isotropic radiel 1213 represents light in which one or more properties of interest, such as radiometric or spectral or polarization, are uniform across the directional range of the radiel. An anisotropic radiel 1215 represents light without such uniformity in the properties of interest. A split radiel 1221 is an anisotropic radiel with two distinct regions of different light content separated by a piecewise continuous one-dimensional manifold (curve). An example of a split radiel 1221 is a radiel containing the edge of a highly collimated light beam. A smoothly varying radiel 1223 represents light in which one or more properties of interest vary smoothly across the directional range of the radiel. An example of a smoothly varying radiance is the light from a pixel on a laptop screen, which shows a decrease in radiance as the exit angle shifts away from normal.

[0151] The image shown in Figure 13 is a rendering of a computerized model of a trivial scene: a real-world kitchen. Two 3D points of the kitchen scene are shown in Figure 13 for use in the following illustrations and discussion. Point 1302 is a typical point in the open space of the kitchen (a dotted line perpendicular to the floor is shown to clarify its placement). Point 1304 is a point on the surface of a marble counter.

[0152] The exemplary embodiments described herein can realistically represent scenes such as that shown in FIG. 13 because the techniques according to the exemplary embodiments model not only the matter field of the scene but also the light field and the interaction between the two. Light entering and leaving a volumetric region of space is represented by one or more radii that enter or exit a specified point within the region representing the space. A collection of radii is therefore called a "point" light field, or PLF. The incoming and outgoing light of the PLF is represented by one or more radii that intersect with a specified region on a "surrounding" cube centered at the representative point.

[0153] This can be visualized by viewing a cube with light passing through the cube's surfaces, displayed on its faces, and to and from a central point. Such a "light cube" is 1401 in the image of FIG. 14. It is centered at point 1302 in FIG. 13. This light cube shows incident light entering point 1302. Thus, the light intensity shown at a point or area on the cube's surface is light from items and light sources in the kitchen (or beyond) passing through that point or area, which also intersects with point 1302, the center of the cube. FIG. 15 shows six additional exterior views of light cube 1401. The light cube can also be seen from inside the cube. The images in FIGS. 16, 17, and 18 show various views from inside light cube 1401.

[0154] A light cube can also be used to visualize light emanating from a point, i.e., an outgoing PLF. Such a light cube is shown in Figure 19, 1902, for point 1304, i.e., a point on a marble counter in a kitchen as shown in Image 13. The surface of the cube shows light emanating from points intersecting with the cube's faces or areas on the cube's faces. This is like looking at a point through a straw from all directions positioned on the surrounding sphere. Note that the surface of the bottom half of the light cube is black. This is because the center of the PLF is on the surface of an opaque material (marble in this case). No light exits the point in directions toward the counter's interior, and therefore those directions are black within the light cube.

[0155] Light cubes can also be used to visualize other phenomena. The image in FIG. 20 shows a light cube 2001. It illustrates the role of the BLIF function in generating an output PLF based on an input PLF. In this case, the input light is a single beam of vertically polarized light at input light element (radial) 2002. The output light resulting from this single light beam is shown on the face of light cube 2001. Depending on the details of the BLIF used, complex patterns of output light appear as shown in light cube 2001.

[0156] Some exemplary embodiments provide techniques for calculating the transport of light and its interaction with materials in a modeled scene. These and other calculations involving spatial information are performed in a spatial processing unit, or SPU. It uses a plenoptic octree, which consists of two types of data structures. The first is an octree. An example is a volumetric octree 2101, as shown in Figures 21 and 22. An eight-way partitioned hierarchical tree structure is used to represent a cubic region of space in a finite cubic universe. The top of the tree structure At the top is a root node 2103 at level 0 that exactly represents the universe 2203. The root node has eight child nodes, such as node 2105 at level 1, which represents a voxel 2205, i.e., one of eight equally sized, disjoint cubes that exactly fill the universe. This process continues at the next level in the same way to subdivide the space. For example, node 2107 at level 2 represents cubic space 2207. The octree portion of a plenoptic octree will be referred to as a "volume octree" or VLO.

[0157] The second data structure used in plenoptic octrees is the Seyel tree. A Seyel is a "solid angle element" used to represent a region of direction space projecting from an "origin". It is typically used as a container for radii, outgoing rays from the origin, or incoming rays striking the origin from the direction represented by the Seyel. A Seyel tree typically represents direction space for some region of volumetric space around the origin (e.g., a voxel).

[0158] The space represented by a Seyer is determined by a square area on the face of a cube. This cube is the "bounding cube" and is centered at the origin of the Seyer tree. This cube can be of any size and surrounds or bounds the Seyer tree. It simply specifies a particular geometric shape of Seyer within the Seyer tree, each of which extends an unlimited distance from the origin (but is typically used only within the volume represented by the plenoptic octree). Similar to an octree, a Seyer tree is a hierarchical tree structure whose nodes represent Seyer trees.

[0159] A Seyer tree 2301 is illustrated in Figures 23 and 24. The root node 2303 is at level 0 and represents all directions emanating from its origin, i.e., the point at the center of the surrounding cube 2403. While Seyer and Seyer trees encompass the unlimited volume of space extending from the origin, they are typically defined and usable only within the universe of their plenoptic octree, which is usually the universe of their VLO. As can be seen in Figure 23, the root node of a Seyer tree has six children, and all nodes in the following subtrees have four children (or no children). Node 2305 is one of the root's six possible child nodes (only one is shown). Node 2305 is at level 1 and represents all the space projecting from the origin that intersects with the face 2405 of the Seyer tree's surrounding cube. Note that when the center of a Seyer tree is at the center of the universe, its definition face becomes the face of the universe. When the Seyel tree is in a different configuration, its origin is in a different configuration in the plenoptic octree, and the surrounding cube moves around the origin. It is no longer a universe; since we only determine the direction of the Seyel tree relative to the origin, it could be any cube of any size with the origin as its center.

[0160] At the next level of subdivision, node 2307 is one of four level-2 child nodes of node 2305 and represents face square 2407, which is one-fourth of the associated face of the universe. At level 3, node 2309 represents the direction space defined by face square 2409, which is one of four equal divisions of square 2407 (one-sixteenth of face 2405). The hierarchical nature of Seyer trees is illustrated below in 2D in FIG. 41 for a Seyer tree 4100 originating at point 4101. Node 4102 is a non-root node at level n of the Seyer tree (the root has six child nodes). It represents a segment of direction space 4103. At the next level down, two of the four level n+1 nodes 4104 represent two Seyer trees 4105 (the other two represent the other two 3D nodes). At level n+2, nodes 4106 represent four regions 4107 in 2D (16 in 3D). The Seyer tree used in the plenoptic octree is called an SLT.

[0161] Note that, as with octrees, subdivision of a Seyel tree terminates (there are no subtrees) when the properties in the subtrees are fully represented by the properties attached to the nodes. This may be true when a sufficient level of resolution is reached, or for other reasons. Seyel trees, like octrees, can be represented in many ways. Nodes are typically connected by unidirectional links (parent to child) or bidirectional links (parent to child, child to parent). In some cases, subtrees of an octree or Seyel tree can be used multiple times within the same tree structure (technically, in this case, it becomes a graph structure). Therefore, storage can be saved by having multiple parent nodes pointing to the same subtree.

[0162] Figure 25 shows the combination of one VLO and three SLTs within a plenoptic octree 2501. The overall structure is that of a VLO where cubic voxel 2505 is represented by a level 1 VLO node and voxel 2502 is represented by a level 2 VLO node. Seyel 2507 is a level 3 Seyel with its origin at the center of the VLO universe (level 0). Seyel 2503, however, has a different origin. It is located at the center of the level 1 VLO node that is used as its surrounding cube. It is a level 2 Seyel because its defining square is one-quarter of the face of the surrounding cube.

[0163] Rather than a single VLO, as explained above, a plenoptic octree can consist of multiple VLOs representing multiple objects or properties that share the same universe and are typically combined using set operations. These are like layers in an image. In this way, multiple sets of properties can be defined for the same region of space and displayed and used as needed. Seyels in multiple Seyel trees can be combined in the same way if their origins are the same point and the nodes have the same alignment. This can be used, for example, to maintain multiple wavelengths of light that can be combined as needed.

[0164] The SLT and VLO in a plenoptic octree have the same coordinate system and the same universe, except that the SLT has its origin located at a different point in the plenoptic octree and is not necessarily located at the node center of the VLO. Thus, the surrounding cube of the SLT is oriented in the same way as the VLO or VLOs in the plenoptic octree, but does not necessarily coincide exactly with the VLO universe or other nodes.

[0165] The use of perspective plenoptic projections (or simply "projections") in a plenoptic octree, as computed by the plenoptic projection engine, is illustrated (in 2D) in FIG. 26. Plenoptic octree 2600 includes three SLTs attached to a VLO. SLT A 2601 has its origin at point 2602. From SLT A 2601, one sievel 2603 is shown projecting through the plenoptic octree in the positive x and positive y directions. SLT B 2604 has a sievel 2606 projecting into the plenoptic octree, and SLT C 2607 has a sievel 2608 projecting in another direction.

[0166] This continues in Figure 27, where two VLO voxels are shown, including VLO voxel 2710. The ray 2603 of SLT A 2601 and the ray 2606 of SLT B 2604 are outgoing ray sigmas. This means they represent light emanating from the center of their respective origins. Only one ray is shown for each SLT. In use, there will typically be many ray sigmas of varying resolution projecting from the origin of each SLT. In this case, the two ray sigmas pass through two VLO nodes. SLT C 2607 does not have a ray that intersects either of the two VLO nodes and is not shown in Figure 27.

[0167] In operation, the intersection of the SLT Seyleigh and the VLO node results in a refinement of the Seyleigh and the VLO node until some resolution limit (e.g., spatial and angular resolution) is reached. In a typical situation, refinement occurs until the Seyleigh projection approximates the size of the VLO node at some level of resolution determined by the characteristics of the data and the immediate needs of the requesting process. It will be held.

[0168] In Figure 28, the portion of outgoing light that hits voxel 2710 is captured from point 2602 via ray 2603 and from the origin of SLT 2604 via ray 2606. There are many ways that light can be captured and it depends on the application. This light that hits voxel 2710 is represented by incident SLT D 2810. This can be created when the light hits the voxel or added to an existing one if one already exists. The result in this case is two incident ray 2811 and 2806. This now represents the light hitting the voxel as represented by the light hitting the center of the node.

[0169] A typical use of SLTs in a plenoptic octree is to use the light entering a voxel, as represented by the input SLT, to calculate the output SLT for that voxel. Figure 29 illustrates this. The BLIF function known or assumed for voxel 2710 is used to generate a second SLT, the output SLT. This is output SLT D2910. Its origin is at the same point as output SLT D2810. Thus, output light from multiple locations in the scene, along with the light hitting the voxel captured by the input SLT, is projected outward and then used to calculate the output SLT for that voxel.

[0170] The functionality of an SPU in generating and operating on a plenoptic octree, according to some example embodiments, is shown in Figure 30. SPU 3001 may include a collection of modules, such as a collection of operations module 3003, a geometric shape module 3005, a shape transformation module 3007, an image generation module 3009, a spatial filtering module 3011, a surface extraction module 3013, a morphological operations module 3015, a connectivity module 3017, a mass property module 3019, a registration module 3021, and a light field operations module 3023. The operations of SPU modules 3003, 3005, 3007, 3009, 3011, 3013, 3015, 3017, 3019, and 3021 on the octree are generally known and understood by those skilled in the art. Those skilled in the art will understand that such modules may be implemented in many ways, including software and hardware.

[0171] Some SPU functionality has been extended to apply to plenoptic octrees and SLTs. Modifying the set operation module 3003 to operate on SLTs is a straightforward extension of the node set operations on octrees. The nodes of multiple SLTs must represent the same seiels (regions of direction space). The nodes are then traversed in the same order, providing the relevant properties contained in the SLTs to the manipulation algorithm. As is well known in the literature, terminal nodes in one SLT are matched with subtrees in another SLT using the "Full-Node Push" (FNP) operation, similar to octrees.

[0172] Due to the nature of the SLT, the operation of the Geometry 3005 process is limited when applied to the SLT. For example, no translation is applied in that the incoming or outgoing Seymour at one point in a plenoptic octree will generally not be the same at another origin. In other words, the light field at one point will typically be different from that at another point and must be recalculated at that point. The Seymour interpolation and extrapolation light field operations performed in the Light Field Operation module 3023 accomplish this. An exception to this is when the same lighting is applied to the entire region (e.g., lighting from beyond the disparity boundary). In such cases, the same SLT can simply be used at any point within the region.

[0173] Geometric scaling within function 3005 also does not apply to SLT. Individual spheres represent directions that extend to infinity and do not have a scalable size. Process 3005 The geometric rotation performed by can be applied to the SLT using the method described below.

[0174] Morphological operations in 3015, such as dilation and erosion, can be applied to spheres in the SLT by extending their bounds, for example, to overlapping illumination. This can be implemented by using undersized or oversized rectangles on the faces of the SLT's surrounding cube. In some situations, the connectivity function 3017 can be extended for SLT incorporation by adding properties to VLO nodes that indicate which spheres, including properties such as illumination, intersect with them. This can then be used in conjunction with connectivity to identify connected components that have a specific relationship with a projection property (e.g., materials illuminated by a particular light source, or materials that are not visible from a particular point in space).

[0175] The operations of the light field operations processor 3023 are divided into specific operations as shown in Figure 31. The position invariant light field generation module 3101 is used to generate an SLT for light from beyond the disparity boundary, and therefore can be used anywhere within the region where the disparity boundary is valid. The light may be sampled (e.g., from an image) or synthetically generated from modeling of the real world (e.g., the sun or moon) or from computerized models of objects and materials that cross the disparity boundary.

[0176] The outgoing light field generation module 3103 is used to generate point light field information in the form of SLTs that are located at specific points within the plenoptic octree scene model. This can be from sampled illumination or synthetically generated. For example, in some cases, pixel values ​​in an image may be traced back to a location on a surface. This illumination is then attached to the surface points as one or more outgoing sigils that are attached to (or contribute to) that location in the direction of the camera viewpoint of the image.

[0177] The outgoing / incoming light field processing module 3105 is used to generate an incident SLT for a point in the scene (e.g., a point on an object), called the "query" point. If it does not already exist, an SLT is generated for the point, and its sigma is populated with illumination information by projecting it onto the scene. When the first object in that direction is found, its outgoing sigma is accessed for information about the illumination projected back to the starting point. If no sigma exists in the direction of interest, neighboring sigma are accessed to generate an interpolated or extrapolated set of illumination values, possibly with the help of known or expected BLIF functions. This process continues for other sigma included in the incident SLT at the query point. Thus, the incident SLT models an estimate of the light landing on the query point from all or a subset of directions (e.g., light from the interior of an opaque object containing the surface of the query point may not be necessary).

[0178] The incident / exit light field processing module 3107 may then be used to generate an exit SLT at a point based on the incident SLT at that point, possibly generated by module 3105. The exit SLT is typically calculated using a BLIF function applied to the incident SLT. The operations of the sub-modules included in the light field operation module 3123 employ the Seyel projection and Seyel rotation methods shown below.

[0179] Figure 32 shows the surrounding cube 3210 of a Seyel tree. The six square faces of the SLT surrounding cube are numbered 1 through 6. The origin of the coordinate system is located at the center of the SLT universe. Face 0 3200 is the SLT face that intersects the -x axis (hidden in the illustration). Face 1 3201 is the face that intersects the +x axis, and Face 2 3202 is the face that intersects the -y axis (hidden). Face 3 3203 intersects the +y axis, Face 4 3204 intersects the -z axis (hidden), and Face 5 3205 intersects the +z axis.

[0180] Level 0 in the SLT contains all of the selves representing the entire area of ​​the sphere surrounding the SLT origin (4π steradians). At Level 1 of the SLT, six selves intersect exactly one of the six faces. At Level 2, each sel represents one-quarter of a face. Figure 33 illustrates the numbering of Face 5 3205. Quarter Face 0 3300 is in the -x, -y direction, Quarter Face 1 3301 is in the +x, -y direction, Quarter Face 2 3302 is in the -x, +y direction, and Quarter Face 3 3303 is in the +x, +y direction. Next, we focus on the quarter face of Face 3 3303 that is in the +x, +y, +z direction, as highlighted in Figure 33. Figure 34 shows Face 5 3205 as viewed from the +z axis toward the origin. From this perspective, quarter face 3401 is seen as a vertical line that is the edge of the quarter face square.

[0181] A seirl that intersects a level 2 quarter face is called an apex seirl. There are six faces per face and four quarter faces, so there are a total of 24 apex seirls. In 3D, an apex seirl is the space bounded by four planes that intersect the origin of the SLT, each of which intersects an edge of a level 2 quarter face. In 2D, this reduces to two rays that intersect the center and both ends of a quarter face, such as 3401. An example of an apex seirl is 3502 in Figure 35, with origin 3501.

[0182] A ray is a region of space that can be used to represent, for example, a light projection. They are determined by planes that enclose the volumetric space. Technically, they are diagonal (or non-right) rectangular pyramids of unlimited height. In 2D, planes appear as rays. For example, ray 3601 is shown in Figure 36. It originates at the SLT origin 3602. A particular ray is defined by its intersection with the origin and the projection plane 3603, which is a plane (a line in 2D) parallel to one face of the ray's surrounding cube (perpendicular to the x-axis in this case). The projection plane is typically attached to a node in the VLO and is used to determine if the ray intersects with that node and, when appropriate, to perform lighting calculations. The intersection point 3604, t(t x ,t y ) is the origin 3605 of the projection plane 3603, usually the center of the VLO to which it is attached, and from the origin 3605 of the projection plane, 3606, t y The intersection 3608 of ray 3601 with Seyer's surface 3607 is point "a" 3608. In the case shown, the distance from the origin to the surface in the x direction is 1, so the slope of the ray in the xy plane is therefore the y value of point 3608, a y This becomes:

[0183] The SLT is anchored to a specific point in the universe, the origin. The anchor point is an explicitly defined point for which the associated projection information is custom calculated. Alternatively, as described here, the SLT origin can start at the center of the universe and be moved to the anchor point using a VLO PUSH operation while maintaining a geometric relationship with the projection plane (which moves around in a similar manner). This has the advantage that multiple SLTs can be attached to a VLO node and share simplified projection calculations as the octree is traversed to identify the SLT center. The VLO octree that identifies the SLT center also contains nodes that represent materials within a single unified data set: the plenoptic octree.

[0184] When implementing plenoptic octree projection, each of the 24 top Seyhels can be processed independently on a separate processor. To reduce the memory access bandwidth of the VLO, each such processor can have a set of half-space generators. These are used to build the Seyhel pyramid that intersects the VLO locally (for each top Seyhel processor). Thus, unnecessary demands on VLO memory are eliminated.

[0185] The center of the lowest level VLO node can be used as the SLT origin, or if greater accuracy is required, the node with the projection correction calculated for the geometric calculation. An offset relative to the center may be specified.

[0186] The SLT is then positioned within the plenoptic octree by traversing the VLO (with or without offset correction) to locate the STL's origin. The STL's projection information, relative to the projection plane attached to the center of the universe, is set up relative to the VLO's root node and then updated with each PUSH until the SLT origin is located. In Figure 37, the center of the SLT (not shown) is placed at point 3702, the center of the VLO root node (only the level 1 VLO node 3701 is shown). To move the SLT, the center 3702 is moved with a PUSH of the VLO node. In the case shown, this is a move to the VLO child node in the +x and +y directions. This then moves to the level 1 node center 3703 in the +x and +y directions (relative to this top ray). The original ray 3704 (representing either edge of the ray) thus becomes 3705 after the PUSH. The slope of this new ray remains the same as the slope of the original ray 3704, but the point of intersection with the projection plane 3706 moves. The original intersection point 3708, t(t x ,t y ) is 3709, t'(t' x ,t' y ) and the x coordinate of the projection plane is t xremains the same, but t y The value of t' y changes to.

[0187] The y step is calculated by taking into account the step to the new origin and the slope of the ray. The edge of the VLO node at level 1 is e for 3701. x 1 as shown by 3710. The magnitude of the edge is the same in all directions of the axis, but they are kept as separate values ​​because the direction changes during the traversal. The y value is e y 3711. When a VLO PUSH occurs, the new edge value e' x 3712 and e' y 3713 becomes half of the original value. As shown in the diagram for this PUSH operation: e' x =e x / 2 and e' y =e y / 2

[0188] The new intersection point 3709 moves in the y direction, which is e' y The x value of the SLT origin, 3712e', is the y value of the origin shifted by 3713 multiplied by the slope of the edge. x This is due to the addition of the movement of t' y =t y +e' y -slope*e' x

[0189] This calculation can be done in a variety of ways. For example, rather than performing the product each time, the product of the slope and edge of the VLO universe can be kept in a shift register and divided by 2 with a right shift operation for each VLO PUSH. This shows that the center of the SLT can be moved by a PUSH operation on the VLO while maintaining the Seyer's projection on the projection plane.

[0190] The next operation moves the projection plane while maintaining its geometric relationship to the SLT. Projection planes are typically attached to the centers of different VLO nodes, generally at different levels of the VLO. When the node to which the projection plane is attached is subdivided, the projection plane and its origin move through the universe. This is shown in Figure 38. Projection plane 3802 is attached to the center of the VLO root node when a PUSH is performed on VLO node 3801 at level 1. Projection plane 3802 moves to a new location, becoming projection plane 3803. The projection plane origin moves from the center of the universe 3804 to point 3805, the center of the child node. The original sieve edge ray intersection point 3806, t(t x ,t y ) is a new intersection point 3807 on the new projection plane 3803, t'(t' x ,t' y ) as shown above. x is divided into two by PUSH and becomes 3812e' x y edge 3811e y Also, it is divided into two parts, 3813e' y This is calculated as follows: e' x =e x / 2 and e' y =e y / 2

[0191] The y component of the intersection point, relative to the new origin, is: t' y =t y -e' y +slope*e' x

[0192] e' yThe subtraction is done because the projection plane origin has moved from 3804 to 3805 in the + direction. And here again, the slope-multiplied edges are in shift registers and can be divided by two with a right shift for each PUSH. Depending on the details of the actual projection method, slope values ​​may need to be calculated separately in cases where two paths in the same tree (SLT origin and projection plane) can be PUSHed and POPed separately. For example, an SLT-placed octree structure may be traversed to the lowest level before the VLO traversal begins, then reusing some registers.

[0193] A "span" is a line segment in the projection plane between two rays that defines the limit of a splay (in one dimension). This is shown in Figure 39 for level 1 node 3901, which hosts splay 3902. It is defined by three points: the origin of the SLT, 3903, the "top" edge 3904, and the "bottom" edge 3905. An edge is defined by where it intersects with the projection plane 3906, which has its origin at point 3907. The intersection points are point t 3909 for the top edge and point b 3910 for the bottom edge.

[0194] The sel is defined only from the SLT origin outwards between the bottom and top edges. It is not defined on the other side of the origin. During processing, this situation can be detected for sel as shown in Figure 40 for a level 1 VLO node 4001 containing a sel 4002 with its origin at 4003. The projection plane 4004 moves to the other side of the SLT origin, resulting in 4005 where no sel exists. y The offset value is b' y 4006, and t y The offset value is 4007t' y After the move, the top offset value will be lower than the bottom offset value, indicating that the sphere is undefined. It no longer projects onto the projection plane, but its use for intersection operations with VLO nodes must be suspended until it returns to the other side of the origin, while the geometric relationship is calculated and maintained.

[0195] The sphere is subdivided into four sub-spheres using a sphere PUSH operation by calculating new top and bottom offsets. The sphere subdivision process is illustrated in Figure 41, as described above. Figure 42 shows a level 1 VLO node 4201 hosting a sphere defined by the origin 4203, the top point t' 4204, and the bottom point b' 4205. Depending on which child the sphere is pushing (typically based on geometric calculations performed at the time of the PUSH), the new sub-sphere can be an upper sub-sphere or a lower sub-sphere. The upper sub-sphere is defined by the origin 4203, the top point 4204, and point 4206, which is the center between the top point 4204 and the bottom point 4205. In the case shown, the lower sub-sphere is the result of the PUSH, defined by the origin 4203, the original bottom point b' 4205, and the new top point t' 4206. The value t' of the new apex point t is calculated as follows: t' y =(t y +b y ) / 2

[0196] The new bottom edge is the same as the original edge and has the same slope. The top edge defined by t' has a new slope, slope_t', which can be calculated as follows: slope_t'=(slope_t+slope_b) / 2

[0197] All seires at a particular level have the same surface area, but do not represent the same solid angular area because the origin moves relative to the surface area. This can be corrected by moving the edges of the rectangles on the surface for each seire at a level. This simplifies lighting calculations, but makes the geometric calculations more complex. The preferred method uses an SLT "template," which is a static, pre-calculated "shadow" SLT that is traversed simultaneously with the SLT. For light projection, it contains precise measurements of the solid area for each seire to use in the lighting transfer calculations.

[0198] A seires represents illumination incident on or emanating from a point in space, the center of the SLT (typically the space represented by a point). While the plenoptic octree can be employed in light transport in many ways, a preferred method is to first initialize the geometric variables at the origin of the SLT, which is at the center of the VLO. The geometric relationships are then maintained as the SLT is moved to its location in the universe. The VLO is then traversed, starting from the root, in forward-backward (FTB) order from the SLT origin, proceeding in the direction of the seires from the origin. In this way, VLO nodes, which typically contain matter in some form, are encountered in a general order of distance from the seires origin and processed accordingly. In general, this may need to be performed multiple times to account for sets of seires in different direction groups (top seires).

[0199] As the VLO is traversed in FTB order corresponding to the spheres projecting from the SLT origin, the first interacting VLO matter node encountered is then examined to determine the next step required. For example, it may be determined that illumination from a sphere is transferred to the VLO node by removing some or all of the illumination from the sphere and attaching it or a portion of it to a sphere attached to the VLO node containing the matter. This is typically an incident SLT attached to the VLO node. The transfer can be from properties that may be generated from an image sampling the light field. Alternatively, it can be from an outgoing sphere attached to the illumination source. The incident illumination can be used with a model of the light interaction properties of the surface to determine, for example, the outgoing light to be attached to an existing or newly created sphere.

[0200] As shown in Figure 43, a propagation between SEIs takes place from an outgoing SEI attached to VLO node 4301 to an incoming SEI attached to VLO node 4302. Propagation begins when the projection of the outgoing SEI 4303 onto its origin 4304 is of the correct size with respect to VLO node 4302, typically encircling its center. If not, the VLO node containing the outgoing SEI or the incoming SLT (or both) is subdivided and the situation is re-examined for the resulting subtree.

[0201] When propagation occurs, the antipodal Seyhels in the incoming Seyhel tree 4305 along the origin-origin segment 4307 are then accessed or generated at some Seyhel resolution. If a VLO node is too large, it is subdivided as necessary to increase the relative size of the projection. If an incoming Seyhel is too large, it is typically subdivided to reduce the size of the projection.

[0202] A specific traverse sequence is used to achieve FTB ordering of VLO node visits. This is shown in Figure 44 for the traversal of VLO node 4401. The traversal has the origin at 4403, and the edges (top and bottom) intersect with quarter-circle (eighth sphere in 3D) region 4404 (between bottom edge limit 4405 and top edge limit 4406). Edge 4407 is a typical edge within this region. The sequence 0 through 3 4408 generates the FTB sequence at VLO node 4401. Other sequences are used for other regions. In 3D, there is an equivalent traverse sequence of eight child nodes. With VLO, traversal is applied recursively. The traverse sequence ordering is not unique in that multiple sequences can generate FTB traversals for a region.

[0203] As the spheres are subdivided, some algorithms require tracking of spheres that contain light that has been consumed (e.g., absorbed or reflected) by material-containing VLO nodes that they encounter. As in octree image generation, a quadtree is used to mark "used" spheres. This is illustrated in Figure 45, where sphere 4502 with origin at 4503 is projected onto VLO node 4501. The quadtree 4504 (only edges are shown) is used to mark the spheres that have been consumed (e.g., absorbed or reflected) by material-containing VLO nodes that it encounters. It is used to keep track of Seyers in the Seyer tree that are not active or have previously been partially or completely used.

[0204] It is also possible for multiple processors to operate simultaneously on different Seyels. For example, 24 processors could each compute a different projection of the top Seyel and its descendants. This can require significant bandwidth from the memory holding the plenoptic octree, especially the VLO. The SLT-centered tree is typically synthetically generated for each processor, and the top Seyel and its descendants may be split into separate memory segments, but the VLO memory will be accessed by multiple Seyel processing units.

[0205] As noted above, memory bandwidth requirements can also be reduced by using a collection of half-space generators for each unit. As shown in Figure 46 in 2D, a half-space octree will be generated locally (within each processor) for the two edges (four planes in 3D) that define the sides of the Seyhel. Edge 4602 is the top Seyhel edge. The region below it is half-space 4603. Edge 4604 is the bottom edge that defines the upper half-space 4605. The Seyhel space is the intersection of two half-spaces 4606 in 2D. In 3D, the Seyhel volume is the intersection of four volume-occupying half-spaces.

[0206] The local Seyel-shaped octree would then be used as a mask that would be intersected with the VLO. If a node in the local generated octree is empty, the VLO octree in memory does not need to be accessed. In Figure 47, this is illustrated by the upper-level VLO node 4701 containing multiple lower-level nodes in its subtree. Node A, 4703, is completely disjoint with Seyel 4702 and does not need to be accessed. Seyel 4702 occupies a portion of the space of node B, 4704. VLO memory would need to be accessed, but all of its child nodes, 0, 2, 3, etc., are disjoint with Seyel and therefore no memory access is required. Node C, 4705, is completely surrounded by Seyel, so it and its descendants are required for processing. They need to be accessed from VLO memory as needed. Memory access issues can be reduced by interleaving the VLO memory in eight segments, corresponding to the eight level-1 octree nodes, as well as in other ways.

[0207] A "frontier" is defined here as a surface at a distance from a region in the plenoptic octree such that anything at an equal or greater distance exhibits no parallax at any point within that region. Thus, light coming from a particular direction does not change regardless of its placement within that plenoptic octree region. For light coming from beyond the frontier, a single SLT for the entire plenoptic octree can be used. In operation, for a given point, the incident SLT is accumulated for the point from outward projections. When all such illuminants have been determined (all illuminants from within the frontier), for any illuminants for which no such illuminant is found, the illuminant from the frontier SLT is used to determine its properties. Illuminants beyond the plenoptic octree but within the frontier can be represented, for example, by SLTs on the faces of the plenoptic octree (rather than a single SLT).

[0208] In many operations, such as calculations using surface properties like BLIF, it may be important to rotate the SLT. This is illustrated in Figure 48 in 2D and can be extended to 3D in a similar fashion. 4801 is the original VLO node containing the sieve 4804. Node 4802 is the rotated VLO node generated from it containing the sieve 4805, which is a rotated version of 4804. The two SLTs share the same origin, which is the VLO center point 4803. The algorithm generates new rotated sieves and subsieves from the original sieves and subsieves. This can be done for all the original sieves, or for example A plenoptic mask, or simply "mask" as used here, can be used to block the generation of some sigils in the new SLT because they are typically not needed for some reason (e.g., directions from surface points to opaque solids, or directions that are not needed, such as BLIFs for specular surfaces where some directions contribute little or nothing to the outgoing light). The mask may also specify property values ​​to look at (e.g., ignore sigils with radiance values ​​below some specified threshold). As shown in the figure, the faces (2D edges) of the new SLT surrounding cube become projection surfaces (2D lines), such as 4806. The spans in the new SLT are projections of the original SLT sigils.

[0209] FIG. 49 shows the intersection of the top edge of the new Seyel with the rotated projection plane 4902, at point t(t x ,t y ) 4901. Similarly, point b(b x ,b y ) 4903 are the intersection points of the bottom edges. These correspond to the endpoints of the edge / face spans in the new SLT's sel. They start at the edges or corners of the SLT octree universe. They are then subdivided as necessary. As shown, the center point 4904 now becomes the new apex point t', as follows: t' x =(t x +b x) / 2 and t' y =(t y +b y ) / 2

[0210] The x distance between the top and bottom points is d x 4905, which is divided into two parts for each PUSH. The change in y is d y At 4906, each push is divided into two. The difference for each subdivision is a function of the slope of the edge, which is also divided into two for each push. The task is to track the sel that is in the original SLT as it is subdivided and projected onto the new sel. At the bottom level (highest resolution in directional space), for nodes needed during processing, the property values ​​of the original sel are used to calculate the value of the new sel. This can be done by selecting the value from the sel with the largest projection, or by some weighted average or calculated value.

[0211] Figure 50 shows how span information is maintained. The original sieve is bounded by a top edge 5005 and a bottom edge (not shown). From point t to t original The distance to, y, is calculated at the start and then maintained as the subdivision continues. t 5010. Also, from point b to point b original There is also an equivalent distance, y, to point b' (not shown). The purpose of this calculation is to calculate the distance from a new point t, i.e. point b, to the associated original edge. This is the new apex point t' 5004 in the diagram. An equivalent method can be used to handle generating the bottom distance to point b'.

[0212] This calculation deals with two slopes: the original Seymour edge and the slope of the projected edge (a 3D plane). In both cases, the change in y distance for each step in x, d x / 2, in this case 5014, is the value determined by the slope and divided by two on each PUSH. These two values ​​can be kept in a shift register. These values ​​are initialized at the start and then shifted as needed on PUSH and POP operations.

[0213] As shown in the figure, the new offset distance dt'5004 is first calculated by x This can be calculated by determining the displacement along the projected edge for steps 5014 of 1 / 2, or in this case the value of "a" 5009. This can then be used to determine the distance from the new apex point t' to the original perpendicular intersection with the original apex edge. This is the value of "e" 5011 in the figure, and is equal to a-dt. The other part is the distance in y from the original intersection point on the apex edge of the original sieve to the new intersection point on the apex edge. This distance is calculated by multiplying the slope of the edge by d. x / 2, or "c" 5007 in the diagram. The new distance, dt' 5006, is thus the sum e+c.

[0214] When extending this to 3D, the tilt information in the new dimension needs to be used to calculate an additional value for the step in z direction, which is a straightforward extension of the 2D SLT rotation.

[0215] The SLT is hierarchical in that higher-level nodes represent directions in space in a larger volume than their descendant nodes. The SLT center of a parent node lies within this volume, but generally does not coincide with the center of any of its child nodes. If the SLT is generated from, for example, an image, a quadtree can be generated for the image. This can then be projected onto the SLT at node centers at multiple levels of resolution.

[0216] In other cases, higher levels are derived from lower levels. SLT reduction is a process used to generate values ​​for higher levels from the information contained in lower-level levels. This can generate averages, minimums, and maximums. In addition, a measure of coverage (e.g., the percentage of directional space in a sub-level that has a value) can be calculated and possibly accumulated. In some implementations, one or more "property vectors" can be used. These are directions along which some property of the level is spatially balanced in some sense.

[0217] It is often assumed that the SLT lies on or near a locally planar surface. If known, the local surface normal vector can be expressed for the SLT as a whole and used to improve its value in the reduction process.

[0218] In some situations, especially when lighting gradients are large, an improved minification process would be to project the lower-level Seyels onto a plane (e.g., through SLT space parallel to the known plane of the surface) or surface, filter the results on the surface (e.g., interpolate about the center of the larger parent Seyels), and then project the new values ​​back onto the SLT. Machine learning (ML) can be employed to analyze the lighting based on a previous training set and improve the minification process.

[0219] The outgoing SLT of a point in space representing a volumetric region containing material that interacts with light can be assembled from light field samples (e.g., images). If there is enough information to determine illumination in various directions, it may be possible to estimate (or "discover") the BLIF for the represented material. This can be facilitated if the incoming SLT can be estimated. ML can be used for BLIF discovery. For example, an "image" containing Seyle's illumination values ​​for a two-dimensional array (two angles) of SLTs can be stacked (from multiple SLTs) and used to recognize the BLIF.

[0220] SLT interpolation is the process of determining a value for an unknown Seyel based on values ​​in some set of other Seyels in the SLT. There are various ways this can be done. If a BLIF is known, can be estimated, or can be discovered, this can be used to intelligently estimate an unknown Seyel value from other Seyels.

[0221] Light sources can often be used to represent real or synthetic lighting. An ideal point light source can typically be represented by a single SLT, perhaps with uniform illumination in all directions. Enclosed point or directional light sources can be represented using mask SLTs to prevent illumination in blocked directions. Directional light sources can be represented using geometric extrusion of the "light" to generate an octree. Extrusion can also represent, for example, orthogonal projections of non-uniform illumination (e.g., an image).

[0222] A possible plenoptic octree projection processor is shown in Figure 51. It implements the projection of an SLT onto a VLO node in a plenoptic octree. Three PUSH operations can be performed: PUSH Center (pushes the center of the SLT onto a child node), PUSH VLO (pushes a VLO node onto a child node), and PUSH Sael (pushes a parent Sael onto a child). POP operations is not explicitly included here. It is assumed that all of the registers are pushed onto the stack at the start of each operation and then simply popped out. Alternatively, only certain PUSH operations (not values) are put onto the stack, and you can pop them back out by reversing the PUSH computation.

[0223] The processor is used for the "top" plane, which is projected onto the surface at x=1. This unit performs the projection calculations in the xy plane. The replication unit performs the computations in the yz plane.

[0224] To simplify the operation, all SLT center PUSH operations are performed first, thereby placing the SLT in its position (while maintaining the projected geometry). The two delta registers are reinitialized, and then a VLO PUSH operation is performed. Then an SLT PUSH operation is performed. These operations can be performed simultaneously, for example, by duplicating the delta registers.

[0225] The top register 5101 holds the y-position of the top surface of the projection plane (parallel to plane 1 in this case). The bottom register 5102 holds the y-position of the bottom plane. The delta shift register holds the slope values: Delta_U 5103 for the top surface and Delta_L 5104 for the bottom surface. These have enough "lev" (for level) bits to the right to maintain precision when a POP operation is performed after a PUSH to the lowest possible level. The delta register is initialized with the slope of the associated plane in the xy plane. This contains the change in y for a step in x of 1. With each PUSH (SLT center or VLO), it is shifted right by 1. This is therefore the change in y for a step to the child node in the x direction.

[0226] The edge shift registers maintain the distances of the VLO node edges. These are VLO_Edge 5105 for the edge of the node during the VLO traversal. SLT_Edge 5106 for the VLO node during the traversal to identify the edge in the plenoptic octree. The two are typically at different levels in the VLO. The edge registers also have a "lev" bit on the right to maintain precision. Other elements are selectors (5107 and 5108) and five adders (5110, 5111, 5112, 5113, 5114). The selectors and adders are controlled by signals A through D according to the following rules: The result is the VLO refinement signal 5109.

[0227] The operation of the SLT projection unit can be implemented in various ways. For example, if the clock speed in a particular implementation is low enough, instances of the processor can be replicated in a serial configuration to form a cascade of PUSH operations that can perform multiple level moves in a single clock cycle.

[0228] An alternative design that allows VLO and SLT PUSH operations to be performed simultaneously is shown in Figure 52. Two new delta registers have been added: V_Delta_U 5214 (for VLO delta, upper) and V_Delta_L 5215 for VLO delta. The delta registers in Figure 51 are currently only used for SLT PUSH operations, which are currently S_Delta_U 5203 and S_Delta_L 5204.

[0229] The processor startup state in Figure 51 is shown in Figure 56. Top sphere 5602 is at the origin (0,0) of universe 5603. A projection plane parallel to face 1 intersects the same point and its origin is at the same point. Note the quadrant numbering 5610 and subsphere numbering 5611.

[0230] The registers are initialized as follows: Upper=Lower=0 (Both the upper and lower edges intersect the projection plane at the origin.) Delta_U=1 (upper edge slope=1) Delta_L=0 (lower edge slope=0) VLO_edge=SLT_edge=1 (both start of edge distance for level 1 nodes)

[0231] The projection unit works as follows. SLT Center PUSH Shift SLT_edge and Delta_U to the right by 1 bit For SLT children 0 or 2: A is +, otherwise - For SLT child 0 or 1: C is -, otherwise + B is 0 D is 1 E is a no-load (Delta_U and Delta_L registers are unchanged) VLO Node PUSH Shift VLO_Edge and Delta_U to the right by 1 bit For VLO child 0 or 2: A is -, otherwise + For VLO child 0 or 1: B is +, otherwise - C is 0 D is 1 E is no-load SLT Sael PUSH New Seyer is upper: D is 1, otherwise 0 E is load

[0232] It may be desirable to place the center of the SLT at a point other than the center of the plenoptic node. This could be used, for example, to place a representative point at a particular point of some underlying structure rather than at the center of a local cubic volume in the space represented by the node. Or, a particular placement relative to a point light source could be desired.

[0233] This can be done by incorporating the SLT alignment into the initialization of the projection processor. This is simplified because the top slope starts at 1 and the bottom slope starts at 0. Thus the initial top projection plane intersection is at y, which is the y value of the Seymour center minus the x value. The bottom value is the y value of the Seymour center.

[0234] The projection calculation then proceeds as before. It is possible to add an offset shift value in the final push to the SLT node center, but this is generally undesirable, at least not when an SLT center push and a VLO push are done simultaneously. The correct span is required during VLO traversal, since the span value is used to select the next SVO child to visit.

[0235] The register values ​​for the three types of multiple PUSHs are contained in the Excel spreadsheets in Figures 53 and 54. The two offset values ​​in row 5 have been set to 0 to simplify the calculations. The Excel formula used is shown in Figure 55 (with the rows and columns reversed for readability). The offset values ​​are located in row 5 (F5 for x and H5 for y).

[0236] The values ​​in the spreadsheet are in floating point format for ease of understanding with geometric diagrams. In a real processor, registers can be scaled to integers using only integer arithmetic. The columns of the spreadsheet are as follows: A. Iteration (sequential number of PUSH operations) B. SLT PUSH (PUSHed to SLT central child node) C.VLO PUSH (VLO pushed to central child node) D. Sael PUSH (pushed by Saelko) E. SLT To Level (New level of SLT placement after PUSH) F.VLO To Level (New level of VLO node after PUSH) G.Sael To Level (New level of Sael after PUSH) H.SLT Edge (size of the node in the octree used to identify the center of the SLT) I. SLT Step x (Current step in x of the PUSH to the child when identifying the SLT center point. Depends on the number of children.) J. SLT Step y (The current step in y of the PUSH to children when identifying the SLT center point. Depends on the number of children. The magnitude is the same as SLT Step x except that the sign depends on the number of children being pushed.) K.SLTx (x location of the current center of the node used to locate the SLT) L.SLTy (the y position of the current center of the node used to locate the SLT) M.VLO Edge (length of current node at push time in VLO) N.VLO Step x (Current step size in x for moving to VLO child nodes. Sign depends on the number of children.) O.VLO Step y (The current step size in y for moving to VLO child nodes. In this implementation, it is the same size as VLO Step x, but the sign depends on the number of children.) P.VLO x (position of VLO node center at x) Q.VLO y (positioning of VLO node center in y) R.TOP Slope (The slope of the top (upper) edge of the sieve. Note: This is the actual slope, not the value relative to the current x-step size.) S.BOT Slope (The slope of the bottom edge of the sieve. Note: This is the actual slope, not the value relative to the current x-step size.) T.t_y (the y-value of the top (upper) end point of the span on the projection plane) U.b_y (bottom y value of the end point of the span on the projection plane) V.comp_t_y (the independently calculated value of t_y to compare with t_y) W.comp_b_y (the independently calculated value of b_y to compare with b_y) X.Notes (Comments about iterations)

[0237] Initialization values ​​in the first column ("(start)" in the first column). The values ​​are as listed above (shown in Figure 56n). Then, 14 iterations of PUSH operations with SLT center, VLO, and SLT SEIEL are performed. The first 7 iterations are shown geometrically in Figures 56-63.

[0238] The first two iterations are an SLT PUSH followed by two VLO PUSHes, then two Seyel PUSHes. This is then followed by two VLO PUSHes (iterations 7 and 8), one SLT PUSH (iteration 9), and finally a VLO PUSH.

[0239] The result of iteration #1 is shown in Figure 57, which is an SLT PUSH from the VLO root to child 3. At VLO node 5701 at level 1, the sieve 5702 is moved from the VLO origin 5706 to the center of child 3 at level 2, point 5703. The new sieve origin is positioned at (0.5,0.5). The projection plane 5707 remains in the same place and the slope does not change (0 for the bottom edge 5704 and 1 for the top edge 5705).

[0240] Iteration #2 is shown in Figure 58, which is an SLT PUSH to child 2. This is similar to the last operation, except that it is to child 2 of level 2, and therefore in a different direction, and the step is half the previous distance. At node 5801, Seyel 5802 is moved to point 5803 at (0.25, 0.75). The slope of bottom edge 5804 remains 0, and the slope of top edge 5805 remains 1. The projection plane 5807 remains in the same position. The intersection of the top edge with the projection plane is below the bottom intersection (not shown). Therefore, at this point, Seyel is actually does not intersect the projection plane and the projection is inactive.

[0241] Iteration #3 is a VLO PUSH from the root VLO node to child 3 (level 1). This is shown in Figure 59. At VLO node 5901, the sieve 5902 does not move. However, the projection plane 5907 moves from the center of the VLO root node to the center of child 3 at level 1, point 5906. Note that the origin of the projection plane has now moved to this point, which is the center of child 3. Because the projection plane moves and its origin changes, the intersection of the edge of 5902 with the projection plane is recalculated.

[0242] Iteration #4 is shown in Figure 60. This is a VLO PUSH to child 1. Seyle 6002 does not move, but the projection plane 6007 does, so the intersection of bottom edge 6004 and top edge 6005 must be recalculated. The projection plane moves in the +x direction with 6006 as the new origin. The slope of the edge does not change.

[0243] Figure 61 illustrates a sieve push to Iteration #5, Child 1. The origin of sieve 6102 is unchanged, but it is split into two subsieves, the lower of which is retained. The bottom edge 6104 remains the same, but the new top edge 6105 moves so that its intersection with the projection plane is halfway between the previous top and bottom intersections, or a distance of 0.75 from the projection plane origin. The bottom distance remains 0.5. The bottom slope remains 0, but the top slope decreases to an average slope of 0.5.

[0244] Iteration #6 is shown in Figure 62. This is a sieve PUSH to child 2. Again, sieve 6202 is split into two sub-sieves, with the upper sieve preserved. Thus, the top edge remains the same, with the same slope. The bottom edge is moved upward, away from the origin of the projection plane, and its slope is reset to the average of 0 and 0.5 or 0.25.

[0245] Iteration #7 is shown in Figure 63. The operation is a VLO push to child 0. The projection plane is moved in the -x direction and its origin is moved in the -x and -y directions. Seyle 6302 remains in the same position and its edges are unchanged except that their intersections with the projection plane change to correspond to the movement.

[0246] The Excel spreadsheet simulation was re-run with the SLT center offset set to a non-zero value (0.125 for the x offset value in cell F5 and 0.0625 for y in H5 in row 5). The results are shown in the spreadsheets in Figure 64 and Figure 65.

[0247] Volumetric techniques are used to represent materials, including objects, in a scene (VLO). They are also used to represent light (light fields) in a scene using SLT. As described above, information needed for high-quality visualization and other applications can be obtained from real-world scenes using a scene reconstruction engine (SRE). This can be combined with synthetically generated objects (generated by the SPU geometry transformation module 3007) to form composite scenes. In some exemplary embodiments, the techniques use a hierarchical, multi-resolution, and spatially sorted volumetric data structure for both materials and light and their interactions within the SPU 3001. This allows for rapid identification of the portions of a scene needed for remote use based on, for example, placement, resolution, visibility, and other characteristics as determined by each user's placement and viewing direction or statistically estimated for a group of users. In other cases, an application may request a subset of the database based on other considerations. By communicating only the necessary portions, channel bandwidth requirements are minimized. The use of volumetric models also facilitates advanced functionality in the virtual world, such as collision detection (e.g., using set operations module 3003) and physics-based simulation (e.g., mass properties easily calculated by mass properties module 3019).

[0248] In some applications, it may be desirable to combine matter field and light field models generated separately by an SRE, or by multiple SREs, to form a composite scene model, for example, for remote visualization and interaction by one or more users (e.g., musicians or dancers located in a remote arena). Because lighting and material properties are modeled, lighting from one scene can be applied to replace lighting in another scene, ensuring that the viewer experiences a uniformly lit scene. The light field operations module 3023 can be used to calculate lighting while the image generation module 3009 is generating images.

[0249] A scene graph or other mechanism is used to represent the spatial relationships between individual scene elements. One or more SREs may generate a real-world model that is stored in the plenoptic scene database 1A07. In addition, the database contains other real-world or synthetic spatial models represented in other formats (not plenoptic octrees). This can be almost any representation that can be easily converted to a plenoptic octree representation by the geometry transformation module 3007. This includes polygonal models, parametric models, solid models (e.g., CSG (Constructive Space Gradient) or boundary representations), etc. The function of the SPU 3001 is to perform the transformation one or more times as the model changes or requirements change (e.g., if the viewer moves closer to an object and a higher resolution transformation is required).

[0250] In addition to light fields and material properties, SRE can also discover a wide variety of additional characteristics within a scene. This can be used, for example, to recognize visual attributes within a scene that can be used to enable previously acquired or synthesized models to be incorporated into the scene. For example, if a remote viewer is visually too close to an object and requires higher resolution than that acquired by SRE from the real world (e.g., a tree), an alternative model (e.g., parametric bark) can be smoothly "switched in" to generate higher resolution visual information for the user.

[0251] The SPU module in 3001 can be used to transform and manipulate models to achieve application requirements, often as the scene graph is modified by an application program, such as in response to a user request. This and other SPU spatial operations can be used to implement advanced functionality, including features requiring mass properties such as mass, weight, and center of gravity, as calculated by the SPU Mass Properties module 3019, in addition to interference and collision detection, as calculated by the Set Operations module 3003. Thus, models in the plenoptic scene database are modified to reflect real-time scene changes as determined by users and application programs.

[0252] Both material (VLO) and light (SLT) types of information can be accessed and transmitted for selected regions of space (directional space in the case of SLT) and specified levels of resolution (angular resolution in the case of SLT). In addition, property values ​​are typically stored in lower-resolution nodes (higher in the tree) in the tree structure that represent properties in the subtree of the node. This could be, for example, the average or min / max color in the subtree of an octree node, or some representative measure of illumination in the subtree of a Seymour tree node.

[0253] Depending on the needs of the remote process (e.g., one or more users), only the necessary subset of the scene model needs to be transmitted. For viewing, this typically means sending high-resolution information about the portion of the scene currently being viewed (or expected to be viewed) by module 3009 at a higher resolution than other areas. Higher-resolution information is transmitted for objects that are visually closer than for objects that are further away. Tracked or predicted motion will be used to predict the portions of the scene that are needed. These are transmitted with high priority. The advanced image generation method of the octree model in 3009 can determine occluded regions as the scene is rendered. This indicates areas of the scene that are not needed or that can be represented with a lower level of fidelity (taking into account potential future viewing). This selective transmission capability is an inherent part of the codec. Only portions of the scene at various resolutions are accessed from storage and transmitted. Control information is transmitted as needed to maintain synchronization with the remote user.

[0254] When multiple remote viewers are moving simultaneously, their viewing parameters can be summarized to set transmission priorities. An alternative would be to model expected viewer preferences probabilistically, perhaps based on experience. Since a version of the model of the entire scene is always available to all viewers at some, possibly limited, level of resolution, unexpected views still result, but at a lower level of image quality.

[0255] The information required for image generation is maintained in a local database that is generally a subset of the source scene model database. The composition of the scene is controlled by a local scene graph, which may be a subset of the global scene graph at the source. Thus, for particularly large "virtual worlds," the local scene graph may maintain only object and light field information, as well as other items that are visible or potentially visible to the user or that may be important to the application (e.g., the user's experience).

[0256] Information communicated between the scene server and clients consists of control information and parts of the model in the form of a plenoptic octree, and possibly other models (e.g., other forms of geometry, BLIF functions). The plenoptic octree contains a matter field in the form of a VLO and a light field in the form of an SLT. Each is a hierarchical, multi-resolution, spatially sorted volumetric tree structure. This allows them to be accessed by specified regions of modeling space, with variable resolution that can be specified by spatial domain (or direction space for Seymour trees). Each user's position in the scene space, viewing direction, and required resolution (typically based on the viewing direction's distance from the viewpoint) and predicted future changes can thus be used to determine the subset of the scene that needs to be transmitted, and the priority for each based on various considerations (e.g., how far and fast the viewpoint can move, the bandwidth characteristics of the communication channel, the relative importance of image quality for various sections of the scene).

[0257] Depending on the computational power available to dedicate at the remote site, functionality associated with the server side of the communication channel can be implemented at the remote site. This allows, for example, for the matter model (VLO) to be transmitted only to the remote site along with the light field information (SLT) reconstructed there, rather than also being transmitted over the channel. The potential communication efficiency, of course, depends on the details of the situation. Transmitting a simple model of the solid material to the remote site and subsequently computing and displaying the light field locally may be more efficient than transmitting the full light field information. This may be particularly true for static objects in the scene. On the other hand, objects that change shape or undergo complex motion may benefit from transmitting only the light field SLT on demand.

[0258] In a plenoptic octree, the SLT is a 5D hierarchical representation of some location in space within the scene (or, in some cases, beyond the scene). The five dimensions are the three location components (x, y, z) of the center where all the selves intersect, plus the two angles that define the selves. The selveve tree can be located at the center of the VLO voxel or specified anywhere within the voxel. Thus, a VLO node can contain material, as defined by its properties, and optionally also contain a selveve tree. Voxels in space that contain substantially non-opaque (transparent) media and are adjacent to the scene boundary (void voxels) can be referred to as "window" voxels in some embodiments.

[0259] A set of Seyhels may be similar at multiple points in a scene (e.g., nearby points on a surface with the same reflectance properties). In such cases, sets of Seyhels with different centers can be represented independently of the center location. If they are identical for multiple center points, they can be referenced from multiple center locations. If the differences are small enough, multiple sets can be represented by individual sets of deviations from a single model Seyhel or set of model Seyhels. Alternatively, they can be generated by applying coefficients to a set of precomputed basis functions (e.g., a Seyhel dataset generated from a representative dataset using principal component analysis). In addition, other transformations, such as rotation around the center, can be used to modify a single Seyhel model to a specific set. Some types of SLTs, such as point light sources, can be replicated simply by providing additional locations to the model (no interpolation or extrapolation is required).

[0260] Scene codecs operate in a data flow mode where there is a data source and a data sink. Typically this takes the form of a request / response pair. The request may be directed to a local codec where a response is generated (e.g., the current state), or may be transmitted to a remote codec where an action is performed and a response is returned, providing the results of the requested action.

[0261] Requests and responses are communicated through the scene codec's Application Programming Interface (API). The core functionality of the Basic Codec API 6601 is summarized in Figure 66. A codec is initialized through the Operational Parameters module 6603. This function can be used to specify or read the codec's operating mode, control parameters, and status. After a link to another scene codec has been established, this function can also be used to control and query the remote codec, given certain permissions.

[0262] When triggered, the codec API establish link module 6605 attempts to establish a communications link to the specified remote scene codec. This typically initiates a "handshake" sequence to establish communication operating parameters (protocol, credentials, expected network bandwidth, etc.). If successful, both codecs report to the calling routine that they are ready for communication operations.

[0263] The next step is to establish a scene session, which is set up through the API open scene module 6607. This involves establishing links to the scene database on both the remote side, and often also the local side, to access or update the remote scene database, for example, to build a local sub-scene database from the remote scene database, or to update the local scene database simultaneously with the remote scene database.

[0264] Once a connection to a scene or multiple scenes has been established, the scene database is accessed. Two code API modules may be used to request and modify information about a remote scene that does not involve moving a sub-scene across a communication channel. The remote scene access module 6609 is used to request information and modifications about a remote scene that does not involve moving a sub-scene across a communication channel. Operations performed on the local scene database are performed using the local scene access module 6611. Scene database queries that involve moving a sub-scene are performed using the query processor module 6613. All actions performed by the codec are logged by the session log module 6615.

[0265] The primary function of the query processor module 6613 is to transmit subscenes from a remote scene database or request subscenes to be incorporated into (or removed from) it. This may involve, for example, queries about the state of plenoptic octree nodes, requests for calculation of mass properties, etc. This typically involves the subscene extraction and transmission of the plenoptic octree's subtree and associated information in a compressed, serialized, and possibly encrypted form. The subscene to be extracted is typically defined as a collection of geometric shapes, octrees, and other geometric entities specified in some form of scene graph that can result in spatial, volumetric, and / or directional space regions. In addition, the required resolution in various regions of volumetric or directional space is specified (e.g., decreasing from the viewpoint in a rendering situation). The type of information is also specified to avoid transmitting unnecessary information. In some situations, subscene extraction can be used to perform some form of spatial query. For example, a request to perform a subscene extraction of a region but only down to level 0 returns a single node, which, if found to be NULL, indicates that there is no material in that region. This can also be extended to search for specific features within the plenoptic scene.

[0266] The sub-functions of the Query Processor module 6613 are shown in Figure 67. It consists of a Status & Property Query module 6703, which is used to obtain information about the plenoptic scene, such as the ability to perform writes, or what properties exist in it, or whether new properties can be defined. The Subscene Mask Control module accepts some form of subscene extraction request and builds a mask plan to fulfill the request. This is typically a collection of evolving masks that incrementally send subscenes to the requesting system as planned by the Plan Subscene Mask module 6705.

[0267] The Subscene Mask Generator 6707 builds and sends back to the requesting system a plenoptic octree mask used to select nodes from the scene database. It continually builds the next mask for extraction. The Subscene Extraction module 6709 performs a traversal of the scene plenoptic octree to select nodes as determined by the mask. These are then serialized, further processed, and input into a stream of packets that are then transmitted to the requesting system. The Subscene Inserter module 6711 is used by the requesting system to modify its local subtree of the scene model using the transmitted stream of plenoptic node requests.

[0268] A codec may perform sub-scene extraction or sub-scene insertion, or both. If only one is implemented, modules or functions required for the other may be omitted. Thus, an encoder-only unit requires a sub-scene extractor 6709 but not a sub-scene inserter 6711. A decoder-only unit requires a sub-scene inserter module 6711 but not a sub-scene extractor module 6709.

[0269] As described above, extracting sub-scenes from a plenoptic scene model allows for efficient transmission to a client of only those portions of the scene database that are needed for immediate or near-term visualization or other uses. In some embodiments, a plenoptic octree is used for the scene database. The properties of such a data structure facilitate efficient extraction of sub-scenes.

[0270] Various types of information may be included in the plenoptic octree, either in auxiliary data structures, in separate databases, or in some other manner, as separate VLOs, or as properties contained in or attached to octree or Seymour tree nodes within the plenoptic octree. The initial subscene extraction request specifies the type of information the client needs. This may be done in a variety of ways specific to the application being served.

[0271] The following is a use case where a client is requesting sub-scene extraction for remote viewing with a display device such as a VR or AR headset. A large plenoptic octree is maintained on the server side. Only a subset is required on the client side to generate the image. A plenoptic octree mask is used here as an example. Many other methods can be used to achieve this. The mask is a form of plenoptic octree that is used to select nodes within the plenoptic octree using a set operation. For example, a subsection of the large octree can be selected using a smaller octree mask and an intersection operation. The two octrees share the exact same universe and orientation. The two trees are traversed simultaneously from the root node. Any nodes in the large octree that do not exist as occupied nodes are simply skipped and ignored in memory. They simply do not appear in the traversal. They effectively disappear. In this way, a subset of nodes can be selected by traversal and serialized for transmission. This subset is then recreated on the receiving side and applied to the local plenoptic octree. This concept can be easily extended to Seyel trees.

[0272] Next, the concept of a mask is extended using incremental masks. Thus, the starting mask may be increased, decreased, or otherwise modified to select additional nodes for transmission to the receiver. The mask may be modified for this purpose in a variety of ways. Morphological operations of dilation and erosion can be applied using the SPU Morphological Operations module 3015. Geometric shapes may be added by transforming them using the SPU Shape Transform module 3007 and the SPU Set Operations module 3003, or used to remove portions of the mask. Typically, the new mask is subtracted from the old mask to produce an incremental mask. These are used to traverse the large scene model and identify new nodes to be serialized and transmitted, either added or otherwise processed at the receiver. Optionally, the opposite subtraction is performed, whereby the new mask is subtracted from the old mask to determine the set of nodes to be removed. This can be serialized and transmitted directly, with removal at the receiver (without sub-scene extraction). A similar method is used on the receiving side to remove nodes that are no longer needed for some reason (e.g. the viewer moves and high resolution information is no longer needed in some areas) and to notify the server side of changes to the current mask.

[0273] The purpose of a plenoptic projection engine (PPE) is to efficiently project light from one location in a plenoptic scene model to another, resulting in light transfer. This can be, for example, from a light source represented by an exit point light field (PLF) to an incident PLF attached to a medium, or it can be an incident PLF such that the exit light is added to the exit PLF.

[0274] Plenoptic projection utilizes a spatially sorted hierarchical multiresolution tree structure to efficiently perform the projection process. Three tree structures are used: (1) a medial-holding VLO or volume octree (which is considered a single octree, but can be multiple octrees merged with UNION); (2) a SOO or Seyel tree origin octree, which is an octree containing the origin of a Seyel tree within a plenoptic octree; and (3) an SLT, several Seyel trees within a plenoptic octree (the origin placement is within a SOO).

[0275] The plenoptic projection engine projects the seires in the SLTs onto nodes in the VLO in a forward-backward sequence, starting from the origin of each SLT. When a seire intersects with a media node, the size of the projection is compared to the size of the media voxel. The analysis is based on many factors, including the currently required spatial or angular resolution, the relative sizes of the media and seire projections above it, the presence of higher-resolution information at lower levels in the tree, and other factors. If necessary, either the media or seire, or both, can be subdivided into regions represented by their children. The same analysis then continues at the higher resolution.

[0276] Once the subdivision process is complete, optical transfer can occur. A ray in a ray tree can result, for example, in the creation or modification of a ray or ray's in a ray tree attached to a medium. In a typical application, incident ray information can be stored in an incident PLF attached to a medium. When the incident SLT is fully filled, a BLIF for the medium can be applied, resulting in an exit PLF for the medium.

[0277] The projection process works by maintaining a projection of the sphere onto a projection plane attached to each VLO node visited in the traversal. The projection plane is perpendicular to the axis according to the apex sphere to which the projected sphere belongs.

[0278] The process begins by constructing the VLO and SOO trees starting from the center of the universe. Thus, the location in the SOO begins at the center of the universe. It is traversed to the location of the first SLT to be projected, as determined by any applied masks and any specified traversal sequence. The projection plane begins as a plane passing through the origin of the universe, perpendicular to the appropriate axis, depending on the first seires. In operation, all three can be defined and tracked, taking into account the orientation of all top seires.

[0279] The primary function of the Plenoptic Projection Engine is to continually maintain the projection of the oblique pyramidal projection, which is the Seyle projection onto the projection plane attached to the medial, as the VLO is traversed. This is done by first initializing the geometry, then traversing the three tree structures, typically by projecting all Seyle in all SLTs into the scene and maintaining it. This may create additional SLTs that may be further traversed during this process or as they are created subsequently.

[0280] Thus, the typical flow of this method is to initialize a tree structure, then traverse the TOO using a series of TOO PUSH operations, placing the first SLT at its origin, maintaining the projection geometry for each step. Next, the VLO is traversed around the origin of the first SLT. The VLO is then traversed in front-to-back order, visiting nodes in a general order of increasing distance from the SLT origin toward the top SLT. At each PUSH of the VLO, the projection onto the projection plane connected to the node is checked to see if the SLT continues to intersect with the VLO node. If not, all subtrees are truncated. Ignored by rubbers.

[0281] When a medial VLO node is encountered, the next action to take is determined by analysis as outlined above, which typically involves visiting subtrees of the VLO and / or SLT. Once complete, the tree is popped back to where the next SLT can be projected in the same manner. Once the final SLT of the first or subsequent SLT has been processed, the tree is popped to the point where processing of the next SLT can begin. This process continues until all SLTs in all SLTs have been processed, or no further SLTs are needed because ancestor SLTs have been processed.

[0282] The overall procedure is shown in the Plenoptic Projection Engine flowchart in Figure 68A. This is a sample procedure out of many possible procedures. The process begins with the initialization of the projection mechanism in operation 68A02. As presented above, the VLO traverse begins at its root. Therefore, the projection plane of interest is attached to the center of the universe (three can actually be tracked). The SSO is also initialized to its root. Therefore, the initial SLT point starts at the origin of the universe and is pushed to the origin of the SLT. The initial seires to be visited are the top seires 0.

[0283] In operation 68A04, the SOO tree structure is traversed using PUSH operations to the origin of the next SLT in the plenoptic octree universe. On the first use, this is from the origin of the universe; at other times, it is from where the last operation left off. The projection of the current top Seyer onto the projection plane attached to the current VLO projection plane (which is initially attached to the center of the universe) is maintained for each operation to reach the origin of the next SLT. If there are no additional SLTs in the SOO (typically detected by attempting a POP from the root node), decision operation 68A06 ends the operation and returns control to the request routine.

[0284] If not, operation 68A08 traverses the spheres of the current SLT to the sphere that represents the first non-null node (non-void voxel) radius. Again, the projection geometry between the spheres and the projection plane is maintained. If no spheres with radius remain, control is returned by decision operation 68A10 to operation 68A04 to find and traverse the next SLT.

[0285] If a seyel needs to be projected, operation 68A12 traverses the VLO tree to the node that encloses the origin of the current SLT. Essentially, it finds the first VLO node whose projection plane intersects the VLO node's intersection with that projection plane. If no such VLO node is found, decision operation 68A14 returns control to operation 68A08 to proceed to the next seyel to be processed. If the seyel's projection does not intersect the node, decision operation 68A16 returns control to operation 68A12 to proceed to the next VLO node to be examined.

[0286] Otherwise, control is passed to operation 68A18, where the projection of the current seyel on the current projection plane is analyzed. If the current rules for such are satisfied, control is passed by decision operation 68A20 to operation 68A22, where radiance transfer or resampling is performed. This generally means that a sufficient level of resolution has been achieved, perhaps based on the variance of the seyel radiances, and that the size of the projection is, in some sense, comparable to the size of the VLO node. In some cases, transfer of some or all of the radiance to the appropriate seyel within the SLT attached to that node (created as necessary) is performed. In other cases, the radiance may be used in some other way within the node.

[0287] If the analysis determines that a higher level of resolution is needed for the Seyle or VLO node, operation 68A24 determines whether the VLO node needs to be subdivided. If so, control passes to operation 68A26, which performs a PUSH. Information about the situation is typically pushed onto the operation stack for later visits to all sibling nodes. If not, control passes to decision 68A28, where the need for subdivision of the current Seyle is addressed. If so, control passes to operation 68A30, where a Seyle PUSH is performed. Otherwise, the Seyle projection onto the current VLO node does not require additional processing, and control passes back to operation 68A12, which traverses to the next VLO node for investigation and possible transfer of radiance.

[0288] The general flow of subscene extraction from a plenoptic octree is shown in the flowchart of Figure 68B. The process begins when a subscene request is received. Initial step 68B02 initializes a null subscene mask, which is typically a single-node plenoptic octree and related parameters. The request is then analyzed in step 68B04. For image generation situations, this could include the viewer's 3D placement and viewing direction within the scene. Additional information would be the field of view, screen resolution, and other viewing and related parameters.

[0289] For viewing, this will then be used to define the initial view frustum for the first image. This may be represented as a geometric shape and converted to an octree using the SPU Shape Transform module 3007. In other situations, a Seymour tree may be generated, with each pixel forming a Seymour as a result. The distance from the viewpoint may be incorporated as part of the mask data structure or calculated in some other way (e.g., distance calculated on the fly during sub-scene extraction). This is used during sub-scene extraction to determine the resolution of the scene model (volume space or direction space) that should be selected for transmission.

[0290] From this analysis by module 68B04, a plan for a series of sub-scene masks is created. A common strategy is to start with a mask that produces an initial sub-scene model at the receiver that results in a usable image for the viewer very quickly. This could, for example, have a low-resolution request for a smaller initial data set for transmission. The next step is typically to add increasingly higher resolution detail. Information in the request from the viewing client can then be used to anticipate future changes in viewing needs. This could include, for example, the direction and speed of the viewer's orientation and rotational movement. This is used to extend the mask to account for future anticipated needs. These extensions are incorporated into scheduled steps for future mask changes.

[0291] This plan is then passed to operation 68B06, which intersects the subscene mask as defined by the current step in the plan with the complete plenoptic scene model. The nodes resulting from a particular traversal of the plenoptic octree are then collected and put into a serialized form for transmission to a requesting program. Compression, encryption, and other processing operations on the nodes may be performed before they are passed on for transmission.

[0292] The next flow operation is a decision performed by accept 68B08 of the next request. If this is for a new subscene that cannot be accommodated by modifying the current subscene mask and plan, the current mask plan is discarded and a new subscene mask is initialized in operation 68B02. On the other hand, if the request is for a subscene that is already anticipated in the current plan, as determined in decision operation 68B10, the next step in the plan is executed in operation 68B12. The subscene mask is modified and control is returned to operation 68B06 to implement the next subscene extraction. If the end of the subscene mask is encountered by decision 68B10, and there is a new subscene extraction request as determined by decision operation 68B14, the next request is used to start a new subscene mask in operation 68B02. If no new requests are pending, subscene extraction operation 68B00 is placed in a wait state until a new request arrives.

[0293] Figure 69 is a flow diagram of a process 6900 for extracting a subscene (model) from a scene database for generating images from multiple viewpoints within a scene, in one embodiment. The first step, operation 6901, establishes a connection to a database containing the complete scene model, from which the subscene is extracted. In operation 6903, the output new subscene is initialized to have a plenoptic field with no primitives. In other words, there are no matter or light fields in the subscene at this point. In operation 6905, a set of "query sel" is determined based on image generation parameters, including the 6DOF pose, intrinsic parameters, and image dimensions of the virtual camera at each viewpoint. A query sel is a sel defined at a level within the SLT of the complete scene, and this sel is used to spatially query (probe) the scene for plenoptic primitives that lie within the solid angular volume of the query sel. The set of query sel is typically a union of the sets of sel per viewpoint. The set of query ellipses for each viewpoint typically covers the FOV (camera frustum) such that each image pixel is included in at least one query ellipses. Query ellipses may be adaptively sized to match the sizes of primitives placed at various distances in the scene. The set of query ellipses may also be purposefully crafted to cover a 5D plenoptic space that is slightly larger or much larger than a tight union of the FOVs. One exemplary reason for such non-minimal plenoptic coverage is that process 6900 can predict the 6DOF path of the virtual camera used by the client for image generation.

[0294] In operation 6907, the primitives in the plenoptic field are accumulated into a new subscene by projecting each query light into the full scene using process 7000, which generally involves a recursive chain of projections such that the light field is resolved to a target accuracy specified by the image generation parameters. The target accuracy may include representations of radiance, spectrum, and polarization target accuracy. A description of process 7000 is provided below with reference to FIG. 70.

[0295] In operation 6909, process 6900 determines a subset of the accumulated primitives to retain in the subscene. Details of this determination are provided below in the description of operation 6915. In one simple yet practical case, primitives are retained if they fall at least partially within one of the camera FOVs specified in the image generation parameters of the subscene request. In operation 6911, an outer scene boundary for the subscene is defined to enclose at least the accumulated primitives that are partially or completely contained within at least one of the FOVs. A radius of interest is projected onto the defined outer scene boundary in operation 6913. This projection can generally be performed from scene regions both inside and outside the boundary. The boundary light field is generally implemented as a window light field at the boundary medial adjacent to the boundary.

[0296] In operation 6915, the process 6900 further simplifies the sub-scene as appropriate for the current use case and QoS thresholds. One notable example of sub-scene simplification is when minimizing sub-scene data size is important, such as when BLIF interactions result in a sparse sub-scene. This is the complete or partial removal of media radii resulting from transport (projection) from other media. That is, by removing radii that are not part of the window or irradiated light field, the subscene is in a more "canonical" form, which typically has a smaller data size than a non-canonical form, especially when a compressed BLIF representation is used. In the context of this specification, a canonical representation of the plenoptic field of a scene model ("canonical form") is one that contains, to some practical extent depending on the use case, a minimal amount of light field information contained in the non-window and non-irradiated portions of the light field. This is achieved by containing sufficiently accurate representations of the matter field (including media BLIF) and the window and irradiated light field radii. Then, when needed, other portions of the total quasi-steady-state light field can be calculated, for example, by a process such as that described with reference to Figures 70 and 71 below.

[0297] Some degree of simplification (compression) can be achieved by adapting the BLIF representation to suit the needs of the sub-scene extraction client. In this example, where the client intends to generate an image of the sub-scene, a glossy BLIF of, say, the car's surface might reflect a very busy tree from one viewpoint and only a homogenous patch of sky from another viewpoint. If only the second viewpoint is included in the image generation parameters in operation 6905, a more compact BLIF representation, although less accurate in its specular lobes, may suffice for the sub-scene model.

[0298] Note that in many use cases, sparseness of the sub-scen...

Claims

1. a plenoptic projection engine comprising a digital data processing circuit having at least one data processor, said digital data processing circuit in communication with a digital data memory; accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; a plenoptic projection engine configured to store the processed output light field in said digital data memory;

2. The plenoptic projection engine of claim 1 , wherein the digital data defining one or more input light fields is indexed by a plenoptic octree.

3. The plenoptic projection engine of claim 1 , wherein the digital data processing circuit determines the output light field using integer arithmetic without general division operations.

4. 2. The plenoptic projection engine of claim 1, wherein the one or more input volume elements are determined using one or more digital data defining a plenoptic selection mask representing characteristics including one or more of occupancy, resolution, and color.

5. configured to operate as part of a sub-scene extractor system, said digital data processing circuitry comprising: accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; The plenoptic projection engine of claim 1 , further configured to store the processed output light field in the digital data memory.

6. 1. A scene codec comprising a plenoptic scene data encoder and / or a plenoptic scene data decoder, an encoder / decoder processor circuit configured to communicate with an associated memory circuit to stream hierarchical multi-resolution and plenoptically sorted sub-scene data packet structures representing light fields and / or matter fields as directionally and / or volumetrically subdivided data into a locally stored three-dimensional (3D) or 4D sub-scene model database as plenoptic increments of a remotely stored plenoptic scene model database of real-world 3D / 4D space; the encoder / decoder configured to communicate with at least one digital data input / output data transport path and configured to encode / decode plenoptic sub-scene data to / from the input / output data transport path; a processor circuit; and a codec control processor circuit connected to control the encoder / decoder processing circuit and the locally stored three-dimensional (3D) sub-scene model database by generating and / or responding to data control packets communicated over the data transport path.

7. 7. The scene codec of claim 6, wherein the codec control processor circuit is configured to operate as a query processor in response to data control packets by accessing identified scene segments in a plenoptic scene database.

8. 8. The scene codec of claim 7, wherein the query processor is configured to communicate with a remote plenoptic scene database via a data transport path, and sub-scene data on a server side of the data transport path is queried to support image generation on a client side of the data transport path.

9. 8. The scene codec of claim 7, wherein the query processor is configured to communicate with a remote plenoptic scene database via a data transport path, and plenoptic sub-scene data on a server side of the data transport path is provided to support plenoptic sub-scene image generation on a client side of the data transport path.

10. 8. The scene codec of claim 7, wherein the query processor is configured to communicate with a data transport path and uses a plenoptic scene database on a server side of the data transport path to support image generation on a client side of the data transport path by inserting plenoptic subscenes into a plenoptic subscene database on the client side of the data transport path.

11. 1. A method for processing plenoptic projections using digital data processing circuitry and associated digital data memory, comprising: accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; and storing the processed output light field in the digital data memory.

12. 12. The method of claim 11, wherein the digital data defining one or more input light fields is indexed by a plenoptic octree.

13. 12. The plenoptic projection processing method of claim 11, wherein the output light field is determined by the digital data processing circuitry using integer arithmetic without general division operations.

14. 12. The plenoptic projection processing method of claim 11, wherein the one or more input volume elements are determined using one or more digital data defining a plenoptic selection mask representing characteristics including one or more of occupancy, resolution, and color.

15. Sub-scene extraction is accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; and storing the processed output light field in the digital data memory.

16. 1. A plenoptic scene codec data processing method, comprising: communicating with an associated memory circuit to stream hierarchical multi-resolution and plenoptically sorted sub-scene data packet structures representing light fields and / or matter fields as directionally and / or volumetrically subdivided data into a locally stored three-dimensional (3D) or 4D sub-scene model database as plenoptic increments of a remotely stored plenoptic scene model database of real-world 3D / 4D space; communicating with at least one digital data input / output data transport path and encoding / decoding plenoptic sub-scene data to / from said input / output data transport path; and controlling the encoding / decoding and the locally stored three-dimensional (3D) sub-scene model database by generating and / or responding to data control packets communicated via the data transport path.

17. 17. The method of claim 16, wherein queries passed over the data transport path are answered by accessing identified scene segments in a plenoptic scene database.

18. 18. The plenoptic scene codec data processing method of claim 17, wherein communication with a remote plenoptic scene database is achieved via a data transport path, and sub-scene data on a server side of the data transport path is queried to support image generation on a client side of the data transport path.

19. 18. The plenoptic scene codec data processing method of claim 17, wherein communication with a remote plenoptic scene database is achieved via a data transport path, and plenoptic sub-scene data on a server side of the data transport path is provided to support plenoptic sub-scene image generation on a client side of the data transport path.

20. Communication with the data transport path is performed by a server side of the data transport path.

18. The plenoptic scene codec data processing method of claim 17, wherein the plenoptic scene codec data processing method is utilized to support image generation on a client side of the data transport path by inserting plenoptic subscenes into a plenoptic subscene database on the client side of the data transport path using the plenoptic scene database.

21. 1. A non-volatile digital computer program storage medium containing a computer program structure that, when executed by at least one digital data computer processor, performs a method for processing plenoptic projections using digital data processing circuitry and associated digital data memory, said method comprising: accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; and storing the processed output light field in the digital data memory.

22. 22. The non-volatile digital computer program storage medium of claim 21, wherein the digital data defining one or more input light fields is indexed by a plenoptic octree.

23. 22. The non-volatile digital computer program storage medium of claim 21, wherein the output light field is determined by a digital data processing circuit using integer arithmetic without general division operations.

24. 22. The non-volatile digital computer program storage medium of claim 21, wherein the one or more input volume elements are determined using one or more digital data defining a plenoptic selection mask representing characteristics including one or more of occupancy, resolution, and color.

25. Sub-scene extraction is accessing digital data defining one or more input light fields representing input solid angle elements arranged within volume elements indexed by a hierarchical, plenoptically sorted, multi-resolution plenoptic tree structure; accessing digital data defining one or more input volume elements indexed by a hierarchical, spatially sorted, multi-resolution volume tree structure; determining an output light field at the input volume element by traversing the tree and evaluating intersections between the input solid angle element and the input volume element; and storing the processed output light field in the digital data memory.

26. When executed by at least one digital data computer processor, A non-volatile digital computer program storage medium containing a computer program structure for executing a plenoptic scene codec data processing method, said method comprising: communicating with an associated memory circuit to stream hierarchical multi-resolution and plenoptically sorted sub-scene data packet structures representing light fields and / or matter fields as directionally and / or volumetrically subdivided data into a locally stored three-dimensional (3D) or 4D sub-scene model database as plenoptic increments of a remotely stored plenoptic scene model database of real-world 3D / 4D space; communicating with at least one digital data input / output data transport path and encoding / decoding plenoptic sub-scene data to / from said input / output data transport path; and controlling the encoding / decoding and the locally stored three-dimensional (3D) sub-scene model database by generating and / or responding to data control packets communicated over the data transport path.

27. 27. The non-volatile digital computer program storage medium of claim 26, wherein queries passed over the data transport path are responded to by accessing identified scene segments in a plenoptic scene database.

28. 28. The non-volatile digital computer program storage medium of claim 27, wherein communication with a remote plenoptic scene database is achieved via a data transport path, and sub-scene data on a server side of the data transport path is queried to support image generation on a client side of the data transport path.

29. 28. The non-volatile digital computer program storage medium of claim 27, wherein communication with the plenoptic scene database is achieved via a data transport path, and plenoptic sub-scene data on a server side of the data transport path is provided to support plenoptic sub-scene image generation on a client side of the data transport path.

30. 28. The non-volatile digital computer program storage medium of claim 27, wherein communication with a data transport path is utilized to support image generation on a client side of the data transport path using a plenoptic scene database on a server side of the data transport path by inserting plenoptic subscenes into a plenoptic subscene database on the client side of the data transport path.

Citation Information

Patent Citations

  • Multi-user system sharing three-dimensional virtual space

    JP2001076179A

  • Method and system for coding decoding image data

    JP2001186516A

  • Optimized plenoptic image encoding

    US20160241855A1

  • Quotidian scene reconstruction engine

    WO2017180615A1