Method, computing device and computer program for streaming scene-based immersive media
The 'smart client' mechanism optimizes immersive media delivery by determining client resource availability and asset reuse, reducing latency and resource consumption in heterogeneous networks.
Patent Information
- Application Number
- JP2024515921
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-08
- Filing Date
- 2023-03-10
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-03-10
AI Technical Summary
Existing media delivery networks struggle to efficiently deliver immersive media to heterogeneous client devices due to the need for significant conversion and streaming processes, which consume substantial network and computational resources without access to information about client cache availability.
A 'smart client' mechanism that communicates with a network server to determine client device resources and availability, allowing for optimized media streaming by deciding whether to convert media on the network or client, or a combination of both, based on cache availability and asset reuse.
This approach reduces latency and resource consumption by intelligently managing media conversion and streaming, enhancing the efficiency of immersive media delivery to diverse client devices.
Smart Images

Figure 0007733224000001 
Figure 0007733224000002 
Figure 0007733224000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 331,152, filed April 14, 2022, entitled "SmartClient for Streaming of Scene-Based Immersive Media," and claims the benefit of a continuation and priority of U.S. Provisional Patent Application No. 18 / 119,074, filed March 8, 2023, entitled "SmartClient for Streaming of Scene-Based Immersive Media," the disclosure of which is incorporated herein by reference in its entirety.
[0002] This disclosure generally describes embodiments relating to architectures, structures, and components for systems and networks that deliver media, including video, audio, geometric (3D) objects, haptics, associated metadata, or other content for client devices. Particular embodiments relate to systems, structures, and architectures for the delivery of media content to heterogeneous immersive and interactive client devices. [Background technology]
[0003] Immersive media generally refers to media that stimulates any or all human sensory systems (e.g., sight, hearing, somatosensation, smell, and possibly taste) to create or enhance the user's perception of being physically present in the media experience, such as beyond that delivered over existing (e.g., "legacy") commercial networks for timed two-dimensional (2D) video and corresponding audio; such timed media is also known as "legacy media." Immersive media may also be defined as media that seeks to create or mimic the physical world through digital simulation of dynamics and the laws of physics, thereby stimulating any or all human sensory systems to create the perception by the user of being physically present in a scene depicting the real or virtual world.
[0004] An immersive media-enabled presentation device may refer to a device with sufficient resources and capabilities to access, interpret, and present immersive media. Such an immersive media-enabled device supports multiple quantities and formats of media and multiple network resources necessary to deliver immersive media at scale. "At scale" may refer to delivery of media by a service provider that achieves delivery equivalent to legacy delivery of video and audio media over a network, e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.
[0005] In contrast, legacy presentation devices such as laptop displays, televisions, and mobile handset displays are uniform in their capabilities, as all of these devices consist of rectangular display screens that consume 2D rectangular video or still images as their primary visual media format. Some of the visual media formats commonly used by legacy presentation devices may include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266.
[0006] The term "frame-based" media refers to the property of visual media that it consists of one or more consecutive rectangular frames of an image. In contrast, "scene-based" media refers to visual media that is organized by "scenes," where each scene refers to individual assets that collectively describe a visual scene.
[0007] A comparison between frame-based and scene-based visual media is illustrated in the case of visual media showing a forest. In a frame-based representation, the forest is captured using a camera device, such as one provided with a mobile phone. The user focuses the camera on the forest, and the frame-based media captured is what the user sees through the camera viewport provided with the phone, including any camera movements initiated by the user. The resulting frame-based representation of the forest is a series of 2D image frames recorded by the camera at a standard rate, typically 30 or 60 frames per second. Each image frame is a collection of pixels where the information stored in each pixel is matched with the next pixel.
[0008] A scene-based representation of a forest is composed of individual assets that describe each object in the forest. For example, the scene-based representation may include individual objects called "trees," with each tree being composed of a collection of smaller assets called "trunks," "branches," and "leaves." Each tree trunk may be further described individually by a mesh that describes the trunk's complete 3D geometry and a texture that is applied to the tree-trunk mesh to capture the trunk's color and radiance characteristics. Furthermore, the trunk may be accompanied by additional information that describes the trunk's surface in terms of its smoothness or roughness, or its ability to reflect light. The individual assets that make up a scene vary in the type and amount of information stored in each asset.
[0009] The delivery of any media over a network may employ media delivery systems and architectures that reformat and / or convert the media from an input or network "ingested" media format to a distributed media format that is not only suitable for ingestion by a target client device and its applications, but also lends itself to being "streamed" over the network. The network typically performs two processes on the ingested media: 1) converting the media from format A to format B suitable for ingestion by the target client, i.e., based on the client's ability to ingest a particular media format, and 2) preparing the media to be streamed.
[0010] "Streaming" media broadly refers to fragmenting and / or packetizing media so that it can be delivered over a network in successive, smaller-sized "chunks" that are logically organized and sequenced according to either or both the media's temporal and spatial structure. Converting media from format A to format B (also known as "transcoding") may be a process typically performed by a network or service provider prior to delivering the media to a client. Such transcoding may consist of converting media from format A to format B based on prior knowledge that format B is the preferred or only format that can be ingested by the intended client, or that is more suitable for delivery over constrained resources, such as a commercial network. One example of media conversion is converting media from a scene-based representation to a frame-based representation. In some cases, both the steps of converting the media and preparing the media for streaming are required before the client can receive and process the media instead.
[0011] The above one-step or two-step process of operating on the ingested media by the network, i.e., before delivering the media to the client, results in a media format referred to as a "delivery media format" or simply "delivery format." Generally, these steps should be performed only once for a given media data object if the network has access to information indicating when the client will need the converted and / or streamed media object on multiple occasions. That is, the processing and transfer of data for converting and streaming media is generally considered a source of latency that must consume a potentially significant amount of network and / or computational resources. Therefore, there is a need for a network design that has access to information indicating when a client may already have a particular media data object stored in its cache or stored locally to the client. Summary of the Invention [Means for solving the problem]
[0012] According to embodiments, methods, systems, and devices are provided for facilitating the process of determining whether a client device already has access to a copy of a media asset and / or media object stored on a local cache managed by the client device, or whether a new copy of the media asset should be generated by a network server and inserted into a media stream session. According to embodiments, the processes disclosed herein may be performed by or together with the server or the client device.
[0013] According to one aspect of the present disclosure, there is provided a method for streaming scene-based media assets during a media streaming session performed by a computing device, the method including: providing a bidirectional interface for communicating information about the scene-based media assets between a network server and a computing device functioning as a client device; and providing client device attributes and corresponding information about availability of client device resources, when requested, to the network server via the interface, wherein the client device attributes and information are used by the computing device to render the scene-based media assets.
[0014] According to another aspect of the present disclosure, there is provided a computing device for streaming media assets during a media streaming session, the computing device including at least one memory configured to store computer program code and at least one processor configured to execute the computer program code to perform the aforementioned method for streaming scene-based media assets during a media streaming session.
[0015] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions that, when executed by at least one processor of a computing device, cause the computing device to perform the aforementioned method for streaming scene-based media assets during a media streaming session.
[0016] Additional embodiments will be set forth in the description that follows, and in part will be obvious from the description, and / or may be realized by practice of the illustrated embodiments of the present disclosure.
[0017] The foregoing and additional embodiments of the present invention will be more clearly understood as a result of the following detailed description of various aspects of the invention when taken in conjunction with the drawings, in which like reference numerals refer to corresponding parts throughout the several views of the drawing. [Brief explanation of the drawings]
[0018] [Figure 1A] 1 is an exemplary diagram of media delivery to client devices in a media streaming network, according to an embodiment. [Figure 1B] 1 is an exemplary workflow illustrating the creation of media in a distribution format and the generation of reuse indicators in a media streaming network, according to an embodiment. [Figure 2A] 1 is an exemplary workflow illustrating streaming media to a client device, according to an embodiment. [Figure 2B] 1 is an exemplary workflow illustrating streaming media to a client device, according to an embodiment. [Figure 2C] An exemplary workflow showing streaming media to a smart client, with the addition of a smart client that determines whether previously streamed media is still available in a nearby cache, locally or otherwise, and reports this to the network, according to an embodiment. [Figure 2D] 10 is an exemplary workflow illustrating streaming media to a Smart Client with the addition of the Smart Client obtaining device status, profile information and resource availability and transmitting them back to the network, according to an embodiment. [Figure 2E] 1 is an exemplary workflow illustrating streaming media to a Smart Client with the addition of a Smart Client that requests and receives streamed media from a network process, according to an embodiment. [Figure 3A]FIG. 1 illustrates an exemplary diagram of a data model for displaying and streaming timed immersive media, according to an embodiment. [Figure 3B] FIG. 1 illustrates an exemplary diagram of a data model for displaying and streaming timed immersive media with visual assets ordered in a list based on descending frequency values, according to an embodiment. [Figure 4A] FIG. 1 illustrates an exemplary diagram of a data model for displaying and streaming non-timed immersive media, according to an embodiment. [Figure 4B] FIG. 10 is an exemplary diagram of a data model for displaying and streaming non-timed immersive media with haptic and audio assets ordered based on increasing frequency values, according to an embodiment. [Figure 5] FIG. 1 illustrates an exemplary workflow illustrating natural media synthesis, according to an embodiment. [Figure 6] 1 is an exemplary workflow illustrating synthetic media capture generation, according to an embodiment. [Figure 7] FIG. 1 is an exemplary diagram of a computer system, according to an embodiment. [Figure 8] 1 is an exemplary diagram of a network media distribution system, according to an embodiment. [Figure 9] 1 is an exemplary workflow illustrating immersive media delivery using redundant caching, according to an embodiment. [Figure 10] FIG. 1 is a system diagram of a media adaptation process system, according to an embodiment. [Figure 11A] 1 is an exemplary workflow illustrating the creation of media in a distribution format, according to an embodiment. [Figure 11B] 1 is an exemplary workflow illustrating the creation of media in a distribution format with assets ordered by asset type and the frequency with which the assets are used across representations based on asset type, according to an embodiment. [Figure 12] 1 is an exemplary workflow illustrating a packetization process, according to an embodiment. [Figure 13] FIG. 2 illustrates an exemplary workflow illustrating communication flow between components, according to an embodiment. [Figure 14A] 1 is an exemplary workflow illustrating media reuse analysis, according to an embodiment. [Figure 14B] 1 is an example of a set of lists of specific assets for a scene being represented, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0019] The following detailed description of the exemplary embodiments refers to the accompanying drawings, in which the same reference numbers in different drawings may identify the same or similar elements.
[0020] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Moreover, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may occur simultaneously (at least in part), and the order of one or more operations may be rearranged.
[0021] It will be apparent that the systems and / or methods described herein may be implemented in different forms, such as hardware, software, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0022] Although particular feature combinations are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or disclosed herein. Although each dependent claim listed below may depend directly on only one claim, each dependent claim in combination with any other claim in the set of claims will be included in the disclosure of possible implementations.
[0023] The proposed features described below may be used individually or combined in any order. Furthermore, the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0024] No element, act, or instruction used in this application should be construed as essential or required unless expressly stated as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar language is used. Also, as used herein, terms such as "has," "have," "having," "include," and "including" are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless otherwise specified. Furthermore, phrases such as "at least one of [A] and [B]" and "at least one of [A] or [B]" should be understood to include A only, B only, or both A and B.
[0025] According to embodiments, an immersive media-enabled presentation device may refer to a device with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices are heterogeneous with respect to the amount and format of media they can support and the amount and type of network resources required to deliver such media on a large scale. "On a large scale" may refer to the delivery of media by a service provider that achieves delivery equivalent to the delivery of legacy video and audio media over a network, e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.
[0026] According to embodiments, client devices that serve as endpoints for the delivery of immersive media over a network are all highly diverse. The delivery of any media over a network may use a media delivery system and architecture that reformats the media from an input or network "ingested" media format into a delivered media format that is not only suitable for ingestion by a target client device and its applications, but also lends itself to being "streamed" over the network. Thus, there may be two processes performed on media ingested by the network: 1) converting the media from format A to format B that is suitable for ingestion by a target client, i.e., based on the client's ability to ingest a particular media format, and 2) preparing the media to be streamed.
[0027] In embodiments, streaming media broadly refers to the fragmentation and / or packetization of media, thereby allowing media to be delivered over a network in successive, smaller-sized chunks logically organized and sequenced according to either or both the temporal or spatial structure of the media. Converting media from format A to format B (sometimes referred to as “transcoding”) may be a process typically performed by a network or service provider prior to delivering the media to a client device. Such transcoding may consist of converting media from format A to format B based on prior knowledge that format B is the preferred or only format that can be captured by the target client device, or that is more suitable for delivery over constrained resources, such as a commercial network. In many, if not all, steps of converting media and preparing the media to be streamed are necessary before the media can be received from the network and processed by the target client device.
[0028] Converting (or transforming) media and preparing media for streaming are process steps that operate on media captured by a network before delivering the media to a client device. The result of the processing (i.e., converting and preparing for streaming) is a media format called a delivery media format, or simply a delivery format. These steps, when performed for a given media data object, should be performed only once if the network has access to information indicating that a client requires the converted and / or streamed media object for multiple occasions to trigger the conversion and streaming of such media multiple times. That is, the processing and transfer of data for media conversion and streaming is generally considered a source of latency that requires the consumption of potentially significant amounts of network and / or computational resources. Thus, network designs that do not have access to information indicating that a client may already have a particular media data object stored in a cache or locally to the client perform less optimally than networks that have access to such information.
[0029] A scene graph may be a general data structure commonly used by vector-based graphics editing applications and modern computer games, which arranges a logical and often (but not necessarily) spatial representation of a graphics scene, or may be a collection of nodes and vertices in a graph structure.
[0030] A scene, in the context of computer graphics, may be a collection of objects (e.g., 3D assets, which may also be known as media assets, media objects, objects, and assets), object attributes, and other metadata including visual, acoustic, and physics-based characteristics that describe a particular setting bounded by either space or time with respect to the interactions of objects within that setting.
[0031] A node may be a basic element of a scene graph, consisting of information related to the logical or spatial or temporal representation of visual, audio, tactile, olfactory, gustatory, or related processing information. Each node shall have at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) connected to it.
[0032] The base layer may be a nominal representation of the media asset, and is typically formulated to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network.
[0033] An enhancement layer may be a set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities not supported in the base layer.
[0034] An attribute may be metadata associated with a node that is used to describe a particular characteristic or feature of that node, either in a canonical form or in a more complex form (e.g., with respect to another node).
[0035] The container may be a serialized format for storing and exchanging information for representing an entire natural scene, an entire synthetic scene, or a combination of synthetic and natural scenes, including a scene graph and all media resources required to render the scene.
[0036] Serialization may be the process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link), and later reconstructed (possibly in a different computing environment). When the resulting series of bits is reread according to the serialized format, it can be used to create a semantically identical clone of the original object.
[0037] A renderer may be a (typically software-based) application or process based on a selective combination of academic disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development that, given an input scene graph and asset container, emits exemplary visual and / or audio signals suitable for presentation on a target device or adapted to desired characteristics specified by attributes of the nodes to be rendered in the scene graph. In the case of visual-based media assets, the renderer may emit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline). In the case of audio-based media assets, the renderer may emit audio signals for presentation over multi-channel loudspeakers and / or binauralized headphones, or for repackaging into another (output) container. Common examples of renderers include the real-time rendering capabilities of game engines such as Unity Engine and Unreal Engine.
[0038] The scripting language may be an interpreted programming language that can be executed by the renderer at runtime to process dynamic inputs and variable state changes applied to scene graph nodes, which changes affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformations, collisions) and energy propagation and transport (light, sound).
[0039] A shader may be a type of computer program, originally used for shading (producing appropriate levels of light and color in an image), but now performing a variety of specialized functions in various areas of computer graphics special effects, or video post-processing unrelated to shading, or even functions completely unrelated to graphics.
[0040] Path following is a computer graphics method for rendering three-dimensional scenes so that the lighting in the scene is realistic.
[0041] Timed media may include media and / or media objects that may be ordered by time, e.g., having start and end times according to a particular clock. Non-timed media may include media and / or media objects that may be organized by spatial, logical, or temporal relationships, e.g., an interactive experience realized according to actions taken by a user.
[0042] A neural network model (NN model) may be a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations that are applied to a visual signal to arrive at an improved visual output, which may include interpolation of new views of the visual signal that were not explicitly provided by the original signal.
[0043] The number of immersive media-enabled devices being introduced into the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and gaming consoles, has exploded over the past decade. Additionally, holographic displays and other forms of volumetric displays are poised to enter the consumer market within the next three to five years. However, despite the immediate or imminent availability of these devices, a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks has not materialized for several reasons.
[0044] One reason a coherent end-to-end ecosystem for the delivery of immersive media over commercial networks has not been realized is that the client devices that serve as endpoints in such delivery networks for immersive displays are all highly diverse. Some client devices support specific immersive media formats, while others do not. Some are capable of creating immersive experiences from legacy raster-based formats, while others are not. Unlike networks designed solely for the delivery of legacy media, networks that must support a variety of display clients require a significant amount of information about the details of each client's capabilities and the format of the media being delivered before such networks can use an adaptation process to convert the media into a format appropriate for each target display and corresponding application. At a minimum, such networks need access to information describing the characteristics of each target display and the complexity of the ingested media in order for the network to ascertain how to meaningfully adapt the input media source to a format appropriate for the target display and application.
[0045] Networks that support heterogeneous client devices should take advantage of the fact that some of the assets adapted from an input media format to a particular target format can be reused across a set of similar display targets. That is, once converted to a format appropriate for the target displays, some assets may be reused across several such displays that have similar adaptation requirements. Therefore, such networks that use caching mechanisms to store adapted assets in a relatively immutable area are more efficient.
[0046] Immersive media may be organized into "scenes" that are described by a scene graph, also known as a scene description. The scope of a scene graph may be to describe the visual, audio, and other forms of immersive assets that comprise a particular setting that is part of the representation, for example, actors and events that take place in a particular location within a building that is part of the representation, such as a movie. A list of all the scenes that comprise a single representation may be formulated in a scene manifest.
[0047] An advantage of the "scene"-based approach is that, for content that is prepared before such content must be delivered, a "bill of materials" can be created that identifies all of the assets used throughout the representation and the frequency with which each asset is used across the various scenes within the representation. The network has knowledge of the existence of cached resources that can be used to satisfy the asset requirements of a particular representation. Similarly, a client device that is rendering a series of scenes may want to have knowledge of the frequency with which any given asset is used across multiple scenes. For example, if a media asset (also known as a media object, asset, or object) is referenced multiple times across multiple scenes that have been or will be processed by the client, the client device should avoid discarding that particular asset from its cached resources until the last scene requiring that asset has been rendered by the client device.
[0048] For legacy rendering devices, the delivery format may be equivalent or sufficiently equivalent to the "representation format" ultimately used by the client rendering device to create the presentation. That is, the presentation media format is a media format whose properties (resolution, frame rate, bit depth, color gamut, etc.) are closely tailored to the capabilities of the client rendering device. Some examples of delivery formats versus presentation formats include a high-definition (HD) video signal (1920 pixel columns by 1080 pixel rows) delivered over a network to an ultra-high-definition (UHD) client device with a resolution (3840 pixel columns by 2160 pixel rows). The UHD client may apply a process called "super-resolution" to the HD delivery format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format presented by the client device is the "representation format," which in this example is a UHD signal, but the HD signal includes the delivery format. In this example, the HD signal distribution format is very similar to the UHD signal representation format since both signals are linear video formats, and the process of converting the HD format to the UHD format is relatively simple and easy to perform on most legacy client devices.
[0049] However, in some embodiments, the preferred representation format for the intended client device may differ significantly from the capture format received by the network. Nevertheless, the client device may have access to sufficient computational, storage, and bandwidth resources to convert the media from the capture format to the required representation format suitable for representation by the client device. In this scenario, the network may bypass steps of reformatting the captured media, such as “transcoding” the media from a first format A to a second format B, because the client has access to sufficient resources to perform all media conversion without the network having to do so beforehand. The network may still perform steps of fragmenting and packaging the captured media so that the media can be streamed to the client.
[0050] However, in some embodiments, the captured media received by the network differs significantly from the client's preferred representation format, and the client device does not have access to sufficient computational, storage, and / or bandwidth resources to convert the media to the preferred representation format. In the absence of access to resources, the network may assist the client by performing some or all of the conversion on behalf of the client from the capture format to a format that is equivalent or nearly equivalent to the client device's preferred representation format. In some embodiments, such assistance provided by the network on behalf of the client is commonly referred to as "split rendering."
[0051] The embodiments of the present disclosure described herein address the need for a "smart client" mechanism or process that responds to network server processes and participates in the architecture described herein on behalf of client devices. Such a "smart client" mechanism is particularly relevant to networks designed to stream scene-based media to immersive, heterogeneous, interactive client devices such that the delivery of the media is performed efficiently and within the constraints of the capabilities of the various components of the overall media streaming network. A "smart client" is associated with a particular client device that responds to network requests for information regarding the current state of its associated client device, including the availability of resources on the client device for rendering and creating representations of the scene-based media.
[0052] Note that the remainder of this disclosure assumes, without loss of generality, that a smart client that can respond on behalf of a particular client device can also respond on behalf of a client device on which one or more applications are active. That is, the problem of responding on behalf of a client device is equivalent to the problem of responding on behalf of a client device on which one or more applications are active.
[0053] Additionally, it should be noted that the terms "media object," "media asset," or "asset" may be used interchangeably and all refer to a specific instance of a particular format of media data. The term "client device" or "client" (without limitation) refers to the device and its components on which the rendering of media ultimately executes. The term "smart client" may refer to the client described in this disclosure or the network interactivity of a client having network interactivity capabilities.
[0054] FIG. 1A is an exemplary diagram of a media delivery process 100 for delivering media from a network cloud, edge device, or server 104 to a computing device functioning as a client device 108, according to an embodiment. As shown in FIG. 1A , media, including immersive media including one or more scenes and one or more media objects in a first format A (hereinafter, “ingested media format A”), is received from a content provider. The process (i.e., media flow process 100) may be executed or performed by a network cloud or edge device (hereinafter, “network device 104”) and delivered to a client, e.g., client device 108. In some embodiments, the same process may be performed in advance by a manual process or client device. The network device 104 may, for example, use a delivery module to ingest the media in the first format 101, generate and / or create delivered media 102 in a second format (hereinafter, “delivered media creation 102”), and deliver the media in the second format 103. The client device 108 may include a rendering module 106 and a presentation module 107.
[0055] According to one aspect, the network device 104 may receive ingested media from a content provider or the like. The media streaming network may obtain the ingested media stored in ingested media format A. Any necessary transformations or conditioning of the ingested media may be used to create and / or generate delivery media. That is, delivery formats of media objects within the ingested media may be created. As described above, the delivery format is a media format that can be delivered to a client by formatting the media into delivery format B. Delivery format B is a format that is ready to be streamed to the client device 108. The delivery media creation 102 may include optimization reuse logic that performs a decision-making process to determine whether a particular media object has already been streamed to the client device 108. Further operations related to the delivery media creation 102 and the optimization reuse logic are described in detail with reference to FIG. 1B.
[0056] Media formats A and B may or may not be representations that follow the same syntax of a particular media format specification, but format B is likely tailored to facilitate delivery of the media over a network protocol. The network protocol may be, for example, a connection-oriented protocol (TCP) or a connectionless protocol (UDP). A delivery module of network device 104 streams the streamable media (i.e., media format B) from network device 104 to client device 108 over network connection 105.
[0057] Client device 108 can receive delivered media and render media for presentation using rendering module 106. Rendering module 106 has access to several rendering capabilities that can be rudimentary or similarly sophisticated, depending on the targeted client device 108. Rendering module 106 can create presentation media in presentation format C. Presentation format C may or may not be presented according to a third format specification. Thus, presentation format C may be the same as or different from media formats A and / or B. Rendering module 106 outputs presentation format C to presentation module 107, which can present the presentation media within a display (or the like) of client device 108.
[0058] Embodiments of the present disclosure facilitate the decision-making process used by a network to calculate a sequential order for packaging and streaming assets from the network to clients. In this case, all assets utilized across a set of one or more scenes that make up a representation are analyzed by a media reuse analyzer to determine how frequently each asset is used across all scenes that make up the representation. Thus, the order in which assets for a particular scene are packaged and streamed to clients can be based on how frequently each asset is used across the set of scenes that make up the representation.
[0059] Embodiments address the need for a mechanism or process to analyze an immersive media scene to obtain sufficient information that can be used to support a decision-making process, when employed by a network or a client, to provide instructions as to whether conversion of media objects from format A to format B should be performed entirely by the network, entirely by the client, or via a mix of both (with instructions as to which assets should be converted by the client or the network). Such an immersive media data complexity analyzer can be used by either the client or the network in an automated situation, or manually, for example, by a human operating a system or device.
[0060] According to embodiments, the process of adapting an input immersive media source to a particular endpoint client device may be the same as or similar to the process of adapting the same input immersive media source to a particular application running on a particular client endpoint device. Thus, the problem of adapting an input media source to the characteristics of an endpoint device is of the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.
[0061] 1B is a workflow for distributed media creation 102, according to an embodiment. More specifically, the workflow of FIG. 1B is a decision-making process for reuse logic that assists in the decision-making process for determining whether a particular media object has already been streamed to a client device 108.
[0062] At 152, a delivery media creation process begins. At 155, conditional logic may be executed to determine whether the current media object has previously been streamed to the client device 108. To determine whether the media object has previously been streamed to the client, a list of unique assets may be accessed for the representation. If the current media object has previously been streamed, the process proceeds to operation 160. At operation 160, an indicator (hereinafter referred to as a "proxy") is created to identify that the client has already received the current media object and should access a copy of the media object from a local cache or other cache. If it is determined that the media object has not previously been streamed, the process proceeds to operation 165. At operation 165, the media object may be prepared for conversion and / or delivery, and a delivery format for the media object is created. The process for the current media object then ends.
[0063] 2A is an exemplary workflow for processing media ingested over a network. The workflow shown in FIG. 2A illustrates a media conversion decision-making process 200 according to an embodiment. The media conversion decision-making process 200 is employed to determine whether the network should convert the media before delivering it to a client device. The media conversion decision-making process 200 may be handled manually or via an automated process within the network.
[0064] Ingested media, represented in format A, is provided to the network by a content provider. At 205, the media is ingested by the media streaming network from the content provider. At 210, attributes of the target client are obtained, if not already known. The attributes describe the processing capabilities of the target client.
[0065] At 215, it is determined whether the network (or client) should support conversion of the ingested media. In some embodiments, at 215, it is specifically determined whether there will be a format conversion for any of the media assets contained within the ingested media (e.g., conversion of one or more media objects from format A to format B) before the media is streamed to the target client. At 215, the determination may be based on a determination of whether the media can be streamed in its original ingested format A or whether it must be converted to a different format B to facilitate presentation of the media by the client. Such decision-making (i.e., to determine whether conversion of the ingested media is necessary before streaming the media to the client, or whether the media needs to be streamed directly to the client in its original ingested format A) may require access to information that describes aspects or characteristics of the ingested media.
[0066] If it is determined that the network (or client) should support conversion of any media assets (“YES” at 215 ), the process 200 proceeds to 220 .
[0067] At 220, the captured media is converted from format A to format B to produce converted media 222. The converted media 222 undergoes a process at 225 to prepare the media for streaming to a client, in which case the converted media 222 (i.e., the input media at 225) is prepared to be streamed.
[0068] Streaming of immersive media may be in its relatively early stages, especially if such media is “scene-based” rather than “frame-based.” For example, streaming of frame-based media may be equivalent to streaming frames of video, with each frame capturing a complete picture of an entire scene or object being presented by a client. The sequence of frames is reconstructed by the client from their compressed form and, when rendered to a viewer, creates a video sequence containing the entire immersive representation or a portion of the representation. In the case of frame-based streaming, the order in which frames are streamed from the network to the client may conform to a predetermined specification (e.g., ITU-T Recommendation H.264 Advanced Video Coding for General Audio-Visual Services). However, unlike frame-based streaming, scene-based streaming of media may consist of individual assets, which may themselves be independent of one another. A given scene-based asset may be used multiple times within a particular scene or across a series of scenes. The amount of time a client, or any given renderer, needs to spend reconstructing a particular asset may depend on many factors, including, but not limited to, the size of the asset, the availability of computational resources to perform the rendering, and other attributes that describe the overall complexity of the asset. A client that supports scene-based streaming may require some or all of the rendering for each asset in a scene to be completed before any of the representations of the scene can begin. Thus, the order in which assets are streamed from the network to clients can affect the overall performance of the system.
[0069] The conversion of media from format A to another format (e.g., format B) may be performed entirely by the network, entirely by the client, or between both the network and the client together. It becomes clear that with split rendering, a dictionary of attributes describing media formats may be needed so that both the client and the network have complete information characterizing the work they must do. Furthermore, a dictionary providing attributes of the client's capabilities, e.g., with respect to available computational resources, available storage resources, and access to bandwidth, may similarly be needed. Furthermore, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of an ingested media format so that the network and client together, or alone, may determine whether or when the network employs a split rendering process to deliver media to the client.
[0070] If it is determined that the network (or client) should not (or does not need to) support conversion of any media assets ("No" at 215), process 200 proceeds to 225. At 225, the media is prepared for streaming. In this case, the ingested data (i.e., the media in its original form) is prepared to be streamed.
[0071] Finally, once the media data is in a streamable format, the prepared media at 225 is streamed to the client at 230. In some embodiments, if conversion and / or streaming of certain media objects needed or required by the client to complete the rendering of the media can be avoided (as described with reference to FIG. 1B ), the network can skip converting and / or streaming of the ingested media (i.e., 215-230), assuming the client still has access to or availability of the media objects that may be needed to complete the rendering of the client's media. To facilitate a client's ability to perform at its full potential, it may be desirable for the network to have sufficient information regarding the order in which scene-based assets are streamed from the network to the client so that the network can determine such an order to improve the client's performance. For example, such a network with sufficient information to avoid repetitive conversion and / or streaming steps of assets used multiple times in a particular rendering can perform more optimally than a network not designed in this way. Similarly, a network that can “intelligently” order the delivery of assets to a client can facilitate a client's ability to perform at its full potential (i.e., to create a more enjoyable experience for the end user).
[0072] 2B illustrates an exemplary media conversion process 250 that includes an asset reuse process for determining media asset reuse and redundant caching, according to an embodiment. Similar to the media conversion decision-making process 200, the media conversion process 250 using the asset reuse and redundant caching process ingests media over a network and determines whether the network should convert the media before delivering the media to a client.
[0073] Ingested media, represented in format A, is provided to the network by a content provider. According to one embodiment, operations 255-260 and 275-286 are performed similarly to operations 205-210 and 215-230 shown in FIG. 2A. At 255, media is ingested by the network from the content provider. Next, at 260, attributes of the target client are obtained, if not already known. The attributes describe the processing capabilities of the target client.
[0074] If it is determined that the network has previously streamed the particular media object or the current media object (“Yes” at 265), the process proceeds to 270. At 270, a proxy is created in place of the previously streamed media object, indicating that the client should use its local copy of the previously streamed object or a copy of the previously streamed object stored in a redundant cache managed by the media streaming network.
[0075] If it is determined that the network has not previously streamed the media object (NO at 265), the process proceeds to 275. At 275, it is determined whether the network or client should perform any format conversion on any of the media assets contained within the ingested media at 255. For example, the conversion may include converting a particular media object from format A to format B before the media is streamed to the client. Operation 275 may be similar to that performed at 215 shown in FIG. 2A.
[0076] If it is determined that the media asset should be converted over the network (YES at 275), the process proceeds to 280. At 280, the media object is converted from format A to format B. The converted media is then prepared to be streamed to the client (286).
[0077] If it is determined that the media asset should not be converted over the network (No at 275), the process proceeds to 285. At 285, the media object is prepared to be streamed to the client. Once the media is in a streamable format, the media prepared at 285 is streamed to the client at 286.
[0078] FIG. 2C shows an example process 2500 of how the network side assists the client-side smart client in determining whether previously streamed media is still available in a local or other nearby cache, or whether to use asset query logic implemented on the client side and the network side. Network-side operations 255-265 are similar to operations 255-265 shown in FIG. 2B, but with an asset query logic process implemented on the client side to illustrate the smart client's role in determining when a particular asset is already accessible to the client. At 255, media is ingested into the network. At 260, the network obtains attributes describing the processing capabilities of the target client (not shown). The network then initiates a request to the client-side asset query logic to determine whether the asset should be streamed to the client. At 265-1, the smart client receives information from the network regarding assets needed by the smart client for rendering. At 265-3, the smart client accesses a list of previously streamed assets stored in a local cache or network-based storage (e.g., a database) based on a request from the network. At 265-2, the smart client determines whether the requested asset is already accessible to the client device and returns information regarding whether the asset should be streamed to the network's asset query logic regarding whether the asset in question should be streamed. If the asset does not need to be streamed, the network creates a proxy for the asset in question at 270 and inserts the proxy into the media at 285 to prepare it for streaming to the client, as shown in Figure 2B. If the asset does need to be streamed to the client, the network then proceeds to determine at 275 whether the network should support any conversion of the media being streamed to the client.If such conversion is necessary ("yes" at 275), the network performs the conversion accordingly on the media asset and then merges the converted media asset into the media streamed to the client at 285. If the network determines that such conversion is not necessary ("no" at 275), the network prepares the media with the original media asset to be streamed to the client. Finally, at 286, the network streams the media to the client.
[0079] FIG. 2D illustrates an exemplary process 2600 of how a smart client obtains its device status, profile information, and resource availability and returns such information to the network in conjunction with the asset reuse logic described above. Process 2600 is similar to process 2500 illustrated in FIG. 2C, but with a client attribute obtainment operation performed by the smart client on the client side to indicate the smart client's role in obtaining information about the client device's capabilities, including the status, profile, and availability of resources used to render the media. At 255, media is ingested into the network and requested by a client device (not shown). The network then initiates a request to the smart client to obtain device attributes and resource status at 260. The smart client then initiates a query to the client device for its device attributes and resource availability at 260-1. The client device then responds to the smart client with attributes about the client device (e.g., memory, GPU, CPU, and their current status) at 260-3 and descriptive information about other available resources, including its profile and resource availability for processing future work, at 260-2. The smart client then delivers to the network at 260-4 the target client's processing capabilities and client device attributes describing the status and availability of resources on the client device. Based on the received information about the client, the network then determines whether the asset should be streamed to the client at 265. If the asset does not need to be streamed, the network creates a proxy for the asset in question at 270 and inserts the proxy into the media at 285 to prepare it for streaming to the client, as shown in Figure 2B.If the asset needs to be streamed to the client, the network then proceeds to determine whether the network should support any conversion of the media streamed to the client at 275. If such conversion is required ("Yes" at 275), the network performs the conversion on the media asset accordingly and then merges the converted media asset with the media streamed to the client at 285. If the network determines that such conversion is not required ("No" at 275), the network prepares the media with the original media asset to be streamed to the client. Finally, the network streams the media to the client at 286, as shown in FIG. 2B.
[0080] FIG. 2E illustrates an exemplary process 2700 of how a Smart Client requests and receives streamed media from a network. Process 2700 is similar to process 2600 illustrated in FIG. 2D, with the addition of client-side Smart Client-managed media requesting and receiving. As illustrated in FIG. 2E, the Smart Client first receives a media request from a user at 255-1. For example, the request may be received by a client device and then forwarded to the Smart Client at 255-1. The Smart Client then forwards the media request to the network server at 255-2. Upon receiving the media request, the network server initiates steps to obtain device attributes and resource status at 260, as illustrated in FIG. 2B. The status and availability of resources on the client device are returned to the network server. The network then initiates asset query logic, as illustrated in FIG. 2B, to determine whether an asset should be streamed to the client at 265. If the asset does not need to be streamed, the network creates a proxy for the asset in question at 270 and inserts the proxy into the media at 285 to prepare it for streaming to the client, as shown in FIG. 2B. If the asset does need to be streamed to the client, the network then proceeds to determine whether the network should support any conversion of the media to be streamed to the client at 275. If such conversion is required ("Yes" at 275), the network performs the conversion on the media asset accordingly and then merges the converted media asset into the media to be streamed to the client at 285. If the network determines that such conversion is not required ("No" at 275), the network prepares the media with the original media asset to be streamed to the client. Finally, the network streams the media to the client at 286, as shown in FIG. 2B.The smart client then stores the media in the media store at 255-3, and the client device can access the media from the media store at 255-4.
[0081] Figures 3A-3B refer to timed scenes for timed media, while Figures 4A-4B refer to untimed scenes for untimed media. Timed and untimed scenes may correspond to different scene representations or scene descriptions. Figures 3A-3B and 4A-4B both employ a single exemplary generic media format adapted from a source ingest media format to match the capabilities of a particular client endpoint. That is, the generic media format is a delivery format that can be streamed to a client device. The generic media format is robust in its structure to accommodate a wide variety of media attributes, each layer being layered based on the amount of salient information each layer contributes to the representation of the media.
[0082] As shown in FIG. 3A, a timed media presentation 300 includes a timed scene manifest 300A that includes a list of scene information 301. Each scene information 301 points to a list of components 302 that separately describe the processing information and types of media assets (e.g., asset lists and other processing information) that make up the scene information 301. The list of components 302 may reference proxy assets 308 that correspond to asset types (e.g., proxy visual and auditory assets, as shown in FIG. 3). The components 302 point to a list of unique assets that have not been previously used in other scenes. For example, the list of unique assets 307 for (timed) Scene 1 is shown in FIG. 3. The components 302 also point to assets 303, which include a base layer 304 and an attribute enrichment layer 305. The base layer is a nominal representation of an asset that can be formulated to minimize computational resources, the time required to render the asset, and / or the time required to transmit the asset over a network. In this exemplary embodiment, each of the base layers 304 points to a numerical frequency metric that indicates the number of times the asset was used across the set of scenes that comprise the presentation. An enhancement layer may be a set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities that may not be supported in the base layer.
[0083] Figure 3B shows a timed media display 3100 sorted in descending order of frequency. This timed media display 3100 is similar to the timed media display 300 shown in Figure 3A, except that the assets shown in Figure 3B are ordered in the list by asset type and descending frequency values within each asset type. The timed scene manifest 303A includes a list of scene information 3031. Each scene information 3031 points to a list of components 3032 that separately describe the processing information and type of media asset that comprise the scene information 3031. The components 3032 reference assets 3033, which in turn reference base layers 3034 and attribute enhancement layers 3035. Each of the base layers 3034 is ordered according to descending values in the corresponding frequency metric. A list of unique assets not previously used in other scenes is provided at 3037.
[0084] As shown in FIG. 4A , untimed media and complexity representation 400 includes scene information 401. Scene information 401 is not associated with a start and end time / duration (according to a clock, timer, etc.). An untimed scene manifest (not shown) may reference scene 1.0 without other scenes that can branch into it. Scene information 401 points to a list of components 402 that separately describe the types of processing information and media assets that make up scene information 401. Components 402 point to visual assets, audio assets, haptic assets, and timed assets (collectively referred to as assets 403). Assets 403 further point to base layer 404 and attribute enrichment layers 405 and 406. In this exemplary embodiment, each of base layers 404 points to a numerical frequency value that indicates the number of times the asset was used across the set of scenes that make up the representation. Scene information 401 can also point to other non-timed scenes that are for non-timed media sources (i.e., referenced in FIG. 4A as non-timed scenes 2.1-2.4), and / or scene information 407 that is for a timed media scene (i.e., referenced in FIG. 4A as timed scene 3.0). In the example of FIG. 4A, the non-timed immersive media includes a set of five scenes (including both timed and non-timed). List of unique assets 408 identifies unique assets associated with a particular scene that have not previously been used in a higher-level (e.g., parent) scene. List of unique assets 408 shown in FIG. 4A includes unique assets for non-timed scene 2.3.
[0085] FIG. 4B illustrates untimed media and an ordered complexity representation 4100. The untimed scene manifest (not shown) references scene 1.0, which does not have other scenes that can branch into it. Scene information 4041 is not associated with a start and end time according to a clock or the like. Scene information 4041 also points to a list of components 4042 that separately describe the processing information and media asset types that comprise scene information 4041. Components 4042 reference assets 4043, which in turn reference base layers 4044 and attribute enrichment layers 4045, 4046. In this embodiment, each base layer 4044 points to a numerical frequency value indicating the number of times the asset is used across the set of scenes that make up the representation. In this embodiment, haptic assets 4043 are organized by increasing frequency values, and audio assets 4043 are organized by decreasing frequency values. Scene information 4041 also points to other scene information 4041 for untimed media. Scene information 4041 also points to scene information 4047 for timed media scenes. List 4046 identifies unique assets associated with a particular scene that have not previously been used in a higher-level (eg, parent) scene.
[0086] Media streamed according to the generic media format is not limited to legacy visual and audio media. The generic media format may include any type of media information that can interact with a machine to generate signals that stimulate a human's senses of sight, hearing, taste, touch, or smell. As shown in Figures 3A-3B and 4A-4B, media streamed according to the generic media format may be timed media, untimed media, or a combination of both. The generic media format is streamable by using a base layer and enhancement layer architecture to enable layered representation of media objects.
[0087] In some embodiments, the separate base and enhancement layers are computed by applying multi-resolution or multi-mosaic analysis techniques to the media objects within each scene, which computation techniques are not limited to raster-based visual formats.
[0088] In some embodiments, the progressive representation of the geometric object may be a multi-resolution representation of the object computed using wavelet analysis techniques.
[0089] In some embodiments, in a layered display media format, enhancement layers may apply different attributes to the base layer. For example, one or more enhancement layers may modify the material properties of the surface of a visual object represented by the base layer.
[0090] In some embodiments, in a layered display media format, the attribute may modify the texture of the surface of the object represented by the base layer, for example, by changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.
[0091] In some embodiments, in a layered presentation media format, the surfaces of one or more visual objects in a scene may be changed from Lambertian surfaces to ray-traceable surfaces.
[0092] In some embodiments, in a layered representation media format, the network may deliver a base layer representation to the client so that the client can create a nominal representation of the scene, and the client awaits the transmission of additional enhancement layers to refine the resolution or other characteristics of the base layer.
[0093] In embodiments, attribute resolution or refinement information in the enhancement layer is not explicitly coupled to the resolution of objects in the base layer. Furthermore, a generic media format may support any type of information media that can be rendered or acted upon by a rendering device or machine, thereby enabling support of heterogeneous media formats to heterogeneous client endpoints. In some embodiments, a network delivering a media format first queries a client endpoint to determine the client's capabilities. If, based on the query, the client is unable to meaningfully capture the media representation, the network may remove layers of attributes not supported by the client. In some embodiments, if the client is unable to meaningfully capture the media representation, the network may adapt the media from its current format to a format appropriate for the client endpoint. For example, the network may adapt the media by converting a stereoscopic visual media asset into a 2D representation of the same visual asset using a network-based media processing protocol. In some embodiments, the network may adapt the media by employing a neural network (NN) process to reformat the media into an appropriate format or, optionally, by synthesizing views required by the client endpoint.
[0094] A scene manifest for a complete (or partially complete) immersive experience (such as a live streaming event, a game, or playback of an on-demand asset) is organized by scenes, each containing the minimum amount of information necessary to render and ingest to create a representation. The scene manifest includes a list of individual scenes to be rendered in their entirety for the immersive experience requested by the client. Each scene is associated with one or more representations of the geometric objects in the scene that correspond to a streamable version of the scene geometry. One embodiment of a scene may reference a low-resolution version of the scene's geometric objects. Another embodiment of the same scene may reference enhancement layers for the low-resolution representation of the scene to add further detail to the geometric objects of the same scene or to increase mosaicking. As described above, each scene may have one or more enhancement layers to progressively increase the detail of the scene's geometric objects. Each layer of a media object referenced in a scene may be associated with a token (e.g., a uniform resource identifier (URI)) that points to an address where the resource can be accessed in the network. Such resources are similar to a content delivery network (CDN) from which content can be fetched by a client. The tokens for representations of geometric objects may point to a location within the network or a location within the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.
[0095] According to embodiments, a scene (timed or untimed) can be addressed by a scene graph as a multi-planar image (MPI) or a multi-spherical image (MSI). Both MPI and MSI technologies are examples of technologies that support the creation of display-independent scene representations for natural content (i.e., real-world images captured simultaneously from one or more cameras). Scene graph technology, on the other hand, can be employed to represent both natural and computer-generated images in the form of a synthetic representation. However, such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. Scene graph representations of naturally captured content are both time- and computationally intensive to create, requiring complex analysis of the natural imagery using photogrammetry and / or deep learning techniques to create a synthetic representation that can later be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations cannot be practically created in real time to accommodate use cases requiring real-time delivery, making them impractical to consider as candidates for displaying natural content. Therefore, the best representation of a computer-generated image is to use a scene graph in a synthetic model, because computer-generated imagery is created using 3D modeling processes and tools, and using a scene graph in a synthetic model results in the best representation of a computer-generated image.
[0096] 5 illustrates an example of a natural media composition process 500, according to an embodiment. The natural media composition process 500 converts a capture format from a natural scene into a representation that can be used as a capture format for a network that serves heterogeneous client endpoints. To the left of the dashed line 510 is the content capture portion of the natural media composition process 500. To the right of the dashed line 510 is the capture format composition (for natural images) of the natural media composition process 500.
[0097] As shown in FIG. 5 , the first camera 501 uses a single camera lens to capture a scene of, for example, a person (i.e., an actor shown in FIG. 5 ). The second camera 502 captures a scene with five diverging fields of view by mounting five camera lenses around a ring-shaped object. The arrangement of the second camera 502 shown in FIG. 5 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. The third camera 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter of a sphere. The arrangement of the third camera 503 is an exemplary arrangement commonly used to capture a light field or a light field for a holographic immersive display. The embodiment is not limited to the configuration shown in FIG. 5 . The second camera 502 and the third camera 503 may include multiple camera lenses.
[0098] Natural image content 509 is output from a first camera 501, a second camera 502, and a third camera 503 and serves as input to a synthesizer 504. The synthesizer 504 may employ NN training 505, which uses a set of training images 506 to generate a capture NN model 508. The training images 506 may be predefined or may be stored from a previous synthesis process. A NN model (e.g., capture NN model 508) is a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations applied to the visual signal to arrive at an improved visual output, which may include interpolation of new views of the visual signal that were not explicitly provided by the original signal.
[0099] In some embodiments, a photogrammetry process may be performed in place of NN training 505. If a capture NN model 508 is created during natural media composition process 500, the capture NN model 508 becomes one of the assets in the capture format 507 of the natural media content. The capture format 507 may be, for example, MPI or MSI. The capture format 507 may also include media assets.
[0100] 6 illustrates an example of a synthetic media capture creation process 600, according to an embodiment. The synthetic media capture creation process 600 creates a captured media format for synthetic media, such as computer-generated images.
[0101] As shown in FIG. 6 , a camera 601 may capture a point cloud 602 of a scene. The camera 601 may be, for example, a LIDAR camera. A computer 603 employs, for example, a common gateway interface (CGI) tool, a 3D modeling tool, or another animation process to create synthetic content (i.e., a representation of the synthetic scene that can be used as a network ingestion format to serve heterogeneous client endpoints). The computer 603 may create CGI assets 604 over the network. Additionally, a sensor 605A may be attached to an actor 605 in the scene. The sensor 605A may be, for example, a motion capture suit with sensors attached. The sensor 605A captures a digital recording of the actor's 605 movements to generate animated motion data 606 (or MoCap data). Data from the point cloud 602, CGI assets 604, and motion data 606 are provided as input to a synthesizer 607, which creates a synthetic media ingestion format 608. In some embodiments, the synthesizer 607 may use the NN and training data to create a NN model and generate the synthetic media capture format 608.
[0102] Both natural and computer-generated (i.e., synthetic) content can be stored in a container. The container can include a serialization format for storing and exchanging information representing all-natural, all-synthetic, or a mix of synthetic and natural scenes, including a scene graph and all media resources needed to render the scene. The content serialization process involves converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link) and later reconstructed in the same or a different computer environment. When the resulting series of bits is reread according to the serialization format, it can be used to create a semantically identical clone of the original object.
[0103] The dichotomy in optimal display of both natural and computer-generated content suggests that the optimal capture format for naturally captured content is different from the optimal capture format for computer-generated (i.e., synthetic) content or natural content that is not essential for real-time distribution applications. Thus, according to embodiments, the network aims to be robust enough to support multiple capture formats for visually immersive media, whether created naturally using, for example, a physical camera, or created by a computer.
[0104] Technologies such as OTOY's ORBX, Pixar's Universal SceneDescription, and the Graphics Language Transmission Format 2.0 (glTF2.0) specification written by the Khronos 3D Group embody the scene graph as a format suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content where deep learning or photogrammetry techniques are used to create corresponding synthetic representations of natural scenes (i.e., not required for real-time delivery applications).
[0105] OTOY's ORBX is one of several scene graph technologies capable of supporting any type of visual media, timed or untimed, including ray-traceable, legacy (frame-based), stereoscopic, and other types of composited or vector-based visual formats. ORBX is different from other scene graphs because it natively supports freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph intentionally designed to facilitate interchange across multiple vendor technologies that operate on the scene graph. Additionally, ORBX offers a rich material system, support for an open shader language, a robust camera system, and support for Lua scripting. ORBX is also the basis for the Immersive Technologies Media Format, released for royalty-free licensing by the Immersive Digital Experience Alliance (IDEA). In the context of real-time media distribution, the ability to create and deliver ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of data captured by cameras and composition of that same data into composite representations.
[0106] Pixar's USD is a scene graph widely used in visual effects and professional content creation. USD is integrated into Nvidia's Omniverse platform, a set of developer tools for 3D model creation and rendering using Nvidia's graphics processing units (GPUs). A subset of USD published by Apple and Pixar is called USDZ, which is supported by Apple's ARKit.
[0107] glTF 2.0 is a version of the Graphics Language Transmission Format specification written by the Khronos 3D Group. This format supports simple scene graph formats, including PNG and JPEG image formats, that are generally capable of supporting static (untimed) objects in a scene. glTF 2.0 supports simple animation, which includes support for translation, rotation, and scaling of basic shapes (i.e., geometric objects) described using glTF primitives. glTF 2.0 does not support timed media, and therefore does not support video or audio media input.
[0108] These designs for immersive visual media scene displays are provided by way of example only and do not limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source to a format appropriate for the particular characteristics of a client endpoint device. Additionally, any or all of the above exemplary media displays may employ or employ deep learning techniques to train and create neural network (NN) models that enable or facilitate the selection of specific views to fill a particular display's viewing frustum based on the particular dimensions of the frustum. The views selected for a particular display's viewing frustum may be interpolated from existing views explicitly provided in the scene display, for example, from MSI or MPI techniques. Views may also be rendered directly from a rendering engine based on specific virtual camera positions, filters, or virtual camera descriptions for these rendering engines.
[0109] The methods and devices of the present disclosure are robust enough to take into account that there is a relatively small but well-known set of immersive media capture formats that can adequately meet the requirements for both real-time or on-demand (e.g., non-real-time) delivery of media that is captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.
[0110] Interpolation of views from immersive media ingest formats using either NN models or network-based rendering engines will become even easier as advanced network technologies (e.g., 5G for mobile networks) and fiber optic cables are deployed in fixed networks. These advanced network technologies increase the capacity and capability of commercial networks, as such advanced network infrastructure can support the transmission and delivery of increasingly large amounts of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networks (SDN), and network functions virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands on specific network resources, for example, to respond to dynamic increases and decreases in demand for network throughput, network speed, round-trip latency, and computational resources. Furthermore, this inherent ability to adapt to dynamic network requirements similarly facilitates the network's ability to adapt immersive media ingest formats to appropriate delivery formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.
[0111] Immersive media applications themselves may also have different requirements for network resources, including gaming applications that require significantly lower network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications that may have increasing demands on downlink resources depending on the type of display at the client endpoint that is consuming the data. In general, any consumer application may be supported by a variety of client endpoints that include different on-board client capabilities for storage, computation, and power, as well as equally different requirements for the particular media display.
[0112] Thus, embodiments of the present disclosure enable well-equipped networks, i.e., networks that utilize some or all of the characteristics of modern networks, to simultaneously support multiple legacy and immersive media-enabled devices according to the characteristics specified within the devices. Thus, the immersive media delivery methods and processes described herein provide the flexibility to leverage practical media ingest formats for both real-time and on-demand use cases for the delivery of media, the flexibility to support natural and computer-generated content for both legacy and immersive media-enabled client endpoints, and support both timed and untimed media. The methods and processes also dynamically adapt the source media ingest format to an appropriate delivery format based on the client endpoint's capabilities and capabilities and the application's requirements. This ensures that the delivery format is streamable over an IP-based network and enables the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media-enabled devices. Furthermore, embodiments provide an exemplary media presentation framework that facilitates the organization of delivered media along scene boundaries.
[0113] An end-to-end implementation of heterogeneous immersive media delivery according to an embodiment of the present disclosure that provides the aforementioned improvements is achieved according to the processes and components described in the detailed description of Figures 7-14, which are described in further detail below.
[0114] The techniques for representing and streaming heterogeneous immersive media described above may be implemented as computer software using computer-readable instructions at both the source and destination, and may be physically stored on one or more non-transitory computer-readable media or by one or more specifically configured hardware processors. Figure 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0115] Computer software may be coded using any suitable machine code or computer language that may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that may be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or via interpretation, microcode execution, etc.
[0116] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0117] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 700.
[0118] The computer system 700 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input. The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0119] The input human interface devices may include one or more of a keyboard 701, a trackpad 702, a mouse 703, a screen 709, which may be, for example, a touchscreen, a joystick 704, a microphone 705, a camera 706, and a scanner 707 (only one of each is shown).
[0120] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via the screen 709, data gloves, or joystick 704, although there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers 708, headphones, etc.), visual output devices (e.g., screens 709, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, and each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or three-dimensional hypervisible output via means such as stereographic output, virtual reality glasses, holographic displays, and smoke tanks), and printers.
[0121] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 711 with CD / DVD or similar media 710, thumb drives 712, removable hard drives or solid state drives 713, legacy magnetic media such as tape and floppy disks, and specialized ROM / ASIC / PLD-based devices such as security dongles.
[0122] Additionally, those skilled in the art will understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0123] The computer system 700 may also include a bidirectional interface 715 to one or more communications networks 714. The network 714 may be, for example, wireless, wired, or optical. The network 714 may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks 714 include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; wired or wireless wide-area digital networks for TV, including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus; and the like. A particular network 714 typically requires an external network interface adapter (e.g., a graphics adapter 725) attached to a particular general-purpose data port (e.g., a USB port on the computer system 700) or peripheral bus 716; other networks are typically integrated into the core of the computer system 700 by attaching to the system bus 748 described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 714, computer system 700 may communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or two-way, e.g., to other computer systems using local-area or wide-area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.
[0124] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core 717 of the computer system 700 .
[0125] The core 717 may include one or more central processing units (CPUs) 718, graphics processing units (GPUs) 719, specialized programmable processing units in the form of field programmable gate arrays (FPGAs) 720, hardware accelerators 721 for specific tasks, etc. These devices may be connected via a system bus 748, along with read-only memory (ROM) 723, random access memory (RAM) 724, and internal mass storage 722, such as an internal non-user-accessible hard drive, SSD, etc. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 748 or via a peripheral bus 716. Architectures for peripheral buses include PCI, USB, etc.
[0126] The CPU 718, GPU 719, FPGA 720, and accelerator 721 may execute specific instructions that, in combination, may constitute the machine code (or computer code) described above. That computer code may be stored in ROM 723 or RAM 724. Also, temporary data may be stored in RAM 724, while persistent data may be stored in, for example, internal mass storage 722. Fast storage and retrieval from any memory device may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 718, GPU 719, mass storage 722, ROM 723, RAM 724, etc.
[0127] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0128] By way of example and not limitation, a computer system having the architecture of computer system 700, and in particular core 717, may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage as described above, as well as media associated with specific storage of core 717 that is non-transitory in nature, such as core internal mass storage 722 or ROM 723. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by core 717. Computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause core 717, and in particular the processors therein (including a CPU, GPU, FPGA, etc.), to perform particular processes, or particular portions of particular processes, described herein, including defining data structures stored in RAM 724 and modifying such data structures in accordance with software-defined processes. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 721), which may operate in place of or in conjunction with software to perform particular processes, or portions of particular processes, described herein. References to software may encompass logic, and vice versa. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0129] The number and arrangement of components shown in Figure 7 are provided as an example. In practice, an input human interface device may include additional components, fewer components, different components, or components arranged differently than those shown in Figure 7. Additionally or alternatively, a set of components (e.g., one or more components) of an input human interface device may perform one or more functions that are described as being performed by another set of components of the input human interface device.
[0130] In an embodiment, any one of the operations or processes of FIGS. 1-6 and 8-15 may be performed by or using any one of the elements shown in FIG.
[0131] 8 illustrates an exemplary network media distribution system 800 that serves multiple heterogeneous client endpoints. That is, the system 800 supports a variety of legacy and heterogeneous immersive media-enabled displays as client endpoints. The system 800 may include a content acquisition module 801, a content preparation module 802, and a transmission module 803.
[0132] A content acquisition module 801 captures or creates source media using, for example, the embodiments described in FIG. 6 and / or FIG. 5. A content preparation module 802 creates an ingest format that is sent to a network media distribution system using a transmission module 803. A gateway 804 may service customer premises equipment (CPE) and provide network access to various client endpoints for the network. A set-top box 805 may also function as CPE to provide access to aggregated content by a network service provider. A wireless demodulator 806 may function as a mobile network access point for mobile devices, as shown, for example, by a mobile handset display 813. In this particular embodiment of system 800, a legacy 2D television 807 is shown connected directly to one of the gateway 804, the set-top box 805, or a Wi-Fi (router) 808. A laptop 2D display 809 (i.e., a computer or laptop with a legacy 2D display) is shown as a client endpoint connected to the Wi-Fi (router) 808. A head-mounted 2D (raster-based) display 810 is also connected to the Wi-Fi (router) 808. A lenticular light field display 811 is shown connected to one of the gateways 804. The lenticular light field display 811 may include one or more GPUs 811A, a storage device 811B, and a visual representation component 811C that creates multiple views using ray-based lenticular optics technology. A holographic display 812 is shown connected to the set-top box 805. The holographic display 812 may include one or more CPUs 812A, a GPU 812B, a storage device 812C, and a visualization component 812D. The visualization component 812D may be a Fresnel pattern, wave-based holographic device / display. An augmented reality (AR) headset 814 is shown connected to the wireless demodulator 806.The AR headset 814 may include a GPU 814A, a storage device 814B, a battery 814C, and a stereoscopic visual representation component 814D. A high-density light field display 815 is shown connected to a Wi-Fi (router) 808. The high-density light field display 815 may include one or more GPUs 815A, a CPU 815B, a storage device 815C, an eye-tracking device 815D, a camera 815E, and a high-density beam-based light field panel 815F.
[0133] The number and arrangement of components shown in Figure 8 are provided as an example. In practice, system 800 may include additional, fewer, different, or differently arranged components than those shown in Figure 8. Additionally or alternatively, a set of components (e.g., one or more components) of system 800 may perform one or more functions that are described as being performed by another set of components in a device or each display.
[0134] Figure 9 illustrates an example workflow of an immersive media delivery process 900 that can serve legacy and heterogeneous immersive media-enabled displays as previously described in Figure 8. The immersive media delivery process 900 performed by the network can, for example, provide adaptation information about particular media represented in a media ingest format prior to the network's process of adapting the media for consumption by a particular immersive media client endpoint (as described with reference to Figure 10).
[0135] The immersive media delivery process 900 may be divided into two parts: immersive media production to the left of dashed line 912 and immersive media network delivery to the right of dashed line 912. The immersive media production and immersive media network delivery may be performed by a network or a client device.
[0136] Media content 901 is created or acquired from a network (or client device) or content source, respectively. Methods for creating or acquiring data may correspond to FIGS. 5 and 6 for natural and synthetic content, respectively. The created content 901 is then converted to an ingest format using a network ingest format creation process 902. The network ingest format creation process 902 may also correspond to FIGS. 5 and 6 for natural and synthetic content, respectively. The ingest format may also be updated to store information about assets potentially reused across multiple scenes, for example, from a media reuse analyzer 911 (described in more detail below with reference to FIGS. 10 and 14A). The ingest format is transmitted to the network and stored in ingest media storage 903 (i.e., a storage device). In some embodiments, the storage device may be within the immersive media content creator's network or may be accessed remotely for immersive media network distribution 920. Client and application specific information is optionally available in a remote storage device, client specific information 904. In some embodiments, the client specific information 904 may reside remotely in an alternative cloud network or be transmitted to the network.
[0137] The network orchestrator 905 then executes. The network orchestrator serves as the primary source and sink of information for performing the primary tasks of the network. The network orchestrator 905 may be implemented in a unified format with other components of the network. The network orchestrator 905 may also be a process that employs a bidirectional message protocol with client devices to facilitate all processing and delivery of media according to the characteristics of the client devices. Furthermore, the bidirectional protocol may be implemented across different delivery channels (e.g., control plane channels and / or data plane channels).
[0138] 9, network orchestrator 905 receives information about the characteristics and attributes of client devices 908. Network orchestrator 905 gathers requirements regarding applications currently running on client devices 908. This information may be obtained from client-specific information 904. In some embodiments, the information may be obtained by directly querying client devices 908. When client devices are directly queried, it is assumed that a bidirectional protocol exists and is operational so that client devices 908 may communicate directly with network orchestrator 905.
[0139] The network orchestrator 905 may also initiate and communicate with the media adaptation and fragmentation module 910 (which is depicted in FIG. 10 ). Once the ingested media has been adapted and fragmented by the media adaptation and fragmentation module 910, the media may be transferred to an inter-media storage device, such as media prepared for delivery 909. If the network is designed to include a cache for assets used multiple times in the context of a presentation, a separate inter-media storage device redundant cache for reused media assets 912 can be utilized to act as a cache for such assets. Once the media for delivery is prepared and stored in the media prepared for delivery 909 storage device, the network orchestrator 905 ensures that the client device 908 can receive the delivered media and description information 906 via a “push” request, or the client device 908 can initiate a “pull” request for the delivered media and description information 906 from the stored media prepared for delivery 909. Information may be “pushed” or “pulled” through the network interface 908B of the client device 908. The "pushed" or "pulled" distributed media and descriptive information 906 may be descriptive information that corresponds to the distributed media.
[0140] In some embodiments, network orchestrator 905 uses a two-way message interface to perform "push" requests or initiate "pull" requests with client devices 908. Client devices 908 may optionally employ GPUs 908C (or CPUs).
[0141] The distributed media format is then stored in a storage device or storage cache 908D included in the client device 908. Finally, the client device 908 visually represents the media via a visualization component 908A.
[0142] Throughout the process of streaming immersive media to client devices 908, network orchestrator 905 monitors the status of client progress via client progress and status feedback channel 907. In some embodiments, status monitoring may be performed via a two-way communication message interface.
[0143] 10 illustrates an example of a media adaptation process 1000, performed, for example, by the media adaptation and fragmentation module 910. By performing the media adaptation process 1000, the captured source media can be appropriately adapted to match the requirements of a client (e.g., client device 908).
[0144] As shown in Figure 10, media adaptation process 1000 includes multiple components that facilitate the adaptation of ingested media to an appropriate delivery format for client device 908. The components shown in Figure 10 should be considered exemplary. In practice, media adaptation process 1000 may include additional components, fewer components, different components, or components in a different arrangement than the components shown in Figure 10. Additionally or alternatively, a set of components (e.g., one or more components) of media adaptation process 1000 may perform one or more functions that are described as being performed by another set of components.
[0145] In FIG. 10 , the adaptation module 1001 receives input network status 1005 to track the current traffic load on the network. As described above, the adaptation module 1001 also receives information from the network orchestrator 905. The information may include a description of the attributes and capabilities of the client device 908, application capabilities and descriptions, the current status of the application, and the client NN model (if available) to help map the client's frustum geometry to the interpolation capabilities of the captured immersive media. Such information may be obtained via a two-way message interface. The adaptation module 1001 ensures that the adapted output, once generated, is stored in the storage device for storing the client adapted media 1006.
[0146] The media reuse analyzer 911 may be an optional process that may be run in advance or as part of a network-automated process for the distribution of media. The media reuse analyzer 911 may store the ingested media formats and assets in the storage device (1002). The ingested media formats and assets may then be transmitted from the storage device (1002) to the adaptation module 1001.
[0147] The adaptation module 1001 may be controlled by the logic controller 1001F. The adaptation module 1001 may also employ a renderer 1001B or a processor 1001C to adapt the particular ingested source media to a format suitable for the client. The processor 1001C may be a NN-based processor. The processor 1001C uses the NN model 1001A. Examples of such processors 1001C include the Deepview NN model generators described in MPI and MSI. If the media is in a 2D format but the client must have a 3D format, the processor 1001C may invoke a process that uses highly correlated images from the 2D video signal to derive a stereoscopic representation of the scene depicted in the media.
[0148] The renderer 1001B may be a software-based (or hardware-based) application or process based on a selective mix of fields related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development, which, given an input scene graph and asset container, emits (usually) visual and / or audio signals suitable for presentation on a target device or that conform to desired characteristics specified by attributes of the nodes to be rendered in the scene graph. In the case of visual-based media assets, the renderer may emit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged in another container and used in the series of rendering processes in a graphics pipeline). In the case of audio-based media assets, the renderer may emit audio signals for presentation over multi-channel loudspeakers and / or binauralized headphones, or for repackaging into another (output) container. The renderer may include, for example, the real-time rendering capabilities of a source and cross-platform game engine. The renderer may include a scripting language (i.e., an interpreted programming language) that can be executed by the renderer at runtime to process dynamic input and variable state changes made to scene graph nodes. The dynamic input and variable state changes can affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformations, collisions), and energy propagation and transport (light, sound). The evaluation of spatial and temporal object topology produces results that move the output from abstract to concrete (e.g., similar to evaluating a document object model of a web page).
[0149] Renderer 1001B may, for example, be a modified version of the OTOY Octane renderer that is modified to interact directly with adaptation module 1001. In some embodiments, renderer 1001B implements a computer graphics method (e.g., path tracing) that renders a three-dimensional scene such that the lighting in the scene is realistic. In some embodiments, renderer 1001B may use shaders (i.e., a type of computer program that was originally used for shading (generating appropriate levels of light, darkness, and color in an image) but now performs various specialized functions in various areas of computer graphics special effects, video post-processing unrelated to shading, and other functions unrelated to graphics).
[0150] The adaptation module 1001 may perform compression and decompression of the media content using a media compressor 1001D and a media decompressor 1001E, respectively, depending on the compression and decompression needs based on the format of the captured media and the format required by the client device 908. The media compressor 1001D may be a media encoder, and the media decompressor 1001E may be a media decoder. After performing compression and decompression (if necessary), the adaptation module 1001 outputs client adaptation media 1006 optimized for streaming or delivery to the client device 908. The client adaptation media 1006 may be stored in a storage device for storing adaptation media.
[0151] FIG. 11 illustrates an example of a delivery format creation process 1100. As shown in FIG. 11, the delivery format creation process 1100 includes a media adaptation module 1101 and an adapted media packaging module 1103 that packages media output from the media adaptation process 1000 and stored as client adapted media 1006. The media packaging module 1103 formats the adapted media from the client adapted media 1006 into a robust delivery format 1104. The delivery format may be, for example, the exemplary format shown in FIGS. 3A-3B or 4A-4B. The information manifest 1104A may provide the client device 908 with a list of scene data assets 1104B. The list of scene data assets 1104B may also include metadata describing how frequently each asset is used across the set of scenes that make up the representation. The list of scene data assets 1104B indicates a list of visual assets, audio assets, and haptic assets, each accompanied by corresponding metadata. In this exemplary embodiment, each asset in the list of scene data assets 1104B references metadata that includes a numerical frequency value that indicates the number of times a particular asset is used across all the scenes that make up the representation.
[0152] The media may be further packetized before steaming. Figure 12 shows an exemplary packetization process 1200. The packetization system 1200 includes a packetizer 1202. The packetizer 1202 may receive a list of scene data assets 1104B (or 1104B) (as shown in Figure 12) as input media 1201. In some embodiments, client adapted media 1006 or delivery format 1104 is input to the packetizer 1202. The packetizer 1202 separates the input media 1201 into individual packets 1203 suitable for display or streaming to client devices 908 over a network.
[0153] 13 is a sequence diagram illustrating an example of data and communication flow between components according to an embodiment. The sequence diagram of FIG. 13 is of a network adapting a particular immersive media in an ingestion format into a streamable and appropriate delivery format for a particular immersive media client endpoint. The data and communication flow may be as follows:
[0154] The client device 908 initiates a media request 1308 to the network orchestrator 905. In some embodiments, the request may be made to a network delivery interface of the client device. The media request 1308 includes information identifying the media requested by the client device 908. The media request may be identified, for example, by a uniform resource name (URN) or another standard nomenclature. The network orchestrator 905 then responds to the media request 1308 with a profile request 1309. The profile request 1309 requests the client to provide information about its currently available resources (including computation, storage, battery charge, and other information to characterize the client's current operating status). The profile request 1309 also requests the client to provide one or more NN models that can be used by the network for NN inference to extract or interpolate the correct media view to match the characteristics of the client's representation system, if such NN models are available at the client endpoint.
[0155] The client device 908 then continues with a response 1310 from the client device 908 to the network orchestrator 905, provided as a client token, an application token, and one or more NN model tokens (if such NN model tokens are available at the client endpoint). The network orchestrator 905 then provides a session ID token 1311 to the client device. The network orchestrator 905 then requests ingested media 1312 from the ingest media server 1303. The ingest media server 1303 may include, for example, the assets of the ingested media storage 903 or the ingested media format and storage device 1002. The request for ingested media 1312 may also include the URN or other standard name of the media identified in the request 1308. The ingest media server 1303 responds to the request for ingested media 1312 with a response 1313 that includes the ingested media tokens. The network orchestrator 905 then provides the media tokens from the response 1313 to the client device 908 in a call 1314. The network orchestrator 905 then initiates the adaptation process for the requested media by providing the ingest media token, client token, application token, and NN model token to the adaptation and fragmentation module 910 in a request 1315. The adaptation and fragmentation module 910 requests access to the ingested media by providing the ingest media token to the ingest media server 1303 in a request 1316 requesting access to the ingested media asset.
[0156] The ingestion media server 1303 responds to the request 1316 with the ingestion media access token in response 1317 to the adaptation and fragmentation module 910. The adaptation and fragmentation module 910 then requests that the media adaptation process 1000 adapt the ingestion media located in the ingestion media access token for the client, application, and NN inference model corresponding to the session ID token created and sent in response 1313. A request 1318 is made from the adaptation and fragmentation module 910 to the media adaptation process 1000. The request 1318 includes the necessary token and session ID. The media adaptation process 1000 provides the adapted media access token and session ID to the network orchestrator 905 in an update response 1319. The network orchestrator 905 then provides the adapted media access token and session ID to the media packaging module 11043 in an interface call 1320. The media packaging module 11043 provides response 1321 to the network orchestrator 905, along with the packaged media access token and the session ID in response 1321. The media packaging module 11043 then provides the packaged asset for the session ID, the URN, and the packaged media access token to the packaged media server 1307 for storage in response 1322. The client device 908 then performs request 1323 to the packaged media server 1307 to begin streaming the media asset corresponding to the packaged media access token received in response 1321. Finally, the client device 908 performs another request and provides a status update in message 1324 to the network orchestrator 905.
[0157] Figure 14A shows the workflow for the media reuse analyzer 911 shown in Figure 9. The media reuse analyzer 911 analyzes metadata relating to the identity of objects in scenes contained in media data.
[0158] At 1405, media data is obtained, for example, from a content provider or content source. At 1410, initialization is performed. Specifically, an iterator "i" is initialized to 0. The iterator may be, for example, a counter. A set of lists of unique assets 1465 (shown in FIG. 14B) for each scene is also initialized, identifying unique assets that occur across all scenes that contain the representation (shown in FIG. 3 and / or FIG. 4).
[0159] At 1415, it is determined whether the value of iterator "i" is less than the total number N of scenes containing the representation. If the value of iterator "i" is equal to (or greater than) the number N of scenes containing the representation (1415 "No"), the process proceeds to 1420, where the reuse analysis ends (i.e., the process terminates). If the value of iterator "i" is less than the number N of scenes containing the representation (1415 "Yes"), the process proceeds to 1425. At 1425, the value of iterator "j" is set to 0.
[0160] Next, at 1430, it is determined whether the value of iterator "j" is less than the total number X of media assets (also called media objects) in the current scene. If the value of iterator "j" is equal to (or greater than) the total number X of media assets in scene s ("No" at 1435), the process proceeds to 1410, where iterator "i" is incremented by 1, and then returns to 1415. If the value of iterator "j" is less than the total number X of media assets in scene s ("Yes" at 1430), the process proceeds to 1440.
[0161] At 1440, the characteristics of the media asset are compared to previously analyzed assets from scenes prior to the current scene (ie, scenes s) to determine if the current media asset has been used previously.
[0162] If the current media asset is identified as a unique asset ("No" at 1440), i.e., the current media asset has not been previously analyzed in a scene associated with a smaller value of the iterator "i," the process proceeds to 1445. At 1445, a unique asset entry is created in the set 1465 of lists of unique assets corresponding to the current scene (i.e., scene s). The unique asset entry is also assigned a unique identifier, and the number of times (e.g., frequency) the asset has been used across scenes 0 to N-1 is set to 1. The process then proceeds to 1455.
[0163] If the current media asset is identified as an asset that was used in one or more scenes prior to scene s (“Yes” at 1440), the process proceeds to 1450. At 1450, the number of times the current media asset was used across scenes 0 through N−1 is incremented by 1 in the set 1465 of lists of unique assets that correspond to the current scene (i.e., scene s). Then, the process proceeds to 1455.
[0164] At 1455, the value of the iterator "j" is incremented by 1. Then, the process returns to 1430.
[0165] In some embodiments, the media reuse analyzer 911 may further signal to a client, e.g., a client device 108, that the client should use a copy of the asset for each instance the asset is used in a set of scenes (after the asset is delivered to the client for the first time).
[0166] It should be noted that the sequence diagrams and workflow steps described with reference to Figures 13-14 are not intended to limit the configuration of data and communication flows in embodiments. For example, one or more of the steps may be performed simultaneously, data may be stored, and / or flow in directions not explicitly shown in the flows of Figures 13-14A.
[0167] 14B is an example of a set 1465 of lists of unique assets initialized in 1410 (and potentially updated in 1445-1450) for all scenes upon completion of a representation, according to an embodiment. The list of unique assets in set 1465 of lists of unique assets may be pre-identified or pre-defined by the network or client device. Set 1465 of lists of unique assets shows a sample list of entries of information describing assets that are unique to the entire representation, including an indicator of the type of media (e.g., mesh, audio, or volume) that comprises the asset, a unique identifier for the asset, and the number of times the asset is used across the set of scenes that comprise the entire representation. As an example, for Scene N-1, no assets are included in its list because all assets required for Scene N-1 have been identified as assets also used in Scene 1 and Scene 2.
[0168] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]
[0169] 100 Media distribution process, media flow process 101 Format 102 Distribution Media Creation 103 Format 104 Network devices, servers 105 Network Connections 106 Rendering Module 107 Expression Module 108 client devices 200 Media Conversion Decision-Making Process 222 Media 250 Media Conversion Process 300 Timed Media Display 300A Timed Scene Manifest 301 Scene Information 302 Components 303 Assets 303A Timed Scene Manifesto 304 Base Layer 305 Attribute enhancement layer 307 List 308 Proxy Assets 400 Complexity Display 401 Scene Information 402 Components 403 Assets 404 Base Layer 405 Attribute enhancement layer 407 Scene Information 408 List 500 Natural Media Synthesis Process 501 Camera 502 Camera 503 Camera 504 Synthesizer 505 NN Training 506 training images 507 Import Format 508 Capture NN model 509 Natural Image Content 510 dashed line 600 Synthetic Media Capture Creation Process 601 Camera 602 point cloud 603 Computer 604 CGI assets 605 Actors 605A Sensor 606 Motion Data 607 Synthesizer 608 Synthetic Media Capture Format 700 Computer Systems 701 Keyboard 702 Trackpad 703 Mouse 704 Joystick 705 Microphone 706 Camera 707 Scanner 708 Speaker 709 screens 710 Medium 712 thumb drive 713 Solid State Drive 714 Communication Network 715 Bidirectional Interface 716 Peripheral Bus 717 cores 720 FPGA 721 Hardware Accelerator 722 Internal Mass Storage 723 ROM 725 graphics adapter 748 System Bus 800 Network Media Distribution System 801 Content Acquisition Module 802 Content Preparation Module 803 Transmit Module 804 Gateway 805 Set-Top Box 806 Radio Demodulator 807 Legacy 2D TV 809 Laptop 2D Display 810 Display 811 Lenticular Light Field Display 811B Storage Device 811C Visual Representation Components 812 Holographic Display 812C Storage Device 812D Visualization Components 813 Mobile Handset Display 814 Augmented Reality (AR) Headset 814A GPU 814B Storage Device 814C Battery 814D Stereoscopic Visual Representation Components 815 High-Density Light Field Display 815C Storage Device 815D Eye Tracking Device 815E Camera 815F Light Field Panel 900 Immersive Media Delivery Process 901 Media Content 902 Network Import Format Creation Process 903 Ingest Media Storage 904 Information 905 Network Orchestrator 906 Descriptive Information 907 Status Feedback Channel 908 client device 908A Visualization Components 908B Network Interface 908D Storage Cache 909 Media 910 Fragmentation Module 911 Media Reuse Analyzer 912 dashed line 920 Immersive Media Network Distribution 1000 Media Adaptation Process 1001 Adaptation Module 1001A NN model 1001B Renderer 1001C processor 1001D Media Compressor 1001E Media Decompressor 1001F Logic Controller 1002 Storage Device 1005 Input Network Status 1006 Client Adapted Media 1100 Delivery Format Creation Process 1101 Media Adaptation Module 1103 Media Packaging Module 1104 Delivery Format 1104A Information Manifesto 1104B Scene Data Asset 1200 Packetization System, Packetization Process 1201 Input Media 1202 Packetizer 1203 packets 1303 Ingest Media Server 1307 Packaged Media Server 1308 Media Request 1309 Profile Request 1310 Response 1311 Session ID Token 1312 Capture Media 1313 Response 1314 Call 1315 request 1316 request 1317 Responses 1318 request 1319 Update Response 1320 Interface Call 1321 Responses 1322 Responses 1323 request 1324 Messages 1465 sets 2500 processes 2600 processes 2700 processes 3031 Scene Information 3032 Components 3033 Assets 3034 Base Layer 3035 Attribute enhancement layer 3100 Timed Media Display 4041 Scene Information 4042 Components 4043 assets, audio assets, tactile assets, 4044 base layer 4045 Attribute enhancement layer 4046 Attribute Enhancement Layer, List 4047 Scene Information 4100 Ordered Complexity Display
Claims
1. 1. A method for streaming scene-based media assets during a media streaming session performed by a computing device, the method comprising: providing a bidirectional interface for communicating information about scene-based media assets between a network server and the computing devices functioning as client devices; receiving a request from the network server via the interface, the request identifying one or more scene-based media assets to be streamed; accessing a list of previously streamed scene-based media assets stored locally on the computing device in response to the request to determine whether the requested one or more scene-based media assets are streamed from the network server; providing client device attributes and corresponding information regarding availability of client device resources, when requested, to the network server via the interface, wherein the client device attributes and the information are used by the computing device to render the scene-based media asset; transmitting the determination to the network server via the interface, wherein the network server is configured to generate a proxy for the requested scene-based media asset and deliver the proxy to the computing device according to a determination that there is no need to stream the requested scene-based media asset from the network server; an indicator for identifying that the proxy should access a copy of a previously streamed scene-based media asset from a local cache of the computing device.
2. 2. The method of claim 1 , wherein the network server is configured to generate the proxy for the requested scene-based media asset and insert the proxy into the media streaming session in accordance with a determination that the requested scene-based media asset does not need to be streamed from the network server.
3. The method of claim 1 , wherein the request is generated at the network server prior to determining that the requested one or more scene-based media assets have been generated by the network server.
4. The method comprises:
4. The method of claim 3, further comprising: transmitting associated information regarding the client device attributes and availability of client device resources to the network server via the interface, wherein the network server is configured to deliver the one or more scene-based media assets or the proxies to the computing device according to the associated information regarding the client device attributes and availability of client device resources.
5. The method comprises: receiving a user request from a user of the client device, the user request identifying the one or more scene-based media assets to be streamed; 2. The method of claim 1, further comprising the step of: forwarding the user request to the network server via the interface.
6. receiving, from the network server via the interface, the one or more scene-based media assets in the media streaming session generated by the network server; The method of claim 5 , further comprising: storing the one or more scene-based media assets in a media store accessible to the client device.
7. A computing device configured to perform the method of any one of claims 1 to 6.
8. A computer program comprising instructions that, when executed by at least one processor of a computing device, cause said computing device to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Remote display protocol, video display system, and terminal
JP2005045357A
Media distribution device
JP2005295467A
Moving image processor
JP2010114815A
Configurable access-based cache policy control
US20220086253A1