Method, apparatus and program for immersive media presentations containing multiple scenes

By determining media asset availability in client caches, the method optimizes media format conversion, addressing the challenge of diverse client devices and reducing latency and resource waste in immersive media delivery.

JP7729690B2Active Publication Date: 2025-08-26TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023560733
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2022-10-25
Publication Date
2025-08-26
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing networks struggle to efficiently deliver immersive media due to the great diversity of client devices, requiring significant computational resources for media conversion and lacking information about client capabilities, leading to latency and inefficient resource usage.

Method used

A method and apparatus that determine whether media assets are available in a client's local cache before converting or streaming, using a decision-making process to optimize media format conversion based on client resources and network traffic.

Benefits of technology

This approach reduces latency and resource consumption by minimizing redundant media conversions, enabling efficient delivery of immersive media to heterogeneous client devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007729690000001
    Figure 0007729690000001
  • Figure 0007729690000002
    Figure 0007729690000002
  • Figure 0007729690000003
    Figure 0007729690000003
Patent Text Reader

Abstract

A method and apparatus for determining that a media asset appears in at least two or more scenes within a scene associated with an immersive media presentation, sending a request to a client inquiring whether the client can access the media asset that appears in at least two or more scenes in a local cache, receiving a response indicating whether the client can access it, signaling the client to use the media asset in a subsequent scene in response to a response indicating the client can access the media asset that appears in at least two or more scenes in the local cache, and delivering the media asset to the client in response to a response indicating the client cannot access the media asset that appears in at least two or more scenes in the local cache.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 276,535, filed November 5, 2021, and U.S. Application No. 17 / 970,109, filed October 20, 2022, the contents of which are incorporated herein by reference in their entireties.

[0002] This disclosure generally describes embodiments relating to architectures, structures, and components for systems and networks that deliver media, including video, audio, geometric (3D) objects, haptics, associated metadata, or other content for client presentation devices. Particular embodiments are directed to systems, structures, and architectures for delivering media content to heterogeneous immersive and interactive client presentation devices. [Background technology]

[0003] "Immersive media" generally refers to media that stimulates some or all of the human sensory systems (sight, hearing, somatosensation, smell, and sometimes taste) to create or enhance the perception that the user is physically present in the media experience, i.e., beyond that delivered over existing (e.g., "legacy") commercial networks for timed two-dimensional (2D) video and corresponding audio; such timed media is also known as "legacy media."

[0004] Yet another definition of "immersive media" is media that attempts to create or mimic the physical world through digital simulation of kinematics and the laws of physics, thereby stimulating some or all of the human sensory systems to create the perception by the user of being physically present in a scene representing a real or virtual world.

[0005] An immersive media-enabled presentation device may refer to a device with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices are heterogeneous in the amount and format of media they can support relative to the media provided by the network. Similarly, media is heterogeneous in the amount and type of network resources required to deliver such media at scale. "At scale" may refer to the delivery of media by a service provider, e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions, that achieves delivery equivalent to the delivery of legacy video and audio media over a network.

[0006] In contrast, legacy presentation devices such as laptop displays, televisions, and mobile phone displays are homogenous in their capabilities, as all of these devices currently consist of rectangular display screens that consume 2D rectangular video or still images as their primary visual media format. Some of the visual media formats commonly used by legacy presentation devices may include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266.

[0007] The delivery of any media over a network may employ media delivery systems and architectures that reformat media from an input or network "ingest" media format into a delivered media format that is not only suitable for ingestion by a target client device and its applications, but also lends itself to being "streamed" over the network. Thus, there may be two processes performed on media ingested by the network: 1) converting the media from format A to format B that is suitable for ingestion by the target client, i.e., based on the client's ability to ingest a particular media format, and 2) preparing the media to be streamed.

[0008] "Streaming" media broadly refers to fragmenting and / or packetizing media so that it can be delivered over a network in successive, smaller-sized "chunks" that are logically organized and sequenced according to either or both the temporal and spatial structure of the media. The "conversion," sometimes referred to as "transcoding" of media from format A to format B, may be a process typically performed by a network or service provider prior to delivery of the media to a client. Such transcoding may consist of converting media from format A to format B based on prior knowledge that format B is the preferred or only format that can be ingested by the intended client, or that format B is more suitable for delivery over constrained resources, such as a commercial network. In many, but not all, cases, both the steps of converting the media and preparing the media for streaming are necessary before the client can receive and process the media from the network.

[0009] The above one-step or two-step process of operating on the ingested media by the network, i.e., before delivering the media to the client, results in a media format referred to as a "delivery media format" or simply a "delivery format." Generally, due to technical constraints, these steps, when performed for a given media data object, should be performed only once, even if the network has access to information indicating that the client requires the converted and / or streamed media object multiple times, otherwise triggering the conversion and streaming of such media multiple times. That is, the processing and transmission of data for media conversion and streaming is generally considered a source of latency that requires the consumption of potentially significant amounts of network and / or computational resources. Therefore, network designs that do not have access to information indicating that a client may already have a particular media data object cached or locally stored to the client perform less optimally than networks that have access to such information.

[0010] For legacy presentation devices, the delivery format may be equivalent or sufficiently equivalent to the “presentation format” ultimately used by the client presentation device to create the presentation. That is, the presentation media format is a media format whose properties (resolution, frame rate, bit depth, color gamut, etc.) are closely tailored to the capabilities of the client presentation device. Some examples of delivery format versus presentation format include a high-definition (HD) video signal (1920 pixel columns by 1080 pixel rows) delivered over a network to an ultra-high-definition (UHD) client device with a resolution (3840 pixel columns by 2160 pixel rows). In this scenario, the UHD client applies a process called “super-resolution” to the HD delivery format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format presented by the client device is the “presentation format,” which in this example is a UHD signal, but the HD signal includes the delivery format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format since both signals are linear video formats, and the process of converting the HD format to the UHD format is relatively simple and easy to perform on most legacy client devices.

[0011] Alternatively, the preferred presentation format for the target client device may be significantly different from the ingest format received by the network. Nevertheless, the client may have access to sufficient computational, storage, and bandwidth resources to convert the media from the ingest format to the required presentation format suitable for presentation by the client. In this scenario, the network may bypass steps of reformatting the ingested media, such as "transcoding" the media from format A to format B, simply because the client has access to sufficient resources to perform all media conversion without the network having to do so beforehand. However, the network may still perform steps of fragmenting and packaging the ingest media so that the media can be streamed to the client.

[0012] Yet another alternative scenario is when the ingested media received by the network is significantly different from the client's preferred presentation format, and the client does not have access to sufficient computational, storage, and / or bandwidth resources to convert the media to the preferred presentation format. In such a scenario, the network can assist the client by performing some or all of the conversion on the client's behalf from the ingest format to a format that is equivalent or nearly equivalent to the client's preferred presentation format. In some architectural designs, such assistance provided by the network on behalf of the client is commonly referred to as "split rendering."

[0013] Considering each of the above scenarios in which the conversion of media from format A to another format may be performed entirely by the network, entirely by the client, or jointly between the network and the client, for example, for split rendering, it becomes clear that a dictionary of attributes describing the media format may be needed so that both the client and the network have complete information to characterize the work that must be done. Furthermore, a dictionary providing attributes of the client's capabilities, for example, with respect to available computational resources, available storage resources, and access to bandwidth, may be needed as well. Furthermore, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of an ingest format so that the network and client, jointly or independently, can determine whether or when the network can use a split rendering step to deliver media to the client. Finally, if the conversion and / or streaming of certain media objects that the client requires to complete the media presentation can be avoided, the network can skip the conversion and streaming steps, assuming that the client has access to or availability of the media objects that the client requires to complete the media presentation. Such networks, which have enough information to avoid repetitive conversion and / or streaming steps for assets that are used multiple times in a particular presentation, can perform more optimally than networks that are not designed to do so. Summary of the Invention [Means for solving the problem]

[0014] 1. A method and apparatus comprising: a memory configured to store computer program code; and one or more hardware processors configured to access the computer program code and to operate as instructed by the computer program code, the program code comprising: a determining code configured to cause the at least one hardware processor to determine that a media asset appears in at least two or more subsequent scenes of a plurality of scenes associated with an immersive media presentation; and a sending code configured to cause the at least one hardware processor to send a request to a client inquiring whether the client has access to the media assets appearing in the at least two or more scenes in a local cache, the sending code being configured to cause the at least one hardware processor to and receiving code configured to cause at least one hardware processor to receive a response from the client indicating whether the client can access the media asset that appears in at least two or more scenes in the local cache; signaling code configured to cause the at least one hardware processor to send a signal in response to the response indicating the client can access the media asset that appears in at least two or more scenes in the local cache, whereby the client uses the media asset in a subsequent scene without further waiting for the media asset to be delivered to the client; and distribution code configured to cause the at least one hardware processor to deliver the media asset to the client in response to the response indicating the client cannot access the media asset that appears in the at least two or more scenes in the local cache.

[0015] According to an exemplary embodiment, the computer program code further includes initialization code configured to cause the at least one hardware processor to perform initialization of a set of lists, each of the lists corresponding to a respective scene of the immersive media presentation including at least one scene and a subsequent scene, and the step of initializing the set of lists includes a step of incrementally respectively assigning a respective unique identifier to each of the assets appearing in the scene, including the asset.

[0016] According to an exemplary embodiment, initializing the set of lists further includes incrementally determining the number of times each asset appears in each scene.

[0017] According to an exemplary embodiment, the computer program code further includes receiving code configured to cause the at least one hardware processor to receive a request from a client for the immersive media presentation, and requesting code configured to cause the at least one hardware processor to request that the client provide an indication of client resources of the client in response to the request, wherein processing of the asset is performed in accordance with the indication of the client resources.

[0018] According to an exemplary embodiment, requesting that the client provide an indication of the client resources includes requesting that the client provide one or more neural network models, and processing the asset includes neural network inference based on the one or more neural network models requested by the client.

[0019] According to an exemplary embodiment, the processing of the asset is further based on determining a current traffic load on a network interfacing the at least one hardware processor and the client.

[0020] According to an exemplary embodiment, the computer program code further includes monitoring code configured to cause the at least one hardware processor to monitor the client's progress in outputting the immersive media presentation, and the step of sending the query is timed based on the progress.

[0021] According to an exemplary embodiment, the immersive media presentation includes instructions to the client to stimulate at least one of the senses of sight, hearing, taste, touch, and smell.

[0022] According to an exemplary embodiment, the query to the client further inquires whether the client has access to assets whose access is local to the client.

[0023] According to an exemplary embodiment, an immersive media presentation includes either a timed presentation or an untimed presentation.

[0024] The technology described herein improves computer technology by facilitating various aspects, such as the decision-making process used by a network and / or client to determine whether the network should convert some or all of the ingested media from format A to format B to further facilitate the client's ability to generate media presentations, potentially in a third format C. To aid in such a decision-making process, a method is postulated for determining, within the context of a presentation, which assets are used multiple times within the presentation or are readily available for the network to use by design. Relying on information from such an analysis, the network may be designed to require the client to maintain a copy of each asset used multiple times in its local cache. However, in this scenario, the network may not have control over the management of the client's local cache, and as a result, the client may encounter a situation where it must remove resources (even reusable resources) from its local cache. To facilitate the design of optimizing the network to minimize the need to perform format A to format B conversions for assets that are used multiple times, or to facilitate the need for the network to stream assets that are used multiple times to the client, the network may first query the client to obtain feedback that ensures the asset in question is still available in the client's local cache. If the client's response indicates that it no longer has a copy of the asset in question, the network can convert the ingest asset from format A to format B and / or stream the original copy of the asset to the client.Such a network ensures that the client can access the asset either through the client's own copy of the asset (previously provided to the client by the network) stored in the client's local cache, or by the network repeating the steps of converting the asset again and / or streaming it to the client. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 2 is a schematic diagram illustrating the flow of media through a network for delivery to a client, according to an exemplary embodiment. [Figure 1-02] This is the same schematic diagram as Figure 1, but with the addition of logic to determine whether a proxy of the original media should be streamed (either converted to another format or in its original format) instead of the original media itself, according to an exemplary embodiment. [Figure 2] 1 is a schematic diagram of media flow through a network in which a decision-making process is used to determine whether the network should convert the media before delivering it to a client, according to an exemplary embodiment. [Figure 2-03] This is the same schematic diagram as Figure 2, but with the addition of logic to determine whether a proxy of the original media should be streamed (either converted to another format or in its original format) instead of the original media itself, according to an exemplary embodiment. [Figure 2-033] Same schematic diagram as Figure 2-03, but with the addition of logic to first query the client to ensure that the client still has access to a copy of the reusable asset (either converted to another format or in its original format), according to an exemplary embodiment. [Figure 3]FIG. 1 is a schematic diagram of one embodiment of a data model for representing and streaming timed immersive media, according to an exemplary embodiment, where such timed immersive media includes a list of assets that are reused across a set of N scenes. [Figure 4] FIG. 1 is a schematic diagram of an embodiment of a data model for representing and streaming non-timed immersive media, according to an exemplary embodiment, where such non-timed immersive media includes a list of assets that are reused across a set of five scenes. [Figure 5] FIG. 1 is a schematic diagram of a process for capturing and converting a natural scene into a representation that can be used as an ingest format for a network serving heterogeneous client endpoints, according to an example embodiment. [Figure 6] FIG. 1 is a schematic diagram of a process for using 3D modeling tools and formats to create a representation of a synthetic scene that can be used as an ingest format for a network serving heterogeneous client endpoints, according to an exemplary embodiment. [Figure 7] FIG. 1 is a system diagram of a computer system in accordance with an illustrative embodiment. [Figure 8] FIG. 1 is a schematic diagram of a network serving multiple heterogeneous client endpoints, according to an example embodiment. [Figure 9] FIG. 1 is a schematic diagram of a network that provides adaptation information about particular media represented in a media ingest format, for example, prior to the network's process of adapting the media for consumption by a particular immersive media client endpoint, according to an exemplary embodiment. [Figure 10] FIG. 1 is a system diagram of a media adaptation process consisting of a media rendering converter that converts source media from its ingest format to a specific format suitable for a particular client endpoint, according to an example embodiment. [Figure 11]FIG. 1 is a schematic diagram of a network that formats adapted source media into a data model suitable for presentation and streaming, according to an exemplary embodiment. [Figure 12] FIG. 13 is a system diagram of a media streaming process that fragments the data model of FIG. 12 into the payload of network protocol packets, according to an exemplary embodiment. [Figure 13] FIG. 1 is a sequence diagram of a network adapting a particular immersive media in an ingest format into an appropriate streamable delivery format for a particular immersive media client endpoint, according to an exemplary embodiment. [Figure 14] FIG. 2 illustrates a logic flow diagram for an immersive media asset reuse analyzer according to an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] Definition: Scene graph: A common data structure typically used by vector-based graphics editing applications and modern computer games that constitutes a logical and often (but not necessarily) spatial representation of a graphical scene; it is a collection of nodes and vertices in a graph structure. Scene: In the context of computer graphics, a scene is a collection of objects (e.g., 3D assets), object attributes, and other metadata, including visual, acoustic, and physics-based characteristics, that describe a particular setting, bounded either by space or time, with respect to the interactions of objects within that setting. Node: A basic element of a scene graph consisting of information related to the logical, spatial, or temporal representation of visual, auditory, tactile, olfactory, gustatory, or related processing information. Each node shall have at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) attached to it. Base Layer: A nominal representation of an asset, typically formulated to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network. Enhancement Layer: A set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities not supported in the base layer. Attribute: Metadata associated with a node that is used to describe a particular characteristic or feature of that node, either in canonical form or in more complex form (e.g., with respect to another node). Container: A serialized format for storing and exchanging information for representing an entire natural scene, an entire synthetic scene, or a combination of synthetic and natural scenes, including a scene graph and all the media resources required to render the scene. Serialization: The process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link), and later reconstructed (possibly in a different computing environment). When the resulting series of bits is reread according to the serialization format, it can be used to create a semantically identical clone of the original object. Renderer: A (typically software-based) application or process based on a selective combination of academic disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development that, given an input scene graph and asset container, emits exemplary visual and / or audio signals suitable for presentation on a target device or adapted to desired properties specified by attributes of the nodes to be rendered in the scene graph. In the case of visual-based media assets, a renderer can emit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline); in the case of audio-based media assets, a renderer can emit audio signals for presentation over multi-channel loudspeakers and / or binaural headphones or for repackaging into another (output) container. Common examples of renderers include the real-time rendering capabilities of game engines such as Unity Engine and Unreal Engine. Evaluate: Generate results that move the output from abstract to concrete (e.g., similar to evaluating a document object model for a web page). Scripting Language: An interpreted programming language that can be executed by the renderer at runtime to process dynamic inputs and variable state changes applied to scene graph nodes, which changes affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformations, collisions) and energy propagation and transport (light, sound). Shader: a type of computer program originally used for shading (producing appropriate levels of light and color in an image), but now performing a variety of specialized functions in various areas of computer graphics special effects, or video post-processing unrelated to shading, or even functions completely unrelated to graphics. Path Tracing: A computer graphics method for rendering three-dimensional scenes so that the lighting in the scene is realistic. Timed Media: Media that is ordered by time, e.g., has a start time and an end time according to a particular clock. Non-timed media: Media that is organized by spatial, logical, or temporal relationships, such as an interactive experience that is realized according to actions taken by a user. Neural network model: a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numbers) used in well-defined mathematical operations that are applied to a visual signal to arrive at an improved visual output, which may include the interpolation of new views of the visual signal that were not explicitly provided by the original signal.

[0027] Over the past decade, several immersive media-enabled devices have been introduced to the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and gaming consoles. Similarly, holographic displays and other forms of volumetric displays are poised to appear on the consumer market within the next three to five years. Despite the immediate or imminent availability of these devices, a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks has not materialized for several reasons.

[0028] Any descriptions herein of ideal technologies, processes, and the like should not be construed as admissions of prior art, but as a disclosure of what has been invented by the inventors and is disclosed by this application. Unless otherwise specified, any statements herein of technical deficiencies and needs should also be construed as having been realized by the inventors and disclosed by this application.

[0029] One of the obstacles to achieving a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks is the great diversity of client devices that serve as endpoints in such delivery networks for immersive displays. Some of these support specific immersive media formats, while others do not. Some of these are capable of creating immersive experiences from legacy raster-based formats, while others are not. Unlike networks designed solely for the delivery of legacy media, networks that must support a variety of display clients require a significant amount of information detailing each client's capabilities and the format of the media being delivered before such networks can use an adaptation process to convert the media into a format appropriate for each target display and corresponding application. At a minimum, such networks need access to information describing the characteristics of each target display and the complexity of the ingested media in order for the network to ascertain how to meaningfully adapt the input media source to a format appropriate for the target display and application.

[0030] Similarly, an ideal network supporting heterogeneous clients should take advantage of the fact that some of the assets adapted from an input media format to a particular target format can be reused across a set of similar display targets. That is, once converted to a format appropriate for the target displays, some assets may be reused across several such displays with similar adaptation requirements. Such an ideal network would therefore employ a caching mechanism to store the adapted assets in a relatively immutable area, i.e., similar to the use of content delivery networks (CDNs) used in legacy networks.

[0031] Furthermore, immersive media may be organized into "scenes" that are described by a scene graph, also known as a scene description. The scope of a scene graph is to describe the visual, audio, and other forms of immersive assets that comprise a particular setting that is part of a presentation, for example, the actors and events that take place in a particular location within a building that is part of a presentation such as a movie. A list of all the scenes that comprise a single presentation may be formulated in a scene manifest.

[0032] An additional benefit of such an approach is that, for content that is prepared before such content must be delivered, a "bill of materials" can be created that identifies all of the assets used throughout the presentation and the frequency with which each asset is used across various scenes within the presentation. An ideal network should have knowledge of the existence of cached resources that can be used to satisfy the asset requirements of a particular presentation. Similarly, a client presenting a series of scenes may want to have knowledge of the frequency with which any given asset is used across multiple scenes. For example, if a media asset (also known as an object) is referenced multiple times across multiple scenes that are or will be processed by the client, the client should avoid discarding the asset from its caching resource until the last scene requiring that particular asset has been presented by the client.

[0033] The disclosed subject matter addresses the need for a mechanism or process to analyze an immersive media scene to obtain sufficient information that, when employed by a network or a client, can be used to support a decision-making process that provides instructions regarding whether conversion of media objects (or media assets) from format A to format B should be performed entirely by the network, entirely by the client, or via a combination of both (along with instructions on which assets should be converted by the client or the network). Such an "immersive media data complexity analyzer" may be employed by either the client or the network in an automated situation, or by a human in a manual situation.

[0034] Note that the remainder of the disclosed subject matter assumes, without loss of generality, that the process of adapting an input immersive media source to a particular endpoint client device is the same as or similar to the process of adapting the same input immersive media source to a particular application running on a particular client endpoint device, i.e., the problem of adapting an input media source to the characteristics of an endpoint device is of the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.

[0035] Furthermore, it should be noted that the terms media object and media asset are sometimes used interchangeably, and both refer to a particular instance of media data in a particular format.

[0036] FIG. 1 is a schematic diagram of the flow of media over a network for delivery to a client. In FIG. 1, processing of ingested media format A is performed by a "cloud" or edge process 104. Note that similar processing may also be performed a priori by a manual process or by the client. Ingested media 101 is obtained from a content provider (not shown). Process 102 performs any necessary transformations or adjustments of the ingested media to create a potential alternative representation of the media as delivery format B. Media formats A and B may or may not be representations that follow the same syntax of a particular media format specification, but format B is likely to be adjusted to a scheme that facilitates delivery of the media over a network protocol such as TCP or UDP. Such "streamable" media is depicted as stream 105, which represents media streamed to client 108. Client 108 has access to some rendering functionality, depicted as process 106. Such rendering process 106 may be rudimentary or similarly advanced, depending on the type of target client 108. The rendering process 106 creates presentation media that may or may not be rendered according to a third format specification, eg, Format C.

[0037] FIG. 1-02 is the same as FIG. 1, but adds logic to support the decision-making process for determining whether a particular media object has already been streamed to the client 108. Step 102A begins a series of steps to support the decision-making process. Conditional logic 102B accesses a list of unique assets for the presentation (not shown in example 100-2 of FIG. 1-2) to determine whether the media object has previously been streamed to the client. If the media object has been previously streamed, an indicator 102C (later referred to as a "proxy") is created to identify that the client has already received this particular media object and should use its local copy of the media object. If the media object has not been previously streamed, step 102D directs processing to step 103, which creates a delivery format for the media object.

[0038] FIG. 2 is a schematic diagram of media flow through a network in which a media conversion decision-making process 200 is used to determine whether the network should convert the media before delivering it to a client. In FIG. 2, ingested media 201, represented by format A, is provided to the network by a content provider (not shown). Process 202 obtains attributes describing the processing capabilities of a target client (not shown). Decision-making process 203 is used to determine whether the network or client should perform any format conversion on any of the media assets included in ingested media 201 before the media is streamed to the client, such as converting a particular media object from format A to format B. If any of the media assets should be converted by the network, the network converts the media object from format A to format B using process 204. Converted media 205 is the output from process 204. The converted media is merged with preparation process 206 to prepare the media to be streamed to the client (not shown). Process 207 streams the media to the client.

[0039] Figure 2-03 is a schematic diagram of the media conversion decision-making process by the asset reuse logic 2030. As media flows through a network, two decision-making processes are used to determine whether the network should convert the media before delivering it to a client. In example 20300 of Figure 2-03, ingested media 2031, represented in format A, is provided to the network by a content provider (not shown). Process 2032 obtains attributes describing the processing capabilities of the target client (not shown). Decision-making process 2033 is used to determine whether the network has previously streamed a particular media object to the client. If the media object has been previously streamed to the client, step 2034 is used to indicate that the client should use a local copy of the previously streamed object, substituting a media proxy. If the media has not been previously streamed, decision-making process 2035 is used to determine whether the network or the client should perform any format conversion on any of the media assets contained within the ingested media 2031, such as converting a particular media object from format A to format B, before the media is streamed to the client. If any of the media assets should be converted by the network, the network uses process 2038 to convert the media objects from format A to format B. Converted media 2039 is output from process 2038. The converted media is merged into preparation process 2036 to prepare the media to be streamed to a client (not shown). Process 2037 streams the media to the client.

[0040] Figure 2-33 shows an example of a client-queried media transformation decision-making process 20330 in the asset reuse logic 20330. As media flows through the network, three decision-making processes are used to determine whether the network should transform the media before delivering it to the client. In Figure 2-33, ingest media 20331, represented by format A, is provided to the network by a content provider (not shown). Process 20332 obtains attributes describing the processing capabilities of the target client (not shown). Decision-making process 20333 is used to determine whether the network has previously streamed a particular media object to the client. If the media object has previously been streamed to the client, decision-making process 20334 is used to query the client to determine whether the client can still access the previously streamed asset. If the client can still access the asset, step 203310 is used to indicate that the client should use a local copy of the previously streamed object, substituting a media proxy. If the media has not been previously streamed, or if the client no longer has a copy of the previously streamed asset, decision making process 20335 is used to determine whether the network or client should perform any format conversion on any of the media assets contained within ingested media 20331, such as converting a particular media object from format A to format B, before the media is streamed to the client. If any of the media assets should be converted by the network, the network uses process 20338 to convert the media object from format A to format B. Converted media 20339 is output from process 20338. The converted media is merged into preparation process 20336 to prepare the media to be streamed to the client (not shown).Process 20337 streams the media to the client.

[0041] Figure 3 is an exemplary representation of a streamable format for timed heterogeneous immersive media, timed media representation 300. Figure 4 is an exemplary representation of a streamable format for non-timed heterogeneous immersive media, non-timed media representation 400. Both figures relate to scenes: Figure 3 relates to a timed media scene 301, and Figure 4 relates to a non-timed media scene 401. In either case, the scene may be embodied by various scene representations or scene descriptions.

[0042] For example, in some immersive media designs, a scene may be embodied by a scene graph, or as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI and MSI technologies are examples of technologies that support the creation of display-independent scene representations for natural content, i.e., real-world images captured simultaneously from one or more cameras. On the other hand, scene graph technologies are sometimes used to represent both natural and computer-generated imagery in the form of synthetic representations, but such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. That is, scene graph representations of naturally captured content are time- and computationally intensive to create, requiring complex analysis of the natural imagery using photogrammetry and / or deep learning techniques to create a synthetic representation that can later be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations are not currently practical to consider as candidates for representing natural content because they cannot actually be created in real time to account for use cases requiring real-time delivery. Nevertheless, because computer-generated images are created using 3D modeling processes and tools, currently the best representation candidate for computer-generated images is through the use of a scene graph with synthetic models.

[0043] This dichotomy in optimal representation of both natural and computer-generated content suggests that the optimal ingest format for naturally captured content may be different from the optimal ingest format for computer-generated or natural content that is not required for real-time delivery applications. Thus, the disclosed subject matter aims to be robust enough to support multiple ingest formats for visually immersive media, whether created naturally through the use of a physical camera or created by a computer.

[0044] Below are exemplary techniques for embodying scene graphs as a format suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content, whereby deep learning or photogrammetry techniques are employed to create a corresponding synthetic representation of the natural scene, i.e., not required for real-time distribution applications.

[0045] 1. ORBX (registered trademark) by OTOY OTOY's ORBX is one of several scene graph technologies capable of supporting any type of visual media, timed or untimed, including ray-traceable, legacy (frame-based), stereoscopic, and other types of composited or vector-based visual formats. ORBX is different from other scene graphs because it natively supports freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph intentionally designed to facilitate interchange across multiple vendor technologies that operate on the scene graph. Additionally, ORBX offers a rich material system, support for an open shader language, a robust camera system, and support for Lua scripting. ORBX is also the basis of the Immersive Technologies Media Format, released for royalty-free licensing by the Immersive Digital Experience Alliance (IDEA). In the context of real-time media distribution, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of camera-captured data and composition of that same data into a composite representation. To date, the availability of sufficient computation for real-time delivery is impractical, but nevertheless not impossible.

[0046] 2. Pixar's Universal Scene Description Pixar's Universal Scene Description (USD) is another well-known, mature scene graph that is popular in the VFX and professional content creation industries. USD is integrated into Nvidia's Omniverse platform, a set of developer tools for 3D modeling and rendering with Nvidia's GPUs. A subset of USD was released by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.

[0047] 3. Khronos glTF 2.0 glTF 2.0 is the latest version of the "Graphics Language Transmission Format" specification written by the 3D group at Khronos. This format supports simple scene graph formats, including "png" and "jpeg" image formats, that can generally support static (untimed) objects in a scene. glTF 2.0 supports simple animation and supports the translation, rotation, and scaling of basic shapes, i.e., geometric objects, described using glTF primitives. glTF 2.0 does not support timed media, and therefore does not support video or audio.

[0048] These known designs for immersive visual media scene representations are provided by way of example only and do not limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source to a format suited to the particular characteristics of a client endpoint device.

[0049] Additionally, any or all of the above exemplary media representations currently employ or can employ deep learning techniques to train and create neural network models that enable or facilitate the selection of specific views to fill a particular display's viewing frustum based on the particular dimensions of the frustum. The views selected for a particular display's viewing frustum may be interpolated from existing views explicitly provided in the scene representation, e.g., from MSI or MPI techniques, or may be rendered directly from a rendering engine based on specific virtual camera positions, filters, or virtual camera descriptions for those rendering engines.

[0050] Thus, the disclosed subject matter is sufficiently robust to consider that there is a relatively small but well-known set of immersive media ingest formats that can adequately meet the requirements for both real-time or "on-demand" (e.g., non-real-time) delivery of media that is captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.

[0051] Interpolation of views from immersive media ingest formats using either neural network models or network-based rendering engines will become even easier as advanced network technologies such as 5G for mobile networks and fiber optic cable for fixed networks are deployed. These advanced network technologies increase the capacity and capability of commercial networks, as such advanced network infrastructure can support the transmission and delivery of increasingly large amounts of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networking (SDN), and network functions virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands on specific network resources, e.g., to accommodate dynamic increases or decreases in network throughput, network speed, round-trip latency, and demand for computing resources. Furthermore, this inherent ability to adapt to dynamic network requirements similarly facilitates the network's ability to adapt immersive media ingest formats to appropriate delivery formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.

[0052] Immersive media applications themselves may also have different requirements for network resources, including gaming applications that require significantly lower network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications that may have increasing demands on downlink resources depending on the display type of the client endpoint that is consuming the data. In general, any consumer application may be supported by a variety of client endpoints that include different on-board client capabilities for storage, computation, and power, as well as equally different requirements for the particular media presentation.

[0053] The disclosed subject matter thus enables a fully equipped network, i.e., a network employing some or all of the characteristics of a modern network, to simultaneously support multiple legacy and immersive media-enabled devices in accordance with the characteristics specified below. 1. Providing the flexibility to leverage a practical media ingest format for both real-time and "on-demand" use cases for the delivery of media. 2. Provides flexibility to support both natural and computer-generated content for both legacy and immersive media-enabled client endpoints. 3. Support both timed and non-timed media. 4. Provide a process for dynamically adapting source media ingest formats to appropriate delivery formats based on the capabilities and capabilities of the client endpoint, as well as based on application requirements. 5. Ensure that the delivery format is streamable over IP-based networks. 6. Allows the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media capable devices. 7. We provide an exemplary media representation framework that facilitates the organization of distributed media along scene boundaries.

[0054] An end-to-end embodiment of the improvements enabled by the disclosed subject matter is achieved according to the processes and components described in the detailed description of FIGS. 3-16 as follows:

[0055] Both Figures 3 and 4 use a single exemplary generic delivery format adapted from an ingest source format to accommodate the capabilities of a particular client endpoint. As noted above, the media shown in Figure 3 is timed, while the media shown in Figure 4 is not. The particular generic format is robust enough in its structure to accommodate a wide variety of media attributes, each of which can be layered based on the amount of salient information each layer contributes to the media presentation. It should be noted that such layering is already well-known in the current state of the art, as exemplified by progressive JPEG and scalable video architectures such as those specified in ISO / IEC 14496-10 (Scalable Advanced Video Coding).

[0056] 1. Media streamed according to the generic media format is not limited to legacy visual and audio media, but may include any type of media information capable of interacting with a machine to produce signals that stimulate the human senses of sight, sound, taste, touch, and smell.

[0057] 2. Media streamed according to the generic media format may be both timed and non-timed media, or a combination of both.

[0058] 3. The generic media format is further streamable by enabling layered representations of media objects through the use of a base and enhancement layer architecture. In one example, separate base and enhancement layers are computed by applying multi-resolution or multi-tessellation analysis techniques to media objects within each scene. This is similar to the progressively rendered image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In an exemplary embodiment, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using wavelet analysis.

[0059] In another example of a layered representation of a media format, the enhancement layer applies different attributes to the base layer, such as modifying the material properties of the surface of the visual object represented by the base layer. In yet another example, the attributes may modify the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.

[0060] In yet another example of a layered representation, the surfaces of one or more visual objects in a scene may be changed from Lambertian surfaces to ray-traceable surfaces.

[0061] In yet another example of a layered representation, the network delivers a base layer representation to a client so that the client can create a nominal presentation of the scene while awaiting the transmission of additional enhancement layers to refine the resolution or other characteristics of the base representation.

[0062] 4. The resolution of the attributes or refinement information in the enhancement layer is not explicitly coupled to the resolution of the objects in the base layer, unlike today's existing MPEG video and JPEG image standards.

[0063] 5. A generic media format supports any type of information media that can be presented or acted upon by a presentation device or machine, thereby enabling support of heterogeneous media formats for heterogeneous client endpoints. In one embodiment of a network that delivers media formats, the network first queries the client endpoint to determine the client's capabilities, and if the client is unable to meaningfully consume the media representation, the network either removes layers of attributes not supported by the client or adapts the media from its current format to a format appropriate for the client endpoint. In one example of such adaptation, the network converts a volumetric visual media asset into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can use neural network processes to reformat the media into an appropriate format or, optionally, synthesize a view required by the client endpoint.

[0064] 6. A manifest for a full or partially full immersive experience (such as a live streaming event, game, or on-demand asset playback) is organized by scenes, which are the minimum amount of information that rendering and game engines can currently capture to create the presentation. The manifest includes a list of individual scenes to be rendered in their entirety for the immersive experience requested by the client. Associated with each scene are one or more representations of the geometric objects in the scene that correspond to a streamable version of the scene geometry. One embodiment of a scene representation relates to a low-resolution version of the scene's geometric objects. Another embodiment of the same scene relates to enhancement layers for the low-resolution representation of the scene to add further detail or increase tessellation to the geometric objects of the same scene. As mentioned above, each scene may have two or more enhancement layers to progressively increase the detail of the scene's geometric objects.

[0065] 7. Each layer of a media object referenced in a scene is associated with a token (e.g., a URI) that points to an address where the resource can be accessed in the network. Such a resource is similar to a CDN where content can be fetched by a client.

[0066] 8. Tokens for representations of geometric objects may point to locations within the network or within the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.

[0067] In the figures described below, the same reference numeral may be indicated for multiple elements of the arrangement, and in such cases, the description may be assumed to relate to any and all of those identically labeled elements, respectively.

[0068] 3 illustrates an embodiment of a generic media format for timed media as follows: A timed scene manifest includes a list of scene information 301. The scene 301 references a list of components 302 that separately describe the processing information and types of media assets that make up the scene 301. The components 302 reference assets 303, which further reference a base layer 304 and an attribute enrichment layer 305. A list of unique assets not previously used in other scenes is provided at 307.

[0069] 4 illustrates an embodiment of a generic media format for non-timed media: Scene information 401 is not associated with start and end times according to a clock. Scene information 401 references a list of components 402 that separately describe the types of media assets and processing information that make up the scene 401. Components 402 reference assets 403, which in turn reference base layer 404 and attribute enrichment layers 405, 406. Additionally, scene 401 references other scenes 401 for non-timed media. Scene 401 also references scene 407 for timed media scenes. List 406 identifies unique assets associated with a particular scene that have not previously been used in a higher-level (e.g., parent) scene.

[0070] FIG. 5 illustrates one embodiment of a process 500 for synthesizing an ingest format from natural content. Camera unit 501 captures a scene of a person using a single camera lens. Camera unit 502 captures a scene with five diverging fields of view by mounting five camera lenses around a ring-shaped object. The arrangement of 502 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter of a sphere. The arrangement of 503 is an exemplary arrangement commonly used to capture a light field for a light field display or a holographic immersive display. Natural image content 509 is provided as input to a synthesis process 504, which can optionally employ a neural network training process 505 that uses a set of training images 506 to generate an optional capture neural network model 508. Another commonly used process in place of the training process 505 is photogrammetry. 5, the model 508 becomes one of the assets of the ingest format 507 for natural content. Example embodiments of the ingest format 507 include MPI and MSI.

[0071] FIG. 6 illustrates one embodiment of a process 600 for creating synthetic media, e.g., an ingest format for computer-generated images. A LIDAR camera 601 captures a point cloud 602 of a scene. A CGI tool, 3D modeling tool, or another animation process for creating synthetic content is used on a computer 603 to create CGI assets over a network (604). A sensored motion capture suit 605A is worn by an actor 605 and captures a digital recording of the actor's 605 movements to generate animated MoCap data 606. Data 602, 604, and 606 are provided as input to a compositing process 607, which can also optionally use a neural network and training data to create a neural network model (not shown in FIG. 6).

[0072] The techniques for rendering and streaming heterogeneous immersive media described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.

[0073] Computer software may be coded using any suitable machine code or computer language that may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that are executable by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or via interpretation, microcode execution, etc.

[0074] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0075] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 700.

[0076] The computer system 700 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0077] The input human interface devices may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touchscreen 710, a data glove (not shown), a joystick 705, a microphone 706, a scanner 707, and a camera 708 (only one of each is shown).

[0078] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 710, data gloves (not shown), or joystick 705, although some haptic feedback devices may not function as input devices), audio output devices (e.g., speakers 709, headphones (not shown)), visual output devices (e.g., screens 710, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities and each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output in greater than three dimensions via means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown).

[0079] The computer system 700 may also include human-accessible storage devices and media associated with the storage devices, such as optical media including CD / DVD ROM / RW 720 having media 721 such as CD / DVDs, thumb drives 722, removable hard drives or solid state drives 723, legacy magnetic media such as tape or floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), and the like.

[0080] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0081] The computer system 700 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus; and the like. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (749) (e.g., a USB port on the computer system 700), while other networks are generally integrated into the core of the computer system 700 by attachment to a system bus as described below (e.g., an Ethernet interface to a personal computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 700 may communicate with other entities. Such communications may be one-way receive-only (e.g., television broadcast), one-way transmit-only (e.g., a CANbus to a particular CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used in each of those networks and network interfaces, as described above.

[0082] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to core 740 of computer system 700 .

[0083] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 743, hardware accelerators 744 for specific tasks, etc. These devices may be connected via a system bus 748, along with read-only memory (ROM) 745, random access memory 746, and internal mass storage 747, such as a non-user-accessible internal hard drive or SSD. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 748 or via a peripheral bus 749. Peripheral bus architectures include PCI, USB, etc.

[0084] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may combine to execute specific instructions that may constitute the aforementioned computer code. That computer code may be stored in ROM 745 or RAM 746. Transient data may also be stored in RAM 746, while permanent data may be stored, for example, in internal mass storage 747. Rapid storage and retrieval to any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc.

[0085] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0086] By way of example and not limitation, a computer system having architecture 700, and specifically core 740, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as described above, as well as media associated with specific storage of core 740 that is non-transitory in nature, such as core internal mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 740. The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause core 740, and specifically the processors in the core (including a CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. Additionally, or alternatively, a computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator 744) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0087] FIG. 8 illustrates an exemplary network media distribution system 800 supporting a variety of legacy displays and heterogeneous immersive media-enabled displays as client endpoints. A content acquisition process 801 captures or creates media using the exemplary embodiments of FIG. 6 or FIG. 5. An ingest format is created in a content preparation process 802 and then transmitted to the network media distribution system using a transmission process 803. A gateway 804 can service customer premises equipment (CPE) to provide network access to various client endpoints in the network. A set-top box 805 can also function as CPE to provide access to aggregated content by a network service provider. A wireless demodulator 806 can function as a mobile network access point for mobile devices, as shown, for example, by a mobile handset display 813. In this particular embodiment of the system 800, a legacy 2D television 807 is shown connected directly to the gateway 804, the set-top box 805, or a WiFi router 808. A computer laptop with a legacy 2D display 809 is shown as a client endpoint connected to the WiFi router 808. A head-mounted 2D (raster-based) display 810 is also connected to the router 808. A lenticular light field display 811 is shown connected to the gateway 804. The display 811 consists of a local compute GPU 811A, a storage device 811B, and a visual presentation unit 811C that creates multiple views using ray-based lenticular optics technology. A holographic display 812 is shown connected to the set-top box 805. The display 812 consists of a local compute CPU 812A, a GPU 812B, a storage device 812C, and a Fresnel pattern, wave-based holographic visualization unit 812D. An augmented reality headset 814 is shown connected to the wireless demodulator 806.Headset 814 is comprised of a GPU 814A, a storage device 814B, a battery 814C, and a volumetric visual presentation component 814D. A high-density light field display 815 is shown connected to WiFi router 808. Display 815 is comprised of multiple GPUs 815A, a CPU 815B, a storage device 815C, an eye-tracking device 815D, a camera 815E, and a high-density beam-based light field panel 815F.

[0088] FIG. 9 illustrates one embodiment of an immersive media distribution process 900 capable of serving legacy and heterogeneous immersive media-enabled displays such as those previously shown in FIG. 8. Content is created or acquired in process 901, which is further embodied in FIGS. 5 and 6 for natural and CGI content, respectively. The content 901 is then converted to an ingest format using a network ingest format creation process 902, which is similarly further embodied in FIGS. 5 and 6 for natural and CGI content, respectively. The ingest media is optionally updated to store information from a media reuse analyzer 911 about assets that may be reused across multiple scenes. The ingest media format is transmitted to a network and stored in a storage device 903. Optionally, the storage device may reside on the immersive media content creator's network and be remotely accessed by an immersive media network distribution process (not numbered), as indicated by the dashed line bisecting 903. Client and application specific information is optionally available in a remote storage device 904, which may optionally reside remotely in an alternative "cloud" network.

[0089] As shown in FIG. 9, the network orchestration process 905 serves as the primary source and sink of information for performing the primary tasks of the distribution network. In this particular embodiment, process 905 may be implemented in an integrated format with other components of the network. Nevertheless, the tasks represented by process 905 in FIG. 9 form essential elements of the disclosed subject matter. The orchestration process 905 may further employ a bidirectional message protocol with clients to facilitate all processing and delivery of media according to client characteristics. Furthermore, the bidirectional protocol may be implemented across different distribution channels, i.e., control plane channels and data plane channels.

[0090] Process 905 receives information about the characteristics and attributes of client 908 and further gathers requirements regarding applications currently running on 908. This information may be obtained from device 904, or in alternative embodiments, by directly querying client 908. In the case of direct querying of client 908, it is assumed that a two-way protocol (not shown in FIG. 9) exists and is operational so that the client may communicate directly with orchestration process 905.

[0091] The orchestration process 905 also initiates and communicates with the media adaptation process 910 described in FIG. 10. Once the ingested media has been adapted and fragmented by process 910, the media is optionally transferred to an intermediate storage device, shown as media storage device 909 prepared for delivery. Once the delivery media is prepared and stored on device 909, the orchestration process 905 ensures that the immersive client 908 receives the delivery media and corresponding description information 906 via its network interface 908B through a "push" request, or the client 908 itself can initiate a "pull" request for the media 906 from the storage device 909. The orchestration process 905 can use a two-way message interface (not shown in FIG. 9) to perform the "push" request or to initiate a "pull" request by the client 908. The immersive client 908 can optionally use a GPU (or CPU, not shown) 908C. The media delivery format is stored on a storage device or storage cache 908D of the client 908. Finally, the client 908 visually presents the media via its visualization component 908A.

[0092] Throughout the process of streaming immersive media to clients 908, the orchestration process 905 monitors the status of the client's progress via a client progress and status feedback channel 907. The status monitoring may be performed by a two-way communication message interface (not shown in FIG. 9).

[0093] FIG. 10 depicts a specific embodiment of a media adaptation process so that ingested source media can be appropriately adapted to match the requirements of the client 908. Controlled by one or more processors, the media adaptation process 1001 is composed of multiple components that facilitate the adaptation of ingested media to an appropriate delivery format for the client 908. These components should be considered exemplary. In FIG. 10, the adaptation process 1001 receives input network status 1005 to track the current traffic load on the network. The client 908 information includes attribute and feature descriptions, application capabilities and descriptions, the current status of the application, and a client neural network model (if available) that assists in mapping the client's frustum geometry to the interpolation capabilities of the ingested immersive media. Such information may be obtained via a two-way message interface (not shown in FIG. 10). The adaptation process 1001 ensures that the adapted output, once created, is stored in the client adaptation media storage device 1006. The media reuse analyzer 1007 is shown in FIG. 10 as an optional process that may be run a priori or as part of a network automation process for the distribution of media.

[0094] The adaptation process 1001 is controlled by the logic controller 1001F. The adaptation process 1001 also employs a renderer 1001B or neural network processor 1001C to adapt the particular ingest source media to a format appropriate for the client. The neural network processor 1001C uses the neural network model of 1001A. Examples of such neural network processors 1001C include deep view neural network model generators such as those described in MPI and MSI. If the media is in a 2D format but the client requires it in a 3D format, the neural network processor 1001C can invoke a process that uses highly correlated images from the 2D video signal to derive a volumetric representation of the scene depicted in the video. An example of a suitable renderer 1001B may be a modified version of the OTOY Octane renderer (not shown) modified to interact directly with the adaptation process 1001. The adaptation process 1001 can optionally use a media compressor 1001D and a media decompressor 1001E depending on the needs of these tools regarding the format of the ingested media and the format required by the client 908.

[0095] Figure 11 shows a delivery format creation process 1100. An adapted media packaging process 1103 packages media from a media adaptation process 1101 (shown as process 1000 in Figure 10) currently resident on a client adaptation media storage device 1102. The packaging process 1103 formats the adapted media from process 1101 into a robust delivery format 1104, such as the exemplary formats shown in Figures 3 or 4. Manifest information 1104A provides the client 908 with a list 1104B of scene data assets that the client 908 can expect to receive, as well as optional complexity metadata that describes the complexity of all assets in the scene. List 1104B shows a list of visual, audio, and haptic assets, each with corresponding metadata.

[0096] 12 shows a packetizer process system 1200. A packetizer process 1202 separates adapted media 1201 into individual packets 1203 suitable for streaming to a client 908.

[0097] The components and communications shown in FIG. 13 for sequence diagram 1300 are described below. Client endpoint 1301 initiates a media request 1308 to network delivery interface 1302. Request 1308 includes information identifying the media the client requests, either by URN or other standard nomenclature. Network delivery interface (also known as client 1302) responds to request 1308 with a profile request 1309, which requests that client 1301 provide information about its currently available resources (including computation, storage, battery charge, and other information characterizing the client's current operating status). Profile request 1309 also requests that the client provide one or more neural network models, if such models are available at the client, that can be used by the network for neural network inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system. Response 1311 from client 1301 to interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). Interface 1302 then provides client 1301 with session ID token 1311. Interface 1302 then makes a request to ingest media server 1303 with Ingest Media Request 1312 that includes the URN or other canonical name of the media identified in request 1308. Server 1303 responds to request 1312 with response 1313 that includes the ingest media token. Interface 1302 then provides the media token from response 1313 to client 1301 in call 1314.Interface 1302 then initiates an adaptation process for the requested media at 1308 by providing the ingest media token, client token, application token, and neural network model token to adaptation interface 1304. Interface 1304 requests access to the ingest media by providing the ingest media token to server 1303 in call 1316 to request access to the ingest media asset. Server 1303 responds to request 1316 with the ingest media access token in response 1317 to interface 1304. Interface 1304 then requests that media adaptation process 1305 adapt the ingest media located in the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created at 1313. A request 1318 from interface 1304 to process 1305 includes the necessary tokens and session ID. Process 1305 provides the adapted media access token and session ID to interface 1302 in update 1319. Interface 1302 provides the adapted media access token and session ID to packaging process 1306 in interface call 1320. Packaging process 1306 provides response 1321 with the packaged media access token and session ID to interface 1302. Process 1306 provides the packaged media access token for the packaged set, URN, and session ID to packaging media server 1307 in response 1322. Client 1301 executes request 1323 to start streaming the media asset corresponding to the packaged media access token received in message 1321.Client 1301 performs other requests and provides status updates to interface 1302 in message 1324 .

[0098] FIG. 14 illustrates the logic flow of the media reuse analyzer 911 of the immersive media data reuse optimizer, shown in FIG. 9 as media reuse analyzer 1400. Process initialization begins at step 1401. Initialization step 1402 initializes an iterator “i” to 0 and also initializes a set of lists 1404 (one list for each scene) that identify unique assets encountered across all scenes that make up the presentation, as shown in FIG. 3 or FIG. 4. List 1404 illustrates sample list entries of information describing assets that are unique with respect to the entire presentation, including an indicator of the type of media that makes up the asset (e.g., mesh, audio, or volume), a unique identifier for the asset, and the number of times the asset is used across the set of scenes that make up the presentation. As an example, for scene N-1, no assets are included in its list because all assets required for scene N-1 have been identified as assets that are also used in scene 1 and scene 2. Step 1403 determines whether iterator “i” is less than the total number of scenes that make up the presentation (as shown in FIG. 3 or FIG. 4). If iterator "i" is equal to the number N of scenes that make up the presentation, the reuse analysis ends at step 1405. Otherwise, if iterator "i" is less than the total number of scenes, processing proceeds to step 1406, where iterator "j" is set to 0. Step 1407 tests iterator "j" to determine whether iterator "j" is less than the total number of media assets (also called media objects) in the current scene "i". If iterator "j" is less than the total number of media assets in scene "i", processing proceeds to step 1408. Otherwise, processing proceeds to step 1412, where iterator "i" is incremented by 1 before returning to step 1403. If the value of "j" is less than the total number of assets in scene "i", processing proceeds to conditional step 1408, where the characteristics of the media asset are compared to previously analyzed assets from scenes prior to the current scene "i".If the asset is identified as an asset used in a scene prior to scene "i", then in step 1411 the number of times the asset has been used across scenes 0 to N-1 is incremented by one. Otherwise, if the asset is a unique asset, i.e., has not been previously analyzed in a scene associated with a lower value of iterator "i", then in step 1409 a unique asset entry is created in list 1404 for scene "i". Step 1409 also creates and assigns a unique identifier to the asset's entry and sets the number of times the asset has been used across scenes 0 to N-1 to one. Following step 1409, processing proceeds to step 1410, where iterator "j" is incremented by one. After step 1410, processing returns to step 1407.

[0099] While this disclosure describes several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0100] 101 ingest media, 102 process, 102B conditional logic, 102C indicator, 104 edge process, 105 stream, 106 rendering process, 108 target client, 200 media transformation decision process, 201 ingest media, 202 process, 203 decision process, 204 process, 205 media, 206 preparation process, 207 process, 300 timed media representation, 301 scene information, 302 components, 303 assets, 304 base layer, 305 attribute enrichment layer, 400 non-timed media representation, 401 scene information, 402 components, 403 assets, 404 base layer, 405 attribute enrichment layer, 406 attribute enrichment layer, 407 scene, 500 process, 501 camera unit, 502 camera unit, 503 camera unit, 504 Synthesis process, 505 Training process, 506 Training images, 507 Ingest format, 508 Model, 509 Natural image content, 600 Process, 601 LIDAR camera, 602 Data, 603 Computer, 604 Data, 605 Actor, 605A Motion Capture Suit with Sensors, 606 MoCap data, 607 Synthesis process, 700 Computer system, 701 Keyboard, 702 Mouse, 703 Trackpad, 705 Joystick, 706 Microphone, 707 Scanner, 708 Camera, 709 Speaker, 710 Screen, 721 Media, 722 Thumb drive, 723 Solid-state drive, 740 Core, 741 Central Processing Unit (CPU), 742 Graphics Processing Unit (GPU), 743 Field Programmable Gate Array (FPGA), 744 Accelerator, 745 Read-only memory (ROM), 746 Random access memory (RAM), 747 Mass storage, 748 System bus, 749 Peripheral bus, 800 Network media distribution system, 801 Content acquisition process, 802 Content preparation process, 803 Transmission process, 804 Gateway, 805 Set-top box, 806 Wireless demodulator, 807 Legacy 2D television, 808WiFi router, 809 Legacy 2D display, 810 Display, 811 Display, 811A GPU, 811B Storage device, 811C Visual presentation unit, 812 Holographic display, 812A CPU, 812B GPU, 812C Storage device, 812D Holographic visualization unit, 813 Mobile handset display, 814 Headset, 814A GPU, 814B Storage device, 814C Battery, 814D Volumetric visual presentation component, 815 High density light field display, 815A GPU, 815B CPU, 815C Storage device, 815D Eye tracking device, 815E Camera, 815F Light field panel, 900 Immersive media delivery process, 901 Content, 901 Process, 902 Process, 902 Process, 903 Storage device, 904 Remote storage device, 905 orchestration process, 906 description information, 907 status feedback channel, 908 immersive client, 908A visualization component, 908B network interface, 908C GPU, 908D storage cache, 909 media storage device, 910 media adaptation process, 911 media reuse analyzer, 1000 process, 1001 media adaptation process, 1001A neural network model, 1001B renderer, 1001C neural network processor, 1001D media compressor, 1001E media decompressor, 1001F logic controller, 1005 input network status, 1006 client adapted media storage device, 1007 media reuse analyzer, 1100 delivery format creation process, 1101 media adaptation process, 1102 client adapted media storage device, 1103 packaging process, 1104 delivery format, 1104A Manifest information, 1104B list, 1200 Packetizer process system, 1201 Media, 1202 Packetizer process, 1203 Packet, 1300 Sequence diagram, 1301Client Endpoint, 1302 Interface, 1303 Ingest Media Server, 1304 Interface, 1305 Media Adaptation Process, 1306 Packaging Process, 1307 Packaging Media Server, 1308 Media Request, 1309 Profile Request, 1311 Session ID Token, 1312 Request, 1313 Response, 1314 Call, 1316 Call, 1317 Response, 1318 Request, 1319 Update, 1320 Interface Call, 1321 Response, 1322 Response, 1323 Request, 1324 Message, 1400 Media Reuse Analyzer, 2030 Asset Reuse Logic, 2031 Ingest Media, 2032 Process, 2033 Decision-Making Process, 2035 Decision-Making Process, 2036 Prepare Process, 2037 Process, 2038 Process, 2039 Media, 20330 Asset Reuse Logic, 20331 Ingest Media, 20332 Process, 20333 Decision Making Process, 20334 Decision Making Process, 20335 Decision Making Process, 20336 Preparation Process, 20337 Process, 20338 Process, 20339 Media

Claims

1. 1. A method for an immersive media presentation including a plurality of scenes, implemented by at least one hardware processor, the method comprising: determining that a media object appears in at least two or more scenes of the plurality of scenes associated with the immersive media presentation; creating a proxy corresponding to the media object in response to determining that the media object appears in at least two or more scenes of the plurality of scenes associated with the immersive media presentation; sending a request to a client inquiring whether the client has access to the media objects appearing in at least two or more scenes in a local cache, the media objects representing particular instances of particular formats of media data used to compose scenes, and the client having sufficient storage resources to store copies of the media objects associated with the immersive media presentation in the local cache; receiving a response from the client indicating whether the client can access the media object that appears in at least two or more scenes in the local cache; signaling the proxy to the client to use the media object in a subsequent scene without further delivery of the media object to the client in response to the response indicating that the client has access to the media object that appears in at least two or more scenes in the local cache; delivering the media object to the client in response to the response indicating that the client cannot access the media object that appears in at least two or more scenes in the local cache; A method comprising:

2. initializing a set of lists, each list corresponding to a respective one of the scenes of the immersive media presentation; wherein initializing the set of lists includes incrementally assigning a unique identifier to each of the media objects, including the media object, that appear in the scene. The method of claim 1.

3. The method of claim 2 , wherein initializing the set of lists further comprises incrementally determining the number of times each media object appears in each of the scenes.

4. receiving a request from the client for the immersive media presentation; requesting, in response to said request, that said client provide an indication of said client's client resources; The method of claim 1 further comprising:

5. wherein requesting that the client provide the indication of the client resources includes requesting that the client provide one or more neural network models; processing the media object includes neural network inference based on the one or more neural network models requested by the client; The method of claim 4.

6. the processing of the media object is based on determining a current traffic load on a network interfacing the at least one hardware processor and the client. The method of claim 5.

7. monitoring the client's progress in outputting the immersive media presentation; the step of sending the request is timed based on the progress. The method of claim 1.

8. The method of claim 1 , wherein the immersive media presentation includes instructions to the client to stimulate at least one of the senses of sight, hearing, taste, touch, and smell.

9. The method of claim 1 , wherein the request to the client further inquires whether the client has access to the media object, and the access is local to the client.

10. The method of claim 1 , wherein the immersive media presentation comprises one of a timed presentation and an untimed presentation.

11. 1. An apparatus comprising: at least one memory configured to store computer program code; and at least one hardware processor configured to access said computer program code and to act as instructed by said computer program code, said computer program code causing said at least one hardware processor to perform the method of any one of claims 1 to 10.

12. A program causing a computer to execute a process, said process causing said computer to execute a method according to any one of claims 1 to 10.

13. The method of claim 1, wherein the immersive media presentation comprises multiple media objects for each of the scenes, and the proxy is a unique identifier for the media object.

Citation Information

Patent Citations

  • Media data providing method, media data reception method, and program

    JP2020047303A

  • System and method for managing status change of multimedia asset in multimedia distribution system

    JP2020198650A

  • Optimizing Audio Delivery for Virtual Reality Applications

    JP2020537418A

  • System and method for providing a virtual immersive environment

    US20140192087A1