Immersive media presentation method, apparatus, and storage medium
By optimizing resource allocation and utilizing neural network models in the decision-making process between the network and the client, the problem of resource waste and latency of heterogeneous client devices in commercial networks is solved, and efficient immersive media distribution is achieved.
Patent Information
- Application Number
- CN202280007882.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2022-10-25
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-10-25
AI Technical Summary
Existing commercial networks struggle to effectively manage the diversity of heterogeneous client devices when distributing immersive media, leading to resource waste and latency, and failing to efficiently convert media formats from ingestion formats to formats suitable for client presentation.
By implementing a decision-making process between the network and the client, it determines which media assets can be reused in the client's local cache, avoiding redundant transformation and streaming steps, and optimizing resource allocation using neural network models and advanced networking techniques.
It improves the efficiency of media distribution, reduces resource waste and delays, and ensures that heterogeneous client devices can efficiently present immersive media.
Smart Images

Figure CN116569535B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application 63 / 276,535, filed November 5, 2021, and U.S. Application 17 / 970,109, filed October 20, 2022, the contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure describes embodiments of architectures, structures, and components generally relating to systems and networks for distributing media, including video, audio, geometric (3D) objects, haptic feedback, associated metadata, or other content for client rendering devices. Specific embodiments relate to systems, structures, and architectures for distributing media content to heterogeneous immersive and interactive client rendering devices. Background Technology
[0004] "Immersive media" generally refers to media that stimulates any or all human sensory systems (visual, auditory, tactile, olfactory, and possibly gustatory) to create or enhance the user's perception of what physically exists in the media experience, that is, media that goes beyond timed two-dimensional (2D) video and corresponding audio distributed on existing (e.g., "traditional") commercial networks; such timed media are also called "traditional media".
[0005] Another definition of “immersive media” is media that attempts to create or mimic the physical world through digital simulations of dynamics and physical laws, thereby stimulating any or all human sensory systems to create a user’s perception of a scene that physically exists within a depiction of the real or virtual world.
[0006] Presentation devices with immersive media capabilities can refer to devices equipped with sufficient resources and capabilities to access, interpret, and present immersive media. In terms of media delivered over the network, such devices are heterogeneous in terms of the quantity and formats of media they can support. Similarly, the media itself is heterogeneous in terms of the quantity and types of network resources required for large-scale distribution of such media. "Large-scale" can refer to the distribution of media by service providers that achieves distribution equivalent to that of traditional video and audio media on the network (e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions).
[0007] In contrast, traditional presentation devices such as laptop displays, televisions, and mobile handheld displays are homogenous in their capabilities in that all of these devices currently include rectangular display screens that use 2D rectangular video or still images as their primary visual media format. Some of the visual media formats commonly used in traditional presentation devices can include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266.
[0008] The distribution of any media over a network can employ a media delivery system and architecture that reformats media from an input or network "ingested" media format to a distribution media format that is not only suitable for ingestion by a target client device and its applications, but also facilitates being "streamed" over the network. Thus, there can be two processes performed by the network on ingested media: 1) converting the media from format A to format B that is suitable for ingestion by the target client, i.e., based on the client's ability to ingest certain media formats, and 2) preparing the media to be streamed.
[0009] "Streaming" of media broadly refers to the segmentation and / or packetization of media so that it can be delivered over a network in successive smaller sized "chunks" that are logically organized and ordered according to one or both of the temporal or spatial structure of the media. "Transcoding" of media from format A to format B (sometimes referred to as "transcoding") can be a process that is typically performed by the network or service provider prior to distributing the media to the client. Such transcoding can include converting the media from format A to format B based on a priori knowledge that format B is the preferred or only format that can be ingested by the target client or is more suitable for distribution over a constrained resource such as a commercial network. In many but not all cases, both the steps of transcoding the media and preparing the media to be streamed are necessary before the client can receive and process the media from the network.
[0010] The above one or two step process acts on the ingested media through the network, i.e., before the media is distributed to the client, a media format is produced that is referred to as the "distribution media format" or simply "distribution format." Typically, due to technical limitations, these steps should only be performed once if they are performed completely for a given media data object if the network has access to information that indicates that the client will need the media object transformed and / or streamed on multiple occasions (which would trigger transformation and streaming of such media multiple times). That is, the processing and transfer of data for transformation and streaming of media is generally considered a source of delay that requires potentially large amounts of network and / or computing resources to expend. Thus, a network that does not have access to information that indicates when a client can already have a particular media data object stored in its cache or locally relative to the client performs sub-optimally to a network that does have access to such information.
[0011] For legacy rendering devices, the distribution format can be identical or substantially identical to the "rendering format" that the client rendering device ultimately uses to create a rendering. That is, the rendering media format is one that is closely tuned in properties (resolution, frame rate, bit depth, color gamut, etc.) to the capabilities of the client rendering device. Some examples of distribution versus rendering formats include a high definition (HD) video signal (1920 pixel columns x 1080 pixel rows) distributed by a network to an ultra-high definition (UHD) client device that has a resolution of (3840 pixel columns x 2160 pixel rows). In this scenario, the UHD client would apply a process called "super resolution" to the HD distribution format in order to increase the resolution of the video signal from HD to UHD. Thus, the final signal format rendered by the client device is the "rendering format," which in this example is the UHD signal, while the HD signal comprises the distribution format. In this example, the HD signal distribution format is very similar to the UHD signal rendering format because both signals are linear video formats and the process of converting the HD format to the UHD format is a relatively straightforward and easy process to perform on most legacy client devices.
[0012] Alternatively, the preferred rendering format of the target client device can be significantly different from the ingest format received by the network. However, the client has access to sufficient computing, storage, and bandwidth resources to transform the media from the ingest format to the necessary rendering format suitable for rendering by the client. In this scenario, the network can bypass the step of reformatting the ingest media (e.g., "transcoding" the media from format A to format B) simply because the client has access to sufficient resources to perform all of the media transformations, and the network does not have to use such a priori practices. However, the network can still perform the steps of segmenting and encapsulating the ingest media so that the media can be streamed to the client.
[0013] Yet another option is that the ingest media received by the network is significantly different from the client's preferred presentation format, and the client does not have access to sufficient computing, storage, and / or bandwidth resources to convert the media to the preferred presentation format. In such scenarios, the network can help the client by performing some or all of the transformations from the ingest format to a format that is identical or nearly identical to the client's preferred presentation format on behalf of the client. In some architectural designs, such help provided by the network on behalf of the client is commonly referred to as "split rendering."
[0014] In each of the above scenarios given the transformation of media from format A to another format can be done entirely by the network, entirely by the client, or jointly between the network and the client (e.g., for split rendering), it is apparent that a lexicon describing the attributes of the media formats can be needed so that both the client and the network have complete information to characterize the work that must be done. Furthermore, it can also be necessary to provide a lexicon of the attributes of the client's capabilities in terms of, for example, available computing resources, available storage resources, and access to bandwidth. Still further, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of the ingest format so that the network and the client can jointly or individually determine whether or when the network can employ split rendering steps to distribute media to the client. Finally, if it can be avoided that the client has to do the transformation and / or streaming of particular media objects needed or that will be needed for the client to do the presentation of its media, then the network can skip the steps of transformation and streaming (assuming the client is able to access or obtain the media objects it can need) in order to do the presentation of the client's media. Such a network has sufficient information to avoid repeated transformation and / or streaming steps for assets that are used more than once, which it can do better than a network not so designed. SUMMARY
[0015] A method and apparatus are included herein, the apparatus comprising a memory configured to store computer program code and one or more hardware processors configured to access the computer program code and operate as instructed by the computer program code, the computer program code comprising determining code configured to cause the at least one hardware processor to determine that a media asset appears in at least two or more scenes of the plurality of scenes associated with the immersive media presentation; sending code configured to cause the at least one hardware processor to send a request to a client, the request querying whether the client has access to the media asset appearing in the at least two or more scenes in a local cache, wherein the client has sufficient storage resources to store copies of media assets associated with the immersive media presentation in the local cache; receiving code configured to cause the at least one hardware processor to receive a reply from the client, the reply indicating whether the client has access to the media asset appearing in the at least two or more scenes in the local cache; signaling code configured to cause the at least one hardware processor to signal the client to use the media asset in subsequent scenes without further distribution of the media asset to the client in response to the reply indicating that the client has access to the media asset appearing in the at least two or more scenes in the local cache; and distributing code configured to cause the at least one hardware processor to distribute the media asset to the client in response to the reply indicating that the client does not have access to the media asset appearing in the at least two or more scenes in the local cache.
[0016] According to example embodiments, the computer program code further comprises initialization code configured to cause the at least one hardware processor to implement initializing a set of lists, wherein each of the lists respectively corresponds to a respective scene of the plurality of scenes of the immersive media presentation, wherein initializing the set of lists comprises incrementally assigning a unique identifier to each of a plurality of media assets appearing in the plurality of scenes respectively, the plurality of media assets including the media asset.
[0017] According to example embodiments, initializing the set of lists further comprises incrementally determining a number of times each asset respectively appears in a respective scene.
[0018] According to example embodiments, the computer program code further comprises receiving code configured to cause the at least one hardware processor to receive a request for the immersive media presentation from the client; and requesting code configured to cause the at least one hardware processor to request the client to provide an indication of client resources of the client in response to the request, wherein the processing of the asset is implemented in reliance on the indication of client resources.
[0019] According to example embodiments, requesting the client to provide the indication of the client resources comprises requesting the client to provide one or more neural network models, and wherein the processing of the media asset comprises neural network inference based on the one or more neural network models requested from the client.
[0020] According to example embodiments, the processing of the media asset is further based on determining a current traffic load on a network connecting the at least one hardware processor and the client.
[0021] According to example embodiments, the computer program code further comprises detection code configured to cause the at least one hardware processor to detect a progress of the client in outputting the immersive media presentation, wherein the timing of sending the query is based on the progress.
[0022] According to example embodiments, the immersive media presentation comprises instructions to stimulate a sense of at least one of vision, sound, and at least one of taste, touch, and smell of the client.
[0023] According to example embodiments, the query to the client further queries whether the client has access to the asset, wherein the access is local to the client.
[0024] According to example embodiments, the immersive media presentation comprises any one of a timed presentation and an untimed presentation.
[0025] The techniques described herein improve computer technology by facilitating various aspects such as decision processes used by a network and / or a client to determine whether some or all ingest media should be transformed from format A to format B to further facilitate the ability of the client to produce a presentation of the media in a possible third format C. To aid in such decision processes, assume that a method exists or can be readily employed by a network by design to determine in the context of a presentation which assets are used more than once in the presentation. Relying on information from such analysis, a network can then be designed such that a requesting client retains a copy of each asset that is used more than once in its local cache. However, in this scenario, the network can not have control over the management of the client's local cache and, as a result, the client can encounter situations where it must delete a resource (even a reusable resource) from its local cache. To facilitate optimization of the network's design so as to minimize the need to perform a transformation from format A to format B on an asset that is used more than once, or to facilitate the network not having to stream an asset that is used more than once to the client, the network can first query the client for feedback to ensure that the asset under consideration is still available in the client's local cache. If the client's reply indicates that it no longer has a copy of the asset under consideration, then the network can transform the ingest asset from format A to format B and / or stream the original copy of the asset to the client. Such a network ensures that the client has access to the asset either through the client's own asset copy stored in the client's local cache (previously provided to the client by the network), or via the network repeating the steps of transforming and / or streaming the asset to the client again. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 is a schematic illustration of a media stream for distribution to a client by a network according to an example embodiment.
[0027] Figure 1-02 is the same schematic illustration as Figure 1 but with added logic to determine whether a substitute for the original media itself (transformed to another format or in its original format) should be streamed instead.
[0028] Figure 2 is a schematic illustration of a media stream by a network according to an example embodiment in which a decision process is employed to determine whether the network should transform the media before distributing it to a client.
[0029] Figure 2-03 is the same schematic illustration as Figure 2 but with added logic to determine whether a substitute for the original media itself (transformed to another format or in its original format) should be streamed instead.
[0030] Figure 2-033 is a schematic illustration of a representation and streaming of timed immersive media according to exemplary embodiments, wherein such timed immersive media contains a list of assets that are reused across a set of N scenes. Figure 2-03
[0031] Figure 3 is a schematic illustration of an embodiment of a data model for a representation and streaming of timed immersive media according to exemplary embodiments, wherein such timed immersive media contains a list of assets that are reused across a set of N scenes.
[0032] Figure 4 is a schematic illustration of an embodiment of a data model for a representation and streaming of untimed immersive media according to exemplary embodiments, wherein such untimed immersive media contains a list of assets that are reused across a set of 5 scenes.
[0033] Figure 5 is a schematic illustration of a process of capturing a natural scene and converting it into a representation that can be used as an ingest format for a network serving a heterogeneous set of client endpoints.
[0034] Figure 6 is a schematic illustration of a process of creating a representation of a synthetic scene using 3D modeling tools and formats that can be used as an ingest format for a network serving a heterogeneous set of client endpoints.
[0035] Figure 7 is a system diagram of a computer system according to exemplary embodiments.
[0036] Figure 8 is a schematic illustration of a network serving a plurality of heterogeneous client endpoints according to exemplary embodiments.
[0037] Figure 9 is a schematic illustration of a network providing adaptation information about a particular media represented in a media ingest format, for example, prior to a process in which the network adapts the media for use by a particular immersive media client endpoint, according to exemplary embodiments.
[0038] Figure 10 is a system diagram of a media adaptation process consisting of a media renderer- converter that converts source media from its ingest format to a particular format suitable for a particular client endpoint, according to exemplary embodiments.
[0039] Figure 11 is a schematic illustration of a network formatting adapted source media into a data model suitable for representation and streaming, according to exemplary embodiments.
[0040] Figure 12 is a media streaming process that segments data models of Figure 12 into network protocol packetized payloads according to exemplary embodiments.
[0041] Figure 13 is a sequence diagram of a network that adapts specific immersive media in ingest format to a streamable and suitable distribution format for a specific immersive media client endpoint according to exemplary embodiments.
[0042] Figure 14 is a logical flow diagram of an immersive media asset reuse analyzer according to exemplary embodiments. DETAILED DESCRIPTION
[0043] Definitions:
[0044] Scene graph: A general data structure commonly used by vector-based graphics editing applications and modern computer games that arranges the logical and often (but not necessarily) spatial representation of a graphics scene; a collection of nodes and vertices in a graph structure.
[0045] Scene: In the context of computer graphics, a scene is a collection of objects (e.g., 3D assets), object properties, and other metadata including visual, acoustic, and physics-based characteristics that describe a particular setting relative to the interaction of objects within that setting, bounded by space or time.
[0046] Node: A basic element of a scene graph, including information related to a logical or spatial or temporal representation of visual, audio, haptic, olfactory, gustatory, or related processing information; each node shall have at most one output edge, zero or more input edges, and at least one edge (input or output) connected to it.
[0047] Base layer: A nominal representation of an asset, often formulated to minimize the computational resources or time required to render the asset, or the time to transmit the asset over a network.
[0048] Augmentation layer: A set of information that, when applied to a base layer representation of an asset, enhances the base layer to include features or capabilities not supported in the base layer.
[0049] Property: Metadata associated with a node to describe a particular characteristic or feature of that node (e.g., in terms of another node) in a normative or more complex form.
[0050] Container: A serialized format that stores and exchanges information to represent all natural scenes, all synthetic scenes, or a mix of synthetic and natural scenes, including a scene graph and all media resources required to render the scene.
[0051] Serialization: The process of transforming the state of a data structure or object into a format that can be stored (e.g., in a file or storage buffer) or transmitted (e.g., over a network connection link) and later reconstructed (possibly in a different computer environment). When the resulting bit series is re-read according to the serialization format, it can be used to create a semantically identical clone of the original object.
[0052] Renderer: (Usually software-based) application or process that, given an input scene graph and asset container, typically emits a visual and / or audio signal suitable for rendering on a target device or conforming to desired properties specified by the attributes of the render target nodes in the scene graph, based on a selective mix of disciplines related to acoustic physics, light physics, visual perception, audio perception, mathematics, and software development. For visual-based media assets, the renderer can emit a visual signal suitable for a target display or suitable for storage as an intermediate asset (e.g., re-encapsulated into another container, i.e., used in a series of rendering processes in a graphics pipeline); for audio-based media assets, the renderer can emit an audio signal for rendering in a multi-channel loudspeaker and / or binauralized headphones, or for re-encapsulation into another (output) container. Popular examples of renderers include the real-time rendering features of the game engines Unity and Unreal Engine.
[0053] Evaluation: Produces a result (e.g., analogous to an evaluation of a document object model of a web page) that moves the output from the abstract to a concrete result.
[0054] Scripting language: An interpreted programming language that can be executed by a renderer at runtime to handle dynamic input and variable state changes made to scene graph nodes that affect the rendering and evaluation of spatial and temporal object topologies (including physics forces, constraints, inverse kinematics, deformations, collisions) and energy propagation and transport (light, sound).
[0055] Shader: A type of computer program that was originally used for shading (producing the right level of light, darkness, and color within an image), but now performs a variety of specialized functions in various areas of computer graphics effects, or does video post-processing unrelated to shading, or even functions completely unrelated to graphics.
[0056] Path tracing: A computer graphics method of rendering a three-dimensional scene so that the lighting of the scene is faithful to reality.
[0057] Timed media: Media that is ordered by time; for example, with start and end times according to a specific clock.
[0058] Timerless media: media organized in spatial, logical, or temporal relationships; for example, as in an interactive experience implemented in response to actions taken by one or more users.
[0059] Neural network model: a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations applied to visual signals to obtain improved visual outputs that can include interpolation of new views of the visual signals not explicitly provided by the original signals.
[0060] Over the past decade, many devices with immersive media capabilities have been introduced to the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and game consoles. Also, holographic displays and other forms of stereoscopic displays are ready to be introduced to the consumer market in the next three to five years. Despite the fact that these devices are currently or will soon be available, a coherent end-to-end ecosystem for distributing immersive media over commercial networks has failed to materialize for several reasons.
[0061] The descriptions herein, such as descriptions of ideal techniques and processes, should not be taken as acknowledgments that any of the existing technologies are prior art. Rather, these descriptions should be taken as disclosures of subject matter that is disclosed by the inventors and disclosed by this application. Unless otherwise specified, descriptions herein of technical problems and needs are also to be construed as having been recognized by the inventors and disclosed by this application.
[0062] One of the obstacles to realizing a coherent end-to-end ecosystem for distributing immersive media over commercial networks is that the client devices that serve as endpoints for such distribution networks for immersive displays are all very diverse. Some of them support certain immersive media formats, while others do not. Some of them are capable of creating immersive experiences from traditional raster-based formats, while others are not. Unlike networks designed for the distribution of traditional media only, a network that must support a variety of display clients requires a large amount of information related to the details of each of the client’s capabilities and the format of the media to be distributed before such a network can employ an adaptation process to transform the media into a format suitable for each target display and corresponding application. Such a network needs at least access to information describing the characteristics of each target display and the complexity of ingesting the media in order for the network to determine how to meaningfully adapt the input media source into a format suitable for the target display and application.
[0063] Similarly, an ideal network that supports heterogeneous clients should exploit the fact that some of the assets that are adapted from an input media format into a specific target format can be reused across a set of similar display targets. That is, some assets, once converted to a format suitable for a target display, can be reused across many such displays that have similar adaptation requirements. Thus, such an ideal network would employ a caching mechanism to store adapted assets into a relatively immutable area, i.e., similar to the use of a content distribution network (CDN) for traditional networks.
[0064] Further, immersive media can be organized into "scenes" described by a scene graph, also referred to as a scene description. The scope of a scene graph is to describe the visual, audio, and other forms of immersive assets that comprise a particular setting as part of a presentation, e.g., actors and events that occur in a particular location in a building as part of a presentation (e.g., a movie). A list that includes all scenes of a single presentation can be formalized as a manifest of scenes.
[0065] An additional benefit of such an approach is that, for content that is prepared in advance of having to distribute such content, a "bill of materials" can be created that identifies all assets that will be used for an entire presentation, as well as the frequency of use of each asset across various scenes within the presentation. An ideal network should be aware of the existence of cache resources that can be used to satisfy the asset requirements of a particular presentation. Similarly, a client that is rendering a series of scenes can want to be aware of the frequency of use of any given asset across multiple scenes. For example, if a media asset (also referred to as an object) is referenced multiple times across multiple scenes that the client is or will be processing, the client should refrain from discarding that asset from its cache resources until the last scene that requires that particular asset has been rendered by the client.
[0066] The disclosed subject matter addresses the need for a mechanism or process to analyze immersive media scenes for sufficient information that can be used to support a decision process that, when employed by a network or a client, provides an indication as to whether a transformation of a media object (or media asset) from format A to format B should be performed entirely by the network, entirely by the client, or via a mix of the two (along with an indication of which assets the client or network should transform). Such an "immersive media data complexity analyzer" can be employed by a client or network in an automated context, or by a human in a manual context.
[0067] Note that, without loss of generality, the remainder of the disclosed subject matter assumes that the process of adapting an input immersive media source to a particular endpoint client device is the same as or similar to the process of adapting the same input immersive media source to a particular application running on that particular client endpoint device. That is, the problem of adapting the characteristics of an input media source to an endpoint device has the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.
[0068] It should be noted that the terms media object and media asset are used interchangeably, both referring to a specific instance of media data in a particular format.
[0069] Figure 1 This is a schematic diagram of a media stream over a network that is distributed to clients. Figure 1 In this process, the processing of ingested media format A is performed by "cloud" or edge process 104. Note that the same processing can also be performed manually or by the client prior. Ingested media 101 is obtained from a content provider (not shown). Process 102 performs any necessary transformations or adjustments on the ingested media to create possible alternative representations of the media as distribution format B. Media formats A and B may or may not be representations that follow the same syntax as a particular media format specification; however, format B may be adapted to facilitate media distribution via network protocols such as TCP or UDP. Such "streamable" media is depicted in bitstream 105 as media to be streamed to client 108. Client 108 has access to some rendering capabilities depicted as process 106. Depending on the type of client 108 targeted, such rendering process 106 can be preliminary or equally complex. Rendering process 106 creates presentation media that may or may not be represented according to a third format specification (e.g., format C).
[0070] Figure 1-02 and Figure 1 The process is the same, but logic has been added to help the decision-making process determine whether a specific media object has been streamed to the client 108. Step 102A initiates a series of steps to aid the decision-making process. Conditional logic 102B accesses the list of unique assets used for presentation ( Figure 1-2 (Not depicted in Example 100-2) to determine whether the media object has been previously streamed to the client. If the media object has been previously streamed, an indicator 102C (hereinafter referred to as "alternative") is created to indicate that the client has received the specific media object and should use its local copy of the media object. If the media object has not been previously streamed, the process continues to step 103 to create the distribution format of the media object via step 102D.
[0071] Figure 2is a schematic illustration of a media stream through a network in which a media transform decision process 200 is employed to determine whether the network should transform the media before it is delivered to a client. In Figure 2 In the example 2000, ingest media 201 represented in format A is provided to the network by a content provider (not depicted). Process 202 acquires attributes that describe the processing capabilities of the target client (not depicted). Decision process 203 is employed to determine whether the network or the client should perform any format conversion on any of the media assets contained within ingest media 201 before streaming the media to the client, e.g., converting a particular media object from format A to format B. If any of the media assets should be transformed by the network, the network employs process 204 to transform the media object from format A to format B. Transformed media 205 is the output from process 204. The transformed media is merged into preparation process 206 to prepare the media to be streamed to the client (not shown). Process 207 streams the media to the client.
[0072] Figure 2-03 is a schematic illustration of a media transform decision process 2030 with asset reuse logic. A media stream through a network employs two decision processes to determine whether the network should transform the media before it is delivered to a client. In Figure 2-03 In the example 20300, ingest media 2031 represented in format A is provided to the network by a content provider (not depicted). Process 2032 acquires attributes that describe the processing capabilities of the target client (not depicted). Decision process 2033 is employed to determine whether the network has previously streamed a particular media object to the client. If the media object has been previously streamed to the client, step 2034 is employed to supersede the media to indicate that the client should use its local copy of the object that was previously streamed. If the media has not been previously streamed, decision process 2035 is employed to determine whether the network or the client should perform any format conversion on any of the media assets contained within ingest media 2031 before streaming the media to the client, e.g., converting a particular media object from format A to format B. If any of the media assets should be transformed by the network, the network employs process 2038 to transform the media object from format A to format B. Transformed media 2039 is the output from process 2038. The transformed media is merged into preparation process 2036 to prepare the media to be streamed to the client (not shown). Process 2037 streams the media to the client.
[0073] Figure 2-033 depicts an example 20330 in which a media transform decision process 20330 with client query in asset reuse logic is shown. A media stream through a network employs three decision processes to determine whether the network should transform the media before it is delivered to a client. In Figure 2-033In this example, ingest media 20331 represented in format A is provided to the network by a content provider (not depicted). Process 20332 obtains attributes that describe the processing capabilities of a target client (not depicted). Decision process 20333 is employed to determine whether the network has previously streamed a particular media object to the client. If the media object has been previously streamed to the client, decision process 20334 is employed to query the client to determine whether the client still has access to the previously streamed asset. If the client still has access to the asset, step 203310 is employed to supersede the media with an alternative to indicate that the client should use its local copy of the previously streamed object. If the media has not been previously streamed, or if the client no longer has a copy of the previously streamed asset, decision process 20335 is employed to determine whether the network or the client should perform any format conversion on any of the media assets contained within ingest media 20331, e.g., converting a particular media object from format A to format B, before the media is streamed to the client. If any of the media assets should be transformed by the network, the network employs process 20338 to transform the media object from format A to format B. Transformed media 20339 is the output from process 20338. The transformed media is merged into preparation process 20336 to prepare the media to be streamed to the client (not shown). Process 20337 streams the media to the client.
[0074] Figure 3 is an example representation of a streamable format of timed heterogeneous immersive media, i.e., timed media representation 300. Figure 4 is an example representation of a streamable format of untimed heterogeneous immersive media, i.e., untimed media representation 400. Both figures relate to a scene; Figure 3 relates to a scene 301 of timed media, and Figure 4 relates to a scene 401 of untimed media. For both cases, the scene can be embodied by various scene representations or scene descriptions.
[0075] For example, in some immersive media designs, a scene can be embodied by a scene graph, or embodied as a multi-planar image (MPI), or embodied as a multi-spherical image (MSI). Both MPI and MSI technologies are examples of techniques that help create display-agnostic scene representations for natural content (i.e., images of the real world captured simultaneously from one or more cameras). On the other hand, scene graph technology can be employed to represent both natural images and computer-generated images in the form of synthetic representations, however, creating such representations is extra computationally intensive for the case when the content is captured by one or more cameras as natural scenes. That is, creating a scene graph representation of natural captured content is time and computationally intensive, requiring complex analysis of natural images using photogrammetry or deep learning or both techniques in order to create a synthetic representation that can later be used to interpolate a sufficient and sufficient number of views to fill the frustum of views of a target immersive client display. Therefore, it is currently impractical to consider such synthetic representations as candidates for representing natural content, as they cannot actually be created in real-time given the use case of real-time distribution. However, currently, the best candidate representation for computer-generated images is the use of scene graphs with synthetic models, as computer-generated images are created using 3D modeling processes and tools.
[0076] This dichotomy in the best representation of both natural content and computer-generated content indicates that the best ingest format for natural captured content is different than the best ingest format for computer-generated content or natural content that is not important for real-time distribution applications. Therefore, it is a goal of the disclosed subject matter to be robust enough to support multiple ingest formats for visual immersive media, whether they are created naturally through the use of physical cameras or created by computers.
[0077] The following are example techniques to embody a scene graph into a format suitable for representing visual immersive media, whether created using computer-generated techniques or natural captured content for which deep learning or photogrammetry techniques are employed to create a corresponding synthetic representation of the natural scene, i.e., that is not important for real-time distribution applications.
[0078] 1. OTOY’s
[0079] OTOY’s ORBX is one of several scene graph technologies that can support any type of timed or untimed visual media, including ray-traceable, traditional (frame-based), stereoscopic, and other types of synthetic or vector-based visual formats. ORBX is different from other scene graphs in that ORBX provides native support for meshes, point clouds, and textures to free and / or open source formats. ORBX is a scene graph designed with the goal of facilitating interchange across multiple vendor technologies operating on scene graphs. In addition, ORBX provides a rich material system, support for the Open Shading Language, a robust camera system, and support for Lua scripting. ORBX is also the basis for the Immersive Technology Media Format licensed by the Immersive Digital Experiences Alliance (IDEA) under royalty-free terms. In the context of real-time distribution of media, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computing resources to perform complex analysis on camera-captured data and to synthesize the same data into synthetic representations. So far, it has been impractical to provide sufficient computing for real-time distribution, but not impossible.
[0080] 2. Pixar’s Universal Scene Description
[0081] Pixar’s Universal Scene Description (USD) is another well-known and mature scene graph that is popular in the VFX and professional content production community. USD is integrated into Nvidia’s Omniverse platform, a set of tools that developers use to create and render 3D models with Nvidia’s GPUs. A subset of USD is published by Apple and Pixar as USDZ. USDZ is supported by Apple’s ARKit.
[0082] 3. Khronos’ glTF 2.0
[0083] glTF 2.0 is the latest version of the “Graphics Language Transmission Format” specification written by the Khronos 3D Group. The format supports a simple scene graph format that can generally support static (untimed) objects in a scene, including “png” and “jpeg” image formats. glTF 2.0 supports simple animations, including support for translation, rotation, and scaling of primitive shapes (i.e., geometric objects) described using glTF primitives. glTF 2.0 does not support timed media, and therefore neither supports video nor audio.
[0084] These known designs for scene representations of immersive visual media are provided by way of example only and do not limit the ability of the disclosed subject matter specifying process to adapt an input immersive media source into a format suitable for the bit characteristics of a client endpoint device.
[0085] Further, any or all of the above example media representations can currently employ or can be adapted to employ deep learning techniques to train and create neural network models that allow or facilitate selection of particular views to fill the viewing frustum of a particular display based on the particular dimensions of the frustum. The views selected for the viewing frustum of a particular display can be interpolated from existing views explicitly provided in the scene representation (e.g., from MSI or MPI techniques), or they can be directly rendered from the rendering engine based on the particular virtual camera locations, filters, or descriptions of the virtual cameras of these rendering engines.
[0086] Accordingly, the disclosed subject matter is sufficiently robust to account for the existence of a relatively small but well-known set of immersive media ingest formats that are fully capable of meeting the requirements for real-time or “on-demand” (e.g., non-real-time) distribution of media that is either naturally captured (e.g., with one or more cameras) or created using computer-generated techniques.
[0087] With the deployment of advanced network technologies such as 5G for mobile networks and fiber cables for fixed networks, interpolation of views from immersive media ingest formats using neural network models or network-based rendering engines is further facilitated. That is, these advanced network technologies increase the capacity and capabilities of commercial networks as such advanced network infrastructure can support the carriage and delivery of an increasing amount of visual information. Network infrastructure management technologies such as Multi-Access Edge Computing (MEC), Software Defined Networks (SDN), and Network Function Virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changes in demand for certain network resources, for example, in response to dynamic increases or decreases in demand for network throughput, network speed, round-trip latency, and computing resources. Moreover, this inherent ability to adapt to dynamic network requirements likewise facilitates the ability of the network to adapt immersive media ingest formats into suitable distribution formats in order to support a variety of immersive media applications with possible heterogeneous visual media formats targeted to heterogeneous client endpoints.
[0088] Immersive media applications themselves can also have varying requirements on network resources, including gaming applications that require significantly lower network latency in response to real-time updates of game state, telepresence applications that have symmetric throughput requirements on both the uplink and downlink portions of the network, and passive viewing applications that can have increased demand on downlink resources depending on the type of client endpoint display that is consuming the data. In general, any consumer-facing application can be supported by a variety of client endpoints with various on-board client capabilities for storage, computation, and power supply, and likewise with various requirements for particular media representations.
[0089] Accordingly, the disclosed subject matter enables a fully equipped network (i.e., a network that employs some or all of the features of modern networks) to support both multiple legacy devices and immersive media capable devices simultaneously according to the features specified within the following conditions:
[0090] 1. Provide flexibility to utilize media ingest formats that are practical for both real-time and "on-demand" use cases of the distribution of media.
[0091] 2. Provide flexibility to support both natural content and computer generated content for both legacy client endpoints and immersive media capable client endpoints.
[0092] 3. Support both timed media and untimed media.
[0093] 4. Provide a process for dynamically adapting source media ingest formats into appropriate distribution formats based on the features and capabilities of the client endpoints and based on the requirements of the application.
[0094] 5. Ensure that the distribution formats are streamable over IP based networks.
[0095] 6. Enable a network to serve multiple heterogeneous client endpoints simultaneously, which can include both legacy devices and immersive media capable devices.
[0096] 7. Provide an exemplary media representation framework that facilitates organizing distributed media along scene boundaries.
[0097] The improved end-to-end embodiments enabled by the disclosed subject matter are implemented in accordance with the processes and components described in the detailed description of Figure 3 to Figure 1 6 below.
[0098] Figure 3 and Figure 4 Both employ a single exemplary encompassing distribution format that has been adapted from an ingest source format to match the capabilities of a particular client endpoint. As noted above, Figure 3 The media shown is timed, and Figure 4 The media shown is untimed. The particular encompassing format is robust in its structure to accommodate a variety of media attributes, each of which can be layered based on the amount of salient information each layer contributes to the presentation of the media. Note that such a layering process is already a well known technology in the current state of the art, as demonstrated by the use of progressive JPEG and the scalable video architecture as specified in ISO / IEC 14496-10 (Scalable Video Coding).
[0099] 1. According to the encompassing media format being streamed, the media is not limited to traditional visual and audio media, but can include any type of media information that can produce a signal that interacts with a machine to stimulate the human senses of sight, sound, taste, touch, and smell.
[0100] 2. According to the encompassing media format being streamed, the media can be timed media or untimed media, or a mixture of both.
[0101] 3. In addition, the layered representation of the media objects is achieved through the use of a base layer and enhancement layer architecture, which is streamable. In one example, the separate base and enhancement layers are computed by applying a multi-resolution or multi-tessellation analysis technique to the media objects in each scene. This is similar to the progressive rendering image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG 2000), but is not limited to raster-based visual formats. In an example embodiment, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using a wavelet analysis.
[0102] In another example of the layered representation of the media format, the enhancement layer applies different attributes to the base layer, such as modifying the material properties of the surface of a visual object represented by the base layer. In yet another example, the attributes can modify the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a shiny surface.
[0103] In yet another example of the layered representation, the surface of one or more visual objects in the scene can change from a Lambertian surface to a light-traceable surface.
[0104] In yet another example of the layered representation, the network distributes the base layer representation to the client so that the client can create a nominal rendering of the scene while the client waits for the transmission of additional enhancement layers to modify the resolution or other characteristics of the base representation.
[0105] 4. The resolution of the attributes or modification information in the enhancement layer is not explicitly coupled to the resolution of the objects in the base layer as it is today in existing MPEG video and JPEG image standards.
[0106] 5. The encompassing media format support can encompass any type of information media that can be rendered or driven by a rendering device or machine, thus enabling support for heterogeneous media formats of heterogeneous client endpoints. In one embodiment of a network that distributes media formats, the network will first query the client endpoint to determine the client's capabilities, and if the client cannot ingest the media representation meaningfully, the network will remove layers of attributes that the client does not support, or adapt the media from its current format to a format suitable for the client endpoint. In one example of such adaptation, the network will convert a stereoscopic visual media asset to a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can employ neural network processing to reformat the media to an appropriate format or alternatively synthesize the view required by the client endpoint.
[0107] 6. The manifest of a complete or partial immersive experience (live event, game, or playback of an on-demand asset) is organized by scenes, which is the minimum amount of information that rendering and game engines can ingest in order to create a presentation. The manifest includes a list of individual scenes to be rendered for the overall immersive experience requested by the client. Associated with each scene is one or more representations of the geometry objects within the scene that correspond to a streamable version of the scene geometry. One embodiment of a scene representation involves a low resolution version of the geometry objects of the scene. Another embodiment of the same scene involves an enhancement layer of the low resolution representation of the scene to add additional detail to the geometry objects of the same scene or increase tessellation. As noted above, each scene can have more than one enhancement layer in order to increase the detail of the geometry objects of the scene in a progressive manner.
[0108] 7. Each layer of a media object referenced within a scene is associated with a token (e.g., URI) that points to an address within the network where the resource can be accessed. Such resources are similar to a CDN where content can be ingested by a client.
[0109] 8. The token for a representation of a geometry object can point to a location within the network or to a location within the client. That is, the client can signal to the network that its resources are available to the network for network-based media processing.
[0110] In the drawings described below, like elements can be indicated by the same reference numbers throughout the several drawings. Where such elements are referenced as "a," "an" or "the" element, this is intended to mean that there are one or more of these elements present.
[0111] Figure 3Embodiments of a timed media encompassing media format are described as follows. A timed scene manifest includes a list of scene information 301. A scene 301 refers to a list of components 302 that respectively describe processing information and type of media assets that comprise the scene 301. A component 302 refers to an asset 303 that further refers to a base layer 304 and an attribute enhancement layer 305. A list of unique assets not previously used in other scenes is provided in 307.
[0112] Figure 4 Embodiments of a non-timed media encompassing media format are described as follows. Scene information 401 is not associated with start and end duration according to a clock. Scene information 401 refers to a list of components 402 that respectively describe processing information and type of media assets that comprise the scene 401. A component 402 refers to an asset 403 that further refers to a base layer 404 and attribute enhancement layers 405 and 406. In addition, a scene 401 refers to other scenes 401 for non-timed media. A scene 401 also refers to a scene 407 for timed media scenes. A list 406 identifies unique assets associated with a particular scene that have not been previously used in a higher order (e.g., parent) scene.
[0113] Figure 5 Embodiments of a process 500 to synthesize an ingest format from natural content are illustrated. Camera unit 501 uses a single camera lens to capture a scene of a person. Camera unit 502 captures a scene with five diverging fields of view by mounting five camera lenses around a ring-shaped object. The arrangement in 502 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter portion of a sphere. The arrangement 503 is an exemplary arrangement commonly used to capture light fields for light field or holographic immersive displays. The natural image content 509 is provided as input to a synthesis process 504 that can alternatively employ a neural network training process 505 that uses a set of training images 506 to produce an alternative capture neural network model 508. Another process commonly used in lieu of the training process 505 is photogrammetry. If the model 508 is created during the depicted process 500, the model 508 becomes one of the assets in the ingest format 507 of natural content. Exemplary embodiments of the ingest format 507 include MPI and MSI. Figure 5 If the model 508 is created during the depicted process 500, the model 508 becomes one of the assets in the ingest format 507 of natural content. Exemplary embodiments of the ingest format 507 include MPI and MSI.
[0114] Figure 6An embodiment of a process 600 to create ingest formats for synthetic media (e.g., computer generated imagery) is illustrated. A LIDAR camera 601 captures a point cloud 602 of a scene. CGI tools to create synthetic content, 3D modeling tools, or another animation process are employed on a computer 603 to create 604 CGI assets on a network. A motion capture suit with sensors 605A is worn on an actor 605 to capture a digital record of the actor’s 605 motion to produce animation motion capture (MoCap) data 606. The data 602, 604, and 606 are provided as input to a synthetic process 607, which can likewise create a neural network model (not depicted in Figure 6
[0115] The above-described techniques for representing and streaming heterogeneous immersive media can be implemented using computer readable instructions, which can be physical stored in one or more computer readable media. For example, Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0116] Computer software can be coded using any suitable machine code or computer language that can be subject to assembly, compilation, linking, or similar
[0117] Instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, internet of things devices, and the like.
[0118] Figure 7 The components shown for computer system 700 are exemplary and not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of this disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of a computer system 700.
[0119] Computer system 700 can include certain human interface input devices. Such a human interface input device can be responsive to user input. The various human interface input devices can respond to different kinds of input. For example, a keyboard can respond to the keystrokes of a user, a mouse can respond to gestures by a user, a touch screen can respond to gestures by a user, and a microphone can respond to the user's voice. Such a human interface input device can be cross- device sensitive. For example, a gesture by the user on a mouse can be understood by the computer system 700 as equivalent to the gesture by the user on a touch screen.
[0120] The input human interface devices can include one or more of (only one of each is depicted): keyboard 701, mouse 702, trackpad 703, touch screen 710, data
[0121] Computer system 700 can also include certain human interface output devices. Such human interface output devices can be stimulating human sense such as, for example, tactile sense, auditory sense, and visual sense. A human interface output device can include a tactile output device, an audio output device, a visual output device, and a somatosensory (haptic) output device. The tactile output device can be, for example, a vibration motor arranged to provide tactile feedback to a user. The audio output device can be, for example, a speaker arranged to produce sound waves. The visual output device can be, for example, a screen for displaying visual information to the user. The somatosensory (haptic) output device can be, for example, a device arranged to deliver a tactile sensation to the user, such as a touch screen, a data glove, or a joystick.
[0122] Computer system 700 can also include human accessible storage devices and their associated media such as optical media including CD / DVD ROM / RW 720 with CD / DVD 721, thumb-drive 722, removable hard drive or solid state drive 723, legacy magnetic media such as tape and floppy disk (not depicted), specialized ROM / ASIC / PLD memory such as security dongles (not depicted), and the like.
[0123] Those skilled in the art will further appreciate that the term "computer-readable medium" as used herein is intended to include a transmission medium (wireless or wired), such as those that carry computer instructions, data structures, program modules, or other data over the Internet, over a network, or through a broadcast transmission. It is not intended that the computer-readable medium be limited to any single medium or medium set, as the term is intended to include both volatile and non-volatile media, both removable and non-removable media.
[0124] Computer system 700 can also include an interface to one or more communication networks. Networks can for example be wireless, wireline, optical. Networks can further be local, wide-area, metropolitan, vehicular and industrial, real-time, delay-tolerant, and so on. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks to include GSM, 3G, 4G, 5G, LTE and the like, TV wireline or wireless wide area digital networks to include cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial to include CANBus, and so forth. Certain networks commonly require external network interface adapters that attached to certain general purpose input / output ports or peripheral buses (749) (such as USB ports of the computer system 700; others are integrated into the core of the computer system 700 by attachment to a system bus as described below (for example, Ethernet adapters into PC computer systems or cellular network adapters into smartphone computer systems). Using any of these networks, computer system 700 can communicate with other entities. Such communication can be uni-directional, receive only (for example, broadcast TV), uni-directional send-only (for example, CANbus to certain CANbus devices), or bi-directional, for example to other computer systems using local or wide area digital networks. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.
[0125] The human interface devices, human-accessible storage devices, and network interfaces described above can be attached to the core 740 of the computer system 700.
[0126] The core 740 can include one or more Central Processing Units (CPU) 741, Graphics Processing Units (GPU) 742, specialized programmable processing units in the form of Field Programmable Gate Areas (FPGA) 743, hardware accelerators for certain tasks 744, and so forth. These devices, along with Read-only memory (ROM) 745, Random-access memory (RAM) 746, internal mass storage such as internal non-user accessible hard drives, SSDs, and the like, can be connected through a system bus 748. In some computer systems, the system bus 748 can be accessible as a non-user accessible internal bus through a physical connector to the computer system 700, such as a system in a chip (SiD connector, and the like. Peripheral devices can be attached either directly to the core’s system bus 748, or through a peripheral bus 749. The peripheral bus architecture can be any of a variety of bus architectures, including PCI, USB, and the like.
[0127] CPUs 741, GPUs 742, FPGAs 743, and accelerators 744 can execute certain instructions that, in combination, can make up the aforementioned computer code. That computer code can be stored in ROM 745 or RAM 746. Transitional data can be also be stored in RAM 746, whereas permanent data can be stored for example, in the internal mass storage 747. Fast storage and retrieval speeds of either of the memory devices can be enabled through the use of cache memory, that can be closely associated with one or more CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, and the like.
[0128] The computer software can be implemented as computer code that is executable by one or more of the processors 741, GPU 742, FPGA 743, and accelerators 744, in conjunction with one or more computer readable media storing computer code thereon.
[0129] By way of example, and not limitation, the computer system 700, and specifically the core 740, can provide functionality as a result of processor(s) executing software embodied in one or more tangible, computer-readable media, such as one or more of the mass storage devices, memory, and the like. In one example, the core 740, and more specifically the processor(s) therein, provide functionality as a result of the software implementing known techniques such as the techniques disclosed herein. In one example, the software is executed upon the core 740, and more specifically one or more of the processors 741, GPU 742, FPGA 743, and accelerators 744, to transform the core 740, and more specifically one or more of the processors 741, GPU 742, FPGA 743, and accelerators 744, to thereby provide functionality.
[0130] Figure 8An example network media distribution system 800 is illustrated that supports a variety of legacy displays and heterogeneous immersive media capable displays as client end points. A content acquisition process 801 uses example embodiments in Figure 6 or Figure 5 to capture or create media. The ingest format is created in a content preparation process 802 and then transmitted to the network media distribution system using a transport process 803. Gateways 804 can serve customer premises equipment to provide network access to various client end points of the network. Set top boxes 805 can also serve as customer premises equipment to provide network service provider access to aggregated content. Wireless modems 806 can serve as mobile network access points for mobile devices, for example, as shown by mobile handset and display 813. In this particular embodiment of system 800, a legacy 2D television 807 is shown connected directly to a gateway 804, set top box 805, or WiFi router 808. A laptop computer 809 with a legacy 2D display is illustrated as a client end point connected to WiFi router 808. A head mounted 2D (raster based) display 810 is also connected to router 808. A lenticular light field display 811 is shown connected to gateway 804. Display 811 includes a local compute GPU 811A, storage device 811B, and a visual rendering unit 811C that creates multiple views using ray based lenticular optics. A holographic display 812 is shown connected to set top box 805. Display 812 includes a local compute CPU 812A, GPU 812B, storage device 812C, and a Fresnel pattern, wave based holographic visualization unit 812D. An augmented reality headset 814 is shown connected to wireless modem 806. Headset 814 includes a GPU 814A, storage device 814B, battery 814C, and stereoscopic visual rendering components 814D. A dense light field display 815 is shown connected to WiFi router 808. Display 815 includes multiple GPUs 815A, CPU 815B, storage device 815C, eye tracking device 815D, camera 815E, and dense ray based light field panel 815F.
[0131] Figure 9 An embodiment of an immersive media distribution process 900 is illustrated that is capable of serving legacy displays and heterogeneous immersive media capable displays as previously depicted in Figure 8 . Content is created or acquired in process 901, which is further embodied for natural content and CGI content in Figure 5 and Figure 6 respectively. Content 901 is then converted to an ingest format using a create network ingest format process 902. Process 902 is likewise embodied in Figure 5 and Figure 6Further embodiments are embodied with respect to natural content and CGI content. The ingest media is alternatively updated to store information from the media reuse analyzer 911 about assets that can be reused across multiple scenes. The ingest media format is transmitted to the network and stored on storage device 903. Alternatively, the storage device can reside in the network of immersive media content producers and be accessed remotely by the immersive media network distribution process (not numbered), as depicted by the dashed lines through 903. Client and application specific information is alternatively available on remote storage device 904, which can alternatively exist remotely in an alternate "cloud" network.
[0132] As Figure 9 depicted, the network orchestration process 905 serves as the primary source and sink of information for performing the primary tasks of the distribution network. In this particular embodiment, the process 905 can be implemented in a format that is uniform with the other components of the network. However, Figure 9 The tasks depicted in process 905 form an important element of the disclosed subject matter. The orchestration process 905 can further employ a two-way messaging protocol with the clients to facilitate all processing and distribution of media according to the characteristics of the clients. Further, the two-way protocol can be implemented across different delivery channels (i.e., control plane channels and data plane channels).
[0133] The process 905 receives information about the features and attributes of the client 908 and also collects requirements about the applications that are currently running on 908. This information can be obtained from the device 904 or, in alternate embodiments, can be obtained by directly querying the client 908. In the case of direct querying of the client 908, it is assumed that a two-way protocol (not shown in Figure 9 ) exists and is operational so that the client can directly communicate with the orchestration process 905.
[0134] The orchestration process 905 also initiates and communicates with the media adaptation process 910 described in Figure 10 When the ingest media is adapted and segmented by the process 910, the media is alternatively transferred to an intermediate storage device, depicted as media 909 ready for distribution storage device. When the distribution media is prepared and stored in the device 909, the orchestration process 905 ensures that the immersive client 908 receives the distribution media and corresponding descriptive information 906 via its network interface 908B by "push" requests, or the client 908 itself can initiate requests to "pull" the media 906 from the storage device 909. The orchestration process 905 can employ a two-way messaging interface Figure 9The immersive client 908 can alternatively employ a GPU (or CPU not shown) 908C. The distribution format of the media is stored in a storage device or storage cache 908D of the client 908. Finally, the client 908 visually presents the media via its visualization component 908A.
[0135] Throughout the process of streaming immersive media to the client 908, the orchestration process 905 will detect the status of the progress of the client via the client progress and status feedback channel 907. The detection of the status can be performed through a bidirectional communication message interface (not shown in Figure 9
[0136] Figure 10 A particular embodiment of a media adaptation process is depicted, such that ingest media can be properly adapted to match the requirements of the client 908. The media adaptation process 1001 controlled by one or more processors includes a number of components that facilitate the adaptation of ingest media into the proper distribution format for the client 908. These components should be considered exemplary. In Figure 10 In particular, the adaptation process 1001 receives input network status 1005 to track the current traffic load on the network; client 908 information including attribute and feature descriptions, application characteristics and descriptions, and application current state, and client neural network model (if available) to help map the client's frustum to the interpolation capabilities of the ingest immersive media. Such information can be obtained through a bidirectional message interface (not shown in Figure 10 The adaptation process 1001 ensures that the adapted output is stored into the client adapted media storage device 1006 as it is created. A media reuse analyzer 1007 is depicted in Figure 10 In particular, the adaptation process 1001 receives input network status 1005 to track the current traffic load on the network; client 908 information including attribute and feature descriptions, application characteristics and descriptions, and application current state, and client neural network model (if available) to help map the client's frustum to the interpolation capabilities of the ingest immersive media. Such information can be obtained through a bidirectional message interface (not shown in
[0137] The adaptation process 1001 is controlled by a logic controller 1001F. The adaptation process 1001 also employs a renderer 1001B or neural network processor 1001C to adapt the particular ingest source media into a format suitable for the client. The neural network processor 1001C uses the neural network model in 1001A. An example of such a neural network processor 1001C is the Deepview neural network model generator described in the MPI and MSI. If the media is in 2D format but the client must have a 3D format, the neural network processor 1001C can invoke a process to use highly correlated images from the 2D video signal to derive a stereoscopic representation of the scene depicted in the video. An example of a suitable renderer 1001B can be a modified version of the OTOY Octane renderer (not shown) that is modified to directly interact with the adaptation process 1001. The adaptation process 1001 can optionally employ a media compressor 1001D and media decompressor 1001E depending on the need for these tools relative to the format of the ingest media and the format required by the client 908.
[0138] Figure 11 A distribution format creation process 1100 is depicted. An adapted media packaging process 1103 packages the media from the media adaptation process 1101 (depicted as process 1000 in Figure 10 ) that now resides on the client adapted media storage device 1102. The packaging process 1103 formats the adapted media from process 1101 into a robust distribution format 1104, such as the example format shown. Figure 3 or Figure 4 The manifest information 1104A provides the client 908 with a list 1104B of the scene data assets that it can expect to receive as well as optional complexity metadata that describes the complexity of all of the assets of the scene. The list 1104B depicts a list of visual assets, audio assets, and haptic assets, each with their corresponding metadata.
[0139] Figure 12 A packager processing system 1200 is depicted. The packager process 1202 separates the adapted media 1201 into individual packages 1203 suitable for streaming to the client 908.
[0140] Figure 13The components and communications of the illustrated sequence diagram 1300 are explained as follows: A client endpoint 1301 initiates a media request 1308 to a network distribution interface 1302. The request 1308 includes information identifying the media requested by the client by URN or other standard nomenclature. The network distribution interface (also referred to as the client 1302) responds to the request 1308 with a profile request 1309 requesting that the client 1301 provide information about its currently available resources (including compute, storage, battery charge percentage, and other information characterizing the current operating state of the client). The profile request 1309 also requests that the client provide one or more neural network models, if such models are available at the client, that can be used by neural network inference to extract or interpolate the correct media view to match the characteristics of the client’s presentation system. A response 1310 from the client 1301 to the interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). The interface 1302 then provides a session ID token 1311 to the client 1301. The interface 1302 then requests an ingest media server 1303 by an ingest media request 1312 that includes the URN or other standard name of the media identified in the request 1308. In response to the request 1312, the server 1303 returns a response 1313 that includes an ingest media token. The interface 1302 then provides the media token from the response 1313 to the client 1301 in a call 1314. Then, in 1315, the interface 1302 initiates an adaptation process for the requested media by providing the ingest media token, the client token, the application token, and the neural network model token to an adaptation interface 1304. The interface 1304 requests access to the ingest media by providing the ingest media token to the server 1303 at a call 1316, thereby requesting access to the ingest media. The server 1303 responds to the request 1316 with an ingest media access token in a response 1317 to the interface 1304. The interface 1304 then requests a media adaptation process 1305 to adapt the ingest media located at the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created at 1313. The request 1318 from the interface 1304 to the process 1305 contains the required tokens and session ID. The process 1305 provides an adapted media access token and session ID to the interface 1302 in an update 1319. The interface 1302 provides the adapted media access token and session ID to a packaging process 1306 in an interface call 1320. The packaging process 1306 provides a response 1321 to the interface 1302 with a packaged media access token and session ID in the response 1321. The process 1306 provides the packaged asset for the session ID, URN, and packaged media access token to a packaged media server 1307 in a response 1322.Client 1301 executes request 1323 to initiate streaming of the media asset corresponding to the packaged media access token received in message 1321. Client 1301 executes other requests and provides a status update to interface 1302 in message 1324.
[0141] Figure 14 A media reuse analyzer 1400 logic flow of an immersive media data reuse optimizer depicted as media reuse analyzer 911 in Figure 9 The initialization step 1402 initializes an iterator "i" to zero and further initializes a set of lists 1404 (one list per scene) that identify the unique assets encountered in all scenes comprising the presentation as described in Figure 3 or Figure 4 The lists 1404 depict a sample list entry of information describing an asset that is unique relative to the entire presentation, including an indicator for the type of media comprising the asset (e.g., Mesh, Audio, or Volume), a unique identifier for the asset, and the number of times the asset is used in the set of scenes comprising the presentation. As an example, for scene N-1, the asset is not included in its list because all of the assets required for scene N-1 have already been identified as assets that are also used in scenes 1 and 2. Step 1403 determines whether the iterator "i" is less than the total number of scenes comprising the presentation (as described in Figure 3 or Figure 4The reuse analysis terminates at step 1405 if the iterator "i" is equal to the number N of scenes presented. Otherwise, if the iterator "i" is less than the total number of scenes, processing continues to step 1406 where the iterator "j" is set to zero. Step 1407 tests the iterator "j" to determine if it is less than the total number of media assets (also referred to as media objects) in the current scene "i". If the iterator "j" is less than the total number of media assets of scene "i", processing continues to step 1408. Otherwise, processing continues to step 1412 where the iterator "i" is incremented by one before returning to step 1403. If the value of "j" is less than the total number of assets of scene "i", processing continues to conditional step 1408 where the characteristics of the media asset are compared to assets previously analyzed from scenes prior to the current scene "i". If the asset has been identified as an asset used in a scene prior to scene "i", then in step 1411 the number of times the asset is used across scenes 0 to N-1 is incremented by one. Otherwise, if the asset is a unique asset, i.e., it has not been previously analyzed in a scene associated with a smaller value of the iterator "i", then in step 1409 a unique asset entry is created for scene "i" in the list 1404. Step 1409 also creates a unique identifier and assigns it to the entry for the asset, and the number of times the asset is used across scenes 0 to N-1 is set to one. After step 1409, processing continues to step 1410 where the iterator "j" is incremented by one. After step 1410, processing returns to step 1407.
[0142] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various take equivalents, which fall within the scope of the disclosure. It should be understood that various of the steps within the foregoing processes can be performed in the reverse order, or can be performed concurrently, that the modulus of the steps can be performed in one or more stages (either consecutively, or with breaks in between stages), and that one or more steps can be skipped altogether. Accordingly, any sequence or order of steps that can be performed in alternate orders or concurrently is performed are properly fall within the scope of the disclosure. Additionally, the description of a particular feature or aspect of the disclosure should not be construed as an abandonment of that feature or aspect or the disclosure.
Claims
1. A method for immersive media presentation, the immersive media presentation including multiple scenes, the method being executed by at least one hardware processor, characterized in that, The method includes: The media asset is determined to appear in at least two or more of the plurality of scenarios, which are associated with the immersive media presentation; Send a request to the client to query whether the client can access the media assets that appear in at least two or more scenarios in a local cache, wherein the client has sufficient storage resources to store a copy of the media assets associated with the immersive media presentation in the local cache; Receive a response from the client, the response indicating whether the client is able to access the media asset that appears in at least two or more scenarios in the local cache; In response to the reply indicating that the client can access the media asset appearing in at least two or more scenarios in the local cache, the system signals the client to use the media asset in subsequent scenarios without further distributing the media asset to the client; and In response to the reply indicating that the client cannot access the media asset that appears in at least two or more scenarios in the local cache, the media asset is distributed to the client.
2. The method according to claim 1, characterized in that, Further includes: Initialize a set of lists, wherein each list corresponds to a scene among the plurality of scenes presented by the immersive media. Initializing the set of lists includes: incrementally assigning a unique identifier to each of the multiple media assets appearing in the multiple scenarios, wherein the multiple media assets include the media assets.
3. The method according to claim 2, characterized in that, Initializing the set of lists further includes: incrementally determining the number of times each media asset appears in each scene.
4. The method according to claim 1, characterized in that, Further includes: Receive a request for the immersive media presentation from the client; as well as In response to the request, the client is requested to provide an indication of the client's client resources.
5. The method according to claim 4, characterized in that, The instruction to request the client to provide the client resources includes: requesting the client to provide one or more neural network models, and The processing of the media assets includes neural network inference based on one or more neural network models requested from the client.
6. The method according to claim 5, characterized in that, The processing of the media assets is further based on determining the current business load on the network, which connects the at least one hardware processor and the client.
7. The method according to claim 1, characterized in that, Further includes: Detect the progress of the client in outputting the immersive media presentation. Specifically, the sending of the request is timed based on the progress.
8. The method according to claim 1, characterized in that, The immersive media presentation includes instructions to the client to stimulate visual, auditory, and at least one of the senses of taste, touch, and smell.
9. The method according to claim 1, characterized in that, The request sent to the client further queries whether the client can access the media asset, wherein the access is local to the client.
10. The method according to claim 1, characterized in that, The immersive media presentation includes either timed presentation or timeless presentation.
11. An immersive media presentation device, characterized in that, include: At least one memory is configured to store computer program code; At least one hardware processor is configured to access the computer program code and operate according to the instructions of the computer program code to implement the method as claimed in any one of claims 1-10.
12. A non-volatile computer-readable medium storing a program that causes a computer to perform the method as described in any one of claims 1-10.
Citation Information
Patent Citations
System and method for providing a virtual immersive environment
US20140192087A1
Deduplicated data transmission
US20210056109A1