Streaming Scene Prioritization for Immersive Media - Patent application
By prioritizing scenes and assets based on network and device capabilities, the method optimizes immersive media delivery, addressing bandwidth and device heterogeneity issues to improve playback quality.
Patent Information
- Application Number
- JP2024522107
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-04-21
- Filing Date
- 2023-05-02
- Publication Date
- 2025-09-24
- Estimated Expiration
- 2043-05-02
AI Technical Summary
Existing technologies face challenges in efficiently delivering immersive media, such as light field and holographic content, due to bandwidth limitations and heterogeneous client devices, leading to buffering and stuttering, especially when using light field-based displays.
A method for media processing that involves assigning priority values to scenes and assets based on factors like network bandwidth, computational complexity, and client device capabilities, and reordering streaming to optimize delivery.
Enhances the delivery of immersive media by reducing latency and improving playback quality on diverse client devices, ensuring smoother experiences even in bandwidth-constrained environments.
Smart Images

Figure 0007743623000001 
Figure 0007743623000002 
Figure 0007743623000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 18 / 137,849, entitled "Streaming Scene Prioritizer for Immersive Media," filed April 21, 2023, which is a continuation of U.S. Provisional Patent Application No. 63 / 341,191, entitled "Streaming Scene Prioritizer for Lightfield Holographic Media," filed May 12, 2022; U.S. Provisional Patent Application No. 63 / 342,532, entitled "Streaming Scene Prioritizer for Lightfield, Holographic, and / or Immersive Media," filed May 16, 2022; and U.S. Provisional Patent Application No. 63 / 342,526, entitled "Streaming Asset Prioritizer for Lightfield, Holographic, and / or Immersive Media," filed May 16, 2022. U.S. Provisional Patent Application No. 63 / 344,907, filed May 23, 2022, entitled "Streaming Asset Prioritizer for Lightfield, Holographic, or Immediate Media Based on Asset Visibility," U.S. Provisional Patent Application No. 63 / 354,071, filed June 21, 2022, entitled "Prioritizing Streaming Assets for Lightfield, Holographic, and / or Immediate Media Based on Multiple Policies," and U.S. Provisional Patent Application No. 63 / 355,768, filed June 27, 2022, entitled "Scene Analyzer for Prioritizing Asset Rendering Based on Asset Visibility in Scene Default" VIEWPORT,” and U.S. Provisional Patent Application No. 63 / 416,390, filed October 14, 2022, entitled “STREAMING ASSET PRIORITIZER FOR LIGHTFIELD,”"HOLOGRAPHIC OR IMMERSIVE MEDIA BASED ON DISTANCE WITHIN FIELD OF VIEW," U.S. Provisional Patent Application No. 63 / 416,395, filed October 14, 2022, entitled "STREAMING ASSET PRIORITIZER FOR LIGHTFIELD, HOLOGRAPHIC OR IMMERSIVE MEDIA BASED ON ASSET COMPLEXITY," U.S. Provisional Patent Application No. 63 / 422,175, filed November 3, 2022, entitled "STREAMING ASSET PRIORITIZER FOR LIGHTFIELD, HOLOGRAPHIC, OR IMMERSIVE MEDIA BASED ON MULTIPLE METADATA ATTRIBUTES," and U.S. Provisional Patent Application No. 63 / 428,698, filed November 29, 2022, entitled "PRIORITIZATION OF MEDIA ADAPTATION BY ATTACHED CLIENT" This application claims the benefit of priority to "Patent Citation 101 of the International Application No. 101001001, filed on May 1, 2009, entitled 'Patent Citation 101001001,' ...
[0002] This disclosure describes embodiments generally related to media processing and distribution, including adaptive streaming of immersive media, such as for light field or holographic immersive displays. [Background technology]
[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.
[0004] Immersive media generally refers to media that stimulates any or all human sensory systems (sight, hearing, somatosensation, smell, and sometimes taste) to create or enhance the user's perception of being physically present in the media experience, such as beyond that delivered over existing commercial networks for timed two-dimensional (2D) video and corresponding audio known as "legacy media." Both immersive and legacy media may be characterized as either timed or non-timed media.
[0005] Timed media refers to media that is structured and presented according to time. Examples include movie features, news reports, episodic content, etc., which are organized according to periods. Traditional video and audio are generally considered to be timed media.
[0006] Non-timed media is media that is not structured by time, but rather by logical, spatial, and / or temporal relationships. Examples include video games in which the user controls the experience created by a gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media may incorporate timed media, for example, in a continuously looped audio or video segment of a video game scene. Conversely, timed media may incorporate non-timed media, such as a video with a fixed still image as a background.
[0007] An immersive media-enabled device can refer to a device with the ability to access, interpret, and present immersive media. Such media and devices are heterogeneous with respect to the number and format of media and the number and type of network resources required to distribute such media at scale, i.e., to achieve distribution equivalent to legacy video and audio media over a network. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogeneous in their capabilities because all of these devices consist of rectangular display screens and consume 2D rectangular video or still images as their primary media format. Summary of the Invention [Means for solving the problem]
[0008] Aspects of the present disclosure provide methods and apparatus for media processing. According to some aspects of the present disclosure, a method of media processing includes receiving, by a network device, scene-based immersive media for playback on a light field-based display. The scene-based immersive media includes a plurality of scenes. The method includes assigning priority values to each of the plurality of scenes of the scene-based immersive media, and determining an order for streaming the plurality of scenes to an end device according to the priority values.
[0009] In some examples, the method includes re-ranking the plurality of scenes according to priority values, and transmitting the re-ranked plurality of scenes to the end device.
[0010] In some examples, the method includes selecting, by the priority-aware network device, a highest priority scene having a highest priority value from a subset of untransmitted scenes in the plurality of scenes, and transmitting the highest priority scene to the end device.
[0011] In some examples, the method includes determining a priority value for the scene based on the likelihood of a need to render the scene.
[0012] In some examples, the method includes determining that available network bandwidth is limited, selecting a highest priority scene from a subset of untransmitted scenes in the plurality of scenes having a highest priority value, and transmitting the highest priority scene in response to the available network bandwidth being limited.
[0013] In some examples, the method includes determining that available network bandwidth is limited, identifying a subset of the plurality of scenes that is unlikely to be required for rendering next based on priority values, and refraining from streaming the subset of the plurality of scenes in response to the available network bandwidth being limited.
[0014] In some examples, the method includes assigning a first priority value to the first scene based on a second priority value of the second scene in response to a relationship between the first scene and the second scene.
[0015] In some examples, the method includes receiving a feedback signal from the end device and adjusting a priority value of at least one scene among the plurality of scenes based on the feedback signal. In an example, the method includes assigning the highest priority to a first scene in response to the feedback signal indicating that the current scene is a second scene related to the first scene. In another example, the feedback signal indicates an adjustment of the priority determined by the end device.
[0016] According to some aspects of the present disclosure, a method of media processing includes receiving, by an end device having a light field-based display, a media presentation description (MPD) of scene-based immersive media for playback by the light field-based display. The scene-based immersive media includes multiple scenes, and the MPD indicates sequential streaming of the multiple scenes to the end device. The method further includes detecting bandwidth availability, determining a reordering of at least one scene based on the bandwidth availability, and transmitting a feedback signal indicating the reordering of the at least one scene.
[0017] In some examples, the feedback signal indicates a next scene to render, in some examples, the feedback signal indicates an adjustment to a priority value of at least one scene, in some examples, the feedback signal indicates a current scene.
[0018] According to some aspects of the present disclosure, a method of media processing includes receiving, by a network device, scene-based immersive media for playback on a light field-based display. The scene of the scene-based immersive media includes a plurality of assets in a first order. The method includes determining a second order for streaming the plurality of assets to an end device, the second order being different from the first order.
[0019] In some examples, the method includes assigning priority values to each of a plurality of assets of a scene according to one or more attributes of the plurality of assets, and determining a second order for streaming the plurality of assets according to the priority values.
[0020] In some examples, the method includes assigning priority values to assets of a scene according to the size of the assets.
[0021] In some examples, the method includes assigning priority values to assets of a scene according to the visibility of the assets associated with a default entry position of the scene.
[0022] In some examples, the method includes assigning a first set of priority values to a plurality of assets based on a first attribute of the plurality of assets. The first set of priority values is used to order the plurality of assets in a first prioritization scheme. The method further includes assigning a second set of priority values to the plurality of assets based on a second attribute of the plurality of assets. The second set of priority values is used to order the plurality of assets in a second prioritization scheme. The method further includes selecting a prioritization scheme for the end device from the first prioritization scheme and the second prioritization scheme according to information of the end device, and ordering the plurality of assets for streaming according to the selected prioritization scheme.
[0023] In some examples, the method includes assigning a first priority value to the first asset and a second priority value to the second asset in response to the first asset having a higher computational complexity than the second asset, the first priority value being higher than the second priority value.
[0024] According to some aspects of the present disclosure, a method for media processing includes determining a position and size of a viewport associated with a camera for viewing a scene of scene-based immersive media. The scene includes a plurality of assets. The method further includes, for each asset, determining whether the asset at least partially intersects with the viewport based on the position and size of the viewport and a geometry of the asset, and, in response to the asset at least partially intersecting with the viewport, storing an identifier of the asset in a list of viewable assets and including the list of viewable assets in metadata associated with the scene.
[0025] According to some aspects of the present disclosure, a method of media processing includes receiving profile information (types and number of each type) of client devices connected to a network to adapt media to one or more media requirements of the client devices for delivery of the media to the client devices, and prioritizing a first adaptation that adapts the media to a first media requirement for delivery to a first subset of the client devices according to the client device profile information.
[0026] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for media processing.
[0027] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief explanation of the drawings]
[0028] [Figure 1] 1 illustrates the media flow process in some examples. [Figure 2] 1 illustrates the media transformation decision-making process in some examples. [Figure 3] 1 shows a representation of a format for timed heterogeneous immersive media in an example. [Figure 4] 1 illustrates a representation of a streamable format for non-timed heterogeneous immersive media in an example. [Figure 5] 1 shows a diagram of a process for synthesizing media from natural content into an ingest format in some examples. [Figure 6] 1 shows a diagram of a process for creating an ingest format for composite media in some examples. [Figure 7] FIG. 1 is a schematic diagram of a computer system, according to an embodiment. [Figure 8]1 illustrates a network media distribution system that supports a variety of legacy and heterogeneous immersive media capable displays as client endpoints in some examples. [Figure 9] 1 illustrates a diagram of an immersive media delivery module capable of servicing legacy and heterogeneous immersive media-enabled displays in some examples. [Figure 10] 1 shows a diagram of the media adaptation process in some examples. [Figure 11] 1 illustrates a delivery format creation process in some examples. [Figure 12] 1 illustrates a packetizer processing system in some examples. [Figure 13] 1 illustrates a diagram of a network sequence that adapts a particular immersive media in an ingest format into a streamable and appropriate delivery format for a particular immersive media client endpoint in some examples. [Figure 14] 1 illustrates a diagram of a media system with a virtual network and client devices for scene-based media processing in some examples. [Figure 15] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 16] 10A-10C illustrate diagrams for adding complexity and priority values to a scene in a scene manifest in some examples. [Figure 17] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 18] 1 shows an illustration of a virtual museum to illustrate scene priorities in some examples. [Figure 19] 1 illustrates a diagram of a scene for scene-based immersive media in some examples. [Figure 20]1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 21] 1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 22] 1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 23] 10 shows a diagram of a scene manifest with scene to asset mappings in an example. [Figure 24] 10A-10C illustrate diagrams of streaming assets for a scene in some examples. [Figure 25] 10A-10C show diagrams for re-ranking assets of a scene in some examples. [Figure 26] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 27] 1 shows a diagram of a scene in some examples. [Figure 28] 10A-10C show diagrams for re-ranking assets of a scene in some examples. [Figure 29] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 30] 1 shows a diagram of a scene in some examples. [Figure 31] 10A-10C show diagrams for assigning priority values to assets in a scene in some examples. [Figure 32] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 33] 1 shows a diagram of a scene in some examples. [Figure 34] 10A-10C show diagrams for re-ranking assets of a scene in some examples. [Figure 35] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 36] 1 shows diagrams of displayed scenes in some examples. [Figure 37] In the example, a diagram of a scene 3701 of scene-based immersive media is shown. [Figure 38] 1 shows a diagram for prioritizing assets of a scene in some examples. [Figure 39] 1 illustrates an example diagram for streaming scene-based immersive media to an end device in some examples. [Figure 40] 1 shows a diagram illustrating metadata attributes of assets in an example. [Figure 41] 1 shows a diagram for prioritizing assets of a scene in some examples. [Figure 42] 1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 43] 1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 44] 10A-10C illustrate diagrams of timed media presentations with asset signaling for default camera viewpoints in some examples. [Figure 45] 1 shows a diagram of a non-timed media presentation with asset signaling in a default viewport in some examples. [Figure 46] 1 illustrates a process flow for a scene analyzer to analyze assets of a scene according to a default viewport. [Figure 47] 1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 48] 1 illustrates a diagram of an immersive media delivery module in some examples. [Figure 49] 1 shows a diagram of a media adaptation process in some examples. [Figure 50] 5 shows a flowchart outlining a process 5000 according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0029] Aspects of the present disclosure provide architectures, structures, components, techniques, systems, and / or networks that deliver media, including video, audio, geometric (3D) objects, haptics, associated metadata, or other content, for client devices. In some examples, the architectures, structures, components, techniques, systems, and / or networks are configured to deliver media content to heterogeneous immersive and interactive client devices, such as game engines. In some examples, the client devices may include light fields or holographic immersive displays. It should be noted that while some embodiments of the present disclosure use immersive media for light fields or holographic immersive displays as examples, the disclosed techniques may be used with any suitable immersive media.
[0030] Immersive media is defined by immersive technologies that attempt to create or mimic the physical world through digital simulation, thereby simulating any or all of the human sensory systems to create the user's perception of being physically present in the scene.
[0031] There are several different types of immersive media technologies currently being developed and used, including virtual reality (VR), augmented reality (AR), mixed reality (MR), and light field / holographic. VR refers to a digital environment that replaces a user's physical environment by placing them in a computer-generated world using a headset. AR, on the other hand, takes digital media and overlays it on the real world around the user, using either clear vision or a smartphone. MR refers to blending the real and digital worlds, thereby creating an environment where technology and the physical world can coexist.
[0032] Holographic and / or light field technology can create virtual environments with accurate depth perception and three-dimensionality without the need for any headset, thus avoiding side effects such as nausea. Light field and / or holographic technology may, in some instances, use light rays in 3D space, with light rays coming from various points and directions. Light field and / or holographic technology is based on the concept that everything visible is illuminated by light coming from any source, traveling through space, and hitting the surface of an object, where the light is partially absorbed and partially reflected by another surface before reaching the human eye. In some instances, light field and holographic technology can recreate a light field that provides users with 3D effects such as binocularity and continuous motion parallax. For example, a light field display may include a large array of projection modules that recreate an approximation of a light field by projecting light rays onto a holographic screen to show different but consistent information in slightly different directions. In another example, a light field display can emit light rays according to a plenoptic function, where the light rays converge in front of the light field display to create a real 3D realistic image. In this disclosure, a light field-based display refers to a display that uses light rays in 3D space, such as a light field display, a holographic display, etc.
[0033] In some examples, light rays may be defined by a five-dimensional plenoptic function, where each light ray may be defined by three coordinates (three dimensions) in 3D space and two angles to specify a direction in 3D space.
[0034] Generally, capturing content for 360-degree video requires a 360-degree camera. However, to capture content for lightfield or holographic-based displays, in some instances, expensive setups including multi-depth cameras or arrays of cameras may be required depending on the field of view (FoV) of the scene being rendered.
[0035] In an example, a conventional camera can capture a 2D representation of the light rays reaching the camera lens at a given position: the image sensor records the sum of the brightness and color of all the light rays reaching each pixel.
[0036] In some examples, a light-field camera (also known as a plenoptic camera) is used to capture the content of a light-field-based display; the light-field camera can capture not only the brightness and color, but also the direction of all light rays reaching the camera sensor. The information captured by the light-field camera can be used to reconstruct a digital scene with an accurate representation of the origin of each light ray, allowing for the digital reconstruction of an accurately captured scene in 3D.
[0037] Two techniques have been developed to capture such volumetric scenes: the first uses a camera or an array of camera modules to capture different light rays / viewpoints from each direction, and the second uses a plenoptic camera (or an array of plenoptic cameras) that can also capture light rays from different directions.
[0038] In some examples, the volumetric scene may also be synthesized using computer-generated imagery (CGI) techniques and then rendered using the same techniques used to render the volumetric scene captured by the camera.
[0039] Whether captured by a plenoptic camera or synthesized by CGI, multimedia content for a light-field-based display is captured and stored on a media server (also called a media server device). Ultimately, the content is converted into a geometric representation that includes explicit information about the geometry of each object in the visual scene, as well as metadata describing surface properties (e.g., smoothness, roughness, ability to reflect, refract, or absorb light). Each object in the visual scene is called an "asset" of the scene. Transmitting this data to a client device requires a significant amount of bandwidth, even after the data is compressed. Therefore, in bandwidth-limited situations, the client device may experience buffering or stuttering, resulting in an unpleasant experience.
[0040] In some examples, the use of CDNs and edge network elements can reduce the latency between a client device's request for a scene and the delivery of the scene (including each visual object in the scene, collectively referred to as the scene's "assets") to the client device. Furthermore, in contrast to video, which is simply decoded and placed into a series of linear frame buffers to be presented to the user, immersive media scenes are computationally intensive and consist of objects and rich metadata that must be rendered by real-time renderers (similar to the renderers featured in game engines such as Unity Technologies' Unity Engine or Epic Game's Unreal Engine) that may require significant resources to fully construct the scene. For such scene-based media, the use of cloud / edge network elements can be used to offload the computational burden of rendering by the client device to more powerful computation engines, such as one or more server devices in the network.
[0041] Multimedia content for light-field-based holographic displays is captured and stored on a media server. The multimedia content can be real-world or synthetic content. A large amount of bandwidth is required to transmit the data of the multimedia content to a client end device (also referred to as a client device or an end device), even after the data is compressed. Therefore, in situations where bandwidth is limited, the client device may experience buffering or interruptions, which can be an unpleasant experience.
[0042] Immersive media generally refers to media that stimulates any or all human sensory systems (visual, auditory, somatosensory, olfactory, and sometimes gustatory) to create or enhance the user's perception of being physically present in the media experience, i.e., beyond that delivered over existing commercial networks for timed two-dimensional (2D) video and corresponding audio known as "legacy media." In some instances, immersive media refers to media that attempt to create or mimic the physical world through digital simulation of dynamics and the laws of physics, thereby stimulating any or all human sensory systems to create the user's perception of being physically present in a scene depicting the real or virtual world. Both immersive and legacy media can be characterized as either timed or non-timed media.
[0043] Timed media refers to media that is structured and presented according to time. Examples include feature films, news reports, and episodic content, all of which are organized according to durations. Traditional video and audio are generally considered to be timed media.
[0044] Non-timed media is media that is not structured by time, but rather by logical, spatial, and / or temporal relationships. Examples include video games in which the user controls the experience created by a gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media may incorporate timed media, for example, in a continuously looped audio or video segment of a video game scene. Conversely, timed media may incorporate non-timed media, such as a video with a fixed still image as a background.
[0045] An immersive media-capable device can refer to a device with sufficient resources and capabilities to access, interpret, and present immersive media. Such media and devices are heterogeneous with respect to the number and format of media and the number and type of network resources required to deliver such media on a large scale, i.e., to achieve delivery equivalent to legacy video and audio media over a network. Similarly, media are heterogeneous in the amount and type of network resources required to deliver such media on a large scale. "On a large scale" can refer to delivery of media by a service provider, e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions, that achieves delivery equivalent to delivery of legacy video and audio media over a network.
[0046] Generally, legacy devices such as laptop displays, televisions, and mobile handset displays are uniform in their capabilities because all of these devices consist of rectangular display screens and consume 2D rectangular video or still images as their primary media format. Similarly, the number of audio formats supported by legacy devices is limited to a relatively small set.
[0047] The term "frame-based" media refers to the property of visual media that it consists of one or more consecutive rectangular frames of an image. In contrast, "scene-based" media refers to visual media that is organized by "scenes," where each scene refers to individual assets that collectively describe the visual characteristics of the scene.
[0048] An example of a comparison between frame-based and scene-based visual media can be described using a visual media showing a forest. In a frame-based representation, the forest is captured using a camera device, such as a camera phone. A user can enable the camera device to focus on the forest, and the frame-based media captured by the camera device is the same as what the user sees through the camera viewport provided to the camera device, including any movements of the camera device initiated by the user. The resulting frame-based representation of the forest is a series of 2D images recorded by the camera device at a standard rate, typically 30 or 60 frames per second. Each image is a collection of pixels where the information stored in each pixel matches the next pixel.
[0049] In contrast, a scene-based representation of a forest consists of individual assets describing each object in the forest and a human-readable scene graph description that presents a myriad of metadata describing the asset or assets and how they are rendered. For example, a scene-based representation may include individual objects called "trees," each composed of a collection of smaller assets called "trunks," "branches," and "leaves." Each tree trunk may be further described individually by a mesh (tree trunk mesh), which describes the complete 3D geometry of the tree trunk and a texture applied to the tree trunk mesh to capture the color and emissive properties of the tree trunk. Furthermore, the tree trunk may be accompanied by additional information describing the surface of the tree trunk in terms of its smoothness or roughness or its ability to reflect light. The corresponding human-readable scene graph description may provide information about where to place the tree trunk relative to the viewport of a virtual camera focused on the forest scene. Furthermore, the human-readable description may include information about how many branches to generate and where to place them in the scene from a single branch asset called a "branch." Similarly, the description may include how many leaves to generate and their position relative to the branches and tree trunk. Additionally, a transformation matrix can provide information about how to scale or rotate leaves so that they do not appear uniform. As a whole, the individual assets that make up a scene vary in the type and amount of information stored in each asset. Each asset is typically stored in its own file, but often an asset is used to create multiple instances of the object it is designed to create, for example, the branches and leaves of a tree.
[0050] In some instances, the human-readable portion of the scene graph is rich in metadata to describe not only the relationship of assets to their positions in the scene, but also instructions on how to render objects, for example, with different types of light sources, or with surface characteristics (to indicate that an object has a metallic versus matte surface) or other materials (porous versus smooth textures). Other information that is often stored in the human-readable portion of the graph is the relationship of assets to other assets, for example, to form a group of assets that are rendered or treated as a single entity, e.g., a tree trunk with branches and leaves.
[0051] An example of a scene graph with human-readable components includes glTF 2.0, where the node tree component is provided in Java Script Object Notation (JSON), a human-readable notation for describing objects. Another example of a scene graph with human-readable components is the Immersive Technologies Media Format, where OCS files are generated using XML, another human-readable notation format.
[0052] Yet another difference between scene-based media and frame-based media is that in frame-based media, the view created of a scene is identical to the view captured by a user through a camera, i.e., when the media was created. When frame-based media is presented by a client, the view of the presented media is the same as the view captured in the media by the camera used to record the video, for example. However, in scene-based media, there may be multiple ways for a user to view a scene using various virtual cameras, for example, a thin-lens camera versus a panoramic camera. In some examples, the visual information of a scene that a viewer can see may change depending on the view angle and position of the viewer.
[0053] A client device that supports scene-based media may be equipped with renderers and / or resources (e.g., GPU, CPU, local media cache storage) whose capabilities and supported features collectively comprise an upper bound or upper limit to characterize the client device's overall ability to ingest various scene-based media formats. For example, a mobile handset client device may be limited in the complexity of the geometric assets it can render, e.g., the number of polygons describing the geometric assets, particularly for support of real-time applications. Such a limit may be established based on the fact that the mobile client is battery-powered, and therefore the amount of computational resources available to perform real-time rendering is similarly limited. In such a scenario, it may be desirable for the client device to inform the network that the client prefers that the client access geometric assets with polygon counts below a specified upper bound. Furthermore, information communicated from the client to the network may best be conveyed using a well-defined protocol that leverages a well-defined vocabulary of attributes.
[0054] Similarly, a media delivery network may have computational resources that facilitate the delivery of immersive media in various formats to various clients with various capabilities. In such a network, it may be desirable for the network to be informed of client-specific capabilities according to a well-defined profile protocol, e.g., a vocabulary of attributes communicated via a well-defined protocol. Such a vocabulary of attributes may include information for describing the media or the minimum computational resources required to render the media in real time, so that the network can better establish priorities for how to serve media to its heterogeneous clients. Furthermore, a centralized data store in which client-provided profile information is collected across client domains can help provide a summary of which types of assets and in which formats are in high demand. Given information about which types of assets are in higher and lower demand, an optimized network can prioritize the task of responding to requests for assets in higher demand.
[0055] In some examples, distribution of media over a network can employ media distribution systems and architectures that reformat media from an input or network "ingested" media format to a distributed media format. In examples, the distributed media format is not only suitable for ingestion by a target client device and its applications, but also lends itself to being "streamed" over the network. In some examples, there can be two processes that occur upon ingestion of media over the network: 1) a process to convert the media from format A to format B suitable for ingestion by the target client device, i.e., a process that converts based on the client device's ability to ingest a particular media format, and 2) a process to prepare the media to be streamed.
[0056] In some examples, "streaming" media broadly refers to fragmenting and / or packetizing media so that the processed media can be delivered over a network in successive, smaller-sized "chunks" that are logically organized and sequenced according to either or both the temporal and spatial structure of the media. In some examples, "converting" media from format A to format B, sometimes referred to as "transcoding," may be a process typically performed by a network or service provider prior to delivering the media to a target client device. Such transcoding may consist of converting media from format A to format B based on prior knowledge that format B is the preferred or only format that can be captured by the target client device at all, or that is more suitable for delivery over constrained resources, such as a commercial network. One example of media conversion is converting media from a scene-based representation to a frame-based representation. In some examples, both converting the media and preparing the media to be streamed are necessary before the media can be received from the network and processed by the target client device. Such prior knowledge of client preferred formats can be obtained through the use of a profile protocol that utilizes an agreed-upon vocabulary of attributes that summarize the characteristics of preferred scene-based media across various client devices.
[0057] In some examples, the above one- or two-step process of operating on the ingested media by the network, i.e., before delivering the media to the target client device, results in a media format referred to as a "delivery media format" or simply a "delivery format." Generally, these steps, if performed at all for a given media data object, may be performed only once if the network has access to information indicating that the target client device requires the converted and / or streamed media object for multiple occasions to trigger the conversion and streaming of such media multiple times. That is, the processing and transfer of data for media conversion and streaming is generally considered a source of latency that requires the consumption of potentially significant amounts of network and / or computational resources. Thus, network designs that do not have access to information indicating that a client device may already have a particular media data object cached or stored locally to the client device will perform less optimally than networks that have access to such information.
[0058] In some examples, for legacy presentation devices, the delivery format may be equivalent or sufficiently equivalent to the “presentation format” ultimately used by a client device (e.g., client presentation device) to create the presentation. For example, a presentation media format is a media format whose properties (resolution, frame rate, bit depth, color gamut, etc.) are closely aligned with the capabilities of the client presentation device. Some examples of delivery format versus presentation format include a high-definition (HD) video signal (1920 pixel columns by 1080 pixel rows) delivered over a network to an ultra-high-definition (UHD) client device with a resolution (3840 pixel columns by 2160 pixel rows). For example, a UHD client device may apply a process called “super-resolution” to the HD delivery format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format presented by the UHD client device is the “presentation format,” which in this example is a UHD signal, but the HD signal includes the delivery format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format since both signals are linear video formats, and the process of converting the HD format to the UHD format is relatively simple and easy to perform on most legacy client devices.
[0059] In some examples, the preferred presentation format for the target client device may be significantly different from the ingest format received by the network. Nevertheless, the target client device may have access to sufficient computational, storage, and bandwidth resources to convert the media from the ingest format to the required presentation format suitable for presentation by the target client device. In this scenario, the network may bypass steps of reformatting the ingested media, such as “transcoding” the media from format A to format B, simply because the client device has access to sufficient resources to perform all media conversion without the network having to do so. However, the network may still perform steps of fragmenting and packaging the ingest media so that the media can be streamed to the target client device.
[0060] In some examples, the ingested media received by the network is significantly different from the target client device's preferred presentation format, and the target client device does not have access to sufficient computational, storage, and / or bandwidth resources to convert the media to the preferred presentation format. In such scenarios, the network may assist the target client device by performing some or all of the conversion from the ingest format to a format equivalent or nearly equivalent to the target client device's preferred presentation format on behalf of the target client device. In some architectural designs, such assistance provided by the network on behalf of the target client device is referred to as "split rendering" or media "adaptation."
[0061] FIG. 1 illustrates a media flow process 100 (also referred to as process 100) in some examples. Media flow process 100 includes a first step that may be performed by a network cloud (or edge device) 104 and a second step that may be performed by a client device 108. In some examples, media in ingest media format A is received over a network from a content provider in step 101. Step 102, a network process step, may prepare the media for delivery to a client device 108 by formatting the media into format B and / or by preparing the media to be streamed to the client device 108. In step 103, the media is streamed from the network cloud 104 to the client device 108 over a network connection using a network protocol such as TCP or UDP. In some examples, such streamable media, depicted as media 105, may be streamed to a media store 110. The client device 108 accesses the media store 110 via a fetch mechanism 111 (e.g., ISO / IEC 23009 Dynamic Adaptive Streaming over HTTP). Client device 108 can receive or fetch the distributed media from the network and prepare the media for presentation via a rendering process as indicated by 106. The output of the rendering process 106 is presentation media in yet another, potentially different format C as indicated by 107.
[0062] FIG. 2 illustrates a media conversion decision-making process 200 (also referred to as process 200) illustrating a network logic flow for processing ingested media within a network (also referred to as a network cloud), for example, by one or more devices within the network. At 201, media is ingested by the network cloud from a content provider. If not already known, attributes of the target client device are obtained at 202. Decision-making step 203 determines whether the network should support conversion of the media, if necessary. If decision-making step 203 determines that the network should support conversion, the ingested media is converted by converting the media from format A to format B, generating converted media 205. At 206, the media, either converted or in its original form, is prepared to be streamed. At 207, the prepared media is appropriately streamed to a target client device, such as a game engine target client device, or to media store 110 of FIG. 1.
[0063] An important aspect to the logic of Figure 2 is decision-making process 203, which may be performed by an automated process. The decision-making step may determine whether the media can be streamed in its original captured format A, or whether the media must be converted to a different format B to facilitate presentation of the media by the target client device.
[0064] In some examples, decision-making process step 203 may require access to information describing aspects or characteristics of the ingest media to assist decision-making process step 203 in making the optimal choice, i.e., to determine whether conversion of the ingest media is required before streaming the media to the target client device, or whether the media can be streamed directly to the target client device in its original ingest format A.
[0065] According to aspects of the present disclosure, streaming of scene-based immersive media may differ from streaming frame-based media. For example, streaming of frame-based media may be equivalent to streaming frames of video, with each frame capturing a complete picture of an entire scene or object being presented by a client device. The sequence of frames is reconstructed by the client device from their compressed format and, when presented to a viewer, creates a video sequence that includes the entire immersive presentation or a portion of the presentation. When streaming frame-based media, the order in which frames are streamed from the network to the client device may conform to a predetermined specification (such as ITU-T Recommendation H.264 Advanced Video Coding for General Audiovisual Services).
[0066] However, scene-based streaming of media differs from frame-based streaming because a scene may be composed of individual assets that may themselves be independent of one another. A given scene-based asset may be used multiple times within a particular scene or across a series of scenes. The amount of time required for a client device, or any given renderer, to create the correct presentation of a particular asset may depend on many factors, including, but not limited to, the size of the asset, the availability of computational resources to perform the rendering, and other attributes that describe the overall complexity of the asset. Client devices that support scene-based streaming may require some or all of the rendering of each asset in a scene to be completed before any presentation of the scene can begin. Therefore, the order in which assets are streamed from the network to the client device may affect overall performance.
[0067] According to aspects of the present disclosure, considering each of the above scenarios in which the conversion of media from format A to another format may be performed entirely by the network, entirely by the client device, or jointly between the network and the client, for example, for split rendering, a vocabulary of attributes describing the media format may be needed so that both the client device and the network have complete information to characterize the conversion operation. Additionally, a vocabulary providing attributes of the client device's capabilities, for example, with respect to available computational resources, available storage resources, and access to bandwidth, may also be needed. Furthermore, a mechanism may be needed to characterize the level of computational, storage, or bandwidth complexity of an ingest format so that the network and client device, jointly or independently, can determine whether or when the network can use a split rendering step to deliver media to the client device. Furthermore, if the network can avoid converting and / or streaming certain media objects that the client device requires to complete the presentation of the media, the network can skip the conversion and ingest media streaming steps, assuming that the client device has access to or availability of media objects (e.g., previously streamed to the client device) that the client device may require to complete the presentation of the media. In some examples, if conversion from format A to another format is determined to be a necessary step to be performed by or on behalf of a client device, a prioritization scheme can be implemented to order the conversion process of individual assets in a scene, which can benefit an intelligent and efficient network architecture.
[0068] To facilitate the ability of client devices to perform at their full potential, it may be desirable for the network to have sufficient information regarding the order in which scene-based assets are streamed from the network to client devices so that the network can determine such an order to improve client device performance. For example, such a network with sufficient information to avoid repetitive conversion and / or streaming steps for assets used multiple times in a particular presentation can perform more optimally than a network not designed for that purpose. Similarly, a network that can “intelligently” order the delivery of assets to clients can facilitate the ability of client devices to perform at their full potential, i.e., creating a more enjoyable experience for end users. Furthermore, the interface between client devices and the network (e.g., a server device of the network) may be implemented using one or more communication channels to convey essential information regarding the operating state characteristics of the client devices, the availability of resources at or local to the client devices, the type of media being streamed, and the frequency of assets used or across multiple scenes. Thus, a network architecture implementing the streaming of scene-based media to heterogeneous clients may require access to a client interface that can provide and update information related to the processing of each scene to a network server process, including current conditions related to the client device's ability to access computational and storage resources. Such a client interface may also interact closely with other processes running on the client device, particularly game engines, which may play an essential role on behalf of the client device's ability to deliver an immersive experience to an end user.An example of an essential role that a game engine can play includes providing an application program interface (API) to enable the delivery of an interactive experience. Another role that can be provided by a game engine on behalf of a client device is the rendering of the precise visual signals required by the client device to provide a visual experience that matches the capabilities of the client device.
[0069] Definitions of some terms used in this disclosure are provided in the following paragraphs.
[0070] Scene graph: A common data structure typically used by vector-based graphics editing applications and modern computer games that constitutes a logical and often (but not necessarily) spatial representation of a graphical scene; it is a collection of nodes and vertices in a graph structure.
[0071] Scene: In the context of computer graphics, a scene is a collection of objects (e.g., 3D assets), object attributes, and other metadata, including visual, acoustic, and physics-based characteristics, that describe a particular setting, bounded either by space or time, with respect to the interactions of objects within that setting.
[0072] Node: A basic element of a scene graph consisting of information related to the logical, spatial, or temporal representation of visual, audio, tactile, olfactory, gustatory, or related processing information. Each node shall have at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) attached to it.
[0073] Base Layer: A nominal representation of an asset, typically formulated to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network.
[0074] Enhancement Layer: A set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities not supported in the base layer.
[0075] Attribute: Metadata associated with a node that is used to describe a particular characteristic or feature of that node, either in canonical form or in more complex form (e.g., with respect to another node).
[0076] Binding LUT: A logical structure that associates metadata from the IMS of ISO / IEC 23090 Part 28 with metadata or other mechanisms used to describe the features or functionality of a particular scene graph format, e.g., ITMF, glTF, Universal Scene Description.
[0077] Container: A serialization format for storing and exchanging information for representing an entire natural scene, an entire synthetic scene, or a mix of synthetic and natural scenes, including a scene graph and all the media resources required to render the scene.
[0078] Serialization: The process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link), and later reconstructed (possibly in a different computing environment). When the resulting series of bits is read again according to the serialization format, it can be used to create a semantically identical clone of the original object.
[0079] Renderer: A (typically software-based) application or process based on a selective blend of academic disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development that, given an input scene graph and asset container, emits exemplary visual and / or audio signals suitable for presentation on a target device or adapted to desired properties specified by attributes of the nodes to be rendered in the scene graph. For visual-based media assets, a renderer can emit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline); for audio-based media assets, a renderer can emit audio signals for presentation over multichannel loudspeakers and / or binaural headphones or for repackaging into another (output) container. Common examples of renderers include the real-time rendering capabilities of game engines Unity and Unreal Engine.
[0080] Evaluation: Generate results that move the output from abstract to concrete results (e.g., similar to evaluating a document object model for a web page).
[0081] Scripting Language: An interpreted programming language that can be executed by the renderer at runtime to process dynamic inputs and variable state changes applied to scene graph nodes, which changes affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformations, collisions) and energy propagation and transport (light, sound).
[0082] Shader: a type of computer program originally used for shading (producing appropriate levels of light and color in an image), but now performing a variety of specialized functions in various areas of computer graphics special effects, or video post-processing unrelated to shading, or even functions unrelated to graphics at all.
[0083] Path Following: A computer graphics method for rendering three-dimensional scenes so that the lighting in the scene is realistic.
[0084] Timed Media: Media that is ordered by time, e.g., has a start time and an end time according to a particular clock.
[0085] Non-timed media: Media that is organized by spatial, logical, or temporal relationships, such as an interactive experience that is realized according to actions taken by a user.
[0086] Neural network model: a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numbers) used in well-defined mathematical operations that are applied to a visual signal to arrive at an improved visual output, which may include the interpolation of new views of the visual signal that were not explicitly presented by the original signal.
[0087] Conversion: Refers to the process of converting one media format or type into another media format or type, respectively.
[0088] Adaptation: Refers to the translation and / or converting of media into multiple representations of bitrates.
[0089] Frame-based media: 2D video, with or without associated audio.
[0090] Scene-based media: Audio, visual, tactile, and other major types of media, as well as media-related information organized logically and spatially using a scene graph.
[0091] Over the past decade, several immersive media-enabled devices have been introduced to the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and gaming consoles. Similarly, holographic displays and other forms of stereoscopic displays are poised to appear on the consumer market within the next three to five years. Despite the immediate or imminent availability of these devices, a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks has not materialized for several reasons.
[0092] One of the obstacles to realizing a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks is the great diversity of client devices that serve as endpoints in such delivery networks for immersive displays. Some of these support specific immersive media formats, while others do not. Some of these are capable of creating immersive experiences from legacy raster-based formats, while others are not. Unlike networks designed solely for the delivery of legacy media, networks that must support a variety of display clients require a significant amount of information detailing each client's capabilities and the format of the media being delivered before such networks can use an adaptation process to convert the media into a format appropriate for each target display and corresponding application. At a minimum, such networks need access to information describing the characteristics of each target display and the complexity of the ingested media in order for the network to ascertain how to meaningfully adapt the input media source to a format appropriate for the target display and application. Similarly, a network optimized for efficiency may wish to maintain a database of the types of media supported by client devices connected to such a network and their corresponding attributes.
[0093] Similarly, an ideal network supporting heterogeneous client devices would exploit the fact that some of the assets adapted from an input media format to a particular target format can be reused across a set of similar display targets. That is, once converted to a format appropriate for the target displays, some assets can be reused across several such displays with similar adaptation requirements. Such an ideal network would therefore employ caching mechanisms to store the adapted assets in a relatively immutable domain, i.e., similar to the use of content delivery networks (CDNs) used in legacy networks.
[0094] Furthermore, immersive media may be organized into "scenes," e.g., "scene-based media," that are described by a scene graph, also known as a scene description. The scope of a scene graph is to describe the visual, audio, and other forms of immersive assets that comprise a particular setting that is part of a presentation, e.g., actors and events that take place in a particular location within a building that is part of a presentation such as a movie. A list of all the scenes that comprise a single presentation may be formulated in a scene manifest (also called a scene manifest).
[0095] An additional benefit of such an approach is that, for content that is prepared before such content must be delivered, a "billing of materials" can be created that identifies all of the assets used throughout the presentation and the frequency with which each asset is used across various scenes within the presentation. An ideal network should have knowledge of the existence of cached resources that can be used to satisfy the asset requirements of a particular presentation. Similarly, a client presenting a series of scenes may want to have knowledge of the frequency with which any given asset is used across multiple scenes. For example, if a media asset (also known as an object) is referenced multiple times across multiple scenes that are or will be processed by the client, the client should avoid discarding the asset from its caching resource until the last scene requiring that particular asset has been presented by the client.
[0096] Furthermore, such a process, which can generate a "billing of materials" for a given scene or collection of scenes, can also annotate the scenes with standardized metadata, for example from IMS in ISO / IEC 23090 Part 28, to facilitate adaptation of the scenes from one format to another.
[0097] Finally, many emerging advanced imaging displays, including but not limited to Oculus Rift, Samsung Gear VR, Magic Leap goggles, all Looking Glass Factory displays, SolidLight by Light Field Lab, Avalon Holographic displays, and Dimenco displays, utilize game engines as the mechanism by which the respective displays can capture content to be rendered and presented on the displays. Currently, the most popular game engines employed across this aforementioned set of displays include Unreal Engine by Epic Games and Unity by Unity Technologies. That is, advanced imaging displays are currently designed and shipped employing one or both of these game engines as the mechanism by which the displays can capture media to be rendered and presented by such advanced imaging displays. Both Unreal Engine and Unity are optimized to capture scene-based media rather than frame-based media. However, existing media distribution ecosystems are only capable of streaming frame-based media. There is a significant "gap" in the current media delivery ecosystem that includes standards (de jure or de facto) and best practices to enable the delivery of scene-based content to emerging advanced imaging displays so that media can be delivered "at scale," e.g., at the same scale at which frame-based media is delivered.
[0098] In some examples, a mechanism or process may be used that responds to a network server process and participates in a combined network and immersive client architecture on behalf of a client device on which a game engine is employed to ingest scene-based media. Such a "smart client" mechanism is particularly appropriate for networks designed to stream scene-based media to immersive, heterogeneous, interactive client devices, where the delivery of media is performed efficiently and within the constraints of the capabilities of the various components that make up the overall network. A "smart client" is associated with a particular client device and responds to network requests for information regarding the current state of that associated client device, including the availability of resources on the client device for rendering and creating presentations of scene-based media. The "smart client" also acts as an "intermediary" between the client device on which the game engine is employed and the network itself.
[0099] It should be noted that the remainder of the disclosed subject matter assumes, without loss of generality, that a smart client that can respond on behalf of a particular client device can also respond on behalf of client devices on which one or more other applications (i.e., not game engine applications) are active. That is, the problem of responding on behalf of a client device is equivalent to the problem of responding on behalf of a client device on which one or more other applications are active.
[0100] Additionally, it should be noted that the terms "media object" and "media asset" are sometimes used interchangeably and both refer to a specific instance of media in a particular format. The term client device or client (without qualification) refers to the device and its components on which the presentation of media ultimately occurs. The term "game engine" refers to Unity or Unreal Engine, or any game engine that plays a role in a delivery network architecture.
[0101] In some examples, a mechanism or process may be employed by a network or client device to analyze an immersive media scene to obtain sufficient information that can be used to support a decision-making process. When employed by a network or client, the mechanism or process may indicate whether conversion of media objects (or media assets) from format A to format B should be performed entirely by the network, entirely by the client, or via a mix of both (along with indicating which assets should be converted by the client or the network). Such an "immersive media data complexity analyzer" (also referred to in some examples as a media analyzer) may be used by either a client device or a network device in an automated context.
[0102] Referring back to FIG. 1 , media flow process 100 illustrates the flow of media over a network 104 or delivery to a client device 108 where a game engine is used. In FIG. 1 , processing of ingest media format A is performed by processing in a cloud or edge device 104. At 101, media is obtained from a content provider (not shown). Process step 102 performs any necessary conversion or conditioning of the ingested media to create a potential alternative representation of the media as delivery format B. Media formats A and B may or may not be representations that follow the same syntax of a particular media format specification, but format B will likely be tailored in a manner that facilitates delivery of the media over a network protocol such as TCP or UDP. Such “streamable” media is shown as streamed media over a network connection 105 to a client device 108. The client device 108 has access to some rendering functionality, depicted as 106. Such rendering functionality 106 may be rudimentary or similarly sophisticated, depending on the client device 108 and the type of game engine running thereon. The rendering process 106 creates presentation media that may or may not be expressed according to a third format specification, e.g., Format C. In some examples, in client devices that use game engines, the rendering process 106 is typically a function provided by the game engine.
[0103] Referring to FIG. 2, a media conversion decision-making process 200 can be employed to determine whether a network needs to convert media before delivering it to a client device. In FIG. 2, ingested media 201, represented in format A, is provided to the network by a content provider (not shown). Process step 202 obtains attributes describing the processing capabilities of the target client (not shown). Decision-making process step 203 is used to determine whether the network or the client should perform any format conversion on any of the media assets included in the ingested media 201 before the media is streamed to the client, such as converting a particular media object from format A to format B. If any of the media assets need to be converted by the network, the network converts the media object from format A to format B using process step 204. Converted media 205 is the output from process step 204. The converted media is merged with preparation process 206 to prepare the media to be streamed to a game engine client (not shown). Process step 207, for example, streams the prepared media to the game engine client.
[0104] Figure 3 shows a representation of a streamable format 300 for heterogeneous immersive media that is timed in an example, and Figure 4 shows a representation of a streamable format 400 for heterogeneous immersive media that is non-timed in an example. In the case of Figure 3, Figure 3 references a scene of timed media 301. In the case of Figure 4, Figure 4 references a scene of non-timed media 401. In either case, the scene may be embodied by various scene representations or scene descriptions.
[0105] For example, in some immersive media designs, a scene may be embodied by a scene graph, or as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI and MSI technologies are examples of technologies that support the creation of display-independent scene representations for natural content, i.e., real-world images captured simultaneously from one or more cameras. On the other hand, scene graph technologies can be used to represent both natural and computer-generated imagery in the form of synthetic representations, but such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. That is, scene graph representations of naturally captured content are time-consuming and computationally intensive to create, requiring complex analysis of the natural imagery using photogrammetry or deep learning, or both, techniques, to create a synthetic representation that can later be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations are not currently practical to consider as candidates for representing natural content because they cannot actually be created in real time to account for use cases requiring real-time delivery. In some instances, the best candidate representation of a computer-generated image is through the use of a scene graph with a synthetic model, as computer-generated images are created using 3D modeling processes and tools.
[0106] This dichotomy in optimal representation of both natural and computer-generated content suggests that the optimal ingest format for naturally captured content may be different from the optimal ingest format for computer-generated or natural content that is not required for real-time delivery applications. Thus, the disclosed subject matter aims to be robust enough to support multiple ingest formats for visually immersive media, whether created naturally through the use of a physical camera or created by a computer.
[0107] Below are exemplary techniques for embodying scene graphs as a format suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content, whereby deep learning or photogrammetry techniques are employed to create a corresponding synthetic representation of the natural scene, i.e., not required for real-time distribution applications.
[0108] 1. ORBX (registered trademark) by OTOY OTOY's ORBX is one of several scene graph technologies capable of supporting any type of visual media, timed or untimed, including ray-traceable, legacy (frame-based), stereoscopic, and other types of synthetic or vector-based visual formats. According to an aspect, ORBX is unique from other scene graphs because it provides native support for freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph purposefully designed to facilitate interchange across multiple vendor technologies that operate on the scene graph. Furthermore, ORBX provides a rich material system, support for an open shader language, a robust camera system, and support for Lua scripting. ORBX is also the basis for an immersive technology media format released for royalty-free licensing by the Immersive Digital Experience Alliance (IDEA). In the context of real-time media distribution, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of camera-captured data and composition of that same data into a synthetic representation. To date, the availability of sufficient computation for real-time delivery is impractical, but nevertheless not impossible.
[0109] 2. Universal Scene Description by Pixar Pixar's Universal Scene Description (USD) is a scene graph that can be used in the visual effects (VFX) and professional content creation communities. USD is integrated into Nvidia's Omniverse platform, a set of developer tools for 3D modeling and rendering using Nvidia's GPUs. A subset of USD was released by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.
[0110] 3. Khronos glTF 2.0 glTF 2.0 is the latest version of the Graphics Language Transmission Format specification written by the Khronos 3D Group. This format supports simple scene graph formats, including "png" and "jpeg" image formats, that are generally capable of supporting static (untimed) objects in a scene. glTF 2.0 supports simple animation and supports the translation, rotation, and scaling of basic shapes, i.e., geometric objects, described using glTF primitives. glTF 2.0 does not support timed media, and therefore does not support video or audio.
[0111] 4.ISO / IEC 23090 Part 14 Scene Description is an extension to glTF2.0 that adds support for timed media, e.g., video and audio. It should be noted that the above scene representations of immersive visual media are provided by way of example only and do not limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source to a format suited to the particular characteristics of a client endpoint device.
[0112] Additionally, any or all of the above exemplary media representations currently employ or can employ deep learning techniques to train and create neural network models that enable or facilitate the selection of specific views to fill a particular display's viewing frustum based on the particular dimensions of the frustum. The views selected for a particular display's viewing frustum may be interpolated from existing views explicitly provided in the scene representation, e.g., from MSI or MPI techniques, or may be rendered directly from a rendering engine based on specific virtual camera positions, filters, or virtual camera descriptions for those rendering engines.
[0113] Thus, the disclosed subject matter is sufficiently robust to consider that there is a relatively small but well-known set of immersive media ingest formats that can adequately meet the requirements for both real-time or "on-demand" (e.g., non-real-time) delivery of media that is captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.
[0114] Interpolation of views from immersive media ingest formats using either neural network models or network-based rendering engines will become even easier as advanced network technologies such as 5G for mobile networks and fiber optic cable for fixed networks are deployed. These advanced network technologies increase the capacity and capability of commercial networks, as such advanced network infrastructure can support the transmission and delivery of increasingly large amounts of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networking (SDN), and network functions virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands on specific network resources, for example, to respond to dynamic increases and decreases in demand for network throughput, network speed, round-trip latency, and computational resources. Furthermore, this inherent ability to adapt to dynamic network requirements similarly facilitates the network's ability to adapt immersive media ingest formats to appropriate delivery formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.
[0115] Immersive media applications themselves may also have different requirements for network resources, including gaming applications that require significantly lower network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications that may have increasing demands on downlink resources depending on the type of display at the client endpoint that is consuming the data. In general, any consumer application may be supported by a variety of client endpoints that include different on-board client capabilities for storage, computation, and power, as well as equally different requirements for the particular media presentation.
[0116] The disclosed subject matter thus enables a fully equipped network, i.e., a network employing some or all of the characteristics of a modern network, to simultaneously support multiple legacy and immersive media-enabled devices in accordance with the characteristics specified below.
[0117] 1. Providing the flexibility to leverage a practical media ingest format for both real-time and on-demand use cases for the delivery of media.
[0118] 2. Provides flexibility to support both natural and computer-generated content for both legacy and immersive media-enabled client endpoints.
[0119] 3. Support both timed and non-timed media.
[0120] 4. Provide a process for dynamically adapting source media ingest formats to appropriate delivery formats based on the capabilities and capabilities of the client endpoint, as well as based on application requirements.
[0121] 5. Ensure that delivery formats are streamable across IP-based networks.
[0122] 6. Allows the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media-enabled devices and applications.
[0123] 7. We provide an exemplary media representation framework that facilitates the organization of distributed media along scene boundaries.
[0124] End-to-end implementation of the improvements enabled by the disclosed subject matter is achieved according to the processes and components described in the detailed description that follows.
[0125] Figures 3 and 4 each employ an exemplary generic delivery format that can be adapted from an ingest source format to match the capabilities of a particular client endpoint. As noted above, the media shown in Figure 3 is timed, while the media shown in Figure 4 is not. The particular inclusion format is robust enough in its structure to accommodate a wide variety of media attributes, each of which can be layered based on the amount of salient information each layer contributes to the media presentation. Note that the layering process can be applied, for example, to progressive JPEG and scalable video architectures (e.g., as specified in ISO / IEC 14496-10 Scalable Advanced Video Coding).
[0126] According to aspects, media streamed according to the encompassing media format is not limited to legacy visual and audio media, but may include any type of media information that can interact with a machine to generate signals that stimulate human sight, sound, taste, touch, and smell.
[0127] According to other aspects, media streamed according to the encompassing media format may be both timed media or non-timed media, or a mixture of both.
[0128] According to other aspects, the incorporating media format can be further streamlined by enabling hierarchical representation of media objects through the use of a base layer and enhancement layer architecture. In one example, separate base and enhancement layers are computed by applying multi-resolution or multi-tessellation analysis techniques to the media objects of each scene. This is similar to the progressively rendered image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In an example, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using wavelet analysis.
[0129] In another example of a layered representation of a media format, the enhancement layer applies different attributes to the base layer, such as modifying the material properties of the surface of the visual object represented by the base layer. In yet another example, the attributes may modify the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.
[0130] In yet another example of layered representation, the surfaces of one or more visual objects in a scene may be changed from Lambertian surfaces to ray-traceable surfaces.
[0131] In yet another example of a layered representation, the network delivers a base layer representation to a client so that the client can create a nominal presentation of the scene while awaiting the transmission of additional enhancement layers to refine the resolution or other characteristics of the base representation.
[0132] According to another aspect, the resolution or refinement information of attributes in the enhancement layer is not explicitly coupled with the resolution of objects in the base layer as is currently the case in existing MPEG video and JPEG image standards.
[0133] According to other aspects, the encompassing media format supports any type of information media that can be presented or acted upon by a presentation device or machine, thereby enabling support of heterogeneous media formats to heterogeneous client endpoints. In one embodiment of a network that delivers media formats, the network first queries the client endpoint to determine the client's capabilities, and if the client is unable to meaningfully consume the media representation, the network either removes layers of attributes not supported by the client or adapts the media from its current format to a format appropriate for the client endpoint. In one example of such adaptation, the network converts a volumetric visual media asset into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can use neural network processes to reformat the media into an appropriate format or, optionally, synthesize a view required by the client endpoint.
[0134] According to another aspect, a manifest for a complete or partially complete immersive experience (such as a live streaming event, game, or on-demand asset playback) is organized by scenes, which are the minimum amount of information currently available to rendering and game engines to create the presentation. The manifest includes a list of individual scenes for which the entire immersive experience requested by the client will be rendered. Associated with each scene are one or more representations of the geometric objects in the scene that correspond to a streamable version of the scene's geometry. One embodiment of a scene representation relates to a low-resolution version of the scene's geometric objects. Another embodiment of the same scene relates to enhancement layers for the low-resolution representation of the scene to add further detail or increase tessellation to the geometric objects of the same scene. As described above, each scene may have two or more enhancement layers to progressively increase the detail of the scene's geometric objects.
[0135] According to another aspect, each layer of a media object referenced in a scene is associated with a token (e.g., a URI) that points to an address where the resource can be accessed in the network. Such a resource is similar to a CDN where content can be fetched by a client.
[0136] According to other aspects, tokens for representations of geometric objects can point to locations within the network or within the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.
[0137] FIG. 3 illustrates a timed media representation 300 in some examples. The timed media representation 300 illustrates an example of a containing media format for timed media. The timed scene manifest 300A includes a list of scene information 301. The scene information 301 points to a list of components 302 that separately describe the processing information and types of media assets present in the scene information 301. The components 302 point to assets 303, which further point to base layers 304 and attribute enhancement layers 305. In the example of FIG. 3, each of the base layers 304 points to a numerical frequency metric indicating the number of times the asset was used across a set of scenes in the presentation. A list of unique assets not previously used in other scenes is provided at 307. The proxy visual assets 306 include information about the reused visual assets, such as unique identifiers for the reused visual assets, and the proxy audio assets 308 include information about the reused audio assets, such as unique identifiers for the reused audio assets.
[0138] FIG. 4 illustrates an example untimed media representation 400. The untimed media representation 400 illustrates an example of a containing media format for untimed media. The untimed scene manifest (not shown) references scene 1.0, which has no other scenes that can branch to it. The scene information 401 is not associated with a start or end time according to a clock. The scene information 401 references a list of components 402 that separately describe the processing information and types of media assets that make up the scene. The components 402 reference assets 403, which in turn reference base layers 404 and attribute enrichment layers 405, 406. In the example of FIG. 4, each of the base layers 404 points to a numerical frequency value indicating the number of times the asset is used across a set of scenes in the presentation. The scene information 401 may also point to other scene information 401 for untimed media. The scene information 401 also points to scene information 407 for timed media scenes. List 408 identifies unique assets associated with a particular scene that have not been previously used in a higher-level (eg, parent) scene.
[0139] 5 shows a diagram of a process 500 for synthesizing an ingest format from natural content. The process 500 includes a first sub-process for content capture and a second sub-process for ingest format synthesis for natural images.
[0140] In the example of FIG. 5 , in the first subprocess, camera units may be used to capture natural image content 509. For example, camera unit 501 may use a single camera lens to capture a scene of a person. Camera unit 502 may mount five camera lenses around a ring-shaped object to capture a scene with five diverging fields of view. The arrangement of camera unit 502 is an exemplary arrangement for capturing omnidirectional content for VR applications. Camera unit 503 may mount seven camera lenses on the inner diameter of a sphere to capture a scene with seven converging fields of view. The arrangement of camera unit 503 is an exemplary arrangement for capturing a light field or a light field for a holographic immersive display.
[0141] In the example of FIG. 5 , in a second sub-process, natural image content 509 is synthesized. For example, the natural image content 509 is provided as input to a synthesis module 504, which in the example may employ a neural network training module 505 that uses a set of training images 506 to generate a capture neural network model 508. Another process commonly used in place of the training process is photogrammetry. When the model 508 is created during the process 500 shown in FIG. 5 , the model 508 becomes one of the assets of the natural content ingest format 507. In some examples, an annotation process 511 may optionally be performed to annotate the scene-based media with IMS metadata. Exemplary embodiments of the ingest format 507 include MPI and MSI.
[0142] FIG. 6 shows a diagram of a process 600 for creating synthetic media 608, e.g., a computer-generated imagery ingest format. In the example of FIG. 6, a LIDAR camera 601 captures a point cloud 602 of a scene. Computer-generated imagery (CGI) tools, 3D modeling tools, or other animation processes for creating synthetic content are used on a computer 603 to create CGI assets 604 over a network. A motion capture suit with sensors 605A is worn by an actor 605 to capture a digital recording of the actor's 605 movements to generate animated motion capture (MoCap) data 606. Data 602, 604, and 606 are provided as inputs to a compositing module 607, which can also create a neural network model (not shown in FIG. 6), e.g., using a neural network and training data. In some examples, the compositing module 607 outputs synthetic media 608 in an ingest format. The composed media in ingest format 6608 may then be input to an optional IMS annotation process 609, the output of which is IMS annotated composed media in ingest format 610.
[0143] The techniques for representing, streaming, and processing heterogeneous immersive media in this disclosure can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0144] Computer software can be coded using any suitable machine code or computer language that is amenable to mechanisms such as assembly, compilation, linking, etc., to create code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or via interpretation, microcode execution, etc.
[0145] The instructions may be executed in various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0146] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 700.
[0147] The computer system 700 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0148] The input human interface devices may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touchscreen 710, a data glove (not shown), a joystick 705, a microphone 706, a scanner 707, and a camera 708 (only one of each is depicted).
[0149] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 710, data gloves (not shown), or joystick 705, although haptic feedback devices that do not function as input devices may also be present), audio output devices (such as speakers 709, headphones (not depicted)), visual output devices (such as screens 710, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output, virtual reality glasses (not depicted), holographic displays, and smoke tanks (not depicted)), and printers (not depicted).
[0150] The computer system 700 may also include human-accessible storage devices and media associated with the storage devices, such as optical media including CD / DVD ROM / RW (720) with media 721 such as CD / DVDs, thumb drives 722, removable hard drives or solid state drives 723, legacy magnetic media such as tape or floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0151] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0152] The computer system 700 may also include an interface 754 to one or more communications networks 755. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet; cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, and the like; television wired or wireless wide-area digital networks including cable, satellite, and terrestrial television; and vehicular and industrial networks including CANBus. Some networks require an external network interface adapter that attaches in a typical manner to a particular general-purpose data port or peripheral bus 749 (e.g., a USB port on the computer system 700); others are typically integrated into the core of the computer system 700 by attaching to the system bus, as described below (e.g., an Ethernet interface for a PC computer system or a cellular network interface for a smartphone computer system). Using any of these networks, the computer system 700 can communicate with other entities. Such communications may be one-way receive-only (e.g., television broadcast), one-way transmit-only (e.g., a CANbus to a particular CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.
[0153] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to core 740 of computer system 700 .
[0154] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing devices in the form of field programmable gate arrays (FPGAs) 743, hardware accelerators for specific tasks 744, graphics adapters 750, etc. These devices may be connected via a system bus 748, along with read-only memory (ROM) 745, random access memory 746, and internal mass storage such as an internal non-user-accessible hard drive, SSD, etc. 747. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 748 or via a peripheral bus 749. In an example, a screen 710 may be connected to the graphics adapter 750. Architectures for peripheral buses include PCI, USB, etc.
[0155] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute certain instructions that, in combination, may constitute the aforementioned computer code. That computer code may be stored in ROM 745 or RAM 746. Transient data may also be stored in RAM 746, while persistent data may be stored, for example, in internal mass storage 747. Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc.
[0156] The computer-readable medium may bear computer code for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0157] By way of example and not limitation, a computer system having architecture 700, and specifically core 740, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as described above, as well as media associated with specific storage of core 740 that is non-transitory in nature, such as core internal mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 740. The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause core 740, and specifically the processors in the core (including a CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator 744) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software can encompass logic, and vice versa. Where appropriate, references to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0158] FIG. 8 illustrates a network media distribution system 800 that supports various legacy and heterogeneous immersive media-enabled displays as client endpoints in some examples. In the example of FIG. 8, a content acquisition module 801 captures or creates media using the exemplary embodiments of FIG. 6 or FIG. 5. An ingest format is created in a content preparation module 802 and then transmitted to one or more client endpoints in the network media distribution system using a transmission module 803. A gateway 804 can service customer premises equipment (CPE) to provide network access to various client endpoints in the network. A set-top box 805 can also function as CPE to provide access to aggregated content by a network service provider. A wireless demodulator 806 can function as a mobile network access point for mobile devices (e.g., similar to a mobile handset and display 813). In one or more embodiments, a legacy 2D television 807 may be connected directly to the gateway 804, the set-top box 805, or a WiFi router 808. A computer laptop with a legacy 2D display 809 may be a client endpoint connected to the WiFi router 808. A head-mounted 2D (raster-based) display 810 may also be connected to the router 808. A lenticular light field display 811 may be connected to the gateway 804. The display 811 may consist of a local computation GPU 811A, a storage device 811B, and a visual presentation unit 811C that creates multiple views using ray-based lenticular optics technology. The holographic display 812 may be connected to the set-top box 805 and may include a local computation CPU 812A, a GPU 812B, a storage device 812C, and a Fresnel pattern, wave-based holographic visualization unit 812D.An augmented reality headset 814 may be connected to the wireless demodulator 806 and may include a GPU 814A, a storage device 814B, a battery 814C, and a stereoscopic visual presentation component 814D. A high-density light field display 815 may be connected to the WiFi router 808 and may include multiple GPUs 815A, a CPU 815B, and a storage device 815C, an eye-tracking device 815D, a camera 815E, and a high-density beam-based light field panel 815F.
[0159] FIG. 9 shows a diagram of an immersive media delivery module 900 capable of serving legacy and heterogeneous immersive media-enabled displays as previously described in FIG. 8. Content is created or acquired in module 901, which is embodied in FIGS. 5 and 6 for natural content and CGI content, respectively. The content is then converted to an ingest format using network ingest format creation module 902. Several examples of module 902 are embodied in FIGS. 5 and 6 for natural content and CGI content, respectively. In some examples, a media analyzer 911 can perform appropriate media analysis on the ingest media, and the captured media can be updated to store information from the media analyzer 911. In an example, the ingest media is updated to be annotated with IMS metadata by a scene analyzer with optional IMS notation in the media analyzer 911. The ingest media format is sent to the network and stored in storage device 903. In some other examples, the storage device may reside within the immersive media content creator's network and be accessed remotely by the immersive media network delivery module 900, as indicated by the bisecting dashed line. Client and application specific information may in some examples be available on a remote storage device 904, which may optionally reside remotely in an alternative cloud network in examples.
[0160] As shown in Figure 9, network orchestrator 905 serves as a primary source and sink of information for performing the primary tasks of the distribution network. In this particular embodiment, network orchestrator 905 may be implemented in a unified format with other components of the network. Nevertheless, the tasks depicted by network orchestrator 905 in Figure 9 form elements of the disclosed subject matter in some instances. Network orchestrator 905 may be implemented in software and executed by processing circuitry to perform processes.
[0161] According to some aspects of the present disclosure, the network orchestrator 905 may further employ a bidirectional messaging protocol for communicating with client devices to facilitate processing and delivery of media (e.g., immersive media) according to the characteristics of the client devices. Further, the bidirectional messaging protocol may be implemented across different delivery channels, i.e., a control plane channel and a data plane channel.
[0162] The network orchestrator 905 receives information about the characteristics and attributes of client devices, such as the client 908 (also referred to as the client device 908) in FIG. 9, and also gathers requirements regarding applications currently running on the client 908. This information may be obtained from the device 904, or in alternative embodiments, by directly querying the client 908. In some examples, a two-way messaging protocol is used to enable direct communication between the network orchestrator 905 and the client 908. For example, the network orchestrator 905 may send queries directly to the client 908. In some examples, the smart client 908E may participate in collecting and reporting client status and feedback on behalf of the client 908. The smart client 908E may be implemented in software that can be executed by processing circuitry to perform processes.
[0163] The network orchestrator 905 also initiates and communicates with the media adaptation and fragmentation module 910 described in FIG. 10. Once the ingested media has been adapted and fragmented by module 910, the media is transferred to an inter-media storage device, shown in some examples as prepared media for delivery storage device 909. Once the delivery media is prepared and stored on device 909, the network orchestrator 905 ensures that the immersive client 908 receives the delivery media and corresponding description information 906 via its network interface 908B through a push request, or the client 908 itself may initiate a pull request for the media 906 from the storage device 909. In some examples, the network orchestrator 905 can use a two-way message interface (not shown in FIG. 9) to perform a "push" request or to initiate a "pull" request by the immersive client 908. In an example, the immersive client 908 can use the network interface 908B, a GPU (or CPU, not shown) 908C, and storage 908D. Additionally, the immersive client 908 may utilize a game engine 908A. The game engine 908A may further utilize a visualization component 908A1 and a physics engine 908A2. The game engine 908A communicates with a smart client 908E to orchestrate the processing of the media via a game engine API and callback functions 908F. The delivery format of the media is stored in a storage device, or storage cache 908D, of the client 908. Finally, the client 908 visually presents the media via a visualization component 908A1.
[0164] Throughout the process of streaming immersive media to the immersive client 908, the network orchestrator 905 can monitor the status of the client's progress via a client progress and status feedback channel 907. The status monitoring can be performed by a two-way communication message interface (not shown in FIG. 9), which can be implemented in the smart client 908E.
[0165] FIG. 10 shows a diagram of a media adaptation process 1000 so that, in some examples, ingested source media can be appropriately adapted to match the requirements of an immersive client 908. The media adaptation and fragmentation module 1001 consists of multiple components that facilitate the adaptation of ingested media to an appropriate delivery format for the immersive client 908. In FIG. 10, the media adaptation and fragmentation module 1001 receives input network status 1005 to track the current traffic load on the network. The immersive client 908 information can include attributes and feature descriptions, application features and descriptions, and the current status of the application, as well as a client neural network model (if available) that helps map the client's frustum geometry to the interpolation functions of the ingested immersive media. Such information can be obtained through a two-way message interface (not shown in FIG. 10) with the assistance of a smart client interface, shown as 908E in FIG. 9. The media adaptation and fragmentation module 1001 ensures that, once the adapted output is created, it is stored in the client-adapted media storage device 1006. The media analyzer 1007 is shown in Figure 10 as a process that may be executed appli- cably or as part of a network automation process for the distribution of media. In some examples, the media analyzer 1007 includes a scene analyzer with optional IMS notation capabilities.
[0166] In some examples, the network orchestrator 1003 initiates the adaptation process 1001. The media adaptation and fragmentation module 1001 is controlled by the logic controller 1001F. In examples, the media adaptation and fragmentation module 1001 uses a renderer 1001B or a neural network processor 1001C to adapt specific ingest source media to a format appropriate for the client. In examples, the media adaptation and fragmentation module 1001 receives client information 1004 from a client interface module 1003, such as a server device in one example. The client information 1004 may include a client description and current status, an application description and current status, and a client neural network model. The neural network processor 1001C uses the neural network model 1001A. Examples of such neural network processors 1001C include deep view neural network model generators such as those described in MPI and MSI. In some examples, the media is in a 2D format but the client requires a 3D format, and then the neural network processor 1001C can invoke a process that uses highly correlated images from the 2D video signal to derive a stereoscopic representation of the scene depicted in the video. An example of such a process could be the neural radiation field from one or a few images process developed at the University of California, Berkley. One example of a suitable renderer 1001B could be a modified version of the OTOY Octane renderer (not shown) that is modified to interact directly with the media adaptation and fragmentation module 1001. The media adaptation and fragmentation module 1001, in some examples, can use a media compressor 1001D and a media decompressor 1001E, depending on the needs of these tools regarding the format of the ingested media and the format required by the immersive client 908.
[0167] FIG. 11 illustrates a delivery format creation process 1100 in some examples. An adaptive media packaging module 1103 packages media from a media adaptation module 1101 (shown as process 1000 in FIG. 10) currently residing on a client adaptive media storage device 1102. The media packaging module 1103 formats the adapted media from the media adaptation module 1101 into a robust delivery format 1104, such as the exemplary formats shown in FIG. 3 or 4. Manifest information 1104A provides the client 908 with a list 1104B of scene data assets it can expect to receive, as well as optional metadata describing the frequency with which each asset is used across the set of scenes comprising the presentation. List 1104B illustrates a list of visual, audio, and haptic assets, each with corresponding metadata. In this exemplary embodiment, each asset in list 1104B references metadata including a numerical frequency value indicating the number of times a particular asset is used across all scenes comprising the presentation.
[0168] Figure 12 illustrates some example packetizer processing system 1200. In the example of Figure 12, packetizer 1202 fragments adapted media 1201 into individual packets 1203 suitable for streaming to immersive client 908, shown as client endpoint 1204 on a network.
[0169] FIG. 13 illustrates a sequence diagram 1300 of a network that, in some examples, adapts particular immersive media in an ingest format into a streamable and appropriate delivery format for a particular immersive media client endpoint.
[0170] The components and communications shown in FIG. 13 are described as follows: Client 1301 (in some examples, also referred to as a client endpoint or client device) initiates a media request 1308 to network orchestrator 1302 (in some examples, also referred to as a network delivery interface). Media request 1308 includes information identifying the media requested by client 1301, either by URN or other standard nomenclature. Network orchestrator 1302 responds to media request 1308 with a profile request 1309, which requests that client 1301 provide information about its currently available resources (including computation, storage, battery charge, and other information to characterize the client's current operating status). Profile request 1309 also requests that the client provide one or more neural network models, if such models are available at the client, that can be used by the network for neural network inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system. Response 1310 from client 1301 to network orchestrator 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available to the client). Network orchestrator 1302 then provides session ID token 1311 to client 1301. Network orchestrator 1302 then requests ingest media server 1303 with ingest media request 1312, which includes the URN or canonical nomenclature name of the media identified in request 1308. Ingest media server 1303 responds to request 1312 with response 1313, which includes the ingest media tokens. Network orchestrator 1302 then provides the media tokens from response 1313 to client 1301 in call 1314.The network orchestrator 1302 then initiates the adaptation process for the media request 1308 by providing the adaptation interface 1304 with the ingest media token, client token, application token, and neural network model token 1315. The adaptation interface 1304 requests access to the ingest media by providing the ingest media token to the ingest media server 1303 in a call 1316 requesting access to the ingest media asset. The ingest media server 1303 responds to the request 1316 with the ingest media access token in a response 1317 to the adaptation interface 1304. The adaptation interface 1304 then requests that the media adaptation module 1305 adapt the ingest media located in the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created in 1313. A request 1318 from the adaptation interface 1304 to the media adaptation module 1305 includes the necessary tokens and session ID. Media adaptation module 1305 provides the adapted media access token and updated session ID 1319 to network orchestrator 1302. Network orchestrator 1302 provides the adapted media access token and session ID to packaging module 1306 in interface call 1320. Packaging module 1306 provides response message 1321 to network orchestrator 1302 with the packaged media access token and session ID in response 1321. Packaging module 1306 provides the packaged asset, URN, and packaged media access token for the session ID to packaged media server 1307 in response 1322.Client 1301 executes request 1323 to begin streaming the media asset corresponding to the packaged media access token received in response message 1321. Client 1301 executes other requests and provides status updates in message 1324 to network orchestrator 1302.
[0171] FIG. 14 shows a diagram of a media system 1400 with a virtual network and client devices 1418 (also referred to as game engine client devices 1418) for scene-based media processing in some examples. A smart client, such as that represented by MPEG Smart Client Process 1401 in FIG. 4, can act as a central coordinator for preparing media to be processed by and for other entities within the game engine client device 1418, as well as for entities residing outside the game engine client device 1418. In some examples, a smart client is implemented as software instructions that can be executed by processing circuitry to run a process such as MPEG Smart Client Process 1401 (also referred to as MPEG Smart Client 1401 in some examples). The game engine 1405 is primarily responsible for rendering media to create a presentation experienced by an end user. A haptic component 1413, a visualization component 1415, and an audio component 1414 can assist the game engine 1405 in rendering haptic, visual, and audio media, respectively. The edge processor or network orchestrator device 1408 can communicate information and system media to, and receive status updates and other information from, the MPEG Smart Client 1401 via a network interface protocol 1420. The network interface protocol 1420 can be divided across multiple communication channels and processes and can use multiple communication protocols. In some examples, the game engine 1405 is a game engine device that includes control logic 14051, a GPU interface 14052, a physics engine 14053, a renderer 14054, a compression decoder 14055, and device-specific plug-ins 14056. The MPEG Smart Client 1401 also serves as the primary interface between the network and the client device 1418.For example, MPEG Smart Client 1401 can interact with game engine 1405 using game engine API and callback functions 1417. In an example, MPEG Smart Client 1401 can be responsible for reconstructing the streaming media communicated at 1420 before invoking game engine API and callback functions 1417 managed by game engine control logic 14051 to have game engine 1405 process the reconstructed media. In such an example, MPEG Smart Client 1401 can utilize client media reconstruction process 1402, which can utilize sequential compression decoder process 1406.
[0172] In some other examples, MPEG Smart Client 1401 may not be responsible for reassembling the packetized media streamed at 1420 before invoking API and callback functions 1417. In such examples, game engine 1405 may decompress and reassemble the media. Further, in such examples, game engine 1405 may decompress the media using compression decoder 14055. Upon receiving the reassembled media, game engine control logic 14051 may render the media via renderer process 14054 using GPU interface 14052.
[0173] In some examples, the rendered media is animated, and the physics engine 14053 may then be used by the game engine control logic 14051 to simulate the laws of physics in the animation of the scene.
[0174] In some examples, throughout the processing of media by client device 1418, neural network model 1421 may be employed by neural network processor 1403 to assist in operations coordinated by MPEG Smart Client 1401. In some examples, reconstruction process 1402 may require fully reconstructing the media using neural network model 1421 and neural network processor 1403. Similarly, client device 1418 may be configured by a user via user interface 1412 to cache media received from the network in client adaptive media cache 1404 after the media has been reconstructed, or to cache rendered media in rendered client media cache 1407 after the media has been rendered. Additionally, in some examples, MPEG Smart Client 1401 may replace system-provided visual / non-visual assets with user-provided visual / non-visual assets from user-provided media cache 1416. In such an embodiment, user interface 1412 may guide an end user to perform steps to load user-provided visual / non-visual assets from user-provided media cache 1419 (e.g., external to client device 1418) into client-accessible user-provided media cache 1416 (e.g., internal to client device 1418). In some embodiments, MPEG Smart Client 1401 may be configured to store rendered assets (for potential reuse or sharing with other clients) in rendered media cache 1411.
[0175] In some examples, the media analyzer 1410 can examine the client adaptive media 1409 (in the network) to determine asset complexity or the frequency with which assets are reused across one or more scenes (not shown) for potential prioritization for rendering by the game engine 1405 and / or reconstruction processing via the MPEG smart client 1401. In such examples, the media analyzer 1410 stores complexity, prioritization, and asset usage frequency information in the media stored in 1409.
[0176] It should be noted that although processes are shown and described in this disclosure, the processes may be implemented as instructions in a software module, and the instructions may be executed by a processing circuit to perform the process. It should also be noted that although modules are shown and described in this disclosure, the modules may be implemented as software modules having instructions, and the instructions may be executed by a processing circuit to perform the process.
[0177] In some related examples, various techniques can address the problem of providing a smooth flow of scenes to clients, including the use of content delivery networks (CDNs) and edge network elements to reduce the latency between a client device's request for a scene and the scene's appearance at the client device, and the use of cloud / edge network elements to offload the computational burden of rendering to more powerful computation engines. While these techniques can reduce the latency between a client device's request for a scene and the scene's appearance at the client device, they rely on the availability of a CDN to achieve, and immersive scene providers, with or without the use of a CDN and network edge devices, can rely on end devices to request a scene before the scene's data can be streamed to the end device.
[0178] According to a first aspect of the present disclosure, a prioritized scene streaming method for immersive media streaming can be used for immersive media streaming for light-field-based displays (e.g., light-field displays, holographic displays, etc.). In some examples, scene priority is determined by the relationship between the scenes.
[0179] In some examples, during operation, a media server may use adaptive streaming to stream media data to an end device, such as an end device having a light field-based display, in which the order in which scenes are made available may not provide an acceptable user experience if the end device requests the entire immersive environment at once.
[0180] FIG. 15 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 15, the scene-based immersive media is represented by a scene manifest 1501 that includes a list of individual scenes to be rendered for the immersive experience. The scene manifest 1501 is provided to the end device 1504 via a cloud 1503 (e.g., a network cloud) by a media server 1502. In the example of FIG. 15, the scenes are retrieved by and sent to the end device 1504 in the order of their scene numbers in the scene manifest 1501. For example, the scenes in the scene manifest 1501 are ordered by scene 1, scene 2, scene 3, scene 4, scene 5, etc. The scenes are captured in the order of scene 1 (indicated by 1505), scene 2 (indicated by 1506), scene 3 (indicated by 1507), scene 4 (indicated by 1508), scene 5 (indicated by 1509), etc. and transmitted to end device 1504.
[0181] In some examples, each scene in a scene manifest may include, for example in the scene's metadata, a complexity value indicating the complexity of the scene and a priority value indicating the priority of the scene.
[0182] FIG. 16 illustrates a diagram for adding complexity values and priority values to scenes in a scene manifest in some examples. In the example of FIG. 16, a scene complexity analyzer 1602 is used to analyze the complexity of each scene and assign each scene a complexity value, and then a scene prioritization unit (1603) is used to assign each scene a priority value. For example, a first scene manifest 1601 includes multiple scenes to be rendered for an immersive experience. The first scene manifest 1601 is processed by the scene complexity analyzer 1602, which determines a complexity value (denoted as SC) for each of the multiple scenes, and then processed by the scene prioritization unit 1603, which assigns a priority value (denoted as Priority) to each of the multiple scenes, resulting in a second scene manifest 1604. The second scene manifest 1604 includes multiple scenes with respective complexity values and priority values. The multiple scenes in the second scene manifest 1604 are reordered according to the complexity values and / or priority values. For example, scene 5 is ordered to be the fifth scene in the first scene manifest 1601 (as shown by 1605) and is reordered to be the third scene in the second scene manifest 1604 according to its priority value (as shown by 1606).
[0183] Additionally, in some examples, a priority-aware media server may be used to stream scenes to an end device in order according to the priority value of the scenes.
[0184] FIG. 17 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 17, the scene-based immersive media is represented by a scene manifest 1701 that includes a list of individual scenes to be rendered for the immersive experience. Each of the multiple scenes has an analyzed complexity value (denoted by SC) and an assigned priority value (denoted by Priority). The scene manifest 1701 is provided to the end device 1704 via a cloud 1703 by a priority-aware media server 1702. In the example, the scenes are retrieved by and sent to the end device 1704 in order according to the priority values in the scene manifest 1701. For example, the scenes in the scene manifest 1701 are ordered according to the priority values. The scenes are captured in the order of scene 7 (indicated by 1705), scene 8 (indicated by 1706), scene 5 (indicated by 1707), scene 2 (indicated by 1708), scene 11 (indicated by 1709), etc. and transmitted to end device 1704.
[0185] According to aspects of the present disclosure, scene priority-based streaming techniques can be used to provide advantages compared to requesting all scenes in an immersive environment simultaneously. This is because there are often reasons why an end device may need to render scenes “out of order.” For example, scenes may be initially ordered in the most useful order typically used in immersive environments, or the immersive environment may contain branches, loops, etc. that may require scenes to be rendered out of the most useful order. In an example, there may be sufficient bandwidth to potentially provide an acceptable user experience when scenes are transferred in the most useful order, with a required scene being queued behind one or more scenes that have already been requested. In another example, when bandwidth is limited, scene priority-based streaming techniques may be more useful for improving the user experience.
[0186] It should be noted that the scene prioritization unit may use any suitable technique for determining the priority values of the scenes, and in some examples, the scene prioritization unit may use relationships between scenes as part of the prioritization.
[0187] 18 shows a diagram of a virtual art museum 1801 to illustrate scene priorities in some examples. The virtual art museum 1801 has a first floor 1802 and an upper floor 1803, and a stairwell 1804 between the first floor 1802 and the upper floor 1803. In the example, the immersive tour begins in a lobby 1805. The lobby 1805 has five exits: a first exit for exiting the virtual art museum 1801, a second exit for accessing the stairwell 1804, a third exit for accessing first floor exhibit E, a fourth exit for accessing first floor exhibit D, and a fifth exit for accessing first floor exhibit H on the first floor 1802. In the example, first floor exhibits E, D, and H are possible "next" scenes to be rendered when the virtual tourist exits lobby 1805, so the scene prioritization unit may reasonably prioritize a first scene in lobby 1805, then a second scene in stairwell 1804, then a third scene in first floor exhibits E, D, and H.
[0188] In some examples, the scene prioritization unit may adjust the priorities of scenes based on previous experiences with the same client or other clients. Referring to FIG. 18 , in an example, virtual museum 1801 opens a particularly popular exhibit (e.g., upper floor exhibit I) to upper floor exhibition room 1808, and the scene prioritization unit may observe that many virtual tourists head straight for the popular exhibit. The scene prioritization unit may then adjust the priorities previously provided. For example, the scene prioritization unit may prioritize a first scene in lobby 1805, a second scene in stairwell 1804, and scenes on upper floor 1803 on the way to upper floor exhibition room 1808, such as the scene of upper floor exhibit J and the scene of upper floor exhibit I.
[0189] In some examples, the content description provided by a media server for immersive media may include two parts: a media presentation description (MPD) that describes a manifest of available scenes, various alternatives, and other characteristics, as well as multiple scenes with various assets. In some examples, an end device may first obtain the MPD of the immersive media to play. The end device may parse the MPD and learn about various scenes with various assets, scene timing, media content availability, media types, various encoding alternatives for the media content, minimum and maximum bandwidths supported, and other content characteristics. Using the information obtained from the MPD, the end device may appropriately select which scene to render at what time and with what available bandwidth. The end device may continuously measure bandwidth fluctuations, and depending on the measurements, the end device may determine how to adapt to the available bandwidth by fetching alternative scenes with fewer or more assets.
[0190] In some examples, the priority value of a scene of scene-based immersive media may be defined by a server device or sender device and changed by an end device during a session for playing the scene-based immersive media.
[0191] According to a second aspect of the present disclosure, scene priority values of scene-based immersive media for a light field-based display can be dynamically adjusted. In some examples, the priority values can be updated based on feedback from end devices (also referred to as client devices). For example, a scene prioritizer can receive feedback from the end devices and dynamically change the scene priority values based on the feedback.
[0192] FIG. 19 shows a diagram of scenes in scene-based immersive media, such as immersive game 1901 in some examples. Scene relationships are indicated by dotted lines connecting the scenes. In the example, game character 1902 appears in spawn area 1903, and the scene with the highest priority is then the scene in spawn area 1903. The scene in spawn area 1903 can be followed by scenes directly connected to spawn area 1903, such as scene 1, scene 11, and scene 20, from which game character 1902 departs. When game character 1902 moves to scene 1 (indicated by 1904), scenes directly connected to scene 1, such as scene 2 (indicated by 1905), now have the highest priority. When game character 1902 moves to scene 2, scenes directly connected by scene 2, such as scene 12 and scene 5, now have the highest priority.
[0193] In some examples, when a game character 1902 moves to a new scene that has not yet been streamed to the end device, the end device can send the game character's current location to the scene prioritization unit so that the current location can be included in the prioritization.
[0194] FIG. 20 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 20, the scene-based immersive media is represented by a scene manifest 2001 that includes a list of individual scenes to be rendered for the immersive experience. Each of the multiple scenes has an analyzed complexity value (denoted by SC) and an assigned priority value (denoted by Priority). The scene manifest 2001 is provided to an end device 2004 via a cloud 2003 by a priority-aware media server 2002. In the example of FIG. 20, the scenes are retrieved by and sent to the end device 2004 in order of priority in the scene manifest 2001.
[0195] Further, taking the immersive game example of Figure 19, the end device 2004 can send the current location 2010 of the game characters to the scene prioritization unit 2011. The scene prioritization unit 2011 can then include the current location information in its calculations for updates to scene priorities in the scene manifest 2001.
[0196] It should be noted that the techniques for adaptive streaming of light field-based media described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media.
[0197] In some examples, a method for adaptive streaming for a light field-based display includes assigning respective priority values to multiple scenes of immersive media for the light field-based display and adaptively streaming the multiple scenes based on the priority values. In examples, the multiple scenes are transferred (e.g., streamed) based on the assigned priority values. In some examples, the priority values of the scenes are assigned based on a likelihood of a need to render the scenes. In some examples, the priority values of the scenes are dynamically adjusted based on feedback from an end device. For example, the priority values of the scenes can be adjusted based on, for example, the position of a character in a game, i.e., feedback from a client device.
[0198] In some examples, in response to available network bandwidth being limited, a scene having a highest priority is selected to be streamed as the next scene. In some examples, in response to available network bandwidth being limited, the method includes identifying a plurality of scenes that are unlikely to be needed and refraining from streaming the identified plurality of scenes until the identified plurality of scenes exceeds a threshold of likelihood that they will be needed.
[0199] 21 shows a flowchart outlining a process 2100 according to an embodiment of the present disclosure. The process 2100 may be performed in a network, such as by a server device of the network. In some embodiments, the process 2100 is implemented in software instructions, and thus, when a processing circuit executes the software instructions, the processing circuit performs the process 2100. The process begins at S2101 and proceeds to S2110.
[0200] At S2110, scene-based immersive media for playback on a light field-based display is received, the scene-based immersive media including a plurality of scenes.
[0201] At S2120, a priority value is assigned to each of a plurality of scenes of the scene-based immersive media.
[0202] At S2130, the order in which the multiple scenes are streamed to the end device in a light field-based display is determined according to the priority values.
[0203] In some examples, multiple scenes have an initial order in the scene manifest, are re-ordered according to priority values, and the re-ordered scenes are transmitted (streamed) to the end device.
[0204] In some examples, the priority-aware network device may select the highest priority scene having the highest priority value from a subset of untransmitted scenes among the plurality of scenes and transmit the highest priority scene to the end device.
[0205] In some examples, the priority value for a scene is determined based on the likelihood of needing to render the scene.
[0206] In some examples, it is determined that available network bandwidth is limited, and a highest priority scene having a highest priority value is selected from a subset of untransmitted scenes of the plurality of scenes, and the highest priority scene is transmitted in response to the limited available network bandwidth.
[0207] In some examples, available network bandwidth is determined to be limited, a subset of the plurality of scenes that is unlikely to be needed for rendering next is identified based on the priority values, and the subset of the plurality of scenes is refrained from streaming in response to the limited available network bandwidth.
[0208] In some examples, a first priority value is assigned to a first scene based on the second scene having a second priority value depending on the relationship between the first scene and the second scene.
[0209] In some examples, a feedback signal is received from the end device. Based on the feedback signal, a priority value of at least one scene among the plurality of scenes is adjusted. In an example, in response to the feedback signal indicating that the current scene is a second scene related to the first scene, the first scene is assigned the highest priority. In another example, the feedback signal indicates an adjustment of the priority determined by the end device.
[0210] The process 2100 then proceeds to step S2199 and ends.
[0211] Process 2100 can be appropriately adapted to various scenarios, and the steps of process 2100 can be adjusted accordingly. One or more of the steps of process 2100 can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform process 2100. Additional steps can be added.
[0212] 22 shows a flowchart outlining a process 2200 according to an embodiment of the present disclosure. The process 2200 may be performed by an end device (also referred to as a client device). In some embodiments, the process 2200 is implemented in software instructions, and thus, when a processing circuit executes the software instructions, the processing circuit performs the process 2200. The process starts at S2201 and proceeds to S2210.
[0213] At S2210, an end device having a light field-based display receives a Media Presentation Description (MPD) of scene-based immersive media for playback by the light field-based display, the scene-based immersive media including multiple scenes, and the MPD indicates that the multiple scenes are to be streamed sequentially to the end device.
[0214] At S2220, bandwidth availability is detected.
[0215] At S2230, a reordering of at least one scene is determined based on bandwidth availability.
[0216] At S2240, a feedback signal is sent indicating a reordering of at least one scene.
[0217] In some examples, the feedback signal indicates the next scene to render.
[0218] In some examples, the feedback signal indicates an adjustment to the priority value of at least one scene.
[0219] In some examples, the feedback signal indicates the current scene.
[0220] Process 2200 then proceeds to S2299 and ends.
[0221] Process 2200 may be appropriately adapted to various scenarios, and the steps of process 2200 may be adjusted accordingly. One or more of the steps of process 2200 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 2200. Additional steps may be added.
[0222] Some aspects of the present disclosure also provide prioritized asset streaming methods for immersive media streaming, such as immersive media streaming for light field-based displays.
[0223] Note that the appearance of a scene at a client device (also referred to as an end device) in response to a client request from the client device may depend on the availability of assets that the scene relies on the client device to enable the client device to render the scene. To minimize the amount of time required to transfer assets to the client device, in some examples, assets may be prioritized.
[0224] According to a third aspect of the present disclosure, scene-based immersive media assets, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), can be prioritized based on asset attributes, such as asset size of the assets.
[0225] In some examples, in operation, a media server can use adaptive streaming for light field-based displays, which relies on the transfer of scenes to end devices, which in turn relies on the transfer of assets that make up each scene.
[0226] 23 shows a diagram of a scene manifest 2301 with a mapping of scenes to assets in an example. Scene manifest 2301 includes multiple scenes 2303, each of which depends on one or more assets in set of assets 2302. For example, scene 2 depends on assets B, D, E, F, and G in set of assets 2302.
[0227] In some examples, as an end device requests assets for a scene, the assets are streamed to the end device in order.
[0228] FIG. 24 shows a diagram of streaming assets for a scene in some examples. In FIG. 24, a scene 2401 depends on multiple assets, such as asset A (shown at 2405), asset B (shown at 2406), asset C (shown at 2407), and asset D (shown at 2408). The multiple assets are provided by a media server 2402 to an end device 2404 via a cloud 2403. In the example of FIG. 24, the multiple assets for the scene 2401 are ordered as follows: asset A, asset B, asset C, and asset D. The multiple assets are retrieved and sent to the end device 1504 according to the order of the assets for the scene 2401, such as the order asset A, asset B, asset C, asset D. The order in which the assets are made available to the end device may not provide an acceptable user experience.
[0229] According to aspects of the present disclosure, before assets for a scene are streamed to an end device, an asset prioritization unit can be used to analyze the assets required for the scene and assign a priority value to each asset.
[0230] FIG. 25 shows a diagram for re-prioritizing assets of a scene in some examples. In the example of FIG. 25, an asset prioritization unit 2502 is used to assign a priority value to each asset in the scene. For example, a scene depends on four assets called asset A (shown at 2504), asset B (shown at 2505), asset C (shown at 2506), and asset D (shown at 2507). Initially, as shown at 2501, the four assets are in asset order: asset A, asset B, asset C, and asset D. The asset prioritization unit 2502 assigns a first priority value (e.g., 2) to asset A, a second priority value (e.g., 4, the highest priority value) to asset B, a third priority value (e.g., 1, the lowest priority value) to asset C, and a fourth priority value (e.g., 3) to asset D. The four assets may be re-ordered according to priority values, such as asset B, asset D, asset A, and asset C, as shown at 2503 .
[0231] In an example, the re-ranked assets may be streamed to an end device by a media server.
[0232] It should be noted that although higher values are used in this disclosure to represent higher priorities, other suitable priority numbering techniques may be used to represent asset priorities.
[0233] 26 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 26, a scene of scene-based immersive media includes four assets re-ordered according to the priority values of the four assets, such as from highest to lowest: asset B, asset D, asset A, and asset C, as shown at 2601. The four assets are provided from a media server 2602 to an end device 2604 via a cloud 2603 in order of the priority values. For example, the four assets are obtained and sent to the end device 2604 in the following order: asset B (shown at 2605), asset D (shown at 2606), asset A (shown at 2607), and asset C (shown at 2608).
[0234] According to aspects of the present disclosure, the asset prioritizer 2502 can assign priority values to assets based on the size in bytes of each asset. In an example, the smallest assets are assigned the highest priority value, so that a series of assets can be rendered in an order that more assets (e.g., smaller assets) are presented to the end user more quickly while the end device is still waiting for the larger assets to arrive. Using FIG. 26 as an example, asset B has the fewest number of bytes (smallest size), and asset C has the most number of bytes (largest size). Comparing the first streaming order of FIG. 24 with the second streaming order of FIG. 26, in FIG. 24, the larger assets are streamed first, and end device 2404 may start presenting assets earlier than end device 2404 because end device 2404 may not need to present anything until the larger assets arrive.
[0235] In some examples, the prioritization assigned by the asset prioritization unit may also be included in the manifest of the MPD. In some examples, the asset priority values of assets in a scene of the scene-based immersive media may be defined by a server device or a sender device, and the asset priority values may be changed by a client device during a session for playing the scene-based immersive media.
[0236] Some aspects of the present disclosure provide methods for dynamically adaptive streaming for a light-field-based display (e.g., a light-field display or a holographic display) according to assigned asset priority values. Assets that have been assigned asset priority values can be transmitted (streamed) based on the asset priority values. In some examples, asset priority values are assigned to assets based on asset size. In some examples, asset priority values are adjusted based on feedback from one or more client devices.
[0237] According to a fourth aspect of the present disclosure, assets of scene-based immersive media, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), may be prioritized based on the visibility of the assets.
[0238] According to aspects of the present disclosure, not all assets in a stream are equally relevant at any given time because some assets may not be visible to an end user at some point in time. In some examples, prioritizing visible assets over hidden assets in a stream can reduce the delay before a relevant asset (e.g., a visible asset) is available to an end device (e.g., a client device). In some examples, the priority value of an asset is determined according to the visibility of the asset from a camera position, such as a default entry position for a scene.
[0239] In some examples, when a virtual person associated with an end user (also referred to as a camera in some examples) enters a scene at an entry position (e.g., the scene's default entry position), some assets of the scene are immediately visible, while other assets may not be visible when viewed from the entry position or when the camera is positioned at the entry position.
[0240] Figure 27 shows a diagram of a scene 2701 in some examples. A camera represents a virtual person at a viewing position corresponding to the end user. The camera enters the scene 2701 at a default entry position 2702. Scene 2701 includes Asset A, Asset B, Asset C, and Asset D. Asset A is a set of floor-to-ceiling bookshelves, as shown at 2703. Asset B is a computer desk, as shown at 2704. Asset C is a conference table, as shown at 2705. Asset D is a doorway to another scene, as shown at 2706.
[0241] In the example of Figure 27, asset B is completely obscured by asset A, so asset B is not visible from the default position 2702. In Figure 27, asset B is shaded to indicate that asset B is not visible to the camera at the default entry position 2702.
[0242] In some examples, before the assets of a scene are streamed to an end device, an asset prioritization unit is used to analyze the assets of the scene and assign a priority value to each asset according to the asset's visibility relative to the scene's default entry location.
[0243] FIG. 28 shows a diagram for re-prioritizing assets of a scene (e.g., scene 2701) in some examples. In the example of FIG. 28, an asset prioritization unit 2802 is used to assign a priority value to each asset of the scene. For example, a scene (e.g., scene 2701 in FIG. 27) depends on asset A (shown as 2804), asset B (shown as 2805), asset C (shown as 2806), and asset D (shown as 2807). Initially, as shown in 2801, the four assets are in the asset order of asset A, asset B, asset C, and asset D. The asset prioritization unit 2802 analyzes the visibility of the assets according to the default entry position for entering the scene. For example, assets A, C, and D are visible from the default entry position, and asset B is not visible from the default entry position. The asset prioritization unit 2802 assigns priority values to the assets according to their visibility. For example, the asset prioritization unit 2802 assigns high priority values to asset A, asset C, and asset D, and assigns a low priority value to asset B. The four assets can be re-ordered according to the priority values, such as in the order asset A, asset C, asset D, and asset B, as shown in 2803.
[0244] In an example, the re-ranked assets may be streamed to an end device by a media server.
[0245] 29 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 29, a scene (e.g., scene 2701) of scene-based immersive media includes four assets re-ordered according to the priority values of the four assets, such as from highest to lowest: asset A, asset C, asset D, and asset B, as shown at 2901. The four assets are provided from a media server 2902 to an end device 2904 via a cloud 2903 in order of the priority values. For example, the four assets are obtained and sent to the end device 2904 in the following order: asset A (shown at 2905), asset C (shown at 2906), asset D (shown at 2907), and asset B (shown at 2908).
[0246] According to aspects of the present disclosure, the asset prioritization unit 2802 can assign priority values to assets in a scene based on the visibility of the assets. In the example, visible assets are assigned the highest priority value, so they are streamed first and can reach the end user for rendering more quickly. Comparing the first streaming order of FIG. 24 with the second streaming order of FIG. 29, the end device 2904 can begin showing the visible assets in a scene earlier than the end device 2404.
[0247] In some examples, the asset prioritization unit 2802 can assign priority values to assets based on whether the asset is blocked from the camera view by another asset. Priority values can be assigned according to a "visible assets first, all visible assets then non-visible assets" policy, so that a series of assets can be rendered in an order that more quickly presents all of the visible assets to the end user.
[0248] In some examples, other assets are between this asset and the light source, so the asset prioritization unit 2802 can assign a priority value to the asset based on whether the other assets occlude this asset. The priority values can be assigned according to a policy of "visible assets first, all visible assets then non-visible assets," so that the set of assets can be rendered in an order that more quickly presents all of the visible assets to the end user.
[0249] FIG. 30 shows a diagram of a scene 3001 in some examples. Scene 3001 is similar to scene 2701 in FIG. 27. A camera represents a virtual person at a viewing position corresponding to an end user. The camera enters scene 3001 at default entry position 3002. Scene 3001 includes asset A, asset B, asset C, and asset D. Asset A is a set of floor-to-ceiling bookshelves, as shown at 3003. Asset B is a computer desk, as shown at 3004. Asset C is a conference table, as shown at 3005. Asset D is a doorway to another scene, as shown at 3006. Asset B is completely obscured by asset A, so asset B is not visible from default entry position 3002. In FIG. 30, asset B is shaded to indicate that asset B is not visible to the camera at default entry position 3002.
[0250] 30 further shows light source 3007. Due to the position of light source 3007, conference table 3005 and doorway 3006 are both blocked by the shadow of bookshelf 3003 (e.g., between 3008 and 3009). Therefore, asset C and asset D are not visible to the camera, and asset prioritization unit 2802 may, in the example, lower the priority of asset C and asset D.
[0251] In some examples, a media server may have two parts: a media presentation description (MPD) that describes a manifest of available scenes, various alternatives, and other characteristics, as well as a content description of multiple scenes with various assets. Prioritization assigned by the asset prioritization unit may also be included in the manifest of the MPD.
[0252] In some examples, a priority value of an asset of a scene of scene-based immersive media may be defined by a server device or a sender device. The priority value of the asset may be provided to a client device in an MPD. In examples, the priority value may be changed by a client device during a session of playing the scene-based immersive media.
[0253] Some aspects of the present disclosure provide a method for dynamically adaptive streaming for a light field-based display according to assigned asset priority values. Assets can be transmitted (e.g., streamed) based on the assigned asset priority values. In some examples, the asset priority value assigned to an asset is lowered in response to the asset being blocked from the camera view by one or more other assets and therefore not visible. In some examples, the asset priority value for an asset is lowered when the asset is occluded from one or more light sources by one or more other assets and is not visible to the camera. In some examples, the priority value (asset priority value) is adjusted based on feedback from a client device, such as an updated camera position.
[0254] According to a fifth aspect of the present disclosure, scene-based immersive media assets, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), may be prioritized based on multiple policies.
[0255] According to aspects of the present disclosure, all assets of a stream do not arrive at the end device at the same moment in time for a variety of reasons, including but not limited to, differences in asset size. Furthermore, not all assets are equally valuable at any given time for a variety of reasons, including but not limited to, some assets being occluded behind other assets. For a variety of reasons, a client may desire to prioritize assets.
[0256] In an example, a client end device, given a sufficiently detailed manifest describing the assets used in a scene, may have sufficient local resources to perform all asset prioritization itself. In many foreseeable situations, this may not be the case for various reasons (e.g., available processor power or the need to offload asset prioritization and preserve available battery life). Some aspects of the present disclosure provide ways to prioritize assets for streaming on behalf of a client end device. For example, asset priority values are assigned externally to the client end device and determined by one or more prioritization schemes.
[0257] In some examples, before the assets of a scene are streamed to an end device, an asset prioritization unit is used to analyze the assets required for the scene and assign a priority to each asset.
[0258] FIG. 31 shows a diagram for assigning priority values to assets of a scene in some examples. In the example of FIG. 31, an asset prioritization unit 3102 is used to assign one or more priority values to each asset of a scene. For example, a scene (e.g., scene 2701 of FIG. 27) depends on asset A (indicated by 3105), asset B (indicated by 3106), asset C (indicated by 3107), and asset D (indicated by 3108). Initially, as shown by 3101, the four assets are in the asset order of asset A, asset B, asset C, and asset D. The asset prioritization unit 3102 assigns priority values to assets according to multiple policies. The assets are re-ordered according to the assigned priority values. In FIG. 31, as shown by 3103, assets A to D are re-ordered according to the assigned priority values in the order of asset A, asset C, asset D, and asset B. Also note that each asset in 3103 now holds one or more priority values assigned by the asset prioritization unit 3102, such as two priority values denoted by +P1+P2 for each asset. In an example, the asset prioritization unit 3102 may assign a first set of priority values, such as P1, to each asset in a scene according to a first priority policy, and a second set of priority values, such as P2, to each asset in a scene according to a second priority policy. The assets may be re-ranked according to the first set of priority values, or may be re-ranked according to the second set of priority values.
[0259] In some examples, after the assets are prioritized, the assets are requested and streamed in order of priority.
[0260] 32 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 32, a scene of scene-based immersive media includes four assets re-ordered according to the priority values of the four assets, such as from highest to lowest: asset A, asset C, asset D, and asset B, as shown at 3201. The four assets are provided from a media server 3202 to an end device 3204 via a cloud 3203 in order of the priority values. For example, the four assets are obtained and sent to the end device 3204 in the order of asset A (shown at 3205), asset C (shown at 3206), asset D (shown at 3207), and asset B (shown at 3208).
[0261] According to aspects of the present disclosure, when an end device requests assets for a scene in prioritized order, the end device may be able to begin rendering the scene more quickly.
[0262] It should be noted that the asset prioritizer 3102 may use any ranking criteria that provides an improved usage experience as the basis for assigning priority values to assets.
[0263] In some examples, the asset prioritization unit 3102 can use multiple prioritization schemes and assign multiple priority values to each asset, allowing the media server 3202 to use the most appropriate prioritization scheme for each end device 3204.
[0264] In an example, the end device 3204 is at the end of a relatively long and / or low bitrate network path from the media server 3202, and the end device 3204 can benefit from a prioritization scheme that ranks assets from smallest to largest file size, so that the end device can begin rendering more assets quickly without waiting for larger assets to arrive. In another example, the end device 3204 is at the end of a shorter, faster network path from the media server 3202, and the asset prioritizer 3102 can assign a separate priority to each asset based on whether the asset is blocked from the camera view by another asset. This priority policy can be as simple as "visible assets first, all visible assets then non-visible assets," and the media server 3202 can provide the series of assets in an order that allows the end device to more quickly present all visible assets to the end user.
[0265] In some examples, the media server 3202 may include the assigned asset priority value in the description (e.g., MPD) transferred to the end device 3204. The client of the end device 3204 may use additional prioritization schemes not supported by the asset prioritization unit 3102, relying on the assigned priority value in the description as a tiebreaker if two or more assets have the same client-assigned priority value (e.g., a prioritization scheme with a small number of distinct possible values such as "visible" / "invisible").
[0266] Some aspects of the present disclosure provide a method for dynamic adaptive streaming for a light field-based display based on assigned asset priority values. In some examples, the order of assets in asset streaming may be based on one of multiple assigned asset priority values assigned using different asset prioritization schemes. In some examples, the same set of assets may be transferred to different clients in different orders based on priority values assigned to each asset in the set using different asset prioritization schemes. In some examples, the priority values assigned to the asset prioritization are transferred to the client along with the asset descriptions, allowing the client to use an asset prioritization scheme not supported by the asset prioritization unit while relying on the priority values assigned to the asset prioritization to prioritize assets with the same asset priority value assigned to the client. In some examples, the priority values may be adjusted by the asset prioritization unit based on feedback from the client.
[0267] According to a sixth aspect of the present disclosure, scene-based immersive media assets, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), may be prioritized based on distance in the field of view.
[0268] According to aspects of the present disclosure, not all assets in a stream are equally relevant at any one time because some assets may be outside the user's field of view when the user enters the scene (e.g., represented by a virtual person or represented by a virtual camera). To minimize the delay before the most visible and salient assets to the end user are available to the end device, assets that are visible within the field of view may be prioritized at the scene's default entry position where the user enters the scene. In some examples, assets in a scene may be prioritized according to whether the assets are within the field of view of a renderer's virtual camera and the asset's distance from the virtual camera position used to render the view presented to the user.
[0269] FIG. 33 shows a diagram of a scene 3301 in some examples. Scene 3301 is similar to scene 2701 in FIG. 27. A camera represents a virtual person at a viewing position corresponding to an end user. The camera enters scene 3301 at a default entry position 3302 (also called the initial position). Scene 3301 includes asset A, asset B, asset C, and asset D. Asset A is a set of floor-to-ceiling bookshelves, as shown at 3303. Asset B is a computer desk, as shown at 3304. Asset C is a conference table, as shown at 3305. Asset D is a doorway to another scene, as shown at 3306.
[0270] When the virtual camera enters the scene 3301 at its default entry position 3302, some assets may be immediately visible, while others may not. The virtual camera's field of view (also called field of view) 3307 is limited and does not include the entire scene. In the example of Figure 33, asset A and asset B are visible from the virtual camera's initial position 3302 because they are within the field of view 3307 from the virtual camera. Note that asset A is farther from the virtual camera's initial position 3302 than asset B.
[0271] In FIG. 33, assets C and D are outside the field of view 3307 of the virtual camera and therefore cannot be seen from the initial position 3302 of the virtual camera.
[0272] In some examples, before assets for a scene are streamed to an end device, an asset prioritization unit is used to analyze the assets required for the scene and assign a priority value to each asset according to the field of view associated with the default entry location of the scene and the distance of the asset to the default entry location of the scene.
[0273] FIG. 34 shows a diagram for re-prioritizing assets of a scene in some examples. In the example of FIG. 34, an asset prioritization unit 3402 is used to assign a priority value to each asset of the scene. For example, a scene (e.g., scene 3301 of FIG. 33) depends on asset A (shown as 3404), asset B (shown as 3405), asset C (shown as 3406), and asset D (shown as 3407). Initially, as shown in 3401, the four assets are in the asset order of asset A, asset B, asset C, and asset D. The asset prioritization unit 3402 analyzes the assets according to the field of view at a default entry position for entering the scene and the distance within the field of view from the default entry position. In some examples, the priority value of an asset can be determined based on whether the asset is visible within the camera field of view. In some examples, the priority value of an asset can be based on the asset's relative distance from the camera. For example, assets A and B are within the field of view and can be seen from the default entry position, while assets C and D are outside the field of view and cannot be seen from the default entry position. Furthermore, asset B is closer to the default entry position than asset A. As shown in 3403, the four assets can be re-ordered according to their priority values, such as in the following order: asset B, asset A, asset C, and asset D.
[0274] In an example, the re-ranked assets may be streamed to an end device by a media server.
[0275] 35 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 35, a scene of scene-based immersive media includes four assets re-ordered according to the priority values of the four assets, such as from highest to lowest: asset B, asset A, asset C, and asset D, as shown at 3501. The four assets are provided from a media server 3502 to an end device 3504 via a cloud 3503 in order of the priority values. For example, the four assets are retrieved and sent to the end device 3504 in the order of asset B (shown at 3505), asset A (shown at 3506), asset C (shown at 3507), and asset D (shown at 3508).
[0276] Comparing the first streaming order of FIG. 24 with the second streaming order of FIG. 35, the visible assets are streamed first, so that the end device 3504 can render the visible assets of the scene faster than the end device 2404.
[0277] In some examples, the asset prioritization unit 3402 can assign priority values to assets based on the field of view of the virtual camera and the distance from the virtual camera, with assets closest to the user within the field of view being rendered first.
[0278] In some examples, assets outside the field of view (also called the field of view) may not be scheduled for transmission to the end device 3504 until all assets within the field of view are scheduled for transmission. There may be a delay before assets outside the field of view are scheduled for transmission. In extreme cases, if a user quickly views a scene and then backs away from the scene to enter another scene, assets that were not scheduled for transmission may not be scheduled until the leaving user re-enters the scene.
[0279] 36 shows a diagram of a displayed scene 3601 in some examples. The displayed scene 3601 corresponds to scene 3301. In the displayed scene 3601, bookshelf 3603 and workstation 3604 have been retrieved and rendered because they are assets of scene 3601 visible to the camera from the default entry position.
[0280] Some aspects of the present disclosure provide a method of dynamic adaptive streaming for light field-based display according to an assigned asset priority value determined by whether an asset is within the field of view (also referred to as the field of view) of a camera (e.g., a virtual camera corresponding to a user or player for an immersive media experience). In some examples, the assigned asset priority value of an asset is determined by the asset's distance from the camera's position. The assets may be transmitted in an order determined based on the asset priority value. In some examples, the asset priority value of an asset is lowered when the asset is outside the field of view of the camera. In some examples, the asset priority value of an asset is lowered if the asset is farther from the camera than another asset within the field of view. In some examples, the priority value is adjusted based on feedback from one or more client devices.
[0281] According to a seventh aspect of the present disclosure, scene-based immersive media assets, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), may be prioritized based on the complexity of the assets.
[0282] According to aspects of the present disclosure, not all assets of a scene are equally complex and do not require the same amount of computation from computing resources (e.g., graphics processing devices (GPUs)) to render the assets to an end user's end device. To minimize the delay before the scene can be presented to an end user, in some examples, the assets that make up a scene can be prioritized for streaming based on the relative complexity of each asset present in the scene description.
[0283] Figure 37 shows a diagram of an example scene-based immersive media scene 3701. Scene 3701 includes four assets: asset A, asset B, asset C, and asset D. In Figure 37, asset A is a sleek modern house as shown at 3705, asset B is a real person as shown at 3706, asset C is a stylized human figure as shown at 3707, and asset D is a highly detailed tree line as shown at 3708.
[0284] In an example, when an end device requests assets for scene 3701, media server 2402 can provide the assets to the end device in any order, such as, for example, asset A, asset B, asset C, and asset D. The order in which the assets are made available to the end device may not provide an acceptable user experience.
[0285] It should be noted that the assets in scene 3701 are not of the same complexity and do not all impose the same computational load on the GPU of the end device. In this disclosure, the number of polygons that make up an asset is used as a proxy for the asset's computational load on the GPU. It should be noted that other characteristics of the asset, such as surface properties, may impose a computational load on the GPU.
[0286] In some examples, the number of polygons that make up an asset is used to measure the complexity of the asset. Of the assets in scene 3701, asset A is a sleek, modern-style house and has the fewest number of polygons, asset C is a stylized human statue and has the second fewest number of polygons, asset B is a realistic person and has the second most polygons, and asset D is a highly detailed tree line and has the most polygons. Based on the number of polygons, the assets in scene 3701 can be ranked from least complex to most complex: asset A, asset C, asset B, and asset D.
[0287] In some examples, before assets for a scene are streamed to an end device, an asset prioritization unit is used to analyze the assets required for the scene and assign a priority value to each asset based on the asset's relative complexity.
[0288] FIG. 38 shows a diagram for prioritizing assets in a scene in some examples. In the example of FIG. 38, an asset prioritization unit 3802 is used to assign a priority value to each asset in the scene based on the relative complexity of the asset. For example, a scene (e.g., scene 3701 in FIG. 37) depends on asset A (shown as 3804), asset B (shown as 3805), asset C (shown as 3806), and asset D (shown as 3807). Initially, as shown in 3801, the four assets are in asset order: asset A, asset B, asset C, and asset D. The asset prioritization unit 3802 prioritizes (e.g., re-ranks) the assets according to their relative complexity, such as based on the number of polygons in each asset. For example, asset D has the highest polygon count and is assigned the highest priority value (e.g., 4). Asset B has the second highest polygon count and is assigned the second highest priority value (e.g., 3). Asset C has the second fewest polygon count and is assigned the second lowest priority value (e.g., 2). Asset A has the fewest polygon count and the lowest priority value (e.g., 1). As shown in 3803, the four assets can be re-ranked according to priority value, such as in the following order: asset D, asset B, asset C, and asset A.
[0289] In the above example, the polygon count of the asset is used to evaluate the complexity of the asset and determine the priority value of the asset, but it should be noted that other attributes of the asset that may affect the computational load of the GPU in the end device may be used to evaluate the complexity and determine the priority value.
[0290] In an example, the re-ranked assets may be streamed to an end device by a media server.
[0291] 39 shows an example diagram for streaming scene-based immersive media to an end device in some examples. In the example of FIG. 39, a scene of scene-based immersive media includes four assets re-ordered according to the priority values of the four assets, such as from highest to lowest: asset D, asset B, asset C, and asset A, as shown at 3901. The four assets are provided from a media server 3902 to an end device 3904 via a cloud 3903 in order of the priority values. For example, the four assets are retrieved and sent to the end device 3904 in the following order: asset D (shown at 3905), asset B (shown at 3906), asset C (shown at 3907), and asset A (shown at 3908).
[0292] Comparing the first streaming order of Figure 24 with the second streaming order of Figure 39, the time required to render the most complex asset acts as the minimum elapsed time for the entire scene rendered by the end device, so that end device 3904 can render the scene faster than end device 2404. In some examples, the GPU of the end device can be configured to be multi-threaded, such that the GPU can render less complex assets in parallel with the most complex asset, making the less complex assets available when the most complex asset is being rendered.
[0293] Some aspects of the present disclosure provide a method for dynamic adaptive streaming for light field-based display based on assigned asset priority values determined by the expected computational load of a GPU. In some examples, the order of assets for streaming may be based on the asset priority values. In some examples, the asset priority values increase if an asset has higher computational complexity than other assets. In some examples, the priority values are adjusted based on feedback from a client device. For example, the feedback may indicate a GPU configuration on the client device or may indicate different attributes for complexity evaluation.
[0294] According to an eighth aspect of the present disclosure, scene-based immersive media assets, such as scene-based immersive media for light-field-based displays (e.g., light-field displays, holographic displays, etc.), may be prioritized based on multiple metadata attributes.
[0295] According to aspects of the present disclosure, each asset in a scene can have various metadata attributes that each enable the renderer to prioritize the order in which the assets are retrieved in order to render the scene in a way that provides a better experience for the user than simply retrieving the assets in the order in which they appear in the scene manifest.
[0296] According to another aspect of the present disclosure, priority values based on a single metadata attribute may not reflect the best strategy for prioritization. In some examples, a set of priority values that reflect multiple metadata attributes may provide a better user experience than another set of priority values based on a single metadata attribute.
[0297] According to another aspect of the present disclosure, priority values based on a single set of priority preferences may not reflect the best strategy for prioritization because end devices may vary greatly in many ways—the characteristics of the path between the end device and the cloud, and the GPU computational capabilities of the end device, to name just two ways. To provide a better user experience, techniques can be used to prioritize assets for streaming based on more than one asset metadata attribute.
[0298] Using scene 3701 in Figure 37 as an example, each of the four assets may have multiple metadata attributes.
[0299] 40 shows a diagram illustrating metadata attributes for an example asset 4001. Asset 4001 may correspond to asset B 3706 in scene 3701, a realistic person. Asset 4001 includes various metadata attributes 4002, such as size in bytes as shown at 4003, polygon count as shown at 4004, distance from default camera as shown at 4005, visibility from default camera (also called virtual camera of default entry position) as shown at 4006, and other metadata attributes as shown at 4007 and 4008.
[0300] It should be noted that metadata attributes 4002 are for illustrative purposes only, and that in some instances other suitable metadata attributes may be present, and that some metadata attributes listed in FIG. 40 may not be present in some instances.
[0301] In an example, when an end device requests assets for scene 3701, media server 2402 can provide the assets to the end device in any order, such as, for example, asset A, asset B, asset C, and asset D. The order in which the assets are made available to the end device may not provide an acceptable user experience.
[0302] The assets of scene 3701 are not prioritized by size, complexity, or immediate value to an end user. While each of these metadata attributes can be used separately as a basis for prioritization, each of these metadata attributes can also be used in combination with other metadata attributes to determine relative asset priority. For example, assets can be prioritized according to a policy that includes multiple metadata attributes, such as a first metadata attribute of visibility from a default camera, as shown at 4006, a second metadata attribute of polygon count, as shown at 4004, and a third metadata attribute of distance from the default camera, as shown at 4005. In the example, the first asset in the scene that is visible from the default camera is assigned the highest priority (e.g., 4 in the example), and the remaining assets are further processed according to the second and third metadata attributes. For example, the second asset of the remaining assets, whose polygon count is greater than a polygon threshold, is assigned the second highest priority (e.g., 3 in the example), and the remaining assets are further processed according to the third metadata attribute. For example, the third asset among the remaining assets whose distance to the default camera is less than the distance threshold is assigned the third highest priority (e.g., 2 in the example).
[0303] In some examples, before assets for a scene are streamed to an end device, an asset prioritization unit is used to analyze the assets required for the scene and assign a priority to each asset according to a prioritization scheme in use, such as a policy that includes multiple metadata attributes. For example, the asset prioritization unit 3802 can be used to apply prioritization using a policy that has a combination of multiple metadata attributes. After the assets are prioritized, the assets are streamed in priority order, as shown in FIG.
[0304] The use of prioritization based on multiple metadata attributes allows an end device to send assets to the end device in a manner appropriate to enhance the user experience. For example, an end device may request assets that are immediately visible to the user (e.g., based on a first metadata attribute), then request assets that are most computationally demanding to render (e.g., based on a second metadata attribute), then request assets that are closest to the user in the rendered scene (e.g., based on a third metadata attribute).
[0305] According to aspects of the present disclosure, not all end devices have the same capabilities, and not all end devices have the same path characteristics, such as available bandwidth between the end device and the cloud, and the best user experience for different users may result from the use of different prioritization schemes. In some examples, end devices can provide priority preference information to the asset pirator to assign priority values accordingly.
[0306] FIG. 41 shows a diagram for prioritizing assets of a scene in some examples. In the example of FIG. 41, an end device 4108 can provide a set of priority preferences 4109 to the asset prioritization unit 4102. For example, the set of priority preferences 4109 includes multiple metadata attributes. The asset prioritization unit 4102 can then assign a priority value to each asset of the scene based on the set of priority preferences. Thus, the assigned priority values of the assets of the scene are optimized for the end device 4108.
[0307] In some examples, the end device 4108 can dynamically adjust its set of priority preferences based on measurements made by the end device. For example, if the end device 4018 has relatively low available bandwidth to the cloud, the end device can prioritize assets based on asset size measured in bytes (e.g., 4003 in FIG. 40) because the transfer time of the largest asset acts as a lower bound on how quickly a scene can be rendered. If the same end device 4108 detects that the path from the cloud to the end device 4108 has abundant bandwidth, the end device 4108 can provide a different set of priority preferences to the asset prioritization unit 4102.
[0308] In some examples, related metadata attributes may be considered composite attributes, and the end device may include the composite attributes in its set of priority preferences. For example, if the available metadata attributes for each asset include the number of polygons and some measure of the level of surface detail, those attributes may be considered together as a complexity calculation measure, and the end device may include that measure in its set of priority preferences.
[0309] Some aspects of the present disclosure provide a method for dynamically adaptive streaming for a light field-based display according to assigned asset priority values. The asset priority values can be assigned based on two or more metadata attributes for each asset to be streamed. In some examples, the asset priority values of the assets of a scene are determined based on a set of priority preferences. In some examples, the order of assets for streaming is based on the asset priority values. In some examples, the set of priority preferences varies between end devices. In some examples, the set of priority preferences is provided by the end device. In some examples, the set of priority preferences is updated by the end device. In some examples, the set of priority preferences includes one or more composite attributes. In some examples, the set of priority preferences provided by the end device includes one or more composite attributes.
[0310] 42 shows a flowchart outlining a process 4200 according to an embodiment of the present disclosure. Process 4200 may be performed in a network, such as by a server device of the network. In some embodiments, process 4200 is implemented in software instructions, and thus, processing circuitry executes the software instructions to perform process 4200. For example, a Smart Client may be implemented in software instructions, and the software instructions may be executed to perform a Smart Client process, including process 4200. The process starts at S4201 and proceeds to S4210.
[0311] At S4210, the server device receives scene-based immersive media for playback on a light field-based display. The scene of the scene-based immersive media includes a first order of a plurality of assets.
[0312] At S4220, the server device determines a second order for streaming the plurality of assets to the end device, the second order being different from the first order.
[0313] In some examples, the server device assigns priority values to each of multiple assets of the scene according to one or more attributes of the multiple assets, and determines a second order for streaming the multiple assets according to the priority values.
[0314] In some examples, the server device includes priority values for the multiple assets in a description of the multiple assets and provides the description to the end device, and the end device requests the multiple assets according to the priority values in the description.
[0315] According to some aspects of the present disclosure, the server device assigns priority values to assets of a scene according to the size of the assets. In an example, the server device assigns a first priority value to a first asset and a second priority value to a second asset in response to a first asset having a size, in bytes, smaller than a second asset, the first priority value being higher than the second priority value.
[0316] According to some aspects of the present disclosure, the server device assigns priority values to assets of a scene according to the visibility of the assets associated with a default entry location of the scene. In some examples, the server device assigns a first priority value to a first asset that is visible from the default entry location of the scene and a second priority value to a second asset that is not visible from the default entry location, the first priority value being higher than the second priority value. In an example, the server device assigns a second priority value to an asset in response to the asset being blocked by another asset of the scene. In another example, the server device assigns a second priority value to an asset in response to the asset being outside of a field of view associated with the default entry location of the scene. In another example, the server device assigns a second priority value to an asset in response to the asset being occluded by a light source in the scene by a shadow cast by another asset in the scene.
[0317] According to some aspects of the present disclosure, the server device assigns a first priority value to the first asset and a second priority value to the second asset in response to the first asset and the second asset being within a field of view associated with a default entry location of the scene, the first asset being closer to the default entry location of the scene than the second asset, and the first priority value being greater than the second priority value.
[0318] In some examples, the server device assigns a first set of priority values to the plurality of assets based on a first attribute of the plurality of assets, and the first set of priority values is used to order the plurality of assets in a first prioritization scheme. Further, the server device assigns a second set of priority values to the plurality of assets based on a second attribute of the plurality of assets, and the second set of priority values is used to order the plurality of assets in a second prioritization scheme. The server device selects a prioritization scheme for the end device from the first prioritization scheme and the second prioritization scheme according to information of the end device, and orders the plurality of assets for streaming according to the selected prioritization scheme.
[0319] In some examples, the server device includes priority values assigned to the assets in a description of the assets. The priority values are used to order the assets in a first prioritization scheme. The server device provides the description of the assets to the end device. In examples, the end device adapts a second prioritization scheme to the requested order of the assets and uses the first prioritization scheme as a tiebreaker in response to a tie due to the second prioritization scheme.
[0320] In some examples, the server device assigns a first priority value to the first asset and a second priority value to the second asset in response to the first asset having a higher computational complexity than the second asset, the first priority value being higher than the second priority value. In examples, the higher computational complexity is determined based on at least one of a number of polygons of the first asset and a surface characteristic of the first asset.
[0321] In some examples, the one or more attributes are metadata attributes associated with each asset and may include at least one of: size in bytes, number of polygons, distance from the default camera, and visibility from the default camera.
[0322] In some examples, the server device receives a set of priority preferences for the end device, the set of priority preferences including a first set of attributes, and determines the second order according to the first set of attributes. In examples, the server device receives an update of the priority preferences for the end device, the update of the priority preferences indicating a second set of attributes different from the first set of attributes. The server device can determine the updated second order according to the second set of attributes.
[0323] In some examples, the server device receives a feedback signal from the end device and adjusts the priority values and the second order according to the feedback signal.
[0324] Process 4200 then proceeds to S4299 and ends.
[0325] Process 4200 may be appropriately adapted to various scenarios, and the steps of process 4200 may be adjusted accordingly. One or more of the steps of process 4200 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 4200. Additional steps may be added.
[0326] 43 shows a flowchart outlining a process 4300 according to an embodiment of the present disclosure. The process 4300 can be performed by an electronic device such as an end device (also referred to as a client device). In some embodiments, the process 4300 is embodied in software instructions, and thus, when a processing circuit executes the software instructions, the processing circuit performs the process 4300. For example, the process 4300 can be implemented as software instructions, and the processing circuit can execute the software instructions to perform a smart controller process. The process starts at S4301 and proceeds to S4310.
[0327] At S4310, a request for scene-based immersive media for playback on a light field-based display is transmitted by the electronic device to a network, the scene of the scene-based immersive media including a first order of a plurality of assets;
[0328] In S4320, information of the electronic device is provided to adjust the order of streaming the multiple assets to the electronic device, and then the order is determined according to the information of the electronic device.
[0329] In some examples, an electronic device receives a description of a plurality of assets of a scene of scene-based immersive media for playback on a light field-based display, the description including priority values of the plurality of assets, and the electronic device requests one or more of the plurality of assets according to the priority values.
[0330] In some examples, the priority value is associated with a first prioritization scheme. The electronic device determines that the first asset and the second asset have the same priority according to a second prioritization scheme that is different from the first prioritization scheme. The electronic device then prioritizes one of the first asset and the second asset over the other according to the first prioritization scheme.
[0331] In some examples, the electronic device provides the set of priority preferences to the network. Further, in examples, the electronic device detects changes in the operating environment, such as changes in network status, and provides updates to the set of priority preferences in response to the changes in the operating environment.
[0332] In some examples, the electronic device detects a change in the operating environment, such as a change in network status, and then provides a feedback signal to the network that indicates the change in the operating environment.
[0333] The process 4300 then proceeds to S4399 and ends.
[0334] Process 4300 may be appropriately adapted to various scenarios, and the steps of process 4300 may be adjusted accordingly. One or more of the steps of process 4300 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 4300. Additional steps may be added.
[0335] According to a ninth aspect of the present disclosure, the scene analyzer can be used to prioritize asset rendering based on the visibility of the assets in the default viewport of the scene.
[0336] According to aspects of the present disclosure, client devices that support scene-based media may be equipped with renderers and / or game engines that support the concept of a default camera and viewport for a scene. The default camera provides camera attributes that the renderer can use to create an interior image of a scene. Such an interior image is thus defined by camera attributes such as horizontal and vertical resolution, x- and y-dimensions, and depth of field, among many other possible attributes. The viewport is defined by the interior image itself and the virtual position of the camera relative to the scene. The virtual position of the camera relative to the scene allows the content creator or user to specify exactly which portions of the scene will be rendered in the interior image. Thus, objects that are not directly visible in the default camera's viewport—for example, objects that are behind the default camera or outside the viewport into the scene—may not be prioritized for rendering by the renderer, i.e., if the user cannot see them anyway. Nevertheless, some objects may be located close enough to the default camera's viewport that shadows cast from their presence and / or proximity to the viewport are likely to actually be inside the viewport. For example, if a point light source that is not visible in the viewport is placed behind other objects that are also not visible in the scene, the light that is reflected into the scene's viewport may be affected by shadows produced by the not visible objects because their presence blocks the light from being able to travel into the scene's viewport. It may therefore be important for a renderer to render such objects that affect the light's ability to travel to the part of the scene captured by the viewport.
[0337] Some aspects of the present disclosure provide techniques for a scene analyzer that facilitates a separate decision-making process for ordering individual assets in a scene for conversion from format A to format B by or on behalf of a client. The scene analyzer can calculate and store priority values for individual assets in the scene and can be determined based on several different factors, including whether the asset is in the region of the scene captured by the default camera's viewport. The calculated priority values can then be stored directly in the metadata describing the scene, such that subsequent processes on either the network or the client device are signaled regarding the potential importance of converting one asset before another, i.e., as indicated by the priority value stored in the scene metadata by the scene analyzer. In some examples, such a scene analyzer can utilize metadata describing camera characteristics that are expected to be used to create a viewport in the scene, i.e., to calculate whether a particular asset is inside or outside the camera's viewport in the scene. The scene analyzer can calculate the region of the viewport based on camera attributes, such as the width and length of the viewport, the depth of field of the object in focus of the camera, etc. In some examples, a first asset that is not in the scene's camera viewport may be prioritized by the scene analyzer for transformation after a second asset that is within the default camera viewport is transformed. Such prioritization may be signaled and stored by the scene analyzer in the scene's metadata. Storing such metadata in the scene may benefit subsequent processes within the distribution network, i.e., the same calculation to determine which assets to transform before other assets may not have to be performed if the prioritization was calculated by the scene analyzer and stored in the scene's metadata.The scene analyzer process may be performed a priori on the media being streamed by the client, or in some instances as a specific step performed by the client device itself.
[0338] Figure 44 shows a diagram of a timed media representation 4400 with signaling of assets for a default camera perspective in some examples. The timed media representation 4400 includes a timed scene manifest 4400A that contains information for a list of scenes 4401. The scene 4401 information points to a list of components 4402 that separately describe processing information and media asset types for the scene 4401. The components 4402 point to assets 4403, which further point to a base layer 4404 and an attribute enrichment layer 4405. A list of assets 4407 located within the scene's default camera viewport is provided for each scene 4401.
[0339] Figure 45 shows a diagram of an untimed media representation 4500 with asset signaling in a default viewport in some examples. The untimed scene manifest (not shown) references Scene 1.0, which has no other scenes that can branch to it. Scene information 4501 is not associated with a start and end time according to a clock. Scene information 4501 points to a list of components 4502 that separately describe the processing information and media asset types of the scene. Components 4502 reference assets 4503, which in turn reference base layer 4504 and attribute enrichment layers 4505, 4506. Scene information 4501 may also point to other scene information 4501 for untimed media. Scene information 4501 also points to scene information 4507 for a timed media scene. List 4508 identifies assets whose geometry is located within the viewport of the scene's default camera.
[0340] 9, in some examples, the media analyzer 911 of FIG. 9 includes a scene analyzer that can perform analysis of assets of a scene. In an example, the scene analyzer can determine which assets of the assets of a scene are in a default viewport. For example, ingest media can be updated to store information about which assets are located within the default viewport of a camera in a scene.
[0341] 14 , in some examples, the media analyzer 1410 of FIG. 14 includes a scene analyzer that can perform analysis of assets of a scene. In examples, the scene analyzer can examine the client adaptive media 1409 to determine which assets are located in the scene's default camera viewport (also referred to as the default viewport for entering the scene) for potential prioritization for rendering by the game engine 1405 and / or reconstruction processing via the MPEG Smart Client 1401. In such embodiments, the media analyzer can store in each scene of the client adaptive media 1409 a list of each scene's assets that are in the default camera viewport.
[0342] Figure 46 shows a process flow 4600 for a scene analyzer to analyze a scene's assets according to a default viewport. The process begins at step 4601, where the scene analyzer obtains attributes (e.g., from the scene or from an external source not shown) of a default camera position and characteristics that describe the default viewport (the initial view into the scene when it is first rendered) based on the attributes of the camera used to define the viewport. Such camera attributes include the type of camera used to create the viewport and the length, width, and depth dimensions of the area of the scene captured by the viewport. In step 4602, the scene analyzer can calculate the actual default viewport based on the attributes obtained in step 4601. Next, in step 4603, the scene analyzer determines whether there are any more assets to examine in the list of assets for a particular scene. If there are no more assets to examine, the process proceeds to step 4606. In step 4606, the scene analyzer writes (to the scene's metadata) a list of assets whose geometry is within (wholly or partially) the scene's default viewport. The analyzed media is shown as 4607. If there are more assets to process in step 4603, the process proceeds to step 4604. In step 4604, the scene analyzer analyzes the geometry of the next asset in the scene, and the next asset becomes the current asset. The geometry and position of the current asset in the scene are compared to the area of the default viewport calculated in step 4602. In step 4604, the scene analyzer examines the asset's geometry and its position in the scene to determine whether any portion of the asset geometry falls within the area of the default viewport. If any portion of the asset geometry falls within the default viewport, the process proceeds to step 4605. In step 4605, the scene analyzer stores the asset identifier in a list of assets located within the default viewport.Following step 4605, the process returns to step 4603 to examine the list of assets for the scene to determine if there are any more assets to process.
[0343] Some aspects of the present disclosure provide a method for asset prioritization. The method may calculate the position and size of a viewport created by a camera used to view a scene and compare the geometry of each asset in the scene with the viewport to determine whether any portion of the asset's geometry intersects with the viewport (e.g., the position and size of the viewport created by the camera). The method may identify assets whose geometry partially or fully intersects with the viewport created by the camera and store the asset's identification information in a list of similar assets whose geometry intersects with an area of the viewport created by the camera. The method then includes ordering and prioritizing the rendering of assets of the scene in a scene-based media presentation based on the list of similar assets whose geometry intersects with an area of the viewport created by the camera.
[0344] 47 shows a flowchart outlining a process 4700 according to an embodiment of the present disclosure. The process 4700 may be performed in a network, such as by a server device of the network. In some embodiments, the process 4700 is implemented in software instructions, and thus the processing circuitry performs the process 4700 when the processing circuitry executes the software instructions. The process starts at S4701 and proceeds to S4710.
[0345] At S4710, attributes of a viewport, such as a position and a size of a viewport associated with a camera for viewing a scene of the scene-based immersive media, are determined. The scene includes multiple assets.
[0346] At S4720, for each asset, whether the asset at least partially intersects with the viewport is determined based on the position and size of the viewport and the geometry of the asset.
[0347] At S4730, an identifier of the asset is stored in a list of visible assets in response to the asset at least partially intersecting the viewport.
[0348] At S4740, the list of visible assets is included in the metadata associated with the scene.
[0349] In some examples, the position and size of the viewport are determined according to the camera position and characteristics of the camera through which the scene is viewed.
[0350] In some examples, the order of converting at least the first asset and the second asset from the first format to the second format is determined according to a list of visible assets. In examples, the list of visible assets is converted from the first format to the second format before converting another asset in the scene that is not on the list of visible assets.
[0351] In some examples, the order for streaming at least the first asset and the second asset to the end device is determined according to a list of visible assets, in examples, the list of visible assets is streamed before streaming another asset of the scene that is not in the list of visible assets.
[0352] Process 4700 then proceeds to S4799 and ends.
[0353] Process 4700 may be appropriately adapted to various scenarios, and the steps of process 4700 may be adjusted accordingly. One or more of the steps of process 4700 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 4700. Additional steps may be added.
[0354] According to a tenth aspect of the present disclosure, techniques for prioritizing media adaptation by connected client type may be used.
[0355] According to aspects of the present disclosure, one of the problems affecting the efficiency of a media distribution network for heterogeneous clients is determining which media to prioritize for conversion, given that there are a large number of clients (e.g., client devices, end devices) connected to the network for which media needs to be converted.
[0356] Without loss of generality, it is assumed that the conversion process may also be referred to as an "adaptation" process, as both names may be used in the industry. A network that performs media adaptation on behalf of its connected clients benefits from information that guides the prioritization of such adaptation with information about the current client types connected to the network.
[0357] Some aspects of the present disclosure provide techniques for media adaptation or media conversion processes that can be performed to maximize the number of clients that can benefit from the adaptation (or conversion) process performed by the network. The disclosed subject matter addresses the need for a network to perform some or all of the media adaptation on behalf of one or more clients and client types to prioritize the adaptation process based on the number and types of clients currently connected to the network.
[0358] FIG. 48 shows a diagram of an immersive media delivery module 4800 in some examples. The immersive media delivery module 4800 is similar to the immersive media delivery module 900 of FIG. 9, but with the addition of a client tracking process indicated by 4812 to illustrate immersive media network delivery with a client tracking process. The immersive delivery module 4800 can serve legacy and heterogeneous immersive media-enabled displays, as previously described in FIG. 8. Content is created or acquired as indicated by 4801, which is further embodied in FIGS. 5 and 6 for natural content and CGI content, respectively. The content of 4801 is then converted to an ingest format using a network ingest format creation process indicated by 4802. The process indicated by 4802 is similarly further embodied in FIGS. 5 and 6 for natural content and CGI content, respectively. The ingest media is optionally annotated with IMS metadata or analyzed to determine complexity attributes by a scene analyzer in media analyzer 4811. The ingest media format is sent to a network and stored in storage device 4803. In some examples, the storage device resides on the immersive media content creator's network and may be remotely accessed by an immersive media network distribution process (not numbered), as indicated by the dashed line bisecting 4803. Client and application specific information is available in some examples on remote storage device 4804, which may reside remotely in an alternative cloud network in some examples.
[0359] As shown in FIG. 48 , the network orchestrator, designated by 4805, serves as the primary source and sink of information for performing the primary tasks of the distribution network. In this particular embodiment, the network orchestrator 4805 may be implemented in a unified format with other components of the network. Nevertheless, the tasks depicted by the network orchestrator 4805 in FIG. 48 form several elements of the disclosed subject matter. The network orchestrator 4805 may further employ a bidirectional message protocol with clients to facilitate all processing and delivery of media according to the client's characteristics. Furthermore, the bidirectional protocol may be implemented across different distribution channels, i.e., the control plane channel and the data plane channel.
[0360] In some examples, network orchestrator 4805 receives information over channel 4807 regarding the characteristics and attributes of one or more client devices 4808, and also gathers requirements regarding applications currently running on 4808. This information may be obtained from device 4804, or in alternative embodiments, by directly querying client device 4808. When querying client device 4808 directly, examples assume that a two-way protocol (not shown in FIG. 48) exists and is operable to allow the client device to communicate directly with network orchestrator 4805.
[0361] In some examples, the channel 4807 may also update a client tracking process 4812 so that a record of the current number and types of client devices 4808 is maintained by the network. The network orchestrator 4805 may calculate one or more priorities for scheduling the media adaptation and fragmentation process 4810 based on the information stored in the client tracking process 4812.
[0362] The network orchestrator 4805 also initiates and communicates with the media adaptation and fragmentation process 4810 described in FIG. 10. Once the ingested media has been adapted and fragmented by process 4810, the media is transferred to an inter-media storage device, shown in some examples as media prepared for distribution storage device 4809. Once the distribution media is prepared and stored on device 4809, the network orchestrator 4805 ensures that the client device 4808 either receives the distribution media and corresponding description information 4806 via its network interface 4808B via a push request, or the client device 4808 itself can initiate a pull request for the media 4806 from the storage device 4809. The network orchestrator 4805 can use a two-way message interface (not shown in FIG. 48) to perform a "push" request or to initiate a "pull" request by the client device 4808. The client 4808 can, in some examples, use a GPU (or CPU, not shown) 4808C. The media delivery format is stored on a storage device or storage cache 4808D of the client device 4808. Finally, the client device 4808 visually presents the media via its visualization component 4808A.
[0363] Throughout the process of streaming immersive media to client devices 4808, network orchestrator 4805 may monitor the status of the client's progress via client progress and status feedback channel 4807. Status monitoring may be performed by a two-way communication message interface (not shown in FIG. 48).
[0364] Figure 49 shows a diagram of a media adaptation process 4900 in some examples. The media adaptation process can adapt ingested source media to fit the requirements of one or more client devices, such as client device 908 (shown in Figure 9). The media adaptation process 4901 can be performed by multiple components that facilitate the adaptation of ingested media to an appropriate delivery format for the client device.
[0365] In some examples, network orchestrator 4903 initiates adaptation process 4901. Database 4912 can, in some examples, assist network orchestrator 4903 in prioritizing adaptation process 4901 by providing information related to the type and number of client devices (e.g., immersive client devices 4808) connected to the network.
[0366] In FIG. 49, adaptation process 4901 receives input network status 4905 to track the network's current traffic load. Client device information includes attribute and feature descriptions, application capabilities and descriptions, and the application's current status, as well as a client neural network model (if available) that assists in mapping the client's frustum geometry to the interpolation capabilities of the ingested immersive media. Such client device information may be obtained through a two-way message interface (not shown in FIG. 49). Adaptation process 4901 ensures that the adapted output, once created, is stored in client adaptation media storage device 4906. In some examples, scene analyzer 4907 may be executed in advance or as part of a network automation process for the delivery of media.
[0367] In some examples, the adaptation process 4901 is controlled by a logic controller 4901F. The adaptation process 4901 also uses a renderer 4901B or a neural network processor 4901C to adapt the particular ingest source media to a format appropriate for the client device. The neural network processor 4901C uses the neural network model of 4901A. Examples of such neural network processors 4901C include deep neural network model generators such as those described in MPI and MSI. If the media is in a 2D format, and the client requires a 3D format, the neural network processor 4901C can then invoke a process that uses highly correlated images from the 2D video signal to derive a stereoscopic representation of the scene depicted in the video. An example of a suitable renderer 4901B may be a modified version of the OTOY Octane renderer (not shown) modified to interact directly with the adaptation process 4901. The adaptation process 4901 can optionally use a media compressor 4901D and a media decompressor 4901E depending on the needs of these tools regarding the format of the ingested media and the format required by the client device.
[0368] Some aspects of the present disclosure provide a method that includes directing a prioritization process for adapting media to media requirements of client devices on a network. The method includes collecting profile information for client devices connected to the network. The media requirements of the client devices, in some examples, include attributes such as one or more of media format and bit rate. The client device profile information, in some examples, includes the type and number of client devices connected to the network.
[0369] 50 shows a flowchart outlining a process 5000 according to an embodiment of the present disclosure. Process 5000 may be performed in a network, such as by a server device of the network. In some embodiments, process 5000 is implemented in software instructions, and thus, when a processing circuit executes the software instructions, the processing circuit performs process 5000. The process starts at S5001 and proceeds to S5010.
[0370] At S5010, profile information for a client device connected to a network is received, the profile information being usable to direct adaptation of media to one or more media requirements of the client device for delivery of the media to the client device.
[0371] At S5020, a first adaptation that adapts media to first media requirements for delivery to a first subset of client devices is prioritized according to client device profile information.
[0372] In some examples, the media requirements in the one or more media requirements include at least one of a format requirement and a bit rate requirement for the media.
[0373] In some examples, the client device profile information includes the type of client device and the number of client devices of each type.
[0374] In an example, a first adaptation that adapts media to a first media requirement for delivery to a first subset of client devices is prioritized over a second adaptation that adapts media to a second media requirement for delivery to a second subset of client devices in response to a first number of client devices in the first subset being greater than a second number of client devices in the second subset.
[0375] Process 5000 then proceeds to S5099 and ends.
[0376] Process 5000 may be appropriately adapted to various scenarios, and the steps of process 5000 may be adjusted accordingly. One or more of the steps of process 5000 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 5000. Additional steps may be added.
[0377] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]
[0378] 100 media flow process, 104 network cloud, 104 edge devices, 104 network, 105 network connectivity, 105 media, 106 rendering process, 106 rendering function, 108 client device, 110 media store, 111 fetching mechanism, 200 media transformation decision process, 200 process, 201 media, 205 media, 206 preparation process, 300 timed media representation, 300 streamable format, 301 scene, 301 scene information, 302 components, 303 assets, 304 base layer, 305 attribute enrichment layer, 306 proxy visual assets, 308 proxy audio assets, 400 streamable format, 400 non-timed media representation, 401 scene, 401 scene information, 402 components, 403 assets, 404 base layer, 405 attribute enrichment layer, 406 attribute enrichment layer, 407 Scene information, 408 List, 500 Process, 501 Camera unit, 502 Camera unit, 503 Camera unit, 504 Synthesis module, 505 Neural network training module, 506 Training images, 507 Ingest format, 508 Capture neural network model, 508 Model, 509 Natural image content, 511 Annotation process, 600 Process, 601 LIDAR camera, 602 Point cloud, 602 Data, 603 Computer, 604 CGI asset, 604 Data, 605 Actors, 606 Data, 607 Synthesis module, 608 Synthetic media, 609 IMS annotation process, 610 IMS annotated synthetic media, 700 Computer system, 700 Architecture, 701 Keyboard, 702 Mouse, 703 Trackpad, 705 Joystick, 706 Microphone, 707 Scanner, 708 Camera, 709 Speaker, 710 Touchscreen, 710 Screen, 721 Media, 722 Thumbdrive, 723 Solid-state drive, 740 Core, 741 Central Processing Unit (CPU), 742 Graphics Processing Unit (GPU), 743Field Programmable Gate Area (FPGA), 744 Accelerator, 744 Hardware Accelerator, 745 ROM, 746 RAM, 746 Random Access Memory, 747 Core Internal Mass Storage, 747 Mass Internal Mass Storage, 748 System Bus, 749 Peripheral Bus, 750 Graphics Adapter, 754 Interface, 755 Communication Network, 800 Network Media Distribution System, 801 Content Acquisition Module, 802 Content Preparation Module, 803 Transmission Module, 804 Gateway, 805 Set-Top Box, 806 Wireless Demodulator, 807 Legacy 2D Television, 808 WiFi Router, 808 Router, 809 Legacy 2D Display, 810 Head-Mounted 2D Raster-Based Display, 811 Display, 811 Lenticular Light Field Display, 812 Holographic Display, 813 Display, 814 Augmented Reality Headset, 815 high density light field display, 900 immersive media distribution module, 900 immersive media network distribution module, 901 module, 902 network ingest format creation module, 902 module, 903 storage device, 904 remote storage device, 904 device, 905 network orchestrator, 906 media, 906 description information, 907 status feedback channel, 908 client, 908 immersive client, 908 client device, 909 storage device, 909 distribution storage device, 909 device, 910 fragmentation module, 910 module, 911 media analyzer, 1000 media adaptation process, 1000 process, 1001 adaptation module, 1001 media adaptation and fragmentation module, 1001 adaptation process, 1001 fragmentation module, 1003 client interface module, 1003 network orchestrator, 1004 client information, 1005 Input network status, 1006 Media storage device, 1007 Media analyzer, 1100Delivery format creation process, 1101 media adaptation module, 1102 client adaptation media storage device, 1103 media packaging module, 1104 delivery format, 1200 packetizer process system, 1201 client, 1201 media, 1202 packetizer, 1203 packet, 1204 client endpoint, 1301 client, 1302 network orchestrator, 1303 ingest media server, 1304 adaptation interface, 1305 media adaptation module, 1306 packaging module, 1307 media server, 1308 media request, 1308 request, 1309 profile request, 1310 response, 1311 session ID token, 1312 request, 1312 ingest media request, 1313 response, 1314 call, 1315 neural network model token, 1316 call, 1316 request, 1317 response, 1318 request, 1320 interface call, 1321 response message, 1321 response, 1322 response, 1323 request, 1324 message, 1400 media system, 1401 MPEG smart client process, 1401 MPEG smart client, 1402 client media reconstruction process, 1402 reconstruction process, 1403 neural network processor, 1404 client adaptive media cache, 1405 game engine, 1406 sequential compression decoder process, 1407 client media cache, 1408 network orchestrator device, 1409 client adaptive media, 1410 media analyzer, 1411 media cache, 1412 user interface, 1413 haptic component, 1414 audio component, 1415 visualization component, 1416 media cache, 1417 callback function, 1418 game engine client device, 1418 client device, 1419 Media cache, 1420 Network interface protocol, 1421 Neural network model, 1501 Scene manifest, 1502Media server, 1503, cloud, 1504, end device, 1601, first scene manifest, 1602, scene complexity analyzer, 1603, scene prioritization unit, 1604, second scene manifest, 1701, scene manifest, 1702, priority-aware media server, 1703, cloud, 1704, end device, 1801, virtual museum, 1803, upper floor, 1804, stairwell, 1805, lobby, 1808, upper floor exhibition room, 1901, immersive game, 1902, person, 1903, spawn area, 1920, high-resolution HD video signal, 2001, scene manifest, 2002, priority-aware media server, 2003, cloud, 2004, end device, 2010, current location, 2011, scene prioritization unit, 2100, process, 2200, process, 2301 Scene manifest, 2302, Asset, 2303, Scene, 2401, Scene, 2402, Media Server, 2403, Cloud, 2404, End Device, 2502, Asset Prioritization Unit, 2602, Media Server, 2603, Cloud, 2604, End Device, 2701, Scene, 2702, Entry Location, 2702, Location, 2802, Asset Prioritization Unit, 2902, Media Server, 2903, Cloud, 2904, End Device, 3001, Scene, 3002, Entry Location, 3003, Bookshelf, 3005, Conference Table, 3006, Doorway, 3007, Light Source, 3102, Asset Prioritization Unit, 3202, Media Server, 3203, Cloud, 3204, End Device, 3301, Scene, 3302, Entry Location, 3302, Initial Location, 3307, Field of View, 3402 Asset prioritization unit, 3502, media server, 3503, cloud, 3504, end device, 3601, scene, 3603, bookshelf, 3604, workstation, 3701, scene, 3802, asset prioritization unit, 3902, media server, 3903, cloud, 3904, end device, 4001, asset, 4002, metadata attribute, 4018, end device, 4055, attribute enrichment layer, 4102, asset prioritization unit, 4108, end device, 4109, set, 4200, process, 4300, process, 4400, timed media representation, 4401Scene, 4402, Component, 4403 Asset, 4404 Base Layer, 4405 Attribute Enrichment Layer, 4407 Asset, 4500 Non-timed Media Representation, 4501 Scene Information, 4502 Component, 4503 Asset, 4504 Base Layer, 4505 Attribute Enrichment Layer, 4506 Attribute Enrichment Layer, 4507 Scene Information, 4508 List, 4600 Process Flow, 4700 Process, 4800 Immersive Delivery Module, 4800 Immersive Media Delivery Module, 4803 Storage Device, 4804 Remote Storage Device, 4804 Device, 4805 Network Orchestrator, 4806 Media, 4806 Description Information, 4807 Channel, 4807 Status Feedback Channel, 4808 Client Device, 4808 Client, 4808 Immersive Client Device, 4809 delivery storage device, 4809 device, 4809 storage device, 4810 fragmentation process, 4810 process, 4811 media analyzer, 4812 client tracking process, 4900 media adaptation process, 4901 adaptation process, 4901 media adaptation process, 4903 network orchestrator, 4905 input network status, 4906 client adaptation media storage device, 4907 scene analyzer, 4912 database, 5000 process, 6608 composite media, 14051 game engine control logic, 14051 control logic, 14052 GPU interface, 14053 physics engine, 14054 renderer, 14054 renderer process, 14055 compression decoder, 14056 device specific plugin, 300A timed scene manifest, 605A sensor, 811B storage device, 811C Visual presentation unit, 812C storage device, 812D holographic visualization unit, 814B storage device, 814C battery, 814D stereoscopic visual presentation component, 815C storage device, 815D eye tracking device, 815E camera, 815F light field panel, 908A game engine, 908BNetwork interface, 908D Storage cache, 908D Storage, 908E Smart client, 908F Callback function, 1001A Neural network model, 1001B Renderer, 1001C Neural network processor, 1001D Media compressor, 1001E Media decompressor, 1001F Logic controller, 11 04A manifest information, 1104B list, 4400A timed scene manifest, 4808A visualization component, 4808B network interface, 4808D storage cache, 4901B renderer, 4901C neural network processor, 4901D media compressor, 4901E media decompressor, 4901F logic controller
Claims
1. 1. A method of immersive media processing, comprising: receiving, by a network device, scene-based immersive media for playback on a light field-based display, the scene-based immersive media including a plurality of immersive media scenes; assigning a priority value to each of the plurality of scenes of the scene-based immersive media; determining an order for streaming the plurality of scenes to an end device according to the priority values; A method comprising:
2. The step of determining the order comprises: re-ranking the plurality of scenes according to the priority values; transmitting the re-ordered scenes to the end device; The method of claim 1 further comprising:
3. The step of determining the order comprises: selecting, by a priority aware network device, a highest priority scene from a subset of untransmitted scenes in the plurality of scenes, the highest priority scene having the highest priority value; transmitting the highest priority scene to the end device; The method of claim 1 further comprising:
4. The step of assigning priority values comprises: determining a priority value for a scene based on the likelihood of needing to render said scene; The method of claim 1 further comprising:
5. determining that available network bandwidth is limited; selecting a highest priority scene from a subset of untransmitted scenes in the plurality of scenes, the highest priority scene having the highest priority value; transmitting the highest priority scene in response to the available network bandwidth being limited; The method of claim 1 further comprising:
6. determining that available network bandwidth is limited; identifying a subset of the plurality of scenes that is unlikely to be required for next rendering based on the priority values; withholding the subset of the plurality of scenes from streaming in response to the available network bandwidth being limited; The method of claim 1 further comprising:
7. The step of assigning priority values comprises: assigning a first priority value to the first scene based on a second priority value of the second scene depending on a relationship between the first scene and a second scene; The method of claim 1 further comprising:
8. receiving a feedback signal from the end device; adjusting a priority value of at least one scene of the plurality of scenes based on the feedback signal; The method of claim 1 further comprising:
9. assigning the highest priority to a first scene in response to the feedback signal indicating that the current scene is a second scene related to the first scene; The method of claim 8 further comprising:
10. The method of claim 8 , wherein the feedback signal indicates an adjustment of a priority determined by the end device.
11. 1. A method of immersive media processing, comprising: receiving, by an end device having a light field-based display, a Media Presentation Description (MPD) of scene-based immersive media for playback by the light field-based display, the scene-based immersive media comprising multiple scenes of immersive media, the MPD indicating sequential streaming of the multiple scenes to the end device based on priority; detecting bandwidth availability; determining a change in order of the at least one scene by determining a change in priority of the at least one scene based on the bandwidth availability; transmitting a feedback signal indicating the change in the order of the at least one scene; A method comprising:
12. The method of claim 11 , wherein the feedback signal indicates a next scene to render.
13. The method of claim 11 , wherein the feedback signal indicates an adjustment to a priority value of the at least one scene.
14. The method of claim 11 , wherein the feedback signal indicates a current scene.
15. 1. A method of immersive media processing, comprising: receiving, by a network device, scene-based immersive media for playback on a light field-based display, the scene of the scene-based immersive media including a first order of a plurality of assets; assigning priority values to each of the plurality of assets in the scene according to one or more attributes of the plurality of assets; determining a second order for streaming the plurality of assets to an end device according to the priority values, the second order being different from the first order; A method comprising:
16. assigning priority values to assets of the scene according to the size of the assets; 16. The method of claim 15, comprising:
17. assigning priority values to assets of the scene according to the visibility of the assets associated with a default entry position of the scene; 16. The method of claim 15, comprising:
18. assigning a first set of priority values to the plurality of assets based on a first attribute of the plurality of assets, the first set of priority values being used to order the plurality of assets in a first prioritization scheme; assigning a second set of priority values to the plurality of assets based on second attributes of the plurality of assets, the second set of priority values being used to order the plurality of assets in a second prioritization scheme; selecting a prioritization scheme for the end device from the first prioritization scheme and the second prioritization scheme according to information of the end device; ordering the plurality of assets for streaming according to the selected prioritization scheme; 16. The method of claim 15, comprising:
19. assigning a first priority value to the first asset and a second priority value to the second asset in response to the first asset having a higher computational complexity than the second asset, the first priority value being higher than the second priority value; 16. The method of claim 15, comprising:
20. 20. Apparatus for carrying out the method of any one of claims 1 to 10, any one of claims 11 to 14, or any one of claims 15 to 19.
21. A computer program comprising instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 10, claims 11 to 14, or claims 15 to 19.
Citation Information
Patent Citations
Communication system and its communication method
JP2001359067A
Data processor, data processing method, program and storage medium
JP2005176094A
Adaptive streaming of an immersive video scene
US20170374411A1
A codec for processing scenes of almost unlimited detail
US20210133929A1