A smart client for streaming scene-based immersive media to game engines
The smart client mechanism addresses the inefficiencies in delivering scene-based media to diverse client devices by optimizing asset reuse and conversion processes, ensuring efficient media delivery and improved user experience.
Patent Information
- Application Number
- JP2024534586
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2023-04-04
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-04-04
AI Technical Summary
Existing media distribution ecosystems are unable to efficiently deliver scene-based immersive media to heterogeneous client devices, such as game engines, due to the lack of a consistent end-to-end ecosystem and insufficient information about client capabilities and media format requirements, leading to suboptimal network resource usage and repetitive conversion processes.
A smart client mechanism that interacts with a network server to provide information about client device capabilities and availability, enabling efficient delivery of scene-based media by optimizing asset reuse and conversion processes, and acting as an intermediary between client devices and the network.
Facilitates efficient delivery of scene-based media to heterogeneous client devices by reducing redundant conversions and optimizing resource usage, thereby enhancing the immersive experience for end-users.
Smart Images

Figure 0007797654000001 
Figure 0007797654000002 
Figure 0007797654000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 332,853, entitled "Smart Client for Streaming Scene-Based Immersive Media to a Game Engine," filed April 20, 2022, and U.S. Provisional Application No. 18 / 126,120, entitled "Smart Client for Streaming Scene-Based Immersive Media to a Game Engine," filed March 24, 2023, and U.S. Provisional Application No. 63 / 345,814, entitled "Smart Client for Streaming Visually Immersive Media Assets for Personalized Experiences," filed May 25, 2022. No. 63 / 346,105, entitled "Non-visual Immersive Media Asset Replacement for Personalized Experiences," filed May 26, 2022; U.S. Provisional Application No. 63 / 351,218, entitled "Smart Controller for Network-Based Media Adaptation," filed June 10, 2022; and U.S. Provisional Application No. 63 / 400,364, entitled "Client Scene-Based Immersive Media Profile for Supporting Heterogeneous Rendering-Based Clients," filed August 23, 2022. The entire disclosures of the prior applications are incorporated herein by reference in their entireties.
[0002] This disclosure describes embodiments that generally relate to media processing and distribution. [Background technology]
[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. To the extent that it is described in this background section, the inventors' work, as well as aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.
[0004] Immersive media generally refers to media that stimulates any or all human sensory systems (sight, hearing, somatosensation, smell, and sometimes taste) to create or enhance the user's perception of being physically present in the media experience, such as beyond that delivered over existing commercial networks for timed two-dimensional (2D) video and corresponding audio known as "legacy media." Both immersive and legacy media can be characterized as either timed or non-timed media.
[0005] Timed media refers to media that is structured and presented according to time. Examples include feature films, news reports, and episodic content, all of which are organized according to periods of time. Traditional video and audio are commonly considered to be timed media.
[0006] Non-timed media is media that is not structured by time, but rather by logical, spatial, and / or temporal relationships. One example includes a video game in which a user controls the experience created by a gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media may incorporate timed media, for example, in a continuously looped audio or video segment of a video game scene. Conversely, timed media may incorporate non-timed media, such as a video with a fixed still image as a background.
[0007] An immersive media-enabled device can refer to a device with the ability to access, interpret, and present immersive media. Such media and devices are heterogeneous with respect to the number and format of media and the number and type of network resources required to distribute such media on a large scale, i.e., to achieve distribution equivalent to legacy video and audio media over a network. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogeneous in their capabilities because all of these devices are configured with rectangular display screens and consume 2D rectangular video or still images as their primary media format. Summary of the Invention [Means for solving the problem]
[0008] Aspects of the present disclosure provide a method and apparatus (electronic device) for media processing. In some examples, the electronic device includes processing circuitry for performing a process of a smart client, which is a client interface of the electronic device. The method for media processing includes transmitting, by the client interface of the electronic device, to a server device in a network (e.g., an immersive media streaming network) information about the capability and availability of the electronic device for playing scene-based immersive media. The method further includes receiving, by the client interface, a media stream carrying adapted media content for the scene-based immersive media. The adapted media content is generated from the scene-based immersive media by the server device based on the capability and availability information. Thereafter, the method includes playing the scene-based immersive media according to the adapted media content.
[0009] In some examples, the method includes determining, by a client interface, that a first media asset associated with a first scene is received for the first time and should be reused in one or more scenes in accordance with the adapted media content, and storing the first media asset in a cache device accessible by the electronic device.
[0010] In some examples, the method includes extracting, by a client interface, a first list of unique assets in the first scene from the media stream, the first list of unique assets identifying the first media asset as a unique asset in the first scene to be used in one or more other scenes.
[0011] In some examples, the method includes transmitting, by the client interface, a signal to the server device indicating availability of the first media asset at the electronic device, the signal causing the server device to substitute a proxy for the first media asset in the adapted media content.
[0012] In some examples, the method includes determining, by a client interface, that a first media asset has previously been stored in a cache device according to a proxy in the adapted media content, and accessing the cache device to retrieve the first media asset.
[0013] In some examples, the method includes receiving a query signal for a first media asset from a server device and transmitting a signal indicating availability of the first media asset at the electronic device in response to the query signal.
[0014] In some examples, the method includes receiving, by a client interface, a request to obtain device attributes and resource status from a server device; querying one or more internal components of the electronic device and / or one or more external components associated with the electronic device regarding the attributes of the electronic device and resource availability for processing the scene-based immersive media; and transmitting the attributes and resource availability of the electronic device to the server device.
[0015] In some examples, the method includes receiving a request for scene-based immersive media from a user interface and forwarding, by the client interface, the request for the scene-based immersive media to a server device.
[0016] In some examples, the method includes generating reconstructed scene-based immersive media based on decoding of the media stream and media reconstruction under control of the client interface, and providing the reconstructed scene-based immersive media to a game engine for playback via an application programming interface (API) of the game engine of the electronic device.
[0017] In some examples, the method includes depacketizing, by a client interface, the media stream to generate depacketized media data; providing, via an application programming interface (API) of the game engine of the electronic device, the depacketized media data to a game engine; and generating, by the game engine, reconstructed scene-based immersive media for playback based on the depacketized media data.
[0018] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for media processing.
[0019] Further features, nature and various advantages of the disclosed subject matter may become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0020] [Figure 1] 1 illustrates the media flow process in some examples. [Figure 2] 1 illustrates the media transformation decision-making process in some examples. [Figure 3] One example shows the display of formats in timed heterogeneous immersive media. [Figure 4] An example shows the presentation of non-timed heterogeneous immersive media in a streamable format. [Figure 5] 1 shows a diagram illustrating the process of synthesizing media from natural content into an ingest format in some examples. [Figure 6] 1 shows a diagram of a process for creating an ingest format for composite media in some examples. [Figure 7] FIG. 1 is a schematic diagram of a computer system according to one embodiment. [Figure 8] 1 illustrates a network media distribution system that supports a variety of legacy and heterogeneous immersive media capable displays as client endpoints in some examples. [Figure 9] 1 illustrates a diagram of an immersive media delivery module capable of servicing legacy and heterogeneous immersive media capable displays in some examples. [Figure 10] 1 shows a diagram of the media adaptation process in some examples. [Figure 11] 1 illustrates a distribution format creation process in some examples. [Figure 12] 1 illustrates a packetizer processing system in some examples. [Figure 13]1 illustrates a sequence diagram of a network that adapts, in some examples, specific immersive media in an ingest format into a streamable and appropriate delivery format for a specific immersive media client endpoint. [Figure 14] 1 illustrates a diagram of a media system with a virtual network and client devices for scene-based media processing in some examples. [Figure 15] FIG. 1 illustrates a media flow process for media delivery to a client device over a network in some examples. [Figure 16] 1 illustrates a diagram of a media flow process with asset reuse into a game engine process in some examples. [Figure 17] 1 illustrates a diagram of a media conversion decision-making process involving asset reuse logic and redundant caching for a client device in some examples. [Figure 18] FIG. 10 is a diagram of asset reuse logic using a Smart Client in some examples. [Figure 19] 10 shows a diagram of a process for capturing device state and profile in a Smart Client in some examples. [Figure 20] 1 shows a process diagram for illustrating a Smart Client on a client device requesting and receiving media streamed from a network on behalf of a game engine on the client device. [Figure 21] 1 shows a diagram of timed media presentations ordered by descending frequency in some examples. [Figure 22] 1 shows a diagram of non-timed media presentations ordered by frequency in some examples. [Figure 23] 1 illustrates a delivery format creation process using ordered frequencies in some examples. [Figure 24] 10 illustrates a flowchart of the logic flow for a media reuse analyzer in some examples. [Figure 25]10 shows a diagram of a network process communicating with a Smart Client and game engine in a client device in some examples. [Figure 26] 1 shows a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 27] 1 shows a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 28] 1 shows a diagram of a timed media presentation with signaling of visual asset replacement in some examples. [Figure 29] 1 shows a diagram of a non-timed media presentation with signaling of visual asset replacement in some examples. [Figure 30] 1 shows diagrams of timed media presentations with signaling of substitutions in non-visual assets in some examples. [Figure 31] 1 shows a diagram of a non-timed media presentation with signaling of non-visual asset replacement in some examples. [Figure 32] 10 illustrates a process for performing asset replacement with a user-provided asset on a client device according to some embodiments of the present disclosure. [Figure 33] 1 illustrates a process for populating a user-provided media cache on a client device according to some embodiments of the present disclosure. [Figure 34] 1 shows a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 35] 1 illustrates a process for performing network-based media conversion in some examples. [Figure 36] 1 illustrates a process for performing network-based media adaptation in a smart controller in a network in some examples. [Figure 37] 1 shows a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 38] 1 shows diagrams of client media profiles in some examples. [Figure 39]1 shows a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 40] 1 shows a flowchart outlining a process according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0021] Aspects of the present disclosure provide architectures, structures, components, techniques, systems, and / or networks for delivering media including video, audio, geometric (3D) objects, haptics, associated metadata, or other content for client devices. In some examples, the architectures, structures, components, techniques, systems, and / or networks are configured to deliver media content to heterogeneous immersive and interactive client devices, such as game engines.
[0022] Immersive media generally refers to media that stimulates any or all human sensory systems (visual, auditory, somatosensory, olfactory, and sometimes gustatory) to create or enhance the user's perception of being physically present in the media experience, i.e., beyond that delivered over existing commercial networks for timed two-dimensional (2D) video and corresponding audio known as "legacy media." In some instances, immersive media refers to media that attempt to create or mimic the physical world through digital simulation of kinetics and the laws of physics, thereby stimulating any or all human sensory systems to create the user's perception of being physically present in a scene depicting the real or virtual world. Both immersive and legacy media can be characterized as either timed or non-timed media.
[0023] Timed media refers to media that is structured and presented according to time. Examples include feature films, news reports, and episodic content, all of which are organized according to periods of time. Traditional video and audio are commonly considered to be timed media.
[0024] Non-timed media is media that is not structured by time, but rather by logical, spatial, and / or temporal relationships. One example includes a video game in which a user controls the experience created by a gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media may incorporate timed media, for example, in a continuously looped audio or video segment of a video game scene. Conversely, timed media may incorporate non-timed media, such as a video with a fixed still image as a background.
[0025] An immersive media-capable device can refer to a device with sufficient resources and capabilities to access, interpret, and present immersive media. Such media and devices are heterogeneous with respect to the number and format of media and the number and type of network resources required to distribute such media on a large scale, i.e., to achieve distribution equivalent to legacy video and audio media over a network. Similarly, media are heterogeneous in the amount and type of network resources required to distribute such media on a large scale. "On a large scale" can refer to distribution of media by a service provider, e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions, that achieves distribution equivalent to distribution of legacy video and audio media over a network.
[0026] In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are uniform in their capabilities because all of these devices are configured with rectangular display screens and consume 2D rectangular video or still images as their primary media format. Similarly, the number of audio formats supported by legacy devices is limited to a relatively small set.
[0027] The term "frame-based" media refers to the property that visual media is composed of one or more consecutive rectangular frames of an image. In contrast, "scene-based" media (e.g., scene-based immersive media) refers to visual media that is organized by "scenes," where each scene, in some instances, refers to individual assets that collectively describe a visual scene.
[0028] A comparison between frame-based and scene-based visual media can be described using visual media that shows a forest. In a frame-based display, the forest is captured using a camera device, such as a camera phone. A user can enable the camera device to focus on the forest, and the frame-based media ingested by the camera device is the same as what the user sees through the camera viewport provided to the camera device, including any user-initiated camera device movements. The resulting frame-based display of the forest is a series of 2D images recorded by the camera device at a standard rate, typically 30 or 60 frames per second. Each image is a collection of pixels where the information stored in each pixel matches the next pixel.
[0029] In contrast, a scene-based representation of a forest is composed of individual assets that describe each of the objects in the forest. For example, a scene-based representation may include individual objects called "trees," with each tree being composed of a collection of smaller assets called "trunks," "branches," and "leaves." Each tree trunk may be further described individually by a mesh (tree trunk mesh) that describes the complete 3D geometry of the tree trunk and a texture that is applied to the tree trunk mesh to capture the color and radiance characteristics of the tree trunk. Furthermore, the tree trunk may be accompanied by additional information that describes the surface of the tree trunk in terms of its smoothness or roughness or its ability to reflect light. The individual assets that make up a scene differ in the type and amount of information stored in each asset.
[0030] Yet another difference between scene-based media and frame-based media is that in frame-based media, the view created of a scene is identical to the view captured by a user via a camera, i.e., when the media was created. When frame-based media is presented by a client, the presented view of the media is the same as the view ingested into the media by, for example, the camera used to record the video. However, in scene-based media, there may be multiple ways for a user to view a scene.
[0031] A client device that supports scene-based media may be equipped with renderers and / or resources (e.g., GPU, CPU, local media cache storage) whose capabilities and supported features collectively comprise an upper limit or cap to characterize the client device's overall ability to ingest various scene-based media formats. For example, a mobile handset client device may be limited in the complexity of the geometric assets it can render, e.g., the number of polygons describing the geometric assets, particularly for support of real-time applications. Such a limit may be established based on the fact that the mobile client is battery-powered, and therefore similarly limits the amount of computational resources available for performing real-time rendering. In such a scenario, it may be desirable for the client device to inform the network that the client prefers to access geometric assets with polygon counts below a specified upper limit. Furthermore, information conveyed from the client to the network may best be conveyed using a well-defined protocol that leverages a well-defined vocabulary of attributes.
[0032] Similarly, a media delivery network may have computational resources that facilitate the delivery of immersive media in various formats to various clients with various capabilities. In such a network, it may be desirable for the network to be informed of client-specific capabilities according to a well-defined profile protocol, e.g., a vocabulary of attributes communicated via the well-defined protocol. Such an attribute vocabulary may include information to describe the media or the minimum computational resources required to render the media in real time, so that the network can better establish priorities for how to serve media to its heterogeneous clients. Furthermore, a centralized data store in which client-provided profile information is collected across client domains can help provide a summary of which types of assets and in which formats are in high demand. By being provisioned with information about which types of assets are in higher and lower demand, an optimized network can prioritize the task of responding to requests for assets in higher demand.
[0033] In some examples, distribution of media over a network can employ media distribution systems and architectures that reformat media from an input or network "ingest" media format to a distributed media format. In one example, the distributed media format is not only suitable for ingestion by a target client device and its applications, but also lends itself to being "streamed" over the network. In some examples, there can be two processes performed on media ingested by the network: 1) converting the media from format A to format B suitable for ingestion by the target client device, i.e., based on the client device's ability to ingest a particular media format, and 2) preparing the media to be streamed.
[0034] In some examples, "streaming" media broadly refers to fragmenting and / or packetizing media so that the processed media can be delivered over a network in successive, smaller-sized "chunks" that are logically organized and sequenced according to either or both the temporal or spatial structure of the media. In some examples, "converting" media from format A to format B (sometimes referred to as "transcoding") may be a process typically performed by a network or service provider prior to delivering the media to a target client device. Such transcoding may involve converting media from format A to format B based on prior knowledge that format B is somehow the preferred or only format that can be ingested by the target client device, or is more suitable for delivery over constrained resources, such as a commercial network. One example of media conversion is converting media from a scene-based representation to a frame-based representation. In some examples, both steps of converting media and preparing the media to be streamed are necessary before the media can be received and processed by the target client device from the network. Such prior knowledge of client-preferred formats can be obtained through the use of well-defined profile protocols that utilize agreed-upon attribute vocabularies that summarize the characteristics of preferred scene-based media across a variety of client devices.
[0035] In some examples, the one- or two-step process described above operates on media ingested by the network, i.e., results in a media format referred to as a "delivery media format" or simply a "delivery format," prior to delivering the media to a target client device. Generally, these steps, if performed for a given media data object, can be performed only once if the network has access to information indicating that the target client device needs the converted and / or streamed media object for multiple occasions that would otherwise trigger the conversion and streaming of such media multiple times. That is, the processing and transfer of data for media conversion and streaming is generally considered a source of latency that requires the consumption of potentially significant amounts of network and / or computational resources. Thus, a network design that does not have access to information to indicate when a client device may already have a particular media data object stored in its cache or locally to the client device will be suboptimal for a network that has access to such information.
[0036] In some instances, for legacy presentation devices, the delivery format may be equivalent or sufficiently equivalent to the "presentation format" ultimately used by a client device (e.g., client presentation device) to create the presentation. For example, a presentation media format is a media format whose properties (resolution, frame rate, bit depth, color gamut, etc.) are closely aligned with the capabilities of the client presentation device. Some examples of delivery formats versus presentation formats include a high-definition (HD) video signal (1920 pixel columns by 1080 pixel rows) delivered over a network to an ultra-high-definition (UHD) client device with a resolution (3840 pixel columns by 2160 pixel rows). For example, the UHD client device may apply a process known as "super-resolution" to the HD delivery format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format presented by the UHD client device is the "presentation format," which in this example is a UHD signal, but the HD signal includes the delivery format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format since both signals are linear video formats, and the process of converting the HD format to the UHD format is relatively simple and easy to perform on most legacy client devices.
[0037] In some instances, the preferred presentation format for the target client device may be significantly different from the ingest format received by the network. Nevertheless, the target client device may have access to sufficient computational, storage, and bandwidth resources to convert the media from the ingest format to the required presentation format suitable for presentation by the target client device. In this scenario, the network may bypass the step of reformatting the ingested media from format A to format B, e.g., "transcoding" the media, simply because the client has access to sufficient resources to perform all media conversions without the network having to do so. However, the network may still perform the steps of fragmenting and packaging the ingested media so that the media can be streamed to the target client device.
[0038] In some instances, the ingested media received by the network is significantly different from the target client device's preferred presentation format, and the target client device does not have access to sufficient computational, storage, and / or bandwidth resources to convert the media to the preferred presentation format. In such scenarios, the network can assist the target client device by performing some or all of the conversion on behalf of the target client device from the ingest format to a format that is equivalent or nearly equivalent to the target client device's preferred presentation format. In some architectural designs, such assistance provided by the network on behalf of the target client device is referred to as "split rendering."
[0039] FIG. 1 illustrates a media flow process 100 (also referred to as process 100) in some examples. Media flow process 100 includes a first step that can be performed by a network cloud (or edge device) 104 and a second step that can be performed by a client device 108. In some examples, media in ingest media format A is received over a network from a content provider in step 101. A network process step, step 102, can prepare the media for delivery to a client device 108 by formatting the media into format B or by preparing the media to be streamed to the client device 108. In step 103, the media is streamed from the network cloud 104 to the client device 108 via a network connection 105. The client 108 can receive the delivery media and prepare the media for presentation via a rendering process, indicated by 106. The output of the rendering process 106 is presentation media in yet another, potentially different, format C, indicated by 107.
[0040] FIG. 2 illustrates a media conversion decision-making process 200 (also referred to as process 200) illustrating a network logic flow for processing ingested media within a network (also referred to as a network cloud), for example, by one or more devices within the network. At 201, media is ingested by the network cloud from a content provider. Attributes of the target client device are obtained at 202, if not already known. Decision-making step 203 determines whether the network should support conversion of the media, if necessary. If decision-making step 203 determines that the network should support conversion, the ingested media is converted by process step 204, which converts the media from format A to format B to generate converted media 205. At 206, the media, either converted or in its original form, is prepared to be streamed. At 207, the prepared media is appropriately streamed to a target client device, such as a game engine client device.
[0041] An important aspect to the logic of Figure 2 is the decision-making process 203, which may be performed by an automated process. That decision-making step may determine whether the media can be streamed in its original ingest format A, or whether the media needs to be converted to a different format B to facilitate presentation of the media by the target client device.
[0042] In some examples, decision-making process step 203 may require access to information describing aspects or characteristics of the ingest media to assist decision-making process step 203 in making the optimal choice, i.e., to determine whether conversion of the ingest media is required before streaming the media to the target client device, or whether the media can be streamed directly to the target client device in its original ingest format A.
[0043] According to one aspect of the present disclosure, streaming of scene-based immersive media may differ from streaming frame-based media. For example, streaming of frame-based media may be equivalent to streaming frames of video, with each frame capturing a complete picture of an entire scene or object presented by a client device. The sequence of frames is reconstructed from their compressed form by the client device and, when presented to a viewer, creates a video sequence that includes the entire immersive presentation or a portion of the presentation. For frame-based media streaming, the order in which frames are streamed from the network to the client device may conform to a predetermined specification, such as ITU-T Recommendation H.264 Advanced Video Coding for General Audiovisual Services.
[0044] However, scene-based streaming of media differs from frame-based streaming because a scene may be composed of individual assets that may themselves be independent of one another. A given scene-based asset may be used multiple times within a particular scene or across a series of scenes. The amount of time a client device, or any given renderer, needs to create the correct presentation of a particular asset may depend on many factors, including, but not limited to, the size of the asset, the availability of computational resources to perform the rendering, and other attributes that describe the overall complexity of the asset. Client devices that support scene-based streaming may require some or all of the rendering of each asset in a scene to be completed before any presentation of the scene can begin. Therefore, the order in which assets are streamed from the network to the client device can affect overall performance.
[0045] According to one aspect of the present disclosure, considering each of the above scenarios in which the conversion of media from format A to another format may be performed entirely by the network, entirely by the client device, or jointly between both the network and the client device, for example, for split rendering, a vocabulary of attributes describing the media format may be needed so that both the client device and the network have complete information characterizing the conversion operation. Furthermore, a vocabulary providing attributes of the client device's capabilities, e.g., with respect to available computational resources, available storage resources, and access to bandwidth, may similarly be needed. Furthermore, a mechanism may be needed to characterize the level of computational, storage, or bandwidth complexity of an ingest format so that the network and client device, together or alone, can decide whether or when the network employs a split rendering step to deliver media to the client device. Furthermore, if conversion and / or streaming of certain media objects required or required by the client device to complete the presentation of the media can be avoided, the network may skip the conversion and streaming steps, assuming the client device has access to or availability of the media objects that the client device may need to complete the client device's presentation of the media. To facilitate the ability of client devices to perform to their full potential, it may be desirable for the network to be equipped with sufficient information regarding the order in which scene-based assets are streamed from the network to client devices so that the network can determine such order to improve the performance of the client devices. For example, such a network with sufficient information to avoid repetitive conversion and / or streaming steps for assets that are used multiple times in a particular presentation may perform more optimally than a network not designed to do so.Similarly, a network that can "intelligently" sequence the delivery of assets to clients can facilitate the ability of client devices to perform to their full potential, i.e., to create a more enjoyable experience for end users. Furthermore, the interface between client devices and the network (e.g., a server device within the network) may be implemented using one or more communication channels that convey information about the operating state of the client device, the availability of resources at or local to the client device, the type of media being streamed, and the frequency of assets used, or across multiple scenes. Thus, a network architecture that implements streaming of scene-based media to heterogeneous clients may require access to a client interface that can provide and update information related to the processing of each scene to a network server process, including current conditions related to the client device's ability to access computational and storage resources. Such a client interface may also interact closely with other processes running on the client device, particularly a game engine, which can play an essential role on behalf of the client device's ability to deliver an immersive experience to an end user. An example of an essential role a game engine can play includes providing an application program interface (API) to enable the delivery of interactive experiences. Another role that may be provided by the game engine on behalf of the client device is the rendering of the precise visual signals required by the client device to provide a visual experience that matches the capabilities of the client device.
[0046] Definitions of some terms used in this disclosure are provided in the following paragraphs.
[0047] Scene graph: A common data structure typically used by vector-based graphics editing applications and modern computer games that constitutes a logical and often (but not necessarily) spatial representation of a graphical scene; it is a collection of nodes and vertices in a graph structure.
[0048] Scene: In the context of computer graphics, a scene is a collection of objects (e.g., 3D assets), object attributes, and other metadata, including visual, acoustic, and physics-based characteristics, that describe a particular setting, bounded either by space or time, with respect to the interactions of objects within that setting.
[0049] Node: A basic element of a scene graph consisting of information related to the logical, spatial, or temporal representation of visual, audio, tactile, olfactory, gustatory, or related processing information; each node has at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) attached to it.
[0050] Base Layer: A nominal representation of an asset, typically formulated to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network.
[0051] Enhancement Layer: A set of information that, when applied to a base layer representation of an asset, augments the base layer to include features or capabilities not supported in the base layer.
[0052] Attribute: Metadata associated with a node that is used to describe a particular characteristic or feature of that node (e.g., with respect to another node) in either a canonical form or a more complex form.
[0053] Container: A serialization format for storing and exchanging information to represent an entire natural scene, an entire synthetic scene, or a combination of synthetic and natural scenes, including a scene graph and all the media resources required to render the scene.
[0054] Serialization: The process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link), and later reconstructed (possibly in a different computing environment). When the resulting series of bits is reread according to the serialized format, it can be used to create a semantically identical clone of the original object.
[0055] Renderer: A (typically software-based) application or process based on a selective combination of academic disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development that, given an input scene graph and asset container, emits exemplary visual and / or audio signals suitable for presentation on a target device or adapted to desired properties specified by attributes of render target nodes in the scene graph. In the case of visual-based media assets, a renderer can emit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline); in the case of audio-based media assets, a renderer can emit audio signals for presentation over multi-channel loudspeakers and / or binaural headphones or for repackaging into another (output) container. Common examples of renderers include the real-time rendering capabilities of game engines such as Unity Engine and Unreal Engine.
[0056] Evaluation: Generate results that move the output from summary to concrete results (e.g., similar to evaluating a document object model for a web page).
[0057] Scripting Language: An interpreted programming language that can be executed by the renderer at runtime to process dynamic inputs and variable state changes applied to scene graph nodes, which affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformations, collisions) and energy propagation and transport (light, sound).
[0058] Shader: a type of computer program originally used for shading (producing appropriate levels of light and color in an image), but now performing a variety of specialized functions in various areas of computer graphics special effects, or video post-processing unrelated to shading, or even functions completely unrelated to graphics.
[0059] Path Tracing: A computer graphics method for rendering three-dimensional scenes so that the lighting in the scene is realistic.
[0060] Timed Media: Media that is ordered by time, e.g., has a start time and an end time according to a particular clock.
[0061] Non-timed media: Media that is organized by spatial, logical, or temporal relationships, such as an interactive experience that is realized according to actions taken by a user.
[0062] Neural Network Model: A collection of parameters and tensors (e.g., matrices) that define weights (i.e., numbers) used in well-defined mathematical operations that are applied to a visual signal to arrive at an improved visual output, which may include the interpolation of new views of the visual signal that were not explicitly provided by the original signal.
[0063] Frame-based media: 2D video with or without associated audio.
[0064] Scene-based media: Audio, visual, tactile, and other major types of media, and media-related information organized logically and spatially using a scene graph.
[0065] Over the past decade, several immersive media-enabled devices have been introduced to the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and gaming consoles. Similarly, holographic displays and other forms of volumetric displays are poised to appear on the consumer market within the next three to five years. Despite the immediate or imminent availability of these devices, a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks has not materialized for several reasons.
[0066] One of the obstacles to achieving a consistent end-to-end ecosystem for the delivery of immersive media over commercial networks is the great diversity of client devices that serve as endpoints in such delivery networks for immersive displays. Some of these support specific immersive media formats, while others do not. Some of these are capable of creating immersive experiences from legacy raster-based formats, while others are not. Unlike networks designed solely for the delivery of legacy media, networks that must support a variety of display clients require a significant amount of information detailing each client's capabilities and the format of the media being delivered before such networks can use an adaptation process to convert the media into a format appropriate for each target display and corresponding application. At a minimum, such networks need access to information describing the characteristics of each target display and the complexity of the ingested media in order for the network to ascertain how to meaningfully adapt the input media source to a format appropriate for the target display and application. Similarly, a network optimized for efficiency may want to maintain a database of the types of media supported by client devices connected to such a network and their corresponding attributes.
[0067] Similarly, an ideal network that supports heterogeneous clients should take advantage of the fact that some of the assets that have been adapted from an input media format to a particular target format can be reused across a set of similar display targets. That is, once converted into a format appropriate for the target displays, some assets may be reused across several such displays that have similar adaptation requirements. Such an ideal network would therefore employ a caching mechanism to store the adapted assets in a relatively immutable area, i.e., similar to the use of content delivery networks (CDNs) used in legacy networks.
[0068] Furthermore, immersive media can be organized into "scenes," e.g., "scene-based media," that are described by a scene graph, also known as a scene description. The scope of a scene graph is to describe the visual, audio, and other forms of immersive assets that comprise a particular setting that is part of a presentation, e.g., actors and events taking place in a particular location within a building that is part of a presentation such as a movie. A list of all the scenes that comprise a single presentation may be formulated in a scene manifest.
[0069] An additional benefit of such an approach is that, for content that is prepared before such content must be delivered, a "bill of materials" can be created that identifies all of the assets used throughout the presentation and the frequency with which each asset is used across various scenes within the presentation. An ideal network should have knowledge of the existence of cached resources that can be used to satisfy the asset requirements of a particular presentation. Similarly, a client presenting a series of scenes may want to have knowledge of the frequency with which any given asset is used across multiple scenes. For example, if a media asset (also known as an "object") is processed by a client or is referenced multiple times across multiple scenes being processed, the client should avoid discarding that particular asset from its caching resources until the last scene requiring that asset has been presented by the client.
[0070] Finally, many emerging advanced imaging displays, including but not limited to Oculus Rift, Samsung Gear VR, Magic Leap goggles, all Looking Glass Factory displays, SolidLight by Light Field Labs, Avalon Holographic displays, and Dimenco displays, utilize game engines as the mechanism by which the respective displays can ingest content to be rendered and presented on the displays. Currently, the most popular game engines employed across this aforementioned set of displays include Unreal Engine by Epic Games and Unity by Unity Technologies. That is, advanced imaging displays are currently designed and shipped employing one or both of these game engines as the mechanism by which the displays can obtain media to be rendered and presented by such advanced imaging displays. Both Unreal Engine and Unity are optimized to ingest scene-based media rather than frame-based media. However, existing media distribution ecosystems are only capable of streaming frame-based media. There is a significant "gap" in the current media distribution ecosystem that includes standards (de jure or de facto standards) and best practices to enable the delivery of scene-based content to emerging advanced imaging displays so that media can be delivered "at scale," e.g., on the same scale as frame-based media is delivered.
[0071] The disclosed subject matter addresses the need for a mechanism or process that responds to a network server process on behalf of client devices for which a game engine is employed to ingest scene-based media and participates in the combination of network and immersive client architectures described herein. Such a "smart client" mechanism is particularly relevant to networks designed to stream scene-based media to immersive, heterogeneous, interactive client devices, where media delivery is efficient and within the constraints of the capabilities of the various components that make up the overall network. A "smart client" is associated with a particular client device and responds to network requests for information regarding the current state of that associated client device, including the availability of resources on the client device for rendering and creating presentations of scene-based media. The "smart client" also acts as an "intermediary" between the client devices for which a game engine is employed and the network itself.
[0072] Note that the remainder of the disclosed subject matter assumes, without loss of generality, that a smart client that can respond on behalf of a particular client device can also respond on behalf of client devices on which one or more other applications (i.e., not game engine applications) are active. That is, the problem of responding on behalf of a client device is equivalent to the problem of responding on behalf of a client device on which one or more other applications are active.
[0073] Additionally, it should be noted that the terms "media object" and "media asset" can be used interchangeably and both refer to a specific instance of media in a particular format. The term "client device" or "client" (without limitation) refers to the device and its components on which the presentation of media ultimately occurs. The term "game engine" refers to the Unity or Unreal engines, or any game engine that plays a role in a delivery network architecture. The term "smart client" refers to the subject matter of this document.
[0074] Referring back to FIG. 1 , media flow process 100 illustrates the flow of media through a network 104 or delivery to a client device 108 where a game engine is used. In FIG. 1 , processing of ingested media format A is performed by processing in a cloud or edge device 104. At 101, media is obtained from a content provider (not shown). Process step 102 performs any necessary conversion or conditioning of the ingested media to create a potential alternative representation of the media as delivery format B. Media formats A and B may or may not be representations that follow the same syntax of a particular media format specification, but format B will likely be tailored in a manner that facilitates delivery of the media over a network protocol such as TCP or UDP. Such “streamable” media is shown as streamed media to a client device 108 over a network connection 105. The client device 108 has access to several rendering functions, shown as 106. Such rendering functions 106 may be rudimentary or similarly sophisticated, depending on the client device 108 and the type of game engine running on the client device. The rendering process 106 creates presentation media that may or may not be displayed according to a third format specification, e.g., Format C. In some examples, in a client device that uses a game engine, the rendering process 106 is typically a function provided by the game engine.
[0075] Referring to FIG. 2, a media conversion decision-making process 200 can be used to determine whether the network needs to convert media before delivering it to a client device. In FIG. 2, ingested media 201, represented in format A, is provided to the network by a content provider (not shown). Process step 202 obtains attributes describing the processing capabilities of the target client (not shown). Decision-making process step 203 is used to determine whether the network or the client should perform format conversion of any of the media assets contained within ingested media 201, such as converting a particular media object from format A to format B, before the media is streamed to the client. If any of the media assets need to be converted by the network, the network uses process step 204 to convert the media object from format A to format B. Converted media 205 is the output from process step 204. The converted media is merged with preparation process 206 to prepare the media to be streamed to a game engine client (not shown). Process step 207, for example, streams the prepared media to the game engine client.
[0076] Figure 3 shows a representation of a streamable format 300 for heterogeneous immersive media that is timed in one example, and Figure 4 shows a representation of a streamable format 400 for heterogeneous immersive media that is non-timed in one example. In the case of Figure 3, Figure 3 references a timed media scene 301. In the case of Figure 4, Figure 4 references a non-timed media scene 401. In either case, the scene may be embodied by various scene representations or scene descriptions.
[0077] For example, in some immersive media designs, a scene may be embodied by a scene graph, or as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI and MSI technologies are examples of technologies that support the creation of display-independent scene representations for natural content, i.e., real-world images captured simultaneously from one or more cameras. On the other hand, while scene graph technology can be used to display both natural and computer-generated imagery in the form of a synthetic representation, such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. That is, scene graph representations of naturally captured content are time-consuming and computationally intensive to create, requiring complex analysis of the natural imagery using photogrammetry or deep learning, or both, techniques, to create a synthetic representation that can then be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations are not currently practical to consider as candidates for displaying natural content because they cannot actually be created in real time to allow for use cases requiring real-time delivery. In some instances, the best candidate representation of a computer-generated image is through the use of a scene graph with a synthetic model, since the computer-generated image is created using 3D modeling processes and tools.
[0078] This dichotomy in optimal presentation of both natural and computer-generated content suggests that the optimal ingest format for naturally captured content may be different from the optimal ingest format for computer-generated or natural content that is not required for real-time delivery applications. Accordingly, the disclosed subject matter aims to be robust enough to support multiple ingest formats for visually immersive media, whether created naturally through the use of a physical camera or created by a computer.
[0079] The following are exemplary techniques for embodying a scene graph as a format suitable for displaying visually immersive media created using computer-generated techniques, or naturally captured content, whereby deep learning or photogrammetry techniques are employed to create a corresponding synthetic representation of the natural scene, i.e., not required for real-time distribution applications.
[0080] 1. ORBX (registered trademark) by OTOY OTOY's ORBX is one of several scene graph technologies capable of supporting any type of visual media, timed or non-timed, including ray-traceable, legacy (frame-based), stereoscopic, and other types of composite or vector-based visual formats. According to one aspect, ORBX is unique from other scene graphs because it provides native support for freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph purposefully designed to facilitate interchange across multiple vendor technologies that operate on the scene graph. Furthermore, ORBX offers a rich material system, support for an open shader language, a robust camera system, and support for Lua scripting. ORBX is also the basis for an immersive technology media format released for license on royalty-free terms by the Immersive Digital Experience Alliance (IDEA). In the context of real-time media distribution, the ability to create and deliver ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of data captured by cameras and composition of that same data into composite representations. To date, the availability of sufficient computation for real-time delivery is impractical, but nevertheless not impossible.
[0081] 2. Pixar's Universal Scene Description Pixar's Universal Scene Description (USD) is a scene graph that can be used in the visual effects (VFX) and professional content creation communities. USD is integrated into Nvidia's Omniverse platform, a set of developer tools for creating and rendering 3D models using Nvidia's GPUs. A subset of USD was released by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.
[0082] 3. Khronos glTF 2.0 glTF 2.0 is the latest version of the Graphics Language Transmission Format specification written by the Khronos 3D Group. This format supports simple scene graph formats, including "png" and "jpeg" image formats, that can generally support static (non-timed) objects in a scene. glTF 2.0 supports simple animation and supports translation, rotation, and scaling of basic shapes, i.e., geometric objects, described using glTF primitives. glTF 2.0 does not support timed media, and therefore does not support video or audio.
[0083] It should be noted that the above scene representations of immersive visual media are provided by way of example only and do not limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source to a format suited to the particular characteristics of a client endpoint device.
[0084] Additionally, any or all of the above exemplary media displays currently use or can use deep learning techniques to train and create neural network models that enable or facilitate the selection of specific views to fill a particular display's viewing frustum based on the particular dimensions of the frustum. The views selected for a particular display's viewing frustum may be interpolated from existing views explicitly provided in the scene representation, for example, from MSI or MPI techniques, or may be rendered directly from a rendering engine based on specific virtual camera positions, filters, or virtual camera descriptions for those rendering engines.
[0085] Thus, the disclosed subject matter is sufficiently robust to consider that there is a relatively small but well-known set of immersive media ingest formats that can adequately meet the requirements for both real-time or "on-demand" (e.g., non-real-time) delivery of media that is captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.
[0086] Interpolation of views from immersive media ingest formats using either neural network models or network-based rendering engines will become even easier as advanced network technologies, such as 5G for mobile networks and fiber optic cable for fixed networks, are deployed. These advanced network technologies increase the capacity and capability of commercial networks, as such advanced network infrastructure can support the transmission and delivery of increasingly large amounts of visual information. Network infrastructure management technologies, such as multi-access edge computing (MEC), software-defined networking (SDN), and network functions virtualization (NFV), enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands on specific network resources, e.g., to respond to dynamic increases and decreases in demand for network throughput, network speed, round-trip delay, and computational resources. Furthermore, this inherent ability to adapt to dynamic network requirements similarly facilitates the network's ability to adapt immersive media ingest formats to appropriate delivery formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.
[0087] Immersive media applications themselves may also have different requirements for network resources, including gaming applications that require extremely low network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications that may have increasing demands on downlink resources depending on the type of display at the client endpoint that is consuming the data. In general, any consumer application may be supported by a variety of client endpoints that include different on-board client capabilities for storage, computation, and power, as well as equally different requirements for the particular media display.
[0088] The disclosed subject matter thus enables a fully equipped network, i.e., a network that employs some or all of the characteristics of a modern network, to simultaneously support multiple legacy and immersive media capable devices in accordance with the characteristics specified below.
[0089] 1. Providing the flexibility to utilize a practical media ingest format for both real-time and "on-demand" use cases for the delivery of media.
[0090] 2. Provides flexibility to support both natural and computer-generated content for both legacy and immersive media-enabled client endpoints.
[0091] 3. Supports both timed and non-timed media.
[0092] 4. Provide a process for dynamically adapting source media ingest formats to appropriate delivery formats based on the capabilities and capabilities of the client endpoint and based on application requirements.
[0093] 5. Ensure that the delivery format is streamable over IP-based networks.
[0094] 6. Allows the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media capable devices.
[0095] 7. Provide an exemplary media presentation framework that facilitates the organization of distributed media along scene boundaries.
[0096] An end-to-end implementation of the improvements enabled by the disclosed subject matter is achieved according to the processes and components described in the detailed description that follows.
[0097] Figures 3 and 4 each employ an exemplary generic delivery format that can be adapted from an ingest source format to match the capabilities of a particular client endpoint. As previously noted, the media shown in Figure 3 is timed, while the media shown in Figure 4 is non-timed. The particular inclusion format is robust enough in its structure to accommodate a wide variety of media attributes, each of which can be layered based on the amount of salient information each layer contributes to the media presentation. Note that the layering process can be applied, for example, to progressive JPEG and scalable video architectures (e.g., as specified in ISO / IEC 14496-10 Scalable Advanced Video Coding).
[0098] According to one aspect, media streamed according to the encompassing media format is not limited to legacy visual and audio media, but may include any type of media information that can interact with a machine to generate signals that stimulate human sight, sound, taste, touch, and smell.
[0099] According to other aspects, media streamed according to the encompassing media format may be both timed media or non-timed media, or a mixture of both.
[0100] According to other aspects, the incorporating media format can be further streamlined by enabling hierarchical representations of media objects through the use of a base layer and enhancement layer architecture. In one example, separate base and enhancement layers are computed by application of multi-resolution or multi-tessellation analysis techniques to the media objects in each scene. This is similar to the progressively rendered image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In one example, the progressive representation of geometric objects can be a multi-resolution representation of the objects computed using wavelet analysis.
[0101] In another example of a layered representation of a media format, the enhancement layer applies different attributes to the base layer, such as modifying the material properties of the surface of the visual object represented by the base layer. In yet another example, the attributes may modify the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.
[0102] In yet another example of a layered display, the surfaces of one or more visual objects in a scene may be changed from Lambertian surfaces to ray-traceable surfaces.
[0103] In yet another example of a layered display, the network delivers a base layer display to a client so that the client can create a nominal presentation of the scene while awaiting the transmission of additional enhancement layers to refine the resolution or other characteristics of the base display.
[0104] According to another aspect, the resolution or refinement information of attributes in the enhancement layer is not explicitly coupled with the resolution of objects in the base layer as is currently the case in existing MPEG video and JPEG image standards.
[0105] According to another aspect, the encompassing media format supports any type of information media that can be presented or acted upon by a presentation device or machine, thereby enabling support of heterogeneous media formats to heterogeneous client endpoints. In one embodiment of a network that delivers media formats, the network first queries the client endpoint to determine the client's capabilities, and if the client is unable to meaningfully ingest the media representation, the network either removes layers of attributes not supported by the client or adapts the media from its current format to a format appropriate for the client endpoint. In one example of such adaptation, the network converts a volumetric visual media asset into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can use neural network processes to reformat the media into an appropriate format or, optionally, synthesize a view required by the client endpoint.
[0106] According to another aspect, a manifest for a complete or partially complete immersive experience (such as a live streaming event, game, or on-demand asset playback) is organized by scenes, which are the minimum amount of information that rendering and game engines can currently ingest to create the presentation. The manifest includes a list of individual scenes to be rendered in their entirety for the immersive experience requested by the client. Associated with each scene are one or more representations of the geometric objects in the scene that correspond to a streamable version of the scene geometry. One embodiment of a scene representation relates to a low-resolution version of the scene's geometric objects. Another embodiment of the same scene relates to enhancement layers for the low-resolution representation of the scene to add further detail or increase mosaicking to the geometric objects of the same scene. As previously mentioned, each scene may have two or more enhancement layers to progressively increase the detail of the scene's geometric objects.
[0107] According to another aspect, each layer of a media object referenced in a scene is associated with a token (e.g., a URI) that points to an address where the resource can be accessed in the network. Such a resource is similar to a CDN where content can be fetched by a client.
[0108] According to other aspects, tokens for representations of geometric objects can point to locations within the network or within the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.
[0109] FIG. 3 illustrates a timed media presentation 300 in some examples. The timed media presentation 300 describes an example of a comprehensive media format for timed media. The timed scene manifest 300A includes a list of scene information 301. The scene information 301 points to a list of components 302 that separately describe the processing information and types of media assets within the scene information 301. The components 302 point to assets 303, which further point to base layers 304 and attribute enhancement layers 305. In the example of FIG. 3, each of the base layers 304 points to a numerical frequency metric indicating the number of times the asset was used across the set of scenes in the presentation. A list of unique assets not previously used in other scenes is provided in 307. Proxy visual assets 306 include information about the reused visual assets, such as unique identifiers for the reused visual assets, and proxy audio assets 308 include information about the reused audio assets, such as unique identifiers for the reused audio assets.
[0110] FIG. 4 illustrates a non-timed media presentation 400 in some examples. The non-timed media presentation 400 describes an example of a generic media format for non-timed media. A non-timed scene manifest (not shown) references scene 1.0, which has no other scenes that can branch to it. Scene information 401 is not associated with a start or end time according to a clock. Scene information 401 points to a list of components 402 that separately describe the processing information and media asset types that make up the scene. Components 402 point to assets 403, which further point to base layer 404 and attribute enhancement layers 405 and 406. In the example of FIG. 4, each of the base layers 404 points to a numeric frequency value indicating the number of times an asset is used across the set of scenes in the presentation. Scene information 401 can also reference other scene information 401 for non-timed media. Scene information 401 can also reference scene information 407 for timed media scenes. List 408 identifies unique assets associated with a particular scene that have not been previously used in a higher-level (eg, parent) scene.
[0111] 5 shows a diagram of a process 500 for synthesizing an ingest format from natural content. The process 500 includes a first sub-process for content capture and a second sub-process for ingest format synthesis for natural images.
[0112] In the example of FIG. 5 , in the first sub-process, camera units may be used to capture natural image content 509. For example, camera unit 501 may use a single camera lens to capture a scene of a person. Camera unit 502 may mount five camera lenses around a ring-shaped object to capture a scene with five diverging fields of view. The arrangement in camera unit 502 is an exemplary arrangement for capturing omnidirectional content for VR applications. Camera unit 503 mounts seven camera lenses on the inner diameter of a sphere to capture a scene with seven converging fields of view. The arrangement in camera unit 503 is an exemplary arrangement for capturing a light field or a light field for a holographic immersive display.
[0113] In the example of Figure 5, in a second sub-process, natural image content 509 is synthesized. For example, the natural image content 509 is provided as input to a synthesis module 504, which in one example may employ a neural network training module 505 that uses a set of training images 506 to generate a captured neural network model 508. Another process commonly used in place of the training process is photogrammetry. When the model 508 is created during the process 500 shown in Figure 5, the model 508 becomes one of the assets of the natural content ingest format 507. Example embodiments of the ingest format 507 include MPI and MSI.
[0114] FIG. 6 shows a diagram of a process 600 for creating synthetic media 608, e.g., a computer-generated imagery ingest format. In the example of FIG. 6, a LIDAR camera 601 captures a point cloud 602 of a scene. Computer-generated imagery (CGI) tools, 3D modeling tools, or other animation processes for creating synthetic content are used on a computer 603 to create CGI assets 604 over a network. A motion capture suit with sensors 605A is worn by an actor 605 to capture a digital recording of the actor's 605 movements to generate animated motion capture (MoCap) data 606. Data 602, 604, and 606 are provided as inputs to a synthesis module 607, which can also create a neural network model (not shown in FIG. 6), for example, using a neural network and training data.
[0115] The techniques for displaying, streaming, and processing heterogeneous immersive media in this disclosure may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0116] Computer software can be coded using any suitable machine code or computer language, which may rely on mechanisms such as assembly, compilation, linking, etc., to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or via interpretation, microcode execution, etc.
[0117] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0118] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 700.
[0119] Computer system 700 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures), olfactory input (not shown), etc. Human interface devices may also be used to ingest certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0120] The input human interface devices may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touchscreen 710, a data glove (not shown), a joystick 705, a microphone 706, a scanner 707, and a camera 708 (only one of each is shown).
[0121] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 710, data gloves (not shown), or joystick 705, although there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers 709, headphones (not shown)), visual output devices (e.g., CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional hypervisible output via means such as stereo output 710, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0122] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (720), including CD / DVD or similar media 721, thumb drives 722, removable hard drives or solid state drives 723, legacy magnetic media (not shown) such as tape and floppy disks, and dedicated ROM / ASIC / PLD-based devices (not shown) such as security dongles.
[0123] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0124] The computer system 700 may also include an interface 754 to one or more communications networks 755. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus; and the like. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus 749 (e.g., a USB port on the computer system 700); others are generally integrated into the core of the computer system 700 by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 700 can communicate with other entities. Such communications may be one-way receive-only (e.g., television broadcast), one-way transmit-only (e.g., CANbus to a particular CANbus device), or two-way, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used for each of these networks and network interfaces, as previously described.
[0125] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to core 740 of computer system 700 .
[0126] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 743, hardware accelerators 744 for specific tasks, graphics adapters 750, etc. These devices may be connected via a system bus 748, along with read-only memory (ROM) 745, random access memory 746, and internal mass storage such as an internal non-user-accessible hard drive, SSD, etc. 747. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 748 or via a peripheral bus 749. In one example, a screen 710 may be connected to the graphics adapter 750. Architectures for peripheral buses include PCI, USB, etc.
[0127] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute specific instructions that, in combination, may constitute the aforementioned computer code. That computer code may be stored in ROM 745 or RAM 746. Transient data may also be stored in RAM 746, while persistent data may be stored, for example, in internal mass storage 747. Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc.
[0128] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0129] By way of example and not limitation, a computer system having architecture 700, and specifically core 740, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as described above, as well as media associated with specific storage of core 740 that is non-transitory in nature, such as core internal mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 740. Computer-readable media can include one or more memory devices or chips, depending on particular needs. Software can cause core 740, and specifically the processors in the core (including a CPU, GPU, FPGA, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to software-defined processes. Additionally or alternatively, a computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator 744) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software may encompass logic, and vice versa. Where appropriate, references to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0130] FIG. 8 illustrates a network media distribution system 800 that supports various legacy and heterogeneous immersive media-enabled displays as client endpoints in some examples. In the example of FIG. 8, a content acquisition module 801 captures or creates media using the exemplary embodiments of FIG. 6 or FIG. 5. An ingest format is created in a content preparation module 802 and then transmitted to one or more client endpoints in the network media distribution system using a transmission module 803. A gateway 804 can service customer premises equipment (CPE) to provide network access to various client endpoints in the network. A set-top box 805 can also function as CPE to provide access to aggregated content by a network service provider. A wireless demodulator 806 can function as a mobile network access point for mobile devices (e.g., similar to a mobile handset and display 813). In one or more embodiments, a legacy 2D television 807 may be connected directly to the gateway 804, the set-top box 805, or a WiFi router 808. A computer laptop with a legacy 2D display 809 may be a client endpoint connected to the WiFi router 808. A head-mounted 2D (raster-based) display 810 may also be connected to the router 808. A lenticular light field display 811 may be connected to the gateway 804. The display 811 may consist of a local computation GPU 811A, a storage device 811B, and a visual presentation unit 811C that creates multiple views using ray-based lenticular optics technology. A holographic display 812 may be connected to the set-top box 805 and may include a local computation CPU 812A, a GPU 812B, a storage device 812C, and a Fresnel pattern, wave-based holographic visualization unit 812D.Augmented reality headset 814 can be connected to wireless demodulator 806 and can include a GPU 814A, a storage device 814B, a battery 814C, and a volumetric visual presentation component 814D. High-density light field display 815 can be connected to WiFi router 808 and can include multiple GPUs 815A, CPUs 815B, and a storage device 815C; an eye-tracking device 815D; a camera 815E; and a high-density ray-based light field panel 815F.
[0131] FIG. 9 shows a diagram of an immersive media delivery module 900 capable of serving legacy and heterogeneous immersive media-enabled displays as previously described in FIG. 8. Content is created or acquired in module 901, which is embodied in FIGS. 5 and 6 for natural and CGI content, respectively. The content is then converted to an ingest format using a create network ingest format module 902. Several examples of module 902 are embodied in FIGS. 5 and 6 for natural and CGI content, respectively. The ingest media is optionally updated to store information from a media reuse analyzer 911 about assets that may be reused across multiple scenes. The ingest media format is transmitted to the network and stored in a storage device 903. In some other examples, the storage device may reside within the immersive media content creator's network and be accessed remotely by the immersive media network delivery module 900, as indicated by the bisecting dashed line. Client and application specific information may in some examples be available on a remote storage device 904, which may optionally reside remotely in an alternative cloud network in examples.
[0132] As shown in Figure 9, network orchestrator 905 serves as a primary source and sink of information for performing the primary tasks of the distribution network. In this particular embodiment, network orchestrator 905 may be implemented in a unified manner with other components of the network. Nevertheless, the tasks illustrated by network orchestrator 905 in Figure 9 may, in some instances, form elements of the disclosed subject matter. Network orchestrator 905 may be implemented in software and executed by processing circuitry to perform processing.
[0133] According to some aspects of the present disclosure, the network orchestrator 905 may further employ a two-way message protocol for communicating with client devices to facilitate processing and delivery of media (e.g., immersive media) according to the characteristics of the client devices. Furthermore, the two-way message protocol may be implemented across different delivery channels, i.e., a control plane channel and a data plane channel.
[0134] The network orchestrator 905 receives information about the characteristics and attributes of client devices, such as the client 908 (also referred to as the client device 908) of FIG. 9, and also gathers requirements for applications currently running on the client 908. This information may be obtained from the device 904, or in alternative embodiments, by querying the client 908 directly. In some examples, a two-way message protocol is used to enable direct communication between the network orchestrator 905 and the client 908. For example, the network orchestrator 905 may send queries directly to the client 908. In some examples, the smart client 908E may participate in collecting and reporting client status and feedback on behalf of the client 908. The smart client 908E may be implemented in software that can be executed by processing circuitry to perform processes.
[0135] The network orchestrator 905 also invokes and communicates with the media adaptation and fragmentation module 910 described in FIG. 10 . Once the ingested media has been adapted and fragmented by module 910, the media is transferred to an inter-media storage device, shown in some examples as prepared media for distribution storage device 909. Once the distribution media is prepared and stored on device 909, the network orchestrator 905 ensures that the immersive client 908 receives the distribution media and corresponding description information 906 via its network interface 908B via a push request, or the client 908 itself can initiate a pull request for the media 906 from the storage device 909. In some examples, the network orchestrator 905 may use a two-way message interface (not shown in FIG. 9 ) to make “push” requests or initiate “pull” requests by the immersive client 908. In one example, the immersive client 908 may use the network interface 908B, a GPU (or CPU, not shown) 908C, and storage 908D. Additionally, the immersive client 908 may utilize a game engine 908A. The game engine 908A may further utilize a visualization component 908A1 and a physics engine 908A2. The game engine 908A communicates with a smart client 908E to orchestrate the processing of the media via a game engine API and callback functions 908F. The delivery format of the media is stored in a storage device or storage cache 908D of the client 908. Finally, the client 908 visually presents the media via a visualization component 908A1.
[0136] Throughout the process of streaming immersive media to an immersive client 908, the network orchestrator 905 can monitor the client's progress via a client progress and status feedback channel 907. Status monitoring can be done by a two-way communication message interface (not shown in FIG. 9) that can be implemented in a smart client 908E.
[0137] FIG. 10 shows a diagram of a media adaptation process 1000 so that, in some examples, ingested source media can be appropriately adapted to fit the requirements of an immersive client 908. The media adaptation and fragmentation module 1001 is composed of multiple components that facilitate the adaptation of ingested media to an appropriate delivery format for the immersive client 908. In FIG. 10, the media adaptation and fragmentation module 1001 receives input network conditions 1005 to track the current traffic load on the network. The immersive client 908 information can include attributes and feature descriptions, application characteristics and descriptions, and the current state of the application, as well as a client neural network model (if available) that helps map the client's frustum geometry to interpolation functions for the ingested immersive media. Such information can be obtained by a two-way message interface (not shown in FIG. 10) using a smart client interface, shown as 908E in FIG. 9. The media adaptation and fragmentation module 1001 ensures that, once the adapted output is created, it is stored in the client adaptation media storage device 1006. The media reuse analyzer 1007 is shown in FIG. 10 as a process that can be performed in advance or as part of a network automated process for the distribution of media.
[0138] In some examples, the media adaptation and fragmentation module 1001 is controlled by a logic controller 1001F. In one example, the media adaptation and fragmentation module 1001 uses a renderer 1001B or a neural network processor 1001C to adapt the particular ingest source media to a format suitable for the client. In one example, the media adaptation and fragmentation module 1001 receives client information 1004 from a client interface module 1003, such as a server device in one example. The client information 1004 may include a client description and current state, an application description and current state, and a client neural network model. The neural network processor 1001C uses the neural network model 1001A. Examples of such neural network processors 1001C include deep view neural network model generators such as those described in MPI and MSI. In some examples, the media is in a 2D format but the client requires a 3D format, and the neural network processor 1001C can then invoke a process that uses highly correlated images from the 2D video signal to derive a stereoscopic representation of the scene depicted in the video. An example of such a process could be the neural radiation field from one or a few images process developed at the University of California, Berkeley. An example of a suitable renderer 1001B could be a modified version of the OTOY Octane renderer (not shown) that is modified to interact directly with the media adaptation and fragmentation module 1001. The media adaptation and fragmentation module 1001 can, in some examples, use a media compressor 1001D and a media decompressor 1001E, depending on the needs of these tools regarding the format of the ingested media and the format required by the immersive client 908.
[0139] Figure 11 illustrates a delivery format creation process 1100 in some examples. An adaptive media packaging module 1103 packages media from a media adaptation module 1101 (shown as process 1000 in Figure 10) currently residing on a client adaptive media storage device 1102. The packaging module 1103 formats the adapted media from the media adaptation module 1101 into a robust delivery format 1104, such as the exemplary formats shown in Figures 3 or 4. Manifest information 1104A provides the client 908 with a list 1104B of scene data assets it can expect to receive, along with optional metadata describing the frequency with which each asset is used across the set of scenes that make up the presentation. List 1104B illustrates a list of visual, audio, and haptic assets, each with corresponding metadata. In this exemplary embodiment, each asset in list 1104B references metadata that includes a numeric frequency value indicating the number of times a particular asset is used across all scenes that make up the presentation.
[0140] Figure 12 illustrates some example packetizer processing system 1200. In the example of Figure 12, packetizer 1202 separates adapted media 1201 into individual packets 1203 suitable for streaming to immersive client 908, shown as client endpoint 1204 on a network.
[0141] FIG. 13 illustrates a network sequence diagram 1300 that, in some examples, adapts particular immersive media in an ingest format into a streamable and appropriate delivery format for a particular immersive media client endpoint.
[0142] The components and communications shown in Figure 13 are described as follows: Client 1301 (also referred to in some examples as a client endpoint or client device) initiates a media request 1308 to network orchestrator 1302 (also referred to in some examples as a network delivery interface). Media request 1308 includes information identifying the media requested by client 1301, either by URN or other standard nomenclature. Network orchestrator 1302 responds to media request 1308 with a profile request 1309, which requests that client 1301 provide information about its currently available resources (including computation, storage, battery charge, and other information to characterize the client's current operating state). Profile request 1309 also requests that the client provide one or more neural network models, if such models are available at the client, that can be used by the network for neural network inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system. Response 1310 from client 1301 to network orchestrator 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available to the client). Network orchestrator 1302 then provides client 1301 with a session ID token 1311. Network orchestrator 1302 then requests ingest media server 1303 with an ingest media request 1312 that includes the URN or canonical name of the media identified in request 1308. Ingest media server 1303 responds to request 1312 with response 1313 that includes the ingest media tokens. Network orchestrator 1302 then provides the media tokens from response 1313 to client 1301 in call 1314.The network orchestrator 1302 then initiates the adaptation process for the media request 1308 by providing the adaptation ingest 1304 with the ingest media token, client token, application token, and neural network model token 1315. The adaptation ingest 1304 requests access to the ingest media by providing the ingest media token 1316 to the ingest media server 1303 in the call requesting access to the ingest media asset. The ingest media server 1303 responds to the request 1316 with the ingest media access token in a response 1317 to the adaptation ingest 1304. The adaptation ingest 1304 then requests that the media adaptation module 1305 adapt the ingest media located in the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created in 1313. A request 1318 from the adaptation ingest 1304 to the media adaptation module 1305 includes the necessary tokens and session ID. Media adaptation module 1305 provides the adapted media access token and session ID to network orchestrator 1302 in update 1319. Network orchestrator 1302 provides the adapted media access token and session ID to packaging module 1306 in interface call 1320. Packaging module 1306 provides response 1321 to network orchestrator 1302 with the packaged media access token and session ID in response message 1321. Packaging module 1306 provides the packaged asset, URN, and packaged media access token for the session ID to packaged media server 1307 in response 1322.Client 1301 executes request 1323 to begin streaming the media asset corresponding to the packaged media access token received in response message 1321. Client 1301 executes other requests and provides status updates in message 1324 to network orchestrator 1302.
[0143] FIG. 14 shows a diagram of a media system 1400 with a virtual network and client devices 1418 (also referred to as game engine client devices 1418) for scene-based media processing in some examples. A smart client, such as that represented by MPEG Smart Client Process 1401 in FIG. 4, can act as a central coordinator for preparing media to be processed by and for other entities within the game engine client device 1418, as well as for entities external to the game engine client device 1418. In some examples, a smart client is implemented as software instructions that can be executed by processing circuitry to perform a process such as MPEG Smart Client Process 1401 (also referred to as MPEG Smart Client 1401 in some examples). The game engine 1405 is primarily responsible for rendering media to create a presentation experienced by an end user. A haptic component 1413, a visualization component 1415, and an audio component 1414 can assist the game engine 1405 in rendering haptic, visual, and audio media, respectively. Edge processor or network orchestrator device 1408 communicates information and system media to MPEG Smart Client 1401 and can receive state updates and other information from MPEG Smart Client 1401 via network interface protocol 1420. Network interface protocol 1420 can be divided across multiple communication channels and processes and can use multiple communication protocols. In some examples, game engine 1405 is a game engine device that includes control logic 14051, GPU interface 14052, physics engine 14053, renderer 14054, compression decoder 14055, and device-specific plug-ins 14056. MPEG Smart Client 1401 also serves as the primary interface between the network and client device 1418.For example, MPEG Smart Client 1401 can interact with Game Engine 1405 using Game Engine API and callback functions 1417. In one example, MPEG Smart Client 1401 can be responsible for reconstructing the streaming media communicated at 1420 before invoking Game Engine API and callback functions 1417 managed by Game Engine Control Logic 14051 to have Game Engine 1405 process the reconstructed media. In such an example, MPEG Smart Client 1401 can utilize Client Media Reconstruction Process 1402, which can utilize Compression Decoder Process 1406.
[0144] In some other examples, MPEG Smart Client 1401 may not be responsible for reassembling the packetized media streamed at 1420 before invoking API and callback functions 1417. In such examples, game engine 1405 may decompress and reassemble the media. Further, in such examples, game engine 1405 may decompress the media using compression decoder 14055. Upon receiving the reassembled media, game engine control logic 14051 may render the media via renderer process 14054 using GPU interface 14052.
[0145] In some examples, the rendered media is animated, and then the physics engine 14053 may be used by the game engine control logic 14051 to simulate the laws of physics in the animation of the scene.
[0146] In some examples, throughout the processing of media by client device 1418, neural network model 1421 may be employed by neural network processor 1403 to assist in operations coordinated by MPEG Smart Client 1401. In some examples, reconstruction process 1402 may require fully reconstructing the media using neural network model 1421 and neural network processor 1403. Similarly, client device 1418 may be configured by a user via user interface 1412 to cache media received from the network in client adaptive media cache 1404 after the media has been reconstructed, or to cache rendered media in rendered client media cache 1407 after the media has been rendered. Furthermore, in some examples, MPEG Smart Client 1401 may replace system-provided visual / non-visual assets with user-provided visual / non-visual assets from user-provided media cache 1416. In such an embodiment, User Interface 1412 may guide an end user through the steps of loading user-provided visual / non-visual assets from User Provided Media Cache 1419 (e.g., external to Client Device 1418) into Client-Accessible User Provided Media Cache 1416 (e.g., internal to Client Device 1418). In some embodiments, MPEG Smart Client 1401 may be configured to store rendered assets in Rendered Media Cache 1411 (for potential reuse or sharing with other clients).
[0147] In some examples, media analyzer 1410 can examine client adaptive media 1409 (in the network) to determine asset complexity or the frequency with which assets are reused across one or more scenes (not shown) for potential prioritization for rendering by game engine 1405 and / or reconstruction processing via MPEG Smart Client 1401. In such examples, media analyzer 1410 stores complexity, prioritization, and asset usage frequency information in the media stored in 1409.
[0148] It should be noted that although processes are shown and described in this disclosure, the processes may be implemented as instructions in a software module, and the instructions may be executed by a processing circuit to perform the process. It should also be noted that although modules are shown and described in this disclosure, the modules may be implemented as software modules having instructions, and the instructions may be executed by a processing circuit to perform the process.
[0149] According to a first aspect of the present disclosure, various techniques can be used to implement a Smart Client on a client device equipped with a game engine to provide scene-based immersive media to the game engine. The Smart Client can be embodied by one or more processes, with one or more communication channels implemented between the client device process and a network server process. In some examples, the Smart Client is configured to receive and communicate media and media-related information to facilitate processing of scene-based media on a particular client device between the network server process and the client device's game engine, which in some examples can function as a media rendering engine within the client device. The media and media-related information can include metadata, command data, client state information, media assets, and information that facilitates optimization of one or more operational aspects within the network. Similarly, in the case of a client device, the Smart Client can use the availability of an application programming interface provided by the game engine to efficiently enable playback of scene-based immersive media streamed from the network. In a network architecture intended to support a heterogeneous collection of immersive media processing devices, the Smart Client described herein is a component of the network architecture that provides the main interface through which the network server process interacts with and delivers scene-based media to a particular client device. Similarly, for client devices, the Smart Client utilizes the programming architecture used by the client device's game engine to effect efficient management and delivery of scene-based assets for rendering by the game engine.
[0150] Figure 15 is a diagram of a media flow process 1530 (also referred to as process 1530) for media delivery to a client device over a network in some examples. Process 1530 is similar to process 100 shown in Figure 1, but with explicit processes for media delivery shown in the "cloud" or "edge." Additionally, the client device in Figure 15 explicitly illustrates the use of a game engine process as one of the processes executing on behalf of the client device. A smart client process is also shown as one of the processes executing on behalf of the client device.
[0151] Process 1530 is similar to process 100 of FIG. 1, except that an explicit media streaming process 1536 is shown, and the logic of smart client process 1539 is shown as part of client device 15312. Smart client process 1539 serves as the primary interface through which the media distribution network communicates with client device 15312. In some examples, client device 15312 uses a game engine to render media, and media presentation format C 15311 is scene-based rather than frame-based media. Thus, such smart client process can play an important role in communicating to the network that frequently reused assets from scene-based media have already been received by client device 15312, eliminating the need for the network to stream them again to client device 15312. Process 1530 includes a series of steps, beginning with process (step) 1531. In process 1531, media is ingested into the network by a content provider (not shown). Next, delivery media process (step) 1532 converts the media to delivery format B. Next, a process (step) 1533 converts the media in delivery format B into another format 1535 that can be packetized and streamed over the network. A media streaming process 1536 then delivers the media 1537 over the network to the game engine client device 15312. A smart client process 1539 receives the media on behalf of the client device 15312 and returns progress, status, and other information to the network over an information channel 1538. The smart client process 1539 coordinates the processing of the streamed media 1537 with components of the client device 15312, which may include converting the media to yet another format, presentation format C 15311, via a rendering process 15310.
[0152] Figure 16 shows a diagram of a media flow process 1640 (also referred to as process 1640) involving asset reuse into a game engine process in some examples. Process 1640 is similar to process 1530 shown in Figure 15, but query logic is explicitly shown to determine if the client device already has access to the asset in question, so the network does not have to stream the asset again. In Figure 16, the client device is shown with a smart client and game engine present as part of the client device.
[0153] Process 1640 is similar to process 1530 shown in FIG. 15 , except that query logic is explicitly shown to determine whether a client device (e.g., having a game engine) 16412 has already accessed the targeted media asset, and therefore the network does not need to stream the asset again. Like process 1530 shown in FIG. 15 , process 1640 shown in FIG. 16 includes a series of steps beginning with process (step) 1641. In process 1641, media is ingested into the network by a content provider (not shown). Process (step) 1642 determines whether the media asset in question has previously been streamed to the client device and therefore does not need to be streamed again. Some of the relevant steps of process 1642 are explicitly shown to illustrate the logic for determining whether the media asset in question has previously been streamed to the client device and therefore does not need to be streamed again. For example, a distribution media creation initiation process 1642A initiates the logical sequence of process 1642. Decision-making process 1642B determines whether media needs to be streamed to client device 16412, depending on the information contained in smart client feedback and status 1648. If media needs to be streamed to client device 16412, process 1642D continues preparing the media to be streamed. If media should not be streamed to client device 16412, process 1642C indicates that the media asset will not be included in streamed media 1647 (and therefore can be obtained from another source, e.g., a resource cache available to the client device). Process 1642E marks the end of process 1642. Process 1643 then converts the media in delivery format B into another format 1645 that can be packetized and streamed over a network.A media streaming process 1646 then delivers the media 1647 over the network to the client device 16412 (which has the game engine). A smart client process 1649 receives the media on behalf of the client device 16412 and returns progress, status, and other information to the network over an information channel 1648. The smart client process 1649 coordinates the processing of the streamed media 1647 with components of the client device 16412, which may include converting the media to yet another format, presentation format C 16411, via a rendering process 16410.
[0154] Figure 17 shows a diagram of a media conversion decision-making process 1730 (also referred to as process 1730) with asset reuse logic and redundant caching for a client device (e.g., a game engine client device). Process 1730 is similar to process 200 of Figure 2, with the addition of logic to determine whether a proxy to the original media needs to be streamed instead of the original media itself (converted to another format or its original format).
[0155] In FIG. 17, the flow of media through a network uses two decision-making processes to determine whether the network needs to convert the media before delivering it to a client device. In FIG. 17, in process step 1731, ingested media, represented by format A, is provided to the network by a content provider (not shown). Process step 1732 obtains attributes describing the processing capabilities of the client device (not shown). Decision-making process step 1733 is used to determine whether the network has previously streamed a particular media object (also referred to as a media asset) to the client device. If the media object has previously been streamed to the client device, process step 1734 uses a proxy (e.g., an identifier) in place of the media object to indicate that the client device can access a copy of the previously streamed media object from its local cache or other cache. If the media object has not previously been streamed, decision-making process step 1735 is used to determine whether the network or the client device needs to perform a format conversion of any of the media assets contained within the ingested media in process step 1731, such as converting a particular media object from format A to format B, before the media object is streamed to the client device. If any of the media assets need to be converted by the network, the network uses process step 1738 to convert the media assets from format A to format B. Converted media 1739 is output from process step 1738. The converted media 1739 is merged with preparation process step 1736 to prepare the media to be streamed to a client device (not shown). Process 1737 streams the media to the client device.
[0156] In some examples, the client device includes a smart client. The smart client can perform process step 17310, which determines whether the received media asset is being streamed for the first time and will be reused. If the received media asset is being streamed for the first time and will be reused, the smart client can perform process step 17311, which creates a copy of the reusable media asset in a cache (also called a redundant cache) accessible by the client device. The cache can be an internal cache of the client device or an external cache device from the client device.
[0157] FIG. 18 is a diagram of asset reuse logic using a smart client for game engine client process 1840 (also referred to as process 1840). Process 1840 is similar to process 1730 shown in FIG. 17, except that asset query logic process 1733 shown in FIG. 17 is implemented by asset query logic 1843 shown in FIG. 18, which illustrates the role of the smart client on a client device in determining whether a particular asset is already accessible to the client device. Media is ingested onto the network in process step 1841. Process step 1842 obtains attributes describing the processing capabilities of the target client (not shown). The network then initiates asset query logic 1843 to determine whether the media asset needs to be streamed to the client device. Smart client process step 1843A receives information from the network (e.g., a network server process) regarding media assets needed for the presentation. The smart client accesses information from a database (e.g., a local cache, network storage accessible by the client device) 1843B to determine whether the media asset in question is already accessible to the client device. The smart client process 2043C returns information to the network server's asset query logic 1843 regarding whether the media asset in question needs to be streamed. If the asset does not need to be streamed, process step 1844 creates a proxy for the media asset in question and inserts the proxy into the media that is prepared for streaming to the client device in process step 1846. If the media asset needs to be streamed to the client device, network server process step 1845 determines whether the network needs to support any conversion of the media being streamed to the client device. If such conversion is required and the conversion needs to be performed by available resources of the network, network process 1848 performs the conversion.The converted media 1849 is then provided to process step 1846 for merging the converted media with the media to be streamed to the client device. If a determination is made in network process step 1845 that the network will not perform such conversion of the media, process step 1846 prepares the media with the original media assets to be streamed to the client device. Following process step 1846, the media is streamed to the client device in process step 1847.
[0158] FIG. 19 shows a diagram of a process 1950 for obtaining device state and profile information (e.g., using a game engine) using a smart client 1952A of a client device 1952B. Process 1950 is similar to process 1840 shown in FIG. 18, but is implemented by process 1952 for obtaining client device attributes from a smart client, to illustrate the role of the smart client in obtaining information about the capabilities of the client device, including the availability of resources for rendering media. Media is ingested onto the network in the process and requested by a game engine client device (not shown) in process step 1951. Process 1952 initiates a request to the smart client to obtain the device attributes and resource state of the client device. For example, smart client 1952A initiates a query to client device 1952B and its additional resources (if any) 1952C to retrieve descriptive attributes about the client device and information about those resources, including resource availability for processing future work, respectively. The client device 1952B delivers client device attributes describing the processing capabilities of the client device (e.g., the target client device for the ingested media). The status and availability of resources on the client device are returned to the smart client 1952A from additional resources 1952C. The network then initiates asset query logic 1953 to determine whether the asset (media asset) needs to be streamed to the client device. If the asset does not need to be streamed, process step 1954 creates a proxy for the asset in question and inserts the proxy into the media in preparation for streaming to the client device in process step 1956.If an asset needs to be streamed to a client device, network server process step 1955 determines whether the network needs to support any conversion of the media being streamed to the client, if such conversion is necessary. If such conversion is necessary and the conversion needs to be performed by available resources of the network, network process step 1958 performs the conversion. The converted media 1959 is then provided to process step 1956 for merging the converted media with the media to be streamed to the client device. If a determination is made in network process step 1955 that the network does not need to perform such media conversion, process step 1956 prepares media with the original media asset to be streamed to the client device. Following process step 1956, the media is streamed to the client device in process step 1957.
[0159] FIG. 20 shows a diagram of a process 2060 to illustrate a smart client on a client device requesting and receiving streamed media from a network on behalf of the client device's game engine. Process 2060 is similar to process 1950 shown in FIG. 19 or process 1730 of FIG. 17, with the explicit addition of the requesting and receiving of media being managed by smart client 20610 instead of client device 20611. In FIG. 20, a user requests specific media in process step 20612. The request is received by client device 20611. Client device 20611 forwards the request for media to smart client 20610. Smart client 20610 forwards the media request to the server in process step 2061. The server in process step 2061 receives the media request. In process step 2062, the server initiates steps to obtain device attributes and resource state (the exact steps for obtaining the attributes and resource state are not shown). The status and availability of resources on the client device are returned to the server in process step 2062. The network (e.g., server) then initiates asset query logic 2063 to determine whether the asset (media asset) needs to be streamed to the client device. If the asset does not need to be streamed, process step 2064 creates a proxy for the asset in question and inserts the proxy into the media in preparation for streaming to the client device in process step 2066. If the asset needs to be streamed to the client device, the network server determines in process step 2065 whether the network needs to support any conversion of the media being streamed to the client device, if such conversion is required. If such conversion is required and the conversion needs to be performed by available resources of the network, the network (e.g., network server) performs the conversion in process step 2068.The converted media 2069 is then provided to process step 2066 for merging the converted media with the media to be streamed to the client device. If a determination is made in network process 2065 that the network will not perform such media conversion, process step 2066 prepares the media with the original media assets to be streamed to the client device. Following process step 2066, the media is streamed to the client device in process 2067. In some examples, the media arrives at a media store 20613 accessible to the client device 20611. The client device 20611 accesses the media from the media store 20613. The media store 20613 may be within the client device 20611 in one example. In another example, the media store 20613 is external to the client device 20611, such as in a device within the same local network as the client device 20611, and the client device 20611 can access the media store 20613.
[0160] FIG. 21 illustrates a diagram 2130 of a timed media display ordered by descending frequency in some examples. The timed media display is similar to the timed media display illustrated in FIG. 3, except that the assets in the timed media display of FIG. 21 are ordered in a list by asset type and descending frequency values within each asset type. Specifically, timed scene manifest 2103A includes a list of scene information 2131. Scene information 2131 points to a list of components 2132 that separately describe the processing information and types of media assets in scene information 2131. Components 2132 point to assets 2133, which further point to base layers 2134 and attribute enhancement layers 2135. In the example of FIG. 21, each of the base layers 2134 is ordered according to the descending value of the corresponding frequency metric. A list of unique assets not previously used in other scenes is provided at 2137. Proxy visual assets 2136 contain information about reused visual assets, and proxy audio assets 2138 contain information about reused audio assets.
[0161] In some examples, ordering by descending frequency can allow client devices to process media assets with high frequency reuse first to reduce delay.
[0162] FIG. 22 illustrates a diagram 2240 of a non-timed media presentation ordered by frequency in some examples. The non-timed media presentation is similar to the non-timed media presentation shown in FIG. 4, except that the assets in the non-timed media presentation of FIG. 22 are ordered in the list by asset type and frequency value within each asset type. A non-timed scene manifest (not shown) references Scene 1.0, which has no other scenes that can branch into it. Scene information 2241 for Scene 1.0 is not associated with a start or end time according to a clock. Scene information 2241 also points to a list of components 2242 that separately describe the processing information and types of media assets within Scene information 2241. Components 2242 point to assets 2243, which further point to base layer 2244 and attribute enhancement layers 2245 and 2246. In the example of FIG. 22, each of the base layers 2244 points to a numeric frequency value that indicates the number of times the asset is used across the set of scenes in the presentation. For example, haptic assets 2243 are organized by increasing frequency value, and audio assets 2243 are organized by decreasing frequency value. Scene information 2241 also references other scene information 2241 for non-timed media. Scene information 2241 also references scene information 2247 for timed media scenes. List 2248 identifies unique assets associated with a particular scene that have not previously been used in a higher-level (e.g., parent) scene.
[0163] It should be noted that the frequency ordering in this disclosure is for illustrative purposes only, and assets may be ordered by any suitable increasing or decreasing frequency value to optimize media delivery. In some examples, the order of assets is determined according to the processing capabilities, resources, and optimization strategies of the client device, which may be provided to the network in a feedback signal from the client device to the network.
[0164] Figure 23 illustrates a process 2340 for creating a distribution format with ordered frequency. Process 2340 is similar to process 1100 of Figure 11, which illustrates network formatting the same adapted source media into a data model suitable for display and streaming. However, the resulting distribution format illustrated in process 2340 illustrates that the assets are ordered first by asset type and then by how frequently the assets are used throughout the presentation, for example, in either ascending or descending order of frequency value based on asset type.
[0165] In Figure 23, adaptive media packaging process 2343 packages media from media adaptation process 2341 (shown as process 1000 in Figure 10) that is now located on storage device 2342, which is configured to store client-adapted media. Packaging process 2343 formats the adapted media from process 2341 into robust delivery format 2344, such as the exemplary format shown in Figure 3 or Figure 4. Manifest information 2344A provides the client device with a list of scene data assets 2344B that it can expect to receive, as well as optional metadata indicating how frequently all assets are used across the set of scenes in the entire presentation. List 2344B shows a list of visual, audio, and haptic assets, each with corresponding metadata. In the example of Figure 23, packaging process 2343 orders the assets in 2344B, i.e., the visual assets in 2344B are ordered by decreasing frequency, while the audio and haptic assets in 2344B are ordered by increasing frequency.
[0166] FIG. 24 shows a logic flow (process) flowchart 2400 for a media reuse analyzer, such as the media reuse analyzer 911 depicted in FIG. 9. The process begins in step 2401 by optimizing a presentation for asset reuse across scenes. Step 2402 initializes an iterator “i” to 0 and further initializes a set of lists 2404 (one list per scene) that identify unique assets encountered across all scenes in the presentation, as shown in FIG. 3 or FIG. 4. List 2404 shows a sample list entry of information describing a unique asset for the entire presentation, including the asset's media type (e.g., mesh, audio, or volume), the asset's unique identifier, and an indicator of the number of times the asset was used across the set of scenes in the presentation. As an example, for scene N-1, no assets are included in its list because all assets required for scene N-1 have been identified as assets also used in scene 1 and scene 2. Step 2403 determines whether the iterator “i” is less than the total number of scenes in the presentation (as shown in FIG. 3 or FIG. 4). If iterator "i" is equal to the number of scenes N in the presentation, the reuse analysis ends at step 2405. Otherwise, if iterator "i" is less than the total number of scenes, the process proceeds to step 2406, where iterator "j" is set to 0. Step 2407 tests iterator "j" to determine if iterator "j" is less than the total number of media assets (also called media objects) in the current scene "i". If iterator "j" is less than the total number of media assets in scene "i", the process proceeds to step 2408. Otherwise, the process proceeds to step 2412, where iterator "i" is incremented by 1 before returning to step 2403. If the value of "j" is less than the total number of assets in scene "i", the process continues to conditional step 2408, where the characteristics of the media asset are compared to previously analyzed assets from scenes previous to the current scene "i".If the asset is identified as an asset used in a scene prior to scene "i", then in step 2411, the number of times (e.g., frequency) the asset has been used across scenes 0 through N-1 is incremented by 1. Otherwise, if the asset is a unique asset, i.e., has not been previously analyzed in a scene associated with a lower value of iterator "i", then in step 2409, a unique asset entry is created in list 2404 for scene "i". Step 2409 also creates and assigns a unique identifier to the asset's entry and sets the number of times the asset has been used across scenes 0 through N-1 to 1. Following step 2409, the process proceeds to step 2410, where iterator "j" is incremented by 1. After step 2410, the process returns to step 2407.
[0167] Figure 25 shows a diagram 2500 of network processes communicating with a smart client and game engine in a client device. A network orchestration process 2501 communicates information and media to, and receives status updates and other information from, a network interface process 2502. Similarly, the network interface process 2502 communicates information and media to, and receives status updates and other information from, a client device 2504. Client device 2504 illustrates a specific embodiment of a game engine client device. A smart client 2504A serves as the primary interface between the network and other media-related processes and / or components within client device 2504. For example, smart client 2504A can use a game engine API and callback functions (not shown) to interact with a game engine 2504G. In some examples, smart client 2504A may be responsible for reconstructing the streaming media delivered at 2503 before invoking a game engine API (not shown) managed by game engine control logic 2504G1 to have game engine 2504G process the reconstructed media. In such examples, smart client 2504A may utilize client media reconstruction process 2504H, which may utilize compression decoder process 2504B. In some other examples, smart client 2504A may be solely responsible for reconstructing the packetized media streamed at 2503 before invoking a game engine API (not shown) to have game engine 2504G decompress and process the media. In such examples, game engine 2504G may decompress the media using compression decoder 2504G6.Upon receiving the reconstructed media and optionally decompressing the media (e.g., the media may have been decompressed on behalf of the client device via smart client 2504A), game engine control logic 2504G1 can render the media via renderer process 2504G3 using GPU interface 2504G5. If the rendered media is animated, physics engine 2504G2 may be used by game engine control logic 2504G1 to simulate the laws of physics in the animation of the scene.
[0168] Throughout the processing of media by the client device 2504, the neural network model 2504E can be used to guide the actions taken by the client device. For example, in some cases, the reconstruction process 2504H may need to fully reconstruct the media using the neural network model 2504E and the neural network processor 2504F. Similarly, the client device 2504 can be configured via the client device control logic 2504J to cache media received from the network after it has been reconstructed or to cache media after it has been rendered. In such an embodiment, the client adaptive media cache 2504D may be utilized to store reconstructed client media, and the rendered client media cache 2504I may be utilized to store rendered client media. Additionally, the client device control logic 2504J can be responsible for completing the presentation of media on behalf of the client device 2504. In such an embodiment, the visual component 2504C can be responsible for creating the final visual presentation by the client device 2504.
[0169] 26 shows a flowchart outlining a process 2600 according to one embodiment of the present disclosure. Process 2600 may be performed in an electronic device, such as a client device having a smart client for interfacing the client device with a network. In some embodiments, process 2600 is implemented in software instructions, such that processing circuitry performs process 2600 when the processing circuitry executes the software instructions. For example, a smart client may be implemented in software instructions, and the software instructions may be executed by the processing circuitry to perform a smart client process, which may include process 2600. The process starts at S2601 and proceeds to S2610.
[0170] At S2610, a client interface (eg, a smart client) of the electronic device sends capability and availability information for playing scene-based immersive media on the electronic device to a server device in the network.
[0171] At S2620, the client interface receives a media stream carrying adapted media content for the scene-based immersive media, the adapted media content being generated from the scene-based immersive media by the server device based on capability and availability information.
[0172] At S2630, the scene-based immersive media is played on the electronic device according to the adapted media content.
[0173] In some examples, the client interface determines from the adapted media content that a first media asset associated with a first scene is received for the first time and is to be reused in one or more scenes. The client interface can store the first media asset in a cache device accessible by the electronic device.
[0174] In some examples, the client interface can extract a first list of unique assets in a first scene from the media stream, the first list identifying the first media asset as a unique asset in the first scene to be reused in one or more other scenes.
[0175] In some examples, the client interface can send a signal to the server device, the signal indicating availability of the first media asset at the electronic device, the signal causing the server device to use a proxy in place of the first media asset in the adapted media content.
[0176] In some examples, to play the scene-based immersive media, the client interface determines, according to a proxy in the adapted media content, that a first media asset has previously been stored in a cache device, and can access the cache device to retrieve the first media asset.
[0177] In one example, the client interface receives a query signal for a first media asset from the server device and, in response to the query signal, transmits a signal indicating the availability of the first media asset at the electronic device when the first media asset is stored in the cache device.
[0178] In some examples, the client interface receives a request from the server device to obtain device attributes and resource status. The client interface queries one or more internal components of the electronic device and / or one or more external components associated with the electronic device about the attributes of the electronic device and resource availability for processing scene-based immersive media. The client interface can transmit information received from the internal components of the electronic device and the external components associated with the electronic device, such as the attributes of the electronic device and resource availability for the electronic device, to the server device.
[0179] In some examples, the electronic device receives a request for scene-based immersive media from a user interface, and the client interface forwards the request for the scene-based immersive media to a server device.
[0180] In some examples, to play the scene-based immersive media, reconstructed scene-based immersive media is generated based on decoding of the media streams and media reconstruction under control of the client interface, and the reconstructed scene-based immersive media is then provided to a game engine of the electronic device for playback via an application programming interface (API) of the game engine.
[0181] In some other examples, the client interface depacketizes the media stream to generate depacketized media data. The depacketized media data is provided to a game engine of the electronic device via an application programming interface (API) of the game engine. The game engine then generates reconstructed scene-based immersive media for playback based on the depacketized media data.
[0182] The process 2600 then proceeds to S2699 and ends.
[0183] Process 2600 may be appropriately adapted to various scenarios, and steps within process 2600 may be adjusted accordingly. One or more of the steps within process 2600 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 2600. Additional steps may be added.
[0184] 27 shows a flowchart outlining a process 2700 according to one embodiment of the present disclosure. The process 2700 may be performed within a network, such as on a server device within the network. In some embodiments, the process 2700 is implemented in software instructions, such that the processing circuitry performs the process 2700 when it executes the software instructions. The process starts at S2701 and proceeds to S2710.
[0185] At S2710, the server device receives, from the client interface of the electronic device, capability and availability information for playing scene-based immersive media on the electronic device.
[0186] At S2720, the server device generates adapted media content of the scene-based immersive media for the electronic device based on the capability and availability information on the electronic device.
[0187] At S2730, the server device transmits a media stream carrying the adapted media content to the electronic device (eg, a client interface of the electronic device).
[0188] In some examples, the server device determines that a first media asset in the first scene has previously been streamed to the electronic device and replaces the first media asset in the first scene with a proxy that represents the first media asset.
[0189] In some examples, the server device extracts a list of unique assets for each scene.
[0190] In some examples, the server device receives a signal indicating availability of the first media asset at the electronic device. In one example, the signal is transmitted from a client interface of the client device. The server device then replaces the first media asset in the first scene with a proxy representing the first media asset.
[0191] In some examples, the server device transmits a query signal for the first media asset to the client device and receives a signal in response to the query signal indicating availability of the first media asset at the electronic device.
[0192] In some examples, the server device sends a request to obtain device attributes and resource status, and then receives the attributes and capability and availability information of the electronic device.
[0193] The process 2700 then proceeds to S2799 and ends.
[0194] Process 2700 may be appropriately adapted to various scenarios, and steps within process 2700 may be adjusted accordingly. One or more of the steps within process 2700 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 2700. Additional steps may be added.
[0195] According to a second aspect of the present disclosure, various techniques disclosed in this disclosure are used for streaming scene-based immersive media, enabling the presentation of a more personalized media experience to end users by substituting user-provided visual assets for content creator-provided visual assets. In some examples, a smart client in a client device is implemented with several technologies, the smart client can be embodied by one or more processes, and one or more communication channels are implemented between the client device process and a network server process. In some examples, metadata in the immersive media stream can signal the availability of visual assets suitable for exchange with user-provided visual assets. The client device can access a repository of user-provided assets, each of which is annotated with metadata to assist in the substitution process. The client device can further use an end-user interface to enable the loading of user-provided visual assets into an accessible cache (also referred to as a user-provided media cache) for subsequent substitution into a scene-based media presentation streamed from a server device in the network.
[0196] According to a third aspect of the present disclosure, various techniques disclosed in this disclosure are used for streaming scene-based immersive media to enable the presentation of a more personalized media experience to end users by substituting user-provided non-visual assets (e.g., audio, somatosensory, olfactory) for content creator-provided (also known as "system") non-visual assets. In some examples, a smart client in a client device is implemented with several technologies, the smart client can be embodied by one or more processes, and one or more communication channels are implemented between the client device process and the network server process. In some examples, metadata in the immersive media stream is used to signal the availability of non-visual assets suitable for exchange with the user-provided non-visual assets. In some examples, the smart client can access a repository of user-provided assets, each of which is annotated with metadata to assist in the substitution process. The client device can further use an end-user interface to enable loading of user-provided non-visual assets into a client-accessible user-provided media cache for subsequent substitution into the scene-based media presentation streamed from the network server.
[0197] Figure 28 shows a diagram of a timed media presentation 2810 with signaling of visual asset replacement in some examples. The timed media presentation 2810 includes a timed scene manifest 2810A containing information for a list of scenes 2811. The scenes 2811 point to a list of components 2812 that separately describe the types of media assets and processing information within the scene 2811. The components 2812 point to assets 2813, which further point to a base layer 2814 and an attribute enhancement layer 2815. The visual asset base layer 2814 is provided with metadata to indicate whether the corresponding asset is a candidate for replacement with a user-provided asset (not shown). A list 2817 of assets that have not previously been used in other scenes 2811 (e.g., previous scenes) is provided for each scene 2811.
[0198] In the example of Figure 28, the base layer of visual asset 1 indicates that visual asset 1 is replaceable, the base layer of visual asset 2 indicates that visual asset 2 is replaceable, and the base layer of visual asset 3 indicates that visual asset 3 is not replaceable.
[0199] FIG. 29 shows a diagram of a non-timed media presentation 2910 with signaling of visual asset replacement in some examples. In some examples, a non-timed scene manifest (not shown) references scene 1.0, which has no other scenes that can branch to it. The scene information 2911 in FIG. 29 does not have an associated start time and end time according to a clock. The scene information 2911 points to a list of components 2912 that separately describe the processing information and type of media asset for the scene. The components 2912 point to assets 2913, which further point to a base layer 2914 and attribute enrichment layers 2915 and 2916. Assets 2913 of type "visual" are further provisioned with metadata to indicate whether the corresponding asset is a candidate for replacement with a user-provided asset (not shown). Additionally, the scene information 2911 can reference other scene information 2911 for non-timed media. The scene information 2911 can also reference scene information 2917 for timed media scenes. List 2918 identifies unique assets associated with a particular scene that have not been previously used in a higher-level (eg, parent) scene.
[0200] Figure 30 shows a diagram of a timed media presentation 3010 with signaling of replacement of non-visual assets in some examples. The timed media presentation 3010 includes a timed scene manifest 3010A containing information for a list of scenes 3011. The scene 3011 information points to a list of components 3012 that separately describe the processing information and types of media assets in the scene 3011. The components 3012 point to assets 3013, which further point to a base layer 3014 and an attribute enrichment layer 3015. The base layer 3014 of non-visual assets is provided with metadata to indicate whether the corresponding asset is a candidate for replacement with a user-provided asset (not shown). A list 3017 of assets not previously used in other scenes 3011 is provided for each scene 3011.
[0201] FIG. 31 shows a diagram of a non-timed media presentation 3110 with signaling of non-visual asset replacement in some examples. A non-timed scene manifest (not shown) references Scene 1.0, which has no other scenes that can branch to it. Scene 3111 information is not associated with a start and end duration in clock terms. Scene 3111 information points to a list of components 3112 that separately describe the processing information and type of media asset within the scene. Components 3112 point to assets 3113, which further point to base layer 3114 and attribute enrichment layers 3115 and 3116. Assets 3113 that are not of the "visual" type are further provisioned with metadata to indicate whether or not the corresponding asset is a candidate for replacement by a user-provided asset (not shown). Additionally, scene 3111 information can reference other scene information 3111 for non-timed media. Scene 3111 information can also reference scene information 3117 for timed media scenes. List 3118 identifies unique assets associated with a particular scene that have not been previously used in a higher-level (eg, parent) scene.
[0202] FIG. 32 illustrates a process (also referred to as replacement logic) 3200 for performing asset replacement with a user-provided asset on a client device according to some embodiments of the present disclosure. In some examples, process 3200 may be performed by a smart client in the client device. Starting at process step 3201, metadata within the asset is examined to determine whether the asset in question is a candidate for replacement. If the asset is not a candidate for replacement, the process continues with the asset (e.g., provided by the content provider) at process step 3207. Otherwise, if the asset is a candidate for replacement, decision-making step 3202 is performed. If the client device does not have the capability to store the user-provided asset (e.g., shown as media cache 1415 in FIG. 14), the process continues with the asset provided by the content provider at process step 3207. Otherwise, if the client device is provisioned with a user-provided asset cache (e.g., user-provided media cache 1416 in FIG. 14), the process continues at process step 3203. Process step 3203 constructs a query of an asset cache (e.g., user-provided media cache 1415 of FIG. 14) to retrieve a suitable user-provided asset in place of the asset provided by the content provider. Decision step 3204 determines whether a suitable user-provided asset is available. If a suitable user-provided asset is available, the process continues to process step 3205, where references to assets provided by the content provider are replaced with references to the user-provided assets in the media. The process then continues to process step 3206, which marks the end of the substitution logic 3200. If a suitable user-provided asset is not available from the asset cache (e.g., user-provided media cache 1416 of FIG. 14), the process continues with the asset provided by the content provider in process step 3207, i.e., no substitution is made.From process step 3207, processing continues to process step 3206, which indicates the end of the replacement logic 3200.
[0203] In some examples, process 3200 is applied to visual assets. In one example, a person may want to replace the visual assets of a character in media with the visual appearance of a person for a better experience. In another example, a person may want to replace the visual appearance of a building in a scene with the visual appearance of the actual building. In some examples, process 3200 is applied to non-visual assets. In one example, for a person with a visual impairment, the person can have a customized haptic asset to replace the haptic asset provided by the content provider. In another example, for a person from a region with an accent, the person can have an audio asset with the accent to replace the audio asset provided by the content provider.
[0204]
[00110] Figure 33 shows a process (also referred to as populating logic) 3300 for populating a user-provided media cache on a client device according to some embodiments of the present disclosure. The process begins at process step 3301, where a display panel on the client device (e.g., client device 1418 of Figure 14) presents the user with an option to load assets into the cache (e.g., user-provided media cache 1416 of Figure 14). Decision-making step 3302 is then performed to determine whether the user has more assets to load onto the client device (e.g., from an external storage device, such as user-provided media storage device 1419 of Figure 14). If the user has more assets to load onto the client device, the process proceeds to decision-making step 3303. If the user does not have any (more) assets to load onto the client device, the process proceeds to process step 3306. Process step 3306 marks the end of the populating logic 3300. Decision step 3303 determines whether there is enough storage on the client device (e.g., user-provided media cache 1416 of FIG. 14) to store the user-provided assets. If there is enough storage on the client device (e.g., user-provided media cache 1416 of FIG. 14), the process proceeds to process step 3304. If there is not enough storage to store the user-provided assets, the process proceeds to process step 3305. Process step 3305 issues a message notifying the user that the client device does not have enough storage to store the user's assets. Once the message is issued in process step 3305, the process proceeds to process step 3306, which marks the end of the populating logic 3300. If there is enough storage for the assets, process step 3304 copies the user's assets to the client device (e.g., user-provided media cache 1416 of FIG. 14).Once the assets are copied in process step 3304, the process returns to step 3302 to query whether the user has more assets to load onto the client device.
[0205] In some examples, the process 3300 is applied to visual assets (e.g., user-provided visual assets). In some examples, the process 3300 is applied to non-visual assets (e.g., user-provided non-visual assets).
[0206] 34 shows a flowchart outlining a process 3400 according to one embodiment of the present disclosure. Process 3400 can be performed on an electronic device, such as a client device having a smart client for interfacing the client device with a network. In some embodiments, process 3400 is implemented with software instructions, and thus, processing circuitry performs process 3400 when the processing circuitry executes the software instructions. For example, a smart client can be implemented with software instructions, and the software instructions can be executed to perform smart client processes, including process 3400. The process starts at S3401 and proceeds to S3410.
[0207] At S3410, a media stream carrying scene-based immersive media is received. The scene-based immersive media includes a plurality of media assets associated with a scene.
[0208] At S3420, a first media asset in the plurality of media assets is determined to be replaceable.
[0209] At S3430, the second media asset is substituted for the first media asset to generate updated scene-based immersive media.
[0210] In some examples, the base layer metadata of the first media asset may indicate that the first media asset is replaceable. In one example, the first media asset is a timed media asset. In another example, the first media asset is a non-timed media asset.
[0211] In some examples, the first media asset is a visual asset. In some examples, the first media asset is a non-visual asset, such as an audio asset, a haptic asset, etc.
[0212] In some examples, a storage device (e.g., a cache) of the client device is accessed to determine whether a second media asset corresponding to the first media asset is available on the client device. In one example, the smart client formulates a query asking whether a second media asset corresponding to the first media asset is available.
[0213] In some examples, the client device may perform a populating process to load a user-provided media asset into the storage device. For example, the client device may load a second media asset into the cache via a user interface. In some examples, the second media asset is a user-provided media asset.
[0214] Then, the process 3400 proceeds to S3499 and ends.
[0215] Process 3400 may be appropriately adapted to various scenarios, and steps within process 3400 may be adjusted accordingly. One or more of the steps within process 3400 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 3400. Additional steps may be added.
[0216] It should be noted that a Smart Client component (e.g., a Smart Client process, software instructions, etc.) can manage the receipt and processing of streamed media on behalf of a client device in a delivery network designed to stream scene-based media to a client device (e.g., an immersive client device). The Smart Client can be configured to perform many functions on behalf of the client device, including: 1) requesting media resources from the delivery network; 2) reporting on the current state or configuration of the client device resources, including attributes describing the client's preferred media format; 3) accessing media that may have been previously converted and stored in a format suitable for the client device, i.e., previously processed by another similar or identical client device and cached in data storage for subsequent reuse; and 4) substituting user-provided media assets for system-provided media assets. Furthermore, in some examples, several techniques can be implemented in a network device to perform functions similar to a Smart Client in a client device, but tailored to focus on adapting media from format A to format B to contribute to the effectiveness of a network that delivers immersive scene-based media to multiple heterogeneous client devices. In some examples, these techniques may be embodied in a network device as software instructions, which may be executed by processing circuitry in the network device, and the software instructions or the processes performed by the processing circuitry in accordance with the software instructions may be referred to as a smart controller in the sense that the smart controller configures, initiates, manages, terminates, and destroys resources for adapting ingested media into a client-specific format before the media is available for delivery to a particular client device.
[0217] According to a fourth aspect of the present disclosure, various techniques disclosed in the present disclosure can manage media adaptation on behalf of a network that converts media to a client-specific format on behalf of a client device. In some examples, a smart controller in a network server device is implemented according to some techniques, and the smart controller may be embodied by one or more processes with one or more communication channels implemented between the smart controller process and the network process. In some examples, some techniques include metadata to fully describe the intended results of the media conversion process. In some examples, some techniques can provide access to a media cache that may contain previously converted media. In some examples, some techniques can provide access to one or more renderer processors. In some examples, some techniques can provide access to sufficient GPU and / or CPU processors. In some examples, some techniques can provide access to storage devices sufficient to store the resulting converted media.
[0218] 35 illustrates a process 3510 for performing network-based media conversion in some examples. The network-based media conversion is performed within a network, such as by a smart controller in a network server device, based on client device information, which in some examples is provided by a smart client in the client device.
[0219] Specifically, in Figure 35, in process step 3511, media requested by the client device is ingested into the network. Next, in process step 3515, the network obtains the device attributes and resource status of the client device. For example, the network initiates a request to the smart client (e.g., as shown in process step 3512) to obtain the device attributes and resource status. In process step 3512, the smart client initiates queries to the client device 3513 and additional resources (if any) 3514 to retrieve information about the client device's resources, including descriptive attributes about the client device and resource availability for processing future work, respectively. The client device 3513 delivers client device attributes describing the processing capabilities of the client device. The status and availability of resources on the client device are returned from the additional resources 3514 to the smart client. Next, the smart client provides the status and availability of resources on the client device to the network, as shown in process step 3515. In process step 3516, the network then determines whether the network needs to transcode the media before it is streamed to the client device, for example, based on the state and availability of resources on the client device. If such transcoding is required, for example, if the client device is battery powered and lacks processing capabilities, smart controller 3517 performs the transcoding. Smart controller 3517 can utilize renderer 3517A, GPU 3517B, media cache 3517C, and neural network processor 3517D to assist in the transcoding. The transcoded media 3519 is then provided to process step 3518 for merging the transcoded media with the media to be streamed to the client device.If a determination is made in process step 3516 that the network should not perform such conversion of the media, for example, if the client device has sufficient processing power, process step 3518 prepares the media with the original media asset to be streamed to the client device. Following process step 3518, the media is streamed to the client in process step 3520.
[0220] FIG. 36 illustrates a process 3600 for performing network-based media adaptation in a smart controller within a network. Process step 3601 instantiates and / or initializes resources for performing media adaptation. Such resources may include separate computational threads for managing renderer and / or decoder instances to return the media to its original state, e.g., returning the media to an uncompressed representation of the same media. Decision-making process step 3602 determines whether the media needs to first undergo a reconstructing and / or decompressing process. If the media needs to be first reconstructed and / or decompressed before being adapted for delivery to the client device, process step 3603 performs the task of reconstructing and / or decompressing the media. The reconstructed and / or decompressed media is returned to the smart controller via a channel. Next, in process step 3604, the smart controller determines whether to further refine the media using a neural network and possibly a neural network model (not shown) provided or stored on the network on behalf of the client device to assist in adapting the media to a client-friendly format. If the smart controller determines that a neural network application is required, then in process step 3606, a neural network processor may apply a neural network model to the media. The media resulting from the neural network processing is returned via a channel to the smart controller. If, in process step 3604, the media does not need to be enhanced using a neural network processor, then the neural network processing is skipped. In process step 3608, the smart controller converts the media (e.g., reconstructed and / or decompressed media or resulting media from the neural network processing) into a format suitable for the particular client device. Such a conversion process may use a renderer tool, a 3D modeling tool, a video decoder and / or a video encoder.The processing performed by process step 3608 may include reducing the polygon count of a particular asset so that the computational complexity for rendering the asset by the client device matches the capabilities of the client device's computing power. Other factors, including, for example, replacing a particular asset with a different texture resolution, may also be addressed as part of conversion process step 3608. For example, if the UV texture of a particular mesh-based asset needs to be provided at a particular resolution, e.g., high definition (HD) rather than ultra high definition (UHD), conversion process 3608 may also create an HD representation of the mesh asset's UV texture. Another example is if a particular asset in general network media is described according to a particular 3D format specification, e.g., FBX, glTF, OBJ, and the client device is not capable of ingesting media displayed according to that particular format specification, conversion process 3608 may use 3D tools to convert the media asset into a format that can be ingested by the client device. Following completion of conversion process 3608, the media is stored in a client adaptive media cache 3609 for subsequent access by the client device (not shown). If necessary, the smart controller then performs a process 3610 of discarding or deallocating the computational and storage resources previously allocated in 3601 for the smart controller.
[0221] 37 shows a flowchart outlining a process 3700 according to one embodiment of the present disclosure. The process 3700 may be performed within a network, such as a server device within the network. In some embodiments, the process 3700 is implemented with software instructions, such that the processing circuitry performs the process 3700 when it executes the software instructions. For example, the process 3700 may be implemented as software instructions within a smart controller, and the processing circuitry may execute the software instructions to perform the smart controller process. The process starts at S3701 and proceeds to S3710.
[0222] At S3710, the server device determines a first media format based on the capability information of the client device, the first media format being processable by the client device having the capability information.
[0223] At S3720, the server device converts the media in the second media format into adapted media in the first media format. In some examples, the media is scene-based immersive media.
[0224] At S3730, the adapted media in the first media format is provided (eg, streamed) to the client device.
[0225] In some examples, the smart controller may include a renderer, a video decoder, a video encoder, a neural network model, etc. In some examples, the second media format is a network ingest media format. The smart controller may convert media in the network ingest media format to an intermediary media format and then convert the media from the intermediary media format to the first media format. In some examples, the smart controller may decode media in the ingest media format to generate reconstructed media, and then the smart controller may convert the reconstructed media to media in the first media format.
[0226] In some examples, the smart controller may cause the neural network processor to apply a neural network model to the reconstructed media to refine the reconstructed media and then convert the refined media into media in the first media format.
[0227] In some examples, the media in the second media format is a universal media type for network transmission and storage.
[0228] In some examples, the client device capability information includes at least one of a capability of computation by the client device and a capability of an accessible storage resource of the client device.
[0229] In some examples, the smart controller can cause the client device to send a request message, where the request message requests the capabilities of the client device, and the smart controller can then receive a response with the capability information of the client device.
[0230] The process 3700 then proceeds to S3799 and ends.
[0231] Process 3700 may be appropriately adapted to various scenarios, and steps within process 3700 may be adjusted accordingly. One or more of the steps within process 3700 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 3700. Additional steps may be added.
[0232] According to a fifth aspect of the present disclosure, various techniques can be used to characterize the capabilities of a client device with respect to the media it can ingest. In some examples, the characterization can be represented by a client media profile that serves to convey a description of scene-based immersive media asset types, the level of detail for each asset type, the maximum size in bytes for each asset type, the maximum number of polygons for each asset type, and other parameters that describe the types of media and characteristics of that media that the client device can ingest directly from the network. A network that receives the client media profile from the client device can then operate more efficiently in terms of preparing the ingested media to be delivered and / or accessed by the client device.
[0233] Figure 38 shows a diagram of some example client media profiles 3810. In the example of Figure 38, information about media formats, media containers, and other attributes related to the media supported by the client device is conveyed uniformly across the network (i.e., according to specifications not shown in Figure 38).
[0234] For example, the types of media formats that a client device can support are shown as list 3811. Data element 3812 conveys the maximum polygon count that the client device can support. Data element 3813 indicates whether the client device supports physically based rendering. List 3814 identifies asset media containers that the client device supports. List 3815 indicates that there are other media related items that may include a complete media profile to characterize the media preferences of the client device. The disclosed subject matter can be considered media based immersive media scenes resulting from information (including supported color formats, color depths, video and audio formats) exchanged via the Specification for High-Definition Multimedia Interface between a source process and a sink process.
[0235] 39 shows a flowchart outlining a process 3900 according to one embodiment of the present disclosure. The process 3900 may be performed on an electronic device, such as a client device having a smart client for interfacing the client device with a network. In some embodiments, the process 3900 is implemented in software instructions, and thus, when a processing circuit executes the software instructions, the processing circuit performs the process 3900. The process starts at S3901 and proceeds to S3910.
[0236] At S3910, the client device receives a request message from a network to stream media (e.g., scene-based immersive media) to the client device, the request message requesting capability information of the client device.
[0237] At S3920, the client device generates a media profile that indicates one or more media formats that can be processed by the client device.
[0238] At S3930, the client device transmits the media profile to the network.
[0239] In some examples, the media profile defines one or more types of scene-based media supported by the client device, as shown at 3811 in FIG.
[0240] In some examples, the media profile includes a list of media parameters that characterize the client device's capabilities to support particular variations of media in a manner consistent with the client device's processing capabilities, as shown by 3812, 3813, and 3814 in FIG. 38.
[0241] Then, the process 3900 proceeds to S3999 and ends.
[0242] Process 3900 may be appropriately adapted to various scenarios, and steps within process 3900 may be adjusted accordingly. One or more of the steps within process 3900 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 3900. Additional steps may be added.
[0243] 40 shows a flowchart outlining a process 4000 according to one embodiment of the present disclosure. The process 4000 may be performed within a network, such as on a server device within the network. In some embodiments, the process 4000 is implemented in software instructions, such that the processing circuitry performs the process 4000 when it executes the software instructions. The process starts at S4001 and proceeds to S4010.
[0244] In S4010, the first media format is determined, for example, by the server device, based on a client media profile indicating one or more media formats that can be processed by the client device. For example, the client media profile may be client media profile 3810 of FIG. 38.
[0245] At S4020, the media in the second media format is converted to adapted media in the first media format. In some examples, the media is scene-based immersive media in a network-ingested media format. The server device can convert the media in the network-ingested media format to the first media format that can be processed by the client device.
[0246] At S4030, the adapted media in the first media format is provided (eg, streamed) to the client device.
[0247] In some examples, the media profile defines one or more types of scene-based media supported by the client device, as shown at 3811 in FIG.
[0248] In some examples, the media profile includes a list of media parameters that characterize the client device's capabilities to support particular variations of media in a manner consistent with the client device's processing capabilities, as shown by 3812, 3813, and 3814 in FIG. 38.
[0249] Then, the process 4000 proceeds to S4099 and ends.
[0250] Process 4000 may be appropriately adapted to various scenarios, and the steps in process 4000 may be adjusted accordingly. One or more of the steps in process 4000 may be adapted, omitted, repeated, and / or combined. Any suitable order may be used to perform process 4000. Additional steps may be added.
[0251] While this disclosure describes several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]
[0252] 100 Media flow process, 104 Network, cloud or edge device, 105 Network connection, 106 Rendering process, rendering function, 108 Client device, 200 Media conversion decision process, 201 Media, 202 Process step, 203 Decision process step, 204 Process step, 205 Media, 206 Preparation process, 207 Process step, 300 Streamable format, timed media presentation, 300A Timed Scene Manifest, 301 Scene information, 302 Components, 303 Assets, 304 Base layer, 305 Attribute enrichment layer, 306 Proxy visual assets, 307 List of unique assets, 308 Proxy audio assets, 400 Streamable format, 401 Scene information, 402 Components, 403 Assets, 404 Base layer, 405 Attribute enrichment layer, 406 Attribute enrichment layer, 407 Scene information, 408 List, 500 Process, 501 Camera unit, 502 Camera unit, 503 Camera unit, 504 Synthesis module, 505 Neural network training module, 506 Training images, 507 Ingest format, 508 Neural network model, 509 Natural image content, 600 Process, 601 LIDAR camera, 602 Point cloud data, 603 Computer, 604 CGI asset data, 605 Actor, 605A Sensor, 606 Motion capture (MoCap) data, 607 Synthesis module, 608 Synthesis media, 700 Computer system, 701 Keyboard, 702 Mouse, 703 Trackpad, 705 Joystick, 706 Microphone, 707 Scanner, 708 Camera, 709 Speaker, 710 Touchscreen, 720 CD / DVD ROM / RW, 721 CD / DVD or similar media, 722 Thumb drive, 723 removable hard drive or solid state drive, 740 core, 741 central processing unit (CPU), 742 graphics processing unit (GPU), 743 field programmable gate array (FPGA), 744 hardware accelerator, 745 read-only memory (ROM), 746Random access memory, 747 internal mass storage, 748 system bus, 749 peripheral bus, 750 graphics adapter, 754 interface, 755 communication network, 800 network media distribution system, 801 content acquisition module, 802 content preparation module, 803 transmission module, 804 gateway, 805 set-top box, 806 wireless demodulator, 807 legacy 2D television, 808 WiFi router, 809 legacy 2D display, 810 head-mounted 2D (raster-based) display, 811 display, 811A local computation GPU, 811B storage device, 811C visual presentation unit, 812 holographic display, 812A local computation CPU, 812B GPU, 812C storage device, 812D visualization unit, 813 mobile handset and display, 814 augmented reality headset, 814A GPU, 814B 908A1 visualization component, 908A2 physics engine, 908B network interface, 908C GPU, 908D storage, 908E smart client, 908F callback function, 909 delivery storage device, 910 media adaptation and fragmentation module, 911 Media reuse analyzer, 1000 media adaptation process, 1001 media adaptation and fragmentation module, 1001A neural network model, 1001B renderer, 1001C neural network processor, 1001DMedia compressor, 1001E media decompressor, 1001F logic controller, 1003 client interface module, 1004 client information, 1005 input network state, 1006 client adaptive media storage device, 1007 media reuse analyzer, 1100 delivery format creation process, 1101 media adaptation module, 1102 client adaptive media storage device, 1103 adaptive media packaging module, 1104 robust delivery format, 1104A manifest information, 1104B scene data asset list, 1200 packetizer process system, 1201 adaptive media, 1202 packetizer, 1203 packet, 1204 client endpoint, 1300 sequence diagram, 1301 client, 1302 network orchestrator, 1303 ingest media server, 1304 adaptive ingest, 1305 media adaptation module, 1306 packaging module, 1307 Packaged Media Server, 1308 Media Request, 1309 Profile Request, 1310 Response, 1311 Session ID Token, 1312 Ingest Media Request, 1313 Response, 1314 Call, 1315 Neural Network Model Token, 1316 Request, 1317 Response, 1318 Request, 1319 Update, 1320 Interface Call, 1321 Response Message, 1322 Response, 1323 Request, 1324 Message, 1400 Media System, 1401 MPEG Smart Client Process, 1402 Client Media Reconstruction Process, 1403 Neural Network Processor, 1404 Client Adaptive Media Cache, 1405 Game Engine, 14051 Control Logic, 14052 GPU Interface, 14053 Physics Engine, 14054 Renderer, 14055 Compression Decoder, 14056 Device-specific plugin, 1406 compression decoder process, 1407 rendered client media cache, 1408 edge processor or network orchestrator device, 1409 client adaptive media, 1410 media analyzer, 1411Rendered media cache, 1412 user interface, 1413 haptic component, 1414 audio component, 1415 visualization component, 1416 user provided media cache, 1417 game engine API and callback function, 1418 game engine client device, 1419 media cache, 1420 network interface protocol, 1421 neural network model, 1530 process, 1531 process, 1532 process, 1533 process, 1535 format, 1536 media streaming process, 1537 media, 1538 information channel, 1539 smart client process, 15310 rendering process, 15311 presentation format C, 15312 game engine client device, 1640 process, 1641 process, 1642 process, 1642A distribution media creation initiation process, 1642B decision making process, 1642C process, 1642D Process, 1642E Process, 1643 Process, 1645 Alternate Format, 1646 Media Streaming Process, 1647 Streamed Media, 1648 Smart Client Feedback and Status, 1649 Smart Client Process, 16410 Rendering Process, 16411 Presentation Format C, 16412 Game Engine Client Device, 1645 Format, 1646 Media Streaming Process, 1647 Media, 1648 Smart Client Feedback and Status, 1649 Smart Client Process, 1730 Media Conversion Decision Process, 1731 Process Step, 1732 Process Step, 1733 Process Step, 1734 Process Step, 1735 Process Step, 1736 Process Step, 1737 Process, 1738 Process Step, 1739 Converted Media, 17310 Process Step, 17311 Process Step, 1840 Game Engine Client Process, 1841 Process Steps, 1842 Process Steps, 1843 Query Logic, 1843AProcess Step, 1843B Database, 1844 Process Step, 1845 Process Step, 1846 Process Step, 1847 Process Step, 1848 Network Process, 1849 Converted Media, 1950 Process, 1951 Process Step, 1952 Process, 1952A Smart Client, 1952B Client Device, 1952C Additional Resources, 1953 Asset Query Logic, 1954 Process Step, 1955 Process Step, 1956 Process Step, 1957 Process Step, 1958 Process Step, 1959 Converted Media, 2043C Smart Client Process, 2060 Process, 2061 Process Step, 20610 Smart Client, 20611 Client Device, 20612 Process Step, 20613 Media Store, 2062 Process Step, 2063 Asset Query Logic, 2064 process steps, 2065 process steps, 2066 process steps, 2067 process, 2068 process steps, 2069 transformed media, 2103A timed scene manifest, 2131 scene information, 2132 components, 2133 assets, 2134 base layer, 2135 attribute enhancement layer, 2136 proxy visual assets, 2137 list, 2138 proxy audio assets, 2240 diagram of non-timed media presentation, 2241 scene information, 2242 components, 2243 assets, 2244 base layer, 2245 attribute enhancement layer, 2246 attribute enhancement layer, 2247 scene information, 2248 list, 2340 distribution format creation process, 2341 media adaptation process, 2342 storage device, 2343 media packaging process, 2344 distribution format, 2344A Manifest information, 2344B List, 2400 Flowchart, 2404 List, 2500 Network process diagram, 2501 Network orchestration process, 2502 Network interface, 2503 Information, 2504 Client device, 2504A Smart client, 2504B Compression decoder process, 2504C Visualization component, 2504DClient Adaptive Media Cache, 2504E Neural Network Model, 2504F Neural Network Processor, 2504G Game Engine, 2504G1 Control Logic, 2504G2 Physics Engine, 2504G3 Renderer Process, 2504G5 GPU Interface, 2504G6 Compression Decoder, 2504H Client Media Reconstruction Process, 2504I Rendered Client Media Cache, 2504J Client Device Control Logic, 2600 Process, 2700 Process, 2810A Timed Scene Manifest, 2811 Scene, 2812 Component, 2813 Asset, 2814 Base Layer, 2815 Attribute Enhancement Layer, 2817 List, 2911 Scene Information, 2912 Component, 2913 Asset, 2914 Base Layer, 2915 Attribute Enhancement Layer, 2916 Attribute Enhancement Layer, 2917 Scene information, 2918 List, 3010 Timed media display, 3010A Timed scene manifest, 3011 Scene, 3012 Component, 3013 Asset, 3014 Base layer, 3015 Attribute enrichment layer, 3017 List, 3111 Scene, 3112 Component, 3113 Asset, 3114 Base layer, 3115 Attribute enrichment layer, 3116 Attribute enrichment layer, 3117 Scene information, 3118 List, 3200 Process, 3201 Process step, 3202 Decision step, 3203 Process step, 3204 Decision step, 3205 Process step, 3206 Process step, 3207 Process step, 3300 Populating logic, 3301 Process step, 3302 Decision step, 3303 Decision step, 3304 Process step, 3305 Process step, 3306 Process step, 3400 Process, 3510 Process, 3511 Process step, 3512 Process step, 3513 Client device, 3514 Additional resources, 3515 Process step, 3516 Process step, 3517 Smart controller, 3517A Renderer, 3517B GPU, 3517C Media cache, 3517D Neural network processor, 3518 Processor process step, 3519 Transformed Media, 3520 process step, 3600 process, 3601 process step, 3602 process step, 3603 process step, 3604 process step, 3606 process step, 3608 process step, 3609 Client Adaptive Media Cache, 3610 process, 3700 process, 3810 Client Media Profile, 3811 list, 3812 data element, 3813 data element, 3814 list, 3815 list, 3900 process, 4000 process
Claims
1. A media processing method performed by an electronic device, comprising: transmitting, via a client interface of the electronic device, information about the capabilities and availability of the electronic device to play scene-based immersive media to a server device in an immersive media streaming network; receiving, by the client interface, a media stream carrying adapted media content for the scene-based immersive media, the adapted media content being generated from the scene-based immersive media by the server device based on the capability and availability information; playing the scene-based immersive media according to the adapted media content; determining, by the client interface, that a first media asset associated with a first scene is received for the first time and should be reused in one or more scenes according to the adapted media content; storing the first media asset in a cache device accessible by the electronic device; Including, The method, wherein the media assets stored in the cache device are ordered by frequency in the cache device.
2. extracting, by the client interface, a first list of unique assets in the first scene from the media stream, the first list of unique assets identifying the first media asset as a unique asset in the first scene to be used in one or more other scenes; The method of claim 1 further comprising:
3. said step of transmitting said capability and availability information further comprising: sending, by the client interface, a signal to the server device indicating availability of the first media asset on the electronic device, the signal causing the server device to substitute a proxy for the first media asset in the adapted media content; 2. The method of claim 1, comprising:
4. The step of playing the scene-based immersive media comprises: determining, by the client interface, that the first media asset has previously been stored in the caching device according to the proxy in the adapted media content; accessing the cache device to retrieve the first media asset; The method of claim 3, further comprising:
5. The step of transmitting the signal indicating the availability of the first media asset comprises: receiving a query signal for the first media asset from the server device; transmitting the signal indicating the availability of the first media asset at the electronic device in response to the query signal; 4. The method of claim 3, comprising:
6. said step of transmitting said capability and availability information further comprising: receiving, by the client interface, a request to obtain device attributes and resource status from the server device; querying one or more internal components of the electronic device and / or one or more external components associated with the electronic device regarding attributes of the electronic device and resource availability for processing the scene-based immersive media; transmitting the attributes and resource availability of the electronic device to the server device; 2. The method of claim 1, comprising:
7. receiving a request for the scene-based immersive media from a user interface; forwarding, by the client interface, the request for the scene-based immersive media to the server device; The method of claim 1 further comprising:
8. The step of playing the scene-based immersive media comprises: generating, under control of the client interface, reconstructed scene-based immersive media based on decoding the media streams and media reconstruction; providing the reconstructed scene-based immersive media to a game engine of the electronic device for playback via an application programming interface (API) of the game engine; The method of claim 1 further comprising:
9. The step of playing the scene-based immersive media comprises: depacketizing, by the client interface, the media stream to generate depacketized media data; providing the depacketized media data to a game engine of the electronic device via an application programming interface (API) of the game engine; generating, by the game engine, reconstructed scene-based immersive media for playback based on the depacketized media data; The method of claim 1 further comprising:
10. An electronic device configured to perform the method according to any one of claims 1 to 9.
11. A computer program for causing a computer to carry out the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Media distribution device
JP2005295467A
Moving image processor
JP2010114815A