Method and related apparatus for streaming media assets

By maintaining redundant caches in the streaming media network to store redundant copies of immersive media assets, the problems of multiple client transformations and streaming transmission are solved, enabling efficient immersive media distribution to heterogeneous client devices and reducing the consumption of network and computing resources.

CN116648899BActive Publication Date: 2026-05-19TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-10-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, client devices need to transform and stream multiple times when using the same media object, resulting in a significant increase in network and computing resources, and commercial networks have difficulty effectively distributing immersive media from heterogeneous client devices.

Method used

By maintaining redundant caches in the streaming network, redundant copies of immersive media assets are stored, and the decision to stream directly or provide media assets from the redundant cache is made based on the client's local caching status, avoiding repeated transformations and streaming.

Benefits of technology

It reduces the demand for network and computing resources, improves the efficiency of immersive media distribution to heterogeneous client devices, and reduces network latency and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116648899B_ABST
    Figure CN116648899B_ABST
Patent Text Reader

Abstract

Methods and related apparatuses for streaming media assets are provided. The method can include receiving, by a streaming media server, an immersive media stream including one or more immersive media assets associated with one or more scenes; determining that a subset of the one or more immersive media assets is included multiple times in the one or more scenes; storing a redundant copy of each immersive media asset in the subset in a redundant cache maintained by a streaming media network such that both the streaming media server and a client have access to each immersive media asset in the subset; and responsive to a local cache of the client not storing at least one media asset in the subset, streaming the at least one media asset.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 276,523, filed November 5, 2021, and U.S. Patent Application No. 17 / 969,468, filed October 19, 2022, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure relates generally to the field of streaming media technology, and more specifically to methods and related apparatus for streaming media assets. Background Technology

[0004] Immersive media generally refers to media that stimulates any or all human sensory systems (e.g., vision, hearing, tactile sensation, smell, and possibly taste) to create or enhance a user's perception of physical presence within the media experience—that is, beyond the perception distributed over existing (e.g., "traditional") commercial networks for temporal two-dimensional (2D) video and corresponding audio; such temporal media is also known as "traditional media." Immersive media can also be defined as media that attempts to create or mimic the physical world through digital simulations of dynamics and physical laws, thereby stimulating any or all human sensory systems to create a user's perception of physical presence within a scene depicting a real or virtual world.

[0005] A presentation device with immersive media-capable capabilities can refer to a device equipped with sufficient resources and capabilities to access, interpret, and present immersive media. Such a device supports multiple media formats and also supports the multiple network resources required for the mass distribution of immersive media. "Mass distribution" can refer to media distribution by a service provider that achieves distribution over the network equivalent to the distribution of traditional video and audio media, such as Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.

[0006] In contrast, traditional presentation devices such as laptop computer monitors, televisions, and mobile handheld displays are homogeneous in their capabilities because all of these devices include rectangular displays that consume 2D rectangular video or still images as their primary visual media format. Some visual media formats commonly used in traditional presentation devices may include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Universal Video Coding / H.266.

[0007] The distribution of any media over a network can employ a media delivery system and architecture that reformats and / or converts media from an input format or a network-ingested media format into a distribution media format, wherein the distribution media format is not only suitable for ingestion by the target client device and its application, but also facilitates "streaming" over the network. Reformatting or streaming can be performed by the network (e.g., a server in a streaming media network), i.e., before the media is distributed to the client, resulting in a media format referred to as a "distribution media format" or simply "distribution format".

[0008] In existing technologies, if a client needs to use the same transformed media object (which can also be called a media asset) and / or streamed media object in multiple scenarios, these multiple uses will trigger multiple transformations and streaming of the media object. This multiple transformation and transmission of the media object is a root cause of network latency, leading to a significant increase in the network and / or computing resources used. Summary of the Invention

[0009] According to some embodiments, methods, systems, and apparatus are provided for facilitating a process for determining whether a client device should access copies of media assets and / or media objects stored in a local cache managed by the client device, or whether the client device should access copies of media assets stored in a redundant cache maintained by a server and / or network. According to some embodiments, the processes disclosed herein can be performed by a server or a client device.

[0010] According to one aspect of this disclosure, a method for streaming media assets is provided. The method includes: receiving an immersive media stream from a streaming media server, the immersive media stream including one or more immersive media assets associated with one or more scenes; determining that a subset of the one or more immersive media assets is included multiple times in one or more scenes; storing a redundant copy of each immersive media asset in the subset in a redundant cache maintained by a streaming media network, such that both the streaming media server and a client can access each immersive media asset in the subset; and streaming the at least one media asset in response to a local cache on the client not storing at least one media asset in the subset.

[0011] According to another aspect of this disclosure, an apparatus for streaming media assets is provided, the apparatus comprising: a receiving module configured to receive an immersive media stream via a streaming media server, the immersive media stream including one or more immersive media assets associated with one or more scenes; a determining module configured to determine that a subset of the one or more immersive media assets is included multiple times in the one or more scenes; a storage module configured to store redundant copies of each immersive media asset in the subset in a redundant cache maintained by the streaming media network, such that both the streaming media server and a client can access each immersive media asset in the subset; and a transmission module configured to stream the at least one media asset in the subset in response to a local cache of the client not storing at least one media asset in the subset.

[0012] According to another aspect of this disclosure, an apparatus for streaming media assets is provided. The apparatus may include: at least one memory configured to store computer program code; and at least one processor configured to read the computer program code and perform operations as instructed by the computer program code to implement the described method.

[0013] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided that stores instructions which, when executed by at least one processor, cause the at least one processor to perform the method described above.

[0014] Other implementations will be set forth in the description below and will be apparent in part from that description, and / or may be achieved by practicing the implementations presented in this disclosure.

[0015] According to the present invention, when a subset of one or more immersive media assets is determined to be included multiple times in one or more scenes, a redundant copy of each immersive media asset in the subset is stored in a redundant cache maintained by the streaming media network. This allows both the streaming media server and the client to access each immersive media asset in the subset, and in response to the client's local cache not storing at least one media asset in the subset, the at least one media asset is streamed. This avoids the client having to perform multiple transformations and / or streaming of specific media objects, which are required or will be required, to complete its media presentation, thus significantly reducing the required network and / or computing resources. Attached Figure Description

[0016] Figure 1A This is an exemplary illustration of media distribution to a client device in a streaming media network according to some implementation methods.

[0017] Figure 1BThis illustrates an exemplary workflow for creating media in a distribution format and generating reuse indicators in a streaming media network, according to some implementation methods.

[0018] Figure 2A This illustrates an exemplary workflow for streaming media to a client device according to some implementations.

[0019] Figure 2B This illustrates an exemplary workflow for streaming media to a client device according to some implementations.

[0020] Figure 2C This illustrates an exemplary workflow for streaming media to a client device according to some implementations.

[0021] Figure 3 This is an exemplary illustration of a data model for the representation and streaming of temporally immersive media according to some implementation methods.

[0022] Figure 4 This is an exemplary illustration of a data model for the representation and streaming of non-temporally immersive media according to some implementation methods.

[0023] Figure 5 This illustrates an exemplary workflow for natural media synthesis according to some implementation methods.

[0024] Figure 6 This illustrates an exemplary workflow for creating synthetic media ingestion according to some implementation methods.

[0025] Figure 7 This is an exemplary illustration of a computer system according to some implementation methods.

[0026] Figure 8 This is an exemplary illustration of a network media distribution system according to some implementation methods.

[0027] Figure 9 This illustrates an exemplary workflow for immersive media distribution using redundant caching according to some implementation methods.

[0028] Figure 10 This is a system diagram of a media adaptation processing system according to some implementation methods.

[0029] Figure 11 This illustrates an exemplary workflow for creating media in a distribution format according to some implementation methods.

[0030] Figure 12 This illustrates an exemplary workflow for packaging processing according to some implementation methods.

[0031] Figure 13This is an exemplary workflow illustrating the communication flow between components according to some implementations.

[0032] Figure 14A This illustrates an exemplary workflow for performing media reuse analysis according to some implementation methods.

[0033] Figure 14B This is an example of a unique list of assets for each scene in the presentation, based on some implementation methods. Detailed Implementation

[0034] The following detailed description of exemplary embodiments is given with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0035] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the implementation to the exact forms disclosed. Modifications and variations can be made based on the foregoing disclosure, or modifications and variations can be obtained from the practice of implementations. Furthermore, one or more features or components in one embodiment may be incorporated into or combined with another embodiment (or one or more features in another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.

[0036] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, software, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limited in its implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0037] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each appended dependent claim may directly refer to only one claim, the disclosure of possible implementations includes combinations of each dependent claim with every other claim in the claim set.

[0038] The proposed features discussed below can be used individually or in any order. Furthermore, implementations can be carried out using a processing circuit system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0039] Unless explicitly stated otherwise, elements, actions, or instructions used herein should not be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “one” are intended to include one or more items and may be used interchangeably with “one or more.” The term “one” or similar language is used where only one item is intended. Similarly, as used herein, the terms “have,” “possess,” “contain,” “include,” “comprise,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on.” Additionally, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0040] According to some implementations, a presentation device with immersive media capabilities can refer to a device equipped with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices are heterogeneous in terms of the number and formats of media they can support, as well as the quantity and type of network resources required for the large-scale distribution of such media. "Large-scale" can refer to media distribution by a service provider that achieves distribution over the network equivalent to the distribution of traditional video and audio media, such as Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.

[0041] According to some implementations, client devices used as endpoints for distributing immersive media over a network are highly diverse. The distribution of any media over a network can employ media delivery systems and architectures that reformat media from an input format or network-ingested media format into a distribution media format, wherein this distribution media format is not only suitable for ingestion by the target client device and its applications but also facilitates streaming over the network. Therefore, there can be two processes performed by the network on the ingested media: 1) converting the media from format A to format B suitable for ingestion by the target client, i.e., based on the client's ability to ingest a specific media format, and 2) preparing the media for streaming.

[0042] In some implementations, streaming media broadly refers to the segmentation and / or packaging of media, enabling its delivery over a network in consecutive, smaller chunks logically organized and ordered according to one or both of the media's temporal or spatial structure. Transforming media from format A (sometimes called "transcoding") to format B can be a process typically performed by the network or service provider before the media is distributed to client devices. Such transcoding may include converting media from format A to format B based on prior knowledge that format B is the preferred or only format that the target client device can ingest, or is more suitable for distribution on resource-constrained networks such as commercial networks. In many, but not all, cases where both media transformation and preparation for streaming are necessary before the target client device can receive and process the media from the network are possible.

[0043] Media conversion (or transformation) and media preparation for streaming are steps in the process by which the network takes action on the ingested media before distributing it to a client device. The result of this process (i.e., conversion and preparation for streaming) is a media format known as the distribution media format (or simply distribution format). If these steps are performed on a given media data object, they should only be performed once, provided the network has access to information indicating which media objects the client will need to convert and / or stream in multiple scenarios; otherwise, multiple conversions and streaming of such media will be triggered. That is, the processing and transmission of data for media conversion and streaming is generally considered a source of latency, potentially consuming significant network and / or computing resources. Therefore, a network is less efficient if it lacks access to information indicating when a client has stored a particular media data object in its cache or relative to local client storage.

[0044] A scene graph can be a general data structure commonly used in vector-based graphics editing applications and modern computer games, which arranges the logical representation of a graphical scene and usually (but not necessarily) arranges the spatial representation of a graphical scene, or it can be a collection of nodes and vertices in a graphical structure.

[0045] In the context of computer graphics, a scene can be a collection of the following: objects (e.g., 3D assets, also referred to as media assets, media objects, objects, and assets), object properties, and other metadata, wherein the other metadata includes visual, acoustic, and physical features describing a particular setting that is spatially or temporally defined relative to the interaction of objects within the setting.

[0046] Nodes can be basic elements of a scene graph, including information or related processing information related to logical, spatial, or temporal representations of vision, audio, touch, smell, or taste; each node should have at most one output edge, zero or more input edges, and at least one edge (input or output) connected to the node.

[0047] The base layer can be a nominal representation of a media asset, typically designed to minimize the computational resources or time required to render the asset or the time required to transmit the asset over a network.

[0048] An enhancement layer can be a set of information that, when applied to the base layer representation of an asset, expands the base layer to include features or capabilities not supported in the base layer.

[0049] Attributes can be metadata associated with a node, used to describe a specific characteristic or feature of that node in a canonical or more complex form (e.g., based on another node).

[0050] Containers can be serialization formats used to store and exchange information representing fully natural scenes, fully synthetic scenes, or a mixture of synthetic and natural scenes, including scene graphs and all media resources required to render the scene.

[0051] Serialization can be the process of converting the state of a data structure or object into a format that can be stored (e.g., stored in a file or storage buffer) or transmitted (e.g., transmitted over a network link) and subsequently (potentially in different computing environments). When the resulting bit sequence is reread according to the serialized format, that bit sequence can be used to create an object that is semantically identical to the original object.

[0052] A renderer can be (typically software-based) an application or process that, based on a selective mix of disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development, emits typical visual and / or audio signals, given an input scene graph and an asset container, suitable for rendering on a target device or conforming to desired characteristics specified by attributes of the rendering target node in the scene graph. For visual-based media assets, a renderer may emit visual signals suitable for the target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in the graphics pipeline); for audio-based media assets, a renderer may emit audio signals for rendering in multi-channel speakers and / or stereo headphones, or for repackaging into another (output) container. Popular examples of renderers include the real-time rendering features of game engines Unity and Unreal Engine.

[0053] The scripting language can be an interpreted programming language that can be executed by the renderer at runtime to handle dynamic inputs to scene graph nodes and variable state changes made to scene graph nodes, which affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse motion, deformation, collisions), as well as energy propagation and transmission (light, sound).

[0054] Shaders can be a class of computer programs that were originally used for coloring (producing appropriate levels of light, darkness, and color within an image), but now they perform various specialized functions in different areas of computer graphics special effects, or perform video post-processing unrelated to coloring, or even functions completely unrelated to graphics.

[0055] Path following is a computer graphics method for rendering 3D scenes to make the scene lighting faithful to reality.

[0056] Timed media can include media and / or media objects that can be ordered by time; for example, media with start and end times based on a specific clock. Untimed media can include media and / or media objects that can be organized by spatial, logical, or temporal relationships; for example, in interactive experiences implemented based on user actions.

[0057] A neural network model (NN model) can be a set of tensors (e.g., matrices) that define weights (i.e., numerical values) and parameters that are applied to well-defined mathematical operations on a visual signal to obtain an improved visual output, which may include interpolation of visual signals in a new view that are not explicitly provided for the original signal.

[0058] Over the past decade, the number of devices with immersive media capabilities—including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and game consoles—has surged in the consumer market. Furthermore, holographic displays and other forms of volumetric displays are expected to enter the consumer market within the next three to five years. However, despite the imminent or near-term availability of these devices, a coherent end-to-end ecosystem for distributing immersive media through commercial networks has failed to materialize for several reasons.

[0059] One reason why a coherent, end-to-end ecosystem for distributing immersive media via commercial networks has not yet been realized is the high diversity of client devices that act as endpoints in such distribution networks for immersive displays. Some client devices support certain immersive media formats, while others do not. Some client devices can create immersive experiences based on traditional raster-based formats, while others cannot. Unlike networks designed solely for the distribution of traditional media, such networks, which must support multiple display clients, require a wealth of information about the capabilities of each client and the format of the media to be distributed before they can employ adaptation processing to convert media into a format suitable for each target display and corresponding application. Such networks will at least need access to information describing the characteristics of each target display and the complexity of the ingested media so that the network can determine how meaningfully the input media source can be adapted to a format suitable for the target display and application.

[0060] Networks supporting such heterogeneous client devices should leverage the fact that some assets adapted from the input media format to a specific target format can be reused across a set of similar display targets. That is, some assets, once converted to a format suitable for the target display, can be reused on several such displays with similar adaptation requirements. Therefore, networks employing caching mechanisms to store adapted assets in relatively immutable areas will be more efficient.

[0061] Immersive media can be organized into “scenes” described by scene diagrams, also known as scene descriptions. A scene diagram can be used to describe visual, audio, and other forms of immersive assets that comprise a specific setting as part of the presentation, such as events occurring at a specific location within a building as part of the presentation (e.g., a film), and the actors. A list of all scenes comprising a single presentation can be compiled into a scene inventory.

[0062] The advantage of a "scenario-based" approach is that, for content prepared before it must be distributed, a "bulk list" can be created that identifies all assets that will be used throughout the presentation, and how frequently each asset will be used in the various scenarios within the presentation. The network possesses knowledge about the existence of cached resources that can be used to meet the asset requirements of a particular presentation. Similarly, a client device presenting a series of scenarios may want to have knowledge about how frequently any given asset will be used in multiple scenarios. For example, if a media asset (also referred to as a media object, asset, or object) is referenced multiple times in multiple scenarios that the client device is processing or will process, the client device should avoid discarding that asset from its cached resources until the last scenario requiring that particular asset has been presented by the client.

[0063] For traditional presentation devices, the distribution format can be equivalent to or fully equivalent to the "presentation format" that the client presentation device ultimately uses to create the presentation. That is, the presentation media format is a media format whose attributes (resolution, frame rate, bit depth, color gamut, etc.) are tuned to approximate the capabilities of the client presentation device. Some examples of distribution formats relative to presentation formats include: a high-definition (HD) video signal (1920 pixel columns x 1080 pixel rows) distributed over a network to an Ultra-high-definition (UHD) client device with a certain resolution (3840 pixel columns x 2160 pixel rows). A UHD client can apply a process called "super-resolution" to the HD distribution format to upscale the resolution of the video signal from HD to UHD. Therefore, the final signal format presented by the client device is the "presentation format," which in this example is a UHD signal, while the HD signal includes the distribution format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format because both signals are linear video formats, and the process of converting HD to UHD is a relatively straightforward and simple process performed on most traditional client devices.

[0064] However, in some implementations, the preferred presentation format of the target client device may differ significantly from the ingested format received by the network. Nevertheless, the client device may have access to sufficient computing, storage, and bandwidth resources to transform the media from the ingested format to the necessary presentation format suitable for presentation by the client device. The network can bypass the step of reformatting the ingested media, for example, "transcoding" the media from a first format A to a second format B, because the client has sufficient resources to perform all media transformations without requiring the network to do so beforehand. The network can still perform the steps of segmenting and packaging the ingested media so that the media can be streamed to the client.

[0065] However, in some implementations, the ingested media received by the network differs significantly from the client's preferred rendering format, and the client device lacks access to sufficient computing, storage, and / or bandwidth resources to convert the media into the preferred rendering format. In cases where the client lacks resource access, the network can assist the client by performing some or all of the transformations from the ingested format to an equivalent or nearly equivalent rendering format of the client device on its behalf. In some implementations, such assistance provided by the network on behalf of the client is often referred to as "split rendering."

[0066] Implementations of the present disclosure described herein enable the determination of whether a network should transform some or all of the ingested media from a first format (e.g., format A) to a second format (e.g., format B) to facilitate the client device's ability to present the media in a potential third format C. This determination can be made by a processor or server of the streaming network or by the client device. To aid in this determination, it may be useful to determine which media assets are used more than once within the context of presentation and to design processes and / or the network to make these media assets readily available for network adoption. Based on information from such analysis, the network can be designed such that it can request client devices (also referred to as "clients") to retain copies of one or more media assets that can be used more than once in their local cache.

[0067] However, if a client device stores a copy of a media asset in its local cache, the network may lack control over the management of that local cache, potentially leading to situations where the client device must remove resources (even reusable ones) from its local cache. For ease of design, and to optimize the network to minimize the need for conversions from a first format to a second format for repeatedly used media assets, or to avoid the network having to re-stream repeatedly used media assets to the client, the network can manage its own cache, separate from any cache maintained by the client. This ensures that at least one redundant copy of each reusable asset is accessible to both the client and the network. For example, a redundant copy of a media asset is a copy of a reused or previously used media asset. The redundant copy contains the essential elements of the media asset without losing features or characteristics and can be streamed in place of the corresponding media asset. The redundant copy is called "redundant" because, for example, the media asset can be stored via a network cache regardless of whether the client's local cache stores it.

[0068] In some implementations, the network may first query the client device to obtain feedback to ensure that the requested media asset is still available in the client's local cache. If the client device's response indicates that it no longer has a copy of the requested media asset, the network may signal to the client that it should access a copy of the media asset in its distribution format from the redundant cache. In some implementations, the query to the client device may be omitted, and the network may signal to the client that it should access a copy of the asset's distribution format from the redundant cache.

[0069] Figure 1A This is an exemplary illustration of a media distribution process 100, according to some implementations, distributing media from a network cloud, edge device, or server 104 to a client device 108. (As...) Figure 1A As shown, media in a first format A (hereinafter referred to as "ingest media format A") including immersive media is received from a content provider. This immersive media includes one or more scenes and one or more media objects. This process (i.e., media flow process 100) can be performed or implemented by a network cloud or edge device (hereinafter referred to as "network device 104") and distributed to clients, such as client device 108. In some implementations, the same process can be performed beforehand manually or by the client device. Network device 104 can ingest media 101 in the first format, generate and / or create distribution media 102 in a second format (hereinafter referred to as "distribution media creation 102"), and distribute media 103 in the second format, for example, using a distribution module. Client device 108 may include a rendering module 106 and a presentation module 107.

[0070] According to one aspect, network device 104 can receive ingested media from a content provider, etc. The streaming network can obtain ingested media stored in ingested media format A. Any necessary transformations or adjustments to the ingested media can be used to create and / or generate distribution media to create potential alternative representations of the media. That is, distribution formats for each media object in the ingested media can be created. As mentioned, a distribution format is a media format that can be distributed to a client by formatting the media into distribution format B. Distribution format B is a format prepared for streaming to client device 108. Distribution media creation 102 may include optimized reuse logic for performing a decision process to determine whether a particular media object has already been streamed to client device 108. (See also...) Figure 1B Describe in detail other operations related to the creation of distribution media 102 and the optimization of reuse logic.

[0071] Media format A and media format B may or may not be representations that follow the same syntax as a specific media format specification; however, format B can be converted into a scheme that facilitates media distribution via a network protocol. The network protocol can be, for example, a connection-oriented protocol (TCP) or a connectionless protocol (UDP). The distribution module streams the streamable media (i.e., media format B) from network device 104 to client device 108 via network connection 105.

[0072] Client device 108 can receive distribution media and can use rendering module 106 to render the media for presentation. Rendering module 106 may have access to rendering capabilities, which may be basic or similar / complex, depending on the client device 108 being targeted. Rendering module 106 can create presentation media in presentation format C. Presentation format C may or may not be represented according to a third format specification. Therefore, presentation format C may be the same as or different from media format A and / or media format B. Rendering module 106 outputs presentation format C to presentation module 107, which can display the presentation media on the display (or similar device) of client device 108.

[0073] Implementations of this disclosure improve upon the network's decision-making process for calculating the order in which assets are packaged and streamed from the network to clients. A media reuse analyzer analyzes all assets used in a set of one or more scenarios to determine the frequency with which each asset is used in all those scenarios. Therefore, the order in which these assets are packaged and streamed to clients can be determined based on the frequency with which each asset is used in the set of scenarios.

[0074] Implementations of this disclosure provide mechanisms or processes for analyzing immersive media scenarios to obtain sufficient information to support a decision-making process. When the network or client employs this decision-making process, it provides instructions regarding whether a transformation of a media object from format A to format B should be performed entirely by the network, entirely by the client, or via a combination of both (along with instructions on which assets should be transformed by the client or network). Such an immersive media data complexity analyzer can be used automatically or manually by the client or network, for example, depending on the user's operating system or device.

[0075] According to some implementations, the process of adapting an input immersive media source to a specific endpoint client device can be the same as or similar to the process of adapting the same input immersive media source to a specific application running on that specific client endpoint device. Therefore, the problem of adapting the characteristics of an input media source to an endpoint device has the same complexity as the problem of adapting a specific input media source to the characteristics of a specific application.

[0076] Figure 1B This is a workflow for creating 102 based on distribution media according to some implementation methods. More specifically, Figure 1B The workflow involves reusing logical decision-making processes, which help the decision-making process determine whether a particular media object has been streamed to the client device 108.

[0077] At point 152, the media creation process begins. At point 155, conditional logic can be executed to determine whether the current media object has previously been streamed to client device 108. A unique asset list used for rendering can be accessed to determine if the media object has previously been streamed to the client. If the current media object has previously been streamed, the process proceeds to operation 160. At operation 160, an indicator (later also referred to as a "proxy") is created to identify that the client has received the current media object and should access a copy of the media object from a local cache or other cache. If it is determined that the media object has not previously been streamed, the process proceeds to operation 165. At operation 165, the media object can be prepared for transformation and / or distribution, and the distribution format of the media object is created. Processing for the current media object then concludes.

[0078] Figure 2A This is an exemplary workflow for processing ingested media via a network. Figure 2A The illustrated workflow depicts a media transformation decision process 200 according to some implementations. The media transformation decision process 200 is used to determine whether the network should transform the media before distributing it to client devices. The media transformation decision process 200 can be handled by manual or automatic processes within the network.

[0079] The content provider supplies the network with ingested media, represented in Format A. At 205, the streaming network ingests media from the content provider. At 210, if the attributes of the target client are not yet known, the attributes of the target client are obtained. These attributes describe the processing capabilities of the target client.

[0080] At 215, it is determined whether the network (or client) should assist in transforming the ingested media. In some implementations, at 215, it may specifically be determined whether any format conversion of the media asset (e.g., conversion of one or more media objects from format A to format B) is required before streaming any media asset contained within the ingested media to the target client. At 215, this determination may be based on whether the media can be streamed in its original ingested format A, or whether the media must be transformed into a different format B for the client to present the media. Such a decision (i.e., determining whether media transformation is required before streaming the ingested media to the client, or whether the media should be streamed directly to the client in its original ingested format A) may require access to information describing aspects or characteristics of the ingested media.

[0081] If it is determined that the network (or client) should assist in transforming any media assets in the media assets ("Yes" at 215), then process 200 proceeds to 220.

[0082] At 220, the captured media is transformed from format A to format B, resulting in transformed media 222. Transformed media 222 is output, and at 225, the input media undergoes a preparation process for streaming to the client. In this case, the transformed media 222 (i.e., the input media) is prepared for streaming.

[0083] Streaming immersive media, especially when such media is “scene-based” rather than “frame-based,” may be relatively new. For example, streaming frame-based media can be equivalent to streaming video frames, where each frame captures a complete view of the entire scene or a complete view of an entire object to be presented by the client. This frame sequence creates a video sequence that includes the entire immersive presentation or a portion of it, as the client reconstructs the frame sequence from a compressed form and presents it to the viewer. For frame-based streaming, the order in which frames are streamed from the network to the client can conform to predefined specifications (e.g., ITU-T Recommendation H.264 Advanced Video Coding for General Audiovisual Services). However, scene-based streaming of media differs from frame-based streaming because a scene can include individual assets that may be independent of each other. A given scene-based asset can be used multiple times within a specific scene or across a series of scenes. The amount of time required for a client or any given renderer to reconstruct a particular asset can depend on several factors, including but not limited to: the size of the asset, the availability of computational resources for rendering, and other properties that describe the overall complexity of the asset. Clients supporting scene-based streaming may need to render some or all of each asset within a scene before any rendering of the scene can begin. Therefore, the order in which assets are streamed from the network to the client can impact the overall performance of the system.

[0084] The transformation of media from format A to another format (e.g., format B) can be done entirely by the network, entirely by the client, or jointly by both the network and the client. For split rendering, it becomes apparent that a lexicon describing the properties of the media format may be needed, so that both the client and the network have complete information to characterize the work that must be done. Furthermore, a lexicon providing attributes of the client's capabilities, such as available computing resources, available storage resources, and access bandwidth, may also be needed. Further still, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of ingesting the media format, allowing the network and client to jointly or separately determine whether or when the network can employ a split rendering process to distribute media to the client.

[0085] If it is determined that the network (or client) should not (or does not need) assist the media in transforming any media assets in the assets ("No" at 215), then process 200 proceeds to 225. At 225, the media is prepared for streaming. In this case, the ingested data (i.e., the media in its original form) is prepared for streaming.

[0086] Finally, once the media data is in a streamable format, at position 230, the media already prepared at position 225 will be streamed to the client. In some implementations (see reference...) Figure 1B As described, if the transformation and / or streaming of specific media objects required or to be required by the client to complete its media presentation can be avoided, the network can skip the transformation and / or streaming of the ingested media while the client still has access to or availability of the media objects that may be needed to complete its media presentation (i.e., 215 to 230). To determine the sequence of scene-based assets streaming from the network to the client to facilitate the client's ability to fully realize its potential, the network needs to be equipped with sufficient information to determine such a sequence to improve client performance. For example, a network that has sufficient information to avoid multiple transformation and / or streaming steps for assets used multiple times in a particular presentation can be more efficient. Similarly, if the network can "intelligently" order the assets delivered to the client, it can enable the client to fully realize its potential (i.e., provide a more enjoyable experience for the end user).

[0087] Figure 2B An example media transformation process 250, according to some implementations, includes determining media asset reuse and redundant caching. Similar to the media transformation decision process 200, the media transformation process 250, which utilizes asset reuse and redundant caching, processes ingested media over the network to determine whether the network should transform the media before distributing it to clients.

[0088] Content providers supply ingested media, represented in format A, to the network. According to some implementations, this is related to... Figure 2A Operations 205 to 210 and 215 to 230 shown are similarly performed as operations 255 to 260 and 275 to 286. At 255, media is retrieved by the network from the content provider. Then, at 260, if the attributes of the target client are not yet known, the attributes of the target client are obtained. These attributes describe the processing capabilities of the target client.

[0089] If it is determined that the network has previously streamed a specific media object or the current media object ("Yes" at 265), the process proceeds to 270. At 270, a proxy is created to replace the previously streamed media object to instruct the client whether to use a local copy of the previously streamed object or a copy of the previously streamed object stored in a redundant cache managed by the streaming media network.

[0090] If it is determined that the network has not previously streamed the media object ("No" at 265), the process proceeds to 275. At 275, it is determined whether the network or the client should perform any format transformation on any media assets contained within the media assets ingested at 255. For example, the transformation may include converting the media object from format A to format B before streaming the specific media object to the client. Operation 275 can be combined with... Figure 2A The operation performed at 215 is similar to the one shown in the diagram.

[0091] If it is determined that the media assets should be transformed by the network ("Yes" at 275), the process proceeds to 280. At 280, the media object is transformed from format A to format B. Then, preparations are made to stream the transformed media to the client (286).

[0092] If it is determined that the media assets should not be transformed by the network ("No" at 275), the process proceeds to 285. At 285, the media object to be streamed to the client is prepared. Once the media is in a streamable format, the media prepared at 285 is streamed to the client at 286.

[0093] Finally, at 288, it is determined whether the media assets streamed to the client should also be stored in the redundant cache. Based on the determination that the media assets streamed to the client will subsequently be reused for rendering, and that the media assets streamed to the client have never been streamed to the client before (or this is the first time the media has been streamed to the client), the media assets that have not been streamed to the client can be stored in the redundant cache at 289.

[0094] Figure 2C An example media transformation process 2500, including client query and asset reuse determination, is illustrated according to some implementations. Similar to media transformation process 250, media transformation process 2500 utilizing asset reuse and client query processes ingested media over a network to determine whether the network should transform the media before distributing it to the client.

[0095] Content providers supply ingested media, represented in format A, to the network. According to some implementations, this is related to... Figure 2BOperations 255 to 265 and 275 to 289 shown are performed similarly. At 255, media is retrieved by the network from the content provider. Then, at 260, if the attributes of the target client are not yet known, the attributes of the target client are obtained. These attributes describe the processing capabilities of the target client.

[0096] If it is determined that the network has previously streamed a specific media object or the current media object ("Yes" at 265), the process proceeds to 290. At 290, it is determined whether the client can access the media asset. If it is determined at 290 that the client can access the media asset, a proxy is created at 295 to replace the previously streamed media object, instructing the client to use a local copy of the previously streamed object. If it is determined at 290 that the client cannot access the media asset, at 293, a proxy is created to replace the previously streamed media object, instructing the client to use a copy of the previously streamed object stored on a redundant cache.

[0097] The streaming format of media can be time-series or non-time-series heterogeneous immersive media. Figure 3 An example of a time-series media representation 300 in a streamable format for heterogeneous immersive media is shown. Time-series immersive media can include a set of N scenes. Time-series media is media content ordered chronologically according to a specific clock, for example, having start and end times. Figure 4 An example of a non-chronological media representation 400 in a streamable format for heterogeneous immersive media is shown. Non-chronological media is media content organized according to spatial, logical, or temporal relationships (e.g., in an interactive experience realized based on actions taken by one or more users).

[0098] Figure 3 It involves timing scenarios for timing media, and Figure 4 This involves non-temporal scenarios for non-temporal media. Temporal and non-temporal scenarios can correspond to various scenario representations or scenario descriptions. Figure 3 and Figure 4 All use a single exemplary encompassing media format that has been adapted from the source ingested media format to match the capabilities of a specific client endpoint. That is, an encompassing media format is a distribution format that can be streamed to client devices. The encompassing media format is structurally robust enough to accommodate a wide variety of media attributes, where each media attribute can be layered based on the amount of significant information contributed by each layer to the presentation of the media.

[0099] like Figure 3As shown, the time-series media representation 300 includes a time-series scene list 300A, which includes a list of scene information 301. Scene information 301 refers to a list of components 302 that individually describe the processing information and type of the media assets constituting scene information 301. For example, an asset list and other processing information. The list of components 302 may involve proxy assets 308 corresponding to the asset type (e.g., such as...). Figure 3 (The proxy visual and audio assets shown). Component 302 involves a list of unique assets that have not been previously used in other scenarios. For example, in Figure 3 The diagram shows a unique list of assets 307 for (temporal) scene 1. Component 302 also involves assets 303 comprising a base layer 304 and an attribute enhancement layer 305. The base layer is a nominal representation of an asset, which can be configured to minimize computational resources, the time required to render the asset, and / or the time required to transmit the asset over a network. In this exemplary embodiment, each base layer in the base layer 304 refers to a digital frequency metric indicating the number of times the asset is used in the set of scenes to be rendered. An enhancement layer can be a set of information that expands the base layer representation when applied to the asset to include features or capabilities that may not be supported in the base layer.

[0100] like Figure 4 As shown, the non-chronological media and complexity representation 400 includes scene information 401. Scene information 401 is not associated with start and end times / durations (based on clocks, timers, etc.). A list of non-chronological scenes (not shown) can be referenced to scene 1.0, for which there are no other scenes that can branch off relative to scene 1.0. Scene information 401 refers to a list of components 402 that individually describe the processing information and type of the media assets constituting scene information 401. Components 402 involve visual assets, audio assets, haptic assets, and chronological assets (collectively referred to as assets 403). Assets 403 also involve a base layer 404 and attribute enhancement layers 405 and 406. In this exemplary embodiment, each base layer in base layer 404 refers to a digital frequency value indicating the number of times the asset is used in the set of scenes including the presentation. Scene information 401 may also involve other non-chronological scenes for non-chronological media sources (i.e., in Figure 4 The references in sections 2.1 to 2.4 (for non-temporal scenarios) and / or section information 407 (for temporal media scenarios) are used in sections 2.1 to 2.4 and / or for temporal media scenarios. Figure 4 It was cited as a timing scenario 3.0 in [the context]. Figure 4 In the example, non-chronological immersive media comprises a set of five scenes (including both chronological and non-chronological scenes). A unique asset list 408 identifies a unique asset associated with a specific scene that was not previously used in a higher-level (e.g., parent) scene. Figure 4The unique asset list 408 shown includes unique assets for non-sequential scenario 2.3.

[0101] Media streamed according to inclusive media formats is not limited to traditional visual and audio media. Inclusive media formats can include any type of media information capable of generating signals that can interact with machines to stimulate human senses of sight, hearing, taste, touch, and smell. Figures 3 to 4 As shown, media streamed according to inclusive media formats can be time-series media, non-time-series media, or a mixture of both. Inclusive media formats are capable of streaming by using a base layer and enhancement layer architecture to implement a hierarchical representation of media objects.

[0102] In some implementations, separate base and enhancement layers are computed by applying multi-resolution or multi-segmentation analysis techniques to media objects in each scene. This computational technique is not limited to raster-based visual formats.

[0103] In some implementations, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using wavelet analysis techniques.

[0104] In some implementations of a layered representation media format, enhancement layers can apply different properties to a base layer. For example, one or more enhancement layers can refine the material properties of the surface of a visual object represented by the base layer.

[0105] In some implementations, in a layered representation media format, attributes can refine the texture of a surface by, for example, changing the surface of an object represented by a base layer from a smooth texture to a porous texture or from a matte surface to a glossy surface.

[0106] In some implementations, in a layered representation media format, the surfaces of one or more visual objects in a scene can be changed from Lambertian surfaces to light-following surfaces.

[0107] In some implementations of a layered representation media format, the network can distribute a base layer representation to a client, allowing the client to create a nominal representation of the scene while the client waits for the transmission of other enhancement layers to refine the resolution or other features of the base layer.

[0108] In some implementations, the resolution of attributes or refinement information in the enhancement layer is not explicitly coupled to the resolution of objects in the base layer. Furthermore, inclusive media formats can support any type of information media that can be rendered or driven by a rendering device or machine, thereby enabling heterogeneous media formats to support heterogeneous client endpoints. In some implementations, the network distributing the media format will first query the client endpoint to determine the client's capabilities. Based on this query, if the client cannot meaningfully ingest the media representation, the network can remove attribute layers that the client does not support. In some implementations, if the client cannot meaningfully ingest the media representation, the network can adapt the media from its current format to a format suitable for the client endpoint. For example, the network can adapt the media by using a network-based media processing protocol to convert stereoscopic media assets into a 2D representation of the same visual asset. In some implementations, the network can adapt the media by employing a Neural Network (NN) process to reformat the media into an appropriate format or optionally synthesize the view required by the client endpoint.

[0109] The scene inventory for a complete (or partially complete) immersive experience (replay of a live streaming event, game, or on-demand asset) is organized by containing the minimum amount of information needed for rendering and ingestion to create the scenes to be presented. The scene inventory includes a list of individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene are one or more representations of geometric objects within the scene corresponding to a streamable version of the scene's geometry. One implementation of a scene may refer to a low-resolution version of the scene's geometry. Another implementation of the same scene may refer to an enhancement layer used for the low-resolution representation of the scene to add additional detail to or increase the subdivision of the same scene's geometry. As described above, each scene may have one or more enhancement layers to progressively increase the detail of the geometric objects within the scene. Each layer of media objects referenced within a scene may be associated with a token (e.g., a Uniform Resource Identifier (URI)) pointing to an address within the network from which a resource can be accessed. Such a resource is analogous to a Content Delivery Network (CDN) from which a client can obtain content. Tokens used to represent geometric objects can point to locations within the network or within the client. That is, a client can signal to the network that its resources are available for network-based media processing.

[0110] According to some implementations, a scene (temporally or non-temporally) can correspond to a scene graph as a multi-plane image (MPI) or a multi-spherical image (MSI). Both MPI and MSI techniques are examples of technologies that facilitate the creation of display-agnostic scene representations for natural content (i.e., images of the real world captured simultaneously by one or more camera devices). On the other hand, scene graph techniques can be employed to represent both natural and computer-generated images in the form of synthetic representations. However, creating such a representation is particularly computationally intensive when the content is captured as a natural scene by one or more camera devices. Creating a scene graph representation of naturally captured content is both temporally and computationally intensive, thus requiring complex analysis of the natural image using photogrammetry or deep learning techniques, or both, to create a synthetic representation that can then be used to interpolate a sufficient number of views to fill the viewing frustum of the target immersive client display. Therefore, considering such a synthetic representation as a candidate for representing natural content is impractical, as natural content cannot actually be created in real time when considering use cases requiring real-time distribution. Thus, the optimal representation for computer-generated images is achieved by using a scene graph along with a synthetic model, since computer-generated images are created using 3D modeling processes and tools, and therefore using a scene graph along with a synthetic model yields the best representation for computer-generated images.

[0111] Figure 5 An example of a natural media compositing process 500 according to some implementations is shown. The natural media compositing process 500 converts an ingestion format from a natural scene into a representation of an ingestion format that can be used as a service for heterogeneous client endpoints on a network. To the left of the dashed line 510 is the content capture portion of the natural media compositing process 500. To the right of the dashed line 510 is the ingestion format compositing (for natural images) of the natural media compositing process 500.

[0112] like Figure 5 As shown, the first camera device 501 uses a single camera lens to capture, for example, a person (i.e., Figure 5 The scene depicted (the actors shown). The second camera device 502 captures the scene with five divergent fields of view by mounting five camera lenses around the circular object. Figure 5The arrangement of the second camera device 502 shown is an exemplary arrangement typically used for capturing omnidirectional content for VR applications. The third camera device 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter portion of the sphere. The arrangement of the third camera device 503 is an exemplary arrangement typically used for capturing light fields or the light field of a holographic immersive display. The implementation is not limited to this. Figure 5 The configuration shown is illustrated. The second camera device 502 and the third camera device 503 may include multiple camera lenses.

[0113] Natural image content 509 is output from the first camera device 501, the second camera device 502, and the third camera device 503 and used as input to the synthesizer 504. The synthesizer 504 may employ neural network training 505, which uses a set of training images 506 to generate a capture neural network model 508. The training images 506 may be predefined or stored through previous synthesis processing. The neural network model (e.g., the capture neural network model 508) is a set of tensors (e.g., matrices) defining weights (i.e., numerical values) and parameters used in well-defined mathematical operations applied to the visual signal to obtain an improved visual output, which may include interpolations of new views of visual signals not explicitly provided by the original signal.

[0114] In some implementations, a photogrammetric process can be performed instead of NN training 505. If the capture NN model 508 is created during the natural media synthesis process 500, then the capture NN model 508 becomes one of the assets of the natural media content in an ingestion format 507. For example, the ingestion format 507 can be MPI or MSI. The ingestion format 507 may also include media assets.

[0115] Figure 6 An example of a synthetic media ingestion creation process 600 according to some embodiments is shown. The synthetic media ingestion creation process 600 creates an ingestion media format for synthetic media such as computer-generated images.

[0116] like Figure 6As shown, camera device 601 can capture point cloud 602 of the scene. For example, camera device 601 can be a LiDAR camera device. Computer 603 uses, for example, Common Gateway Interface (CGI) tools, 3D modeling tools, or other animation processes to create composite content (i.e., a representation of the composite scene, which can be used as an ingestion format for a network of heterogeneous client endpoint services). Computer 603 can create CGI assets 604 over the network. Additionally, sensor 605A can be worn on actor 605 in the scene. Sensor 605A can be, for example, a motion capture suit with attached sensors. Sensor 605A captures digital records of the actor 605's motion to generate animation motion data 606 (or MoCap data). Data from point cloud 602, CGI assets 604, and motion data 606 are provided as input to compositor 607 to create composite media ingestion format 608. In some embodiments, compositor 607 can use a neural network (NN) and training data to create an NN model to generate composite media ingestion format 608.

[0117] Both natural content and computer-generated (i.e., composite) content can be stored in containers. Containers can include serialization formats used to store and exchange information representing entirely natural scenes, entirely composite scenes, or a mixture of composite and natural scenes, including scene graphs and all media resources required to render the scene. The content serialization process involves converting data structures or object states into a format that can be stored (e.g., stored in a file or storage buffer) or transmitted (e.g., transmitted over a network link) and subsequently reconstructed in the same or different computing environments. When the resulting bit sequence is reread according to the serialization format, that bit sequence can be used to create an object that is semantically identical to the original object.

[0118] The dichotomy between the optimal representations of natural content and computer-generated (i.e., synthetic) content suggests that the optimal ingestion format for naturally captured content differs from the optimal ingestion format for computer-generated content or natural content that is not necessary for real-time distribution applications. Therefore, according to some implementations, the goal of the network is to be robust enough to support multiple ingestion formats for visually immersive media, regardless of whether these ingestion formats are created naturally using, for example, physical camera devices or by computer.

[0119] Techniques such as OTOY's ORBX, Pixar's Universal Scene Description, and the Graphical Language Transport Format 2.0 (glTF2.0) specification written by the Khronos 3D group represent scene graphs in a format suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content that uses deep learning or photogrammetry techniques to create corresponding synthetic representations of natural scenes (i.e., not required for real-time distribution applications).

[0120] OTOY's ORBX is one of several scene graph technologies that supports any type of temporal or non-temporal visual media, including ray-following visual formats, traditional (frame-based) visual formats, stereoscopic visual formats, and other types of composite or vector-based visual formats. ORBX differs from other scene graphs because it provides native support for freely available and / or open-source formats of meshes, point clouds, and textures. ORBX is a scene graph intentionally designed to facilitate exchange between various vendor technologies that manipulate scene graphs. Furthermore, ORBX offers a rich material system, support for open shader languages, robust camera systems, and Lua scripting. ORBX is also the foundation for an immersive technology media format licensed royalty-free by the Immersive Digital Experiences Alliance (IDEA). In the context of real-time media distribution, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of data captured by camera devices and to composite this data into synthetic representations.

[0121] Pixar's USD is a widely used scene graph for visual effects and professional content creation. USD is integrated into Nvidia's Omniverse platform, a suite of tools for developers to create and render 3D models using Nvidia's Graphics Processing Unit (GPU). A subset of USD released by Apple and Pixar is called USDZ, which is supported by Apple's ARKit.

[0122] glTF 2.0 is a version of the graphics language transport format specification written by the Khronos 3D group. This format supports simple scene graph formats, typically capable of supporting static (non-temporal) objects in a scene, including PNG and JPEG image formats. glTF 2.0 supports simple animations, including translation, rotation, and scaling of basic shapes (i.e., geometric objects) described using glTF primitives. glTF 2.0 does not support temporal media and therefore does not support either video or audio media input.

[0123] These designs for scene representations of immersive visual media are provided as examples only and do not limit the ability of the disclosed subject matter to specify a process for adapting an input immersive media source to a format suitable for the specific characteristics of a client endpoint device. Furthermore, any or all of the aforementioned example media representations employ or can employ deep learning techniques to train and create neural network models that enable or facilitate the selection of a specific view to fill a specific view frustum based on the specific size of that frustum for a particular display. The view selected for a particular display frustum can be interpolated based on existing views explicitly provided in the scene representation (e.g., according to MSI or MPI techniques). Views can also be rendered directly from rendering engines based on specific virtual camera locations, filters, or descriptions of virtual camera devices used for rendering engines.

[0124] The methods and apparatus of this disclosure are robust enough to account for the existence of a relatively small but well-known set of immersive media capture formats that are adequate for the real-time or on-demand (e.g., non-real-time) distribution of media captured naturally (e.g., using one or more camera devices) or created using computer-generated techniques.

[0125] The deployment of advanced networking technologies (e.g., 5G for mobile networks) and fiber optic cables in fixed networks has further facilitated the interpolation of views based on immersive media ingestion formats using neural network models or network-based rendering engines. These advanced networking technologies enhance the capacity and capabilities of commercial networks, as such advanced network infrastructures can support the transmission and delivery of increasingly large volumes of visual information. Network infrastructure management technologies such as Multi-access Edge Computing (MEC), Software Defined Networking (SDN), and Network Functions Virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands for certain network resources, such as dynamic increases or decreases in demand for network throughput, network speed, round-trip latency, and computing resources. Furthermore, this inherent ability to adapt to dynamic network demands also facilitates the network's ability to adapt immersive media ingestion formats to suitable distribution formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.

[0126] Immersive media applications can also have different requirements for network resources. These include: gaming applications that require significantly lower network latency to respond to real-time updates in the game state; telepresence applications that have symmetrical throughput requirements for both the uplink and downlink portions of the network; and passive viewing applications that can increase their downlink resource requirements depending on the type of client endpoint display consuming data. Typically, any consumer-facing application can be supported by a variety of client endpoints that have various onboard client capabilities for storage, computation, and power, and also have various requirements for specific media representations.

[0127] Therefore, embodiments of this disclosure enable well-equipped networks, i.e., networks employing some or all of the characteristics of modern networks, to simultaneously support multiple legacy devices and devices with immersive media capabilities, based on specified features within the device. Thus, the immersive media distribution methods and processes described herein provide the flexibility to distribute media using media ingestion formats applicable to both real-time and on-demand use cases, the flexibility to simultaneously support both natural and computer-generated content for both legacy and immersive media-capable client endpoints, and support for both temporal and non-temporal media. The methods and processes also dynamically adapt the source media ingestion format to a suitable distribution format based on the characteristics and capabilities of the client endpoints and the requirements of the application. This ensures that the distribution format is streamable over IP-based networks and enables the network to simultaneously serve multiple heterogeneous client endpoints, including legacy devices and devices with immersive media capabilities. Furthermore, embodiments provide exemplary media representation frameworks that facilitate the organization of media distribution along scene boundaries.

[0128] according to Figures 7 to 14B The processes and components described in the detailed description implement an end-to-end implementation of the improved heterogeneous immersive media distribution described above, according to embodiments of this disclosure, which are further described in detail below.

[0129] The aforementioned techniques for representing and streaming heterogeneous immersive media can be implemented in both the source and destination as computer software using computer-readable instructions and physically stored in one or more non-transitory computer-readable media, or through one or more specially configured hardware processors. Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0130] Computer software can be coded using any suitable machine code or computer language, which can be subjected to assembly, compilation, linking or similar mechanisms to create code that includes instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.

[0131] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0132] Figure 7The components shown for computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any of the components shown or a combination thereof in the exemplary embodiments of computer system 700.

[0133] Computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input from one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), and olfactory input. The human-machine interface device can also be used to capture certain media that are not necessarily directly related to human conscious input: said conscious input includes, for example, audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0134] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard 701, touchpad 702, mouse 703, screen 709 which may be, for example, a touch screen, data glove, joystick 704, microphone 705, camera device 706, and scanner 707.

[0135] Computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback via screen 709, data gloves, or joystick 704, but tactile feedback devices that are not used as input devices may also exist), audio output devices (e.g., speaker 708, headphones), visual output devices (e.g., screen 709, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which are capable of outputting two-dimensional or more than three-dimensional visual output via devices such as stereoscopic image output; virtual reality glasses, holographic displays, and smoke boxes), and printers.

[0136] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 711 with media such as CD / DVD 710, thumb drives 712, removable hard disk drives or solid-state drives 713, conventional magnetic media such as magnetic tapes and floppy disks, and devices based on dedicated ROM / ASIC / PLD such as security dongles.

[0137] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0138] Computer system 700 may also include interfaces 715 to one or more communication networks 714. For example, network 714 may be wireless, wired, or optical. Network 714 may also be local area, wide area, metropolitan area, vehicle-mounted and industrial, real-time, latency-tolerant, etc. Examples of network 714 include: local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle-mounted and industrial networks including CANBus, etc. Some networks 714 typically require external network interface adapters (e.g., graphics adapter 725) attached to certain general-purpose data ports or peripheral buses 716 (e.g., USB ports of computer system 700); other networks are typically integrated into the core of computer system 700 via system buses 748 as described below (e.g., integrated into a PC computer system via an Ethernet interface or integrated into a smartphone computer system via a cellular network interface). Using any of these networks 714, computer system 700 can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., to a CANbus device), or bidirectional (e.g., to other computer systems via a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0139] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core 717 of the computer system 700.

[0140] The core 717 may include one or more Central Processing Units (CPUs) 718, Graphics Processing Units (GPUs) 719, dedicated programmable processing units in the form of Field Programmable Gate Arrays (FPGAs) 720, hardware accelerators 721 for certain tasks, etc. These devices, along with read-only memory (ROM) 723, random-access memory (RAM) 724, and internal mass storage devices (such as internal non-user-accessible hard disk drives, SSDs, etc.) 722, can be connected via the system bus 748. In some computer systems, the system bus 748 can be accessed via one or more physical connectors to allow for expansion through additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the core's system bus 748, or attached to the core's system bus 748 via peripheral buses 716. Peripheral bus architectures include PCI, USB, etc.

[0141] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can execute certain instructions that, when combined, constitute the aforementioned machine code (or computer code). This computer code can be stored in ROM 723 or RAM 724. Transient data can also be stored in RAM 724, while permanent data can be stored, for example, in an internal mass storage device 722. Fast storage and retrieval of any memory device within the memory apparatus can be achieved using a cache memory, which can be closely associated with one or more CPUs 718, GPUs 719, mass storage devices 722, ROM 723, RAM 724, etc.

[0142] Computer-readable media may contain computer code for performing various computer-implemented operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0143] By way of example and not limitation, a computer system having a computer system 700 architecture, and in particular a core 717, can be functionally provided by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with the user-accessible mass storage devices described above, as well as certain storage devices of the core 717 having a non-transitory nature, such as the internal mass storage device 722 or ROM 723. Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core 717. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core 717, and in particular the processors therein (including CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 724 and modifying such data structures according to the software-defined processes. Alternatively or as an alternative, the computer system may be provided with functionality by hardwired or otherwise embodied logic in circuitry (e.g., accelerator 721), which may replace or operate with software to perform the specific processing or a specific portion of the specific processing described herein. Where appropriate, references to software may include logic, and conversely, references to logic may include software. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits, ICs) storing software for execution, circuitry implementing logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0144] Figure 7 The number and arrangement of components shown are provided as examples. In practice, input human-machine interface devices can be compared to... Figure 7 The components shown include additional components, fewer components, different components, or components arranged differently. Alternatively or additionally, a set of components (e.g., one or more components) of the input human-machine interface device may perform one or more functions described as being performed by another set of components of the input human-machine interface device.

[0145] In some implementations... Figures 1A to 6 and Figures 8 to 14A Any operation or process in the operation or process can be or used Figure 7 Implemented by any of the components shown.

[0146] Figure 8An exemplary network media distribution system 800 serving multiple heterogeneous client endpoints is illustrated. Specifically, system 800 supports various conventional displays and displays with heterogeneous immersive media capabilities as client endpoints. System 800 may include a content acquisition module 801, a content preparation module 802, and a transmission module 803.

[0147] Content acquisition module 801 uses, for example Figure 6 and / or Figure 5 The implementation described herein is used to capture or create source media. Content preparation module 802 creates an ingestion format, which is then transmitted to a network media distribution system using transmission module 803. Gateway 804 can serve as a customer premises device to provide network access to various client endpoints. Set-top box 805 can also be used as a customer premises device to provide network service providers with access to aggregated content. Radio demodulator 806 can be used as a mobile network access point for mobile devices, such as the mobile handset display 813 shown. In this particular implementation of system 800, a conventional 2D television 807 is shown as directly connected to one of gateway 804, set-top box 805, or WiFi (router) 808. A laptop computer 2D display 809 (i.e., a computer or laptop computer with a conventional 2D display) is shown as a client endpoint connected to WiFi (router) 808. Head-mounted 2D (grid-based) display 810 is also connected to WiFi (router) 808. A lenticular light field display 811 is shown connected to one of the gateways 804. The lenticular light field display 811 may include one or more GPUs 811A, memory 811B, and a visual presentation component 811C that creates multiple views using light-based lenticular optics. A holographic display 812 is shown connected to a set-top box 805. The holographic display 812 may include one or more CPUs 812A, GPUs 812B, storage devices 812C, and a visualization component 812D. The visualization component 812D may be a Fresnel pattern, a wave-based holographic device / display. An augmented reality (AR) headset 814 is shown connected to a radio demodulator 806. The AR headset 814 may include a GPU 814B, memory 814C, a battery 814A, and a stereoscopic visual presentation component 814D. A dense light field display 815 is shown connected to a WiFi (router) 808. The dense light field display 815 may include one or more GPUs 815A, CPUs 815B, memory 815C, eye-tracking devices 815D, camera devices 815E, and a dense light field panel 815F.

[0148] Figure 8The number and arrangement of components shown are provided as examples. In practice, with Figure 8 Compared to the components shown, system 800 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of system 800 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of the device or corresponding display.

[0149] Figure 9 An exemplary workflow of an immersive media distribution process 900 is shown, which is capable of serving applications such as those previously described in [the document / document / etc.]. Figure 8 The text describes both traditional displays and displays with heterogeneous immersive media capabilities. For example, it mentions adapting media on a network for consumption by specific immersive media client endpoints (as seen in the reference). Figure 10 Prior to the process described, the immersive media distribution process 900 performed by the network can provide adaptation information about a specific media represented in a media ingestion format.

[0150] The immersive media distribution process 900 can be divided into two parts: immersive media production to the left of the dotted line 912 and immersive network distribution to the right of the dotted line 912. Immersive media production and immersive media network distribution can be performed by a network or client device.

[0151] Media content 901 is created by a network (or client device) or obtained from a content source. The methods used to create or obtain the data may correspond to those for natural content and synthetic content, respectively. Figure 5 and Figure 6 Then, the created content 901 is converted into an ingestion format using the web ingestion format creation process 902. The web ingestion format creation process 902 can also be adapted for natural content and synthetic content respectively. Figure 5 and Figure 6 It can also update the ingestion format to store data from, for example, the Media Reuse Analyzer 911 (see later). Figure 10 and Figure 14A (Detailed description) Information about assets that can potentially be reused in multiple scenarios. The ingested format is transmitted to the network and stored in ingested media storage 903 (i.e., storage device). In some embodiments, the storage device may be located in the network of the immersive media content creator and may be remotely accessed for immersive media network distribution. Client and application-specific information may optionally be available in the remote storage device, i.e., client-specific information 904. In some embodiments, client-specific information 904 may exist remotely in an alternative cloud network and may be transmitted to the network.

[0152] Then, the network orchestrator 905 is executed. The network orchestrator serves as the primary information source and sink for performing the network's main tasks. The network orchestrator 905 can be implemented in a format consistent with other components of the network. The network orchestrator 905 can also employ a bidirectional messaging protocol with client devices to facilitate all processing and distribution of media according to the characteristics of the client devices. Furthermore, this bidirectional protocol can be implemented across different delivery channels (e.g., control plane channels and / or data plane channels).

[0153] like Figure 9 As shown, network orchestrator 905 receives information about the characteristics and attributes of client device 908. Network orchestrator 905 collects requests regarding applications currently running on client device 908. This information can be obtained from client-specific information 904. In some embodiments, this information can be obtained by directly querying client device 908. In the case of direct querying of the client device, it is assumed that a bidirectional protocol exists and is operational, allowing client device 908 to communicate directly with network orchestrator 905.

[0154] The network orchestrator 905 can also initiate and communicate with the media adapter and segmentation module 910 (this is in Figure 10 (As described in the text). When the ingested media is adapted and segmented by the media adaptation and segmentation module 910, the media can be transmitted to an intermediate storage device, such as media 909 ready for distribution. If the network is designed to include a cache for assets that are used multiple times in the presentation context, another intermediate storage device, a redundant cache 912 for reusing media assets, can be used as a cache for such assets. When the distribution media is ready and the media 909 ready for distribution is stored in the storage device, the network orchestrator 905 ensures that the client device 908 receives the distribution media and descriptive information 906 via a "push" request, or the client device 908 can initiate a "pull" request for the distribution media and descriptive information 906 based on the stored media 909 ready for distribution. This information can be "pushed" or "pulled" via the network interface 908B of the client device 908. The "pushed" or "pulled" distribution media and descriptive information 906 can be descriptive information corresponding to the distribution media.

[0155] In some implementations, the network orchestrator 905 uses a bidirectional messaging interface to execute "push" requests or initiate "pull" requests from the client device 908. The client device 908 may optionally employ a GPU 908C (or a CPU).

[0156] The media format is then stored in a storage device or storage cache 908D included in the client device 908. Finally, the client device 908 visually presents the media via a visualization component 908A.

[0157] Throughout the streaming of immersive media to client device 908, network orchestrator 905 detects the client's progress status via client progress and status feedback channel 907. In some embodiments, the detection of this status may be performed via a bidirectional communication message interface.

[0158] Figure 10 An example of a media adaptation process 1000 performed, for example, by a media adaptation and segmentation module 910, is shown. By performing the media adaptation process 1000, the ingested source media can be appropriately adapted to match the requirements of a client (e.g., client device 908).

[0159] like Figure 10 As shown, the media adaptation process 1000 includes multiple components that help adapt the ingested media into an appropriate distribution format for the client device 908. Figure 10 The components shown should be considered exemplary. In practice, with Figure 10 Compared to the components shown, the media adaptation process 1000 may include additional components, fewer components, different components, or components with different arrangements. Alternatively or additionally, a set of components of the media adaptation process 1000 (e.g., one or more components) may perform one or more functions described as being performed by another set of components.

[0160] exist Figure 10 In this process, the adaptation module 1001 receives input from the network status 1005 to detect the current traffic load on the network. As described above, the adaptation module 1001 also receives information from the network orchestrator 905. This information may include attribute and feature descriptions of the client device 908, application features and descriptions, the current state of the application, and the client NN model (if available) to help map the geometry of the client's frustum to the interpolation capabilities of the ingested immersive media. Such information can be obtained through a bidirectional messaging interface. The adaptation module 1001 ensures that the adapted output is stored in a storage device for storing the client adaptation media 1006 when it is created.

[0161] The media reuse analyzer 911 can be an optional process that can be executed before or as part of an automated network process for distributing media. The media reuse analyzer 911 can store the ingested media format and assets in storage device 1002. The ingested media format and assets can then be transferred from storage device 1002 to adapter module 1001.

[0162] The adaptation module 1001 can be controlled by a logic controller 1001F. The adaptation module 1001 can also employ a renderer 1001B or a processor 1001C to adapt a specific ingested source media into a format suitable for the client. The processor 1001C can be a neural network-based processor. The processor 1001C uses a neural network model 1001A. Examples of such a processor 1001C include the Deepview neural network model generator as described in MPI and MSI. If the media is in 2D format, but the client must have a 3D format, the processor 1001C can invoke a process that uses a highly correlated image from the 2D video signal to obtain a stereoscopic representation of the scene depicted in the media.

[0163] Renderer 1001B can be a software-based (or hardware-based) application or process based on a selective hybrid of disciplines related to: acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development. That is, given an input scene graph and asset container, renderer 1001B emits visual and / or audio signals suitable for rendering on a target device or conforming to the desired characteristics specified by the attributes of the rendering target nodes in the scene graph. For visual-based media assets, renderer can emit visual signals suitable for the target display or suitable for storage as intermediate assets (e.g., repackaged into another container and used in a series of rendering processes in the graphics pipeline). For audio-based media assets, renderer can emit audio signals for rendering in multi-channel speakers and / or stereo headphones or for repackaging into another (output) container. For example, renderer includes real-time rendering features of source and cross-platform game engines. Renderer can include scripting languages ​​(i.e., interpreted programming languages) that can be executed by renderer at runtime to handle dynamic input to scene graph nodes and variable state changes made to scene graph nodes. The dynamic inputs and variable state changes can affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse motion, deformation, and collisions), as well as energy propagation and transmission (light, sound). The evaluation of spatial and temporal object topology produces results that transform the output from an abstract outcome to a concrete one (e.g., similar to the evaluation of a document object model for a webpage).

[0164] For example, renderer 1001B may be a modified version of the OTOY Octane renderer that interacts directly with adapter module 1001. In some implementations, renderer 1001B implements computer graphics methods (e.g., path following) to render a 3D scene so that the scene's lighting is faithful to reality. In some implementations, renderer 1001B may employ shaders (i.e., computer programs originally used for shading (producing appropriate levels of light, dark, and color within an image), but now performing various specialized functions in various domains such as computer graphics effects, video post-processing unrelated to shading, and other functions unrelated to graphics).

[0165] Depending on the compression and decompression requirements based on the format of the ingested media and the format required by the client device 908, the adaptation module 1001 can use a media compressor 1001D and a media decompressor 1001E to perform compression and decompression of the media content, respectively. The media compressor 1001D can be a media encoder, and the media decompressor 1001E can be a media decoder. After performing compression and decompression (if necessary), the adaptation module 1001 outputs the client-adapted media 1006, optimal for streaming or distribution, to the client device 908. The client-adapted media 1006 can be stored in a storage device for storing the adaptation media.

[0166] Figure 11 An exemplary distribution format creation process 1100 is illustrated. For example... Figure 11 As shown, the distribution format creation process 1100 includes a media adaptation module 1101 and an adaptation media packaging module 1103. The adaptation media packaging module 1103 packages the media output from the media adaptation process 1000 and stores it as client-side adaptation media 1006. The media packaging module 1103 formats the adaptation media in the client-side adaptation media 1006 into a robust distribution format 1104. For example, the distribution format could be... Figure 3 or Figure 4 The exemplary format shown is illustrated. Information list 1104A can provide scene data asset list 1104B to client device 908. Scene data asset list 1104B may also include metadata describing the frequency with which each asset is used in the set of scenes included in the presentation. Scene data asset list 1104B depicts a list of visual assets, an audio asset list, and a haptic asset list, each with its corresponding metadata. In this exemplary embodiment, each asset in scene data asset list 1104B references metadata containing a digital frequency value indicating the number of times a particular asset is used in all scenes including the presentation.

[0167] The media can be packaged separately before streaming. Figure 12An exemplary packaging process 1200 is illustrated. The packaging system 1200 includes a packer 1202. The packer 1202 can receive a scene data asset list 1104B (or 11044B) as input media 1201 (such as...). Figure 12 (As shown). In some embodiments, client-adapted media 1006 or distribution format 1104 is input to packetizer 1202. Packetizer 1202 separates the input media 1201 into individual packets 1203 suitable for presentation and streaming to client device 908 on the network.

[0168] Figure 13 This is a sequence diagram illustrating examples of data and communication flows between components according to some implementation methods. Figure 13 The sequence diagram is how a network adapts a specific immersive media ingested in a particular format to a suitable streaming and distribution format for a specific immersive media client endpoint. Data and communication flows can be as follows.

[0169] Client device 908 initiates media request 1308 to network orchestrator 905. In some implementations, the request may be made to the network distribution interface of the client device. Media request 1308 includes information identifying the media requested by client device 908. Media request can be identified by, for example, a Uniform Resource Name (URN) or other standard naming conventions. Network orchestrator 905 then responds to media request 1308 with profile request 1309. Profile request 1309 requests the client to provide information about currently available resources (including compute, storage, battery charge percentage, and other information characterizing the client's current operating state). Profile request 1309 also requests the client to provide one or more NN models, which, if available at the client endpoint, the network can use for NN inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system.

[0170] Then, client device 908 provides network orchestrator 905 with a response 1310, which is provided as a client token, an application token, and one or more NN model tokens (if such NN model tokens are available at the client endpoint). Network orchestrator 905 then provides session ID tokens 1311 to client device 908. Network orchestrator 905 then requests ingest media 1312 from ingest media server 1303. Ingest media server 1303 may include, for example, ingest media storage 903 or ingest media format and asset storage device 1002. The request for ingest media 1312 may also include the URN or other standard name of the media identified in request 1308. Ingest media server 1303 responds to ingest media 1312 request with a response 1313 including an ingest media token. Network orchestrator 905 then provides the media token from response 1313 to client device 908 in call 1314. Then, the network orchestrator 905 initiates the adaptation process for the media requested in request 1315 by providing the adaptation and segmentation module 910 with an ingest media token, a client token, an application token, and a NN model token. The adaptation and segmentation module 910 requests access to the ingested media by providing the ingest media server 1303 with an ingest media token at request 1316 to request access to the ingested media asset.

[0171] In response 1317 to the adaptation and segmentation module 910, the ingest media server 1303 responds to request 1316 with an ingest media access token. The adaptation and segmentation module 910 then requests the media adaptation process 1000 to adapt the ingested media located at the ingest media access token for the client, application, and NN inference model corresponding to the session ID token created and transmitted at response 1313. Request 1318 is issued from the adaptation and segmentation module 910 to the media adaptation process 1000. Request 1318 includes the required token and session ID. The media adaptation process 1000 provides the adapted media access token and session ID to the network orchestrator 905 in an updated response 1319. The network orchestrator 905 then provides the adapted media access token and session ID to the media packaging module 11043 in interface call 1320. The media packaging module 11043 provides a response 1321 to the network orchestrator 905, in which a media access token and a session ID are packaged. Then, in response 1322, the media packaging module 11043 provides the packaged asset, URN, and packaged media access token for the session ID to the packaging media server 1107 for storage. Subsequently, the client device 908 executes a request 1323 to the packaging media server 1307 to initiate streaming of the media asset corresponding to the packaged media access token received in response 1321. Finally, the client device 908 executes other requests and provides a status update to the network orchestrator 905 in message 1324.

[0172] Figure 14A Showing the target Figure 9 The workflow of the media reuse analyzer 911 is shown. The media reuse analyzer 911 analyzes metadata related to the uniqueness of objects included in the scene in the media data.

[0173] At 1405, media data is obtained from, for example, a content provider or content source. At 1410, initialization is performed. Specifically, the iterator "i" is initialized to zero. The iterator can be, for example, a counter. A unique list of assets for each scene is also initialized at 1465 (e.g., ...). Figure 14B As shown), it identifies the unique asset encountered in all scenes, including those presented (such as...). Figure 3 and / or Figure 4 (As shown).

[0174] At 1415, it is determined whether the value of iterator "i" is less than the total number N of scenes included in the presentation. If the value of iterator "i" is equal to (or greater than) the number N of scenes included in the presentation ("No" at 1415), the process proceeds to 1420, where the reuse analysis terminates (i.e., the process ends). If the value of iterator "i" is less than the number N of scenes included in the presentation ("Yes" at 1415), the process proceeds to 1425. At 1425, the value of iterator "j" is set to zero.

[0175] Then, at 1430, it is determined whether the value of iterator "j" is less than the total number Xs of media assets (also known as media objects) in the current scene. If the value of iterator "j" is equal to (or greater than) the total number Xs of media assets for scene s ("No" at 1430), the process proceeds to 1410, where iterator "i" is incremented by 1 before returning to 1410. If the value of iterator "j" is less than the total number Xs of media assets for scene s ("Yes" at 1430), the process proceeds to 1440.

[0176] At 1440, the characteristics of the media asset are compared with the characteristics of the asset previously analyzed based on the current scenario (i.e., scenario s) to determine whether the current media asset has been used before.

[0177] If the current media asset is not identified as a unique asset ("No" at 1440), meaning it has not been previously analyzed in a scenario associated with a smaller value of iterator "i", then processing proceeds to 1445. At 1445, a unique asset entry is created in the unique asset list set 1465 corresponding to the current scenario (i.e., scenario s). A unique identifier is also assigned to this unique asset entry, and the number of times the asset has been used between scenario 0 and scenario N-1 (e.g., frequency) is set to 1. Processing then proceeds to 1455.

[0178] If the current media asset has already been identified as an asset used in one or more previous scenarios ("Yes" at 1440), then the process proceeds to 1450. At 1450, in the unique asset list set 1465 corresponding to the current scenario (i.e., scenario s), the number of times the current media asset has been used between scenario 0 and scenario N-1 is incremented by 1. Then, the process proceeds to 1455.

[0179] At 1455, increment the value of iterator "j" by 1. Then, the process returns to 1430.

[0180] In some implementations, the media reuse analyzer 911 may also signal to the client (e.g., client device 108) that the client should use a copy of the asset for each instance of the asset used in the scene set (after the asset is first distributed to the client).

[0181] Note, refer to Figures 13 to 14A The steps in the described sequence diagrams and workflows are not intended to limit the configuration of data and communication flows in the implementation. For example, one or more of these steps may be performed simultaneously, and data may be stored and / or flow along with... Figures 13 to 14A The direction of flow is not explicitly shown in the process.

[0182] Figure 14B This is an example of a unique asset list set 1465 that, according to some implementations, is initialized at 1410 for all scenes upon completion of rendering (and potentially updated at 1445-1450). The unique asset list in the unique asset list set 1465 may be identified a priori or predefined by the network or client device. The unique asset list set 1465 illustrates an example list of information entries describing assets that are unique relative to the entire rendering, including: an indicator of the type of media (e.g., grid, audio, or volume) of the asset, a unique identifier for the asset, and the number of times the asset is used in the set of scenes encompassing the entire rendering. For example, for scene N-1, the asset is not included in its list because all assets required for scene N-1 have already been identified as assets also used in scenes 1 and 2.

[0183] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it will be understood that those skilled in the art will be able to design various systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

Claims

1. A method for streaming media assets, characterized in that, The method includes: An immersive media stream is received by a streaming media server, the immersive media stream comprising one or more immersive media assets associated with one or more scenes; Determine that a subset of the one or more immersive media assets is included multiple times in the one or more scenarios; A redundant copy of each immersive media asset in the subset is stored in a redundant cache maintained by the streaming network, enabling both the streaming server and the client to access each immersive media asset in the subset; and In response to the fact that the client's local cache does not store at least one media asset in the subset, the at least one media asset is streamed, wherein the redundant cache is different from the client's local cache; The network first queries the client device to obtain feedback to ensure that the requested media asset is still available in the client's local cache. If the client device's response indicates that it no longer has a copy of the requested media asset, the network signals the client to access a copy of the media asset presenting its distribution format from the redundant cache.

2. The method according to claim 1, characterized in that, The method further includes: The streaming media network determines whether the media assets have been streamed to the client device; If it is determined that the media asset has been streamed to the client device, the streaming media network determines whether a copy of the media asset is stored in a local cache managed by the client device. If it is determined that a copy of the media asset is not stored in a local cache managed by the client device, the streaming media network generates a first alternative proxy for the media asset, wherein the first alternative proxy instructs the client device to use a copy of the media asset stored in the redundant cache managed by the streaming media network; and Immersive media is streamed to the client device via the streaming network, the immersive media including a first alternative proxy of the media assets but excluding the media assets.

3. The method according to claim 1, characterized in that, The method further includes: If it is determined that a copy of the media asset is stored on a local cache managed by the client device, then the streaming media network generates a second alternative proxy for the media asset, wherein the second alternative proxy instructs the client device to use a copy of the media asset stored on a local cache managed by the client device; and Immersive media is streamed to the client device via the streaming network, wherein the immersive media includes a second alternative proxy of the media assets but does not include the media assets.

4. The method according to claim 1, characterized in that, The method further includes: If it is determined that the media asset has not yet been streamed to the client device, the streaming media network determines one or more format conversions to be performed on the media asset. The streaming media network performs the one or more format conversions on the media assets; The streaming media network will stream immersive media, including converted media assets, to the client device; and If it is determined that the converted media asset is to be reused in the streaming session, the streaming network generates a copy of the media asset to be stored on the redundant cache managed by the streaming network.

5. The method according to claim 4, characterized in that, Generating a copy of the media asset to be stored includes generating the converted media asset.

6. The method according to claim 4, characterized in that, The method further includes: The converted media assets are stored on a redundant cache managed by the streaming media network.

7. The method according to claim 1, characterized in that, The method further includes: If it is determined that the media asset has not yet been streamed to the client device, then the streaming media network will stream immersive media including the media asset to the client device; and The streaming media network stores copies of the already streamed media assets on a redundant cache managed by the streaming media network.

8. The method according to claim 1, characterized in that, The redundant cache is different from any local cache maintained by any client device.

9. An apparatus for streaming media assets, characterized in that, The device includes: The receiving module is configured to receive an immersive media stream via a streaming media server, the immersive media stream comprising one or more immersive media assets associated with one or more scenes; The determination module is configured to determine that a subset of the one or more immersive media assets is included multiple times in the one or more scenes; A storage module is configured to store a redundant copy of each immersive media asset in the subset in a redundant cache maintained by the streaming network, enabling both the streaming server and the client to access each immersive media asset in the subset; and A transmission module is configured to stream the at least one media asset in the subset in response to the client’s local cache not storing at least one media asset in the subset, wherein the redundant cache is different from the client’s local cache. The network first queries the client device to obtain feedback to ensure that the requested media asset is still available in the client's local cache. If the client device's response indicates that it no longer has a copy of the requested media asset, the network signals the client to access a copy of the media asset presenting its distribution format from the redundant cache.

10. The apparatus according to claim 9, characterized in that, The device further includes: The streaming media network determines whether the media assets have been streamed to the client device; If it is determined that the media asset has been streamed to the client device, the streaming media network determines whether a copy of the media asset is stored in a local cache managed by the client device. If it is determined that a copy of the media asset is not stored in a local cache managed by the client device, the streaming media network generates a first alternative proxy for the media asset, wherein the first alternative proxy instructs the client device to use a copy of the media asset stored in the redundant cache managed by the streaming media network; and Immersive media is streamed to the client device via the streaming network, the immersive media including a first alternative proxy of the media assets but excluding the media assets.

11. The apparatus according to claim 9, characterized in that, The device further includes: If it is determined that a copy of the media asset is stored on a local cache managed by the client device, then the streaming media network generates a second alternative proxy for the media asset, wherein the second alternative proxy instructs the client device to use a copy of the media asset stored on a local cache managed by the client device; and Immersive media is streamed to the client device via the streaming network, wherein the immersive media includes a second alternative proxy of the media assets but does not include the media assets.

12. The apparatus according to claim 9, characterized in that, The device further includes: If it is determined that the media asset has not yet been streamed to the client device, the streaming media network determines one or more format conversions to be performed on the media asset. The streaming media network performs the one or more format conversions on the media assets; The streaming media network will stream immersive media, including converted media assets, to the client device; and If it is determined that the converted media asset is to be reused in the streaming session, the streaming network generates a copy of the media asset to be stored on the redundant cache managed by the streaming network.

13. The apparatus according to claim 12, characterized in that, Generating a copy of the media asset to be stored includes generating the converted media asset.

14. The apparatus according to claim 12, characterized in that, The device further includes: The converted media assets are stored on a redundant cache managed by the streaming media network.

15. The apparatus according to claim 9, characterized in that, The device further includes: If it is determined that the media asset has not yet been streamed to the client device, then the streaming media network will stream immersive media including the media asset to the client device; and The streaming media network stores copies of the already streamed media assets on a redundant cache managed by the streaming media network.

16. The apparatus according to claim 9, characterized in that, The redundant cache is different from any local cache maintained by any client device.

17. An apparatus for streaming media assets, characterized in that, The device includes: At least one memory configured to store computer program code; and At least one processor is configured to read the computer program code and operate as instructed by the computer program code to implement the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable medium storing instructions, characterized in that, When executed by at least one processor, the instructions cause the at least one processor to perform the method according to any one of claims 1 to 8.