Immersive Media Streaming with Priority Determined by Asset Complexity

By analyzing the complexity of media assets and sequencing them for efficient packaging and streaming, the method optimizes media delivery in media streaming networks, addressing inefficiencies related to repeated media conversion and streaming.

JP7695047B2Active Publication Date: 2025-06-18TENCENT AMERICA LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023553547
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2022-10-25
Publication Date
2025-06-18
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing media streaming networks face inefficiencies due to repeated conversion and streaming of media assets, leading to delays and increased resource usage, as they lack the ability to access information about client-side caching of media objects.

Method used

A method and system for optimizing media delivery by analyzing the complexity of media assets and ordering them in a sequence for packaging and streaming, allowing clients to access cached media objects and reducing redundant processing and transfer of data.

Benefits of technology

This approach enhances network efficiency by minimizing redundant media processing and transfer, reducing latency, and optimizing resource usage by leveraging client-side caching and intelligent asset sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695047000001
    Figure 0007695047000001
  • Figure 0007695047000002
    Figure 0007695047000002
  • Figure 0007695047000003
    Figure 0007695047000003
Patent Text Reader

Abstract

A media packaging method for optimizing media delivery in a media streaming network is provided, the method including the steps of: a media streaming server receiving an immersive media stream including one or more immersive media assets related to one or more scenes; identifying a subset of the one or more immersive media assets that include essential elements of the respective scenes; ordering the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that include essential elements of the respective scenes in the one or more scenes; and streaming the one or more immersive media assets in the ordered sequence from the media streaming server to a client device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Cross - Reference to Related Applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 276,545 filed on November 5, 2021 and U.S. Patent Application No. 17 / 971,037 filed on October 21, 2022, the disclosures of which are hereby incorporated by reference in their entirety.

[0002]

[0002] Field This disclosure generally describes embodiments related to the architecture, structure, and components of systems and networks for delivering media, including video, audio, geometric (3D) objects, tactile, associated metadata, or other content for client presentation devices. Some embodiments relate to systems, structures, and architectures for delivering media content to heterogeneous immersive and interactive client presentation devices.

Background Art

[0003]

[0003] Immersive Media generally refers to media that stimulates any or all of the human sensory systems (e.g., vision, hearing, somatosensation, smell, and possibly taste) to create or enhance the user's perception physically experienced during the media experience, i.e., beyond what is delivered over existing (e.g., "legacy") commercial networks for time - specified 2 - dimensional (2D) video and corresponding audio; such time - specified media is known as "legacy media". Immersive Media may also be defined as media that attempts to generate or simulate the physical world through digital simulation of the laws of physical motion, thereby stimulating any or all of the human sensory systems to create the user's perception physically provided within a scene depicting a real or virtual world.

[0004]

[0004] An immersive media - compatible presentation device may refer to a device with sufficient resources and functions for access, interpretation, and presentation of immersive media. Such an immersive media - compatible device supports media in multiple quantities and formats and also supports multiple network resources required for large - scale distribution of immersive media. "At scale" may refer to media distribution by a service provider that realizes distribution equivalent to that of conventional video and audio media via a network, such as Netflix, Hulu, Comcast subscribers, and Spectrum subscribers.

[0005]

[0005] In contrast, conventional presentation devices such as laptop displays, televisions, and mobile handset displays are homogenous in their functions because all of these devices include a rectangular display screen that uses 2D rectangular videos or still images as their primary visual media format. Some of the visual media formats commonly used in conventional presentation devices may include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266.

[0006]

[0006] The delivery of some media via a network may use a media delivery system and architecture that reformats and / or converts the media from an input format or network “ingest” media format to a delivery media format, where the delivery media format is suitable not only for being captured by the targeted client device and its application, but also for being “streamed” over the network. The reformatting or streaming is performed by the network (e.g., a server in a media streaming network), i.e., before delivering the media to the client, and may result in a media format called the “delivery media format” or simply the “delivery format”.

[0007]

[0007] When the network has access to information indicating that the client will need the converted media object (which may also be called a media asset) and / or the streamed media object on multiple occasions, in related technologies, such multiple uses will trigger the conversion and streaming of such media multiple times. That is, this constant reprocessing and transfer of data for media conversion and streaming causes delays in the network and may significantly increase the amount of network and / or computing resources used.

[0008] On the other hand, a network design that allows a client to access information indicating that the client may already have a particular media data object stored in a cache or stored locally for the client will perform more efficiently than a network that cannot access such information. Therefore, a network design that includes allowing a client to access information indicating that the client may have a media object stored locally in a cache may be required.

SUMMARY OF THE INVENTION

[0009]

[0009] According to an embodiment, a method, system, and apparatus are provided for facilitating the calculation of a sequence order for packaging and streaming assets from a network to a client. A media asset that includes an essential element of a scene is analyzed by a complexity analyzer to determine which asset for a particular scene will take the most time to process. The order in which the assets for a particular scene are packaged and streamed to the client is based on the complexity of each asset of the particular scene.

[0010] According to one aspect of the disclosure, it is possible to provide a media packaging method for optimizing media delivery in a media streaming network. The method includes: a media streaming server receiving an immersive media stream including one or more immersive media assets related to one or more scenes; identifying a subset of the one or more immersive media assets, the subset including essential elements of individual scenes in the one or more scenes; ordering the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes; and streaming the one or more immersive media assets from the media streaming server to a client device in the ordered sequence.

[0011] According to another aspect of the disclosure, it is possible to provide a device (or apparatus) for optimizing media delivery in a media streaming network. The device can include at least one memory configured to store computer program code; and at least one processor configured to read the computer program code and operate as directed by the computer program code. The computer program code includes: first reception code configured to cause the at least one processor to receive an immersive media stream including one or more immersive media assets related to one or more scenes; identification code configured to cause the at least one processor to identify a subset of the one or more immersive media assets, the subset including essential elements of individual scenes in the one or more scenes; ordering code configured to cause the at least one processor to order the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets including essential elements of individual scenes in the one or more scenes; and streaming code configured to cause the at least one processor to stream the one or more immersive media assets from a media streaming server to a client device in the ordered sequence.

[0012] According to another aspect of the disclosure, it is possible to provide a non-transitory computer-readable storage medium storing instructions which, when executed by a processor of a device for optimizing media delivery in a media streaming network, cause the at least one processor to: receive an immersive media stream including one or more immersive media assets associated with one or more scenes; identify a subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes; order the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes; and stream the one or more immersive media assets from a media streaming server to a client device in the ordered sequence.

[0013] Additional embodiments will be described in the following description, will become apparent to some extent from the description, and / or may be realized by practicing the presented embodiments of the disclosure.

Brief Description of the Drawings

[0014]

Figure 1A

Figure 1B

[0015] FIG. 1B is an exemplary workflow showing the generation of reuse indicators in a media streaming network according to an embodiment and the creation of media in a delivery format.

Figure 2A

[0016] FIG. 2A is an exemplary workflow for streaming media to a client device according to an embodiment.

Figure 2B

[0017] Figure 2B is an exemplary workflow for streaming media to a client device according to an embodiment.

Figure 3A

[0018] Figure 3A is an exemplary diagram of a data model related to the representation and streaming of time - stamped immersive media according to an embodiment.

Figure 3B

[0019] Figure 3B is an exemplary diagram of a data model related to the representation and streaming of time - stamped immersive media according to an embodiment.

Figure 4A

[0020] Figure 4A is an exemplary diagram of a data model related to the representation and streaming of non - time - stamped immersive media according to an embodiment.

Figure 4B

[0021] Figure 4B is an exemplary diagram of a data model related to the representation and streaming of non - time - stamped immersive media according to an embodiment.

Figure 5

[0022] Figure 5 is an exemplary workflow illustrating natural media synthesis according to an embodiment.

Figure 6

[0023] Figure 6 is an exemplary workflow showing the creation of synthetic media capture according to an embodiment.

Figure 7

[0024] Figure 7 is an exemplary diagram of a computer system according to an embodiment.

Figure 8

[0025] Figure 8 is an exemplary diagram of a network media distribution system according to an embodiment.

Figure 9

[0026] Figure 9 is an exemplary workflow illustrating immersive media delivery according to an embodiment.

Figure 10

[0027] Figure 10 is a system diagram of a media adaptation processing system according to an embodiment.

Figure 11A

[0028] Figure 11A is an exemplary workflow showing the creation of media in a delivery format according to an embodiment.

Figure 11B

[0029] Figure 11B is an exemplary workflow showing the sequential creation of media in a delivery format according to an embodiment.

Figure 12

[0030] Figure 12 is an exemplary workflow showing the packetization process according to an embodiment.

Figure 13

[0031] Figure 13 is an exemplary workflow showing the communication flow between components according to an embodiment.

Figure 14A

[0032] Figure 14A is an exemplary workflow showing the analysis of the complexity of immersive media according to an embodiment.

Figure 14B

[0033] Figure 14B is an example of a set of complexity attributes of scenes in a presentation according to an embodiment.

DETAILED DESCRIPTION OF THE INVENTION

[0015]

[0034] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numbers in the various figures may identify the same or similar elements.

[0016]

[0035] The foregoing disclosure provides examples and descriptions, but is not intended to be exhaustive or to limit the implementation to the specific forms disclosed. Modifications and variations are possible from the perspective of the above disclosure and may be made in the practice of implementation. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and operation descriptions provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be executed simultaneously (at least partially), and the order of one or more operations may be rearranged.

[0017]

[0036] It will be apparent that the systems and / or methods described herein may be implemented in various forms of hardware, software, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit the implementation. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0018]

[0037] Specific combinations of functions are recited in the claims and / or disclosed in the specification, but these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these functions may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Each of the dependent claims listed below may be directly dependent on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0019]

[0038] The proposed features discussed below can be used separately or in combination in any order. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0020]

[0039] Elements, acts, or instructions used in this disclosure should not be construed as important or essential unless explicitly described as such. Also, as used in this disclosure, the terms “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” The term “one” or similar words are used when only one item is intended. Also, as used in this disclosure, terms such as “has,” “have,” “having,” “include,” “including,” etc. are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “at least partially based on” unless explicitly stated otherwise. Further, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0021]

[0040] According to an embodiment, an immersive media - capable presentation device may refer to a device with sufficient resources and functions for access, interpretation, and presentation of immersive media. Such devices are heterogeneous in terms of the quantities and formats they can support, and in terms of the quantity and type of network resources required to distribute such media at scale. "At scale" may refer to the distribution of media by a service provider that realizes a distribution equivalent to that of conventional video and audio media via a network, such as Netflix, Hulu, Comcast subscribers, Spectrum subscribers.

[0022]

[0041] According to an embodiment, all client devices that function as end - points for the distribution of immersive media via a network are highly diverse. The distribution of any media via a network may use a media distribution system and architecture that reformats the media from an input or network "ingest" media format to a delivery media format, where the delivery media format is suitable not only for being ingested by the target client device and its applications, but also for being "streamed" via the network. Thus, there may be two processes that are performed on media ingested by the network: 1) converting the media from format A to format B, which is suitable for being ingested by the target client, i.e., based on the capabilities of the client that ingests a particular media format, and 2) preparing the media to be streamed.

[0023]

[0042] In an embodiment, streaming media broadly refers to fragmenting and packetizing media into successive small-sized chunks that are logically organized and ordered according to either or both of the temporal or spatial structure of the media so that they can be delivered over a network. The conversion of media from format A to format B (sometimes also called "transcoding") can be a process that is typically performed by a network or a service provider before delivering the media to a client device. Such transcoding can be done by converting the media from format A to format B based on prior knowledge, which can be that format B is a preferred format or the only format that can be taken in by the target client device, or is more suitable for delivery over a constrained resource such as a commercial network. In many cases, although not all, it is necessary to both convert and prepare the media for streaming so that the media can be received and processed by the target client device from the network.

[0024]

[0043] Converting (or transforming) media and preparing media for streaming is part of the process performed on media captured by a network before the media is delivered to a client device. The result of the process (i.e., the conversion and preparation for streaming) is a delivery media format, or simply a media format called a delivery format. These processes should be performed only once if they are performed on all of a given media data object and if the network has access to information indicating that the client will need the converted and / or streamed media object multiple times; otherwise, such media conversion and streaming will be triggered multiple times. That is, the processing and transfer of data for media conversion and streaming are generally considered to be a cause of latency, which involves the need to consume potentially large amounts of network and / or computing resources. Therefore, a network design that cannot access information indicating whether a client has cached a particular media data object or may have stored it locally with respect to the client will perform second-best performance compared to a network that can access such information.

[0025]

[0044] A scene graph is a common data structure commonly used in vector-based graphics editing applications and modern computer games, which may be something that arranges the logical and often (but not necessarily) spatial representation of a graphical scene, or may be a collection of nodes and vertices within a graph structure.

[0026]

[0045] A scene, in the context of computer graphics, may be a collection of objects (which may also be called, for example, 3D assets - media assets, media objects, objects, and assets). An object is an essential element of media data, object attributes, and other metadata, including visual, acoustic, and physical-based characteristics that describe a specific setting restricted by space or time regarding the interaction of objects within its setting.

[0027]

[0046] A node may be a basic element of a scene graph composed of information related to the logical or spatial or temporal representation of visual, auditory, tactile, olfactory, gustatory, or related processing information; each node shall have at most one output edge, zero or more input edges, and at least one edge (either input or output) connected thereto.

[0028]

[0047] The base layer may be a nominal representation of a media asset typically formulated to minimize the computing resources or time required to render an asset, or the time to transmit an asset over a network.

[0029]

[0048] The extended layer can be considered a set of information that, when applied to the base layer representation of an asset, extends the base layer to include features and functions not supported in the base layer.

[0030]

[0049] An attribute may be metadata associated with a node, and that node is used to describe a specific feature or function in either a normal form or a more complex form (e.g., regarding other nodes).

[0031]

[0050] The container may be a serialized format for storing and exchanging information representing all natural, all synthetic, or a mixture of synthetic and natural scenes, including all media resources and scene graphs necessary for rendering the scene.

[0032]

[0051] Serialization can be considered a process of converting the state of a data structure or object into a format that can be saved (e.g., in a file or memory buffer) or transmitted (e.g., via a network connection link) and later reconstructed (possibly in another computer environment). The resulting sequence of bits can be used to create a semantically identical clone of the original object when reread according to the serialization format.

[0033]

[0052] A renderer may be an (typically, software-based) application or process based on a selective mix of fields related to acoustical physics, optical physics, visual recognition, speech recognition, mathematics, and software development. Given an input scene graph and an asset container, it outputs typically visual and / or auditory signals suitable for presentation on a target device or compliant with desired properties specified by the attributes of a render target node within the scene graph. In the case of vision-based media assets, the renderer can output appropriate visual signals to a target display or, as an intermediate asset, to storage (e.g., repackaged into another container, i.e., used in a series of rendering processes of a graphics pipeline); in the case of audio-based media assets, the renderer can output audio signals for presentation on multi-channel loudspeakers and / or binauralized headphones or for repackaging into another (output) container. General examples of renderers include the real-time rendering capabilities of game engines Unity and Unreal Engine.

[0034]

[0053] A script language may be an interpreter-type programming language that is executed by the renderer at runtime and can process changes in variable state and dynamic inputs applied to scene graph nodes, which affects the rendering and evaluation of spatial and temporal object topologies (including physical forces, constraints, inverse kinematics, deformations, collisions) and the propagation and transport of energy (light, sound).

[0035]

[0054] Although originally used for shading (generating appropriate levels of light, darkness, and color within an image), a shader can now be a type of computer program that performs various special functions in different fields of computer graphics special effects, performs video post-processing unrelated to shading, or even performs functions completely unrelated to graphics.

[0036]

[0055] Path tracing is a computer graphics method for rendering a 3D scene such that the lighting of the scene is realistically faithful.

[0037]

[0056] Timed media is media and / or media objects that can potentially be ordered by time and may include, for example, those having a start time and an end time according to a specific clock. Untimed media is media and / or media objects that are organized by spatial, logical, or temporal relationships and may include, for example, those in an interactive experience realized according to actions performed by a user.

[0038]

[0057] A neural network model (NN model) can be considered a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations applied to visual signals to reach an improved visual output, and that improved visual output can include interpolation of a new perspective or view for visual signals not explicitly provided by the original signal.

[0039]

[0058] The number of immersive media-enabled devices introduced into the consumer market, including head-mounted displays, augmented reality glasses, hand-held controllers, multi-view displays, haptic gloves, and game consoles, has increased explosively over the past decade. Furthermore, holographic displays and other forms of volumetric displays are poised to enter the consumer market within the next three to five years. However, despite the near or imminent availability of these devices, a consistent end-to-end ecosystem for the delivery of immersive media in commercial networks has not been successfully realized for several reasons.

[0040]

[0059] One reason a consistent end-to-end ecosystem for the delivery of immersive media on commercial networks has not been realized is that all client devices that function as endpoints of such delivery networks for immersive displays are highly diverse. Some client devices support a particular immersive media format, while others do not. Some can create an immersive experience from conventional raster-based formats, while others cannot. Unlike networks designed only for the delivery of legacy media, networks that must support the diversity of display clients require a large amount of information about the details of each client's capabilities and the format of the media being delivered, and then such networks can use an adaptation process to convert the media into a format suitable for each target display and corresponding application. Such networks need to access at least information that describes the complexity of the media being ingested into the network and the characteristics of each target display, and determine a way to meaningfully adapt the input media source to a format suitable for the target display and application.

[0041]

[0060] A network that supports such heterogeneous client devices should take advantage of the fact that some assets that are adapted from an input media format to a specific target format may be reused across a set of similar display targets. That is, some assets, once converted to a format suitable for the target display, may be reused among several such displays with similar adaptation requirements. Therefore, a network that uses a caching mechanism to store the adapted assets in a relatively immutable area will be more efficient.

[0042]

[0061] Immersive media can be composed into "scenes" described by a scene graph, which is also called scene descriptions. The purpose of the scene graph may be to describe visual, auditory, and other forms of immersive assets that are composed from specific settings that are part of a presentation, for example, events that take place at a specific location within a building or actors that are part of a presentation such as a movie. A list of all scenes that make up one presentation may be formulated into a manifest of scenes.

[0043]

[0062] The advantages of a "scene"-based approach are for pre-prepared content that requires delivering such content, and it is possible to create a "bill of materials" that identifies all the assets that will be used throughout the presentation, a bill of materials that identifies how often each asset is used across the various scenes within the presentation. The network has knowledge about the existence of cached resources that can be used to meet the asset requirements of a particular presentation. Similarly, a client device presenting a series of scenes may wish to have knowledge about the frequency of a given asset used across multiple scenes. For example, if a media asset (also called a media object, asset, or object) is referenced multiple times across multiple scenes that are processed or will be processed by the client device, the client device should avoid discarding the asset from its cached resources until the last scene that requires that particular asset is presented by the client.

[0044]

[0063] In the case of a conventional presentation device, the delivery format may be equivalent or sufficiently equivalent to the "presentation format" that is ultimately used by the client presentation device to create the presentation. That is, the presentation media format is a media format whose properties (resolution, frame rate, bit depth, color gamut, etc.,...) are precisely adjusted to the capabilities of the client presentation device. Some examples of the delivery format and the presentation format include a high-definition (HD) video signal (1920 pixel columns x 1080 pixel rows) delivered over a network to an ultra-high-definition (UHD) client device having a resolution of (3840 pixel columns x 2160 pixel rows). The UHD client may apply a process called "super-resolution" to the HD delivery format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format presented by the client device is the "presentation format", which is a UHD signal in this example, while the HD signal constitutes the delivery format. In this example, the delivery format of the HD signal is very similar to the presentation format of the UHD signal because both signals are in a rectilinear video format and the process of converting the HD format to the UHD format is a relatively straightforward and simple process in most legacy client devices.

[0045]

[0064] However, in some embodiments, the preferred presentation format for the targeted client device may be significantly different from the capture format received over the network. Nevertheless, the client device may have access to sufficient computing, storage, and bandwidth resources to convert the media from the capture format to the necessary presentation format suitable for presentation by the client device. The network may bypass reformatting the captured media from a first format A to a second format B, e.g., "transcoding" the media, because the client can access sufficient resources to perform all media conversions without the network having to do so a priori. The network can fragment and package the captured media so that the media can be streamed to the client.

[0046]

[0065] However, in some embodiments, the captured media received over the network is significantly different from the client's preferred presentation format, and the client device does not have access to sufficient computing, storage, or bandwidth resources to convert the media to its preferred presentation format. In the absence of access to resources, the network can assist the client by performing all or part of the conversion from the capture format to a format that is equivalent or nearly equivalent to the client device's preferred presentation format. In some embodiments, such assistance provided by the network on behalf of the client is generally referred to as "split rendering."

[0047]

[0066] Embodiments of the present disclosure as described herein enable a network to convert all or part of an ingest media from a first format (e.g., Format A) to a second format (e.g., Format B) and to determine whether to assist the capabilities of a client device to potentially generate a presentation of the media in a third format C. This determination may be made by a processor or server of the media streaming network or by the client device. To assist in the determination, it may be useful to determine which media assets are used multiple times within the context of the presentation and to design a process and / or network that makes those media assets readily available to the network for use. Based on information from such an analysis, it may then be possible to design the network such that a client device (also referred to as a "client") is requested to hold copies of one or more media assets that are likely to be used multiple times in its local cache.

[0048]

[0067] However, when the client device saves a copy of the media asset to its local cache, the network may not have any control over the management of the client device's local cache. As a result, the client device may encounter a situation where it has to delete resources (resources that are still reusable) from its local cache. By assisting with the design, in order to optimize the network to minimize the need to perform a conversion of a media asset from a first format to a second format for a media asset that is used multiple times, or the need to promote re-streaming a media asset that is used multiple times to the client, the network can manage its own cache separately from any cache held by the client. As a result, the network guarantees that at least one redundant copy of the reusable asset is accessible to both the client and the network respectively.

[0049]

[0068] In an embodiment, the network can first query the client device to obtain feedback and ensure that the media asset in question is still available in the client's local cache. If the response from the client device indicates that it no longer has a copy of the media asset in question, the network can signal to the client that the client should access a copy of the media asset in its delivery format from the redundant cache. In some embodiments, the query to the client device may be omitted, and the network can signal to the client that the client should access a copy of the delivery format of the asset from the redundant cache.

[0050]

[0069] FIG. 1A is an exemplary diagram of a media delivery streaming system 100 according to an embodiment for delivering media from a network cloud, edge device, or server 104 to a client device 108. As shown in FIG. 1A, media including immersive media including one or more scenes and one or more media objects is received from a content provider in a first format A (hereinafter referred to as "capture media format A"). This process is performed or carried out by a network cloud or edge device (hereinafter referred to as "network device 104") and is delivered to a client such as, for example, client device 108. In some embodiments, the same process may be performed empirically by a manual process or by the client device. The network device 104 can, for example, capture the media in a first format 101, generate and / or create the delivery media 102 in a second format 102 (hereinafter referred to as "delivery media creation 102"), and deliver the media in the second format 103, for example, using a delivery module. The client device 108 can include a rendering module 106 and a presentation module 107.

[0051]

[0070] According to one aspect, the network device 104 may receive the captured media from a content provider or the like. The media streaming network can obtain the captured media stored in the captured media format A. The delivered media can be created and / or generated using any necessary conversion or conditioning of the captured media to create a potential alternative representation of the media. That is, the delivery format of the media object in the captured media can be created. As described above, the delivery format is a media format that can be delivered to the client by formatting the media into the delivery format B. The delivery format B is a format prepared for streaming to the client device 108. The delivered media creation 102 can include optimization reuse logic that executes a decision-making process to determine whether a particular media object is already being streamed to the client device 108. Further operations related to the delivered media creation 102 and the optimization reuse logic will be described in detail with reference to FIG. 1B.

[0052]

[0071] The media formats A and B may or may not be representations that follow the same syntax of a particular media format specification, but format B is likely to be conditioned on a scheme that facilitates the delivery of media by a network protocol. The network protocol may be, for example, a connection-oriented protocol (TCP) or a connectionless protocol (UDP). The delivery module streams the streamable media (i.e., the media format B) from the network device 104 to the client device 108 via the network connection 105.

[0053]

[0072] The client device 108 can receive the delivered media and use the rendering module 106 to render the media for presentation. The rendering module 106 can access any rendering function, which may be basic or advanced, depending on the target client device 108. The rendering module 106 can create presentation media in presentation format C. The presentation format C may or may not be expressed according to a third format specification. Therefore, the presentation format C may be the same as or different from the media formats A and / or B. The rendering module 106 outputs the presentation format C to the presentation module 107, and the presentation module can present the presentation media on the display (or the like) of the client device 108.

[0054]

[0073] Embodiments of the present disclosure facilitate the decision-making process used by a network to calculate the sequence order for packaging assets and streaming them from the network to a client. In this case, all assets used across a set of one or more scenes constituting a presentation are analyzed by a media complexity analyzer to determine the complexity associated with each asset throughout all the scenes constituting the presentation. Therefore, the order in which the assets of a particular scene are packaged and streamed to the client can be based on the complexity with which each asset is used across the set of scenes constituting the presentation.

[0055]

[0074] Embodiments address the need for a mechanism or process to analyze immersive media scenes in order to obtain sufficient information that can be used to support a decision-making process, where the decision-making process provides an indication as to whether the conversion of media objects from format A to format B should be performed completely by the network, completely by the client, or by a mix of both (along with an indication of which assets should be converted by the client or the network) when used by the network or the client. Such an immersive media data complexity analyzer can be used by the client or the network in an automated situation or manually, for example, by a human operating a system or device.

[0075] According to embodiments, the process of adapting an input immersive media source to a particular endpoint client device may be the same as or similar to the process of adapting the same input immersive media source to a particular application running on the particular client endpoint device. Thus, the problem of adapting an input media source to the characteristics of an endpoint device is of the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.

[0056]

[0076] Figure 1B is a workflow of delivery media creation 102 according to an embodiment. More specifically, the workflow of Figure 1B includes the generation of a reuse indicator in a media streaming network that assists in a decision-making process to determine whether a particular media object is already being streamed to client device 108.

[0057]

[0077] At 152, the delivery media creation process is started. At 155, it is possible to execute conditional logic to determine whether the current media object has previously been streamed to the client device 108. To determine whether the media object has previously been streamed to the client, a list of unique assets can be accessed for presentation. If the current media object has previously been streamed, the process proceeds to process 160. In process 160, since the client has already received the current media object, an indicator (hereinafter also referred to as a "proxy") is created to identify that the client should access a copy of the media object from the local cache or other cache. If it is determined that the media object has not previously been streamed, the process proceeds to process 165. In process 165, the media object may be prepared for conversion and / or delivery, and a delivery format for the media object is created. Then, the processing of the current media object ends.

[0058]

[0078] FIG. 2A is an exemplary workflow for processing captured media over a network. The workflow shown in FIG. 2A illustrates a media conversion decision-making process 200 according to an embodiment. The media conversion decision-making process 200 is used to determine whether the network should convert the media before delivering the media to the client device. The media conversion decision-making process 200 may be processed through manual or automatic processes within the network.

[0059]

[0079] The captured media represented in Format A is provided to the network by a content provider. At 205, the media is captured from the content provider by a media streaming network. At 210, if the attributes of the target client are not yet known, they are obtained. The attributes describe the processing capabilities of the target client.

[0060]

[0080] At 215, it is determined whether the network (or client) should assist in the conversion of the captured media. In some embodiments, at 215, it is specifically determinable whether any format conversion (e.g., conversion of one or more media objects from Format A to Format B) is required for any media asset included in the captured media before the media is streamed to the target client. At 215, the determination may be based on whether the media can be streamed in the original captured Format A or whether it needs to be converted to another Format B to facilitate the presentation of the media by the client. Such a determination (i.e., whether conversion of the captured media is required before streaming the media to the client or whether the media should be streamed directly to the client in the original captured Format A) may require accessing information describing the aspects or characteristics of the captured media.

[0061]

[0081] If it is determined that the network (or client) should assist in the conversion of any media asset (YES at 215), process 200 proceeds to 220.

[0062]

[0082] At 220, the captured media is media-converted from format A to format B to generate the converted media (222). The converted media is output (222), and at 225, the input media executes a preparation process for streaming the media to the client. In this case, the converted media 222 (i.e., the input media) is prepared to be streamed.

[0063]

[0083] Immersive media streaming may be in a relatively nascent stage, especially when such media is "scene - based" rather than "frame - based". For example, frame - based media streaming may be equivalent to streaming video frames, where each frame captures a complete image of the entire scene or a complete picture of the entire object presented by the client. When the sequence of frames is reconstructed by the client from their compressed format and presented to the viewer, it creates a video sequence that constitutes the entire immersive presentation or a part of the presentation. In the case of frame - based streaming, the order in which frames are streamed from the network to the client may conform to a given specification (such as ITU - T Recommendation H.264 Advanced Video Coding for general audio - visual services). However, scene - based streaming of media is different from frame - based streaming because scenes may be composed of individual assets that may be independent of each other. A given scene - based asset may be used multiple times within a particular scene or across a series of scenes. The time required for a client or any given renderer to use to reconstruct a particular asset may depend on a number of factors including, but not limited to, the size of the asset, the availability of computing resources for performing the rendering, and other attributes representing the overall complexity of the asset. A client that supports scene - based streaming may require that all or part of the rendering of each asset within a scene be complete before any presentation of the scene can begin. Therefore, the order in which assets are streamed from the network to the client can affect the overall performance of the system.

[0064]

[0084] The conversion of media from format A to another format (e.g., format B) may be performed entirely by the network, entirely by the client, or jointly between both the network and the client. In the case of split rendering, it may be obvious that a lexicon of attributes that describe the media format is required so that both the client and the network have complete information to characterize the work that must be done. Further, a dictionary that provides the attributes of the client's capabilities may similarly be required in terms of available computing resources, available storage resources, access to bandwidth, etc. Further, a mechanism is required to characterize the level of computational, memory, or bandwidth complexity of the captured media format so that the network and the client can jointly or independently determine whether or when to use the split rendering process for the network to deliver the media to the client.

[0065]

[0085] If it is determined that the network (or client) should not (or does not need to) assist in the conversion of any media assets (NO at 215), process 200 proceeds to 225. At 225, the media is prepared for streaming. In this case, the captured data (i.e., the media in its original form) is prepared for streaming.

[0066]

[0086] Finally, once the media data is in a streaming format, the media prepared at 225 is streamed to the client at 230. In some embodiments (as described with reference to FIG. 1B), if the conversion and streaming of certain media objects that may be required or will be required by the client to complete the media presentation are avoided, the network may skip the conversion and streaming of the ingest media (i.e., 215-230), assuming that the client can still access or utilize the media objects that the client may need to complete the media presentation. To maximize the capabilities of the client, it may be desirable for the network to have sufficient information such that the network can determine an order for streaming scene-based assets from the network to the client that improves the client's performance. For example, in a particular presentation, such a network having sufficient information to avoid repetitive conversion and / or streaming steps for assets that are used multiple times may function optimally compared to a network not designed as such. Similarly, a network that can "intelligently" order the delivery of assets to the client can encourage the client to maximize its capabilities (i.e., create a more enjoyable experience for the end user).

[0067]

[0087] FIG. 2B shows an exemplary media conversion process 250 that includes determining media asset reuse. Similar to the media conversion decision process 200, the media conversion process 250 that performs media asset reuse processes the ingest media via the network and determines whether the network should convert the media before delivering the media to the client.

[0068]

[0088] The capture media represented in format A is provided to the network by a content provider. According to an embodiment, processes 255-260 and 275-286 are executed in the same manner as 205-210 and 215-230 shown in FIG. 2A. At 255, the media is captured from the content provider by the network. Then, at 260, the attributes of the target client are obtained if they are not yet known. The attributes describe the processing capabilities of the target client.

[0069]

[0089] If it is confirmed that the network has previously streamed a specific media object or the current media object (YES at 265), the process proceeds to 270. At 270, a proxy is created in place of the previously streamed media object to indicate that the client should use a local copy of the previously streamed object or a copy of the previously streamed object stored in another cache.

[0090] If it is determined that the network has not previously streamed a media object (NO at 265), the process proceeds to 275. At 275, it is determined whether the network or the client should perform any format conversion on any media asset included in the capture media at 255. For example, the conversion may include converting a specific media object from format A to format B before the media is streamed to the client. It is possible to assume that process 275 is similar to the process executed at 215 shown in FIG. 2A.

[0070]

[0091] If it is determined that the media asset should be converted by the network (YES at 275), the process proceeds to 280. At 280, the media object is converted from format A to format B. The converted media is then prepared to be streamed to the client (286).

[0071]

[0092] If it is determined that the media asset should not be converted by the network (NO at 275), the process proceeds to 285. At 285, the media object is then prepared to be streamed to the client. When the media is in a streamable format, the media prepared at 285 is streamed to the client at 286.

[0072]

[0093] The streamable format of the media can be timed or untimed heterogeneous immersive media. FIG. 3A shows an example of a timed media representation 300 of a streamable format of heterogeneous immersive media. The timed immersive media can include a set of N scenes. The timed media is media content ordered by time according to a specific clock, for example, by a start time and an end time. FIG. 4 shows an example of an untimed media representation 400 of a streamable format of heterogeneous immersive media. The untimed media is media content organized by spatial, logical, or temporal relationships (such as an interactive experience realized according to actions performed by one or more users).

[0073]

[0094] FIGS. 3A-B represent time-specified scenes of time-based media, and FIGS. 4A-B represent non-time-specified scenes of non-time-based media. The time-specified scenes and the non-time-specified scenes may correspond to various scene representations, or descriptions of the scenes. FIGS. 3B, 3B, 4A, and 4B all use a single exemplary comprehensive media format adapted from the source capture media format to conform to the functions of a particular client end point. That is, the comprehensive media format is a delivery format streamable to the client device. The comprehensive media format is robust enough in its structure to be able to accommodate various media attributes, in which case each attribute can be hierarchically arranged based on the amount of significant information contributed by each layer to the presentation of the media.

[0074]

[0095] As shown in FIG. 3A, the timed media representation 300 includes a time-specified scene manifest 300A that contains a list of scene information 301. The scene information 301 refers to a list of components 302 that separately describe the type of media asset that makes up the scene information 301 and the processing information, for example, an asset list and other processing information. The list of components 302 can refer to proxy assets 308 corresponding to the type of asset (e.g., proxy visual and audio assets as shown in FIG. 3A). Component 302 refers to a list of unique assets that have not been used in other scenes previously. For example, a list of unique assets 307 for (time-specified) scene 1 is shown in FIG. 3A. Component 302 also refers to an asset 303 that includes a base layer 304 and an attribute enhancement layer 305. The base layer is a nominal representation of an asset that can be formulated to minimize the computational resources, the time required to render the asset, and / or the time required to transmit the asset over the network. In this exemplary embodiment, each base layer 304 refers to a numerical complexity metric that characterizes the effort that a client may need to expend from the perspective of the time or resources to process the asset. The extended layer can be considered a set of information that extends the base layer to include features or functions that may not be supported by the base layer when applied to the base layer representation of the asset.

[0075]

[0096] Figure 3B shows a time-ordered media representation with the complexities 3030 sorted in descending order. This time-ordered media representation is the same as the one shown in Figure 3A, but the assets 3033 in Figure 3B are ordered by asset type and the descending complexity scale values within each asset type. The time-specified scene manifest 303A includes a list of scene information 3031. The list of scene information 3031 refers to a list of components 3032 that individually describe the type of media asset and the processing information that includes the list of scene information 3031, for example, an asset list and other processing information. Component 3032 may refer to a proxy asset 3038 corresponding to the type of asset (e.g., proxy visual and audio assets as shown in Figure 3B). Component 3032 also refers to an asset 3033 that further references a basic layer 3034 and an attribute enhancement layer 3035. Each basic layer 3034 is ordered according to the descending values of the corresponding complexity scale. A list of unique assets 3037 that have not been used in other scenes previously is also provided.

[0076]

[0097] As shown in FIG. 4A, the non-time-based media and complexity representation 400 includes scene information 401. The scene information 401 is not associated with start and end times / durations (e.g., by a clock, timer, etc.). A non-time-specified scene manifest (not shown) can refer to scene 1.0, and there are no other scenes that can branch to scene 1.0. The scene information 401 refers to a list of components 402 that individually describe the types and processing information of the media assets that make up the scene information 401. The components 402 refer to visual assets, auditory assets, tactile assets, and time-specified assets (collectively referred to as assets 403). The assets 403 further refer to a base layer 404 and attribute enhancement layers 405 and 406. In this exemplary embodiment, each base layer 404 refers to a numerical complexity metric that characterizes the effort that a client may need to expend in terms of time or resources to process the assets across the scene including the presentation. The scene information 401 can also refer to scene information 407 regarding other non-time-specified scenes for non-time-based media sources (i.e., those referred to as non-time-specified scenes 2.1 - 2.4 in FIG. 4A) and / or time-specified media scenes (i.e., those referred to as time-specified scene 3.0 in FIG. 4A). In the example of FIG. 4A, the non-time-based immersive media includes a set of five scenes (including both time-specified and non-time-specified ones). A list 408 of unique assets identifies unique assets associated with a particular scene that have not been previously used in a higher-ranked (e.g., parent) scene. The list 408 of unique assets shown in FIG. 4 includes unique assets for non-time-specified scene 2.3.

[0077]

[0098] Figure 4B shows the complexity representation 4040 ordered with non-time-specified media. A non-time-specified scene manifest (not shown) refers to scene 1.0 where there are no other scenes that can branch to scene 1.0. Scene information 4041 is not associated with start and end durations according to a clock. Scene information 4041 also refers to a list of components 4042 that individually describe the type and processing information of the media assets that make up scene information 4041. Component 4042 refers to visual assets, auditory assets, tactile assets, and time-specified assets (collectively referred to as assets 4043). Assets 4043 further refer to a base layer 4044 and attribute enhancement layers 4045 and 4046. In this exemplary embodiment, the tactile assets 4043 are composed by increasing the complexity scale value, while the auditory assets 4043 are composed by decreasing the complexity scale value, or vice versa. Further, scene information 4041 refers to other non-time-specified scene information 4041 for non-time-specified media. Scene information 4041 also refers to other time-specified scene information 4047 for time-specified media. A list 4048 of unique assets identifies unique assets associated with a particular scene that have not been previously used in a higher-ranked (e.g., parent) scene.

[0078]

[0099] Media streamed according to an inclusive media format is not limited to conventional visual and audio media. The inclusive media format may include any type of media information capable of generating signals that interact with a machine to stimulate human senses for vision, hearing, taste, touch, and smell. As shown in FIGS. 3A-B and 4A-B, media streamed according to an inclusive media format may be time-specified or unspecified media, or a mixture of both. The inclusive media format uses a base layer and enhancement layer architecture to enable a hierarchical representation of media objects, making it streamable.

[0079]

[0100] In some embodiments, the individual base layer and enhancement layer are computed by applying multi-resolution or multi-tessellation analysis techniques to the media objects of each scene. This computing technique is not limited to raster-based visual formats.

[0080]

[0101] In some embodiments, the progressive representation of geometric objects may be a multi-resolution representation of the object computed using wavelet analysis techniques.

[0081]

[0102] In some embodiments, in a hierarchical representation media format, the enhancement layer can apply various attributes to the base layer. For example, one or more enhancement layers can refine the material properties of the surface of the visual object represented by the base layer.

[0082]

[0103] In some embodiments, in a hierarchical representation media format, an attribute can refine the texture of the surface of an object represented by a base layer, for example, by changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.

[0083]

[0104] In some embodiments, in a hierarchical representation media format, the surface of one or more visual objects within a scene can be changed from a Lambertian surface to a surface that is ray-traceable.

[0084]

[0105] In some embodiments, in a hierarchical representation media format, a network can deliver a base layer representation to a client, such that the client can create a nominal presentation of the scene, while the client can wait for the transmission of additional enhancement layers to refine the resolution or other characteristics of the base layer.

[0085]

[0106] In an embodiment, the resolution or sophistication information of attributes in the enhancement layer is not explicitly associated with the resolution of the objects in the base layer. Further, an inclusive media format can support any type of information media that may be presented or acted upon by a presentation device or machine, thereby enabling support for heterogeneous media formats for heterogeneous client endpoints. In some embodiments, the network that distributes the media format first queries the client's endpoint to determine the client's capabilities. Based on the query, if the client cannot meaningfully ingest the media representation, the network can remove the layer of attributes not supported by that client. In some embodiments, if the client cannot meaningfully ingest the media representation, the network can adapt the media from its current format to a format suitable for the client's endpoint. For example, the network can adapt the media by using a network-based media processing protocol to convert a volumetric visual media asset to a 2D representation of the same visual asset. In some embodiments, the network can adapt the media by using a neural network (NN) process to reformat the media into an appropriate format or optionally synthesize the views required at the client endpoint.

[0086]

[0107] The manifest of a scene for a complete (or partially complete) immersive experience (live streaming event, game, or playback of an on-demand asset) is composed of scenes that contain the minimum information required for rendering and capture to create a presentation. The scene manifest includes a list of individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene are one or more representations of geometric objects within the scene that correspond to a streamable version of the scene geometry. One embodiment of a scene may refer to a low-resolution version of the geometric objects of the scene. Another embodiment of the same scene may refer to an enhancement layer of the low-resolution representation of the scene to add additional detail or increase tessellation of the geometric objects of the same scene. As described above, each scene can have one or more enhancement layers to incrementally increase the detail of the geometric objects of the scene. Each layer of a media object referenced within a scene can be associated with a token (e.g., a Uniform Resource Identifier (URI)) that points to the address of where the resource is accessible within the network. Such resources are similar to content delivery networks (CDNs) where a client can acquire content. Tokens for representations of geometric objects can point to locations within the network or within the client. That is, the client can signal to the network that a resource is available to the network for network-based media processing.

[0087]

[0108] According to an embodiment, a (time-specified or unspecified) scene may correspond to a scene graph as a plurality of plane images (Multi-Plane Image, MPI) or a plurality of spherical images (Multi-Spherical Image, MSI). The technologies of both MPI and MSI are examples of technologies that assist in creating display-agnostic scene representations of natural content (i.e., images of the real world captured simultaneously from one or more cameras). On the other hand, scene graph technology may be used to represent both natural images and computer-generated images in the form of a synthetic representation. However, such a representation is particularly computationally burdensome to create for cases where the content is captured as a natural scene by one or more cameras. The scene graph representation of naturally captured content is computationally and time-consuming to create, and requires complex analysis of natural images using photogrammetry or deep learning or both technologies to create a synthetic representation that can be subsequently used to interpolate a sufficient and appropriate number of views to satisfy the viewing frustum of the target immersive client display. As a result, such synthetic representations are not realistically considered as candidates for representing natural content, because they cannot be created in real time for use cases that require real-time delivery. Therefore, the optimal representation of computer-generated images is to adopt the use of scene graphs in a synthetic model, because computer-generated images are created using 3D modeling processes and tools, and adopting the use of scene graphs in a synthetic model results in the best representation of computer-generated images.

[0088]

[0109] Figure 5 shows an example of the natural media synthesis process 500 according to the embodiment. The natural media synthesis process 500 converts the capture format from a natural scene to a representation (a representation that can be used as a capture format for a network that caters to heterogeneous client endpoints). The left side of the dashed line 510 is the content capture part of the natural media synthesis process 500. The right side of the dashed line 510 is the capture format synthesis (for natural images) of the natural media synthesis process 500.

[0089]

[0110] As shown in Figure 5, the first camera 501 uses a single camera lens to capture a scene of, for example, a person (i.e., the actor shown in Figure 5). The second camera 502 captures a scene with five diverging fields of view by attaching five camera lenses around a ring-shaped object. The arrangement of the second camera 502 shown in Figure 5 is an exemplary arrangement commonly used for capturing omnidirectional content of VR applications. The third camera 503 captures a scene with seven converging fields of view by attaching seven camera lenses to the inner diameter part of a sphere. The arrangement of the third camera 503 is an exemplary arrangement commonly used for capturing a light field for a light field or holographic immersive display. The embodiment is not limited to the configuration shown in Figure 5. The second camera 502 and the third camera 503 may include multiple camera lenses.

[0090]

[0111] The natural image content 509 is output from the first camera 501, the second camera 502, and the third camera 503 and provided as an input to the synthesizer 504. The synthesizer 504 can perform NN training 505 using a collection of training images 506 to generate a capture NN model 508. The training images 506 may be predefined or may be saved from previous synthesis processes. The NN model (e.g., the capture NN model 508) is a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations applied to visual signals to reach an improved visual output, and the improved visual output can include interpolation of new views for visual signals not explicitly provided by the original signal.

[0091]

[0112] In some embodiments, a photogrammetry process may be performed instead of the NN training 505. If the capture NN model 508 is created during the natural media synthesis process 500, the capture NN model 508 becomes one of the assets in the natural media content capture format 507. The capture format 507 may be, for example, MPI or MSI. The capture format 507 may include media assets.

[0092]

[0113] FIG. 6 shows an example of a synthetic media capture creation process 600 according to an embodiment. The synthetic media capture creation process 600 creates a capture media format for synthetic media, such as computer-generated images.

[0093]

[0114] As shown in FIG. 6, camera 601 can capture point cloud 602 of a scene. Camera 601 may be, for example, a LIDAR camera. Computer 603 employs, for example, a common gateway interface (CGI) tool, a 3D modeling tool, or another animation process to create synthetic content (i.e., a representation of a synthetic scene that can be used as an import format for a network to handle heterogeneous client endpoints). Computer 603 can create CGI asset 604 via a network. Further, sensor 605A may be attached to actor 605 within the scene. Sensor 605A may be, for example, a motion capture suit with sensors attached. Sensor 605A captures a digital record of the movement of actor 605 and generates activity motion data 606 (or MoCap Data). The data from point cloud 602, CGI asset 604, and motion data 606 are provided as inputs to synthesizer 607 that creates synthetic media import format 608. In some embodiments, synthesizer 607 can create an NN model that uses an NN and training data to generate synthetic media import format 608.

[0094]

[0115] Both natural content and computer-generated (i.e., synthetic) content can be stored in a container. The container can include a serialized format for storing and exchanging information representing all natural, all synthetic, or a mixture of synthetic and natural scenes, including a scene graph and all media resources (necessary for rendering the scene). The serialization process of the content includes converting the data structure or the state of an object into a format that can later be stored (e.g., in a file or memory buffer) or transmitted (e.g., via a network connection link) and reconfigured in the same or a different computer environment. When the resulting series of bits is reread according to the serialized format, it can be used to create a semantically identical clone of the original object.

[0095]

[0116] The dichotomy in the optimal representation of both natural content and computer-generated (i.e., synthetic) content suggests that the optimal capture format for naturally captured content is different from the optimal capture format for computer-generated content or for natural content that is not essential for real-time delivery applications. Thus, according to an embodiment, the network is intended to be robust enough to support multiple capture formats for visually immersive media, regardless of whether it is created naturally, e.g., by using a physical camera, or created by a computer.

[0096]

[0117] Technologies such as ORBX by OTOY, Universal Scene Description by Pixar, and the Graphics Language Transmission Format 2.0 (glTF2.0) specification described by the Khronos 3D Group embody a scene graph as a format suitable for representing visually immersive media created using computer-generated techniques or naturally captured content. For naturally captured content, deep learning or photogrammetry techniques are used to create corresponding synthetic representations of natural scenes (i.e., those that are not essential in the case of real-time delivery applications).

[0097]

[0118] ORBX by OTOY is one of several scene graph technologies capable of supporting any type of visual media, either time-specified or not, including ray-traceable, legacy (frame-based), volumetric, and other types of synthetic or vector-based visual formats. ORBX is a unique one different from other scene graphs because ORBX provides native support for freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph intentionally designed to facilitate interactions across multiple vendor technologies operating on the scene graph. Further, ORBX provides a rich material system, support for the Open Shader Language, a robust camera system, and support for Lua scripts. ORBX is also the basis of the Immersive Technologies Media Format, which has been licensed royalty-free by the IDEA (Immersive Digital Experiences Alliance). In the context of real-time delivery of media, the ability to create and deliver an ORBX representation of a natural scene is a function of the availability of computing resources for performing complex analysis of data captured by a camera and synthesis of the same data into a synthetic representation.

[0098]

[0119] USD by Pixar is a scene graph commonly used in visual effects and specialized content production. USD is integrated into NVIDIA's Omniverse platform, a set of developer tools for creating and rendering 3D models using NVIDIA's graphics processing units (GPUs). A subset of USD, published by Apple and Pixar, is called USDZ and is supported by Apple's ARKit.

[0099]

[0120] glTF2.0 is a version of the Graphics Transmission Format specification written by the Khronos 3D Group. This format supports a simple scene graph format (including PNG and JPEG image formats) that can generally support static (not time-specified) objects within a scene. glTF2.0 supports simple animations, which includes supporting translation, rotation, and scaling of basic shapes described using glTF primitives (i.e., for geometric objects). glTF2.0 does not support time-specified media, and thus does not support video or audio media input.

[0100]

[0121] These designs for the scene representation of immersive visual media are provided for illustrative purposes only and do not limit the subject matter disclosed in the function of specifying a process of adapting an input immersive media source to a format suitable for the specific characteristics of a client endpoint device. Further, any or all of the above-exemplified media representations may employ or may be likely to employ deep learning techniques and train and create an NN model that enables or facilitates the selection of a specific view so as to satisfy the frustum of a view of a specific display based on specific dimensions of the frustum. The view selected for the frustum of a specific display can be interpolated from among the existing views explicitly provided in the scene representation (e.g., by MSI or MPI techniques). The view can also be directly rendered from a rendering engine based on the location of a specific virtual camera, filter, or the description of the virtual camera of these rendering engines.

[0101]

[0122] The methods and devices of the present disclosure are robust enough to account for a relatively small but well-known set of immersive media capture formats, which can sufficiently meet the requirements for real-time or on-demand (e.g., non-real-time) delivery of media that has been naturally captured (e.g., using one or more cameras) or created using computer-generated techniques.

[0102]

[0123] Interpolation of views from immersive media capture formats using either an NN model or a network-based rendering engine is further facilitated by advanced network technologies (e.g., those related to 5G of mobile networks), and optical fiber cables are deployed for fixed networks. These advanced network technologies improve the capacity and capabilities of commercial networks because such advanced network infrastructure can support the increasing transfer and delivery of visual information. Network infrastructure management technologies such as MEC (Multi-access Edge Computing), SDN (Software Defined Networks), and NFV (Network Functions Virtualization) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changes in the requirements for specific network resources, e.g., to respond to dynamic increases and decreases in requirements for network throughput, network speed, round-trip latency, and computing resources. Furthermore, this inherent ability to adapt to dynamic network requirements facilitates the network's ability to adapt immersive media capture formats to appropriate delivery formats in order to support various immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.

[0103]

[0124] The immersive media application itself may also have various requirements for network resources, including game applications that require significantly less network latency to respond to real-time updates within the game state, telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network, and passive display applications that may increase downlink resource requirements depending on the type of client endpoint display using the data. Generally, any consumer application may be supported by various client endpoints with various on-board client functions related to memory, computing, and power, as well as various requirements related to specific media representations.

[0104]

[0125] Accordingly, embodiments of the present disclosure enable a fully equipped network, i.e., a network that employs all or some of the features of a state-of-the-art network, to simultaneously support a plurality of legacy and immersive media compliant devices according to the functions specified within the device. As such, the immersive media delivery methods and processes described herein provide the flexibility to utilize practical media capture formats for both real-time and on-demand use cases for media delivery, the flexibility to support both natural and computer-generated content for both legacy and immersive media compliant client endpoints, and the flexibility to support both time-specified and non-time-specified media. This method and process also dynamically adapt the source media capture format to an appropriate delivery format based on the characteristics and capabilities of the client endpoint, and further based on the requirements of the application. This ensures that the delivery format can be streamed over an IP-based network and that the network can simultaneously handle a plurality of heterogeneous client endpoints that may include both legacy and immersive media compliant devices. Additionally, the embodiments provide an exemplary media presentation framework that facilitates the composition of the delivered media across scene boundaries.

[0105]

[0126] The end-to-end implementation of heterogeneous immersive media delivery according to embodiments of the present disclosure that provide the aforementioned improvements is achieved according to the processes and components described in the detailed description of FIGS. 7-16, which are described in detail below.

[0106]

[0127] The technology for representing and streaming the above-described heterogeneous immersive media can be physically stored in one or more non-transitory computer-readable media as computer software using computer-readable instructions, or can be specifically constructed by one or more specially constructed hardware processors, and can be implemented at both the source and the destination. FIG. 7 shows a computer system 700 suitable for implementing a particular embodiment of the disclosed subject matter.

[0107]

[0128] The computer software can be coded using any suitable machine code or computer language, and can be processed by assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation and micro-code execution.

[0108]

[0129] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0109]

[0130] The components shown in FIG. 7 for the computer system 700 are essentially exemplary and are not intended to suggest any limitation as to the scope or functionality of the computer software for implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system 700.

[0110]

[0131] The computer system 700 can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), auditory input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input. Also, the human interface device can be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).

[0111]

[0132] The input human interface device can be one or more of a keyboard 701, a trackpad 703, a mouse 703, a screen 709 (although only one of each is depicted), and can include, for example, a touch screen, a data glove, a joystick 704, a microphone 705, a camera 706, and a scanner 707.

[0112]

[0133] Computer system 700 may also include certain human interface output devices. Such human interface output devices can, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., touch screen 709, data glove, tactile feedback by joystick 704, although there may be tactile feedback devices that are not useful as input devices), auditory output devices (e.g., speaker 708, headphones), visual output devices (e.g., screens 709 including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may be capable of outputting two-dimensional visual output or output in three or more dimensions by means such as stereoscopic output; virtual reality glasses (not shown), holographic display, and smoke tank), and may include printers.

[0113]

[0134] Computer system 700 may also include optical media 711 including CD / DVD ROM / RW that uses media 710 such as CD / DVD, thumb drive 712, removable hard drive or solid state drive 713, legacy magnetic media such as tape and floppy disk, and human-accessible storage devices and associated media such as specialized ROM / ASIC / PLD-based devices such as security dongles.

[0114]

[0135] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.

[0115]

[0136] Computer system 700 may also include an interface 715 to one or more communication networks 714. Network 714 can be, for example, wireless, wired, or optical. Network 714 can further be related to local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of network 714 include Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital networks for TV (including cable TV, satellite TV, and terrestrial broadcast TV), vehicle and industrial including CANBus, etc. Certain networks 714 generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus 716 (e.g., the USB port of computer system 700); others are generally integrated into the core of computer system 700 by attaching to a system bus as described below (e.g., an Ethernet interface is integrated within a PC computer system, and a cellular network interface is integrated within a smartphone computer system). Using any of these networks 714, computer system 700 can communicate with other entities. Such communication can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus for certain CANbus devices), or two-way, e.g., for other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0116]

[0137] The aforementioned human interface devices, human accessible storage devices, and network interfaces can be attached to the core 717 of computer system 700.

[0117]

[0138] The core 717 can include one or more central processing units (CPUs) 718, a graphics processing device (GPU) 719, a special programmable processing device in the form of a field programmable gate array (FPGA) 720, a hardware accelerator 721 for specific tasks, etc. These devices can be connected via a system bus 748 together with a read only memory (ROM) 723, a random access memory (RAM) 724, and an internal mass storage device (e.g., an internal user-inaccessible hard drive, SSD, etc.) 722. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the core system bus 748 or attached either way via a peripheral bus 716. The architecture of the peripheral bus includes PCI, USB, etc.

[0118]

[0139] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can be combined to execute specific instructions that can constitute the aforementioned machine code (or computer code). The computer code can be stored in the ROM 723 or RAM 724. Temporary data can also be stored in the RAM 724, while persistent data can be stored, for example, in the internal mass storage 722. Fast storage and retrieval for any memory device may be made possible by using cache memory, which can be closely associated with one or more CPUs 718, GPUs 719, mass storage 722, ROM 723, RAM 724, etc.

[0119]

[0140] A computer-readable medium can have computer code thereon for performing operations implemented on various computers. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well known and available to those having skill in the art of computer software.

[0120]

[0141] By way of illustration and not limitation, a computer system having an architecture 700 of a computer system, and specifically a core 717, can provide a function of executing software embodied in one or more tangible computer-readable media as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable media can be a media related to user-accessible mass storage as described above, as well as specific storage of the core 717 of a non-transitory nature such as the mass storage 722 or ROM 723 inside the core. The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core 717. The computer-readable media can include one or more memory devices or chips according to specific needs. The software includes defining a data structure stored in the RAM 724 and modifying such a data structure according to a process defined by the software, and causing the core 717 and particularly the processor (including a CPU, GPU, FPGA, etc.) therein to execute a specific process or a specific part of a specific process described in the present application. Further or alternatively, the computer system can provide a function as a result of logic wired in a circuit (e.g., accelerator 721) or otherwise embodied, and the logic can execute a specific process or a specific part of a specific process described in the present application instead of or together with the software. References to software include logic and, if necessary, vice versa. References to computer-readable media can include a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or, where appropriate, both. The present disclosure encompasses any suitable combination of hardware and software.

[0121]

[0142] The number and arrangement of the components shown in FIG. 7 are given as an example. In reality, the input human interface device may include additional components, fewer components, different components, or components arranged differently compared to those shown in FIG. 7. Additionally or alternatively, a set of components (e.g., one or more components) of the input human interface device may perform one or more functions described to be performed by another set of components of the input human interface device.

[0122]

[0143] In an embodiment, any of the operations or processes of FIGS. 1-6 and FIGS. 8-15 can be implemented by or using any of the elements shown in FIG. 7.

[0123]

[0144] FIG. 8 shows an exemplary network media delivery system 800 that caters to a plurality of heterogeneous client end points. That is, system 800 supports various legacy and heterogeneous immersive media compliant displays as client end points. System 800 can include a content acquisition module 801, a content preparation module 802, and a transmission module 803.

[0124]

[0145] The content acquisition module 801 captures or creates source media using, for example, the embodiments described in FIGS. 6 and / or 5. The content preparation module 802 then creates an ingestion format to be sent to the network media delivery system using the transmission module 803. The gateway 804 can respond to subscriber premise equipment to provide network access to various client endpoints of the network. The set-top box 805 functions as subscriber premise equipment and provides access to content aggregated by a network service provider. The wireless demodulator 806 can function as a mobile network access point for a mobile device, as shown, for example, with the mobile handset display 813. In this particular embodiment of the system 800, the legacy 2D TV 807 is shown as being directly connected to either the gateway 804, the set-top box 805, or the WiFi (router) 808. The laptop 2D display 809 (i.e., a computer or laptop with a legacy 2D display) is shown as a client endpoint connected to the WiFi (router) 808. The head-mounted 2D (raster-based) display 810 is also connected to the WiFi (router) 808. The lenticular light field display 811 is shown as being connected to either of the gateways 804. The lenticular light field display 811 can include one or more GPUs 811A, a storage device 811B, and a visual presentation component 811C that creates multiple views using light-ray-based lenticular optics. The holographic display 812 is shown as being connected to the set-top box 805. The holographic display 812 can include one or more CPUs 812A, a GPU 812B, a storage device 812C, and a visualization component 812D.The visualization component 812D may be a Fresnel pattern, a wave-based holographic device / display. The augmented reality (AR) headset 814 is shown as being connected to the wireless demodulator 806. The AR headset 814 can include a GPU 814A, a storage device 814B, a battery 814C, and a volumetric visual presentation component 814D. The high-density light field display 815 is shown as being connected to a WiFi (router) 808. The high-density light field display 815 can include one or more GPUs 815A, a CPU 815B, a storage device 815C, an eye tracking device 815D, a camera 815E, and a high-density ray-based light field panel 815F.

[0125]

[0146] The number and arrangement of the components shown in FIG. 8 are given as an example. In practice, the system 800 may include additional components, fewer components, different components, or components arranged differently than those shown in FIG. 8. Additionally or alternatively, a set of components of the system 800 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of a device or an individual display.

[0126]

[0147] FIG. 9 shows an exemplary workflow of an immersive media delivery process 900 that can handle legacy and heterogeneous immersive media compliant displays, as previously described in FIG. 8. The immersive media delivery process 900 executed by the network can provide adaptation information regarding a particular media represented in a media capture format, for example, prior to the process of the network that adapts the media to be consumed by a particular immersive media client endpoint (as described with reference to FIG. 10).

[0127]

[0148] The immersive media delivery process 900 can be divided into two parts: immersive media production on the left side of the dashed line 912 and immersive media network delivery on the right side of the dashed line 912. Immersive media production and immersive media network delivery can be executed by a network or a client device.

[0128]

[0149] Media content 901 is created or acquired by a network (or a client device) or from a content source. The methods of creating or acquiring data may respectively correspond to FIGS. 5 and 6 regarding natural and synthetic content. The created content 901 is then converted into an import format using the network import format creation process 902. The network import creation process 902 may also respectively correspond to FIGS. 5 and 6 regarding natural and synthetic content. Also, the import format can be updated to store information regarding assets that may be reused across multiple scenes, for example, from a media complexity analyzer 911 (to be detailed later with reference to FIGS. 10 and 14A). The import format is sent to the network and stored in the import media storage 903 (i.e., a storage device). In some embodiments, the storage device is on the network of the immersive media content producer and may be remotely accessed for immersive media network delivery 920. Client and application specific information is optionally the client specific information 904 available on a remote storage device. In some embodiments, the client specific information 904 may exist remotely in an alternative cloud network and may be sent to the network.

[0129]

[0150] Next, the network orchestrator 905 is executed. Network orchestration functions as the main source and sink of information for executing the network's major tasks. The network orchestrator 905 may be implemented in a unified format together with other components of the network. The network orchestrator 905 may be a process that further uses a bi - directional message protocol with the client device to facilitate all processing and delivery of media according to the characteristics of the client device. Further, the bi - directional protocol may be implemented across different delivery channels (e.g., the control plane channel and / or the data plane channel).

[0130]

[0151] As shown in FIG. 9, the network orchestrator 905 receives information regarding the functions and attributes of the client device 908. The network orchestrator 905 collects requirements regarding the applications currently operating on the client device 908. This information may be obtained from the client - specific information 904. In some embodiments, the information may be obtained by directly querying the client device 908. When the client device is directly queried, it is assumed that a bi - directional protocol exists and operates so that the client device 908 can communicate directly with the network orchestrator 905.

[0131]

[0152] The network orchestrator 905 can also start and communicate with a media adaptation fragmentation module 910 (described in FIG. 10). When the captured media is adapted and fragmented by the media adaptation fragmentation module 910, the media may be transferred to an intermediate storage device such as the media 909 prepared for distribution. If the network is designed to include a cache for assets that are used multiple times in a presentation scenario, another intermediate storage device redundancy cache for the reusable media asset 912 may be used as a cache for such assets. When the delivery media is prepared and stored in a media storage device prepared for delivery 909, the network orchestrator 905 ensures that the client device 908 can receive the delivery media and the description information 906 through a "push" request, or that the client device 908 can initiate a "pull" request from the stored media prepared for delivery 909 for the delivery media and the description information 906. The information can be "pushed" or "pulled" through the network interface 908B of the client device 908. The "pushed" or "pulled" delivery media and description information 906 may be description information corresponding to the delivery media.

[0132]

[0153] In some embodiments, the network orchestrator 905 uses a bi-directional message interface to execute a "push" request or initiate a "pull" request by the client device 908. The client device 908 may optionally use a GPU 908C (or CPU).

[0133]

[0154] The delivery media format is then stored in the storage device or storage cache 908D included in the client device 908. Finally, the client device 908 visually presents the media through the visualization component 908A.

[0134]

[0155] Through a process of streaming immersive media to client device 908, network orchestrator 905 monitors the status of the client's progress via the client's progress and status feedback channel 907. In some embodiments, the status monitoring may be performed via a bi-directional communication message interface.

[0135]

[0156] FIG. 10 shows, for example, an example of a media adaptation process 1000 executed by media adaptation fragmentation module 910. By executing media adaptation process 1000, the captured source media can be appropriately adapted to meet the requirements of a client (e.g., client device 908).

[0157] As shown in FIG. 10, media adaptation process 1000 includes a plurality of components that facilitate adapting the captured media to an appropriate delivery format for client device 908. The components shown in FIG. 10 should be considered as examples. In practice, media adaptation process 1000 may include additional components, fewer components, different components, or components arranged differently than those shown in FIG. 10. Additionally or alternatively, a set of components (e.g., one or more components) of media adaptation process 1000 may perform one or more functions described as being performed by another set of components.

[0136]

[0158] In FIG. 10, the adaptation module 1001 receives an input network status 1005 to track the current traffic load on the network. As described above, the adaptation module 1001 also receives information from the network orchestrator 905. The information includes descriptions of the attributes and characteristics of the client device 908, descriptions of the application characteristics, the current status of the application, and the client NN model (if available), which may help map the frustum geometry of the client to the interpolation function of the capture immersive media. Such information can be obtained by using a bi-directional message interface. The adaptation module 1001 ensures that the adapted output is stored in a storage device for storing the client adaptation media 1006 (if it is created).

[0137]

[0159] The media complexity analyzer 911 can be considered as part of the network automation process for media delivery or as an optional process that may be executed in advance. The media complexity analyzer 911 can save the capture media format and assets to a storage device (1002). Then, the capture media format and assets can be transmitted from the storage device (1002) to the adaptation module 1001.

[0138]

[0160] The adaptation module 1001 can be controlled by a logic controller 1001F. The adaptation module 1001 can use a renderer 1001B or a processor 1001C to adapt a specific capture source media to a format suitable for the client. The processor 1001C may be an NN-based processor. The processor 1001C uses an NN model 1001A. Examples of such a processor 1001C include a Deepview NN model generator as described in MPI and MSI. Although the media is in a 2D format, if the client needs to have a 3D format, the processor 1001C can initiate a process of deriving a volumetric representation of the scene depicted in the media using highly correlated images from the 2D video signal.

[0139]

[0161] Renderer 1001B may be a software-based (or hardware-based) application or process based on a selective mix of fields related to acoustic physics, optical physics, visual recognition, speech recognition, mathematics, and software development, which, when given an input scene graph and an asset container, delivers (typically) visual and / or auditory signals suitable for presentation on a target device or compliant with desired characteristics specified by the attributes of the render target nodes in the scene graph. In the case of visual-based media assets, the renderer is capable of delivering appropriate visual signals for the target display or for storage such as an intermediate asset (e.g., repackaged into another container and used in a series of rendering processes in a graphics pipeline). In the case of acoustic-based media assets, the renderer is capable of delivering audio signals for presentation on a multi-channel loudspeaker or binauralized headphones or for repackaging into another (output) container. The renderer includes, for example, the real-time rendering capabilities of source and cross-platform game engines. The renderer may include a scripting language (i.e., an interpreted type programming language) that may be executed by the renderer at runtime to process dynamic inputs and variable state changes applied to the scene graph nodes. Dynamic inputs and variable state changes may affect the rendering and evaluation of spatial and temporal object topologies (including physical forces, constraints, inverse kinematics, deformations, collisions), and the energy propagation and transmission (light, sound). The evaluation of spatial and temporal object topologies produces a result, which in turn produces a result that moves the output from the abstract to the concrete (e.g., similar to the evaluation of a document object model of a web page).

[0140]

[0162] The renderer 1001B may be, for example, a modified version of the OTOY Octane renderer that is modified to communicate directly with, for example, the adaptation module 1001. In some embodiments, the renderer 1001B implements computer graphics techniques (e.g., path tracing) for rendering a three-dimensional scene such that the lighting of the scene is realistic. In some embodiments, the renderer 1001B may be a shader (i.e., a type of computer program that was originally used for shading (generating appropriate levels of light, darkness, and color within an image), but now performs various special functions in various fields of computer graphics special effects, performs post-processing of videos not related to shading, or performs other functions not related to graphics).

[0141]

[0163] The adaptation module 1001 can compress and decompress media content by using a media compression unit 1001D and a media decompression unit 1001E, respectively, according to the need for compression and decompression, based on the format of the captured media and the format required by the client device 908. The media compression unit 1001D may be a media encoder, and the media decompression unit 1001E may be a media decoder. After performing (if necessary) compression and decompression, the adaptation module 1001 outputs client-adapted media 1006 that is optimal for streaming or distribution to the client device 908. The client-adapted media 1006 can be stored in a storage device that stores the adapted media.

[0142]

[0164] FIG. 11A shows an exemplary delivery format creation process 1100. As shown in FIG. 11A, the delivery format creation process 1100 includes a media adaptation module 1101 and an adaptive media packaging module 1103 that package the media output from the media adaptation process 1000 and store it as client-adapted media 1006. The media packaging module 1103 formats the adapted media from the client-adapted media 1006 into a robust delivery format 1104. The delivery format may be, for example, the exemplary format shown in FIG. 3A or FIG. 4A. The information manifest 1104A can provide a list of scene data assets 1104B to the client device 908. The list of scene data assets 1104B can also include metadata that describes the complexity with which each asset is used across the set of scenes that make up the presentation. The list of scene data assets 1104B represents a list of visual assets, auditory assets, and tactile assets, each having its respective corresponding metadata. In this exemplary embodiment, each asset in the list of scene data assets 1104B references metadata that includes a numerical complexity scale value that characterizes the amount of time and resources that the client will expend to process the asset across all the scenes that make up the presentation.

[0143]

[0165] Figure 11B shows the creation of a delivery format using the ordered complexity process 1110. The creation of the delivery format using the ordered complexity process 1110 includes components similar to those in Figure 11A. Therefore, repetitive descriptions related to Figure 11B are omitted here. The media packaging module 11043 is similar to the media packaging module 1103. However, the media packaging module 11043 orders the assets within the list of scene data assets 11044B based on a complexity metric. For example, the visual assets within the list of scene data assets 11044B can be ordered in ascending order of complexity, while the auditory and tactile assets within the list of scene data assets 1104B can be ordered in descending order of complexity, and vice versa. As shown in Figure 11, the assets of the delivery format 11044 are ordered first by asset type and then by the complexity with which the asset is used throughout the presentation, for example, in ascending or descending order of the complexity value based on the asset type (i.e., visual, tactile, auditory, etc.). The information manifest 11044A provides the client device 908 with a list of scene data assets 11044B that can be expected to be received and optional metadata indicating the complexity with which all assets are used across the set of scenes including the entire presentation. The list of scene data assets 11044B includes lists of visual assets, auditory assets, and tactile assets, each having their respective corresponding metadata.

[0144]

[0166] The media may be further packetized before streaming. FIG. 12 shows an exemplary packetization process 1200. The packetization system 1200 includes a packetization unit 1202. The packetization unit 1202 may receive a list of scene - data assets 1104B (or 11044B) as input media 1201 (shown in FIG. 12). In some embodiments, the client - adaptive media 1006 or the delivery format 1104 is input to the packetization unit 1202. The packetization unit 1202 separates the input media 1201 into individual packets 1203 suitable for streaming and presentation to a client device 908 on the network.

[0145]

[0167] FIG. 13 is a sequence diagram showing an example of data and communication flow between components according to an embodiment. The sequence diagram of FIG. 13 is of a network that adapts a particular immersive media in a capture format to a delivery format that is streamable and suitable for a particular immersive - media client endpoint. The data and communication flow may be as follows.

[0146]

[0168] Client device 908 initiates media request 1308 to network orchestrator 905. In some embodiments, the request may be made to the network delivery interface of the client device. Media request 1308 includes information for identifying the media requested by client device 908. The media request may be identified, for example, by a Uniform Resource Name (URN) or another standard nomenclature. The network orchestrator 905 then responds to media request 1308 using profile request 1309. Profile request 1309 requests that the client provide information regarding the resources currently available to the client (including computing, memory, battery charge rate, and other information characterizing the current operating state of the client). Profile request 1309 also requests that the client provide one or more NN models that may be used on the network for NN inference to extract or interpolate the correct media view that matches the capabilities of the client's presentation system, if such NN models are available at the client's endpoint.

[0147]

[0169] Next, following response 1310 from client device 908 to network orchestrator 905, client device 908 provides a client token, an application token, and one or more NN model tokens (if such NN model tokens are available at the client endpoint). Next, network orchestrator 905 provides a session ID token 1311 to the client device. Network orchestrator 905 then requests ingest media 1312 from ingest media server 1303. Ingest media server 1303 can include, for example, ingest media storage 903 or an ingest media format and asset storage device 1002. The request for ingest media 1312 can also include the URN or other standard name of the media specified in request 1308. Ingest media server 1303 responds to the ingest media 1312 request with a response 1313 that includes an ingest media token. Network orchestrator 905 then provides the media token from response 1313 to client device 908 in call 1314. Network orchestrator 905 then starts the adaptation process of the requested media in request 1315 by providing the ingest media token, the client token, the application token, and the NN model token to adaptation fragmentation module 910. Adaptation fragmentation module 910 requests access to the ingest media by providing the ingest media token to ingest media server 1303 in request 1316 and requesting access to the ingest media asset.

[0148]

[0170] In response 1317 to the adaptive fragmentation module 910, the capture media server 1303 responds to request 1316 using the capture media access token. Next, the adaptive fragmentation module 910 requests that the media adaptation process 1000 adapt the capture media found in the capture media access token regarding the client, application, and NN inference model corresponding to the session ID token created and sent in response 1313. A request 1318 is made from the adaptive fragmentation module 910 to the media adaptation process 1000. Request 1318 includes the necessary tokens and session ID. The media adaptation process 1000 provides the network orchestrator 905 with a compliant media access token and session ID in update response 1319. Next, the network orchestrator 905 provides the media packaging module 11043 with a compliant media access token and session ID in interface call 1320. The media packaging module 11043 provides a response 1321 to the network orchestrator 905 using the packaged media access token and session ID in response 1321. Next, the media packaging module 11043 provides the packaged media access token of the packaged asset, URN, and session ID to the packaged media server 1307 for storage in response 1322. Thereafter, the client device 908 executes a request 1323 against the packaged media server 1107 and starts streaming the media asset corresponding to the packaged media access token received in response 1321. Finally, the client device 908 executes other requests and provides a status update in message 1324 to the network orchestrator 905.

[0149]

[0171] Figure 14A shows the workflow of the media complexity analyzer 911 shown in FIG. 9. The media complexity analyzer 911 analyzes metadata regarding the uniqueness of objects in a scene included in media data.

[0150]

[0172] In process 1401, it is possible to start initialization. In process 1402, media asset data is potentially read. In 1403, it is determined whether the media asset data has been successfully read. If the media asset data has not been successfully read, the process ends in 1409. However, if the media asset data has been successfully read, the process continues to 1404. In 1404, a read or data search is executed, and complexity attributes are obtained from the media asset data. Next, the data obtained in 1404 is analyzed to access the attributes that describe the media asset. Each attribute is provided as an input to process 1405. In 1405, the data from 1404 is inspected to determine whether it is any of a list of pre-specified complexity attributes 1410. Such complexity attributes 1410 are the size of the object (which affects the amount of storage required to process the object); the number of polygons for the object, which may be an indicator of the amount of processing required by the GPU; the fixed point _vs._ floating point numerical representation, which may be an indicator of the amount of processing required by the GPU or CPU; the bit depth, which may be an indicator of the amount of processing required by the GPU or CPU; the single precision _vs._ double precision floating point numerical representation, which may be an indicator of the size of the data values during processing; The existence of a light distribution function that may indicate that the scene needs to do something in the process of modeling the physics of how light is distributed in the scene; The type of light distribution function that may indicate the complexity of the light distribution function for modeling the physics of light; A conversion process that may be required when the complexity of the way an object is placed (by rotation, translation, and scaling) within a scene is indicated; May include. If the attribute is a complexity attribute, the process proceeds to 1406; otherwise, it moves to 1407. In 1406, obtain the value of the complexity attribute and store the value in the complexity summary area of the media asset. In 1407, it is possible to determine whether there are further attributes to be read from the media asset. If there are no further attributes to be read, the process proceeds to 1408, and write the complexity summary of the object to the area identified for storing the complexity data of the scene containing the object.

[0151]

[0173] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, it will be recognized that those skilled in the art will be able to devise numerous systems and methods that embody the principles of the present disclosure and thus fall within its spirit and scope, although not explicitly illustrated or described herein.

[0152]

[0174] <Appendix> (Appendix 1) A media packaging method executed by at least one processor to optimize media delivery in a media streaming network, comprising: A step in which a media streaming server receives an immersive media stream including one or more immersive media assets related to one or more scenes; Identifying a subset of the one or more immersive media assets that includes the essential elements of the individual scenes in the one or more scenes; Ordering the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes the essential elements of the individual scenes; and Streaming the one or more immersive media assets from the media streaming server to the client device in the ordered sequence; A method comprising.

[0153] (Appendix 2) In the method according to Appendix 1, further: Determining individual complexity metrics and individual asset types associated with the one or more immersive media assets based on a determination that at least one of the one or more individual attributes is present in a complexity attribute list; and Ordering the one or more immersive media assets in a sequence based on the individual complexity metrics and individual asset types associated with the one or more immersive media assets; A method comprising.

[0154] (Appendix 3) In the method according to Appendix 2, the sequence of the one or more immersive media assets is first ordered by the asset type and then by the individual complexity metrics, method.

[0155] (Appendix 4) In the method according to Appendix 2, the method further: Based on a determination that more than one of the one or more individual attributes is present in the complexity attribute list, the media streaming server performs an asset complexity analysis related to the one or more scenes; A method including

[0156] (Appendix 5) In the method according to Appendix 2, the one or more individual attributes related to the one or more immersive media assets include at least one of the size of the immersive media asset, the number of polygons in the immersive media asset, the bit depth, or the type of conversion required.

[0157] (Appendix 6) In the method according to Appendix 5, the individual complexity metrics associated with each of the one or more immersive media assets are based on one or more individual attributes associated with each of the one or more immersive media assets.

[0158] (Appendix 7) In the method according to Appendix 5, the individual complexity metrics associated with the immersive media assets among the one or more immersive media assets are based on one or more individual attributes of the immersive media assets that exist in the complexity attribute list.

[0159] (Appendix 8) A media streaming server for optimizing media delivery in a media streaming network, comprising: At least one memory configured to store computer program code; and At least one processor configured to read the computer program code and operate as instructed by the computer program code; Including, the computer program code is: First receiving code configured to cause the at least one processor to execute a step of receiving an immersive media stream including one or more immersive media assets related to one or more scenes; Identification code configured to cause the at least one processor to identify a subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes; Sequencing code configured to cause the at least one processor to sequence the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes essential elements of the individual scenes; and Streaming code configured to cause the at least one processor to stream the one or more immersive media assets from the media streaming server to the client device in the sequenced sequence; A media streaming server comprising.

[0160] (Appendix 9) In the media streaming server according to Appendix 8, the computer program code further: Determination code configured to cause the at least one processor to determine individual complexity metrics and individual asset types associated with the one or more immersive media assets based on a determination that at least one of the one or more individual attributes is present in a complexity attribute list; A media streaming server, wherein the sequence is sequenced based on individual complexity metrics and individual asset types associated with the one or more immersive media assets.

[0161] (Appendix 10) In the media streaming server according to Appendix 9, the sequence of the one or more immersive media assets is first sequenced by the asset type and then by individual complexity metrics.

[0162] (Appendix 11) In the media streaming server according to Appendix 9, the computer program code further: Generating code configured to cause the at least one processor to perform a step of performing an asset complexity analysis related to the one or more scenes based on a determination that more than one of the one or more individual attributes are present in the complexity attribute list; A media streaming server comprising:

[0163] (Appendix 12) In the media streaming server according to Appendix 9, the one or more individual attributes related to the one or more immersive media assets include at least one of the size of the immersive media asset, the number of polygons in the immersive media asset, the bit depth, or the type of conversion required. A media streaming server.

[0164] (Appendix 13) In the media streaming server according to Appendix 12, the individual complexity metrics associated with each of the one or more immersive media assets are based on one or more individual attributes associated with each of the one or more immersive media assets. A media streaming server.

[0165] (Appendix 14) In the media streaming server according to Appendix 12, the individual complexity metrics associated with the immersive media assets among the one or more immersive media assets are based on one or more individual attributes of the immersive media assets present in the complexity attribute list. A media streaming server.

[0166] (Appendix 15) A non-transitory computer-readable storage medium storing instructions, which, when executed by at least one processor of a media streaming server that packages media to optimize media delivery in a media streaming network, cause the at least one processor to: Receive an immersive media stream including one or more immersive media assets related to one or more scenes; Identify a subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes; Order the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes the essential elements of the individual scenes; and Stream the one or more immersive media assets in the ordered sequence from the media streaming server to a client device. A non-transitory computer-readable storage medium that causes the above to be performed.

[0167] (Appendix 16) In the non-transitory computer-readable storage medium according to Appendix 15, the instructions further cause the at least one processor to: Execute a step of determining individual complexity metrics and individual asset types related to the one or more immersive media assets based on a determination that at least one of one or more individual attributes exists in a complexity attribute list; A non-transitory computer-readable storage medium, wherein the sequence is ordered based on the individual complexity metrics and individual asset types related to the one or more immersive media assets. (Appendix 17) In the non-transitory computer-readable storage medium according to Appendix 16, the sequence of the one or more immersive media assets is first ordered by the asset type and then ordered by individual complexity metrics, a non-transitory computer-readable storage medium.

[0168] (Appendix 18) In the non-transitory computer-readable storage medium according to Appendix 16, the instructions further cause the at least one processor to: Based on a determination that more than one of the one or more individual attributes are present in the complexity attribute list, cause the media streaming server to perform a step of performing an asset complexity analysis related to the one or more scenes, a non-transitory computer-readable storage medium.

[0169] (Appendix 19) In the non-transitory computer-readable storage medium according to Appendix 15, the one or more individual attributes related to the one or more immersive media assets include at least one of the size of the immersive media asset, the number of polygons in the immersive media asset, the bit depth, or the type of transformation required, a non-transitory computer-readable storage medium.

[0170] (Appendix 20) In the non-transitory computer-readable storage medium according to Appendix 19, the individual complexity metrics associated with each of the one or more immersive media assets are based on one or more individual attributes associated with each of the one or more immersive media assets, a non-transitory computer-readable storage medium.

Claims

1. A media packaging method executed by at least one processor to optimize media delivery in a media streaming network, comprising: receiving, by a media streaming server, an immersive media stream including one or more immersive media assets associated with one or more scenes; identifying a subset of the one or more immersive media assets that includes essential elements of individual scenes in the one or more scenes, wherein the essential elements are elements that visually, auditorily, or physically characterize a particular scene that is spatially or temporally limited; ordering the one or more immersive media assets in a sequence based on the identified subset of the one or more immersive media assets that includes essential elements of the individual scenes; and streaming the one or more immersive media assets from the media streaming server to a client device in the ordered sequence; including, wherein the ordering step comprises: determining individual complexity metrics and individual asset types associated with the one or more immersive media assets; and ordering the one or more immersive media assets by using a sequence determined based on the individual complexity metrics and individual asset types associated with the one or more immersive media assets; including, a method.

2. The method according to claim 1, wherein the determining step is performed based on a determination that at least one of one or more individual attributes is present in a complexity attribute list. A method.

3. The method according to claim 2, wherein the sequence of the one or more immersive media assets is first ordered by the asset type and then by individual complexity metrics.

4. The method according to claim 2, wherein the method further comprises: Based on a determination that more than one of the one or more individual attributes are present in the complexity attribute list, the media streaming server performing an asset complexity analysis for the one or more scenes; A method comprising the steps of:

5. The method according to claim 2, wherein the one or more individual attributes associated with the one or more immersive media assets include at least one of the size of the immersive media asset, the number of polygons in the immersive media asset, the bit depth, or the type of transformation required.

6. The method according to claim 5, wherein the individual complexity metrics associated with each of the one or more immersive media assets are based on one or more individual attributes associated with each of the one or more immersive media assets.

7. The method according to claim 5, wherein the individual complexity metrics associated with an immersive media asset among the one or more immersive media assets are based on one or more individual attributes of the immersive media asset that are present in the complexity attribute list.

8. A media streaming server for optimizing media delivery in a media streaming network, comprising: At least one memory configured to store computer program code; and At least one processor configured to read the computer program code and operate as instructed by the computer program code; A media streaming server including the same, wherein the computer program code causes the at least one processor to execute the method according to any one of claims 1 to 7.

9. A computer program that causes a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tile selection and bandwidth optimization for delivering 360° immersive video

    JP2021527356A

  • Video quality and throughput control

    US20190191168A1

  • Systems and Methods For Content Delivery Acceleration of Virtual Reality and Augmented Reality Web Pages

    US20190238648A1

  • System and method for editing video contents automatically technical field

    US20190364211A1

  • Texture based immersive video coding

    WO2021211173A1