Methods for packaging media, media streaming servers, and media.
By performing complexity analysis on immersive media assets and optimizing the streaming sequence, the problem of resource waste in immersive media distribution is solved, achieving efficient media distribution and network performance improvement, and adapting to diverse client devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, the network distribution of immersive media suffers from wasted network and computing resources due to multiple conversions and streaming, resulting in latency and inefficiency. In particular, the diversity and complexity of client devices make it impossible to efficiently distribute immersive media.
By performing complexity analysis on immersive media assets through a media streaming server, necessary elements are identified, the media streaming sequence is optimized, unnecessary conversions and streaming are reduced, and conversion and streaming tasks are rationally allocated by utilizing network and client caching mechanisms.
It improves the efficiency and network performance of immersive media distribution, reduces resource waste, optimizes the media distribution process, and adapts to the needs of diverse client devices.
Smart Images

Figure CN116711285B_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application is based on and claims the priority interests of U.S. Provisional Patent Application No. 63 / 276,545, filed November 5, 2021, and U.S. Patent Application No. 17 / 971,037, filed October 21, 2022, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0002] This disclosure generally describes embodiments of the architecture, structure, and components of systems and networks for distributing media, including video, audio, geometric (3D) objects, haptic feedback, associated metadata, or other content for client-side rendering devices. Some embodiments relate to systems, structures, and architectures for distributing media content to heterogeneous immersive and interactive client-side rendering devices. Background Technology
[0003] Immersive media generally refers to media that stimulates any or all human sensory systems (e.g., vision, hearing, tactile sensation, smell, and possibly taste) to create or enhance the user's perception of physically being present in the media experience; that is, beyond the perception of timing two-dimensional (2D) video and corresponding audio distributed on existing (e.g., "traditional") commercial networks; such timing media is also known as "traditional media." Immersive media can also be defined as media that attempts to create or mimic the physical world through digital simulations of dynamics and physical laws, thereby stimulating any or all human sensory systems to create the user's perception of physically being present within a scene depicting a real or virtual world.
[0004] Immersive media-enabled presentation devices refer to devices equipped with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices support a wide variety of media quantities and formats, and also support the diverse network resources required for large-scale distribution of immersive media. "Large-scale" can refer to media distribution by service providers over the network comparable to traditional video and audio media distribution, such as Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.
[0005] In contrast, traditional presentation devices such as laptop monitors, televisions, and mobile phone displays are functionally homogeneous (i.e., identical in function) because all of these devices include rectangular displays that use 2D rectangular video or still images as their primary visual media format. Some visual media formats commonly used in traditional presentation devices can include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266.
[0006] The distribution of any media over a network can employ a media delivery system and architecture that reformats and / or transforms the media from an input format or a network-ingested media format into a distribution media format. This distribution media format is not only suitable for ingestion by the target client devices and their applications but also facilitates "streaming" over the network. Reformatting or streaming can be performed by the network (e.g., a server in a media streaming network), that is, before distributing the media to clients, a media format known as a "distribution media format" or simply a "distribution format" is generated.
[0007] In related technologies, when a network can access information to indicate that a client will require converted media objects (also known as media assets) and / or streamed media objects on multiple occasions, this multiple use will trigger such media conversion and streaming multiple times. In other words, this continuous reprocessing and transmission of data used for media conversion and streaming is a source of latency within the network, leading to a potentially significant increase in the amount of network and / or computing resources being used.
[0008] In contrast, a network design that can access information to indicate when a client may already have a specific media data object stored in its cache or relative to the client's local storage will be more efficient than a network that can only access such information. Therefore, it may be necessary to include network designs that include access to information indicating when a client can locally store a media object in its cache. Summary of the Invention
[0009] According to embodiments, methods, systems, and apparatus are provided to facilitate the computation of the sequence order in which assets are packaged from a network and streamed to a client. A complexity analyzer analyzes media assets including the necessary elements of a scene to determine which assets in a particular scene will take the most time to process. The order in which assets for a particular scene are packaged and streamed to the client is based on the complexity of each asset in that particular scene.
[0010] According to one aspect of this disclosure, a method for packaging media to optimize media distribution in a media streaming network can be provided. The method may include: a media streaming server receiving an immersive media stream comprising one or more immersive media assets associated with one or more scenes; identifying a subset of the one or more immersive media assets, the subset comprising necessary elements of a corresponding scene in the one or more scenes; sorting the one or more immersive media assets in sequence based on the identified subset, the subset comprising the necessary elements of the corresponding scene in the one or more scenes; and streaming the one or more immersive media assets from the media streaming server to a client device in the sorted order.
[0011] According to another aspect of this disclosure, an apparatus (or device) for optimizing media distribution in a media streaming network can be provided. The apparatus may include at least one memory configured to store computer program code; and at least one processor configured to read the computer program code and execute it according to the instructions of the computer program code. The computer program code may include: first receiving code configured to cause the at least one processor to: receive an immersive media stream comprising one or more immersive media assets associated with one or more scenes; identification code configured to cause the at least one processor to: identify a subset of the one or more immersive media assets, the subset comprising necessary elements of a corresponding scene in the one or more scenes; sorting code configured to cause the at least one processor to: sort the one or more immersive media assets in order based on the identified subset of the one or more immersive media assets, the subset comprising the necessary elements of a corresponding scene in the one or more scenes; and streaming code configured to cause the at least one processor to: stream the one or more immersive media assets from the media streaming server to a client device in the sorted order.
[0012] According to another aspect of this disclosure, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor of a device for packaging media to optimize media distribution in a media streaming network, cause the at least one processor to: receive an immersive media stream comprising one or more immersive media assets associated with one or more scenes; identify a subset of the one or more immersive media assets, the subset comprising necessary elements of a corresponding scene in the one or more scenes; sort the one or more immersive media assets in order based on the identified subset of the one or more immersive media assets, the subset comprising the necessary elements of a corresponding scene in the one or more scenes; and stream the one or more immersive media assets from the media streaming server to a client device in the sorted order.
[0013] Additional embodiments will be set forth in the following description and will be apparent in part from the description, and / or may be implemented by practice of the embodiments presented in this disclosure. Attached Figure Description
[0014] Figure 1A This is an exemplary illustration of a media distribution streaming network according to an embodiment.
[0015] Figure 1B This is an exemplary workflow according to an embodiment, which illustrates the creation of media in a distribution format and the generation of one or more reuse indicators in a media streaming network.
[0016] Figure 2A This is an exemplary workflow according to an embodiment, which illustrates the streaming of media to a client device.
[0017] Figure 2B This is an exemplary workflow according to an embodiment, which illustrates the streaming of media to a client device.
[0018] Figure 3A This is an exemplary illustration of a data model for the representation and streaming of timed immersive media according to an embodiment.
[0019] Figure 3B This is an exemplary illustration of a data model for the representation and streaming of timed immersive media according to an embodiment.
[0020] Figure 4A This is an exemplary illustration of a data model for the representation and streaming of untimed immersive media according to an embodiment.
[0021] Figure 4BThis is an exemplary illustration of a data model for the representation and streaming of intermittently immersive media according to an embodiment.
[0022] Figure 5 An exemplary workflow for natural media synthesis is illustrated according to an embodiment.
[0023] Figure 6 An exemplary workflow for creating synthetic media ingestion is illustrated according to an embodiment.
[0024] Figure 7 This is an exemplary illustration of a computer system according to an embodiment.
[0025] Figure 8 This is an exemplary illustration of a network media distribution system according to an embodiment.
[0026] Figure 9 An exemplary workflow for immersive media distribution is illustrated according to an embodiment.
[0027] Figure 10 This is a system diagram of the media adaptation process system according to an embodiment.
[0028] Figure 11A The illustrations illustrate an exemplary workflow for creating media in a distribution format, according to an embodiment.
[0029] Figure 11B The illustrations, based on embodiments, depict an exemplary workflow for creating media sequentially in a distribution format.
[0030] Figure 12 An exemplary workflow of the packaging process is illustrated according to an embodiment.
[0031] Figure 13 This is an exemplary workflow illustrating the communication process between components according to an embodiment.
[0032] Figure 14A The illustration illustrates an exemplary workflow for immersive media complexity analysis according to an embodiment.
[0033] Figure 14B These are examples of a set of complexity attributes of a scene presented according to an embodiment. Detailed Implementation
[0034] The following detailed description of the example embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0035] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. As will be apparent from the foregoing disclosure, modifications and variations are possible, or modifications and variations may arise from practice of the embodiments. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be interchanged.
[0036] Clearly, the systems and / or methods described herein can be implemented in various forms of hardware, software, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions herein.
[0037] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly refer to only one claim, the disclosure of possible embodiments includes combinations of each dependent claim with every other claim in the claim set.
[0038] The features discussed below can be used individually or in any combination in any order. Furthermore, embodiments can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.
[0039] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” If only one item is intended to be used, the term “one” or similar wording is used. Additionally, as used herein, the terms “has,” “have,” “having,” “include,” “including,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the word “based on” means “at least partially based on.” Furthermore, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.
[0040] According to embodiments, a presentation device with immersive media capabilities can refer to a device equipped with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices are heterogeneous in terms of the number and formats of media they can support, as well as the quantity and type of network resources required for large-scale distribution of such media. "Large-scale" can refer to media distribution over a network by a service provider comparable to the distribution of traditional video and audio media, such as Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions.
[0041] According to embodiments, client devices that act as endpoints for distributing immersive media over a network are highly diverse. The distribution of any media over a network can employ a media delivery system and architecture that reformats the media from an input or network-ingested media format to a distribution media format, wherein this distribution media format is not only suitable for ingestion by the target client device and its applications but also facilitates "streaming" over the network. Therefore, there can be two processes performed by the network on the ingested media: 1) converting the media from format A to format B suitable for ingestion by the target client (i.e., based on the client's ability to ingest certain media formats), and 2) preparing the media for streaming.
[0042] In this embodiment, streaming media broadly refers to the fragmentation and / or packetizing of media, such that the media can be transmitted over a network in logically organized and ordered consecutive smaller chunks according to one or both of the media's temporal or spatial structure. Converting media from format A (sometimes called "transcoding") to format B can be a process typically performed by the network or service provider before distributing the media to client devices. This transcoding may include converting media from format A to format B based on prior knowledge that format B is the preferred or unique format, which may be ingested by the target client device or is more suitable for distribution on limited resources such as commercial networks. In many, but not all, cases require both media conversion and preparation for streaming before the target client device can receive and process the media from the network.
[0043] Transforming (or converting) media and preparing it for streaming is part of the process performed by the network on the ingested media before it is distributed to client devices. The result of this process (i.e., transformation and preparation for streaming) is a media format known as the distribution media format, or simply the distribution format. If these operations are performed on a given media data object, and if the network has access to information indicating that the client will need a transformed and / or streaming media object for multiple occasions, these operations should only be performed once; otherwise, such media transformation and streaming will be triggered multiple times. That is, the processing and transmission of data for media transformation and streaming are generally considered sources of latency due to the potentially large amount of network and / or computing resources required. Therefore, a network design that cannot access information indicating when a client may already have a particular media data object stored in its cache or relative to the client's local storage will be suboptimal (i.e., the network design is not optimal) than one that can access such information.
[0044] A scene graph can be a general data structure commonly used by vector-based graphics editing applications and modern computer games, which arranges the logical and usually (but not necessarily) spatial representation of a graphical scene, or it can be a collection of nodes and vertices in a graphical structure.
[0045] In the context of computer graphics, a scene can be a collection of objects (e.g., 3D assets—also referred to as media assets, media objects, objects, and assets). Objects include essential elements of media data, object attributes, and additional metadata that describes the visual, auditory, and physical characteristics of a particular setting, which is spatially or temporally constrained with respect to the interaction of objects within that setting.
[0046] Nodes can be basic elements of a scene graph, including information related to the logical, spatial, or temporal representation of visual, auditory, tactile, olfactory, gustatory, or related processing information; each node should have at most one output edge, zero or more input edges, and at least one (input or output) edge connected to it.
[0047] The base layer can be a nominal representation of media assets, and is typically designed to minimize the computational resources or time required to render the assets, or the time required to transmit the assets over a network.
[0048] An enhancement layer can be a set of information that, when applied to the base layer representation of an asset, enhances the base layer to include features or capabilities not supported in the base layer.
[0049] Attributes can be metadata associated with a node, used to describe a specific characteristic or feature of that node in a canonical or more complex form (e.g., from the perspective of another node).
[0050] Containers can be serialized formats used to store and exchange information to represent all natural scenes, all composite scenes, or a mixture of composite and natural scenes, including scene graphs and all media resources required for rendering scenes.
[0051] Serialization can be the process of converting the state of a data structure or object into a format that can be stored (e.g., in a file or storage buffer) or transmitted (e.g., via a network connection link) and subsequently (potentially in a different computing environment). When the resulting bit sequence is reread according to the serialization format, that bit sequence can be used to create a semantically identical clone of the original object.
[0052] A renderer can be a (typically software-based) application or process based on a selective mix of disciplines involving acoustic physics, optical physics, visual perception, auditory perception, mathematics, and software development. Given an input scene graph and asset containers, the renderer emits typical visual and / or audio signals suitable for rendering on a target device or conforming to the desired performance specified by the properties of the rendering target nodes in the scene graph. For visual-based media assets, the renderer may emit visual signals suitable for target display or storage as intermediate assets (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline); for audio-based media assets, the renderer may emit audio signals for rendering in multi-channel speakers and / or dual-audio headphones, or for repackaging into another (output) container. Common examples of renderers include the real-time rendering capabilities of game engines Unity and Unreal Engine.
[0053] Scripting languages can be interpreted programming languages that can be executed by the renderer at runtime to handle dynamic inputs and variable state changes to scene graph nodes. This affects the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, inverse kinematics, deformation, and collisions) as well as energy propagation and transmission (light, sound).
[0054] A shader can be a computer program that was originally used for coloring (producing the appropriate levels of light, dark and color in an image), but now performs various specialized functions in all areas of computer graphics effects, or performs video post-processing unrelated to coloring, or even functions completely unrelated to graphics.
[0055] Path tracing is a computer graphics method used to render 3D scenes so that the lighting of the scene is faithful to reality.
[0056] Timed media can include media and / or media objects that can be ordered by time; for example, having start and end times based on a specific clock. Untimed media can include media and / or media objects that can be organized by spatial, logical, or temporal relationships; for example, as in an interactive experience implemented based on actions taken by one or more users.
[0057] A neural network model (NN Model) can be a collection of parameters and tensors (such as matrices) that define the weights (i.e., numerical values) used in well-defined mathematical operations applied to a visual signal to obtain an improved visual output that may include interpolations of new views of the visual signal not explicitly provided by the original signal.
[0058] Over the past decade, the number of devices with immersive media capabilities introduced to the consumer market has surged, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and game consoles. Furthermore, holographic displays and other forms of volumetric displays are poised to enter the consumer market within the next three to five years. However, despite the immediate or imminent availability of these devices, a coherent end-to-end ecosystem for distributing immersive media across commercial networks has failed to materialize for several reasons.
[0059] One reason why a coherent, end-to-end ecosystem for distributing immersive media over commercial networks has not yet been realized is the highly diverse range of client devices used as endpoints in such distribution networks for immersive displays. Some client devices support certain immersive media formats, while others do not. Some can create immersive experiences from traditional raster-based formats, while others cannot. Unlike networks designed solely for distributing traditional media, such networks, which must support the diversity of display clients, require a wealth of information relating to the details of each client's capabilities and the format of the media to be distributed before they can employ an adaptation process to convert media into a format suitable for each target display and corresponding application. At a minimum, such networks need access to information describing the characteristics of each target display and the complexity of the ingested media so that the network can determine how meaningfully to adapt the input media source to a format suitable for the target display and application.
[0060] Networks supporting such diverse client devices should leverage the fact that some assets adapted from input media formats to a specific target format can be reused across a set of similar display targets. In other words, once some assets are converted to a format suitable for the target display, they can be reused among several such displays with similar adaptation requirements. Therefore, networks employing caching mechanisms to store adaptation assets in a relatively immutable region will be more efficient.
[0061] Immersive media can be organized into “scenes” described by scene diagrams, also known as scene descriptions. Scene diagrams can describe visual, auditory, and other forms of immersive assets, including specific settings as part of a presentation, such as actors and events taking place in specific locations within a building as part of a presentation, like in a film. A list of all scenes encompassing a single presentation can be represented as a list of scenes.
[0062] The advantage of a "scenario-based" approach is that, for content prepared before it must be distributed, a "bill of materials" can be created that identifies all assets that will be used throughout the presentation, and how frequently each asset will be used in the various scenarios within the presentation. A network is aware of the existence of cached resources that can be used to satisfy the asset requirements of a specific presentation. Similarly, a client device presenting a series of scenarios might want to know how frequently any given asset will be used in multiple scenarios. For example, if a media asset (also known as a media object, asset, or object) is referenced multiple times in multiple scenarios being or to be processed by a client device, the client device should avoid discarding that asset from its cached resources until the client has presented the last scenario that required that particular asset.
[0063] For traditional presentation devices, the distribution format can be equivalent to or fully equivalent to the "presentation format" that the client presentation device ultimately uses to create the presentation. That is, the presentation media format is a media format whose performance (resolution, frame rate, bit depth, color gamut, etc.) is closely related to the capabilities of the client presentation device. Some examples of distribution and presentation formats include: a high-definition (HD) video signal (1920 pixels in columns x 1080 pixels in rows) distributed over a network to an Ultra-high-definition (UHD) client device with a resolution of (3840 pixels in columns x 2160 pixels in rows). A UHD client can apply a process called "super-resolution" to the HD distribution format to upscale the video signal from HD to UHD. Therefore, the final signal format presented by the client device is the "presentation format," which in this example is the UHD signal, while the HD signal includes the distribution format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format because both signals are rectilinear video formats, and the process of converting HD to UHD is a relatively simple and easy process performed on most traditional client devices.
[0064] However, in some embodiments, the preferred presentation format of the target client device may differ significantly from the ingested format received by the network. Nevertheless, the client device has access to sufficient computing, storage, and bandwidth resources to convert the media from the ingested format to the necessary presentation format suitable for presentation by the client device. The network can bypass reformatting the ingested media from a first format A (e.g., "transcoding" the media) to a second format B because the client has sufficient resources to perform all media conversions without requiring the network to do so beforehand. The network can still segment and package the ingested media so that it can be streamed to the client.
[0065] However, in some embodiments, the ingested media received by the network differs significantly from the client's preferred presentation format, and the client device cannot access sufficient computational, storage, and / or bandwidth resources to convert the media into the preferred presentation format. In cases where resources are unavailable, the network can assist the client by performing some or all of the conversions on its behalf from the ingested format to a format equivalent to or nearly equivalent to the client device's preferred presentation format. In some embodiments, this assistance provided by the network on behalf of the client is often referred to as "split rendering."
[0066] Embodiments of this disclosure as described herein can determine whether a network should convert some or all ingested media from a first format (e.g., format A) to a second format (e.g., format B) to facilitate the ability of client devices to produce media presentation in a potential third format C. This determination can be made by a processor or server of the media streaming network, or it can be made by the client device. The following may be useful in assisting this determination: determining which media assets are used more than once within the context of presentation, and designing processes and / or the network to make those media assets readily available for network use. Based on information from such analysis, the network can then be designed such that it can request client devices (also referred to as “clients”) to retain copies of one or more media assets that can be used more than once in their local cache.
[0067] However, if client devices store copies of media assets in their local caches, the network may have no control over the management of these local caches, potentially leading to situations where client devices must remove resources (even reusable resources) from their local caches. To facilitate a network design that minimizes the need to perform format-to-format conversions on frequently used media assets, or to prevent the network from re-streaming frequently used media assets to clients, the network can manage its own cache, separate from any cache maintained by the client. This ensures that the network has access to at least one redundant copy of each reusable asset for both the client and the network.
[0068] In one embodiment, the network may first query the client device for feedback to ensure that the media asset in question is still available in the client's local cache. If the client device's response indicates that it no longer has a copy of the media asset in question, the network may signal the client to access a copy of the media asset in its distribution format from the redundant cache. In some embodiments, the query to the client device may be omitted, and the network may signal the client to access a copy of the asset's distribution format from the redundant cache.
[0069] Figure 1A This is an exemplary illustration of a media distribution streaming system 100 according to an embodiment for distributing media from a network cloud, edge device, or server 104 to a client device 108. (See also...) Figure 1AAs shown, media in a first format A (hereinafter referred to as "ingestion media format A") is received from a content provider. This media includes immersive media comprising one or more scenes and one or more media objects. This process can be performed or implemented by a network cloud or edge device (hereinafter referred to as "network device 104") and distributed to a client, such as client device 108. In some embodiments, the same process can be performed in advance manually or by the client device. Network device 104 can ingest media 101 in the first format, generate and / or create distribution media 102 in a second format (hereinafter referred to as "distribution media creation 102"), and distribute media 103 in the second format, for example, using a distribution module. Client device 108 may include a rendering module 106 and a presentation module 107.
[0070] According to one aspect, network device 104 can receive ingested media from a content provider, etc. A media streaming network can obtain ingested media stored in ingested media format A. Distribution media can be created and / or generated using any necessary conversions or adjustments to the ingested media to create potential alternative representations of the media. That is, a distribution format for media objects in the ingested media can be created. As described, the distribution format is a media format that can be distributed to clients by formatting the media into distribution format B. Distribution format B is a format prepared for streaming to client device 108. Distribution media creation 102 may include optimized reuse logic to perform a decision process to determine whether a particular media object has already been streamed to client device 108. (See also...) Figure 1B Describe in detail the further operations associated with the distribution media creation 102 and optimized reuse logic.
[0071] Media format A and media format B may or may not be representations that follow the same syntax as a specific media format specification; however, format B may be adapted to facilitate media distribution via a network protocol. The network protocol may be, for example, a connection-oriented protocol (TCP) or a connectionless protocol (UDP). The distribution module streams the streamable media (i.e., media format B) from network device 104 to client device 108 via network connection 105.
[0072] Client device 108 can receive distribution media and can use rendering module 106 to render the media for presentation. Depending on the client device 108 being targeted, rendering module 106 can acquire some rendering capabilities, which can be basic or equally complex. Rendering module 106 can create presentation media in presentation format C. Presentation format C may or may not be represented according to a third format specification. Therefore, presentation format C may be the same as or different from media format A and / or media format B. Rendering module 106 outputs presentation format C to presentation module 107, which can present the presentation media on the display (or similar device) of client device 108.
[0073] Embodiments of this disclosure facilitate a decision-making process adopted by the network to calculate the sequence order in which assets are packaged and streamed from the network to a client. In this context, a media complexity analyzer analyzes all assets utilized in one or more scenes comprising the presentation to determine the complexity associated with each asset throughout all scenes comprising the presentation. Thus, the order in which assets for a particular scene are packaged and streamed to the client can be based on the complexity of using each asset in the set of scenes comprising the presentation.
[0074] The embodiments address the need for a mechanism or process for analyzing immersive media scenarios to obtain sufficient information to support a decision-making process that, when adopted by a network or client, provides indications as to whether the conversion of media objects from format A to format B should be performed entirely by the network, entirely by the client, or a combination of both (and indications of which assets should be converted by the client or network). This immersive media data complexity analyzer can be adopted by the client or network in an automated environment, or manually by a person, such as by an operating system or device.
[0075] According to embodiments, the process of adapting an input immersive media source to a specific endpoint client device can be the same as or similar to the process of adapting the same input immersive media source to a specific application running on a specific client endpoint device. Therefore, the characteristics of adapting an input media source to an endpoint device have the same complexity as the characteristics of adapting a specific input media source to a specific application.
[0076] Figure 1B This describes the workflow for creating distribution media 102 according to the embodiment. More specifically, Figure 1B The workflow includes generating one or more reuse indicators in the media streaming network, which assist the decision-making process in determining whether a particular media object has been streamed to the client device 108.
[0077] At operation 152, the media creation process begins. At operation 155, conditional logic can be executed to determine whether the current media object has previously been streamed to client device 108. A list of unique assets used for rendering can be accessed to determine whether the media object has previously been streamed to the client. If the current media object has previously been streamed, the process proceeds to operation 160. At operation 160, an indicator (hereinafter also referred to as a "proxy") is created to identify that the client has received the current media object and should access a copy of the media object from the local cache or other cache. If it is determined that the media object has not previously been streamed, the process proceeds to operation 165. At operation 165, the media object can be prepared for conversion and / or distribution, and the distribution format of the media object is created. Processing of the current media object then ends.
[0078] Figure 2A This is an exemplary workflow for processing ingested media over a network. According to an embodiment, Figure 2A The workflow illustrated in the diagram depicts the media conversion decision process 200. This process determines whether the network should convert media before distributing it to client devices. The media conversion decision process 200 can be handled through manual or automated procedures within the network.
[0079] The ingested media, represented in Format A, is provided to the network by the content provider. At operation 205, the media streaming network ingests media from the content provider. At operation 210, the attributes of the target client are obtained (if not already known). These attributes describe the processing capabilities of the target client.
[0080] At operation 215, it is determined whether the network (or client) should assist in the conversion of the ingested media. In some embodiments, at operation 215, it may specifically be determined whether any format conversion of any media assets included in the ingested media (e.g., conversion of one or more media objects from format A to format B) is required before the media is streamed to the target client. At operation 215, this determination may be based on whether the media can be streamed in its original ingestion format A, or whether it must be converted to a different format B to facilitate the client's presentation of the media. Such a decision (i.e., determining whether conversion of the ingested media is required before streaming the media to the client, or whether the media should be streamed directly to the client in its original ingestion format A) may require access to information describing aspects or characteristics of the ingested media.
[0081] If it is determined that the network (or client) should assist in the conversion of any media assets ("Yes" at operation 215), then process 200 proceeds to operation 220.
[0082] At operation 220, the ingested media is converted from format A to format B to produce converted media 222. Converted media 222 is output, and at operation 225, the input media undergoes a preparation process for streaming to the client. In this case, converted media 222 (i.e., the input media) is prepared for streaming.
[0083] Streaming immersive media, especially when the media is “scene-based” rather than “frame-based,” can be relatively new. For example, frame-based media streaming can be equivalent to video frame streaming, where each frame captures the entire scene or a complete image of an object to be presented by the client. A video sequence comprising the entire immersive presentation or a portion of it is created when the client reconstructs the frame sequence from a compressed form and presents it to the viewer. For frame-based streaming, the order in which frames are streamed from the network to the client can conform to predefined specifications (e.g., Advanced Video Coding for Universal Audiovisual Services, such as ITU-T Recommendation H.264). However, scene-based media streaming differs from frame-based streaming because a scene can consist of individual assets that can be independent of each other. A given scene-based asset can be used multiple times within a specific scene or across a series of scenes. The amount of time required for a client or any given renderer to reconstruct a particular asset can depend on many factors, including, but not limited to, the size of the asset, the availability of computational resources to perform the rendering, and other properties that describe the overall complexity of the asset. Clients that support scene-based streaming can require partial or full rendering of each asset in a scene before any rendering of the scene can begin. Therefore, the order in which assets are streamed from the network to the client can impact the overall performance of the system.
[0084] The conversion of media from format A to another format (e.g., format B) can be done entirely by the network, entirely by the client, or jointly by both the network and the client. For split rendering, it is clear that a dictionary of attributes describing the media format may be needed so that both the client and the network have complete information to characterize the work that must be done. Furthermore, a dictionary of attributes providing client capabilities (e.g., in terms of available compute resources, available storage resources, and access to bandwidth) may also be needed. Going further, a mechanism is needed to characterize the level of compute, storage, or bandwidth complexity of ingesting the media format so that the network and the client can jointly or individually determine whether or when the network can employ a split rendering process to distribute the media to the client.
[0085] If it is determined that the network (or client) should not (or does not need to) assist in the conversion of any media assets ("No" at operation 215), process 200 proceeds to operation 225. At operation 225, media is prepared for streaming. In this case, ingested data (i.e., media in its original form) is prepared for streaming.
[0086] Finally, once the media data is in a streamable format, the media prepared at operation 225 is streamed to the client at operation 230. In some embodiments, (as referenced) Figure 1B (As described) If the conversion and / or streaming of specific media objects required or to be required by the client for its media presentation can be avoided, assuming the client still has access to or can use the media objects it might need to complete its media presentation, then the network can skip the conversion and / or streaming of ingested media (i.e., operations 215-230). Regarding the order in which scene-based assets are streamed from the network to the client to facilitate the client's full potential, it may be desirable for the network to be equipped with sufficient information to determine such an order to improve client performance. For example, a network with sufficient information to avoid repeated conversion and / or streaming steps of assets used more than once in a particular presentation can perform better than a network not designed in this way. Similarly, a network capable of delivering assets to the client "intelligently" in sequence can facilitate the client's full potential (i.e., creating a potentially more enjoyable experience for the end user).
[0087] Figure 2B An exemplary media conversion process 250, including determining media asset reuse, is illustrated according to an embodiment. Like the media conversion decision process 200, the media conversion process 250, which has a media asset reuse process, ingests media from the network to determine whether the network should convert the media before distributing it to clients.
[0088] The ingested media, represented in Format A, is provided to the network by a content provider. According to one embodiment, similar to... Figure 2A Operations 205-210 and 215-230 are shown to perform operations 255-260 and 275-286. At operation 255, the network retrieves media from the content provider. Subsequently, at operation 260, the attributes of the target client are obtained (if not already known). These attributes describe the processing capabilities of the target client.
[0089] If it is determined that the network has previously streamed a specific media object or the current media object ("Yes" at operation 265), the process proceeds to operation 270. At operation 270, a proxy is created to replace the previously streamed media object, instructing the client to use either a local copy of its previously streamed object or a copy of the previously streamed object stored in another cache.
[0090] If it is determined that the network has not previously streamed any media objects ("No" at operation 265), the process proceeds to operation 275. At operation 275, it is determined whether the network or the client should perform any format conversion on any media assets contained within the media ingested at operation 255. For example, conversion could include changing a specific media object from format A to format B before the media is streamed to the client. Operation 275 could be similar to... Figure 2A The operations performed at position 215 shown.
[0091] If it is determined that the media assets should be converted by the network (Yes at operation 275), the process proceeds to operation 280. At operation 280, the media object is converted from format A to format B. Then, preparations are made to stream the converted media to the client (operation 286).
[0092] If it is determined that the media assets should not be converted over the network ("No" at operation 275), the process proceeds to operation 285. At operation 285, preparation is then made to stream the media object to the client. Once the media is in a streamable format, the media prepared at operation 285 is streamed to the client at operation 286.
[0093] The streaming format of media can be heterogeneous immersive media that is scheduled or not scheduled. Figure 3A An example of a timed media representation 300 in a streamable format for heterogeneous immersive media is shown. Timed immersive media can include a set of N scenes. Timed media is media content ordered by time, for example, having start and end times based on a specific clock. Figure 4A An example of an opaque media representation 400 in a streamable format for heterogeneous immersive media is shown. Opaque media is media content organized by spatial, logical, or temporal relationships (e.g., in an interactive experience implemented based on actions taken by one or more users).
[0094] Figures 3A-3B This refers to timing scenarios used for timed media. Figures 4A-4B This refers to the untimed scenes of untimed media. Timed scenes and untimed scenes can correspond to various scene representations or scene descriptions. Figure 3A , Figure 3B , Figure 4A and Figure 4BAll employ a single, exemplary encompassing media format that has been adapted from the source media format to match a specific client endpoint. In other words, an encompassing media format is a distribution format capable of streaming to client devices. The encompassing media format is structurally robust enough to accommodate a wide variety of media attributes, each of which can be layered based on the significant amount of information each layer contributes to media presentation.
[0095] like Figure 3A As shown, the timed media representation 300 includes a timed scene list 300A, which includes a list of scene information 301. Scene information 301 refers to a list of components 302 that individually describe the processing information and the types of media assets constituting the scene information 301. For example, an asset list and other processing information. The list of components 302 can refer to proxy assets 308 corresponding to the asset type (e.g., proxy visual and audio assets, such as...). Figure 3A (As shown). Component 302 refers to a list of unique assets that have not been used in other scenarios before. For example, Figure 3A The diagram shows a list 307 of unique assets for (timing) scenario 1. Component 302 also refers to asset 303, which includes a base layer 304 and an attribute enhancement layer 305. The base layer is a nominal representation of an asset that can be configured to minimize the computational resources, time required to render the asset, and / or the time required to transmit the asset over a network. In this exemplary embodiment, each of the base layers 304 refers to a numerical complexity metric that indicates the effort a client might need to expend in terms of time or resources to process the asset. The enhancement layer can be a set of information that, when applied to the base layer representation of the asset, enhances the base layer to include features or capabilities that might not be supported in the base layer.
[0096] Figure 3B The timing media representation 3030 is shown, ordered in descending order of complexity. This timing media representation is... Figure 3A The representation shown is the same, however, Figure 3B Assets 3033 in the list are sorted in descending order by asset type and complexity metric within each asset type. The timing scene list 303A includes a list of scene information 3031. The list of scene information 3031 refers to a list of components 3032 that individually describe processing information and include the list of media assets of the type. For example, an asset list and other processing information. Component 3032 may refer to proxy assets 3038 corresponding to the asset type (e.g., proxy visual and audio assets, such as...). Figure 3B(As shown). Component 3032 refers to asset 3033, which further refers to base layer 3034 and attribute enhancement layer 3035. Each base layer 3034 is sorted according to a decreasing value of the corresponding complexity metric. A list 3037 of unique assets that have not been used in other scenarios is also provided.
[0097] like Figure 4A As shown, the timed media and complexity representation 400 includes scene information 401. Scene information 401 is not associated with start and end times / durations (based on clocks, timers, etc.). A list of timed scenes (not depicted) may refer to scene 1.0, for which no other scene can branch off. Scene information 401 refers to a list of components 402 that individually describe the types of information processed and the media assets constituting scene information 401. Components 402 refer to visual assets, audio assets, haptic assets, and timing assets (collectively referred to as assets 403). Assets 403 also refer to the base layer 404 and attribute enhancement layers 405 and 406. In this exemplary embodiment, each of the base layers 404 refers to a numerical complexity value that characterizes the effort a client may need to expend in terms of time or resources to process assets in the scene including the presentation. Scene information 401 may also refer to other timed scenes used for timed media sources (i.e., in...). Figure 4A This is referred to as the untimed scenario (2.1-2.4) and / or used for timed media scenarios (i.e., in...). Figure 4A Scene information 407 (referred to as Timed Scene 3.0) is in the format. Figure 4A In the example, the timed immersive media comprises a set of five scenes (including timed and timed scenes). The list of unique assets 408 identifies unique assets associated with a specific scene that have not been previously used in higher-order (e.g., parent) scenes. Figure 4A The list of unique assets 408 shown includes unique assets for occasional scenario 2.3.
[0098] Figure 4BIrregular media and a sorted complexity representation 4040 are shown. The list of irregular scenes (not depicted) refers to scene 1.0, for which no other scene can branch off. Scene information 4041 is not associated with start and end durations based on a clock. Scene information 4041 also refers to a list of components 4042 that individually describe the types of information processed and the media assets constituting scene information 4041. Components 4042 refer to visual assets, audio assets, haptic assets, and timing assets (also collectively referred to as assets 4043). Assets 4043 also refer to the base layer 4044 and attribute enhancement layers 4045 and 4046. In this exemplary embodiment, haptic assets 4043 are organized by increasing complexity metrics, while audio assets 4043 are organized by decreasing complexity metrics, and vice versa. Furthermore, scene information 4041 refers to other irregular scene information 4041 used for irregular media. Scene information 4041 also refers to other timing scene information 4047 used for timing media. The list of unique assets 4048 identifies unique assets associated with a specific scenario that have not been used in higher-level (e.g., parent) scenarios before.
[0099] Media streamed according to its included media format is not limited to traditional visual and audio media. The included media format can include any type of media information capable of generating signals that can interact with machines to stimulate human vision, hearing, taste, touch, and smell. For example... Figures 3A-3B and Figures 4A-4B As shown, the media to be streamed, depending on the included media format, can be timed media, untimed media, or a mixture of both. The hierarchical representation of media objects is made possible by using a base layer and enhancement layer architecture, making the included media format streamable.
[0100] In some embodiments, separate base and enhancement layers are computed by applying multi-resolution or multi-mosaic analysis techniques to media objects in each scene. This computational technique is not limited to raster-based visual formats.
[0101] In some embodiments, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using wavelet analysis techniques.
[0102] In some embodiments of a hierarchical representation media format, enhancement layers can apply different properties to a base layer. For example, one or more enhancement layers can refine the material properties of the surface of a visual object represented by a base layer.
[0103] In some embodiments, in a layered representation medium format, attributes can refine the texture of an object surface represented by a base layer by, for example, changing the surface from a smooth texture to a porous texture, or from a rough surface to a glossy surface.
[0104] In some embodiments, in a layered representation media format, the surface of one or more visual objects in a scene can be changed from a Lambertian surface to a ray-traceable surface.
[0105] In some embodiments, in a layered representation media format, the network can distribute a base layer representation to a client, allowing the client to create a nominal representation of the scene while the client waits for the transmission of additional enhancement layers to refine the resolution or other features of the base layer.
[0106] In this embodiment, the resolution of attributes or refinement information in the enhancement layer is not explicitly related to the resolution of objects in the base layer. Furthermore, the included media format can support any type of information media that can be rendered or driven by a rendering device or machine, thus enabling support for heterogeneous media formats to heterogeneous client endpoints. In some embodiments, the network distributing the media format will first query the client endpoint to determine the client's capabilities. Based on this query, if the client cannot meaningfully ingest the media representation, the network can remove attribute layers that the client does not support. In some embodiments, if the client cannot meaningfully ingest the media representation, the network can adapt the media from its current format to a format suitable for the client endpoint. For example, the network can adapt the media by using a network-based media processing protocol to transform volumetric visual media assets into a 2D representation of the same visual asset. In some embodiments, the network can adapt the media by employing neural network (NN) processing to reformat the media to an appropriate format or optionally synthesize the view required by the client endpoint.
[0107] A complete (or partially complete) immersive experience (replay of a live stream event, game, or video-on-demand asset) scene inventory is organized by scenes containing the minimum amount of information needed for rendering and ingestion to create the presentation. The scene inventory includes a list of individual scenes to be rendered for the overall immersive experience requested by the client. Associated with each scene is one or more representations of the geometric objects within the scene, corresponding to a streamable version of the scene geometry. One embodiment of a scene may refer to a low-resolution version of the scene's geometry. Another embodiment of the same scene may refer to an enhancement layer used for the low-resolution representation of the scene to add additional detail or tessellation to the geometry of the same scene. As described above, each scene may have one or more enhancement layers to progressively increase the detail of the scene's geometry. Each layer of media objects referenced within a scene may be associated with a token (e.g., a uniform resource identifier (URI)) that points to an address where a resource can be accessed within the network. This resource is analogous to a content delivery network (CDN), where content can be retrieved by a client. The token used to represent the geometry may point to a location within the network or to a location within the client. In other words, a client can signal to the network that its resources are available for network-based media processing.
[0108] According to embodiments, a scene (timed or untimed) can correspond to a scene graph as a multi-plane image (MPI) or a multi-spherical image (MSI). Both MPI and MSI techniques are examples of technologies that facilitate the creation of display-independent scene representations (i.e., images of the real world captured simultaneously from one or more cameras) for natural content. On the other hand, scene graph techniques can be employed to represent both natural and computer-generated images in the form of synthetic representations. However, creating such a representation is computationally very demanding when content is captured as a natural scene by one or more cameras. Creating a scene graph representation of naturally captured content is both time-consuming and computationally intensive, requiring complex analysis of the natural image using photogrammetry or deep learning techniques, or both, to create a synthetic representation that can then be used to interpolate a sufficient number of views to fill the view frustum of the target immersive client display. As a result, considering such synthetic representations as candidates for representing natural content is impractical, as they cannot actually be created in real-time for use cases requiring real-time distribution. Therefore, the best representation of computer-generated images is to use scene graphs with composite models, because computer-generated images are created using 3D modeling processes and tools, and using scene graphs with composite models produces the best representation of computer-generated images.
[0109] Figure 5 An example of a natural media compositing process 500 according to an embodiment is shown. The natural media compositing process 500 transforms the ingestion format from a natural scene into a representation of an ingestion format that can be used as a network serving heterogeneous client endpoints. To the left of the dashed line 510 is the content capture portion of the natural media compositing process 500. To the right of the dashed line 510 is the ingestion format compositing (for a natural image) of the natural media compositing process 500.
[0110] like Figure 5 As shown, the first camera 501 uses a single camera lens to capture, for example, a person (i.e., Figure 5 The scene depicts the actors shown. The second camera 502 captures the scene with five diverging fields of view by mounting five camera lenses around the circular object. Figure 5 The arrangement of the second camera 502 shown is an exemplary arrangement typically used for capturing omnidirectional content for VR applications. The third camera 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter portion of the sphere. The arrangement of the third camera 503 is an exemplary arrangement typically used for capturing light fields for light fields or holographic immersive displays. The embodiments are not limited to this. Figure 5 The configuration shown. The second camera 502 and the third camera 503 may include multiple camera lenses.
[0111] Natural image content 509 is output from the first camera 501, the second camera 502, and the third camera 503 and used as input to the synthesizer 504. The synthesizer 504 can use a set of training images 506 to train an neural network (NN) 505 to generate a captured NN model 508. The training images 506 can be predefined or stored based on previous synthesis processing. The NN model (e.g., the captured NN model 508) is a set of parameters and tensors (e.g., matrices) that define the weights (i.e., numerical values) used in well-defined mathematical operations applied to the visual signal to obtain an improved visual output that may include interpolations of new views of the visual signal not explicitly provided by the original signal.
[0112] In some embodiments, a photogrammetric process may be implemented instead of NN training 505. If the captured NN model 508 is created during the natural media synthesis process 500, then the captured NN model 508 becomes one of the assets in the ingestion format 507 of the natural media content. The ingestion format 507 may be, for example, MPI or MSI. The ingestion format 507 may also include media assets.
[0113] Figure 6An example of a composite media ingestion creation process 600 according to an embodiment is shown. The composite media ingestion creation process 600 creates an ingestion media format for composite media such as, for example, computer-generated images.
[0114] like Figure 6 As shown, camera 601 can capture point cloud 602 of the scene. Camera 601 can be, for example, a LiDAR camera. For example, computer 603 uses common gateway interface (CGI) tools, 3D modeling tools, or another animation process to create composite content (i.e., a representation of the composite scene in an ingestion format that can be used as a network serving heterogeneous client endpoints). Computer 603 can create CGI assets 604 over the network. Furthermore, sensor 605A can be worn by actor 605 in the scene. Sensor 605A can be, for example, a motion capture suit with attached sensors. Sensor 605A captures digital records of the actor 605's movements to generate animation motion data 606 (or MoCap data). Data from point cloud 602, CGI assets 604, and motion data 606 are provided as input to compositor 607, which creates a composite media ingestion format 608. In some embodiments, compositor 607 can use a neural network (NN) and training data to create an NN model to generate the composite media ingestion format 608.
[0115] Both natural and computer-generated (i.e., composite) content can be stored in a container. This container can include a serialization format for storing and exchanging information to represent all natural scenes, all composite scenes, or a mixture of composite and natural scenes, including scene graphs and all media resources required for rendering the scene. The content serialization process involves converting data structures or object states into a format that can be stored (e.g., in a file or storage buffer) or transmitted (e.g., via a network connection link) and subsequently reconstructed in the same or different computing environments. When the resulting bit sequence is reread according to the serialization format, this bit sequence can be used to create a semantically identical clone of the original object.
[0116] The dichotomy between optimal representations of natural content and computer-generated (i.e., synthetic) content suggests that the optimal ingestion format for naturally captured content differs from the optimal ingestion format for computer-generated content, or even from the optimal ingestion format for natural content that is not necessary for real-time distribution applications. Therefore, according to embodiments, the goal of the network is to be robust enough to support multiple ingestion formats for visually immersive media, regardless of whether these ingestion formats are naturally created using, for example, physical cameras or computer-generated.
[0117] Technologies such as OTOY's ORBX, Pixar's Universal Scene Description, and the Graphics Language Transmission Format 2.0 (glTF2.0) specification written by the Khronos 3D group represent scene graphics in formats suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content that uses deep learning or photogrammetry techniques to create corresponding synthetic representations of natural scenes (i.e., not required for real-time distribution applications).
[0118] OTOY's ORBX is one of several scene graph technologies that supports any type of timed or timed visual media, including ray-traced, traditional (frame-based), volumetric, and other types of composite or vector-based visual formats. ORBX differs from other scene graphs because it provides native support for free and / or open-source formats of meshes, point clouds, and textures. ORBX is a carefully designed scene graph intended to facilitate the exchange between multiple vendor technologies operating on scene graphs. Furthermore, ORBX offers a rich material system, support for open shader languages, robust camera systems, and support for Lua scripting. ORBX is also the foundation for the immersive technology media format released by the Immersive Digital Experiences Alliance (IDEA) under royalty-free terms. In the context of real-time media distribution, the ability to create and distribute ORBX representations of natural scenes is a feature of the availability of computational resources that perform complexity analysis of camera-captured data and composite the same data into a composite representation.
[0119] Pixar's USD is a widely used scene graph for visual effects and professional content creation. USD is integrated into Nvidia's Omniverse platform, a suite of tools for developers to create and render 3D models using Nvidia's graphics processing units (GPUs). A subset of USD released by Apple and Pixar is called USDZ, powered by Apple's ARKit.
[0120] glTF 2.0 is a version of a graphics language transport format specification written by the Khronos 3D group. This format supports simple scene graphics formats (including PNG and JPEG image formats) that typically support static (non-timed) objects in a scene. glTF 2.0 supports simple animations, including translation, rotation, and scaling of basic shapes (i.e., geometric objects) described using glTF primitives. glTF 2.0 does not support timed media, therefore it does not support video or audio media input.
[0121] These designs for scene representations of immersive visual media are provided as examples only and do not limit the ability of the disclosed subjects to specify a process for adapting an input immersive media source to a format suitable for the specific characteristics of a client endpoint device. Furthermore, any or all of the exemplary media representations described above employ or can employ deep learning techniques to train and create neural network models that enable or facilitate the selection of a specific view to fill the view frustum of a specific display based on a specific size of the frustum. The view selected for the view frustum of a specific display can be interpolated from existing views explicitly provided in the scene representation, such as interpolation based on MSI or MPI techniques. Alternatively, the view can be rendered directly from these rendering engines based on the specific virtual camera position, filters, or description of the virtual camera.
[0122] The methods and apparatus disclosed herein are robust enough to account for the existence of a relatively small but well-known set of immersive media ingestion formats that are sufficient to meet the requirements for real-time or on-demand (e.g., non-real-time) distribution of media captured naturally (e.g., with one or more cameras) or media created using computer-generated techniques.
[0123] Advanced network technologies, such as 5G for mobile networks, have further facilitated the use of neural network models or network-based rendering engines to interpolate views from immersive media ingestion formats, and the deployment of fiber optic cables for fixed networks. These advanced network technologies increase the capacity and capability of commercial networks, as this advanced network infrastructure can support the transmission and delivery of ever-increasing volumes of visual information. Network infrastructure management technologies such as Multi-access Edge Computing (MEC), Software Defined Networking (SDN), and Network Functions Virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changes in certain network resource requirements, such as dynamic increases or decreases in network throughput, network speed, round-trip latency, and computing resource demands. Furthermore, this inherent ability to adapt to dynamic network demands also facilitates the network's ability to adapt immersive media ingestion formats to suitable distribution formats to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.
[0124] Immersive media applications can have varying requirements for network resources. These include gaming applications that require significantly lower network latency to respond to real-time updates in the game state, telepresence applications with symmetrical throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications whose downlink resource requirements can increase depending on the type of client endpoint display consuming the data. Generally, any consumer-facing application can be supported by a variety of client endpoints with diverse onboard client capabilities for storage, computation, and power, as well as varying requirements for specific media representations.
[0125] Therefore, embodiments of this disclosure enable a fully equipped network, i.e., a network employing some or all of the characteristics of a modern network, to simultaneously support multiple traditional and immersive media-enabled devices based on specified features within the device. Thus, the immersive media distribution methods and processes described herein provide flexibility in utilizing media ingestion formats applicable to both real-time and on-demand use cases for media distribution, flexibility in supporting natural and computer-generated content from both traditional and immersive media-capable client endpoints, and support for both timed and untimed media. The methods and processes also dynamically adapt the source media ingestion format to a suitable distribution format based on the characteristics and capabilities of the client endpoints and application requirements. This ensures that the distribution format is streamable over an IP-based network and enables the network to simultaneously serve multiple heterogeneous client endpoints, which may include traditional devices and devices with specific immersive media capabilities. Furthermore, embodiments provide exemplary media representation frameworks that facilitate the organization of media distribution along scene boundaries.
[0126] The end-to-end implementation of heterogeneous immersive media distribution according to the aforementioned improved embodiments of this disclosure is based on... Figures 7-14A The processes and components described in the detailed description will be further described below.
[0127] The aforementioned techniques for representing and streaming heterogeneous immersive media can be implemented as computer software in both the source and destination, using computer-readable instructions and physically stored in one or more non-transitory computer-readable media, or implemented by one or more hardware processors with a specific configuration. Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0128] Computer software can be coded using any suitable machine code or computer language. This machine code or computer language can be assembled, compiled, linked, or similar mechanisms to create code containing instructions that can be executed directly by the computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.
[0129] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.
[0130] Figure 7The components shown for computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any component or combination of components illustrated in the exemplary embodiments of computer system 700.
[0131] Computer system 700 may include certain human-computer interface input devices. Such human-computer interface input devices can respond to input from one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as speech, clapping), visual input (such as gestures), olfactory input, etc. Human-computer interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images, photographic images obtained from still image cameras), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0132] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard 701, touchpad 702, mouse 703, screen 709 (which may be, for example, a touch screen), data glove, joystick 704, microphone 705, camera 706, and scanner 707.
[0133] Computer system 700 may also include certain human-machine interface (HMI) output devices. Such HMI output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. These HMI output devices may include tactile output devices (e.g., tactile feedback from screen 709, data gloves, or joystick 704, but tactile feedback devices not used as input devices may also exist), audio output devices (e.g., speakers 708, headphones), visual output devices (e.g., screen 709, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be able to output two-dimensional or more than three-dimensional visual output via devices such as stereoscopic output devices; virtual reality glasses, holographic displays, and smoke canisters), and printers.
[0134] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media 711 including CD / DVD ROM / RW with CD / DVD or similar media 710, thumb drives 712, removable hard disk drives or solid-state drives 713, conventional magnetic media such as magnetic tapes and floppy disks, devices based on dedicated ROM / ASIC / PLD such as security dongles, etc.
[0135] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0136] Computer system 700 may also include an interface 715 to one or more communication networks 714. Network 714 may be, for example, wireless, wired, or optical. Network 714 may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of network 714 include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANbus, etc. Some networks 714 typically require an external network interface adapter (e.g., graphics adapter 725) attached to some general-purpose data port or peripheral bus 716 (such as, for example, a USB port of computer system 700); others are typically integrated into the core of computer system 700 via attachment to system bus 748 as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks 714, computer system 700 can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcasting TV), send-only (e.g., to a CANbus device), or bidirectional (e.g., using a local or wide area digital network to other computer systems). As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.
[0137] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the kernel 717 of the computer system 700.
[0138] The core 717 may include one or more central processing units (CPUs) 718, graphics processing units (GPUs) 719, dedicated programmable processing units 720 in the form of field-programmable gate arrays (FPGAs), hardware accelerators 721 for certain tasks, and so on. These devices, along with read-only memory (ROM) 723, random access memory (RAM) 724, and internal mass storage 722 such as internal non-user-accessible hard disk drives (SDs), can be connected via a system bus 748. In some computer systems, the system bus 748 may be accessed as one or more physical connectors to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus 748 directly or via a peripheral bus 716. Peripheral bus architectures include PCI, USB, etc.
[0139] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can execute certain instructions, which, when combined, constitute the aforementioned machine code (or computer code). This computer code can be stored in ROM 723 or RAM 724. Transient data can also be stored in RAM 724, while permanent data can be stored, for example, in internal mass storage 722. Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs 718, GPU 719, mass storage 722, ROM 723, RAM 724, etc.
[0140] Computer-readable media may contain computer code for performing operations of various computer implementations. For the purposes of this disclosure, the media and computer code may be specially designed and constructed, or they may be of a type well known and available to those skilled in the art of computer software.
[0141] As an example, and not a limitation, a computer system having a computer system 700 architecture, particularly a kernel 717, can provide functionality as a result of executing software embodied in one or more tangible computer-readable media (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be associated with user-accessible mass storage as described above, and with certain non-transitory storage of the kernel 717 (such as internal mass storage 722 or ROM 723). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the kernel 717. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the kernel 717 and specifically cause the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific portions of specific processes described herein, including defining and modifying data structures stored in RAM 724 according to software-defined processes. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 721), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to a computer-readable medium may include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.
[0142] Figure 7 The number and arrangement of components shown are provided as an example. In practice, input human-machine interface devices may include... Figure 7 Compared to those components shown, there are additional components, fewer components, different components, or components with different arrangements. Additionally or alternatively, a set of components (e.g., one or more components) of the input human-machine interface device can perform one or more functions described as being performed by another set of components of the input human-machine interface device.
[0143] In an embodiment, Figures 1A-6 and Figures 8-14B Any operation or process can be accessed or used Figure 7 Implemented by any of the components shown in the diagram.
[0144] Figure 8 An exemplary network media distribution system 800 serving multiple heterogeneous client endpoints is illustrated. Specifically, system 800 supports various conventional and heterogeneous displays with immersive media capabilities as client endpoints. System 800 may include a content acquisition module 801, a content preparation module 802, and a transmission module 803.
[0145] The content acquisition module 801 uses, for example... Figure 6 and / or Figure 5The embodiments described herein are used to capture or create source media. The content preparation module 802 creates an ingestion format, which is then transmitted to a network media distribution system using the transmission module 803. A gateway 804 can serve customer premises equipment to provide network access to various client endpoints of the network. A set-top box 805 can also be used as customer premises equipment to provide access to content aggregated by a network service provider. A radio demodulator 806 can be used as a mobile network access point for a mobile device (e.g., a mobile phone display 813 shown). In a particular embodiment of system 800, a conventional 2D television 807 is shown directly connected to one of the gateway 804, set-top box 805, or WiFi (router) 808. A laptop computer 2D display 809 (i.e., a computer or laptop computer with a conventional 2D display) is illustrated as a client endpoint connected to WiFi (router) 808. A head-mounted 2D (raster-based) display 810 is also connected to WiFi (router) 808. A lenticular light field display 811 is shown connected to one of the gateways 804. A lenticular light field display 811 may include one or more GPUs 811A, a storage device 811B, and a visual presentation component 811C that creates multiple views using light-based lenticular optics. A holographic display 812 is shown connected to a set-top box 805. The holographic display 812 may include one or more CPUs 812A, one or more GPUs 812B, a storage device 812C, and a visualization component 812D. The visualization component 812D may be a Fresnel pattern-based, wave-based holographic device / display. An augmented reality (AR) headset 814 is shown connected to a radio demodulator 806. The AR headset 814 may include a GPU 814B, a storage device 814C, a battery 814A, and a volumetric visual presentation component 814D. A dense light field display 815 is shown connected to a WiFi (router) 808. The dense light field display 815 may include one or more GPUs 815A, one or more CPUs 815B, a storage device 815C, an eye-tracking device 815D, a camera 815E, and a dense light field panel 815F.
[0146] Figure 8 The number and arrangement of components shown are provided as an example. In practice, system 800 may include components with... Figure 8 Compared to those components shown, there are additional components, fewer components, different components, or components with different arrangements. Additionally or optionally, a set of components of system 800 (e.g., one or more components) can perform one or more functions described as being performed by another set of components of a device or corresponding display.
[0147] Figure 9 An exemplary workflow of an immersive media distribution process 900 is shown, which is capable of delivering media as previously described. Figure 8 As described, it serves both traditional displays and heterogeneous displays with immersive media capabilities. The immersive media distribution process 900, performed over the network, can, for example, adapt media over the network for consumption by specific immersive media client endpoints (as described in the reference). Figure 10 Prior to the process described above, adaptation information about the specific media represented in the media ingestion format is provided.
[0148] The immersive media distribution process 900 can be divided into two parts: immersive media production to the left of the dashed line 912 and immersive media network distribution to the right of the dashed line 912. Immersive media production and immersive media network distribution can be performed by a network or client devices.
[0149] Media content 901 is created or retrieved from a network (or client device) or from a content source. The methods used to create or retrieve the data can correspond to... Figure 5 as well as Figure 6 The content consists of both natural and synthetic content. Then, the created content 901 is transformed into an ingestion format using the web ingestion format creation process 902. The web ingestion format creation process 902 can also correspond to... Figure 5 as well as Figure 6 The ingested content includes both natural and synthetic content. The ingestion format can also be updated to store information about assets that may be reused across multiple scenarios, such as those from a media complexity analyzer (see later). Figure 10 and Figure 14A (Detailed description). The ingested format is transmitted to the network and stored in the ingested media storage 903 (i.e., storage device). In some embodiments, the storage device may be located in the network of the immersive media content creator and may be remotely accessed for immersive media network distribution 920. Client-specific information 904 may optionally be available in the remote storage device. In some embodiments, the client-specific information 904 may reside remotely in an alternate cloud network and may be transmitted to that network.
[0150] Then, the network orchestrator 905 is executed. Network orchestration serves as the primary source and aggregation of information for performing the network's main tasks. The network orchestrator 905 can be implemented in a uniform format with other network components. The network orchestrator 905 can further employ a bidirectional messaging protocol with client devices to facilitate all media processing and distribution based on the characteristics of the client devices. Furthermore, the bidirectional protocol can be implemented across different transport channels (e.g., control plane channels and / or data plane channels).
[0151] like Figure 9 As shown, network orchestrator 905 receives information about the characteristics and attributes of client device 908. Network orchestrator 905 collects requests regarding applications currently running on client device 908. This information can be obtained from client-specific information 904. In some embodiments, the information can be obtained by directly querying client device 908. When directly querying the client device, it is assumed that a bidirectional protocol exists and is operational, allowing client device 908 to communicate directly with network orchestrator 905.
[0152] The network orchestrator 905 can also be started and work with the media adapter and segmentation module 910 (which in Figure 10 (Described in the text) Communication. When the ingested media is adapted and segmented by the media adapting and segmenting module 910, the media can be transferred to an inter-media storage device, such as media 909 prepared for distribution. If the network is designed to include a cache for assets used multiple times in the context of presentation, a redundant cache 912 of another inter-media storage device for reused media assets can be utilized as a cache for such assets. When the distribution media is prepared and stored in the storage device 909 prepared for distribution, the network orchestrator 905 ensures that the client device 908 receives the distribution media and descriptive information 906 via a "push" request, or the client device 908 can initiate a "pull" request for the distribution media and descriptive information 906 from the storage media 909 prepared for distribution. Information can be "push" or "pull" via the network interface 908B of the client device 908. The distribution media and descriptive information 906 after being "pushed" or "pulled" can be descriptive information corresponding to the distribution media.
[0153] In some embodiments, the network orchestrator 905 employs a bidirectional messaging interface to execute "push" requests or "pull" requests initiated by the client device 908. The client device 908 may optionally employ GPUs 908C (or CPUs).
[0154] The distribution media format is then stored in a storage device or storage cache 908D included in the client device 908. Finally, the client device 908 visualizes the media via a visualization component 908A.
[0155] Throughout the process of streaming immersive media to client device 908, network orchestrator 905 monitors the client's progress status via client progress and status feedback channel 907. In some embodiments, status monitoring can be performed via a bidirectional communication message interface.
[0156] Figure 10An example of a media adaptation process 1000 performed by, for example, a media adaptation and segmentation module 910 is shown. By performing the media adaptation process 1000, the ingested source media can be appropriately adapted to match the requirements of the client (e.g., client device 908).
[0157] like Figure 10 As shown, the media adaptation process 1000 includes multiple components that facilitate the adaptation of ingested media into an appropriate distribution format for client device 908. Figure 10 The components shown should be considered exemplary. In practice, the media adaptation process 1000 may include... Figure 10 Compared to the components shown, there are additional components, fewer components, different components, or components with different arrangements. Additionally or optionally, a set of components (e.g., one or more components) of the media adaptation process 1000 may perform one or more functions described as being performed by another set of components.
[0158] exist Figure 10 In this process, the adaptation module 1001 receives input network status 1005 to track the current traffic load on the network. As described above, the adaptation module 1001 also receives information from the network orchestrator 905. This information may include attribute and characteristic descriptions of the client device 908, application characteristics and descriptions, the current state of the application, and the client NN model (if available) to assist in mapping the geometry of the client's frustum to the interpolation capability for ingesting immersive media. This information can be obtained via a bidirectional messaging interface. The adaptation module 1001 ensures that the output of the adaptation is stored in the storage device 1006 for storing the client adaptation media when it is created.
[0159] The media complexity analyzer 911 can be an optional process that can be executed first or as part of a network automation process for distributing media. The media complexity analyzer 911 can store the ingested media formats and assets in a storage device (1002). The ingested media formats and assets can then be transferred from the storage device (1002) to the adaptation module 1001.
[0160] The adaptation module 1001 can be controlled by a logic controller 1001F. The adaptation module 1001 can also employ a renderer 1001B or a processor 1001C to adapt a specific ingested source media to a format suitable for the client. The processor 1001C can be a neural network-based processor. The processor 1001C uses a neural network model 1001A. Examples of such a processor 1001C include depth-view neural network model generators as described in MPI and MSI. If the media is in 2D format, but the client must have a 3D format, the processor 1001C can invoke a process to derive a volumetric representation of the scene depicted in the media using highly correlated images from the 2D video signal.
[0161] A renderer 1001B can be a software-based (or hardware-based) application or process based on a selective mix of disciplines involving acoustic physics, optical physics, visual perception, auditory perception, mathematics, and software development. Given an input scene graph and asset containers, the renderer emits (typical) visual and / or audio signals suitable for rendering on a target device or conforming to the desired performance specified by the properties of the rendering target nodes in the scene graph. For visual-based media assets, the renderer may emit visual signals suitable for the target display or stored as intermediate assets (e.g., repackaged into another container and used in a series of rendering processes in the graphics pipeline). For audio-based media assets, the renderer may emit audio signals for rendering in multi-channel speakers and / or dual-audio headphones, or for repackaging into another (output) container. The renderer includes, for example, real-time rendering features of source and cross-platform game engines. The renderer may include a scripting language (i.e., an interpreted programming language) that can be executed by the renderer at runtime to handle dynamic input and variable state changes to scene graph nodes. Dynamic inputs and changes in variable states can affect the rendering and evaluation of spatial and temporal object topology (including physics forces, constraints, inverse kinematics, deformation, and collisions) as well as energy propagation and transmission (light, sound). Evaluation of spatial and temporal object topology produces results that lead to a shift in output from abstract to concrete (e.g., similar to the evaluation of a document object model for a webpage).
[0162] Renderer 1001B may be a modified version of, for example, the OTOY Octane renderer, which will be modified to interact directly with adapter module 1001. In some embodiments, renderer 1001B implements computer graphics methods (e.g., path tracing) for rendering 3D scenes, such that the scene's lighting is realistic. In some embodiments, renderer 1001B may employ a shader (i.e., a computer program originally used for shading (producing appropriate levels of light, dark, and color in an image), but now this shader performs various specialized functions in various areas of computer graphics special effects, video post-processing unrelated to shading, and other functions unrelated to graphics).
[0163] The adaptation module 1001 can perform compression and decompression of media content using a media compressor 1001D and a media decompressor 1001E, depending on the compression and decompression requirements based on the format of the ingested media and the format required by the client device 908. The media compressor 1001D can be a media encoder, and the media decompressor 1001E can be a media decoder. After performing compression and decompression (if necessary), the adaptation module 1001 outputs client-adapted media 1006, optimized for streaming or distribution, to the client device 908. The client-adapted media 1006 can be stored in a storage device for storing the adaptation media.
[0164] Figure 11A An exemplary distribution format creation process 1100 is illustrated. For example... Figure 11A As shown, the distribution format creation process 1100 includes a media adaptation module 1101 and an adapted media packaging module 1103. The adapted media packaging module 1103 packages the media output from the media adaptation process 1000 and stores it as client-adapted media 1006. The media packaging module 1103 formats the adapted media from the client-adapted media 1006 into a robust distribution format 1104. The distribution format can be, for example, in... Figure 3A or Figure 4A The exemplary format shown is illustrated. Information list 1104A can provide a list 1104B of scene data assets to client device 908. The list 1104B of scene data assets may also include metadata describing the complexity of each asset, each asset being used across a set of scenes including the presentation. The list 1104B of scene data assets depicts a list of visual assets, audio assets, and haptic assets, each asset having its corresponding metadata. In this exemplary embodiment, each asset reference in the list 1104B of scene data assets includes metadata containing a numerical complexity metric characterizing the amount of time and resources the client will spend processing the asset across all scenes including the presentation.
[0165] Figure 11BThe process 1110 for creating a distribution format with ordered complexity is described. The process 1110 for creating a distribution format with ordered complexity includes steps similar to... Figure 11A The components. Therefore, refer to this. Figure 11B The description repeats itself. Media packaging module 11043 is similar to media packaging module 1103. However, media packaging module 11043 sorts the assets in scene data asset list 11044B based on a complexity metric. For example, visual assets in scene data asset list 11044B can be sorted by increasing complexity, while audio and haptic assets in scene data asset list 1104B can be sorted by decreasing complexity, and vice versa. Figure 11B As shown, the assets in distribution format 11044 are first sorted by asset type, and then by the complexity of how the assets are used across the entire presentation, for example, according to an increasing or decreasing complexity value based on asset type (i.e., visual, haptic, audio, etc.). Information list 11044A provides the client device 908 with a list 11044B of scene data assets it expects to receive, along with optional metadata indicating the complexity of using all assets across a set of scenes encompassing the entire presentation. The list 11044B of scene data assets includes lists of visual assets, audio assets, and haptic assets, each with its corresponding metadata.
[0166] The media can be further packaged before streaming. Figure 12 An exemplary packaging process 1200 is illustrated. The packaging system 1200 includes a packer 1202. The packer 1202 can receive a list of scene data assets 1104B (or 11044B) as input media 1201 (such as...). Figure 12 (As shown). In some embodiments, client-adapted media 1006 or distribution format 1104 is input to packetizer 1202. Packetizer 1202 divides the input media 1201 into individual data packets 1203 suitable for representation and streaming over the network to client device 908.
[0167] Figure 13 This is a sequence diagram illustrating an example of the data and communication flow between components according to an embodiment. Figure 13 The sequence diagram shows how the network adapts a specific immersive media ingested in a particular format to a streamable and suitable distribution format for a specific immersive media client endpoint. The data and communication flow can be illustrated as follows.
[0168] Client device 908 initiates a media request 1308 to network orchestrator 905. In some embodiments, this request may be made to the network distribution interface of the client device. Media request 1308 includes information identifying the media requested by client device 908. The media request may be identified by, for example, a uniform resource name (URN) or other standard naming conventions. Then, in response to media request 1308, network orchestrator 905 sends a profile request 1309 to client device 908. Profile request 1309 requests the client to provide information about currently available resources (including compute, storage, battery charge percentage, and other information characterizing the client's current operating state). Profile request 1309 also requests the client to provide one or more neural network (NN) models, which, if available at the client endpoint, the network can use to perform NN inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system.
[0169] Then, client device 908 sends a response 1310 from client device 908 to network orchestrator 905, which is provided as a client token, an application token, and one or more NN model tokens (if such NN model tokens are available on the client endpoint). Network orchestrator 905 then provides a session ID token 1311 to the client device. Network orchestrator 905 then requests media ingestion 1312 from ingest media server 1303. Ingest media server 1303 may include, for example, ingest media storage 903 or storage device 1002 for ingesting media formats and assets. The request for media ingestion 1312 may also include the URN or other standard name of the media identified in request 1308. Ingest media server 1303 responds to the media ingestion 1312 request with a response 1313 including an ingest media token. Network orchestrator 905 then provides the media token from response 1313 to client device 908 in call 1314. Then, network orchestrator 905 initiates an adaptation process for the requested media in request 1315 by providing the adaptation and segmentation module 910 with an ingest media token, a client token, an application token, and an NN model token. The adaptation and segmentation module 910 requests access to the ingested media by providing the ingest media token to the ingest media server 1303 in request 1316 to request access to the ingested media asset.
[0170] In response to request 1316, the ingest media server 1303 sends a response 1317 containing an ingest media access token to the adaptation and segmentation module 910. The adaptation and segmentation module 910 then requests the media adaptation process 1000 to adapt the ingested media located at the ingest media access token for the client, application, and NN inference model corresponding to the session ID token created and transmitted at response 1311. The adaptation and segmentation module 910 sends a request 1318 to the media adaptation process 1000. This request 1318 contains the required token and session ID. The media adaptation process 1000 provides the adapted media access token and session ID to the network orchestrator 905 in an update response 1319. The network orchestrator 905 then provides the adapted media access token and session ID to the media packaging module 11043 in an interface call 1320. The media packaging module 11043 provides a response 1321 to the network orchestrator 905, which includes the packaged media access token and session ID. Then, in response 1322, the media packaging module 11043 provides the packaged asset, URN, and packaged media access token for the session ID to the packaged media server 1107 for storage. Subsequently, the client device 908 sends a request 1323 to the packaged media server 1107 to initiate streaming of the media asset corresponding to the packaged media access token received in response 1321. Finally, the client device 908 executes other requests and provides a status update to the network orchestrator 905 in message 1324.
[0171] Figure 14A It shows Figure 9 The workflow of the media complexity analyzer 911 is shown. The media complexity analyzer 911 analyzes metadata related to the uniqueness of objects included in the media data.
[0172] At operation 1401, initialization can begin. At operation 1402, media asset data is potentially read. At operation 1403, it is determined whether the media asset data was successfully read. If the media asset data was not successfully read, the process ends at operation 1409. However, if the media asset data was successfully read, the process continues to operation 1404. At operation 1404, a read or data retrieval is performed to obtain complexity attributes from the media asset data. The data retrieved at operation 1404 is then parsed to access attributes describing the media assets. Each attribute is provided as input to operation 1405. At operation 1405, the data from operation 1404 is examined to determine if it is part of the list 1410 of the previously identified complexity attributes (see reference). Figure 14BOne of the following is a list of complexity attributes. This list 1410 may include the size of the object, which affects the amount of storage required to process the object; the number of polygons of the object, which may be an indicator of the processing power required by the GPU; fixed-point and floating-point numerical representations, which may be an indicator of the processing power required by the GPU or CPU; bit depth, which may also be an indicator of the processing power required by the GPU or CPU; single-float and double-float numerical representations, which may be an indicator of the size of the data value being processed; the existence of a light distribution function, which may indicate the physical process by which the scene needs to undergo to simulate how light is distributed in the scene; the type of light distribution function, which may indicate the complexity of the light distribution function used to simulate the physical process of light; and the required transformation process (if any), which may indicate the complexity of how the object needs to be placed (via rotation, translation, and scaling) into the scene. If the attribute is a complexity attribute, the process continues to operation 1406; otherwise, the process moves to operation 1407. At operation 1406, the value(s) of the complexity attribute(s) are retrieved and such values(s) are stored in the complexity summary area of the media asset. At operation 1407, it can be determined whether there are more attributes to read from the media asset. If no more attributes are to be read, the process continues to operation 1408, which writes a summary of the object's complexity into an area identified as being used to store complexity data of the scene containing the object.
[0173] Although several exemplary embodiments have been described in this disclosure, there are variations, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
Claims
1. A method for packaging media to optimize media distribution in a media streaming network, characterized in that, The method is performed by at least one processor, the method comprising: a media streaming server receiving an immersive media stream comprising one or more immersive media assets associated with one or more scenes; identifying a subset of the one or more immersive media assets, the subset comprising essential elements of a respective scene of the one or more scenes; ordering the one or more immersive media assets in an order based on the identified subset of the one or more immersive media assets, the subset comprising the essential elements of the respective scene; and determining a respective complexity measure and a respective asset type associated with the one or more immersive media assets based on a determination that at least one of the one or more respective properties exists in a list of complexity properties; and ordering the one or more immersive media assets in an order based on the respective complexity measure and the respective asset type associated with the one or more immersive media assets; streaming the one or more immersive media assets from the media streaming server to a client device in the ordered sequence.
2. The method of claim 1, wherein, The order of the one or more immersive media assets is first ordered by the asset type and then ordered by the respective complexity measure.
3. The method of claim 1, wherein, The method further comprises: generating, by the media streaming server, an asset complexity analysis associated with the one or more scenes based on a determination that more than one of the one or more respective properties exists in the list of complexity properties.
4. The method of claim 1, wherein, The one or more respective properties associated with the one or more immersive media assets comprise at least one of: a size of an immersive media asset, a number of polygons in the immersive media asset, a bit depth, and a type of conversion required.
5. The method of claim 4, wherein, The respective complexity measure associated with each of the one or more immersive media assets is based on the one or more respective properties associated with each of the one or more immersive media assets.
6. The method of claim 4, wherein, The respective complexity measure associated with the immersive media asset of the one or more media immersive assets is based on the one or more respective properties of the immersive media asset existing in the list of complexity properties.
7. A media streaming server for packaging media to optimize media distribution in a media streaming network, characterized in that, The media streaming server comprises: at least one memory configured to store computer program code; and at least one processor configured to read the computer program code and operate as instructed by the computer program code, the computer program code comprising: first receiving code configured to cause the at least one processor to receive an immersive media stream comprising one or more immersive media assets associated with one or more scenes; identifying code configured to cause the at least one processor to identify a subset of the one or more immersive media assets, the subset comprising essential elements of a respective scene of the one or more scenes; determining code configured to cause the at least one processor to determine respective complexity metrics and respective asset types associated with the one or more immersive media assets based on a determination that at least one of the one or more respective properties exists in a list of complexity properties; ordering code configured to cause the at least one processor to order the one or more immersive media assets in an order based on the respective complexity metrics and the respective asset types associated with the one or more immersive media assets based on the identified subset of the one or more immersive media assets, the subset including the necessary elements of the respective scene, and streaming code configured to cause the at least one processor to stream the one or more immersive media assets from the media streaming server to a client device in the ordered sequence.
8. The media streaming server of claim 7, wherein, The order of the one or more immersive media assets is ordered first by the asset type and then by the respective complexity metrics.
9. The media streaming server of claim 7, wherein, The computer program code further includes: generating code configured to cause the at least one processor to generate an asset complexity analysis associated with the one or more scenes based on a determination that more than one of the one or more respective properties exists in the list of complexity properties.
10. The media streaming server of claim 7, wherein, The one or more respective properties associated with the one or more immersive media assets include at least one of a size of an immersive media asset, a number of polygons in the immersive media asset, a bit depth, and a type of conversion required.
11. The media streaming server of claim 10, wherein, The respective complexity metric associated with each of the one or more immersive media assets is based on the one or more respective properties associated with each of the one or more immersive media assets.
12. The media streaming server of claim 10, wherein, The respective complexity metric associated with the immersive media asset of the one or more immersive media assets is based on the one or more respective properties of the immersive media asset existing in the list of complexity properties.
13. A non-transitory computer-readable medium, comprising: The non-transitory computer-readable medium stores instructions that, when executed by at least one processor of a media streaming server for packaging media to optimize media distribution in a media streaming network, cause the at least one processor to: receive an immersive media stream including one or more immersive media assets associated with one or more scenes; identify a subset of the one or more immersive media assets, the subset including necessary elements of a respective scene of the one or more scenes; order the one or more immersive media assets in an order based on the identified subset of the one or more immersive media assets, the subset including the necessary elements of the respective scene; and determine respective complexity metrics and respective asset types associated with the one or more immersive media assets based on a determination that at least one of the one or more respective properties exists in a list of complexity properties; and based on the respective complexity metrics and the respective asset types associated with the one or more immersive media assets, the one or more immersive media assets are ordered in a sequence; streaming the one or more immersive media assets from the media streaming server to a client device in the ordered sequence.
14. The non-transitory computer-readable medium of claim 13, wherein, the sequence of the one or more immersive media assets is ordered first by the asset type and then by the respective complexity metrics.
15. The non-transitory computer-readable medium of claim 13, wherein, the instructions further cause the at least one processor to: based on determining that more than one of the one or more respective properties exist in the complexity property list, the media streaming server generates an asset complexity analysis associated with the one or more scenes.
16. The non-transitory computer-readable medium of claim 13, wherein, the one or more respective properties associated with the one or more immersive media assets include at least one of a size of an immersive media asset, a number of polygons in the immersive media asset, a bit depth, and a type of conversion required.
17. The non-transitory computer-readable medium of claim 16, wherein, the respective complexity metric associated with each of the one or more immersive media assets is based on the one or more respective properties associated with each of the one or more immersive media assets.
Citation Information
Patent Citations
Prioritized Data Synchronization with Host Device
US20080168526A1
Systems and Methods For Content Delivery Acceleration of Virtual Reality and Augmented Reality Web Pages
US20200292013A1