Method for streaming immersive media, and computer system and computer program therefor

JP2024105424A5Active Publication Date: 2025-05-07TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024075250
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-20
Filing Date
2024-05-07
Publication Date
2025-05-07
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

Existing systems lack a coherent end-to-end ecosystem for delivering immersive media over commercial networks, as there is no single standard representation that can address real-time and non-real-time delivery of immersive media to diverse client endpoints with varying capabilities and resources.

Method used

A network-based media distribution system that adapts immersive media sources to the specific characteristics of client endpoints using neural networks, embedding scene-specific models in the coded bitstream or signaling their use, enabling flexible delivery formats for both legacy and immersive displays.

Benefits of technology

The system effectively supports multiple client endpoints with varying capabilities, improving visual quality and ensuring seamless delivery of immersive media by dynamically adapting media formats based on endpoint characteristics and requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for streaming immersive media and a computer system therefor, and a computer program.SOLUTION: An immersive media distribution module 900 includes a network capturing format creation module 902 capturing a content using a first two-dimensional format or a first three-dimensional format to refer to a neural network using the format, converting the captured content to a second two-dimensional format or a second three-dimensional format on the basis of the neural network referred to, and streaming the converted content to an immersive client 908 such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a dense light field display.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 127,036, filed in the U.S. Patent and Trademark Office (December 17, 2020), and U.S. Patent Application No. 17 / 407,816, filed in the U.S. Patent and Trademark Office (August 20, 2021), the entire contents of which are incorporated herein by reference. [Technical field] FIELD OF THE DISCLOSURE The present disclosure relates generally to the field of data processing, and more specifically, to video coding. [Background technology]

[0002] "Immersive media" generally refers to media that stimulates any or all of the human sensory systems (vision, hearing, somatosensation, smell, and sometimes taste) to create or enhance the user's perception of being physically present in the media experience, i.e., other than what is known as "legacy media," delivered over existing commercial networks for timed two-dimensional (2D) video and corresponding audio. Both immersive and legacy media can be characterized as timed or non-timed.

[0003] Timed media refers to media that is structured and presented according to time. Examples include feature films, news reports, and episodic content, which are all organized according to durations. Legacy video and audio are commonly considered timed media.

[0004] Non-timed media is media that is structured by logical, spatial and / or temporal relationships rather than by time. An example is a video game where the user has control over the experience created by the gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media also does not incorporate timed media, for example, in the continuous repeating audio or video segments of a video game scene. Conversely, timed media also does not incorporate non-timed media, such as, for example, a video with a fixed still image as a background.

[0005] An immersive media-enabled device refers to a device that has the capability to access, interpret, and present immersive media. Such media and devices are heterogeneous in terms of the amount and format of media and the number and type of network resources required to deliver such media at scale, i.e., over a network comparable to the delivery of legacy video and audio media. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogenous in capabilities since they all consist of rectangular display screens and use 2D rectangular video or still images as their primary media format. Summary of the Invention [Means for solving the problem]

[0006] Embodiments relate to methods, systems and computer readable media for streaming immersive media. In one aspect, a method for streaming immersive media is provided. The method may include capturing content in a first two-dimensional format or a first three-dimensional format, the format may refer to a neural network. Based on the referenced neural network, converting the captured content into a second two-dimensional format or a second three-dimensional format. Streaming the converted content to a client endpoint, such as a television, a computer, a head mounted display, a lenticular light field display, a holographic display, an augmented reality display or a high density light field display.

[0007] In another aspect, a computer system for streaming immersive media is provided. The computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices that at least one of the one or more processors executes via at least one of the one or more memories, thereby performing a method. The method may include capturing content in a first two-dimensional format or a first three-dimensional format, the format referencing a neural network. Based on the referenced neural network, converting the captured content into a second two-dimensional format or a second three-dimensional format. Streaming the converted content to a client endpoint, such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a high-density light field display.

[0008] In yet another aspect, a computer-readable medium for streaming immersive media is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored in at least one of the one or more tangible storage devices that can be executed by a processor. The program instructions are executable by a processor to perform a method that may include capturing content in a first two-dimensional format or a first three-dimensional format, the format correspondingly including referencing a neural network. Based on the referenced neural network, converting the captured content into a second two-dimensional format or a second three-dimensional format. Streaming the converted content to a client endpoint, such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a high-density light field display.

[0009] These and other objects, features and advantages will become apparent from the following detailed description of illustrative embodiments, which is to be read in connection with the accompanying drawings. Because the figures are intended to facilitate understanding of the invention by those skilled in the art together with the detailed description, various features of the drawings are not drawn to scale. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a schematic diagram of an end-to-end process for timed legacy media delivery. [Diagram 2] FIG. 1 is a schematic diagram of a standard media format used for streaming timed legacy media. [Diagram 3] FIG. 1 is a schematic diagram of an embodiment of a data model for representing and streaming timed immersive media. [Figure 4] FIG. 1 is a schematic diagram of an embodiment of a data model for representing and streaming non-timed immersive media. [Diagram 5]1 is a schematic diagram of a process for capturing a natural scene and converting it into a representation that can be used as an ingestion format for a network that serves heterogeneous client endpoints. [Figure 6] 1 is a schematic diagram of a process using 3D modeling tools and formats to create a representation of a synthetic scene that can be used as an ingestion format for a network that serves heterogeneous client endpoints. [Figure 7] FIG. 1 is a system diagram of a computer system. [Figure 8] 1 is a schematic diagram of a network providing services to multiple heterogeneous client endpoints; [Figure 9] For example, a schematic diagram of a network that provides adaptation information regarding particular media represented in a media capture format prior to the network's process of adapting the media for consumption by a particular immersive media client endpoint. [Figure 10] FIG. 1 is a schematic diagram of a media adaptation process that consists of a media render converter that converts source media from an ingest format to a specific format appropriate for a particular client endpoint. [Figure 11] 1 is a schematic diagram of a network that formats adapted source media into a data model suitable for presentation and streaming. [Figure 12] FIG. 12 is a flow diagram of a media streaming process for fragmenting the data model of FIG. 11 into network protocol packet payloads. [Figure 13] FIG. 1 is a sequence diagram of a network that adapts a particular immersive media in an ingest format to a streamable and appropriate delivery format for a particular immersive media client endpoint. [Figure 14] FIG. 10 is a schematic diagram of the captured media formats and assets 1002 of FIG. 10, consisting of both immersive and legacy content formats, i.e., 2D video formats only, or both immersive and 2D video formats. [Figure 15] 1 illustrates the transmission of neural network model information along with a coded video stream. [Figure 16] 1 illustrates the transmission of neural network model information along with input immersive media and assets. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Detailed embodiments of the claimed structures and methods are disclosed herein. However, it is understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which may be embodied in various forms. However, those structures and methods may be embodied in many different forms and should not be construed as being limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In the description, details of well-known features and techniques may be omitted so as not to unnecessarily obscure the presented embodiments.

[0012] FIELD OF THE DISCLOSURE The embodiments relate generally to the field of data processing, and more specifically to video coding. The techniques described herein allow a network to ingest a 2D video source of media containing one or more (usually a small number of) views, adapt the source of 2D media into one or more streamable "delivery formats" and signal scene-specific neural network models to accommodate the requirements of various heterogeneous client endpoint devices, their various features and capabilities, and the applications used by the client endpoints, before actually delivering the formatted media to various client endpoints. The network models may be embedded directly into the scene-specific coded video streams of the coded bitstream using an SEI structured field, or the SEI may signal the use of a particular model that is stored elsewhere in the delivery network, but accessible to the neural network process. The ability to reformat 2D media sources into various streamable delivery formats allows the network to simultaneously serve various client endpoints with different capabilities and available computational resources, enabling support of new immersive client endpoints such as holographic and light field displays in commercial networks. Additionally, the ability to adapt scene-specific 2D media sources based on scene-specific neural network models improves the final visual quality. Such an ability to adapt 2D media sources is particularly important when no immersive media sources are available and when clients cannot support delivery formats based on 2D media. In this scenario, neural network-based approaches can be more optimally used with specific scenes present in 2D media by retaining scene-specific neural network models trained with priors that are generally similar to objects in a particular scene or the context of a particular scene.This improves the network's ability to infer depth-based information about a particular scene, and allows it to adapt 2D media into a scene-specific volume format appropriate for the target client endpoint.

[0013] As previously mentioned, "immersive media" generally refers to media that stimulates any or all of the human sensory systems (vision, hearing, somatosensation, smell, and sometimes taste) to create or enhance the user's perception of being physically present in the media experience, i.e., other than what is distributed over existing commercial networks for timed two-dimensional (2D) video and corresponding audio, known as "legacy media." Both immersive and legacy media can be characterized as either timed or non-timed. Timed media refers to media that is structured and presented according to time. Examples include feature films, news reports, and episodic content, which are all organized according to durations. Legacy video and audio are commonly considered timed media. Non-timed media is media that is structured by logical, spatial and / or temporal relationships rather than by time. An example is a video game where the user has control over the experience created by the gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media also does not incorporate timed media, for example, in the continuous repeating audio or video segments of a video game scene. Conversely, timed media also does not incorporate non-timed media, such as, for example, a video with a fixed still image as a background.

[0014] An immersive media-enabled device refers to a device that has the capability to access, interpret, and present immersive media. Such media and devices are heterogeneous in terms of the amount and format of media and the number and type of network resources required to deliver such media at scale, i.e., over a network comparable to the delivery of legacy video and audio media. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogenous in capabilities since they all consist of rectangular display screens and use 2D rectangular video or still images as their primary media format.

[0015] The delivery of any media over a network may use media delivery systems and architectures that reformat the media from an input or network "ingest" format into a final delivery format that is suitable for the target client device and its application as well as conducive to streaming over the network. "Streaming" media broadly refers to the fragmentation and packetization of source media so that it can be delivered over a network in successive, small-sized "chunks" that are logically organized and ordered according to either or both of the media's temporal or spatial structure. In such delivery architectures and systems, the media may undergo a compression or layering process so that only the most salient media information is delivered to the client initially. In some cases, the client must receive all of the significant media information of a portion of media before presenting any of the same media portions to the end user.

[0016] The process of reformatting the input media to match the capabilities of the target client endpoint may use a neural network process that uses a network model that may encapsulate some prior knowledge of the particular media being reformatted. For example, a particular model may be tuned to recognize an outdoor park scene (with trees, plants, grass, and other objects commonly found in park scenes), while another particular model may be tuned to recognize an indoor dinner scene (with a dinner table, cooking utensils, people sitting at the table, etc.). Those skilled in the art will recognize that a neural network process with a network model tuned to recognize objects from a particular context, e.g., objects in a park scene, and a network model tuned to match the content of the particular scene, will produce better visual results than a less tuned network model. Thus, there is an advantage to providing a scene-specific network model to a neural network process tasked with reformatting the input media to match the capabilities of a target client endpoint.

[0017] A mechanism for associating a neural network model with a particular scene of 2D media can be achieved by optionally compressing the network model and inserting it directly into the 2D coded bitstream of the visual scene using the Supplemental Enhancement Information (SEI) structured field commonly used to attach metadata to coded video streams in the H.264, H.265 and H.266 video compression formats. The presence of an SEI message containing a particular neural network model within the context of a portion of a coded video bitstream may be used to indicate that the network model is used to interpret and adapt the video content within the portion of the bitstream in which the model is embedded. Alternatively, the SEI message can be used to signal which neural network model can be used in the absence of the actual model itself, via an identifier for the network model.

[0018] The mechanism for associating an appropriate neural network with an immersive media may be achieved by the immersive media itself referencing the appropriate neural network model to use. This referencing may be achieved by directly embedding the network model and its parameters on a per-object, per-scene, or combination thereof. Alternatively, rather than embedding one or more neural network models within the media, a media object or scene may reference a particular neural network model by an identifier.

[0019] Yet another alternative mechanism for referencing an appropriate neural network for adaptation of media for streaming to a client endpoint is for the particular client endpoint itself to provide at least one neural network model and corresponding parameters to the adaptation process to be used. Such a mechanism may be implemented by the client providing the neural network model in communication with the adaptation process, for example, when the client connects itself to the network. After adapting the video to the target client endpoint, the adaptation process in the network may choose to apply a compression algorithm to the result, which may optionally separate the adapted video signal into layers corresponding to the most to least salient portions of the visual signal.

[0020] An example of a compression and layering process is the progressive format of the JPEG standard (ISO / IEC 10918 Part 1), which first separates the image into layers so that the entire unfocused image is presented with only basic shapes and colors, i.e. from the low-order DCT coefficients of the entire image scan, and then separates additional layers of detail so that the image is focused, i.e. from the higher-order DCT coefficients of the image scan.

[0021] The process of breaking media into smaller pieces, organizing them into payload portions of successive network protocol packets, and delivering these protocol packets is called "streaming" the media, whereas the process of converting the media into a format suitable for presentation at one of a variety of heterogeneous client endpoints operating one of a variety of heterogeneous applications is known as "adapting" the media.

[0022] definition A scene graph is a common data structure commonly used by vector-based graphics editing applications and modern computer games that arranges a logical and often (but not necessarily) spatial representation of a graphical scene; it is also a collection of nodes and vertices in a graph structure.

[0023] A node is a basic element of a scene graph consisting of information related to the logical, spatial or temporal representation of visual, auditory, tactile, olfactory, gustatory or related processing information; each node must have at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) connected to it. The base layer is typically a nominal representation of the asset that is created to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network.

[0024] An enhancement layer is a set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities not supported in the base layer. An attribute is metadata associated with a node that is used to describe a particular characteristic or feature of the node in a standard or more complex manner (eg, in terms of another node).

[0025] A container is a serialized format for storing and exchanging information representing an entire natural scene, an entire synthetic scene, or a combination of synthetic and natural scenes, including the scene graph and all the media resources required to render the scene.

[0026] Serialization is the process of converting a data structure or the state of an object into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link) and then reconstructed (e.g., in another computing environment). The resulting series of bits, when reread according to the serialized format, can be used to create a semantically identical clone of the original object.

[0027] A renderer is a (usually software-based) application or process that, given an input scene graph and an asset container, transmits typically visual and / or audio signals suitable for presentation on a target device or conforming to desired characteristics specified in the attributes of the render target nodes in the scene graph, based on a selective combination of disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development. In the case of visual-based media assets, the renderer may transmit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged in another container, i.e., used in the sequence of rendering processes in a graphics pipeline); in the case of audio-based media assets, the renderer may transmit audio signals for presentation on multi-channel speakers and / or binaural headphones, or repackaged in another (output) container. Common examples of renderers include Unity, Unreal.

[0028] Evaluation involves generating results that change the output from an abstract to a concrete result (eg, similar to evaluating a Document Object Model of a web page).

[0029] A scripting language is an interpreted programming language that can process dynamic inputs and mutable state changes made to scene graph nodes that are executed by the renderer at run-time to affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, IK, transformations, collisions) and energy propagation and transfer (light, sound).

[0030] A shader is a type of computer program originally used for shading (producing the correct levels of light, darkness, and color in an image), but now used to perform a variety of special functions for various areas of computer graphics special effects, as well as for video post-processing unrelated to shading, and even for functions completely unrelated to graphics.

[0031] Path tracing is a computer graphics method of rendering three-dimensional scenes such that the lighting in the scene is realistic. Timed media is media that is ordered by time, for example having a start and end time according to a particular clock.

[0032] Non-timed media is media that is organized by spatial, logical, or temporal relationships, such as an interactive experience that is realized according to actions performed by a user. A neural network model is a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numbers) used in well-defined mathematical operations that are applied to a visual signal to arrive at an improved visual output, including the interpolation of new views of the visual signal that were not explicitly provided by the original signal.

[0033] Immersive media can be considered as one or more types of media that, when presented to a human by an immersive media-enabled device, stimulate any of the five senses: sight, sound, taste, touch, and hearing, in a manner that is more realistic and consistent with the human understanding of experiences in the natural world, i.e., other than the stimuli that would be achieved with legacy media presented by a legacy device. In this context, the term "legacy media" refers to two-dimensional (2D) visual media, still or video frames, and / or corresponding audio whose user interaction capabilities are limited to pausing, playing, fast-forwarding, or rewinding, and "legacy devices" refers to televisions, laptops, displays, and mobile devices whose capabilities are limited to presenting only legacy media. In consumer application scenarios, a presentation device for immersive media (i.e., an immersive media-enabled device) is a consumer hardware device that is specifically equipped with the capabilities to exploit certain information embodied by immersive media to be able to create a presentation that more closely approximates human understanding and interaction with the physical world, i.e., other than the capabilities of legacy devices to do so. Legacy devices are constrained in their ability to present only legacy media, whereas immersive media devices are similarly unconstrained.

[0034] Over the past decade, many immersive media-enabled devices have been introduced into the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, haptic gloves, and gaming consoles. Similarly, holographic displays and other forms of volumetric displays are poised to emerge within the next decade. Despite the immediate or imminent availability of these devices, a coherent end-to-end ecosystem for delivering immersive media over commercial networks has not materialized for several reasons.

[0035] One of these reasons is the lack of a single standard representation of immersive media that can address the two main use cases associated with current large-scale media distribution over commercial networks: 1) real-time distribution of live-action events, i.e., content is created and delivered to client endpoints in real-time or near real-time, and 2) non-real-time distribution, where content does not necessarily need to be delivered in real-time, i.e., content is physically captured or created. Respectively, these two use cases may be compared equally to currently existing "broadcast" and "on-demand" distribution formats.

[0036] For real-time delivery, content can be captured by one or more cameras or created using computer-generated techniques. Content captured by cameras is referred to herein as "natural" content, and content created using computer-generated techniques is referred to herein as "synthetic" content. Media formats representing synthetic content can be formats used in the 3D modeling, visual effects, and CAD / CAM industries and can include object formats and tools such as meshes, textures, point clouds, structured volumes, amorphous volumes (e.g., for fire, smoke, fog), shaders, procedurally generated shapes, materials, lighting, virtual camera definitions, animations, and the like. Although synthetic content is computer-generated, synthetic media formats can be used for both natural and synthetic content. However, the process of converting natural content into synthetic media formats (e.g., synthetic representations) can be a time-consuming and computationally intensive process, and thus may be impractical for real-time applications and use cases.

[0037] For real-time delivery of natural content, the content captured by the camera can be delivered in a raster format, which is suitable for legacy display devices since many legacy display devices are designed to display raster formats as well, i.e., delivery of the raster format is best suited for displays that can only display raster formats, since legacy displays are designed to uniformly display raster formats.

[0038] However, immersive media-capable displays are not necessarily limited to displaying raster-based formats. Furthermore, some immersive media-capable displays are not capable of presenting media that is available only in raster-based formats. The availability of displays optimized to create immersive experiences based on formats other than raster-based formats is another important reason why there is not yet a coherent end-to-end ecosystem for the delivery of immersive media. Yet another challenge in creating a coherent delivery system for multiple different immersive media devices is that current and new immersive media-enabled devices themselves can differ significantly. For example, some immersive media devices, e.g., head-mounted displays, are explicitly designed to be used by only one user at a time. Other immersive media devices are designed to be used by multiple users simultaneously, e.g., the "Looking Glass Factory 8K Display" (hereinafter referred to as a "Lenticular Light Field Display") can display content that can be viewed by up to 12 users simultaneously, where each user experiences their own unique perspective (i.e., view) of the displayed content.

[0039] Further complicating the development of coherent delivery systems is that the number of unique views each display can produce can vary significantly. Legacy displays are often only capable of creating a single view of the content. Lenticular light field displays, on the other hand, can support multiple users, each of whom can experience their own unique view of the same visual scene. To achieve the creation of multiple views of the same scene, lenticular light field displays create a specific volumetric viewing frustum that requires 45 unique views of the same scene as input to the display. This means that 45 slightly different unique raster representations of the same scene must be captured and delivered to the display in a format specific to one particular display, i.e., its viewing frustum. In contrast, the viewing frustum of legacy displays is limited to a single two-dimensional plane, which means that multiple viewing perspectives of the content cannot be presented through the viewing frustum of the display, regardless of the number of viewers experiencing the display simultaneously.

[0040] In general, immersive media displays may vary significantly depending on the characteristics of all displays: the dimensions and volume of the viewing frustum, the number of viewers supported simultaneously, the optical technology used to fill the viewing frustum, which may be point-based, ray-based, or wave-based technology, the density of light units (either points, rays, or waves) that occupy the viewing frustum, the availability of computing power, the type of computing (CPU or GPU), the source and availability of power (batteries or wires), the amount of local storage or caching, and access to auxiliary resources such as cloud-based computing and storage. These characteristics contribute to the heterogeneity of immersive media displays, which, in contrast to the homogeneity of legacy displays, complicates the development of a single delivery system that can support all displays, including both legacy and immersive types of displays.

[0041] The disclosed subject matter addresses the development of a network-based media delivery system capable of supporting both legacy and immersive media displays as client endpoints within the context of a single network. Specifically, presented herein is a mechanism for adapting an input immersive media source to a format suitable for the particular characteristics of a client endpoint device, including an application currently running on the client endpoint device. Such a mechanism for adapting an input immersive media source includes matching the characteristics of the input immersive media with the characteristics of a target endpoint client device, including an application running on the client device, and adapting the input immersive media to a format suitable for the target endpoint and its application. Additionally, the adaptation process may include interpolating additional views from the input media, such as novel views, to create additional views required by the client endpoint. Such interpolation may be performed utilizing a neural network process.

[0042] It should be noted that the remainder of the disclosed subject matter assumes, without loss of generality, that the process of adapting an input immersive media source to a particular endpoint client device is the same as or similar to the process of adapting the same input immersive media source to a particular application running on a particular client endpoint device, i.e., the problem of adapting an input media source to the characteristics of an endpoint device has the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.

[0043] Legacy devices supported by legacy media achieve broad consumer adoption because they are similarly supported by an ecosystem of legacy media content providers that generate standards-based representations of the legacy media, and commercial network service providers that provide the network infrastructure to connect legacy devices to sources of standard legacy content. In addition to their role in delivering legacy media over their networks, commercial network service providers may also facilitate pairing of legacy client devices with access to legacy content on a content delivery network (CDN). When paired with access to the appropriate form of content, the legacy client device can request or "pull" the legacy content from the content server to the device for presentation to the end user. Nevertheless, an architecture in which a network server "pushes" the appropriate media to the appropriate client is similarly relevant without introducing additional complexity into the overall architecture and solution design.

[0044] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0045] The exemplary embodiments described below relate to system and network architectures, structures, and components for delivering media including video, audio, geometric (3D) objects, haptics, associated metadata, or other content to client devices. Particular embodiments are directed systems, structures, and architectures for delivering media content to heterogeneous immersive and interactive client devices.

[0046] Figure 1 is an example of an end-to-end process of timed legacy media distribution. In Figure 1, timed audiovisual content is captured by a camera or microphone at 101A or computer generated at 101B creating a sequence 102 of 2D images and associated audio that is input to a preparation module 103. The output of 103 is edited content (e.g. for post-production including language translation, subtitling and other editing functions) and is called master format, ready to be converted by a converter module 104 into a standard mezzanine format, e.g. for on-demand media, or into a standard contribution format, e.g. for live events. Media is "ingested" by a commercial network service provider and an adaptation module 105 packages the media into various bit rates, temporal resolutions (frame rates) or spatial resolutions (frame sizes) packaged into a standard distribution format. The resulting adaptations are stored in a content delivery network 106, from which various clients 108 make pull requests 107 to fetch the media for presentation to end users. It is important to note that the master format may be comprised of a hybrid of media from both 101A or 101B, and format 101A may be obtained in real time from media obtained from, for example, a live sporting event. Additionally, the client 108 is responsible for selecting the particular adaptation 107 that best suits the client's configuration and / or current network conditions, although it is equally possible for a network server (not shown in FIG. 1) to determine and then "push" the appropriate content to the client 108.

[0047] FIG. 2 is an example of a standard media format used for the delivery of legacy timed media, such as video, audio, and supporting metadata (including timed text as used for subtitling). As described in item 106 of FIG. 1, the media is stored in a standards-based delivery format on a CDN 201. The standards-based format is shown as an MPD 202 composed of multiple parts including timed periods 203 with start and end times corresponding to a clock. Each period 203 points to one or more adaptation sets 204. Each adaptation set 204 is generally used for a single type of media, such as video, audio, or timed text. For any given period 203, multiple adaptation sets 204 may be provided, for example one adaptation set for video and multiple adaptation sets for audio, such as used for translation into various languages. Each adaptation set 204 points to one or more representations 205 that provide information about the frame resolution (for video), frame rate, and bit rate of the media. The multiple representations 205 may be used to provide access to a representation 205 for each of, for example, ultra-high definition, high definition, or standard definition video. Each representation 205 points to one or more segment files 206 where the media is actually stored for fetching by a client (shown as 108 in FIG. 1) or delivery (in a "push-based" architecture) by a network media server (not shown in FIG. 1).

[0048] Figure 3 is an example representation of a streamable format for timed heterogeneous immersive media. Figure 4 is an example representation of a streamable format for non-timed heterogeneous immersive media. Both figures refer to a scene, Figure 3 to a scene 301 for timed media, and Figure 4 to a scene 401 for non-timed media. In both cases, the scene can be embodied by various scene representations or descriptions.

[0049] For example, in some immersive media designs, a scene may be materialized by a scene graph, as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI and MSI techniques are examples of techniques that help create view-agnostic scene representations for natural content, i.e., real-world images captured simultaneously by one or more cameras. On the other hand, scene graph techniques may be used to represent both natural and computer-generated images in the form of synthetic representations, but such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. That is, scene graph representations of naturally captured content are time- and computationally intensive to create, requiring complex analysis of the natural imagery by photogrammetry, deep learning, or both techniques, in order to create a synthetic representation that can later be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations are currently unrealistic to be considered as candidates for representing natural content, since they cannot be created in real time to consider use cases that require real-time delivery. Nevertheless, currently the best candidate representation for computer-generated imagery is to use a scene graph in the synthetic model, as computer-generated imagery is created using 3D modeling processes and tools.

[0050] This dichotomy in optimal representations of both natural and computer-generated content suggests that the optimal capture format for naturally captured content will be different from the optimal capture format for computer-generated content or natural content that is not critical to real-time delivery applications. Thus, the disclosed subject matter aims to be robust enough to support multiple capture formats of visually immersive media, regardless of whether the content is naturally or computer-generated.

[0051] The following are examples of techniques that embody scene graphs as a format suitable for representing visually immersive media created using computer-generated techniques, or corresponding synthetic representations of natural scenes using deep learning or photogrammetry techniques, i.e., naturally captured content that is not essential for real-time delivery applications. 1. ORBX® by OTOY ORBX by OTOY is one of several scene graph technologies that can support any type of visual media, timed or untimed, including ray-traceable, legacy (frame-based), volumetric and other types of synthetic or vector-based visual formats. ORBX is different from other scene graphs because it provides native support for freely available and / or open source formats for meshes, point clouds and textures. ORBX is a scene graph purposefully designed with the goal of facilitating interchange between multiple vendor technologies that work with scene graphs. In addition, ORBX provides a rich material system, support for open shading languages, a robust camera system and support for Lua scripting. ORBX is also the basis of an immersive technical media format released for license under royalty-free terms by the Immersive Digital Experience Alliance (IDEA). In the context of real-time delivery of media, the ability to create and deliver ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of camera-captured data and composition of that same data into a synthetic representation. To date, the availability of sufficient computation for real-time delivery is not realistic, but it is still not impossible.

[0052] 2. Universal Scene Description by Pixar Pixar's Universal Scene Description (USD) is another well-known and mature scene graph that is popular in the VFX and professional content creation communities. USD is integrated into Nvidia's Omniverse platform, a toolset for developers to create and render 3D models using Nvidia's GPUs. A subset of USD was published by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.

[0053] 3. glTF 2.0 by Khronos glTF2.0 is the latest version of the "Graphics Language Transmission Format" specification created by the Khronos3D group. The format supports a simple scene graph format that can generically support static (non-timed) objects in a scene, including "png" and "jpeg" image formats. glTF2.0 supports simple animation, including support for moving, rotating, and scaling basic shapes, or geometric objects, described using glTF primitives. glTF2.0 does not support timed media, and therefore does not support video or audio. These known designs for immersive visual media scene representations are provided by way of example only and are not intended to limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source into a format suitable for the particular characteristics of a client endpoint device.

[0054] Additionally, any or all of the above example media representations currently use or may use deep learning techniques to train and create neural network models that enable or facilitate the selection of specific views to fill a particular display's view frustum based on particular dimensions of the frustum. The views selected for a particular display's view frustum may be interpolated from existing views explicitly provided in the scene representation, i.e., from MSI or MPI techniques, or may be rendered directly from these rendering engines based on particular virtual camera positions, filters, or virtual camera descriptions of the rendering engines.

[0055] Thus, the disclosed subject matter is robust enough to consider that there is a relatively small but well-known set of immersive media capture formats that can adequately meet the requirements of both real-time or "on-demand" (e.g., non-real-time) delivery of media that is either naturally captured (e.g., by one or more cameras) or created using computer-generated techniques.

[0056] With the introduction of advanced network technologies such as 5G for mobile networks and fiber optic cable for fixed networks, the interpolation of views from immersive media capture formats using either neural network models or network-based rendering engines will become even easier. That is, these advanced network technologies will increase the capacity and capabilities of commercial networks, as such advanced network infrastructure can support the transport and delivery of increasingly large amounts of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networking (SDN), and network function virtualization (NFV) allow commercial network service providers to flexibly deploy their network infrastructure to adapt to changes in demand for certain network resources, for example, to respond to dynamic increases or decreases in demand for network throughput, network speed, round-trip delay, and computational resources. Furthermore, this inherent ability to adapt to dynamic network requirements to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints will likewise facilitate the network's ability to adapt immersive media capture formats to appropriate delivery formats.

[0057] Immersive media applications themselves may have different requirements for network resources, including gaming applications that require significantly lower network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements on both the uplink and downlink portions of the network, and passive viewing applications that may have increased demands on downlink resources depending on the type of client endpoint display that is consuming the data. Typically, consumer applications are supported by a variety of client endpoints with different on-board client capabilities in terms of storage, computation and power, and different requirements for the particular media presentation.

[0058] The disclosed subject matter thus enables a fully equipped network, i.e., a network that uses some or all of the characteristics of a modern network, to simultaneously support multiple legacy and immersive media capable devices in accordance with features specified therein, which features are as follows: 1-7.

[0059] 1. It provides the flexibility to leverage realistic media ingest formats for both real-time and "on-demand" media delivery use cases. 2. It provides the flexibility to support both natural and computer-generated content for both legacy and immersive media-enabled client endpoints. 3. Support both timed and non-timed media. 4. Provide a process that dynamically adapts the capture format of source media to an appropriate delivery format based on the characteristics and capabilities of the client endpoint and the requirements of the application. 5. Ensure that the delivery format is capable of being streamed over IP-based networks. 6. Allows the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media-enabled devices. 7. We provide an exemplary media representation framework that facilitates the organization of distributed media along scene boundaries. The improved end-to-end implementation enabled by the disclosed subject matter is accomplished according to the processes and components described in the detailed description of FIGS. 3-16 as follows.

[0060] Both Figures 3 and 4 use a single exemplary generic delivery format adapted from the ingest source format to match the capabilities of a particular client endpoint. As noted above, the media shown in Figure 3 is timed and the media shown in Figure 4 is non-timed. The particular generic format is sufficiently robust in its structure to accommodate a wide variety of media attributes where each attribute may be layered based on the amount of salient information each layer contributes to the presentation of the media. Note that such layering processes are already well known in the current state of the art, as exemplified by progressive JPEG and scalable video architectures such as those specified in ISO / IEC 14496-10 (Scalable Advanced Video Coding).

[0061] 1. Media streamed in accordance with the generic media format is not limited to legacy visual and audio media, but may include any type of media information capable of interacting with a machine to produce signals that stimulate the human senses of sight, hearing, taste, touch, and smell. 2. Media streamed according to the generic media format may be timed media, non-timed media, or a combination of both. 3. The generic media format is further streamable by allowing layered representations of media objects using a base and enhancement layer architecture. In one example, separate base and enhancement layers are computed by applying multi-resolution or multi-tessellation analysis techniques to the media objects of each scene. This is similar to the progressively rendered image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In an exemplary embodiment, the progressive representation of a geometric object may be a multi-resolution representation of the object computed using wavelet analysis. In another example of a layered representation of a media format, the enhancement layer applies various attributes to the base layer, such as improving the material properties of the surface of the visual object represented by the base layer. In yet another example, the attributes may improve the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface. In yet another example of a layered representation, the surfaces of one or more visual objects in a scene may be modified from Lambertian to ray-traceable. In yet another example of layered representations, the network delivers a base layer representation to a client so that the client can create a nominal presentation of the scene while waiting for the transmission of additional enhancement layers to improve on the resolution or other properties of the base representation.

[0062] 4. The resolution of the enhancement layer attributes or refinement information is not explicitly tied to the resolution of the base layer objects, as in currently existing MPEG video and JPEG image standards. 5. Generic media formats enable support of heterogeneous media formats to heterogeneous client endpoints by supporting any type of information media that can be presented or acted upon by a presentation device or machine. In one embodiment of a network that delivers media formats, the network first queries the client endpoint to determine the client's capabilities, and then, if the client is unable to meaningfully consume the media representation, the network either removes layers of attributes not supported by the client or adapts the media from its current format to a format appropriate for the client endpoint. In one example of such adaptation, the network would convert a volumetric visual media asset into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can use neural network processes to reformat the media into an appropriate format or optionally synthesize a view required by the client endpoint.

[0063] 6. The manifest of a full or partially full immersive experience (live streaming event, game or on-demand asset playback) is organized by scenes, which are the minimum information that rendering and game engines can currently ingest to create the presentation. The manifest contains a list of the individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene are one or more representations of the geometric objects in the scene, which correspond to a streamable version of the scene geometry. One embodiment of a scene representation refers to a low-resolution version of the geometric objects of the scene. Another embodiment of the same scene refers to an enhancement layer for the low-resolution representation of the scene to add additional detail or increase tessellation to the geometric objects of the same scene. As mentioned above, each scene may have multiple enhancement layers to increase the detail of the geometric objects of the scene in a progressive manner.

[0064] 7. Each layer of a media object referenced in a scene is associated with a token (e.g., a URI) that points to an address where the resource can be accessed in the network. Such resources are similar to CDNs from which content may be fetched by clients. 8. Tokens in representations of geometric objects may point to locations in the network or to locations in the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.

[0065] 3 illustrates an embodiment of a generic media format for timed media as follows: A timed scene manifest contains a list of scene information 301. The scene 301 points to processing information and a list of components 302 that individually describe the types of media assets that make up the scene 301. The components 302 point to assets 303 which further point to a base layer 304 and an attribute enrichment layer 305.

[0066] 4 illustrates an embodiment of a generic media format for non-timed media as follows: Scene information 401 is associated with a start time and an end time according to a clock. Scene information 401 points to processing information and a list of components 402 that individually describe the types of media assets that make up the scene 401. Components 402 point to assets 403 (e.g. visual, audio and haptic assets) which further point to a base layer 404 and an attribute enhancement layer 405. Furthermore, the scene 401 points to other scenes 401 for non-timed media. The scenes 401 also point to timed media scenes.

[0067] FIG. 5 illustrates an embodiment of a process 500 for synthesizing a capture format from natural content. Camera unit 501 captures a human scene using a single camera lens. Camera unit 502 captures a scene with five divergent fields of view by mounting five camera lenses around a ring-shaped object. The arrangement in 502 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven convergent fields of view by mounting seven camera lenses on the inner diameter of a sphere. The arrangement 503 is an exemplary arrangement commonly used to capture light fields for light field or holographic immersive displays. Natural image content 509 is provided as input to a synthesis module 504, which may optionally generate an arbitrary capture neural network model 508 using a neural network training module 505 using a set of training images 506. Another process commonly used instead of the training process 505 is photogrammetry. When the model 508 is created during the process 500 shown in Figure 5, the model 508 becomes one of the assets of the ingestion format 507 for natural content. Example embodiments of the ingestion format 507 include MPI and MSI.

[0068] FIG. 6 illustrates an embodiment of a process 600 for creating synthetic media, such as a capture format for computer generated images. A LIDAR camera 601 captures a point cloud 602 of the scene. A CGI tool, 3D modeling tool, or another animation process for creating synthetic content is used on a computer 603 to create a CGI asset 604 over a network. A motion capture suit 605A equipped with sensors is worn by an actor 605 to capture a digital recording of the actor's 605 motion to generate animated motion capture data 606. Data 602, 604, and 606 are provided as inputs to a synthesis module 607, which may also optionally use a neural network and training data to create a neural network model (not shown in FIG. 6).

[0069] The heterogeneous immersive media rendering and streaming techniques described above may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, FIG. 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.

[0070] Computer software can be coded using any suitable machine code or computer language, or similar mechanisms, that can be assembled, compiled, linked by a computer central processing unit (CPU), graphics processing unit (GPU), or the like to produce code comprising instructions that can be executed directly or via interpretation, microcode execution, or the like.

[0071] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0072] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system 700.

[0073] The computer system 700 may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly associated with conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0074] The input human interface devices may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touch screen 710, a data glove (not shown), a joystick 705, a microphone 706, a scanner 707, and a camera 708 (only one of each is shown).

[0075] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen 710, data gloves (not shown), or joystick 705, but may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 709, headphones (not shown)), visual output devices (such as screens 710 including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, haptic feedback capability, some of which may output two-dimensional visual output or three or more dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0076] The computer system 700 may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 with CDs / DVDs or similar media 721, thumb drives 722, and removable hard drives or solid state drives 723, legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0077] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transitory signals.

[0078] The computer system 700 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, or optical networks. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant networks, and the like. Examples of networks include local area networks such as Ethernet and wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicular and industrial networks including CANBus. Certain networks typically require an external network interface adapter connected to a specific general-purpose data port or peripheral bus 749 (e.g., a USB port of the computer system 700, etc.). Other networks are typically integrated into the core of the computer system 700 by connecting to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 700 may communicate with other entities. Such communications may be, for example, one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., from the CANbus to a particular CANbus device), or bidirectional, to other computer systems using local or wide area digital networks. As noted above, specific protocols and protocol stacks may be used for each of these networks and network interfaces.

[0079] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to core 740 of computer system 700 .

[0080] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 743, hardware accelerators 744 for specific tasks, etc. These devices may be connected via a system bus 748, along with read only memory (ROM) 745, random access memory 746, and internal mass storage 747, such as an internal hard drive, SSD, etc., that is not accessible to the user. In some computer systems, the system bus 748 is accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus 748 or via a peripheral bus 749. Peripheral bus architectures include PCI, USB, etc.

[0081] The CPU 741, GPU 742, FPGA 743 and accelerator 744 may combine to execute certain instructions that may constitute the aforementioned computer code. The computer code may be stored in ROM 745 or RAM 746. Transient data may also be stored in RAM 746, while permanent data may be stored, for example, in internal mass storage 747. A cache memory, which may be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc., may be used to enable fast storage and retrieval from any memory device.

[0082] The computer-readable medium may bear computer code for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0083] By way of example only and not of limitation, the architecture 700, and in particular a computer system having the core 740, may provide functionality as a result of the processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer readable media. Such computer readable media may be media associated with mass storage accessible to the user as introduced above, and with specific storage of the core 740 that is non-transitory in nature, such as the core internal mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core 740. The computer readable media may include one or more memory devices or chips, depending on the particular needs. The software may cause the core 740, and in particular the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain parts of certain processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to the processes defined by the software. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 744) that may operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. References to software may include logic, where appropriate, and vice versa. References to computer-readable media may include circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.

[0084] FIG. 8 illustrates an exemplary network media distribution system 800 supporting a variety of legacy and heterogeneous immersive media capable displays as client endpoints. A content acquisition module 801 captures or creates media using the exemplary embodiment of FIG. 6 or FIG. 5. Ingest formats are created in a content preparation module 802 and then transmitted using a transmission module 803 to one or more client endpoints 804 in the network media distribution system. A gateway may serve customer premises equipment to provide network access to various client endpoints of the network. A set-top box may also serve as customer premises equipment to provide access to aggregated content by a network service provider. A wireless demodulator may act as a mobile network access point for mobile devices (e.g., similar to mobile handsets and displays). In one or more embodiments, a legacy 2D television may be directly connected to a gateway, a set-top box, or a WiFi router. A laptop computer with a legacy 2D display may be a client endpoint connected to a WiFi router. A head-mounted 2D (raster-based) display may also be connected to the router. A lenticular light field display may be to the gateway. The display may be comprised of a local computation GPU, a storage device, and a visual presentation unit that uses ray-based lenticular optics techniques to create multiple views. The holographic display may be connected to a set-top box and may include a local computation CPU, a GPU, a storage device, and a Fresnal pattern wave-based holographic visualization unit. The augmented reality headset may be connected to a wireless demodulator and may include a GPU, a storage device, a battery, and a volumetric visual presentation component. The high-density light field display may be connected to a WiFi router and may include multiple GPUs, CPUs, storage devices, eye trackers, cameras, and high-density ray-based light field panels.

[0085] FIG. 9 illustrates an embodiment of an immersive media delivery module 900 capable of servicing legacy and heterogeneous immersive media enabled displays as previously illustrated in FIG. 8. Content is created or acquired in module 901, which is further embodied for natural and CGI content in FIG. 5 and FIG. 6, respectively. The content is then converted to an ingest format using a network ingest format creation module 902, which is similarly further embodied for natural and CGI content in FIG. 5 and FIG. 6, respectively. The ingested media format is sent to the network and stored in a storage device 903. Optionally, the storage device may reside in the immersive media content producer's network and be accessed remotely by an immersive media network delivery module (not numbered), as indicated by the dashed line bisecting 903. Client and application specific information is optionally available in a remote storage device 904, which may optionally reside remotely in an alternative "cloud" network.

[0086] As shown in Figure 9, a client interface module 905 may act as the primary source and sink of information to perform the primary tasks of the distribution network. In this particular embodiment, module 905 may be implemented in an integrated format with other components of the network. Nevertheless, the tasks represented by module 905 of Figure 9 form essential elements of the disclosed subject matter.

[0087] Module 905 receives information regarding the characteristics and attributes of clients 908, and also gathers requirements regarding applications currently running on 908. This information may be obtained from device 904, or in alternative embodiments, by directly querying the client 908. When directly querying a client 908, it is assumed that a two-way protocol (not shown in FIG. 9) is present and operational, such that the client may communicate directly with interface module 905.

[0088] The interface module 905 also initiates and communicates with the media adaptation and fragmentation module 910 described in FIG. 10. Once the captured media has been adapted and fragmented by the module 910, the media is optionally transferred to an intermediate storage device shown as prepared media for delivery storage device 909. Once the delivery media is prepared and stored on the device 909, the interface module 905 ensures that the immersive client 908 receives the delivery media and corresponding description information 906 via a "pull" request via its network interface 908B, or the client 908 itself may initiate a "pull" request for the media 906 from the storage device 909. The immersive client 908 may optionally use a GPU (or CPU, not shown) 908C. The delivery format of the media is stored in the client's 908 storage device or storage cache 908D. Finally, the client 908 visually presents the media via its visualization component 908A.

[0089] Throughout the process of streaming immersive media to a client 908 , the interface module 905 monitors the status of the client's progress via a client progress and status feedback channel 907 .

[0090] FIG. 10 illustrates a particular embodiment of a media adaptation process in which captured source media may be appropriately adapted to match the requirements of the client 908. The media adaptation module 1001 is comprised of multiple components that facilitate adapting the captured media to the appropriate delivery format of the client 908. These components should be considered as exemplary. In FIG. 10, the adaptation module 1001 receives input network state 1005 to track the current traffic load on the network, client 908 information including attribute and feature description, application characteristics, description and current state, and client neural network model (if available) to help map the shape of the client frustum to the interpolation capabilities of the ingestible immersive media. The adaptation module 1001 ensures that the adapted output is stored in the client adaptation media storage device 1006 as it is created.

[0091] The adaptation module 1001 adapts the particular captured source media to a format suitable for the client using a renderer 1001B or a neural network processor 1001C. The neural network processor 1001C uses the neural network model 1001A. Examples of such neural network processors 1001C include the Deep View Neural Network Model Generator as described in MPI and MSI. If the media is in a 2D format but the client needs it in a 3D format, the neural network processor 1001C can invoke a process that uses highly correlated images from the 2D video signal to derive a volumetric representation of the scene depicted in the video. An example of such a process may be the Neural Radiance Fields from one or several images developed at the University of California, Berkeley. An example of a suitable renderer 1001B may be a modified version of the OTOY Octane renderer (not shown) that is modified to interact directly with the adaptation module 1001. The adaptation module 1001 may optionally use a media compressor 1001D and a media decompressor 1001E depending on the needs of these tools regarding the format of the captured media and the format required by the client 908.

[0092] Figure 11 shows an adaptation media packaging module 1103 that ultimately converts the adaptation media from the media adaptation module 1101 from Figure 10 that is now residing on the client adaptation media storage device 1102. The packaging module 1103 formats the adaptation media from module 1101 into a robust distribution format, such as the exemplary formats shown in Figure 3 or Figure 4. The manifest information 1104A provides the client 908 with a list of scene data it can expect to receive, as well as a list of visual assets and corresponding metadata, and audio assets and corresponding metadata.

[0093] FIG. 12 shows a packetizer module 1202 that “fragments” the adapted media 1201 into individual packets 1203 suitable for streaming to a client 908 . The components and communications depicted in FIG. 13 of sequence diagram 1300 are described as follows: Client endpoint 1301 initiates a media request 1308 to network delivery interface 1302. Request 1308 includes information to identify the media requested by the client, by URN or other standard nomenclature. Network delivery interface 1302 responds to request 1308 with a profile request 1309, which requests that client 1301 provide information about its currently available resources, including computation, storage, battery charge rate, and other information that characterizes the client's current operating state. Profile request 1309 also requests that the client provide one or more neural network models that can be used by the network for neural network inference, and, if such models are available at the client, to extract or interpolate the correct media view to match the characteristics of the client's presentation system. Response 1311 from client 1301 to interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). The interface 1302 then provides the client 1301 with a session ID token 1311. The interface 1302 then requests the capture media server 1303 with a capture media request 1312 that includes the URN or canonical nomenclature name of the media identified in the request 1308. The server 1303 responds to the request 1312 with a response 1313 that includes the capture media token. The interface 1302 then provides the media token from the response 1313 to the client 1301 in a call 1314. The interface 1302 then begins the adaptation process for the requested media at 1308 by providing the capture media token, the client token, the application token, and the neural network model token to the adaptation interface 1304.The interface 1304 requests access to the captured media by providing the captured media token to the server 1303 in a call 1316 to request access to the captured media asset. The server 1303 responds to the request 1316 with the captured media access token in a response 1317 to the interface 1304. The interface 1304 then requests that the media adaptation module 1305 adapt the captured media located in the captured media access token for the client, application and neural network inference model corresponding to the session ID token created in 1313. A request 1318 from the interface 1304 to the module 1305 includes the necessary tokens and session ID. The module 1305 provides the adapted media access token and session ID to the interface 1302 in an update 1319. The interface 1302 provides the adapted media access token and session ID to the packaging module 1306 in an interface call 1320. The packaging module 1306 provides a response 1321 to the interface 1302 with the packaged media access token and the session ID in response 1321. The module 1306 provides the packaged media access token for the packaged asset, the URN, and the session ID to the packaged media server 1307 in response 1322. The client 1301 executes a request 1323 to start streaming the media asset corresponding to the packaged media access token received in message 1321. The client 1301 executes other requests and provides status updates to the interface 1302 in message 1324.

[0094] Figure 14 shows the ingested media format and assets 1002 of Figure 10 optionally composed of two parts of immersive media and assets in a 3D format 1401 and a 2D format 1402. The 2D format 1402 may be a coded video stream containing a single view, for example ISO / IEC 14496 Part 10 advanced video coding, or may be a coded video stream containing multiple views, for example a multi-view compression modification of ISO / IEC 14496 Part 10.

[0095] Figure 15 illustrates the transmission of neural network model information along with a coded video stream. In this figure, coded video stream 1501 includes a neural network model and corresponding parameters carried directly by one or more SEI messages 1501A. In contrast, in coded video stream 1502, one or more SEI messages carry an identifier for the neural network model and its corresponding parameters. In the 1502 scenario, the neural network model and parameters are stored outside the coded video stream, e.g., in 1001A of Figure 10.

[0096] FIG. 16 illustrates the transmission of neural network model information in a captured immersive media asset 1601 (originally shown as item 1401 in FIG. 14) in a 3D format. The media 1601 refers to scenes 1-N, shown as 1602. Each scene 1602 refers to a geometry 1603 and processing parameters 1604. The geometry 1603 may include a reference 1603A to a neural network model. The processing parameters 1604 may also include a reference 1604A to a neural network model. Both 1604A and 1603A may refer to a network model stored directly with the scene, or may refer to an identifier that points to a neural network model that exists outside of the captured media, for example, a network model stored in 1001A in FIG. 10. Some embodiments relate to systems, methods and / or computer-readable media at any possible level of technical detail integration. The computer-readable media may include a computer-readable non-transitory storage medium having computer-readable program instructions thereon that cause a processor to perform operations.

[0097] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), static random access memories (SRAMs), portable compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards or mechanically encoded devices such as ridge structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage media, as used herein, should not be construed as being, per se, a transitory signal, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over an electrical wire.

[0098] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.

[0099] The computer readable program code / instructions for carrying out the operations may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or object oriented programming languages ​​such as Smalltalk, C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer readable program instructions to perform aspects or operations by utilizing state information of the computer readable program instructions to customize the electronic circuitry.

[0100] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, generate means for performing the functions / acts specified in the flowcharts and / or block diagrams or blocks. These computer readable program instructions may be stored in a computer readable storage medium capable of directing a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer readable storage medium on which the instructions are stored comprises an article of manufacture including instructions implementing aspects of the functions / acts specified in the flowcharts and / or block diagrams or blocks. The computer readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device such that the instructions operate on the computer, other programmable apparatus, or other device to perform the functions / operations specified in the flowcharts and / or block diagrams or blocks, and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to produce a computer-implemented process.

[0101] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions that includes one or more executable instructions that perform the specified logical function. The methods, computer systems, and computer-readable media may include additional, fewer, different, or differently arranged blocks than those shown in the drawings. In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the drawings. For example, two blocks shown in succession may in fact be executed simultaneously or substantially simultaneously, or the blocks may be executed in reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, or combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or implements a combination of dedicated hardware and computer instructions.

[0102] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specified control hardware or software code used to implement these systems and / or methods is not limiting of implementation. Thus, the operation and behavior of the systems and / or methods have been described herein without reference to specific software code. It will be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0103] No element, act, or instruction used herein should be construed as critical or essential unless expressly stated. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Furthermore, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." When only one item is intended, the term "a" or similar language is used. Also, as used herein, terms such as "having," "containing," or "having" are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless expressly stated otherwise.

[0104] The description of various aspects and embodiments is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Although combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically recited in the claims and / or specifically disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible embodiments includes each dependent claim in combination with every other claim in the set of claims. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used in this specification have been selected to best explain the principles of the embodiments, practical applications or technical improvements to the technology found in the marketplace, or to enable other skilled in the art to understand the embodiments disclosed herein. [Explanation of symbols]

[0105] 101A Camera or microphone 101B Computer 102 2D image and associated audio sequences 103 Preparation Module 104 Converter Module 105 Adaptation Module 106 Content Delivery Network 107 Pull Requests 108 clients 202 MPD 203 Time Limit 204 Adaptation Set 205 Expression 206 Segment File 301 Scene Information 302 Components 303 Assets 304 Base Layer 305 Attribute enhancement layer 401 Scene Information 402 Components 403 Assets 404 Base Layer 405 Attribute enhancement layer 500 processes 501 Camera unit 502 Camera unit 503 Camera Unit 504 Synthesis Module 505 Training Process 505 Neural Network Training Module 506 Training Images 507 Import Format 508 Capture Neural Network Model 509 Natural Image Content 600 processes 601 LIDAR Camera 602 point cloud data 603 Computer 604 CGI assets 605 Actor 605A Motion Capture Suit 606 Motion Capture Data 607 Synthesis Module 608 Synthetic Media Ingest Format 700 Computer Systems 700 Architecture 701 Keyboard 702 Mouse 703 Trackpad 705 Joystick 706 Microphone 707 Scanner 708 Camera 709 Speaker 710 Touchscreen 720 CD / DVD ROM / RW 721 Media 722 Thumb Drive 723 Solid State Drive 740 cores 741 Central Processing Unit (CPU) 742 Graphics Processing Unit (GPU) 743 Field Programmable Gate Array (FPGA) 744 Hardware Accelerator 745 Read-Only Memory (ROM) 746 Random Access Memory 747 Massive Storage 748 System Bus 749 Surrounding Bus 800 Network Media Distribution System 801 Content Acquisition Module 802 Content Preparation Module 803 Transmitting Module 804 client endpoint 900 Immersive Media Delivery Module 901 Module 901 Content Acquisition / Creation Module 902 Network Import Format Creation Module 903 Import storage device 904 Remote Storage Device 905 Module 905 Client Interface Module 906 Media and written information 907 Client Progress and Status Feedback Channel 908 Immersive Client 908A Visualization Components 908B Network Interface 908D Memory Cache 909 Distribution storage device 910 Media Adaptation and Fragmentation Module 1001 Adaptation Module 1001A Neural Network Model 1001B Renderer 1001C Neural Network Processor 1001D Media Compressor 1001E Media Decompressor 1002 Assets 1005 Input network state 1006 Client-adaptive media storage device 1101 Media Adaptation Module 1102 Current Client Compatible Media Storage Device 1103 Adaptive Media Packaging Module 1104A Manifest Information 1201 Adaptable Media 1202 Packetizer Module 1203 packets 1204 client endpoint 1300 Sequence Diagram 1301 client endpoint 1302 Network Distribution Interface 1303 Ingest Media Server 1304 Adaptation Interface 1305 Media Adaptation Module 1306 Packaging Module 1307 Packaged Media Server 1401 3D Immersive Media and Assets 1402 2D Immersive Media and Assets 1501 Coded Video Stream 1501A SEI Message 1502 coded video stream 1502A SEI Message 1601 3D Immersive Media and Assets 1602 scenes 1603 Shape See 1603A 1604 Processing parameters See 1604A

Claims

1. 1. A processor-executable method for streaming immersive media, comprising: obtaining information indicative of a characteristic of a client endpoint; capturing content in a first two-dimensional format or a first three-dimensional format, the content including one or more SEI messages; the one or more SEI messages include a reference to a neural network model for transforming the content; the reference to the neural network model indicates whether the neural network model is associated with a scene in the content; and converting the content into a second two-dimensional format or a second three-dimensional format adapted to the characteristics of the client endpoint based on the neural network consulted; and streaming the converted content to a client endpoint.

2. The reference to the neural network model indicates that the neural network model is stored with the scene in the content; or The method of claim 1 , wherein the reference to the neural network model indicates whether the neural network model exists outside the content.

3. The method of claim 1 , wherein the reference to the neural network model includes metadata identifying a location of the neural network model.

4. The method of claim 3 , wherein the reference to the neural network model includes a universal resource identifier that corresponds to the metadata that describes the content.

5. The method of claim 1 , wherein the neural network is trained prior to ingesting the content based on a prior distribution corresponding to objects in the content.

6. The method of claim 1, wherein the content is transformed based on characteristics of the client endpoint.

7. 10. The method of claim 1, wherein the one or more client endpoints include one or more of a television, a computer, a head mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a high density light field display.

8. A computer system configured to cause one or more computer processors to carry out the method according to any one of claims 1 to 7.

9. A computer program for causing a computer to carry out the method according to any one of claims 1 to 7.