Method for streaming immersive media, computer system and computer program therefor
A network-based media delivery system adapts immersive media using neural networks to convert and interpolate formats, addressing the heterogeneity of immersive devices and enabling efficient delivery to diverse endpoints.
Patent Information
- Application Number
- JP2024075250
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-20
- Filing Date
- 2024-05-07
- Publication Date
- 2025-12-17
- Estimated Expiration
- 2041-09-01
AI Technical Summary
Current systems lack a coherent end-to-end ecosystem for delivering immersive media over commercial networks due to the heterogeneity of immersive media-capable devices and the lack of a single standard representation that can accommodate real-time and non-real-time distribution of both natural and computer-generated content, as well as varying display capabilities and requirements.
A network-based media delivery system that adapts immersive media sources to formats suitable for specific client endpoint devices by employing neural networks to interpolate and convert media formats, using scene-specific models and layering processes, enabling support for both legacy and immersive displays.
The system efficiently delivers immersive media to heterogeneous client endpoints, improving visual quality and supporting both real-time and non-real-time delivery, while accommodating diverse display capabilities and requirements.
Smart Images

Figure 0007787939000001 
Figure 0007787939000002 
Figure 0007787939000003
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 127036, filed with the United States Patent and Trademark Office (December 17, 2020), and U.S. Patent Application No. 17 / 407816, filed with the United States Patent and Trademark Office (August 20, 2021), the entire contents of which are incorporated herein by reference. [Technical field] FIELD OF THE DISCLOSURE This disclosure relates generally to the field of data processing, and more particularly to video coding. [Background technology]
[0002] "Immersive media" generally refers to media that stimulates any or all of the human sensory systems (visual, auditory, somatosensory, olfactory, and sometimes gustatory) to create or enhance the user's perception of being physically present in the media experience, i.e., other than what is distributed over existing commercial networks for timed two-dimensional (2D) video and corresponding audio, known as "legacy media." Both immersive and legacy media can be characterized as timed or non-timed.
[0003] Timed media refers to media that is structured and presented according to time. Examples include movie features, news reports, and episodic content, all of which are organized according to durations. Legacy video and audio are commonly considered timed media.
[0004] Non-timed media is media that is structured by logical, spatial, and / or temporal relationships rather than by time. An example is a video game where the user has control over the experience created by the gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media also does not incorporate timed media, for example, in the continuously repeating audio or video segments of a video game scene. Conversely, timed media also does not incorporate non-timed media, such as a video with a fixed still image as a background.
[0005] Immersive media-capable devices are devices that have the capability to access, interpret, and present immersive media. Such media and devices are heterogeneous in terms of the amount and format of media and the number and type of network resources required to deliver such media at scale, i.e., over a network comparable to the delivery of legacy video and audio media. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogenous in capabilities because they all consist of rectangular display screens and use 2D rectangular video or still images as their primary media format. Summary of the Invention [Means for solving the problem]
[0006] Embodiments relate to methods, systems, and computer-readable media for streaming immersive media. In one aspect, a method for streaming immersive media is provided. The method may include capturing content in a first two-dimensional format or a first three-dimensional format, where the format may reference a neural network. Based on the referenced neural network, converting the captured content into a second two-dimensional format or a second three-dimensional format. Streaming the converted content to a client endpoint, such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a high-density light field display.
[0007] In another aspect, a computer system for streaming immersive media is provided. The computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices, executed by at least one of the one or more processors via at least one of the one or more memories, thereby performing a method. The method may include capturing content in a first two-dimensional format or a first three-dimensional format, the format referencing a neural network. Based on the referenced neural network, converting the captured content into a second two-dimensional format or a second three-dimensional format. Streaming the converted content to a client endpoint, such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a high-density light field display.
[0008] In yet another aspect, a computer-readable medium for streaming immersive media is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored on at least one of one or more tangible storage devices that are executable by a processor. The program instructions are executable by the processor to perform a method that captures content in a first two-dimensional format or a first three-dimensional format, where the format may accordingly include referencing a neural network. Based on the referenced neural network, converts the captured content into a second two-dimensional format or a second three-dimensional format. Streams the converted content to a client endpoint, such as a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, or a high-density light field display.
[0009] These and other objects, features and advantages will become apparent from the following detailed description of illustrative embodiments, which is to be read in connection with the accompanying drawings. Because the figures are intended to facilitate understanding of the invention by those skilled in the art together with the detailed description, various features of the drawings are not drawn to scale. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram of an end-to-end process for timed legacy media delivery. [Figure 2] FIG. 1 is a schematic diagram of a standard media format used for streaming timed legacy media. [Figure 3] FIG. 1 is a schematic diagram of an embodiment of a data model for representing and streaming timed immersive media. [Figure 4] FIG. 1 is a schematic diagram of an embodiment of a data model for representing and streaming non-timed immersive media. [Figure 5]1 is a schematic diagram of a process for capturing a natural scene and converting it into a representation that can be used as an ingestion format for a network that serves heterogeneous client endpoints. [Figure 6] 1 is a schematic diagram of a process using 3D modeling tools and formats to create a representation of a synthetic scene that can be used as an ingestion format for a network serving heterogeneous client endpoints. [Figure 7] FIG. 1 is a system diagram of a computer system. [Figure 8] 1 is a schematic diagram of a network serving multiple heterogeneous client endpoints; [Figure 9] For example, this is a schematic diagram of a network that provides adaptation information about particular media represented in a media capture format prior to the network's process of adapting the media for consumption by a particular immersive media client endpoint. [Figure 10] FIG. 1 is a flow diagram of a media adaptation process consisting of a media render converter that converts source media from an ingest format into a specific format suitable for a particular client endpoint. [Figure 11] 1 is a schematic diagram of a network that formats adapted source media into a data model suitable for presentation and streaming. [Figure 12] FIG. 12 is a flow diagram of a media streaming process that fragments the data model of FIG. 11 into the payload of network protocol packets. [Figure 13] FIG. 1 is a sequence diagram of a network adapting a particular immersive media in an ingestion format to a streamable and appropriate delivery format for a particular immersive media client endpoint. [Figure 14] FIG. 10 is a schematic diagram of the ingested media formats and assets 1002 of FIG. 10, consisting of both immersive and legacy content formats, i.e., 2D video formats only, or both immersive and 2D video formats. [Figure 15] 1 illustrates the transmission of neural network model information along with a coded video stream. [Figure 16] 1 illustrates the transmission of neural network model information along with input immersive media and assets. DETAILED DESCRIPTION OF THE INVENTION
[0011] Detailed embodiments of the claimed structures and methods are disclosed herein. However, it is understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which may be embodied in various forms. However, these structures and methods may be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In the description, details of well-known features and techniques may be omitted so as not to unnecessarily obscure the presented embodiments.
[0012] Embodiments relate generally to the field of data processing, and more specifically, to video coding. The techniques described herein enable a network to ingest a 2D video source of media containing one or more (typically a small number of) views, adapt the source of 2D media into one or more streamable "delivery formats," and signal a scene-specific neural network model for the 2D coded video stream to accommodate the requirements of various heterogeneous client endpoint devices, their varying features and capabilities, and the applications used by the client endpoints, before actually delivering the formatted media to various client endpoints. The network model may be embedded directly into the scene-specific coded video stream of the coded bitstream using an SEI structured field, or the SEI may signal the use of a specific model stored elsewhere in the delivery network but accessible to the neural network process. The ability to reformat a 2D media source into various streamable delivery formats enables a network to simultaneously serve various client endpoints with different capabilities and available computational resources, enabling support for new immersive client endpoints, such as holographic and light field displays, on commercial networks. Additionally, the ability to adapt scene-specific 2D media sources based on scene-specific neural network models improves final visual quality. This ability to adapt 2D media sources is particularly important when no immersive media sources are available and when clients cannot support delivery formats based on 2D media. In this scenario, neural network-based approaches can be more optimally used with specific scenes present in 2D media by retaining scene-specific neural network models trained with priors that are generally similar to the objects in a particular scene or the context of a particular scene.This improves the network's ability to infer depth-based information about a particular scene, and allows it to adapt 2D media into scene-specific volume formats appropriate for the target client endpoint.
[0013] As previously mentioned, "immersive media" generally refers to media that stimulates any or all of the human sensory systems (visual, auditory, somatosensory, olfactory, and sometimes gustatory) to create or enhance the user's perception of being physically present in the media experience, i.e., other than what is distributed over existing commercial networks for timed two-dimensional (2D) video and corresponding audio, known as "legacy media." Both immersive and legacy media can be characterized as timed or non-timed. Timed media refers to media that is structured and presented according to time. Examples include movie features, news reports, and episodic content, all of which are organized according to durations. Legacy video and audio are commonly considered timed media. Non-timed media is media that is structured by logical, spatial, and / or temporal relationships rather than by time. An example is a video game where the user has control over the experience created by the gaming device. Another example of non-timed media is a still image photograph taken by a camera. Non-timed media also does not incorporate timed media, for example, in the continuously repeating audio or video segments of a video game scene. Conversely, timed media also does not incorporate non-timed media, such as a video with a fixed still image as a background.
[0014] Immersive media-capable devices are devices that have the capability to access, interpret, and present immersive media. Such media and devices are heterogeneous in terms of the amount and format of media and the number and type of network resources required to deliver such media at scale, i.e., over a network comparable to the delivery of legacy video and audio media. In contrast, legacy devices such as laptop displays, televisions, and mobile handset displays are homogenous in capabilities because they all consist of rectangular display screens and use 2D rectangular video or still images as their primary media format.
[0015] The distribution of any media over a network may use media distribution systems and architectures that reformat the media from an input or network "ingest" format into a final distribution format that is not only suitable for the target client device and its applications, but also lends itself to streaming over the network. "Streaming" media broadly refers to the fragmentation and packetization of source media so that it can be delivered over the network in successive, small-sized "chunks" that are logically organized and ordered according to either or both the temporal and spatial structure of the media. In such distribution architectures and systems, the media may undergo a compression or layering process so that only the most salient media information is delivered to the client initially. In some cases, the client must receive all of the significant media information of a portion of media before presenting any of the same media portions to the end user.
[0016] The process of reformatting input media to match the capabilities of a target client endpoint may employ a neural network process that uses a network model that may encapsulate some prior knowledge of the particular media being reformatted. For example, a particular model may be tuned to recognize outdoor park scenes (comprising trees, plants, grass, and other objects commonly found in park scenes), while another particular model may be tuned to recognize indoor dinner scenes (comprising a dinner table, cooking utensils, people sitting at the table, etc.). Those skilled in the art will recognize that a neural network process with a network model tuned to recognize objects from a particular context, e.g., park scene objects, and a network model tuned to match the content of the particular scene, will produce better visual results than a less tuned network model. Thus, there is an advantage to providing a scene-specific network model for the neural network process tasked with reformatting input media to match the capabilities of a target client endpoint.
[0017] A mechanism for associating a neural network model with a particular scene of 2D media can be achieved by optionally compressing the network model and inserting it directly into the 2D coded bitstream of the visual scene using the Supplemental Enhancement Information (SEI) structured field commonly used to attach metadata to coded video streams in H.264, H.265, and H.266 video compression formats. The presence of an SEI message containing a particular neural network model within the context of a portion of the coded video bitstream may be used to indicate that the network model is to be used to interpret and adapt the video content within the portion of the bitstream in which the model is embedded. Alternatively, the SEI message can be used to signal, via a network model identifier, which neural network model can be used in the absence of the actual model itself.
[0018] The mechanism for associating an appropriate neural network with immersive media may be achieved by the immersive media itself referencing the appropriate neural network model to use. This referencing may be achieved by directly embedding the network model and its parameters on a per-object, per-scene basis, or a combination thereof. Alternatively, rather than embedding one or more neural network models within the media, a media object or scene may reference a particular neural network model by an identifier.
[0019] Yet another alternative mechanism for referencing an appropriate neural network for adaptation of media for streaming to a client endpoint is for the particular client endpoint itself to provide at least one neural network model and corresponding parameters to the adaptation process to use. Such a mechanism may be implemented by the client providing the neural network model in communication with the adaptation process, for example, when the client connects itself to the network. After adapting the video to the target client endpoint, the adaptation process in the network may choose to apply a compression algorithm to the result, which may optionally separate the adapted video signal into layers corresponding to the most to least salient portions of the visual signal.
[0020] An example of a compression and layering process is the progressive format of the JPEG standard (ISO / IEC 10918 Part 1), which separates the image into layers so that the entire image is first out of focus and presented with only basic shapes and colors, i.e., from the low-order DCT coefficients of the entire image scan, and then separates additional layers of detail so that the image is brought into focus, i.e., from the high-order DCT coefficients of the image scan.
[0021] The process of breaking media into smaller parts, organizing them into payload portions of successive network protocol packets, and delivering these protocol packets is called "streaming" the media, whereas the process of converting the media into a format suitable for presentation at one of a variety of heterogeneous client endpoints operating one of a variety of heterogeneous applications is known as "adapting" the media.
[0022] definition A scene graph is a generic data structure commonly used by vector-based graphics editing applications and modern computer games that arranges a logical and often (but not necessarily) spatial representation of a graphical scene; it is also a collection of nodes and vertices in a graph structure.
[0023] A node is a basic element of a scene graph that consists of information related to the logical, spatial or temporal representation of visual, auditory, tactile, olfactory, gustatory or related processing information; each node must have at most one outgoing edge, zero or more incoming edges, and at least one edge (either incoming or outgoing) connected to it. The base layer is typically a nominal representation of the asset that is created to minimize the computational resources or time required to render the asset or the time to transmit the asset over a network.
[0024] An enhancement layer is a set of information that, when applied to a base layer representation of an asset, extends the base layer to include features or capabilities not supported in the base layer. Attributes are metadata associated with a node that are used to describe particular properties or characteristics of the node in standard or more complex forms (eg, in terms of another node).
[0025] A container is a serialized format for storing and exchanging information representing an entire natural scene, an entire synthetic scene, or a combination of synthetic and natural scenes, including a scene graph and all the media resources required to render the scene.
[0026] Serialization is the process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link) and then reconstructed (e.g., in another computing environment). The resulting series of bits, when reread according to the serialized format, can be used to create a semantically identical clone of the original object.
[0027] A renderer is a (usually software-based) application or process that, given an input scene graph and asset container, transmits, typically, visual and / or audio signals suitable for presentation on a target device or conforming to desired characteristics specified by attributes of the render target nodes in the scene graph, based on a selective combination of disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development. In the case of visual-based media assets, a renderer may transmit visual signals suitable for a target display or suitable for storage as an intermediate asset (e.g., repackaged into another container and used in the series of rendering processes in a graphics pipeline); in the case of audio-based media assets, a renderer may transmit audio signals for presentation over multi-channel speakers and / or binaural headphones, or repackaged into another (output) container. Common examples of renderers include Unity and Unreal.
[0028] Evaluation is the generation of a result that changes the output from abstract to concrete (e.g., similar to evaluating a document object model of a web page).
[0029] A scripting language is an interpreted programming language that can process dynamic inputs and mutable state changes applied to scene graph nodes that are executed by the renderer at runtime to affect the rendering and evaluation of spatial and temporal object topology (including physical forces, constraints, IK, deformations, collisions) and energy propagation and transfer (light, sound).
[0030] A shader is a type of computer program originally used for shading (producing appropriate levels of light, darkness, and color in an image), but now used to perform various special functions in various areas of computer graphics special effects, video post-processing unrelated to shading, and functions completely unrelated to graphics.
[0031] Path tracing is a computer graphics method of rendering three-dimensional scenes so that the lighting in the scene is realistic. Timed media is media that is ordered by time, for example, a start and end time according to a particular clock.
[0032] Non-timed media is media that is organized by spatial, logical, or temporal relationships, such as an interactive experience that is realized according to actions performed by a user. A neural network model is a collection of parameters and tensors (e.g., matrices) that define weights (i.e., numbers) used in well-defined mathematical operations applied to a visual signal to arrive at an improved visual output, including the interpolation of new views of the visual signal that were not explicitly provided by the original signal.
[0033] Immersive media can be considered one or more types of media that, when presented to humans by an immersive-media-enabled device, stimulate any of the five senses—sight, sound, taste, touch, and hearing—in a manner that is more realistic and consistent with human understanding of experiences in the natural world, i.e., beyond what would be achieved with legacy media presented by a legacy device. In this context, the term "legacy media" refers to two-dimensional (2D) visual media, still or video frames, and / or corresponding audio whose user interaction capabilities are limited to pausing, playing, fast-forwarding, or rewinding, and "legacy devices" refer to televisions, laptops, displays, and mobile devices whose capabilities are limited to presenting only legacy media. In consumer application scenarios, immersive media presentation devices (i.e., immersive-media-enabled devices) are consumer hardware devices specifically equipped with the capabilities to leverage the specific information embodied by immersive media to create presentations that more closely approximate human understanding and interaction with the physical world—capabilities beyond those of legacy devices. Legacy devices are constrained in their ability to present only legacy media, whereas immersive media devices are similarly unconstrained.
[0034] Over the past decade, many immersive media-enabled devices have been introduced to the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, haptic gloves, and gaming consoles. Similarly, holographic displays and other forms of volumetric displays are poised to emerge within the next decade. Despite the immediate or imminent availability of these devices, a coherent end-to-end ecosystem for delivering immersive media over commercial networks has not materialized for several reasons.
[0035] One of these reasons is the lack of a single standard representation of immersive media that can address the two primary use cases associated with current large-scale media distribution over commercial networks: 1) real-time distribution of live-action events, i.e., content is created and delivered to client endpoints in real-time or near real-time, and 2) non-real-time distribution, where content does not necessarily need to be delivered in real-time, i.e., content is physically captured or created. These two use cases may be compared equally to currently existing "broadcast" and "on-demand" distribution formats, respectively.
[0036] For real-time delivery, content can be captured by one or more cameras or created using computer-generated technology. Content captured by a camera is referred to herein as “natural” content, and content created using computer-generated technology is referred to herein as “synthetic” content. Media formats representing synthetic content can be formats used in the 3D modeling, visual effects, and CAD / CAM industries and can include object formats and tools such as meshes, textures, point clouds, structured volumes, amorphous volumes (e.g., for fire, smoke, and fog), shaders, procedurally generated shapes, materials, lighting, virtual camera definitions, and animations. Although synthetic content is computer-generated, synthetic media formats can be used for both natural and synthetic content. However, the process of converting natural content into synthetic media formats (e.g., synthetic representations) can be a time-consuming and computationally intensive process, making it impractical for real-time applications and use cases.
[0037] For real-time delivery of natural content, the content captured by the camera can be delivered in a raster format, which is suitable for legacy display devices because many legacy display devices are designed to display raster formats as well, i.e., delivery of a raster format is best suited for displays that can only display raster formats, since legacy displays are designed to uniformly display raster formats.
[0038] However, immersive media-capable displays are not necessarily limited to displaying raster-based formats. Furthermore, some immersive media-capable displays are not capable of presenting media that is available only in raster-based formats. The availability of displays optimized to create immersive experiences based on formats other than raster-based formats is another important reason why there is not yet a coherent end-to-end ecosystem for the delivery of immersive media. Yet another challenge in creating a coherent distribution system for multiple different immersive media devices is that current and emerging immersive media-enabled devices themselves can vary significantly. For example, some immersive media devices, such as head-mounted displays, are explicitly designed to be used by only one user at a time. Other immersive media devices are designed for simultaneous use by multiple users; for example, the "Looking Glass Factory 8K Display" (hereinafter referred to as a "lenticular light field display") can display content that can be viewed by up to 12 users simultaneously, where each user experiences their own unique perspective (i.e., view) of the displayed content.
[0039] Further complicating the development of coherent delivery systems is the fact that the number of unique views each display can generate can vary significantly. Legacy displays are often only capable of creating a single view of content. Lenticular light field displays, on the other hand, can support multiple users, each experiencing their own unique view of the same visual scene. To achieve the creation of multiple views of the same scene, lenticular light field displays create a specific volumetric viewing frustum, requiring 45 unique views of the same scene as input to the display. This means that 45 slightly different, unique raster representations of the same scene must be captured and delivered to the display in a format specific to that particular display—that viewing frustum. In contrast, the viewing frustum of legacy displays is limited to a single two-dimensional plane, preventing the presentation of multiple viewing perspectives of content through the display's viewing frustum, regardless of the number of viewers experiencing the display simultaneously.
[0040] In general, immersive media displays can vary significantly depending on the characteristics of all displays, such as the dimensions and volume of the viewing frustum, the number of simultaneous viewers supported, the optical technology used to fill the viewing frustum, which can be point-based, ray-based, or wave-based technology, the density of light units (either points, rays, or waves) that occupy the viewing frustum, the availability of computing power, the type of computing (CPU or GPU), the source and availability of power (battery or wire), the amount of local storage or caching, and access to auxiliary resources such as cloud-based computing and storage. These characteristics contribute to the heterogeneity of immersive media displays, in contrast to the homogeneity of legacy displays, which complicates the development of a single delivery system that can support all displays, including both legacy and immersive types of displays.
[0041] The disclosed subject matter addresses the development of a network-based media delivery system capable of supporting both legacy and immersive media displays as client endpoints within the context of a single network. Specifically, presented herein is a mechanism for adapting an input immersive media source to a format appropriate for the specific characteristics of a client endpoint device, including an application currently running on the client endpoint device. Such a mechanism for adapting an input immersive media source includes matching the characteristics of the input immersive media with the characteristics of a target endpoint client device, including the application running on the client device, and adapting the input immersive media to a format appropriate for the target endpoint and its application. Furthermore, the adaptation process may include interpolating additional views, such as novel views, from the input media to create additional views required by the client endpoint. Such interpolation may be performed utilizing a neural network process.
[0042] It should be noted that the remainder of the disclosed subject matter assumes, without loss of generality, that the process of adapting an input immersive media source to a particular endpoint client device is the same as or similar to the process of adapting the same input immersive media source to a particular application running on a particular client endpoint device, i.e., the problem of adapting an input media source to the characteristics of an endpoint device has the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.
[0043] Legacy devices supported by legacy media achieve widespread consumer adoption because they are similarly supported by an ecosystem of legacy media content providers that produce standards-based representations of the legacy media, and by commercial network service providers that provide the network infrastructure for connecting legacy devices to sources of standard legacy content. In addition to their role in delivering legacy media over their networks, commercial network service providers may also facilitate pairing of legacy client devices with access to legacy content on a content delivery network (CDN). When paired with access to the appropriate form of content, legacy client devices can request or "pull" the legacy content from a content server to the device for presentation to the end user. Nevertheless, an architecture in which a network server "pushes" the appropriate media to the appropriate client is equally relevant without introducing additional complexity into the overall architecture and solution design.
[0044] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0045] The exemplary embodiments described below relate to system and network architectures, structures, and components for delivering media including video, audio, geometric (3D) objects, haptics, associated metadata, or other content to client devices. Particular embodiments are directed systems, structures, and architectures for delivering media content to heterogeneous immersive and interactive client devices.
[0046] Figure 1 illustrates an example of the end-to-end process for timed legacy media distribution. In Figure 1, timed audiovisual content is captured by a camera or microphone at 101A or computer-generated at 101B, creating a sequence of 2D images and associated audio 102 that is input into a preparation module 103. The output of 103 is edited content (e.g., for post-production, including language translation, subtitling, and other editing functions) and is ready to be converted by a converter module 104 into a standard mezzanine format, e.g., for on-demand media, or a standard contribution format, e.g., for live events. Media is "ingested" by a commercial network service provider, and an adaptation module 105 packages the media into various bit rates, temporal resolutions (frame rates), or spatial resolutions (frame sizes) packaged into a standard distribution format. The resulting adaptations are stored in a content delivery network 106, from which various clients 108 make pull requests 107 to fetch the media and present it to end users. It is important to note that the master format may be comprised of a hybrid of media from both 101A or 101B, and format 101A may be obtained in real time from media obtained from, for example, a live sporting event. Additionally, while the client 108 is responsible for selecting the particular adaptation 107 that best suits the client's configuration and / or current network conditions, it is equally possible for a network server (not shown in FIG. 1) to determine and then "push" the appropriate content to the client 108.
[0047] FIG. 2 illustrates an example of a standard media format used for the distribution of legacy timed media, such as video, audio, and supporting metadata (including timed text, such as used for subtitling). As described in item 106 of FIG. 1, the media is stored on a CDN 201 in a standards-based distribution format. The standards-based format is illustrated as an MPD 202 composed of multiple parts, each containing a timed period 203 with a start and end time corresponding to a clock. Each period 203 points to one or more adaptation sets 204. Each adaptation set 204 is typically used for a single type of media, such as video, audio, or timed text. For any given period 203, multiple adaptation sets 204 may be provided, e.g., one adaptation set for video and multiple adaptation sets for audio, such as those used for translation into various languages. Each adaptation set 204 points to one or more representations 205 that provide information about the frame resolution (for video), frame rate, and bitrate of the media. Multiple representations 205 may be used to provide access to a representation 205 for each of ultra-high definition, high definition, or standard definition video, for example. Each representation 205 points to one or more segment files 206 where the media is actually stored for fetching by a client (shown as 108 in FIG. 1) or delivery (in a "push-based" architecture) by a network media server (not shown in FIG. 1).
[0048] Figure 3 is an example representation of a streamable format for timed heterogeneous immersive media. Figure 4 is an example representation of a streamable format for non-timed heterogeneous immersive media. Both figures refer to a scene: Figure 3 refers to a timed media scene 301, and Figure 4 refers to a non-timed media scene 401. In both cases, the scene can be embodied by various scene representations or scene descriptions.
[0049] For example, in some immersive media designs, a scene may be embodied by a scene graph, as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI and MSI technologies are examples of technologies that support creating view-agnostic scene representations for natural content, i.e., real-world images captured simultaneously by one or more cameras. On the other hand, scene graph technologies may be used to represent both natural and computer-generated imagery in the form of synthetic representations, but such representations are particularly computationally intensive to create when the content is captured as a natural scene by one or more cameras. That is, scene graph representations of naturally captured content are time- and computationally intensive to create, requiring complex analysis of the natural imagery via photogrammetry, deep learning, or both techniques, to create a synthetic representation that can later be used to interpolate a sufficient and appropriate number of views to fill the viewing frustum of the target immersive client display. As a result, such synthetic representations are currently unrealistically considered as candidates for representing natural content because they cannot be created in real time to consider use cases requiring real-time delivery. Nevertheless, currently the best candidate representation for computer-generated images is to use a scene graph in the synthetic model, as computer-generated images are created using 3D modeling processes and tools.
[0050] This dichotomy in the optimal representation of both natural and computer-generated content suggests that the optimal capture format for naturally captured content will be different from the optimal capture format for computer-generated content or natural content that is not essential to real-time delivery applications. Accordingly, the disclosed subject matter aims to be robust enough to support multiple capture formats for visually immersive media, regardless of whether the content is naturally or computer-generated.
[0051] The following are examples of techniques that embody scene graphs as a format suitable for representing visually immersive media created using computer-generated techniques, or corresponding synthetic representations of natural scenes using deep learning or photogrammetry techniques, i.e., naturally captured content that is not essential for real-time distribution applications. 1. ORBX® by OTOY ORBX by OTOY is one of several scene graph technologies capable of supporting any type of visual media, timed or untimed, including ray-traceable, legacy (frame-based), volumetric, and other types of synthetic or vector-based visual formats. ORBX differs from other scene graphs because it provides native support for freely available and / or open-source formats for meshes, point clouds, and textures. ORBX is a scene graph purposefully designed to facilitate interchange between multiple vendor technologies that operate on the scene graph. Additionally, ORBX offers a rich material system, support for open shading languages, a robust camera system, and support for Lua scripting. ORBX is also the basis for an immersive technical media format released for royalty-free licensing by the Immersive Digital Experience Alliance (IDEA). In the context of real-time distribution of media, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of camera-captured data and composition of that same data into a synthetic representation. To date, the availability of sufficient compute for real-time distribution is impractical, but not impossible.
[0052] 2. Pixar's Universal Scene Description Pixar's Universal Scene Description (USD) is another well-known and mature scene graph that is popular in the VFX and professional content creation communities. USD is integrated into Nvidia's Omniverse platform, a toolset for developers to create and render 3D models using Nvidia's GPUs. A subset of USD was released by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.
[0053] 3. glTF 2.0 by Khronos glTF 2.0 is the latest version of the "Graphics Language Transmission Format" specification created by the Khronos3D group. This format supports a simple scene graph format that can generally support static (untimed) objects in a scene, including image formats such as "png" and "jpeg." glTF 2.0 supports simple animation, including support for moving, rotating, and scaling basic shapes, or geometric objects, described using glTF primitives. glTF 2.0 does not support timed media, and therefore does not support video or audio. These known designs for immersive visual media scene representations are provided as examples only and do not limit the disclosed subject matter in its ability to specify a process for adapting an input immersive media source to a format suited to the particular characteristics of a client endpoint device.
[0054] Additionally, any or all of the above example media representations currently use or may use deep learning techniques to train or create neural network models that enable or facilitate the selection of specific views to fill a particular display's viewing frustum based on the particular dimensions of the frustum. The views selected for a particular display's viewing frustum may be interpolated from existing views explicitly provided in the scene representation, i.e., from MSI or MPI techniques, or may be rendered directly from these rendering engines based on particular virtual camera positions, filters, or virtual camera descriptions of those rendering engines.
[0055] Thus, the disclosed subject matter is sufficiently robust to consider that there is a relatively small but well-known set of immersive media capture formats that can adequately meet the requirements for both real-time or "on-demand" (e.g., non-real-time) delivery of media that is captured naturally (e.g., by one or more cameras) or created using computer-generated techniques.
[0056] With the introduction of advanced network technologies such as 5G for mobile networks and fiber optic cable for fixed networks, the interpolation of views from immersive media ingest formats using either neural network models or network-based rendering engines will become even easier. These advanced network technologies will improve the capacity and capabilities of commercial networks, as such advanced network infrastructures can support the transport and delivery of increasingly large amounts of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networking (SDN), and network functions virtualization (NFV) enable commercial network service providers to flexibly deploy their network infrastructure to adapt to changes in demand for certain network resources, for example, to respond to dynamic increases or decreases in demand for network throughput, network speed, round-trip delay, and computational resources. Furthermore, this inherent ability to adapt to dynamic network requirements to support a variety of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints will similarly facilitate the network's ability to adapt immersive media ingest formats to appropriate delivery formats.
[0057] Immersive media applications themselves may have varying requirements for network resources, including gaming applications that require significantly low network latency to respond to real-time updates on the state of the game, telepresence applications that have symmetric throughput requirements on both the uplink and downlink portions of the network, and passive viewing applications that may have increased demands on downlink resources depending on the type of client endpoint display that is consuming the data. Consumer-oriented applications are typically supported by a variety of client endpoints with different on-board client capabilities in terms of storage, computation, and power, and different requirements for the particular media presentation.
[0058] Thus, the disclosed subject matter enables a fully equipped network, i.e., a network that uses some or all of the characteristics of a modern network, to simultaneously support multiple legacy and immersive media capable devices in accordance with the features specified therein, which features are as follows: 1-7.
[0059] 1. Providing the flexibility to leverage realistic media ingest formats for both real-time and "on-demand" media delivery use cases. 2. Provides flexibility to support both natural and computer-generated content for both legacy and immersive media-enabled client endpoints. 3. Support both timed and non-timed media. 4. Provide a process that dynamically adapts the ingest format of source media to an appropriate delivery format based on the characteristics and capabilities of the client endpoint and the requirements of the application. 5. Ensure delivery formats are streamable over IP-based networks. 6. Allows the network to simultaneously serve multiple heterogeneous client endpoints, which may include both legacy and immersive media capable devices. 7. We provide an exemplary media representation framework that facilitates the organization of distributed media along scene boundaries. The improved end-to-end implementation enabled by the disclosed subject matter is achieved according to the processes and components described in the detailed description of FIGS. 3-16 as follows.
[0060] Both Figures 3 and 4 use a single exemplary generic delivery format that is adapted from an ingestion source format to match the capabilities of a particular client endpoint. As noted above, the media shown in Figure 3 is timed, and the media shown in Figure 4 is non-timed. The particular generic format is sufficiently robust in its structure to accommodate a wide variety of media attributes, where each attribute can be layered based on the amount of salient information each layer contributes to the media presentation. Note that such layering processes are already well-known in the current state of the art, as demonstrated by progressive JPEG and scalable video architectures such as those specified in ISO / IEC 14496-10 (Scalable Advanced Video Coding).
[0061] 1. Media streamed pursuant to the generic media format is not limited to legacy visual and audio media, but may include any type of media information that can interact with a machine to produce signals that stimulate the human senses of sight, hearing, taste, touch, and smell. 2. Media streamed according to the generic media format may be timed media, untimed media, or a combination of both. 3. The generic media format is further streamable by enabling layered representations of media objects using a base and enhancement layer architecture. In one example, separate base and enhancement layers are computed by applying multi-resolution or multi-tessellation analysis techniques to the media objects in each scene. This is similar to the progressively rendered image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In an exemplary embodiment, the progressive representation of a geometric object may be a multi-resolution representation of the object computed using wavelet analysis. In another example of a layered representation of a media format, the enhancement layer applies various attributes to the base layer, such as improving the material properties of the surface of the visual object represented by the base layer. In yet another example, the attributes may improve the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface. In yet another example of a layered representation, the surfaces of one or more visual objects in a scene may be modified from Lambertian to ray-traceable. In yet another example of layered representations, the network delivers a base layer representation to a client, so that the client can create a nominal presentation of the scene while waiting for the transmission of additional enhancement layers to improve the resolution or other properties of the base representation.
[0062] 4. The resolution of the enhancement layer attributes or refinement information is not explicitly coupled to the resolution of the base layer objects, as in the currently existing MPEG video and JPEG image standards. 5. A generic media format enables support of heterogeneous media formats to heterogeneous client endpoints by supporting any type of information media that can be presented or acted upon by a presentation device or machine. In one embodiment of a network delivering a media format, the network first queries the client endpoint to determine the client's capabilities, and then, if the client cannot meaningfully consume the media representation, the network either removes layers of attributes not supported by the client or adapts the media from its current format to a format appropriate for the client endpoint. In one example of such adaptation, the network would convert a volumetric visual media asset into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network could use neural network processes to reformat the media into an appropriate format or, optionally, synthesize a view required by the client endpoint.
[0063] 6. The manifest of a full or partially full immersive experience (live streaming event, game, or on-demand asset playback) is organized by scenes, which are the minimum information that rendering and game engines can currently capture to create the presentation. The manifest includes a list of individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene are one or more representations of the geometric objects in the scene, which correspond to a streamable version of the scene shape. One embodiment of a scene representation refers to a low-resolution version of the scene's geometric objects. Another embodiment of the same scene refers to an enhancement layer for the low-resolution representation of the scene to add additional detail or increase tessellation to the same scene's geometric objects. As mentioned above, each scene may have multiple enhancement layers to increase the detail of the scene's geometric objects in a progressive manner.
[0064] 7. Each layer of a media object referenced in a scene is associated with a token (e.g., a URI) that points to the address where the resource can be accessed in the network. Such a resource is similar to a CDN from which content may be fetched by a client. 8. Tokens in representations of geometric objects may point to locations within the network or within the client, i.e., the client may signal to the network that its resources are available to the network for network-based media processing.
[0065] 3 illustrates an embodiment of a generic media format for timed media: A timed scene manifest contains a list of scene information 301. The scene 301 points to processing information and a list of components 302 that individually describe the types of media assets that make up the scene 301. The components 302 point to assets 303, which further point to base layers 304 and attribute enrichment layers 305.
[0066] 4 illustrates an embodiment of a generic media format for non-timed media as follows: Scene information 401 is associated with a start and end time according to a clock. Scene information 401 points to processing information and a list of components 402 that individually describe the types of media assets that make up the scene 401. Components 402 point to assets 403 (e.g., visual, audio, and haptic assets) which further point to a base layer 404 and an attribute enhancement layer 405. Furthermore, scene 401 points to other scenes 401 for non-timed media. Scene 401 also points to a timed media scene.
[0067] FIG. 5 illustrates an embodiment of a process 500 for synthesizing a capture format from natural content. Camera unit 501 captures a human scene using a single camera lens. Camera unit 502 captures a scene with five divergent fields of view by mounting five camera lenses around a ring-shaped object. The configuration in 502 is an exemplary configuration commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven convergent fields of view by mounting seven camera lenses on the inner diameter of a sphere. Configuration 503 is an exemplary configuration commonly used to capture light fields for light field or holographic immersive displays. Natural image content 509 is provided as input to a synthesis module 504, which may optionally use a neural network training module 505 using a set of training images 506 to generate an arbitrary capture neural network model 508. Another process commonly used in place of the training process 505 is photogrammetry. 5, the model 508 becomes one of the assets of the ingestion format 507 for natural content. Example embodiments of the ingestion format 507 include MPI and MSI.
[0068] Figure 6 shows an embodiment of a process 600 for creating synthetic media, such as a computer-generated image capture format. A LIDAR camera 601 captures a point cloud 602 of a scene. A CGI tool, 3D modeling tool, or another animation process for creating synthetic content is used on a computer 603 to create CGI assets 604 over a network. A motion capture suit 605A equipped with sensors is worn by an actor 605 to capture a digital recording of the actor's 605 motion to generate animated motion capture data 606. Data 602, 604, and 606 are provided as inputs to a synthesis module 607, which may also optionally use a neural network and training data to create a neural network model (not shown in Figure 6).
[0069] The techniques for presenting and streaming heterogeneous immersive media described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 7 illustrates a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0070] Computer software can be coded using any suitable machine code or computer language, or similar mechanism, that can be assembled, compiled, linked by a computer central processing unit (CPU), graphics processing unit (GPU), etc. to create code comprising instructions that can be executed directly or via interpretation, microcode execution, etc.
[0071] The instructions may be executed by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0072] 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system 700.
[0073] The computer system 700 may also include certain human interface input devices that can respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly associated with conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0074] The input human interface devices may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touchscreen 710, a data glove (not shown), a joystick 705, a microphone 706, a scanner 707, and a camera 708 (only one of each is shown).
[0075] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 710, data gloves (not shown), or joystick 705, but may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers 709, headphones (not shown)), visual output devices (e.g., screens 710, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input and haptic feedback capabilities, some of which may output two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0076] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 or similar media 721 with CDs / DVDs, thumb drives 722, and removable hard drives or solid state drives 723, legacy magnetic media such as tape or floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0077] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transitory signals.
[0078] The computer system 700 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, or optical networks. The networks may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, or the like. Examples of networks include local area networks such as Ethernet and wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CANBus. Particular networks typically require an external network interface adapter connected to a particular general-purpose data port or peripheral bus 749 (e.g., a USB port on the computer system 700). Other networks are typically integrated into the core of the computer system 700 by connecting to the system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 700 can communicate with other entities. Such communications may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., from a CANbus to a particular CANbus device), or bidirectional, for example, to other computer systems using local or wide area digital networks. As noted above, specific protocols and protocol stacks may be used for each of these networks and network interfaces.
[0079] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to core 740 of computer system 700 .
[0080] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 743, and hardware accelerators 744 for specific tasks. These devices may be connected via a system bus 748, along with read-only memory (ROM) 745, random access memory 746, and internal mass storage 747, such as an internal hard drive or SSD, that is not user accessible. In some computer systems, the system bus 748 is accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus 748 or via a peripheral bus 749. Peripheral bus architectures include PCI, USB, etc.
[0081] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may combine to execute specific instructions that may constitute the aforementioned computer code. That computer code may be stored in ROM 745 or RAM 746. Transient data may also be stored in RAM 746, while permanent data may be stored, for example, in internal mass storage 747. Cache memory, which may be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc., may be used to enable fast storage and retrieval from any memory device.
[0082] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0083] By way of example only and not limitation, architecture 700, and specifically a computer system having core 740, may provide functionality as a result of processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with user-accessible mass storage as introduced above, as well as specific storage of core 740 that is non-transitory in nature, such as core internal mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by core 740. Computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause core 740, and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 744) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may, where appropriate, include logic, and vice versa. References to computer-readable media may, where appropriate, include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0084] FIG. 8 illustrates an exemplary network media distribution system 800 supporting a variety of legacy and heterogeneous immersive media-enabled displays as client endpoints. A content acquisition module 801 captures or creates media using the exemplary embodiments of FIG. 6 or FIG. 5. Ingest formats are created in a content preparation module 802 and then transmitted to one or more client endpoints 804 in the network media distribution system using a transmission module 803. A gateway may serve customer premises equipment (CPE) to provide network access to various client endpoints in the network. A set-top box may also serve as CPE, providing access to aggregated content by a network service provider. A wireless demodulator may act as a mobile network access point for mobile devices (e.g., similar to mobile handsets and displays). In one or more embodiments, a legacy 2D television may be directly connected to a gateway, a set-top box, or a WiFi router. A laptop computer with a legacy 2D display may be a client endpoint connected to a WiFi router. A head-mounted 2D (raster-based) display may also be connected to a router. A lenticular light field display may be connected to a gateway. The display may be comprised of a local computational GPU, a storage device, and a visual presentation unit that creates multiple views using ray-based lenticular optics technology. The holographic display may be connected to a set-top box and may include a local computational CPU, a GPU, a storage device, and a Fresnal pattern wave-based holographic visualization unit. The augmented reality headset may be connected to a wireless demodulator and may include a GPU, a storage device, a battery, and a volumetric visual presentation component. The high-density light field display may be connected to a WiFi router and may include multiple GPUs, a CPU, a storage device, an eye tracker, a camera, and a high-density ray-based light field panel.
[0085] FIG. 9 illustrates an embodiment of an immersive media delivery module 900 capable of servicing legacy and heterogeneous immersive media-enabled displays, as previously illustrated in FIG. 8. Content is created or acquired in module 901, which is further embodied for natural content and CGI content in FIGS. 5 and 6, respectively. The content is then converted to an ingestion format using network ingestion format creation module 902, which is similarly further embodied for natural content and CGI content in FIGS. 5 and 6, respectively. The ingestion media format is transmitted to a network and stored in storage device 903. Optionally, the storage device may reside on the immersive media content producer's network and be accessed remotely by an immersive media network delivery module (not numbered), as indicated by the dashed line bisecting 903. Client and application specific information is optionally available in remote storage device 904, which may optionally reside remotely in an alternative "cloud" network.
[0086] As shown in Figure 9, client interface module 905 may function as the primary source and sink of information to perform the primary tasks of the distribution network. In this particular embodiment, module 905 may be implemented in an integrated format with other components of the network. Nevertheless, the tasks represented by module 905 in Figure 9 form essential elements of the disclosed subject matter.
[0087] Module 905 receives information about the characteristics and attributes of client 908 and also gathers requirements for applications currently running on 908. This information may be obtained from device 904, or in alternative embodiments, by directly querying client 908. When directly querying client 908, it is assumed that a two-way protocol (not shown in FIG. 9) is present and operational, so that the client may communicate directly with interface module 905.
[0088] The interface module 905 also initiates and communicates with the media adaptation and fragmentation module 910 described in FIG. 10 . Once the ingested media has been adapted and fragmented by module 910, the media is optionally transferred to an intermediate storage device, shown as prepared media for delivery storage device 909. Once the delivery media is prepared and stored on device 909, the interface module 905 ensures that the immersive client 908 receives the delivery media and corresponding description information 906 via its network interface 908B via a “pull” request, or the client 908 itself can initiate a “pull” request for the media 906 from the storage device 909. The immersive client 908 may optionally use a GPU (or CPU, not shown) 908C. The delivery format of the media is stored in the client's 908 storage device or storage cache 908D. Finally, the client 908 visually presents the media via its visualization component 908A.
[0089] Throughout the process of streaming immersive media to a client 908 , the interface module 905 monitors the status of the client's progress via a client progress and status feedback channel 907 .
[0090] FIG. 10 illustrates a specific embodiment of a media adaptation process in which ingested source media may be appropriately adapted to match the requirements of the client 908. The media adaptation module 1001 is comprised of several components that facilitate adapting the ingested media to the appropriate delivery format of the client 908. These components should be considered exemplary. In FIG. 10, the adaptation module 1001 receives input network conditions 1005 and tracks client 908 information, including current traffic load on the network, attribute and feature descriptions, application characteristics, descriptions, and current state, and a client neural network model (if available), to help map the shape of the client's frustum to the interpolation capabilities of the ingestable immersive media. The adaptation module 1001 ensures that the adapted output, as it is created, is stored in the client adaptation media storage device 1006.
[0091] The adaptation module 1001 uses a renderer 1001B or a neural network processor 1001C to adapt the particular captured source media to a format suitable for the client. The neural network processor 1001C uses a neural network model 1001A. Examples of such neural network processors 1001C include deep view neural network model generators such as those described in MPI and MSI. If the media is in a 2D format but the client requires it in a 3D format, the neural network processor 1001C can invoke a process that uses highly correlated images from the 2D video signal to derive a volumetric representation of the scene depicted in the video. An example of such a process might be the neural radiance field from one or several images developed at the University of California, Berkeley. An example of a suitable renderer 1001B might be a modified version of the OTOY Octane renderer (not shown), modified to interact directly with the adaptation module 1001. The adaptation module 1001 may optionally use a media compressor 1001D and a media decompressor 1001E depending on the needs of these tools regarding the format of the captured media and the format required by the client 908.
[0092] Figure 11 shows an adaptation media packaging module 1103 that ultimately converts the adapted media from media adaptation module 1101 from Figure 10 that is currently residing on client adaptation media storage device 1102. Packaging module 1103 formats the adapted media from module 1101 into a robust distribution format, such as the exemplary formats shown in Figure 3 or Figure 4. Manifest information 1104A provides the client 908 with a list of scene data it can expect to receive, as well as a list of visual assets and corresponding metadata, and audio assets and corresponding metadata.
[0093] FIG. 12 shows a packetizer module 1202 that “fragments” adapted media 1201 into individual packets 1203 suitable for streaming to a client 908 . The components and communications shown in FIG. 13 of sequence diagram 1300 are described as follows: Client endpoint 1301 initiates media request 1308 to network delivery interface 1302. Request 1308 includes information identifying the media requested by the client by URN or other standard nomenclature. Network delivery interface 1302 responds to request 1308 with profile request 1309, which requests that client 1301 provide information about its currently available resources (including computation, storage, battery charge, and other information characterizing the client's current operating state). Profile request 1309 also requests that the client provide one or more neural network models that can be used by the network for neural network inference and, if such models are available at the client, to extract or interpolate the correct media view to match the characteristics of the client's presentation system. Response 1311 from client 1301 to interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). Interface 1302 then provides client 1301 with session ID token 1311. Interface 1302 then requests ingest media server 1303 with ingest media request 1312, which includes the URN or canonical nomenclature name of the media identified in request 1308. Server 1303 responds to request 1312 with response 1313, which includes the ingest media token. Interface 1302 then provides client 1301 with the media token from response 1313 in call 1314. Interface 1302 then begins the adaptation process for the requested media at 1308 by providing adaptation interface 1304 with the ingest media token, client token, application token, and neural network model token.Interface 1304 requests access to the ingested media by providing the ingested media token to server 1303 in call 1316 to request access to the ingested media asset. Server 1303 responds to request 1316 with the ingested media access token in response 1317 to interface 1304. Interface 1304 then requests that media adaptation module 1305 adapt the ingested media located in the ingested media access token for the client, application, and neural network inference model corresponding to the session ID token created in 1313. Request 1318 from interface 1304 to module 1305 includes the necessary token and session ID. Module 1305 provides the adapted media access token and session ID to interface 1302 in update 1319. Interface 1302 provides the adapted media access token and session ID to packaging module 1306 in interface call 1320. Packaging module 1306 provides response 1321 to interface 1302 with the packaged media access token and session ID in response 1321. Module 1306 provides the packaged media access token for the packaged asset, URN, and session ID to packaged media server 1307 in response 1322. Client 1301 executes request 1323 to begin streaming the media asset corresponding to the packaged media access token received in message 1321. Client 1301 executes other requests and provides status updates to interface 1302 in message 1324.
[0094] Figure 14 shows the ingested media format and asset 1002 of Figure 10 optionally composed of two parts of immersive media and assets in a 3D format 1401 and a 2D format 1402. The 2D format 1402 may be a coded video stream containing a single view, for example ISO / IEC 14496 Part 10 Advanced Video Coding, or may be a coded video stream containing multiple views, for example the multi-view compression modification of ISO / IEC 14496 Part 10.
[0095] Figure 15 illustrates the transmission of neural network model information along with a coded video stream. In this figure, coded video stream 1501 includes a neural network model and corresponding parameters carried directly by one or more SEI messages 1501A. In contrast, in coded video stream 1502, one or more SEI messages carry identifiers for the neural network model and its corresponding parameters. In scenario 1502, the neural network model and parameters are stored outside the coded video stream, for example, in 1001A of Figure 10.
[0096] FIG. 16 illustrates the transmission of neural network model information in a captured immersive media asset 1601 (originally shown as item 1401 in FIG. 14) in a 3D format. Media 1601 refers to scenes 1-N, shown as 1602. Each scene 1602 refers to geometry 1603 and processing parameters 1604. Geometry 1603 may include a reference 1603A to a neural network model. Processing parameters 1604 may also include a reference 1604A to a neural network model. Both 1604A and 1603A may refer to a network model stored directly with the scene, or may refer to an identifier that points to a neural network model that exists outside of the captured media, such as a network model stored in 1001A in FIG. 10. Some embodiments relate to systems, methods and / or computer-readable media at any possible level of technical detail. The computer-readable media may include a computer-readable non-transitory storage medium having computer-readable program instructions thereon that cause a processor to perform operations.
[0097] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction-execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, should not be construed as being, per se, a transitory signal such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over an electrical wire.
[0098] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or can be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0099] The computer-readable program code / instructions for carrying out the operations may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions to perform aspects or operations by utilizing state information of the computer-readable program instructions to customize the electronic circuitry.
[0100] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to manufacture a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, generate means for performing the functions / acts specified in the flowcharts and / or block diagrams or blocks. These computer-readable program instructions may be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in the flowcharts and / or block diagrams or blocks. The computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device, and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions operating on the computer, other programmable apparatus, or other device perform the functions / operations specified in the flowcharts and / or block diagrams or blocks.
[0101] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions that perform the specified logical function(s). The methods, computer systems, and computer-readable media may include additional, fewer, different, or differently arranged blocks than those shown in the figures. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be executed concurrently or substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, or combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or implements a combination of dedicated hardware and computer instructions.
[0102] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specified control hardware or software code used to implement these systems and / or methods is not limiting of implementation. Thus, the operation and behavior of the systems and / or methods have been described herein without reference to specific software code. It will be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0103] As used herein, no element, act, or instruction should be construed as critical or essential unless expressly stated. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Furthermore, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." When only one item is intended, the term "a" or similar language is used. Also, as used herein, terms such as "have," "contain," or "having" are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless expressly stated otherwise.
[0104] The descriptions of various aspects and embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. While combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically recited in the claims and / or specifically disclosed in the specification. While each dependent claim listed below may depend directly on only one claim, the disclosure of possible embodiments includes each dependent claim in combination with every other claim in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein was selected to best explain the principles of the embodiments, their practical application to commercially available technology, or technical improvements thereon, or to enable others skilled in the art to understand the embodiments disclosed herein. [Explanation of symbols]
[0105] 101A Camera or microphone 101B Computer 102 2D image and associated audio sequences 103 Preparation Module 104 Converter Module 105 Adaptation Module 106 Content Delivery Network 107 pull requests 108 clients 202 MPD 203 Time Limit 204 Adaptation Set 205 Expression 206 Segment File 301 Scene Information 302 Components 303 Assets 304 Base Layer 305 Attribute enhancement layer 401 Scene Information 402 Components 403 Assets 404 Base Layer 405 Attribute enhancement layer 500 processes 501 Camera Unit 502 Camera Unit 503 Camera Unit 504 Synthesis Module 505 Training Process 505 Neural Network Training Module 506 training images 507 Import Format 508 Capture Neural Network Model 509 Natural Image Content 600 processes 601 LIDAR camera 602 point cloud data 603 Computer 604 CGI assets 605 Actors 605A Motion Capture Suit 606 Motion Capture Data 607 Synthesis Module 608 Synthetic Media Ingest Format 700 Computer Systems 700 Architecture 701 Keyboard 702 Mouse 703 Trackpad 705 Joystick 706 Microphone 707 Scanner 708 Camera 709 Speaker 710 Touchscreen 720 CD / DVD ROM / RW 721 Media 722 thumb drive 723 Solid State Drive 740 cores 741 Central Processing Unit (CPU) 742 Graphics Processing Unit (GPU) 743 Field Programmable Gate Array (FPGA) 744 Hardware Accelerator 745 Read-Only Memory (ROM) 746 Random Access Memory 747 Mass Storage 748 System Bus 749 Peripheral Bus 800 Network Media Distribution System 801 Content Acquisition Module 802 Content Preparation Module 803 Transmit Module 804 client endpoint 900 Immersive Media Delivery Module 901 Module 901 Content Acquisition / Creation Module 902 Network Import Format Creation Module 903 Capture Storage Device 904 Remote Storage Device 905 Module 905 Client Interface Module 906 Media and Description Information 907 Client Progress and Status Feedback Channel 908 Immersive Client 908A Visualization Components 908B Network Interface 908D Memory Cache 909 Distribution Storage Device 910 Media Adaptation and Fragmentation Module 1001 Adaptation Module 1001A Neural Network Model 1001B Renderer 1001C Neural Network Processor 1001D Media Compressor 1001E Media Decompressor 1002 Assets 1005 Input Network State 1006 Client-adaptive media storage device 1101 Media Adaptation Module 1102 Current Client Adaptable Media Storage Device 1103 Adaptive Media Packaging Module 1104A Manifest Information 1201 Adaptable Media 1202 Packetizer Module 1203 packets 1204 client endpoint 1300 Sequence Diagram 1301 client endpoint 1302 Network Distribution Interface 1303 Ingest Media Server 1304 Adaptation Interface 1305 Media Adaptation Module 1306 Packaging Module 1307 Packaged Media Server 1401 3D Immersive Media and Assets 1402 2D Immersive Media and Assets 1501 coded video stream 1501A SEI Message 1502 coded video stream 1502A SEI Message 1601 3D Immersive Media and Assets 1602 scenes 1603 Shape See 1603A 1604 Processing parameters See 1604A
Claims
1. 1. A processor-executable method for streaming immersive media, comprising: obtaining information indicative of a characteristic of a client endpoint; capturing content in a first two-dimensional format or a first three-dimensional format, the content including one or more SEI messages; the one or more SEI messages include a reference to a neural network model for transforming the content; the reference to the neural network model indicates whether the neural network model is associated with a scene in the content; and converting the content into a second two-dimensional format or a second three-dimensional format suited to the characteristics of the client endpoint based on the referenced neural network model; and streaming the converted content to a client endpoint.
2. the reference to the neural network model indicates whether the neural network model is stored with the scene in the content; or The method of claim 1 , wherein the reference to the neural network model indicates whether the neural network model exists outside the content.
3. The method of claim 1 , wherein the reference to the neural network model includes metadata identifying a location of the neural network model.
4. The method of claim 3 , wherein the reference to the neural network model includes a universal resource identifier that corresponds to the metadata that describes the content.
5. The method of claim 1 , wherein the neural network model is trained prior to ingesting the content based on prior distributions corresponding to objects in the content.
6. The method of claim 1 , wherein the content is transformed based on characteristics of the client endpoint.
7. 10. The method of claim 1, wherein the one or more client endpoints comprise one or more of a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a high-density light field display.
8. A computer system configured to cause one or more computer processors to carry out the method of any one of claims 1 to 7.
9. A computer program for causing a computer to carry out the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Converting video data according to a 3D input format.
JP2013501475A
Depth map generation technique for converting 2D video data to 3D video data
JP2013509104A
Mobile media server
JP2013513319A
Model-based crosstalk reduction in stereoscopic and multi-view displays.
JP2014529954A
Context-Based Priors for Object Detection in Images
JP2018526723A