Method and apparatus for streaming immersive media
Through scene-specific neural network models and hierarchical compression algorithms, the immersive media source is adapted into a format suitable for client endpoints, solving the compatibility problems of traditional and immersive media displays in network distribution systems, and realizing unified distribution and visual quality improvement of heterogeneous client endpoints.
Patent Information
- Application Number
- CN202180009467.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-20
- Filing Date
- 2021-09-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-09-01
AI Technical Summary
Existing network distribution systems are difficult to effectively support traditional displays and displays with immersive media capabilities as client endpoints for a single network, and the lack of single standard representation and device heterogeneity leads to complex distribution.
Using a scenario-specific neural network model, the neural network model and parameters are embedded through SEI messages, combined with compression algorithms and hierarchical processes, the immersive media source is adapted into a specific characteristic format suitable for client endpoint devices and streamed over the network.
It realizes unified distribution between traditional and immersive media displays, supports diversified needs of heterogeneous client endpoints, improves visual quality and optimizes network resource utilization.
Smart Images

Figure CN114981822B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 126,188, filed on December 16, 2020, and U.S. Patent Application No. 17 / 407,711, filed on August 20, 2021, both of which are hereby incorporated by reference in their entirety. Technical Field
[0003] The present disclosure generally relates to the field of data processing, and more particularly to video coding. Background Art
[0004] "Immersive media" generally refers to media that stimulates any or all of the human sensory systems (vision, hearing, somatosensation, smell, and possibly taste) to create or enhance the perception of a user physically present in the experience of the media, i.e., beyond the media known as "traditional media" that is distributed over existing commercial networks for timed two - dimensional (2D) video and corresponding audio. Immersive media and traditional media can both be characterized as timed or untimed.
[0005] Time - based media refers to media that is structured and presented according to time. Examples include movie features, news reports, and episodic content, all of which are organized according to time periods. Traditional video and audio are generally considered time - based media.
[0006] Non - time - based media is media that is not structured according to time; rather, it is structured by logical, spatial, and / or temporal relationships. Examples include video games, where the user can control the experience created by the gaming device. Another example of non - time - based media is a still - image photograph taken by a camera. Non - time - based media can contain time - based media, e.g., a continuous loop of audio or video segments in a video game scene contains time - based media. Conversely, time - based media can contain non - time - based media, e.g., a video has a fixed still image as a background.
[0007] Devices with immersive media capabilities can refer to devices equipped with the ability to access, interpret, and present immersive media. Such media and devices are not uniform in terms of the quantity and format of the media, nor in terms of the quantity and type of network resources required for large - scale distribution of such media, i.e., the quantity and type of network resources required to achieve a distribution equivalent to that of traditional video and audio media over a network. In contrast, traditional devices such as laptop monitors, televisions, and mobile handheld displays are homogeneous in their capabilities because all of these devices include rectangular displays and consume 2D rectangular video or still images as their primary media format. Summary of the Invention
[0008] A method, apparatus, and computer-readable medium for streaming immersive media are provided.
[0009] According to a first aspect of the present disclosure, a method for streaming immersive media, executable by a processor, includes: ingesting content in a two-dimensional format, where the two-dimensional format references at least one neural network; converting the ingested content into a three-dimensional format based on the at least one referenced neural network; and streaming the converted content to a client endpoint.
[0010] The at least one neural network may include a scene-specific neural network that corresponds to a scene included in the ingested content.
[0011] Converting the ingested content may include: using the scene-specific neural network to infer depth information related to the scene; and adapting the ingested content to a scene-specific volumetric format associated with the scene.
[0012] The at least one neural network may be trained based on priors corresponding to objects within the scene.
[0013] The at least one neural network may be referenced in a Supplemental Enhancement Information (SEI) message included in an encoded video bitstream corresponding to the ingested content.
[0014] The neural network model and at least one parameter corresponding to the at least one neural network may be directly embedded in the SEI message.
[0015] The location of the neural network model corresponding to the at least one neural network may be written in the SEI message.
[0016] The client endpoint may include one or more of a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a dense light field display.
[0017] According to a second aspect of the present disclosure, an apparatus for streaming immersive media includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code, the program code including: ingestion code configured to cause the at least one processor to ingest content in a two-dimensional format, where the two-dimensional format references at least one neural network; conversion code configured to cause the at least one processor to convert the ingested content into a three-dimensional format based on the at least one referenced neural network; and streaming code configured to cause the at least one processor to stream the converted content to a client endpoint.
[0018] At least one neural network may include a scene-specific neural network that corresponds to a scene included in the ingested content.
[0019] The conversion code may include: inference code configured to cause at least one processor to use the scene-specific neural network to infer depth information related to the scene; and adaptation code configured to cause at least one processor to adapt the ingested content into a scene-specific volumetric format associated with the scene.
[0020] At least one neural network may be trained based on priors corresponding to objects within the scene.
[0021] At least one neural network may be referenced in a Supplemental Enhancement Information (SEI) message included in an encoded video bitstream corresponding to the ingested content.
[0022] The neural network model and at least one parameter corresponding to at least one neural network may be directly embedded in the SEI message.
[0023] The location of the neural network model corresponding to at least one neural network may be written in the SEI message.
[0024] The client endpoint may include one or more of a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a dense light field display.
[0025] According to a third aspect of the present disclosure, a non-transitory computer-readable medium stores instructions that include one or more instructions that, when executed by at least one processor of a device for streaming immersive media, cause the at least one processor to perform the method of streaming immersive media described in the first aspect above.
[0026] The disclosed subject matter addresses the development of a network-based media distribution system that can support traditional displays and immersive media displays as client endpoints within the context of a single network. Specifically, a mechanism is proposed herein that adapts an input immersive media source into a format suitable for the specific characteristics of a client endpoint device (including applications currently executing on that client endpoint device). This mechanism for adapting the input immersive media source includes coordinating the characteristics of the input immersive media with the characteristics of the target endpoint client device (including applications executing on the client device), and then adapting the input immersive media into a format suitable for the target endpoint and its applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1It is a schematic diagram of the end-to-end process of timed traditional media distribution.
[0028] Figure 2 It is a schematic diagram of the standard media format for streaming timed traditional media.
[0029] Figure 3 It is a schematic diagram of an embodiment of a data model for representing and streaming timed immersive media.
[0030] Figure 4 It is a schematic diagram of an embodiment of a data model for representing and streaming untimed immersive media.
[0031] Figure 5 It is a schematic diagram of the process of capturing a natural scene and converting the natural scene into a representation in an ingestion format that can be used as a network serving heterogeneous client endpoints.
[0032] Figure 6 It is a schematic diagram of the process of creating a synthetic scene in a 3D modeling tool and format in a representation in an ingestion format that can be used as a network serving heterogeneous client endpoints.
[0033] Figure 7 It is a system diagram of a computer system.
[0034] Figure 8 It is a schematic diagram of a network serving multiple heterogeneous client endpoints.
[0035] Figure 9 It is a schematic diagram of a network providing adaptation information related to a specific media represented in a media ingestion format, for example, before the network process of adapting the media for consumption by a specific immersive media client endpoint.
[0036] Figure 10 It is a system diagram of a media adaptation process consisting of a media render-converter that converts the source media from its ingestion format into a specific format suitable for a specific client endpoint.
[0037] Figure 11 It is a schematic diagram of a network formatting the adapted source media into a data model suitable for representation and streaming.
[0038] Figure 12 It is to Figure 12 A system diagram of a media streaming process that divides the data model into the payload of network protocol packets.
[0039] Figure 13 It is a sequence diagram of a network that adapts a specific immersive media in an ingestion format into a distributable format that can be streamed and is suitable for a specific immersive media client endpoint.
[0040] Figure 14 It consists of an immersive content format and a traditional content format (i.e., only 2D video format), or consists of an immersive and 2D video format Figure 10 Schematic diagram of the ingestion media format and asset 1002
[0041] Figure 15 Depicts the carrying of neural network model information along with the encoded video stream Detailed implementation
[0042] This document discloses detailed embodiments of the claimed structures and methods; however, it can be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, and the claimed structures and methods can be implemented in various forms. However, these structures and methods can be implemented in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. On the contrary, these exemplary embodiments are provided so that the present disclosure will be thorough and complete and will fully convey the scope to those skilled in the art. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments
[0043] Embodiments generally relate to the field of data processing, and more particularly to video coding. The techniques described herein allow a 2D encoded video stream to signal a scene-specific neural network model so that a network ingests a 2D video source of media that includes one or more (usually a relatively small number) of views and adapts the source of the 2D media into one or more streamable "distribution formats" to accommodate various heterogeneous client endpoint devices, the different characteristics and capabilities of the heterogeneous client endpoint devices, and the requirements of applications used on the client endpoints, and then actually distributes the formatted media to the various client endpoints. The network model can be directly embedded into the scene-specific encoded video stream of the encoded bitstream via an SEI structured field, or the SEI can signal the use of a specific model stored elsewhere on the distribution network but accessible by the neural network process. The ability to reformat the 2D media source into various streamable distribution formats enables the network to simultaneously serve various client endpoints with various capabilities and available computing resources and enables support for emerging immersive client endpoints such as holographic and light field displays in commercial networks. In addition, the ability to adapt a scene-specific 2D media source based on a scene-specific neural network model improves the final visual quality. This ability to adapt the 2D media source is particularly important when no immersive media source is available and when the client cannot support the distribution format based on the 2D media. In this scenario, the neural network-based method can be more optimally used for a specific scene present in the 2D media by carrying a scene-specific neural network model, where the scene-specific neural network model is trained using priors that are typically similar to objects within the specific scene or objects for the context of the specific scene. This improves the network's ability to infer depth-based information related to the specific scene, enabling the network to adapt the 2D media into a scene-specific volumetric format suitable for the target client endpoint.
[0044] As previously described, "immersive media" generally refers to media that stimulates any or all of the human sensory systems (visual, auditory, somatosensory, olfactory, and possibly gustatory) to create or enhance the perception of a user physically present in the experience of the media, i.e., beyond the media known as "traditional media" that is distributed over existing commercial networks for timed two-dimensional (2D) video and corresponding audio. Immersive media and traditional media can both be characterized as timed or untimed.
[0045] Time media refers to media that is structured and presented according to time. Examples include movie features, news reports, and episodic content, all of which are organized according to time periods. Traditional video and audio are generally considered time media.
[0046] Non-temporal media are media that are not structured by time; rather, they are structured by logical, spatial, and / or temporal relationships. Examples include video games, where the user can control the experience created by the gaming device. Another example of non-temporal media is a still image photograph taken by a camera. Non-temporal media can contain temporal media, for example, in a continuous loop of audio or video segments in a video game scenario. Conversely, temporal media can contain non-temporal media, for example, a video has a fixed still image as a background.
[0047] A device with immersive media capabilities can refer to a device equipped with the ability to access, interpret, and present immersive media. Such media and devices are not uniform in terms of the quantity and format of the media, and the quantity and type of network resources required to distribute such media on a large scale, that is, the quantity and type of network resources required to achieve a distribution equivalent to that of traditional video and audio media on the network are not uniform. In contrast, traditional devices such as laptop monitors, televisions, and mobile handheld displays are homogeneous in their capabilities because all of these devices include a rectangular display screen and consume 2D rectangular video or still images as their primary media format.
[0048] The distribution of any media on the network can adopt a media delivery system and architecture that reformats the media from an input or network "ingest" format into a final delivery format that is not only suitable for the target client device and its applications but also facilitates streaming over the network. The "streaming" of media generally refers to the segmentation and grouping of the source media so that the source media can be distributed over the network in logically organized and sequenced continuous smaller-sized "chunks" according to either the temporal or spatial structure of the media or both the temporal and spatial structures of the media. In such a delivery architecture and system, the media can undergo a compression or layering process so that only the most significant media information is initially distributed to the client. In some cases, the client must receive all of the significant media information for certain media portions before the client can present any media portion within the same media portion to the end user.
[0049] The process of reformatting input media to match the capabilities of a target client endpoint can employ a neural network process that uses a network model that may have some prior knowledge of the specific media to be reformatted. For example, a specific model can be adjusted to recognize an outdoor park scene (with trees, plants, grass, and other objects common to park scenes), while a different specific model can be adjusted to recognize an indoor dinner scene (with a dinner table, serving utensils, people sitting at the table, etc.). Those skilled in the art will recognize that a network model adjusted to recognize objects from a specific context (such as park scene objects) will recognize that a neural network process equipped with a network model adjusted to match specific scene content will produce better visual results compared to a network model not so adjusted. Thus, the benefit is to provide a scene-specific network model to a neural network process whose task is to reformat input media to match the capabilities of a target client endpoint.
[0050] The mechanism for associating a neural network model with a specific scene can be implemented by optionally compressing the network model and directly inserting the network model into the encoded bitstream of the visual scene through a Supplemental Enhancement Information (SEI) structured field, which is commonly used to attach metadata to the encoded video stream of H.264, H.265, and H.266 video compression formats. The presence of an SEI message containing a specific neural network model in the context of a portion of the encoded video bitstream can be used to indicate that the network model will be used to interpret and adapt the video content within that portion of the bitstream in which the model is embedded. Alternatively, the SEI message can be used to signal which neural network model can be used in the absence of the actual model itself through an identifier of the network model.
[0051] After adapting the video to the target client endpoint, the adaptation process within the network can then optionally choose to apply a compression algorithm to the result. Additionally, optionally, the compression algorithm can divide the adapted video signal into multiple layers that correspond to the most significant part to the least significant part of the visual signal.
[0052] An example of the compression and layering process is the progressive development format of the JPEG standard (ISO / IEC 10918 Part 1), which divides an image into multiple layers such that the entire image is first presented only in basic shapes and colors that are initially not in focus, i.e., the lower-order DCT coefficients from the entire image scan, and subsequently the entire image is presented with additional detail layers that can bring the image into focus, i.e., the higher-order DCT coefficients from the image scan.
[0053] The process of dividing media into smaller parts, organizing them into the payload part of consecutive network protocol packets, and distributing these protocol packets is called "streaming" of media, while the process of converting media into a format suitable for presentation on one of various heterogeneous client endpoints running one of various heterogeneous applications is called "adapting" the media.
[0054] Definition
[0055] Scene graph: A general data structure typically used by vector-based graphics editing applications and modern computer games, which arranges the logic of a graphical scene and its usual (but not necessarily) spatial representation; a collection of nodes and vertices in a graphical structure.
[0056] Node: The basic element of a scene graph, including information related to the logical or spatial or temporal representation of visual, audio, tactile, olfactory, gustatory, or related processing information; each node should have at most one output edge, zero or more input edges, and at least one edge (input or output) connected to the node.
[0057] Base layer: The nominal representation of an asset, typically formulated to minimize the computational resources or time required to render the asset, or to minimize the time required to transmit the asset over a network.
[0058] Enhancement layer: A set of information when applied to the base layer representation of an asset, enhancing the base layer to include features or capabilities not supported in the base layer.
[0059] Property: Metadata associated with a node, used to describe a specific characteristic or feature of the node in a typical or more complex form (e.g., from the perspective of another node).
[0060] Container: A serialization format used to store and exchange information to represent all natural scenes, all synthetic scenes, or a mixture of synthetic and natural scenes, including all media resources required for a scene graph and a rendered scene.
[0061] Serialization: The process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link) and later reconstructed (possibly in a different computer environment). When the resulting bit sequence is reread according to the serialization format, the bit sequence can be used to create a semantically identical clone of the original object.
[0062] Renderer: An application or process (usually software - based) of selective mixing based on disciplines, where the disciplines involve: acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development. That is, given an input scene graph and an asset container, the renderer emits a typical visual and / or audio signal that is suitable for presentation on a target device or conforms to the desired properties specified by the attributes of the render target node in the scene graph. For visual - based media assets, the renderer can emit a visual signal that is suitable for the target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in a graphics pipeline); for audio - based media assets, the renderer can emit an audio signal for presentation in multi - channel speakers and / or binaural headphones or for repackaging into another (output) container. Popular examples of renderers include: Unity, Unreal.
[0063] Evaluation: Produces a result (e.g., similar to the evaluation of a web page's document object model) such that the output moves from abstract to a concrete result.
[0064] Scripting language: An interpreted programming language that can be executed by the renderer at runtime to handle dynamic inputs and variable state changes made to the scene graph nodes, which affect the rendering and evaluation of spatial and temporal object topologies (including physical forces, constraints, IK, deformations, collisions) and energy propagation and transmission (light, sound).
[0065] Shader: A type of computer program that was originally used for shading (producing appropriate levels of light, darkness, and color within an image), but now, in various fields of computer graphics special effects, it performs various specialized functions, or performs video post - processing unrelated to shading, or even performs functions not related to graphics at all.
[0066] Path tracing: A computer graphics method for rendering three - dimensional scenes such that the illumination of the scene is more realistic.
[0067] Temporal media: Media that is ordered by time; for example, having a start time and an end time according to a specific clock.
[0068] Non - temporal media: Media that is organized by spatial, logical, or temporal relationships; for example, in an interactive experience implemented based on actions taken by a user.
[0069] Neural network model: A collection of parameters (i.e., numerical values) and tensors (e.g., matrices) that define weights used in well - defined mathematical operations, such mathematical operations being applied to visual signals to obtain an improved visual output, where the improved visual output may include interpolation of new views for the visual signal that are not explicitly provided by the original signal.
[0070] Immersive media can be considered as one or more types of media that, when presented to a human by a device with immersive media capabilities, stimulate any of the five senses (i.e., vision, sound, taste, touch, and hearing) in a more realistic and consistent manner with the human's understanding of experiences within the natural world, i.e., beyond what can be achieved using traditional media presented via traditional devices. In this context, the term "traditional media" refers to two-dimensional (2D) visual media (still or moving picture frames) and / or corresponding audio, for which the ability to interact with the user is limited to pausing, playing, fast-forwarding, or rewinding; and "traditional devices" refer to televisions, laptops, monitors, and mobile devices, whose capabilities are limited to presenting only traditional media. In consumer-oriented application scenarios, the presentation device for immersive media (i.e., the device with immersive media capabilities) is a consumer-oriented hardware device, particularly equipped with the ability to utilize the specific information embodied by immersive media such that the presentation created by the device more closely conforms to the human's understanding of and interaction with the physical world, i.e., beyond the ability of traditional devices to do so. The capabilities of traditional devices are limited to presenting only traditional media, while immersive media devices are not similarly constrained.
[0071] Over the past decade, many devices with immersive media capabilities have been introduced into the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, haptic gloves, and game consoles. Similarly, holographic displays and other forms of volumetric displays are expected to emerge within the next decade. While these devices are immediately available or soon will be, for several reasons, the realization of a coherent end-to-end ecosystem for distributing immersive media over a commercial network has failed.
[0072] One of these reasons is the lack of a single standard representation for immersive media that can address two main use cases related to the current distribution of large-scale media over a commercial network: 1) the real-time distribution of live-action events, i.e., content is created and distributed to client endpoints in real-time or near real-time, and 2) non-real-time distribution, which does not require the distribution of content in real-time, i.e., at the time the content is physically captured or created. Correspondingly, these two use cases can be equivalently compared to the "broadcast" and "on-demand" formats of distribution that exist today.
[0073] For real-time distribution, content can be captured by one or more cameras or created using computer generation techniques. Content captured by cameras is referred to herein as "natural" content, while content created using computer generation techniques is referred to herein as "synthetic" content. The media formats representing synthetic content can be those used in the 3D modeling, visual effects, and CAD / CAM industries and can include object formats and tools such as meshes, textures, point clouds, structured volumes, amorphous volumes (e.g., for fire, smoke, and fog), shaders, procedurally generated geometry, materials, lighting, virtual camera definitions, and animations. Although synthetic content is computer-generated, the synthetic media formats can be used for both natural and synthetic content. However, the process of converting natural content into a synthetic media format (e.g., a synthetic representation) can be time- and computationally intensive and thus may not be practical for real-time applications and use cases.
[0074] For the real-time distribution of natural content, the content captured by cameras can be distributed in a raster format, which is suitable for traditional display devices because many such devices are similarly designed to display raster formats. That is, assuming that traditional displays are uniformly designed to display raster formats, the distribution in raster format is preferably suitable for displays that can only display raster formats.
[0075] However, displays with immersive media capabilities are not necessarily limited to displaying raster-based formats. In addition, some displays with immersive media capabilities cannot present media that is only available in a raster-based format. The availability of displays optimized to create immersive experiences based on formats other than raster-based formats is another important reason why there is no coherent end-to-end ecosystem for distributing immersive media.
[0076] Another problem with creating a coherent distribution system for multiple different immersive media devices is that current and emerging immersive media-capable devices themselves vary widely. For example, some immersive media devices are explicitly designed to be used by only one user at a time, such as head-mounted displays. Other immersive media devices are designed such that they can be used by more than one user simultaneously, such as the "Lumus 8K Display" (hereinafter referred to as the "lenticular light field display") that can display content that can be viewed simultaneously by up to 12 users, where each user is experiencing his or her own unique perspective (i.e., view) of the content being displayed.
[0077] Further complicating the development of a coherent distribution system is that the number of unique perspectives or views that each display can produce can vary widely. In most cases, traditional displays can only create a single perspective of the content. However, a lenticular light field display can support multiple users, where each user experiences a unique perspective of the same visual scene. To enable the creation of multiple views of the same scene, the lenticular light field display creates a specific volumetric viewing volume, where 45 unique perspectives or views of the same scene are required as input to the display. This means that 45 slightly different unique raster representations of the same scene need to be acquired and distributed to the display in a format specific to this one particular display (i.e., its viewing volume). In contrast, the viewing volume of a traditional display is limited to a single two-dimensional plane, so there is no way to present more than one viewing perspective of the content through the viewing volume of the display (regardless of the number of simultaneous viewers experiencing the display).
[0078] Generally, immersive media displays can vary significantly according to the following characteristics of all displays: the size and volume of the viewing volume, the number of viewers supported simultaneously, the optical technology used to fill the viewing volume (which can be point-based, ray-based, or wave-based technologies), the density of the light units (points, rays, or waves) occupying the viewing volume, the availability of computing power and type of computing (CPU or GPU), the power source and availability (battery or line), the amount of local storage or caching, and access to auxiliary resources such as cloud-based computing and storage. These characteristics contribute to the heterogeneity of immersive media displays, as opposed to the uniformity of traditional displays, making the development of a single distribution system that can support all displays (including traditional displays and immersive types of displays) complicated.
[0079] The disclosed subject matter addresses the problem of the development of a network-based media distribution system that can support traditional displays and immersive media displays as client endpoints in the context of a single network. Specifically, a mechanism is proposed herein that adapts an input immersive media source into a format suitable for the specific characteristics of the client endpoint device (including the applications currently executing on that client endpoint device). This mechanism for adapting the input immersive media source includes coordinating the characteristics of the input immersive media with the characteristics of the target endpoint client device (including the applications executing on the client device), and then adapting the input immersive media into a format suitable for the target endpoint and its applications.
[0080] In addition, the adaptation process can include interpolating additional views (such as new views) from the input media to create the additional views required by the client endpoint. Such interpolation can be performed by means of a neural network process.
[0081] Note that, without loss of generality, the remainder of the disclosed subject matter assumes that the process of adapting an input immersive media source to a particular endpoint client device is the same as or similar to the process of adapting the same input immersive media source to a particular application executing on the particular client endpoint device. That is, the problem of adapting an input media source to the characteristics of an endpoint device has the same complexity as the problem of adapting a particular input media source to the characteristics of a particular application.
[0082] Traditional devices supported by traditional media have achieved wide consumer adoption because traditional devices are also supported by an ecosystem of traditional media content providers (which produce standard-based representations of traditional media) and commercial network service providers (which provide network infrastructure to connect traditional devices to standard traditional content sources). In addition to their role in distributing traditional media over the network, commercial network service providers can also facilitate the pairing of access to traditional client devices with traditional content on a content delivery network (CDN). Once paired with access to the appropriate form of content, the traditional client device can request or "pull" traditional content from a content server to the device for presentation to the end user. However, an architecture that "pushes" the appropriate media to the appropriate client is equally suitable without adding additional complexity to the overall architecture and solution design.
[0083] Aspects are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0084] The exemplary embodiments described below relate to the architecture, structure, and components of systems and networks for distributing media (including video, audio, geometric (3D) objects, haptics, associated metadata, or other content of a client device). Particular embodiments relate to systems, structures, and architectures for distributing media content to heterogeneous immersive and interactive client devices.
[0085] Figure 1 is an example diagram of the end-to-end process for timed traditional media distribution. In Figure 1In it, the timed audio-visual content is collected by the camera or microphone 101A or generated by the computer 101B to create a 2D image sequence 102 and the associated audio. The associated audio of the 2D image sequence 102 is input into the preparation module 103. The output of 103 is the edited content (e.g., for post-production including language translation, subtitles, other editing functions), which is referred to as the main format ready to be converted into a standard sandwich format (e.g., for on-demand media), or is converted into a standard contribution format (e.g., for live events) by the converter module 104. The media is "ingested" by a commercial network service provider, and the adaptation module 105 packs the media into various bitrates, temporal resolutions (frame rates), or spatial resolutions (frame sizes), which are packed into a standard distribution format. The resulting adaptation is stored on the content delivery network 106, and various clients 108 make pull requests 107 from the content delivery network 106 to obtain the media and present the media to the end user. It is important to note that the main format can consist of a mixture of media from 101A or 101B, and the format 101A can be obtained in real time, e.g., media obtained from a live sports event. Additionally, the client 108 is responsible for selecting the specific adaptation 107 that is most suitable for the client's configuration and / or the current network conditions, but it is also possible that the network server ( Figure 1 not shown) can determine the appropriate content and then push the appropriate content to the client 108.
[0086] Figure 2 is an example of a standard media format for distributing traditional time media (e.g., video, audio, and supporting metadata including, for example, timed text for subtitles). As Figure 1 shown by item 106 in, the media is stored on the CDN 201 in a standard-based distribution format. The standard-based format is shown as MPD 202, which consists of multiple parts that contain timed periods 203 with start times and end times corresponding to a clock. Each period 203 refers to one or more adaptation sets 204. Each adaptation set 204 is typically used for a single type of media, e.g., video, audio, or timed text. For any given period 203, multiple video sets 204 can be provided, e.g., one video set for video and multiple video sets for audio (e.g., for conversion into various languages). Each video set 204 refers to one or more representations 205 that provide information related to the frame resolution (for video), frame rate, and bitrate of the media. Multiple representations 205 can be used to provide access, e.g., each representation 205 for ultra-high definition video, high definition video, or standard definition video. Each representation 205 refers to one or more segment files 206, and the media is actually stored in one or more segment files 206 for the client to extract (as Figure 1 shown by 108 in) or by the network media server (Figure 1 is distributed (in a “push-based” architecture) which is not shown.
[0087] Figure 3 is an example representation of a streamable format for timed heterogeneous immersive media. Figure 4 is an example representation of a streamable format for untimed heterogeneous immersive media. Both figures refer to a scene; Figure 3 refers to scene 301 for temporal media, while Figure 4 refers to scene 401 for non-temporal media. For both cases, the scene can be embodied by various scene representations or scene descriptions.
[0088] For example, in some immersive media designs, the scene can be embodied by a scene graph, or as a Multi-Plane Image (MPI) or a Multi-Spherical Image (MSI). MPI technology and MSI technology are examples of technologies that contribute to creating a display-independent scene representation of natural content (i.e., images of the real world captured simultaneously from one or more cameras). On the other hand, scene graph technology can be employed to represent natural images and computer-generated images in the form of a synthetic representation. However, in cases where the content is captured as a natural scene by one or more cameras, creating such a representation is particularly computationally intensive. That is, creating a scene graph representation of naturally captured content is time-consuming and computationally intensive, requiring complex analysis of natural images using photogrammetry techniques or deep learning techniques or both, in order to create a synthetic representation that can then be used to interpolate a sufficient and appropriate number of views to fill the frustum of a target immersive client display. Therefore, it is currently impractical to consider such a synthetic representation as a candidate for representing natural content, as such a synthetic representation cannot actually be created in real time to account for usage scenarios that require real-time distribution. However, currently, the best candidate representation for computer-generated images is to use a scene graph in conjunction with a synthetic model, because computer-generated images are created using 3D modeling processes and tools.
[0089] This dichotomy of the best representations for natural content and computer-generated content implies that the best ingestion format for naturally captured content is different from the best ingestion format for computer-generated content or from the best ingestion format for natural content that is not necessary for real-time distribution applications. Therefore, an object of the disclosed subject matter is to support multiple ingestion formats for visually immersive media robustly enough, regardless of whether they are created naturally or by a computer.
[0090] The following is an example technique for embodying a scene graph in a format suitable for representing visually immersive media created using computer-generated techniques or for representing naturally captured content, for which deep learning or photogrammetry techniques are used to create a corresponding synthetic representation of the natural scene, i.e., not necessary for real-time distribution applications.
[0091] 1. OTOY's
[0092] OTOY's ORBX is one of several scene graph techniques that can support any type of timed or untimed visual media, including ray-traced, traditional (frame-based), volumetric, and other types of synthetic or vector-based visual formats. ORBX is unique relative to other scene graphs in that it provides native support for the free use and / or open-source formats of meshes, point clouds, and textures. ORBX is a deliberately designed scene graph intended to facilitate the exchange of multiple vendor technologies operating on the scene graph. Additionally, ORBX provides a rich material system, supports the Open Shader Language, a robust camera system, and supports Lua scripting. ORBX is also the basis for an immersive technology media format released by the Immersive Digital Experience Alliance (IDEA) and licensed under royalty-free terms. In the context of the real-time distribution of media, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computing resources to perform complex analysis of camera-captured data and synthesize the same data into a synthetic representation. To date, the availability of computing sufficient for real-time distribution has been impractical but still impossible.
[0093] 2. Pixar's Universal Scene Description
[0094] Pixar's Universal Scene Description (USD) is another well-known and mature scene graph that is popular in the VFX and professional content production communities. USD is integrated into NVIDIA's Omniverse platform, which is a toolset for facilitating developers to create and render 3D models using NVIDIA's GPUs. A subset of USD is released by Apple and Pixar as USDZ. USDZ is supported by Apple's ARKit.
[0095] 3. Khronos's glTF 2.0
[0096] glTF 2.0 is the latest version of the "Graphics Language Transmission Format" specification written by the Khronos 3D group. This format supports a simple scene graph format, which generally can support static (non-temporal) objects in a scene, including "png" and "jpeg" image formats. glTF 2.0 supports simple animations of basic shapes (i.e., for geometric objects) described using glTF primitives, including support for translation, rotation, and scaling. glTF 2.0 does not support temporal media, and thus does not support video or audio.
[0097] For example, these known designs are represented only by way of example as scenarios that provide only immersive visual media, and do not limit the capabilities of the disclosed subject matter to a specified process for adapting an input immersive media source into a format suitable for the specific characteristics of a client endpoint device.
[0098] In addition, any or all of the above example media representations currently employ or may employ deep learning techniques to train and create a neural network model that can select a specific view or facilitate the selection of a specific view based on the specific dimensions of a view frustum to fill the view frustum of a specific display. The view selected for the view frustum of a specific display can be interpolated from existing views explicitly provided in the scene representation (e.g., according to MSI or MPI techniques), or can be rendered directly from these rendering engines based on a description of the virtual camera of a specific virtual camera position, filter, or rendering engine for the view selected for the view frustum of a specific display.
[0099] Therefore, the disclosed subject matter is robust enough to consider that a relatively small but well-known set of immersive media ingestion formats can adequately meet the needs of real-time or "on-demand" (e.g., non-real-time) distribution of media that is naturally captured (e.g., using one or more cameras) or created using computer-generated techniques.
[0100] With the deployment of advanced network technologies, such as 5G for mobile networks and fiber optic cables for fixed networks, it has further facilitated the interpolation of views from immersive media capture formats by using neural network models or network-based rendering engines. That is to say, these advanced network technologies increase the capacity and capabilities of commercial networks because such advanced network infrastructures can support the transmission and distribution of an increasing amount of visual information. Network infrastructure management technologies (such as Multi-access Edge Computing (MEC), Software Defined Networks (SDN), and Network Function Virtualization (NFV)) enable commercial network service providers to flexibly configure their network infrastructures to adapt to changes in the demand for certain network resources. For example, in response to dynamic increases or decreases in the demand for network throughput, network speed, round-trip latency time, and computing resources. In addition, this inherent ability to adapt to dynamic network demands also promotes the network's ability to adapt immersive media capture formats into suitable distribution formats to support various immersive media applications with possible heterogeneous visual media formats for heterogeneous client endpoints.
[0101] Immersive media applications themselves can also have different requirements for network resources. Immersive media applications include: game applications that require significantly lower network latency times to respond to real-time updates of game states; telepresence applications that have symmetric throughput requirements for both the uplink and downlink portions of the network; and passive viewing applications that can have increased demands for downlink resources depending on the type of client endpoint display that consumes data. Generally, any consumer-oriented application can be supported by a variety of client endpoints, which have a variety of on-board client capabilities for storage, computing, and power, as well as various similar requirements for specific media representations.
[0102] Therefore, the disclosed subject matter enables a fully equipped network, i.e., a network that incorporates some or all of the features of modern networks, to support multiple traditional devices and immersive media-capable devices simultaneously based on specified characteristics. Thus:
[0103] 1. Provide flexibility in using media capture formats, which are practical for both real-time and "on-demand" usage scenarios of distributing media.
[0104] 2. Provide flexibility to traditional client endpoints and immersive media-capable client endpoints in supporting natural content and computer-generated content.
[0105] 3. Support temporal media and non-temporal media.
[0106] 4. Provide a process for dynamically adapting a source media ingestion format into a suitable distribution format based on the characteristics and capabilities of client endpoints and application-based requirements.
[0107] 5. Ensure that the distribution format is streamable over an IP-based network.
[0108] 6. Enable the network to serve multiple heterogeneous client endpoints that may include traditional devices and devices with immersive media capabilities simultaneously.
[0109] 7. Provide an exemplary media representation framework that facilitates organizing distributed media along scene boundaries.
[0110] Examples of improved end-to-end embodiments that the disclosed subject matter can achieve are implemented according to the processes and components of the following detailed description Figures 3 to 14 as implemented.
[0111] Figure 3 and Figure 4 adopt a single exemplary containment distribution format that has been adapted from an ingestion source format to match the capabilities of a specific client endpoint. As described above, Figure 3 the media shown is temporal media, Figure 4 the media shown is non-temporal media. The specific containment format is robust enough in its structure to accommodate a large number of media attributes, each of which can be layered based on the amount of significant information, and each layer contributes to media presentation. It should be noted that this layering process is already a technique known in the prior art, as demonstrated by scalable video architectures such as those specified in the evolving JPEG and scalable video architectures (e.g., ISO / IEC 14496-10 (Scalable High Efficiency Video Coding)).
[0112] 1. Media streamed according to a containment media format is not limited to traditional visual and audio media, but can include any type of media information capable of generating a signal that interacts with a machine to stimulate human senses of sight, sound, taste, touch, and smell.
[0113] 2. Media streamed according to a containment media format can be temporal media, non-temporal media, or a mixture of temporal and non-temporal media.
[0114] 3. The layered representation of media objects can be achieved by using a base layer and an enhancement layer architecture, whereby media formats can also be streamed. In one example, separate base and enhancement layers are calculated by applying multi-resolution or multi-subdivision analysis techniques to the media objects in each scene. This is similar to the progressive rendering image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG 2000), but is not limited to raster-based visual formats. In an example embodiment, the progressive representation of a geometric object can be a multi-resolution representation of the object calculated using wavelet analysis.
[0115] In another example of the layered representation of media formats, the enhancement layer applies different attributes to the base layer, such as refining the material properties of the surface of a visual object represented by the base layer. In yet another example, the attributes can refine the texture of the surface of the base layer object, such as changing the surface from smooth to porous texture or from a matte surface to a glossy surface.
[0116] In yet another example of the layered representation, the surface of one or more visual objects in a scene can change from Lambertian to ray-traced.
[0117] In yet another example of the layered representation, the network will distribute the base layer representation to the client, so that the client can create a nominal rendering of the scene while the client waits for the transmission of additional enhancement layers to refine the resolution or other characteristics of the base representation.
[0118] 4. The resolution of the attributes or refinement information in the enhancement layer is not explicitly coupled to the resolution of the objects in the base layer in current existing MPEG video and JPEG image standards.
[0119] 5. The media format includes any type of information media that can be rendered or actuated by a rendering device or machine, enabling heterogeneous client endpoints to support heterogeneous media formats. In one embodiment of the network distributing the media format, the network will first query the client endpoint to determine the capabilities of the client. If the client cannot meaningfully ingest the media representation, the network will remove the attribute layers not supported by the client, or transcode the media from its current format into a format suitable for the client endpoint. In one example of such transcoding, the network can convert volumetric visual media assets into a 2D representation of the same visual assets by using a network-based media processing protocol.
[0120] 6. A manifest for a complete or partially complete immersive experience (live streaming event of an on-demand asset, game, or playback) is organized by scenes, and the manifest is the minimum amount of information that the rendering and game engines can currently ingest to create a presentation. The manifest includes a list of independent scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene is one or more representations of geometric objects corresponding to a streamable version of the scene geometry within the scene. In one embodiment, the scene representation refers to a low-resolution version of the geometric objects of the scene. In another embodiment, the same scene refers to an enhancement layer that is used for the low-resolution representation of the scene to add additional details or increase the tessellation to the geometric objects of the same scene. As described above, each scene can have more than one enhancement layer to incrementally add details to the geometric objects of the scene in a line-by-line manner.
[0121] 7. Each layer of a media object referenced within a scene is associated with a token (e.g., a URI) that points to an address of a location where the resource can be accessed within the network. Such a resource is similar to a location in a CDN where the content can be fetched by the client.
[0122] 8. The token used to represent a geometric object can point to a location within the network or a location within the client. That is, the client can signal to the network that its resources are available to the network for network-based media processing.
[0123] Figure 3 Embodiments for media formats that include media for time media are described as follows. The time scene manifest includes a list of scene information 301. Scene 301 refers to a list of components 302 that respectively describe processing information and the type of media asset that includes scene 301. Component 302 refers to asset 303, and asset 303 further refers to a base layer 304 and an attribute enhancement layer 305.
[0124] Figure 4 Embodiments for media formats that include media for non-time media are described as follows. Scene information 401 is not associated with a start duration and an end duration according to a clock. Scene information 401 refers to a list of components 402 that respectively describe processing information and the type of media asset that includes scene 401. Component 402 refers to asset 403 (e.g., visual asset, audio asset, and tactile asset), and asset 403 further refers to a base layer 404 and an attribute enhancement layer 405. Additionally, scene 401 refers to other scenes 401 for non-time media. Scene 401 also refers to time media scenes.
[0125] Figure 5An embodiment of a process 500 for synthesizing an ingestion format from natural content is shown. Camera unit 501 uses a single camera lens to capture a scene of a person. Camera unit 502 captures a scene with five diverging fields of view by mounting five camera lenses around an annular object. The arrangement of 502 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter portion of a sphere. The arrangement 503 is an exemplary arrangement commonly used to capture the light field for a light field or holographic immersive display. Natural image content 509 is provided as input to a synthesis module 504, and the synthesis module 504 may optionally employ a neural network training module 505 that trains a set of images 506 to produce an optional acquisition neural network model 508. Another process commonly used instead of the training process 505 is photogrammetry. If model 508 is created during the process 500 depicted in Figure 5 then model 508 becomes one of the assets in the ingestion format 507 for natural content. Exemplary embodiments of the ingestion format 507 include MPI and MSI.
[0126] Figure 6 An embodiment of a process 600 for creating an ingestion format for synthetic media (e.g., computer-generated images) is shown. A LIDAR camera 601 captures a point cloud 602 of a scene. CGI tools, 3D modeling tools, or another animation process are employed on a computer 603 to create synthetic content to create 604 CGI assets on a network. A motion capture kit 605A with sensors is worn on an actor 605 to capture a digital record of the motion of the actor 605 to produce animated motion capture data 606. Data 602, 604, and 606 are provided as inputs to a synthesis module 607, and the synthesis module 607 may also optionally use neural networks and training data to create a neural network model ( Figure 6 not shown).
[0127] Techniques for representing and streaming the heterogeneous immersive media described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 7 FIG. shows a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0128] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code that includes instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0129] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0130] Figure 7 The components of the illustrated computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system 700.
[0131] The computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as, for example: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to collect certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0132] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard 701, mouse 702, touchpad 703, touch screen 710, data glove (not depicted), joystick 705, microphone 706, scanner 707, and camera 708.
[0133] The computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen 710, data glove (not depicted), or joystick 705, but may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speaker 709, headphones (not depicted)), visual output devices (e.g., screen 710 including CRT screen, LCD screen, plasma screen, OLED screen, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality - some of which screens are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).
[0134] The computer system 700 may also include human-accessible storage devices and their associated media, such as, for example, optical media including CD / DVD ROM / RW 720 with media 721 such as CD / DVD, thumb drives 722, removable hard disk drives or solid state drives 723, conventional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0135] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0136] The computer system 700 may also include an interface to one or more communication networks. The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system 700) connected to certain common data ports or peripheral buses (749); as described below, other network interfaces are typically integrated into the core of the computer system 700 by connecting to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system 700 may communicate with other entities using any of these networks. Such communication may be only unidirectional reception (e.g., broadcast television), only unidirectional transmission (e.g., CANbus connected to certain CANbus devices), or bidirectional, e.g., using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.
[0137] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core 740 of the computer system 700.
[0138] The kernel 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) 743, a hardware accelerator 744 for certain tasks, etc. These devices, as well as a read-only memory (ROM) 745, a random access memory (RAM) 746, an internal mass storage 747 such as an internal hard disk drive, SSD, etc. that is not user-accessible, may be connected via a system bus 748. In some computer systems, the system bus 748 may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus 748 of the kernel or connected to the system bus 748 of the kernel via a peripheral bus 749. The architecture of the peripheral bus includes PCI, USB, etc.
[0139] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute certain instructions, which may be combined to form the aforementioned computer code. The computer code may be stored in the ROM 745 or RAM 746. Transitional data may also be stored in the RAM 746, while permanent data may be stored, for example, in the internal mass storage 747. Fast storage and retrieval to any storage device may be performed by using a cache, which may be closely associated with one or more CPUs 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc.
[0140] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the art of computer software.
[0141] As a non - limiting example, a computer system having architecture 700, particularly kernel 740, can provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer - readable media. Such computer - readable media can be media associated with user - accessible mass storage as described above, as well as certain non - transient memories of kernel 740, such as on - kernel mass storage 747 or ROM 745. Software implementing embodiments of the present disclosure can be stored in such devices and executed by kernel 740. Depending on specific needs, the computer - readable media can include one or more storage devices or chips. The software can cause kernel 740, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hard - wired or otherwise embodied in circuitry (e.g., accelerator 744) that can replace software or operate in conjunction with software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, portions referring to software can include logic and vice versa. In appropriate cases, portions referring to computer - readable media can include circuitry (e.g., integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or including both. The present disclosure encompasses any suitable combination of hardware and software.
[0142] Figure 8 An exemplary network media distribution system 800 is shown. Network media distribution system 800 supports various traditional displays and heterogeneous immersive - media - capable displays as client endpoints. Content acquisition module 801 uses Figure 6 or Figure 5Exemplary embodiments are used to capture or create media. An ingestion format is created in the content preparation module 802 and then sent using the transport module 803 to one or more client endpoints 804 in the network media distribution system. The gateway can serve client devices to provide network access to various client endpoints to the network. The set-top box can also be used as a client device to provide access to aggregated content through a network service provider. The wireless demodulator can be used as a mobile network access point for mobile devices (e.g., like a mobile handset and display). In one or more embodiments, a traditional 2D TV can be directly connected to the gateway, set-top box, or WiFi router. A laptop with a traditional 2D display can be a client endpoint connected to the WiFi router. A head-mounted 2D (raster-based) display can also be connected to the router. A lenticular light field display can be connected to the gateway. The display can include a local computing GPU, a storage device, and a visual presentation unit that uses ray-based lenticular optics to create multiple views. A holographic display can be connected to the set-top box and can include a local computing CPU, GPU, storage device, and a holographic visualization unit based on Fresnel mode waves. An augmented reality head-mounted headset can be connected to the wireless demodulator and can include a GPU, storage device, battery, and volumetric visual presentation components. A dense light field display can be connected to the WiFi router and can include multiple GPUs, CPUs, and storage devices; an eye-tracking device; a camera; and a dense-ray-based light field panel.
[0143] Figure 9 An embodiment of the immersive media distribution module 900 is shown, which is capable of serving traditional displays and heterogeneous immersive media-capable displays as previously depicted in Figure 8 Content is created or acquired in module 901, which is further embodied in Figure 5 and Figure 6 for natural content and CGI content, respectively. Then, the content 901 is converted into an ingestion format using the create network ingestion format module 902. Module 902 is also further embodied in Figure 5 and Figure 6 for natural content and CGI content, respectively. The ingested media format is sent to the network and stored on the storage device 903. Optionally, the storage device can reside in the network of the immersive media content producer and be remotely accessed by an immersive media network distribution module (not numbered), as depicted by the dashed line bisecting 903. Client information and application-specific information are optionally available on the remote storage device 904, which can optionally be remotely present in an alternative "cloud" network.
[0144] As Figure 9As depicted, the client interface module 905 serves as the primary source and repository of information for performing the main tasks of the distribution network. In this particular embodiment, module 905 may be implemented in a unified format with other components of the network. However, the tasks depicted by module 905 in Figure 9 form an essential element of the disclosed subject matter.
[0145] Module 905 receives information related to the characteristics and attributes of client 908 and, in addition, collects requirements regarding the applications currently running on 908. This information may be obtained from device 904 or, in an alternative embodiment, may be obtained by directly querying client 908. In the case of directly querying client 908, it is assumed that a two-way protocol exists and is operable ( Figure 9 not shown), such that the client can communicate directly to interface module 905.
[0146] Interface module 905 also initiates communication with and Figure 10 communicates with the media adaptation and segmentation module 910 described in
[0147] When the ingested media is adapted and segmented by module 910, the media is optionally transferred to an inter-media storage device depicted as the media storage device 909 ready for distribution. When the distribution media is prepared and stored in device 909, interface module 905 ensures that the immersive client 908 receives the distribution media and corresponding description information 906 via its network interface 908B through a "push" request, or the client 908 itself may initiate a "pull" request to obtain media 906 from storage device 909. The immersive client 908 may optionally employ a CPU (or CPUs, not shown) 908C. The distribution format of the media is stored in the storage device or cache 908D of client 908. Finally, client 908 visually presents the media through its visualization component 908A.
[0148] Figure 10 A particular embodiment of the media adaptation process is depicted such that the ingested source media can be appropriately adapted to match the requirements of client 908. Media adaptation module 1001 includes a plurality of components facilitating the adaptation of the ingested media into a suitable distribution format for client 908. These components should be considered exemplary. In Figure 10Among them, the adaptation module 1001 receives the input network state 1005 to trace the current traffic load on the network; client 908 information (including attributes and feature descriptions, application features and descriptions, and the current state of the application); and the client neural network model (if available) to assist in mapping the geometry of the client's view volume to the interpolation capabilities of the captured immersive media. The adaptation module 1001 ensures that when creating the adapted output, the adapted output is stored in the client-adapted media storage device 1006.
[0149] The adaptation module 1001 employs a renderer 1001B or a neural network processor 1001C to adapt a specific captured source media into a format suitable for the client. The neural network processor 1001C uses the neural network model 1001A. Examples of such neural network processors 1001C include the deep view neural network model generators described in, for example, MPI and MSI. If the media is in 2D format but the client must have 3D format, the neural network processor 1001C can invoke a process that uses highly correlated images from the 2D video signal to derive a volumetric representation of the scene depicted in the video. An example of such a process can be the neural radiance field from one or more image processes developed by the University of California, Berkeley. An example of a suitable renderer 1001B can be a modified version of the OTOY Octane renderer (not shown), which can be modified to interact directly with the adaptation module 1001. Depending on the need for these tools with respect to the format of the captured media and the format required by the client 908, the adaptation module 1001 can optionally employ a media compressor 1001D and a media decompressor 1001E.
[0150] Figure 11 Depicts an adapted media encapsulation module 1103, which finally converts the adapted media from the media adaptation module 1101 that now resides on the client-adapted media storage device 1102. Figure 10 The encapsulation module 1103 formats the adapted media from module 1101 into a robust distribution format, for example, Figure 3 or Figure 4 the exemplary formats shown. The manifest information 1104A provides the client 908 with a list of the scene data it can expect to receive, and also provides a list of visualization assets and corresponding metadata, as well as audio assets and corresponding metadata.
[0151] Figure 12 Depicts a grouper module 1202, which divides the adapted media 1201 into independent packets 1203 suitable for streaming to the client 908.
[0152] Figure 13The components and communications shown in sequence diagram 1300 are explained as follows: The client endpoint 1301 initiates a media request 1308 to the network distribution interface 1302. The request 1308 includes information for identifying the media requested by the client via a URN or other standard naming. The network distribution interface 1302 responds to the request 1308 with a profile request 1309, which requests the client 1301 to provide information related to its currently available resources (including computing, storage, battery charge percentage, and other information characterizing the current operating state of the client). The profile request 1309 also requests the client to provide one or more neural network models that can be used by the network for neural network inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system (if such models are available at the client). The response 1311 from the client 1301 to the interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). Then, the interface 1302 provides the session ID token 1311 to the client 1301. Then, the interface 1302 uses an ingest media request 1312 to request the ingest media server 1303. The ingest media request 1312 includes the URN or standard naming name of the media identified in the request 1308. The server 1303 replies to the request 1312 with a response 1313 that includes an ingest media token. Then, the interface 1302 provides the media token from the response 1313 to the client 1301 in an invocation 1314. Then, the interface 1302 initiates an adaptation process for the media requested in 1308 by providing the ingest media token, the client token, the application token, and the neural network model token to the adaptation interface 1304. The interface 1304 requests access to the ingest media by providing the ingest media token to the server 1303 at an invocation 1316 to request access to the ingest media asset. The server 1303 responds to the request 1316 with an ingest media access token in the response 1317 that arrives at the interface 1304. Then, the interface 1304 requests the media adaptation module 1305 to adapt the ingest media located at the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created at 1313. The request 1318 from the interface 1304 to the response module 1305 contains the required tokens and the session ID. The module 1305 provides the adapted media access token and the updated session ID 1319 to the interface 1302. The interface 1302 provides the adapted media access token and the session ID to the encapsulation module 1306 in an interface invocation 1320. The encapsulation module 1306 provides a response 1321 to the interface 1302, in which there is an encapsulated media access token and the session ID.Module 1306 provides an encapsulated asset, a URN, and an encapsulated media access token for the session ID to the encapsulated media server 1307 in response 1322. The client 1301 executes request 1323 to initiate the streaming of the media asset corresponding to the encapsulated media access token received in message 1321. The client 1301 executes other requests and provides a status update to interface 1302 in message 1324.
[0153] Figure 14 depicts Figure 10 the ingest media format and asset 1002, which optionally consists of two parts: immersive media and assets in 3D format 1401 and immersive media and assets in 2D format 1402. The 2D format 1402 can be a single-view encoded video stream, such as ISO / IEC 14496 Part 10 Advanced Video Coding, or can be an encoded video stream containing multiple views (e.g., a multi-view compression modification of ISO / IEC 14496 Part 10).
[0154] Figure 15 depicts the carrying of neural network model information together with the encoded video stream. In this figure, the encoded bitstream 1501 includes a neural network model and corresponding parameters directly carried by one or more SEI messages 1501A and the encoded video stream 1501B. While in the encoded bitstream 1502 including one or more SEI messages 1502A and the encoded video stream 1502B, one or more SEI messages 1502A can carry the identifier of the neural network model and its corresponding parameters. In the scenario of the encoded bitstream 1502, the neural network model and parameters can be stored outside the encoded video stream, for example, stored in Figure 10 1001A of
[0155] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of technical detail of integration. The computer-readable media may include computer-readable non-transitory storage media (or media) having computer-readable program instructions thereon that cause a processor to perform operations.
[0156] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to: an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing devices. A non-exhaustive list of more specific examples of computer-readable storage media includes the following items: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punched card or raised structures in grooves on which instructions are recorded), and any suitable combination of the foregoing items. As used herein, a computer-readable storage medium should not be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (such as the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or an external storage device. The network can include copper transmission cables, transmission optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0158] The computer-readable program code / instructions for performing the operations can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, such programming languages including object-oriented programming languages such as Smalltalk, C++, etc., and including procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be run entirely on the user's computer, partly on the user's computer, run as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can run the computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions to perform various aspects or operations.
[0159] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions run by the processor of the computer or other programmable data processing device create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing device, and / or other devices to act in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0160] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other device, such that a series of operational steps are performed on the computer, other programmable device, or other device to produce a computer-implemented process, such that the instructions run on the computer, other programmable device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. The methods, computer systems, and computer-readable media may include more blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in the figures. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed concurrently or substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0162] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special-purpose control hardware or software code used to implement these systems and / or methods is not a limitation on the implementations. Thus, without reference to specific software code, the operations and behaviors of the systems and / or methods are described herein—it should be understood that the software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0163] The elements, acts, or instructions used herein should not be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Additionally, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with “one or more.” Where only one item is intended, the term “one” or similar language is used. Additionally, as used herein, the terms “have,” “having,” “include,” etc. are intended to be open-ended terms. Also, the phrase “based on” is intended to mean “at least partially based on” unless otherwise explicitly stated.
[0164] The description of various aspects and embodiments has been presented for illustrative purposes, but the various aspects and embodiments are not intended to be exhaustive or limited to the disclosed embodiments. Although combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may only directly depend on one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms chosen herein are intended to best explain the principles of the embodiments, the practical application or technical improvement of the techniques found in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for streaming immersive media, capable of being executed by a processor, characterized in that, The method includes: Ingesting content in a two-dimensional format, where the two-dimensional format references at least one neural network, and the at least one neural network includes a scene-specific neural network that corresponds to a scene included in the ingested content; Converting the ingested content into a three-dimensional format based on the at least one referenced neural network, where the ingested content includes: using the scene-specific neural network to infer depth information related to the scene, and adapting the ingested content to a scene-specific volume format associated with the scene; and Streaming the converted content to a client endpoint, where the at least one neural network is referenced in a supplementary enhancement information SEI message that is included in an encoded video bitstream corresponding to the ingested content.
2. The method according to claim 1, characterized in that, The at least one neural network is trained based on a prior corresponding to an object within the scene.
3. The method according to claim 1, wherein The neural network model and at least one parameter corresponding to the at least one neural network are directly embedded in the SEI message.
4. The method according to claim 1 or 3, characterized in that, Writing the location of the neural network model corresponding to the at least one neural network in the SEI message.
5. The method according to any one of claims 1 and 3, characterized in that The client endpoint includes one or more of a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a dense light field display.
6. A device for streaming immersive media, characterized in that, The device includes: At least one memory configured to store program code; and At least one processor configured to read the program code and operate according to the instructions of the program code, where the program code includes: Ingesting code configured to cause the at least one processor to ingest content in a two-dimensional format, where the two-dimensional format references at least one neural network, and the at least one neural network includes a scene-specific neural network that corresponds to a scene included in the ingested content; Converting code configured to cause the at least one processor to convert the ingested content into a three-dimensional format based on the at least one referenced neural network, where the ingested content includes: using the scene-specific neural network to infer depth information related to the scene, and adapting the ingested content to a scene-specific volume format associated with the scene; and Streaming code configured to cause the at least one processor to stream the converted content to a client endpoint, where the at least one neural network is referenced in a supplementary enhancement information SEI message that is included in an encoded video bitstream corresponding to the ingested content.
7. The device according to claim 6, characterized in that, The at least one neural network is trained based on a prior corresponding to an object within the scene.
8. The device according to claim 6, characterized in that, The neural network model and at least one parameter corresponding to the at least one neural network are directly embedded in the SEI message.
9. The device according to claim 6 or 8, characterized in that, Writing the location of the neural network model corresponding to the at least one neural network in the SEI message.
10. The device according to any one of claims 6 and 8, characterized in that, The client endpoint includes one or more of a television, a computer, a head-mounted display, a lenticular light field display, a holographic display, an augmented reality display, and a dense light field display.
11. A non-transitory computer-readable medium, characterized in that, Stored with instructions, the instructions including one or more instructions which, when executed by at least one processor of a device for streaming immersive media, cause the at least one processor to perform the method of streaming immersive media according to any one of claims 1 to 5.
Citation Information
Patent Citations
Employing three-dimensional (3D) data predicted from two-dimensional (2D) images using neural networks for 3D modeling applications and other applications
US20190026956A1
Supplemental enhancement information messages for neural network based video post processing
US20200304836A1
Hybrid video and feature coding and decoding
WO2020055279A1