Architecture for interactive multi-view multi-plane imaging video streaming

By adopting the MPEG-DASH and ISO-BMFF standards to build an end-to-end interactive multi-view MPI video streaming system, the problem of low viewpoint switching efficiency in video streaming transmission of multi-plane imaging technology is solved, realizing efficient viewpoint switching and immersive experience.

CN121569490APending Publication Date: 2026-02-24DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480048648.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2024-06-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing multi-plane imaging technologies struggle to achieve seamless interactive viewpoint switching during video streaming, resulting in high deployment costs and low efficiency.

Method used

Using MPEG-DASH and ISO-BMFF standards, combined with ISO/IEC 14496-12, 14496-15, 23000-19 and 23009-1 standards, an end-to-end interactive multi-view MPI video streaming system is constructed. It efficiently stores and transmits MPI video bitstreams on the server side and supports dynamic viewpoint switching on the client side.

Benefits of technology

It enables seamless interactive perspective switching for multi-view MPI video streaming, improves the utilization of network and computing resources, reduces deployment costs, and provides a better immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569490A_ABST
    Figure CN121569490A_ABST
Patent Text Reader

Abstract

Methods and apparatus for multi-plane image (MPI) video streaming. According to an example embodiment, a method for video streaming includes providing to a client device a media presentation description of MPI streaming content stored in a storage container accessible by a server device, and a corresponding initialization segment from the storage container. The respective initialization segment is configured to inform a client device of a selection of a perspective of the MPI streaming content for which a media segment for rendering is requested. The method also includes receiving a request from a client device, the request identifying the selection and indicating respective recommended values for at least one of bit rate, resolution, codec type, and frame rate; and transmitting, to the client device, one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 510,571, filed June 27, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] Various exemplary embodiments generally relate to multiplanar imaging (MPI), and more specifically, but not exclusively, to the transmission of multiplanar images. Background Technology

[0004] Multiplane imaging embodies a relatively new approach for storing volumetric content. MPI can be used to render both still images and videos, representing three-dimensional (3D) scenes within the view frustum by using, for example, 8, 16, or 32 texture and transparency (alpha) information planes per camera. Exemplary applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. Summary of the Invention

[0005] This document discloses various embodiments of an end-to-end interactive MPI video streaming system. Various examples of MPI storage formats and transport protocols used in the MPI video streaming system are described, including descriptions of Media Presentation Description (MPD), initialization segments, and media segments. At least some embodiments of the MPI video streaming system provide seamless interactive perspective switching for a multi-view MPI video streaming experience. Some embodiments aim to provide an interactive, immersive volumetric experience by utilizing relevant components of available technologies, such as two-dimensional (2D) video codecs, Dynamic Adaptive Streaming over MPEG HTTP (DASH), and ISO-based Media File Format (BMFF) and its extensions. At least some of the disclosed solutions can be deployed in a relatively short time by modifying related existing solutions according to the various embodiments disclosed herein.

[0006] According to an exemplary embodiment, a method for transmitting an MPI video stream is provided, comprising: providing a client device with a media presentation description of MPI stream content stored in a storage container accessible by a server device; for a period of time, providing the client device with a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, wherein media segments for rendering are requested for the one or more views; receiving a request from the client device, the request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and transmitting to the client device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0007] According to another exemplary embodiment, a non-transitory computer-readable medium is provided that stores instructions which, when executed by an electronic processor of a server device, cause the server device to perform operations including the MPI video streaming method described above.

[0008] According to yet another exemplary embodiment, an apparatus for MPI video streaming transmission is provided, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to combine with the at least one processor such that the apparatus performs at least the following operations: providing a client device with a media presentation description of MPI stream content stored in a storage container accessible via a server device; providing, for a period of time, a corresponding initialization segment from the storage container to the client device, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, wherein media segments for rendering are requested for the one or more views; receiving a request from the client device, the request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and transmitting to the client device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0009] According to yet another exemplary embodiment, a method for transmitting an MPI video stream is provided, comprising: receiving at a client device a media rendering description of MPI stream content stored in a storage container accessible via a server device; for a period of time, receiving at the client device, via the server device, a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, for which a media segment is requested for rendering; transmitting to the server device a request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and receiving from the server device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0010] According to yet another exemplary embodiment, a non-transitory computer-readable medium is provided that stores instructions, when executed by an electronic processor of a client device, to cause the client device to perform operations including the MPI video streaming method described above.

[0011] According to yet another exemplary embodiment, an apparatus for MPI video stream transmission is provided, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to combine with the at least one processor such that the apparatus performs at least the following operations: receiving at a client device a media rendering description of MPI stream content stored in a storage container accessible via a server device; for one cycle, receiving at the client device via the server device a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, for which a media segment is requested for rendering; transmitting to the server device a request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and receiving from the server device one or more bitstreams, the one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values. Attached Figure Description

[0012] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent from the following detailed description and accompanying drawings by way of example:

[0013] Figure 1 illustrates an example process of a video / image delivery pipeline.

[0014] Figure 2 schematically illustrates a 3D scene representation using multi-planar images according to an embodiment.

[0015] Figure 3 schematically illustrates the process of generating a new perspective of a 3D scene based on an example.

[0016] Figure 4 is a block diagram showing how the set of active viewpoints changes over time according to an example.

[0017] Figure 5 is a block diagram illustrating an example communication system that can implement MPEG DASH streaming according to some embodiments.

[0018] Figure 6 is a block diagram illustrating an advanced hierarchical data model that can be used in DASH according to some embodiments.

[0019] Figure 7 is a block diagram illustrating an example practical application of the various hierarchical levels assigned to the data model of Figure 6 according to some embodiments.

[0020] Figure 8 shows pseudocode for a generic tag used to describe a Media Presentation Description (MPD) in XML format, based on some examples.

[0021] Figure 9 shows an example of a restricted video sample entry that can be used in some embodiments.

[0022] Figure 10 is a block diagram showing the overall structure of an ISO BMFF file for timing data, based on some examples.

[0023] Figure 11 is a block diagram showing the overall segment structure of a fragmented MP4 according to some examples.

[0024] Figure 12 is a block diagram showing the overall structure of the DASH representation based on some examples.

[0025] Figure 13 is a block diagram illustrating a communication system configured to support MPI transmission using DASH and ISO BMFF according to some embodiments.

[0026] Figure 14 is a block diagram illustrating the data structure of the MPI container used in the communication system of Figure 13 according to some embodiments.

[0027] Figure 15 shows an example of an MPD that can be used in the communication system of Figure 13 according to some embodiments.

[0028] Figure 16 shows another example of an MPD that can be used in the communication system of Figure 13 according to some embodiments.

[0029] Figure 17 shows yet another example of an MPD that can be used in the communication system of Figure 13 according to some embodiments.

[0030] Figure 18 shows yet another example of an MPD that can be used in the communication system of Figure 13 according to some embodiments.

[0031] Figures 19A-19F provide example definitions related to the viewpoint identifiers used in the communication system of Figure 13, according to some embodiments.

[0032] Figures 20A and 20B provide example definitions related to camera intrinsics used in the communication system of Figure 13, according to some embodiments.

[0033] Figures 21A and 21B provide example definitions related to camera extrinsics used in the communication system of Figure 13, according to some embodiments.

[0034] Figure 22 provides example definitions related to the restricted scheme information used in the communication system of Figure 13, according to some embodiments.

[0035] Figures 23A and 23B provide example definitions related to the SEI information box used in the communication system of Figure 13 according to some embodiments.

[0036] Figure 24 illustrates an example of an MPI video arrangement method used in the communication system of Figure 13 according to some embodiments.

[0037] Figure 25 provides example definitions related to the MPI video arrangement used in the communication system of Figure 13 according to some embodiments.

[0038] Figure 26 illustrates an example of another MPI video arrangement method used in the communication system of Figure 13 according to some embodiments.

[0039] Figure 27 illustrates an example of specifying multi-view camera settings in the communication system of Figure 13 according to some embodiments.

[0040] Figures 28A-28C provide example definitions related to the prefetch operation used in the communication system of Figure 13, according to some embodiments.

[0041] Figure 29 is a block diagram illustrating the encapsulation of multi-view MPI encoded data in an ISO BMFF file for storage, according to some embodiments.

[0042] Figure 30 is a block diagram illustrating a computing device used in the communication system of Figure 13 according to some embodiments. Detailed Implementation

[0043] This disclosure and its aspects may take many forms, including hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces and application programming interfaces; as well as hardware-implemented methods, signal processing circuits, memory arrays, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. The foregoing is intended only to give a general concept of the various aspects of this disclosure and does not limit the scope of this disclosure in any way.

[0044] In the following description, numerous details, such as device configuration, timing, and operation, are set forth to provide an understanding of one or more aspects of this disclosure. These specific details will be apparent to those skilled in the art that they are merely exemplary and are not intended to limit the scope of this application.

[0045] Example video / image delivery pipeline

[0046] Figure 1 depicts an example process of a video delivery pipeline (100), illustrating the various stages from video / image capture to video / image content display according to an embodiment. A series of video / image frames (102) can be captured or generated using an image generation block (105). Frames (102) can be digitally captured (e.g., by a digital camera) or computer-generated (e.g., using computer animation) to provide video and / or image data (107). Alternatively, frames (102) can be captured on film by a film camera. The film can then be converted to a digital format to provide video / image data (107). In some examples, the image generation block (105) includes generating MPI images or videos.

[0047] During the production phase (110), data (107) can be edited to provide a video / image production stream (112). The data in the video / image production stream (112) can be provided to a processor (or one or more processors, such as a central processing unit (CPU)) at the post-production block (115) for post-production editing. Post-production editing at the block (115) may include, for example, adjusting or modifying the color or brightness of specific areas of an image to enhance image quality or achieve a specific look for the image, depending on the creative intent of the video creator. This part of post-production editing is sometimes referred to as “color timing” or “color grading”. Other edits may be performed at the block (115) (e.g., scene selection and sorting, image cropping, adding computer-generated visual effects, removing artifacts, etc.) to produce a “final” version of the production for distribution (117). In some examples, operations performed at the block (115) include enhancing the texture and / or alpha channel in a multi-planar image / video. During post-production editing (115), the video and / or image can be viewed on a reference monitor (125).

[0048] After post-production (115), the data of the final version (117) can be transmitted to the encoding block (120) for further downstream transmission to decoding and playback devices, such as televisions, set-top boxes, cinemas, etc. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray, and other transmission formats, to generate an encoded bitstream (122). In the receiver, the encoded bitstream (122) is decoded by the decoding unit (130) to generate a copy or closely approximate corresponding decoded signal (132) representing the signal (117). The receiver may be connected to a target display (140), the characteristics of which may be somewhat different or completely different from those of the reference display (125). In this case, a display management (DM) block (135) can be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display mapping signal (137). According to an embodiment, the decoding unit (130) and the display management block (135) may include their respective processors, or may be based on a single integrated processing unit.

[0049] The codecs used in the encoding block (120) and / or decoding block (130) enable video / image data processing and compression / decompression. Compression is used in the encoding block (120) to reduce the size of the corresponding file or stream. The decoding process performed by the decoding block (130) typically involves decompressing the received video / image data file or stream into a form suitable for playback and / or further editing. Example encoding / decoding operations that may be used in the encoding block (120) and decoding unit (130) according to various embodiments are described in more detail below.

[0050] Multiplanar imaging

[0051] A multi-plane image comprises multiple image planes, each a "snapshot" of the 3D scene at a specific depth relative to the camera position. Information stored in each plane includes texture information (e.g., represented by R, G, B values) and transparency information (e.g., represented by alpha (A) values). In this paper, the abbreviations R, G, and B represent red, green, and blue, respectively. In some examples, the three texture components can be (Y, Cb, Cr) or (I, Ct, Cp) or another functionally similar set of values. Multi-plane images can be generated in various ways. For example, two or more input images from two or more cameras located at different known viewpoints can be collaboratively processed to generate corresponding multi-plane images. Alternatively, a multi-plane image can be generated using a source image captured by a single camera.

[0052] Figure 2 schematically illustrates a 3D scene representation using a multi-plane image (200) according to an embodiment. The multi-plane image (200) has D planes or layers (P0, P1, ..., P(D-1)), where D is an integer greater than 1. Typically, the planes (layers) are indexed such that the layer furthest from the reference camera position (RCP) is indexed as layer 0, and its distance (or depth) from the RCP along the Z dimension of the 3D scene is d0. For each subsequent layer closer to the RCP, the index is incremented by 1. The plane (layer) closest to the RCP has an index value of (D-1) and a distance (or depth) from the RCP along the Z dimension of d0. (D-1) Each of the planes (P0, P1, ..., P(D-1)) is orthogonal to the reference plane (202) which is parallel to the XZ coordinate plane. The RCP is located at a vertical height h above the reference plane (202). Figure 2 The XYZ triplet shown indicates the overall orientation of the multiplane image (200) and planes (P0, P1, ..., P(D-1)) relative to the X, Y, and Z dimensions of the 3D scene. In various examples, the quantity D can be 32, 16, 8, or any other suitable integer greater than 1.

[0053] The color component (e.g., RGB) value of the i-th layer at camera position s is represented as... The layer's horizontal dimensions are H × W, where H is the layer's height (Y dimension) and W is the layer's width (X dimension). Color channel c is located at... The pixel value at that location is represented as The α value of the i-th layer is Location in the Alpha layer The pixel value at that location is represented as The depth distance from the i-th layer to the reference camera position is... The image from the original reference viewpoint (without camera movement) is represented as R, and the texture pixel value is... Therefore, the static MPI image at camera position s can be represented as:

[0054]

[0055] If the camera position s remains static, this static MPI image representation can be directly extended to the video representation. This video representation is given by equation (2):

[0056]

[0057] Where t represents time.

[0058] As previously mentioned, a multiplanar image (e.g., a multiplanar image (200)) can be generated from a single source image R or from two or more source images. This generation can, for example, be performed during the production phase (110). The corresponding MPI generation algorithm can typically output a multiplanar image (200) containing XYZ parsed pixel values, in the form of… .

[0059] By processing The multiplanar image (200) represents a viewable image that the MPI rendering algorithm can generate corresponding to or different from the RCP (Representational Plane Point). Example MPI rendering algorithms (often referred to as "MPI viewers") that can be used for this purpose may include warping and compositing steps. Other suitable MPI viewers may also be used. The rendered multiplanar image (200) can be viewed, for example, on a reference display (125).

[0060] During the warp step of the MPI rendering algorithm, each layer of the multi-plane image (200) ( , ) can be viewed from the RCP viewpoint position ( ) warp to a new viewpoint position ( For example, as follows:

[0061] (3)

[0062] (4)

[0063] in σ is the distortion function; σ is the consistency ratio (to minimize error). In an exemplary embodiment, the distortion function... It can be represented as follows:

[0064] (5)

[0065] in and Through (5), the position of each pixel on the target viewpoint of a specific MPI plane ( It can be mapped to its corresponding pixel position in the source view ( .function and Let represent the camera intrinsic parameter models of the reference viewpoint and the target viewpoint, respectively. Functions R and t represent the camera extrinsic parameter models of rotation and translation, respectively. n represents the normal vector [0 0 1]. T 'a' indicates the depth. The distance between the point and the plane parallel to the front of the source camera.

[0066] During the compositing step of the MPI rendering algorithm, a new viewable image can be generated, for example, using processing operations corresponding to the following formula. :

[0067] (6)

[0068] Among them, weight It can be represented as:

[0069] (7)

[0070] disparity map corresponding to the source view It can be calculated as follows:

[0071] (8)

[0072] Among them, weight Represented as:

[0073] (9)

[0074] The MPI rendering algorithm can also be used to generate viewable images corresponding to RCP. In this case, the distortion step is omitted, and the image... The calculation is as follows:

[0075] (10)

[0076] In single-camera transmission scenarios, only one MPI is fed through the bitstream. The goal in this case is to optimize the layers that merge the original MPIs to maintain the quality of the MPI after local distortion. In multi-camera transmission scenarios, multiple MPIs captured at different camera positions are encoded in a compressed bitstream. Information from these MPIs is combined to generate a new global perspective for positions located between the original camera positions. There are also scenarios where information from multiple cameras can be combined to generate a single MPI for transmission. Multi-camera transmission scenarios are commonly used for MPI video transmission, as described below.

[0077] Figure 3 The process of generating a new perspective of a 3D scene (302) according to an example is illustrated. In the example shown, the 3D scene (302) is captured using forty-two RCPs (1, 2, ..., 42). The new perspective being generated corresponds to the camera position (50). The four RCPs closest to the camera position (50) are RCPs (11, 12, 18, 19). The corresponding multiplane image is a multiplane image (200). 11 200 12 200 18 200 19 By correspondingly distorting the multiplanar image (200) 11 200 12 200 18 200 19 Then, the resulting distorted multiplanar images are merged to generate a multiplanar image (200) corresponding to the camera position (50). 50 Finally, the synthesis step (310) of the MPI rendering algorithm is applied to the multi-plane image (200). 50 ), generating a viewable image (312) of the 3D scene (302).

[0078] Typically, any appropriate number of RCPs can be used to capture a 3D scene (e.g., a 3D scene (302)). The positions of these RCPs can also be chosen differently, for example, based on creative intent. In a typical practical example, when rendering a new viewpoint (e.g., a viewable image (312)), only a few adjacent RCPs are used for rendering. This adjacent viewpoint is then referred to as the “active view”. In the example shown in Figure 3, the number of active views is four. In other examples, a different number (other than four) of active views can be used similarly. Therefore, the number of active views is an optional parameter. For illustrative purposes and without any implied limitations, some exemplary embodiments are described below with reference to four active views. In some examples, the set of active views may change over time as the camera position (50) moves. In some examples, the number of active views may change over time as the camera position (50) moves.

[0079] Figure 4 is a block diagram illustrating the change of the active view set over time according to an example. In the example shown, a 3D scene is captured using a rectangular array of forty RCPs arranged in five rows and eight columns. The dashed arrow (402) represents the movement trajectory of the new camera position (50) during the time interval from time t0 to time t1 used for virtual view synthesis. At time t0, the active view set includes four views enclosed by the dashed box (410). At time t1, the active view set includes four views enclosed by the dotted box (420). At time t... n (where t0 < t) n < t1), the set of active viewpoints changes from the set in the dashed box (410) to the set in the dotted box (420).

[0080] The various embodiments disclosed herein are intended to provide an end-to-end interactive multi-view MPI video streaming system, including:

[0081] ● Components located on the server side and the content storage formats used therein;

[0082] ● Components that support network distribution protocols; and

[0083] ● Components that support interactive playback solutions on the terminal client side.

[0084] Some embodiments use the MPEG-DASH and ISO-BMFF standards and / or their extensions, and additionally provide new tools to support new features. One objective of this embodiment is to enable seamless interactive perspective switching in multi-view MPI video streaming by building upon and extending certain legacy technologies / infrastructure, thereby enabling the rapid and / or cost-effective deployment of interactive immersive volumetric video experiences. Some embodiments may benefit from at least some of the features disclosed in one or more of the following documents: (1) ISO / IEC 14496-12:2020, Information technology—Encoding of audiovisual objects—Part 12: ISO basic media file format; (2) ISO / IEC 14496-15:2021, Information technology—Encoding of audiovisual objects—Part 15: Video carrying a Network Abstraction Layer (NAL) unit structure in ISO basic media file format; (3) ISO / IEC 23000-19:2018, Information technology—Formats for Multimedia Applications (MPEG-A)—Part 19: Common Media Application Format (CMAF) for segmented media; and (4) ISO / IEC 23009-1 (2019), Information technology—Dynamically Adaptive Streaming over HTTP (DASH)—Part 1: Media presentation description and segment format. Each of these documents is incorporated herein by reference in its entirety.

[0085] In at least some examples, multi-view MPI representation can convey one or both of the following benefits:

[0086] ● Improved single MPI view space rendering. In some examples, a maximum of four views are sufficient. In other examples, the number of views can be reduced, for example, to a single view, through different MPI generation processes.

[0087] ● It can render new perspectives in both space and time over time, such as providing more rendering pose options and a better experience.

[0088] The above figure 4 illustrates an example of spatial and temporal switching of the reference viewpoint.

[0089] In some examples, multi-view MPI systems operate by independently encoding each single-view MPI using one of the frame packing schemes disclosed in commonly owned U.S. Patent Application No. 63 / 495,715, filed April 12, 2023, entitled “Transmission of Volume Images in Multi-Plane Imaging Format,” which is incorporated herein by reference in its entirety. In some use cases, users will request different new view poses at different times. Therefore, the client side of the system needs to have the ability to dynamically switch to different MPI bitstreams. One approach is to download the MPI bitstreams of all views to the client side. However, this approach may not be bandwidth efficient. Another approach is to stream a smaller number of appropriately selected view MPI bitstreams to improve network and / or computational resource utilization. Examples of challenges in implementing the latter approach include:

[0090] ● How to efficiently store MPI video bitstreams on the server side;

[0091] ● How to efficiently transmit MPI video bitstreams between the server and client sides; and

[0092] ● How to efficiently manage MPI bitstream switching and decode / render new perspectives based on user and / or network interactivity.

[0093] The embodiments disclosed below provide schemes for constructing end-to-end interactive multi-view MPI video streaming systems. For storage formats, various embodiments are based on the ISO Basic Media File Format (commonly referred to as ISO BMFF) (ISO / IEC 14496-12) and its extension ISO / IEC 14496-15, which provides the carrying of Network Abstraction Layer (NAL) unit structure video (e.g., AVC, HEVC, VVC, etc.). In this document, the abbreviation ISO stands for International Organization for Standardization. For data transmission, various embodiments are based on Dynamic Adaptive Streaming over HTTP (DASH) (ISO / IEC 23009-1). In some examples, a similar approach is applied to the Common Media Application Format (CMAF). It should be noted that the disclosed video streaming systems can also support other multi-view coding schemes, such as MVC, SHVC, and multi-layer VVC. Exemplary embodiments are described with reference to the HEVC codec for illustrative purposes and without any implied limitations. Based on the provided description, those skilled in the art will be able to make and use other embodiments compatible with other suitable codecs without excessive experimentation.

[0094] MPEG DASH

[0095] Figure 5 is a block diagram illustrating an example communication system (500) capable of implementing MPEG DASH streaming according to some embodiments. MPEG DASH is a media streaming technology that enables adaptive streaming of high-quality media content (e.g., video, audio, etc.) over the Internet via HTTP web servers and proxies. In the illustrated example, the communication system (500) includes a media capture device (510) connected to a computing device (520). The computing device (520) is further networked to multiple servers, including a Digital Rights Management (DRM) encryption server (530), a media source server (540), and an HTTP caching server (550). Multiple client devices (560) can connect to receive digital content from the HTTP caching server (550), provided that appropriate licenses for the digital content are available from one of the DRM license servers (570).

[0096] From the client's perspective, an example of the DASH streaming process includes the following operation blocks:

[0097] 1) The client device (560) obtains a Media Presentation Description (MPD) of the streaming content (e.g., a movie). In some examples, the MPD includes information about different alternative representations of the streaming content (e.g., bitrate, video resolution, frame rate, audio language) and the URL of the HTTP resources (initialization segment and media segment).

[0098] 2) The client device (560) requests one or more of the desired representations at a time, one segment or (part of it), based on relevant information in the MPD and local information of the client (such as network bandwidth, decoding / display capabilities and / or user preferences).

[0099] 3) When the client device detects a change in network bandwidth, the client device (560) requests a segment with a different representation that better matches the bit rate. In some cases, it starts with a segment that begins with a random access point.

[0100] Figure 6 is a block diagram illustrating an advanced hierarchical data model (600) that can be used in DASH according to some embodiments. In one dimension (schematically a horizontal domain), the model (600) includes the time series corresponding to the media presentation. In another dimension (schematically a vertical domain), the model (600) includes the choices provided in the media presentation, which are selected by the DASH client in a static or dynamic manner.

[0101] In some examples, model (600) depends on the following definitions of the various levels used in the DASH hierarchy:

[0102] ● Media Presentation Description (MPD): DASH media presentation is described by the MPD document (602). This document describes a series of periods (610) that constitute the media presentation over time.

[0103] ● Period: Period (610) usually refers to the media content period during which a consistent set of encoded versions of the media content is obtained. Within period (610), the material is placed into the adaptation set (620).

[0104] ● Adaptation set: An adaptation set (620) represents a set of interchangeable encoded versions of one or more media content components and contains a set of representations (630). In some examples, the concept of pre-selection (612) is used to combine different adaptation sets (630) into a single decoding instance and user experience.

[0105] ● Presentation: A representation (630) describes a transportable coded version of one or more media content components and includes one or more media streams (e.g., one media stream for each media content component in a multiplex). An example way to use a representation (630) is to apply different bit rates. Within a representation (630), content can be temporally divided into segments (632, 634, 636) for appropriate accessibility and transport.

[0106] ● Segment: In some examples, a segment is the largest unit of data that can be retrieved through a single HTTP request. To access a segment, a URL is provided for each segment. Two types of segments are used to represent them:

[0107] ○ The initialization section (632) contains static metadata representing (630);

[0108] ○ Media segment (634) contains media samples and advances the timeline.

[0109] ● Self-initialized segment: Some representations (630) can also be organized as having a single self-initialized segment (636) which contains both initialization information and media data.

[0110] Figure 7 is a block diagram illustrating example practical applications of the various hierarchical levels of the data model of Figure 6 according to some embodiments. In the example shown, the corresponding MPD has four cycles (610). The cycle (610) can be used for a first practical application (702) involving splicing of arbitrary content (e.g., ad insertion, etc.). Each cycle (610) includes multiple adapter sets (620), which can be different types of multimedia. The corresponding second practical application (710) involves selecting components or tracks based on attributes. In each adapter set (620), there are different representations (630), each representation using a different corresponding bit rate. The corresponding third practical application (720) involves selecting and / or switching representations (630) based on communication channel bandwidth. In each representation 630, there are multiple segments, including, for example, an initialization segment (632) and one or more media segments (634). The corresponding fourth practical application (730) provides a well-defined media format for each segment. Each segment (632, 634) contains a Uniform Resource Locator (URL) address pointing to an mp4 bitstream. The corresponding fifth practical application (740) involves providing a media transmission format and data blocks with unique addresses and associated timings.

[0111] Figure 8 is pseudocode showing common tags (800) for describing MPD in XML format according to some examples.

[0112] ISO Basic Media File Format (BMFF)

[0113] ISO BMFF uses "boxes" to organize data into different categories. Each box is guided by a 4-character type (4CC). Boxes can be nested and organized hierarchically; for example, one box can be inside another. The following pseudocode illustrates the box concept.

[0114]

[0115] ● "size" is an integer that specifies the number of bytes in this box (including all its fields and the boxes it contains); if "size" is 1, the actual size is in the "largesize" field; if "size" is 0, this box is the last box in the file and its contents extend to the end of the file (typically only used for Media Data Boxes).

[0116] ● The “Type” field identifies the bin type; standard bins use a compact type (typically four printable characters) for easy identification, as shown in the bins below. User extensions use an extended type; in this case, the “Type” field is set to 'uuid'.

[0117] In addition to the "box" described above, a "full box" with more information can be defined as follows:

[0118]

[0119] Two different definitions are used to separate timed data from non-timed data.

[0120] ● Track: Image sequences are stored as tracks. Image sequence tracks are used when there are encoding dependencies between images or when image playback is timed. Unlike video tracks, timing in image sequence tracks is advisory.

[0121] ● Item: Still images are stored as items. All image items are encoded independently, and their decoding does not depend on any other item. Any number of image items can be contained in the same file.

[0122] Timing data in ISO BMFF files is divided into samples, which are described using a logical concept called a "track". Track bins, identified by 4CC 'trak', are stored within the movie bin 'moov' and contain multiple bins that provide information about where the samples are found in the file and how to interpret them. The samples themselves are stored in the media data bin, identified by 4CC 'mdat'. Each track contains at least one sample entry within the sample description bin 'stsd'. The sample entry describes the encoding and container format used in the track sample. Sample entries are also identified by 4CC and can contain more bins for further configuration. In some examples, non-timing data is stored in the metadata bin 'meta', which is represented by a sample by a logical concept called an item. Non-timing sample data can be stored in the media data bin or in the item data bin 'idat' within the metadata bin.

[0123] ISO BMFF defines the top-level box as follows:

[0124]

[0125] It should be noted that one sample entry type defined by the ISO BMFF specification is a restricted video sample entry identified by 4CC 'resv'. This sample entry is used when decoded samples of a track require post-processing for rendering or other purposes. The post-processing operation is identified by the scheme type of the signaling delivery within the restricted sample entry itself. Regarding signaling supplementary enhancement information (SEI), the file author can list the SEI message IDs that appear and categorize them into two types: those the file author deems necessary for proper playback, and others. The appearance of any type of SEI message can be achieved using the 'seii' signaling delivery under the scheme information box ('schi') in the SEI information box. The scheme used for signaling delivery SEI uses 'aSEI' as the scheme type.

[0126] Figure 9 Table 1 shows an example of the restricted video sample entry 'resv'. Restricted sample entries are useful, for example, for MPI video, because MPI uses post-processing based on MPI information, SEI messages, and camera information.

[0127] Figure 10 This is a block diagram showing the overall structure of an ISO BMFF file (1000) for timing data, based on some examples. Figure 10 The illustrations provided further details of the various boxes in the description document (1000). Arrows indicate several boxes that are considered particularly relevant to the exemplary embodiments disclosed herein.

[0128] Figure 11 This is a block diagram illustrating the overall segment structure of a fragmented MP4 used in DASH, based on some examples. The example shown contains an initialization segment (1102) and a media segment (1104). The initialization segment (1102) contains a 'moov' box. In some examples, the 'moov' box is mandatory and there should be exactly one. This box contains other boxes that have detailed metadata information about the representation. The media segment (1104) contains a 'moof' box and a 'mdat' box. In some examples, the 'moov' box contains the metadata for the entire movie, but the 'moof' box is used to allow tracks (such as audio, video, and text) to be broken down into segments and contains corresponding metadata for each track. Each segment has one 'moof' box. The media data box 'mdat' stores the actual media sample of the representation. Each media segment contains one or more complete self-contained movie segments. A complete self-contained movie clip consists of a movie clip ('moof') box and a media data ('mdat') box, which contains all media samples of external data references that are not used in the movie clip box for track runs. Each 'moof' box contains at least one track clip.

[0129] Figure 12 is a block diagram illustrating the overall structure of a DASH representation (1200) according to some examples. In the example shown, the representation (1200) includes an initialization segment (1202) and multiple media segments (1204). The initialization segment (1202) contains static metadata of the representation (1200). The media segments (1204) contain media samples and advance the timeline. As previously mentioned, ISO BMFF supports segmented designs, such as the segmented design of the representation (1200).

[0130] MPI transfer using DASH and ISO BMFF

[0131] Figure 13 is a block diagram illustrating a communication system (1300) configured to support MPI transmission using DASH and ISO BMFF according to some embodiments. The communication system (1300) includes a server (1302) and a client (1306) connected via a communication channel (1304). The server (1302) includes a data storage device storing one or more MPI containers (1400). In operation, the client (1306) requests and receives, via the communication channel (1304), a relevant portion of the data stored in the MPI container (1400), as described in more detail below.

[0132] Figure 14 is a block diagram illustrating the data structure of an MPI container (1400) used in a communication system (1300) according to some embodiments. The MPI container (1400) includes multiple segments that are logically organized (e.g., using appropriate indexes) to form three plane sets in a 3D space, the dimensions of which are the fit set, representation, and segment time. There are two types of segments: initialization segments (632) and media segments (634). The end user at the client (1306) typically requests an initialization segment (632) for each period (610). In one example, the initialization segment (632) contains camera information and metadata related to media segment components (e.g., video, audio, captions, etc.) so that the end user knows which viewpoint(s)(s) to use. Each media segment(634) contains audio and / or video content for a different corresponding viewpoint at a different corresponding bitrate in the corresponding period (610). Through the client (1306), the end user typically requests only the media segment(634) required.

[0133] Each adapter set includes a representation (630) indexed as being logically located in a plane orthogonal to the adapter set axis in the logical 3D space of the MPI container (1400). In one example, each adapter set has an MPI from one camera viewpoint. If we have fifty camera views, the MPI container (1400) has fifty corresponding adapter sets. The end user selects a specific adapter set by choosing the camera MPI to render. In various examples, depending on the rendering algorithm and other relevant conditions, the end user can request segments from one adapter set for single-view MPI rendering or request segments from multiple adapter sets for multi-view MPI rendering.

[0134] Each representation comprises segments (632, 634) indexed as being logically located in a plane orthogonal to the representation axes of the logical 3D space of the MPI container (1400). In one example, each representation is associated with a corresponding encoded MPI bit rate. Different representations typically correspond to different corresponding bit rates. The segments (632, 634) of the various representations are selected based on the current conditions of the communication channel (1304) between the server (1302) and the client (1306).

[0135] Each segment time plane has a segment (634) indexed as being logically located in a plane orthogonal to the segment time axis of the logical 3D space of the MPI container (1400). In one example, each segment time plane is associated with a different corresponding playback time within the video. Different segment time planes typically correspond to different corresponding playback times. In the view presented in Figure 14, each segment (634) is shown as a bin, whose position in the 3D space of the MPI container (1400) reflects the playback time, viewpoint, and bitrate of that segment.

[0136] In some examples, in order to enable end users to select the media segment (634) to be downloaded via the client (1306), the metadata transmitted through the communication channel (1304) includes:

[0137] ● MPD (1312): MPD (1312) provides information about the MPI container (1400) and the URLs of different segments corresponding to different adapter sets, representations, and segment time planes. MPD (1312) enables end users to retrieve sufficient MPI scene information and determine the media bitstream to download.

[0138] ● Initialization Section (1314): The initialization section (1314) is sent for each cycle (610) and provides information about the number of views, camera intrinsic / extrinsic parameters, and indicators of post-decoder operations along the time dimension. It should be noted that different MPI scenes may have different corresponding number of cameras and different corresponding camera pose settings. The transmitted initialization section (1314) contains those time-related information.

[0139] At the client (1306), three inputs are used to determine which media segment(s)(s) to request from the server (1302):

[0140] ● New viewing posture (1336):

[0141] In some examples, the client (1306) infers the user’s new choice of viewing posture from the user input device (eye tracking, touch screen, sensor, etc.) (1336).

[0142] ○ The client (1306) obtains information about the camera view selection from the initialization section (1314).

[0143] Based on the above information, the client (1306) operates to determine which view (adapter set) of the MPI container (1400) to request the segment from.

[0144] Note that, depending on the rendering algorithm, segments can be requested from one or more adapter sets of the MPI container (1400).

[0145] In some examples, a dedicated algorithm can be used to request a larger number of adjacent views, for example at a lower bit rate due to bandwidth considerations. This dedicated algorithm is configured to provide a smooth playback transition when the end user requests a relatively large change in a new viewing posture (1336).

[0146] ● Network conditions (1332): Based on the network conditions experienced by the communication channel (1304) (e.g., bandwidth, packet loss and packet delay), the client (1306) operates to determine which representation of the MPI container (1400) to request the segment from.

[0147] ● Current time (1334): Based on the current image sequence number (POC), the client (1306) "knows" how many frames are left in the current segment, and which segment to request from the server (1302) next for continuous playback.

[0148] Using the inputs described above (1332, 1334, 1336), the client (1306) operates to generate a corresponding request (1326) for the desired media segment (634) required for the end user to continue seamless MPI playback. The request (1326) is then passed to the server (1302), for example, via an HTTP GET message (1316). In response to message (1316), the server (1302) transmits an MPI bitstream (1318) carrying the desired media segment (634) for rendering (1338) at the client (1306).

[0149] In various other embodiments, the MPI container (1400) can be further extended to include additional dimensions such as resolution, codec type, frame rate, etc. In some examples, the bitrate dimension can be reduced to a single bitrate value.

[0150] MPD

[0151] Some embodiments of the MPD (1312) have been described above with reference to Figures 5-8. Further embodiments of the MPD (1312) are described below with reference to Figures 15-18. These further embodiments cover two example use cases: (i) a single-view MPI use case and (ii) a multi-view MPI use case.

[0152] Figure 15 shows Table 2, which provides an example of MPD (1312) for single-view MPI. MPI-related information can be extracted from the initialization segment (632) (“init.mp4”), as shown in Figure 12, or from 'moov' when there is only one segment in the representation dimension of the MPI container (1400). In the example shown, the value of @codecs is set to “resv.mpiv.XXXX”, where XXXX corresponds to the 4CC of the video codec in the original_format field of the RestrictedSchemeInfoBox of the sample entry (e.g., 'avc1' or 'hvc1'). This feature can be used to relatively quickly inform the client player that the video is in MPI package format so that the player can convert to the appropriate configuration. In another example, the value of @codec is set to “XXXX”, and parsing of the initialization segment (632) is performed.

[0153] In some examples, it is desirable to provide users with a free-view rendering experience in a multi-view MPI system. To achieve this, the client (1306) needs to "know" the multi-camera (reference viewpoint) settings from the MPD. Then, the client (1306) can request the corresponding reference viewpoint MPI video bitstream based on the input (1336) described above. Because all the information required for free-view rendering is captured in at least some embodiments of the ISO BMFF design described above, the following basic format can be used for encapsulation and signaling in DASH:

[0154] - Each MPI reference view video is represented as a separate adaptation set in DASH MPD.

[0155] - Each adapter set contains a corresponding initialization section (632).

[0156] When the client (1306) receives the MPD (1312), it can construct the camera setup based on metadata information from the initialization segments (1314) from all viewpoints. In some examples, the client (1306) can use 'mvcg' to identify prefetched neighbors and request low-resolution / low-bitrate video to reduce latency in free-view rendering (1338). In some examples, the client (1306) can also be configured to use the Viewpoint element in DASH to achieve this, for example, as follows:

[0157]

[0158] Figure 16 illustrates Table 3, which provides an example of an MPD (1312) for multi-view MPI. For simplicity, some parts of the MPD (1312) are not explicitly shown. Based on the description provided herein, those skilled in the art will be able to add those parts without excessive experimentation. In the example shown, viewpoint 0, viewpoint 1, etc., are placed in different corresponding adapter sets, identified by their respective viewpoint values, and contain different corresponding bit rates. Each adapter set contains a corresponding initialization segment (632).

[0159] Figure 17 illustrates Table 4, which provides another example of an MPD (1312) for multi-view MPI. The example in Figure 17 can be used, for example, when the user does not wish to interact with any free-view rendering and is therefore intended to view a director-recommended fly-through perspective. In some cases, this example can also be used to provide an initial viewpoint.

[0160] Figure 18 illustrates Table 5, which provides another example of an MPD (1312) for multi-view MPI. In this example, an adapter set is included, which has only one representation. This representation includes only an initialization segment (632), but no media segment (634). The initialization segment (632) contains multi-view global information. In this way, the client (1306) only needs to download the initialization segment (632) and obtain the global camera settings.

[0161] Initialize segment bitstream

[0162] To enable new perspective rendering using MPI scene representation, camera information, including both intrinsic and extrinsic parameters, is required in addition to MPI side information. Because MPI uses post-processing based on MPI information SEI messages and camera information—which is part of the post-decoder input required for media processing—some embodiments disclosed herein involve the storage of video component tracks. In some examples, this storage leverages the existing capabilities of ISO BMFF and the corresponding ISO / IEC 14496-15 mechanism. This mechanism allows players to inspect files to determine their ability to render bitstreams, thus preventing traditional players from decoding and rendering files for which they lack the appropriate processing capabilities. Therefore, the key information contained in the initialization section (632) is: (i) camera pose, including both internal and external camera models; and (ii) an indicator indicating the need for post-decoder operations using SEI. In the example scenario considered below, there are N cameras. This information will first be discussed in more detail, followed by how it can be organized across different adapter sets.

[0163] For carrying MPI-encoded data in ISO BMFF, the various embodiments described below address the following issues:

[0164] ● Viewpoint identifier (related to multi-camera settings);

[0165] ● The carrier of internal camera information;

[0166] ● The carrier of external camera information; and

[0167] ● Signaling delivery after using the SEI decoder.

[0168] As used herein, the term "camera" should be interpreted to encompass both real and virtual cameras. In some examples, for virtual cameras, MPI data may be rendered from camera interpolation, 3D model rendering (e.g., neural radiation field (NeRF)), or using another suitable method. Some embodiments utilize existing ISO BMFF capabilities (e.g., those described in ISO / IEC 14496-15), while other embodiments may employ new boxes, schemes, and / or types of enhancements or new features specifically implemented for MPI video streaming.

[0169] Figures 19A-19F Definitions related to view identifiers used in the communication system (1300) are provided according to some embodiments. The view identifier is represented by a "vwid" box, which can be as follows: Figure 19A and 19B The definition is shown. Figure 19C Example syntax for view_id (view ID) is provided, which is generally suitable for MPI transmission purposes in communication systems (1300). However, Figure 19C The semantics of `view_id` shown indicate the value of the `view_id` syntax element in the NAL unit header extension, which may not exist for single-layer codecs. To address this issue and related problems with syntax not being available in some configurations, Figure 19D-19F Additional syntax for temporarily marking version (version) = 1 is defined, and further sample entries are added to allow inclusion of HEVC and VVC use cases. In some applications, when backward compatibility is not required or cannot be used, some alternative embodiments can use a newly defined bin (e.g., 'mvid') instead of 'vwid' version = 1.

[0170] Figure 20A and 20B Example definitions related to camera intrinsics used in the communication system (1300) are provided according to some embodiments. In the illustrated example, the camera intrinsic box 'icam' is compatible with ISO / IEC 14496-15.

[0171] Figures 21A and 21B provide example definitions related to camera extrinsics used in the communication system (1300) according to some embodiments. In the examples shown, the camera intrinsic box 'ecam' is compatible with ISO / IEC 14496-15. For MPI-encoded video, the 'icam' and 'ecam' boxes can be reused by expanding the container Sample Entry to include HEVC and VVC, for example, as follows:

[0172]

[0173] It should be noted that in some examples, the value of the syntax 'view_id' in the 'vwid' box, and the syntax 'ref_view_id (reference view ID)' in the 'ecam' and 'icam' boxes are set to be equal to the MPI_view_id in the MPI information SEI message of the same camera view.

[0174] Figure 22 provides example definitions related to restricted scheme information used in the communication system (1300) according to some embodiments. In some examples, the initialization segment (632) incorporates metadata related to the video component. This metadata primarily describes MPI packing information and is placed under the bin 'resv'. MPI-encoded video invokes the decoder post-mechanism. Therefore, an MPI video component track can be represented as a restricted video in a file and can use the generic restricted sample entry 'resv'. The restricted scheme information bin 'rinf' can contain a raw format bin (frma) to record the raw sample entry type and a scheme type bin (schm). In some examples, a scheme information bin (schi) can also be used depending on the restricted scheme.

[0175] In some examples, two MPI video placement methods (2400, 2600) are implemented in the 'resv' box. These two methods are described in more detail below with reference to Figures 23-26.

[0176] In some examples, the MPI information SEI packing arrangement supports spatial packing of texture packing images and alpha map packing images into frames, or temporally interleaved texture packing images and alpha map packing images. For the temporally interleaved case, texture packing images and alpha map packing images can be placed on a single track or on different tracks. In some examples, the system (1300) is configured to use the mechanism described in ISO / IEC 14496-15 to signal MPI-related SEIs. For example, scheme type 'aSEI' can be used. Then, SEI information box 'seii' can be used under 'schi'.

[0177] Figures 23A and 23B provide example definitions related to the SEI information box 'seii' used in the communication system (1300) according to some embodiments. In 'seii', numRequiredSEI is greater than 0. A requiredSEI_ID is set to the value of the MPI information SEI message "payloadType".

[0178] Figure 24 illustrates Table 6, which provides examples of an MPI video arrangement method (2400) with MPI-related SEI messages used in a communication system (1300) according to some embodiments. For illustrative purposes and without any implied limitations, the method (2400) is represented in Extensible Markup Language (XML).

[0179] Figure 25 provides example definitions related to an MPI video arrangement (MPI video box) used in a communication system (1300) according to some embodiments. These definitions are compatible with the design of stereoscopic video arrangement schemes in ISO / IEC 14496-12, section 8.1.1. It should be noted that these definitions are intended to keep the MPI video box relatively concise, for example, containing a concise set of information. In the example shown, only mpi_indication_type (MPI indication type) is used as the highest level information. In various examples, the number of MPI layers and whether there is a constant depth distance between MPI layers can be added to the MPIVideoBox. For other details, the MPI information SEI message can be used. Also note that in some examples, there can be two 'resv' entries under 'stsd', for example, one for the MPI information SEI and another for 'mpiv'.

[0180] Figure 26 Table 7 is shown, which provides examples of an MPI video arrangement method (2600) using MPIVideoBox used in a communication system (1300) according to some embodiments. For illustrative purposes and without any implied limitations, the method (2400) is represented in Extensible Markup Language (XML).

[0181] Each of AVC, HEVC, and VSEI provides the use of Multi-View Acquisition Information SEI (MAI SEI) messages, which contain intrinsic and extrinsic parameters for perspective projection. In some examples, the system (1300) is also configured to reuse MAI SEI messages to describe camera information. Therefore, Table 6 (Figure 24) can be modified as follows: remove the 'icam' and 'ecam' entries, change the value of numRequiredSEI (number of required SEIs) to 2, and add requiredSEI_ID (required SEI ID) for the MAE SEI payload type. Comparing the two schemes above, using the 'icam' and 'ecam' bins may prove more general, for example, because those bins are at a higher level and do not require parsing the bitstream to obtain camera information.

[0182] The following description relates to the organization of information in the initialization section (632) of some examples. More specifically, the above information can be organized in different adapter sets using several methods.

[0183] ● Single-view (local) information: Camera information and MPI-related information for each view are stored in the initialization segment (632) of its corresponding view. In this case, in each cycle (610), the client (1306) operates to request N initialization segments from the server (1302) to obtain information from all relevant cameras.

[0184] ● Multi-view (global) information: Camera information and MPI-related information for all views are stored in an initialization segment (632) for one view. In this case, the client (1306) operates to request an initialization segment for each cycle to obtain information for those N cameras.

[0185] In the case of single-view (local) information, each view contains its own 'icam' and 'ecam'. The client (1306) operates to request all initialization segments from all views in the MPI container (1400). In the case of multi-view (global) information, the "number" of 'icam' and 'ecam' can be 0 or more, depending on ISO-BMFF. Therefore, multiple 'icam' and 'ecam' bins can be used to carry multiple camera information. For example, in bin 'vwid', we specify the syntax num_view (number of views) and view_id (view ID) information. In bins 'ecam' and 'icam', the syntax 'ref_view_id (reference view ID)' is specified to map 'ecam' and 'icam' to the corresponding syntax 'view_id' in bin 'vwid'. All their values ​​are equal to mpi_view_id (MPI view ID) in the MPI information SEI message for the same camera view.

[0186] Figure 27 illustrates Table 8, which provides an example of multi-view camera settings in a communication system (1300) according to some embodiments. For illustrative purposes and without any implied limitations, the example is represented in Extensible Markup Language (XML).

[0187] Media segment bitstream

[0188] In various examples, the media segment (634) contains one or more MPI bitstreams and corresponding MPI SEI information. In the case of video and / or audio bitstreams, the media segment (634) contains video samples (bitstreams), audio samples (bitstreams), etc. The MPI-related SEI is carried in the video bitstream. In some examples, the MPI SEI information provides certain MPI information, such as the number of MPI layers, the packing of texture and alpha data for the MPI layers, and the depth value for each layer. Furthermore, it specifies the mpi_view_id corresponding to a multi-camera setup.

[0189] Stream prefetching

[0190] In some cases, a prefetch strategy is employed for streaming to reduce latency. Therefore, it is important for the client (1306) to take action to identify the viewpoints to be prefetched for the prefetch strategy to be implemented. ISO / IEC 14496-15 provides a multiview group box ('mvcg') contained within a multiview information box ('mvci', contained within 'minf'). In some examples, the system (1300) can be configured to use 'mvcg' to group adjacent viewpoints. In one example, for each viewpoint, the system (1300) can indicate the current viewpoint and eight adjacent viewpoints as a multiview group box 'mvcg'. The box 'mvcg' is contained within the 'mvci' (multiview information box), and the 'mvci' is contained within the 'minf' (media information box).

[0191] The 'mvcg' bin was originally designed to specify multi-view groups for the views of the output MVC and MVD streams. The target view can be indicated based on the track_id (track ID) within the multi-view group bin (e.g., entry type equals 0). When using multi-view sample grouping, and the hierarchy covers more than one view or some hierarchies contain temporal subsets of the bitstream, it is recommended to use the tier_id (hierarchy ID) within the multi-view group bin (i.e., entry type equals 1). Otherwise, it is recommended to use one of the indications based on view_id (view ID) (i.e., entry type equals 2 or 3).

[0192] Figures 28A-28CExample definitions related to prefetching operations used in a communication system (1300) according to some embodiments are provided. In the illustrated example, 'mvcg' is reused for the purposes described above. More specifically, the syntax output_view_id (output view ID) is specified using an entry type equal to 2. output_view_id corresponds to a neighboring camera view_id of the current camera view. For example, to specify 8 neighboring views, the multi-view group ID can be set to the current camera view_id, and num_entry can be set to 8.

[0193] Storage format

[0194] In some examples, for streaming (DASH) scenarios, the MPI for each view is stored in a different corresponding ISOBMFF file. Figure 29 The block diagram (2900) illustrates the encapsulation of multi-view MPI encoded data in a single ISO BMFF for storage, according to some embodiments. In some examples, for storage purposes (as opposed to streaming), to encapsulate multi-view MPI encoded data in a single ISO BMFF file, the communication system (1300) is configured to place each view's MPI encoded data into a separate track (2902-2908), indicated by the syntax track_ID (track ID) under 'tkhd' contained in 'trak', which is contained in a 'moov' bin. This approach provides significant flexibility in rendering new perspectives both spatially and temporally.

[0195] Example hardware

[0196] Figure 30 is a block diagram illustrating a computing device (3000) used in a communication system (1300) according to one embodiment. The device (3000) can be used, for example, as a server (1302) or a client (1306). The computing device (3000) includes an input / output (I / O) device (3010), a processing engine (3020), and a memory (3030). The I / O device (3010) can be used to enable the device (3000) to receive various input signals (3002) and output various output signals (3004). For example, the I / O device (3010) can be operatively connected to send and receive signals via a communication channel (1304).

[0197] The memory (3030) may have a buffer for receiving data. Once data is received, the memory (3030) may provide a portion of the data to the processing engine (3020) for processing. The processing engine (3020) includes a processor (3022) and a memory (3024). The memory (3024) may store program code therein, which, when executed by the processor (3022), enables the processing engine (3020) to perform various data processing operations, including but not limited to at least some of the operations of the MPI method described above.

[0198] According to the exemplary embodiments disclosed above, such as in the Summary of the Invention and / or with reference to any one or all of Figures 1-30, an apparatus for MPI video streaming is provided, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and program code are configured to combine with the at least one processor such that the apparatus at least: provides a client device with a media rendering description of MPI stream content stored in a storage container accessible by a server device; provides, for a period, a corresponding initialization segment from the storage container to the client device, the corresponding initialization segment being configured to inform the client device of the selection of one or more views of the MPI stream content, wherein media segments for rendering are requested for the one or more views; receives a request from the client device, the request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from the group consisting of bitrate, resolution, codec type, and frame rate; and transmits to the client device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended parameter values. As used herein, the term “view” should be interpreted to encompass both captured view and new view. In some examples, new viewpoints can be generated for the MPI stream based on the captured viewpoint (e.g., via a NeRF3D model). One benefit of using new viewpoint locations in at least some examples of the MPI stream is that it makes the camera pose more uniform, which, for example, helps to perform more efficient multi-MPI rendering at the decoder.

[0199] According to another exemplary embodiment disclosed above, such as in the Summary of the Invention section and / or with reference to any one or any combination of some or all of Figures 1-30, a server-implemented MPI video streaming method is provided, comprising: providing a client device with a media rendering description of MPI stream content stored in a storage container accessible by a server device; for a period of time, providing the client device with a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, wherein a media segment for rendering is requested for the one or more views; receiving a request from the client device, the request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and transmitting to the client device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0200] In some embodiments of the method implemented by the server described above, for the period, the storage container has multiple media segments logically organized according to different perspectives and further logically organized according to one or more of different bit rates, different resolutions, different codec types, and different frame rates.

[0201] In some embodiments of the methods implemented by any of the servers described above, for each of the different capture perspectives and for each of the different bitrates, the storage container has a corresponding sequence of media segments corresponding to the different corresponding video segment times.

[0202] In some embodiments of the method implemented by any of the servers described above, the method further includes: switching from a first response sequence of media segments to a second response sequence of different media segments when the request indicates a change in the identified selection or a change in the recommended bit rate.

[0203] In some embodiments of the method implemented by any of the above servers, the transmission includes: transmitting a first bitstream carrying a first corresponding sequence of media segments corresponding to a first viewpoint in different captured viewpoints; and transmitting a second bitstream carrying a second corresponding sequence of different media segments corresponding to a second viewpoint in different captured viewpoints, wherein the first bitstream and the second bitstream have corresponding media segments corresponding to the same video segment time in different corresponding video segment time.

[0204] In some embodiments of the method implemented by any of the above servers, the media presentation description includes one or more sets of information selected from the group consisting of: a set of captured viewpoint information; a set of bitrate information; a set of video resolution information; a set of frame rate information; a set of audio language information; and a set of information identifying one or more Uniform Resource Locator addresses of an initialization segment or media segment stored in the storage container.

[0205] In some embodiments of the method implemented by any of the servers described above, the media presentation description includes information corresponding to two or more cycles of the MPI stream content.

[0206] In some embodiments of the method implemented by any of the servers described above, the media presentation description includes information corresponding to two or more different captured perspectives.

[0207] In some embodiments of the method implemented by any of the above servers, the corresponding initialization segment includes camera information and metadata related to one or more components of the media segment (e.g., video, audio, subtitles, etc.).

[0208] In some embodiments of the method implemented by any of the above servers, the corresponding initialization segment has a box format compatible with at least one of the MPEGDASH specification and ISO BMFF specification and / or their extended formats, such as CMAF (using ISO BMFF to define containers), etc.

[0209] In some embodiments of the method implemented by any of the above servers, the corresponding initialization segment contains multi-view information or single-view information (e.g., camera information or viewpoint rendered from a model such as Neural Radiation Field (NeRF)).

[0210] In some embodiments of the method implemented by any of the servers described above, the method further includes transmitting a Supplemental Enhancement Information (SEI) message or MPI video box that specifies a packing arrangement of the constituent frames, the packing arrangement being selected from: a first packing arrangement comprising one or more texture-packed and alpha-map-packed images; and a second packing arrangement comprising texture-packed images and alpha-map-packed images that are time-interleaved in a single track or in two or more different tracks.

[0211] In some embodiments of the method implemented by any of the servers described above, the method further includes, for the period, providing a client device with two or more corresponding initialization segments from the storage container, each of the two or more corresponding initialization segments corresponding to a different corresponding view of the MPI stream content.

[0212] In some embodiments of the method implemented by any of the servers described above, the one or more bitstreams include: a sequence of media segments containing at least video and audio samples; and supplemental enhancement information for one or more MPI parameters specifying the MPI stream content.

[0213] In some embodiments of the method implemented by any of the servers described above, the method further includes providing a view identifier box, the view identifier box including one or more of the following: an indication of views contained in a track or hierarchy; an indication of the view order index for each listed view; an indication of the minimum and maximum values ​​of the time parameters for the track or hierarchy; an indication of a reference view for decoding views contained in the track or hierarchy; and an indication of whether texture or depth exists in the track for each contained view.

[0214] In some embodiments of the method implemented by any of the servers described above, the method further includes providing a view identifier box or another box that includes one or both of the following: an indication of a view contained in a track or hierarchy; and an indication of a view identifier for each listed view.

[0215] In some embodiments of the method implemented on any of the servers described above, the method further includes providing an MPI packing information box using restricted sample entries as defined in the ISO BMFF specification.

[0216] In some embodiments of the method implemented by any of the servers described above, the method further includes providing an MPI video box to indicate whether the decoded frame contains a representation of two spatially packed frames or a representation of two temporally interleaved constituent frames that form an MPI representation.

[0217] In some embodiments of the method implemented on any of the servers described above, the method further includes providing a multi-view box indicating the current view and a fixed number of adjacent views.

[0218] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a server device, cause the server device to perform operations including any of the server implementations described above for MPI video streaming methods.

[0219] According to yet another exemplary embodiment disclosed above, such as in the summary section and / or references Figure 1-30Any one or more of the following, or any combination thereof, provides an apparatus for MPI video streaming transmission, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured, via the at least one processor, such that the apparatus performs at least the following operations: receiving at a client device a media rendering description of MPI stream content stored in a storage container accessible via a server device; for one cycle, receiving at the client device, via the server device, a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, requesting media segments for rendering for the one or more views; transmitting to the server device a request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and receiving from the server device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0220] According to another exemplary embodiment disclosed above, such as in the Summary of the Invention section and / or with reference to any one or all of Figures 1-30, a client-implemented MPI video streaming method is provided, comprising: receiving at a client device a media rendering description of MPI stream content stored in a storage container accessible by a server device; for one cycle, receiving at the client device, via the server device, a corresponding initialization segment from the storage container, the corresponding initialization segment being configured to inform the client device of a selection of one or more views of the MPI stream content, for which a media segment is requested for rendering; transmitting to the server device a request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from a group consisting of bitrate, resolution, codec type, and frame rate; and receiving from the server device one or more bitstreams carrying media segments selected in the storage container based on the identified selection and further based on one or more of the corresponding recommended values.

[0221] In some embodiments of the client-side implementation of the above method, for the period, the storage container has multiple media segments logically organized according to different perspectives and further logically organized according to one or more of different bit rates, different resolutions, different codec types, and different frame rates.

[0222] In some embodiments of the method implemented by any of the clients described above, for each different captured viewpoint and each different bitrate, the storage container has a corresponding sequence of media segments corresponding to different corresponding video time segments.

[0223] In some embodiments of the method implemented by any of the clients described above, the method further includes requesting a change in the identified selection or a change in the recommended bit rate to cause the server device to switch from a first corresponding media segment sequence to a different second corresponding media segment sequence.

[0224] In some embodiments of the method implemented by any of the above clients, the method further includes receiving a first bitstream carrying a first corresponding sequence of media segments corresponding to a first viewpoint in different viewpoints; and receiving a second bitstream carrying a second corresponding sequence of different media segments corresponding to a second viewpoint in different viewpoints, wherein the first bitstream and the second bitstream have corresponding media segments corresponding to the same video segment time in the different corresponding video segment time.

[0225] In some embodiments of the method implemented by any of the clients described above, the media presentation description includes one or more sets of information selected from the group consisting of: a set of captured viewpoint information; a set of bitrate information; a set of video resolution information; a set of frame rate information; a set of audio language information; and a set of information identifying one or more Uniform Resource Locator addresses of an initialization segment or media segment stored in a storage container.

[0226] In some embodiments of the method implemented by any of the clients described above, the media presentation description includes information corresponding to two or more cycles of the MPI stream content.

[0227] In some embodiments of the method implemented in any of the above clients, the media presentation description includes information corresponding to two or more different captured perspectives.

[0228] In some embodiments of the method implemented by any of the clients described above, the corresponding initialization segment includes camera information and metadata related to one or more components of the media segment.

[0229] In some embodiments of the methods implemented by any of the clients described above, the corresponding initialization segment has a box format compatible with at least one of the MPEG DASH, ISO BMFF, and CMAF specifications.

[0230] In some embodiments of the method implemented by any of the above clients, the corresponding initialization segment contains multi-view information or single-view information.

[0231] In some embodiments of the method implemented by any of the clients described above, the method further includes receiving a Supplemental Enhancement Information (SEI) message or MPI video box specifying a packing arrangement of constituent frames, the packing arrangement being selected from: a first packing arrangement comprising one or more texture-packed and alpha-map-packed images; and a second packing arrangement comprising texture-packed images and alpha-map-packed images that are time-interleaved in a single track or in two or more different tracks.

[0232] In some embodiments of the method implemented by any of the clients described above, the method further includes, for the period, receiving two or more corresponding initialization segments from a storage container via a server device, each of the two or more corresponding initialization segments corresponding to a different corresponding viewpoint of the MPI stream content.

[0233] In some embodiments of the method implemented by any of the clients described above, one or more bitstreams include: a sequence of media segments containing at least video and audio samples; and supplementary enhancement information for one or more MPI parameters specifying the MPI stream content.

[0234] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a client device, cause the client device to perform operations including any of the client implementations described above for MPI video streaming methods.

[0235] Regarding the processes, systems, methods, heuristics, etc., described herein, it should be understood that although the steps of such processes, etc., are described as occurring in a specific ordered sequence, such processes can be implemented by performing the described steps in a different order than that described herein. It should also be understood that some steps may be performed simultaneously, other steps may be added, or some steps described herein may be omitted. In other words, the process descriptions herein are for illustrative purposes and should in no way be construed as limiting the claims.

[0236] Therefore, it should be understood that the above description is intended to be illustrative rather than restrictive. Many embodiments and applications, in addition to the examples provided, will become apparent upon reading the above description. The scope should not be determined by reference to the above description, but rather by reference to the appended claims and the full scope of their equivalents. Future developments are anticipated and intended for the techniques discussed herein, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that modifications and variations are possible with this application.

[0237] All terms used in the claims are intended to be interpreted in the broadest reasonable sense and as commonly understood by those skilled in the art, unless expressly indicated otherwise herein. In particular, singular articles such as “a,” “the,” and “the” should be understood to refer to one or more of the indicated elements, unless expressly limited to the contrary in the claims.

[0238] This abstract is provided to allow the reader a quick understanding of the nature of the technical disclosure. The abstract is submitted on the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, in the preceding detailed description section, various features are grouped in various embodiments to simplify the disclosure. This method of disclosure should not be construed as reflecting an intention that the claimed embodiments include more features than expressly recited in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in fewer than all the features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the detailed description section, each claim being a separate claimed subject matter.

[0239] Although this disclosure contains references to illustrative embodiments, this specification is not intended to be construed as limiting. Various modifications to the described embodiments and other embodiments within the scope of this disclosure will be apparent to those skilled in the art, as set forth in the following claims, and such modifications and embodiments are considered to fall within the principles and scope of this disclosure.

[0240] Some embodiments can be implemented as a circuit-based process, including possibly on a single integrated circuit.

[0241] Some embodiments may be embodied in the form of methods and means for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium, such as a magnetic recording medium, optical recording medium, solid-state storage, floppy disk, CD-ROM, hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into and executed by a machine (e.g., a computer), the machine becomes means for practicing the licensed invention. Some embodiments may also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium, including being loaded into and / or executed by a machine, wherein when the program code is loaded into and executed by a machine (e.g., a computer or processor), the machine becomes means for practicing the licensed invention. When implemented on a general-purpose processor, program code segments are combined with the processor to provide unique means whose operation resembles a specific logic circuit.

[0242] Unless explicitly stated otherwise, each value and range should be interpreted as an approximation, as if preceded by the words “approximately” or “about”.

[0243] The use of figure numbers and / or figure reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to aid in the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding figures.

[0244] Although the elements in the following method claims (if any) are described in a particular order and with corresponding markings, these elements are not necessarily intended to be limited to being implemented in that particular order unless the description of the claims otherwise implies a particular order for implementing some or all of these elements.

[0245] The reference to "an embodiment" or "an embodiment" in this document means that a particular feature, structure, or characteristic described in association with that embodiment may be included in at least one embodiment of this disclosure. The phrase "in an embodiment" appearing in different places in the specification does not necessarily refer to the same embodiment, nor does it mean that a single or alternative embodiment is necessarily mutually exclusive with other embodiments. The same applies to the term "implementation".

[0246] Unless otherwise specified herein, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to refer to one of a plurality of similar objects merely indicates a different instance of such similar object being referred to, and is not intended to imply that the similar objects referred to must be in a corresponding order or sequence, whether temporally, spatially, in rank, or in any other way.

[0247] Unless otherwise specified herein, the conjunction “if” can also be interpreted, or alternatively, as “when”, “in response to determination”, or “in response to detection”, in addition to its ordinary meaning, depending on the specific context. For example, the phrase “if determination” or “if detection [the condition]” can be interpreted as “when determination”, “in response to determination”, “when [the condition or event] is detected”, or “in response to detection [the condition or event]”.

[0248] Similarly, for the purposes of this description, the terms “coupled,” “coupled,” “connected,” and “linked” refer to any manner known in the art or subsequently developed, in which energy is allowed to be transferred between two or more elements, and the insertion of one or more additional elements may be considered, though not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0249] As used in the referenced elements and standards herein, the term "compatible" means that the element communicates with other elements in a manner fully or partially specified by the standard, and will be recognized by other elements as sufficient to communicate with other elements in the manner specified by the standard. A compatible element is not required to operate internally in the manner specified by the standard.

[0250] The functionality of the various elements shown in the accompanying figures, including any functional blocks labeled “processor” and / or “controller,” can be provided using dedicated hardware and hardware capable of executing software in conjunction with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, a single shared processor, or multiple independent processors (some of which may be shared). Furthermore, the explicit use of the terms “processor” or “controller” should not be construed as referring only to hardware capable of executing software, and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile memory. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the accompanying figures are conceptual only. Their functionality can be performed through the operation of program logic, the operation of dedicated logic, the interaction of program control and dedicated logic, or even manually, the specific techniques of which may be chosen by the implementer based on a more specific understanding of the context.

[0251] As used herein, the terms “circuit” and “circuit system” may refer to one or more or all of the following: (a) a hardware circuit implementation only (e.g., an implementation only in analog and / or digital circuitry); (b) a combination of hardware circuitry and software, such as (if applicable): (i) a combination of analog and / or digital hardware circuitry with software / firmware; and (ii) any portion of a hardware processor (including a digital signal processor) combined with software, software, and memory, which work together to enable a device (e.g., a mobile phone or a server) to perform various functions; and (c) hardware circuitry and / or a processor (e.g., a microprocessor or a portion thereof) that requires software (e.g., firmware) to operate, but the software may be absent when operation is not required. This definition of “circuit” applies to all uses of the term in this application, including in any claim. As another example, as used herein, the term “circuit system” also encompasses an implementation of only hardware circuitry or a processor (or multiple processors) or a portion thereof and its accompanying software and / or firmware. The term "circuit system" also covers, for example and if applicable to certain claim elements, baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices or other computing or networking devices.

[0252] Those skilled in the art will understand that any block diagram herein represents a conceptual view of an illustrative circuit embodying the principles of this disclosure. Similarly, it should be understood that any flowchart, state transition diagram, pseudocode, etc., represents various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, whether or not such a computer or processor is explicitly shown.

[0253] The "Brief Overview of Some Specific Embodiments" in this specification is intended to describe some exemplary embodiments, with additional embodiments described in the "Detailed Description" and / or with reference to one or more accompanying drawings. The "Brief Overview of Some Specific Embodiments" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

1. A method for transmitting multi-plane image (MPI) video streams, comprising: Provides client devices with a media presentation description of MPI streaming content stored in a storage container accessible through server devices; For each cycle, a corresponding initialization segment from the storage container is provided to the client device, the corresponding initialization segment being configured to inform the client device of the selection of one or more views of the MPI stream content, wherein a media segment for rendering is requested for the one or more views; A request is received from a client device, the request identifying the selection and indicating a recommended value for at least one parameter selected from the group consisting of bit rate, resolution, codec type, and frame rate; as well as Transmit to a client device one or more bitstreams carrying media segments selected in a storage container based on the identified selection and further based on one or more of the corresponding recommended values.

2. The method according to claim 1, wherein, For the period, the storage container has multiple media segments logically organized according to different perspectives and further logically organized according to one or more of different bit rates, different resolutions, different codec types, and different frame rates.

3. The method according to claim 2, wherein, For each of the different viewpoints and for each of the different bitrates, the storage container has a corresponding media segment sequence corresponding to the different corresponding video time segments.

4. The method of claim 3, further comprising: When the request indicates a change in the identified selection or a change in the corresponding recommended value, the system switches from the first corresponding sequence of media segments to the second corresponding sequence of different media segments.

5. The method of claim 3, wherein the transmission comprises: Transmit a first bitstream carrying a first corresponding sequence of media segments corresponding to a first viewpoint in the different viewpoints; and Transmit a second bitstream carrying a second corresponding sequence of media segments corresponding to the second perspective in the different perspectives. The first bitstream and the second bitstream have corresponding media segments corresponding to the same video segment time in the different corresponding video segment time.

6. The method according to claim 1, wherein, The media presentation description includes one or more sets of information selected from a group consisting of: Perspective information set; Bit rate information set; Video resolution information set; Frame rate information set; Audio language information set; as well as, A set of information that identifies one or more Uniform Resource Locator addresses of an initialization segment or media segment stored in the storage container.

7. The method according to claim 1, wherein, The media presentation description includes information corresponding to two or more cycles of the MPI streaming content.

8. The method according to claim 1, wherein, The media presentation description includes information corresponding to two or more different perspectives.

9. The method of claim 1, wherein the corresponding initialization segment includes metadata and camera information related to one or more components of the media segment.

10. The method of claim 1, wherein the corresponding initialization segment has a box format compatible with at least one of the MPEG DASH, ISOBMFF, and CMAF specifications.

11. The method according to claim 1, wherein the corresponding initialization segment includes multi-view information or single-view information.

12. The method according to claim 1, further comprising: The transmission specifies a Supplemental Enhancement Information (SEI) message or MPI video box that defines the packet arrangement of the constituent frames, the packet arrangement being selected from: A first packing arrangement, comprising one or more texture-packed and alpha-map-packed images; as well as The second packing arrangement includes texture packing images and alpha packing images that are staggered in time in a single track or in two or more different tracks.

13. The method according to claim 1, further comprising: For the period, the client device is provided with two or more of the corresponding initialization segments from the storage container, each of the two or more corresponding initialization segments corresponding to a different corresponding view of the MPI stream content.

14. The method of claim 1, wherein the one or more bit streams comprise: A sequence of media segments containing at least video and audio samples; and Additional enhancements to one or more MPI parameters that specify the MPI stream content.

15. The method according to claim 1, further comprising: Provide a viewpoint identifier box or another box, which includes one or both of the following: Indications of views including those within the orbit or layer; and An indication of the viewpoint identifier for each listed viewpoint.

16. The method according to claim 1, further comprising: The MPI packing information box is provided using restricted sample entries defined in the ISO BMFF specification.

17. The method of claim 1, further comprising providing an MPI video box to indicate whether the decoded frame is a representation containing two spatially packed frames or a representation containing two temporally interleaved constituent frames forming an MPI representation.

18. The method of claim 1, further comprising providing a multi-view box indicating the current view and a fixed number of adjacent viewpoints.

19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a server device, cause the server device to perform operations including the method according to any one of claims 1-18.

20. An apparatus for transmitting multi-plane image (MPI) video streams, the apparatus comprising: At least one processor; as well as At least one memory containing program code; and The at least one memory and program code are configured, via the at least one processor, to cause the device to perform at least the following operations: Provides client devices with a media presentation description of MPI streaming content stored in a storage container accessible through server devices; For each cycle, a corresponding initialization segment from the storage container is provided to the client device, the corresponding initialization segment being configured to inform the client device of the selection of one or more views of the MPI stream content, wherein a media segment for rendering is requested for the one or more views; A request is received from a client device, the request identifying the selection and indicating a recommended value for at least one parameter selected from the group consisting of bit rate, resolution, codec type, and frame rate; as well as Transmit to a client device one or more bitstreams carrying media segments selected in a storage container based on the identified selection and further based on one or more of the corresponding recommended values.

21. A method for transmitting a multi-plane image (MPI) video stream, comprising: A media presentation description that receives MPI stream content stored in a storage container accessible through a server device at the client device. For one cycle, the client device receives a corresponding initialization segment from the storage container via the server device. This corresponding initialization segment is configured to inform the client device of the selection of one or more views of the MPI stream content and to request media segments for rendering for the one or more views. A request is transmitted to the server device, which identifies the selection and indicates a recommended value for at least one parameter selected from the group consisting of bit rate, resolution, codec type, and frame rate; as well as Receive one or more bitstreams from a server device, the bitstreams carrying media segments selected in a storage container based on an identified selection and further based on corresponding recommended values.

22. The method according to claim 21, wherein, For the period, the storage container has multiple media segments logically organized according to different perspectives and further logically organized according to one or more of different bit rates, different resolutions, different codec types, and different frame rates.

23. The method according to claim 22, wherein, For each of the different viewpoints and for each of the different bitrates, the storage container has a corresponding media segment sequence corresponding to the different corresponding video time segments.

24. The method of claim 23, further comprising: The request is to change the identified selection or the recommended bit rate so that the server device switches from a first response sequence of a media segment to a second response sequence of a different media segment.

25. The method of claim 23, further comprising: Receive a first bitstream carrying a first corresponding sequence of media segments corresponding to a first viewpoint among the different viewpoints; as well as Receive a second bitstream carrying a second corresponding sequence of media segments corresponding to the second perspective in the different perspectives. The first bitstream and the second bitstream have corresponding media segments corresponding to the same video segment time in the different corresponding video segment time.

26. The method according to claim 21, wherein, The media presentation description includes one or more sets of information selected from a group consisting of: Perspective information set; Bit rate information set; Video resolution information set; Frame rate information set; Audio language information set; as well as, A set of information that identifies one or more Uniform Resource Locator addresses of an initialization segment or media segment stored in the storage container.

27. The method according to claim 21, wherein, The media presentation description includes information corresponding to two or more cycles of the MPI streaming content.

28. The method according to claim 21, wherein, The media presentation description includes information corresponding to two or more different perspectives.

29. The method of claim 21, wherein the corresponding initialization segment includes metadata and camera information related to one or more components of the media segment.

30. The method of claim 21, wherein the corresponding initialization segment has a box format compatible with at least one of the MPEG DASH, ISOBMFF, and CMAF specifications.

31. The method of claim 21, wherein the corresponding initialization segment includes multi-view information or single-view information.

32. The method of claim 21, further comprising: Receive a Supplemental Enhancement Information (SEI) message or an MPI video box specifying the packet arrangement of the constituent frames, wherein the packet arrangement is selected from: A first packing arrangement, comprising one or more texture-packed and alpha-map-packed images; as well as The second packing arrangement includes texture packing images and alpha packing images that are staggered in time in a single track or in two or more different tracks.

33. The method of claim 21, further comprising: For the period, two or more of the corresponding initialization segments are received from the storage container via the server device, each of the two or more corresponding initialization segments corresponding to a different corresponding viewpoint of the MPI stream content.

34. The method of claim 21, wherein the one or more bit streams comprise: A sequence of media segments containing at least video and audio samples; and Additional enhancements to one or more MPI parameters that specify the MPI stream content.

35. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a client device, cause the client device to perform operations including the method according to any one of claims 21-34.

36. An apparatus for transmitting multi-plane image (MPI) video streams, the apparatus comprising: At least one processor; as well as At least one memory containing program code; The at least one memory and program code are configured, via the at least one processor, to cause the device to perform at least the following operations: A media presentation description that receives MPI stream content stored in a storage container accessible through a server device at the client device. For one cycle, the client device receives a corresponding initialization segment from the storage container via the server device. This corresponding initialization segment is configured to inform the client device of the selection of one or more views of the MPI stream content and to request media segments for rendering for the one or more views. A request is transmitted to the server device, which identifies the selection and indicates a recommended value for at least one parameter selected from the group consisting of bit rate, resolution, codec type, and frame rate; as well as Receive one or more bitstreams from a server device, the bitstreams carrying media segments selected in a storage container based on an identified selection and further based on corresponding recommended values.