Multi-view multi-plane imaging video streaming

By using ISO BMFF and MPEG DASH in combination with the MPEG immersive video coding standard, the problem of seamless interactive view switching in video streaming of multi-view multi-plane imaging technology was solved, realizing a fast and cost-effective immersive volumetric video experience.

CN121569487APending Publication Date: 2026-02-24DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480048958.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-22
Filing Date
2024-06-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing multi-view, multi-plane imaging technologies lack seamless interactive view switching in video streaming, making it difficult to achieve a fast and cost-effective immersive volumetric video experience.

Method used

MPI video is encoded, stored, and transmitted using the ISO Basic Media File Format (BMFF) and MPEG Dynamic Adaptive Streaming over HTTP (DASH). Combined with the MPEG Immersive Video Coding Standard, seamless interactive view switching is achieved by defining new tools.

Benefits of technology

It enables seamless interactive view switching for multi-view MPI videos, supporting a fast and cost-effective immersive volumetric video experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569487A_ABST
    Figure CN121569487A_ABST
Patent Text Reader

Abstract

Methods and apparatus for multi-plane imaging (MPI) video streaming. According to an example embodiment, a method for streaming MPI video includes generating a sequence of video frames, each of the video frames including a respective plurality of tiles representing a texture layer and a transparency layer of one or more multi-plane images of the MPI video; and applying video compression to the sequence of video frames to generate a video substream. The method further includes generating a representation sequence of atlas frames corresponding to the sequence of video frames to specify at least a packed arrangement of tiles, and applying a compression to the representation sequence to generate an atlas substream. The method further includes multiplexing the video substream and the atlas substream to generate a first encoded bitstream encoding at least a portion of the MPI video.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1. Cross-references to related applications This patent application claims priority to the following applications: U.S. Provisional Patent Application No. 63 / 556,461, filed February 22, 2024; U.S. Provisional Patent Application No. 63 / 588,337, filed October 6, 2023; and U.S. Provisional Patent Application No. 63 / 510,571, filed June 27, 2023, each of which is incorporated herein by reference in its entirety. 2. Technical Field Various example embodiments generally relate to multiplanar imaging (MPI), and more specifically, but not exclusively, to the transmission of multiplanar images. 3. Background Technology Multi-plane imaging embodies a relatively new approach for storing volumetric content. MPI can be used to render both still images and videos, and uses, for example, 8, 16, or 32 texture and transparency (alpha) information planes per camera to represent a three-dimensional (3D) scene within the view frustum. Example applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. Summary of the Invention

[0004] The example embodiments disclosed herein provide formats for encoding, storing, and transmitting multi-view MPI video, as well as corresponding MPI video systems. Some examples use the ISO Basic Media File Format (BMFF) for storage and MPEG-based HTTP Dynamic Adaptive Streaming (DASH) for MPI video streaming to provide an immersive volumetric video experience. Some examples use modifications to available technologies and infrastructure (e.g., by defining new tools designed to fill the technological gaps described above) to achieve seamless interactive view switching for multi-view MPI video streaming. At least some of the disclosed solutions can be deployed in a relatively short time after modifications to relevant existing solutions are implemented according to the various embodiments disclosed herein.

[0005] According to an example embodiment, a method for streaming MPI video is provided, the method comprising: generating a sequence of video frames, each video frame including corresponding plurality of tiles of texture layers and transparency layers representing one or more multiplanar images of the MPI video; applying video compression to the video frame sequence to generate a video substream; generating a representation sequence of atlas frames corresponding to the video frame sequence to at least specify a packing arrangement of tiles; applying compression to the representation sequence to generate an atlas substream; and multiplexing the video substream and the atlas substream to generate a first coded bitstream that encodes at least a portion of the MPI video.

[0006] According to another example embodiment, a non-transitory computer-readable medium is provided that stores instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the methods described above.

[0007] According to yet another example embodiment, an apparatus for streaming MPI video is provided, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured, together with the at least one processor, to cause the apparatus to at least: generate a sequence of video frames, each video frame including a plurality of tiles of texture layers and transparency layers representing one or more multiplanar images of MPI video; apply video compression to the video frame sequence to generate a video substream; generate a representation sequence of atlas frames corresponding to the video frame sequence to at least specify a packing arrangement of tiles; apply compression to the representation sequence to generate an atlas substream; and multiplex the video substream and the atlas substream to generate a first coded bitstream that encodes at least a portion of the MPI video. Attached Figure Description

[0008] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent by way of example, in conjunction with the following detailed description and accompanying drawings, in which: Figure 1 An example process for a video / image transmission pipeline is described.

[0009] Figure 2 The accompanying image illustrates a 3D scene representation using multi-plane images according to an embodiment.

[0010] Figure 3 The image illustrates the process of generating a new view of a 3D scene based on an example.

[0011] Figure 4 It is a diagram illustrating how a set of active views changes over time, based on an example.

[0012] Figure 5 The illustration shows that, according to the embodiment, it is possible to... Figure 1 A block diagram of the MPI video system used in the transmission pipeline.

[0013] Figure 6 The illustration shows that, according to the embodiment, it is possible to... Figure 5 A block diagram of the MPI encoder used in an MPI video system.

[0014] Figure 7 The illustration shows a method according to an embodiment that can be derived from... Figure 6 A block diagram of an example structure of the output V3C bitstream generated by an MPI encoder.

[0015] Figures 8A to 8F The illustration shows a method according to some embodiments. Figure 5 A diagram illustrating an example of how the texture and transparency graph spaces of an MPI layer are packed into video frames in an MPI video system.

[0016] Figure 9 The illustration shows examples that can be found here. Figure 5 A block diagram of single-track encapsulation of V3C bitstream implemented in an MPI video system.

[0017] Figures 10A to 10B The illustration shows examples that can be found here. Figure 5 A block diagram of the multitrack encapsulation of V3C bitstreams implemented in an MPI video system.

[0018] Figure 11 It is illustrated with Figure 10A The example shown corresponds to a block diagram with further details of the multi-track package.

[0019] Figure 12 It is illustrated with Figure 10B The example shown corresponds to a block diagram with further details of the multi-track package.

[0020] Figure 13 The illustration shows examples that can be found here. Figure 5 A block diagram of a single-track encapsulation of a 2D video stream (MPI bitstream) implemented in an MPI video system.

[0021] Figure 14 The illustration shows examples that can be found here. Figure 5 A block diagram of the DASH configuration implemented in the MPI video system for grouping atlases and packed video components.

[0022] Figure 15 The illustrations show various examples that can be found Figure 5 The flowchart illustrates the client-server communication implemented in the MPI video system and the corresponding processing operations performed at the client.

[0023] Figure 16 The illustration shows examples that can be found here. Figure 5 A block diagram of the DASH configuration used in the MPI video system to support seamless switching between MPI views.

[0024] Figure 17 This is a block diagram of an example computing device, which can be used to implement one or more instances of the computing device according to various examples. Figure 5 MPI video system. Detailed Implementation

[0025] Example video / image delivery pipeline Figure 1 An example process of a video transmission pipeline (100) according to an embodiment is depicted, illustrating the various stages from video / image capture to video / image content display. An image generation block (105) can be used to capture or generate a sequence of video / image frames (102). Frames (102) can be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and / or image data (107). Alternatively, frames (102) can be captured on film by a film camera. The film can then be converted to a digital format to provide video / image data (107). In some examples, the image generation block (105) includes generating MPI images or videos.

[0026] During the production phase (110), data (107) can be edited to provide a video / image production stream (112). The data from the video / image production stream (112) can be provided to a processor (or one or more processors, such as a central processing unit CPU) at the post-production block (115) for post-production editing. Post-production editing at the block (115) can include, for example, adjusting or modifying the color or brightness of specific areas of an image to enhance image quality or achieve a specific look for the image according to the creative intent of the video creator. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other edits (e.g., scene selection and sorting, image cropping, adding computer-generated visual effects, removing artifacts, etc.) can be performed at the block (115) to produce a “final” version (117) of the work for release. In some examples, operations performed at the block (115) include enhancing the texture and / or alpha channel in a multi-planar image / video. During post-production editing (115), the video and / or image can be viewed on a reference monitor (125).

[0027] After post-production (115), the data of the final version (117) can be transmitted to the encoding block (120) for further downstream transmission to decoding and playback devices such as televisions, set-top boxes, and cinemas. In some embodiments, the encoding block (120) may include audio and video encoders (such as audio and video encoders defined by ATSC, DVB, DVD, Blu-ray, and other transmission formats) to generate an encoded bitstream (122). In the receiver, the encoded bitstream (122) is decoded by the decoding unit (130) to generate a corresponding decoded signal (132) representing a copy or a near-approximate version of the signal (117). The receiver may be attached to a target display (140), which may have slightly different or completely different characteristics from the reference display (125). In this case, the display management (DM) block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display mapping signal (137). According to an embodiment, the decoding unit (130) and the display management block (135) may include separate processors or may be based on a single integrated processing unit.

[0028] The codecs used in the encoding block (120) and / or decoding block (130) are capable of video / image data processing and compression / decompression. Compression is used in the encoding block (120) to reduce the size of the corresponding file(s) or stream(s). The decoding process performed by the decoding block (130) typically involves decompressing the received video / image data(s) or stream(s) into a form that can be played and / or further edited. Example encoding / decoding operations that can be used in the encoding block (120) and decoding unit (130) according to various embodiments will be described in more detail below.

[0029] Multiplanar imaging A multiplane image comprises multiple image planes, each of which is a "snapshot" of the 3D scene at a specific depth relative to the camera position. Information stored in each plane includes texture information (e.g., represented by R, G, and B values) and transparency information (e.g., represented by alpha (A) values). In this paper, the acronyms R, G, and B represent red, green, and blue, respectively. In some examples, the three texture components can be (Y, Cb, Cr), or (I, Ct, Cp), or another set of values ​​that are functionally similar. There are different ways to generate multiplane images. For example, two or more input images from two or more cameras located at different known viewpoints can be processed together to generate a corresponding multiplane image. Alternatively, single-view composition of multiplane images can be performed using source images captured by a single camera.

[0030] Figure 2The image illustrates a 3D scene representation using a multi-plane image (200) according to an embodiment. The multi-plane image (200) has D planes or layers (P0, P1, ..., P(D-1)), where D is an integer greater than one. Typically, the planes (layers) are indexed such that the layer furthest from the reference camera position (RCP) is indexed as layer 0, and maintains a certain distance (or depth) from the RCP along the Z dimension of the 3D scene. d 0. For each subsequent layer closer to the RCP, the index increments by one. The plane (layer) closest to the RCP has an index of (D-1) and maintains a certain distance (or depth) from the RCP along the Z dimension. d D-1 Each of the planes (P0, P1, ..., P(D-1)) is orthogonal to the base plane (202) parallel to the XZ coordinate plane. The RCP is located at a vertical height above the base plane (202). h Place. Figure 2 The XYZ triplet shown indicates the overall orientation of the multiplane image (200) and planes (P0, P1, ..., P(D-1)) relative to the X, Y, and Z dimensions of the 3D scene. In various examples, the number D can be 32, 16, 8, or any other suitable integer greater than one.

[0031] We will position the camera s First i The color component (e.g., RGB) values ​​of the layer are represented as The lateral dimension of this layer is H × W ,in, H The height of the layer (Y dimension), and W The width of the layer (X dimension). Location ( x , y (color channel) c The pixel value is represented as . No. i The α value of the layer is The pixel values ​​in the alpha layer ( x , y ) represents . No. i The depth distance between the layer and the reference camera position is The image from the original reference view (where the camera is not moving) is represented as Its texture pixel value Therefore, camera position s A static MPI image can be represented as: MPI ( s ) = { , }, i= 0, …, D -1(1) As long as the camera position s By keeping the image static over time, this static MPI image representation can be directly extended to a video representation. This video representation is given by equation (2): MPI ( s , t ) = { , }, i = 0, …, D -1(2) in, t Indicates time.

[0032] As indicated above, a single source image can be used. Single-view compositing or multi-view compositing using two or more source images can be used to generate multiplanar images, such as multiplanar images (200). This compositing can be performed, for example, during the production phase (110). The corresponding (multiple) MPI compositing algorithms can typically be implemented using {( , ),in i =0, …, D Output a multi-plane image (200) containing XYZ resolution pixel values ​​in the form of -1}.

[0033] By processing {( , ),in i =0, …, D The multiplanar image (200) represented by {-1} can be used by an MPI rendering algorithm to generate a visual image corresponding to the RCP or to a new virtual camera position different from the RCP. An example MPI rendering algorithm (often referred to as an "MPI viewer") that can be used for this purpose may include steps of warping and compositing. Other suitable MPI viewers may also be used. The rendered multiplanar image (200) can be viewed, for example, on a reference display (125).

[0034] During the warp step of the MPI rendering algorithm, each layer of the multi-plane image (200) ( , All can be viewed from the RCP viewpoint ( ) warped to a new viewpoint position ( For example, as shown below: (3) (4) in, For the distortion function; and σ For a consistent scaling (to minimize error). In the example embodiment, the twist function This can be expressed as follows: (5) in, and By using (5), the position of each pixel on the target view of a certain MPI plane can be determined. ) mapped to its corresponding pixel position on the source view ( ).function and Let represent the intrinsic camera models used for the reference view and the target view, respectively. Functions R and t represent the extrinsic camera models used for rotation and translation, respectively. n represents the normal vector [0 0 1]. T . a Indicates to depth The distance between the point and the plane parallel to the front of the source camera.

[0035] During the compositing step of the MPI rendering algorithm, new visual images can be generated, for example, using processing operations corresponding to the following equation. : (6) Among them, the weight Expressed as: (7) disparity map corresponding to the source view It can be calculated as: (8) Among them, the weight Expressed as: (9) The MPI rendering algorithm can also be used to generate visual images corresponding to RCP. In this case, the distortion step is omitted, and the image... Calculated as: (10) In single-camera transmission scenarios, only one MPI is transmitted via a bitstream. The goal in this case is to optimally merge the layers of the original MPI so that the quality of the MPI is maintained after local distortion. In multi-camera transmission scenarios, multiple MPIs captured at different camera positions are encoded into a compressed bitstream. Information from these MPIs is combined to generate a new global view of the location between the original camera positions. There are also scenarios where information from multiple cameras can be combined to generate a single MPI to be transmitted. Multi-camera transmission scenarios are commonly used for the transmission of MPI video, as explained below.

[0036] Figure 3 The process of generating a new view of a 3D scene (302) according to an example is illustrated in the image. In the example shown, the 3D scene (302) is captured using forty-two RCPs (1, 2, ..., 42). The new view being generated corresponds to the camera position (50). The four RCPs closest to the camera position (50) are RCPs (11, 12, 18, 19). The corresponding multiplane image is multiplane image (200). 11 200 12 200 18 200 19 ). By correspondingly applying multi-planar images (200 11 200 12 200 18 200 19 The images are distorted and then merged to generate a multi-plane image (200) corresponding to the camera position (50). 50 Finally, the synthesis step (310) of the MPI rendering algorithm is applied to the multi-plane image (200). 50 ) to generate a visual image (312) of the 3D scene (302).

[0037] Generally, any suitable number of RCPs can be used to capture a 3D scene, such as a 3D scene (302). The position of these RCPs can also be chosen in various ways, for example, based on creative intent. In a typical practical example, when rendering a new view, such as a visual image (312), only a few adjacent RCPs are used for rendering. These adjacent views are referred to as "active views" below. Figure 3In the example shown, the number of active views is four. In other examples, a different number (other than four) of active views can be used similarly. Thus, the number of active views is an optional parameter. For illustrative purposes and without any implied limitation, the example embodiments are described below with reference to four active views. In some examples, when the camera position (50) moves, the set of active views can change over time. In some examples, when the camera position (50) moves, the number of active views can change over time.

[0038] Figure 4 is a block diagram illustrating the change of a set of active views over time according to an example. In the example shown, a 3D scene is captured using a rectangular array of forty RCPs arranged in five rows and eight columns. The dashed arrow (402) represents the movement trajectory of the new camera position (50) for virtual view synthesis during a time interval (starting at time t0 and ending at time t1). At time t0, the set of active views includes the four views covered by the dashed box (410). At time t1, the set of active views includes the four views covered by the dotted box (420). At time t n (where t0 < t n < t1), the set of active views changes from the set in the dashed box (410) to the set in the dotted box (420).

[0039] Various embodiments disclosed herein relate to providing an end-to-end interactive multi-view MPI video streaming system that includes: • Components located on the server side and the content storage format used therein; • Components that support a network distribution protocol; and • Components that support an interactive playback solution on the terminal client side.

[0040] Some embodiments use the MPEG-DASH and ISO-BMFF standards and / or their extensions, and additionally provide new tools to support new features. One objective of this embodiment is to enable the rapid and / or cost-effective deployment of interactive, immersive volumetric video experiences by building upon and extending certain legacy technologies / infrastructure to achieve seamless interactive view switching for multi-view MPI video streaming. Some embodiments may benefit from at least some of the features disclosed in one or more of the following documents: (1) ISO / IEC 14496-12:2020, Information technology – Coding of audiovisual objects – Part 12: ISO basic media file format; (2) ISO / IEC 14496-15:2021, Information technology – Coding of audiovisual objects – Part 15: Bearing of structured video in network abstraction layer (NAL) units in ISO basic media file format; (3) ISO / IEC 23000-19:2018, Information technology – Multimedia application format (MPEG-A) – Part 19: Common media application format (CMAF) for segmented media; and (4) ISO / IEC 23009-1(2019), Information technology – Dynamic adaptive streaming based on HTTP (DASH) – Part 1: Media presentation description and segment format. Each of these documents is incorporated herein by reference in its entirety.

[0041] For storage formats, some embodiments are based on the ISO Basic Media File Format, commonly known as ISO BMFF (ISO / IEC 14496-12), and its extension ISO / IEC 14496-15, which provides the carrying of Network Abstraction Layer (NAL) unit structured video (such as AVC, HEVC, VVC, etc.). In this document, the acronym ISO stands for International Organization for Standardization. For data transmission, some embodiments are based on HTTP-based Dynamic Adaptive Streaming (DASH) (ISO / IEC 230090-1). In some examples, similar approaches can be applied to the Common Media Application Format (CMAF). Note that the disclosed video streaming system can also support other multi-view codec solutions, such as MVC, SHVC, and multi-layer VVC. For illustrative purposes and without any implied limitation, the example embodiments are described with reference to the HEVC codec. Based on the provided description, those skilled in the art will be able to create and use other embodiments compatible with other suitable codecs without any unnecessary experimentation.

[0042] Transmission of encoded MPI video Figure 5This is a block diagram illustrating an MPI video system (500) that can be used in a transport pipeline (100) according to various embodiments. In different embodiments, the system (500) can be configured for real-time and / or on-demand use cases. The system (500) includes an MPI video source (502) where a scene is captured by a set of cameras, and scene acquisition produces MPI video (504). Each MPI image of the MPI video (504) includes a frontal parallel plane of the corresponding scene sampled at a different depth from a reference coordinate system. Information stored in each plane includes texture and opacity / transparency, for example, as described above.

[0043] The MPI encoder (506) operates to encode the MPI video (504) into an encoded bitstream (508), for example, as described in more detail below. In some examples, the MPI encoder (506) uses a packing process that spatially packs and / or temporally interleaves the texture and alpha maps of the MPI representation. The MPI encoder (506) further operates to encode the resulting packed video using a two-dimensional (2D) codec, such as AVC, HEVC, VVC, etc. In some examples, the MPI representation can be encoded using the MPEG Immersive Video Coding Standard. The encapsulation module (514) operates to encapsulate the encoded bitstream (508) into a media file for file playback (516) or into a sequence of segments for streaming (518), depending on a selected media container file format. In some examples, the media container file format is the ISO BMFF described above. The encapsulation module (514) also includes metadata (512) in its output (516 or 518). Metadata (512) is generated using the processing / editing module (510) and can specify how texture maps and alpha maps are packaged into video frames carried by media files or segments. The segments are transmitted to the player (590) using the transmission mechanism (520). The player (590) includes a decapsulation module (524) configured to perform the opposite operation of the encapsulation module (514).

[0044] To enable the player (590) to request a specific video segment from the available video segments, the corresponding metadata is signaled in a streaming manifest file containing sufficient information for the player (590) to download and render the given content segment. In some examples, the streaming manifest file is a Media Presentation Description (MPD) of the MPEG-DASH standard. The player (590) receives the streaming manifest file and determines which video segments are suitable for the current configuration / situation, such as viewing location and orientation, device capabilities, bandwidth, etc. The player (590) then operates to acquire the video segments accordingly.

[0045] The decapsulation module (524) operates to process the video file or segment and reconstruct the encoded bitstream (508). The decapsulation module (524) further operates to parse the metadata (512) and extract packing information (how the texture map and alpha map are packed into the video frames on the encoder side). In some examples, a bitstream of an MPI representation can be carried on multiple tracks. The MPI decoder (528) operates to decode the encoded bitstream (508) into multiple frames, and in the reconstruction module (532), based on the MPI video-related information and the packing metadata of the texture map and alpha map, generates a reconstructed signal (530) carrying the MPI representation from the corresponding decoded video signals (530). The output viewport (538) of the MPI representation is rendered by the rendering module (536) according to the viewport (538) including the viewing position and orientation, and displayed on the display device (540).

[0046] For a multi-view MPI experience, the source (502) generates a multi-view MPI representation comprising a set of MPI representations, each MPI representation originating from a different corresponding reference coordinate system. Each single-view MPI representation is independently encoded as described below and encapsulated into a media file for file playback (516) or into an initialization segment and media segment sequence for streaming (518). When the user changes the viewport (538) (including viewing position and / or orientation), the player (590) needs to dynamically switch to the appropriate MPI bitstream accordingly. That is, the player (590) operates to determine the video segments to be received for the relevant MPI representation based on the viewport, dynamically requests these segments, and decodes the received segments accordingly. The output MPI is then reconstructed, and the output viewport of the MPI representation is displayed, matching the changed viewport. In the example shown, these functions of the player (590) are implemented using a viewport tracking module (550).

[0047] As described above, an MPI representation comprises a set of planar layers corresponding to different depths from a given reference viewpoint. Each layer has texture (color) channels and alpha channels, which are obtained by projecting portions of the 3D scene contained around the layer location onto the same reference camera positioned at the given reference viewpoint. In various examples, sequences of texture and alpha maps from layers of an MPI representation can be encoded using MPEG immersive video (MIV) or 2D video codecs, for example, as described in more detail below.

[0048] The MIV standard is specifically designed for the compression of immersive video content (also known as volumetric video). The MIV standard enables the storage and distribution of immersive video content on existing and future networks for playback within a limited viewing space with six degrees of freedom (6 DoF) of viewing position and orientation, and with different fields of view depending on the capture setup. In some examples, the camera array can be a planar setup with coaxially aligned cameras, linearly arranged to allow significant lateral displacement, or divergently oriented to achieve a large field of view suitable for head-mounted displays (HMDs).

[0049] The MIV standard is part of ISO / IEC 23090 MPEG-I (a set of standards for the digital representation of immersive media). The MIV standard is also consistent with the Video-Based Point Cloud Compression (V-PCC) MPEG standard (ISO / IEC 23090-5). Therefore, the MIV standard references the common parts of ISO / IEC 23090-5 (Revision 2) on Visual Volumetric Video Coding (V3C). V3C provides extension mechanisms for both V-PCC and MIV. The SC 29 / WG 03 MPEG system also developed a system standard (ISO / IEC 23090-10) based on ISO BMFF for carrying V3C data.

[0050] Figure 6 This is a block diagram illustrating an MPI encoder (506) according to an embodiment. In this embodiment, the MPI encoder (506) is configured to encode MPI video (504) using the MIV standard. The resulting output is a V3C bitstream (672), which is an example of an MPI bitstream (508) (see also...). Figure 5 ).

[0051] In the example shown, the MPI encoder (506) includes a tile generation and packing module (610), a texture video generation module (630), and a transparency video generation module (640). Modules (630, 640) operate to process the MPI representation input (504), generating sequences (632, 642) of texture atlases and transparency atlases. The tile generation and packing module (610) generates associated atlas data (612), which is subsequently used by the player (590) to reconstruct the texture and transparency maps of the MPI layer. The frame packing module (650) operates to combine the texture atlases (632) and transparency atlases (642) into packed video data, thereby generating a sequence of video frames (652) including the packed video data. The video compression module (660) operates to encode the video sequence (652) using a suitable video encoding method, thereby generating a corresponding compressed video substream (662). The tile sequence compression module (620) operates to encode the atlas data (612) into a compressed atlas data substream (620). The multiplexer (670) operates to multiplex the compressed atlas data substream (620) and the compressed video substream (662) to generate an output V3C bitstream (672).

[0052] Figure 7 This is a block diagram illustrating an example structure of a V3C bitstream (672) according to an embodiment. The MIV specification defines a set of extensions and tiers referencing V3C. For example, MIV supports multiple atlases and view parameters. For each atlas, geometry, occupancy, and / or attribute information is carried in a video sub-bitstream, which runs in parallel with an atlas sub-bitstream that includes tile parameters. View parameters (e.g., camera extrinsic, intrinsic, resolution, etc.) are carried in a common atlas sub-bitstream shared between atlases.

[0053] The MIV standard also defines a set of grades to accommodate different levels of bandwidth or decoding resources on the client side (e.g., the player (590)). In addition to the main grade, which is based on geometric information embedded with occupancy information and texture, and is commonly referred to as MVD (Multi-Video + Depth), another extended grade can separate occupancy information from geometric information and can activate transparency attributes through a restricted geometric subgrade to allow MPI format transmission. Furthermore, a geometrically missing grade allows for the generation of geometric information on the client side.

[0054] exist Figure 7In the example shown, the V3C bitstream (672) comprises multiple V3C units (702, 704). The first V3C unit (702) includes a V3C parameter set (VPS) (712). The VPS (712) is configured to announce the existence of the sub-bitstream, thereby enabling the corresponding decoder to initialize the sub-decoder for all existing atlas and video sub-bitstreams. Subsequent V3C units (704) carry atlas sub-streams (714) and video sub-streams (716). The atlas sub-stream (714) contains common atlas data (CAD) and atlas data (AD). The video sub-stream (716) contains packed video data (PVD). The AD unit contains an atlas sub-bitstream, which is a Network Abstraction Layer (NAL) unit stream, but its content is not video frames, but rather a NAL unit called the Atlas Tile Layer (ATL), which carries a list of tile data. The CAD unit also contains common information applicable to all atlases, such as a list of view parameters including extrinsic and intrinsic view parameters. The PVD unit contains packaged video data of the texture and transparency of each atlas.

[0055] Figures 8A to 8F This is an illustration of examples of packing the texture map and transparency map space of an MPI layer into a video frame, which can be implemented in an MPI encoder (506) according to various embodiments. In these examples, the MPI encoder (506) is configured to use a 2D video codec. Figures 8A to 8F Each of these examples illustrates a corresponding single video frame. In different examples, texture maps and alpha maps from the same MPI representation layer can be spatially packed into the same frame, or they can be temporally interleaved in various arrangements. For example, Figure 8A and Figure 8D This arrangement is shown as follows: all textures of the layer Figure 1 The grouping begins in the first consecutive block, and all transparency of the layer is... Figure 1 The first group is in the second consecutive block. Then, these two consecutive blocks can be grouped as follows: Figure 8A Placed side by side as shown, or as... Figure 8D Place them vertically as shown. Figure 8B , Figure 8C and Figure 8E The illustration shows an example of how the texture map and transparency map of the middle layer are arranged in alternating rows or columns within a video frame. Figure 4 Figure F illustrates an example where different layers of images are packed into regions of video frames of different sizes (i.e., these images can have different resolutions). In different examples, the 2D video codec used can be either HEVC or VVC. One or more SEI messages can contain MPI video information, such as the number of MPI layers, the layer depth, and the specific arrangement of texture and transparency maps (e.g., ...). Figures 8A to 8FThe identifier is one of the arrangements shown. In this document, the acronym SEI stands for Supplemental Enhancement Information.

[0056] To help the player (590) select which tracks(s) to retrieve from the video file or which segments(s)(s) to download via a channel(s), appropriate metadata can be provided. In some examples, the metadata specifies: • The structure of the entire MPI representation, such as whether it is a single-view or multi-view MPI representation.

[0057] • In the case of multi-view MPI (for multi-camera setups). ○ Information corresponding to each reference camera view, such as: Refer to the camera view identifier; External and internal camera information; ○ Adjacent reference views, so that the player (590) can recognize the view to be prefetched.

[0058] • MPI video information, such as: ○ The number of MPI layers; ○ Depth information of the MPI layer, including whether the MPI has a constant depth increment between layers; Packing information for the texture and transparency maps of the MPI layer; ○ MPI video decoding specifications, including codec and profile identification and level information; ○ Information specific to operations after MPI decoding.

[0059] At the player (590), the following inputs can be used to determine which track(s)(s) to retrieve from the video file or which representation(s) to request. Using this information, the player (590) operates to retrieve or request the desired track(s)(s), allowing the end user to seamlessly continue consuming MPI content.

[0060] • New viewports, including viewing position and orientation, and field of view: ○ We know the user's newly selected viewport from the user's input device (eye tracking, touch screen, sensor, etc.), including the viewing position and orientation, and the vertical and horizontal fields of view.

[0061] • Equipment characteristics: ○ Video decoding capabilities, including supported codecs.

[0062] ○ Supported post-processing / rendering features, such as more advanced algorithms that allow smooth playback when the end user makes a significant change to a new viewing pose, for example, being configured to request more surrounding views (potentially at a lower bit rate to improve bandwidth utilization).

[0063] • Network conditions: Based on the network conditions experienced (such as bandwidth, packet loss, packet latency), the player (590) can determine which bitrate version to download (indicated).

[0064] • Current time: In some examples, we know how many frames are left in the current segment and which segment needs to be requested from the server / origin side.

[0065] MPI video storage and metadata signaling using ISO BMFF As described above, when an MPI representation is encoded as a V3C bitstream, the V3C bitstream can be stored in a single track or multiple tracks within the file. When an MPI representation is encoded as a 2D video bitstream, it can be stored in a track defined in ISO / IEC 14496-15. At least some of the metadata described above is signaled in the track so that the player (590) can appropriately select and retrieve the relevant bitstream(s) from the file. Below, we first refer to Figures 9 to 12 An example embodiment of the storage and metadata signaling for a V3C bitstream (518) carrying MPI video is described. Then, we refer to... Figure 13 An example embodiment of storage and metadata signaling for a 2D video bitstream (518) carrying MPI video is described.

[0066] This section describes the encapsulation of a single V3C bitstream (518) of an encoded MPI bitstream in the track. Only one of the following encapsulations can be used at a time.

[0067] • Single-track packaging, in which one track carries the entire V3C bitstream; • Multi-track encapsulation, wherein one or two tracks of a plurality of tracks carry atlas substreams (714), and the other tracks of the plurality of tracks carry pack video substreams (716).

[0068] The types of boxes described in this article are listed in bold in Table 1.

[0069] Table 1: Types of boxes described in this document (indicated in bold) and their relationship to traditional boxes (not explicitly described in this specification).

[0070] Figure 9This is a block diagram illustrating a single-track encapsulation of a V3C bitstream (672) according to some examples. For a single-track encapsulation, the V3C bitstream (672) is stored directly as a V3C bitstream track (902). The V3C unit header remains unchanged in the bitstream (902). Each sample (920) in the V3C bitstream track (902) includes a corresponding subsample belonging to the sample presentation time, and each subsample contains a corresponding V3C unit containing one of Common Atlas Data (CAD), Atlas Data (AD), and Packed Video Data (PVD).

[0071] The movie header (910) of the V3C bitstream track (902) contains decoding-specific information and subsample information for the V3C bitstream. Therefore, the V3C bitstream track (902) contains 2D video decoder information for configuring and initializing the 2D video decoder. The header (910) also contains a V3CConfigurationBox (912), which includes decoding-specific information for the V3C bitstream (e.g., parameter sets and SEI messages). To enable support for subsample-level access, the V3C bitstream track (902) also contains a SubSampleInformationBox (916), which lists the subsamples of the track. The 32-bit unit header representing the V3C unit of the subsample is copied to the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox (916). By parsing the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox (916), the V3C unit type of each subsample can be identified.

[0072] In some examples, V3CConfigurationBox(912) is defined as follows:

[0073] In some examples, the track uses a VolumetricVisualSampleEntry with sample entry types 'v3e1' or 'v3eg', for example, as described in ISO / IEC 23090-10. This entry contains a V3CConfigurationBox (912) and may contain optional boxes for 2D video decoder configuration information (such as HEVCconfigurationBox and VVCconfigurationBox as defined in ISO / IEC 14496-15) to signal 2D video decoder configuration and initialization information.

[0074] Sample entry types: 'v3e1', 'v3eg' Container: SampleDescriptionBox Requirement: The 'v3e1' or 'v3eg' sample entry is required. Quantity: one or more

[0075] Figures 10A to 10B This is a block diagram illustrating a multitrack encapsulation of a V3C bitstream (672) according to some examples. In the example shown, multitrack encapsulation is implemented using two types of tracks: a V3C atlas track (1002) and a V3C video component track (1004). The V3C atlas track (1002) contains the V3C parameter set and may also contain the atlas parameter set in sample entries, as well as atlas component bitstream NAL units in the samples. The V3C video component track (1004) contains access units in the samples for packing texture and transparency data into the video encoded basic stream. Therefore, one or two V3C atlas tracks (1002) and at least one packed video track can exist in a video file. Figure 10A In the example shown, the atlas substream is carried across two V3C atlas tracks (10021, 10022), one for common atlas data and the other for independent atlas data. Figure 10B In the example shown, the entire atlas substream is carried in a single track (1002).

[0076] In some examples, the V3C atlas track (1002) uses a V3CAtlasSampleEntry, which extends the VolumetricVisualSampleEntry with sample entry types 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag'. A V3C atlas track sample entry contains a V3CConfigurationBox (912), and an optional V3CUnitHeaderBox contains a V3C unit header describing the data carried by the corresponding track.

[0077] Sample entry type: 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag' Container: SampleDescriptionBox Requirement: Sample entries 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag' are required. Quantity: one or more

[0078] In some examples, the V3CUnitHeaderBox also exists in the scheme_specific_data box array of the SchemeInformationBox of the V3C packaged video track. The V3CUnitHeaderBox contains the V3C unit header describing the data carried by the video track.

[0079]

[0080] The header contains a single instance of a 32-bit V3C unit header.

[0081] In some examples, to link tracks to other tracks, the track reference tool defined in ISO / IEC 14496-12 can be used. To link a V3C atlas track carrying only common atlas data to a V3C atlas track carrying independent atlas data, the 4CC for these track reference types is 'v3cs'. To link a V3C atlas track carrying independent atlas data to a V3C packaged video track, the 4CC for these track reference types is 'v3cp'.

[0082] Figure 11 It is illustrated with Figure 10A The example shown is a block diagram corresponding to a multi-track encapsulation with further details. Atlas substreams (620) are carried in atlas tracks (10021, 10022). Atlas track (10021) is used for common atlas data, and atlas track (10022) is used for independent atlas data. Video substreams (662) are carried in a separate track (1004). Figure 11 As shown, each of the atlas tracks (10021, 10021) contains a V3CConfigurationBox and a V3CUnitHeaderBox, where the V3CUnitHeaderBox has a 32-bit unit header representing the V3C unit of the sample in the track. The V3C packed video track (1004) contains video configuration information for 2D video decoder configuration and initialization, as well as a V3CUnitHeaderBox containing the unit header representing the V3C unit of the sample in the track.

[0083] Figure 12 It is illustrated with Figure 10B The example shown is accompanied by a block diagram with further details of the corresponding multi-track encapsulation. In this example, the entire atlas sub-stream (620) is carried within a single track (1002). Figure 12As shown, each sample in the V3C atlas track (1002) corresponds to one or more ACL NAL units, and each subsample contains one ACL NAL unit. The V3C atlas track (1002) contains a V3CConfigurationBox (which includes decoding-specific information for the V3C bitstream) and a SubSampleInformationBox (which lists the subsamples of the track). The 32-bit unit header representing the V3C unit of a subsample is copied to the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox. By parsing the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox, the V3C unit type of each subsample can be identified.

[0084] The following sections of this specification describe the storage and metadata signaling of 2D video bitstreams carrying MPI video. Specifically, we describe the encapsulation of MPI bitstreams encoded by a suitable 2D video codec in the track, as described above. The relevant box types are listed in bold in Table 2.

[0085] Table 2: The types of boxes described in this section and their relationship to boxes not specified in this document.

[0086]

[0087] In some examples, storing video bitstreams (MPI bitstreams) to tracks utilizes some existing capabilities of the ISO Basic Media File Format (e.g., as specified in ISO / IEC 14496-12 and 14496-15), but also defines extensions to support the following features of MPI video: - The decoded image is a packaged image containing a texture map and a transparency map of the MPI layer corresponding to the MPI representation; - Packing information for MPI-encoded video frames, such as representations of two spatially packed component frames or two temporally interleaved component frames.

[0088] In various examples, a player that cannot recognize the scheme type in the SchemeTypeBox can be configured to ignore the corresponding track. A player that recognizes the scheme type in the SchemeTypeBox (e.g., player (590)) operates to parse all boxes contained in the SchemeInformationBox to determine if it has the capability required to correctly handle the track. Unless the player supports all boxes contained in the SchemeInformationBox and all syntax elements and syntax element values ​​present in those boxes, the player will be configured to ignore the track.

[0089] Using the MPI video scheme for the restricted video sample entry type 'resv' indicates that the decoded image is a packaged image containing MPI video content. The use of the MPI video scheme is indicated by setting the scheme_type to 'mpiv' in the SchemeTypeBox within the RestrictedSchemeInfoBox.

[0090] Box type: 'mpiv' Container: SchemeInformationBox Requirement: Yes, when scheme_type equals 'mpiv' Quantity: zero or one The MPIVideoBox is used to indicate MPI video-specific information as described above. In various examples, this box is a container box that contains boxes indicating the following information: (1) Decode the image packing information, that is, the decoded frame either contains the representation of two spatially packed frames or the representation of two temporally interleaved frames that form the MPI representation.

[0091] (2) Regional packaging information (where applicable).

[0092] The MPIVideoBox box contains an MPIPackingFormatBox box, which specifies the arrangement type of the texture and alpha maps of the MPI layers in the decoded image, and whether this arrangement is applied to each layer or to each image that makes up the image. The MPIVideoBox box optionally contains an MPIRegionWisePackingBox box, used for mapping between packed regions and corresponding MPI layers. The relevant definitions are as follows:

[0093] In this article, `packing_type` indicates the type of arrangement of the texture map and alpha map of the MPI layer in the decoded image. The following values ​​are specified, and one of them is set: Table 3: Packaging Types

[0094] When the value of packing_type indicates that time-interleaved frames are packed together, the player needs to implicitly set the composite timestamp of the frame 0 to be the same as the composite timestamp of the frame 1.

[0095] `packing_unit` indicates whether the arrangement type, indicated by `packing_type`, applies to each image in the composition or to each layer. A `packing_unit` of 0 indicates that the arrangement type, indicated by `packing_type`, applies to each image in the composition. When `packing_type` equals 1, it should be 0.

[0096] A packing_unit equal to 1 indicates that the arrangement type indicated by packing_type is applied to each layer.

[0097] Note 1: Figure 8A The illustration shows an example where packing_type equals 3 and packing_unit equals 0. Figure 8D The illustration shows an example where packing_type equals 2 and packing_unit equals 0.

[0098] Note 2: Figure 8B The illustration shows an example where packing_type equals 3 and packing_unit equals 1. Figure 8C The illustration shows an example where packing_type equals 2 and packing_unit equals 0.

[0099] The MPIRegionWisePackingBox box specifies the mapping between the packing regions and their corresponding layers, and specifies the position and size of each region. The relevant definitions are as follows: Box type: 'mprw' Container: MPIVideoBox Necessity: No Quantity: zero or one

[0100] num_mpi_layers specifies the number of MPI layers corresponding to the packetized regions notified by signaling in this syntax structure.

[0101] mpi_layer_width and mpi_layer_height specify the width and height of the MPI layer region in the original mapped image before packaging.

[0102] mpi_layer_id[i] specifies the identifier of the MPI layer corresponding to the i-th packing region.

[0103] `mpi_pic_type[i]` indicates whether the i-th packed region corresponds to transparency or texture. A value of 0 for `mpi_pic_type[i]` specifies that the i-th packed region corresponds to texture of the MPI layer indicated by `mpi_layer_id[i]`. A value of 1 for `mpi_pic_type[i]` specifies that the i-th packed region corresponds to transparency of the MPI layer indicated by `mpi_layer_id[i]`.

[0104] mpi_packed_region[i] specifies the region-level packing between the i-th packed region and the MPI layer indicated by mpi_layer_id[i].

[0105]

[0106] The MPIRegionPacking structure specifies the MPI layer region and its corresponding packing region. The parameters mpi_reg_width, mpi_reg_height, mpi_reg_top, and mpi_reg_left specify the width, height, top offset, and left offset of the MPI layer region in the original mapped image before packing, respectively. The parameters packed_reg_width, packed_reg_height, packed_reg_top, and packed_reg_left specify the width, height, top offset, and left offset of the packing region, which lies within either the spatially packed texture and transparency composition image or within each component image of the temporally interleaved texture and transparency composition image.

[0107] Figure 13 This is a block diagram illustrating a single-track encapsulation of a 2D video stream (MPI bitstream) based on some examples. For example... Figure 13 As shown, the MPI bitstream (508) is carried in a single track (1302). Track (1302) contains 2D video configuration information for 2D video decoder configuration and initialization, and also contains an MPIVideoBox box, which contains packing information of the texture map and transparency map of the MPI layer in the decoded image, as well as the mapping between the packed regions and the corresponding MPI layers. Access units for packing texture and transparency data in the video encoded basic stream are stored in the sample.

[0108] MPI video encapsulation and signaling using DASH A key component on the server side is the MPD. In various examples, the DASH MPD generator includes MPI video-specific descriptors. These descriptors include metadata as described above, such as specifying how the texture map and alpha map of the MPI layer are packaged into the decoded frame. When a user joins a session to begin MPI playback, the MPD is transmitted to the player (590) / client. By parsing the metadata from the MPD (e.g., for packaging information), the player (590) determines which adaptive set and representation contains the supported MPI video bitstream, and which adaptive set and representation covers the current viewing orientation at the highest quality and bitrate that is acceptable for the current estimated network throughput. The player (590) then issues a segment request accordingly.

[0109] As mentioned above, the single-track mode in DASH supports streaming of segments, where the V3C bitstream is stored using a single-track encapsulation method.

[0110] A single-track mode in DASH can be represented as an adaptive set with one or more representations. Each representation is associated with a specific encoded MPI bitrate. The initialization fragment of the representation contains 2D video decoder information for configuring and initializing the 2D video decoder, as well as all the parameter sets required to initialize the V3C decoder (including the V3C parameter set and other parameter sets for sub-bitstreams). The media fragment of the representation contains one or more track fragments of the V3C bitstream track (e.g., as described in the previous section), and subsample information (which lists the subsamples in the fragment to support subsample-level access).

[0111] The following provides an example of MPEG DASH signaling when encoding an MPI video representation using HEVC in single-track mode:

[0112] In multitrack mode, the common atlas, independent atlases, and packed video components are represented as separate adaptive sets in the DASH manifest (MPD) file. The adaptive set used for common atlas information is called the master adaptive set. In some examples, the master adaptive set contains a single initialization fragment at the adaptive set level. This initialization fragment contains all the parameter sets required to initialize the V3C decoder, including the V3C parameter set and other parameter sets for the component sub-bitstreams. The media fragment represented by the master adaptive set contains one or more track fragments of the V3C atlas track. The media fragment represented by the video component adaptive set contains one or more track fragments of the corresponding packed video track as described above.

[0113] Figure 14This is a block diagram illustrating a DASH configuration based on some examples, used to group atlases and packed video components belonging to the same MPI content in an MPEG-DASHMPD file. For example... Figure 14 As shown, the preselection (1402) is used to combine multiple adaptive sets (e.g., 1404, 1406, 1408) into a single decoding instance and MPI user experience. In some examples, the preselection (1402) can be signaled in the MPD, for example, using a PreSelection element or PreSelection descriptor within a Period element. Each of the public atlas track, independent atlas track, and packaged video track is presented as a separate corresponding adaptive set. The preselection (1402) indicates the grouping of the three adaptive sets (1404, 1406, 1408) to provide the user with a single MPI experience.

[0114] To identify the type of the video component adaptive set, the V3CVideoComponent descriptor is used. The V3CVideoComponent descriptor is an EssentialProperty element with its @schemeIdUri set to "urn:mpeg:mpegI:v3c:2020:videoComponent".

[0115] At the adaptive set level, the V3C video component is signaled to the V3CVideoComponent descriptor for each V3C video component present in the representation of the video component adaptive set. The @value of the V3CVideoComponent descriptor may not exist. In some examples, the V3CVideoComponent descriptor includes the elements and attributes specified in Table 4.

[0116] Table 4: Elements and properties of the V3CVideoComponent descriptor.

[0117]

[0118] The following is an example XML schema for the V3CVideoComponent descriptor:

[0119] At the adaptive set level for packed videos, a V3CVideoComponent descriptor is signaled with its videoComponent@type set to "packed". Alternatively, to indicate which attribute components are packed into the adaptive set, two V3CVideoComponent descriptors are signaled: one with videoComponent@type set to "packed" and videoComponent@attribute_type set to "ATTR_TEXTURE", and the other with videoComponent@type set to "packed" and videoComponent@attribute_type set to "ATTR_TRANSPARENCY".

[0120] The following MPD example illustrates the signaling components of each adaptive set (such as a packed video adaptive set) and how to use the preselected element (1402) to indicate an adaptive set group representing a single MPI video.

[0121]

[0122] The next part of this section describes the streaming (518) of segments when the corresponding MPI bitstream is stored using the track encapsulation described above (508). In some examples, each MPI representation is associated with a specific encoded MPI video with a specific bit rate. The initialization segment of the representation contains 2D video decoding-specific information for configuring and initializing the 2D video decoder and MPIVideoBox box described herein. The media segment of the representation contains one or more track segments of the tracks described above.

[0123] The DASH MPD generator includes MPI video-specific descriptors. These descriptors include packing type and region-level packing information. This information can be generated based on equivalent information in the clips.

[0124] To identify the packing arrangement of texture and alpha maps in the MPI representation within an adaptive set, the MPIPackingFormat (MPF) descriptor is used. The MPF descriptor is an EssentialProperty descriptor element with its @schemeIdUri attribute set to "urn:mpeg:dash:mpi:2023:pf.". In some examples, an MPF ​​descriptor exists at the adaptive set level. An MPF ​​descriptor may also exist at the representation level.

[0125] In some examples, the @value attribute of the MPF descriptor is not present. An MPF ​​descriptor can include the elements and attributes specified in Table 5.

[0126] Table 5. Elements and attributes of the MPIPackingFormat (MPF) descriptor.

[0127]

[0128] In some examples, the XML schema for an MPF ​​descriptor can be as follows:

[0129] The following MPD example illustrates MPD signaling when the decoded output image representation of each representation consists of textures and transparency arranged in a top-to-bottom packing order.

[0130]

[0131] The multi-view MPI experience can be referenced as above. Figure 4 This is achieved by switching reference views. The 3D scene has a total of... N A reference camera view. Figure 4 In the example shown, N = 40. In order to [achieve a specific time instance] t 0. Render a new view, using only the four adjacent reference views within block (410). The new view's position moves over time along the dotted path (402). Along path (402) at time instances... t n The nearest neighbor of the new view has become the reference view within block (420).

[0132] In a multi-view MPI system, we assume that each single-view MPI is independently encoded and encapsulated in a track or segment, for example, as described above. Since users can request different new viewports along the timeline, the player (590) needs to be configured to dynamically switch between different MPI bitstreams. Therefore, in multi-view MPI, the player (590) is configured to identify which MPI views to prefetch and switch accordingly to the appropriate MPI view when the viewport changes. The information to be signaled in the MPD for this switching operation can be described as follows.

[0133] • The entire MPI structure, including the number of MPI views.

[0134] • Information corresponding to each reference camera view.

[0135] ○ Refer to the camera view identifier; ○ External and internal reference camera information; ○ Refer to the vertical and horizontal field of view information of the camera view.

[0136] • Adjacent reference views used for joint rendering.

[0137] In some examples, we map each level of the multi-view MPI streaming system to a hierarchical data structure in the MPD. In some examples, this structure may look like this.

[0138] • InitializationSet: The InitializationSet contains the URLs of initialization segments, which include all extrinsic and intrinsic camera information for each time interval, as well as necessary post-decoding operation-specific information (such as packing information). Note that different MPI scenes may have different numbers of cameras and different camera pose settings over time. In this case, there are multiple InitializationSet properties, and each InitializationSet property contains the corresponding URL of an initialization segment containing the extrinsic and intrinsic camera information for that time interval. Each initialization set has a link to the corresponding InitializationSet for that time interval.

[0139] • Preselection: Each preselection represents a combination of adaptive sets that can be selected for joint rendering within a given time period. If we have 20 camera views and at most four adjacent views that can improve rendering of a new viewport, the preselection indicates the corresponding adjacent reference view groups for joint rendering. When preselections exist in the MPD, the player (590) can seamlessly switch to the relevant view by identifying which combinations of adaptive sets need to be prefetched and requesting multiple adaptive sets related to the viewport change.

[0140] • Adaptation Set: Each adaptation set represents an MPI from a reference camera view. Each reference camera view is associated with extrinsic and intrinsic camera information. If we have 20 camera views, we will have 20 corresponding adaptation sets.

[0141] • Representation: Each representation is associated with a specific bitrate version of the encoded MPI. Different appropriate representations are selected based on the current network conditions between the server and the client.

[0142] • Segment: There are two types of segments: initialization segments and media segments. These can be encapsulated as described above.

[0143] Figure 15 This is a flowchart (1500) illustrating the communication between a server (1520) and a client (1530) according to various examples, and the corresponding processing operations performed at the client (1530). In some examples, the client (1530) may be or include a player (590). The final operation box (1511) in the flowchart (1500) includes the client generating an output view and rendering the generated view according to the user viewport.

[0144] When a user initiates or joins a server-client session to begin MPI playback, the MPD (1501) is sent from the server (1520) to the client (1530). The client (1530) parses the MPD (1501) in a first operation box (1502) and determines a preselection based on the current viewport in a second operation box (1503). In a third operation box (1504), the client (1530) selects an adaptive set, sends a request for the corresponding initialization segment (1505), and receives the requested initialization segment (1506). In a fourth operation box (1507), the client (1530) determines one or more representations and sends one or more corresponding requests to the server (1520) (1508). In a fifth operation box (1510), the client (1530) receives the requested one or more media segments (1509) and decodes the corresponding bitstream. In the final box (1511), the client (1530) uses the decoded bitstream to generate (multiple) views suitable for the current view pose and renders the generated views according to the user viewport.

[0145] The InitializationSet contains the URL of an initialization fragment (e.g., 1506) that includes extrinsic and intrinsic camera information for all MPI views and specifies MPI post-decoding operation-specific information (such as packing information) required for that time period. Multiple InitializationSet (1506) instances can exist as this initialization information is updated over time, and each such InitializationSet instance contains the URL of a corresponding initialization fragment containing the updated initialization information for that time period, including extrinsic and intrinsic camera information and MPI post-decoding operation-specific information. Each initializationSet has an @InitializationSetRef to be associated with the appropriate InitializationSet for that time period.

[0146] For a preselection of a combination of adaptive sets that can be selected for joint rendering within this time period, it can be notified by a PreSelection element within the Period element in MPD(1501) or by signaling in the Preselection descriptor at the adaptive set level. The PreSelection element is signaled via a list of IDs of the @preselectionComponents attribute, which includes the IDs of the adaptive sets used for adjacent reference MPI views. The PreSelection element is associated with a 3D Spatial Region Coverage (SRC) descriptor to indicate the 3D spatial region covered by the combination of preselected MPI views.

[0147] A SupplementalProperty element with the @schemeIdUri attribute equal to "urn:mpeg:dash:mpi:2023:src" is called a 3D Spatial Region Coverage (SRC) descriptor. In various examples, an SRC descriptor exists at the preselected level. An SRC descriptor may not exist at the MPD, Adaptive Set, or Representation level. The SRC descriptor indicates the 3D spatial region covered by the combination of preselected MPI views.

[0148] In some examples, the @value attribute of the SRC descriptor may not exist. In various examples, the SRC descriptor may include some or all of the elements and attributes specified in Table 6.

[0149] Table 6. Elements and Attributes of SRC Descriptors

[0150] In some examples, the XML schema for the SRC descriptor is as follows:

[0151] In some examples, an MPIView descriptor is used to identify the MPI view information employed. An MPIView descriptor is an EssentialProperty or SupplementalProperty element with its @schemeIdUri attribute set to "urn:mpeg:dash:mpi:2023:view". At most one MPIView descriptor exists at the adaptive set level.

[0152] In some examples, the @value attribute of the MPIView descriptor may not exist. In various examples, the MPIView descriptor includes some or all of the elements and attributes specified in Table 7.

[0153] Table 7: Elements and attributes of the MPIView descriptor.

[0154]

[0155] In some examples, the XML schema for the MPIView descriptor is as follows.

[0156]

[0157] Figure 16 This is a block diagram illustrating, according to some examples, a DASH configuration that can be used in an MPI video system (500) to support seamless switching between MPI views within an MPEG-DASH MPD file. In the example shown, the MPD (1501) includes an initialization set (1601) and a time segment (1602). The initialization set (1601) includes URLs of initialization segments containing all the necessary initialization information. The time segment (1602) includes one or more preselected segments (1603). k ) and one or more adaptive sets (1604) n Preliminary selection (1603) k This indicates a combination of adaptation sets corresponding to (multiple) adjacent reference views, which can be used to improve the generation of new views through joint rendering. The 3D spatial region coverage information of such a combination of adaptation sets is indicated by a spatial region coverage descriptor. Each adaptation set (1604) n It is also indicated by MPI view information, which includes the corresponding extrinsic and intrinsic camera information.

[0158] Example hardware Figure 17 This is a block diagram illustrating example computing devices (1700) according to various examples. In some examples, instances of two or more computing devices (1700) are used in the MPI video system (500).

[0159] Figure 17 The computing device (1700) shown is illustrated as having multiple components, but any one or more of these components may be omitted or repeated as appropriate for application and setup. In some embodiments, some or all of the components included in the computing device (1700) may be attached to one or more motherboards and encapsulated in a housing. In some embodiments, some of these components may be manufactured onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices (1702) and one or more storage devices (1704)). Additionally, in various embodiments, the computing device (1700) may not include... Figure 17 One or more of the components shown may be included, but may include interface circuitry for coupling to the one or more components using any suitable interface, such as a Universal Serial Bus (USB) interface, a High Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other suitable interface. For example, the computing device (1700) may not include a display device (1710), but may include display device interface circuitry (e.g., connector and driver circuitry) to which an external display device (1710) may be coupled.

[0160] The computing device (1700) includes a processing device (1702) (e.g., one or more processing devices). The terms "electronic processor device" and "processing device" are interchangeable to refer to any device or part of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory. In various embodiments, the processing device 1702 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing device.

[0161] The computing device (1700) also includes a storage device (1704) (e.g., one or more storage devices). In various embodiments, the storage device 1704 may include one or more memory devices, such as random access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive bridged RAM (CBRAM) devices), hard disk drive-based memory devices, solid-state memory devices, network drives, cloud drives, or any combination of memory devices, etc. In some embodiments, the storage device (1704) may include memory that shares a die with the processing device (1702). In such embodiments, the memory may be used as cache memory and may include, for example, embedded dynamic random access memory (eDRAM) or spin-transfer torque magnetic random access memory (STT-MRAM). In some embodiments, the storage device (1704) may include a non-transitory computer-readable medium having instructions thereon that, when executed by one or more processing devices (e.g., processing device (1702)), cause the computing device (1700) to perform any suitable method or portion thereof of the disclosed methods.

[0162] The computing device (1700) further includes an interface device (1706) (e.g., one or more interface devices (1706)). In various embodiments, the interface device (1706) may include one or more communication chips, connectors, and / or other hardware and software to manage communication between the computing device (1700) and other computing devices. For example, the interface device (1706) may include circuitry for managing wireless communication to enable data transmission to and from the computing device (1700). The term “wireless” and its derivatives can be used to describe circuits, devices, systems, methods, techniques, communication channels, etc., that can transmit data via modulated electromagnetic radiation through a non-solid medium. This term does not imply that the associated devices do not contain any wires, although in some embodiments they may not contain wires. The circuitry included in the interface device (1706) for managing wireless communications can implement any of a variety of wireless standards or protocols, including but not limited to IEEE standards (including Wi-Fi (IEEE 802.11 series, IEEE 802.16 standards), Long Term Evolution (LTE) projects, and any amendments, updates, and / or revisions (e.g., Advanced LTE project, Ultra Mobile Broadband (UMB) project (also known as “3GPP2”), etc.). In some embodiments, the circuitry included in the interface device (1706) for managing wireless communications can operate according to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed ​​Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. In some embodiments, the circuitry included in the interface device (1706) for managing wireless communications can operate according to GSM Evolved Enhanced Data (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, the circuitry included in the interface device (1706) for managing wireless communications may be based on Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications System (DECT), Evolved Data Optimization (EV-DO) and its derivative protocols, as well as any other wireless protocol designated as 3G, 4G, 5G, etc. In some embodiments, the interface device (1706) may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0163] In some embodiments, the interface device (1706) may include circuitry for managing wired communications, such as electrical communication protocols, optical communication protocols, or any other suitable communication protocols. For example, the interface device (1706) may include circuitry supporting communications based on Ethernet technology. In some embodiments, the interface device (1706) may support both wireless and wired communications, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry for the interface device (1706) may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry for the interface device (1706) may be dedicated to longer-range wireless communications such as GPS, EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry for the interface device (1706) may be dedicated to wireless communications, and a second set of circuitry for the interface device (1706) may be dedicated to wired communications.

[0164] The computing device (1700) also includes a battery / power circuit (1708). In various embodiments, the battery / power circuit (1708) may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device (1700) to an energy source (e.g., AC line power) separate from the computing device (1700).

[0165] The computing device (1700) also includes a display device (1710) (e.g., one or more separate display devices). In various embodiments, the display device (1710) may include any visual indicator, such as a head-up display, computer monitor, projector, touch screen display, liquid crystal display (LCD), light-emitting diode display, or flat panel display, etc.

[0166] The computing device (1700) also includes additional input / output (I / O) devices (1712). In various embodiments, the I / O devices (1712) may include one or more data / signal transmission interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earpieces, sirens, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human-machine interface devices (e.g., keyboards, cursor control devices such as mice, styluses, trackballs, or touchpads), etc.

[0167] According to specific embodiments, various components of the interface device (1706) and / or I / O device (1712) can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface device (1706) and / or I / O device (1712) includes one or more analog-to-digital converters (ADCs) for converting received analog signals into a form suitable for operation performed by the processing device (1702) and / or storage device (1704). In some additional examples, the interface device (1706) and / or I / O device (1712) includes one or more digital-to-analog converters (DACs) for converting digital signals provided by the processing device (1702) and / or storage device (1704) into an analog form suitable for transmission over a communication channel.

[0168] Based on the exemplary embodiments disclosed above, for example, in the summary section and / or references Figures 1 to 17 Any one or any combination of some or all of the above provides an apparatus for streaming MPI video, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured, together with the at least one processor, to cause the apparatus to at least: generate a sequence of video frames, each of the video frames including a plurality of tiles of texture layers and transparency layers representing one or more multiplanar images of MPI video; apply video compression to the sequence of video frames to generate a video substream; generate a representation sequence of atlas frames corresponding to the sequence of video frames to at least specify a packing arrangement of tiles; apply compression to the representation sequence to generate an atlas substream; and multiplex the video substream and the atlas substream to generate a first coded bitstream that encodes at least a portion of the MPI video.

[0169] According to another example embodiment disclosed above, for example, in the summary section and / or references Figures 1 to 17 A method for streaming MPI video is provided, comprising: generating a sequence of video frames, each video frame including corresponding plurality of tiles of texture layers and transparency layers representing one or more multiplanar images of the MPI video; applying video compression to the video frame sequence to generate a video substream; generating a representation sequence of atlas frames corresponding to the video frame sequence to at least specify a packing arrangement of tiles; applying compression to the representation sequence to generate an atlas substream; and multiplexing the video substream and the atlas substream to generate a first coded bitstream that encodes at least a portion of the MPI video.

[0170] In some embodiments of the above method, the first encoded bitstream is a bitstream based on visual volumetric video coding (V3C) according to the ISO.IEC.23090-5 standard.

[0171] In some embodiments of any of the methods described above, the method further includes encapsulating a first encoded bitstream into a media file configured for file playback.

[0172] In some embodiments of any of the methods described above, the method further includes: encapsulating a first encoded bitstream into a sequence of segments according to a selected media container file format; and streaming the sequence of segments to a client device via a communication channel for playback.

[0173] In some embodiments of any of the methods described above, the first encoded bitstream is configured to be encapsulated in a single V3C bitstream track.

[0174] In some embodiments of any of the methods described above, the first encoded bitstream is configured to be encapsulated in two or more V3C bitstream tracks.

[0175] In some embodiments of any of the methods described above, two or more V3C bitstream tracks include: a first track that carries at least a portion of the atlas substream; and a second track that carries at least a portion of the video substream.

[0176] In some embodiments of any of the methods described above, texture tiles and transparency tiles from the respective plurality of tiles are packaged into the corresponding video frame using a packing arrangement selected from the group consisting of: side-by-side arrangement; top-to-bottom arrangement; vertical staggered arrangement; and horizontal staggered arrangement.

[0177] In some embodiments of any of the methods described above, texture tiles and transparency tiles include tiles of a first size and tiles of different second sizes.

[0178] In some embodiments of any of the methods described above, the corresponding plurality of tiles have fewer tiles than all tiles of the corresponding multiplanar image. (Partially packaged.) In some embodiments of any of the methods described above, the method further includes: providing a client device with a Media Presentation Description (MPD) of MPI streaming content stored in a storage container accessible via a server device, the MPI streaming content including a plurality of encoded bitstreams, the plurality of encoded bitstreams including a first encoded bitstream; providing the client device with a corresponding initialization segment from the storage container for a given time period, the corresponding initialization segment being configured to notify the client device to select one or more views of the MPI streaming content to request a media segment for rendering; receiving a request from the client device that identifies a selection and indicates a corresponding recommended value selected from at least one parameter of a group consisting of bitrate, resolution, codec type, and frame rate; and transmitting one or more of the plurality of encoded bitstreams to the client device based on the identified selection and further based on one or more of the corresponding recommended values, the one or more encoded bitstreams carrying the media segment selected in the storage container.

[0179] In some embodiments of any of the methods described above, for that time period, the storage container has multiple media segments logically organized according to different views, and further logically organized according to one or more of different bitrates, different resolutions, different codec types, and different frame rates.

[0180] In some embodiments of any of the methods described above, in multitrack mode, the common atlas, independent atlases, and packaged video components are represented as separate adaptive sets in the MPD file; and wherein the adaptive set of the common atlas contains a corresponding initialization fragment having a set of parameters for initializing the V3C decoder at the client device.

[0181] In some embodiments of any of the methods described above, the media segment representing the adaptive set of the public atlas includes one or more track segments of the V3C atlas track; and wherein the media segment representing the adaptive set of the video component includes one or more track segments corresponding to the packaged video track.

[0182] In some embodiments of any of the methods described above, the first encoded bitstream is encapsulated in a single V3C bitstream track stored in a storage container.

[0183] In some embodiments of any of the methods described above, the first encoded bitstream is encapsulated in two or more V3C bitstream tracks stored in a storage container.

[0184] In some embodiments of any of the methods described above, the method further includes switching from a first media segment sequence to a different second media segment sequence when a request indicates that the identified selection has changed.

[0185] In some embodiments of any of the methods described above, the transmission includes: transmitting a first encoded bitstream carrying a sequence of first media segments corresponding to a first view in different views; and transmitting a second encoded bitstream carrying a sequence of different second media segments corresponding to a second view in different views, wherein the first encoded bitstream and the second encoded bitstream have corresponding media segments corresponding to the same video segment time in different corresponding video segment times.

[0186] In some embodiments of any of the methods described above, the MPI video is a multi-view MPI video.

[0187] In some embodiments of any of the methods described above, the method further includes: providing a media presentation description corresponding to a multi-view MPI video to a client device; providing a corresponding initialization segment to the client device for a time period to notify two or more camera views of the multi-view MPI video selected at the client device to request a media segment for rendering; receiving a request for identification selection from the client device; and transmitting one or more coded bitstreams carrying the media segment to the client device according to the identification selection.

[0188] In some embodiments of any of the methods described above, the method further switches from transmitting a first media segment sequence to transmitting a different second media segment sequence when a request indicates that at least one of the identified selections of two or more camera views has changed.

[0189] In some embodiments of any of the methods described above, the transmission includes: transmitting a first media segment sequence corresponding to a first camera view of the scene; and transmitting different second media segment sequences corresponding to different second camera views of the scene, wherein the first media segment sequence and the second media segment sequence correspond to the same time interval of the multi-view MPI video.

[0190] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including any of the methods described above.

[0191] Regarding the processes, systems, methods, heuristics, etc., described herein, it should be understood that although the steps of these processes, etc., have been described as being performed in a specific ordered sequence, these processes can be practiced using the described steps performed in a different order than that described herein. Furthermore, it should be understood that some steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claims.

[0192] Accordingly, it should be understood that the above description is intended to be illustrative rather than restrictive. Many embodiments and applications beyond the examples provided will become apparent upon reading the above description. The scope should not be determined by reference to the above description, but rather by reference to the appended claims and the full scope of their equivalents. It is anticipated and desired that the techniques discussed herein will evolve in the future, and the disclosed systems and methods will be incorporated into such future embodiments. In conclusion, it should be understood that modifications and changes are possible with this application.

[0193] All terms used in the claims are intended to be given the broadest reasonable interpretation and common meaning as understood by one skilled in the art described herein, unless expressly indicated otherwise herein. In particular, the use of singular articles such as “a,” “the,” and “the” should be understood to refer to one or more of the indicated elements, unless the claims expressly limit this to the contrary.

[0194] An abstract of this disclosure is provided to allow the reader to quickly determine the nature of this technical disclosure. This abstract is submitted on the premise that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen in the foregoing detailed description, various features are combined together in various embodiments for the purpose of making this disclosure a coherent whole. The method of this disclosure should not be construed as reflecting an intention to incorporate more features than expressly recited in each claim in the claimed embodiments. Rather, as reflected in the appended claims, the inventive subject matter lies in fewer than all the features of a single disclosed embodiment. Therefore, the appended claims are thus incorporated into the detailed description, each claim being an independent subject matter claimed.

[0195] While this disclosure includes references to illustrative embodiments, this specification is not intended to be limiting. Various modifications to the described embodiments, as well as other embodiments within the scope of this disclosure, will be apparent to those skilled in the art to which this disclosure pertains, and are considered to be within the principles and scope of this disclosure, such as those expressed in the claims.

[0196] Some embodiments can be implemented as circuit-based processes, including possible implementations on a single integrated circuit.

[0197] Some embodiments may be embodied in the form of methods and apparatus for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium, such as a magnetic recording medium, optical recording medium, solid-state memory, floppy disk, CD-ROM, hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into and executed by a machine (such as a computer), the machine becomes an apparatus for practicing the patented invention(s). Some embodiments may also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium (including being loaded into and / or executed by the machine), wherein when the program code is loaded into and executed by a machine (such as a computer or processor), the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments are combined with the processor to provide a unique device that operates similarly to a particular logic circuit. All pseudocode examples presented herein are strictly for illustrative purposes to illustrate applicable coding concepts that can be implemented in various components of the disclosed MPI video system. Based on the provided examples, those skilled in the art will be able to create and use functionally similar code blocks for various specific implementations of the corresponding MPI video system without performing any unnecessary experiments.

[0198] Unless otherwise expressly stated, each value and range should be interpreted as approximate as if preceded by the words “about” or “approximately”.

[0199] The use of figure numbers and / or reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding figures.

[0200] Although the elements in the method claims (if any) are recited in a particular order with corresponding markings, these elements are not necessarily intended to be limited to being recited in a particular order unless the recitation of the claims otherwise implies a particular order for implementing some or all of these elements.

[0201] References to “one embodiment” or “an embodiment” herein mean that a particular feature, structure, or characteristic described in connection with an embodiment may include at least one embodiment of this disclosure. In this specification, the phrase “in one embodiment” appearing in various places does not necessarily refer to the same embodiment, nor is it necessarily a separate or alternative embodiment that is mutually exclusive with other embodiments. The same applies to the term “implementation.”

[0202] Unless otherwise specified herein, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to refer to objects among multiple similar objects merely indicates different instances of such similar objects being mentioned, and is not intended to imply that such similar objects must be in a corresponding order or sequence in time, space, ranking, or any other way.

[0203] Unless otherwise specified herein, the conjunction “if” may also be interpreted, in addition to its simple meaning, as meaning “when…”, “at…”, “in response to determination”, or “in response to detection,” depending on the specific context. For example, the phrases “if it is determined…” or “if [the stated condition] is detected” may be interpreted as meaning “after it is determined…”, “in response to determining…”, “after [the stated condition or event] is detected”, or “in response to [the stated condition or event] being detected.”

[0204] For the purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or subsequently developed that allows energy to be transferred between two or more elements, and envision the insertion of one or more additional elements, although this is not required. Conversely, the terms “direct coupling,” “direct connection,” etc., imply the absence of such additional elements.

[0205] As used in this document with respect to components and standards, the term "compatible" means that the component communicates with other components in a manner wholly or partially specified in the standard, and is recognized by other components as capable of communicating with other components in a manner specified in the standard. Compatible components do not need to operate internally in a manner specified in the standard.

[0206] The functions of the various elements shown in the figure, including any functional blocks labeled “processor” and / or “controller”, can be provided by using dedicated hardware and hardware capable of executing the software, associated with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, a single shared processor, or multiple separate processors (some of which may be shared). Furthermore, the explicit use of the terms “processor” or “controller” should not be construed as exclusively referring to hardware capable of executing software, and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the figure are conceptual only. Their functionality can be performed by the operation of program logic, by dedicated logic, by the interaction of program control and dedicated logic, or even manually, the specific technology of which may be chosen by the implementer, as understood more specifically from the context.

[0207] As used herein, the terms “circuit” and “circuitry” may refer to one or more of the following: (a) a purely hardware circuit implementation (such as implementations in purely analog and / or digital circuits); (b) a combination of hardware circuitry and software, such as (if applicable): (i) a combination of (multiple) analog hardware circuitry and / or digital hardware circuitry with software / firmware, and (ii) any portion of (multiple) hardware processors with software (including (multiple) digital signal processors), software, and (multiple) memories, which work together to enable a device such as a mobile phone or server to perform various functions; and (c) (multiple) hardware circuitry and / or (multiple) processors, such as (multiple) microprocessors or portions thereof, which require software (e.g., firmware) to operate, but may be absent when operation is not required. This definition of circuitry applies to all uses of the term herein (including all claims). As a further example, as used in this application, the term "circuit" also covers implementations of a single hardware circuit or a processor (or multiple processors), or a portion thereof, and its accompanying software and / or firmware. The term "circuit" also covers, for example, baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices, where applicable to specific claim elements.

[0208] Those skilled in the art will recognize that any block diagram herein represents a conceptual view of an illustrative circuit embodying the principles of this disclosure. Similarly, it will be appreciated that any flow chart, flow diagram, state transition diagram, pseudocode, etc., represents various processes that can be substantially represented in a computer-readable medium and therefore executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0209] The "Summary" in this specification is intended to introduce some exemplary embodiments, and additional embodiments are described in the "Detailed Description" and / or with reference to one or more accompanying drawings. The "Summary" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

1. A method for streaming multiplane image (MPI) video, the method comprising: Generate a sequence of video frames, each of which includes a plurality of corresponding tiles of a texture layer and a transparency layer representing one or more multiplanar images of the MPI video; Video compression is applied to the video frame sequence to generate a video substream; Generate a representation sequence of atlas frames corresponding to the video frame sequence, to at least specify the packing arrangement of the tiles; Compression is applied to the representation sequence to generate atlas substreams; as well as The video substream and the atlas substream are multiplexed to generate a first coded bitstream that encodes at least a portion of the MPI video.

2. The method as described in claim 1, wherein, The first encoded bitstream is a bitstream based on Visual Volumetric Video Coding (V3C) according to the ISO.IEC.23090-5 standard.

3. The method of claim 2, further comprising encapsulating the first encoded bitstream into a media file configured for file playback.

4. The method of claim 2, further comprising: According to the selected media container file format, the first encoded bitstream is encapsulated into a sequence of fragments; as well as The segment sequence is streamed to the client device via a communication channel for playback.

5. The method of claim 2, wherein, The first encoded bitstream is configured to be encapsulated in a single V3C bitstream track.

6. The method of claim 2, wherein, The first encoded bitstream is configured to be encapsulated in two or more V3C bitstream tracks.

7. The method of claim 6, wherein, The two or more V3C bitstream tracks include: A first track, the first track carrying at least a portion of the atlas substream; and The second track carries at least a portion of the video sub-stream.

8. The method of claim 1, wherein, The texture tiles and transparency tiles from the corresponding plurality of tiles are packaged into the corresponding video frames using a packing arrangement selected from the following groups: Arranged side by side; Arranged from top to bottom; Vertically staggered arrangement; as well as Arranged horizontally in an alternating pattern.

9. The method of claim 8, wherein, The texture tiles and the transparency tiles include tiles of a first size and tiles of different second sizes.

10. The method of claim 1, wherein, The corresponding plurality of tiles have fewer tiles than all the tiles in the corresponding multiplanar image.

11. The method of claim 1, further comprising: Provide a media presentation description (MPD) to a client device of MPI streaming content stored in a storage container accessible via a server device, the MPI streaming content comprising multiple encoded bitstreams, the multiple encoded bitstreams including the first encoded bitstream; For a given time period, a corresponding initialization fragment from the storage container is provided to the client device. The corresponding initialization fragment is configured to notify the client device to select one or more views of the MPI streaming content to request media segments for rendering. Receive a request from the client device, the request identifying the selection and indicating a corresponding recommended value for at least one parameter selected from the group consisting of bit rate, resolution, codec type, and frame rate; as well as Based on the identified selection and further based on one or more of the corresponding recommended values, one or more of the plurality of encoded bitstreams are transmitted to the client device, the one or more encoded bitstreams carrying media segments selected in the storage container.

12. The method of claim 11, wherein, For the specified time period, the storage container has multiple media segments, which are logically organized according to different views, and further logically organized according to one or more of different bit rates, different resolutions, different codec types, and different frame rates.

13. The method of claim 11, in, In multitrack mode, common atlases, independent atlases, and packaged video components are represented as separate adaptive sets in the MPD file; and The adaptive set of the public atlas includes the corresponding initialization fragment, which has a parameter set for initializing the V3C decoder at the client device.

14. The method of claim 13, in, The media segments represented by the adaptive set of the public atlas contain one or more track segments of the V3C atlas track; and The media segment represented by the adaptive set of the video components includes one or more track segments of the corresponding packaged video track.

15. The method of claim 11, wherein, The first encoded bitstream is encapsulated in a single V3C bitstream track stored in the storage container.

16. The method of claim 11, wherein, The first encoded bitstream is encapsulated in two or more V3C bitstream tracks stored in the storage container.

17. The method of claim 11, further comprising switching from a first media segment sequence to a different second media segment sequence when the request indicates a change in the identified selection.

18. The method of claim 17, wherein, The transmission includes: Transmit a first encoded bitstream carrying a sequence of first media segments corresponding to the first view in the different views; and Transmit a second coded bitstream carrying sequences of different second media segments corresponding to the second views in the different views. Wherein, the first encoded bitstream and the second encoded bitstream have corresponding media segments that correspond to the same video segment time in the different corresponding video segment times.

19. The method of claim 1, wherein, The MPI video is a multi-view MPI video.

20. The method of claim 19, further comprising: Provide the client device with a media presentation description corresponding to the multi-view MPI video; For a given time period, a corresponding initialization segment is provided to the client device to notify the client device to select two or more camera views of the multi-view MPI video to request media segments for rendering; Receive a request from the client device to identify the selection; as well as Based on the identified selection, one or more encoded bitstreams carrying media segments are transmitted to the client device.

21. The method of claim 20, further comprising switching from transmitting a first media segment sequence to transmitting a different second media segment sequence when the request indicates that at least one of the identified selections of the two or more camera views has changed.

22. The method of claim 20, wherein, The transmission includes: Transmit a first media segment sequence corresponding to a first camera view of the scene; and Transmit different sequences of second media segments corresponding to different second camera views of the scene. The first media segment sequence and the second media segment sequence correspond to the same time interval of the multi-view MPI video.

23. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the method of claim 1.

24. An apparatus for streaming multiplane image (MPI) video, the apparatus comprising: At least one processor; as well as At least one memory containing program code; and The at least one memory and the program code are configured, together with the at least one processor, to cause the device to perform at least the following operations: Generate a sequence of video frames, each of which includes a plurality of corresponding tiles of a texture layer and a transparency layer representing one or more multiplanar images of the MPI video; Video compression is applied to the video frame sequence to generate a video substream; Generate a representation sequence of atlas frames corresponding to the video frame sequence, to at least specify the packing arrangement of the tiles; Compression is applied to the representation sequence to generate a graph atlas substream; and The video substream and the atlas substream are multiplexed to generate a first coded bitstream that encodes at least a portion of the MPI video.