Imaging file format for multi-plane images with HEIF
Patent Information
- Application Number
- KR1020267026642
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-16
- Filing Date
- 2025-01-14
- Publication Date
- 2026-09-21
Smart Images

Figure P1020267026642_ABST
Abstract
Description
Technology Field
[0001] Cross-reference regarding related applications
[0002] The present application claims priority to U.S. Partial Continuation Application (CIP) No. 18 / 917,891 filed on October 16, 2024, which claims priority to PCT Application No. PCT / US2024 / 024133 filed on April 11, 2024, which claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 621,455 filed on January 16, 2024; and the present application also claims priority to U.S. Partial Continuation Application (CIP) No. 18 / 671,633 filed on May 22, 2024, which claims priority to PCT Application No. PCT / US2024 / 024017 filed on April 11, 2024. The entire disclosures of the above applications are incorporated by reference.
[0003] technology
[0004] This document generally relates to images and videos. More specifically, embodiments of the present invention relate to imaging formats for multi-planar images. Background Technology
[0005] Multiplane Imaging (MPI) implements a relatively new approach to storing volumetric content. MPI can be used to render both still images and videos, for example, to represent a three-dimensional (3D) scene within a view frustum using 8, 16, 32, or more planes of texture and transparency (or opacity) (alpha) information per camera. This representation stores parallel planes of the scene at fixed ranges of depths discretely sampled from a reference coordinate frame. The information stored in each plane includes texture (e.g., in terms of RGB values) and opacity (in terms of the alpha (A) channel). Exemplary applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality.
[0006] The High Efficiency Image File Format (HEIF) (Reference [1]) enables the encapsulation of images and image sequences, as well as their associated metadata, into container files. HEIF is compatible with the ISO Base Media File Format (ISOBMFF) (Reference [3]). HEIF includes specifications for encapsulating images and image sequences in accordance with the High Efficiency Video Coding specification (HEVC, ISO / IEC 23008-2 | ITU-T Rec. H.265). As recognized by the inventors, a new file format for representing MPI images using HEIF is described herein.
[0007] In this specification, the term “metadata” relates to any auxiliary information transmitted as part of or together with a coded bitstream to assist a decoder in rendering or interpreting one or more decoded images. While the examples presented in this specification refer to the HEIF format, those skilled in the art will understand that the techniques discussed in this specification are applicable to any file container that supports the transmission of image and video content.
[0008] The approaches described in this section are approaches that may be pursued, but are not necessarily approaches that have been previously conceived or pursued. Accordingly, unless otherwise indicated, none of the approaches described in this section should be assumed to qualify as prior art merely by their inclusion within this section. Similarly, problems identified in relation to one or more of the approaches should not be assumed to be recognized in any prior art based on this section, unless otherwise indicated. Brief explanation of the drawing
[0009] Embodiments of the present invention are illustrated in the accompanying drawings by way of example rather than by way of limitation, and similar reference numerals in the accompanying drawings refer to similar elements. In the drawings: Figure 1 illustrates an exemplary process for a HEIF player given an input HEIF file. Figure 2 illustrates an exemplary MPI representation using D layers of RGBA images for a single camera view. Figure 3 illustrates an example of spatial packing of K x M layers of an MPI image. FIG. 4 illustrates an exemplary process flow for rendering a HEIF file including an MPI representation according to one embodiment of the present invention. FIG. 5 illustrates an exemplary workflow for a HEIF player that supports the existence of alternative media representations within the same file according to one embodiment of the present invention. FIG. 6 illustrates the encapsulation of an MPI image as a V3C bitstream within a file according to an embodiment of the present invention. FIG. 7 illustrates an example of the operation of a player when receiving an MPI bitstream encapsulated within a HEIF file that supports V3C representation according to an embodiment of the present invention. Specific details for implementing the invention
[0010] Description of exemplary embodiments
[0011] Exemplary embodiments relating to imaging file formats for MPI are described herein. In the following description, for the purposes of explanation, a number of specific details are provided to provide a complete understanding of various embodiments of the invention. However, it will be apparent that various embodiments of the invention can be practiced without these specific details. In other cases, well-known structures and devices are not described in full detail to avoid unnecessarily obscuring, obscuring, or confusing embodiments of the invention.
[0012] summation
[0013] The exemplary embodiments described herein relate to imaging file formats for MPI. The embodiments discuss ways to extend the HEIF file format to support MPI images. The present disclosure provides examples in which MPI texture and opacity information can be coded as a single packed image or as two separate images using both HEVC and VVC codecs. Exemplary embodiments for parsing and decoding MPI images in HEIF are also provided. Examples of carrying MPI metadata according to the T.35 protocol and Visual Volumetric Video-based Coding (V3C), and exemplary HEIF-based players that support these media representations are also presented.
[0014] High Efficiency Image File Format (HEIF)
[0015] The High Efficiency Image File Format (HEIF) (References [1-2]) enables the encapsulation of images and image sequences, as well as their associated metadata, into container files. HEIF is compatible with the ISO-based media file format (Reference [3]). HEIF includes specifications for encapsulating images and image sequences that conform to the High Efficiency Video Coding Standard, also known as HEVC or H.265.
[0016] In ISOBMFF, continuous or timed media or metadata streams form tracks, while static media or metadata is stored as items. Consequently, in HEIF, still images are stored as items. All image items are coded independently and do not depend on any other items for their decoding. Any number of image items can be included in the same file. Image sequences are stored as tracks. Image sequence tracks are used when there are coding dependencies between images or when the playback of images is timed. In contrast to video tracks, timing in image sequence tracks is advisory.
[0017] Table 1 describes the hierarchy of boxes in HEIF. The handler type for a MetaBox will be 'pict' to indicate to the reader that this MetaBox handles images. In HEIF, item attributes that can be used to describe image items or influence the generation of output images, such as spatial ranges of image items and color information, are stored.
[0018]
[0019] Table 2 illustrates an example of a single coded image item containing Interchangeable Image File Format (Exif) metadata stored in HEIF. File metadata for the items is stored within a meta box ('meta'). The handler type is set to 'pict', which indicates to the reader that this meta box handles images. The coded images are stored as items of "hvc1" representing HEVC-coded data. The coded data for the images is contained in a media data box ('mdat') or an item data box ('idat'). The syntax of the 'hvc1' item consists of Network Abstraction Layer (NAL) units of the HEVC bitstream, and the bitstream contains exactly one access unit. All configuration information required to initialize the decoder (e.g., parameter sets and information regarding the coding itself) is stored as an item of type 'hvcC' (in the case of HEVC-coded images). The width and height of the associated image are stored as an item of type 'ispe'. The associations between items and attributes are displayed in the ItemPropertyAssociationBox("ipma"). Each attribute association can be marked as required or non-required. The reader should not process items marked as required for an item that are associated with attributes not recognized or supported by the reader. The reader may ignore associated item attributes marked as non-required for an item. Exif metadata for an image is optionally included in the file as an item of type "Exif" and is linked to the image item using the 'cdsc' reference type within the item reference box("iref").
[0020]
[0021] FIG. 1 illustrates an exemplary process of how a HEIF-compatible player processes coded images (110) and derived images (115) contained in a file. The HEIF player decodes the coded image (110) into a reconstructed image using decoder configuration (130) information. Similarly, the HEIF player obtains each reconstructed image by applying the derivation operation characteristics (105) of the derived image (115) to one or more displayed input images. Descriptive image characteristics generally describe the reconstructed image, excluding decoder configuration and initialization information associated with the coded image. Transformed image characteristics (120), if present, are applied to the reconstructed image to obtain an output image. The output image may be displayed when the coded image or the derived image is not a hidden image. The output image may also serve as an input image (125) for the derived images (115).
[0022] Multiplane Image (MPI) Scene Representation and Packing
[0023] A multi-plane image comprises multiple image planes, each of which is a "snapshot" of a 3D scene at a specific depth relative to the camera position. The information stored in each plane includes texture information (e.g., represented by R, G, and B values) and transparency (or opacity) information (e.g., represented by alpha (A) values). Here, the acronyms R, G, and B represent red, green, and blue, respectively. In some examples, the three texture components may be (Y, Cb, Cr), or (I, Ct, Cp), or a set of other functionally similar values. There are different ways in which a multi-plane image can be generated. For example, two or more input images from two or more cameras located at different known viewpoints can be co-processed to generate a corresponding multi-plane image. Alternatively, a multi-plane image can be generated using a source image captured by a single camera.
[0024] FIG. 2 illustrates a 3D scene representation using a multi-plane image (200) according to one embodiment. The multi-plane image (200) has D planes or layers (P0, P1, ..., P(D-1)), where D is an integer greater than 1. Typically, the planes (layers) are indexed such that the layer furthest from the reference camera position (RCP) is indexed as the 0th layer and is indexed so that it is at a distance (or depth) d0 from the RCP along the Z dimension of the 3D scene. The index is incremented by 1 for each next layer located closer to the RCP. The plane (layer) closest to the RCP has an index value (D-1) and is at a distance (or depth) d0 from the RCP along the Z dimension. D-1Each of the planes (P0, P1, ..., P(D-1)) is orthogonal to the base plane (202) parallel to the XZ coordinate plane. The RCP is located at a vertical height h above the base plane (202). The XYZ triad shown in FIG. 2 represents the general orientation of the planes (P0, P1, ..., P(D-1)) and the multi-plane image (200) for the X, Y, and Z dimensions of the 3D scene. In various examples, the number D can be 32, 16, 8, or any other suitable integer greater than 1.
[0025] The color component (e.g., RGB) values for the i-th layer at camera position s It is denoted as such, where the lateral size of the hierarchy is HxW, H is the height of the hierarchy (Y dimension), and W is the width of the hierarchy (X dimension). The pixel value at position (x, y) for color channel c is It is expressed as. The α value for the i-th layer is The pixel value (x, y) in the alpha layer is It is expressed as. The depth distance between the i-th layer and the reference camera position is d i is. The image from the original reference view (without camera movement) is It is displayed as, and the texture pixel value is. Therefore, a still MPI image for a camera position(s) can be expressed as follows.
[0026]
[0027] If the camera position s remains static over time, extending this still MPI image representation to a video representation is straightforward. This video representation is given by Equation 2:
[0028]
[0029] Here, t represents time.
[0030] As already shown above, a multi-plane image such as a multi-plane image (200) is a single source image It can be generated from or from two or more source images. Such generation can be performed, for example, during a production phase. The corresponding MPI generation algorithm(s) typically generate a multi-plane image (200) containing XYZ-resolved pixel values { , where i = 0, ..., D-1} can be output in the form.
[0031] { By processing a multi-plane image (200) represented by {where i = 0, ..., D-1}, the MPI rendering algorithm can generate a viewable image corresponding to a new virtual camera position different from or to the RCP. An exemplary MPI-rendering algorithm (often referred to as an "MPI viewer") that can be used for this purpose may include warping and compositing steps. Other suitable MPI viewers may also be used. The rendered multi-plane image (200) can be viewed on a display.
[0032] In one embodiment, given an MPI representation, the texture and opacity maps of the MPI layers are first spatially packed into a K x M array to form 2D compositional pictures as shown in FIG. 3. These two compositional pictures may be spatially packed into the same frame as side-by-side or top-bottom images, or they may be two separate texture and opacity map pictures.
[0033] Saving MPI images within HEIF
[0034] In an exemplary embodiment, the texture and opacity maps of the MPI layers are encoded using a 2D video codec such as HEVC (H.265) or VVC (Versatile Video Coding) (or H.266), and the encoded images are stored as items in HEIF. To enable the player to reconstruct the volumetric MPI representation from the decoded images, the following MPI metadata is also stored in HEIF.
[0035] Number of layers to be used in the MPI representation
[0036] Packing information describing how textures and opacity maps of MPI layers are packed into pictures
[0037] Depth information of MPI layers
[0038] MPI post-processing specific information when it needs to be applied
[0039] Exogenous and endogenous camera information (optional)
[0040] This MPI metadata is stored as image item attributes defined below.
[0041] Storage of MPI metadata within item attributes
[0042] In one embodiment, MPIInformationProperty is used to describe MPI image items or to influence the generation of output MPI images. MPIInformationProperty describes the number of MPI layers within the decoded pictures, the depth of each layer, and texture and opacity map packing and arrangement information.
[0043] CameraExtrinsicMatrixProperty (specified in reference [2]) is optionally used to describe the position in the Cartesian representation and the orientation of the camera capturing the associated image item. CameraIntrinsicMatrixProperty is used to describe the properties of the camera capturing the associated image item.
[0044] MPI Information Characteristics
[0045] definition
[0046] Box type: 'mpii'
[0047] Attribute Type: Description Item Attribute
[0048] Container: ItemPropertyContainerBox
[0049] Required (per item): Yes
[0050] Quantity (per item): 1
[0051] MPIInformationProperty contains MPI metadata of the associated image item. This includes the number of MPI layers within the pictures, the depth of each layer, and texture and opacity map packing and arrangement information. In one embodiment, without limitation, exemplary syntax is given by the following:
[0052] Exemplary syntax
[0053]
[0054] The following semantics are defined:
[0055] mpii_num_layers_minus1 + 1 specifies the number of texture and opacity layers for MPI representation.
[0056] mpii_layer_packing_order indicates the order of the two configuration decoding pictures.
[0057] mpii_layer_packing_order being 0 indicates that the first decoded picture of the constituent pictures is the texture map of the MPI layers and the second picture is the opacity map of the MPI layers. mpii_layer_packing_order being 1 indicates that the first picture of the constituent pictures is the opacity map of the MPI layers and the second picture is the texture map of the MPI layers.
[0058] mpii_layer_packing_type represents the scheme of the packing array of texture and opacity maps of MPI layers within the decoded pictures as defined in Table 3.
[0059]
[0060] When mpii_layer_packing_order is 0 and mpii_layer_packing_type is 0, it indicates that texture and opacity map packing is top-down. When mpii_layer_packing_order is 0 and mpii_layer_packing_type is 1, it indicates that texture and opacity map packing is side-by-side. When mpii_layer_packing_order is 0 and mpii_layer_packing_type is 2, it indicates two separate pictures, where the first picture is a texture map and the second picture is an opacity map.
[0061] mpii_pic_num_layers_in_height_minus1 + 1 specifies the number of spatially packed layers in height for picture 0 and picture 1.
[0062] mpii_pic_num_layers_in_width_minus1 + 1 specifies the number of spatially packed layers in width for Picture 0 and Picture 1. This is equivalent to (mpii_num_layers_minus1 + 1) / (mpii_pic_num_layers_in_height_minus1 + 1).
[0063] mpii_layer_depth_equal_distance_flag being 0 indicates that depth information for each layer is signaled as follows. mpii_layer_depth_equal_distance_flag being 1 indicates that equal distances are used to generate MPI layers, the nearest depth and the farthest depth are signaled, and depth information for each layer Z[i] can be derived using the nearest depth value (ZNear) and the farthest depth value (ZFar).
[0064] Depth value Z[i] for the i-th MPI layer =
[0065] = i * (ZFar - ZNear) / (mpi_num_layers_minus1) + ZNear
[0066] The depth_rep_info_element(OutSign, OutExp, OutMantissa, OutManLen) syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen, which represent floating-point values.
[0067] A da_sign_flag of 0 indicates that the sign of the floating-point value is positive. A da_sign_flag of 1 indicates that the sign is negative. The variable OutSign is set to be the same as da_sign_flag.
[0068] da_exponent specifies the exponent of a floating-point value. The value of da_exponent is 0 to 2 7 It must be in the range of -2. The variable OutExp is set to be the same as da_exponent.
[0069] da_mantissa_len_minus1 + 1 specifies the number of bits within the da_mantissa syntax element. The variable OutManLen is set to be equal to da_mantissa_len_minus1 + 1.
[0070] da_mantissa specifies the mantissa of a floating-point value. The variable OutMantissa is set to be the same as da_mantissa.
[0071] Note that, without limitation, some of the MPI metadata syntax elements proposed in this specification may match the names of the syntax elements proposed in reference [4] for transmitting MPI metadata via complementary enhancement information (SEI) messaging. The proposed syntax parameters may be adapted to other metadata formats.
[0072] Storage of MPI metadata within metadata items
[0073] XMP metadata is stored as an item with the item_type value 'mime' and the content type 'application / rdf+xml'. The body of the item is a valid XMP document containing the elements previously described under the "MPI information attribute" and the optional elements described in the XML-form CameraExtrinsicMatrixProperty (specified in reference [2]).
[0074] XMP metadata items are linked to image items by item references of type 'cdsc'.
[0075] Table 4 provides an example of an XMP file describing metadata for an MPI scene with 16 MPI layers using a side-by-side packing array in a 4x4 configuration.
[0076]
[0077] <mpi:depthsign> , <mpi:depthexponent> , <mpi:depthmantissa>Note that it corresponds to the parameters da_sign_flag, da_exponent, and da_mantissa as previously defined.
[0078] Storage of MPI data within ITU T.35
[0079] T.35 metadata items carry ITU-T T.35 messages (Reference [5]). When T.35 metadata is stored as a metadata item, the item_type value is the same as 'it35'.
[0080]
[0081] itu_t_t35_country_code is a byte with a value specified as a country code by Rec. ITU-T T.35 Annex A, or a country code extended value 0xFF.
[0082] itu_t_t35_country_code_extension_byte is a byte that has a value specified as a country code by Rec. ITU-T T.35 Annex B, if present.
[0083] itu_t_t35_payload contains a payload containing data. The ITU-T T.35 terminal provider code and terminal provider-oriented code are included in the first one or more bytes of itu_t_t35_payload in a format specified by the administrator who issued the terminal provider code. Any remaining itu_t_t35_payload data is data having syntax and semantics as specified by the entity identified by the ITU-T T.35 country code, terminal provider code, and terminal provider-oriented code. itu_t_t35_payload contains the elements described above or the SEI messages below. The length of this field is the number of bytes remaining in the item. An exemplary MPI SEI message and its semantics are given below (see Reference [4]).
[0084]
[0085] mpii_num_layers_minus1 + 1 specifies the number of texture and opacity layers for MPI representation.
[0086] mpii_layer_depth_equal_distance_flag being 1 indicates that equal distances are used to generate MPI layers and depth parameters for each layer.
[0087] mpii_texture_opacity_interleave_flag being 1 indicates that the decoded output pictures correspond to texture and opacity composition pictures that are temporally interleaved in the output order. mpii_texture_opacity_interleave_flag being 0 indicates that the decoded output pictures correspond to texture and opacity composition pictures that are spatially packed.
[0088] mpii_texture_opacity_arrangement_flag being 0 indicates that the decoded output pictures represent texture and opacity composition pictures in a top-to-bottom packed array. mpii_texture_opacity_arrangement_flag being 1 indicates that the decoded output pictures represent texture and opacity composition pictures in a side-to-side packed array.
[0089] mpii_picture_num_layers_in_height_minus1 + 1 specifies the number of spatially packed layers in height for picture 0 and picture 1.
[0090] Encapsulation of packed textures and opacity maps in HEIF
[0091] This section describes a format for encapsulating the coded images of packed textures and opacity maps of MPI layers, and the related MPI metadata described in "Storage of MPI metadata within item attributes" in HEIF.
[0092] Spatially packed textures and opacity maps of MPI layers are encoded using a 2D video codec, e.g., HEVC, VVC, etc., and the encoded images are stored as items. 2D video decoder configuration and initialization are stored in decoder configuration properties and are set to be mandatory. When a specific complement enhancement information (SEI) message containing MPI-specific information is present in the bitstream (e.g., see reference [4]), the SEI message is carried in decoder configuration properties. MPI metadata is stored in the MPIInformationProperty (defined earlier in "Storage of MPI metadata in item properties") or as metadata items (defined earlier in "Storage of MPI metadata in metadata items").
[0093] An exemplary format for encapsulating HEVC-coded MPI images
[0094] This section describes an exemplary format for encapsulating HEVC-coded images containing spatially packed textures and opacity maps of MPI layers within HEIF. HEVC-coded images are stored as items of 'hvc1' representing HEVC-coded data. HEVC image items of type 'hvc1' contain independently coded HEVC images of spatially packed textures and opacity maps of MPI layers arranged vertically or side-by-side. Items of type 'hvc1' consist of NAL units of the coded HEVC image bitstream containing exactly one access unit. All configuration information required to initialize the decoder (e.g., parameter sets and information regarding the coding itself) is stored as 'hvcC' attributes. Each HEVC image item of type 'hvc1' must have an association with the 'hvcC' attribute. When MPI metadata is stored in the MPIInformationProperty, HEVC image items of type 'hvc1' must have an association with the MPIInformationProperty. For MPIInformationProperty, essential must be 1. Optionally, CameraExtrinsicMatrixProperty (specified in Reference [2]) exists to describe the position in Cartesian representation and the orientation of the camera capturing the associated image item. CameraIntrinsicMatrixProperty exists to describe the characteristics of the camera capturing the associated image item. When both exist, both are associated with HEVC image items. Table 5 illustrates the encapsulation of a single HEVC-coded image in HEIF. HEVC-coded images are stored as items of 'hvc1'. The coded data for the image is contained in the media data box ('mdat') or the item data box ('idat').The width and height of the associated image are stored as item properties of type 'ispe'. The MPI metadata of the associated image, the number of MPI layers, the depth of each layer, and the packing and arrangement of texture and opacity maps within the picture are stored as item properties of type 'mpii'. Exogenous and endogenous camera information are stored as item properties of types 'cmex' and 'cmin', respectively. The association between the image item and the image properties is indicated in ItemPropertyAssociationBox('ipma'). Since decoder configuration and MPI metadata must be processed, image properties of 'hvcC' and 'mpii' are marked as essential. The player must not process items associated with properties marked as essential that are unrecognized or unsupported.
[0095]
[0096] As previously described in "Storage of MPI metadata within metadata items," when MPI metadata is stored in XMP metadata items as items with the item_type value 'mime' and content type 'application / rdf+xml', the XMP metadata items are linked to image items by item references of type 'cdsc'. A HEIF file containing a single coded image item and XMP metadata is structured as shown in Table 6:
[0097]
[0098] As previously explained, when MPI metadata is stored in a T.35 metadata item as an item with the item_type value 'it35', the T.35 metadata item is linked to image items by item references of type 'cdsc'. A HEIF file containing a single coded image item and a T.35 metadata item is structured as follows:
[0099]
[0100] An exemplary format for encapsulating VVC-coded MPI images with MPI metadata
[0101] This section describes an exemplary format for encapsulating a VVC-coded image containing spatially packed textures and opacity maps of MPI layers within HEIF.
[0102] VVC-coded images are stored as items of 'vvc1' representing VVC-coded data. VVC image items of type 'vvc1' contain independently coded VVC images of spatially packed textures and opacity maps of MPI layers arranged vertically or side-by-side. Items of type 'vvc1' consist of NAL units of the coded VVC image bitstream containing the entire VVC access unit. All VVC decoder configuration information required to initialize the decoder (e.g., parameter sets and information regarding the coding itself) is stored as 'vvcC' attributes. Each VVC image item of type 'vvc1' must have an association with a 'vvcC' attribute.
[0103] When MPI metadata is stored in MPIInformationProperty, each VVC image item of type 'vvc1' must have an association for MPIInformationProperty. essential must be 1 for the MPIInformationProperty associated with the image item of type 'vvc1'.
[0104] Optionally, CameraExtrinsicMatrixProperty and CameraIntrinsicMatrixProperty exist to describe the properties of the camera capturing the associated image item. When both exist, essential is 0.
[0105] Table 7 illustrates the encapsulation of a single VVC-coded image within HEIF. The VVC-coded image is stored as items of 'vvc1'. MPI metadata is stored as item properties of type 'mpii'. Exogenous and endogenous camera information is stored as item properties of types 'cmex' and 'cmin', respectively. The association between the VVC image item and the image properties is indicated in ItemPropertyAssociationBox('ipma'). Since decoder configuration and MPI metadata must be processed, image properties of 'hvcC' and 'mpii' are marked as required. The player must not process items associated with properties marked as required that are unrecognized or unsupported.
[0106]
[0107] As previously described in "Storage of MPI metadata within metadata items," when MPI metadata is stored in XMP metadata items as an item with the item_type value 'mime' and content type 'application / rdf+xml', the XMP metadata items are linked to image items by item references of type 'cdsc'. A HEIF file containing a single VVC-coded image item and XMP metadata is structured as shown in Table 8.
[0108]
[0109] As previously explained, when MPI metadata is stored in a T.35 metadata item as an item with the item_type value 'it35', the T.35 metadata item is linked to image items by item references of type 'cdsc'. A HEIF file containing a single coded image item and a T.35 metadata item is structured as follows:
[0110]
[0111] Exemplary encapsulation in HEIF using separate texture and opacity images
[0112] From an MPI representation, two pictures of a texture and an opacity map can be generated as described in FIG. 3. This section describes a format for encapsulating two coded images of textures and opacity maps of MPI layers and their associated MPI metadata, as described above in “Storage of MPI metadata within item attributes” in HEIF.
[0113] Two distinct texture and opacity maps of the MPI layers are encoded independently using 2D video codecs, such as HEVC, VVC, etc., and the encoded images are stored as items. The texture map is stored as the master image, and the opacity map is stored as an auxiliary image indicating that it contains an alpha plane relative to the master image. The auxiliary image and the master picture are linked from the auxiliary image to the master image using the item reference of 'auxl'. The auxiliary image of the opacity map is associated with an AuxiliaryTypeProperty (specified in reference [1]) that identifies the type of the auxiliary image as an alpha plane.
[0114] As described above in "Encapsulation of Packed Textures and Opacity Maps in HEIF," decoder configuration and initialization are stored in decoder configuration properties and marked as essential. When a specific SEI message containing MPI-specific information exists within the bitstream, the SEI message is carried within the decoder configuration properties. MPI metadata is stored in the MPIInformationProperty or the previously defined metadata items.
[0115] An exemplary format for encapsulating two HEVC-coded MPI-related images
[0116] This section describes a format for encapsulating two HEVC-coded images, one containing texture maps and the other containing opacity maps of MPI layers within HEIF.
[0117] Both HEVC-coded texture images and opacity images are stored as items of 'hvc1'. Each HEVC image item of type 'hvc1' contains a separate coded HEVC image bitstream, each containing exactly one access unit for the texture map and the opacity map, respectively. All decoder configuration information required to initialize the decoder (e.g., parameter sets and information regarding the coding itself) is stored as an 'hvcC' attribute. Each HEVC image item of type 'hvc1' must have an association with an 'hvcC' attribute.
[0118] The texture map is stored as a master HEVC image item of type 'hvc1', and the opacity map is stored as a HEVC auxiliary image item of 'hvc1', indicating that it contains an alpha plane for the master image. The auxiliary opacity image and the master texture image are linked from the auxiliary image to the master image using an item reference of 'auxl'. The auxiliary image of the opacity map is associated with an AuxiliaryTypeProperty 'auxC', which identifies the type of the auxiliary image as an alpha plane, for example, by "urn:mpeg:mpegB:cicp:systems:auxiliary:alpha" as the aux_type value.
[0119] When MPI metadata is stored in MPIInformationProperty, each HEVC image item of type 'hvc1' must have an association with MPIInformationProperty. essential must be 1 for MPIInformationProperty. Optionally, CameraExtrinsicMatrixProperty and CameraIntrinsicMatrixProperty exist, both of which are associated with the master texture HEVC image item.
[0120] A HEIF file containing two independently HEVC-coded images—one containing a texture map and the other containing an opacity map of MPI layers—is structured as follows. Individual pictures are encoded as HEVC-coded images and stored as items of 'hvc1'. The opacity map is stored as a HEVC auxiliary image item by indicating that it contains an alpha plane via the 'auxC' image properties. The auxiliary opacity image and the master texture image are linked from the auxiliary image to the master image using the item reference of 'auxl'. MPI metadata is stored as an item property of type 'mpii' and is marked as essential as it needs to be processed. Exogenous and endogenous camera information is stored as item properties of types 'cmex' and 'cmin', respectively. The association between image items and image properties is indicated in ItemPropertyAssociationBox('ipma'). Since MPI metadata needs to be processed for both the master texture map and the auxiliary opacity map, the 'mpii' item property needs to be associated with both image items. An exemplary explanation is given in Table 9.
[0121]
[0122] When MPI metadata is stored in XMP metadata items, the HEIF file is structured as shown in Table 10.
[0123]
[0124] As previously mentioned, when MPI metadata is stored in a T.35 metadata item as an item with the item_type value 'it35', the T.35 metadata item is linked to two image items containing a texture map or an opacity map by item references of type 'cdsc'. A HEIF file having two coded image items and a T.35 metadata item is structured as follows:
[0125]
[0126] An exemplary format for encapsulating two VVC-coded MPI-related images
[0127] This section describes a format for encapsulating two VVC-coded images, one containing texture maps and the other containing opacity maps of MPI layers within HEIF.
[0128] Each VVC image item of type 'vvc1' contains an individual coded VVC image bitstream, each containing exactly one access unit for a texture map and an opacity map. All decoder configuration information required to initialize the decoder (e.g., parameter sets and information regarding the coding itself) is stored as the 'vvcC' attribute. Each VVC image item of type 'vvc1' must have an association with the 'vvcC' attribute.
[0129] The texture map is stored as a master VVC image item of type 'hvc1', and the opacity map is stored as a VVC auxiliary image item of 'vvc1', indicating that it contains an alpha plane relative to the master image. The auxiliary opacity image and the master texture image are linked from the auxiliary image to the master image using an item reference of 'auxl'. The auxiliary image of the opacity map is associated with an AuxiliaryTypeProperty 'auxC', which identifies the type of the auxiliary image as an alpha plane, for example, by "urn:mpeg:mpegB:cicp:systems:auxiliary:alpha" as the aux_type value.
[0130] When MPI metadata is stored in MPIInformationProperty, each VVC image item of type 'vvc1' must have an association with MPIInformationProperty. essential must be 1 for MPIInformationProperty. Optionally, CameraExtrinsicMatrixProperty and CameraIntrinsicMatrixProperty exist, both of which are associated with the master texture VVC image item.
[0131] Table 11 illustrates the encapsulation of two VVC-coded images, a texture and an opacity map, within HEIF. Each individual picture is coded as a VVC-coded image and stored as items of 'vvc1'. The opacity map is stored as a VVC auxiliary image item, indicating that it contains an alpha plane via the 'auxC' image properties. The auxiliary opacity image and the master texture image are linked from the auxiliary image to the master image using the item reference of 'auxl'. MPI metadata is stored as an item property of type 'mpii' and is marked as essential as it needs to be processed. Exogenous and endogenous camera information is stored as item properties of types 'cmex' and 'cmin', respectively. The association between the image items and image properties is indicated in ItemPropertyAssociationBox('ipma'). Since MPI metadata needs to be processed for both the master texture map and the auxiliary opacity map, the 'mpii' item property needs to be associated with both image items.
[0132]
[0133] Table 12 illustrates an example of a HEIF file when MPI metadata is stored in XMP metadata items.
[0134]
[0135] As previously mentioned, when MPI metadata is stored in a T.35 metadata item as an item with the item_type value 'it35', the T.35 metadata item is linked to two image items containing texture maps and / or opacity maps by item references of type 'cdsc'. A HEIF file having two coded image items and a T.35 metadata item is structured as follows:
[0136]
[0137] Exemplary player actions
[0138] This section provides an exemplary embodiment of a player action having input from a HEIF file containing coded MPI image(s) and associated MPI metadata. Figure 4 illustrates an example of a player action and how it generates a rendered output suitable for the user viewport from the input of the HEIF file.
[0139] At step 405, the player begins by determining whether the input HEIF file is fully supported. The player must not process image items associated with features marked as essential that are unrecognized or unsupported. When the player supports the essential features within the input file, at step 410, the player begins parsing the input HEIF file and extracts coded image items and decoder configuration information from the input HEIF file. The player initializes the decoder using the extracted decoder configuration information and, accordingly, decodes the coded images into a single decoded item or two decoded images (420). If transformation image features exist, the player applies the transformation features to the decoded images (425). Additionally, by using the MPI metadata extracted from the input HEIF file, the player obtains the texture and opacity maps of the MPI layers (427).
[0140] When the decoded image is a spatially packed texture and opacity map, the player recognizes the locations of the texture and opacity maps of the MPI layers within the decoded frame based on the MPI metadata. Additionally, when two decoded images are decoded, the player recognizes which one is a texture map and which is an opacity map of the MPI layers by using associated image properties, and obtains the texture and opacity maps of the MPI layers. Subsequently, an MPI scene representation is reconstructed from the texture and opacity maps of the MPI layers, and the MPI metadata includes the depth of each layer. To preserve real-world coordinates and synchronize multiple cameras, camera information regarding the implicit and extrinsic matrices from the MPI metadata is used in the warping process. When neither is available, the renderer can perform view composition using a predefined universal camera eigenmatrix. After reconstruction, a rendered output suitable for the user viewport is generated and displayed.
[0141] Carrying multiple representations in a single HEIF
[0142] A HEIF file may contain image items representing replacements of the same source within the same replacement group. In this case, the following EntityToGroupBox with grouping_type 'altr' exists in the GroupsListBox because the GroupsListBox contains EntityToGroupBoxes that each specify a single entity group.
[0143]
[0144] An EntityToGroupBox with grouping_type 'altr' represents a set of images that are substitutes for each other, of which only one is selected for display or processing.
[0145] The table below describes the case where 2D HEVC images and MPI images encoded in HEVC are carried in the same HEIF file. MPI images are stored as HEVC image items, and HEVC image items are associated with the aforementioned MPI metadata item attribute ('mpii' item attribute). MPI metadata can be carried in XMP metadata items or T.35 metadata items as previously described. In this case, the XMO metadata items or T.35 metadata items are linked to the HEVC image items by item references of type 'cdsc'. Subsequently, these two HEVC image items are represented as substitutes for each other by using an EntityToGroupBox with grouping_type 'altr'.
[0146]
[0147] FIG. 5 illustrates an exemplary workflow for a HEIF player that supports the existence of alternate media representations within the same file according to one embodiment. The HEIF player recognizes the existence of alternate representations within the file through the 'altr' grouping type, and if alternate representations exist, the player appropriately selects one of the alternate media representations according to the player's capabilities. For example, when the player supports MPI playback, the player checks for the existence of alternates and selects an MPI representation from among the alternates. If alternates do not exist, the player first checks whether the input file is a supported MPI representation. If the input file is supported, the player decodes the coded image, reconstructs the MPI using MPI metadata extracted from the file, and renders it appropriately according to the current viewport. Otherwise, when the player supports only 2D image playback, the player recognizes the existence of multiple alternates, and if alternates exist, selects an appropriate 2D image representation from among the alternates and processes it accordingly. If alternates do not exist, the player checks whether the current input file can be processed and, accordingly, processes or terminates the current input file.
[0148] Transport of V3C-coded MPI images in HEIF
[0149] Visual Volumetric Video-based Coding (V3C) provides a mechanism for coding visual volumetric frames. Visual volumetric frames are coded by converting 3D volumetric information into a collection of 2D images and associated data. The converted 2D images are coded using widely available video and image coding specifications and associated data; that is, video data can be coded using HEVC or VVC, and atlas data is coded according to ISO / IEC 23090-5. The coded images and coded atlas data are multiplexed to form a V3C bitstream.
[0150] This section specifies exemplary embodiments of a format for encapsulating untimed multi-plane image (MPI) data into a HEIF file. FIG. 6 illustrates an exemplary embodiment of the encapsulation of an MPI image as a V3C bitstream within a HEIF file (600). The MPI image is coded as a V3C bitstream consisting of one or more V3C units, each V3C unit comprising a video data unit containing the coded MPI image or coded atlas data containing associated MPI metadata. The V3C bitstream is stored in a V3C item (605). As defined in References [1-2], associated V3C decoder configuration information is carried in a V3C configuration item characteristic (610), while 2D video decoder configuration information is carried in a corresponding video decoder configuration item characteristic (615), and subsample information is carried in a subsample item characteristic (620).
[0151] The handler type for the MetaBox must be 'volv' to indicate the presence of V3C items. A V3C item is an item representing a single visual volumetric video frame of a coded MPI image. A V3C item contains one or more V3C units of a coded MPI image. Items of type 4CC codes 'v3e1' identify V3C items.
[0152] Items of type 'v3e1' must be associated with a single V3CConfigurationProperty. Items of type 'v3e1' can be associated with a 2D video decoder configuration item property, such as a subsample item property of type 'subs', an HEVC configuration item property of type 'hvcC', or a VVC configuration item property of type 'vvcC'.
[0153] If PrimaryItemBox exists, the item_ID within this box should be set to display V3C items of type 'v3e1'. Exemplary syntax:
[0154]
[0155] In the syntax above, the following applies:
[0156] The value of item_size is equal to the sum of the extent_length values of each extent of the item, as specified in ItemLocationBox.
[0157] v3c_config represents the record within the associated V3CConfigurationProperty.
[0158] v3c_unit_size specifies the size of the ss_v3c_unit array in bytes. This size is equivalent to the sample stream V3C unit size ssnu_v3c_unit_size as defined in ISO / IEC 23090-5, Annex C.
[0159] ss_v3c_unit contains a single V3C unit in the V3C unit sample stream format as defined in ISO / IEC 23090-5:2021, Annex C.
[0160] V3C Item Characteristics
[0161] common
[0162] Two description item attributes are defined: the V3C configuration item attribute carries V3C decoder configuration and initialization information, and the subsample item attribute contains subsample information such as the offset of each subsample containing a V3C unit to enable subsample-by-subsample access.
[0163] V3C component item characteristics
[0164] definition
[0165] Box types: 'v3cC'
[0166] Attribute Type: Description Item Attribute
[0167] Container: ItemPropertyContainerBox
[0168] Required (per item): Yes, for V3C items of type 'v3e1'
[0169] Quantity (per item): 0 or more for coded image items
[0170] V3CConfigurationProperty contains a V3C decoder configuration record that provides decoding-specific information (i.e., parameter sets and SEI messages) of the V3C bitstream for further configuration and initialization of the V3C decoder, as described later. V3CConfigurationProperty must be associated with the 'v3e1' V3C item. The V3C configuration item attribute is an essential attribute, and the corresponding essential flag in ItemProperyAssociationBox must be set to 1 for the 'v3cC' item attribute.
[0171] Syntax
[0172]
[0173] Semantics
[0174] v3c_config includes a single instance of V3CDecoderConfigurationRecord that provides decoding-specific information of the V3C bitstream (i.e., parameter sets and SEI messages) for the additional configuration and initialization of the V3C decoder.
[0175] Subsample Item Characteristics
[0176] definition
[0177] Box type: 'subs'
[0178] Attribute Type: Description Item Attribute
[0179] Container: ItemPropertyContainerBox
[0180] Required (per item): No
[0181] Quantity (per item): 0 or 1 for coded image items
[0182] Subsample information for a coded V3C image can be provided for the coding format of the associated coded image item using the exact same associated item properties as SubSampleInformationBox as defined later.
[0183] The entry_count field of SubSampleInformationBox must be 1, and the sample_delta field of SubSampleInformationBox must be 0.
[0184] The 32-bit unit header of the V3C unit representing the subsample must be copied to the 32-bit codec_specific_parameters field of the subsample entry within the SubSampleInformationBox. The V3C unit type of each subsample is identified by parsing the codec_specific_parameters field of the subsample entry within the SubSampleInformationBox.
[0185] Encapsulation example
[0186] The example below illustrates the encapsulation of a V3C bitstream containing a single HEVC-coded image within an HEIF. The HEVC-coded image is stored as items in 'hvc1'. The coded data for the image is contained in the media data box ('mdat') or the item data box ('idat'). The association between the V3C item and the image properties is indicated in the ItemPropertyAssociationBox ('ipma'). The image properties in 'v3cC' and 'hvcC' are marked as required because the V3C decoder configuration and the corresponding HEVC decoder configuration metadata must be processed. The player must not process items associated with properties marked as required that are unrecognized or unsupported. To enable subsample-level access, subsample information is carried in 'subs'. If subsample-level access is not required, 'subs' is marked as non-required.
[0187]
[0188] Exemplary player actions
[0189] FIG. 7 illustrates an exemplary process for player operation and a method for generating a rendered output suitable for a user viewport from a HEIF file of an MPI image as proposed. Given an input HEIF file, first (705), the player determines whether the input file is supported. The player must not process image items associated with features marked as essential that are not recognized or supported. When the player supports the essential features within the input file, the player begins parsing the input file and extracts a V3C bitstream (708) from the V3C item, along with V3C decoder configuration information (610) and 2D video decoder configuration information (615) from the image item features (706) within the input HEIF file.
[0190] Next, the player initializes the decoder using the extracted decoder configuration information and decodes the V3C bitstream into a single decoded image and atlas data (712) (710). The decoded image includes spatially packed textures and opacity maps. The atlas data includes the locations of textures and opacity of each MPI layer within the decoded frame, the depth of each MPI layer, and exogenous and endogenous camera information. The player uses the atlas data to recognize the locations of the textures and opacity maps of the MPI layers within the decoded frame. Subsequently, an MPI scene representation is reconstructed from the textures and opacity maps of the MPI layers and the depth information of each layer within the atlas data (720). Camera information regarding the endogenous and exogenous matrices from the atlas data can be used in a warping process to preserve real-world coordinates and synchronize multiple cameras. After reconstruction, a rendered output suitable for the user's viewport is generated and displayed (730).
[0191] References
[0192] Each of the references listed in this specification is incorporated by reference in its entirety.
[0193]
[0194] Exemplary computer system implementation
[0195] Embodiments of the present invention may be implemented as computer systems, systems composed of electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or devices comprising one or more of such systems, devices, or components. Computers and / or ICs may perform, control, or execute instructions related to imaging file formats for MPI, such as those described herein. Computers and / or ICs may calculate any of the various parameters or values related to imaging file formats for MPI described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0196] Specific embodiments of the present invention include computer processors that execute software instructions that cause the processors to perform the method of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., may implement methods related to imaging file formats for MPI as described above by executing software instructions in program memory accessible by the processors. Embodiments of the present invention may also be provided in the form of a program product. A program product may include any non-transient and tangible medium containing a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the method of the present invention. Program products according to the present invention may be any of a wide variety of non-transient and tangible forms. A program product may include physical media such as, for example, magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, and electronic data storage media including flash RAM. Computer-readable signals on the program product may optionally be compressed or encrypted. Where a component (e.g., software module, processor, assembly, device, circuit, etc.) is mentioned above, unless otherwise indicated, a reference to such component (including a reference to "means") shall be interpreted as including any component that performs the function of the described component (e.g., functionally equivalent) as equivalents of such component, including components that are not structurally equivalent to the disclosed structure performing the function in the exemplary embodiments of the present invention.
[0197] Equivalents, extensions, substitutes, and others
[0198] Accordingly, exemplary embodiments relating to imaging file formats for MPI are described. In the foregoing specification, embodiments of the invention have been described with reference to a number of specific details that may vary by implementation. Accordingly, the sole and exclusive indicator of what constitutes the invention and what is intended to be the invention by the applicants is the set of claims issued in this application, which shall have a specific form in which such claims are issued, including any subsequent corrections. Any definitions expressly provided in this specification for terms included in such claims shall govern the meaning of such terms as used in the claims. Accordingly, any limitation, element, characteristic, feature, advantage, or attribute not expressly described in the claims shall never limit the scope of such claims. Accordingly, the specification and drawings should be regarded as exemplary rather than restrictive.
[0199] supplement
[0200] This appendix provides copies of relevant syntax and semantics from ISO / IEC 23090-10 (Reference [6]) and ISO / IEC 14496-12 (Reference [3]).
[0201] From ISO / IEC 23090-10.
[0202] V3C Decoder Configuration Record
[0203] definition
[0204] The V3C decoder configuration record provides decoding-specific information of the V3C bitstream (i.e., parameter sets and SEI messages) for the additional configuration and initialization of the V3C decoder.
[0205] Syntax
[0206]
[0207] Semantics
[0208] unit_size_precision_bytes_minus1 + 1 specifies the precision in bytes of the sample stream NAL unit or sample stream V3C unit to which this configuration record applies. The value of this field must depend on the 4CC-code of the sample entry. For V3C atlas tracks, unit_size_precision_bytes_minus1 must be equal to ssnh_unit_size_precision_bytes_minus1 in sample_stream_nal_header(). For V3C bitstream tracks, unit_size_precision_bytes_minus1 must be equal to ssvh_unit_size_precision_bytes_minus1 in sample_stream_v3c_header(). num_of_v3c_parameter_sets specifies the number of V3C parameter set units signaled in the decoder configuration record.
[0209] v3c_parameter_set_length represents the size of the v3c_parameter_set array in bytes. The signaled value must not be equal to 0.
[0210] Note: The v3c_parameter_set_length syntax element defined in ISO / IEC 23090-5 can be represented by up to 64 bits, and this document limits the representation of information to 16 bits because that is sufficient in actual implementations.
[0211] v3c_parameter_set is an array of data containing all v3c_units of type V3C_VPS as defined in ISO / IEC 23090-5.
[0212] num_of_setup_unit_arrays indicates the number of arrays of atlas NAL units of the indicated type(s).
[0213] array_completeness is equal to 1, indicating that all atlas NAL units of the given type are in the following array and none are in the stream; equal to 0, indicating that additional atlas NAL units of the indicated type may be in the stream; default and allowed values are constrained by the sample entry name.
[0214] nal_unit_type indicates the type of atlas NAL units in the following arrays (this must be all of such types); takes a value as defined in ISO / IEC 23090-5; and is limited to taking one of the values representing NAL_ASPS, NAL_AAPS, NAL_AFPS, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI atlas NAL units.
[0215] num_nal_units represents the number of atlas NAL units of type nal_unit_type included in the configuration record for the stream to which this configuration record applies.
[0216] setup_unit_length represents the size of the setup_unit array in bytes. The signaled value must not be equal to 0.
[0217] A setup_unit is an array of data containing all nal_units as defined in ISO / IEC 23090-5. The included NAL units must be of the same type as specified by nal_unit_type. When present in a setup_unit, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI contain SEI messages of a 'declarative' nature, i.e., those that provide information about the stream as a whole.
[0218] From ISO / IEC 14496-12
[0219] Subsample Information Box
[0220] definition
[0221] Box type: 'subs'
[0222] Container: SampleTableBox or TrackFragmentBox
[0223] Required: No
[0224] Quantity: 0 or more
[0225] This box is designed to contain subsample information. The subsample information item characteristics include subsample information stored in the V3C item. This includes the number of subsamples, which type of V3C unit is carried in each subsample, and the subsample offset in the V3C item. This information enables access to subsamples containing V3C units and effective decoding of specific types of V3C units from the V3C item. That is, the subsample information enables the player to use the offset of the V3C video unit to extract V3C video data units and decode them using a 2D video decoder, or to extract V3C atlas data units and decode them by a V3C atlas decoder.
[0226] A subsample is a continuous range of bytes of a sample. A specific definition of a subsample must be provided for a given coding system (e.g., ISO / IEC 14496-10:2014, for advanced video coding). If such a specific definition is not provided, this box should not apply to samples using such a coding system.
[0227] If subsample_count is 0 for any entry, such samples do not have subsample information and are not followed by an array. The table is sparsely coded; the table identifies which samples have a subsample structure by recording the difference in sample numbers between each entry. The first entry in the table records the sample number of the first sample that has subsample information.
[0228] Note: It is possible to combine subsample_priority and discardable so that discardable is set to 1 when subsample_priority is less than a certain value. However, since different systems may use different scales of priority values, it is safer to separate them to have a clean solution for discardable subsamples.
[0229] When more than one SubSampleInformationBox exists in the same container box, the values of the flags must differ for each of these SubSampleInformationBoxes. The semantics of the flags must be supplied to the given coding system if they exist. If the flags do not have semantics for the given coding system, the flags must be 0.
[0230] Syntax
[0231]
[0232] Semantics
[0233] version is an integer (0 or 1 in this document) that specifies the version of this box.
[0234] entry_count is an integer that provides the number of entries in the following table.
[0235] sample_delta is an integer representing a sample having a subsample structure. This is coded as the difference in decoding order between the desired sample number and the sample number represented in the previous entry. If the current entry is the first entry within the track, the value represents the sample number of the first sample containing subsample information, i.e., the value is the difference between the sample number and zero (0). If the current entry is the first entry within a track fragment containing preceding non-empty track fragments, the value represents the difference between the sample number of the first sample containing subsample information and the sample number of the last sample within the previous track fragment. If the current entry is the first entry within a track fragment that does not have any preceding track fragments, the value represents the sample number of the first sample containing subsample information, i.e., the value is the difference between the sample number and zero (0). This implies that sample_delta for the first entry describing the first sample within the track or within the track fragment is always 1.
[0236] subsample_count is an integer specifying the number of subsamples for the current sample. If there is no subsample structure, this field takes a value of 0.
[0237] subsample_size is an integer that specifies the size of the current subsample in bytes.
[0238] subsample_priority is an integer that specifies the degradation priority for each subsample. Higher values of subsample_priority indicate subsamples that are important to the decoded quality and have a greater impact on it.
[0239] A discardable value of 0 means that the subsample is needed to decode the current sample, whereas a value of 1 means that the subsample is not needed to decode the current sample but can be used for enhancements, for example, the subsample consists of complementary enhancement information (SEI) messages.
[0240] codec_specific_parameters is defined by the codec in use. If such a definition is not available, this field should be set to 0.< / mpi:depthmantissa> < / mpi:depthexponent> < / mpi:depthsign>
Claims
Claim 1 A method for storing a multi-plane image (MPI) scene according to a High Efficiency Image File (HEIF) file container, comprising: generating an MPI image including two or more image layers, wherein each image layer includes texture information and opacity information; packing the texture information layers to generate a 2D texture image; packing the opacity information layers to generate a 2D opacity image; packing the 2D texture image and the 2D opacity image according to an image packing format to generate a packed image; coding the packed image to generate a coded MPI image; generating MPI metadata for the coded MPI image; and generating a file representation of the MPI image by combining the coded MPI image and the MPI metadata according to the syntax semantics of the HEIF file container, wherein the coded media representation of the coded MPI image and the MPI metadata conforms to the Visual Volumetric Video-based Coding (V3C) specification. Claim 2 A method according to claim 1, wherein the MPI metadata comprises one or more of the number of two or more image layers within the MPI image; a description of the packing format; and depth information for the image layers within the MPI image. Claim 3 A method according to claim 1, wherein in the HEIF file, the presence of V3C items is indicated by a MetaBox of type "volv"; the metabox includes V3C configuration item characteristics, 2D video decoder configuration item characteristics, and subsample item characteristics; and the mdat box includes a V3C item representing a single visual volumetric video frame of the coded MPI image and associated MPI metadata. Claim 4 In paragraph 3, the V3C items are identified by the "v3e1" 4CC code type, method. Claim 5 A method in which, in paragraph 4, items of type 'v3e1' can be associated with 2D video decoder item characteristics such as V3C component item characteristics, subsample item characteristics of type 'subs', and HEVC component item characteristics having type 'hvcC' or VVC component item characteristics having type 'vvcC'. Claim 6 In paragraph 4, if a PrimaryItemBox exists, the item_ID within this box is set to display a V3C item of type 'v3e1', and the syntax of the 'v3e1' item includes the following: A method in which item_size is equal to the sum of extent_length values of each extent of the item as specified in ItemLocationBox, v3c_config represents a record within the associated V3CConfigurationProperty, v3c_unit_size specifies the size in bytes of the ss_v3c_unit array, this size is equivalent to the sample stream V3C unit size ssnu_v3c_unit_size as defined in ISO / IEC 23090-5, Annex C, and ss_v3c_unit contains a single V3C unit of the V3C unit sample stream format as defined in ISO / IEC 23090-5:2021, Annex C. Claim 7 In paragraph 3, the above V3C component item characteristics have a box type of 'v3cC' having the following syntax: A method in which v3c_config includes a single instance of V3CDecoderConfigurationRecord that provides decoding-specific information for a V3C bitstream. Claim 8 In paragraph 3, the method wherein the subsample item characteristic has a box type of 'subs'. Claim 9 In paragraph 3, coding is performed according to HEVC, and combining the packed image and the MPI metadata according to the HEIF representation of a single MPI image is, A method including Claim 10 In paragraph 3, the method further comprises the step of decoding the HEIF file of the coded MPI image with a player, wherein the decoding step determines whether all essential characteristics of the HEIF file are supported by the player, and if true: Parsing the HEIF file above to extract the MPI-coded image and the MPI metadata of the MPI image; Decoding and unpacking the above-mentioned coded MPI image to generate the above-mentioned texture image and the above-mentioned opacity image; A method comprising the step of generating an output image based on the MPI metadata, the texture image, the opacity image, and the user viewport, given a user viewport. Claim 11 A non-transient computer-readable storage medium, the non-transient computer-readable storage medium storing computer-executable instructions for executing a method according to any one of claims 1 through 10 with one or more processors. Claim 12 A device comprising a processor and configured to perform any one of the methods described in claims 1 through 10.