Transmitting volumetric images in multiplane image formats

By packing MPI layers into 2D video frames and using compatible encoders/decoders, the method addresses transmission challenges, ensuring high-quality rendering of 3D scenes for applications like computer vision and virtual reality.

JP2026513465APending Publication Date: 2026-04-27DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-04-11
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently transmitting and rendering multiplane images (MPI) due to the complexity of handling multiple image planes and transparency layers, which affects the quality and flexibility of 3D scene representation in applications like computer vision and virtual reality.

Method used

The method involves packing the texture and alpha layers of MPI into 2D video frames, applying video compression, and multiplexing with metadata to facilitate efficient transmission and decoding, using compatible encoders and decoders to reconstruct the multiplane images based on specified packing arrangements.

Benefits of technology

This approach enables high-quality rendering of 3D scenes by maintaining image quality and flexibility in rendering new views, supporting various applications such as computer vision and virtual reality with efficient data transmission and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513465000001_ABST
    Figure 2026513465000001_ABST
Patent Text Reader

Abstract

A method and apparatus for transmitting volumetric images in MPI format. According to an exemplary embodiment, the texture and alpha layers of a multiplane image are packed as tiles into a sequence of video frames. The sequence of video frames is then compressed to generate a video bitstream, which is transmitted along with a metadata bitstream specifying parameters for at least the packing arrangement of the tiles in the sequence of video frames. Exemplary packing arrangements include various selectable spatial and temporal arrangements for the texture layer, alpha layer, and camera view. In some examples, the metadata bitstream is implemented using SEI messages and includes parameters selected from a group including the size of the reference view, the number of layers in the multiplane image, the number of simultaneous views, one or more characteristics of the packing arrangement, layer merge information, dynamic range adjustment information, and reference view information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related applications] This application claims priority to U.S. Provisional Patent Applications No. 63 / 495,715 filed on 12 April 2023, No. 63 / 510,204 filed on 26 June 2023, No. 63 / 586,232 filed on 28 September 2023, and No. 63 / 613,374 filed on 21 December 2023, all of which are incorporated in their entirety by reference into this specification.

[0002] [Related fields] Various exemplary embodiments generally relate to multiplane imaging (MPI), and more specifically, to the transmission of multiplane imaging, but are not limited to these. [Background technology]

[0003] Multiplane imaging embodies a relatively new approach to storing volumetric content. MPI can be used to render both still images and videos, for example, using 8, 16, or 32 planes per camera with texture and transparency (alpha) information to represent a three-dimensional (3D) scene within a view cone. Exemplary applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. [Overview of the Initiative]

[0004] This specification discloses various embodiments of methods and apparatuses for transmitting volume images in the MPI format. According to an exemplary embodiment, the texture layer and the alpha layer of a video sequence of multi-plane images are packed as tiles into a sequence of two-dimensional (2D) video frames. The sequence of 2D video frames is then compressed to generate a video bitstream, which is transmitted together with a metadata bitstream specifying related MPI parameters, such as parameters specifying the packing arrangement for the tiles in the sequence of 2D video frames. Selectable packing arrangements include, but are not limited to, (i) spatially packed texture and alpha layers with temporally packed views, (ii) spatially packed views with temporally packed texture and alpha layers, and (iii) spatially packed texture layers and spatially packed alpha layers that are temporally interleaved with temporally packed views. In some examples, the metadata bitstream includes parameters selected from a group including the size of a reference view, the number of layers in a multi-plane image, the number of simultaneous views, characteristics of the packing arrangement, layer merge information, dynamic range adjustment information, and reference view information. In some examples, the metadata bitstream includes one or more supplemental enhancement information (SEI) messages.

[0005] According to an exemplary embodiment, an apparatus for encoding a sequence of multi-plane images, the apparatus comprising: at least one processor; at least one memory including program code; comprising the at least one memory and the program code, together with the at least one processor, cause the apparatus to at least Generate a sequence of video frames, each of the video frames including a plurality of tiles each representing a layer of one or more of the multi-plane images, Generate a metadata bitstream for specifying at least a packing arrangement of the tiles within the sequence of video frames, Generate a video bitstream by applying video compression to the sequence of video frames, Multiplex the video bitstream and the metadata bitstream for transmission, A device is provided that is configured as such.

[0006] According to another exemplary embodiment, a method for encoding a sequence of multi-plane images, the method comprising: Generating a sequence of video frames, each of the video frames including a plurality of tiles each representing a layer of one or more of the multi-plane images, Generating a metadata bitstream for specifying at least a packing arrangement of the tiles within the sequence of video frames, Generating a video bitstream by applying video compression to the sequence of video frames, Multiplexing the video bitstream and the metadata bitstream for transmission, A method is provided that includes the above.

[0007] According to yet another exemplary embodiment, a device for decoding an encoded received bitstream of a sequence of multi-plane images, the device comprising: At least one processor, At least one memory including program code, Including, The at least one memory and the program code, together with the at least one processor, cause the device to at least The received bitstream is demultiplexed to obtain a video bitstream encoded with a sequence of video frames, and a metadata bitstream is obtained that specifies at least the packing arrangement of tiles within the sequence of video frames, wherein the tiles represent layers of the multiplane image. By applying video decompression to the video bitstream, the sequence of video frames is reconstructed. Using the tiles from the sequence of video frames, the sequence of multiplane images is reconstructed based on the metadata bitstream. The equipment is provided, configured in such a way.

[0008] According to yet another exemplary embodiment, a method for decoding a received bitstream encoding a sequence of multiplane images, wherein the method is: The steps include: demultiplexing the received bitstream to obtain a video bitstream encoding a sequence of video frames; and obtaining a metadata bitstream specifying at least the packing arrangement of tiles in the sequence of video frames, wherein the tiles represent layers of the multiplane image; The steps include: reconstructing the sequence of video frames by applying video decompression to the video bitstream; The steps of: using the tiles from the sequence of video frames to reconstruct the sequence of multiplane images based on the metadata bitstream; A method including this is provided.

[0009] In some embodiments of the above method, when executed by at least one processor, a non-temporary computer-readable medium is provided that stores instructions causing the at least one processor to perform an operation including one of the above methods. [Brief explanation of the drawing]

[0010] Other aspects, features, and advantages of the various embodiments disclosed will become more fully apparent, for example, from the following detailed description and accompanying drawings.

[0011] [Figure 1] This illustrates an exemplary process for a video / image delivery pipeline.

[0012] [Figure 2] A schematic representation of a 3D scene using a multiplane image according to one embodiment is shown.

[0013] [Figure 3] This diagram illustrates the process of generating a new view of a 3D scene as an example.

[0014] [Figure 4] This block diagram shows an example of how the set of active views changes over time.

[0015] [Figure 5] This is a block diagram showing an MPI encoder that can be used in the distribution pipeline of Figure 1 according to one embodiment.

[0016] [Figure 6] This is a block diagram showing an MPI decoder that can be used in the distribution pipeline of Figure 1 according to one embodiment.

[0017] [Figure 7] This table shows some examples of the specific constraints imposed on the design and interoperability of the MPI encoder in Figure 5 and the MPI decoder in Figure 6.

[0018] [Figure 8] This figure shows the packing operation of an MPI encoder as an example.

[0019] [Figure 9]This figure shows the packing operation of the MPI encoder in another example (Figure 5).

[0020] [Figure 10] Figure 5 shows the packing operation of the MPI encoder as another example.

[0021] [Figure 11] Figure 5 shows the packing operation of the MPI encoder as another example.

[0022] [Figure 12] This is a flowchart showing an MPI encoding method according to one embodiment.

[0023] [Figure 13] This flowchart shows an MPI decoding method according to one embodiment.

[0024] [Figure 14A] The following provides illustrative syntax for SEI messages configured to transmit MPI metadata, using several examples. [Figure 14B] The following provides illustrative syntax for SEI messages configured to transmit MPI metadata, using several examples.

[0025] [Figure 15] The following provides an exemplary syntax for SEI messages configured to transmit MPI metadata, using several other examples.

[0026] [Figure 16] The following provides an exemplary syntax for SEI messages configured to transmit MPI metadata, using several other examples.

[0027] [Figure 17] Block diagram showing a computing device according to one embodiment. [Modes for carrying out the invention]

[0028] This disclosure and its embodiments include hardware, devices, or circuits, computer program products, computer systems and networks, user interfaces and application programming interfaces, as well as hardware-implemented methods, signal processing circuits, memory arrays, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc., which can be embodied in various forms and controlled by computer-implemented methods. The foregoing is intended merely to give an overall idea of ​​the various embodiments of this disclosure and does not limit the scope of this disclosure in any way.

[0029] The following description includes numerous details, such as device configuration, timing, operation, etc., to provide an understanding of one or more aspects of the present disclosure. It will be immediately apparent to those skilled in the art that these specific details are merely examples and are not intended to limit the scope of the application.

[0030] <Example video / image delivery pipeline> Figure 1 shows an exemplary process of a video distribution pipeline 100, illustrating various stages from video / image capture to video / image content display according to an embodiment. A sequence of video / image frames 102 may be captured or generated using an image generation block 105. The video frames 102 may be captured by a digital camera or generated by a computer (e.g., using computer animation) to provide video and / or image data 107. Alternatively, the frames 102 may be captured on film by a film camera. The film is then converted to a digital format to provide video / image data 107. In some examples, the image generation block 105 includes generating MPI images or videos.

[0031] In production stage 110, data 107 can be edited to provide a video / image production stream 112. The data of the video / image production stream 112 may then be provided to a processor (or one or more processors such as a central processing unit (CPU)) in post-production block 115 for post-production editing. Post-production editing in block 115 may include, for example, intermediate adjustments or changes to the color or brightness of specific areas of an image to improve image quality or achieve a particular appearance in accordance with the video producer's intentions. This part of post-production editing is sometimes called "color timing" or "color grading". Other edits (e.g., scene selection and ordering, image cropping, addition of computer-generated visual special effects, artifact removal, etc.) may be performed in block 115 to produce a final version 117 of the production for distribution. In some examples, the operations performed in block 115 include expanding textures and / or alpha channels in a multiplane image / video. During post-production editing 115, the video and / or images may be viewed on a reference display 125.

[0032] Following post-production 115, the final version of the data 117 may be delivered to a coding block 120 for further downstream distribution to decoding and playback devices such as television sets, set-top boxes, and movie theaters. In some embodiments, the coding block 120 may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray®, and other distribution formats, to generate a coded bitstream 122. At the receiver, the coded bitstream 122 is decoded by a decoding unit 130 to generate a corresponding coded signal 132 representing a copy or exact approximation of the signal 117. The receiver may be mounted on a target display 140, which may have some or entirely different characteristics from the reference display 125. In this case, a display management (DM) block 135 may be used to map the coded signal 132 to the characteristics of the target display 140 by generating a display-mapped signal 137. According to one embodiment, both the decoding unit 130 and the display management block 135 may include individual processors or may be integrated into a single processing unit.

[0033] The codecs used in coding block 120 and / or decoding block 130 enable video / image data processing and compression / decompression. Compression is used in coding block 120 to reduce the size of the corresponding file or stream. The decoding process performed by decoding block 130 typically involves decompressing the received video / image data file or stream into a format usable for playback and / or further editing. Examples of coding / decoding operations that can be used in coding block 120 and decoding unit 130 in various embodiments are described in more detail below.

[0034] Multiplane image A multiplane image contains multiple image planes, each of which is a "snapshot" of the 3D scene at a specific depth relative to the camera position. The information stored in each plane includes texture information (e.g., represented by R, G, B values) and transparency information (e.g., represented by alpha (A) values), where the acronyms R, G, and B represent red, green, and blue, respectively. In some examples, the three texture components may be (Y, Cb, Cr), or (I, Ct, Cp), or another set of functionally similar values. There are various ways to generate a multiplane image. For example, two or more input images from two or more cameras located at different known viewpoints can be co-processed to generate a corresponding multiplane image. Alternatively, a single-view composite of multiplane images can be performed using the original image captured by a single camera.

[0035] Figure 2 schematically illustrates a 3D scene representation using a multiplane image 200 according to one embodiment. The multiplane image 200 has D planes or layers (P0, P1, ..., P(D-1)), where D is an integer greater than 1. Typically, the planes (layers) are indexed along the Z dimension of the 3D scene, with the layer furthest from the reference camera position (RCP) being the 0th layer, and the layers being indexed so that they are at a distance (or depth) d0 from the RCP. The index is incremented by 1 for each subsequent layer closer to the RCP. The plane (layer) closest to the RCP has an index value of (D-1), and the distance (or depth) d from the RCP along the Z dimension is d D-1 It is located at a vertical height h above the base plane 202. Each plane (P0, P1, ..., P(D-1)) is orthogonal to the base plane 202 which is parallel to the XZ coordinate plane. The RCP is located at a vertical height h above the base plane 202. The XYZ triad shown in Figure 2 shows the general orientation of the multiplane image 200 and the planes (P0, P1, ..., P(D-1)) with respect to the X, Y, and Z dimensions of the 3D scene. In various examples, the number D can be 32, 16, 8, or any other suitable integer greater than 1.

[0036] Let the color component (e.g., RGB) value of the i-th layer at the camera position s be C i (S) The horizontal size of the layer is H×W, where H is the height (Y dimension) of the layer and W is the width (X dimension) of the layer. Let the pixel value at the position (x, y) of the color channel c be C i (S) be represented as C(x,y,c). The α value of the i-th layer is A i (S) Let the pixel value (x, y) in the alpha layer be A i (S) (x,y). The depth distance between the i-th layer and the reference camera position is d i The image from the original reference view (when there is no camera movement) is represented as R, and the texture pixel is R (S) (x,y,c). Thus, the static MPI image at the camera position s can be expressed as follows: MPI(s)={C i (S) (x,y,c),A i (S) (x,y)},i = 0,…,D - 1 (1) If the camera position s is kept static over time, it is easy to extend this static MPI image representation to a video representation. This video representation is given by Equation (2): MPI(s,t)={C i (S) (x,y,c,t),A i (S) (x,y,t)},i = 0,…,D - 1 (2) Here, t represents time.

[0037] As already mentioned above, multiplane images such as multiplane image 200 can be generated using single-view synthesis from a single source image R, or using multi-view synthesis from two or more source images. Such synthesis may be performed, for example, during the production stage 110. The corresponding MPI synthesis algorithm typically outputs a multiplane image 200 containing XYZ decomposed pixel values ​​in the form {(Ci,Ai) for i=0,.,D-1}.

[0038] By processing the multiplane image 200 represented by {(Ci,Ai) for i=0,.,D-1}, the MPI rendering algorithm can generate a visible image corresponding to a new virtual camera position that is either RCP or different from RCP. An exemplary MPI rendering algorithm that can be used for this purpose (often called an "MPI viewer") may include warping and compositing steps. Other suitable MPI viewers may also be used. The rendered multiplane image 200 can be viewed, for example, on a reference display 125.

[0039] During the warping step of the MPI rendering algorithm, each layer (Ci, Ai) of the multiplane image 200 is set to the RCP viewpoint position (v S ) to a new viewpoint position (v t ) can be warped as follows, for example:

number

number

[0001] T This represents the depth σd. i This indicates the distance to a plane that is frontal-parallel to the source camera.

[0040] During the synthesis step of the MPI rendering algorithm, for example, a new visible image C is created using the processing operation corresponding to the following equation. t It can generate:

number

number

number

number

number

[0041] In a single-camera transmission scenario, only one MPI is supplied via the bitstream. The goal in this situation is to optimally merge the layers of the original MPI so that the quality of this MPI is maintained after local warping. In a multi-camera transmission scenario, multiple MPIs captured at different camera positions are encoded into a compressed bitstream. The information in these MPIs is used together to generate a new global view of the positions located between the original camera positions. There may also be scenarios where information from multiple cameras can be used together to generate a single MPI to be transmitted. For MPI video transmission, multi-camera transmission scenarios are typically used, as described below, for example.

[0042] Figure 3 schematically illustrates the process of generating a new view of 3D scene 302 as an example. In the example shown, 3D scene 302 is captured using 42 RCPs (1,2,.,42). The new view being generated corresponds to camera position 50. The four RCPs closest to camera position 50 are RCP(11,12,18,19). The corresponding multiplane image is multiplane image(200 11 , 200 12 , 200 18 , 200 19 ) is the multiplane image 200 corresponding to camera position 50. 50 This is a multiplane image (200 11 , 200 12 , 200 18 , 200 19 ) are correspondingly warped and then the resulting warped multiplane images are merged to produce the final product. Finally, the synthesis step 310 of the MPI rendering algorithm is performed on the multiplane image 200. 50 By applying this, a visible image 312 of the 3D scene 302 is generated.

[0043] In general, a 3D scene, such as 3D scene 302, can be captured using any suitably selected number of RCPs. The positions of such RCPs can also be selected in various ways, for example, based on the creative intention. In a typical embodiment, when a new view, such as visible image 312, is rendered, only some of the neighboring RCPs are used for rendering. Hereinafter, such neighboring views are referred to as "active views". In the example shown in FIG. 3, the number of active views is 4. In other examples, a different number of active views may be used as well. Therefore, the number of active views is a selectable parameter. For illustrative purposes and without any implicit limitation, exemplary embodiments will be described hereinafter in this specification with reference to four active views. In some examples, the set of active views can change over time as camera position 50 moves. In some examples, the number of active views can change over time as camera position 50 moves.

[0044] FIG. 4 is a block diagram showing the change of the set of active views over time according to an example. In the illustrated example, a 3D scene is captured using a rectangular array of 40 RCPs arranged in 5 rows and 8 columns. The dashed arrow 402 represents the movement trajectory of the new camera position 50 during the time interval starting at time t 0 and ending at time t 1 for virtual view synthesis. At time t 0, the set of active views includes the four views surrounded by the dashed box 410. At time t 1, the set of active views includes the four views surrounded by the dotted box 420. At time t n (t 0 < t n < t 1), the set of active views changes from the set within the dashed box 410 to the set within the dotted box 420.

[0045] Transmission of the Encoded MPI Video Figure 5 is a block diagram of an MPI encoder 500 according to one embodiment. During operation, the MPI encoder 500 converts an MPI video 502 into a coded bitstream 542. The MPI video 502 has a sequence of multiplane images 200 corresponding to a sequence of time. In various examples, the MPI video 502 has two or more components corresponding to a single view or multiple views (see also Figures 3 and 4). In some examples, the MPI video 502 is transmitted via a video / image stream 112 or 117. The coded bitstream 542 is sent to the corresponding MPI decoder via a coded bitstream 122.

[0046] The MPI video 502 is preprocessed in preprocessing block 510 to obtain preprocessed MPI video 512. Exemplary preprocessing operations performed in preprocessing block 510 include, but are not limited to, normalization, reshaping, padding, scaling, and refinement applied to at least one of the texture channel and alpha channel. A typical example of preprocessing operations that may be performed in preprocessing block 510 is described, for example, in U.S. Provisional Patent Application No. 63 / 357,669 filed on 1 July 2022 (also filed as PCT Patent Application PCT / US2023 / 69096 filed on 26 June 2023) by G-MSu and P. Ying, which is incorporated in its entirety by reference in this specification. In some embodiments, a “masking” process can be used during preprocessing to generate a “masked” texture channel that stores only partial texture information according to a predefined binary mask M(M(u,v)) at the sample position (u,v). If M(u,v) is true, C(u,v) is set to a constant value (e.g., zero or mid-gray). The mask M can be created by binarizing the alpha channel, i.e., if A(u,v)==0, then M(u,v)=1, and otherwise M(u,v)=0. When generating the binary mask, the following morphological dilation process is performed.

number

number

[0047] The MPI video 512 is converted into a packed 2D video 522 in the packing block 520. The video 522 has a format compatible with the video encoder 530. Exemplary selectable packing options and the corresponding packing operations performed in the packing block 520 are described in more detail below with reference to, for example, Figures 8 to 11. The packing block 520 also operates to generate an MPI metadata stream 524. The MPI metadata stream 524 is configured to inform the corresponding MPI decoder about the selected packing options and to provide various relevant MPI parameters. Various non-exclusive examples of the MPI metadata stream 524 are described in more detail below with reference to, for example, Figures 14A to 14B and Figure 15.

[0048] The video encoder 530 operates to convert the 2D video 522 into a video bitstream 532 and a corresponding video metadata stream 534, for example, by applying appropriate video compression. In various examples, the video encoder 530 may be a High Efficiency Video Coding (HEVC) encoder, an MPEG-4 Advanced Video Coding (AVC) encoder, a FLOSS encoder, or any other suitable video encoder. The multiplexer (MUX) 540 operates to produce a coded bitstream 542 by appropriately multiplexing the video bitstream 532, the video metadata stream 534, and the MPI metadata stream 524. In some other examples, the MPI metadata stream 524 may be incorporated into or part of the video metadata stream 534.

[0049] Figure 6 is a block diagram showing an MPI decoder 600 according to one embodiment. The MPI decoder 600 is designed to be compatible with the corresponding MPI encoder 500. During operation, the MPI decoder 600 receives a coded bitstream 542 generated by the corresponding MPI encoder 500 as input and produces an MPI video 602 corresponding to the camera position 606 as output. In a typical example, the camera position 606 is specified to the MPI decoder 600 by the viewer and can be the same as one of the RCPs (see also Figure 3) or different from any of the RCPs. In some examples, the camera position 606 can be a function of time (see also Figure 4).

[0050] During operation, the demultiplexer (DMUX) 640 demultiplexes the received coded bitstream 542 to reconstruct the video bitstream 532, video metadata stream 534, and MPI metadata stream 524. In some examples, the MPI metadata stream 524 is part of the video metadata stream 534, as described above. In such examples, the operation of the DMUX 640 is adjusted accordingly. The video decoder 630 is compatible with the video encoder 530 and operates to decompress the video bitstream 532 using the video metadata stream 534, thereby producing the 2D video 622. When lossy compression is used, the 2D video 622 is not an exact copy of the 2D video 522, but rather a relatively close approximation. When lossless compression is used, the 2D video 622 is a copy of the 2D video 522. In either case, the 2D video 622 is useful for the depacking operation, which is configured to be the reverse of the packing operation performed in the packing block 520. Such depacking operations for the 2D video 622 are performed in the depacking block 620 based on the MPI metadata stream 524, resulting in the generation of the MPI video 612 at the output of the depacking block 620. The post-processing block 610 operates to apply post-processing operations to the MPI video 612 to generate the MPI video 608. Based on the camera position 606, the compositing block 604 renders the MPI video 608 to generate a visible video 602 corresponding to the camera position 606. In various examples, the rendering operations performed in the compositing block 604 include some or all of the following steps: warping multiplane images corresponding to one or more of the active RCPs; merging the warped multiplane images; and compositing the relevant sequences of MPI images to generate a viewable video 602.

[0051] As already stated above, blocks 520 and 530 of the MPI encoder 500 and the corresponding blocks 630 and 620 of the MPI decoder 600 operate in a compatible manner. For example, the design and configuration of packing block 520 depends on the selected type of video encoder 530. Furthermore, the configuration of the corresponding blocks 630 and 620 of the MPI decoder 600 must be compatible with the selection / configuration made for blocks 520 and 530 of the MPI encoder 500. For illustrative purposes only, and without any implicit limitations, codec parameters affecting the design and interoperability of blocks 520, 530, 630, and 620 are described below with reference to HEVC encoders / decoders. From the provided description, those skilled in the art will readily understand how to guide the design of blocks 520, 530, 630, and 620 for other types of video encoders / decoders 530, 630, and how to ensure interoperability.

[0052] Many HEVC encoding tools allow the user to choose between the Main or Main10 profile. The Main profile supports 8 bits per sample, which allows for 256 shades per primary color, or 16.7 million colors, in video. In contrast, the Main10 profile supports up to 10 bits per sample, which allows for up to 1024 shades and over 1 billion colors. Easily available (e.g., off-the-shelf) video encoders / decoders generally support HEVC Main or Main10 profiles up to level 6.2. Level 5.1 coding, for example, is relatively common in hardware-implemented decoders. Therefore, the following explanation will focus on levels 5.1 and above, up to level 6.2.

[0053] Figure 7 is a table showing exemplary constraints imposed on the design and interoperability of blocks 520, 530, 630, and 620 by the 5.1, 5.2, and 6.x level specifications, with several examples. More specifically, the maximum luma sample rate (MaxLumaSr), maximum luma picture size (MaxLumaPs), maximum DPB size (MaxDpbSize), and maximum number of tiles (in rows and columns) are all selected based on this table. In some examples, the bitrate and compression ratio values ​​can be met using rate control (QP adaptation). Also note that the luma sample rate (samples / second) is equal to the product of the luma picture size (samples / picture) and the picture rate (pictures / second). In some examples, the picture rate and frame rate are the same.

[0054] In various examples, tiered representations of MPI images are spatially and / or temporally packed or concatenated to create input for the HEVC video codec. The following description provides some relevant details regarding level / profile constraints from the HEVC specification concerning tiers and levels. The corresponding sections in the HEVC specification are "A.4.1: General tier and level limits" and "A.4.2: profile-specific level limits for the video profiles".

[0055] Regarding tier and level restrictions, some or all of the following characteristics may be considered: - Let access unit n be the nth access unit in the decoding order, and let the first access unit be access unit 0 (i.e., the 0th access unit). -Picture n is the coded picture of access unit n or the corresponding decoded picture. In some examples, a bitstream conforming to a profile at a given tier and level is subject to the following constraints: a) PicSizeInSamplesY is less than or equal to MaxLumaPs. b) The value of pic_width_in_luma_samples is less than or equal to Sqrt(MaxLumaPs*8). c) The value of pic_height_in_luma_samples is less than or equal to Sqrt(MaxLumaPs*8). d) For levels 5 and above, the value of CtbSizeY is equal to 32 or 64. e) The value of NumPicTotalCurr shall be 8 or less. f) The value of num_tile_columns_minus1 shall be less than MaxTileCols, and num_tile_rows_minus1 shall be less than MaxTileRows.

[0056] Regarding the profile-specific level limitations of a video profile, some or all of the following characteristics may be considered: In some cases, the value of sps_max_dec_pic_buffering_minus1[HighestTid]+1 is less than or equal to MaxDpbSize, and can be derived as follows:

number

[0057] In some cases, the maximum frame rate supported by the codec is 300 frames per second (fps). MaxDpbSize, the maximum number of pictures in the decoded picture buffer for the maximum luma picture size of that level, is 6 for all levels. If the luma picture size of the video is smaller than the maximum luma picture size of that level, MaxDpbSize can be increased up to 16 frames in increments of 4 / 3×, 2×, or 4×.

[0058] In some examples, the following low-pixel-rate test condition constraints apply. - The combined maximum luma sample rate across all decoders is up to 1,069,547,520 samples per second (e.g., 32MP@30fps, corresponding to HEVC Main10 profile@Level5.2). -Each decoder instantiation is constrained to a maximum Luma picture size of 8,912,896 pixels (e.g., 4096×2048, corresponding to HEVC Main10 profile@Level5.2). - The maximum number of simultaneous decoder instances is 4.

[0059] In some examples, the following high-pixel-rate test condition constraints apply. - The combined maximum luma sample rate across all decoders is up to 4,278,190,080 samples per second (e.g., 128MP@30fps, corresponding to HEVC Main10 profile@Level6.2). -Each decoder instantiation is constrained to a maximum Luma picture size of 35,651,584 pixels (e.g., 8192×4096, corresponding to HEVC Main10 profile@Level6.2). - The maximum number of simultaneous decoder instances is 4.

[0060] In some examples, the maximum number of simultaneous decoders is 4 for levels 5.2 and 6.2. The pixel rate specifications for multiple decoders are as follows: - Multiple decoder instances (e.g., four) are collectively constrained by a single set of hardware-level limits.

[0061] Exemplary MPI video coding solution In some examples, a multiplane image 200 has 32 layers for each frame in a single camera view. In some examples, an adaptive layer merging method is used to reduce the 32 layers to 16 layers while substantially maintaining the subjective quality of the synthesized new view, as described, for example, in U.S. Provisional Patent Applications No. 63 / 429,875 and 63 / 429,878, filed on 2 December 2022, which are incorporated entirely by reference in this specification. For illustrative purposes and without any implicit limitation, several representative examples are described below in this specification with reference to 16-layer MPI representations.

[0062] For example, MPI allocation capacity is defined using the following parameters: -Picture size: Maximum 720p (1280 x 720) • Picture size can be padded to a multiple of the CTU size (e.g., 64x64) to fill the HEVC tile structure. This feature allows for more flexibility for encoder control, preventing loop filtering across MPI picture boundaries and enabling later transcoding when needed. • For new view renderings, cropping is used to remove boundary artifacts. In this case, tiling may not be used for coding. - Number of MPI layers: 8 or 16 layers. For 8-layer designs, multi-CTU / tile-based MPI decomposition and / or adaptive layer merging may be used. In CTU / tile-based MPI designs, post-processing is used to correct multi-CTU / tile boundary artifacts. - The number of camera views to send simultaneously used to render a new view: for example, up to 4 nearest neighbors. • You can also use one, two, or three views. The fewer views you use, the more noticeable residual artifacts may be. • In some examples, you can use more than four views. - Video frame rate at each camera position: 30fps (also known as picture rate). - Latency: (This parameter affects the coding structure and intra-refresh period). When relatively low latency is required, the coding structure may be configured to avoid the use of rearranged pictures (e.g., B-frames). When the application commands frequent changes to the "active" view for new view rendering, the intra-refresh period (I / IDR frame) may be used more frequently to mitigate the possible latency increase. - Pose Generation: In the example application, new poses are generated by the director (called auto-generation), move slowly or quickly, or are controlled interactively by the user. This feature can affect how many neighboring views are selected to generate a new view. • When displayed directly, a single MPI video can be encoded. Furthermore, please note that in the following explanation, we will assume that the following two cases cannot be distinguished for a single hardware decoder. - Each decoder instance is constrained by a 1HW level limit. - Multiple decoder instances (e.g., four) are collectively constrained by a single hardware level limit. For illustrative purposes, different solutions are presented using a single decoder instance. Those skilled in the art will readily understand how to adapt these solutions to multiple decoder instances.

[0063] In this specification, the term “coding tree unit” (CTU) refers to the basic processing unit of the High Efficiency Video Coding (HEVC) standard, and structurally and conceptually corresponds to various macroblock units used in several earlier video standards. In some literature, the CTU is also called the largest coding unit (LCU). In various examples, the CTU has a size ranging from 16x16 pixels to 64x64 pixels, with larger sizes typically resulting in increased coding efficiency.

[0064] In various examples, spatial packing, temporal packing, or a combination of spatial and temporal packing may be used to pack the texture and alpha layers of a multiplane image 200 into HEVC frames. In the case of spatial packing, the MPI encoder 500 works to pack both the texture and alpha layers together so that the picture size is twice the spatial resolution of the original camera view, i.e., luma sample rate = 2 × luma picture size × frame rate.

[0065] Tables 1 and 2 below show exemplary picture sizes for video resolutions of 360p, 480p, 540p, and 720p. In Table 1, the CTU size is 64x64 pixels. In Table 2, the CTU size is 32x32 pixels. In Table 3, the picture size is not limited to an integer multiple of the CTU size, and no padding is performed. For compression in the video encoder 530, the texture layer is converted from RGB to YCbCr4:2:0, 8-bit or 10-bit format. The alpha layer is quantized to 8 / 10 bits and loaded as the Y component. The corresponding Cb and Cr components are loaded with dummy (e.g., constant) values. [Table 1] Example of picture resolution (CTUSize=64) [Table 1] [Table 2] Example of picture resolution (CTUSize=32) [Table 2] [Table 3] Examples of picture resolution (without padding) [Table 3]

[0066] According to one selectable configuration (hereinafter referred to as "Option 1"), the packing block 520 is configured to generate a 2D video 522 by spatially packing the texture and alpha layers of the multiplane image 200 into a single video frame, with different video frames carrying the multiplane image 200 corresponding to different RCPs (or views of the scene) at the corresponding time t. According to one embodiment of Option 1, the packing block 520 supports the following operations and features: • Spatially arrange the texture and alpha layers of 200 multiplane images corresponding to a single view. • Multiplane images 200, which support multiple views, are arranged temporally, for example, using an interleaved method. Each texture / alpha layer may be included in the tile structure when CTU-based padding is used. In such cases, loop filtering can be disabled at tile boundaries, and Motion Constraint Tile Set (MCTS) coding is not required. All texture layers of the multiplane image 200 are grouped into one rectangular region of the video frame, and all alpha layers of the same multiplane image 200 are grouped into another rectangular region. The two regions can be contained in two independent slices to allow for different QP settings for each. For example, a lower QP value can be used for the texture slice and a higher QP value can be used for the alpha slice. Low-latency P-coding structures can be used with periodic IDR refreshes to support on-the-fly view changes. In some examples, random-access coding structures may be used for higher coding efficiency when the latency parameter is relatively insignificant. Views can be coded independently through multiple codec instances, and this feature can be used, for example, to support on-the-fly changes in the number of views (e.g., using 3 views / 2 views / 1 view). According to various examples, the placement of texture or alpha tiles from all layers of the multiplane image 200 is flexible. For example, if 16 MPI layers are used, i.e., D=16, the tile layout can be selected from 1×16, 2×8, 4×4, 8×2, and 16×1 layouts. In different examples, the placement of texture slices and alpha slices is side-by-side or top-and-bottom. Note that this placement does not violate the tile row / column constraints specified in the table in Figure 7. Another option is to first pack the texture and alpha of one MPI layer (top-and-bottom, or side-by-side, or pixel interleaved) and thus pack all the remaining MPI layers of the multiplane image 200. Here, the term "IDR frame" refers to a special type of I-frame in H.264. More specifically, an IDR frame signals that a frame following it cannot reference any frame that precedes it.

[0067] Figure 8 shows an example of the operation of packing block 520. The illustrated example is an example of Option 1 described above, in which four views (corresponding to four different views) are transmitted. The set of transmitted views changes over time, as described above with reference to Figure 4, for example. For example, for MPI video time t0, the set of transmitted views includes four views (V0, V1, V2, V3). In contrast, for MPI video time t nThe transmitted set of views includes four views (V0, V1, V4, V5). Depending on the temporal position in the MPI video sequence 890, the transmitted frame 800 may be an I-frame or a P-frame. In some examples, a B-frame 800 may be transmitted at several temporal positions in the MPI video sequence 890.

[0068] An enlarged view of one of the transmitted frames 800 shows its tile structure in more detail. In the illustrated example, frame 800 contains texture slices 810 and alpha slices 850 packed side-by-side within the frame. The 16 tiles within each slice 810,850 carry the corresponding (texture or alpha) channels of each of the 16 layers (D=16, Figure 2) of the multiplane image 200. Each set of 16 tiles is arranged using a 4x4 spatial layout. The corresponding MPI decoder 600 can be implemented using four decoding instances. As a result, different of the four views of each transmitted view set can be queued and processed in a simpler way. For illustrative purposes, the following description of level constraints assumes that the four corresponding bitstreams are assembled into a single bitstream. Furthermore, it should be noted that if the MPI encoder 500 operates under the constraint that only one reference picture may be used, the assembly operation can be performed with relatively minor modifications implemented at a higher level, such as appropriate updating of the Picture Order Counter (POC) number.

[0069] In HEVC, the decoded picture buffer (DPB) is a buffer that holds decoded pictures for reference, output reorder, or output delay, as specified for the virtual reference decoder in Annex C of the HEVC specification. The current decoded picture is also stored in the DPB. The minimum DPB size that the decoder must allocate to decode a particular bitstream is signaled by sps_max_dec_pic_buffering_minus1. The maximum number of pictures in the decoded picture buffer for the maximum luma picture size of that level is 6 for all levels. If the luma picture size of the video is smaller than the maximum luma picture size of that level, the maximum DPB size can be increased up to 16 frames in increments of 4 / 3×, 2×, or 4×.

[0070] Table 4 below shows the pictures in a DPB based on the Picture Group (GOP) structure shown in Figure 8. At any given time, a DPB is configured to hold a maximum of four pictures and not exceed the maximum allowable value of six. The corresponding sufficient DPB size depends on the number of supported views. In some examples, the minimum DPB size may be set to be equal to the number of supported views when only one reference picture is used from the same view. [Table 4] DPB analysis of the example shown in Figure 8 [Table 4]

[0071] Using four neighboring views (see also Figure 4), the MPI decoder 600 can be configured to perform MPI new view rendering, for example, as follows: For a given source view Vi (i=0,.,3), warping is first applied to each MPI plane, and then the target new view NVi is obtained by alpha compositing the color images from back to front. The final new view is a weighted sum of the new views NV0, NV1, NV2, and NV3. This means that after view Vi is decoded, there is no need to wait for other decoded views, and the view can be immediately sent to the GPU engine for rendering. Therefore, the decoding described above is considered to impose no additional burden on the DPB. For example, one corresponding new view can be quickly rendered for every four decoded frames.

[0072] Exemplary parameter combinations for 360p, 480p, 540p, and 720p resolutions with D=16 and D=8 are shown in Table 5 below, where FPS rate = 30 × number of supported views. The parameters shown in Table 5 are applicable to both CTUSize=64 and CTUSize=32. [Table 5] Example of MPI transmission scenario for HEVC levels 5.x and 6.x: Option 1 [Table 5]

[0073] According to another selectable configuration (hereinafter referred to as "Option 2"), the packing block 520 is configured to generate a 2D video 522 by spatially packing a view (texture and alpha channel) into a set of video frames, where different video frames in the set carry different layers of a multiplane image 200 corresponding to a view of the scene at the corresponding time t. According to one embodiment of Option 2, the packing block 520 supports the following operations and features: - Spatially arrange the textures and alphas of specific layers from a multiplane image 200 corresponding to multiple RCPs within a video frame. Temporarily arrange the layers by placing them in different video frames of a set. Similar to Option 1, texture and alpha layers from multiple views each take their own tile. Alpha tiles are grouped into alpha slices, and texture tiles are grouped into texture slices of the frame. Option 2 may be useful for levels 5.1 and 5.2 to enable 720p video transmission. Option 1 does not support 720p transmission with level 5.x because spatial resolution constraints only allow for four layers of MPI. • Multiple views are spatially packed, which may add complexity to operations targeting modifications to the views (e.g., shrinking, replacing, etc.) compared to Option 1. The maximum frame rate supported by HEVC is 300 frames per second (fps). If the reference view has a rate of 30 fps, the maximum number of layers supported by Option 2 is 300 / 30 = 10 layers.

[0074] Figure 9 shows the operation of packing block 520 in another example. The illustrated example is an example of Option 2 described above, in which four views (V0, V1, V2, V3) are transmitted. Depending on the temporal position in the MPI video sequence 990, the transmitted frame 900 may be an I frame or a P frame. In some examples, B frames 900 may be transmitted at some temporal position in the MPI video sequence 990.

[0075] In the illustrated example, frame 900 contains texture slices 910 and alpha slices 950 stacked vertically (from top to bottom). The four tiles within texture slice 910 are packed using a 1x4 layout, each carrying the texture channels of the corresponding layers for the four views (V0, V1, V2, V3). The four tiles within alpha slice 950 are also packed using a 1x4 layout, each carrying the alpha channels of the corresponding layers for the four views (V0, V1, V2, V3). The eight layers (D=8, Figure 2) of the corresponding multiplane image 200 are arranged in different consecutive video frames 900 of sequence 990, as shown in Figure 9.

[0076] In the example shown in Figure 9, eight layers (D=8, Figure 2) are interleaved in time. The corresponding DPB size is 8 pictures. Therefore, the MaxDpbSize parameter is set to >=8. In this case, the DPB size does not depend on the number of views. Table 6 lists examples of MPI delivery scenarios based on HEVC profiles 5.0–6.2 for Option 2. In some examples, Option 2 is used to support 720p video at level 5.x. [Table 6] Example of MPI transmission scenario for HEVC levels 5.x and 6.x: Option 2 [Table 6]

[0077] According to another selectable configuration (hereinafter referred to as "Option 3"), the packing block 520 is configured to generate a 2D video 522 by spatially packing the texture and alpha layers of the multiplane image 200 into pairs of video frames, with different pairs carrying the multiplane image 200 corresponding to each different view of the scene at the corresponding time t. Option 3 differs from Option 1 in that the texture and alpha layers are packed into different time-interleaved video frames. Therefore, the frame rate of Option 3 is twice the frame rate of the original camera view, but the corresponding luma_picture_size is halved. The total frame rate in this example is 2 × 30 × number_of_views. For four views, the frame rate is 240fps, which is lower than the 300fps constraint.

[0078] Figure 10 shows the operation of packing block 520 in yet another example. The illustrated example is an example of Option 3 described above, in which four views (V0, V1, V2, V3) are transmitted. Depending on the temporal position in the MPI video sequence 1090, the transmitted frame 1000 may be an I frame or a P frame. In some examples, B frames 1000 may be transmitted at several temporal positions in the MPI video sequence 1090.

[0079] A magnified view of the transmitted pair of video frames (1000a, 1000b) shows their tile structure in more detail. In the illustrated example, frame 1000a contains a texture slice, and frame 1000b contains an alpha slice. The 16 tiles in video frame 1000a carry the texture channels of the 16 layers (D=16, Figure 2) of the multiplane image 200 corresponding to view V0. The 16 tiles in video frame 1000b carry the alpha channels of the 16 layers (D=16, Figure 2) of the multiplane image 200 corresponding to view V0. Each set of 16 tiles is arranged using a 4x4 spatial layout. The next two video frames 1000 have a similar structure and carry the texture channels and alpha channels of the multiplane image 200 corresponding to view V1, respectively. The same applies to subsequent frames. In another example, the first frame 1000a contains all alpha layers, and the second frame 1000b contains all texture layers.

[0080] In yet another example, auxiliary pictures, as defined in the H.264 / AVC fidelity range extension or the Multiview-HEVC extension, may be used to mimic the temporally interleaved transmission of alpha and texture layers. The packed alpha layer can be compressed into an auxiliary picture corresponding to the primary coded picture, carrying the packed texture layer. To reconstruct the multiplane image, the corresponding decoder must be appropriately configured to decode the auxiliary picture.

[0081] Compared to Option 1, Option 3 reduces the picture size by half and doubles the total frame rate. For DPB analysis, the minimum DPB size is 2 × number_of_views. Table 7 below shows an exemplary MPI transmission scenario for Option 3. The parameters shown in Table 7 are applicable to both CTUSize=64 and CTUSize=32. [Table 7] Example of MPI transmission scenario for HEVC levels 5.x and 6.x: Option 3 [Table 7]

[0082] A challenging factor in designing the packing operation of the packing block 520 is ensuring compliance with the relevant MaxLuMAP constraints. In some embodiments, at least some of the packing modifications listed below can be applied in addition to options 1-3 described above to make the packing relatively more compact for such compliance. 1) Reduce the resolution of some layers: For example, the alpha layer may be downsampled by 2x (2x) horizontally, vertically, or both, for all alpha layers or several selected layers. The upsampling filter used by the decoder to restore the alpha layer to its original size is specified in the metadata. In another example, a subset of texture layers may be downsampled by a factor of two horizontally, vertically, or both. The upsampling filter used by the decoder to restore the texture layers to their original size is specified in the metadata. In some examples, texture / alpha layers can be downsampled either in a paired manner (e.g., by selecting a subset of layers and applying the same downsampling coefficient to the corresponding texture and alpha layers) or in a non-paired manner (e.g., downsampling can be applied independently to any selected layer). 2) Use an unequal number of layers for texture layers and alpha layers: For example, reduce the number of texture layers while maintaining the same number of alpha layers, merging 16 texture layers into 4 layers (for example, by merging 4 neighboring layers together), and still using 16 alpha layers. In rendering stages 610 and 606, the merged texture layers are used along with each of the 4 corresponding alpha layers. 3) Reducing the bit depth of the alpha layer: • Alpha layers may not require the full bit depth (e.g., 10 bits) used in alpha layers. Reducing the bit depth of alpha layers and packing them into a bit plane can reduce the spatial resolution required to store the alpha layers. For example, if alpha layers are quantized using 5 bits and two adjacent alpha layers in the bit plane are packed to form a 10-bit alpha signal, the effective spatial resolution of the alpha layers can be reduced by half. In some cases, such packing may introduce additional high frequencies into the alpha signal, which can cause corresponding artifacts after lossy compression, so caution is needed with this technique. 4) Utilize the dummy Cb / Cr component of the alpha layer: In some examples of the packing options above, the alpha layer has meaningful values ​​in the Y component, while the CbCr component is filled with dummy constant values. These dummy CbCr components can then be used for more useful purposes. Considering the case YCbCr4:2:0 as an exemplary example, the CbCr component can be used to carry the downsampled alpha plane. For example, with 16 layers, eight layers can be selected (based on some appropriate metric) to preserve the original resolution and place in the Y component. The other eight layers are downsampled and placed in the CbCr components of the eight full-resolution alpha layers. In some cases, the CbCr component usually has very different properties compared to the Y component, which can cause difficulties in predictive / motion compensation during HEVC coding operation, so caution is needed with this technique. 5) Block-based MPI generation: One approach to reduce the number of layers is to enable block-based MPI. This technique relies on the assumption that each block can typically have a different depth range. As a result, it is possible to reduce the number of layers for some blocks. The block size can be chosen to be an integer multiple of the CTU size to facilitate compression operations. The complexity of MPI generation can also be reduced by using larger block sizes. Furthermore, larger block sizes result in a corresponding reduction in metadata overhead. For example, a single 720p picture can be divided into several large blocks, e.g., a 5x3 large block, each with a size of 256x256. MPI generation is then performed individually for each such block. Coding gains are likely to be realized because the number of layers required to produce satisfactory rendering for a large block is typically less than the number of layers required to guarantee similar quality for the entire 720p picture. In some cases, four layers are sufficient to achieve such coding gains. In some cases, the number of MPI layers may differ for different large blocks. For example, blocks with relatively complex scene content may require more MPI layers than blocks with simpler scene content. The latter may require very few MPI layers to achieve good rendering quality. In some cases, there may be too many large blocks, which prevents each of the large blocks from being placed in a single common tile for the aggregated picture (due to a potential violation of the constraint on the maximum number of tile rows / columns), so caution is needed with this technique. Also, several large blocks with layers of different depths may cause additional boundary artifacts. Therefore, additional post-processing may be required. 6) The tiles are not aligned: In some of the embodiments described above, each MPI texture or alpha layer is padded to an integer multiple of the CTU size to facilitate compression. However, such padding may not be necessary in at least some applications, such as those where new view renderings are noticeably cropped by a factor of, for example, 0.8 or less. This feature can also be used to reduce memory usage.

[0083] In various cases, at least some of the above variations may be applied in a combined form. For example, a combination of variations 1) and 2) is compatible with level 5.x and delivers 720p with option 1 packing using the parameters listed in Table 8. [Table 8] Example of MPI transmission scenario for HEVC level 5.x: Option 1 [Table 8]

[0084] In some cases, the original image of the reference camera view may also be transmitted along with the MPI layer representation. The original image can then be used for post-processing to improve the quality of view synthesis. In some cases, the original image may be packed as an additional texture layer (in which case the total number of layers will be D+1). The corresponding alpha layer may be filled with (dummy) constant values. In some other cases, the original image may replace an existing texture layer (e.g., one with the smallest cumulative weight). The corresponding alpha layer is also replaced with (dummy) constant values. In both cases, metadata is signaled to enable the decoder to properly process the received transmission.

[0085] Figure 11 shows the operation of packing block 520 in yet another example. The illustrated example combines modifications 1) and 2) above to produce a 2D video frame 1100 in which four texture layers and sixteen downsampled alpha layers are packed. Frame 1100 contains texture slices 1110 and alpha slices 1150 stacked vertically (from top to bottom). The four tiles in texture slice 1110 are packed using a 1x4 layout. The sixteen tiles in alpha slice 1150 are packed using a 2x8 layout. The picture size in this example is the same as the picture size in option 1, which has four layers.

[0086] Table 9 shows additional examples to support the reduced 720p use case. The corresponding multiplane image 200 has eight layers (D=8). The 2D video frame has eight texture layers at the original resolution and eight alpha layers downsampled with a factor of 2. Option 1 is used for packing. [Table 9] Example of MPI transmission scenario for HEVC level 5.x: 8 layers: Option 1 [Table 9]

[0087] Figure 12 is a flowchart of an MPI encoding method 1200 according to one embodiment. The MPI encoding method 1200 may be implemented in an MPI encoder 500 using options 1 to 3 described above. Method 1200 includes, in block 1202, receiving an MPI representation MPI(s,t) of one camera view s for time t=0, .,T-1. The received MPI representation is an example of an MPI video 502. Method 1200 also includes performing MPI preprocessing in block 1204. The operation in block 1204 can be performed using a preprocessing block 510. Method 1200 further includes an option-specific set (12101, 12102, 12103) of packing operations in block 1206. A decision block 1208 is used to direct the processing flow of block 1206 to one of the selected option-specific sets (12101, 12102, 12103). The packing operation for set 12101 performs option 1 described above. The packing operation for set 12102 performs option 2 described above. The packing operation for set 12103 performs option 3 described above. The operation of block 1206 can be performed using packing block 520.

[0088] The MPI coding method 1200 also includes a video compression operation in block 1212. The video compression operation is applied to the packed 2D video frames generated in block 1206 and can be performed using the video encoder 530. The MPI coding method 1200 also includes multiplexing the compressed video bitstream and MPI metadata in block 1212. The multiplexing operation in block 1212 can be performed using the multiplexer 540. In some examples, for example, if the MPI metadata is static throughout the bitstream duration, the metadata is sent once and block 1212 may be omitted or bypassed. The multiplexing operation in block 1212 is performed in examples where the MPI metadata is different for each picture. The decision block 1216 of the MPI coding method 1200 controls the termination from the loop (1206, 1212, 1214) at the end of the video sequence. Upon such termination, the operation of the final block 1218 is performed and the MPI coding method 1200 terminates.

[0089] Figure 13 is a flowchart of an MPI decoding method 1300 according to one embodiment. The MPI encoding method 1300 is generally compatible with the MPI encoding method 1200 and can be implemented in an MPI decoder 600. Method 1300 includes receiving a bitstream of one camera view s in block 1302. In some examples, the received bitstream is an example of a coded bitstream 542. The determination block 1304 of Method 1300 is used to determine whether the received bitstream is an MPI video bitstream. Depending on the "No" or "Yes" determination in determination block 1304, the operation of block 1306 or block 1308 is then performed. The operation of block 1306 represents conventional video decoding. In contrast, the operation of block 1308 belongs to a processing loop that includes blocks (1308-1316) configured to perform MPI video decoding.

[0090] The operation of block 1308 involves parsing MPI metadata. The parsing operation of block 1308 can be performed using the demultiplexer 640. The parsing operation allows the decoder to obtain relevant MPI information and packing parameters, such as the number (M) of DPB output pictures required to reconstruct one complete MPI representation, packing arrangement, number and depth of layers, post-processing parameters, and camera parameters. As described above, in some cases the texture layer and alpha layer may be interleaved in time. In such cases, the decoder needs to have multiple easily accessible pictures (video frames) to reconstruct one corresponding multiplane image 200 at time t. For example, M=1 for packing option 1, M=D for packing option 2, and M=2 for packing option 3.

[0091] The operation of block 1310 involves decoding a portion of the bitstream corresponding to M pictures, including the texture and alpha layers necessary to reconstruct the image MPI(s,t) at time t. If the bitstream contains only the data for the still image of view s, the decoder operates to decode the entire bitstream. Otherwise, for each time t, the decoder operates to decode the portion of the bitstream containing the output picture necessary to reconstruct the multiplane image 200 representing time t.

[0092] The operation of block 1312 includes unpacking and post-processing the texture and alpha layers from the decoded output picture, and assembling the layers to reconstruct the image MPI(s,t) at time t. The operation of block 1314 includes performing view composition to render image I(t) using the image MPI(s,t), layer depth information, and camera parameters. In various cases, the new view may be the reference view s itself or any virtual view specified by the view input 1313. The determination block 1316 controls the termination from the loop (1308-1314) at the end of the video sequence. Upon such termination, the operation of the final block 1318 is performed and the MPI decoding method 1300 terminates.

[0093] If multiple views are submitted, the decoder operates to execute multiple instances of method 1300 in parallel. The outputs generated by each block 1314 of those multiple instances of method 1300 are merged during the image rendering operation by calculating a weighted sum of their outputs, for example, as described above with reference to Figure 3. If new views are not stationary, the weights become a function of time and are recalculated continuously, for example, as described above with reference to Figure 4.

[0094] Metadata design This section describes the MPI metadata used to properly configure and support various MPI video decoding operations in various examples and scenarios. As described above, the MPI metadata is transmitted by the MPI video encoder 500 via the MPI metadata stream 524. In various examples, the MPI metadata stream 524 may carry one or more of the following categories of metadata: • Basic MPI information: Width and height of the reference view; Number of MPI layers (D); Number of simultaneous views. • Packing / Placement Information: Packing options: For example, option 1, 2, or 3; Texture / Alpha Placement: Side-by-side or top-and-bottom; Is texture layer merging enabled? If yes, scale ratio; Is alpha layer downsampling effective? If yes, what are the downsampling coefficients? Block-based MPI deployment? The texture / alpha, view ID, layer ID, and layer depth for each tile. In some examples, the layer order may be an implicit order from furthest to closest relative to the reference view, or vice versa. In some other examples, the layer order may be explicitly signaled. Packing / sending the original reference image. • Related MPI preprocessing: Did you use adaptive layer merging? If yes, what is the output depth for each layer (precisely quantized)? Was dynamic range adjustment for the texture / alpha channel used? If yes, please specify the adjustment method (linear stretching, non-linear reshaping, etc.) and the corresponding parameters. • Related MPI post-processing: Alpha normalization after decoding? When block-based MPI is used, parameters are required to configure boundary artifact reduction / mitigation. • Camera and reference view-related metadata: (inherent / external matrices, field of view, depth, etc.). This metadata category can typically be used for novel view synthesis. However, in some cases, this category may be optional.

[0095] For illustrative purposes only, without any implicit limitations, syntax examples are presented for the following categories: 1) Basic MPI information, 2) Packing / Placement information, and 3) MPI preprocessing information. Based on the provided examples, those skilled in the art will readily understand how to handle the remaining categories of metadata listed above. Furthermore, with respect to camera-related information, the MPEG Immersive Video (MIV) specification provides syntax examples for both camera external syntax (Section 8.3.2.6.6) and camera internal syntax (Section 8.3.2.6.7). Versatile Supplemental Enhancement Information (VSEI) describes examples of multiview acquisition information SEI (MAI SEI) messages, including internal and external parameters for perspective projection. In some examples, such SEI messages are adapted to describe camera information. The following corresponding documents are incorporated in their entirety by reference into this specification: (1) ISO / IEC 23090-12: Information technology -- Coded representation of immersive media -- Part 12: MPEG Immersive video, and (2) H.274: VSEI; ITU-TH274, Versatile supplemental enhancement information messages for coded video bitstreams (08 / 2020).

[0096] For depth-related information, the VSEI includes a depth representation information SEI message. This SEI message contains the element depth_rep_info_element(OutSign, OutExp, OutMantissa, OutManLen). In some examples, this element is reused for MPI metadata purposes to signal depth information. An example of the corresponding syntax is as follows: [Table 10] Definition of depth_rep_info_element() [Table 10]

[0097] Figures 14A and 14B show exemplary syntax of SEI messages 1400 configured to transmit MPI metadata, using several examples. SEI messages 1400 are configured to cover MPI information prior to packing operations and include basic MPI information (Category 1) and MPI preprocessing information (Category 3). The semantics of SEI messages 1400 are described as follows: A value of 1 for `mpi_is_one_view_among_multiple_flag` indicates that there is only one camera view in the camera setup. A value of 0 for `mpi_is_one_view_among_multiple_flag` indicates that there are two or more camera views in the camera setup. mpi_view_id specifies the view identifier of the current camera view. Note: The mpi_view_id is used to identify camera parameters for multiview camera setup in the multiview acquisition information SEI message. `mpi_layer_width_in_luma_samples` specifies the width in units of luma samples for the original MPI texture and alpha mapping layers. mpi_layer_height_in_luma_samples specifies the height in units of luma samples for the original MPI texture and alpha mapping layers. Note: This is one exemplary way to signal the cropped decoded MPI layer size. Another way is to signal the cropping window offset. mpi_log2_ctu_size_minus5+5 specifies the luma coding tree block size for each CTU. Note: This value is a hint indicating what padding will be used to enable tiling of the MPI layer. mpi_bit_depth_texture_minus8+8 specifies the bit depth of the samples for the luma and chroma arrays for the texture layer. mpi_bit_depth_alpha_minus4+4 specifies the bit depth of samples in the Luma Array for the alpha map layer. mpi_num_layers_minus1+1 specifies the number of texture layers and opacity layers for the MPI scene representation. mpi_num_regions_minus1+1 specifies the number of regions for texture and opacity layers for the MPI scene representation. num_region_rows_minus1+1 specifies the number of rows in the region. num_region_cols_minus1+1 specifies the number of region columns. mpi_depth_equal_distance_flag[i]=1 indicates that equal distances are used to generate the MPI layers and depth parameters for each layer in the i-th region. Z[i][j] can be derived using the nearest depth value ZNear[i] and the farthest depth value ZFar[i]. Note: The depth value of the i-th MPI layer in the j-th region is given by the following equation: Z[i][j]=j*(ZFar[i]-ZNear[i]) / (mpi_num_layers_minus1)+ZNear[i] When mpi_layer_depth_equal_distance_flag is equal to 0, it indicates depth information for each layer in the i-th region following the SEI. The variables in column x of Table 11 are derived from the variables in columns s, e, n, and v of Table 11 as follows: If the value of e is within the range of 0 to 127 (excluding 0), then x is (-1). s *2 e-31 *(1+n÷2 v It is set to be equal to ). In other cases (when e is equal to 0), x is (-1). s *2 -(30+v) Set *n to equal to n. [Table 11] Association between depth parameter variables and syntactic elements [Table 11] If mpi_alpha_mapping_flag[i] is equal to 1, it specifies that reshaping will be applied to the alpha map in the i-th region. If mpi_alpha_mapping_flag[i] is equal to 0, it specifies that reshaping will not be applied to the alpha map in the i-th region. alpha_quant_precision_minus11[i]+11 specifies the precision for quantizing the maximum alpha map value in the i-th region. max_alpha[i][j] specifies the maximum alpha value in the j-th MPI layer of the i-th region. Note: In some examples, the following syntax and semantics may be moved to the MPI packing information SEI (see Figure 15). A value of 1 for `mpi_texture_layer_merge_flag` indicates that the texture layers will be merged. A value of 0 for `mpi_texture_layer_merge_flag` indicates that the texture layers will not be merged. log2_mpi_num_layers_in_one_merged_texture_layer_minus1+1 specifies the number of texture layers in one merged texture layer. mpi_layer_id_in_one_merged_texture_layer[i][j] specifies the texture layer ID for the j-th layer in the i-th merged texture layer.

[0098] Figure 15 shows exemplary syntax of SEI message 1500 configured to convey MPI metadata, with several examples. SEI message 1500 is configured to cover MPI packing. In some examples, option 1 or option 3 is preferred, for example, because it provides flexibility and is a simpler implementation. Therefore, the examples shown in Figure 15 cover options 1 and 3 for illustrative purposes. Furthermore, option 1 allows for scaling of the alpha map and merging of texture layers. SEI message 1500 is configured for a single camera view, assuming that textures are processed before the alpha map. The semantics of SEI message 1500 are described as follows: A value of 0 for `mpi_arrangement_type` indicates that the spatial arrangement of frame 0 and frame 1 is applied. A value of 1 for `mpi_arrangement_type` indicates that the temporal interleaving of frame 0 and frame 1 is applied. Note: For each specified frame packing arrangement scheme, there are two configuration frames, referred to as frame 0 and frame 1, where frame 0 is associated with a spatially packed texture image and frame 1 is associated with a spatially packed alpha map image. If mpi_arrangement_type is equal to 0, the configuration frame associated with the top-left sample of the decoded frame is considered to be configuration frame 0, and the other configuration frame is considered to be configuration frame 1. If mpi_arrangement_type is equal to 1, the first decoded frame in the current coded layered video sequence (CLVS) is configuration frame 0, the next decoded frame in output order is configuration frame 1, and the display time of configuration frame 0 is delayed to match the display time of configuration frame 1. Note: Other placement types may also be used. For example, textures and alpha maps can be interleaved on a pixel-by-pixel basis. Alternatively, the textures and alpha maps of each layer can be packed first, and then all MPI layers can be packed. mpi_alpha_scale_factor_x_minus1+1 specifies the alpha map scale factor in the x direction. mpi_alpha_scale_factor_y_minus1+1 specifies the scale factor for the alpha map in the y direction. A value of 0 for `mpi_spatial_arrangement_type` indicates that top-bottom packing is used for frames 0 and 1. A value of 1 for `mpi_spatial_arrangement_type` indicates that side-by-side packing is used for frames 0 and 1. mpi_num_texture_layers_in_height_minus1+1 specifies the number of spatially packed and merged texture layers in height within frame 0. mpi_num_alpha_layers_in_height_minus1+1 specifies the number of spatially packed alpha layers in height within frame 1. mpi_num_alpha_layers_in_height_minus1+1 specifies the number of spatially packed layers in height for merged texture layers in frame 0 and alpha layers in frame 1.

[0099] Figure 16 shows exemplary syntax of an SEI message 1600 configured to convey MPI metadata, as in several other examples. The SEI message 1600 specifies MPI scene representation information that can be used for view compositing. In some examples, the SEI message 1600 can work together with a multiview acquisition information SEI message for view compositing. The multiview acquisition information SEI message specifies intrinsic and extrinsic parameters for all of the reference camera views. When multiple video bitstreams are available, the reconstructed new view can be rendered from a nearby multiview MPI.

[0100] Table 12 below shows an exemplary SEI message for MPI messaging according to another embodiment with a simpler syntax structure. Table 12 also includes two new syntax elements, namely mpi_layer_depth_or_disparity_values_flag and mpi_depth_equal_distance_type_flag. [Table 12] Exemplary syntax for MPI information SEI messages [Table 12]

[0101] The use of SEI message 1600 and Table 12 depends on the definition of the following variables. In this specification, the width and height of the cropped decoded output picture in a unit of luma sample are indicated by CroppedWidth and CroppedHeight, respectively. A chroma format indicator, indicated by ChromaFormatIdc in this specification. The cropped decoded picture array decPicCurr0[cIdx][x][y], where, cIdx=0..(ChromaFormatIdc==0)?0:2, x=0..(cIdx==0)?CroppedWidth:CroppedWidth / SubWidthC-1, y=0..(cIdx==0)?CroppedHeight:CroppedHeight / SubHeightC-1 - In the output order, the cropped decoded picture array that follows in time: decPicCurr1[cIdx][x][y], Here, cIdx=0..(ChromaFormatIdc==0)?0:2, x=0..(cIdx==0)?CroppedWidth:CroppedWidth / SubWidthC-1, y=0..(cIdx==0)?CroppedHeight:CroppedHeight / SubHeightC-1 The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified.

[0102] The semantics of SEI message 1600 and the messages in Table 12 are explained as follows: A value of 1 for `mpi_cancel_flag` indicates that the MPI SEI message cancels the persistence of any previous MPI SEI message in the output order applied to the current layer. A value of 0 for `mpi_cancel_flag` indicates that the MPI will continue. The `mpi_persistence_flag` flag specifies the persistence of MPI SEI messages for the current layer. A value of 0 for `mpi_persistence_flag` indicates that MPI SEI messages apply only to the current decrypted picture. When mpi_persistence_flag is equal to 1, it specifies that the MPI SEI message is applied to the current decoded picture and persists in output order for all subsequent pictures on the current layer until one or more of the following conditions are true: A new CLVS for the current layer will begin. The bitstream is ending. In the output order, the picture in the current layer within the AU associated with the MPI SEI message, following the current picture, is output. mpi_view_id specifies the view identifier of the current camera view. Note: mpi_view_id is used in the multiview acquisition information SEI message to identify camera parameters for multiview camera setup. The view identifier for the i-th view in the current CVS is equal to ViewId[i]. This is described in Section 8.19.2 of ITU-TH.274,(VSEI)(05 / 2022) in the semantics of the Scalability Dimension Information (SDI) SEI message, which is incorporated into this specification by reference. mpi_num_layers_minus1+1 specifies the number of texture layers and opacity layers for the MPI scene representation. A value of 0 for `mpi_layer_depth_or_disparity_values_flag` indicates that the depth information signaled in the MPI SEI message is interpreted as a depth value. A value of 1 for `mpi_layer_depth_or_disparity_values_flag` indicates that the depth information signaled in the SEI message is interpreted as a disparity value. The relationship between the disparity value D and the depth value Z is D = 1 ÷ Z. The value of mpi_layer_depth_equal_distance_flag being equal to 1 indicates that equal distances are used to generate the MPI layers, and the depth parameter of each layer Z[i] can be derived using the nearest depth value ZNear and the farthest depth value ZFar. Alternatively, the disparity parameter D[i] of each layer can be derived using the disparity values ​​DNear and DFar. A value of 0 for `mpi_depth_equal_distance_type_flag` indicates that the depth values ​​are equal in distance. A value of 1 for `mpi_depth_equal_distance_type_flag` indicates that the depth values ​​are equal in distance in parallax. If mpi_layer_depth_or_disparityvalues_flag is equal to 0, If mpi_depth_equal_distance_type_flag is equal to 0: Depth value Z[mpi_num_layers_minus1-i] =i*(ZFar-ZNear)÷(mpi_num_layers_minus1)+ZNear, (m1) Parallax value D[i]=1÷Z[i] (m2) If mpi_layer_depth_or_disparityvalues_flag is equal to 0, If mpi_depth_equal_distance_type_flag is equal to 1: Depth value Z[i] =1÷(i*(1÷ZNear-1÷ZFar)÷(mpi_num_layers_minus1)+1÷ZFar), (m3) Parallax value D[i] = 1 ÷ Z[i] (m4) If mpi_layer_depth_or_disparityvalues_flag is equal to 1, If mpi_depth_equal_distance_type_flag is equal to 0, Parallax value D[mpi_num_layers_minus1-i] =1÷(i*(1÷DFar-1÷DNear)÷(mpi_num_layers_minus1)+1÷DNear), (m5) Depth value Z[i] = 1 ÷ D[i] (m6) If mpi_layer_depth_or_disparityvalues_flag is equal to 1, If mpi_depth_equal_distance_type_flag is equal to 1: Parallax value D[i] =i*(DNear-DFar)÷(mpi_num_layers_minus1)+DFar, (m7) Depth value Z[i] = 1 ÷ D[i] (m8) A value of 0 for `mpi_layer_depth_equal_distance_flag` indicates that the depth information for each layer follows in the SEI message. Layer index 0 is associated with the layer with the farthest depth value or the smallest disparity value. Layer index `mpi_num_layers_minus1` is associated with the layer with the closest depth value or the largest disparity value. The depth value Z[i] or disparity value D[i] is assumed to be monotonic. The variables ZNear, ZFar, and Z[i] are derived from the variables in columns s, e, n, and v of Table 11, as shown above. Note: In some applications, parallax is used instead of depth (the relationship between the parallax value D and the depth value Z is D = 1 / Z). The value of mpi_texture_opacity_interleave_flag equal to 1 indicates that the component planes of the output cropped decoded picture in the output order form a temporal interleaving of alternating first and second constituent frames, as shown in Figure 10. The mpi_texture_opacity_arrangement_flag identifies the indicated interpretation of the sample array of the output cropped decoded picture, as specified in Table 13. [Table 13] Definition of mpi_texture_opacity_arrangement_flag [Table 13] For each specified frame packing arrangement, there are two configuration frames, referred to as frame 0 and frame 1. If mpi_texture_opacity_interleave_flag is equal to 0, the configuration frame associated with the top-left sample of the decoded frame is considered configuration frame 0, and the other configuration frame is considered configuration frame 1. If mpi_texture_opacity_interleave_flag is equal to 1, the first decoded frame in the current CLVS is configuration frame 0, and the next decoded frame in the output order is configuration frame 1. The display time of configuration frame 0 is delayed to match the display time of configuration frame 1. The two configuration frames form the spatially packed texture and opacity images of the MPI, with frame 0 associated with the spatially packed texture image and frame 1 associated with the spatially packed opacity image. `mpi_frame_num_layers_minus1_in_height+1` specifies the number of spatially packed layers at the height of frame 0 and frame 1. The number of spatially packed layers at the width of frame 0 and frame 1 is set to be equal to (mpi_num_layers_minus1+1) divided by (mpi_frame_num_layers_minus1_in_height+1). Let the variables hLayers and wLayers be the number of spatially packed layers in terms of height and width, respectively. A value of 1 for `mpi_layer_crop_window_flag` indicates that the texture and opacity layer cropping window offset parameters are present in the SEI. A value of 0 for `mpi_layer_crop_window_flag` indicates that the texture and opacity layer crop window parameters are not present in the SEI. mpi_layer_crop_win_left_offset, mpi_layer_crop_win_right_offset, mpi_layer_crop_win_top_offset, and mpi_layer_crop_win_bottom_offset specify that the cropping window is applied to the picture using left, right, top, and bottom offset values ​​in units of Luma samples for each decoded MPI metadata texture and opacity layer. If mpi_layer_crop_win_flag is equal to 0, the values ​​of mpi_layer_crop_win_left_offset, mpi_layer_crop_win_right_offset, mpi_layer_crop_win_top_offset, and mpi_layer_crop_win_bottom_offset are assumed to be 0. The variables layerWidth and layerHeight specify the width and height of the decoded MPI layer, respectively. If mpi_texture_opacity_interleave_flag is equal to 1, layerWidth=CroppedWidth / wLayers, layerHeight=CroppedHeight / hLayers In other cases, If mpi_texture_opacity_arrangement_flag is equal to 0, the following applies: layerWidth=CroppedWidth / wLayers, layerHeight=CroppedHeight / (hLayers*2) In other cases, if mpi_texture_opacity_arrangement_flag is equal to 1, the following applies: layerWidth=CroppedWidth / (wLayers*2), layerHeight=CroppedHeight / (hLayers) For any i within the range of 0 to mpi_num_layers_minus1, the cropping window for the i-th MPI texture layer is specified in frame 0, and the cropping window for the i-th opacity layer is specified in frame 1. Let the variables k = i % wLayers and m = (ik) / hLayers, respectively. The cropping window for the i-th MPI layer includes a luma sample with horizontal picture coordinates from k*layerWidth+SubWidthC*mpi_layer_crop_win_left_offset to (k+1)*layerWidth-(SubWidthC*mpi_layer_crop_win_right_offset+1) and vertical picture coordinates from m*layerHeight+SubHeightC*mpi_layer_crop_win_top_offset to (m+1)*layerHeight-(SubHeightC*mpi_layer_crop_win_bottom_offset+1). If ChromaFormatIdc is not equal to 0, the corresponding specified samples of the two chroma arrays are samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified chroma sample.

[0103] In another example, consider the following semantics. mpi_picture_num_layers_minus1_in_height+1 specifies the number of spatially packed layers at the height of picture0 and picture1. The variable hLayers is set to equal mpi_picture_num_layers_minus1_in_height+1, and the variable wLayers is set to equal (mpi_num_layers_minus1+1) / hLayers.

[0104] The variables fWidth and fHeight specify the width and height of picture0 and picture1, respectively, and are derived as follows: If mpi_texture_opacity_interleave_flag is equal to 1, the following applies: fWidth = CroppedWidth fHeight=CroppedHeight In other cases (when mpi_texture_opacity_interleave_flag is equal to 0) If mpi_texture_opacity_arrangement_flag is equal to 0, the following applies: fWidth=CroppedWidth,fHeight=CroppedHeight / 2 In other cases (when mpi_texture_opacity_arrangement_flag is equal to 1), the following applies: fWidth=CroppedWidth / 2,fHeight=CroppedHeight

[0105] Let the variables cWidth = fWidth / subWidthC and cHeight = fHeight / subHeightC. The array picture0[cIdx][x][y] specifies the samples in picture0, and the array picture1[cIdx][x][y] specifies the samples in picture1, where cIdx = 0.ChromaFormatIdc == 0)? 0 : 2. x = 0..(cIdx == 0)? fWidth : cWidth - 1, y = 0..(cIdx == 0)? fHeight : cHeight - 1, and is derived as follows: If mpi_texture_opacity_interleave_flag is equal to 1, the following applies: picture0[cIdx][x][y]=decPicCurr0[cIdx][x][y] picture1[cIdx][x][y]=decPicCurr1[cIdx][x][y] In other cases (where mpi_texture_opacity_interleave_flag is equal to 0) Let the variable cW = (cIdx == 0) ? fWidth : cWidth. Let the variable cH = (cIdx == 0) ? fHeight : cHeight. If mpi_texture_opacity_arrangement_flag is equal to 0, the following applies: picture0[cIdx][x][y]=decPicCurr0[cIdx][x][y] picture1[cIdx][x][y]=decPicCurr0[cIdx][x][y+cH] Otherwise (where mpi_texture_opacity_arrangement_flag is equal to 1), the following applies: picture0[cIdx][x][y]=decPicCurr0[cIdx][x][y] picture1[cIdx][x][y]=decPicCurr0[cIdx][x+cW][y]

[0106] The variables layerWidth and layerHeight specify the width and height of the decoded MPI layer, respectively. The variables are derived as follows: layerWidth = fWidth / wLayers layerHeight = fHeight / hLayers

[0107] In one embodiment, an example of the MPI reconfiguration process is described as follows: The output of this process is as follows: • 4Dmpi texture layer array recTextureLayer[i][cIdx][w][h], where i=0..mpi_num_layers_minus1, cIdx=0..(ChromaFormatIdc==0)?0:2, w=0..(cIdx==0)?layerWidth:layerWidth / SubWidthC-1, and h=0..(cIdx==0)?layerHeight:layerHeight / SubHeightC-1. • 3Dmpi opacity layer array recOpacityLayer[i][w][h], where i=0..mpi_num_layers_minus1, x=0..layerWidth-1, and y=0..layerHeight-1. The arrays recTextureLayer and recOpacityLayer are derived as follows:

number

[0108] In various additional examples, other appropriate syntax may be used as well. In some examples, syntax is used that allows coverage of both MPI scene information and MPI packing information.

[0109] VUI signaling considerations Since the packed MPI format is not intended to be viewed directly by the end user, signaling is required to inform the playback device of this suboptimal viewing information. One way to do this is to overload the vui_non_packed_constraint_flag semantics. The revised semantics are shown below, with the added syntax in italics (or represented as "<<...>>"): That vui_non_packed_constraint_flag is equal to 1 specifies that there is no <<MPI information SEI>> present in the bitstream applied to CLVS, or that there is no frame packing arrangement SEI message present. That vui_non_packed_constraint_flag is equal to 0 places no such constraint.

[0110] Improvement of depth value coding As defined previously, the structure of depth_rep_info_element() in Table 10 is defined as follows:

Table 14

Table 15

[0111] In an exemplary embodiment, the modifications (shown in italics or represented as “<<...>>”) are proposed as follows: When it is necessary to signal an array of these elements (e.g., to specify the depth value of each layer in the coding loop), predictive-based coding may be introduced to further reduce the bit overhead of the syntax elements in this structure. In one embodiment, the delta value of the elements within the loop may be coded instead of the absolute value, and variable length coding (such as ue(v) or se(v)) may be used instead of fixed length coding. For example,

Table 16

Table 17

[0112] In the exemplary implementation shown here, for a given value in the 16-layer depth representation, the bits used to signal the exponent can be reduced from 112 bits to 32 bits using a prediction-based method. [Table 18]

[0113] MPI transmission in the MIV coding standard MPEG Immersive Video (MIV) specification (ISO / IEC23090-12:2021(E) / AMD.1:2022, <<Information technology-Coded representation of immersive media> >-Part12:<<MPEGImmersive video> >) is the V3C specification (ISO / IEC23090-5:2023(E), <<Information technology-Coded representation of immersive media> This is an extension of Part 5 (Visual volumetric video-based coding (V3C) and video-based point cloud compression (V-PCC)), both of which are incorporated in their entirety by reference into this specification and define a profile called the “MIV Extended Restricted Geometry Profile” intended for the delivery of MPI / MSI content. MPI / MSI video is associated only with texture and transparency attributes. The two attributes are expected to be carried either in two independent V3C_AVD units or frame-packed and carried in one V3C_PVD unit. In the first case, two independent element video decoders (e.g., HEVC, VVC, etc.) are used to decode a multiplexed MIV bitstream containing one atlas substream and two video substreams. In the latter case, a single 2D conventional video decoder is used to decode the frame-packed attributes. However, the current profile definition in Table A-1 of the MIV specification does not appear to support the latter case, as shown in the Appendix. As used in V3C, the term "atlas" refers to "a collection of 2D bounding boxes and their associated information that correspond to volumes in 3D space, which are arranged on a rectangular frame and whose volumetric data is rendered."

[0114] Table 16 shows an example of the proposed revised MIV Table A-1, which provides several exemplary edits to the existing MIV Extended Restricted Geometry profile and also proposes a new MIV Extended Restricted Geometry Packed profile to properly support both the two-stream and (packed) one-stream cases. The proposed modifications to MIV Table A-1 are as follows, with italicized (or represented as "<<...>>") text indicating proposed changes to existing syntax parameters and italicized bold text indicating new additions: 1) To support MPI using frame packing, add a column to define a new "MIV Extended Restricted Geometry Packed" profile. 2) Add vps_attribute_video_present_flag[atlasID] to the syntax element column and set its value appropriately for the two MPI profiles. 3) Add a pin attribute syntax element and set the appropriate value for the packed MPI packed profile. 4) Set the appropriate values ​​for the original syntax elements. [Table 16] Example of a modified MIV table A-1 with an extended, limited geometry packed profile [Table 19-1] [Table 19-2] A copy of the edited new semantics description from the V3C specification is provided in the Appendix.

[0115] MIV metadata for MPI information When MPI video is encoded according to the MIV coding standard, it is necessary to generate atlas data containing patches of information. Each patch contains a 2D bounding box, and its associated information is placed on a rectangular frame corresponding to a volume in 3D space. As a result, redundant patch information may be repeated, and the size of the atlas data increases with a large number of patches. Since constant patch information can be applied across MPI layers, a novel method for reducing the size of the atlas data is proposed. In one embodiment, a new flag (asps_patch_constant_flag) can be added to indicate that the same width, height, and patch mode apply to all patch syntax elements within atlas_sequence_parameter_set_rbsp(). For example, [Table 20] A value of 1 for `asps_patch_constant_flag` indicates that constraints present in the atlas style header, namely patch mode, width, and height, apply to all patches in the current atlas. A value of 0 for `asps_patch_constant_flag` indicates that each patch may have different constraints for the current atlas, such as width and height.

[0116] Consider the atlas_tile_layer_rbsp() function defined as follows. [Table 21] Next, in one embodiment, examples of newly proposed syntax elements in atlas_tile_header() and atlas_tile_data_unit(), shown in the following two tables in italics (or represented as "<<...>>"), may be defined as follows:

[0117] Regarding atlas_tile_header(): [Table 22] Here, the new semantics can be defined as follows: ath_num_patch_minus1 specifies the number of patches within the current atlas style. This specifies the number of textures and opacity layers for the MPI representation. ath_num_patch_in_height_minus1 + 1 specifies the number of patches in height. ath_patch_size_x_minus1 + 1 specifies the quantized width value of the patch. ath_patch_size_y_minus1 + 1 specifies the quantized height value of the patch. If there is a single tile within the atlas frame, it is as follows. num_patch_in_width = (ath_num_patch_minus1 + 1) / (ath_num_patch_in_height_minus1 + 1) ath_patch_size_x_minus1 = asps_frame_width / (num_patch_in_width * PatchSizeXQuantizer) - 1 ath_patch_size_y_minus1 = asps_frame_height / ((ath_num_patch_in_height_minus1 + 1) * PatchSizeYQuantizer) - 1 ath_patch_mode indicates the patch mode of the patch. That ath_patch_equal_3d_offset_d_flag is equal to 1 indicates that equal distances are used to generate the patch and the depth parameters of each patch.

[0118] Regarding atlas_tile_data_unit():

Table 23

[0119] When asps_patch_constant_flag is equal to 1, the patch_information_data structure does not exist in atlas_tile_data_unit(tileID), and similar information can be derived, for example, by using the information in the atlas style header, as described below.

number

[0120] Support for time-interleaved packing in V3C specifications V3C supports spatial domain packing of attributes (e.g., side-by-side or top-and-bottom) via V3C-packed video extensions. However, temporally interleaved packing is not supported. In exemplary embodiments, such support can be added to the specification by adding two new flags to Section 8.3.4.7, “Packing information syntax,” as shown in Table 17 below. The proposed additions are shown in italics.

[0121] As shown in Table 17, in exemplary embodiments, the syntax first uses a first flag (e.g., pin_attribute_same_dimension_flag) to check whether the dimensions of the attributes to be packed are the same. If the dimensions are not the same, then time interleaved packing is not enabled, as only VVC RPR can support this type of single-stream video. Otherwise, a second flag (e.g., pin_attribute_temporal_interleave_flag) is read to check whether time interleaving is enabled. At the same time, in exemplary embodiments, the syntax allows pin_region_xxx information (such as position (x,y) coordinates, width, and height) to be skipped, thus saving 64 bits. [Table 17] Example of the "Packing information syntax" table in the revised V3C specification table (Section 8.3.4.7). [Table 24-1] [Table 24-2] Newly proposed syntax elements The value of pin_attribute_same_dimension_flag[j] equal to 1 indicates that the attributes present in the packed video frames of the atlas with atlas ID j have the same spatial dimension. A value of 1 for pin_attribute_temporal_interleave_flag[j] indicates that attributes present in the packed video frame are packed in a temporally interleaved manner.

[0122] As illustrated, if the attributes have the same dimensions and temporally interleaved packing is used (i.e., pin_attribute_temporal_interleave_flag[j]=1), then signaling for the position and size (width and height) of (0,0) can be skipped.

[0123] MPI transmission using a scalable codec In one embodiment, scalable video coding (e.g., SVC, SHVC, etc.) can be used for MPI video transmission. For example, the base coding layer is a conventional 2D picture from the source camera, and the extension layers include packed MPI layers and their associated MPI metadata. Level constraints apply only to those coding layers. Alternatively, multiple extension layers may exist, each corresponding to a specific MPI layer.

[0124] MPI Reconfiguration with Partial Access to Layers In another embodiment, MPI rendering may use only a subset of layers necessary for partial decoding / access of coded layers within a packed picture. For example, rendering only the background requires only a subset of layers containing background information. Alternatively, rendering the foreground without the background may require only a subset of layers containing foreground information. Therefore, C S =Σ i C i S W i S Here, i is the index of the selected layer from 0 to D-1 (11) In such cases, the decoder may simply decode the partial bitstream corresponding to a subset of the layer and perform rendering. To support partial decoding, the tile / slice and / or subpicture coding features of conventional 2D video coding may need to be enabled. The decoder can also decode and render a “viewport” corresponding to a sub-region of the original full image dimension by appropriately performing the tile / slice / subpicture features. From the MPI information metadata stream, the decoder should be able to understand which spatial regions within the frame correspond to the selected layer and decode the bitstream of those regions.

[0125] Exemplary Hardware Figure 17 is a block diagram of a computing device 1700 according to one embodiment. The device 1700 may be used, for example, in a coding block 120 or a decoding block 130. The device 1700 includes an input / output (I / O) device 1710, a processing engine 1720, and a memory 1730. The I / O device 1710 may be used to enable the device 1700 to receive at least a portion of data stream 117 or 122 and output at least a portion of data stream 122 or 132.

[0126] Memory 1730 may have a buffer for receiving various inputs described above, for example, via a corresponding data stream. Once an input is received, memory 1730 provides various parts thereof to processing engine 1720, where it can be processed. Processing engine 1720 includes processor 1722 and memory 1724. Memory 1724 can store program code, when executed by processor 1722, that enables processing engine 1720 to perform various coding, decoding, image processing, and metadata operations described above. The program code may, among other things, include program code that embodies the various methods described above.

[0127] According to the above exemplary embodiments, which refer, for example, to one or any combination of part or all of the summary and / or FIGS. 1 to 17, a device for encoding a sequence of multi-plane images, the device comprising: at least one processor; at least one memory including program code; including the at least one memory and the program code, together with the at least one processor, cause the device to at least generate a sequence of video frames, each of the video frames including a plurality of tiles each representing a layer of one or more of the multi-plane images; generate a metadata bitstream for specifying at least a packing arrangement of the tiles within the sequence of video frames; generate a video bitstream by applying video compression to the sequence of video frames; multiplex the video bitstream and the metadata bitstream for transmission. There is provided a device configured as such. In the present specification, the term "tile" does not necessarily refer to a shape and / or size specified in the HEVC specification, nor is it limited to an integer multiple of CTUs, but refers to a portion of an image frame.

[0128] In some embodiments of the above device, the first frame of the sequence of video frames has tiles corresponding to the first multi-plane image, and the second frame of the sequence of video frames has tiles corresponding to the second multi-plane image.

[0129] In some embodiments of any of the above devices, the first and second multi-plane images are images of a scene from different camera positions.

[0130] In some embodiments of any of the above devices, the first and second multiplane images are images of the scene at different times.

[0131] In some embodiments of the above-described device, a frame in a sequence of video frames comprises a first tile set representing the texture layer of a first multiplane image and a second tile set representing the alpha layer of the first multiplane image.

[0132] In some embodiments of the above-described apparatus, the first and second tile sets each have a different number of tiles.

[0133] In some embodiments of the above-described device, the frames in a sequence of video frames include a first tile set representing a first multiplane image and a second tile set representing a second multiplane image.

[0134] In some embodiments of any of the above devices, the first and second multiplane images are images of the scene from different camera positions.

[0135] In some embodiments of any of the above devices, the first tile set includes tiles representing the texture layer of the first multiplane image and other tiles representing the alpha layer of the first multiplane image, and the second tile set includes tiles representing the texture layer of the second multiplane image and other tiles representing the alpha layer of the second multiplane image.

[0136] In some embodiments of any of the above devices, the frames of the video frame sequence further include a third tile set representing a third multiplane image and a fourth tile set representing a fourth multiplane image.

[0137] In some embodiments of the above-described device, the metadata bitstream includes supplemental extended information messages. In some embodiments of the above-described device, the frames of a sequence of video frames have tiles representing reference images.

[0138] In some embodiments of any of the above devices, the metadata bitstream includes parameters selected from a group including the size of the reference view, the number of layers in a multiplane image, the number of simultaneous views, one or more characteristics of the packing arrangement, layer merge information, dynamic range adjustment information for the texture channel or alpha channel, and reference view information.

[0139] For example, according to another exemplary embodiment described above, which references one or any combination of all or part of Figures 1-17 in the summary, a method is provided for encoding a sequence of multiplane images, the method comprising: generating a sequence of video frames, each video frame comprising a plurality of tiles representing one or more layers of multiplane images; generating a metadata bitstream for specifying at least the packing arrangement of the tiles in the sequence of video frames; generating a video bitstream by applying video compression to the sequence of video frames; and multiplexing the video bitstream and the metadata bitstream for transmission.

[0140] In some embodiments of the above method, when executed by at least one processor, a non-temporary computer-readable medium is provided that stores instructions causing at least one processor to perform an operation including the above method for encoding a sequence of multiplane images.

[0141] For example, according to yet another exemplary embodiment described above, referencing one or any combination of parts or all of Figures 1-17 in the Abstract, there is a device for decoding an received bitstream encoded with a sequence of multiplane images, the device comprising at least one processor and at least one memory containing program code, the at least one memory and program code together with the at least one processor, the device being configured to at least: demultiplex the received bitstream to obtain a video bitstream encoded with a sequence of video frames; obtain a metadata bitstream specifying at least the packing arrangement of tiles in the sequence of video frames, where the tiles represent layers of a multiplane image; reconstruct the sequence of video frames by applying video decompression to the video bitstream; and reconstruct the sequence of multiplane images based on the metadata bitstream using the tiles from the sequence of video frames.

[0142] In some embodiments of the above-described device, at least one memory and program code, together with at least one processor, are configured to cause the device to generate a sequence of visible images by rendering a sequence of multiplane images.

[0143] In some embodiments of the above-described device, rendering operations aimed at generating a composite visible image corresponding to a new view include: applying warping to layers of a multiplane image set corresponding to different reference camera positions, which is performed according to the new view; compositing the layers of the multiplane image set after warping to generate a corresponding set of individual visible images corresponding to the new view; and generating a composite visible image as a weighted sum of the individual visible images.

[0144] In some embodiments of any of the above devices, the multiplane image set includes one, two, three, or four multiplane images. In some other embodiments, the multiplane image set includes more than four multiplane images.

[0145] In some embodiments of the above-described device, the first frame of a sequence of video frames has a tile corresponding to a first multiplane image, and the second frame of a sequence of video frames has a tile corresponding to a second multiplane image.

[0146] In some embodiments of any of the above devices, the first and second multiplane images are images of the scene from different camera positions.

[0147] In some embodiments of any of the above devices, the first and second multiplane images are images of the scene at different times.

[0148] In some embodiments of the above-described device, a frame in a sequence of video frames comprises a first tile set representing the texture layer of a first multiplane image and a second tile set representing the alpha layer of the first multiplane image.

[0149] In some embodiments of the above-described apparatus, the first and second tile sets each have a different number of tiles.

[0150] In some embodiments of the above-described device, the frames in a sequence of video frames include a first tile set representing a first multiplane image and a second tile set representing a second multiplane image.

[0151] In some embodiments of any of the above devices, the first and second multiplane images are images of the scene from different camera positions.

[0152] In some embodiments of any of the above devices, the first tile set includes tiles representing the texture layer of the first multiplane image and other tiles representing the alpha layer of the first multiplane image, and the second tile set includes tiles representing the texture layer of the second multiplane image and other tiles representing the alpha layer of the second multiplane image.

[0153] In some embodiments of any of the above devices, the frames of the video frame sequence further include a third tile set representing a third multiplane image and a fourth tile set representing a fourth multiplane image.

[0154] In some embodiments of the above-described device, the metadata bitstream includes supplemental extended information messages. In some embodiments of the above-described device, the frames of a sequence of video frames have tiles representing reference images.

[0155] In some embodiments of any of the above devices, the metadata bitstream includes parameters selected from a group including the size of the reference view, the number of layers in a multiplane image, the number of simultaneous views, one or more characteristics of the packing arrangement, layer merge information, dynamic range adjustment information for the texture channel or alpha channel, and reference view information.

[0156] For example, according to yet another exemplary embodiment referring to one or any combination of parts or all of Figures 1-17 in the Abstract, a method is provided for decoding an incoming bitstream encoding a sequence of multiplane images, the method comprising: demultiplexing the incoming bitstream to obtain a video bitstream encoding a sequence of video frames; obtaining a metadata bitstream specifying at least the packing arrangement of tiles in the sequence of video frames, where the tiles represent layers of a multiplane image; reconstructing the sequence of video frames by applying video decompression to the video bitstream; and reconstructing the sequence of multiplane images based on the metadata bitstream using the tiles from the sequence of video frames.

[0157] In some embodiments of the above method, when executed by at least one processor, a non-temporary computer-readable medium is provided that stores instructions causing at least one processor to perform an operation including a corresponding one of the above methods for decoding an incoming bitstream.

[0158] For example, according to yet another exemplary embodiment referring to one or any combination of parts or all of Figures 1-17 in the Abstract, a method for decoding multiple received bitstreams, each received bitstream encoding a sequence of each multiplane image corresponding to each different camera position, the method comprising: demultiplexing a first received bitstream to obtain a first video bitstream encoding a sequence of first video frames; obtaining a first metadata bitstream specifying at least a first packing arrangement of tiles in the sequence of first video frames, where the tiles represent layers of the first multiplane image corresponding to the first camera position; reconstructing the sequence of first video frames by applying video decompression to the first video bitstream; and using the tiles from the sequence of first video frames, based on the first metadata bitstream, the sequence of the first multiplane image A method is provided that includes the steps of: reconstructing a bitstream; demultiplexing a second received bitstream to obtain a second video bitstream encoding a sequence of second video frames; and obtaining a second metadata bitstream specifying at least a second packing arrangement of tiles in the sequence of second video frames, wherein the tiles represent layers of a second multiplane image corresponding to a second camera position; reconstructing the second video frame sequence by applying video decompression to the second video bitstream; reconstructing a second multiplane image sequence from the second video frame sequence based on the second metadata bitstream; and generating a sequence of visible images by rendering a multiplane image set including at least one image from the sequence of first multiplane images and at least one image from the sequence of second multiplane images.

[0159] According to yet another exemplary embodiment described above, for example, referring to one or any combination of all or part of Figures 1-17 in the summary, a method for decoding a bitstream, the method comprising: receiving a coded bitstream comprising a sequence of multiplane images and metadata, wherein the metadata comprises profile parameters for decoding a coded bitstream according to a packed profile of MPEG immersive video (MIV); and decoding a coded bitstream according to the metadata, wherein the metadata comprises vps_attribute_video_present_flag[atlasID] set to 0, pin_attribute_present_flag[atlasID] set to 1, pin_attribute_count[atlasID] set to 2, pin_attribute_type_id[atlasID][attrIdx] set to ATTR_TEXTURE, ATTR_TRANSPARENCY, and vps_packed_video_present_flag[atlasID] set to 1, where vps_attribute_video_present_flag[j] is the atlas ID A method is provided that indicates whether an atlas having j has associated attribute video data, pin_attribute_present_flag[j] indicates whether a packed video frame of an atlas having atlas ID j contains a region with attribute data, pin_attribute_count[j] indicates the number of attributes with unique attribute types present in a packed video frame of an atlas having atlas ID j, pin_attribute_type_id[j][i] indicates the attribute type of the attribute having index i of an atlas having atlas ID j, and vps_packed_video_present_flag[j] indicates whether an atlas having atlas ID j has associated packed video data.

[0160] In the embodiment of the method described above, the MIV metadata further includes flags indicating whether the patch mode, patch width, and patch height apply to all patches in the atlas sequence.

[0161] For example, according to yet another exemplary embodiment described above, which references one or any combination of all or part of Figures 1-17 in the summary, a method for processing a volumetric bitstream is provided, the method comprising: receiving a coded bitstream comprising volumetric data and metadata, wherein the metadata comprises a packing information syntax for decoding data coded in accordance with the Visual Volumetric Video-Based Coding (V3C) specification, and the metadata comprises a first flag for checking whether the dimensions of the attributes to be packed are the same, and a second flag for checking whether time interleaving is enabled; and if the dimensions of the attributes to be packed are the same and time interleaving is enabled, decoding the volumetric data using time interleaving.

[0162] In any embodiment of the above-described device, the metadata bitstream includes a first syntax element (mpi_num_layers_minus1) used to determine the total number of MPI layers, a second syntax element (mpi_layer_depth_or_disparity_values_flag) signaling whether depth information is interpreted as depth values ​​or disparity values, a third syntax element (mpi_layer_depth_equal_distance_flag) signaling whether depth information values ​​have equal distances in depth or equal values ​​in disparity, and the decoded output picture in time in the output order. The syntax may include one or more of the following: a fourth syntax element (mpi_texture_opacity_interleave_flag) that signals whether it corresponds to an interleaved texture and opacity composition picture or a spatially packed texture and opacity composition picture; if the fourth syntax element indicates a spatially packed picture, the fifth syntax element (mpi_texture_opacity_arrangement_flag) indicates a top-bottom or side-by-side arrangement, and the sixth syntax element indicates the number of spatially packed layers at the height of picture 0 and picture 1.

[0163] In some embodiments of the above-described device, if the third syntax element signals that the depth information values ​​are equal in distance, the processor reads the seventh syntax element (mpi_depth_equal_distance_type_flag) which signals whether the depth values ​​are equal in distance or parallax, reads the depth information for the nearest depth (ZNear) and the farthest depth (ZFar) or nearest parallax (DNear) or farthest parallax (DFar), the depth information being applicable to all MPI layers, and otherwise reads the depth information for the nearest depth (ZNear) and the farthest depth (ZFar) or nearest parallax (DNear) or farthest parallax (DFar) for each MPI layer.

[0164] In some embodiments of the above-described device, If mpi_layer_depth_or_disparityvalues_flag is equal to 0 and mpi_depth_equal_distance_type_flag is equal to 0, Depth value Z[mpi_num_layers_minus1-i] =i*(ZFar-ZNear)÷(mpi_num_layers_minus1)+ZNear, Parallax value D[i] = 1 ÷ Z[i]; If mpi_layer_depth_or_disparityvalues_flag is equal to 0 and mpi_depth_equal_distance_type_flag is equal to 1, Depth value Z[i] =1÷(i*(1÷ZNear-1÷ZFar)÷(mpi_num_layers_minus1)+1÷ZFar), Parallax value D[i] = 1 ÷ Z[i]; If mpi_layer_depth_or_disparityvalues_flag is equal to 1 and mpi_depth_equal_distance_type_flag is equal to 0, Parallax value D[mpi_num_layers_minus1-i] =1÷(i*(1÷DFar-1÷DNear)÷(mpi_num_layers_minus1)+1÷DNear), Depth value Z[i] = 1 ÷ D[i]; If mpi_layer_depth_or_disparityvalues_flag is equal to 1 and mpi_depth_equal_distance_type_flag is equal to 1, Parallax value D[i] =i*(DNear-DFar)÷(mpi_num_layers_minus1)+DFar, Depth value Z[i] = 1 ÷ D[i] mpi_layer_depth_or_disparity_values_flag indicates the second syntax element, mpi_depth_equal_distance_type_flag indicates the fourth syntax element, and mpi_num_layers_minus1 indicates the first syntax element.

[0165] While processes, systems, methods, heuristics, etc., are described in this specification, it should be understood that such steps, etc., are described as occurring in a specific ordered sequence, but such processes may be performed in conjunction with described steps that are performed in an order different from that described in this specification. It should be further understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described in this specification may be omitted. In other words, the descriptions of processes in this specification are provided for the purpose of illustrating specific embodiments and should not be considered as limiting the claims.

[0166] Therefore, it should be understood that the above description is intended to be illustrative and not restrictive. Reading the above description will reveal many embodiments and applications beyond those provided. The scope should be determined without reference to the above description, but instead by reference to the attached claims, along with the entire equivalent scope of the claims granted. It is anticipated and intended that future developments will occur in the technology discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is subject to modification and alteration.

[0167] All terms used in the claims are intended to give their broadest, most reasonable form and ordinary meaning so that they may be understood by those familiar with the art described herein. In particular, the use of singular articles such as “a,” “the,” and “said” should be read to refer to one or more of the elements shown, unless the claim expressly states the opposite limitation.

[0168] This summary of the disclosure is provided to enable readers to quickly assess the characteristics of the technical disclosure. It is understood that it is not to be used to interpret or limit the scope or meaning of the claims. Furthermore, it is found that in the prior detailed description, various features are grouped together into various embodiments for the purpose of streamlining the disclosure. This method of the disclosure should not be interpreted as reflecting an intention that the claimed embodiments incorporate more features than those expressly described in each claim. Rather, as reflected in the following claims, the subject matter of the invention lies in fewer features than all of a single disclosed embodiment combined. Accordingly, the following claims are incorporated herein into the detailed description, and each claim stands alone as separately claimed subject matter.

[0169] This disclosure includes references to exemplary embodiments, but this specification is not intended to be constrained. Various modifications of the embodiments described, and other embodiments within the scope of the disclosure that are apparent to those skilled in the art to whom the disclosure relates, are deemed to be within the principles and scope of the disclosure, as set forth, for example, in the following claims.

[0170] Some embodiments may be implemented in the form of methods and apparatus for carrying out these methods. Some embodiments may also be implemented in the form of program code recorded on a tangible medium such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy diskette, a CD-ROM, a hard drive, or any other non-temporary machine-readable storage medium, where the program code is loaded into and executed by a machine such as a computer, which then becomes an apparatus for carrying out the patented invention. When implemented on a general-purpose processor, the program code segment is combined with the processor to provide a unique device that operates similarly to a particular logic circuit.

[0171] Unless otherwise explicitly stated, each number and range should be interpreted as approximate, as if the words "about" or "approximately" preceded the value or range.

[0172] The use of figure numbers and / or figure reference labels in a claim is intended to identify one or more possible embodiments of the subject matter to be claimed, in order to facilitate the interpretation of the claim. Such use is not necessarily construed as limiting the scope of these claims to the embodiments shown in the corresponding figures.

[0173] Where applicable, the elements in the claims of the following methods are described in a specific order by corresponding labels; however, unless the reference to the claim specifically indicates a particular order for the implementation of some or all of these elements, these elements are not necessarily intended to be limited to being implemented in that specific order.

[0174] In this specification, any reference to “one embodiment” or “embodiment” means that certain features, structures, or characteristics described in relation to an embodiment may be included in at least one embodiment of this disclosure. The occurrence of the phrase “in one embodiment” in various places in this specification does not necessarily all refer to the same embodiment, nor does a separate or alternative embodiment necessarily exclude other embodiments. The same applies to the term “implementation.”

[0175] Unless otherwise provided in this specification, the use of ordinal adjectives such as “first,” “second,” and “third” to refer to multiple similar objects simply indicates that different instances of such similar objects are being referred to, and does not imply that the similar objects referred to in this manner must exist in a corresponding order or sequence in time, space, rank, or otherwise.

[0176] Unless otherwise provided in this specification, the conjunction "if," in addition to its simple meaning, may be interpreted, depending on the corresponding specific context, as meaning "when," "upon," "in response to determining," or "in response to detecting." For example, the phrases "if it is determined" or "if [a stated condition] is detected" may be interpreted as meaning "upon determining," "in response to determining [the stated condition or event]," or "in response to determining [the stated condition or event]."

[0177] Furthermore, for the purposes of this specification, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any method known or subsequently developed in the art in which energy can be transferred between two or more elements, and the intervention of one or more additional elements is intended, though not essential. Conversely, terms such as “directly coupled” or “directly connected” mean that no such additional elements exist.

[0178] The functionality of the various elements shown in the diagram, including any functional blocks labeled or referenced as including “Processor” and / or “Controller,” may be provided through the use of dedicated hardware, as well as hardware capable of executing software in conjunction with appropriate software. Where provided by a processor, functionality may be provided by a single dedicated processor, a single shared processor, or multiple individual processors, some of which may be shared. Furthermore, the explicit use of the terms “Processor” or “Controller” should not be interpreted as referring only to hardware capable of executing software, but implicitly including, but not limited to, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM), random-access memory (RAM), and non-volatile storage for storing software. Other conventional and / or custom hardware may also be included. Similarly, all switches shown in the diagram are conceptual. Their functionality may be performed through the operation of programmable logic, dedicated logic, the interaction of programmable control and dedicated logic, or manually. Specific techniques are at the implementer's discretion, as they are better understood from the context.

[0179] As used in this application, the term “circuit” or “circuitry” can mean one or more of the following: (a) a hardware-only circuit implementation (such as an implementation of analog and / or digital circuits only); (b) a combination of hardware circuits and software (where appropriate): (i) a combination of analog and / or digital hardware circuits and software / firmware; (ii) any part of a hardware processor, software, and memory having software (including a digital signal processor) that works together to enable a device such as a mobile phone or server to perform various functions; and (c) a hardware circuit and / or processor, such as a microprocessor or part of a microprocessor, which requires software (such as firmware) to operate but may not exist if software is not required to operate. This definition of circuit applies to all use of the term in this application, including in the claims. As a further example, as used in this application, the term “circuit” also includes a mere hardware circuit or processor (or more processors), or a part of a hardware circuit or processor, and its (or their) accompanying software and / or firmware implementation. The term “circuit” also includes, for example, a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, where applicable to the components of a particular claim.

[0180] Those skilled in the art will understand that the block diagrams in this specification represent conceptual diagrams of exemplary circuits embodying the principles of disclosure. Similarly, flowcharts, flow charts, state transition diagrams, pseudocode, etc., whether or not a computer or processor is explicitly indicated, will be understood to represent various processes performed by a computer or processor, substantially represented in a computer-readable medium.

[0181] The “Summary of the Invention” in this specification is intended to present several exemplary embodiments, and additional embodiments are described with reference to the “Modes for Carrying Out the Invention” and / or one or more drawings. The “Summary of the Invention” is not intended to identify or limit the scope of the claimed subject matter.

[0182] Appendix This Appendix provides a partial copy of Table A-1 and relevant syntax information from the previously cited MIV (ISO / IEC23090-12:2021(E) / AMD.1:2022) and V3C (ISO / IEC23090-5:2023) specifications. MIV Table A-1: ​​Tolerances for Syntax Element Values ​​for MIV Toolset Profile Components [Table 25-1] [Table 25-2] Target syntax parameters from the V3C specification If vps_attribute_video_present_flag[j] is equal to 0, it indicates that the atlas with atlas ID j does not have any associated attribute video data. If vps_attribute_video_present_flag[j] is equal to 1, it indicates that the atlas with atlas ID j should have at least one associated attribute video data. If vps_attribute_video_present_flag[j] does not exist, it is presumed to be equal to 1. A requirement for bitstream compatibility is that if vps_attribute_video_present_flag[j] is equal to 1 for an atlas with atlas ID j, then pin_attribute_present_flag[j] is equal to 0 for all other atlases with the same atlas ID j. A value of 0 for `pin_attribute_present_flag[j]` indicates that the packed video frame of the atlas with atlas ID j does not contain any regions with attribute data. A value of 1 for `pin_attribute_present_flag[j]` indicates that the packed video frame of the atlas with atlas ID j contains regions with attribute data. If `pin_attribute_present_flag[j]` does not exist, its value is presumed to be equal to 0. The bitstream compatibility requirement is that for an atlas with atlas ID j, if vps_attribute_video_present_flag[j] is equal to 1, then for an atlas with the same atlas ID j, pin_attribute_present_flag[j] is equal to 0. pin_attribute_count[j] indicates the number of attributes with unique attribute types present in the packed video frames of the atlas with atlas ID j. pin_attribute_count[j] must be within the range of 1 to 127, including both ends. pin_attribute_type_id[j][i] indicates the attribute type of the attribute that has index i for the atlas with atlas ID j. Table 4 describes the list of supported attribute types. pin_attribute_dimension_minus1[j][i]+1 represents the total number of dimensions (i.e., the number of channels) of the region containing the attribute with index i for the atlas with atlas ID j. pin_attribute_dimension_minus1[j][i] is assumed to be within the range of 0 to 63, including both ends. pin_attribute_dimension_partitions_minus1[j][i]+1 indicates the number of partition groups to which the attribute channels of the region containing the attribute with index i should be grouped for an atlas with atlas ID j. pin_attribute_dimension_partitions_minus1[j][i] is assumed to be within the range of 0 to 63, including both ends. pin_attribute_msb_align_flag[j][i] indicates how, for an atlas with atlas ID j, the decoded region containing the attribute with attribute index i is converted to samples at the nominal attribute bit depth, as specified in Annex B. ai_attribute_count[j] indicates the number of attributes associated with the atlas having atlas ID j. ai_attribute_count[j] must be within the range of 0 to 127, including both ends. ai_attribute_type_id[j][i] indicates the attribute type of the Attribute Video Data unit having index i in the atlas having atlas ID j. Table 4 describes the list of supported attributes and their relationships with ai_attribute_type_id[j][i]. ai_attribute_dimension_minus1[j][i]+1 represents the total number of dimensions (i.e., the number of channels) of the attribute with index i for the atlas with atlas ID j. ai_attribute_dimension_minus1[j][i] is assumed to be within the range of 0 to 63, including both ends. ai_attribute_dimension_partitions_minus1[j][i]+1 indicates the number of partition groups to which the attribute channels of the attribute with index i should be grouped for the atlas with atlas ID j. ai_attribute_dimension_partitions_minus1[j][i] must be in the range of 0 to 63, including both ends. ai_attribute_msb_align_flag[j][i] indicates how, for an atlas having atlas ID j, a decoded attribute video sample having attribute index i is converted to a sample with nominal attribute bit depth, as specified in Annex B. If vps_packed_video_present_flag[j] is equal to 0, it indicates that the atlas with atlas ID j does not have any associated packed video data. If vps_packed_video_present_flag[j] is equal to 1, it indicates that the atlas with atlas ID j has some associated packed video data. If vps_packed_video_present_flag[j] does not exist, it is presumed to be equal to 0. For an atlas having atlas ID j, if vps_packed_video_present_flag[j] is equal to 1, then the bitstream conformance requirement is that at least one of pin_occupancy_present_flag[j], pin_geometry_present_flag[j], or pin_attribute_present_flag[j] is equal to 1.

Claims

1. A device for encoding a sequence of multiplane images, wherein the device is At least one processor, At least one memory location containing program code, Includes, The at least one memory and the program code, together with the at least one processor, are provided to the device at least, A sequence of video frames is generated, each of which video frames includes a plurality of tiles, each representing one or more layers of each multiplane image in the multiplane image, A metadata bitstream is generated to specify at least the packing arrangement of the tiles in the sequence of video frames. By applying video compression to the aforementioned sequence of video frames, a video bitstream is generated. The video bitstream and the metadata bitstream are multiplexed for transmission. A device that is configured in such a way.

2. The apparatus according to claim 1, wherein the first frame of the sequence of video frames has a tile corresponding to a first multiplane image, and the second frame of the sequence of video frames has a tile corresponding to a second multiplane image.

3. The apparatus according to claim 2, wherein the first and second multiplane images are images of scenes from different camera positions.

4. The apparatus according to claim 2, wherein the first and second multiplane images are images of a scene at different times.

5. The frames of the aforementioned sequence of video frames are The first tileset represents the texture layer of the first multiplane image, A second tile set representing the alpha layer of the first multiplane image, The apparatus according to claim 1, having the following features.

6. The apparatus according to claim 5, wherein the first tile set and the second tile set each have a different number of tiles.

7. The frames of the aforementioned sequence of video frames are A first tile set representing the first multiplane image, The apparatus according to claim 1, comprising a second tile set representing a second multiplane image.

8. The apparatus according to claim 7, wherein the first multiplane image and the second multiplane image are images of a scene from different camera positions.

9. The first tile set includes a tile representing the texture layer of the first multiplane image and another tile representing the alpha layer of the first multiplane image, The apparatus according to claim 7, wherein the second tile set includes a tile representing the texture layer of the second multiplane image and another tile representing the alpha layer of the second multiplane image.

10. The frames of the aforementioned sequence of video frames are A third tile set representing the third multiplane image, The apparatus according to claim 7, further comprising a fourth tile set representing a fourth multiplane image.

11. The metadata bitstream includes supplemental extended information messages, The apparatus according to claim 1, wherein the frames of the sequence of video frames have tiles representing reference images.

12. The metadata bitstream is Reference view size, The number of layers in the aforementioned multiplane image, Number of simultaneous views, One or more characteristics of the packing arrangement, Layer merge information, Dynamic range adjustment information for the texture channel or alpha channel, and Reference view information, The apparatus according to any one of claims 1 to 11, comprising a parameter selected from a group including the group.

13. A method for encoding a sequence of multiplane images, wherein the method is A step of generating a sequence of video frames, wherein each video frame includes a plurality of tiles representing one or more layers of the multiplane image, The steps include generating a metadata bitstream for specifying at least the packing arrangement of the tiles in the sequence of video frames, The steps include: generating a video bitstream by applying video compression to the sequence of video frames; The steps of multiplexing the video bitstream and the metadata bitstream for transmission, A method that includes this.

14. A non-temporary computer-readable medium that, when executed by at least one processor, stores instructions causing the at least one processor to perform the method according to claim 13.

15. A device for decoding a received bitstream encoded with a sequence of multiplane images, wherein the device comprises: At least one processor, At least one memory location containing program code, Includes, The at least one memory and the program code, together with the at least one processor, are provided to the device at least, The received bitstream is demultiplexed to obtain a video bitstream encoded with a sequence of video frames, and a metadata bitstream is obtained that specifies at least the packing arrangement of tiles within the sequence of video frames, wherein the tiles represent layers of a multiplane image. By applying video decompression to the video bitstream, the sequence of video frames is reconstructed. Using the tiles from the sequence of video frames, the sequence of multiplane images is reconstructed based on the metadata bitstream. A device configured in such a way.

16. The apparatus according to claim 15, wherein the at least one memory and the program code are configured together with the at least one processor to cause the apparatus to generate a sequence of visible images by rendering the sequence of multiplane images.

17. Rendering operations that aim to generate a composite visible image corresponding to a new view are: Applying warping to layers of a multiplane image set corresponding to different reference camera positions, and performing warping according to the new view, To generate a corresponding set of individual visible images corresponding to the new view, the layers of the multiplane image set are combined after the warping, The composite visible image is generated as a weighted sum of the individual visible images, The apparatus according to claim 16, including the apparatus described in claim 16.

18. The apparatus according to claim 17, wherein the multiplane image set includes two, three, or four multiplane images.

19. The first frame of the sequence of video frames has a tile corresponding to the first multiplane image, The apparatus according to claim 15, wherein the second frame of the sequence of video frames has a tile corresponding to a second multiplane image.

20. The apparatus according to claim 19, wherein the first multiplane image and the second multiplane image are images of a scene from different camera positions.

21. The apparatus according to claim 19, wherein the first multiplane image and the second multiplane image are images of scenes at different times.

22. The frames of the aforementioned sequence of video frames are The first tileset represents the texture layer of the first multiplane image, A second tile set representing the alpha layer of the first multiplane image, The apparatus according to claim 15, having the following features.

23. The apparatus according to claim 22, wherein the first tile set and the second tile set each have a different number of tiles.

24. The frames of the aforementioned sequence of video frames are A first tile set representing the first multiplane image, The apparatus according to claim 15, comprising a second tile set representing a second multiplane image.

25. The apparatus according to claim 24, wherein the first multiplane image and the second multiplane image are images of a scene from different camera positions.

26. The first tile set includes a tile representing the texture layer of the first multiplane image and another tile representing the alpha layer of the first multiplane image, The apparatus according to claim 24, wherein the second tile set includes a tile representing the texture layer of the second multiplane image and another tile representing the alpha layer of the second multiplane image.

27. The frames of the aforementioned sequence of video frames are A third tile set representing the third multiplane image, The apparatus according to claim 24, further comprising a fourth tile set representing a fourth multiplane image.

28. The metadata bitstream includes supplemental extended information messages, The apparatus according to claim 15, wherein the frames of the sequence of video frames have tiles representing reference images.

29. The metadata bitstream is Reference view size, The number of layers in the aforementioned multiplane image, Number of simultaneous views, One or more characteristics of the packing arrangement, Layer merge information, Dynamic range adjustment information for the texture channel or alpha channel, and Reference view information, The apparatus according to any one of claims 15 to 28, comprising a parameter selected from a group including the group.

30. A method for decoding a received bitstream encoded with a sequence of multiplane images, wherein the method is: The steps include: demultiplexing the received bitstream to obtain a video bitstream encoding a sequence of video frames; and obtaining a metadata bitstream specifying at least the packing arrangement of tiles in the sequence of video frames, wherein the tiles represent layers of a multiplane image; The steps include: reconstructing the sequence of video frames by applying video decompression to the video bitstream; The steps of: using the tiles from the sequence of video frames to reconstruct the sequence of multiplane images based on the metadata bitstream; A method that includes this.

31. A non-temporary computer-readable medium that, when executed by at least one processor, stores instructions causing the at least one processor to perform the method according to claim 30.

32. A method for decoding multiple received bitstreams, wherein each received bitstream encodes a sequence of multiplane images corresponding to different camera positions, and the method is: Steps include: demultiplexing a first received bitstream to obtain a first video bitstream encoded with a first sequence of video frames; and obtaining a first metadata bitstream specifying at least a first packing arrangement of tiles in the first sequence of video frames, wherein the tiles represent layers of a first multiplane image corresponding to a first camera position; The steps include: reconstructing the sequence of the first video frames by applying video decompression to the first video bitstream; The steps of: reconstructing a sequence of first multiplane images based on the first metadata bitstream using tiles from the first sequence of video frames; Steps include: demultiplexing a second received bitstream to obtain a second video bitstream encoding a second sequence of video frames; and obtaining a second metadata bitstream specifying at least a second packing arrangement of tiles in the second sequence of video frames, wherein the tiles represent layers of a second multiplane image corresponding to a second camera position; The steps include: reconstructing the sequence of the second video frames by applying video decompression to the second video bitstream; The steps of: reconstructing a second sequence of multiplane images based on the second metadata bitstream using tiles from the second sequence of video frames; A step of generating a sequence of visible images by rendering a set of multiplane images including at least one image from the first sequence of multiplane images and at least one image from the second sequence of multiplane images, A method that includes this.

33. A method for decoding a bitstream, wherein the method is A step of receiving a coded bitstream comprising a sequence of multiplane images and metadata, wherein the metadata includes profile parameters for decoding the coded bitstream according to a packed profile of MPEG immersive video (MIV). The steps include: decoding the coded bitstream according to the metadata; Includes, The aforementioned metadata is vps_attribute_video_present_flag[atlasID] set to 0, The pin_attribute_present_flag[atlasID] set to 1, The pin_attribute_count[atalasID] set in 2, The pin_attribute_type_id[atlasID][attrIdx] set in ATTR_TEXTURE and ATTR_TRANSPARENCY, The vps_packed_video_present_flag[atlasID] set to 1, Includes, A method in which vps_attribute_video_present_flag[j] indicates whether an atlas with atlas ID j has associated attribute video data, pin_attribute_present_flag[j] indicates whether a packed video frame of an atlas with atlas ID j contains a region with attribute data, pin_attribute_count[j] indicates the number of attributes with unique attribute types present in a packed video frame of an atlas with atlas ID j, pin_attribute_type_id[j][i] indicates the attribute type of the attribute having index i of an atlas with atlas ID j, and vps_packed_video_present_flag[j] indicates whether an atlas with atlas ID j has associated packed video data.

34. A method for processing a volumetric bitstream, wherein the method is A step of receiving a coded bitstream including volumetric data and metadata, wherein the metadata includes packing information syntax for decoding the coded data according to the Visual Volumetric Video-Based Coding (V3C) specification, and the metadata is A first flag to check whether the dimensions of the attributes to be packed are the same, A second flag to check whether time interleaving is enabled, Steps including, If the dimensions of the attributes to be packed are the same and time interleaving is enabled, the steps include decoding the volumetric data using time interleaving, A method that includes this.

35. The metadata bitstream is The first syntax element (mpi_num_layers_minus1) is used to determine the total number of MPI layers, A second syntax element (mpi_layer_depth_or_disparity_values_flag) signals whether depth information is interpreted as depth values ​​or disparity values, A third syntax element (mpi_layer_depth_equal_distance_flag) signals whether the depth information values ​​have equal distances in depth or equal values ​​in parallax, A fourth syntax element (mpi_texture_opacity_interleave_flag) signals whether the decoded output picture corresponds to a texture and opacity configuration picture that is temporally interleaved in the output order, or to a texture and opacity configuration picture that is spatially packed. Includes one or more of the following: The apparatus according to claim 1, wherein, when the fourth syntax element indicates spatially packed pictures, the fifth syntax element (mpi_texture_opacity_arrangement_flag) indicates a top-bottom or side-by-side arrangement, and the sixth syntax element indicates the number of spatially packed layers at the heights of picture 0 and picture 1.

36. When the third syntax element signals that the depth information values ​​are at equal distances, the processor: The seventh syntax element (mpi_depth_equal_distance_type_flag) is read, which signals whether the depth values ​​have equal distances in depth or parallax. Depth information is read for the nearest depth (ZNear) and the farthest depth (ZFar) or the nearest parallax (DNear) or the farthest parallax (DFar), and this depth information is applicable to all of the MPI layers. In other cases, for each of the MPI layers, The apparatus according to claim 35, which reads depth information for the nearest depth (ZNear) and the farthest depth (ZFar) or the nearest parallax (DNear) or the farthest parallax (DFar).

37. If mpi_layer_depth_or_disparityvalues_flag is equal to 0 and mpi_depth_equal_distance_type_flag is equal to 0, Depth values Z[mpi_num_layers_minus1−i]=i*(ZFar−ZNear)÷(mpi_num_layers_minus1)+ZNear, The parallax value is D[i] = 1 ÷ Z[i]. If mpi_layer_depth_or_disparityvalues_flag is equal to 0 and mpi_depth_equal_distance_type_flag is equal to 1, The aforementioned depth value is Z[i] = 1 ÷ (i * (1 ÷ ZNear - 1 ÷ ZFar) ÷ (mpi_num_layers_minus1) + 1 ÷ ZFar), The parallax value is D[i] = 1 ÷ Z[i]. If mpi_layer_depth_or_disparity values_flag is equal to 1 and mpi_depth_equal_distance_type_flag is equal to 0, The aforementioned disparity value is D[mpi_num_layers_minus1-i] = 1÷(i*(1÷DFar-1÷DNear)÷(mpi_num_layers_minus1)+1÷DNear), The aforementioned depth value is Z[i]=1÷D [i] is If mpi_layer_depth_or_disparity values_flag is equal to 1 and mpi_depth_equal_distance_type_flag is equal to 1, The aforementioned disparity value is D[i] = i * (DNear - DFar) ÷ (mpi_num_layers_minus1) + DFar, The aforementioned depth value is Z[i]=1÷D [i] is The apparatus according to claim 36, wherein mpi_layer_depth_or_disparity values_flag represents the second syntax element, mpi_depth_equal_distance_type_flag represents the fourth syntax element, and mpi_num_layers_minus1 represents the first syntax element.

38. The method according to claim 33, wherein the metadata further includes flags to indicate whether the patch mode, patch width, and patch height apply to all patches in the atlas sequence.

39. The metadata in the atlas header is The current number of patches in Atlas Style, Number of patches in height, The quantized width value of the aforementioned patch, The quantized height value of the aforementioned patch, and Patch mode for patching, The method according to claim 33, further comprising one or more of the parameters.