Dynamic structure for volumetric data coding

By generating and packaging a two-dimensional representation patch of the volume data set, combining dynamic sub-pictures and two-dimensional video encoder, the problem of low encoding efficiency of dynamic volume data in the prior art is solved, and a more efficient encoding effect is achieved.

CN119948864APending Publication Date: 2025-05-06INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066678.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-19
Filing Date
2023-09-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When existing 2D video encoders encode dynamic volume data, their efficiency is limited by the constraints of the data structure, resulting in low efficiency in writing atlas code.

Method used

By generating a patch of corresponding two-dimensional representation of the volume data set, it is packaged into one or more dynamic sub-pictures associated with the video frame, and these sub-pictures are encoded by a two-dimensional video encoder.

Benefits of technology

The encoding efficiency of dynamic volume data is improved, and the volume data set can be processed and compressed more efficiently through the combination of dynamic sub-pictures and two-dimensional video encoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948864A_ABST
    Figure CN119948864A_ABST
Patent Text Reader

Abstract

Technologies are disclosed, including apparatuses and methods for encoding dynamic volume data, including receiving a sequence of volume data sets representative of dynamic volume data and encoding the sequence. Encoding a volumetric data set in the sequence comprises: generating patches comprising respective two-dimensional representations of the volumetric data set; packing the patch into one or more dynamic sub-pictures associated with the video frame; and then encoding, by a two-dimensional video encoder, the one or more dynamic sub-pictures into a bitstream of a coded volumetric data set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of European application number EP22306373.6 filed on September 19, 2022, which is incorporated herein by reference in its entirety. Background Art

[0003] The coding of dynamic volumetric data (such as immersive content or point cloud data) involves the coding of intermediate two-dimensional (2D) video data (so-called atlases). These atlases contain patches of geometric data and texture data required for the three-dimensional (3D) reconstruction of the volumetric data. The latest standards for the coding of volumetric data (such as the MPEG Immersive Video (MIV) standard and the Video-based Point Cloud Compression (V-PCC) standard developed by ISO / IEC MPEG) rely on encoding the atlas using a traditional 2D video encoder. To this end, data structures used by 2D video encoders (such as subpictures (recently introduced by the VVC standard)) can be used to contain atlases that describe the volumetric data to be encoded. However, the constraints on the data structure defined in the standards followed by current 2D video encoders limit the efficiency with which atlases can be coded. Summary of the invention

[0004] Aspects disclosed in the present disclosure describe a method for encoding dynamic volume data. The method includes receiving a sequence of volume data sets representing the dynamic volume data and encoding the sequence. For the volume data sets in the sequence, the encoding includes: generating patches including corresponding two-dimensional representations of the volume data sets; packing the patches into one or more dynamic sub-pictures associated with a video frame; and then encoding the one or more dynamic sub-pictures into a bitstream of code-written video data by a two-dimensional video encoder. A method for decoding the dynamic volume data is also disclosed. The method includes receiving a bitstream of code-written video data; coding a sequence of volume data sets representing the dynamic volume data; and decoding the sequence. For the volume data sets in the sequence, the decoding includes: decoding one or more dynamic sub-pictures associated with a video frame from the bitstream by a two-dimensional video decoder; extracting patches from the decoded one or more dynamic sub-pictures, the patches including corresponding two-dimensional representations of the volume data sets; and reconstructing the volume data sets based on the patches.

[0005] Aspects disclosed in the present disclosure describe a device for encoding dynamic volume data. The device includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the device to receive a sequence of volume data sets representing the dynamic volume data and encode the sequence. For the volume data sets in the sequence, the encoding includes: generating patches including corresponding two-dimensional representations of the volume data sets; packing the patches into one or more dynamic sub-pictures associated with a video frame; and encoding the one or more dynamic sub-pictures into a bitstream of code-written video data by a two-dimensional video encoder. A device for decoding dynamic volume data is also disclosed. The device includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the device to receive a bitstream of code-written video data; code the sequence of volume data sets representing the dynamic volume data; and decode the sequence. For a volume data set in the sequence, the decoding includes: decoding one or more dynamic sub-pictures associated with a video frame from the bit stream by a two-dimensional video decoder; extracting patches from the decoded one or more dynamic sub-pictures, wherein the patches include corresponding two-dimensional representations of the volume data set; and reconstructing the volume data set based on the patches.

[0006] Another aspect disclosed in the present disclosure describes a non-transitory computer-readable medium including instructions executable by at least one processor to perform a method for encoding dynamic volume data. The method includes receiving a sequence of volume data sets representing the dynamic volume data and encoding the sequence. For the volume data sets in the sequence, the encoding includes: generating patches including corresponding two-dimensional representations of the volume data sets; packing the patches into one or more dynamic sub-pictures associated with a video frame; and then encoding the one or more dynamic sub-pictures into a bitstream of code-written video data by a two-dimensional video encoder. A non-transitory computer-readable medium is also disclosed, including instructions executable by at least one processor to perform a method for decoding dynamic volume data. The method includes receiving a bitstream of code-written video data; coding a sequence of volume data sets representing the dynamic volume data; and decoding the sequence. For the volume data sets in the sequence, the decoding includes: decoding one or more dynamic sub-pictures associated with a video frame from the bitstream by a two-dimensional video decoder; extracting patches from the decoded one or more dynamic sub-pictures, the patches including corresponding two-dimensional representations of the volume data sets; and reconstructing the volume data sets based on the patches.

[0007] This summary is provided to introduce in a simplified form selected concepts that will be further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that address any or all of the disadvantages mentioned in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a block diagram of an example system according to which aspects of the present embodiments can be implemented.

[0009] Figure 2 is a block diagram of an example video encoder according to which aspects of the present embodiments can be implemented.

[0010] Figure 3 is a block diagram of an example video decoder according to which aspects of the present embodiments can be implemented.

[0011] Figure 4 is a block diagram of an example system for encoding dynamic volumetric data according to which aspects of the present embodiments can be implemented.

[0012] Figure 5 is a block diagram of an example system for decoding dynamic volume data according to which aspects of the present embodiments can be implemented.

[0013] Fig. 6A -C illustrates a texture atlas 6A, a corresponding occupancy atlas 6B, and a fill texture atlas 6C according to which aspects of the present embodiment can be implemented.

[0014] Figure 7 An example segmentation of a picture is illustrated according to which aspects of the present embodiment can be implemented.

[0015] Figure 8 An atlas including texture and geometry data is illustrated according to which aspects of the present embodiment can be implemented.

[0016] Fig. 9 Point cloud representations are illustrated according to which aspects of the present embodiments can be implemented, including occupancy atlases, geometry atlases, and texture atlases.

[0017] Fig.10 The packing of geometry atlases, texture atlases, and occupancy atlases into respective sub-pictures of a video picture is illustrated according to which aspects of the present embodiment can be implemented.

[0018] Fig.11 Dynamic sub-picture sizing is illustrated according to which aspects of the present embodiments can be implemented.

[0019] Fig.12 A one-dimensional layout of sub-pictures according to which aspects of the present embodiment can be implemented is illustrated.

[0020] Fig.13 The diagram illustrates sub-picture dependencies according to which aspects of the present embodiment can be implemented.

[0021] Fig.14 A reference picture list is illustrated according to which aspects of the present embodiment can be implemented.

[0022] Fig.15 A reference sub-picture list is illustrated according to which aspects of the present embodiment can be implemented.

[0023] Fig.16 Illustrated is a reference to a block in a reference sub-picture according to which aspects of the present embodiment can be implemented.

[0024] Fig.17 Relative references to blocks in a reference sub-picture are illustrated according to which aspects of the present embodiment can be implemented.

[0025] Fig.18 Inter- and intra-frame referencing through sub-pictures is illustrated according to which aspects of the present embodiment can be implemented.

[0026] Fig.19 is a flow chart of an example method for encoding dynamic volumetric data according to which aspects of the present embodiments can be implemented.

[0027] Fig. 20 is a flow chart of an example method for decoding dynamic volume data according to which aspects of the present embodiments can be implemented. DETAILED DESCRIPTION

[0028] Next reference Figure 1-3 A conventional system and method for video coding of 2D video frames (video streams) is described.

[0029] Figure 1A block diagram of an example system 100 is illustrated. System 100 may be embodied as a device including various components described below, and may be configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100 may be embodied in a single integrated circuit, multiple integrated circuits, and / or discrete components singly or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed over multiple integrated circuits and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.

[0030] The system 100 includes at least one processor 110, which can be configured to execute instructions loaded therein to implement, for example, various aspects described in the present application. The processor 110 may include embedded memory, input and output interfaces, and various other circuits as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include a non-volatile memory and / or a volatile memory, including, for example, an EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, a magnetic disk drive, and / or an optical disk drive. For example, the storage device 140 may be an internal storage device, an attached storage device, and / or a network accessible storage device.

[0031] The system 100 includes an encoder / decoder module 130, which is configured to process data to provide encoded video data or decoded video data. The encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module (or modules) that may be included in a device to perform encoding and / or decoding functions. In addition, as known to those skilled in the art, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated into the processor 110 as a combination of hardware and software.

[0032] Program code to be loaded into the processor 110 or the encoder / decoder 130 to perform various aspects described in the present application may be stored in the storage device 140 and subsequently loaded into the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in the present application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0033] In several embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing functions required during encoding or decoding. However, in other embodiments, memory external to the processing device (where, for example, the processing device may be the processor 110 or the encoder / decoder module 130) may be used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, which may include, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations.

[0034] Inputs to the elements of system 100 may be provided through various input devices, as indicated in block 105. Such input devices include, but are not limited to: (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcast device, (ii) a composite input terminal (COMP), (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0035] In various embodiments, the input devices of block 105 have associated respective input processing elements, as known in the art. For example, the RF portion may be associated with elements adapted to: (i) select a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) again band-limit to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in some embodiments, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired data packet stream. The RF portion of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs some of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter, down-convert and filter to the desired frequency band again to perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of similar or different functions. The element added can include inserting element between existing elements, for example inserting amplifier and analog to digital converter. In various embodiments, the RF part includes antenna.

[0036] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed Solomon error correction) may be implemented, for example, within a separate input processing integrated circuit or within the processor 1010, as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface integrated circuit or within the processor 1010, as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130, to operate in conjunction with memory and storage elements to process the data streams as desired for presentation on an output device.

[0037] The various components of system 100 may be disposed within an integrated housing. Within the integrated housing, the various components may be interconnected and transmit data between them using a suitable connection arrangement 1140 (eg, an internal bus as known in the art, including an I2C bus, wiring, and printed circuit boards).

[0038] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card. The communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0039] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data through an HDMI connection of the input block 105 to provide streaming data to the system 100. Still other embodiments use an RF connection of the input block 105 to provide streaming data to the system 100.

[0040] The system 100 may provide output signals to various output devices, including a display device 165, an audio device (e.g., speaker(s) 175), and other peripheral devices 185. In various examples of embodiments, the other peripheral devices 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 100. In various embodiments, the system 100 may be controlled using signaling such as AV Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. The display interface 1070 may communicate control signals between the system 100 and the display device 165, the audio device 175, or other peripheral devices 185. The output devices may be communicatively coupled to the system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to the system 100 using a communication channel 190 via a communication interface 150. The display device 165 and the audio device 175 may be integrated into a single unit with other components of the system 100 in an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.

[0041] For example, if the RF portion of input 105 is part of a stand-alone set-top box, the display device 165 and the audio device 175 may alternatively be separate from one or more of the other components. In various embodiments where the display device 165 and the audio device 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0042] Figure 2 FIG. 2 is a block diagram of an example video encoder 200. Figure 1 The described system 100 may employ the video encoder 200. For example, the video encoder 200 may be an encoder operating according to a coding standard such as Advanced Video Coding (AVC, H.264 / MPEG-4 | ISO / IEC 14496-10), High Efficiency Video Coding (HEVC, ITU-T H.265 | ISO / IEC 23008-2), or Versatile Video Coding (VVC, ITU-T H.266 | ISO / IEC 23090-3).

[0043] like Figure 2 As shown, before encoding, the video data can be pre-processed by a pre-encoding processor 201. Such pre-processing can include applying a color model transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr4:2:0) or performing a mapping of the components of the input picture to obtain a signal distribution that is more resilient to compression (e.g., applying a histogram equalizer and / or a denoising filter to one or more of the components of the picture). Pre-processing can also include associating metadata with the video data that can be attached to the video bitstream written by the code.

[0044] In the encoder 200, pictures of a video frame are encoded by encoder elements, as generally described below. The pictures to be encoded are segmented into coding units (CUs) by an image segmentor 202. Typically, a CU contains a luminance block and a corresponding chrominance block, and therefore the operations described herein as applied to a CU are applied relative to the luminance block and the corresponding chrominance block of the CU. After segmentation 202, each CU can be encoded using an intra or inter prediction mode. In the intra prediction mode, intra prediction is performed by an intra predictor 260. In the intra prediction mode, the content of a CU in a frame is predicted based on the content of one or more other CUs from the same frame, using reconstructed versions of the other CUs. In the inter prediction mode, motion estimation and motion compensation are performed by a motion estimator 275 and a motion compensator 270, respectively. In the inter prediction mode, the content of a CU in a frame is predicted based on the content of one or more other CUs from adjacent frames, using reconstructed versions of other CUs that can be extracted from a reference picture buffer 280. The encoder decides 205 which prediction result (the prediction result obtained by operating in intra prediction mode 260 or the prediction result obtained by operating in inter prediction mode 270, 275) to use to encode the CU, and indicates the selected prediction mode by, for example, a prediction mode flag. After the prediction operation, residual data is calculated for each CU, for example, by subtracting 210 the predicted CU from the original CU.

[0045] The respective residual data of the CU are then transformed and quantized by the transformer 225 and the quantizer 230, respectively. The entropy encoder 245 then encodes the quantized transform coefficients as well as the motion vectors and other syntax elements, thereby outputting a bitstream of coded video data. The encoder 200 can skip the transform operation 225 and directly quantize 230 the untransformed residual data. The encoder 200 can bypass both the transform operation and the quantization operation, that is, the residual data can be coded directly by the entropy encoder 245.

[0046] The encoder 200 reconstructs the encoded CU to provide a reference for future predictions. Therefore, the quantized transform coefficients (output of the quantizer 230) are dequantized by the inverse quantizer 240 and then inversely transformed by the inverse transformer 250 to decode the residual data of the corresponding CU. The decoded residual data is combined 255 with the corresponding predicted CU to produce the corresponding reconstructed CU. The in-loop filter 265 can then be applied to the reconstructed picture (formed by the reconstructed CU) to perform, for example, deblocking filtering and / or sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered reconstructed picture can then be stored in a reference picture buffer (280).

[0047] Figure 3 A block diagram of an example video decoder 300 is shown. Figure 1 The described system 100 may employ a video decoder 300. Generally speaking, aspects of the operation of the video decoder 300 are the inverse of aspects of the operation of the video encoder 200. Figure 2 As depicted, the encoder 200 also performs decoding operations 240, 250 by which the encoded pictures are reconstructed. The reconstructed pictures may then be stored in a reference picture buffer 280 and used to facilitate motion estimation 275 and compensation 270, as explained above.

[0048] In the decoder 300, the bitstream of video data written in the code generated by the video encoder 200 is first entropy decoded by the entropy decoder 330, thereby decoding the quantized transform coefficients, motion vectors and other control data encoded in the bitstream (such as data indicating how the picture is partitioned and the prediction mode selected by the CU) from the bitstream. The quantized transform coefficients are dequantized by the inverse quantizer 340 and then inversely transformed by the inverse transformer 350 to decode the residual data corresponding to the CU. The decoded residual data is combined 355 with the corresponding predicted CU to produce the corresponding reconstructed CU. Depending on the prediction mode selected by the CU, the predicted CU can be obtained 370 from the intra-frame predictor 360 or from the motion compensator 375. The in-loop filter 365 can be applied to the reconstructed picture (formed by the reconstructed CU). The filtered reconstructed picture can then be stored in the reference picture buffer 380 for motion compensation 375.

[0049] The decoded pictures may be further processed by the post-decoding processor 385. For example, the post-decoding processing may include an inverse color model transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse mapping to reverse the mapping process performed in the pre-encoding processor 201. The post-decoding processor 385 may use metadata derived by the pre-encoding processor 201 and / or signaled in the decoded video bitstream.

[0050] The introduction of immersive content with six degrees of freedom (where the viewer has both translational and rotational freedom of movement, and where motion parallax is supported) increases the amount of volumetric data required to describe dynamic scenes. Similarly, describing animated objects or 3D scenes through point clouds (i.e., dynamic 3D surface points and corresponding attributes) requires a large amount of volumetric data. In order to efficiently code dynamic volumetric data, the Moving Picture Experts Group (MPEG) (ISO / IEC 23090-5, 2021) has developed the Visual Volumetric Video Coding (V3C) standard. The V3C standard provides a platform for coding different types of volumetric data, such as immersive content and point clouds specified by the MPEG Immersive Video (MIV, ISO / IEC 23090-12, 2021) and Video-based Point Cloud Compression (V-PCC, ISO / IEC23090-5, 2021) standards, respectively.

[0051] The MIV standard addresses the compression of data describing a (real / virtual) scene captured by multiple (real / virtual) cameras. For further details, see JM Boyce et al., MPEG Immersive Video Coding Standard, Proceedings of the IEEE, Vol. 109, No. 9, pp. 1521-1536, September 2021, DOI: 10.1109 / JPROC.2021.3062590. The V-PCC standard addresses the compression of dynamic point cloud sequences, which can represent computer-generated objects for virtual / augmented reality applications, or can represent surroundings to enable autonomous driving. For further details, see Graziosi, D. et al., Ongoing Point Cloud Compression Standardization Activities: An Overview of Video-Based (V-PCC) and Geometry-Based (G-PCC), APSIPA Transactions on Signal and Information Processing, Vol. 9, 2020, No. 13, DOI: 10.1017 / ATSIP.2020.12. References in this article Figure 4 and Figure 5 Describes the encoding and decoding of dynamic volumetric data (usually according to the above standards).

[0052] Figure 44 is a block diagram of an example system 400 for encoding dynamic volume data. The system 400 includes a camera system 410 (including one or more cameras), a preprocessor 420, and an encoder 430. One or more cameras of the camera system 410 can be real cameras and / or virtual cameras, which are configured to capture real-world scenes and / or virtual scenes from multiple views, thereby generating corresponding video streams. The image content captured in each video stream is a projection of the scene on a projection plane of the corresponding camera. The camera system 410 may include other sensors configured to, for example, measure depth data of the corresponding video stream. Data associated with multiple views 415 (including the captured video streams (and optionally, the corresponding depth data) together with the parameters of the camera (e.g., intrinsic and extrinsic camera parameters)) are fed to the preprocessor 420. The preprocessor 420 processes the data associated with the multiple views 415 and generates a corresponding depth map therefrom. For example, a depth map associated with a video frame captured by a camera contains distance values ​​between a point at the scene and its projection at the projection plane of the camera. The preprocessor 420 provides the video view 422 (captured video stream) and the corresponding depth map 424 and camera parameters 426 to the encoder 430 to generate a bitstream 475 of coded video data therefrom.

[0053] The encoder 430 includes a dynamic volume data (DVD) encoder 440, multiple sets of traditional video encoders 450, 460, and a multiplexer 470. The DVD encoder 440 generates a property atlas 442 and a geometry atlas 444 based on the provided video view 422, the depth map 424, and the camera parameters 426. The geometry atlas contains geometry patches, each of which describes spatial data associated with the content captured in the video view 422 and its corresponding depth map 424, for example, the geometry patches in the geometry atlas 444 can represent occupancy data and depth data. The property atlas 442 contains attribute patches, each of which describes the properties of the content samples captured in the video view 422, for example, the attribute patches in the attribute atlas 442 can represent the texture, transparency, surface normal, or reflection data of the content samples. The DVD encoder 440 also encodes metadata (using a code writer defined by the MIV standard) to generate a metadata bitstream 446. The metadata describes the atlases 442, 444, and the camera parameters 426.

[0054] By design, the corresponding geometry patches and attribute patches generated by the DVD encoder 440, together with the camera parameters 426, can be used for 3D reconstruction of the scene captured by the camera system 410. The main purpose of the DVD encoder 440 is to generate 2D video data, including a stream of one or more attribute atlases 442 and a corresponding stream of one or more geometry atlases 444. These video streams 442, 444 can then be encoded by the video encoders 450, 460, see Figure 2 The output of the video encoders 452 , 462 (ie, the bitstream of coded video data) is then combined with the metadata bitstream 446 by a multiplexer 470 to form an output bitstream 475 .

[0055] Figure 5 5 is a block diagram of an example system 500 for decoding dynamic volume data. The system 500 includes a decoder 510 and a renderer 550. The decoder 510 is generally associated with Figure 4 The operations of the encoder 430 are opposite to each other. The decoder 510 includes a demultiplexer 520, a plurality of groups of conventional video decoders 530, 540 and a DVD decoder 550. The demultiplexer extracts the received bit stream 505 (i.e., Figure 4 475) extracts the bit stream of the video data 522, 524 written in the code (ie, Figure 4 452, 462) and metadata bitstream 526 (ie, Figure 4 The video data 522, 524 written in the code are then decoded by video decoders 530, 540, respectively, referring to Figure 3 Its operation is generally described. The video decoders 530, 540 recover the attribute atlas 532 and the geometry atlas 542, which are fed into the DVD decoder 550 together with the metadata bitstream 526. The DVD decoder 550 generates a video view 552 and its corresponding depth map 554 and camera parameters 556 therefrom. Using these data 552, 554, 556, the renderer 550 can create immersive content. To this end, for example, based on the pose 554 (viewing position and orientation) of the viewer 560, the renderer can provide the viewer with a viewport 552 of the scene, i.e., a field of view of the scene that can be viewed from the viewer's viewing perspective.

[0056] As described above, the 2D video atlas (i.e., the attribute atlas 442 and the geometry atlas 444) is compressed using a legacy 2D video codec, which may be implemented according to standards such as AVC, HEVC, VVC, or other video formats such as AV1. The 2D video atlas consists of attribute (e.g., texture) patches and geometry patches that are assembled (or packed) into a 2D picture. Fig. 6A-C illustrates an atlas containing texture patches and occupancy patches derived from a point cloud.

[0057] Fig. 6A -C illustrates a texture atlas 6A, a corresponding occupancy atlas 6B, and a fill texture atlas 6C. The texture atlas 6A and the corresponding occupancy atlas 6B are derived from the 3D PCC content "Long Dress." Fig. 6A and Figure 6B , each texture patch (e.g., 610) and its corresponding occupancy patch (e.g., 620) correspond to a view of the point cloud, which can be generated by projecting the point cloud onto an image plane (i.e., the projection image of the virtual camera). In general, to generate an atlas from content captured from multiple views, the DVD encoder 440 can perform a pruning operation and a packing operation. In the pruning operation, inter-view redundancy is removed, where objects captured in more than one view are retained in one view and removed ("pruned") from the other views. In the packing operation, the pruned views are segmented and packed to form a compact representation, such as Fig. 6A and Figure 6B The spatial transformations (e.g., translations and / or rotations) used by the DVD encoder 440 in packing each patch are stored in the metadata bitstream 446 to allow the DVD decoder 550 to extract (unpack) the packed patches. The atlas can be further processed to improve the compression efficiency of the video encoders 450, 460. To this end, for example, Figure 6C As shown, the areas between patches are typically padded to attenuate strong transitions between patches and avoid high frequencies which are typically more costly to code.

[0058] In general, improving 2D compression 450, 460 of atlases 442, 444 can be accomplished by facilitating intra-coding of patches within an atlas (e.g., by increasing spatial correlation) and by facilitating inter-coding of patches across consecutive atlases (e.g., by increasing temporal correlation). The present disclosure proposes additional means to improve coding efficiency of patches within an atlas when using a 2D video encoder, as discussed further below.

[0059] Various coding structures are used in the coding of 2D video. For example, the coding structure used when coding according to the VVC standard (or standards such as AVC or HEVC, and AV1 format) includes: a coded video sequence (CVS), a layer, a picture, a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU) and a transform unit (TU). Therefore, according to the color model used, a video sequence consists of pictures associated with a corresponding video frame, which may include several channels, such as a luminance channel (e.g., Y) and a chrominance channel (e.g., U and V). CVS may include one or more layers, each of which provides a different representation of the video content. For example, a picture of a video frame at time t (or a video frame with a value of p for a picture order count (POC)) may be encoded into several layers, such as a base layer containing a low-resolution representation of the picture and one or more enhancement layers, each of which adds additional details to the low-resolution representation of the picture. Therefore, each layer can provide a representation of the video content at different resolutions, quality levels, or viewing angles. The layer may also provide a supplementary representation such as a depth map or a transparency map. A coded layer video sequence (CLVS) is a layer-by-layer CVS containing coded data for a sequence of pictures across one layer.

[0060] In block-based video code writing (e.g., see Figure 2In the video coding described in the 1990s and 2000s, a picture is partitioned into basic processing units, each of which includes a luminance block and (optionally) a corresponding chrominance block. The size of the basic processing unit can be up to 128×128 (in the VCC standard), up to 64×64 (in the HEVC standard) or 16×16 (in AVC and previous standards, where such units are called macroblocks). For example, a basic processing unit of size 64×64 typically includes a 64×64 luminance block and corresponding two 32×32 chrominance blocks. Further partitioning of the basic processing units of a picture is represented by a corresponding tree syntax, namely the CTU structure introduced in the HEVC and VCC standards. The leaves of the CTU tree can be associated with coding units (CUs) of various sizes (e.g., 8×8, 16×16 or 32×32 blocks) that partition the corresponding basic processing units. A CU (including a luminance block and a corresponding chrominance block) is a coding entity for which the encoder selects a prediction mode (intra-frame prediction mode or inter-frame prediction mode (i.e., motion compensated prediction mode)). If the CU is encoded in inter prediction mode, a subsequent splitting of the CU can be performed to form prediction units (PUs), where the pixels of the luminance blocks and chrominance blocks of the PU share the same set of motion parameters. The CU of a CTU can be further split into transform units (TUs), where the same transform is applied to code the residual data (i.e., the difference between the prediction block and the corresponding original block to be encoded) of the luminance blocks and chrominance blocks of the TU.

[0061] Furthermore, pictures can be segmented into tiles, slices, and sub-pictures (where sub-pictures are introduced in the VVC standard), each of which contains many complete CTUs. Figure 7 An example segmentation of a picture 700 is illustrated. As shown, the picture 700 is segmented into 18 tiles, 24 slices, and 24 sub-pictures. Typically, a tile covers a rectangular area of ​​the picture, confined within the horizontal and vertical boundaries that split the picture into columns and rows of tiles, each tile containing a CTU. The slicing of a picture can be determined according to the following two modes: rectangular slicing and raster scan slicing. Rectangular slices cover rectangular areas of the picture, typically containing one or more complete tiles or one or more complete rows of CTUs within a tile. Raster scan slices can contain one or more complete tiles in tile raster scan order, and therefore are not necessarily rectangular in shape. A sub-picture contains one or more rectangular slices covering a rectangular area of ​​the picture. Note that in Figure 7In the example of , each sub-picture contains a slice. Sub-pictures can be extractable (i.e., independently encodable and decodable) or non-extractable. In each of these cases, the encoder can set cross-border filtering to on or off for each sub-picture. For example, in the VVC standard, even when the sub-picture is extractable, the motion vectors of the coding blocks in the sub-picture can point outside the sub-picture. Importantly, for sub-pictures, either or both of the following conditions need to be met: 1) all CTUs of the sub-picture must be part of the same tile, or 2) all CTUs in the tile must be part of the same sub-picture. The layout of the sub-picture is typically signaled in the sequence parameter set (SPS), and therefore, the same layout is signaled within the CLVS, as described in the example of Table 1 according to the VVC standard.

[0062] Table 1: SPS signaling of sub-picture layout.

[0063]

[0064]

[0065] For example, sub-pictures are useful for extracting viewports from omnidirectional or immersive video, for scalable rendering of content from a single bitstream, for parallel encoding, and for reducing the number of streams and corresponding decoder instances (thereby simplifying the synchronization task when decoding and rendering data from different streams).

[0066] In Santamaria et al., Coding Volumetric Content with MIV Using VVC Sub-Pictures, MMSP2021 ("Santamaria"), coding atlases using sub-pictures is described. Santamaria provides a method for coding atlases using sub-pictures as described in the present document. Figure 8 An example of description. Figure 8 An atlas 800 is illustrated, which includes texture data and geometry data. The atlas 800 is created such that one sub-picture 810 of the atlas represents texture data, i.e., an equirectangular projection (ERP) of the video content; and the remaining sub-pictures 820-840 of the atlas represent geometry data, i.e., a depth map 820 and patches 830, 840, which can be used to achieve six degrees of freedom viewing. In another example, the geometry data and the texture data can be separately packed into a geometry atlas and a texture atlas, respectively, and can be encoded independently.

[0067] Similarly, multiple 2D video streams can be used to encode (one or more) occupancy atlases, (one or more) geometry atlases, and (one or more) texture atlases representing point clouds. To allow denser point clouds, multiple 3D point layers per frame can be created, resulting in a large number of maps. In the V-PCC standard, a point cloud can be represented by an occupancy map, a geometry map, and a texture map. Fig. 9 The figure shows a point cloud representation, which includes an occupancy atlas, a geometry atlas, and a texture atlas. Fig. 9 In the example of , the point cloud is described by (a) a frame containing an occupancy atlas, (b) two frames containing the corresponding geometry (depth) atlas, and (c) two frames containing the corresponding texture atlas. (d) The reconstructed point cloud is also shown. In this case, five video streams are used, that is, 2·N+1 video streams, where N is the number of frames for the geometry atlas and the number of frames for the texture atlas. Note that in this case, the frame rate used by the geometry atlas and texture atlas streams is twice the frame rate of the occupancy atlas stream.

[0068] Therefore, in Fig. 9 In the example of , the geometry atlas, texture atlas and occupancy atlas are coded in separate video streams. However, the frame packing mode is defined in the V3C standard for MIV (i.e., its syntax structure is signaled in the vps_extention parameter), and it can further be used to store V-PCC atlases in one video stream, thus avoiding the need to align decoded data from multiple streams (outputs of multiple decoders). Fig.10 The diagram shows that the geometry atlas, texture atlas and occupancy atlas are packed into corresponding sub-pictures of the video picture. Therefore, when the frame packing mode of the V3C standard is used, 2·N+1 video streams are packed into one video stream.

[0069] As described above, including atlases in a sub-picture structure limits the efficiency with which these atlases can be encoded by a 2D video encoder conforming to, for example, the VVC standard. This is because: 1) the lack of flexibility in the currently defined sub-pictures leads to inefficient compression of patches; 2) the layout of sub-pictures is fixed for the entire sequence; and 3) sub-pictures defined within a rectangular picture are not allowed to dynamically change their size.

[0070] Aspects disclosed herein provide an enhanced sub-picture structure, i.e., a dynamic sub-picture, which overcomes the limitations mentioned above. As disclosed herein, features of dynamic sub-pictures include: 1) dynamic size (or resolution) changes of sub-pictures at the sub-picture level using reference picture resampling (RPR) support; 2) sub-picture layout of one-dimensional vectors of sub-pictures, i.e., 1D layout, so that sub-pictures do not need to be constrained to exist inside the picture structure; and 3) inter-picture prediction using (spatial or temporal) inter-picture sub-picture references. The dynamic sub-pictures disclosed herein are also referred to as sub-pictures hereinafter.

[0071] Supports adaptive sub-image size.

[0072] In one aspect, the size (or resolution) of a sub-picture can be changed dynamically at the sub-picture level. This is in contrast to the approach defined in the VVC standard, where the layout of sub-pictures, typically signaled in the SPS, is constrained to be constant within the CLVS. To support dynamically changing sub-picture sizes, the RPR function can be applied at the sub-picture level. Fig.11 The diagram shows sub-pictures, some of which have changed size within the frame sequence 0-383. Fig.11 In the example of , the resolution of some sub-pictures changes dynamically every 128 frames. For example, sub-picture SP2 is initially at full resolution, and then starting from frame 128, its resolution is changed to 2 / 3 of the full resolution. Note that RPR is needed to predict this sub-picture starting from frame 128 and forward while referencing sub-pictures from frames before frame 128. Similarly, sub-picture SP15, which is initially at full resolution, has its resolution changed to 1 / 2 of the full resolution starting from frame 128, and then returns to full resolution starting from frame 256 and forwards.

[0073] To enable this feature, new syntax elements are needed to dynamically indicate the size change of each sub-picture. In the example of Table 2, a first flag (sps_subpic_dynamic_size_flag) is added to signal whether the size of the sub-picture can be changed dynamically. When this flag is set to 1, for each sub-picture of index i (also referred to as sub-picture i for notation simplicity), the width and height ratios (sps_subpic_width_ratio[i] and sps_subpic_height_ratio[i]) are signaled. In this example, these syntax elements are inserted into the SPS. Therefore, pictures whose associated sub-pictures have changed resolution (size) require the above syntax elements to be inserted before the SPS. Fig.11In the example of , since the resolution of some sub-pictures changes three times, three SPSs should be set with the above syntax elements. In one aspect, when sps_subpic_dynamic_size_flag is not present, it can be inferred to be equal to 0, indicating that the dynamic sub-picture size feature is disabled.

[0074] Table 2: SPS signaling of dynamic sub-picture size modification.

[0075]

[0076]

[0077]

[0078] In one aspect, the scaling ratio of each sub-picture may be indicated in a picture parameter set (PPS), as shown in the example of Table 3. In this case, sub-picture scaling control is provided at the picture level (rather than at the sequence level). In another aspect, the scaling ratio of each sub-picture may be indicated in a picture header.

[0079] Table 3: PPS signaling of sub-picture ratios

[0080]

[0081]

[0082] Supports adaptive one-dimensional layout.

[0083] In one aspect, sub-pictures associated with a video frame may be organized in a list of sub-pictures (i.e., a one-dimensional (1D) layout). That is, instead of placing the sub-pictures in a picture structure of a video frame (e.g., Fig.11 ), but instead place the sub-pictures in a list containing sub-pictures of variable size (resolution). Fig.12 1D layout of sub-pictures SP1-SP8 associated with a video frame is illustrated. Thus, in one aspect, the sub-pictures may be arranged within a picture structure defined by their positions within the picture and their sizes, while in another aspect, the sub-pictures may be arranged in a 1D layout defined only by their sizes. Advantageously, patches may be assigned to different sub-pictures of different resolutions (such as 4-bit x-ray discs) based on the content of the patches or based on application-based priorities. Fig.11 or Fig.12 For example, patches containing uniform, continuous, or similar textures can be packed into a single sub-image to maintain spatial continuity across patches.

[0084] To enable the 1D layout feature, a new syntax element defined in this document by the flag: sps_subpic_1D_layout is introduced. When this flag is set to 1, signaling of the position of the sub-picture of index i within the 2D picture (i.e., sps_subpic_ctu_top_left_x[i] and sps_subpic_ctu_top_left_y[i]) is not required. Only the dimensions of the sub-picture are signaled. The signaling of sps_subpic_info_present_flag must be moved before indicating the picture size; when sps_subpic_1D_layout is set to 0, the size of the picture is signaled. These changes are shown in the example of Table 4 (the moved syntax is indicated by italic and strikethrough fonts, and the added syntax is indicated by a gray background). In one aspect, when sps_subpic_1D_layout is not present, it is inferred to be equal to 0, indicating that the adaptive sub-picture layout feature is disabled.

[0085] Table 4: Modified SPS signaling for ID layout of sub-pictures.

[0086]

[0087]

[0088] Note that the resolution of sub-pictures in 1D layout, which is usually signaled in SPS, is not constrained to be constant in CLVS. Note that when 1D layout is used, no signaling related to picture size is required in other structures such as PPS. The changes in PPS are shown in the example of Table 5.

[0089] Table 5: Modified PPS signaling for sub-picture 1D layout.

[0090]

[0091] Supports referencing sub-images within the same image or the same 1D layout.

[0092] In one aspect, the content of a sub-picture may be predicted using the content of another sub-picture contained in the same picture or in the same 1D layout. The prediction (i.e., intra-sub-picture prediction) may be coded or inferred (e.g., by template matching) and may be based on a reference sub-picture contained in the same picture or in the same 1D layout to which the current sub-picture belongs.

[0093] To enable this feature, it is necessary to know the reference sub-pictures for a given sub-picture (all belonging to the same picture or 1D layout). A signaling mechanism is introduced in this article to indicate for each sub-picture which sub-pictures are used as references for prediction. Table 6 provides an example syntax for sub-picture references. A new syntax element sps_max_num_reference_subpics is added. When this syntax element is equal to 0, it means that prediction from another sub-picture is not enabled. When it is greater than 0, prediction from another sub-picture is enabled. In this case, for each sub-picture of index i, for j=0 to j=(sps_max_num_reference_subpics–1), the index of the reference sub-picture sps_ref_subpic_id[i][j] is indicated.

[0094] Table 6: SPS signaling of reference sub-pictures.

[0095]

[0096]

[0097] In one aspect, a value of sps_subpic_id greater than the maximum value of the variable sps_ref_subpic_id[i][j] means that for the jth element, no reference subpic_id exists. In another aspect, the number of reference subpics is indicated for each subpicture. An example of a syntax for signaling the number of reference subpictures is provided in Table 7, showing a new syntax element sps_num_reference_subpics[i] that indicates the number of reference subpictures used by the subpicture of index i. The value of sps_num_reference_subpics[i] is in the range of 0 to (sps_max_num_reference_subpics–1).

[0098] Table 7: SPS signaling for multiple reference sub-pictures.

[0099]

[0100] In one aspect, in sps_ref_subpic_id[i][j], the reference sub-picture index j is constrained to be lower than the index i of the current sub-picture. This means that the sub-picture refers to the previous (already decoded) sub-picture. This enforces the decoding order so that sub-pictures are decoded in the order of increasing id numbers.

[0101] Fig.13 Diagram showing sub-image dependencies. Fig.13 , using the syntax of Table 7, the following values ​​are obtained:

[0102] ·sps_max_num_reference_subpics=2

[0103] For sub-image 0

[0104] o sps_num_reference_subpics[0]=0

[0105] For sub-image 1

[0106] o sps_num_reference_subpics[1]=1

[0107] o sps_ref_subpic_id[1][0] = 0

[0108] For sub-image 2

[0109] o sps_num_reference_subpics[2]=0

[0110] For sub-image 3

[0111] o sps_num_reference_subpics[3]=2

[0112] o sps_ref_subpic_id[3][0]=1

[0113] o sps_ref_subpic_id[3][1] = 2

[0114] For sub-image 4

[0115] o sps_num_reference_subpics[4]=1

[0116] o sps_ref_subpic_id[4][0] = 0

[0117] The above can also be written as a list "INTRA_REF_LIST" defined as follows:

[0118] INTRA_REF_LIST:

[0119] sps_ref_subpic_id[0] = {}

[0120] sps_ref_subpic_id[1] = {0}

[0121] sps_ref_subpic_id[2] = {}

[0122] ·sps_ref_subpic_id[3]={1,2}

[0123] ·sps_ref_subpic_id[4]={0}.

[0124] On the other hand, instead of indicating the index of the reference sub-picture (sps_ref_subpic_id[i][j]) used by the sub-picture of index i, a flag (sps_ref_subpic_flag[i][j]) is signaled for all sub-pictures except the current sub-picture (i.e., sub-pictures with index j=0 to sps_num_subpics_minus1, where j≠i) to indicate whether the sub-picture of index j is used as a reference for the sub-picture of index i, as shown in the example of Table 8.

[0125] Table 8: SPS signaling of reference sub-pictures using flags.

[0126]

[0127] On the other hand, the current sub-picture of index i can only refer to the previous sub-picture, as shown in Table 9. Table 9 shows the modifications made with respect to the syntax in Table 8.

[0128] Table 9: SPS signaling of reference sub-pictures using flags.

[0129]

[0131] In one aspect, a sub-picture implicitly depends on N previous sub-pictures (in signaling order), where N can be signaled in the SPS, specified in the profile, or specified by default (e.g., N is set to 6). In another aspect, signaling is done in the PPS instead of the SPS to have greater flexibility in controlling the reference of sub-pictures. In yet another aspect, signaling is done in an adaptation parameter set (AP) instead of the SPS to have even greater flexibility in controlling the reference of sub-pictures than the PPS.

[0132] Supports referencing sub-pictures from previous pictures or previous 1D layouts.

[0133] A sub-picture may be predicted using another sub-picture from another picture or another 1D layout as a reference. This prediction (i.e., inter-sub-picture prediction) is similar to temporal prediction and may be based on motion compensation (where motion data is encoded or inferred, e.g., by template matching) and using a reference sub-picture from a picture or 1D layout associated with a previous (in decoding order) video frame.

[0134] As a default (and as currently defined in the VVC standard), a sub-picture of index k in the current picture can use a reference sub-picture of the same index k from a reference picture. This means that the content of the sub-picture is assumed to be more related to the content of the temporally corresponding sub-picture than to the content of the temporally non-corresponding sub-picture. Furthermore, since all motion vectors (MVs) are calculated relative to the current sub-picture, MVs pointing outside the reference sub-picture need to be padded with the content pointed to.

[0135] Similar to other standards, in VVC, referencing of pictures is facilitated by reference picture lists (RPLs). There are two RPLs—RPL 0 (L0) and RPL 1 (L1)—that are directly signaled and derived.

[0136] • Reference picture marking is directly based on L0 and L1, thereby utilizing both active and inactive entries in the RPL, while only the active entry can be used as a reference index in inter prediction of a CTU.

[0137] The information required to derive the two RPLs is signaled by syntax elements and syntax structures in the SPS, PPS, picture header (PH), and slice header (SH). A predefined RPL structure is signaled in the SPS for reference in the PH or SH.

[0138] These two RPLs are generated for all slice types: B, P, and I.

[0139] Two RPLs are constructed without using the RPL Initialization procedure or the RPL Modification procedure.

[0140] Fig.14 The reference picture list is illustrated. For example, when decoding a picture with POC 6, a conventional decoder uses a reference picture list. Each index in the list (L0 or L1) points to a reference picture in the decoded picture buffer (DPB) corresponding to the POC of a picture already decoded in the sequence. In this example, a given picture with POC 6 can reference pictures with POC: 0, 2, 4, 8 from the list.

[0141] In one aspect, Fig.14 The two reference picture lists L0 and L1 of are modified to enable reference to sub-pictures. Fig.15A reference sub-picture list is illustrated. As shown, a given sub-picture of a given picture may refer to sub-picture 3 of a picture with POC 0, sub-picture 2 of a picture with POC 2, sub-pictures 1 and 3 of a picture with POC 4, or sub-picture 1 of a picture with POC 8. Thus, for each reference picture index, the POC and sub-picture index are used to reference the correct sub-picture to use as a reference for decoding the current picture or current sub-picture in the layout.

[0142] In one aspect, the list of available sub-pictures to be used as references for inter-sub-picture prediction (i.e., for predicting sub-pictures within a different picture or layout) can be the same as the list of sub-pictures to be used as references for intra-sub-picture prediction (i.e., for predicting sub-pictures within the same picture or layout). Thus, the same reference sub-picture list can be used for both intra and inter-sub-picture prediction. In another aspect, the current sub-picture id is implicitly added to the reference sub-picture list for inter-sub-picture prediction. For each element in the RPL (L0 or L1), both the index of the picture in the DPB and the index of the reference sub-picture(s) as indicated in sps_ref_subpic_id[i] (where i is the index of the current sub-picture) are signaled or derived from the signaled syntax elements. For example, the reference sub-picture id is implicitly added to the reference sub-picture list for inter-sub-picture prediction. Fig.13 , sps_ref_subpic_id is filled into the list "INTER_REF_LIST" defined as follows:

[0143] INTER_REF_LIST:

[0144] sps_ref_subpic_id[0] = {0}

[0145] ·sps_ref_subpic_id[1]={1,0}

[0146] sps_ref_subpic_id[2] = {2}

[0147] ·sps_ref_subpic_id[3]={3,1,2}

[0148] ·sps_ref_subpic_id[4]={4,0}

[0149] It can be noted that INTER_REF_LIST corresponds to INTRA_REF_LIST, where the reference list of each sub-picture of index k (k=0 to 4) is completed by index k.

[0150] As an example, see Fig.15, assuming that the DPB has the following reference pictures: DBP = {0, 2, 4, 8} and the index of the current sub-picture is 3 (where the reference sub-picture list sps_ref_subpic_id[3] = {3, 1, 2}), the following picture / sub-picture reference table is obtained:

[0151]

[0152] In one aspect, the reference sub-picture list can be extended. When inter-frame sub-picture prediction is applied, the sub-picture can access all sub-pictures associated with previously decoded frames in the DPB. For this reason, for inter-frame coded sub-pictures, the constraint of having a reference sub-picture index that is less than the current sub-picture index is no longer required. In one aspect, the reference sub-picture list is extended by symmetrizing the reference sub-picture list already described. This means that when a sub-picture with index i has a sub-picture with index j in its sub-picture reference list, the reference sub-picture with index i is automatically added to the reference sub-picture list of sub-picture j. Using the same example as above,

[0153] INTRA_REF_LIST:

[0154] sps_ref_subpic_id[0] = {}

[0155] sps_ref_subpic_id[1] = {0}

[0156] sps_ref_subpic_id[2] = {}

[0157] ·sps_ref_subpic_id[3]={1,2}

[0158] sps_ref_subpic_id[4] = {0}

[0159] For the sub-picture written in the intra-frame code given above (corresponding to Fig.13 The reference sub-picture list of ) is transformed into the following list for inter-coded sub-pictures:

[0160] sps_ref_subpic_id[0]={0,1,4}

[0161] ·sps_ref_subpic_id[1]={1,0,3}

[0162] ·sps_ref_subpic_id[2]={2,3}

[0163] ·sps_ref_subpic_id[3]={3,1,2}

[0164] ·sps_ref_subpic_id[4]={4,0}

[0165] This means, for example, that sub-picture 1 of the current picture can refer to sub-pictures 1, 0, 3 of the reference picture (sps_ref_subpic_id[1]={1, 0, 3}).

[0166] On the other hand, for inter-coded sub-pictures, the list is not derived implicitly, but is explicitly signaled using the same syntax as before without constraints on the sub-picture index, as shown in Table 10. In the case where transmission is not required, the current sub-picture index is implicitly derived as the first reference sub-picture index in the list.

[0167] Table 10: SPS signaling of reference sub-pictures.

[0168]

[0169] A motion vector used to reference a block in a different sub-picture (of a different picture or layout) is usually relative to the top left corner of the referenced sub-picture, ie, pixel coordinate (0, 0). Fig.16 The diagram shows a reference of a block in a reference sub-picture 1610 of a reference picture 1620 by a current sub-picture 1650 of a current picture 1660. Fig.16 As shown in the example of , the same pixel coordinate system is used when accessing the motion compensation block 1670 from the reference sub-picture.

[0170] In one aspect, when a reference sub-picture list is transmitted, additional spatial information may be transmitted for the corresponding index in the list. For example, an offset (translation) value may be transmitted to indicate the relative position of a block from the current sub-picture and a corresponding block from the reference sub-picture (used for prediction), such as Fig.17 as shown in the example. Fig.17 The diagram shows the relative reference of the current sub-picture 1750 of the current picture 1760 to the block in the reference sub-picture 1710 of the reference picture 1720. Fig.17 As shown in the example of , the offset value is used to associate the motion compensated block 1730 from the reference sub-picture 1710 with the block in the current sub-picture 1750.

[0171] In one aspect, the spatial information (transmitted for the corresponding index in the reference sub-picture list) may include a geometric transformation. For example, a geometric transformation that can be applied to a reference sub-picture may include a translation operation (e.g., for a reference sub-picture such as Fig.17The information in the image may be a relative reference to the image (interpreted as a relative reference), a mirror operation along the x or y axis, or a rotation operation. In the latter, for example, the rotation angle may be restricted to the set of {90, 180, 270} degrees, allowing geometric transformations only in pixel coordinates. The following table shows an example for spatial information:

[0172]

[0173]

[0174] Where H and W are the height and width of the reference sub-picture. The syntax for using all spatial transforms is shown in the example of Table 11.

[0175] Table 11: Syntax for signaling spatial transforms.

[0176]

[0177] Note that the possible values ​​of the variable sps_ref_subpic_transfo_rotate can be set to {0, 1, 2}, encoding rotations of {90, 180, 270} degrees accordingly. The order in which the geometric transformations are applied is fixed and the order is the same at the encoder and decoder, for example: rotation, mirror, mirror and translation.

[0178] Fig.18 Inter- and intra-frame referencing through sub-pictures is illustrated. Fig.18 Sub-pictures SP1-SP8 of the current 1D layout of sub-picture 1810 and sub-pictures SP1-SP6 of the reference 1D layout of sub-picture 1820 are shown. As shown, sub-picture SP7 belonging to the current 1D layout 1810 refers to sub-picture SP4 from the same 1D layout 1810 (i.e., intra-frame sub-picture reference), and also refers to sub-pictures SP3 and SP6 belonging to the reference 1D layout 1820 (i.e., inter-frame sub-picture reference). Because SP7 refers to SP4, if SP4 has a different image resolution from SP7, the prediction process may require resampling. In addition, the prediction process may require additional spatial transformations (e.g., mirroring or rotation). For temporal prediction, SP7 refers to SP3 and SP6 from the reference 1D layout 1820. As shown, prediction based on SP6 may require spatial transformation, but when SP6 has the same image resolution as SP7, no resampling is required. When SP3 and SP7 do not have the same image resolution, prediction based on SP3 may require resampling.

[0179] For by Fig.18 In the illustrated case, the values ​​of the different syntax elements mentioned above can be set as follows:

[0180] · For the INTRA_REF_LIST of SP7 to SP4

[0181] o sps_ref_subpic_id[7][0] = 4

[0182] o sps_ref_subpic_transfo_mirrorX[7][0] = 1

[0183] o sps_ref_subpic_transfo_mirrorY[7][0] = 0

[0184] o sps_ref_subpic_transfo_rotation[7][0] = 0

[0185] · For the INTER_REF_LIST of SP7 to SP3

[0186] o sps_ref_subpic_id[7][0] = 3

[0187] o sps_ref_subpic_transfo_mirrorX[7][0] = 0

[0188] o sps_ref_subpic_transfo_mirrorY[7][0] = 0

[0189] o sps_ref_subpic_transfo_rotation[7][0] = 0

[0190] · For the INTER_REF_LIST of SP7 to SP6

[0191] o sps_ref_subpic_id[7][1] = 6

[0192] o sps_ref_subpic_transfo_mirrorX[7][1] = 0

[0193] o sps_ref_subpic_transfo_mirrorY[7][0] = 0

[0194] o sps_ref_subpic_transfo_rotation[7][1] = 1.

[0195] Fig.191 is a flow chart of an example method 1900 for encoding dynamic volume data. The method 1900 begins in step 1910 by receiving a sequence of volume data sets representing dynamic volume data to be encoded. The volume data sets may include (one or more) image views, (one or more) corresponding depth maps, and metadata including projection parameters of (one or more) corresponding image views. For the volume data sets in the sequence, encoding may be performed by steps 1920, 1930, and 1940 of the method 1900. In step 1920, a patch may be generated. The patch includes a corresponding two-dimensional representation of the volume data set, for example, the patch may contain geometric data or attribute data derived from the volume data set. In step 1930, the generated patches may be packed into one or more dynamic sub-pictures associated with a video frame, for example, the packing may be based on the content of the patches, wherein patches containing similar content are packed into the same sub-picture. Then, in step 1940, the two-dimensional video encoder encodes the one or more dynamic sub-pictures into a bitstream of coded video data.

[0196] Fig. 20 2000 for decoding dynamic volume data. The method 2000 begins in step 2010 by receiving a bitstream and coding a sequence of volume data sets representing the dynamic volume data. For the volume data sets in the sequence, decoding can be performed by steps 2020, 2030 and 2040 of the method 2000. In step 2020, a two-dimensional video decoder can decode one or more dynamic sub-pictures associated with a video frame from the bitstream. In step 2030, a patch is extracted from the decoded one or more dynamic sub-pictures, the patch including a corresponding two-dimensional representation of the volume data set. Then, in step 2040, the volume data set can be reconstructed based on the patch.

[0197] We have described several aspects and embodiments in this disclosure. These aspects and embodiments provide at least the following outputs and results, including all combinations across different claim categories and types:

[0198] • According to any aspect described herein, encoding into the coded video data a syntax element capable of enabling a decoder to decode the coded video data.

[0199] • A bitstream comprising one or more of the described syntax elements, or variants thereof. A bitstream may be any collection of data, whether transmitted, stored, or otherwise available.

[0200] Creation, transmission, reception and / or decoding of bitstreams.

[0201] An electronic device (e.g., a TV, set-top box, mobile phone, or tablet) that tunes (e.g., using a tuner) a channel to receive the bitstream or receives the bitstream over the air (e.g., using an antenna). The electronic device decodes the syntax elements from the bitstream and optionally displays (e.g., using a monitor, screen, or any other type of display) the resulting image.

[0202] Various other generalized and specific outputs, results, implementations, and claims are also supported and contemplated throughout this disclosure.

[0203] Various methods are described herein, and each method includes one or more steps or actions for realizing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. In addition, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding". Unless specifically required, the use of such terms does not mean the sequencing of the operation to modification. Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can occur in a time period such as before, during, or overlapping with the second decoding.

[0204] The various methods and other aspects described in this application can be used to modify modules, such as Figure 2 and Figure 3 Modules of the video encoder 200 and the video decoder 300 shown. In addition, the present aspects are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, and extensions of any such standards and recommendations. Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.

[0205] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0206] Various embodiments relate to decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refers to a broader decoding process will be clear and is considered to be well understood by those skilled in the art.

[0207] Various embodiments relate to encoding. In a manner similar to the discussion above regarding "decoding", "encoding" as used in this application can encompass, for example, all or part of the processes performed on input video data to produce an encoded bitstream. In addition, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "encoded" or "coded" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.

[0208] Note that syntax elements as used herein are descriptive terms. Therefore, they do not exclude the use of other syntax element names.

[0209] This disclosure has described various pieces of information that may be transmitted or stored, such as syntax, for example. This information may be packaged or arranged in various ways, including, for example, those common in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or a slice header), or SEI message. Other ways are also available, including, for example, those common to system-level or application-level standards, such as signaling the information into one or more of the following:

[0210] a. SDP (Session Description Protocol), a format for describing multimedia communication sessions, used for the purpose of session announcement and session invitation (eg described in RFC) and used in conjunction with RTP (Real-time Transport Protocol) transport.

[0211] b. DASH MPD (Media Presentation Description) descriptor, such as that used in DASH and delivered over HTTP. A descriptor is associated with a representation or a set of representations to provide additional characteristics to the content representation.

[0212] c. RTP header extension, eg as used during RTP streaming.

[0213] d. ISO Base Media File Format, such as the format used in OMAF, and uses boxes which are object-oriented building blocks defined by a unique type identifier and a length (also referred to as "atoms" in some specifications).

[0214] e. HLS (HTTP Live Streaming) manifest delivered over HTTP. A manifest may be associated with, for example, a version or set of versions of content to provide characteristics of the version or set of versions.

[0215] The embodiments and aspects described herein can be implemented in, for example, methods or processes, devices, software programs, data streams or signals. Even if only discussed in the context of a single embodiment form (e.g., discussed only as a method), the embodiments of the features discussed can also be implemented in other forms (e.g., devices or programs). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant (PDA), and other devices that facilitate the communication of information between terminal-users.

[0216] Reference to "one aspect / aspects" or "one embodiment / embodiment" or "one implementation / implementation", and other variations thereof, means that a particular feature, structure, characteristic, etc. described in connection with that aspect / embodiment / implementation is included in at least one embodiment. Thus, the appearance of the phrase "in one aspect / aspects" or "in one embodiment / embodiment" or "in one implementation / implementation", and any other variations thereof throughout this application, are not necessarily all referring to the same embodiment.

[0217] Additionally, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0218] Additionally, the present application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0219] Furthermore, the present application may refer to "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is generally referred to in one way or another during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0220] It is to be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many of the listed items as are apparent to one of ordinary skill in this and related arts.

[0221] In addition, as used herein, the word "signaling" especially indicates something to the corresponding decoder. For example, in some embodiments, the encoder signals the quantization parameter used for dequantization. In this way, in one embodiment, the same parameters are used at both the encoder side and the decoder side. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual data, bit savings are achieved in various embodiments. It is to be understood that signaling can be done in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signaling" is mentioned above, the word "signal" can also be used as a noun in this article.

[0222] It will be apparent to one of ordinary skill in the art that embodiments may generate various signals that are formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. A signal may be stored on a processor readable medium.

Claims

1. A method for encoding, comprising: receiving a sequence of volume data sets representing dynamic volume data; as well as Encoding the sequence, for a volume data set in the sequence, the encoding includes: generating a patch comprising a corresponding two-dimensional representation of the volumetric data set; Packing the patches into one or more dynamic sub-pictures associated with a video frame; and The one or more dynamic sub-pictures are encoded by a two-dimensional video encoder into a bitstream of coded video data.

2. The method according to claim 1, wherein: The volumetric dataset comprises at least one image view, a corresponding depth map, and metadata comprising projection parameters of the at least one image view.

3. The method according to claim 1 or 2, wherein: A patch of the generated patches contains geometric data or attribute data derived from a corresponding volume data set.

4. The method according to any one of claims 1 to 3, wherein: Packing the patches into one or more dynamic sub-pictures is based on the content of the patches, wherein patches containing similar content are packed into the same sub-picture.

5. The method according to any one of claims 1 to 4, further comprising: During encoding of the sequence, a size of a sub-picture of the one or more dynamic sub-pictures is changed.

6. The method according to any one of claims 1 to 5, further comprising: The change in size of the sub-picture is signaled in the bitstream using the following syntax elements: sps_subpic_dynamic_size_flag, sps_subpic_width_ratio, and sps_subpic_height_ratio.

7. The method according to any one of claims 1 to 6, further comprising: The one or more dynamic sub-pictures are placed within a picture structure, wherein each of the sub-pictures is defined by a size of the sub-picture and an offset relative to the picture structure.

8. The method according to any one of claims 1 to 6, further comprising: The one or more dynamic sub-pictures are placed in a one-dimensional (1D) layout, wherein each of the sub-pictures is defined by a sub-picture size.

9. The method according to any one of claims 1 to 6 or 8, further comprising: The placement of the one or more dynamic sub-pictures in the 1D layout is signaled in the bitstream using the syntax element sps_subpic_1D_layout.

10. The method according to any one of claims 1 to 9, further comprising: The two-dimensional video encoder performs intra-frame sub-picture prediction on the content of the sub-pictures of the one or more dynamic sub-pictures according to the content of at least one reference sub-picture of the one or more dynamic sub-pictures.

11. The method according to claim 10, further comprising: The reference sub-picture is signaled in the bitstream using the syntax elements sps_max_num_reference_subpics and sps_ref_subpic_id.

12. The method according to claim 10, further comprising: The reference sub-pictures are signaled in the bitstream using the syntax element sps_num_reference_subpics.

13. The method according to claim 10, further comprising: The reference sub-picture is signaled in the bitstream using the syntax element sps_ref_subpic_flag.

14. The method according to any one of claims 1 to 9, further comprising: The two-dimensional video encoder performs inter-sub-picture prediction on content of sub-pictures of the one or more dynamic sub-pictures based on content of at least one reference sub-picture of a dynamic sub-picture associated with at least one other video frame.

15. The method according to claim 14, further comprising: The reference sub-picture is signaled in the bitstream using the syntax elements sps_max_num_inter_reference_subpics and sps_inter_ref_subpic_id.

16. The method according to any one of claims 1 to 9, 10 or 14, further comprising: A reference sub-picture list including indices to reference sub-pictures is signaled in the bitstream, which is used for prediction by the two-dimensional video encoder.

17. The method according to claim 16, further comprising: The reference sub-picture list is extended, wherein the reference sub-picture list is symmetric.

18. The method according to claim 16, wherein: The reference sub-picture list further includes spatial information specifying at least one of a translation operation, a mirror operation, or a rotation operation associated with the reference sub-pictures in the reference sub-picture list.

19. A method for decoding, comprising: receiving a bitstream of coded video data generated by an encoder, and coding a sequence of volumetric data sets representing dynamic volumetric data; as well as The sequence is decoded, and for a volume data set in the sequence, the decoding includes: Decoding, by a two-dimensional video decoder, one or more dynamic sub-pictures associated with a video frame from the bitstream; extracting a patch from the decoded one or more dynamic sub-pictures, the patch comprising a corresponding two-dimensional representation of the volumetric dataset; and The volumetric dataset is reconstructed based on the patch.

20. The method according to claim 19, wherein: The one or more dynamic sub-pictures are placed by the encoder within a picture structure, wherein each of the sub-pictures is defined by a size of the sub-picture and an offset relative to the picture structure.

21. The method according to claim 19, wherein: The one or more dynamic sub-pictures are placed by the encoder in a one-dimensional (1D) layout, wherein each of the sub-pictures is defined by a sub-picture size.

22. The method according to any one of claims 19 to 21, further comprising: The two-dimensional video decoder performs intra-frame sub-picture prediction on the content of the sub-pictures of the one or more dynamic sub-pictures according to the content of at least one reference sub-picture of the one or more dynamic sub-pictures.

23. The method according to any one of claims 19 to 21, further comprising: Content of a sub-picture of the one or more dynamic sub-pictures is inter-sub-picture predicted by the two-dimensional video decoder based on content of at least one reference sub-picture of a dynamic sub-picture associated with at least one other video frame.

24. An apparatus for encoding, comprising: at least one processor; as well as a memory storing instructions that, when executed by the at least one processor, cause the apparatus to: receiving a sequence of volume data sets representing dynamic volume data; as well as Encoding the sequence, for a volume data set in the sequence, the encoding includes: generating a patch comprising a corresponding two-dimensional representation of the volumetric data set; Packing the patches into one or more dynamic sub-pictures associated with a video frame; and The one or more dynamic sub-pictures are encoded by a two-dimensional video encoder into a bitstream of coded video data.

25. The device according to claim 24, wherein: The encoding of the one or more dynamic sub-pictures further comprises: The one or more dynamic sub-pictures are placed within a picture structure, wherein each of the sub-pictures is defined by a size of the sub-picture and an offset relative to the picture structure.

26. The device according to claim 24, wherein: The encoding of the one or more dynamic sub-pictures further comprises: The one or more dynamic sub-pictures are placed in a one-dimensional (1D) layout, wherein each of the sub-pictures is defined by a sub-picture size.

27. The device according to any one of claims 24 to 26, wherein The encoding of the one or more dynamic sub-pictures further comprises: The two-dimensional video encoder performs intra-frame sub-picture prediction on the content of the sub-pictures of the one or more dynamic sub-pictures according to the content of at least one reference sub-picture of the one or more dynamic sub-pictures.

28. The device according to any one of claims 24 to 26, wherein The encoding of the one or more dynamic sub-pictures further comprises: The two-dimensional video encoder performs inter-sub-picture prediction on content of sub-pictures of the one or more dynamic sub-pictures based on content of at least one reference sub-picture of a dynamic sub-picture associated with at least one other video frame.

29. An apparatus for decoding, comprising: at least one processor; as well as a memory storing instructions that, when executed by the at least one processor, cause the apparatus to: receiving a bitstream of coded video data generated by an encoder, and coding a sequence of volumetric data sets representing dynamic volumetric data; as well as The sequence is decoded, and for a volume data set in the sequence, the decoding includes: Decoding, by a two-dimensional video decoder, one or more dynamic sub-pictures associated with a video frame from the bitstream; extracting a patch from the decoded one or more dynamic sub-pictures, the patch comprising a corresponding two-dimensional representation of the volumetric dataset; and The volumetric dataset is reconstructed based on the patch.

30. The device according to claim 29, wherein: The one or more dynamic sub-pictures are placed by the encoder within a picture structure, wherein each of the sub-pictures is defined by a size of the sub-picture and an offset relative to the picture structure.

31. The device according to claim 29, wherein: The one or more dynamic sub-pictures are placed by the encoder in a one-dimensional (1D) layout, wherein each of the sub-pictures is defined by a sub-picture size.

32. The apparatus according to any one of claims 29 to 31, further comprising: The two-dimensional video decoder performs intra-frame sub-picture prediction on the content of the sub-pictures of the one or more dynamic sub-pictures according to the content of at least one reference sub-picture of the one or more dynamic sub-pictures.

33. The apparatus according to any one of claims 29 to 31, further comprising: Content of a sub-picture of the one or more dynamic sub-pictures is inter-sub-picture predicted by the two-dimensional video decoder based on content of at least one reference sub-picture of a dynamic sub-picture associated with at least one other video frame.

34. A non-transitory computer readable medium comprising instructions executable by at least one processor to perform a method for encoding, the method comprising: receiving a sequence of volume data sets representing dynamic volume data; as well as Encoding the sequence, for a volume data set in the sequence, the encoding includes: generating a patch comprising a corresponding two-dimensional representation of the volumetric data set; Packing the patches into one or more dynamic sub-pictures associated with a video frame; and The one or more dynamic sub-pictures are encoded by a two-dimensional video encoder into a bitstream of coded video data.

35. A non-transitory computer readable medium comprising instructions executable by at least one processor to perform a method for decoding, the method comprising: receiving a bit stream of coded video data and coding a sequence of volume data sets representing dynamic volume data; as well as The sequence is decoded, and for a volume data set in the sequence, the decoding includes: Decoding, by a two-dimensional video decoder, one or more dynamic sub-pictures associated with a video frame from the bitstream; extracting a patch from the decoded one or more dynamic sub-pictures, the patch comprising a corresponding two-dimensional representation of the volumetric dataset; and The volumetric dataset is reconstructed based on the patch.