Encoding formats for optimized encoding of volumetric video

By clustering and encoding the graph blocks of volume videos, the problem of high time redundancy utilization and poor compression performance when encoding volume videos in the prior art is solved, and more efficient video compression and flexible 3D scene description are achieved.

CN120153397APending Publication Date: 2025-06-13INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076581.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-04
Filing Date
2023-10-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art cannot effectively utilize the high time redundancy in 3D scenes when encoding volume videos, and the traditional 2D video encoder's assumptions of the patch map do not conform to, resulting in poor compression performance.

Method used

By obtaining a list of graph blocks, cluster patches according to similarity and continuity criteria, and metadata of volume scenarios are generated, including the number, size, and location of patches. The list of graph blocks and metadata are then encoded in the data stream.

Benefits of technology

More efficient video compression is achieved, improving encoding efficiency by maximizing intra-frame correlation, and supporting compact and flexible 3D scene descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153397A_ABST
    Figure CN120153397A_ABST
Patent Text Reader

Abstract

Methods and apparatus are disclosed for encoding and decoding a volumetric scene in a data stream, the volumetric scene being formatted as a list of map blocks. Embodiments of a syntax describing metadata of such a data stream are also disclosed. For example, by setting the size of the primary map to zero, an indication is provided that the volume scene is formatted as a list of map blocks. The number of atlas tiles in the list and data describing each atlas tile are also provided. In decoding, a list of map blocks is retrieved by using these metadata.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to European Application No. 22306665.5, filed on November 4, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0002] This principle generally relates to the field of three-dimensional (3D) scenes and volumetric video content as a sequence of 3D scenes. This document is also understood in the context of encoding, formatting, and decoding data representing the texture and geometry of 3D scenes for presenting volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). Specifically, this document relates to encoding a sequence of atlas lists prepared for optimizing their encoding with a video encoder. Background Art

[0003] This section is intended to introduce the reader to aspects of the field that may be relevant to aspects of the principle described and / or claimed below. This discussion is considered to be helpful in providing background information to the reader to facilitate a better understanding of the aspects of the principle. Accordingly, it should be understood that these statements are to be read from this perspective and not as an admission of prior art.

[0004] Advances in 3D capture and rendering technologies have made volumetric video an integral part of virtual / augmented / mixed reality (VR / AR / MR) applications. Volumetric video can be defined as a sequence of 3D frames and can be in various 3D formats such as point clouds, meshes, or multi-view plus depth video. For example, the Moving Picture Experts Group (MPEG) has put forward the need for a high coding efficiency standard for compressing visual volumetric data.

[0005] A method of encoding volumetric frames includes converting 3D volumetric information into a set of 2D images and associated data. Then, the converted 2D images can be encoded using a 2D video encoder such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or Versatile Video Coding (VVC), and the associated data can be encoded in an additional metadata stream. The encoded images and associated metadata can then be decoded and used to reconstruct the 3D volumetric information. The Visual Volumetric Video (V3C) coding standard developed by MPEG belongs to this 2D video-compatible method.

[0006] A video-based volume encoder converts an input 3D frame into a set of 2D videos with an image format compatible with traditional 2D video encoders. However, a patch atlas packs together a set of patches that convey a part of a 3D scene captured from several camera viewpoints. Traditional 2D video coding such as HEVC or VVC exploits assumptions about the statistics of the input images to achieve high compression performance. The main underlying assumption is that the 2D frames to be compressed are projections (perspective or equirectangular in the case of 360° content) of a 3D scene from a unique viewpoint, thus generating multiple intra- and inter-frame correlations. However, the patch atlas output by the volume encoder does not satisfy such an assumption, which includes packing multiple patches without adjacent correlations in consecutive frames of the same atlas. Summary of the Invention

[0007] A simplified overview of the present principle is given below to provide a basic understanding of some aspects of the present principle. This summary of the invention is not an extensive review of the present principle. It is not intended to identify the key or critical elements of the present principle. The following overview merely presents some aspects of the present principle in a simplified form as a prelude to the more detailed description provided below.

[0008] The present principle relates to a method for encoding a volume scene. The method includes obtaining a list of atlas tiles. The atlas tiles pack patches that are clustered according to similarity and continuity criteria. In one embodiment, the video component of the atlas tile list interleaves the video component of the atlas tile with the video component of the depth atlas tile. The method includes generating metadata for the volume scene. The metadata includes an indication of whether the volume scene is formatted as a list of atlas tiles, the number of atlas tiles in the list of atlas tiles, and for each atlas tile in the list: size, the position within the atlas tile of each patch packed in the atlas tile. In one embodiment, the indication is provided by setting the size of the main atlas to zero. Then, according to the present principle, the list of atlas tiles and the metadata are encoded in a data stream.

[0009] The present principle relates to a device that includes a memory associated with a processor configured to implement the above method.

[0010] The present principle relates to a data stream generated by the above device.

[0011] The present principle also relates to a method for decoding a volumetric scene from a data stream. The method includes decoding metadata from the data stream. The metadata includes an indication of whether the volumetric scene is formatted as a list of atlas blocks, the number of atlas blocks in the list of atlas blocks, and for each atlas block in the list: a size, and the position within the atlas block of each patch packed in the atlas block. In one embodiment, the indication is provided by setting the size of the main atlas to zero. Then, according to the present principle, a list of atlas blocks is decoded from the data stream based on the decoded metadata.

[0012] The present principle relates to an apparatus that includes a memory associated with a processor configured to implement the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The present disclosure will be better understood and other specific features and advantages will become apparent when reading the following description with reference to the accompanying drawings, in which: - Figure 1 is an example of a patch atlas that packs the central view of a 3D scene together with a set of smaller patches that transmit occluded parts from other camera viewpoints in the same video frame; - Figure 2 shows a frame that packs two atlas components (texture atlas and downsampled depth atlas) in a unique video frame; - Figure 3 shows an example architecture of a device that can be configured to implement a method for encoding or decoding a volumetric scene from a data stream according to the present principle; - Figure 4 shows an example of an embodiment of the syntax of a stream when data is transmitted via a packet-based transport protocol; - Figure 5 shows a patch atlas method with an example of four projection centers; - Figure 6 shows different solutions for encoding a video patch atlas in a more efficient manner. DETAILED DESCRIPTION

[0014] The present principle will be described more fully hereinafter with reference to the accompanying drawings, in which examples of the present principle are shown. However, the present principle may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, while the present principle is amenable to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the present principle is not intended to be limited to the particular forms disclosed, but on the contrary, the present disclosure will cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principle as defined by the claims.

[0015] The terms used herein are for the purpose of describing particular examples only and are not intended to limit the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "includes" and / or "including", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Further, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other elements, no intervening elements are present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0016] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of the present principles, a first element may be referred to as a second element and, similarly, a second element may be referred to as a first element.

[0017] Although some of the figures include arrows on communication paths to indicate the primary direction of communication, it should be understood that communication can occur in a direction opposite to that shown by the arrows.

[0018] Some examples are described with reference to block diagrams and operational flowcharts, where each block represents a circuit element, a module, or a portion of code that includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other embodiments, the functions noted in the blocks may not occur in the order described. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending upon the functions involved.

[0019] References herein to "according to an example" or "in an example" mean that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The phrases "according to an example" or "in an example" that appear in different places in the specification are not necessarily all referring to the same example, nor are they necessarily separate or alternative examples that are mutually exclusive of other examples.

[0020] Reference numerals appearing in the claims are for illustration purposes only and shall have no limiting effect on the scope of the claims.The present examples and variations may be employed in any combination or sub-combination although not explicitly described.

[0021] A video-based volumetric encoder transforms the input 3D frames into a collection of 2D videos with an image format compatible with conventional 2D video encoders. The resulting picture content has very specific properties.

[0022] Figure 1 is an example of a patch map that packs a central view of a 3D scene 11 together with a collection of smaller patches 12 that convey de-occluded parts from other camera viewpoints in the same video frame. A corresponding patch map 13 with the same layout is provided to encode the depth component. For example, such an encoding format is the format adopted in standards such as MPEG Immersive Video (MIV).

[0023] 2D video coding such as HEVC or VVC exploits assumptions about the statistics of the input image to achieve high compression performance. The main underlying assumption is that the 2D frames to be compressed are projections of the 3D scene from a unique viewpoint (either perspective or equirectangular in the case of 360° content), resulting in multiple intra- and inter-frame dependencies. Figure 1 As shown, the patch map output by the volume encoder does not meet such assumptions, and the patch map includes packing multiple patches without obvious correlation between adjacent patches.

[0024] Figure 5 An example patch atlas method with four projection centers is shown. A 3D scene 50 includes characters. For example, projection center 51 is a perspective camera and camera 53 is an orthographic camera. The camera may also be an omnidirectional camera with, for example, a spherical mapping (e.g., an equirectangular mapping) or a cubic mapping. According to the projection operation described in the projection data of the metadata, the 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the projection center. Figure 5 In the example of , the projection of the points captured by camera 51 is mapped onto patch 52 according to perspective mapping, and the projection of the points captured by camera 53 is mapped onto patch 54 according to orthogonal mapping.

[0025] The clustering of projected pixels produces a plurality of 2D patches, which are packed in a rectangular atlas 55. The organization of the patches within the atlas defines the atlas layout. In one embodiment, two atlases have the same layout: one for texture (i.e., color) information and one for depth information. Two patches captured by the same camera or two different cameras may include information representing the same portion of a 3D scene, such as, for example, patches 54 and 56.

[0026] The packing operation generates patch data for each generated patch. The patch data includes a reference to the projection data (e.g., an index in a projection data table or a pointer to the projection data (i.e., an address in memory or in a data stream)) and information describing the position and size of the patch within the atlas (e.g., top-left coordinates, size in pixels, and width). The patch data items are added to the metadata to be encapsulated in the data stream in association with the compressed data of one or two atlases.

[0027] Existing formats for encoding atlas-based representations of 3D scenes, such as the MIV format (Text of ISO / IEC FDIS23090-12MPEG Immersive Video, ISO / IEC JTC 1 / SC 29 / WG 4, N00270), do not provide tools or features to exploit the high temporal redundancy of most 3D scenes. For example, the MIV standard allows the patch-based 3D scene description to be split into multiple atlases (the atlases themselves can be split into multiple blocks). The patch packing layout associated with those atlases and the associated projection parameters are transmitted to a separate "atlas data" sub-bitstream. By sending atlas frames at given consecutive moments, a complete and self-contained ("intra-coded") refresh of the entire atlas data is only allowed by the MIV profile. While the corresponding geometry and property (e.g., texture, transparency) samples of the patch atlases are transmitted in the video sub-bitstream at the full video frame rate. The TMIV reference software (Test Model 14 for MPEG Immersive Video, ISO / IEC JTC 1 / SC 29 / WG 4, N00242) implements a periodic regular refresh of the atlas data every 32 video frames, which corresponds to the intra-coded period of the video bitstream for optimized video coding.

[0028] The V3C specification ([2]ISO / IEC DIS23090-5(2E)Text of Coding of Visual Volume Video (V3C) and Video-Based Point Cloud Compression Second Edition, ISO / IEC JTC 1 / SC 29 / WG 7, N00297), of which MIV is an extension, provides alternative prediction coding modes, namely "inter", "merge", or "skip", for the patch data within the atlas data sub-bitstream, which are not activated by MIV. However, this alternative patch coding mode can only reduce the bitrate of the atlas data sub-bitstream, which is negligible compared to the other video sub-bitstreams that make up the MIV bitstream.

[0029] Figure 6illustrates different solutions for encoding a video patch atlas in a more efficient way. In embodiments of these techniques, the patches are no longer assembled into a rectangular frame, but are organized within a set of rectangular sub-pictures of different sizes arranged in a one-dimensional (1D) vector layout. Patches presenting strong similarity or continuity are packed into the same sub-picture, and the size of the sub-picture is further dynamically adapted. Then inter-sub-picture prediction is allowed. In Figure 6 the example of Figure 5 a set of seventeen patches is obtained from the patch techniques described regarding Figure 6 . By comparing the similarity and continuity between the patches in this set (and in one embodiment, between this set of patches and successive sets of patches of the video sequence), the patches are distributed in different atlas blocks. In

[0030] the V3C format, each patch has a geometry component (depth map) containing information about the exact position of the 3D data in space, an occupancy component that informs the rendering system which samples in the 2D component are associated with data in the final 3D representation, and several attribute components providing additional attributes such as texture (color) or transparency. Additional information about the 3D-to-2D projection is also included in the bitstream to enable inverse reconstruction.

[0031] The format for encoding such volumetric video representations must get rid of the constraint of packing the patches representing a 3D scene at a given moment into a rectangular atlas frame of fixed dimensions. According to this principle, the patches are distributed and packed into several smaller rectangular frames of different sizes, which are referred to herein as "atlas blocks". Here, the term "block" has a different meaning compared to the classical image tiling concept. According to this principle, an atlas block is a patch atlas in a list of atlas blocks (also called a 1D vector). The atlas block is accessed by its index in the 1D vector.

[0032] In a first embodiment of this format, the same layout (i.e., patches are packed within the block) is used for one of the different atlas components (geometry, occupancy, and attributes (color, transparency, etc.)) of a given atlas block.

[0033] The patches are distributed between different blocks according to similarity and continuity criteria to maximize intra-frame correlation, thus enabling efficient video compression. Since this optimal distribution and packing of each block may vary depending on the atlas component (e.g., texture and depth patches may not exhibit the same spatial correlation), in a second embodiment, the patch layout of a given block may vary for each atlas component.

[0034] According to this principle, the term "atlas" defines a set of patches associated with the volume of a 3D space, not necessarily associated with placement on a rectangular frame: "Spectrum: A collection of 2D bounding boxes and their associated information, corresponding to a volume in 3D space in which volume data is rendered.” The definitions of "2D rectangular atlas" and "1D vector of block atlases" are introduced: “ Rectangular spectrum : An atlas whose 2D bounding box is placed on a rectangular frame.” “ 1D vector of the block spectrum : An atlas whose 2D bounding box is placed on a number of rectangular blocks organized in a 1D layout”.

[0035] The following example of the syntax of the format proposed by this principle is inspired by the V3C format for compatibility with the V3C format. It should be understood that other syntaxes can capture equivalent semantic features. In this example, there is a boolean value "vps_1d_tile_vector_atlas_flag". When equal to 1, the syntax elements signaling fixed values for the atlas frame width and height are meaningless and thus not signaled. When vps_1d_tile_vector_atlas_flag is equal to 1, the 1D vector replacing the blocks packed in the rectangular frame can be dynamically refreshed at the frame level and is thus specified in the atlas frame parameter set (AFPS).

[0036] Signaling is repeated at the atlas sequence parameter set (ASPS) level, where the syntax elements defining the atlas frame dimensions (asps_frame_width, asps_frame_height) are no longer meaningful when there is no rectangular atlas packing.

[0037] In one embodiment, an explicit flag asps_1d_tile_vector_atlas_flag is introduced in the ASPS syntax to distinguish between rectangular atlases and 1D vectors of block atlases. However, such a solution is not backward compatible. This means that bitstreams compatible with the current V3C standard cannot be decoded by a decoder implemented according to this principle.

[0038] In another embodiment, a backward-compatible format is proposed. This includes using the values 0 for asps_frame_width and asps_frame_height to signal a 1D vector of the atlas map layout. In fact, such zero-sized frames are meaningless for traditional rectangular maps. The V3C variable AspsFrameSize (set to be equal to asps_frame_height * asps_frame) associated with the ASPS syntax can be used for this purpose: Condition Spectrum type AspsFrameSize > 0 Rectangular spectrum AspsFrameSize = 0 1d vector of the block spectrum .

[0039] An alternative specification of the blocks belonging to the 1D vector of the atlas blocks is provided within the atlas_frame_tile_information() syntax structure of the atlas frame parameters (AFPS) of V3C. In this case, only the atlas width and height are signaled because the spatial position within the fixed rectangular atlas frame is no longer meaningful.

[0040] Volume video standards such as V3C provide a "frame packing" function that enables several atlas components (geometry, occupancy, attributes) to be combined (packed) in the same video frame to feed a single video bitstream to a unique video encoder / decoder. The packing_information(atlasID) syntax structure within the VPS specifies which components are packed in this unique video frame. According to such a standard, the packing_information(atlasID) syntax structure also specifies the spatial layout of the various packed block components within the frame. In the case of the "1D vector of the atlas blocks", as in this principle, the packing information bypasses this useless spatial position. The pin_regions_count_minus1 syntax element signals the total number of block components, i.e., the total number of video sub-pictures of various types (depth, occupancy, occupancy, attributes) that they map in the combined bitstream, and each block is accessed by its ID pin_region_tile_id[j][i].

[0041] Figure 2Shows the frame packing of two atlas components (texture atlas 71 and downsampled depth atlas 72) in a unique video frame. The two atlas components 71 and 72 are arranged in a rectangular frame. According to this principle, in the unique 1D vector of sub-pictures, texture atlas blocks 73a - 73d are interleaved with depth atlas blocks 74a - 74d. The interleaving is not restrictive, and any other arrangement in the 1D vector is possible. For example, in one variant, all texture atlas blocks can be listed first, followed by all depth atlas blocks.

[0042] The following table provides a possible syntax that embodies this organization:

[0043] Figure 3 Shows an example architecture of device 30, which can be configured to implement a method for encoding or decoding a volumetric scene from a data stream according to this principle. The encoder and / or decoder can implement this architecture. Alternatively, each circuit of the encoder and / or decoder can be a device according to the Figure 3 architecture, linked together, for example, via their bus 31 and / or via the I / O interface 36.

[0044] Device 30 includes the following elements linked together by data and address bus 31: - A microprocessor 32 (or CPU), which is, for example, a DSP (or digital signal processor); - A ROM (or read-only memory) 33; - A RAM (or random access memory) 34; - A storage interface 35; - An I / O interface 36 for receiving data from an application for transmission; and - A power supply, such as a battery.

[0045] According to one example, the power supply is external to the device. In each of the mentioned memories, the term "register" used in the specification can correspond to a region of small capacity (a few bits) or a very large region (e.g., the entire program or a large amount of received or decoded data). The ROM 33 includes at least programs and parameters. The ROM 33 can store algorithms and instructions to perform the techniques according to this principle. When powered on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0046] The RAM 34 includes in registers a program executed by the CPU 32 and uploaded after the device 30 is turned on, input data in the registers, intermediate data of different states of the method in the registers, and other variables in the registers for executing the method.

[0047] The implementations described herein can be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method or a device), the implementation of the features discussed can also be implemented in other forms (e.g., a program). The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. For example, the method can be implemented in an apparatus such as, for example, a processor, which refers to a processing device and generally includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device such as, for example, a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.

[0048] According to an example, the device 30 belongs to a set including the following: - A mobile device; - A communication device; - A gaming device; - A tablet computer (or tablet); - A laptop computer; - A still picture camera; - A video camera; - An encoding chip; - A server (e.g., a broadcast server, a video-on-demand server, or a network server).

[0049] Figure 4 An example showing an embodiment of the syntax of a stream when data is transmitted via a packet-based transport protocol. Figure 4 An example structure 4 of a volumetric video stream is shown. This structure exists in a container that organizes the stream into independent syntax elements. The structure can include a header part 41, which is a set of data common to each syntax element of the stream. For example, the header part includes some metadata about the syntax element, describing the nature and role of the syntax element. The header part can also include a part of the metadata described in the tables of this document. The structure includes a payload, which includes syntax elements 42 and at least one syntax element 43. The syntax element 42 includes data representing color and depth frames, which is formatted as a list of atlas blocks. The image may have been compressed according to a video compression method.

[0050] The syntax element 43 is part of the payload of the data stream and may include metadata on how the frames of the syntax element 42 are encoded, such as an indication of whether the volumetric scene is formatted as a list of atlas blocks, the number of atlas blocks in the list of atlas blocks, and for each atlas block of the list, the size of each patch packed in the atlas block and the position within the atlas block.

[0051] According to this principle, a transmission format is provided to effectively support a compact and flexible description of most 3D scenes with constant or only slowly evolving geometry and appearance. In addition, for entities that are static in the physical world of the 3D scene, encoding and decoding techniques are proposed when the camera rig moves and / or when the lighting conditions evolve over time.

[0052] When the camera moves, the static 3D scene portion is considered a moving portion in the reference frame of the camera rig. According to this principle, on the encoder side, the camera rig motion (pose parameters = position and orientation) is estimated and transmitted to the decoder. In doing so, the patches of the transmitted static scene portion are used at a later time on the decoder side with compensation for the camera motion.

[0053] In a 3D scene sequence (even a fully CGI 3D scene), the lighting conditions change very frequently. In such cases, due to changes in lighting or shadows, the geometry does not change, but the appearance does. According to this principle, the texture (i.e., color attributes) of the static patches is updated more frequently than the geometric attributes. In another embodiment, a compact representation of the texture changes in the form of parametric mathematical functions is encoded in the data stream and transmitted to the decoder.

[0054] The main content of the proposed solution is as follows: · At the encoder level, static or quasi-static 3D scene portions are identified and their patch-based description is separated from the description of the rest of the scene; · The static patches are clustered into a set of long-term persistent entities; · The static patch description is refreshed at the entity granularity only when the 3D scene development requires it; · The decoder maintains in memory a data structure in a memory (such as a database) of the decoded patches to render the static portion of the scene by updating the entities (erasing, rewriting, adding) when needed.

[0055] The implementations described herein can be implemented, for example, in a method or process, apparatus, computer program product, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method or device), the implementation of the features discussed can be implemented in other forms (e.g., a program). The apparatus can be implemented in suitable hardware, software, and firmware. For example, the method can be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices such as, for example, a smartphone, a tablet computer, a computer, a mobile phone, a portable / personal digital assistant (“PDA”), and other devices facilitating information communication between end users.

[0056] The implementations of the various processes and features described herein can be implemented in a variety of different devices or applications, particularly, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture information and / or depth information. Examples of such devices include an encoder, a decoder, a post-processor that processes the output from the decoder, a pre-processor that provides input to the encoder, a video encoder, a video decoder, a video codec, a network server, a set-top box, a laptop computer, a personal computer, a cellular phone, a PDA, and other communication devices. It should be clear that the device can be mobile and even installed in a moving vehicle.

[0057] Furthermore, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a processor-readable medium, such as, for example, an integrated circuit, a software carrier, or other storage devices, such as a hard disk, a compact disc (“CD”), an optical disc (such as, for example, a DVD, commonly referred to as a digital versatile disc or digital video disc), a random access memory (“RAM”), or a read-only memory (“ROM”). The instructions can form an application program tangibly embodied on the processor-readable medium. The instructions can be, for example, hardware, firmware, software, or a combination thereof. The instructions can be found, for example, in an operating system, a separate application, or a combination of both. Thus, the processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (e.g., a storage device) having instructions for executing the process. In addition, in addition to or instead of the instructions, the processor-readable medium can store data values generated by the implementation.

[0058] It will be apparent to those skilled in the art that implementations can generate a variety of signals that are formatted to carry information that can, for example, be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry as data rules for writing or reading the syntax of the described embodiments, or to carry actual syntax values written by the described embodiments as data. Such signals can be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different known wired or wireless links. The signal can be stored on a processor-readable medium.

[0059] Numerous time limits have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. In addition, those of ordinary skill in the art will understand that other structures and processes can replace those disclosed, and the resulting implementations will perform at least substantially the same functions in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, this application contemplates these and other implementations.

Claims

1. A method for encoding a volumetric scene, the method comprises: - obtaining a list of atlas blocks, where the atlas blocks package patches clustered according to similarity and continuity criteria; - generating metadata, including: · an indication of whether the volumetric scene is formatted as a list of atlas blocks; · the number of atlas blocks in the list of atlas blocks; · for each atlas block in the list: - size; - the position of each patch packed in the atlas block within the atlas block; and - encoding the list of atlas blocks and the generated metadata in a data stream.

2. The method according to claim 1, wherein the indication of whether the volumetric scene is formatted as a list of atlas blocks is the size of the main atlas set to zero.

3. The method according to claim 1 or 2, wherein the video component of the atlas block list interleaves the video components of the atlas blocks and the video components of the depth atlas blocks.

4. A device for encoding a volumetric scene and comprising a memory associated with a processor, the processor being configured to: - obtain a list of atlas blocks, where the atlas blocks package patches clustered according to similarity and continuity criteria; - generate metadata, including: · an indication of whether the volumetric scene is formatted as a list of atlas blocks; · the number of atlas blocks in the list of atlas blocks; · for each atlas block in the list: - size; - the position of each patch packed in the atlas block within the atlas block; and - encoding the list of atlas blocks and the generated metadata in a data stream.

5. The device according to claim 4, wherein the indication of whether the volumetric scene is formatted as a list of atlas blocks is the size of the main atlas set to zero.

6. The device according to claim 4 or 5, wherein the video component of the atlas block list interleaves the video components of the atlas blocks and the video components of the depth atlas blocks.

7. A method for decoding a volumetric scene from a data stream, the method comprises: - decoding metadata from the data stream, the metadata including: · an indication of whether the volumetric scene is formatted as a list of atlas blocks; · the number of atlas blocks in the list of atlas blocks; · for each atlas block in the list: - size; - the position of each patch packed in the atlas block within the atlas block; and - decoding the list of atlas blocks from the data stream according to the metadata.

8. The method according to claim 7, wherein the indication of whether the volumetric scene is formatted as a list of atlas blocks is the size of the main atlas set to zero.

9. The method according to claim 7 or 8, wherein the video component of the atlas block list interleaves the video components of the atlas blocks and the video components of the depth atlas blocks.

10. A device for decoding a volumetric scene from a data stream and comprising a memory associated with a processor, the processor being configured to: - decode metadata from the data stream, the metadata comprises: · an indication of whether the volumetric scene is formatted as a list of atlas blocks; · the number of atlas blocks in the list of atlas blocks; · For each atlas block of the list: - Size; - The position of each patch packed in the atlas block within the atlas block; And - A list of atlas blocks decoded from the data stream according to the metadata.

11. The apparatus according to claim 10, wherein the indication of whether the volume scene is formatted as a list of atlas blocks is the size of the main atlas set to zero.

12. The apparatus according to claim 10 or 11, wherein the video component interleaving property of the list of atlas blocks interleaves the video components of the atlas blocks and the video components of the depth atlas blocks.

13. A data stream representing a volume scene, and comprising: - Metadata, including: · An indication of whether the volume scene is formatted as a list of atlas blocks; · A plurality of atlas blocks in the list of atlas blocks; · For each atlas block of the list: - Size; - The position of each patch packed in the atlas block within the atlas block; and - A list of atlas blocks.

14. The data stream according to claim 13, wherein the indication of whether the volume scene is formatted as a list of atlas blocks is the size of the main atlas set to zero.

15. The data stream according to claim 13 or 14, wherein the video component interleaving property of the list of atlas blocks interleaves the video components of the atlas blocks and the video components of the depth atlas blocks.