Methods, devices, and computer program for encapsulation of encoded volumetric data as multiple spatial tracks

WO2026201814A1PCT designated stage Publication Date: 2026-10-01CANON KK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/057962
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-09-29
Filing Date
2026-03-20
Publication Date
2026-10-01

Smart Images

  • Figure EP2026057962_01102026_PF_FP_ABST
    Figure EP2026057962_01102026_PF_FP_ABST
Patent Text Reader

Abstract

According to a first aspect of the disclosure, there is provided a method of encapsulating a visual volumetric data bitstream into a media file, the method being implemented by a processing device, the visual volumetric data bitstream comprising a plurality of encoded sub-bitstreams corresponding to a plurality of components of the visual volumetric data, at least one component sub-bitstream enabling independent decoding, the method comprising: determining regions of the visual volumetric data associated with subparts of the at least one component sub-bitstream enabling independent decoding; generating, for each region, a dedicated track associating the subparts of the at least one component sub-bitstream related to this region of the visual volumetric data, each dedicated track being referenced by a base track; generating information associating each dedicated track to its associated region; generating a media file comprising the base track, each dedicated track and said information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS, DEVICES, AND COMPUTER PROGRAM FOR ENCAPSULATION OF ENCODED VOLUMETRIC DATA AS MUTLIPLE SPATIAL TRACKS

[0002] FIELD OF THE DISCLOSURE

[0003] The present disclosure relates to the technical field of encapsulation of encoded volumetric data, in particular of a volumetric data bitstream comprising encoded sub-bitstreams.

[0004] BACKGROUND OF THE DISCLOSURE

[0005] Part-5 of the international standard for Coded representation of immersive media (MPEG-I), called “Visual volumetric video-based coding (V3C) and video-based point cloud compression (V-PCC)” specifies a generic mechanism for visual volumetric video coding, i.e. visual volumetric video-based coding. This standard is commonly denoted ISO / IEC 23090-5 or MPEG-I Part 5. In short, MPEG-I Part-5 defines V3C bitstreams.

[0006] A V3C bitstream (i.e., a visual volumetric video-based coding bitstream) is a sequence of bits that forms the representation of coded volumetric frames and associated data forming one or more Coded V3C Sequences (CVSs). The generic mechanism may be used by applications targeting volumetric content, such as point clouds, immersive video with depth, mesh representations of visual volumetric frames, etc. MPEG-I Part-5 also comprises specific definitions dedicated to point clouds called V-PCC for Video-based Point Cloud Compression. A second part of MPEG-I, Part-12 (ISO / IEC 23090-12) is directed to volumetric media encoded as MPEG Immersive Video (MIV) that can also be described within the generic V3C bitstream structure.

[0007] Another part (Part-29, under definition) of the international standard for Coded representation of immersive media (MPEG-I), called “Video-based dynamic mesh coding (V-DMC)” specifies syntax, semantics, and decoding for video based dynamic mesh coding (V-DMC) methods. Furthermore, Part-29 specifies processes that may be needed for reconstruction of visual volumetric media and may also specify additional processes such as post decoding, pre-reconstruction, post reconstruction, and adaptation. In short, Part-29 defines V-DMC bitstreams. The syntax and semantics for the Part-29 are specified as an extension of Part-5.

[0008] While a V3C bitstream mixes several kinds of bitstreams or sub-bitstreams (corresponding to different components), for example atlas sub-bitstreams with videosub-bitstreams (for geometry, attributes and occupancy), V-DMC is defining additional types of bitstreams or sub-bitstreams, thus requiring modifications in the V3C bitstream description and additional parameters to provide additional items of information to media players and decoders. For the sake of illustration, V-DMC is considering two additional sub-bitstreams, respectively a base-mesh (or basemesh) sub-bitstream (or component) and a displacement sub-bitstream. The base-mesh sub-bitstream is a simplified low-resolution approximation of the original mesh, and the displacement sub-bitstream provides displacement vectors, to refine the base mesh to better fit the original mesh. Encoding of the displacement sub-bitstream is specified either using any video codec (such as HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding) for example) or arithmetic codec. For the base-mesh sub-bitstream, the V-DMC specification introduces a new bitstream format. This new bitstream format, as for video coding format such as HEVC or the atlas sub-bitstream used in V3C, also uses NAL (Network Abstraction Layer) units. High Level Syntax (HLS) structures or syntax structures such as base-mesh sequence parameter sets (BMSPS), base-mesh frame parameter sets (BMFPS), or a submesh layer raw byte sequence payload (RBSP) syntax structure representing the content of some of the NAL units of a V-DMC bitstream are also specified.

[0009] In addition to the coding standards, standard ISO / IEC 14496-12, defining the ISO Base Media File Format (ISOBMFF), allows creating files that encapsulate media data with metadata. Such encapsulated files may be called presentations.

[0010] An ISO Base media file is object-oriented and structured into “boxes” that are sequentially or hierarchically organized.

[0011] Boxes are data structures provided to describe the data in the files. Boxes (also denoted objects, atoms, structure-data, or data structures) are building blocks starting with a header which gives both size and a unique type identifier (typically a four-character code (32-bit), also noted FourCC or 4CC). Some boxes, called ‘FullBox’, also contain a version number and flags field. All data in a file (media data and metadata describing the media data) is contained in boxes. There is no other data within the file. File-level boxes are boxes that are not contained in other boxes.

[0012] Within a presentation, media data are generally represented as tracks. A track is a timed sequence of related samples, a sample corresponding to the data associated with a single time or period of time. ISOBMFF provides a sample description, mainly through sample entry. A sample entry is a box structure which defines and describes the format of some number of samples in a track. A sample entry usuallycontains a configuration box providing information for decoder initialization, sometimes called “decoder configuration” or “decoder initialisation”.

[0013] The ISO / IEC 23008-12 standard, defining the Image File Format (or High Efficiency Image File format - HEIF), is built on tools defined in ISO / IEC 14496-12 and allows creating files that encapsulate media data with metadata for single image, a collection of images and sequences of images. The media data are generally represented as entities that can be either tracks or items. An item is data that does not require timed processing, as opposed to sample data, and is described by the boxes contained in a MetaBox box. HEIF provides a description about an item in item property array. The item properties are stored in a box structure which may describes the format of one or more items in a the HEIF file. For example, the item properties may contain a configuration box or decoder configurations for the items.

[0014] Another derived specification of ISO / IEC 14496-12, known as ISO / IEC 23090-10, defines the carriage of timed visual volumetric video-based coding data defined either by MPEG-I part 5, part-12 or part 29. In particular, it describes single-track encapsulation and multi-track encapsulation modes and defines ISOBMFF structures applicable to the description of volumetric media. For example, a description dedicated to carriage of V3C encoded data defines structures such as V3CBitstreamSampleEntry (with different four-character codes (4CC) such as ‘v3e1’, ‘3eg’, etc.). This structure describes the samples of the track and information to setup or initialize corresponding decoder(s). For multi-track encapsulation modes, ISO / IEC 23090-10 also defines different kinds of tracks and their relationships. One multi-track encapsulation mode allows to organize tracks according to sub-bitstream types as “component tracks” and another multi-track encapsulation mode allows to further organize each or a subset of component tracks into multiple tracks, for example atlas tile tracks or submesh tracks. For video component tracks, they may be split into multiple tracks following specification in ISO / IEC 14496-15; for example HEVC tile tracks or VVC subpicture tracks.

[0015] As can be seen, there may be a proliferation of tracks in a media file encapsulating encoded volumetric data. While this offers flexibility, it also introduces some complexity and possibly overhead in the description of the encoded volumetric data.

[0016] Besides, spatial access of the visual volumetric data in a V3C bitstream can be advantageous. Spatial access refers to an ability to efficiently decode and retrieve only a specific region of the volumetric scene represented by the V3C bitstream, withoutprocessing the entire bitstream. This means that a spatial access allows retrieving the visual volumetric data of the bitstream associated with the specific region to be accessed.

[0017] Due to the proliferation of tracks in a media file encapsulating encoded volumetric data, it may be difficult to access the visual volumetric data associated to a specific region. Indeed, the visual volumetric data associated to a specific region can be dispatched in several tracks with no direct association between them (for example between an atlas tile track and its corresponding submesh track or attribute video track; where corresponding means, for example, encapsulating data corresponding to the atlas tile). Thus, a complex track inspection is required to identify the visual volumetric data associated to a specific region of the volumetric scene.

[0018] In such circumstances, there is a need to improve encapsulation, to simplify spatial access in a visual volumetric data bitstream.

[0019] The following terminology is extracted from ISO / IEC 23090-5 and ISO / I EC 23090-29 and defines several elements disclosed in the present disclosure.

[0020] An attribute is a scalar or vector property optionally associated with each point in a volumetric frame such as colour, reflectance, surface normal, transparency, material ID, etc. An attribute access unit is a collection of attribute maps and auxiliary attribute frames, if available, for a specific attribute that correspond to the same time instance. An attribute frame is a 2D rectangular array created through the aggregation of patches containing values of a specific attribute. An attribute map is an attribute frame containing attribute patch information projected at a particular depth indicated by the corresponding geometry map.

[0021] An atlas is a collection of 2D bounding boxes and their associated information placed onto a rectangular frame and corresponding to a volume at step 3D space on which volumetric data is rendered

[0022] An atlas bitstream is a sequence of bits that forms the representation of atlas frames and associated data forming one or more CASs (Coded Atlas Sequence). An atlas frame is a 2D rectangular array of atlas samples onto which patches are projected and of additional information related to the patches, corresponding to a volumetric frame. An atlas sample is a position on the rectangular frame onto which patches that are associated with an atlas are projected. An atlas sub-bitstream is an extracted subbitstream from the V3C bitstream containing a part of an atlas NAL bitstream

[0023] A basemesh is a mesh structure containing possibly a reduced number of Cartesian coordinates and their associated connectivity, along with scalar or vector property optionally associated with each coordinate and / or set of connectivity. Abasemesh bitstream is a bitstream conforming to a mesh specification that may represent a V3C component. A basemesh sub-bitstream is an extracted sub-bitstream from the V3C bitstream containing a part of a basemesh bitstream.

[0024] Displacement is a set of 3D vectors that are added to the vertices of the subdivided mesh to closely approximate the input mesh surface. Displacement access unit is a set of data units of displacement bitstream that are associated with each other according to a specified classification rule, are consecutive in decoding order, and pertaining to one particular output time. A displacement bitstream is a bitstream conforming to a mesh specification that may represent a V3C component of type displacement. Displacement data is a unit syntax structure containing displacement data information, such as a NAL unit in the context of ISO / IEC 23090-29. Displacement subbitstream is an extracted sub-bitstream from the V3C bitstream containing a part of a displacement bitstream.

[0025] A geometry is a set of Cartesian coordinates associated with a volumetric frame. A geometry access unit is a collection of geometry maps and auxiliary geometry frames, if present, corresponding to the same time instance. A geometry frame is a 2D array created through the aggregation of the geometry information associated with each patch. A geometry map is a geometry frame containing geometry patch information projected at a particular depth

[0026] An occupancy is values that indicate whether atlas samples correspond to associated samples at step 3D space. An occupancy access unit is set of video data units of occupancy bitstream that are associated with each other according to a specified classification rule, are consecutive in decoding order, and pertaining to one particular output time. An occupancy frame is a collection of occupancy values that constitute a 2D array and represents the entire occupancy information of a single atlas frame.

[0027] A patch is a rectangular region within an atlas associated with volumetric information.

[0028] A submesh is an independently decodable region of a basemesh. A visual volumetric video-based coding component V3C component is an atlas, a basemesh, an occupancy, a geometry, a displacement or an attribute of a particular type that is associated with a V3C volumetric content representation (this list being non limitative).

[0029] A visual volumetric video-based coding parameter set is a V3C parameter set VPS syntax structure containing syntax elements that may be referred to by syntaxelements found in the V3C unit header (e.g. the syntax element vuh_v3c_parameter_set_id).

[0030] SUMMARY OF THE DISCLOSURE

[0031] The present disclosure has been devised to address one or more of the foregoing concerns.

[0032] According to a first aspect of the disclosure, there is provided a method of encapsulating a visual volumetric data bitstream into a media file, the method being implemented by a processing device, the visual volumetric data bitstream comprising a plurality of sub-bitstreams corresponding to a plurality of components of the visual volumetric data, the method comprising:

[0033] - obtaining a subpart of at least one component sub-bitstream enabling independent decoding of a region of the at least one component sub-bitstream, the region of the at least one component sub-bitstream being associated with a region of the visual volumetric data;

[0034] - generating a dedicated track associated with the region of the visual volumetric data and comprising the obtained subpart of the at least one component subbitstream;

[0035] - generating a base track comprising data that is common to a plurality of the component sub-bitstreams of the visual volumetric data, the base track referencing the dedicated track; and

[0036] - generating a media file comprising the base track and the dedicated track. One would understand that a bitstream is an ordered series of bits that forms the coded representation of the data. Thus, a sub-bitstream is a portion of a bitstream.

[0037] Such method may allow simplifying a multi-track encapsulation of visual volumetric data for easier spatial access, by associating the visual volumetric data related to a region of the volumetric scene to a dedicated track associated to this region. Such method offers a better trade-off between spatial access to a region of the volumetric scene and description overhead.

[0038] In an embodiment, the method further comprises: generating, in the media file, information associating the dedicated track with the region of the visual volumetric data.

[0039] In an embodiment, a sub-bitstream of the at least one component subbitstream corresponds to an atlas sub-bitstream. In an embodiment, a subpart of the atlas sub-bitstream is at least one atlas tile.In an embodiment, a sub-bitstream of the at least one component subbitstream corresponds to a submesh sub-bitstream. A subpart of the submesh subbitstream is a submesh.

[0040] In an embodiment, a sub-bitstream of the at least one component subbitstream corresponds to an attribute sub-bitstream. A subpart of the attribute subbitstream is at least one attribute tile.

[0041] In an embodiment, each region of visual volumetric data covers an object or a region of interest of a volumetric scene represented by the visual volumetric data.

[0042] In an embodiment, the base track includes a visual volumetric parameter set indication of the decoders that are needed to process the bitstream.

[0043] In an embodiment, the base track is an atlas base track.

[0044] In an embodiment, the atlas base track includes an atlas parameter set. In an embodiment, the dedicated track associated to a region of the visual volumetric data is a track comprises -notably multiplexes- the subparts of the at least one component sub-bitstream related to this region of the visual volumetric data.

[0045] In an embodiment, such dedicated track allows embedding the visual volumetric data associated to a specific region of a volumetric scene in a single track. This simplify the spatial access to such visual volumetric data, by parsing reducing the number of tracks to be parsed.

[0046] In an embodiment, the dedicated track includes atlas tile data with submesh data and / or attribute tile data.

[0047] In an embodiment, the dedicated track contains information indicating that samples of the dedicated track correspond to a multiplex of data units of visual volumetric data comprising subparts of component sub-bitstreams from different components.

[0048] In an embodiment, when several dedicated tracks multiplexing subparts of at least two component sub-bitstreams of a visual volumetric data bitstream are generated, the subparts of the at least two component sub-bitstreams enabling independent decoding are time-aligned between the different dedicated tracks.

[0049] In an embodiment, each dedicated track is an atlas tile track including atlas tile data related to the region associated to this dedicated track;

[0050] wherein the method further comprises, for each region, generating a component track including subpart of the component sub-bitstream enabling independent decoding. Each dedicated track can include a track reference referencing each component track associated with the region associated with this dedicated track. In an embodiment, a component track associated with a region is a submesh track includingsubmesh data related to a region of the visual volumetric data; and / or wherein a component track associated with the region is an attribute tile track including attribute tile data related to the region of the visual volumetric data.

[0051] In an embodiment, the plurality of sub-bitstreams includes at least one component sub-bitstream that does not enable independent decoding of a region of this at least one component sub-bitstream, and the base track further comprises this at least one component sub-bitstream that does not enable independent decoding.

[0052] In an embodiment, the plurality of sub-bitstreams includes at least one component sub-bitstream that does not enable independent decoding of a region of this at least one component sub-bitstream; and

[0053] wherein the method also comprises;

[0054] generating, in the media file, at least one common track including the at least one component sub-bitstream that does not enable independent decoding, generating, in the media file, at least one track reference referencing the at least one common track, to associate the subparts of the at least one component sub-bitstream enabling independent decoding to the at least one component sub-bitstream that does not enable independent decoding.

[0055] In an embodiment, a single common track is generated, this single common track including the at least one component sub-bitstream that does not enable independent decoding.

[0056] In an embodiment, the track reference referencing the single common track is included in each dedicated track.

[0057] In an embodiment, the single common track includes basemesh parameter sets and / or displacement data and / or video parameter sets.

[0058] In an embodiment, when the bitstream includes several component subbitstreams that do not enable independent decoding, a common track is generated for each component sub-bitstream that do not enable independent decoding, each common track including a single component sub-bitstream that do not enable independent decoding,

[0059] In an embodiment, the track reference referencing a common track is included in a component track.

[0060] In an embodiment, a common track is a basemesh track including basemesh parameter sets, and / or wherein a common track is a displacement track including displacement data, and / or wherein a common track is a attribute track including video parameter sets.In an embodiment, determining the regions of the visual volumetric data includes retrieving partitioning information from the visual volumetric data bitstream, partitioning information indicating the partitioning of the visual volumetric data of the at least one component sub-bitstream enabling independent decoding.

[0061] In an embodiment, said information associating each dedicated track to its associated region is included in the base track. This allows retrieving the information without parsing the other tracks of the media file.

[0062] In an embodiment, the region is a dynamic region whose related subpart of the at least one component sub-bitstream is modified over the time, the method further comprising :

[0063] generating a dynamic region description associated with an identifier of the dynamic region.

[0064] In an embodiment, at least a part of the dynamic region description is contained in the base track.

[0065] In an embodiment, at least a part of the dynamic region description is contained in a timed metadata track, the media file further comprising the timed metadata track.

[0066] In an embodiment, at least a part of the dynamic region description is contained in the dedicated track.

[0067] In an embodiment, the dynamic region description is declared in a sample group description. In an embodiment, the dynamic region description is declared within a sample table box of the media file.

[0068] In an embodiment, the modification over the time of the dynamic region may comprises :

[0069] an addition of a subpart of at least one component sub-bitstream to the dynamic region ; or

[0070] a removal of a subpart of at least one component sub-bitstream from the dynamic region.

[0071] In an embodiment, the modification over the time of the dynamic region may comprises :

[0072] - a modification of the identifier of the dynamic region, and, if the dynamic region description comprises a description of a plurality of dynamic regions, each of the dynamic region description being associated with an identifier, a re-allocation of an identifier of another dynamic region to the dynamic region.In an embodiment, the modification over the time of the dynamic region may comprises :

[0073] - a modification of at least one parameter of the dynamic region.

[0074] In an embodiment, the at least one parameter is selected from at least one of:

[0075] a description of a bounding box of the dynamic region,

[0076] a description of a mapping of the identifier of the dynamic region to atlas tiles identifiers, and

[0077] a list of objects included in the dynamic region.

[0078] According to another aspect of the disclosure, there is provided a method of parsing a media file encapsulating a visual volumetric data bitstream, the method being implemented by a processing device, the visual volumetric data bitstream comprising a plurality of sub-bitstreams corresponding to a plurality of components of the visual volumetric data,

[0079] the visual volumetric data bitstream including a region of visual volumetric data associated with a subpart of at least one component sub-bitstream enabling independent decoding of a region of this at least one component sub-bitstream,

[0080] the media file including:

[0081] - a dedicated track associated with the region of the visual volumetric data and comprising the subpart of the at least one component sub-bitstream;

[0082] - a base track comprising data that is common to a plurality of the component sub-bitstreams of the visual volumetric data, the base track referencing the dedicated track; and

[0083] wherein the method comprises:

[0084] - parsing the base track to access the dedicated track associated with at least one region of the visual volumetric data to decode..

[0085] According to another aspect of the disclosure, there is provided a computer program product for a programmable apparatus, the computer program product comprising a sequence of instructions for implementing a method of encoding or a method of parsing as previously described, when loaded into and executed by the programmable apparatus.

[0086] According to another aspect of the disclosure, there is provided a computer-readable storage medium storing instructions of a computer program product as previously described.According to another aspect of the disclosure, there is provided a device for carrying out each of the steps of the methods described hereinabove.

[0087] There is provided a device for encapsulating a visual volumetric data bitstream into a media file according to the method described hereinabove.

[0088] According to another aspect of the disclosure, there is provided a device for parsing a media file encapsulating a visual volumetric data bitstream according to the method described hereinabove.

[0089] At least parts of the methods according to some embodiments of the disclosure may be computer implemented. Accordingly, some embodiments of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit", a "module", or a "system". Furthermore, some embodiments of the present disclosure may take the form of a computer program product embodied in any tangible medium of expression having computer usable program code embodied in the medium.

[0090] Since some embodiments of the present disclosure can be implemented in software, some embodiments of the present disclosure can be embodied as computer readable code for provision to a programmable apparatus on any suitable carrier medium. A tangible carrier medium may comprise a storage medium such as a floppy disk, a CD-ROM, a hard disk drive, a magnetic tape device or a solid state memory device, and the like. A transient carrier medium may include a signal such as an electrical signal, an electronic signal, an optical signal, an acoustic signal, a magnetic signal or an electromagnetic signal, e.g., a microwave or RF signal.

[0091] BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Other features and advantages of the disclosure will become apparent from the following description of non-limiting exemplary embodiments, with reference to the appended drawings, in which:

[0093] Figure 1 schematically illustrates encapsulating and parsing volumetric media presentations, according to some embodiments of the disclosure;

[0094] Figure 2a and 2b illustrate examples of steps of an encapsulation process according to some embodiments of the disclosure;

[0095] Figure 3 illustrates an example of steps of a parsing process according to some embodiments of the disclosure, in order to decode the content of a media file;Figures 4a, 4b, 4c, 4d and 4e illustrate different examples of a media file obtained after carrying out an encapsulation process according to some embodiments of the disclosure;

[0096] Figure 5 illustrates another example of a media file obtained after carrying out an encapsulation process according to some embodiments of the disclosure;

[0097] Figure 6 schematically illustrates a processing device configured to implement at least a part of one or more embodiments of the present disclosure.

[0098] DETAILED DESCRIPTION OF EMBODIMENTS OF THE DISCLOSURE

[0099] According to some embodiments of the disclosure, a solution is provided to simplify a multi-track encapsulation of visual volumetric data for easier spatial access, by associating the visual volumetric data related to a region of the volumetric scene to a dedicated track associated to this region. Such method offers a better trade-off between spatial access to a region of the volumetric scene and description overhead. In particular, ISO / IEC 23090-10 allows the sample entries of an atlas track to declare a collection of 3D regions, using the V3CSpatialRegionCollectionBox. This box provides a mapping between 3D regions and atlas tiles. Then, for region-based access, it makes sense to organize tracks according to these regions. The atlas can be split in atlas tiles, with one or more tiles per region, as indicated in the TileMapping structure of the V3CSpatialRegionCollectionBox. Along with each set of atlas tiles for a region, the corresponding data for the other components can be multiplexed in a single track as explained hereafter. ISO / IEC 23090-10 also supports dynamic spatial region information by associating a timed-metadata track with a sample entry type 'dyvm' to the V3C atlas track or the V3C bitstream track. This track contains either empty samples (representing a period of non-zero duration in which there is no updates to spatial region information); or one or more spatial region descriptions whose position and I or dimensions are being updated.

[0100] Figure 1 schematically illustrates a system for encapsulating and parsing volumetric media presentations, according to some embodiments of the disclosure.

[0101] As illustrated, a server 100 comprises an encapsulation module 105. The server 100 may be connected, via a network interface (not represented), to a communication network 110 to which is also connected, via a network interface (not represented), a client 115 comprising a parser (or de-encapsulation module) 120 or a storage device (not represented).According to the given example, server 100 processes media data 125, for example data representing a 3D sequence, for streaming or for storage. Media data 125 may correspond to a set of items of information representing positions of points in the space and different attributes (for example Texture attribute) to apply at each position. The positions of points may be provided either as video images representing geometry, occupancy and attributes for V-PCC or as a set of 3D points and attributes for V-DMC, the 3D points corresponding for example to vertices of polygons forming a 3D mesh. If media data 125 are not coded, in particular not compressed, they are called raw or uncompressed data. Media data 125 may be encoded using different compression methods, for example V-PCC or V-DMC (3D point cloud or mesh) methods. These compression methods use a V3C bitstream format, including several sub-bitstreams using video or dedicated codecs, for example as specified respectively in the ISO / IEC 23090-5 standard (for V-PCC) or in the ISO / IEC 23090-29 standard (for V-DMC). If media data 125 are encoded, they are called compressed bitstream data.

[0102] The encoding may be done within server 100, for example by compression module 140, or may be done remotely (i.e. , outside the server) in which case the media data 125 are provided in a compressed bitstream. Server 100 may encapsulate media data 125 into media file 135 or into segment files (containing one or more segments), as they are processed, for example for live recording or live transmission. The media file or the segment files comprise the encoded media data and an index or description of media data 125. This index or description comprises items of information allowing locating samples in time, determining their type (e.g. if the sample is a synchronization sample or sync sample or a sample entry indicating which codec was used to encode the data for this sample) and determining their size, (data) container, and offset into the container.

[0103] Media file 135 (or the generated segment files) may be stored in a local or remote storage device or may be transmitted to a client, for example to client 115.

[0104] Client 115 may be configured to process data received from communication network 110, for example to process media file 135, or to process encapsulated data read from a storage device. After the received or the read data have been parsed in parser 120 (also known as a de-encapsulation module or a reader, or even a player or a media player), the parsed data may be stored, displayed or output. According to the given example, the parser outputs the media data referenced 160, as encoded volumetric bitstream. Depending on the configuration and on the capabilities of the client 115, in particular whether decompression module 150 is present, media data 160 maycorrespond to decoded 3D media content 125, representing position of points in space and the different attributes.

[0105] It is observed that since an encoded volumetric bitstream may be composed of several sub-bitstreams using different encoders, decompression module 150 in client 115 requires to instantiate several types of encoders to process the data contained in the media file 135. This requires decompression module 150 to identify the number of decoders and their configurations to generate decoded 3D media content 160, corresponding to media data 125. An example of steps for decompressing all the subbitstreams is described by reference to Figure 3.

[0106] It is observed that server 100 and client 115 may be user devices but may also be network nodes acting on media files being transmitted or stored.

[0107] It is also noted that media file 135 received or read by client 115 may be communicated to parser 120 in different ways. In particular, encapsulation module 105 may generate media file 135 with a media description (e.g., a DASH MPD, i.e. a media presentation description (MPD) of the dynamic adaptive streaming over HTTP (DASH) protocol) and communicate (or stream) it directly to parser 120 upon receiving a request from client 115. Media file 135 may also be downloaded, at once or progressively, by client 115 and stored locally.

[0108] For the sake of illustration, media file 135 may encapsulate media data into boxes according to ISO Base Media File Format (ISOBMFF, ISO / IEC 14496-12) or Image File Format (HEIF, ISO / IEC 23008-12).

[0109] In such a case, media file 135 may correspond to one or several media files or segments (indicated by a FileTypeBox ‘ftyp’ or a SegmentTypeBox ‘styp’). According to ISOBMFF or HEIF, media file 135 may include two kinds of boxes, one or several “media data boxes” (e.g. ‘mdat’ or ‘imda’), containing the media data, and “metadata boxes” or “structure-data wrapper” (e.g. ‘moov’ or ‘moot’ or ‘meta’), containing metadata defining the position of the media data in the media data box(es) and temporal position (if any) of the media data. Different embodiments of encapsulation of compressed bitstream data are described by reference to Figures 6 and 7.

[0110] Figure 2a illustrates an example of steps of an encapsulation method according to some embodiments of the disclosure. For the sake of illustration, the steps may be carried out in server 100 in Figure 1, for example in encapsulation module 105.

[0111] As illustrated, a first step (step 200) is directed to obtaining compressed bitstream data to be encapsulated, for example by encapsulation module 105. The compressed bitstream data may be generated by compression module 140. In this case,the compression module to be used may be configured by a user through a graphical user interface or through a command line. Such a configuration may comprise setting the number of sample frames to encode, whether the different video sub-bitstreams use a single codec type and, if several video codecs are used, the set of video codecs and to which sub-bitstream each codec applies and also the dedicated codecs used for nonvideo sub-bitstreams.

[0112] For example, compressed bitstream data may be encoded according to the V3C sample stream format (Annex C of ISO / IEC 23090-5 specification). The V3C sample stream format defines V3C Units, each identified by a type (vuh_unit_type) defined in a V3C Unit Header field (also named v3c_unit_header). Each V3C unit contains encoded data corresponding to one sub-bitstream, the data being encoded using either a standard video codec (such as AVC (Advanced Video Coding), HEVC, VVC) or a dedicated codec. The dedicated codec may be identified by a specific four-character code registered by a registration authority (e.g., mp4 registration authority, also called mp4ra). The video codec in use may be configured to generate independent decodable parts of the video like motion-constrained HEVC tiles or VVC subpictures, or any bitstream structure allowing independent decoding and reconstruction of a region within an image.

[0113] Still for the sake of illustration, among the V3C units as specified in V-PCC, units with vuh_unit_type may be the following:

[0114] - V3C_AD (for Atlas Data), that contains atlas data, giving information on how to reconstruct the 3D encoded data from the other V3C units present in the V3C bitstream. Atlas data are encoded using NAL (Network Abstraction Layer) unit format specified in ISO / IEC 23090-5,

[0115] - V3C_GVD (for Geometry Video Data), that contains geometry data, giving 3D position information. Geometry data may be encoded using any type of video codec, - V3C_AVD (for Attribute Video Data) that contains attribute data of a single attribute type information (for example Texture). The attribute data may be encoded using any type of video codec,

[0116] - V3C_OVD (for Occupancy Video Data), that contains occupancy data, giving information on significant video area(s) described in geometry / attribute(s) subbitstreams. The occupancy data may be encoded using any type of video codec, and - V3C_PVD (for Packed Video Data), that, when present, contains packed data describing a set of geometry, attribute and occupancy encoded in different regions of a picture. The packed data may be encoded using any type of video codec.Among the V3C units as specified in V-DMC, units with vuh_unit_type may be the following:

[0117] - V3C_AD, that contains atlas encoded data as for V-PCC,

[0118] - V3C_BMD (for base-mesh data), that contains base-mesh data, that gives information on a simplified mesh of the encoded 3D content. Base mesh data may be encoded using one of the coding scheme specified in ISO / IEC 23090-29, annex H or I. For example, Annex H describes a NAL unit -based sub-bitstream format for a base mesh.

[0119] - V3C_GVD or V3C_ADD (for arithmetic coded displacement data), that contains data describing displacement information, applying to points described by the simplified base mesh. V3C_GVD data are encoded using any type of video codec. V3C_ADD data may be encoded using NAL (Network Abstraction Layer) unit format specified in ISO / IEC 23090-29, and

[0120] - V3C_AVD (for Attribute Video Data), that contains attribute data of a single attribute type information (for example Texture). The attribute data may be encoded using any type of video codec.

[0121] Among the V3C units as specified in V3C, some units, for example with vuh_unit_type equal to V3C_VPS (for V3C parameters set), provide parameters allowing determination of how many different decoders are needed to process the V3C bitstream. It may also make it possible to identify the codec type of each sub-bitstream.

[0122] These parameters may also include the number of atlas sub-bitstreams (parameter vps_atlas_count_minus1) and the presence of V3C units containing subbitstreams, with one presence flag per V3C unit type (parameters vps_occupancy_video_present_flag, vps_geometry_video_present_flag, vps_attribute_video_present_flag, vps_auxiliary_video_present_flag, vps_packed_video_present_flag and vpve_ac_displacement_present_flag) and additional parameters to determine the number of attributes (parameter ai_attribute_count) and geometry decoders per atlas (parameters vps_map_count_minus1 and vps_multiple_map_streams_present_flag). This may be useful when each component is encapsulated in its own track.

[0123] Turning back to Figure 2a and after having obtained the compressed bitstream data at step 200, the encapsulation module parses each sub-bitstream identified as present in the encoded data, at step 210, to determine which sub-bitstreams are encoded with spatial access at step 215. This step 215 may apply different spatial access policies:a strict spatial access, meaning that all components provide independently decodable parts allowing region-based decoding or independently decoding of regions;

[0124] a lax spatial access, meaning that at least two components provide independently decodable parts

[0125] a minimal spatial access, meaning that there is one component providing independently decodable parts

[0126] The encapsulation module may be configured to allow one or a combination of these modes when considering the step 215. These modes may correspond to encapsulation profiles or brands for V3C carriage. They may be indicated at top-level of the media file, for example in a ‘ftyp’ box, at least in the list of compatible brands.

[0127] A spatial access corresponds to the ability to decode and retrieve only a subpart of a sub-bitstream that may correspond to a specific region or to an object of a volumetric scene. A sub-bitstream is encoded with spatial access when the subbitstream includes data associated with specific spatial regions of the volumetric scene. For example, an atlas sub-bitstream is encoded with spatial access when atlas data is encoded using more than one tile in the atlas sub-bitstream. As another example, a video-coded component sub-bitstream is encoded with spatial access when it contains motion-constrained HEVC tiles or VVC subpictures that are independently decodable. As another example, a basemesh sub-bitstream is encoded with spatial access, when it contains one or more submeshes (a submesh being an independently decodable region of a basemesh).

[0128] As soon as there is at least one component sub-bitstream encoded with spatial access, the test 215 returns true and the encapsulation module starts encapsulation by creating a base track at step 220 (i.e. append a new ‘trak’ box in the media file). For example for V3C, it can be an atlas base track storing common information for all the coded tiles of the atlas, in particular the V3C VPS unit, possibly with atlas parameter sets. In another example, it can be also a spatial base track that would contain the information common to the spatial tracks for any of the sub-bitstreams. This case may be suitable for the strict spatial access mode The information common to the spatial tracks for any of the sub-bitstream can include parameter sets NAL units for the video sub-bitstreams and also parameter sets for the non-video sub-bitstreams such as the parameter sets for basemesh sub-bitstream or the atlas sub-bitstream.The base track may include information regarding a list of available regions that can be accessed. In particular, the base track can include a sample entry that can declare a collection of regions or objects with their identifiers.

[0129] Then, at step 225, the encapsulation module creates a shared (or common) track (i.e. append a new ‘trak’ box in the media file). This track is used to collect coded data units and also parameter sets or configuration information that is common or complementary to the components with spatial access. For example, in V-DMC encoded data, the displacement information may be stored in this shared (or common) track (as illustrated on Fig.4a). It may contain data for the sub-bitstreams that do not provide independently decodable parts (i.e. that do not allow decoding of part of a sample or frame but only the whole sample or frame). The step 225 may be optional when the encapsulation is configured to generate a minimum number of tracks: in this case, the base and the common track may consist in one track (as on Fig. 4b). There may be several common or shared tracks, as illustrated on Figure 4c, especially for the minimum spatial access mode.

[0130] Then at step 235, the encapsulation module determines the granularity of spatial access (i.e. a number of spatial parts of the volumetric data that can be accessed). Step 235 may be optional. For example, when an encapsulation module is configured to consider as spatial part one atlas tile, there is no need to determine the granularity of the spatial access. Indeed, in this case, by default, the number of spatial parts corresponds to the number of atlas tiles (that can be obtained from some parameter sets for the atlas sub-bitstream).

[0131] The determination of the granularity of the spatial access can include, for example, obtaining atlas tile information from the bitstream (tile position and sizes) with their corresponding tiles (if any) in video-coded components. The encapsulation module may use, for example, the attribute tile information in the bitstream, possibly in combination with some SEI (Supplemental Enhancement Information) messages exposing the tiles or subpictures in use for the video-coded component to determine the granularity of the spatial access. There may be also some SEI messages providing information about object or region of interest onto which the spatial access may be aligned. The spatial access may also consist in an arbitrary, or regular, partitioning of the 3D scene corresponding to one or more atlas tiles. An atlas tile may then be part of more than one region defined in step 235. This may be useful to produce overlapping regions for viewport-based streaming, especially to provide smooth transitions between two viewports (using the overlap region and clipping of useless other region at display). Theencapsulation module may set some indication of the spatial access mode in the media file, for example as a parameter of the atlas base track or of the spatial base track. It can be an identifier or a code from a predetermined list of pre-defined granularity levels. For example:

[0132] “single tile-based” access when the spatial access matches one atlas tile; “multi-tile-based” access when the spatial access corresponds to a combination of tiles;

[0133] “region-based” access when it corresponds to an identified region of interest;

[0134] “object-based” access when it corresponds to an object of interest. The actual description of these different access may be provided per spatial track as explained in reference to Figure 4, or it may be provided in the base track itself. The encapsulation module may also put an indication on whether the spatial tracks provide overlapping regions or not, for example as a parameter of the atlas base track or of the spatial base track. This parameter may be in the track description or in the sample description, for example close to the description of the regions when present. For example, a vscspatiaiRegioncoiiectionBox may be present in the sample entry of the base track.

[0135] Once the number of spatial parts (i.e. regions) is determined, the encapsulation module creates, for each of these parts (iterating onto steps 240, 245 and 250), a dedicated track. In this embodiment, each dedicated track created at step 240 contains independently decodable data units corresponding to the spatial part for the components for which a spatial access has been determined at step 215. Such dedicated track is named spatial track herein. It is up to the encapsulation module to dispatch regions, tiles or objects or group of regions, tiles or objects within a spatial track.

[0136] Then, at step 245, the spatial tracks and the base track are associated. For example, the spatial tracks can be associated to a base track via a track reference from the base track to the spatial tracks, the type of the reference being identified by a specific 4CC. As well, each spatial track may be associated to the shared track. For example, each spatial track can be associated to the shared track via a track reference of a dedicated type, identified as a 4CC (one track reference type 'v3cp’ on Fig. 4a or different track reference types ‘v3va’, ‘vdmd’ on Fig.4c). It is to be noted that, in a variant, the content of the shared track can be described in the base track, to limit the number of tracks and track references.Then, a test 250 is performed to determine, based on the number of spatial parts determined at step 235, whether there is another spatial part for which a spatial track needs to be created. If there is another spatial part for which a spatial track needs to be created, then the steps 240, 245 and 250 are iterated.

[0137] When all the tracks are initialized (i.e. when the corresponding ‘trak’ boxes are created in the media file), the bitstream is read on V3C unit basis (as part of step 240).

[0138] When a V3C unit is pertaining to a sub-bitstream providing spatial access, the atlas tile corresponding to such spatial access is determined from SEI messages or high-level syntax indication like submesh to tile mapping for mesh-based geometry encoded with independent submeshes or like attribute tile information in the high-level syntax with an SEI exposing the video tiles or subpictures or by locating the SPS of the video sub-bitstream and selecting coded data units for tiles or subpictures corresponding to the atlas tiles for the current spatial part.

[0139] For example, to select an atlas tile data to append in a spatial track containing an atlas tile with ID=x, V3C units with a type indicating atlas data (e.g. V3C_AD) are parsed to obtain the atlas NAL units having in their header a value of athjd equal to x.

[0140] Similarly, to select geometry data to append in a spatial track containing an atlas tile with ID=x, V3C units with a type indicating geometry data (e.g. V3C_GVD or V3C_BMD) are inspected to obtain the geometry NAL units corresponding to the tile with ID=x. When geometry is video-coded, the selection is performed according to the selection of video data hereafter. When the geometry is encoded with a non-video codec, the data units of the geometry are inspected to obtain the ones having a submeshjd associated to the tile with ID=x.

[0141] The association between atlas tiles and geometry submeshes may be obtained by inspecting the meshpatch data units or simply by parsing the SEI message called tile_submesh_mapping when present in the bitstream. This SEI message specifies the mapping between the atlas tiles signaled in the Atlas Frame Parameter Set and the submeshes of the basemesh sub-bitstream. It is to be noted that this SEI may also provide a mapping between atlas tiles and video codec tiles. At last, to select video data to append in a spatial track containing an atlas tile with ID=x, V3C units with a type indicating video data (e.g. V3C_AVD or other V3C types for video-coded components), two steps are performed. First, the region in the video corresponding to the atlas tile with ID=x is obtained, let’s call it Rv(x).Then, the video data units are inspected to obtain the NAL units with an indication, for example in their slice header, that the coded data for the NAL unit correspond to coded blocks within the positions of Rv(x). This may depend on the actual video codec in use. With HEVC, the HEVC Temporal motion-constrained tile sets SEI message provides the position of MCTS and the tiles inside this set, the grid of tiles being declared in PPS. To get tile I D without MCTS encapsulation module can inspect the slice header to obtain the slice_segment_address.

[0142] When VVC is used as video codec, the VVC Sequence Parameter Set defines a grid of subpictures with their identifiers and slices indicates in their header a subpicturejd. It is to be noted that in V-DMC the first step may also be performed using the Atlas frame tile attribute information encoded in the atlas Frame Parameter Set. It provides for the video-coded attribute sub-bitstreams the mapping of atlas tiles into regions of the coded video (that we generically call tiles, as a 2D rectangular region of an image). The regions within the video may be indicated by an SEI or be obtained by parsing the parameter sets of the corresponding attribute sub-bitstream.

[0143] The data units for sub-bitstream without spatial access (determined at step 215) are stored in another track, named common or shared track herein. In a variant, these data units may be stored within the base track.

[0144] Information regarding whether the media file provides spatial access can be generated in the media file. Such information can be indicated at top-level of the media file by a brand indication in the ‘ftyp’ box or in the sample entries of the tracks.

[0145] When the test 215 returns false (i.e. when a sub-bitstream is not encoded with spatial access), the encapsulation module checks at step 216 whether there are independent atlas tiles in the encoded bitstream. When the minimum access mode is active, this step 216 may not be performed and the atlas tiles would be part of a spatial track. The step 216 may be performed when the spatial access check in step 215 expects a strict spatial access or a lax spatial access mode.

[0146] If there are no independent tiles, the encapsulation module generates a multi-track encapsulation as component tracks, the atlas track referencing the other component tracks at step 255.

[0147] If there are independent tiles, the encapsulation module creates a base track at step 260 and create component tracks based on the presence flags in the V3C_VPS (e.g. vps_geometry_video_present_flag or vps_attribute_video_present_flag syntax elements) in step 265, one component track per sub-bitstream (e.g. possibly severalvideo tracks) or one component track per sub-bitstream type, for example packing the video-coded attributes in one video track.

[0148] Then, similarly to step 235, at step 270, the encapsulation module determines the number of atlas parts to consider as spatial access granularity. There may be different modes as explained for step 235. This may also consider some complexity, data size criterion, to merge some atlas tiles into an atlas (spatial) part.

[0149] Then, at step 275, the atlas tile tracks are created, one per atlas part determined at step 270. In other words, an atlas tile track may contain more than one tile and their corresponding data units. In this case, the track description provides the list of tiles contained in the track, for example in a sample entry.

[0150] Finally at step 280, each atlas tile track is associated to the atlas base track, for example via the standard-defined track reference type ‘v3ct’. Each tile track is also associated to the component tracks, for example via a dedicated track reference type ‘vttc’ (V3C tile track to component track), so that the data can be pulled from the atlas base track, then from one or more tile tracks, then from the component tracks.

[0151] With these created tracks, the encapsulation module reads the V3C units from the obtained bitstream at step 200 and according to the V3C unit type dispatches the NAL units contained in the V3C unit in the appropriate track; i.e. atlas data in one of the atlas tile track depending on their atlas tile id; geometry data in a geometry track if any, attribute data of a given attribute in the corresponding component track for attribute as indicated in a V3C unit header box (‘vunt’) in the track description.

[0152] Once generated, after one of 255, 280 or 250 steps, the media file is available for storage, reading or streaming to an application using it.

[0153] Figure 2b illustrates with an example detailed embodiment of the encapsulation method according to Figure 2a that highlight the different types of spatial tracks that may be created.

[0154] At step 200, the compressed bitstream data to be encapsulated is obtained; At step 216, the encapsulation module checks at step 216 whether there are independent atlas tiles in the encoded bitstream.

[0155] If there are no independent tiles, the encapsulation module generates a single track or a multi-component tracks depending on whether the atlas tiles are split in components.

[0156] If there are independent tiles, the encapsulation module generates spatial tracks from step 2010 or step 2020.Step 2010 is an optional step and can be absent in the encapsulation method. The optional step 2010 corresponds to the step 235 on Figure 2a. In particular, at step 2010, the encapsulation module determines the granularity of spatial access (i.e. a number of spatial parts of the volumetric data that can be accessed). In particular, the encapsulation module can group tiles per region. When step 2010 is not performed or not absent, the step 2010 is replaced by considering that a region corresponds to a tile.

[0157] Then, the encapsulation method includes tests conducted by the encapsulation module as steps 2020, 2030 and 2040. A test 2020 checks whether the geometry sub-bitstream provides independently decodable parts or not. A test 2030 checks whether the bitstream read at step 200 contains attribute sub-bitstreams with independently decodable parts. Another test 2040 checks whether the geometry (test 2020 true) and some attribute sub-bitstreams have independently decodable parts. Depending on the results of these tests, the created spatial track may contain:

[0158] (step 2035) Atlas data and video-coded data for a part of a 3D scene determined at step 2010, or

[0159] (step 2041) Atlas data and geometry data for a part of a 3D scene determined at step 2010, or

[0160] (step 2045) Atlas data and geometry data and video-coded data for a part of a 3D scene determined at step 2010

[0161] The remaining data (that does not include independently decodable data units corresponding to a region defined at step 2010) are stored in another track (step 2042 or step 2036), named common or shared track herein. This common or shared track is associated to the created spatial tracks at step 2070. There are as many spatial tracks as regions defined at step 2010. There may be a single (common or shared) track collecting data units without independently decodable data units. In a variant, there may be as many (shared or common) tracks collecting data units without independently decodable data units as component type.

[0162] When the test 2030 returns false, the following steps correspond to standard encapsulation as single track or multi component tracks as defined in the working draft for lSO / IEC 23090-10.

[0163] Figure 3 illustrates an example of steps of a parsing process according to some embodiments of the disclosure, in order to decode the content of a media file, for example to decode the content of media file 135, by parser 120, and decompression module 150 in Figure 1.As illustrated, a first step is directed to receiving a media file (step 300), for example media file 135, or at least the portion of the media file comprising an initialization segment, making it possible for the parser to read and interpret metadata, for example the MovieBox and the different TrackBoxes. From, these metadata a number of tracks and their nature and their relationships are determined, by counting the number of TrackBoxes and by inspecting their TrackReferenceBox(es), when present.

[0164] The parser checks whether the media file provides spatial access into the encoded bitstream at step 310. This can be determined either at top-level of the file by a brand indication in the ‘ftyp’ box or by inspecting the sample entries of the tracks. As soon as one base track or spatial track can be found, it is deduced that spatial access is available. If not, the media file is processed at step 390 as currently specified in working draft of ISO / IEC 23090 2ndedition. If the test 310 is true, the media file is processed according to the invention.

[0165] At step 315 a user, or an application and at the end the parser module 120 determines from the base track the available list of regions or objects for which spatial access is provided. The base track can be determined for example as the first track declared in the file or the one indicated as a base track or with track_in_movie set to 1 in its track header box (‘tkhd’). The list of available regions may be determined from the base track and in particular from its sample entry. This sample entry may declare a collection of regions or objects with their identifiers, as explained in reference to Figure 4 below.

[0166] At step 315, the parser selects one or a combination of regions or the complete scene to decode according to a requested access. The combination of regions or the complete scene depends on the configuration of the base track. Indeed, the base track may describe the whole scene as a set of spatial tracks. The base track may also describe some regions (a subset) of the scene, not necessarily the whole scene. There may also be cases where a media file contains several base tracks, these base tracks offering alternatives, for example for different viewports, or for different sets of regions. These base tracks may contain viewport information in their description or be associated to a metadata track providing this viewport information. Such base tracks may be useful for viewport-based streaming application, for example in head-mounted display or glasses devices. Taking as example one of the figures 4a, 4b or 4c, there may be an additional base track referencing a set of spatial tracks that is different than the set of spatial tracks referenced by the base track 410, 4100 or 4600, thus offering alternative access and reconstruction into the 3D scene 401.Then, from the selection from step 315, the corresponding spatial tracks is identified at step 320. This can be done from the track description, for example from the track reference from base track to spatial tracks or from the track description of spatial tracks or a combination of these information. For example, the sample entry of a spatial track may provide information about the region covered by or corresponding to the spatial track at step 325, as explained in reference to Figure 4 later on. From the determined selection, the parser decides whether rewriting is needed. This may be the case when not playing the whole scene or when playing a combination of spatial tracks described in a base track.

[0167] If rewriting is needed (test 330 is true) the parser obtains rewriting instructions from the selected spatial tracks at step 335 and apply the rewriting instructions at step 340. The rewriting instructions may consist in updating parameter sets per sub-bitstream and possibly the V3C VPS to indicate the appropriate number of atlas tiles and number of sub-bitstreams. Especially, when the component bitstreams declare some identifiers for subpictures or tiles or submeshes, the encapsulation needs to make sure, in each sub-bitstream, that these identifiers will not conflict with each other in a reconstructed bitstream from the base mesh. To avoid conflicts, rewriting instructions may provide an identifier of the parameter set to rewrite, possibly with new identifier values to apply.

[0168] Then decoders per sub-bitstream are setup at step 345, according to decoder configuration information obtained from the main or base track or from a common or shared track referenced by the spatial tracks or a combination of both. Typically, the V3C_VPS is in the atlas base track, while parameter sets for the different sub-bitstreams are in the common or shared (e.g. 410 on Figure 4) track referenced by the spatial tracks.

[0169] Once the decoders are initialized, the parser can send data units to appropriate decoders by demultiplexing at step 350 the data in each spatial track and by collecting the data without independently decodable data units from the common track. Data units when obtained as V3C units are parsed to extract one or a sequence of NAL units corresponding to a sub-bitstream, at step 355, these sequences of NAL units being decoded at step 360 to reconstruct the region(s) of the 3D scene selected at step 315.

[0170] Besides, a parser can be configured to select a single spatial track (step 315, 320) to reconstruct a single region. In particular, the single region that can be reconstructed is the region that can be conveyed by the selected spatial track. In particular, A description of this region may be present in the sample entry of the spatialtrack, for example as a V3CRegionCollectionBox, or as a list of one or more atlas tile indexes.

[0171] More particularly, the parser can be configured to obtain parameter sets from the base track. The parser can also be configured to rewrite the parameter sets included in the base track, by implementing rewriting instructions from the spatial track, if present. When rewriting instructions are not present in the media file, no rewriting of the parameter sets is needed. Only the parameter sets (SPS, FPS or PPS) of the components contained in the spatial track may be rewritten.

[0172] In particular, the parser can be configured to reconstruct a bitstream by building a sequence of V3C units to provide to a decoder either as one V3C bitstream or alternatively as several component sub-bitstreams, each as a sequence of NAL units provided to the appropriate decoder (a mesh decoder for basemesh in V-DMC, a video decoder for the texture, an arithmetic decoder for some arithmetic coded displacement ...).

[0173] Then, the parser can be configured to obtain data units from the spatial track sample by sample and to concatenate these data units in the reconstructed bitstream. Then, the parser can be configured to obtain the data units from shared tracks if present, or from base track if any. These data units are obtained sample by sample, and are concatenated in the reconstructed bitstream. For the reconstruction of component subbitstreams, the parser is configured to demultiplex the data units according to their type and for video coded sub-bitstreams according to the attribute index, to obtain NAL units of different types. These NAL units are then concatenated in their respective subbitstreams.

[0174] Besides, a media player can be configured to select multiple spatial tracks to reconstruct several regions of a 3D scene. Notably, the media player can be configured to select all the spatial tracks to the reconstruct the whole 3D scene.

[0175] In particular, the parser can be configured to follow the track reference from the base track to reconstruct a V3C bitstream or a set of component sub-bitstreams, depending on the application processing the data.

[0176] More particularly, the parser can be configured to put the parameter sets from the base track in the reconstructed bitstream or to dispatch them according to their type into one the component sub-bitstreams.

[0177] Then, the parser can be configured to, following the track reference order between base track and spatial tracks, obtain the data for each sample and to concatenate these data in the reconstructed bitstream or dispatch them (the NAL unitswithin the read V3C units) according to their type into one the component sub-bitstreams. Then, if the media file includes shared tracks or if the base track of the media file contains data, the parser can be configured to append them in the same way in the one bitstream or multiple sub-bitstreams. Optionally, when a V3C bitstream contains a zippering SEI message, the content of the message can be used for the reconstruction of the mesh, specifically at the border between regions of the spatial tracks.

[0178] Figure 4a illustrates a first example of a media file 400, for example corresponding to media file 135 in Figure 1, obtained after carrying out an encapsulation process according to some embodiments of the disclosure, for example the one described by reference to Figure 2a or the one described by reference to Figure 2b (when atlas, video and geometry have independently decodable data for a region). The media file 400 encapsulates volumetric data corresponding to a 3D scene 401 into multiple tracks. This corresponds to the steps executed when the test 215 on Figure 2a is true, (or tests 216, 2020 and 2040 are true on Figure 2b).

[0179] On this example, a 3D scene is recorded by one or more cameras as 2 objects of interest 401-1 and 401-2. The 3D regions 401-1 and 401-2 may also correspond to arbitrary partitions or split of the 3D scene 401. These 3D regions may also be called 3D tiles or 3D bounding boxes.

[0180] Figure 4a represents an ISOBMFF encapsulation based on spatial tracks 430 and 440 where a first spatial track 430 corresponds to the description and data for the 3D object or 3D region 401-1 and where the spatial track 440 corresponds to the description and data for the 3D object or 3D region 401-2. The spatial tracks 430 and 440 multiplex data units for three components of a volumetric bitstream, in this example (a spatial track can comprise data units of at least one component of a volumetric bitstream).

[0181] For example, when the volumetric data is encoded with V-DMC, a spatial atlas track can contain multiplexed V3C units for atlas data or for geometry data and optionally for attribute data.

[0182] For example, when the geometry consists in mesh data with independent submeshes, the submeshes in a spatial track match the atlas tiles present in the atlas data for this same track. This avoids having to describe a tile to submesh mapping in any track. It would instead be handled by track references.

[0183] As another example, a spatial track may multiplex data units for atlas tile(s) and for independently decodable video data that correspond to these atlas tile(s). Independently decodable video data can be for example motion-constrained HEVC tileor independent VVC subpicture. This data multiplexing (or interleaving) is illustrated by the example of samples 430-1 and 430-2 for the spatial track 430 and similarly by example of samples 440-1 and 440-2 for spatial track 440.

[0184] As the sample stores V3C units, the order between the different data types may not be constrained. However, there may be advantages to repeat a same interleaving pattern. To distinguish the two organisations (constrained versus not constrained), distinct sample entry types may be defined. When constrained, the atlas data is preferably the first part of a sample.

[0185] While not illustrated here, each spatial track also contains a track description and sample description, in particular a specific sample entry indicating that the samples of this track correspond to a multiplex of data units from different components, these data units corresponding to a spatial part of the different components. A spatial track is identified by a specific sample entry type, through a reserved 4CC code.

[0186] A spatial track may further contain in its sample entries one or more boxes, notably a “V3CUnitHeader box” (a container for a V3C unit header), to indicate what kind of component are actually multiplexed in this spatial track. For example, the tracks 430 or 440 could have three “V3CUnitHeader” boxes, one indicating that atlas units are present, another one indicating that geometry data is present and a third one indicating that video data is present. There could be several instances of this third one, if there are multiple video-coded components. When multiple occurrences, they could be distinguished by indicating an attribute index in the “V3CUnitHeader box”.

[0187] Sample entries for a spatial track may be specified as follows, for example in V3C Carriage specification (ISO / IEC 23090-10):

[0188] Definition

[0189] Sample Entry Type: 'vssT

[0190] Container: SampleDescriptionBox

[0191] Mandatory: Yes, in spatial tracks

[0192] Quantity: One or more

[0193] A spatial track 430 or 440 can contain data units corresponding to one or more 3D regions that are independently decodable. A sample of a spatial track sample 430-1, 430-2 or 440-1, 440-2 contains at least one V3C unit. It may be atlas V3C unit or geometry V3C unit. It may also contain V3C units for one or more attributes or displacement V3C units ora combination of these V3C units. All the data units of a spatial track belong to a same set of atlas tiles. The V3C units are preferably organised according to their increasing time, and for a given time multiplexing the V3C units for thedifferent components. The set of atlas tiles may be described in the sample entry of the spatial track. The parent base track 410 is indicated by a track reference of type 'v3bs' from the base track to the spatial track. A track reference type 'v3cp' may be used to indicate the relation from the spatial track to the associated common or base track 420, when present. A spatial track may contain rewriting indications for its parameter sets of the different sub-bitstreams it contains, when it is played alone (not in combination with other spatial tracks). A spatial track may have a track reference to its base track through a track reference type ‘v3sb’ (from spatial to base track). This may be useful when the parameter sets are stored in the base track or when the base track also stores the components without independently decodable data. For example, for V-DMC, V-PCC and MIV, the following statements may apply:

[0194] each sample should comprise at least one atlas V3C unit; all the data units of a spatial track belong to a same set of atlas tiles. The set of atlas tiles may be described in the sample entry of the spatial track. An example Syntax for a sample entry of a spatial track can de defined as follows:

[0195] aligned(8) class V3CSpatialSampleEntry (the name here is just an example) extends VolumetricVisualSampleEntry('vssT) { V3CSpatialConfigurationBox config; / / optional

[0196] }

[0197] class V3CSpatialConfigurationBox extends FullBox('vssC, version = 0, 0) { V3CSpatialRegionCollectionBox regions;

[0198] }

[0199] With the following Semantics

[0200] “compressorname” in the base class “VolumetricVisualSampleEntry” indicates the name of the compressor used with the value "\012V3C Coding" being recommended; the first byte is a count of the remaining bytes, here represented by \012, which (being octal 12) is 10 (decimal), the number of bytes in the rest of the string.

[0201] “config” contains a single instance of “V3CSpatialConfigurationBox” as defined in subclause 6.2.2. When present, it contains at least a “V3CSpatialConfigurationBox” that describes the set of atlas tiles contained in the spatial track, for example as a “V3CSpatialRegionCollectionBox”, as the example syntax above. The regions described within this “V3CSpatialRegionCollectionBox” should be a subset of a “V3CSpatialRegionCollectionBox” in the parent base track (for example atlas base track or spatial base track, whatever its name).The format of a V3C spatial track sample may be defined as follows, with as syntax:

[0202] aligned(8) class V3CSpatialSample {

[0203] / / sample_size size of sample from SampleSizeBox

[0204] for (int i=0; i < sample_size; ) {

[0205] unsigned int(v3c_config.unit_size_precision_bytes_minus1 + 1)*8) v3c_unit_size;

[0206] bit(8) ss_v3c_unit[v3c_unit_size];

[0207] i += v3c_unit_size + v3c_config.unit_size_precision_bytes_minus1 + 1; }

[0208] }

[0209] With the following semantics

[0210] v3c_unit_size specifies the size, in bytes, of the ss_v3c_unit array. The size is equivalent to the sample stream v3c unit size ssvu_v3c_unit_size as defined in ISO / IEC 23090-5, Annex C.

[0211] ss_v3c_unit contains a single V3C unit in V3C unit sample stream format as defined in ISO / IEC 23090-5:2021 in Annex C.

[0212] In a more constrained variant of spatial track, the sample may correspond to one region of the atlas track. In this case, the “V3CSpatialConfigurationBox” may then contain a “V3CSpatialRegion” instead of a “V3CSpatialRegionCollectionBox”. The regionjd for this “V3CSpatialRegion” should correspond to one region declared in the base track. For this constrained variant, the “V3C spatial region” may even be omitted when the track references from base track to the spatial tracks follows the same order as the declaration of spatial regions in the base track 410. The “regionjd” is then deduced from the index of the track reference of type ‘v3bs’. For example, the track reference from atlas base track 410 to spatial track 430 would correspond to region 1 (401-1) and the track reference from atlas base track 410 to spatial track 440 would correspond to region 2 (401-2). Alternatively, to avoid some repetition with information from the atlas base track 410, a “V3CSpatialConfigurationBox” may consist in a list of region identifiers, the identifiers corresponding to those declared in base track. As another variant, the “V3CSpatialConfigurationBox” may consist in the “V3CAtlasTileConfigurationBox”, providing the list of atlas tiles through their identifier.

[0213] When there are several spatial tracks in a media file, they are considered time-aligned to allow reconstruction of a subset of spatial tracks (for example reconstruction of at least one or several spatial tracks) or of all the spatial tracks fromone atlas base track. However, some spatial tracks may contain empty sample for a given output time, but for any output time corresponding to an encoded volumetric frame, there shall be at least one non-empty sample in at least one of the spatial tracks representing the 3D scene. Optionally, the spatial tracks representing a same 3D scene may be recorded in a PlayoutTrackGroupBox with their base track and common base track, if present. As well, if different encoding versions are available in the encoded volumetric media for a same 3D region, each encoded version may be encapsulated as a spatial track and the so-encapsulated spatial tracks may be recorded in a same alternate group, for example in their track header or through an ‘alte’ track group.

[0214] For seeking or random accessing in spatial tracks, each sync sample of a spatial track may be indicated in a “SyncSampleBox”. Preferably the sync samples are aligned across spatial tracks and aligned with the sync samples declared in the base track. A sync sample in a V3C base track can be a sample that provides random access point for the spatial tracks it references. For some application where this alignment would not be suitable (for example to smooth the bitrate when retrieving and combining several spatial tracks), the indication of synchronised tracks in terms of random access may be declared in the base track, or in an entity group in a top-level box of the media file. There may also be sync samples declared in a spatial track, for example in a SyncSampleBox or in “TrackFragmentHeaderBox” or “TrackRunBox”.

[0215] A sync sample in a V3C spatial track is a sample for which all sub-bitstream composition units are sub-bitstream I RAP composition units as defined in ISO / IEC 23090-5.

[0216] A spatial track may also allow retrieving one type of component among the multiplexed component, for example through a “SubsamplelnformationBox” (‘subs’) box. A V3C spatial track sub-sample is a V3C unit which is contained in a V3C spatial track sample.

[0217] A V3C spatial track may contain one “SubSamplelnformationBox” in its “SampleTableBox”, or in the “TrackFragmentBox” of each of its “MovieFragmentBoxes”, which lists the V3C spatial track sub-samples.

[0218] At encapsulation time, the 32- bit unit header of the V3C unit which represents the sub-sample should be copied to the 32-bit “codec_specific_parameters” field of the sub-sample entry in the “SubsamplelnformationBox”. At parsing time, the V3C unit type of each sub-sample can be identified by parsing the “codec_specific_parameters” field of the sub-sample entry in the “SubsamplelnformationBox”. Since the properties for thesubsamples may be repetitive in a spatial track, using a subs box with the “SubsampleReferenceTableBox” (‘ssrt’) is more compact.

[0219] Figure 4b is another embodiment of a media file obtained after carrying out an encapsulation process according to some embodiments of the disclosure. The media file includes a base track 4100 and spatial tracks 4101 and 4102. In this embodiment, all the components have independently decodable data units, as can be seen in sample 4300 or 4400 (with atlas, video and geometry data). In this case the base track 4100 may contain parameter sets and / or decoder configuration information for all the components (e.g. atlas, geometry, video, etc...). Decoder configuration may be provided in a sample entry of the base track, for example as an exhaustive V3C configuration box ‘v3cC’ providing the exhaustive list of decoder configurations, possibly with their mapping to the sub-bitstreams they apply to. Alternatively, decoder configurations may be provided as several configuration boxes, for example one per sub-bitstream or one per sub-bitstream type. Possibly, the media file can include an additional box describing the mapping between these decoder configurations and the sub-bitstream they apply to. In this embodiment, the media file 4000 does not include a common track, as all the components have independently decodable data units. The base track does not contain any coded data units but can contain parameter set updates or SEI messages.

[0220] A sample entry for a base track like 4100 may be defined as follows:

[0221] Sample Entry Type: 'v3b1'

[0222] Container: SampleDescriptionBox

[0223] Mandatory: Yes, when spatial tracks are present

[0224] Quantity: Zero or more (in a file)

[0225] A V3C base track references one or more spatial tracks using track reference with a track reference type equal to ‘v3bs’ (V3C base track to spatial track). It contains the parameter sets or SEI messages for one or more referenced spatial tracks. It should not contain any coded-layer data units (e.g. ACL NAL units, VCL NAL units for video sub-bitstreams, BMC NAL units or DCL NAL units within V3C units), except if playing the role of the common track. When the V3C spatial base track contains coding layer data units, it should also contain the corresponding decoder configuration boxes in its sample entry. V3C spatial base track uses V3CSpatialBaseSampleEntry which extends VolumetricVisualSampleEntry with a sample entry type of 'v3bT.

[0226] An example syntax can be:

[0227] aligned class V3CSpatialBaseSampleEntry extends VolumetricVisualSampleEntry('v3bT) {V3CConfigurationBox config; / / v3cC with full decoder configurations V3CSpatialRegionCollectionBox regions; / / optional

[0228] }

[0229] With the following semantics:

[0230] "compressorname" in the base class ""config" contains a single instance of “V3CConfigurationBox” providing the exhaustive list of decoder configurations.

[0231] “regions” , when present, describes the collection of regions covered by the spatial tracks referenced from the spatial base track with this sample entry. When the number of regions declared in the “V3CSpatialRegionCollectionBox” is equal to 0, it means that spatial tracks correspond to one or more atlas tiles and not to an identified region. In a variant, the number of regions declared in the V3CSpatialRegionCollectionBox should not be equal to 0 and shall be greater than (when some regions overlap) or equal (no overlap) to the sum of the number of regions carried in each V3C spatial track associated to this base track.

[0232] The V3CSpatialBaseSampleEntry may also extend the V3CBitstreamSampleEntry and then inherit the configuration boxes.

[0233] Figure 4c is another embodiment of a media file 4500 obtained after carrying out an encapsulation process according to some embodiments of the disclosure. The media file 4500 contains a base track 4600, as the track 4100 on Figure 4b. The media file 4500 also contains spatial track 4601 and 4602 referred by the base track 4600. Each spatial track 4601 and 4602 includes only a subset of components, because only atlas and geometry provide spatial access in this example, as shown on samples 4700 or 4800. The media file 4500 contains additional tracks 4603 and 4604 (shared or common tracks) that includes the components that do not include independently decodable data units. These additional tracks 4603 and 4604 may be referenced from the base track directly. An alternative would be that each spatial track 4601 and 4602 references these additional shared tracks 4603 and 4604. Usual track reference types may be used from a base track to the additional shared tracks. It is to be noted that the attribute video track 4603 may also consist in a parameter set track for some video attributes, when the data for these video attributes are contained in a spatial track. This avoids repeating the parameter sets. In such case, the attribute video track may contain only empty samples, except to update some parameter sets or SEI messages.

[0234] Various embodiments have been described hereinabove wherein the collection box referred to as ‘V3CSpatialRegionCollectionBox’ provides a mapping between a collection of 3D regions and atlas tiles. The mapping namely enables toorganize tracks according to the 3D-regions. In the various embodiments, the collection of 3D regions has not been described with time-variation. One can consider that the previous embodiments apply to either a static mapping (i.e. with no variation of the collection of 3D-regions over time), or to the mapping during a single period of time.

[0235] In the following, an extension of the encapsulation according to one of the hereinabove various embodiments is provided, wherein the encapsulation is compliant with a dynamicity of the collection of 3D regions. In other words, the following embodiment enables to handle a change of the spatial information.

[0236] Namely, V3C tracks (single, atlas or timed metadata) have a description of a collection of 3D regions, each region being identified by a regionjdentifier. The dynamicity (that is to say, the change over time) of the collection of the 3D regions can result from different modification types:

[0237] an addition of a new 3D region to the collection;

[0238] a removal of a 3D region from the collection;

[0239] a reallocation of the region identifiers (regionjdentifier) or

[0240] a modification of at least a parameter of one of the 3D regions of the collection.

[0241] In the following, the 3D regions involved in a dynamic collection of 3D regions are qualified as ‘dynamic’.

[0242] Figure 4d represents an example of a collection of dynamic 3D regions along time represented by a timeline 4900. The timeline 4900 is split in periods 4901 , 4902 and 4903, the periods being separated by timepoints t1, t2.

[0243] The collection of 3D regions during the first period of time 4901 (until timepoint t1) comprises two 3D regions. One 3D region is defined as corresponding to atlas tile 1, and the other 3D region corresponds to atlas tiles 2 and 3.

[0244] An example of an addition of a new 3D region to the collection is represented on the second period of time 4902. Indeed, the collection of 3D regions is modified such that, during the second period of time 4902 (from timepoint t1 to timepoint t2), the collection comprises three 3D regions. One of the 3D regions is defined as corresponding to atlas tile 1, one other to atlas tiles 2 plus 3 and the last other to atlas tiles 3 plus 4.

[0245] An example of a removal of a 3D region from the collection is represented on the third period of time 4903. Indeed, during the third period of time 4903, the collection comprises two 3D regions. One is defined as corresponding to atlas tiles 2 plus 3, and the other is defined as corresponding to atlas tiles 3 plus 4.In general, it is to be noted that region identifiers may be reallocated differently from one period of time to another. Some region identifiers may disappear or appear along time. There may be a timed metadata track (not represented) associated to the base track when the 3D regions are dynamic. This metadata track may be a timed-metadata track with a sample entry type 'dyvm' as defined in ISO / IEC 23090-10. There may also be a viewport information timed-metadata track associated to the base track, for example with a sample entry type ‘6vpt’ as defined in ISO / IEC 23090-10. The viewport information timed-metadata track can be associated to the base track with a track reference type equal to ‘cdsc’. The viewport information may be useful when alternative base tracks with their spatial tracks are available for content selection. A player may for example select a base track and a set of spatial tracks corresponding to the recommended viewport by the content creator. Some base track allowing a partial reconstruction of a 3D scene may be associated to a viewport information timed-metadata track with a viewport type indicated a recommended viewport for an associated spatial region. Optionally, when the base track describes the whole 3D scene, the viewport information timed-metadata track with viewport type indicating a recommended viewport for an associated spatial region may be associated directly to a spatial track corresponding to this spatial region, for example via a track reference type ‘cdsc’ or via a dedicated track reference type (for example to allow the timed metadata track to be associated to both base and spatial track).

[0246] An example of a reallocation of the region identifiers is represented with reference to Figure 4d.

[0247] When a new 3D region is added to the collection, a new region identifier is allocated to the new 3D region. For example, the region identifier region_id=3 is added on the second period of time 4902.

[0248] When a 3D region is removed from the collection, the region identifiers are re-allocated to the remaining 3D regions.

[0249] For example, on the third period of time 4903, the 3D region corresponding to atlas tile 1 is removed from the collection, the region identifier region_id=3 is removed from the region identifiers, and the region identifiers region_id=1 and region_id=2 are reallocated to the two remaining 3D regions.

[0250] In general, a 3D region can have a variety of parameters that may be modified along time, such that the collection of 3D region is modified.

[0251] The parameters of the 3D regions may comprise any parameter from the following non exhaustive list:a description of its bounding box, for example as a 3D position and sizes along the 3 axes. The 3D position may be omitted when corresponding to the origin of the 3D scene, i.e. the 3D point with (0,0,0) coordinates in a 3D point in the Cartesian coordinate system associated to the 3D scene.

[0252] a description of its mapping to atlas tiles, possibly mixed tiles from different atlases. For example on a first time period a region with identifier “R” covers an atlas tile with ID “A” and on a second time period, this region with identifier “R” covers an atlas tile with ID “B”.

[0253] a list of objects included in the region that may be described by a bounding box and / or a mapping to atlas tiles. An object may have decoding dependencies to other object(s): it is recommended to store in a spatial track either the set of objects that are inter-dependent or objects that have no dependency to other objects.

[0254] Hence, to handle any one of these modification types (addition, removal, identifier reallocation or parameter modification) that may apply to the 3D region collection, it is proposed to extend the spatial tracks described in previous embodiments to dynamic spatial tracks, wherein a dynamic spatial track is a spatial track that carries a dynamic region. When 3D spatial regions defined for the volumetric media stream are considered as dynamic regions, the collection of regions within a spatial track should be a subset of these 3D spatial regions, meaning that the scope for the region identifiers “regionjd” is unique between a spatial base track, its related spatial tracks and associated timed_metadata track when present.

[0255] A new metadata structure is introduced within the dynamic spatial tracks. The new metadata structure is referred to as ‘dynamic region descriptor’ and indicates which 3D regions or combination of 3D regions are dynamic. For instance, as represented on Figure 4d, a dynamic region descriptor 4961 is contained in the spatial track 4960, and a dynamic region descriptor 4971 is contained in the spatial track 4970.

[0256] Having this possibility for time variation allows for example to create a spatial track that tracks an object or a specific area in a 3D scene along time.

[0257] A dynamic 3D region (for example the 3D region with the region_identifier=2 on Figure 4d) is carried in (or covered by) the spatial track (for example the spatial track 4961) along time. The region covered by the dynamic spatial track could then consist along time in a different set of independently decodable units (for example within one ormore samples 4962 of the spatial track 4960, or within one or more samples 4972 of the spatial track 4070) for the different sub-bitstreams, each set corresponding to a part of the scene where the object is located for the given time. Such a dynamic spatial track 4960 or 4970 with its base track 4950 may be available in a media presentation as an alternative view of the whole scene, to focus on, or to follow a particular object in a 3D scene.

[0258] The base track 4950 or spatial base track contains parameter sets 4951 for the whole volumetric media sequence and optionally data 4952 for one or more subbitstream that are common to all regions, for example data from a subbitstream without independently decodable units.

[0259] With dynamic spatial track, a 3D region with a given region identifier may be carried in a first dynamic spatial track on a first period of time, and carried in a second dynamic spatial track on a second period of time.

[0260] Different embodiments and variants of dynamic spatial tracks with at least a dynamic region descriptor are detailed hereafter.

[0261] It is to be noted that a media file may contain a mix of dynamic spatial and static (i.e. non dynamic) spatial tracks. If static and dynamic spatial tracks are present in a same media file, distinct sample entry types may be defined for players to distinguish and quickly identify the kind of spatial tracks. However, they may use the same sample entry and the player may identify dynamic or static by checking the presence or not of a dynamic region descriptor in the description of the track. When present, the spatial track is considered dynamic spatial track. When not present, the spatial track is considered static spatial track. Alternatively the dynamic region descriptor itself may support a static mode, for example as a parameter of this dynamic region description indicating a static configuration or dynamic configuration.

[0262] In an embodiment of the dynamic spatial track, the dynamic spatial track is self-contained. It means that it contains the description of the regions that it covers and may not rely on an external track like a timed metadata track providing the description of the regions along time. As well, the base track may not contain any description of the regions. This keeps on reducing the number of tracks for a volumetric media presentation carried in ISOBMFF.

[0263] According to this embodiment, a dynamic spatial track can be defined as follows:

[0264] A spatial track is identified by a specific sample entry type, for example ‘vssT.The sample entry may contain a V3CSpatialConfigurationBox providing configuration information for the spatial track. For example, it may indicate the region granularity like set of regions, one region only, tiles, object.... It may indicate the maximum number of regions carried in the spatial track, or the bounding box corresponding to the covered 3D scene along time. Optionally, the V3CSpatialConfigurationBox may contain the initial list of regions covered by the spatial track as a V3CSpatialRegionCollectionBox. A spatial track contains data units corresponding to one or more 3D regions that are independently decodable. A sample of a spatial track contains at least one V3C unit.

[0265] The parent spatial base track is indicated by a track reference of type 'v3bs' from the spatial base track to the spatial track. In a variant the association may be from the spatial track to the base track, using for example the track reference type ‘v3sb’, or in both direction (then using distinct track reference types, each with a specific 4CC). For V-DMC, V-PCC and MIV, the following statements apply:

[0266] each sample comprises at least one atlas V3C unit;

[0267] all the data units of a spatial track belong to a same set of atlas tiles. The set of atlas tiles is described in the dynamic region descriptor 4961 or 4971 of the spatial track 4960 or 4970.

[0268] A sample of a V3C spatial track can be defined as follows:

[0269] aligned ( 8 ) class V3CSpatialSample {

[0270] / / sample size size of sample from SampleSizeBox

[0271] for ( int i=0; i < sample size; ) {

[0272] unsigned int (v3c config . unit size precision bytes minusl + 1 ) *8 ) v3c unit size;

[0273] bit ( 8 ) ss v3c unit [v3c unit size] ;

[0274] i += v3c unit size +

[0275] v3c config . unit size precision bytes minusl + 1;

[0276] }

[0277] }

[0278] With the following semantics:

[0279] v3c_unit_size specifies the size, in bytes, of the ss_v3c_unit array. The size is equivalent to the sample stream v3c unit size ssvu_v3c_unit_size as defined in ISO / IEC 23090-5, Annex C.ss_v3c_unit contains a single V3C unit in V3C unit sample stream format as defined in ISO / IEC 23090-5:2021 in Annex C

[0280] To allow random access or seeking into a spatial track, a sync sample in a V3C spatial track is a sample for which all sub-bitstream composition units are sub-bitstream I RAP composition units as defined in ISO / IEC 23090-5. It is recommended that the description of the regions carried in the track is available for a sync sample.

[0281] As specified for static spatial track described in previous embodiments, a dynamic spatial track may have subsamples.

[0282] A dynamic spatial track contains in its description a dynamic region descriptor (4971 or 4961). In this embodiment the dynamic region descriptor may comprise or consists in a sample group called 3D region information sample group and identified by the grouping type ‘3dri’ (this four-character code is just an example, any reserved and non conflicting 4CC can be used). This sample group can be defined as follows:

[0283] Group Type: '3dri'

[0284] Container: SampleGroupDescriptionBox ('sgpd')

[0285] Mandatory: Yes (in dynamic spatial track)

[0286] Quantity: Zero or more

[0287] The 3DSpatialRegionlnformationGroupEntry may be used to describe one or more 3D spatial regions present in a V3C spatial track. The entries of the SampleGroupDescriptionBox list the regions covered by the samples of a V3C spatial track mapped to this entry.

[0288] The grouping_type_parameter is not defined for the SampleToGroupBox with grouping type '3dri'.

[0289] The Syntax for a sample group description entry of this group can be defined as follows:

[0290] class 3DSpatialRegionlnformationGroupEntry() extends VolumetricVisualSampleGroupEntry ('3dri')

[0291] {

[0292] V3CSpatialRegionCollectionBox spatial_regions;

[0293] }

[0294] With the following semantics:

[0295] spatial_regions describe the regions covered by (or carried in) the spatial track. Each region in the collection may be described using the V3CSpatialRegion descriptor fromISO / IEC 23090-10. In a variant each region may be defined as a region identifier and a 3D bounding box, or as a region identifier and a set of atlas tile identifiers, or as a list of object identifiers provided that the definition of the objects is available, for example in the parent base track.

[0296] The definition of the V3CSpatialRegionCollectionBox is then updated to indicate that it may apply to spatial tracks:

[0297] Container (for a V3CSpatialRegionCollectionBox) can be the sample entries for Single track V3CBitstreamSampleEntry,

[0298] Atlas track V3CAtlasSampleEntry,

[0299] Spatial track V3CSpatialSampleEntry ( 'vssl' ) , or

[0300] Timed-metadata track DynamicVolumetricMetadataSampleEntry ( ' dyvm' ) or

[0301] in a dynamic region descriptor 3DSpatialRegionInf ormationGroupEntry ( '3dri' ) .

[0302] Hence, in this embodiment of the dynamic spatial track, the dynamic region descriptor consists in a sample group mutualizing the description (object (OBJ), region (REG), tile (TILE)) that may apply to the dynamic regions. It is to be noted that the dynamic spatial tracks may be constructed with overlapping regions, as depicted on Fig.

[0303] 4b with region 2 and region 3 on time period 4902 or region 1 and region 2 on time period 4903.

[0304] In general, as a variant of embodiments of the dynamic spatial track, the dynamic region descriptor (for example illustrated by reference number 4961 or 4971 on Figure 4d) contained in spatial tracks (for example illustrated by reference number 4960 or 4970) may rely on a dynamic region description that is described outside the dynamic spatial tracks.

[0305] An embodiment of the dynamic spatial track according to this variant is illustrated with reference to Figure 4e. The metadata structures that are identical to Figure 4d are identified by the same reference numbers.

[0306] In this embodiment of the dynamic spatial track, a dynamic region description is comprised in the base track 4950 associated to the dynamic spatial tracks 4991 , 4992. In this embodiment, the base track 4950 provides the whole description list.Hence, as represented, the base track 4950 contains a common dynamic region descriptor 4990 which comprises the whole list of descriptors that may apply to the different dynamic regions.

[0307] In this embodiment of the dynamic spatial track, the dynamic spatial tracks 4960, respectively 4970, contain a dedicated metadata structure for enabling reference to the common dynamic region descriptor 4990. On Figure 4e, the dedicated metadata structure is referred to as ‘dynamic region descriptor ref’ 4991, respectively 4992.

[0308] Similarly to the previous embodiment of the dynamic spatial track, these features advantageously allow reducing the number of tracks since a timed metadata track describing dynamic regions is not mandatory. Moreover, it has some benefits when spatial tracks provide region overlap: a region may be declared once in the base track and can be referenced in one or more spatial tracks associated to this base track, thus reducing the description cost.

[0309] In this embodiment of the dynamic spatial track, the dynamic region descriptor for the base track may comprise or consists in a sample group called 3D region information sample group and is identified by the grouping type ‘3dri’ (this four-character code is just an example, any reserved and non-conflicting 4CC can be used). This sample group can be defined as follows:

[0310] Group Type: '3dri'

[0311] Container: SampleGroupDescriptionBox ('sgpd')

[0312] Mandatory: Yes, in base track, No in other tracks

[0313] Quantity: Zero or more

[0314] The 3DSpatialRegionGroupEntry may be used to describe one or more 3D spatial regions present in a V3C base track. The entries of the SampleGroupDescriptionBox list the complete set of regions covered by the samples of all the V3C spatial tracks referenced by this base track.

[0315] The grouping_type_parameter is not defined for the SampleToGroupBox with grouping type '3dri'.

[0316] The Syntax for a sample group description entry of this group can be defined as follows:

[0317] class 3DSpatialRegionGroupEntry() extends VolumetricVisualSampleGroupEntry ('3dri')

[0318] {V3CSpatialRegionCollectionBox spatial_regions;

[0319] }

[0320] With the following semantics:

[0321] spatial_regions describe the regions covered by (or carried in) the spatial tracks referenced by the base track. Each region in the collection may be described using the V3CSpatialRegion descriptor from ISO / IEC 23090-10. In a variant each region may be defined as a region identifier and a 3D bounding box, or as a region identifier and a set of atlas tile identifiers, or as a list of object identifiers provided that the definition of the objects is available, for example in the parent base track.

[0322] In this embodiment of the dynamic spatial track, the dynamic region descriptor for a spatial track may comprise or consists in a sample group called 3D region reference sample group and is identified by the grouping type ‘3drr’ (this four-character code is just an example, any reserved and non-conflicting 4CC can be used). This sample group can be defined as follows:

[0323] Group Type: '3drr'

[0324] Container: SampleGroupDescriptionBox ('sgpd')

[0325] Mandatory: Yes, in dynamic spatial track, No in other tracks Quantity: Zero or more

[0326] The 3DSpatialRegionReferenceGroupEntry may be used to describe one or more 3D spatial regions present in a V3C dynamic spatial track, by reference to regions declared in its parent base track. The entries of the SampleGroupDescriptionBox list, by their region identifiers, the region or set of regions covered by the samples of the V3C spatial track mapped to this entry.

[0327] The grouping_type_parameter is not defined for the SampleToGroupBox with grouping type '3drr'.

[0328] The Syntax for a sample group description entry of this group can be defined as follows:

[0329] class 3DSpatialRegionReferenceGroupEntry() extends VolumetricVisualSampleGroupEntry ('3drr')

[0330] {

[0331] unsigned int (8) num_regions;

[0332] unsigned int (16) regionjdentifier [num_regions];

[0333] }With the following semantics:

[0334] num_regions indicates the number of regions from the list declared in the dynamic region descriptor of the base track that are contained in the samples of a spatial track mapped to this sample group entry

[0335] regionjdentifier is the region identifier of a region carried in (or covered by) the samples of spatial track mapped to this sample group entry. It takes as value one of the regionjd values declared in the associated (or parent) base track.

[0336] In case the region collection is static for the volumetric media sequence but their allocation in spatial tracks is changing along time, then the base track may define its dynamic region descriptor as being static by setting appropriately the static_mapping and static_group_description flags of the SampleGroupDescriptionBox with ‘3dri” grouping type. Alternatively, the dynamic region descriptor is replaced by the presence of a V3CSpatialRegionCollectionBox in the sample entry of the base track.

[0337] The definition of the V3CSpatialRegionCollectionBox is then updated to indicate that it may apply to spatial tracks:

[0338] Container (for a V3CSpatialRegionCollectionBox) can be the sample entries for Single track V3CBitstreamSampleEntry,

[0339] Atlas track V3CAtlasSampleEntry,

[0340] Spatial base track V3CSpatialBaseSampleEntry ( 'v3bl' ) , or Timed-metadata track DynamicVolumetricMetadataSampleEntry ( ' dyvm' ) or

[0341] in a dynamic region descriptor 3DSpatialRegionInf ormationGroupEntry ( '3dri' ) .

[0342] In another variant, the dynamic region descriptor present in the base track may comprise or consists in samples containing a V3CSpatialRegionCollectionBox instead of a sample group. This V3CSpatialRegionCollectionBox updates the description of the regions carried in the spatial tracks referenced by the base track. Then, the regionjdentifier in the 3DSpatialRegionReferenceGroupEntry corresponds to a region identifier declared in the timed_aligned sample if any or in the preceding sync sample in the base track referencing this spatial track.To allow random access or seeking, a sync sample in a base track is a sample for which a region description is available and is a sample that provides random access point for the V3C spatial tracks it references.

[0343] In this embodiment of the dynamic spatial track, the metadata structure ‘dynamic region descriptor ref’ 4991 or 4992 contained in spatial tracks 4960 or 4970 relies on a dynamic region description that is described outside the spatial track 4960 or 4970, for example in a timed metadata track (not represented on Fig. 4e). The sample entry of the timed metadata track contains initial information for the 3D regions and its samples describe the 3D regions updates along time.

[0344] In a variant, wherein the dynamic region description is contained in a timed metadata track, the samples contain the complete description of the 3D regions present in the volumetric media presentation, especially on synchronized samples (‘sync samples’) to guarantee random access or seeking into the presentation. The timed-metadata track may be associated to the spatial base track via a track reference type ‘cdsc’.

[0345] Advantageously, the use of a timed metadata track may be combined with a base track containing a dynamic region descriptor. Indeed, when samples of the timed metadata track only provide updates for some regions, these updates may be used to compute and describe in the dynamic region descriptor of the base track the complete list of regions that may be carried in the different spatial tracks referenced by this base track. Then, spatial track keeps on referencing a region identifier from the dynamic region descriptor of its parent (or associated) base track.

[0346] When the use of a timed metadata track is combined with a base track that does not contain a dynamic region descriptor, then the dynamic region descriptor of a spatial track is defined as follows: the dynamic region descriptor for a spatial track consists in a sample group called 3D region reference sample group and is identified by the grouping type ‘3drr’ (this four-character code is just an example, any reserved and non conflicting 4CC can be used). This sample group can be defined as follows:

[0347] Group Type: '3drr'

[0348] Container: SampleGroupDescriptionBox ('sgpd')

[0349] Mandatory: Yes, in dynamic spatial track, No in other tracks Quantity: Zero or more

[0350] The 3DSpatialRegionReferenceGroupEntry may be used to describe one or more 3D spatial regions present in a V3C spatial track, by reference to regions declaredin a timed metadata track associated to its parent base track. The entries of the SampleGroupDescriptionBox list, by their region identifiers, the region or set of regions covered by the samples of the V3C spatial track mapped to this entry.

[0351] The grouping_type_parameter is not defined for the SampleToGroupBox with grouping type '3drr'.

[0352] The Syntax for a sample group description entry of this group can be defined as follows:

[0353] class 3DSpatialRegionReferenceGroupEntryO extends VolumetricVisualSampleGroupEntry ('3drr')

[0354] {

[0355] unsigned int (8) num_regions;

[0356] unsigned int (16) regionjdentifier [num_regions];

[0357] }

[0358] With the following semantics:

[0359] num_regions indicates the number of regions that are carried in (or covered by) the samples of a spatial track mapped to this sample group entry regionjdentifier is the region identifier of a region carried in (or covered by) the samples of spatial track mapped to this sample group entry. It takes as value one of the regionjd values declared in the timed-metadata track associated to its parent base track.

[0360] The sync samples of a spatial track may be aligned with the sync samples of the timed-metadata track, especially when some region identifiers are removed or added from the list of regions declared in the timed metadata track associated to its parent base track (for example using a ‘cdsc’ track reference).

[0361] In the case where the base track also contains a dynamic region descriptor, the sync samples of a spatial track are rather aligned on the sync samples of the base track and do need to be aligned with those of the timed metadata track.

[0362] In this embodiment of the dynamic spatial track, a flexible mechanism is provided.

[0363] In the flexible mechanism, the dynamic region descriptor is a generic dynamic region descriptor which is adapted to contain indications in order to determine whether a dynamic region description is inline (i.e. within the dynamic spatial track) or external (i.e. within a base track or a timed metadata track) or both. Depending on these indications, the dynamic region descriptor may contain:either dynamic region descriptions,

[0364] or references to external dynamic region descriptions,

[0365] or a mix of region descriptions and references.

[0366] In a variant the dynamic region descriptor comprises or consists in a sample group called 3D region information sample group and identified by the grouping type ‘3dri’ (this four-character code is just an example, any reserved and non-conflicting 4CC can be used). This sample group can be defined as follows:

[0367] Group Type: '3dri'

[0368] Container: SampleGroupDescriptionBox ('sgpd')

[0369] Mandatory: Yes (in dynamic spatial track)

[0370] Quantity: Zero or more

[0371] The 3DSpatialRegionlnformationGroupEntry may be used to describe one or more 3D spatial regions present in a V3C spatial track. The entries of the SampleGroupDescriptionBox list the regions covered by the samples of a V3C spatial track mapped to this entry.

[0372] The syntax for a sample group description entry of this group can be defined as follows:

[0373] class 3DSpatialRegionlnformationGroupEntry() extends VolumetricVisualSampleGroupEntry ('3dri')

[0374] {

[0375] bit(1) inline_regions_flag;

[0376] bit(1) external_regions_flag;

[0377] unsigned int (6) reserved;

[0378] if (inline_regions flag) {

[0379] unsigned int(16) num_region_definitions;

[0380] V3CSpatialRegion inline_regions [num_region_definitions]

[0381] }

[0382] if (external_regions_flag) {

[0383] unsigned int(16) num_region_references;

[0384] for(int i=1; i<=num_region references; i++) {

[0385] unsigned int(16) region_id[i];

[0386] unsigned int (32) trackj D[i];

[0387] }

[0388] }With the following semantics:

[0389] inline_regions_flag indicates whether the sample group entry contains region definitions or not.

[0390] external_regions_flag (or referenced_regions_flag) indicates whether the sample group entry contains references to region definitions defined outside the spatial track or not. At least one of these 2 flags should be set to true. However, in some use cases for object tracking, these 2 flags could be both set to false to indicate that there is no data for an object on a time period, corresponding to a group of samples (i.e. the object may have disappeared from the scene).

[0391] num_region_definitions indicates the number of regions definitions contained in the sample group entry.

[0392] inline_regions provide a region definition, for example using the V3CSpatialRegion descriptor from ISO / IEC 23090-10. In a variant each region may be defined as a region identifier and a 3D bounding box, or as a region identifier and a set of atlas tile identifiers, or as a list of object identifiers provided that the definition of the objects is available, for example in the parent base track.

[0393] num_region_references indicates the number of regions references contained in the sample group entry.

[0394] region_id[i] is the region identifier of the referenced region with index i

[0395] track_ID[i] is the identifier of the track (trackJD) that contains the definition of the referenced region with index i. In a variant, a 4CC may be used to indicate the sample entry type of the track where the referenced region is defined; for example ‘v3b1’ for the base track associated to the spatial track of ‘dyvm’ for a timed metadata track. The base track and the timed metadata track can be associated through a track reference type ‘cdsc’.

[0396] In a variant, the grouping type parameter of the SampleToGroupBox with grouping_type=’3dri’ may indicate whether the region is inline or referenced, so that a sample may be mapped to a set of inline regions or / and to a set of referenced regions, for example with two sample to group box with grouping_type = ‘3dri’ but with different values for their “grouping_type_parameter” field. The sample group description entry may then contain one or more inline region definitions or one or more description of a referenced regions. The sample group description entry in this variant would not need the inline_regions_flag and external_regions_flag anymore.In a general variant of the dynamic spatial track, the dynamic region descriptor may comprise or consists in a sample to region definition mapping sample group. The sample to region definition mapping ( ' srdm' ) sample grouping provides a flexible way to associate regions inside samples of a spatial track, with region definitions. It provides the mapping between region identifiers in samples of the spatial track and region definitions carried in various types of containers: tracks, items or sample group descriptions. The region in the spatial track and the region definition are mapped via their respective regionjd.

[0397] The syntax of the corresponding sample group description entry can be defined as follows:

[0398] class SampleToRegionDef initionMappingEntry ( )

[0399] extends SampleGroupDescriptionEntry ( ' srdm' ) {

[0400] unsigned int ( 32 ) entry count;

[0401] for ( i=l ; i<= entry count ; i++) {

[0402] unsigned int ( 32 ) region id;

[0403] unsigned int ( 3 ) region container type;

[0404] unsigned int ( 32 ) region reference type;

[0405] unsigned int ( 32 ) region container idx;

[0406] With the following semantics:

[0407] entry_count is an integer that gives the number of entries in the mapping table, i.e., the number of regions in the mapping table.

[0408] region_id is an integer that specifies the identifier of the region carried by the spatial track.

[0409] region_container_type specifies the type of container carrying the region definition. Following values for region_container_type are defined:

[0410] — 0: the region definition is carried in samples of a track or in an item. When the region definition is carried in samples of the track identified by region reference type and region container idx, the sample in that track that is temporally aligned, or the nearest preceding in the media presentation timeline, comprises the region definition.— 1: the region definition is in a SampleGroupDescriptionEntry corresponding to the index region_container_idx in the SampleGroupDescriptionBox with grouping type equal to region reference type.

[0411] — 2: the region definition is carried in a SampleGroupDescriptionEntry of the base track associated to the spatial track

[0412] — 3-7: reserved.

[0413] region_reference_type specifies the four-character code type of the container carrying the region definition depending on the region_container_type value.

[0414] If the value of region container type equals 0, it specifies the sample entry type or the item type of entity (i.e., track or item) containing the region definition.

[0415] If the value of region_container_type equals 1 or 2, it specifies the grouping type of the SampleGroupDescriptionBox containing the region definition.

[0416] The meaning of region_ref erence_type is undefined for other values of region container type.

[0417] region_container_idx specifies the index to retrieve the container carrying the region definition depending on the region_container_type value.

[0418] If the value of region_container_type equals 0 or 2, it specifies the identifier of the track (track_ID) or of the item (item_ID) from which to retrieve the region definition.

[0419] If the value of region_container_type equals 1 , it specifies the index of the SampleGroupDescriptionEntry in the SampleGroupDescriptionBox with grouping type equal to region reference type from which to retrieve the region definition.

[0420] The meaning of region_container_idx is undefined for other values of region container type.

[0421] The sample entry, subsamples, sync samples for spatial track follow the same rules as in previous embodiments.It is to be noted that spatial tracks either static or dynamic may be described with the dynamic region descriptor when this descriptor has means to indicate that there is no variation. This is the case with sample groups. Indeed, sample group can handle both static and dynamic spatial regions by setting the flags of the SampleGroupDescriptionBox appropriately: the sample group description for dynamic regions may be set to static by using the appropriate static_mapping and static_group_description flags. When set, these two flags mean that the initial list of regions (either in sample entry or in sample group entry) does not change along time. It is also to be noted that when regions are updated on a sequence basis, an alternative way to describe dynamic spatial track can consist, for the dynamic region descriptor, in sample entries with an updated V3CSpatialRegionCollectionBox describing the actual regions for the coming sequence. This requires to also update the sample_description_index in the sample description to these new sample entries. This may be a possibility, especially if a V3C VPS is inserted between consecutive volumetric sequences that would also require to update the V3C decoder configuration and thus the sample entry of the base track. The first sample of a new sequence may be declared as sync sample in the base track and its referenced spatial tracks. When producing fragmented mp4 files with spatial tracks, the change of regions may be a criterion for fragmentation, i.e. a new fragment may start when regions are updated.

[0422] In the case where a spatial track contains more that one spatial regions, i.e. several region identifiers, it may be useful to have a further description of which parts of the samples correspond to which region identifier. This can be used by advanced parsers to filter data for one or more regions within a spatial track or to do extraction of a subset of regions carried in a spatial track. Since the regions carried in a track may be dynamic, this further description should also be dynamic. It is then proposed as a “virtual” sample group description with a dedicated grouping_type. The sample group is said to be virtual because this sample to group first maps samples into groups having the same V3C units ranges and ranges of V3C units are then mapped to region identifiers directly (and not to region descriptions in a SampleGroupDescriptionBox). These region descriptions may be found through one of the static or dynamic region descriptor explained according to previous embodiments. The sample to group box and sample group description box for the V3C unit mapping may have a dedicated grouping type ‘v3um” for “V3C Unit Mapping”. The grouping type parameter may be set to ‘3dri’ to indicate that V3C units are mapped to 3D regions identifiers. For example, a spatial track may consist in arepetitive pattern of 1 V3C unit for atlas followed by 1 V3C unit for geometry or one or more V3C unit for attributes. Samples with the same V3C unit patterns are grouped in a group of samples and this group is associated, through the group_description_index field of the sample to group box with grouping_type equal to ‘v3um’, to an entry in a sample group description box of grouping type equal to ‘v3um’. For example, this entry may indicate that a range of 3 consecutive V3C units (from a given v3C unit start number) are corresponding to a given region. This may be relevant when data in a spatial track have one V3C unit of each sub-bitstream per region and that these data unit follow a “region-first” order. This means that all V3C units of different types for a first region precede that all V3C units of different types for a first region, etc...

[0423] This sample group description box contains one or more V3C unit map entries that may be defined as follows:

[0424] class V3CUnitMapEntry() extends VisualSampleGroupEntry ('v3um') {

[0425] bit(6) reserved = 0;

[0426] unsigned int(1) large_size;

[0427] unsigned int(1) rle;

[0428] if (large_size) {

[0429] unsigned int(16) entry_count;

[0430] } else {

[0431] unsigned int(8) entry_count;

[0432] }

[0433] for (i=1; i<= entry_count; i++) {

[0434] if (rle) {

[0435] if (large_size) unsigned int(16) v3C_unit_start_number;

[0436] else unsigned int(8) v3C_unit_start_number;

[0437] }

[0438] unsigned int(16) groupID;

[0439] }

[0440] }

[0441] With the following semantics:

[0442] large_size indicates whether the number of V3C units entries in the track samples is represented on 8 or 16 bits.

[0443] rle indicates whether run-length encoding is used (1) to assign groupID to V3C units or not (0).entry_count specifies the number of entries in the map. Note that when rle is equal to 1 , the entry_count corresponds to the number of runs where consecutive V3C units are associated with the same group. When rle is equal to 0, entry_count represents the total number of V3C units.

[0444] v3c_unit_start_number is the 1 -based V3C unit index in the sample of the first NAL unit in the current run associated with groupID.

[0445] groupID specifies the unique identifier of the group. All V3C units mapped to the same group with a particular groupID value have the same properties in all the sample groups that indicate a mapping to that particular groupID value and are contained in this track. In case of a virtual sample group, the group ID provides the value for the descriptor associated to the mapped V3C units, i.e. a region identifier in this example.

[0446] More generally, when non virtual, information about the group is provided by the sample group description entry with this groupID and grouping_type equal to the grouping_type_parameter of the SampleToGroupBox of type 'v3um'. In case of a virtual group, this additional sample group description entry with grouping_type equal to the grouping_type_parameter of the SampleToGroupBox of type 'v3um' is not present. It is to be noted that this further description may apply to single track encapsulation when the single track is associated to a timed-metadata track describing the dynamic regions. Then, the V3C units of the single track can be mapped to region identifiers declared in this timed-metadata track describing the dynamic regions.

[0447] It is to be noted that V3C items may also benefit from the data organization as described for spatial tracks: a base item may be defined containing the common data and common parameter sets or configuration information and may reference one or more V3C spatial items, each with a property describing the regions covered by this item. The data for a spatial item consist in independently decodable data units corresponding to a given region. Independently decodable data units corresponding to a given region from the different sub-bitstreams are multiplexed into a single item. Each sub-bitstream may be represented as an extent allowing more flexibility in the data organization. As for tracks, the granularity of the region may be tile, object, one spatial region or a set of spatial regions. The regions, tiles or object may be declared a v3cspatiaiRegionCoiiectionBox declared as an item property box associated to the base item. As for spatial tracks, the spatial items may be associated to a property indicating the regions covered by this item. The syntax for such property may be defined as follows:aligned ( 8 ) class V3CSpatialRegionCollectionProperty extends ItemFullProperty ( ' v3sc ' , version, flags ) {

[0448] unsigned int ( 16) num regions;

[0449] for ( int i=l; i<=num regions; i++) {

[0450] V3CSpatialRegion region;

[0451] }

[0452] }

[0453] With the following semantics:

[0454] num_regions indicates the number of 3D spatial regions in the V3C media.

[0455] region describes attributes related to 3D spatial region belonging to the V3C media.

[0456] Version and flags may be set to 0 by default if there is no need for extension or additional fields.

[0457] This property can be declared in an ItemPropertyContainer box and associated to V3C base item or V3C spatial item through the ItemPropertyAssociationBox. V3C base item and V3C spatial item may be identified by specific item types like ‘v3bT and ‘vssT respectively. A viewport information property may also be associated to a V3C base item or directly to a V3C spatial item when the viewport type indicates a recommended viewport suggested for an associated spatial region. A viewport information property may be defined as follows:

[0458] aligned ( 8 ) class V3CViewportInf ormationPropertyBox extends ItemFullProperty ( ' 6vpt ' , version, flags ) {

[0459] unsigned int (7 ) viewport type;

[0460] bit ( l ) reserved = 0;

[0461] Viewportinfo viewport info;

[0462] string viewport description; / / optional

[0463] }

[0464] With the following semantics:

[0465] viewport_type specifies the type of the viewport, for example in a pre-defined list as specified in the different parts (e.g. 2, 7, 10 or 18) of MPEG immersive specifications ISO / IEC 23090.

[0466] viewportjnfo specifies the viewport information, for example the associated intrinsic and extrinsic camera information for the 3D content.viewport_descri ption is null-terminated UTF-8 string that provides a textual description of the viewport

[0467] Version and flags may be set to 0 by default if there is no need for extension or additional fields.

[0468] Figure 5 illustrates another embodiment of a media file 500, for example corresponding to media file 135 in Figure 1, obtained after carrying out an encapsulation process according to some embodiments of the disclosure, for example the one described by reference to Figure 2 or the one described by reference to Figure 2b. In this embodiment, the dedicated track is not a spatial track multiplexing the data units for independently decodable data units. Instead, the data units for a given spatial part are kept in component tracks. An atlas tile track is used as a dedicated track to reference set of component tracks that have independently decodable data units and that correspond to a part of a 3D scene.

[0469] The set of component tracks corresponding to a part of a 3D scene are themselves referencing one or more component tracks containing common data for the 3D scene. With this track organisation, the correspondence between atlas tile tracks and corresponding spatial video or spatial geometry tracks are expressed through the track references and avoids looking for the correspondence between the parts when selecting one or more regions to decode. Simply “pulling” from the atlas base track and the atlas tile track will provide the data units necessary to decode a selected region.

[0470] The tracks are then organized as follows: a main or base track 510 contains parameter sets describing the whole bitstream. It may also contain atlas information common to all atlas tiles. Atlas tile track 520-1, 520-2 contain the data units for one or more tiles and are referenced by a track reference type ‘v3ct’ from the atlas base track 510. Atlas tile tracks have sample entry type ‘v3t1’.

[0471] Each atlas tile track then references, for example through a specific track reference type identified by a reserved 4CC, one or more submesh tracks (e.g. 520-11 or 520-21, 520-22) or one or more video-coded attribute tracks (e.g. 520-15, 520-16 or 520-25).

[0472] For sub-bitstreams that do not provide spatial access (i.e. independently decodable data units), they may be referenced from the atlas tile tracks, for example the displacement track 530 via a track reference type ‘vttd’ for V3C atlas Tile track To Displacement track).The displacement track 530 may contain video-coded displacement data or arithmetic coded displacement data.

[0473] For sub-bitstreams that do provide spatial access (i.e. independently decodable data units), they may reference a common track providing parameter sets and possibly rewriting instructions (when not reconstructing the whole scene).

[0474] For example, the submesh tracks 520-11 and 520-21 and 520-22 reference the basemesh track 550. As well, the attribute tile tracks 520-15, 520-16 or 520-25 (tile here is used as a generic term for spatial access to a video like motion-constrained HEVC tiles or VVC subpicture for example) reference a common attribute video track 540. Depending on the codec in use, they may reuse the HEVC tile tracks or the VVC subpicture tracks with their corresponding track references to a base track (respectively ‘tbas’ or ‘vvcN’).

[0475] In order to indicate possible direct relationships between the atlas tile tracks and the subsmesh tracks, especially when they may change along time, a table of correspondence between submesh identifiers and atlas tile identifiers may be described, for example in the atlas base track 510. Preferably, since atlas tile identifiers and submesh identifiers may vary along a 3D coded sequence, this table is described as a sample group, allowing to describe various correspondences along time on a group of samples basis . A dedicated grouping type is defined and reserved, for example ‘ttsm’ for “Tiles to Sub Meshes”. The corresponding volumetric sample group entry may be defined as follows:

[0476] aligned(8) class V3CTilesToSubMeshGroupEntry()

[0477] extends VolumetricVisualSampleGroupEntry('ttsm') {

[0478] unsigned int(16) num_tiles;

[0479] for(int i=0; i < num_tiles; i++) {

[0480] unsigned int(16) tilejd;

[0481] unsigned int(16) num_submeshes;

[0482] for(int j=0; j < num_submeshes; j++) {

[0483] unsigned int(16) submeshjd;

[0484] }

[0485] }

[0486] }

[0487] With the following semantics:

[0488] num_tiles indicates the number of atlas tiles associated with the samples mapped to this sample group entry.ti I e_i d specifies the atlas tile I D of an atlas tile associated with the submeshes declared in this sample group entry. The value of tilejd may for example be equal to the value of a syntax element in atlas frame tile information in the atlas sub-bitstream (for example afti_tile_id as defined in ISO / IEC 23090-5 or derived specifications).

[0489] num_submeshes indicates the number of submeshes associated with the the atlas tile.

[0490] submeshjd is the identifier for the submeshes declared in this sample group entry. The value of submeshjd may be equal to the value of a syntax element in a base mesh sub-bitstream (for example the bmsi_submeshjd syntax element in bmesh_submesh_information(), defined in ISO / IEC 23090-29:Annex H).

[0491] A media player or reader may determine from this sample group the submesh identifiers associated to one or more atlas tiles identifiers. It may then deduce which submesh track(s) is (are) corresponding to a given atlas tile track, by parsing the subMeshconf igurationBox of the submesh tracks (520-11 or 520-21, 520-22) and the vscAtias iieconf igurationBox of the atlas tile track (520-1 or 520-2). There may an advantage to store this sample group in each atlas tile track: by doing so, for a given atlas tile track, a list of corresponding submesh identifiers can be found and inspection of the sample entries of the submesh tracks allow determining which submesh track(s) correspond to the atlas tile track. As well an atlas tile to video tiles may also be indicated in a similar manner to associate atlas tile tracks to attribute tile tracks dynamically. This could replace the track references ‘vtts’ and keep on using the track references defined for mutli-track encapsulation mode in ISO / IEC 23090-10.

[0492] Besides, it is possible for spatial tracks to store data units per component as split layer and data being NAL units. This allows simplifying the bitstream reconstruction for a spatial part that is further sent to a set of decoders (one decoder per component type) by a media file parser.

[0493] A spatial track, instead of having one sample entry, like the example “V3CSpatialSampleEntry” previously described, may have a hierarchy of sample entries, one sample entry per component.

[0494] A first sample entry, for example corresponding to the atlas component may be the one declared in a “SampleTableBox” of the spatial track. This sample entry may contain a “SplitSampleDescriptionsBox” containing other sample entries, each corresponding to a component. This means that a sample is split in sample parts (or split-layer) according to a component type.There may be several sample parts (or split-layers) for a given component (for example for video coded ones). A split layer configuration box may be present in the sample description of the spatial track to indicate the number of split-layers, thus of components.

[0495] Then, a container box describing each split layer may indicate the position and sizes of the data for each chunk of samples of each component. With this track organization, the encapsulation module can directly store in spatial track NAL units and no more V3C units.

[0496] The type of NAL units and component can be indicated in the sample entry of each split layer. Optionally, the sample entry of each split layer contains an indication of the original V3C unit type and when needed of the original attribute index.

[0497] Figure 6 is a schematic block diagram of a computing device 600 for implementing at least a part of one or more embodiments of the disclosure. The computing device 600 may be a device such as a micro-computer, a workstation or a light portable device. The computing device 600 comprises a communication bus 602 connected to one, several or all the following elements:

[0498] - a central processing unit 604, such as a microprocessor, denoted CPU; - a random access memory 608, denoted RAM, for storing the executable code of at least a part of the method of embodiments of the disclosure as well as the registers adapted to record variables and parameters necessary for implementing the method according to embodiments of the disclosure, the memory capacity thereof can be expanded by an optional RAM connected to an expansion port for example;

[0499] - a read only memory 606, denoted ROM, for storing computer programs for implementing embodiments of the disclosure;

[0500] - a network interface 612 that may be connected to a communication network over which digital data to be processed may be transmitted or received. The network interface 612 can be a single network interface, or composed of a set of different network interfaces (for instance wired and wireless interfaces, or different kinds of wired or wireless interfaces). Data packets are written to the network interface for transmission or are read from the network interface for reception under the control of the software application running in the CPU 604;

[0501] - a graphical user interface 616 may be used for receiving inputs from a user or to display information to a user;- a hard disk 610 denoted HD may be provided as a mass storage device; and

[0502] - an I / O module 618 may be used for receiving / sending data from / to external devices such as a 3D mediadata source or display.

[0503] The executable code may be stored either in read only memory 606, on the hard disk 610 or on a removable digital medium such as for example a disk. According to a variant, the executable code of the programs can be received by means of the communication network 614, via the network interface 612, in order to be stored in one of the storage means of the communication device 600, such as the hard disk 610, before being executed.

[0504] The central processing unit 604 is adapted to control and direct the execution of the instructions or portions of software code of the program or programs according to embodiments of the disclosure, which instructions are stored in one of the aforementioned storage means. After powering on, the CPU 604 is capable of executing instructions from main RAM memory 608 relating to a software application after those instructions have been loaded from the program ROM 606 or the hard-disc (HD) 610 for example. Such a software application, when executed by the CPU 604, causes steps of the flowcharts of the disclosure to be performed.

[0505] Any step of the algorithms of the disclosure may be implemented in software by execution of a set of instructions or program by a programmable computing machine, such as a PC (“Personal Computer”), a DSP (“Digital Signal Processor”) or a microcontroller; or else implemented in hardware by a machine or a dedicated component, such as an FPGA (“Field-Programmable Gate Array”) or an ASIC (“Application-Specific Integrated Circuit”).

[0506] Although the present disclosure has been described hereinabove with reference to specific embodiments, the present disclosure is not limited to the specific embodiments, and modifications will be apparent to a skilled person in the art which lie within the scope of the present disclosure.

[0507] Many further modifications and variations will suggest themselves to those versed in the art upon making reference to the foregoing illustrative embodiments, which are given by way of example only and which are not intended to limit the scope of the disclosure, that being determined solely by the appended claims. In particular the different features from different embodiments may be interchanged, where appropriate.Each of the embodiments of the disclosure described above can be implemented solely or as a combination of a plurality of the embodiments. Also, features from different embodiments can be combined where necessary or where the combination of elements or features from individual embodiments in a single embodiment is beneficial.

[0508] In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously used.

Claims

CLAIMS1. A method of encapsulating a visual volumetric data bitstream into a media file, the method being implemented by a processing device, the visual volumetric data bitstream comprising a plurality of sub-bitstreams corresponding to a plurality of components of the visual volumetric data, the method comprising:- obtaining a subpart of at least one component sub-bitstream enabling independent decoding of a region of the at least one component sub-bitstream, the region of the at least one component sub-bitstream being associated with a region of the visual volumetric data;- generating a dedicated track associated with the region of the visual volumetric data and comprising the obtained subpart of the at least one component sub-bitstream;- generating a base track comprising data that is common to a plurality of the component sub-bitstreams of the visual volumetric data, the base track referencing the dedicated track; and- generating a media file comprising the base track and the dedicated track.

2. The method according to claim 1, wherein the method further comprises:- generating, in the media file, information associating the dedicated track with the region of the visual volumetric data.

3. The method according to claim 1 or 2, wherein a sub-bitstream of the at least one component sub-bitstream corresponds to an atlas sub-bitstream, and wherein a subpart of the atlas sub-bitstream is at least one atlas tile.

4. The method according to claim 3, wherein the dedicated track includes atlas tile data with submesh data and / or attribute tile data.

5. The method according to any one of claim 1 to 4, wherein the dedicated track contains information indicating that samples of the dedicated track correspond to a multiplex of data units of visual volumetric data comprising subparts of component sub-bitstreams from different components.

6. The method according to any one of claims 1 to 5, wherein when several dedicated tracks multiplexing subparts of at least two component sub-bitstreams of a visual volumetric data bitstream are generated, the subparts of the at least two component sub-bitstreams enabling independent decoding are time-aligned between the different dedicated tracks.

7. The method according to any one of claims 1 to 6, wherein the plurality of subbitstreams includes at least one component sub-bitstream that does not enable independent decoding of a region of this at least one component sub-bitstream, and the base track further comprises this at least one component sub-bitstream that does not enable independent decoding.

8. The method according to any one of claims 1 to 6, wherein the plurality of subbitstreams includes at least one component sub-bitstream that does not enable independent decoding of a region of this at least one component sub-bitstream; andwherein the method also comprises;- generating, in the media file, at least one common track including this at least one component sub-bitstream that does not enable independent decoding, - generating, in the media file, at least one track reference referencing the at least one common track, to associate the subpart of the at least one component sub-bitstream enabling independent decoding to the at least one component subbitstream that does not enable independent decoding.

9. The method according to any one of claims 1 to 8, wherein the region is a dynamic region whose related subpart of the at least one component subbitstream is modified over the time, the method further comprising: generating a dynamic region description associated with an identifier of the dynamic region.

10. The method according to claim 9, wherein at least a part of the dynamic region description is contained in the base track.

11. The method according to claim 9 or 10, wherein at least a part of the dynamic region description is contained in a timed metadata track, the media file further comprising the timed metadata track.

12. The method according to any one of claim 9 to claim 11, wherein at least a part of the dynamic region description is contained in the dedicated track.

13. The method according to any one of claim 9 to claim 12, wherein the dynamic region description is declared in a sample group description.

14. The method according to any one of claim 9 to claim 13, wherein the modification over the time of the dynamic region comprises:an addition of a subpart of at least one component sub-bitstream to the dynamic region; ora removal of a subpart of at least one component sub-bitstream from the dynamic region.

15. The method according to any one of claim 9 to claim 14, wherein the modification over the time of the dynamic region may comprises:a modification of the identifier of the dynamic region, and, if the dynamic region description comprises a description of a plurality of dynamic regions, each of the dynamic region description being associated with an identifier, a re-allocation of an identifier of another dynamic region to the dynamic region.

16. The method according to any one of claim 9 to claim 15, wherein the modification over the time of the dynamic region comprises:a modification of at least one parameter of the dynamic region, the least one parameter being from at least one of:a description of a bounding box of the dynamic region,a description of a mapping of the identifier of the dynamic region to atlas tiles identifiers, anda list of objects included in the dynamic region.

17. A method of parsing a media file encapsulating a visual volumetric data bitstream, the method being implemented by a processing device, the visual volumetric data bitstream comprising a plurality of sub-bitstreams corresponding to a plurality of components of the visual volumetric data, the visual volumetric data bitstream including a region of visual volumetric data associated with a subpart of at least one component sub-bitstream enabling independent decoding of a region of thisat least one component sub-bitstream,the media file including:- a dedicated track associated with the region of the visual volumetric data and comprising the subpart of the at least one component sub-bitstream;- a base track comprising data that is common to a plurality of the component sub-bitstreams of the visual volumetric data, the base track referencing the dedicated track; andwherein the method comprises:- parsing the base track to access the dedicated track associated with at least one region of the visual volumetric data to decode.

18. A device comprising a processing unit configured for carrying out each of the steps of the method according to any one of claims 1 to 17.

19. A computer program product for a programmable apparatus, the computer program product comprising a sequence of instructions for implementing a method according to any one of claims 1 to 17, when loaded into and executed by the programmable apparatus.

20. A computer-readable storage medium storing instructions of a computer program product according to claim 18.