Method, apparatus and computer program for improving transmission and / or storage of point cloud data

By segmenting point cloud data into multiple subsets and generating corresponding input and derivative items, encapsulation and parsing using the ISOBMFF media file format, the problem of low encapsulation and decoding efficiency in the prior art is solved, and efficient parallel processing is achieved.

CN120051799APending Publication Date: 2025-05-27CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073116.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is less efficient when encapsulating and decoding point cloud data, and it is difficult to effectively improve the processing efficiency of encoding and decoding.

Method used

By segmenting point cloud data into multiple subsets and generating input and derivative items describing these subsets, the ISOBMFF media file format is used for encapsulation and parsing, and independent processing of the subsets of point cloud data is achieved.

Benefits of technology

The processing efficiency of point cloud data encoding and decoding is improved, and the subset of point cloud data can be processed in parallel, thereby improving the overall processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051799A_ABST
    Figure CN120051799A_ABST
Patent Text Reader

Abstract

At least one embodiment of a method for encapsulating point cloud data into an ISOBMFF-based media file is provided. After a plurality of subsets of the point cloud data have been obtained, a plurality of entries are generated, each entry describing a subset of the point cloud data of the plurality of subsets, and a derived entry is generated, the derived entry comprising descriptive data of a spatial composition of the plurality of entries, the derived entry being associated with the plurality of entries by an entry reference. The plurality of input items, item references, and derived items are then encapsulated in a media file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods, devices, and computer programs for improving the transmission and / or storage of point cloud data, for example, by defining a volume term as a grid of multiple volume terms or by using the fusion of several volume terms. Background Art

[0002] The Moving Picture Experts Group (MPEG) is standardizing the compression and storage of point cloud data (also known as volumetric media data) information. Point cloud information consists of a set of 3D points with associated attribute information such as color, reflectance, and frame index.

[0003] In one aspect, MPEG-I Part-9 (ISO / IEC 23090-9) specifies Geometry-based Point Cloud Compression (G-PCC) and specifies the bitstream syntax for point cloud information. According to MPEG-I Part-9, a point cloud is an unordered list of points that includes geometric information, optional attributes, and associated metadata. The geometric information describes the position of the points in a three-dimensional Cartesian coordinate system. Attributes are typed properties of the points, such as color or reflectance. Metadata is an information item used to interpret the geometric information and attributes. The G-PCC compression specification (MPEG-I Part-9) defines specific attributes such as the frame index attribute or the frame number attribute, which have reserved attribute tag values (3 for indicating the frame index and 4 for indicating the frame number attribute). It should be noted that, according to MPEG-I Part-9, a point cloud frame is a set of points at a specific time instance. A point cloud frame can be divided into one or more ordered sub-frames. Still in MPEG-I Part-9, the point cloud frame is indicated by the FrameCtr variable, which may use frame boundary markers or parameters in some data unit headers (the frame_ctr_lsb syntax element).

[0004] On the other hand, MPEG-I Part-18 (ISO / IEC 23090-18) specifies the following media format based on ISO / IEC 14496-12 (ISOBMFF), which enables the storage and delivery of geometry-based point cloud compression data. It also supports the flexible extraction of geometry-based point cloud compression data during delivery and / or decoding. According to MPEG-I Part-18, a point cloud frame is encapsulated in one or more G-PCC tracks, and the samples in the G-PCC track correspond to a single point cloud frame. Each sample includes one or more G-PCC units belonging to the same presentation time. A G-PCC unit is a type-length-value (TLV) encapsulation structure that contains one of SPS, GPS, APS, tile inventory, frame boundary marker, geometry data unit, and attribute data unit. The syntax of the TLV encapsulation structure is defined in Appendix B of ISO / IEC 23090-9.

[0005] Closely related to these standards, MPEG-I Part-5 (ISO / IEC 23090-7) specifies visual volume video-based coding (V3C) and video-based point cloud compression (V-PCC). It specifies a general mechanism for encoding visual volume frames by converting 3D volume information into a set of 2D images and associated data. Additionally, it specifies the application of this general mechanism to the point cloud representation of visual volume frames. MPEG-I Part-10 (ISO / IEC 23090-10) specifies a media format for storing and delivering visual volume video-based coded data in a file based on ISO / IEC 14496-12 (ISOBMFF).

[0006] Both MPEG-I Part-18 and MPEG-I Part-10 allow the storage of timed or untimed data. One or more tracks defined by ISOBMFF are used to store timed data. One or more items defined by ISOBMFF are used to store untimed data.

[0007] Although the ISO base media file format has proven to be efficient for encapsulating point cloud data, there is a need to improve the encapsulation efficiency, such as improving the efficiency of encoding the point cloud data to be encapsulated and / or the efficiency of decoding the encapsulated point cloud data. Summary of the Invention

[0008] The present invention has been designed to solve one or more of the foregoing problems. The present invention describes derived items that can be used to combine 3D items, for example, using a regular grid or by specifying the positions of each 3D item, such that point cloud data can be encoded and / or decoded, particularly independently (e.g., in parallel).

[0009] According to a first aspect of the present invention, there is provided a method for encapsulating point cloud data into an ISOBMFF-based media file, the method comprising:

[0010] Obtaining a plurality of subsets of the point cloud data;

[0011] Generating a plurality of input items, each input item describing a subset of the point cloud data in the plurality of subsets;

[0012] Generating a derived item, the derived item including descriptive data of the spatial composition of the plurality of input items, the derived item being associated with the plurality of input items by item references;

[0013] Encapsulating the plurality of input items, the item references, and the derived item in the media file.

[0014] Thus, the method of the present invention enables the processing efficiency of processing point cloud data (e.g., encoding point cloud data) to be improved by enabling independent processing of subsets of the point cloud data.

[0015] According to some embodiments, the descriptive data further includes an indicator indicating whether the point clouds corresponding to the point cloud data in each subset of the point cloud data use the same reference system.

[0016] According to some embodiments, the descriptive data further includes an indicator indicating whether the bounding boxes of the point clouds corresponding to the point cloud data described by each of at least two input items among the plurality of input items overlap.

[0017] According to some embodiments, at least one indicator is associated with at least one input item among the plurality of input items, the at least one indicator indicating that the bounding box corresponding to the at least one input item overlaps with the bounding box associated with another input item among the plurality of input items.

[0018] According to some embodiments, the descriptive data further includes an indicator for describing a three-dimensional array of cells, each input item among the plurality of input items corresponding to one cell in the cells.

[0019] According to some embodiments, the descriptive data further includes an indicator indicating the order of the plurality of input items in the three-dimensional array of cells.

[0020] According to some embodiments, the descriptive data further includes an indicator for describing the size of at least one cell in the cells.

[0021] According to some embodiments, the size of the cell is determined based on the bounding box of the point cloud corresponding to the point cloud data described by each of the plurality of input items.

[0022] According to some embodiments, the descriptive data further includes an indicator indicating that at least one of the cells does not include any points defined in the point cloud data.

[0023] According to some embodiments, at least one item property is associated with at least one of the input item and the derived item, and the at least one item property is used to describe at least one spatial operation to be performed on the points defined in the point cloud data.

[0024] According to some embodiments, the derived item includes at least one indicator for describing at least one spatial operation to be performed on the points defined in the point cloud data.

[0025] According to some embodiments, the method further includes: obtaining point cloud data and dividing the obtained point cloud data into the plurality of subsets of the point cloud data.

[0026] According to a second aspect of the present invention, there is provided a method for parsing an ISOBMFF-based media file encapsulating point cloud data, the method including:

[0027] Obtaining a derived item from the media file, the derived item including descriptive data of the spatial composition of a plurality of input items, and the derived item is associated with the plurality of input items through item references;

[0028] Obtaining the plurality of input items from the media file according to the item references, each input item describing a subset of the point cloud data of the plurality of subsets;

[0029] Obtaining the plurality of subsets of the point cloud data from the media file according to the plurality of input items; and

[0030] Generating point cloud data according to the descriptive data, the point cloud data including the point cloud data of the plurality of subsets.

[0031] Therefore, the method of the present invention enables the processing efficiency of processing point cloud data (e.g., decoding point cloud data) to be improved by enabling independent processing of subsets of the point cloud data.

[0032] According to some embodiments, the method further includes: determining whether the point clouds corresponding to the point cloud data of each subset of the point cloud data use the same reference system according to the indicator of the descriptive data.

[0033] According to some embodiments, the method further includes: determining, based on an indicator of the descriptive data, whether bounding boxes of point clouds corresponding to respective ones of at least two of the plurality of input items overlap.

[0034] According to some embodiments, at least one indicator is associated with at least one of the plurality of input items, the at least one indicator indicating that a bounding box corresponding to the at least one input item overlaps a bounding box associated with another input item of the plurality of input items.

[0035] According to some embodiments, the method further includes: determining, based on an indicator of the descriptive data, a three-dimensional array of cells, with each of the plurality of input items corresponding to one of the cells.

[0036] According to some embodiments, the method further includes: determining, based on an indicator of the descriptive data, an order of the plurality of input items in the three-dimensional array of cells.

[0037] According to some embodiments, the method further includes: determining, based on an indicator of the descriptive data, a size of at least one of the cells.

[0038] According to some embodiments, the method further includes: determining, based on an indicator of the descriptive data, that at least one of the cells does not include any points defined in the point cloud data.

[0039] According to some embodiments, the method further includes: applying a spatial operation to points defined in the point cloud data based on at least one item property associated with at least one of the input item and the derived item.

[0040] According to some embodiments, the method further includes: applying at least one spatial operation to be performed on points defined in the point cloud data based on at least one indicator of the derived item.

[0041] According to other aspects of the present invention, there is provided a processing device including a processing unit configured to perform the respective steps of the above method. Other aspects of the present disclosure have optional features and advantages similar to those of the above first aspect and second aspect.

[0042] At least part of the method according to the present invention may be computer-implemented. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, all of which may be commonly referred to herein as a "circuit", "module" or "system". Furthermore, the present invention may take the form of a computer program product embodied in any tangible expression medium having computer-usable program code embodied in the medium.

[0043] Since the present invention may be implemented in software, the present invention may be embodied as computer-readable code for providing to a programmable device on any suitable carrier medium. The tangible carrier medium may include storage media such as a floppy disk, CD-ROM, hard drive, tape device or solid state memory device, etc. The transient carrier medium may include signals such as electrical, electronic, optical, acoustic, magnetic or electromagnetic signals (e.g., microwave or RF signals), etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Embodiments of the present invention will now be described, by way of example only and with reference to the following drawings, in which:

[0045] Figure 1a and Figure 1b illustrates examples of steps and file structures for storing and / or transmitting point cloud data (or 3D data) and for accessing (e.g., displaying) the stored and / or transmitted point cloud data (or a portion of the stored and / or transmitted point cloud data) according to some embodiments of the present invention;

[0046] Figure 2 illustrates an example of dividing a 3D volume associated with point cloud data into a plurality of 3D sub-volumes according to a 3D mesh;

[0047] Figure 3 illustrates examples of steps for storing or transmitting point cloud data as a mesh of input G-PCC terms according to some embodiments of the present invention;

[0048] Figure 4 illustrates examples of steps for decoding and rendering point cloud data stored as a mesh of input G-PCC terms according to some embodiments of the present invention;

[0049] Figure 5 illustrates examples of steps for storing and / or transmitting point cloud data as a 3D fusion term of input G-PCC terms according to some embodiments of the present invention;

[0050] Figure 6Illustrates an example of steps for decoding and rendering point cloud data stored as an input G-PCC item for 3D fusion according to some embodiments of the present invention; and

[0051] Figure 7 Is a schematic representation of an example of a data processing apparatus configured to implement in whole or in part some embodiments of the present disclosure. Detailed description

[0052] The limitation of the above-mentioned specification is that it does not solve the problem of combining several items representing 3D data into another item representing 3D data. Examples of such combinations include, but are not limited to: combining several items into a 3D mesh, where each item represents the content of a cell from the mesh; or combining several items into a 3D merged item, where each item is located at a position specified by the 3D merged item. The advantage of defining an item representing 3D data as a combination of several items is that the combined items can be encoded or decoded independently in a parallel manner.

[0053] Figure 1a Illustrates an example of steps for storing and / or transmitting point cloud data (or 3D data) and for accessing (e.g., displaying) the stored and / or transmitted point cloud data (or a part of the stored and / or transmitted point cloud data) according to some embodiments of the present invention.

[0054] As illustrated, the point cloud data 100 is segmented into multiple subsets of point cloud data (step 105) or obtained as multiple subsets of point cloud data (e.g., different spatial subsets or subsets obtained from different sensors or corresponding to different viewpoints), and each subset is processed in parallel (steps 110-1 to 110-m) to, for example, perform parallel encoding. For each subset of point cloud data, the processed data is described in an input item that can be referred to as a point cloud item. The point cloud item describes point cloud data that can be encoded, for example, using MPEG-I Part-5 or MPEG-I Part-9 and can be referred to as a V-PCC (or V3C) point cloud item or a G-PCC point cloud item, respectively. When used as a derived input or within a group of items, the point cloud item can also be referred to as an input point cloud, a V-PCC (or V3C) input item, or a G-PCC input item. When only "input item" or "input point cloud" is used in the context of the present invention, regardless of its encoding format, it means a point cloud item. Then, the resulting set of input items is encapsulated (step 115) to generate an encapsulated data file 120, which can be stored locally or transmitted to a remote server for storage and / or processing, for example, to display the point cloud data (or a portion of the point cloud data). The encapsulated data file includes input items that describe subsets of the processed point cloud data, the subsets of the processed point cloud data, and derived items that include descriptive data that enables the combination of the point cloud data of the respective subsets. More generally, a derived item is an item that defines an operation to be performed on an input item associated with the derived item via an item reference of the 'dimg' type. For example, such descriptive data can represent the spatial composition of the point cloud data described by the input item. Additionally, the encapsulation file includes item references for associating the derived items with the input items.

[0055] The encapsulated data file can be in the ISOBMFF format.

[0056] To access the point cloud data (or a portion of the point cloud data) encapsulated in the encapsulated data file, the encapsulated data file is parsed (step 125) to obtain the derived items. Using the information from the derived items, the input items that describe subsets of the processed point cloud data can be accessed, such that the point cloud data can be accessed and processed in parallel (steps 130-1 to 130-n) to, for example, perform parallel decoding. Then, the processed point cloud data is combined (step 135) based on the information from the derived items to generate valid point cloud data 140. As another example, the point cloud data can be decoded in parallel and stored in a memory structure suitable for 3D rendering (steps 130-1 to 130-n), and then combined (step 135) by 3D rendering to generate a 3D rendering of the point cloud data.

[0057] It should be noted that steps 105, 110-1 to 110-m, and 115 can be executed in server 145 or in a number of servers (or processing devices). Similarly, steps 125, 130-1 to 130-n, and 135 can be executed in server 150, the client, or a number of servers or processing devices.

[0058] Figure 1b Illustrates an example of a file structure for storing and / or transmitting point cloud data (or 3D data) and for accessing (e.g., displaying) the stored and / or transmitted point cloud data (or a portion of the stored and / or transmitted point cloud data) according to some embodiments of the present invention.

[0059] As illustrated, the encapsulated data file 120 includes a derived item 155 (e.g., an item including a GPCCGrid structure in its body) identified in an item reference 160 (e.g., a SingleItemTypeReferenceBox of type 'dimg' in an ItemReferenceBox 'iref' as described below), and the item reference also identifies a plurality of input items 165-1 to 165-n (e.g., input G-PCC items, V-PCC items, or V3C items) for describing subsets 170-1 to 170-n of the point cloud data. Thus, the point cloud can be stored and / or transmitted as multiple subsets of the point cloud data and can be rendered in a manner that enables parallel processing of the subsets of the point cloud data.

[0060] For illustrative purposes, two use cases are described below. According to the first use case, the point cloud data is divided into multiple spatial subsets according to a 3D grid, and according to the second use case, the point cloud data is divided into multiple subsets based on another criterion (e.g., the sensor used to acquire the point cloud data) (i.e., according to a 3D fusion criterion).

[0061] 3D grid

[0062] According to a particular embodiment, a 3D volume associated with the point cloud data is divided into multiple 3D sub-volumes, and the point cloud data of each of these 3D sub-volumes is independently processed and described as a particular item referred to as an input item.

[0063] Note that for illustrative purposes, the following description mainly focuses on point cloud data encoded as a G-PCC stream according to MPEG-I Part-9 and stored as a G-PCC item according to MPEG-I Part-18. However, it can be easily extended to point cloud data encoded as a V-PCC stream according to MPEG-I Part 5 and stored according to MPEG-I Part 10, or more generally to any 3D content stored according to ISOBMFF.

[0064] Figure 2 An example is illustrated of dividing a 3D volume associated with point cloud data into multiple 3D sub-volumes according to a 3D grid.

[0065] According to the illustrated example, the 3D grid 200 includes 12 cells arranged in a 3×2×2 array. Naturally, the 3D volume associated with the point cloud data can be divided differently into the same or different numbers of cells.

[0066] The coordinates of the points within the 3D volume are defined as a function of a reference frame, such as the reference frame 205 defined by three axes (X-axis 205-1, Y-axis 205-2, and Z-axis 205-3). The origin of the reference frame 205 can correspond to a corner of the grid cell or can be located elsewhere (as illustrated). The reference frame 205 defines a global reference frame shared by all points in the 3D volume. Alternatively, the coordinates can be represented by three angles such as (azimuth, elevation, and tilt angle) or (yaw, pitch, and roll angle).

[0067] Additionally, a local reference frame can be used within each cell of the 3D grid to identify the points within that cell. For illustrative purposes, the origin of such a local reference frame can correspond to a corner of the cell, such as the corner closest to the origin of the global reference frame or the corner with the smallest coordinates in the global reference frame. The axes of the local reference frame can be parallel to the axes of the global reference frame. An example of a local reference frame associated with a cell of the 3D grid is illustrated using reference numeral 210. Still for illustrative purposes, either the global reference frame 205 or the local reference frame 210 can be used to define the position of the point 215.

[0068] According to a particular embodiment, the 3D grid G-PCC item can be represented as a derived point cloud item with an item_type of a specific value (e.g., the value 'grd3'). The 3D grid can be used to combine one or more than one input G-PCC items in the grid. Thus, an item with an item_type value of 'grd3' defines a derived point cloud item whose reconstructed point cloud is formed from one or more than one input point clouds in a given 3D grid order. The data associated with the derived point cloud item can specify some characteristics of the grid and can have the following structure:

[0069]

[0070]

[0071] Among them, the version field signals the version of the structure, and the flags field is a mapping of flags for some options that can signal the structure. If version or flags are not defined separately, version or flags shall be equal to 0.

[0072] The width_minus_one, height_minus_one, and depth_minus_one fields represent the number of cells along the X, Y, and Z axes in the grid, respectively.

[0073] In this grid, the cell located at index i x 、i y and i z has coordinates that extend from c x *i x to c x *(i x +1) along the X axis, from c y *i y to c y *(i y +1) along the Y axis, and from c z *i z to c z *(i z +1) along the Z axis, where c x 、c y and c z are the sizes of the cells along the X, Y, and Z axes.

[0074] The same_origin flag indicates whether the point cloud data associated with the input G-PCC items to be combined in the grid uses the same reference system. If the value of the same_origin flag is set to the first value (e.g., 1), the same reference system or coordinate system (e.g., the global reference system 205) is used to define all input G-PCC items. In this case, the input G-PCC items can be combined in the grid without applying translation or transformation to them to insert them into the grid. Otherwise, if the value of the same_origin flag is set to the second value (e.g., 0), each input G-PCC item uses its own reference system (e.g., the local reference system of the cell where the point cloud data of the input G-PCC item is located) to define. In the latter case, translation can be applied to the coordinates of the points of the input G-PCC items to calculate their coordinates in the global reference system associated with the grid for insertion into the grid. Such translation can be as follows:

[0075] x g =x l +c x *i x

[0076] y g = y l + c y * i y

[0077] z g = z l + c z * i z

[0078] where x g , y g and z g are the coordinates of the point in the global reference frame, and x l , y l and z l are the coordinates of the point in the local reference frame of the input G-PCC item. As described above, c x , c y and c z are the sizes of the cells along the X, Y, and Z axes, and i x , i y and i z are the zero-based indices of the position of the input G-PCC item in the grid.

[0079] The translation defined above transforms the coordinates of a point represented in the local reference frame of the input G-PCC item containing the point into coordinates represented in the local reference frame of the grid cell located at position (0, 0, 0) (which can be used as the global reference frame of the grid). Naturally, another global reference frame can be used, for example, by defining a translation from the local reference frame of the grid cell located at position (0, 0, 0) to another global reference frame by leveraging the item properties associated with the 3D grid-derived item. As another example, such a translation can be specified as a 3D vector within the GPCCGrid structure, such as:

[0080]

[0081] where the origin_precision field specifies the precision used to encode the position of the global origin, and the global_origin field is a 3D vector representing the position of the global origin, which can be represented using the Vector3 structure specified by MPEG-I Part-7 (ISO / IEC 23090-7). As a variant, the position of the origin of the local reference frame of the grid cell located at position (0, 0, 0) in the global reference frame can be encoded.

[0082] Possibly, when the value of the same_origin flag is set to the second value (0 in the above example), the input G-PCC item can be centered within its grid cell or aligned with one or more faces of its grid cell.

[0083] Possibly, when the value of the same_origin flag is set to the first value (1 in the above example), the input G-PCC item can have a size smaller than the size of the grid cell. The position of the entire grid can be determined based on the bounding box specified in the item properties associated with the 3D grid item. The position of the entire grid can be determined by finding the input G-PCC item with the largest size along each axis and determining the position of the grid cell containing the input G-PCC item along that axis from the range of the G-PCC item along that axis. For example, the position of the grid cell along that axis can be determined such that the input G-PCC item is centered within the grid cell along that axis, or such that the input G-PCC item is aligned with one of the boundaries of the grid cell along that axis. Also, for example, the position of the entire grid can be specified as a 3D vector within the GPCCGrid structure using the following structure:

[0084]

[0085] where the grid_precision field specifies the precision used to encode the position of the grid, and the grid_position field is a 3D vector representing the position of the corner of the grid with the minimum coordinates.

[0086] The same_origin field highlights the difference between a 2D image and a 3D point cloud. In a 2D image, the coordinates of the pixels are implicitly specified: the pixels are signaled according to a row-major scan order from top left to bottom right. In a 3D point cloud, the coordinates of the points are explicitly specified by encoding the position of the points as x, y, z. This difference lies in the nature of the pixels in a 2D image, where each pixel corresponds to a square or rectangular region associated with an information item (usually color and possibly transparency). In a 3D point cloud, the points have no volume and are thus point-by-point. Additionally, 3D point clouds are typically sparse: there are many empty regions in a 3D point cloud. Therefore, in order to achieve an accurate and efficient encoding of a 3D point cloud, the position of the points or point sets can be explicitly specified. As a result, for a 2D image, the pixel coordinates always start at (0,0) of the top left pixel and end at (width, height) of the bottom right pixel, while for a 3D point cloud, the point coordinates can use any reference frame.

[0087] In the GPCC Grid structure, the morton_order field indicates the sorting of the input G-PCC items. According to some embodiments, the input G-PCC items are signaled in a SingleItemTypeReferenceBox of type 'dimg' within the ItemReferenceBox 'iref'. In this SingleItemTypeReferenceBox, the value of the from_item_ID field identifies the derived point cloud item. In these embodiments, the input G-PCC items can be retrieved by looking up a SingleItemTypeReferenceBox of type 'dimg', where the value of the from_item_ID field of this SingleItemTypeReferenceBox of type 'dimg' is the identifier associated with the 3D mesh G-PCC item. The value of reference_count is preferably equal to (width_minus_one + 1) * (depth_minus_one + 1) * (height_minus_one + 1). Still according to some embodiments, the value of the to_item_ID field identifies the input G-PCC item. The SingleItemTypeReferenceBox can be referred to as an item reference.

[0088] If the value of the morton_order field is set to a first value (e.g., 0), the input G-PCC items are sorted in line scan order by first incrementing X, then incrementing Y, and then incrementing Z. In some variations, the axis arrangement of the line scan order can be different. If the value of the morton_order field is set to a second value (e.g., 1), the input G-PCC items are sorted in Morton order (or Z scan order). Using Morton order enables better preservation of the local characteristics of the data within the grid: with line scan order, the input G-PCC items corresponding to adjacent cells may be far apart, while with Morton order, the positions of the G-PCC items are closer.

[0089] The sorting of the input G-PCC items is the order in which they are listed in the to_item_ID field of the SingleItemTypeReferenceBox. Preferably, this is also the order in which their associated data is stored in a file or transmitted over a network.

[0090] According to some embodiments, the input G-PCC items have the same size. This size is the size of the cells in the grid. This size can be specified by the item properties of one or more than one input G-PCC item in the input G-PCC items (e.g., the bounding box item property or the size item property). This size can be specified in the GPCCGrid structure. If the size is specified both in the item properties of one or more than one input G-PCC item in the input G-PCC items and in the GPCCGrid structure, the specified sizes are preferably the same. If the value of the has_cell_size field of the GPCCGrid structure is set to a first value (e.g., 0), the size of the cells of the grid is not specified within the GPCCGrid structure. In this case, all the bounding boxes of the input point cloud should have the same size, which is the size of the cells. If it is desired that the input point cloud is not of a consistent size, derived point cloud items (e.g., the 'iden' type representing the identity transformation) that scale or crop them as needed to make them consistent can be used; however, other specifications can limit whether derived point cloud items are admissible as inputs to point cloud grid derived items.

[0091] If the value of the has_cell_size field of the GPCCGrid structure is set to a second value (e.g., 1), the size of the cells of the grid is specified within the GPCCGrid structure by the cell_size field. The precision field specifies the precision used to encode the size of the cells. The cell_size field is a 3D vector representing the size of the cells, which can be represented using the Vector3 structure specified by MPEG-I Part-7 (ISO / IEC 23090-7).

[0092] When removing the items of the input point cloud that are marked as point cloud grid items, it may be necessary to rewrite the content of the point cloud grid items.

[0093] Preferably, all the points from the input G-PCC item are located within the boundaries of the grid cell corresponding to the input G-PCC item. When the same_origin flag is set to a first value (e.g., 1), all the points from the input G-PCC item preferably have coordinates within the boundaries of the grid cell where the input G-PCC item is located. When the same_origin flag is set to a second value (e.g., 0), after applying the above translation, all the points from the input G-PCC item preferably have coordinates within the boundaries of the grid cell where the input G-PCC item is located.

[0094] Figure 3An example of steps for storing or transmitting point cloud data as a grid of input G-PCC items according to some embodiments of the present invention is illustrated.

[0095] As illustrated, the first step involves obtaining point cloud data for storage or transmission (step 300). This point cloud data can be obtained directly from sensors such as LiDAR, or can be obtained by processing one or more point clouds acquired by one or several sensors. For example, several point clouds can be captured, registered, and transformed by LiDAR so that they use the same reference system, and then combined into a single point cloud.

[0096] Next, the target characteristics of the grid to be used are obtained (step 305). According to some embodiments, the size of each cell of the grid is obtained, and using this cell size and the spatial size of the point cloud obtained in step 300, the number of cells in each dimension is calculated. As a variant, the number of cells in each dimension can be obtained, and the size of each cell can be calculated using the spatial size of the point cloud.

[0097] In addition, the sorting of the input items in the grid can be obtained. As described above, this sorting can be, for example, line scan order or Morton order. Note that if no sorting is provided, a default sorting can be defined, such as line scan sorting.

[0098] Furthermore, an indication can be obtained for indicating whether the same origin is to be used for all input G-PCC items. Possibly, when the same origin is used for all input G-PCC items, the translation to be applied to these input G-PCC items before encoding all input G-PCC items can be obtained. According to some embodiments, in the case where the same origin is not used for all input G-PCC items, the translation to be applied to the reconstructed grid can be obtained.

[0099] It can be observed that step 305 can be performed before step 300 or simultaneously with step 300.

[0100] Next, according to the cells determined based on the grid characteristics obtained in step 305, the point cloud data obtained in step 300 is divided into subsets of point cloud data (step 310). As a result, point clouds corresponding to the subsets of the point cloud obtained in step 300 are obtained for each cell of the grid or for some cells of the grid, because the grid can be sparse (i.e., it can include some cells without any points from the point cloud data).

[0101] Next, each subset of the point cloud data is encoded (step 315). According to some embodiments, this encoding step is performed independently (e.g., in parallel) for each subset of the point cloud data. Still according to some embodiments, the encoding step is based on the MPEG-I Part-9 specification.

[0102] During this step, if the same origin is used for all input G-PCC items, the translation that may be obtained at step 305 can be applied to the points of the point cloud data before or during encoding the point cloud data. If the same origin is not used for the input G-PCC items, the points of each subset of the point cloud data can be translated according to the local reference frame associated with the corresponding cell of the grid before or during encoding the point cloud data.

[0103] Next, a structure for encapsulating the grid and the encoded point cloud data is created (step 320). For example, a GPCCGrid structure as described above can be created to describe the grid and store its characteristics. Additionally, G-PCC items can be created to describe each subset of the encoded point cloud data. Possibly, each subset of the encoded point cloud data can be described in a single G-PCC item or within several G-PCC items. As described above, a SingleItemTypeReferenceBox of type 'dimg' can be created to associate the input G-PCC items that describe the subsets of the encoded point cloud data with the item representing the grid (i.e., the derived item). The sorting of the input G-PCC items within the SingleItemTypeReferenceBox can depend on the sorting obtained at step 305.

[0104] Note that step 320 (or some of its sub-steps) can be performed in parallel with the encoding step 315 or as part of the encoding step 315.

[0105] Next, the subsets of the point cloud data encoded in step 315 and the encapsulation structure created in step 320 (which includes item references (e.g., a SingleItemTypeReferenceBox of type 'dimg') that establish associations between the derived item and the input items) are stored in an encapsulated data file (e.g., an ISOBMFF file). Before being stored, the subsets of the encoded point cloud data can be sorted according to the sorting obtained at step 305. Possibly, the ISOBMFF file is stored in a memory unit (e.g., random access memory) or a hard disk drive. Possibly, the ISOBMFF file is sent via a communication network.

[0106] It can be observed that encoding the point cloud data as a grid can advantageously improve the encoding speed of the point cloud data because several encoding operations can be performed in parallel.

[0107] According to some embodiments, an identity derived point cloud item may be represented as an item with an item_type having a specific value (e.g., the value 'iden'). This identity derived point cloud item can be used when it is desired to derive point cloud items using transform properties. The derived point cloud item has no item body (i.e., no extent), and the reference_count referenced by the 'dimg' item for the 'iden' derived point cloud item is equal to 1.

[0108] According to some embodiments, the point cloud data is obtained in chunks in step 300 (and not all chunks are obtained simultaneously). According to these embodiments, one or more chunks of the point cloud data are processed independently of the remaining chunks. In step 310, one or more chunks are divided into several subsets of the point cloud data. Then, for each subset of the point cloud data, a test is performed to determine whether the subset of the point cloud data contains all the data of its corresponding grid cell or whether some of the cell data is missing. If the subset of the point cloud data contains all the data of its corresponding grid cell, the subset of the point cloud data can be encoded as described above with respect to step 315, a next encapsulation structure can be created as previously described with respect to step 320 to signal the subset, and the subset of the encoded point cloud data and its encapsulation structure can be stored as described with respect to step 325.

[0109] Preferably, in these embodiments, the boundaries between the grid cells are selected such that the boundaries between the chunks match the boundaries of the grid. For example, if the point cloud data is obtained from a rotating LiDAR and each chunk corresponds to a quadrant of the volume captured by the rotating LiDAR, the boundaries corresponding to these quadrants may define the boundaries between the grid cells.

[0110] Still according to some embodiments, the point cloud data to be stored as a grid of the point cloud data has been divided into subsets of the point cloud data, and each subset of the point cloud data has been encoded.

[0111] In such an embodiment, a subset of the encoded point cloud data is obtained in step 300, and grid characteristics are obtained in step 305.

[0112] The size of each cell of the grid can be obtained from the spatial size of one of the subsets of the encoded point cloud data obtained in step 300. Possibly, if the subsets of the encoded point cloud data have different spatial sizes, the size of the cells of the grid can be defined as the maximum spatial size of the subsets of the encoded point cloud data. For example, the size of the cells of the grid along the X-axis can be determined by identifying the maximum spatial size of the subset of the encoded point cloud data along the X-axis. The sizes of the grid along the Y-axis and along the Z-axis can be determined similarly. The spatial size of the subset of the encoded point cloud data can be obtained from the bounding box associated with the subset of the encoded point cloud data. It can be calculated according to the content of the encoded point cloud data.

[0113] Possibly, if the subsets of the encoded point cloud data have different spatial sizes, the size of the cells of the grid can be determined as the minimum spatial size, the average spatial size, or the median spatial size of the subsets of the encoded point cloud data.

[0114] Possibly, if the subsets of the encoded point cloud have different spatial sizes, each point cloud can be centered within its cell by applying a translation to it. This translation can be applied by re-encoding the point cloud or by defining the transformation properties associated with the input G-PCC item used to encapsulate the subset of the point cloud data.

[0115] Possibly, if the spatial size of the subset of the encoded point cloud data is larger than the size of the cell, it can be cropped to fit the cell. This cropping can be performed before or after applying the translation to the points of the subset of the encoded point cloud data.

[0116] The sorting of the input items can be obtained explicitly as described above. It can also be obtained implicitly (e.g., using the order in which the subsets of the encoded point cloud data are obtained in step 300).

[0117] An indication of whether to use the same origin for all input G-PCC items can be obtained explicitly as described above. It can also be determined based on the subsets of the encoded point cloud data obtained. For example, calculate the minimum coordinate values and the maximum coordinate values along each axis of each subset of the encoded point cloud data, and if the minimum coordinate values and the maximum coordinate values are the same for all encoded point clouds, it can be concluded that they do not have the same origin. Otherwise, it is determined that they have the same origin. Still for illustrative purposes, the sizes of each subset of the encoded point cloud data along each axis can be calculated. If the difference between the maximum minimum coordinate value of a subset of the encoded point cloud data along an axis and the minimum minimum coordinate value of another subset of the encoded point cloud data along the same axis is greater than or equal to the maximum size of the subset of the encoded point cloud data along the same axis, it can be concluded that the point clouds have the same origin.

[0118] Possibly, the translation can be obtained as described above.

[0119] Possibly, the positions of the points of each subset of the encoded point cloud data within the grid can be explicitly obtained. In this case, an indication of whether to use the same origin for the input G-PCC item can be obtained by comparing the coordinates used by the points of the subsets of the encoded point cloud data located in adjacent cells. For example, in the case where the points of the first subset of the encoded point cloud data are located at index i along an axis and the points of the second subset of the encoded point cloud data are located at index i + 1 along the same axis, if the maximum coordinate value of the points of the first subset along this axis is lower than the minimum coordinate value of the points of the second subset along this axis, it is concluded that the same origin is used for the different point clouds.

[0120] Possibly, when the input G-PCC items have the same origin, the positions of the points of each subset of the encoded point cloud data within the grid can be implicitly obtained. Using the size of the cells of the grid, the coordinate ranges of each cell of the grid can be calculated. Then, for the points of each subset of the encoded point cloud data, the minimum coordinate value and the maximum coordinate value along each axis are calculated. These minimum coordinate values and maximum coordinate values can be used to obtain the positions of the cells corresponding to these coordinates.

[0121] In this alternative, steps 310 and 315 are not performed. Depending on how the subsets of the encoded point cloud data are obtained in step 300, step 320 can be performed fully or partially. If the encoded point cloud data is obtained as a compressed stream (e.g., as specified in MPEG-I Part-9), input G-PCC items can be created to describe the subsets of the encoded point cloud data. If the subsets of the encoded point cloud data are obtained as input G-PCC items, these items can be used directly. The other parts of step 320 can be performed as described above.

[0122] In this alternative, step 325 is performed as described above.

[0123] Figure 4 Examples of steps for decoding and rendering point cloud data stored in a grid as input G-PCC items according to some embodiments of the present invention are illustrated.

[0124] As illustrated, the first step involves obtaining the encapsulated data file of the grid storing the input item (step 400), such as the ISOBMFF file of the grid storing the input G-PCC item. Additionally, an indication of the first item to be decoded and rendered (i.e., the grid item) can be obtained. If no indication of the first item to be decoded and rendered is obtained, the first item indicated in the encapsulated data file can be considered the first item to be decoded and rendered.

[0125] Next, obtain the mesh item to be decoded and rendered from the ISOBMFF file (step 405). This mesh item may be encapsulated in the ISOBMFF file according to the GPCCGrid structure described above. It can be parsed to obtain, for example, the size of the mesh, whether the input G-PCC items use the same origin, whether these input G-PCC items are sorted in line scan order or Morton order, and whether the size of the cells is included in the structure. If the size of the cells is included in the structure, it can be parsed from the ISOBMFF file to obtain this size.

[0126] Additionally, the SingleItemTypeReferenceBox of type 'dimg' associated with the mesh can be parsed to identify the input G-PCC items used by the mesh. If the size of the cells is not included in the GPCCGrid structure, it can use the size of the input G-PCC items to obtain it. For example, in a particular embodiment, the sizes of all input G-PCC items are the same, and the size of any input G-PCC item is used to obtain the size of the cells. In a variant, the size of the cells along an axis can be calculated as the maximum size of the input G-PCC items along that axis.

[0127] Then obtain a subset of the encoded point cloud data of the input G-PCC items identified in the SingleItemTypeReferenceBox of type 'dimg' associated with the mesh from the ISOBMFF file.

[0128] Next, decode and render the subset of the encoded point cloud data (step 410). Decoding of the subset of the encoded point cloud data can be performed, for example, according to the MPEG-I Part-9 specification, i.e., decoding the point cloud. Rendering may include generating structures in a memory unit (e.g., random access memory) to represent the decoded point cloud, transferring these structures to the graphics card, and generating a display of the point cloud.

[0129] If the same_origin field of the GPCCGrid structure used is set to a second value (i.e., 0 according to the example given above), the positions of the points of the decoded point cloud can be translated, for example, using the translation described above, into the reference frame corresponding to the mesh. Optionally, a further translation can be applied to the decoded points to change the origin of the reference frame used by the mesh. This translation can be applied while decoding the point cloud, after decoding the point cloud and before rendering, or as part of the rendering process.

[0130] If the points of the subset of the encoded point cloud data do not use the same origin, the points of the decoded point cloud can be translated before rendering. Optionally, this translation can be performed as part of decoding or as part of rendering.

[0131] Preferably, decoding and rendering are performed as independently as possible for each subset of the encoded point cloud data. For example, decoding can be performed in parallel for several subsets of the encoded point cloud data (i.e., for several point clouds), generating in a memory unit a structure for representing the decoded point cloud, and transferring these structures to a graphics card. Advantageously, the memory structures for representing the decoded point cloud can be organized in such a way that the memory structures in the most frequently used memory structures are only used for the decoding and rendering of one decoded point cloud. For example, the representation of the 3D space can be divided into several sub-volumes, each sub-volume being represented using one or more memory structures. These sub-volumes can be constructed such that each sub-volume corresponds to a cell of a grid. In this way, only the memory structures representing the associated sub-volume are accessed during the decoding of a subset of the encoded point cloud data, and only these memory structures are accessed during the decoding of this subset of the encoded point cloud data. In this example, some global memory structures can be used to organize the memory structures corresponding to each sub-volume. However, these global memory structures are mostly accessed at step 405 while parsing the grid, rather than while decoding a subset of the encoded point cloud data. Possibly, several sub-volumes can correspond to a cell of the grid, and each sub-volume only corresponds to a cell of the grid.

[0132] Decoding the point cloud of the grid that was previously encoded as a subset of point cloud data can advantageously increase the decoding speed of the point cloud, because several decoding operations can be performed in parallel. Additionally, a part of the rendering process can also be performed in parallel.

[0133] In some embodiments, step 410 can further optimize the decoding and rendering of the point cloud encoded as a grid. In these embodiments, the description of the grid contained in the GPCCGrid structure can be used to calculate the positions of the cells of the grid. Thus, the cells of the grid can be filtered according to their positions, so that only the cells relevant for rendering are selected. For example, only the cells corresponding to the point cloud data visible within the viewport used for rendering can be selected. As another example, only the cells corresponding to the point cloud data visible and close to the viewpoint within the viewport are selected. Memory structures for rendering the point cloud are created. These structures can be adapted to the selected cells. For example, the rendering space can be divided along boundaries corresponding to the boundaries between the grid cells. Thus, only the subsets of the encoded point cloud data corresponding to the selected cells are parsed from the ISOBMFF file. These subsets of the encoded point cloud data are decoded in parallel and stored within the memory structures, and then rendered again in parallel.

[0134] According to other embodiments, one or more input G-PCC items may extend beyond the boundaries of their grid cells. This may be useful, for example, when storing the output of a rotating LiDAR as a grid. In fact, a rotating LiDAR captures a point cloud by using an array of vertically aligned lasers. This array allows the capture of points in a vertical plane. To capture points in a volume, the array rotates along the vertical axis. Thus, a full rotation of the LiDAR can be divided into a grid of 4 cells, each cell corresponding to a quadrant of the volume scanned by the LiDAR. This allows points in each quadrant to be encoded as soon as they are captured, without waiting for the complete capture of the entire volume.

[0135] However, for some rotating LiDARs, the lasers are not perfectly aligned in the same vertical plane: there is a slight offset between the lasers. This may be due to an improvement in the compactness of the laser array. This means that the points captured during a quarter rotation of the LiDAR may extend beyond the corresponding quadrant. Although the captured points can be filtered and reassigned to the correct quadrant, this makes the encoding process more complex and introduces some latency in the encoding process. To solve this problem, the grid as described above can be modified to allow points of the input G-PCC item to extend beyond the boundaries of its grid cell. The structure describing the grid can be modified as follows to support this feature:

[0136]

[0137] where the overlap field indicates whether the input G-PCC item fits within its grid cell or can extend beyond the grid cell. If the value of this field is set to a first value (e.g., 1), then the points of the input G-PCC item can extend beyond its grid cell. Alternatively, if the value of this field is set to a second value (e.g., 0), then the points of the input G-PCC item are fully contained within its grid cell.

[0138] According to these embodiments, the encoder determines whether the input G-PCC item extends beyond the boundaries of its grid cell. This determination can be performed simultaneously with obtaining the grid characteristics at step 305 in Figure 3 When constructing a grid using the encoded point cloud, this determination can be performed by examining the characteristics of the encoded point cloud. Thus, if it is determined that some input G-PCC items extend beyond the boundaries of their grid cells, then when creating the GPCCGrid structure at step 320 in Figure 3 the overlap field is set to the first value (i.e., 1 according to the example given above). Otherwise, the flag is set to the second value (i.e., 0 according to the example given above).

[0139] Still according to these embodiments, the decoder is at Figure 4At step 405 in, it is determined whether the input G-PCC item extends beyond the grid cell by checking the value of the overlap field. If it is determined that the value of the overlap field is set to the first value (i.e., 1 according to the example given above), then at Figure 4 At step 410 in, the points of the input G-PCC item that are outside the boundaries of its cell can be moved to the structure corresponding to the region where these points are located. Optionally, the points of the input G-PCC item that are outside the boundaries of its cell can be discarded, which means that the input G-PCC item is cropped to fit the boundaries of its grid cell.

[0140] If it is determined that the value of the overlap field is set to the second value (i.e., 0 according to the example given above), then Figure 4 Step 410 in is implemented as described above (specifically, by performing decoding and rendering as independently as possible for each subset of the encoded point cloud data).

[0141] In a variation of these embodiments, the structure describing the grid can indicate whether an input G-PCC item that extends beyond the boundaries of its grid cell is to be cropped.

[0142] In another variation, the fact that an input G-PCC item extends beyond the boundaries of its grid cell can be specified with a finer granularity. For example, the overlap field can be associated with each input G-PCC item to indicate whether the input G-PCC item extends beyond the boundaries of its grid cell. As another example, the overlap field can be associated with each input G-PCC item for each adjacent grid cell in its adjacent grid cells to indicate whether the input G-PCC item extends into that adjacent grid cell.

[0143] In some embodiments, the boundaries of the cells can be defined more strictly: the face shared between two cells belongs to only one of the two cells, for example, the cell with the largest coordinates. Similarly, an edge or vertex shared between two or more than two cells can belong to only one of these cells, for example, the cell with the largest coordinates. In these embodiments, if an input G-PCC item contains points located on a face that does not belong to its cell, the input G-PCC item can extend beyond its grid cell.

[0144] Still in some embodiments, the grid can be sparse, with one or more than one cell not containing any points. G-PCC items that do not contain any points can be used to represent empty cells. As an optimization, the GPCCGrid structure can indicate which cells contain points and which cells are empty, for example, as follows:

[0145]

[0146] Among them, the sparse field indicates whether the grid is sparse. If the value of this field is set to the first value (e.g., 1), the grid is sparse, and for each cell, the occupied field indicates whether the cell contains one or more points. If the value of the occupied field corresponding to a given cell is set to the first value (e.g., 1), then the cell is associated with a G-PCC item in the list of input G-PCC items signaled in the SingleItemTypeReferenceBox of type 'dimg'. Otherwise, if the value of the occupied field corresponding to a given cell is set to the second value (e.g., 0), then the cell is not associated with any G-PCC item in the list of input G-PCC items signaled in the SingleItemTypeReferenceBox of type 'dimg' (i.e., the cell is empty).

[0147] The occupied values for different input G-PCC items can be listed according to their sorting.

[0148] Still according to some embodiments, the reference frame for the input G-PCC items can be different from one item to another: some input G-PCC items can use the same reference frame as the grid, while other input G-PCC items can use a reference frame local to their cells. Due to the following structure, these different reference frames can be indicated:

[0149]

[0150]

[0151] where width is equal to width_minus_one plus one; where height is equal to height_minus_one plus one; where depth is equal to depth_minus_one plus one; the same_origin field can have three different values. If the value of this field is set to the first value (e.g., 0), each input G-PCC item uses its own reference system to define. If the value of this field is set to the second value (e.g., 1), all input G-PCC items use the same reference system to define. If the value of this field is set to the third value (e.g., 2), different reference systems can be used for different input G-PCC items. For each cell, the global_cell_origin field indicates whether the corresponding input G-PCC item uses its own reference system or a common reference system. If the value of the global_cell_origin field associated with the cell is set to the first value (e.g., 0), the corresponding input G-PCC item uses its own reference system, and if the value of the global_cell_origin field associated with the cell is set to the second value (e.g., 1), the corresponding input G-PCC item uses the common reference system.

[0152] For an input G-PCC item using its own reference system, translations can be applied to its points to transform these points into the global reference system using the following relationships:

[0153] x g = x l + c x * i x

[0154] y g = y l + c y * i y

[0155] z g = z l + c z * i z

[0156] Additionally, a second translation can be applied to use a reference system different from that of the cell located at position (0, 0, 0). This translation can be defined in the structure by the global_origin field.

[0157] The global_cell_origin field can be sorted according to the input G-PCC item order.

[0158] Still in some embodiments, when one or more than one input G-PCC item is defined using its own reference frame, e.g., when the value of the same_origin flag is set to a second value (e.g., 0), one of the input G-PCC items (the reference input G-PCC item) can be used to define the reference frame of the grid. The reference input G-PCC item can be the input G-PCC item located at position (0, 0, 0). It can be another input G-PCC item signaled in the grid structure.

[0159] The translation between the reference frame of the grid and the reference frame of the cell containing the reference input G-PCC item can be calculated by aligning the reference input G-PCC item with its grid cell and calculating the translation between them. If the size of the reference input G-PCC item is the same as the size of the grid cell, this alignment can be achieved directly. If the size of the reference input G-PCC item is smaller than the size of the grid cell, the reference input G-PCC item can be aligned with the grid cell by centering the reference input G-PCC item within the grid cell, or by aligning the reference input G-PCC item with one or more than one face of the grid cell, or by a combination of both.

[0160] Still in some embodiments, the input G-PCC items of the grid can have different sizes, and the input G-PCC items can correspond to several cells of the grid. In these embodiments, the grid can be described by the following structure:

[0161]

[0162]

[0163] where the item_width_minus_one, item_depth_minus_one, and item_height_minus_one fields can be used to signal the positions and sizes of different input G-PCC items. For this signaling, the cells of the grid are scanned according to the sorting specified by the morton_order field. For each cell, if it corresponds to the previously described input G-PCC item, it is skipped. Otherwise, the cell is the current cell located at position (i x , i y , i z ). The next entry in the SingleItemTypeReferenceBox of type 'dimg' associated with the grid corresponds to the current input G-PCC item. The number of cells spanned by the current input G-PCC item is specified by the s x , s y , and s z values, and this sx , s y and s z The values are specified by the current item_width_minus_one, item_height_minus_one, and item_depth_minus_one fields. The current input G-PCC item spans cells along the X-axis from i x to i x + s x - 1, along the Y-axis from i y to i y + s y - 1, and along the Z-axis from i z to i z + s z - 1.

[0164] Still according to some embodiments, the bounding box of the mesh-derived item can be specified and can be different from the bounding box of the mesh. For example, the bounding box of the mesh can be smaller than the mesh, which means that the point cloud obtained by combining the input G-PCC items in the mesh is cropped to obtain the point cloud corresponding to the mesh-derived item. As another example, the bounding box of the mesh can have a larger range than the mesh itself. In this variant, the mesh can be described by the following structure:

[0165]

[0166]

[0167] wherein, the bounding box of the mesh is specified by the output_position and output_size fields. The output_precision field specifies the precision for the output_position and output_size fields.

[0168] Possibly, item properties associated with the mesh-derived item (e.g., bounding box item property or size item property) can be used to specify the bounding box of the mesh. Possibly, a cropping transformation item property associated with the mesh-derived item can be used to specify the bounding box of the mesh.

[0169] 3D Fusion

[0170] It has been observed that there are cases where a 3D volume can be segmented according to an irregular manner (not corresponding to a mesh). For example, a scene can be captured by LiDAR continuously located at different positions, and the resulting point cloud data can be segmented into several subsets of point cloud data centered at different capture positions. The resulting point cloud can be represented as a fusion of several subsets of point cloud data.

[0171] The 3D fusion G-PCC item can be represented as an item with the item_type being a specific value (e.g., the value 'fus3'). This 3D fusion item is intended to combine one or more input G-PCC items arranged in an irregular manner. Thus, the item with the item_type value of 'fus3' defines the following derived point cloud item, and the reconstructed point cloud of this derived point cloud item is formed by combining one or more input point clouds.

[0172] Insert the input point clouds in the order of the SingleItemTypeReferenceBox of the 'dimg' type for this derived point cloud item within the ItemReferenceBox. In the SingleItemTypeReferenceBox of the 'dimg' type, the value of the from_item_ID identifies the derived point cloud item of the "fus3" type, and the value of the to_item_ID identifies the input point cloud.

[0173] When removing the items of the input point clouds marked as point cloud fusion items, it may be necessary to rewrite the content of the point cloud fusion items.

[0174] The parameters associated with the derived point cloud item specify some characteristics of the 3D fusion item and can have the following structure:

[0175]

[0176] Among them, the version field signals the version of the structure, and the flags field can signal some options of the structure. If the version or flags are not defined separately, the version or flags should be equal to 0.

[0177] The overlap field indicates whether some input G-PCC items of the 3D fusion item can overlap. If the value of the overlap field is set to the first value (e.g., 1), then some input G-PCC items of the 3D fusion item can overlap or may not overlap, that is, the intersection of the bounding boxes of any pair of input point clouds can have or may not have an empty volume. If the value of the overlap field is set to the second value (e.g., 0), then none of the input G-PCC items of the 3D fusion item overlap, that is, the intersection of the bounding boxes of any pair of input point clouds has an empty volume. If the intersection of the bounding boxes of two input G-PCC items is an empty volume, then these two input G-PCC items can be considered non-overlapping. In a variant, if the intersection of the bounding boxes of two input G-PCC items is an empty volume and their bounding boxes share a face, an edge, or a vertex, then these two input G-PCC items can be considered non-overlapping.

[0178] The same_origin flag indicates whether the input G-PCC items combined by the 3D fusion item use the same reference system. If the value of the same_origin flag is set to the first value (e.g., 1), the same reference system or coordinate system is used to define all input G-PCC items. In this case, the input G-PCC items can be combined without applying translation to them. Otherwise, if the value of the same_origin flag is set to the second value (e.g., 0), each input G-PCC item uses its own reference system to be defined. In the latter case, translation can be applied to the coordinates of the points of the input G-PCC item to calculate their coordinates in the global reference system of the 3D fusion item. For each input G-PCC item, the translation from its local reference system to the common reference system can be signaled by the anchor field corresponding to that input G-PCC item, i.e., the i-th anchor is applied to the i-th occurrence in the order of the SingleItemTypeReferenceBox of type 'dimg' for the derived point cloud item within the ItemReferenceBox. The global coordinates of the points of the input G-PCC item can be calculated from their decoded coordinates as follows:

[0179] x g = x l + a x

[0180] y g = y l + a y

[0181] z g = z l + a z

[0182] where x g 、y g and z g are the coordinates of the point in the global reference system, x l 、y l and z l are the coordinates of the point in the local reference system of the input G-PCC item, and a x 、a y and a z are the coordinates along the X, Y, and Z axes of the anchor point of the input G-PCC item as signaled by the anchor field corresponding to that input G-PCC item.

[0183] The precision field indicates the precision used to encode the anchor points of the input G-PCC items. For each input G-PCC item, the anchor field signals the coordinates of its anchor point, which are used to transform the coordinates of the points of the input G-PCC item from the local reference system used by the input G-PCC item to the common reference system used by the 3D fusion item. The anchor fields can be sorted in the order of the input G-PCC items.

[0184] The input G-PCC items can be signaled in a SingleItemTypeReferenceBox of type 'dimg'. In this SingleItemTypeReferenceBox, the value of the from_item_Id field identifies the derived point cloud item. The value of the reference_count field indicates the number of input G-PCC items. The reference_count is obtained from the SingleItemTypeReferenceBox of type 'dimg', where the item is identified by the from_item_ID field. The value of the to_item_ID field identifies the input G-PCC item. Again, the SingleItemTypeReferenceBox can be referred to as an item reference.

[0185] According to some embodiments, each input G-PCC item can have an associated same_origin flag. In such an embodiment, the 3D fusion item can have the following structure:

[0186]

[0187] Figure 5 An example of steps for storing and / or transmitting point cloud data as a 3D fusion item of input G-PCC items according to some embodiments of the present invention is illustrated.

[0188] As illustrated, the first step involves obtaining the point cloud data for storage and / or transmission (step 500). These point cloud data can be obtained directly from a sensor such as LiDAR. It can also be the result of processing one or more point clouds captured by a sensor. For example, several point clouds can be captured by LiDAR, registered to transform them so that they use the same reference system, and then combined into a single point cloud. The point cloud can also be obtained as a set of several point clouds to be combined together. For example, it can be a set of point clouds captured by LiDARs at different positions in a scene or by different LiDARs at different positions in a scene.

[0189] Next, the target characteristics of the 3D fusion item are obtained (step 505).

[0190] In this step, an indication on how to split the point cloud data obtained at step 500 into subsets of point cloud data can be obtained. The indication can be based on 3D regions for splitting the point cloud data. It can be based on a list of points corresponding to the subsets of the point cloud data. It can also be based on the characteristics of the points of the point cloud data. For example, the point cloud data can be split based on the capture time of the points such that the subsets of the point cloud data correspond to different time periods. As another example, the point cloud data can be split by grouping points around a set of center points, with each point grouped with the nearest center point, such that the subsets of the point cloud data include points that are close to each other.

[0191] Additionally, an indication on whether to use the same origin for all input G-PCC items can be obtained. Possibly, when using the same origin for all input G-PCC items, a translation that is applied to all input G-PCC items before encoding all input G-PCC items can be obtained. Similarly, when not using the same origin for the input G-PCC items, a translation that is to be applied to the 3D fusion item can be obtained.

[0192] According to some embodiments, it is determined whether point clouds corresponding to subsets of the point cloud data can overlap. This can include obtaining an indication of whether these point clouds can overlap. It can also include analyzing the indication on how to split the point cloud data. For example, if the indication is based on 3D regions, it can be determined from these 3D regions whether the point clouds corresponding to the subsets of the point cloud data overlap. For example, if the 3D regions are axis-aligned cubes with respect to a reference frame, the point clouds corresponding to the subsets of the point cloud data do not overlap. Conversely, if the 3D regions are not axis-aligned with respect to the reference frame, the point clouds may overlap. As another example, if the indication is based on the characteristics of the points, it can be determined that the point clouds can overlap. The determination can be performed by directly checking whether the point clouds overlap. For example, the determination can be performed by calculating the bounding boxes of the point clouds corresponding to the subsets of the point cloud data and by checking whether these bounding boxes overlap. In the latter case, the determination can be performed at the end of step 510 once the point cloud data has been split into subsets of point cloud data.

[0193] Additionally, an ordering of the input G-PCC items in the 3D fusion item can be obtained. The ordering can be obtained in association with the indication on how to split the point cloud data. For example, if the indication is based on the 3D regions used to split the point cloud data, the ordering can be associated with these 3D regions to indicate the ordering of the input items. As another example, if the splitting is based on the characteristics of the points of the point cloud data, the ordering of the input items can be based on these characteristics or on an indication related to these characteristics. For example, if the point cloud data is split based on the capture time of the points, the input items can be ordered according to that capture time.

[0194] It should be noted that step 505 can be performed before step 500 or simultaneously with step 500.

[0195] Next, according to the indication obtained in step 505, the point cloud data obtained in step 500 is segmented into subsets of the point cloud data (step 510). It can be observed that if the point cloud data is obtained as multiple subsets of the point cloud data in step 500, then step 510 is skipped.

[0196] Next, the point cloud data of each subset of the point cloud data is processed (step 515), for example, encoded. Preferably, the processing of the point cloud data of each subset of the point cloud data is performed independently for each subset. For example, the point cloud data of each subset of the point cloud data can be encoded in parallel. Still by way of illustration, the MPEG-I Part-9 specification can be used to perform the encoding.

[0197] During this step, if the same origin is used for all input G-PCC items, the translation that may be obtained at step 505 can be applied to the points of each subset of the point cloud data before or as part of the encoding of the points of each subset of the point cloud data. If the same origin is not used for all input G-PCC items, the points of each subset of the point cloud data can be translated according to its own reference frame determined by its anchor point before or while encoding it.

[0198] Next, a structure is created for encapsulating the 3D fusion item and the subsets of the encoded point cloud (step 520). For example, a GPCCFusion structure as described above can be created to describe the 3D fusion item and store its characteristics. Additionally, G-PCC items can be created to describe each subset of the encoded point cloud data. Possibly, each subset of the encoded point cloud data can be described by using a single G-PCC item or by using several G-PCC items. A SingleItemTypeReferenceBox of the 'dimg' type can be created to associate the input G-PCC item that describes the subset of the encoded point cloud data with the 3D fusion item. The sorting of the input G-PCC items within the SingleItemTypeReferenceBox can depend on the sorting obtained in step 505.

[0199] It should be noted that step 520 or some sub-steps of step 520 can be performed simultaneously with step 515 or some sub-steps of step 515.

[0200] Next, store the subset of the encoded point cloud data and the encapsulation structure created in step 520 (including item references that establish associations between derived items and input items, such as SingleItemTypeReferenceBox of type 'dimg') in an encapsulated data file (e.g., an encapsulated data file conforming to the ISOBMFF format) (step 535). Optionally, the encapsulated data file is stored in a memory unit or a hard disk drive. Optionally, the encapsulated data file is sent via a communication network.

[0201] Encoding the point cloud data into 3D fusion items can advantageously improve the encoding speed of the point cloud because several encoding operations can be performed in parallel.

[0202] According to some embodiments, obtain the point cloud data as multiple subsets of the point cloud data (step 500), and do not obtain different subsets of the point cloud data simultaneously. In such embodiments, one or more subsets of the point cloud data can be processed independently of other subsets. Encode one or more subsets of the point cloud data as described with respect to step 515, create an encapsulation structure to represent the subset as described with respect to step 520, and store and / or transmit the subset of the encoded point cloud data and its encapsulation structure as described with respect to step 525.

[0203] Figure 6 Illustrates an example of steps for decoding and rendering point cloud data stored as 3D fusion of input G-PCC items according to some embodiments of the present invention.

[0204] In a first step, obtain an encapsulated data file (e.g., an encapsulated data file conforming to the ISOBMFF format) that stores the 3D fusion of G-PCC items (step 600). Additionally, an indication of the first item (i.e., the fusion item) to be decoded and rendered can be obtained. If no indication of the first item to be decoded and rendered is obtained, the first item signaled in the ISOBMFF file can be considered the item to be decoded and rendered.

[0205] Next, obtain the 3D fusion item to be decoded and rendered from the encapsulated data file (step 605). The 3D fusion item can be encoded according to the GPCCFusion structure described above. It can be parsed to obtain, for example, an indication of whether the input G-PCC items use the same origin and / or whether these input G-PCC items can overlap.

[0206] In addition, the SingleItemTypeReferenceBox of type 'dimg' associated with the 3D fusion item can be parsed to identify the input G-PCC items used by the 3D fusion item.

[0207] Additionally, the encoded point cloud data of each input G-PCC item can be obtained from an encapsulated data file (e.g., an ISOBMFF file). Each input G-PCC item can contain a subset of the encoded point cloud data.

[0208] Next, the encoded point cloud data of the input G-PCC item (or some of the input G-PCC items) is decoded and rendered. For illustrative purposes, decoding of the encoded point cloud data can be performed according to the MPEG-I Part-9 specification. The rendering step can include generating structures in a memory unit for representing the decoded point cloud data, transferring these structures to a graphics card, and generating a display of the point cloud data.

[0209] If the same_origin field of the GPCCFusion structure is set to a second value (i.e., 0 according to the previous example) (which means that each subset of the point cloud data uses its own reference frame), then the positions of the points of the subset of the decoded point cloud data can be translated, for example, using the above-mentioned translation, into the reference frame corresponding to the 3D fusion item. Optionally, a further translation can be applied to the decoded points to change the origin of the reference frame used by the 3D fusion item. This translation can be applied while decoding the subset of the point cloud data, after decoding the subset of the point cloud data and before rendering, or as part of the rendering process.

[0210] If the same_origin field of the GPCCFusion structure is set to a first value (i.e., 1 according to the previous example) (which means that the subset of the encoded point cloud data of the subset uses the same origin), then the points of the subset of the decoded point cloud data can be translated before rendering. Optionally, this translation can be performed as part of the decoding or as part of the rendering.

[0211] Preferably, decoding and rendering are performed as independently as possible for each subset of the encoded point cloud data. For example, for several subsets of the encoded point cloud data, decoding, generating structures in a memory unit for representing the decoded point cloud, and transferring these structures to a graphics card can be performed in parallel.

[0212] During the decoding and rendering of the subset of the encoded point cloud data, information on whether the corresponding point clouds can overlap can be used to identify the steps of the process that can be performed independently on the encoded point cloud data.

[0213] In fact, if the point clouds corresponding to the decoded point cloud data subsets are signaled as non-overlapping, the memory structures used to represent these point clouds can be organized in such a way that the memory structures in the most commonly used memory structure are only used for the decoding and rendering of one of these point clouds. For example, the representation of the 3D space can be divided into several sub-volumes, and each sub-volume is represented using one or more memory structures. These sub-volumes can be constructed such that each sub-volume corresponds to the bounding box of the points of a subset of the point cloud data. In this way, during the decoding of a subset of the encoded point cloud data, only the memory structures representing the associated sub-volume are accessed, and only these memory structures are accessed during the decoding of this subset. In this example, some global memory structures can be used to organize the memory structures corresponding to each sub-volume. However, these global memory structures are mostly accessed at step 605 while parsing the 3D fusion item, rather than during the decoding of the encoded point cloud data. Possibly, several sub-volumes can correspond to a subset of the encoded point cloud data, and each sub-volume only corresponds to a subset of the encoded point cloud data.

[0214] Conversely, if the point clouds corresponding to a subset of the point cloud data are signaled as overlapping, the decoded point cloud data subset can be reorganized before being rendered. For example, a subset of the encoded point cloud data can be decoded independently, and then the resulting data can be merged into the memory structures representing different point clouds. As another example, a subset of the encoded point cloud data can be decoded simultaneously, and the resulting data can be directly stored into the memory structures representing all the decoded point cloud data subsets, while taking care to keep the contents of the memory structures coherent in the case of simultaneous access. The organization of the memory structures can be arranged such that some memory structures only contain data from a single subset of the point cloud data, while some other memory structures contain data from several subsets of the point cloud data. In this way, there is no need to merge the data or to take care of the simultaneous access to the memory structures that only contain data from a single subset of the point cloud data.

[0215] Decoding the point cloud previously encoded as a 3D fusion item can advantageously increase the decoding speed of the point cloud because several decoding operations can be performed in parallel. Additionally, a part of the rendering process can also be performed in parallel.

[0216] According to some embodiments, if the points of one input G-PCC item among these input G-PCC items are not included within the bounding box of another input G-PCC item, then these two input G-PCC items can be considered non-overlapping. During step 610, the 3D space can be divided into sub-volumes such that each sub-volume containing one or more points from the point cloud represented by the 3D fusion item corresponds to only a subset of the encoded point cloud data. In these embodiments, the sub-volumes corresponding to the intersection of the bounding boxes of two or more input G-PCC items are empty and do not contain any points from the point cloud represented by the 3D fusion item.

[0217] Still according to some embodiments, the overlap of input G-PCC items can be specified with a finer granularity. For example, an overlap field can be associated with each input G-PCC item to indicate whether this input G-PCC item overlaps with another input G-PCC item. In combination with the foregoing embodiments, this overlap field can indicate whether the bounding box of the input G-PCC item contains points from other input G-PCC items.

[0218] As another example, an overlap field can be associated with each pair of input G-PCC items, which indicates whether these two input G-PCC items overlap. In combination with the foregoing embodiments, this overlap field can indicate whether the bounding box of one input G-PCC item among these input G-PCC items contains points from another input G-PCC item.

[0219] Possibly, the bounding box of the 3D fusion derived item can be calculated as a combination of the bounding boxes of the input G-PCC items. If the value of the same_origin field of the GPCCFusion structure is set to a second value (i.e., 0 according to the previous example), which means that each subset of the point cloud data uses its own reference frame, then this calculation can be performed by translating the bounding boxes of the input G-PCC items. Then, the bounding box of the 3D fusion derived item can be calculated as spanning from the minimum coordinate values of the bounding boxes of the input items to the maximum coordinate values of these bounding boxes along each coordinate axis.

[0220] In a variant, the bounding box of the 3D fusion derived item can be specified and can be different from the combination of the bounding boxes of the input G-PCC items of the 3D fusion item. For example, the bounding box of the 3D fusion derived item can be smaller than the combination of the bounding boxes of the input G-PCC items, which means that the point cloud obtained by combining the input G-PCC items into a 3D fusion item is cropped to obtain the point cloud corresponding to the 3D fusion derived item. As another example, the bounding box of the 3D fusion derived item can have a larger range than the combination of the bounding boxes of the input G-PCC items. In this variant, the 3D fusion derived item can be described with the following structure:

[0221]

[0222]

[0223] The bounding box of the 3D fusion derivative can be specified by the output_position and output_size fields. The output_precision field specifies the precision used for the output_position and output_size fields.

[0224] Still as a variant, the 3D fusion item can be called 3D merge and represented as a derivative item with an item_type value of'mrg3'. The GPCCFusion structure can be called GPCCMerge.

[0225] In some embodiments, one or more input items for the 3D mesh item or for the 3D fusion item can be 3D mesh items and / or 3D fusion items.

[0226] In some embodiments, several embodiments of the 3D mesh item can be used simultaneously. Different values of the version field or different values of the item_type can be used to signal different embodiments.

[0227] In some embodiments, several embodiments of the 3D fusion item can be used simultaneously. Different values of the version field or different values of the item_type can be used to signal different embodiments.

[0228] Structure

[0229] The Vector3 structure used in the above different structures can be the Vector3 structure defined by MPEG - IPart - 7 as follows:

[0230]

[0231] The Vector3 structure can also use floating - point numbers of fixed - point numbers to represent the x, y, and z coordinates as in the following structure:

[0232] aligned(8)class Vector3(unsigned int precision_bytes_minus1){

[0233] signed int((precision_bytes_minus1 + 1)*16)x;

[0234] signed int((precision_bytes_minus1 + 1)*16)y;

[0235] signed int((precision_bytes_minus1 + 1) * 16) z;

[0236] }

[0237] Among them, the x, y, and z coordinates are represented using fixed-point numbers, where the integer part has precision_bytes_minus1 + 1 bytes and the fractional part has precision_bytes_minus1 + 1 bytes.

[0238] Item property

[0239] Several item properties can be associated with G-PCC items, identity point cloud items, 3D mesh-derived point cloud items, and / or 3D fusion-derived point cloud items.

[0240] The BoundingBox item property specifies the bounding box of a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, and / or a 3D fusion-derived point cloud item. This item property can have the following structure:

[0241] aligned(8) class BoundingBox

[0242] extends ItemFullProperty('box3', version = 0, flags = 0) {

[0243] unsigned int(8) precision;

[0244] Vector3 position(precision);

[0245] Vector3 size(precision);

[0246] }

[0247] In this structure, the position field is the reference point of the bounding box and specifies the minimum values of the x, y, and z coordinates of the points contained in the associated or derived item. The size field specifies the ranges along the x, y, and z coordinates of the bounding box.

[0248] The bounding box can also be specified using two points (for example, two points corresponding to two opposite corners of the bounding box).

[0249] The bounding box can also be specified using the center of the bounding box and its size.

[0250] The size item property specifies the size of a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, and / or a 3D fusion-derived point cloud item. This item property can have the following structure:

[0251] aligned(8) class Size

[0252] extends ItemFullProperty('siz3', version = 0, flags = 0) {

[0253] unsigned int(8) precision;

[0254] Vector3 size(precision);

[0255] }

[0256] In this structure, the size field specifies the size of the item along the x, y, and z coordinates.

[0257] The Translation transform item property specifies the translation to be applied to a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, and / or a 3D fusion-derived point cloud item. This transform item property can have the following structure:

[0258] aligned(8) class Translation

[0259] extends ItemFullProperty('tra3', version = 0, flags = 0) {

[0260] unsigned int(8) precision;

[0261] Vector3 translation(precision);

[0262] }

[0263] In this structure, the translation field specifies the translation vector.

[0264] The Scaling transform item property specifies the scaling to be applied to a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, or a 3D fusion-derived point cloud item. This transform item property can have the following structure:

[0265] aligned(8) class Scaling

[0266] extends ItemFullProperty('sca3', version = 0, flags = 0) {

[0267] unsigned int(8) precision;

[0268] Vector3 scaling(precision);

[0269] }

[0270] In this structure, the scaling field specifies the scaling to be applied to the x, y, and z coordinates. Optionally, a simpler scaling transformation term property can be defined where the same scaling is applied to all coordinates.

[0271] The Rotation transformation term property specifies the rotation to be applied to the G-PCC term, identity point cloud term, 3D mesh-derived point cloud term, and / or 3D fusion-derived point cloud term. This transformation term property can have the following structure:

[0272]

[0273] In this structure, the rotation field specifies the rotation as a quaternion. This quaternion is a unit quaternion, and its three imaginary components can be specified, and its real component can be calculated based on these three imaginary components.

[0274] The extended Rotation transformation term property can be defined as follows:

[0275]

[0276] In this structure, the center field specifies the center of rotation.

[0277] The Transformation transformation term property specifies a general 3D transformation that can combine translation, scaling, and / or rotation to be applied to the G-PCC term, identity point cloud term, 3D mesh-derived point cloud term, or 3D fusion-derived point cloud term. This transformation term property can have the following structure:

[0278]

[0279] In this structure, the coeff field specifies the coefficients of the transformation matrix. The transformation matrix can be a 4×4 homogeneous matrix, where the last row of coefficients is (0, 0, 0, 1). In this case, the Transformation structure can specify only 12 coefficients.

[0280] The Crop transformation term property specifies the cropping to be applied to the G-PCC term, identity point cloud term, 3D mesh-derived point cloud term, and / or 3D fusion-derived point cloud term. This transformation term property can have the following structure:

[0281] aligned(8)class Crop

[0282] extends ItemFullProperty('crp3', version = 0, flags = 0) {

[0283] unsigned int(8) precision;

[0284] Vector3 position(precision);

[0285] Vector3 size(precision);

[0286] }

[0287] The position and size fields specify the boundaries of the clipping. Any point in the item that lies outside the clipping boundaries is not part of the transformed point cloud.

[0288] Two points (e.g., two points corresponding to two opposite corners of the clipping region) can also be used to specify the clipping.

[0289] The TransformedBoundingBox item property specifies the bounding box of a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, and / or a 3D fusion-derived point cloud item after the transformation item properties have been applied.

[0290] The TransformedBoundingBox item property can specify the bounding box of the associated item after all the transformation item properties listed before the application of this TransformedBoundingBox item property and before any transformation item properties listed after the application of this TransformedBoundingBox item property.

[0291] In fact, calculating the bounding box of a G-PCC item, an identity point cloud item, a 3D mesh-derived point cloud item, and / or a 3D fusion-derived point cloud item after the transformation item properties have been applied may require calculating the bounding box based on all the transformed points. For example, after applying a translation, the bounding box of the transformed point cloud is a translation of the bounding box of the original point cloud. However, after applying a rotation, the rotated bounding box of the original point cloud is not aligned with the axes of the reference frame. The bounding box constructed from the rotated bounding box contains all the points of the rotated point cloud but may not tightly fit the rotated point cloud. Therefore, indicating the bounding box of the rotated point cloud can help decode and render the rotated point cloud by precisely indicating the extent of the transformed point cloud.

[0292] This item property can have the following structure:

[0293] aligned(8) class TransformedBoundingBox

[0294] extends ItemFullProperty('bot3', version = 0, flags = 0) {

[0295] unsigned int(8) precision;

[0296] Vector3 position(precision);

[0297] Vector3 size(precision);

[0298] }

[0299] In this structure, the position field is the reference point of the transformed bounding box and specifies the minimum x, y, and z coordinates of the points contained in the associated or derived item after applying the transformation item property to the associated or derived item. The size field specifies the ranges along the x, y, and z coordinates of the transformed bounding box.

[0300] Two points (e.g., two points corresponding to two opposite corners of the bounding box) can also be used to specify the transformed bounding box.

[0301] The transformed bounding box can also be specified using the center of the bounding box and its size.

[0302] Possibly, several transformed bounding boxes can be associated with an item or a derived item corresponding to the application of different sets of transformation item properties.

[0303] The TransformedSize item property specifies the size of a G-PCC item, an identity point cloud item, a 3D mesh derived point cloud item, and / or a 3D fusion derived point cloud item after the transformation item property has been applied.

[0304] The TransformedSize item property can specify the size of the item associated with it after all the transformation item properties listed before applying this TransformedSize item property and before any transformation item properties listed after applying this TransformedSize item property. This item property can have the following structure:

[0305] aligned(8) class TransformedSize

[0306] extends ItemFullProperty('szt3', version = 0, flags = 0) {

[0307] unsigned int(8) precision;

[0308] Vector3 size(precision);

[0309] }

[0310] In this structure, the size field specifies the ranges along the x, y, and z coordinates of the transformed bounding box.

[0311] Variant

[0312] For illustrative purposes, the descriptions of the above different embodiments and variants focus on using G-PCC terms. In other embodiments, V3C terms defined by MPEG-I Part-10 can be used. In still other embodiments, other terms for describing encoded 3D data can be used. In still other embodiments, derived terms can combine different types of input terms. For example, a mesh-derived term can combine G-PCC terms and V3C terms.

[0313] In various embodiments, the sorting of input terms in a SingleItemTypeReferenceBox of type 'dimg' can be used to select which term prevails when several input terms define 3D content at the same location. For example, the latest input term in the sorting within the SingleItemTypeReferenceBox may prevail. For example, if two input G-PCC terms define points at the same location, the point corresponding to the latest input G-PCC term among the two input G-PCC terms within the SingleItemTypeReferenceBox order is retained, while the other point is discarded. As another example, when several input terms define 3D content at the same location, these contents can be merged. For example, if two input G-PCC terms define points at the same location, the colors of these points can be combined by calculating the average color and the points can be merged by obtaining the latest timestamp associated with these points. As yet another example, when several input terms define 3D content at the same location, all of these contents can be retained. For example, if two input G-PCC terms define points at the same location, all of these points or some of these points can be retained.

[0314] It should be noted that the different 4cc values used to indicate item types, item property types, or other information are given only as examples. Other values can be used to indicate these types.

[0315] In the different item structures described above, some fields are used to signal values encoded in 1, 2, or several bits. The flags field can also be used to signal some or all of these values. Some or all of the options signaled using these fields can also be signaled using the version field.

[0316] Hardware examples of steps of methods for implementing embodiments of the present disclosure

[0317] Figure 7 Is a schematic representation of an example of a data processing device configured to implement some embodiments of the present disclosure in whole or in part.

[0318] The data processing device 700 can be a device such as a microcomputer, a workstation, or a lightweight portable device. As illustrated, the data processing device 700 includes a communication bus 713, which is preferably connected to:

[0319] - A central processing unit 711 represented as a CPU (such as a microprocessor, etc.) or a GPU (representing a graphics processing unit),

[0320] - A read-only memory 707 represented as a ROM, for storing computer programs for implementing the present disclosure in whole or in part,

[0321] - A random access memory 712 represented as a RAM, for storing executable code of methods according to embodiments of the present disclosure and registers suitable for recording variables and parameters required to implement the methods according to embodiments of the present disclosure, and

[0322] - At least one communication interface 702 connected to a communication network, for transmitting data to and / or receiving data from a remote device.

[0323] Optionally, the data processing device 700 may further include one or several of the following components:

[0324] - A data storage component 704 such as a hard disk, for storing computer programs for implementing the methods according to one or more embodiments of the present disclosure,

[0325] - A disk drive 705 for a disk 706, which is adapted to read data from or write data to the disk 706, and

[0326] - A screen 709, for serving as a graphical interface with a user by means of a keyboard 710 or any other indicating component.

[0327] The data processing device 700 may optionally be connected to various peripheral devices including sensors 708 (such as digital cameras and / or LiDAR, etc.), and each peripheral device is connected to an input / output card (not shown) to supply data to the data processing device 700.

[0328] Preferably, a communication bus provides communication and interoperability between various elements included in or connected to the data processing device 700. The representation of the bus is not restrictive, and in particular, the central processing unit is operable to communicate instructions to any element of the data processing device 700 either directly or by means of other elements of the data processing device 700.

[0329] The disk 706 may optionally be replaced by any information medium (e.g., such as a rewritable or non-rewritable optical disk (CD-ROM), ZIP disk, USB key, or memory card, etc.), and in general, by an information storage component readable by a microcomputer or microprocessor, which may or may not be integrated into the device, may be removable and is adapted to store one or more programs, the execution of which enables the implementation of the method according to the present invention.

[0330] The executable code may optionally be stored in the read-only memory 707, on the hard disk 704, or on a removable digital medium (e.g., such as the disk 706 as described above). According to an alternative variant, the executable code of the program may be received via a communication network through the interface 702 for storage in one of the storage components of the data processing device 700 (such as the hard disk 704, etc.) before execution.

[0331] The central processing unit 711 is preferably adapted to control and direct the execution of instructions or software code portions of one or more programs according to the present invention, which instructions are stored in one of the above storage components. Upon power-on, one or more programs stored in a non-volatile memory (e.g., on the hard disk 704 or in the read-only memory 707) are transferred to the random access memory 712, and then the random access memory 712 contains the executable code of one or more programs as well as registers for storing variables and parameters required to implement the present invention.

[0332] In a preferred embodiment, the device is a programmable device that uses software to implement the present invention. However, optionally, the present invention may be implemented in hardware (e.g., in the form of an application-specific integrated circuit or ASIC).

[0333] Although the present invention has been described above by reference to specific embodiments, the present invention is not limited to these specific embodiments, and those skilled in the art will appreciate modifications within the scope of the present invention.

[0334] In referring to the foregoing exemplary embodiments, those skilled in the art will envision many further modifications and variations, which are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, where appropriate, different features from different embodiments may be interchanged.

[0335] Certain embodiments of the present invention described above may be implemented singly or as a combination of elements of multiple embodiments. Additionally, features from different embodiments may be combined when necessary or where it is advantageous to combine elements or features from various embodiments in a single embodiment.

[0336] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

Claims

1. A method for encapsulating point cloud data into an ISOBMFF-based media file, the method comprises: obtaining a plurality of subsets of the point cloud data; generating a plurality of input items, each input item describing a subset of the point cloud data in the plurality of subsets; generating a derived item, the derived item including descriptive data of the spatial composition of the plurality of input items, the derived item being associated with the plurality of input items by an item reference; encapsulating the plurality of input items, the item reference, and the derived item in the media file.

2. The method according to claim 1, wherein, the descriptive data further includes an indicator indicating whether the point clouds corresponding to the point cloud data in each subset of the point cloud data use the same reference system.

3. The method according to claim 1 or 2, wherein, the descriptive data further includes an indicator indicating whether the bounding boxes of the point clouds corresponding to each of at least two of the plurality of input items overlap.

4. The method according to any one of claims 1 to 3, wherein, the descriptive data further includes an indicator for describing a three-dimensional array of cells, each of the plurality of input items corresponding to one of the cells.

5. The method according to claim 4, wherein, the descriptive data further includes an indicator indicating the order of the plurality of input items in the three-dimensional array of cells.

6. The method according to claim 4 or 5, wherein, the descriptive data further includes an indicator for describing the size of at least one of the cells.

7. The method according to claim 4 or 5, wherein, the size of the cell is determined according to the bounding box of the point cloud corresponding to the point cloud data described by each of the plurality of input items.

8. The method according to any one of claims 4 to 7, wherein, the descriptive data further includes an indicator indicating that at least one of the cells does not include any points defined in the point cloud data.

9. The method according to any one of claims 1 to 8, wherein, at least one item property is associated with at least one of the input item and the derived item, the at least one item property being used to describe at least one spatial operation to be performed on the points defined in the point cloud data.

10. The method according to any one of claims 1 to 9, wherein, the derived item includes at least one indicator for describing at least one spatial operation to be performed on the points defined in the point cloud data.

11. The method according to any one of claims 1 to 10, further comprises: obtaining the point cloud data and splitting the obtained point cloud data into the plurality of subsets of the point cloud data.

12. A method for parsing an ISOBMFF-based media file encapsulating point cloud data, the method comprises: obtaining a derived item from the media file, the derived item including descriptive data of the spatial composition of a plurality of input items, the derived item being associated with the plurality of input items by an item reference; Obtain the plurality of input items from the media file according to the item reference, where each input item describes a subset of the point cloud data of multiple subsets; Obtain the multiple subsets of the point cloud data from the media file according to the plurality of input items; And Generate point cloud data according to the descriptive data, where the point cloud data includes the point cloud data of the multiple subsets.

13. The method according to claim 12, further comprising: Determine whether the point clouds corresponding to the point cloud data of each subset of the point cloud data use the same reference system according to the indicator of the descriptive data.

14. The method according to claim 12 or 13, further comprising: Determine whether the bounding boxes of the point clouds corresponding to the point cloud data described by each of at least two input items among the plurality of input items overlap according to the indicator of the descriptive data.

15. The method according to any one of claims 12 to 14, further comprising: Determine a three-dimensional array of cells according to the indicator of the descriptive data, where each of the plurality of input items corresponds to one cell in the cells.

16. The method according to claim 15, further comprising: Determine the order of the plurality of input items in the three-dimensional array of cells according to the indicator of the descriptive data.

17. The method according to claim 15 or 16, further comprising: Determine the size of at least one cell in the cells according to the indicator of the descriptive data.

18. The method according to any one of claims 15 to 17, further comprising: Determine that at least one cell in the cells does not include any points defined in the point cloud data according to the indicator of the descriptive data.

19. The method according to any one of claims 12 to 18, further comprising: Apply a spatial operation to the points defined in the point cloud data according to at least one item property associated with at least one of the input item and the derived item.

20. The method according to any one of claims 12 to 19, further comprising: Apply at least one spatial operation to be performed on the points defined in the point cloud data according to at least one indicator of the derived item.

21. A computer program product for a programmable device, the computer program product includes instructions for performing the steps of the method according to any one of claims 1 to 20 when the programmable device loads and executes the program.

22. A non-transitory computer-readable storage medium that stores instructions of a computer program for implementing the method according to any one of claims 1 to 20.

23. A processing device, which includes a processing unit configured to perform the steps of the method according to any one of claims 1 to 20.