Method, apparatus, and computer program for improving transmission and / or storage of point cloud data
By encapsulating point cloud data into an ISOBMFF-based media file, the method enables independent processing of subsets, improving encoding and decoding efficiency.
Patent Information
- Application Number
- JP2025513079
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-16
- Publication Date
- 2026-02-19
- Estimated Expiration
- 2043-10-16
AI Technical Summary
Existing encapsulation methods for point cloud data, such as those defined by ISO Base Media File Format (ISOBMFF), are inefficient and require improvement in encoding and/or decoding efficiency, particularly in the context of improving the encapsulation of point cloud data, such as those defined by ISOBMFF), are inefficient and require improvement in encoding and/or decoding efficiency.
The method involves dividing the encapsulation of point cloud data into an ISOBMFF-based media file, generating a plurality of subsets of the point cloud data, generating a plurality of subsets of the point cloud data, generating a plurality of input items, and generating a plurality of subsets of the point cloud data, generating a plurality of subsets of the volumetric items, and encapsulating the volumetric items in the media file.
This method allows for improved processing efficiency for encoding and decoding point cloud data by enabling independent processing of subsets, thereby enhancing the efficiency of encoding and decoding operations.
Smart Images

Figure 0007818144000027 
Figure 0007818144000028 
Figure 0007818144000029
Abstract
Description
[Technical Field]
[0001] Technical Field The present invention relates to a method, device and computer program for improving the transmission and / or storage of point cloud data, for example by defining a volumetric item as a grid of several volumetric items or by using fusion of several volumetric items. [Background technology]
[0002] Background of the Invention The Moving Picture Experts Group (MPEG) has standardized the compression and storage of point cloud data (also known as volumetric media data), which consists of a set of 3D points associated with attribute information such as color, reflectance, and frame index.
[0003] First, MPEG-I Part-9 (ISO / IEC 23090-9) specifies G-PCC (Geometry-based Point Cloud Compression) and a bitstream syntax for point cloud information. According to MPEG-I Part-9, a point cloud is an unordered list of points that includes geometry information, optional attributes, and associated metadata. Geometry information describes the location of the points in a three-dimensional Cartesian coordinate system. Attributes are typed properties of each point, such as color or reflectance. Metadata are items of information used to interpret the geometry information and attributes. The G-PCC compression specification (MPEG-I Part-9) defines specific attributes, such as the frame index attribute or frame number attribute, and reserved attribute label values (3 for frame index and 4 for frame number attribute) are used. According to MPEG-I Part-9, a point cloud frame is a set of points at a particular time instance. A point cloud frame can be partitioned into one or more ordered subframes. In MPEG-I Part-9, point cloud frames are also indicated by the FrameCtr variable using a parameter (frame_ctr_lsb syntax element) in the frame boundary marker data unit or in some data unit headers.
[0004] Second, MPEG-I Part-18 (ISO / IEC 23090-18) specifies a media format based on ISO / IEC 14496-12 (ISOBMFF) that enables the storage and distribution of geometry-based point cloud compressed data. It also supports flexible extraction of geometry-based point cloud compressed data during distribution and / or decoding. According to MPEG-I Part-18, a point cloud frame is encapsulated within one or more G-PCC tracks, and a sample within a G-PCC track corresponds to a single point cloud frame. Each sample contains one or more G-PCC units belonging to the same presentation time. A G-PCC unit is a type-length-value (TLV) encapsulation structure containing one of the following data units: SPS, GPS, APS, tile inventory, frame boundary marker, geometry data unit, and attribute data unit. The syntax of the TLV encapsulation structure is defined in Annex B of ISO / IEC 23090-9.
[0005] Closely related to these standards, MPEG-I Part-5 (ISO / IEC 23090-7) specifies V3C (Visual volumetric video-based coding) and V-PCC (Video-based Point Cloud Compression). It specifies a general mechanism for encoding visual volumetric frames by converting 3D volumetric information into a set of 2D images and associated data. In addition, it specifies an application of this general mechanism targeting point cloud representations of visual volumetric frames. MPEG-I Part-10 (ISO / IEC 23090-10) specifies a media format for storing and delivering visual volumetric video-based coding data in files, based on ISO / IEC 14496-12 (ISOBMFF).
[0006] Both MPEG-I Part-18 and MPEG-I Part-10 allow for the storage of timed or non-timed data. Timed data is stored using one or more tracks defined in ISOBMFF. Non-timed data is stored using one or more items defined by ISOBMFF. Summary of the Invention [Problem to be solved by the invention]
[0007] While the ISO Base Media File Format has proven efficient for encapsulating point cloud data, there is a need to improve the encapsulation efficiency, e.g., to improve the efficiency of encoding the encapsulated point cloud data and / or to improve the efficiency of decoding the encapsulated point cloud data. [Means for solving the problem]
[0008] Summary of the Invention The present invention has been devised to address one or more of the aforementioned problems. It describes derived items that can be used to combine 3D items, for example using a regular grid or by specifying the location of each 3D item, and in particular allows, for example, to encode and / or decode point cloud data independently in parallel.
[0009] According to a first aspect of the present invention, there is provided a method for encapsulating point cloud data into an ISOBMFF-based media file, the method comprising: obtaining a plurality of subsets of the point cloud data; generating a plurality of input items each describing a subset of point cloud data of said plurality of subsets; generating a derived item that includes descriptive data of the spatial configuration of the plurality of input items, the derived item being related to the plurality of input items by item references; The plurality of input items, the item references, and the derived items are encapsulated in the media file.
[0010] Thus, the method of the present invention allows for improved processing efficiency for processing, e.g., encoding, point cloud data by allowing for independent processing of subsets of point cloud data.
[0011] According to some embodiments, the descriptive data further comprises an indicator of whether the point clouds corresponding to the point cloud data of each subset of point cloud data use the same frame of reference.
[0012] According to some embodiments, the description data further includes an indicator indicating whether bounding boxes of the point clouds corresponding to the point cloud data described by each of at least two input items of the plurality of input items overlap.
[0013] According to some embodiments, at least one indicator is associated with at least one input item of the plurality of input items, and the at least one indicator indicates that a bounding box corresponding to the at least one input item overlaps with a bounding box associated with another input item of the plurality of input items.
[0014] According to some embodiments, the descriptive data further comprises an indicator for describing a three-dimensional array of cells, each input item of the plurality of input items corresponding to one of the cells.
[0015] According to some embodiments, the descriptive data further includes an indicator indicating the order of the plurality of input items within the three-dimensional array of cells.
[0016] According to some embodiments, the descriptive data further comprises an indicator for describing a size of at least one of the cells.
[0017] According to some embodiments, the size of the cell is determined as a function of the bounding box of the point cloud corresponding to the point cloud data described by each of the plurality of input items.
[0018] According to some embodiments, the descriptive data further comprises an indicator that at least one of the cells does not contain a point defined in the point cloud data. According to some embodiments, at least one item property is associated with at least one of the input items and / or derived items for describing at least one spatial operation to be performed on points defined in the point cloud data.
[0019] According to some embodiments, the derived item includes at least one indicator for describing at least one spatial operation to be performed on points defined in the point cloud data.
[0020] According to some embodiments, the method further includes acquiring point cloud data and dividing the acquired point cloud data into a plurality of subsets of point cloud data.
[0021] According to a second aspect of the present invention, there is provided a method for parsing an ISOBMFF based media file encapsulating point cloud data, the method comprising: obtaining, from the media file, a derived item that includes descriptive data of a spatial configuration of a plurality of input items, the derived item being associated with the plurality of input items by item references; obtaining input items from the media file, each input item describing a subset of point cloud data of a plurality of subsets, the input items as a function of the item references; obtaining the subsets of the point cloud data from the media file as a function of the input items; generating the point cloud data as a function of the descriptive data, the point cloud data including the point cloud data of the plurality of subsets.
[0022] Thus, the method of the present invention allows for improved processing efficiency for processing, e.g., decoding, point cloud data by allowing for independent processing of subsets of point cloud data.
[0023] According to some embodiments, the method further comprises determining, as a function of the indicator in the descriptive data, whether the point clouds corresponding to the point cloud data of each subset of point cloud data use the same frame of reference.
[0024] According to some embodiments, the method further comprises determining whether bounding boxes of point clouds corresponding to point cloud data described by each of at least two input items of the plurality of input items overlap as a function of the indicator of the description data.
[0025] According to some embodiments, at least one indicator is associated with at least one input item of the plurality of input items, and the at least one indicator indicates that a bounding box corresponding to the at least one input item overlaps with a bounding box associated with another input item of the plurality of input items.
[0026] According to some embodiments, the method further comprises determining a three-dimensional array of cells as a function of the indicator of the descriptive data, each input item of the plurality of input items corresponding to one of the cells.
[0027] According to some embodiments, the method further comprises determining an order of the plurality of input items within the three-dimensional array of cells as a function of the indicator of the descriptive data.
[0028] According to some embodiments, the method further comprises determining a size of at least one of the cells as a function of the indicator of the descriptive data.
[0029] According to some embodiments, the method further comprises determining, as a function of the indicator in the descriptive data, that at least one of the cells does not contain a point defined in the point cloud data.
[0030] According to some embodiments, the method further includes applying a spatial operation to points defined within the point cloud data as a function of at least one item property associated with at least one of the input items and / or the derived items.
[0031] According to some embodiments, the method further comprises applying at least one spatial operation to be performed on points defined within the point cloud data as a function of the at least one indicator of the derived item.
[0032] According to another aspect of the present invention, there is provided a processing device comprising a processing unit configured to perform the steps of the above-described method. The other aspects of the present disclosure have any of the same features and advantages as the first and second aspects described above.
[0033] At least part of the methods according to the present invention may be computer-implemented. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," or "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer-usable program code embodied in the medium.
[0034] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code for provision to a programmable apparatus on any suitable carrier medium. The tangible carrier medium may comprise a storage medium such as a floppy disk, CD-ROM, hard disk drive, magnetic tape device, or solid-state memory device. The transient carrier medium includes signals such as electric, electronic, optical, acoustic, magnetic, or electromagnetic signals, e.g., microwave or RF signals. [Brief explanation of the drawings]
[0035] BRIEF DESCRIPTION OF THE DRAWINGS Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings: [Figure 1A]FIG. 1A shows example steps and file structures for storing and / or transmitting point cloud data (or 3D data) and, for example, for accessing stored and / or transmitted point cloud data (or portions of stored and / or transmitted point cloud data), according to some embodiments of the present invention. [Figure 1B] FIG. 1B shows example steps and file structures for storing and / or transmitting point cloud data (or 3D data) and, for example, for accessing stored and / or transmitted point cloud data (or portions of stored and / or transmitted point cloud data), according to some embodiments of the present invention. [Figure 2] FIG. 2 illustrates an example of dividing a 3D volume associated with point cloud data into multiple 3D sub-volumes according to a 3D grid. [Figure 3] FIG. 3 illustrates example steps for storing or transmitting point cloud data as a grid of input G-PCC items, according to some embodiments of the present invention. [Figure 4] FIG. 4 illustrates example steps for decoding and rendering point cloud data stored as a grid of input G-PCC items, according to some embodiments of the present invention. [Figure 5] FIG. 5 illustrates example steps for storing and / or transmitting point cloud data as a 3D fusion item of an input G-PCC item, according to some embodiments of the present invention. [Figure 6] FIG. 6 illustrates example steps for decoding and rendering point cloud data stored as a 3D fusion of input G-PCC items, according to some embodiments of the present invention. [Figure 7] FIG. 7 is a schematic diagram of an example data processing apparatus configured to implement, in whole or in part, some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0036] MODE FOR CARRYING OUT THE INVENTION A limitation of the above specification is that it does not address combining several items representing 3D (three-dimensional) data into one other item representing 3D data. Examples of such combinations include, but are not limited to, combining several items into a 3D grid (each item representing the contents of a cell from the grid), or combining several items into a 3D merged item, where each item is placed in a position specified by the 3D merged item. An advantage of defining an item representing 3D data as a combination of several items is that the combined items can be encoded or decoded independently in a parallel manner.
[0037] FIG. 1A illustrates example steps for storing and / or transmitting point cloud data (or 3D data) and, for example, for accessing stored and / or transmitted point cloud data (or portions of stored and / or transmitted point cloud data), according to some embodiments of the present invention.
[0038] As shown, point cloud data 100 is divided into or acquired as multiple subsets of point cloud data (step 105), corresponding to different spatial subsets or subsets acquired from different sensors or different viewpoints, and each subset is processed in parallel, e.g., to be encoded in parallel (steps 110-1 through 110-m). For each point cloud data subset, the processed data is described in an input item, which may be referred to as a point cloud item. A point cloud item may be encoded, for example, in MPEG-I Part-5 or MPEG-I Part-9, and describes point cloud data that may be referred to as a V-PCC (or V3C) point cloud item or a G-PCC point cloud item, respectively. When used as input for derivation or within a group of items, it may also be referred to as an input point cloud, a V-PCC (or V3C) input item, or a G-PCC input item. When using only "input item" or "input point cloud" within the context of the present invention, it refers to a point cloud item, whatever its encoding format. The resulting set of input items are then encapsulated (step 115) to generate an encapsulated data file 120 that can be stored locally or transmitted to a remote server to be stored and / or processed, e.g., representing the point cloud data (or portions of the point cloud data). The encapsulated data file includes input items describing the processed point cloud data subsets and derived items containing descriptive data that allows combining the processed point cloud data subsets and the point cloud data of each subset. More generally, derived items are items that define, via item references of type "dimg," operations to be performed on input items associated with the derived item. For example, such descriptive data may represent the spatial configuration of the point cloud data described by the input items.Additionally, the encapsulated file contains item references to associate the derived items with the input items.
[0039] The encapsulated data file may be in ISOBMF format.
[0040] To access the point cloud data (or portions of the point cloud data) encapsulated within the encapsulated data file, the latter is parsed (step 125) to obtain derived items. Information from the derived items can be used to access input items describing subsets of the processed point cloud data, e.g., to access point cloud data that can be processed in parallel so that they can be decoded in parallel (steps 130-1 to 130-n). The processed point cloud data is then combined (step 135) as a function of the information from the derived items to generate valid point cloud data 140. As another example, the point cloud data can be decoded and stored in a memory structure adapted for 3D rendering in parallel (steps 130-1 to 130-n) and then combined by 3D rendering (step 135) to generate a 3D rendering of the point cloud data.
[0041] Note that steps 105, 110-1 through 110-m, and 115 may be performed on server 145, or on several servers (or processing devices). Similarly, steps 125, 130-1 through 130-n, and 135 may be performed on server 150, a client, or on several servers or processing devices.
[0042] FIG. 1B illustrates an example file structure for storing and / or transmitting point cloud data (or 3D data) and, for example, accessing stored and / or transmitted point cloud data (or portions of stored and / or transmitted point cloud data), according to some embodiments of the present invention.
[0043] As shown, encapsulated data file 120 comprises a derived item 155 (e.g., an item comprising a GPCCGrid structure within its body) identified in item reference 160 (e.g., a “SingleItemTypeReferenceBox” of type “dimg” within an “ItemReferenceBox” “iref,” as described below), which also identifies multiple input items 165-1 through 165-n (e.g., input G-PCC items, V-PCC items, or V3C items) that describe subsets 170-1 through 170-n of point cloud data. Thus, a point cloud may be stored and / or transmitted as multiple subsets of point cloud data and rendered while enabling the subsets of point cloud data to be processed in parallel.
[0044] For illustrative purposes, two use cases are described below: according to the first use case, the point cloud data is divided into multiple spatial subsets according to a 3D grid, and according to the second use case, the point cloud data is divided into multiple subsets according to another criterion, for example according to the sensor used to acquire the point cloud data, i.e., according to a 3D fusion criterion.
[0045] 3D Grid According to certain embodiments, the 3D volume associated with the point cloud data is divided into multiple 3D sub-volumes, and the point cloud data for each of these 3D sub-volumes is processed independently and described in specific items called input items.
[0046] Note that for purposes of illustration, the following description focuses primarily on point cloud data encoded as a G-PCC stream according to MPEG-I Part-9 and stored as G-PCC items according to MPEG-I Part-18. However, it can easily be extended to point cloud data encoded as a V-PCC stream according to MPEG-I Part-5 and stored according to MPEG-I Part-10, or more generally any 3D content stored according to ISOBMFF.
[0047] FIG. 2 illustrates an example of dividing a 3D volume associated with point cloud data into multiple 3D sub-volumes according to a 3D grid.
[0048] According to the illustrated example, the 3D grid 200 includes 12 cells arranged in a 3 x 2 x 2 array. Of course, the 3D volume associated with the point cloud data may be divided into the same or a different number of cells, as desired.
[0049] The coordinates of points within this 3D volume are defined as a function of a reference frame, e.g., reference frame 205, defined by three axes: X-axis 205-1, Y-axis 205-2, and Z-axis 205-3. The origin of reference frame 205 may correspond to a corner of a grid cell (as shown) or may be located elsewhere. Reference frame 205 defines a global reference frame shared by all points within the 3D volume. Alternatively, coordinates may be expressed in terms of three angles, such as (azimuth, elevation, and tilt) or (yaw, pitch, and roll).
[0050] Additionally, a local reference frame may be used within each cell of the 3D grid to identify points within this cell. For illustrative purposes, the origin of such a local reference frame may correspond to a corner of the cell, e.g., the one closest to the origin of the global reference frame or the one with the smallest coordinate in the global reference frame. The axes of the local reference frame may be parallel to the axes of the global reference frame. An example of a local reference frame associated with a cell of the 3D grid is shown at 210. For further illustrative purposes, the location of point 215 may be defined using global reference frame 205 or using local reference frame 210.
[0051] According to certain embodiments, a 3D grid G-PCC item may be represented as a derived point cloud item with an "item_type" having a particular value, for example, the value 'grd3'. This 3D grid may be used to combine one or more input G-PCC items in a grid. Thus, an item with an item_type value of 'grd3' defines a derived point cloud item whose reconstructed point cloud is formed from one or more input point clouds in a given 3D grid order. Data associated with the derived point cloud item may specify some properties of the grid and may have the following structure:
[0052] TIFF0007818144000001.tif7692
[0053] where the "version" field signals the version of the structure and the "flags" field is a map of flags that may signal some options of the structure. "version" or "flags" must be equal to 0 if the version is not defined or the flags are not defined, respectively.
[0054] The "width_minus_one", "height_minus_one", and "depth_minus_one" fields represent the number of cells in the grid along the X, Y, and Z axes.
[0055] In this grid, the 0-based index i x , i y , and i z The cell located at is c along the X axis. x *i x From c x *(i x +1) and extends along the Y axis to c y *i y From c y *(i y +1) and extends along the Z axis to c z *i z From c z *(i z +1), and c x , c y , and c z are the sizes of the cells along the X, Y, and Z axes.
[0056] The "same_origin" flag indicates whether the point cloud data associated with the input G-PCC items to be combined in the grid use the same reference frame. If the value of the "same_origin" flag is set to a first value, e.g., 1, all input G-PCC items are defined using the same reference frame or coordinate system, e.g., the global reference frame 205. In this case, the input G-PCC items can be combined in the grid without applying any transformations or transformations to them to insert them into the grid. Alternatively, if the value of the "same_origin" flag is set to a second value, e.g., 0, each input G-PCC item is defined using its own reference frame, e.g., the local reference frame of the cell in which the point cloud data of the input G-PCC item is located. In this latter case, a transformation can be applied to the coordinates of the points of the input G-PCC items to calculate their coordinates in the global reference frame associated with the grid for insertion into the grid. This transformation is as follows:
[0057] TIFF0007818144000002.tif1761
[0058] where x g , y g , and z g are the coordinates of the point in the global reference frame, and x l , y l , and z l are the coordinates of the point in the local reference frame of the input G-PCC item. As mentioned above, cx, cy, and cz are the size of the cell along the X, Y, and Z axes, and i x , i y , and i z is the 0-based index of the input G-PCC item's position in the grid.
[0059] The transformation defined above converts the coordinates of a point expressed in the local reference frame of the input G-PCC item containing this point into coordinates expressed in the local reference frame of the grid's cell located at position (0,0,0), which can be used as the grid's global reference frame. Of course, another global reference frame can be used, for example, by defining a transformation from the local reference frame of the grid's cell located at position (0,0,0) to this other global reference frame through an item property associated with the 3D grid-derived item. As another example, this transformation may be specified inside the GPCCGrid structure as a 3D vector such as:
[0060] TIFF0007818144000003.tif3195
[0061] Here, the "origin_precision" field specifies the precision used to encode the position of the global origin, and the "global_origin" field is a 3D vector representing the position of the global origin, which may be represented using the Vector3 structure specified by MPEG-I Part-7 (ISO / IEC 23090-7). Alternatively, one may encode the position of the origin of the local reference frame of a grid cell located at position (0,0,0) in the global reference frame.
[0062] Perhaps when the value of the "same_origin" flag is set to the second value (0 in the example above), the input G-PCC item may be centered inside the grid cell or aligned with one or more faces of the grid cell.
[0063] Conceivably, when the value of the "same_origin" flag is set to the first value (1 in the example above), the input G-PCC item may have a size smaller than the size of a grid cell. The location of the entire grid may be determined from the bounding box specified in the item properties associated with the 3D grid item. For each axis, the location of the entire grid may be determined by finding the input G-PCC item with the largest size along this axis, and from the extent of this G-PCC item along this axis, determining the location along this axis of the grid cell containing this input G-PCC item. For example, the location of the grid cell along this axis may be determined such that the input G-PCC item is centered inside the grid cell along this axis, or such that the input G-PCC item is aligned with one of the boundaries of the grid cell along this axis. The location of the entire grid may be specified inside the GPCCGrid structure as a 3D vector, for example, using the following structure:
[0064] TIFF0007818144000004.tif115104
[0065] Here, the "grid_precision" field specifies the precision used to encode the grid position, and the "grid_position" field is a 3D vector representing the position of the corner of the grid with the smallest coordinate.
[0066] The "same_origin" field highlights the difference between 2D images and 3D point clouds. In 2D images, pixel coordinates are implicitly specified, with pixels signaled from top-left to bottom-right according to row-first scanning order. In 3D point clouds, point coordinates are explicitly specified by encoding their location as x, y, z. This difference lies in the nature of pixels in 2D images: each pixel corresponds to a square or rectangular area associated with an information item (usually color and sometimes transparency). In 3D point clouds, points do not have volume and are therefore punctiform. Furthermore, 3D point clouds are typically sparse, with many empty areas in 3D point clouds. Therefore, to achieve accurate and efficient encoding of 3D point clouds, the location of a point or set of points can be explicitly specified. As a result, for 2D images, pixel coordinates always start at (0,0) for the top-left pixel and end at (width, height) for the bottom-right pixel, while for 3D point clouds, point coordinates can use any reference frame.
[0067] In the GPCCGrid structure, the "morton_order" field indicates the order of the input G-PCC items. According to some embodiments, the input G-PCC items are signaled in a "SingleItemTypeReferenceBox" of type "dimg" within the "ItemReferenceBox" "iref". In this "SingleItemTypeReferenceBox", the value of the "from_item_ID" field identifies the derived point cloud item. In these embodiments, the input G-PCC items can be retrieved by looking for a "SingleItemTypeReferenceBox" of type "dimg" whose "from_item_ID" field value is the identifier associated with the 3D grid G-PCC item. The value of "reference_count" is preferably equal to (width_minus_one + 1) * (depth_minus_one + 1) * (height_minus_one + 1). Furthermore, according to some embodiments, the value of the "to_item_ID" field identifies the input G-PCC items. The "SingleItemTypeReferenceBox" may be referred to as an item reference.
[0068] If the value of the "morton_order" field is set to a first value, e.g., 0, the input G-PCC items are ordered in line scan order, with X increasing first, then Y increasing, then Z increasing. In some variations, the line scan order may arrange the axes differently. If the value of the "morton_order" field is set to a second value, e.g., 1, the input G-PCC items are ordered in Morton order (or Z scan order). The use of Morton order allows for better preservation of local characteristics of data within a grid; using line scan order, input G-PCC items corresponding to adjacent cells may be far apart, while using Morton order, the G-PCC items are placed closer together.
[0069] The order of input G-PCC items is the order in which they are listed in the to_item_ID field of the SingleItemTypeReferenceBox. Preferably, this is also the order in which their associated data is stored in a file or transmitted over a network.
[0070] According to some embodiments, the input G-PCC items have the same size. This size is the size of a cell in the grid. This size may be specified by one or more item properties of the input G-PCC items, such as a bounding box item property or a size item property. This size may be specified in the GPCCGrid structure. If a size is specified in both one or more item properties of the input G-PCC items and in the GPCCGrid structure, the specified sizes are preferably all the same. If the value of the has_cell_size field of the GPCCGrid structure is set to a first value, such as 0, the size of the grid's cells is not specified in the GPCCGrid structure. In such a case, all bounding boxes of the input point clouds shall have the same size, which is the size of a cell. If the desired input point clouds are not of a consistent size, a derived point cloud item (e.g., of type "iden," which represents an identifier transformation) can be used to scale or crop them as necessary to make them consistent, although other specifications may limit whether a derived point cloud item is acceptable as input to a point cloud grid derived item.
[0071] If the value of the "has_cell_size" field of the GPCCGrid structure is set to a second value, for example 1, the size of the cells of the grid is specified inside the GPCCGrid structure by the "cell_size" field. The "precision" field specifies the precision used to encode the cell size. The "cell_size" field is a 3D vector representing the size of the cell, which can be represented using the Vector3 structure specified by MPEG-I Part-7 (ISO / IEC 23090-7).
[0072] If you delete an item that is marked as an input point cloud for a point cloud grid item, you may need to rewrite the content of the point cloud grid item.
[0073] Preferably, all points from an input G-PCC item are located inside the boundary of the grid cell corresponding to that input G-PCC item. When the "same_origin" flag is set to a first value, for example 1, all points from the input G-PCC item preferably have coordinates within the boundary of the grid cell in which the input G-PCC item is located. When the "same_origin" flag is set to a second value, for example 0, after applying the above transformation, all points from the input G-PCC item preferably have coordinates within the boundary of the grid cell in which the input G-PCC item is located.
[0074] FIG. 3 illustrates example steps for storing or transmitting point cloud data as a grid of input G-PCC items, according to some embodiments of the present invention.
[0075] As shown, the first step is directed to acquiring point cloud data for storage or transmission (step 300). These point cloud data may be acquired directly from a sensor, such as a LiDAR, or may result from processing one or more point clouds acquired by one or more sensors. For example, several point clouds may be captured by a LiDAR, registered, transformed to use the same frame of reference, and then combined into a single point cloud.
[0076] Next, target characteristics of the grid to be used are obtained (step 305). According to some embodiments, the size of each cell of the grid is obtained and the number of cells in each dimension is calculated using this cell size and the spatial size of the point cloud obtained in step 300. Alternatively, the number of cells in each dimension can be obtained and the size of each cell can be calculated using the spatial size of the point cloud.
[0077] Additionally, the ordering of the input items in the grid can be obtained. As mentioned above, this ordering may be, for example, line scan order or Morton order. Note that if no order is provided, a default ordering may be defined, for example line scan ordering.
[0078] Additionally, an indication may be obtained to indicate whether the same origin should be used for all input G-PCC items. In some cases, when using the same origin for all input G-PCC items, a transformation to be applied to all input G-PCC items before encoding them may be obtained. According to some embodiments, a transformation to be applied to the reconstructed grid may be obtained if the same origin is not used for all input G-PCC items.
[0079] It is observed that step 305 may be performed before or simultaneously with step 300.
[0080] Next, the point cloud data obtained in step 300 is divided (step 310) into subsets of point cloud data as a function of the cells determined according to the grid characteristics obtained in step 305. As a result, since the grid may be sparse (i.e., it may include some cells that do not contain any points from the point cloud data), a point cloud corresponding to the subset of the point cloud obtained in step 300 is obtained for each of the cells of the grid, or for some cells of the grid.
[0081] Each subset of the point cloud data is then encoded (step 315). According to some embodiments, this encoding step is performed independently, e.g., in parallel, for each subset of the point cloud data. Further, according to some embodiments, the encoding step is based on the MPEG-I Part-9 specification.
[0082] During this step, if the same origin is used for all input G-PCC items, the transformation obtained in step 305 may be applied to each point of the point cloud data before or during the encoding of the point cloud data. If the same origin is not used for the input G-PCC items, the points of each subset of the point cloud data may be transformed as a function of the local reference frame associated with the corresponding cell of the grid before or during the encoding of the point cloud data.
[0083] Next, a structure for encapsulating the grid and the encoded point cloud data is generated (step 320). For example, the GPCCGrid structure described above may be created to describe the grid and store its characteristics. Additionally, a G-PCC item may be created to describe each subset of the encoded point cloud data. In some cases, each subset of the encoded point cloud data may be described within either a single G-PCC item or several G-PCC items. As described above, a "SingleItemTypeReferenceBox" of type "dimg" may be created to link the input G-PCC items describing the subsets of the encoded point cloud data with items representing the grid (i.e., derived items). The ordering of the input G-PCC items within the "SingleItemTypeReferenceBox" may depend on the ordering obtained in step 305.
[0084] It should be noted that step 320 (or some of its sub-steps) may be performed in parallel with or as part of encoding step 315 .
[0085] The subset of point cloud data encoded in step 315 and the encapsulation structure created in step 320 containing item references (e.g., "SingleItemTypeReferenceBox" of type "dimg") establishing the link between the derived item and the input item are then stored in an encapsulated data file, e.g., an ISOBMFF file. Before being stored, the encoded subset of point cloud data may be ordered according to the order obtained in step 305. Optionally, the ISOBMFF file is stored in a memory unit (e.g., random access memory) or a hard disk drive. Optionally, the ISOBMFF file is transmitted over a communications network.
[0086] It is observed that encoding the point cloud data as a grid can advantageously increase the speed of encoding the point cloud data, since several encoding processes can be performed in parallel.
[0087] According to some embodiments, an identifier derived point cloud item may be represented as an item with an "item_type" having a particular value, for example, the value "iden". This identifier derived point cloud item may be used when it is desired to use transformation properties to derive a point cloud item. A derived point cloud item has no item body (i.e., no extent), and the "reference_count" of the 'dimg' item reference for the 'iden' derived point cloud item is equal to 1.
[0088] According to some embodiments, the point cloud data is acquired in chunks at step 300 (though the chunks are not all acquired at the same time). According to these embodiments, one or more chunks of the point cloud data are processed independently of the remaining chunks. The one or more chunks are divided into subsets of point cloud data at step 310. Then, for each subset of point cloud data, a test is performed to determine whether the subset of point cloud data includes all of the data for its corresponding grid cell or whether some of the cell's data is missing. If the subset of point cloud data does include all of the data for its corresponding grid cell, it can be encoded as described above in connection with step 315, a next encapsulation structure can be created and signaled as described above in connection with step 320, and this encoded subset of point cloud data and its encapsulation structure can be stored as described above in connection with step 325.
[0089] Preferably, in these embodiments, the boundaries between cells of the grid are selected such that the boundaries between chunks coincide with the boundaries of the grid. For example, if the point cloud data is obtained from a rotating LiDAR and each chunk corresponds to a quadrant of the volume captured by the rotating LiDAR, the limits corresponding to these quadrants can define the limits between cells of the grid.
[0090] Furthermore, according to some embodiments, the point cloud data to be stored as a grid of point cloud data has already been divided into subsets of point cloud data, and each subset of point cloud data has already been encoded.
[0091] In such an embodiment, a subset of the encoded point cloud data is obtained in step 300 and grid characteristics are obtained in step 305 .
[0092] The size of each cell of the grid may be obtained from the spatial size of one of the subsets of encoded point cloud data obtained in step 300. In some cases, if the subsets of encoded point cloud data have different spatial sizes, the size of a cell of the grid may be defined as the maximum spatial size of the subset of encoded point cloud data. For example, the size of a cell along the X-axis of the grid may be determined by identifying the maximum spatial size of the subset of encoded point cloud data along the X-axis. The sizes along the Y- and Z-axes of the grid may be determined similarly. The spatial size of the subset of encoded point cloud data may be obtained from a bounding box associated with the subset of encoded point cloud data, which may be calculated from the contents of the encoded point cloud data.
[0093] In some cases, where the subsets of the encoded point cloud data have different spatial sizes, the size of the cells of the grid may be determined as the smallest spatial size, the average spatial size, or the median spatial size of the subsets of the encoded point cloud data.
[0094] In some cases, when the encoded subsets of point clouds have different spatial sizes, each point cloud can be centered within its cell by applying a transformation to that cell. This transformation can be applied by re-encoding the point cloud or by defining a transformation property associated with the input G-PCC item used to encapsulate the subset of point cloud data.
[0095] In some cases, if the spatial size of the subset of encoded point cloud data is larger than the size of the cell, it may be cropped to fit into the cell. This cropping may be performed before or after applying the transformation to the points of the subset of encoded point cloud data.
[0096] The ordering of the input items may be obtained explicitly as described above, or it may also be obtained implicitly, for example, using the order in which the subsets of encoded point cloud data are obtained in step 300.
[0097] An indication of whether the same origin should be used for all input G-PCC items may be explicitly obtained, as described above. It may also be determined from the obtained subsets of encoded point cloud data. For example, the minimum and maximum coordinate values along each axis for each subset of encoded point cloud data may be calculated, and if the minimum and maximum coordinate values are the same for all encoded point clouds, it may be concluded that they do not have the same origin. Otherwise, it may be determined that they have the same origin. For further explanation, the size along each axis for each subset of encoded point cloud data may be calculated. If the difference between the maximum minimum coordinate value along an axis for a subset of encoded point cloud data and the minimum minimum coordinate value along the same axis for another subset of encoded point cloud data is greater than or equal to the maximum size along the same axis for the subsets of encoded point cloud data, it may be concluded that the point clouds have the same origin.
[0098] In some cases, the transformation can be obtained as described above.
[0099] In some cases, the location of points in each subset of encoded point cloud data within a grid can be explicitly obtained. In this case, an indication as to whether the same origin is used for input G-PCC items can be obtained by comparing the coordinates used by points of subsets of encoded point cloud data located in adjacent cells. For example, if a point in a first subset of encoded point cloud data is located at index i along an axis and a point in a second subset of encoded point cloud data is located at index i+1 along the same axis, it can be concluded that the same origin is used for the different point clouds if the maximum coordinate value along this axis for a point in the first subset is smaller than the minimum coordinate value along this axis for a point in the second subset.
[0100] In some cases, when the input G-PCC items have the same origin, the location of points in each subset of the encoded point cloud data within a grid can be implicitly obtained. The size of the grid's cells can be used to calculate the range of coordinates for each cell in the grid. Then, for points in each subset of the encoded point cloud data, the minimum and maximum coordinate values along each axis are calculated. These minimum and maximum coordinate values can be used to obtain the location of the cell corresponding to these coordinates.
[0101] In this alternative, steps 310 and 315 are not performed. Depending on how the subsets of encoded point cloud data are obtained in step 300, step 320 may be performed in full or in part. If the encoded point cloud data is obtained as a compressed stream, such as specified by MPEG-I Part-9, input G-PCC items may be created to describe each subset of encoded point cloud data. If subsets of encoded point cloud data are obtained as input G-PCC items, these items may be used directly. Other portions of step 320 may be performed as described above.
[0102] In this alternative, step 325 is performed as described above.
[0103] FIG. 4 shows example steps for decoding and rendering point cloud data stored as a grid of input G-PCC items, according to some embodiments of the present invention.
[0104] As shown, the first step is directed to obtaining an encapsulated data file storing a grid of input items, e.g., an ISOBMFF file storing a grid of input G-PCC items (step 400). In addition, an indication of the first item (i.e., grid item) to decode and render may be obtained. If an indication of the first item to decode and render is not obtained, the first item indicated in the encapsulated data file may be considered to be the first item to decode and render.
[0105] Next, the grid items to be decoded and rendered are obtained from the ISOBMFF file (step 405). The grid items may be encapsulated in the ISOBMFF file according to the GPCCGrid structure described above. For example, they may be parsed to obtain the size of the grid, whether the input G-PCC items use the same origin, whether the input G-PCC items are ordered using line scan order or Morton order, and whether the cell size is included in the structure. If the cell size is included in the structure, it may be parsed from the ISOBMFF file to obtain this size.
[0106] Additionally, the "SingleItemTypeReferenceBox" of type "dimg" associated with the grid can be parsed to identify the input G-PCC items used by the grid. If the cell size is not included in the GPCCGrid structure, it can be obtained using the size of the input G-PCC items. For example, in certain embodiments, all input G-PCC items have the same size, and the cell size is obtained using the size of any of the input G-PCC items. In a variant, the cell size along an axis can be calculated as the maximum size of the input G-PCC items along this axis.
[0107] A subset of the encoded point cloud data of the input G-PCC items identified by the “SingleItemTypeReferenceBox” of type “dimg” associated with the grid is then retrieved from the ISOBMFF file.
[0108] Next, the subset of the encoded point cloud data is decoded and rendered (step 410). Decoding the subset of the encoded point cloud data, i.e., decoding the point cloud, may be performed, for example, according to the MPEG-I Part-9 specification. Rendering may include generating structures in a memory unit (e.g., random access memory) to represent the decoded point cloud, sending these structures to a graphics card, and generating a display of the point cloud.
[0109] If the "same_origin" field of the GPCCGrid structure used is set to the second value, i.e., 0 according to the example given above, the positions of the points of the decoded point cloud may be transformed into the reference frame corresponding to the grid, for example, using the transformations described above. In some cases, a further transformation may be applied to the decoded points to change the origin of the reference frame used by the grid. This transformation may be applied during decoding of the point cloud, after decoding it, before rendering it, or as part of the rendering process.
[0110] If the points in the subset of encoded point cloud data do not use the same origin, the points of the decoded point cloud may be transformed before being rendered. In some cases, this transformation can be performed as part of the decoding or as part of the rendering.
[0111] Preferably, decoding and rendering are performed independently for each subset of the encoded point cloud data, to the extent possible. For example, decoding, generating structures in a memory unit to represent the decoded point cloud, and transmitting these structures to the graphics card may be performed in parallel for several subsets of the encoded point cloud data, i.e., for several point clouds. Advantageously, the memory structures used to represent the decoded point clouds may be organized so that the most frequently used memory structures are only used for decoding and rendering one decoded point cloud. For example, the representation of 3D space may be divided into several subvolumes, each represented using one or more memory structures. These subvolumes may be structured so that each subvolume corresponds to a cell of a grid. In this way, during decoding of a subset of the encoded point cloud data, only the memory structures representing the relevant subvolume are accessed, and these memory structures are accessed only during decoding of this subset of the encoded point cloud data. In this example, several global memory structures may be used to organize the memory structures corresponding to each subvolume. However, these global memory structures are mostly accessed in step 405 while parsing the grid, not while decoding the subset of the encoded point cloud data. Possibly, several sub-volumes may correspond to one cell of the grid, with each sub-volume corresponding to only one cell of the grid.
[0112] Decoding a point cloud that has previously been encoded as a grid of subsets of point data may advantageously increase the speed of decoding the point cloud, as some decoding operations may be performed in parallel, and portions of the rendering process may also be performed in parallel.
[0113] In some embodiments, step 410 can further optimize the decoding and rendering of the point cloud encoded as a grid. In these embodiments, the grid description contained in the GPCCGrid structure can be used to calculate the location of each cell of the grid. Thus, the grid cells can be filtered according to their location to select only those cells relevant for rendering. For example, only those cells corresponding to point cloud data visible within the viewport used for rendering can be selected. As another example, only those cells corresponding to point cloud data visible within the viewport and close to the viewpoint can be selected. Memory structures for rendering the point cloud are created. These structures can be adapted to the selected cells. For example, the rendering space can be divided along limits corresponding to the boundaries between the cells of the grid. Thus, only subsets of the encoded point cloud data corresponding to the selected cells are parsed from the ISOBMFF file. These subsets of encoded point cloud data are decoded and stored in memory structures before being rendered in parallel, again in parallel.
[0114] According to other embodiments, one or more input G-PCC items may extend outside the boundaries of their grid cells. This may be useful, for example, when storing the output of a rotating LiDAR as a grid. In practice, a rotating LiDAR captures a point cloud by using an array of vertically aligned lasers. This array allows for capturing points in a vertical plane. To capture points within a volume, the array is rotated along the vertical axis. Thus, a complete rotation of the LiDAR can be divided into a grid of four cells, each corresponding to a quadrant of the volume scanned by the LiDAR. This allows the points of each quadrant to be coded as soon as they are captured, without waiting for the complete capture of the entire volume.
[0115] However, in some rotating LiDARs, the lasers are not perfectly aligned in the same vertical plane, and there is a slight offset between the lasers. This can be attributed to improving the compactness of the laser array. This means that points captured during a quarter rotation of the LiDAR can extend outside of their corresponding quadrants. While it is possible to filter the captured points and reassign them to their correct quadrants, this makes the encoding process more complex and introduces some latency into the encoding process. To address this issue, the above grid can be modified to allow points in the input G-PCC items to extend outside the boundaries of their grid cells. The structure describing the grid can be modified as follows to support this feature:
[0116] TIFF0007818144000005.tif82103
[0117] Here, the "overlap" field indicates whether the input G-PCC items fit within their grid cells or may extend outside of them. If the value of this field is set to a first value, e.g., 1, the points of the input G-PCC items may extend outside of their grid cells. Alternatively, if the value of this field is set to a second value, e.g., 0, the points of the input G-PCC items are completely contained within their grid cells.
[0118] According to these embodiments, the encoder determines whether input G-PCC items extend outside their grid cells. This determination may be performed while obtaining grid characteristics in step 305 of FIG. 3. When building a grid using an already encoded point cloud, this determination may be performed by checking the encoded point cloud characteristics. Thus, if an input G-PCC item is determined to extend outside the grid cells, the "overlap" field is set to a first value (i.e., 1 in the above example) when creating the GPCCGrid structure in step 320 of FIG. 3. Otherwise, this flag is set to a second value (i.e., 0 in the above example).
[0119] Further, according to these embodiments, the decoder determines whether an input G-PCC item extends outside a grid cell by checking the value of the "overlap" field in step 405 of Figure 4. If it is determined that the value of the "overlap" field is set to a first value (i.e., 1 according to the example given above), then in step 410 of Figure 4, points of the input G-PCC item located outside the boundary of that cell may be moved to a structure corresponding to the area in which they are located. Possibly, points of the input G-PCC item located outside the boundary of its cell may be discarded, which means that the input G-PCC item is cropped to fit within the boundary of that grid cell.
[0120] If it is determined that the value of the "overlap" field is set to the second value (i.e., 0 according to the example above), step 410 of FIG. 4 is implemented as described above, specifically by performing decoding and rendering as independently as possible for each subset of the encoded point cloud data.
[0121] In variations of these embodiments, the structure describing the grid may indicate whether input G-PCC items that extend outside the boundaries of the cells of that grid should be cropped. In another variation, the extension of input G-PCC items outside the boundaries of a grid cell can be specified with finer granularity. For example, an "overlap" field can be associated with each input G-PCC item to indicate whether the input G-PCC item extends outside the boundaries of that grid cell. As another example, an "overlap" field can be associated with each input G-PCC item for each of its neighboring grid cells to indicate whether the input G-PCC item extends within that neighboring grid cell.
[0122] In some embodiments, the cell limits may more precisely define a face shared between two cells that belongs to only one of those cells, e.g., the cell with the highest coordinate. Similarly, an edge or vertex shared between two or more cells may belong to only one of those cells, e.g., the cell with the highest coordinate. In these embodiments, an input G-PCC item may extend outside a cell of that grid if it contains a point located on a face of that cell that does not belong to that cell.
[0123] Additionally, in some embodiments, the grid may be sparse, with one or more of the cells not containing any points. Empty cells may be represented using G-PCC items that do not contain any points. As an optimization, the GPCCGrid structure may indicate which cells contain points and which cells are empty, for example:
[0124] TIFF0007818144000006.tif106109
[0125] Here, the "sparse" field indicates whether the grid is sparse. If the value of this field is set to a first value, e.g., 1, the grid is sparse, and for each cell, the "occupied" field indicates whether the cell contains one or more points. If the value of the "occupied" field corresponding to a given cell is set to a first value, e.g., 1, this cell is associated with a G-PCC item in the list of input G-PCC items signaled in the "SingleItemTypeReferenceBox" of type "dimg". Otherwise, if the value of the occupied field corresponding to a given cell is set to a second value, e.g., 0, this cell is not associated with any G-PCC item in the list of input G-PCC items signaled in the "SingleItemTypeReferenceBox" of type "dimg" (i.e., the cell is empty).
[0126] The "occupied" values of the different input G-PCC items may be listed according to their order.
[0127] Furthermore, according to some embodiments, the reference frame used for the input G-PCC items may vary from item to item, with some input G-PCC items using the same reference frame as the grid, while other input G-PCC items may use a reference frame that is local to their cell. These different reference frames may be indicated thanks to the following structure:
[0128] TIFF0007818144000007.tif11197
[0129] "width" is equal to "width_minus_one" plus 1, "height" is equal to "height_minus_one" plus 1, "depth" is equal to "depth_minus_one" plus 1, and the "same_origin" field has three different values. If the value of this field is set to a first value, e.g., 0, each input G-PCC item is defined using its own reference frame. If the value of this field is set to a second value, e.g., 1, all input G-PCC items are defined using the same reference frame. If the value of this field is set to a third value, e.g., 2, different reference frames may be used for different input G-PCC items. For each cell, the "global_cell_origin" field indicates whether the corresponding input G-PCC item uses its own reference frame or a common reference frame. If the value of the "global_cell_origin" field associated with a cell is set to a first value, e.g., 0, the corresponding input G-PCC item uses its own reference frame, and if the value of the "global_cell_origin" field associated with a cell is set to a second value, e.g., 1, the corresponding input G-PCC items use a common reference frame.
[0130] For input G-PCC items that use their own reference frame, a transformation can be applied to their points to convert them to the global reference frame using the following relationship:
[0131] TIFF0007818144000008.tif2032
[0132] Additionally, a second transformation may be applied to use a different reference frame than the one of the cells located at position (0,0,0). This transformation may be defined in the structure by the "global_origin" field.
[0133] The "global_cell_origin" fields may be ordered according to the order of the input G-PCC items. Furthermore, in some embodiments, when one or more input G-PCC items are defined using their own reference frame, e.g., when the value of the "same_origin" flag is set to a second value, e.g., 0, one of the input G-PCC items, a reference input G-PCC item, can be used to define the reference frame of the grid. This reference input G-PCC item can be the input G-PCC item located at position (0,0,0). It can be another input G-PCC item signaled in the grid structure.
[0134] The transformation between the reference frame of the grid and the reference frame of the cell containing the reference input G-PCC item can be calculated by aligning the reference input G-PCC item with that grid cell and calculating the transformation between them. If the size of the reference input G-PCC item is the size of the grid cell, this alignment can be achieved directly. If the size of the reference input G-PCC item is smaller than the size of the grid cell, the reference input G-PCC item can be aligned with the grid cell by being centered inside the grid cell, aligned to one or more faces of the grid cell, or a combination of the two.
[0135] Furthermore, in some embodiments, the input G-PCC items of the grid may have different sizes, and an input G-PCC item may correspond to several cells of the grid. In these embodiments, the grid can be described by the following structure:
[0136] TIFF0007818144000009.tif10287
[0137] Here, the locations and sizes of the different input G-PCC items can be signaled by the "item_width_minus_one, item_depth_minus_one", and "item_height_minus_one" fields. For this signaling, the cells of the grid are scanned according to the order specified by the "morton_order" field. For each cell, if it corresponds to the aforementioned input G-PCC item, it is skipped. Otherwise, the cell is the current cell and is located at position (i x , i y , i z ) The next entry in the “SingleItemTypeReferenceBox” of type “dimg” associated with the grid corresponds to the current input G-PCC item. The number of cells that span this current input G-PCC item is specified by the current “item_width_minus_one”, “item_height_minus_one”, and “item_depth_minus_one” fields. x , s y and s z The current input G-PCC item is specified by the value i along the X axis. x From i x +s x along the Y axis to -1 y From i y +s y -1, and along the Z axis i z From i z +s z It spans cells up to -1.
[0138] Furthermore, according to some embodiments, the bounding box of a grid-derived item may be specified and may differ from the bounding box of the grid. For example, the bounding box of the grid may be smaller than the grid, meaning that the point cloud obtained by combining the input G-PCC items in the grid is cropped to obtain the point cloud corresponding to the grid-derived item. As another example, the bounding box of the grid may have an extent larger than the grid itself. In this variant, the grid may be described by the following structure:
[0139] TIFF0007818144000010.tif9392
[0140] Here, the bounding box of the grid is specified by the "output_position" and "output_size" fields. The "output_precision" field specifies the precision used for the "output_position" and "output_size" fields.
[0141] In some cases, the bounding box of the grid may be specified using an item property associated with the grid derived item, such as a bounding box item property or a size item property. In some cases, the bounding box of the grid may be specified using a crop transform item property associated with the grid derived item.
[0142] 3D fusion It has been observed that there are cases where a 3D volume can be divided according to an irregular method that does not correspond to a grid. For example, a scene may be captured by LiDARs placed successively at different locations, and the resulting point cloud data may be divided into several subsets of point cloud data centered at the different capture locations. This resulting point cloud can be represented as a fusion of the several subsets of point cloud data.
[0143] A 3D fused G-PCC item may be represented as an item with an item_type having a specific value, for example, the value "fus3". This 3D fused item aims to combine one or more irregularly placed input G-PCC items. Thus, an item with an "item_type" value of 'fus3' defines a derived point cloud item from which a reconstructed point cloud is formed by combining one or more input point clouds.
[0144] The input point clouds are inserted in the order of the “SingleItemTypeReferenceBox” of type “dimg” of this derived point cloud item within the “ItemReferenceBox”. In the “SingleItemTypeReferenceBox” of type “dimg”, the value of “from_item_ID” identifies the derived point cloud item of type “fus3”, and the value of “to_item_ID” identifies the input point cloud.
[0145] When deleting an item that was marked as an input point cloud of a point cloud fusion item, the contents of the point cloud fusion item may need to be rewritten.
[0146] The parameters associated with the derived point cloud item specify some properties of the 3D fusion item and may have the following structure:
[0147] TIFF0007818144000011.tif6399
[0148] Here, the "version" field signals the version of the structure, and the flags field signals some options of the structure. "version" or "flags" must be equal to 0 if there is no defined version or flags, respectively.
[0149] The "overlap" field indicates whether some input G-PCC items of a 3D fusion item may overlap. When the value of the "overlap" field is set to a first value, e.g., 1, some input G-PCC items of a 3D fusion item may or may not overlap, i.e., the intersection of the bounding boxes of any pair of input point clouds may or may not have an empty volume. When the value of the "overlap" field is set to a second value, e.g., 0, none of the input G-PCC items of a 3D fusion item overlap, i.e., the intersection of the bounding boxes of any pair of input point clouds has an empty volume. Two input G-PCC items may be considered non-overlapping if the intersection of their bounding boxes is an empty volume. In a variant, two input G-PCC items may be considered non-overlapping if the intersection of their bounding boxes is an empty volume and their bounding boxes share a face, edge, or vertex.
[0150] The "same_origin" flag indicates whether the input G-PCC items combined by the 3D fusion item use the same reference frame. If the value of the "same_origin" flag is set to a first value, e.g., 1, all input G-PCC items are defined using the same reference frame or coordinate system. In this case, the input G-PCC items can be combined without applying any transformations to them. Otherwise, if the value of the "same_origin" flag is set to a second value, e.g., 0, each input G-PCC item is defined using its own reference frame. In this latter case, a transformation can be applied to the coordinates of the points of the input G-PCC items to calculate their coordinates in the global reference frame for the 3D fusion item. For each input G-PCC item, the transformation from its local reference frame to the common reference frame may be signaled by the "anchor" field corresponding to this input G-PCC item, i.e., the i-th anchor is applied to the i-th occurrence in the ordering of the "SingleItemTypeReferenceBox" of type "dimg" for this derived point cloud item in the "ItemReferenceBox". The global coordinates for a point in the input G-PCC item may be calculated from its decoded coordinates as follows:
[0151] TIFF0007818144000012.tif1732
[0152] where x g , y g , z g are the coordinates of the point in the global reference frame, and x l , y l , z l are the coordinates of the point in the local reference frame of the input G-PCC item, and a x , a y , a zare the coordinates of the anchor of the input G-PCC item along the X, Y, and Z axes, signaled by the "anchor" field corresponding to this input G-PCC item.
[0153] The "precision" field indicates the precision used to encode the anchors of the input G-PCC items. For each input G-PCC item, the "anchor" field signals the coordinates of its anchor, which are used to transform the coordinates of the points of the input G-PCC items from the local reference frame used by the input G-PCC items to the common reference frame used by the 3D fusion items. The "anchor" fields may be ordered in the order of the input G-PCC items.
[0154] The input G-PCC items may be signaled in a "SingleItemTypeReferenceBox" of type "dimg", where the value of the "from_item_Id" field identifies the derived point cloud item. The value of the "reference_count" field indicates the number of input G-PCC items. The "reference_count" is taken from the "SingleItemTypeReferenceBox" of type "dimg", and this item is identified by the "from_item_ID" field. The value of the "to_item_ID" field identifies the input G-PCC items. Again, the "SingleItemTypeReferenceBox" is sometimes referred to as an item reference.
[0155] According to some embodiments, each input G-PCC item may have an associated "same_origin" flag. In such embodiments, a 3D fusion item may have the following structure:
[0156] TIFF0007818144000013.tif7292
[0157] FIG. 5 illustrates example steps for storing and / or transmitting point cloud data as a 3D fusion item of an input G-PCC item, according to some embodiments of the present invention.
[0158] As shown, the first step is to acquire point cloud data for storage and / or transmission (step 500). These point cloud data can be acquired directly from a sensor, such as a LiDAR. It can also be the result of processing one or more point clouds captured by a sensor. For example, several point clouds can be captured by a LiDAR, registered and transformed to use the same reference frame, and then combined into a single point cloud. Point clouds can also be acquired as a set of several point clouds to be combined together. For example, it can be a set of point clouds captured by a LiDAR at different locations in a scene, or by different LiDARs at different locations in a scene.
[0159] Next, the target properties of the 3D fusion item are obtained (step 505).
[0160] In this step, instructions may be obtained regarding how to divide the point cloud data acquired in step 500 into subsets of point cloud data. The instructions may be based on 3D regions for dividing the point cloud data. They may also be based on a list of points corresponding to each subset of point cloud data. They may also be based on characteristics of the points in the point cloud data. For example, the point cloud data may be divided based on the capture time of the points, such that each subset of point cloud data corresponds to a different time period. As another example, the point cloud data may be divided by grouping points around a set of centers, such that each subset of point cloud data includes points that are close to each other, with each point grouped with the nearest center.
[0161] Additionally, an indication of whether to use the same origin for all input G-PCC items can be obtained. Perhaps when the same origin is used for all input G-PCC items, a transformation to be applied to all input G-PCC items before encoding them can be obtained. Similarly, when the same origin is not used for the input G-PCC items, a transformation to be applied to the 3D fused items can be obtained.
[0162] According to some embodiments, it is determined whether point clouds corresponding to subsets of point cloud data overlap. This may include obtaining an indication of whether these point clouds may overlap. It may also include analyzing an indication of how the point cloud data is divided. For example, if the indication is based on 3D regions, it may be determined from these 3D regions whether point clouds corresponding to the subsets of point cloud data overlap. For example, if the 3D regions are rectangular prisms aligned with the axes of a reference frame, the point clouds corresponding to the subsets of point cloud data do not overlap. Conversely, if the 3D regions are not aligned with the axes of a reference frame, the point clouds may overlap. As another example, if the indication is based on point characteristics, it may be determined that the point clouds may overlap. This determination may be performed by directly checking whether the point clouds overlap. For example, this determination may be performed by calculating bounding boxes of the point clouds corresponding to the subsets of point cloud data and checking whether these bounding boxes overlap. In this latter case, the determination may be performed at the end of step 510 once the point cloud data has been divided into subsets of point cloud data.
[0163] In addition, an ordering of the input G-PCC items in the 3D fusion item can be obtained. This ordering can be linked to instructions on how to divide the point cloud data. For example, if the instructions are based on 3D regions for dividing the point cloud data, ordering can be associated with these 3D regions to indicate the order of the input items. As another example, if the division is based on characteristics of the points in the point cloud data, the ordering of the input items can be based on these characteristics or on an instruction link to these characteristics. For example, if the point cloud data is divided based on the capture time of the points, the input items can be ordered according to this capture time.
[0164] Step 505 may be executed before step 500 or simultaneously with step 500.
[0165] Next, the point cloud data acquired in step 500 is divided into subsets of point cloud data (step 510) according to the instructions acquired in step 505. It is observed that if the point cloud data is acquired as multiple subsets of point cloud data in step 500, step 510 is skipped.
[0166] The point cloud data for each subset of the point cloud data is then processed (step 515), e.g., encoded. Preferably, the processing of the point cloud data for each subset of the point cloud data is performed independently for each subset. For example, the point cloud data for each subset of the point cloud data may be encoded in parallel. For illustrative purposes, the encoding may be performed using the MPEG-I Part-9 specification.
[0167] During this step, if the same origin is used for all input G-PCC items, the possible transformation obtained in step 505 may be applied to the points of each subset of point cloud data before encoding them or as part of the encoding. If the same origin is not used for all input G-PCC items, the points of each subset of point cloud data may be transformed as a function of its own reference frame, as determined by its anchors, before encoding them or while encoding them.
[0168] Next, a structure is created to encapsulate the 3D fusion item and the subset of the encoded point cloud (step 520). For example, the GPCCFusion structure described above may be created to describe the 3D fusion item and store its characteristics. Additionally, a G-PCC item may be created to describe each subset of the encoded point cloud data. In some cases, each subset of the encoded point cloud data may be described using a single G-PCC item or several G-PCC items. A "SingleItemTypeReferenceBox" of type "dimg" may be created to link the input G-PCC items describing the subsets of the encoded point cloud data with the 3D fusion item. The ordering of the input G-PCC items in the "SingleItemTypeReferenceBox" may depend on the ordering obtained in step 505.
[0169] It should be noted that step 520 or some sub-steps of step 520 may be performed simultaneously with step 515 or some sub-steps of step 515 .
[0170] The encoded subset of point cloud data and encapsulation structure created in step 520, including an item reference (e.g., a "SingleItemTypeReferenceBox" of type "dimg") that establishes a link between the derived item and the input item, are then stored in an encapsulated data file (step 535), e.g., an encapsulated data file conforming to the ISOBMF format. Optionally, the encapsulated data file is stored in a memory unit or a hard drive. Optionally, the encapsulated data file is transmitted over a communications network.
[0171] Encoding point cloud data as a 3D fusion item may advantageously increase the speed of encoding the point cloud, since several encoding processes may be performed in parallel.
[0172] According to some embodiments, the point cloud data is acquired as multiple subsets of point cloud data (step 500), where different subsets of point cloud data are not acquired simultaneously. In such embodiments, one or more subsets of point cloud data may be processed independently of others. One or more subsets of point cloud data are encoded as described in connection with step 515, encapsulation structures are created to represent them as described in connection with step 520, and the encoded subsets of point cloud data and their encapsulation structures are stored and / or transmitted as described in connection with step 525.
[0173] FIG. 6 illustrates example steps for decoding and rendering point cloud data stored as a 3D fusion of input G-PCC items, according to some embodiments of the present invention.
[0174] In a first step, an encapsulated data file storing a 3D fusion of G-PCC items, for example, an encapsulated data file conforming to the ISOBMF format, is obtained (step 600). In addition, an indication of the first item to be decoded and rendered (i.e., the fusion item) may be obtained. If an indication of the first item to be decoded and rendered is not obtained, the first item signaled in the ISOBMFF file may be considered as the item to be decoded and rendered.
[0175] Next, the 3D fusion item to decode and render is obtained from the encapsulated data file (step 605). This 3D fusion item may be encoded according to the GPCCFusion structure described above. For example, it may be parsed to obtain an indication of whether the input G-PCC items use the same origin and / or whether these input G-PCC items may overlap.
[0176] Additionally, the "SingleItemTypeReferenceBox" of type "dimg" associated with the 3D fusion item can be parsed to identify the input G-PCC items used by the 3D fusion item.
[0177] Additionally, the encoded point cloud data for each input G-PCC item may be obtained from an encapsulated data file (e.g., an ISOBMFF file), and each input G-PCC item may include a subset of the encoded point cloud data.
[0178] The encoded point cloud data of the input G-PCC item (or some of the input G-PCC items) is then decoded and rendered. By way of example, the decoding of the encoded point cloud data may be performed in accordance with the MPEG-I Part-9 specification. The rendering step may include generating structures in a memory unit to represent the decoded point cloud data, transmitting these structures to a graphics card, and generating a display of the point cloud data.
[0179] If the "same_origin" field of the GPCCFusion structure is set to a second value (i.e., 0 according to the previous example), meaning that each subset of point cloud data uses its own reference frame, the positions of the points of the decoded subset of point cloud data may be transformed into the reference frame corresponding to the 3D fusion item, for example, using the transform described above. In some cases, a further transform may be applied to the decoded points to change the origin of the reference frame used by the 3D fusion item. This transform may be applied while decoding the point cloud data of the subset, after decoding them and before rendering them, or as part of the rendering process.
[0180] If the "same_origin" field of the GPCCFusion structure is set to a first value (i.e., 1 according to the previous example), meaning that the encoded point cloud data of the subset uses the same origin, the points of the decoded point cloud data subset may be transformed before being rendered. In some cases, this transformation may be performed as part of the decoding or as part of the rendering.
[0181] Preferably, decoding and rendering is performed independently for each subset of the encoded point cloud data, where possible, for example, decoding, generating structures in a memory unit to represent the decoded point cloud, and transmitting these structures to a graphics card may be performed in parallel for several subsets of the encoded point cloud data.
[0182] During decoding and rendering of the subset of encoded point cloud data, information about whether corresponding point clouds may overlap can be used to identify steps of this processing that may be performed independently on the encoded point cloud data.
[0183] In fact, if the point clouds corresponding to the decoded point cloud data subsets are signaled as non-overlapping, the memory structures used to represent these point clouds may be organized so that the most frequently used memory structures are used only for decoding and rendering one of these point clouds. For example, the representation of 3D space may be divided into several subvolumes, with each subvolume represented using one or more memory structures. These subvolumes may be constructed so that each subvolume corresponds to a bounding box of the points of one subset of the point cloud data. In this way, during decoding of a subset of the encoded point cloud data, only the memory structures representing the relevant subvolumes are accessed, and these memory structures are accessed only during the decoding of this subset. In this example, several global memory structures may be used to organize the memory structures corresponding to each subvolume. However, these global memory structures are accessed mostly during parsing the 3D fusion item in step 605, but not during decoding of the encoded point cloud data. Perhaps several subvolumes correspond to one subset of the encoded point cloud data, with each subvolume corresponding to only one subset of the encoded point cloud data.
[0184] Conversely, if point clouds corresponding to subsets of point cloud data are signaled as duplicates, the decoded point cloud data subsets may be reorganized before being rendered. For example, the encoded subsets of point cloud data may be decoded independently, and the resulting data may then be merged into memory structures representing different point clouds. As another example, the encoded subsets of point cloud data may be decoded simultaneously, and the resulting data may be stored directly into memory structures representing all decoded point cloud data subsets, while keeping the contents of the memory structures coherent in the event of simultaneous access. The organization of the memory structures may be configured such that some memory structures contain only data from a single subset of point cloud data, while some other memory structures contain data from several subsets of point cloud data. In this way, there is no need to merge data or handle simultaneous accesses to memory structures containing only data from a single subset of point cloud data.
[0185] Decoding a point cloud that has previously been encoded as a 3D fusion item may advantageously increase the speed of decoding the point cloud, as some decoding operations may be performed in parallel, and some of the rendering process may also be performed in parallel.
[0186] According to some embodiments, two input G-PCC items may be considered non-overlapping if no points of one of these input G-PCC items are contained within the bounding box of the other input G-PCC item. During step 610, the 3D space may be divided into subvolumes such that each subvolume containing one or more points from the point cloud represented by the 3D fused item corresponds to only one subset of the encoded point cloud data. In these embodiments, the subvolume corresponding to the intersection of the bounding boxes of two or more input G-PCC items is empty and does not contain any points from the point cloud represented by the 3D fused item.
[0187] Furthermore, according to some embodiments, the overlap of input G-PCC items can be specified at a finer granularity. For example, an "overlap" field can be associated with each input G-PCC item indicating whether this input G-PCC item overlaps with another input G-PCC item. In combination with the previous embodiment, this "overlap" field can indicate whether the bounding box of the input G-PCC item contains points from the other input G-PCC item.
[0188] As another example, an "overlap" field may be associated with each pair of input G-PCC items that indicates whether these two input G-PCC items overlap. In combination with the previous embodiment, this "overlap" field may indicate whether the bounding box of one of these input G-PCC items contains points from the other input G-PCC item.
[0189] In some cases, the bounding box of the 3D fused derived item may be calculated as a combination of the bounding boxes of the input G-PCC items. If the value of the "same_origin" field of the GPCCFusion structure is set to the second value (i.e., 0 according to the previous example), meaning that each subset of point cloud data uses its own reference frame, this calculation may be performed by transforming the bounding boxes of the input G-PCC items. The bounding box of the 3D fused derived item may then be calculated as spanning from the minimum coordinate value of the bounding boxes of the input items to the maximum coordinate value of these bounding boxes along each coordinate axis.
[0190] In a variant, the bounding box of the 3D fusion derived item may be specified and may be different from the combination of the bounding boxes of the input G-PCC items of the 3D fusion item. For example, the bounding box of the 3D fusion derived item may be smaller than the combination of the bounding boxes of the input G-PCC items, which means that the point cloud obtained by combining the input G-PCC items as the 3D fusion derived item is cropped to obtain a point cloud corresponding to the 3D fusion derived item. As another example, the bounding box of the 3D fusion derived item may have a larger extent than the combination of the bounding boxes of the input G-PCC items. In this variant, the 3D fusion derived item may be described by the following structure:
[0191] TIFF0007818144000014.tif5287
[0192] The bounding box of a 3D fusion derived item can be specified by the "output_position" and "output_size" fields. The "output_precision" field specifies the precision used for the "output_position" and "output_size" fields.
[0193] As a further variation, a 3D fusion item may be referred to as a 3D merge and represented as a derived item with an "item_type" value of "mrg3." The GPCCFusion structure may be referred to as a GPCCMerge.
[0194] In some embodiments, one or more input items for a 3D grid item or a 3D fusion item may be a 3D grid item and / or a 3D fusion item.
[0195] In some embodiments, several embodiments of a 3D grid item can be used simultaneously. The various embodiments can be signaled using different values for the "version" field or different values for "item_type".
[0196] In some embodiments, several embodiments of a 3D fusion item can be used simultaneously. The various embodiments can be signaled using different values for the "version" field or different values for "item_type".
[0197] structure The Vector3 structure used in the different structures described above may be the Vector3 structure defined by MPEG-I Part-7 as follows:
[0198] TIFF0007818144000015.tif25122
[0199] The Vector3 structure can also represent x, y, and z coordinates using fixed-point floating-point numbers, such as in the following structure:
[0200] TIFF0007818144000016.tif26123
[0201] Here, the x, y, and z coordinates are represented using fixed-point numbers with "precision_bytes_minus1"+1 bytes for the integer part and "precision_bytes_minus1"+1 bytes for the fractional part.
[0202] Item Properties Several item properties may be associated with a G-PCC item, an identifier point cloud item, a 3D grid derived point cloud item, and / or a 3D fusion derived point cloud item.
[0203] The "BoundingBox" item property specifies the bounding box of a G-PCC item, an identifier point cloud item, a 3D grid derived point cloud item, and / or a 3D fused derived point cloud item. This item property can have the following structure:
[0204] TIFF0007818144000017.tif31109
[0205] In this structure, the "position" field is the reference point of the bounding box and specifies the minimum x, y, and z coordinates of any point contained in the associated item or derived item. The "size" field specifies the extent of the bounding box along the x, y, and z coordinates.
[0206] A bounding box can also be specified using two points, for example two points corresponding to two opposite corners of the bounding box.
[0207] The bounding box can also be specified using the center of the bounding box and its size.
[0208] The "size" item property specifies the size of the G-PCC item, the identifier point cloud item, the 3D grid derived point cloud item, and / or the 3D fused derived point cloud item. This item property can have the following structure:
[0209] TIFF0007818144000018.tif27113
[0210] In this structure, the size field specifies the size of the item along the x, y, and z coordinates.
[0211] The "Translation" transformation item property specifies the transformation to apply to the G-PCC item, the identifier point cloud item, the 3D grid derived point cloud item, and / or the 3D fused derived point cloud item. This transformation item property can have the following structure:
[0212] TIFF0007818144000019.tif24111
[0213] In this structure, the "translation" field specifies the translation vector.
[0214] The "Scaling" transformation item property specifies the scaling to apply to a G-PCC item, an identifier point cloud item, a 3D grid derived point cloud item, or a 3D fused derived point cloud item. This transformation item property can have the following structure:
[0215] TIFF0007818144000020.tif26109
[0216] In this structure, the "scaling" field specifies the scaling to apply to the x, y, and z coordinates. In some cases, a simpler scaling transform item property can be defined, and the same scaling will be applied to all coordinates.
[0217] The "Rotation" transformation item property specifies the rotation to apply to the G-PCC item, the identifier point cloud item, the 3D grid derived point cloud item, and / or the 3D fused derived point cloud item. This transformation item property can have the following structure:
[0218] TIFF0007818144000021.tif26112
[0219] In this structure, the "rotation" field specifies the rotation as a quaternion, which is a unit quaternion whose three imaginary components can be specified and whose real component can be calculated from these three imaginary components.
[0220] An extended "rotation" transformation item property can be defined as follows:
[0221] TIFF0007818144000022.tif30111
[0222] In this structure, the center field specifies the center of rotation.
[0223] The "Transformation" transformation item property specifies a generic 3D transformation that can combine translation, scaling, and / or rotation to apply to a G-PCC item, an identifier point cloud item, a 3D grid derived point cloud item, or a 3D fused derived point cloud item. This transformation item property can have the following structure:
[0224] TIFF0007818144000023.tif35110
[0225] In this structure, the "coeff" field specifies the coefficients of the transformation matrix. The transformation matrix may be a 4x4 uniform matrix where the last row of coefficients is (0,0,0,1). In this case, the "Transformation" structure can specify only 12 coefficients.
[0226] The "Crop" transformation item property specifies the crop to apply to the G-PCC item, the identifier point cloud item, the 3D grid derived point cloud item, and / or the 3D fused derived point cloud item. This transformation item property can have the following structure:
[0227] TIFF0007818144000024.tif30109
[0228] The "position" and "size" fields specify the boundaries of the crop. Any points from items placed outside the boundaries of the crop will not be part of the transformed point cloud.
[0229] A crop may also be specified using two points, for example, two points corresponding to two opposite corners of the crop region.
[0230] The "TransformedBoundingBox" item property specifies the bounding box of a G-PCC item, a discriminated point cloud item, a 3D grid derived point cloud item, and / or a 3D fused derived point cloud item after the transformed item properties have been applied to it.
[0231] The "TransformedBoundingBox" item property allows you to specify the bounding box of the associated item after the application of all transformation item properties listed before this "TransformedBoundingBox" item property, and before the application of any transformation item properties listed after it.
[0232] In fact, calculating the bounding box of a G-PCC item, a discriminated point cloud item, a 3D grid-derived point cloud item, and / or a 3D fused-derived point cloud item after transformation item properties have been applied to it may require calculating the bounding box from all transformed points. For example, after applying a transformation, the bounding box of the transformed point cloud is a transformation of the bounding box of the original point cloud. However, after applying a rotation, the rotated bounding box of the original point cloud is not aligned to the axes of the reference frame. A bounding box constructed from the rotated bounding box contains all points of the rotated point cloud, but may not fit the rotated point cloud exactly. Therefore, indicating the bounding box of the rotated point cloud can aid in decoding and rendering it by accurately indicating the extent of the transformed point cloud.
[0233] This item property can have the following structure:
[0234] TIFF0007818144000025.tif30112
[0235] In this structure, the "position" field is the reference point for the transformed bounding box and specifies the minimum values for the x, y, and z coordinates of a point contained in the associated or derived item after the transformation item properties have been applied to this associated or derived item. The "size" field specifies the extent of the transformed bounding box along the x, y, and z coordinates.
[0236] The transformed bounding box can also be specified using two points, for example two points corresponding to two opposite corners of the bounding box.
[0237] The transformed bounding box can also be specified using the center of the bounding box and its size.
[0238] In some cases, several transformed bounding boxes may be associated with an item or derived item corresponding to the application of different sets of transformation item properties.
[0239] The "TransformedSize" item property specifies the size of a G-PCC item, an identifier point cloud item, a 3D grid derived point cloud item, and / or a 3D fused derived point cloud item after the transformed item properties have been applied to it.
[0240] The "TransformedSize" item property can specify the size of the associated item after the application of all transformation item properties listed before this "TransformedSize" item property and before the application of any transformation item properties listed after it. This item property can have the following structure:
[0241] TIFF0007818144000026.tif25110
[0242] In this structure, the "size" field specifies the extents of the transformed bounding box along the x, y, and z coordinates.
[0243] Variations For purposes of illustration, the above description of various embodiments and variations has focused on the use of G-PCC items. In other embodiments, V3C items defined by MPEG-I Part-10 may be used. In still other embodiments, other items that describe the encoded 3D data may be used. In still other embodiments, a derived item may combine input items of different types. For example, a grid derived item may combine a G-PCC item and a V3C item.
[0244] In various embodiments, the ordering of input items in a "SingleItemTypeReferenceBox" of type "dimg" may be used to select which item takes precedence when several input items define 3D content at the same location. For example, the most recent input item in the ordering in the "SingleItemTypeReferenceBox" may take precedence. For example, if two input G-PCC items define points at the same location, the point corresponding to the most recent of the two input G-PCC items in the "SingleItemTypeReferenceBox" order is kept, and the other point is discarded. As another example, when several input items define 3D content at the same location, these contents may be merged. For example, if two input G-PCC items define points at the same location, these points may be merged by combining their colors by calculating an average color and by taking the most recent timestamp associated with these points. As yet another example, when several input items define 3D content at the same location, these contents may all be kept. For example, if two input G-PCC items define points at the same location, all of these points or some of these points may be kept.
[0245] Note that the different 4cc values used to indicate item types, item property types, or other information are given as examples only; other values may be used to indicate these types.
[0246] In the different item structures described above, some fields are used to signal values coded with one, two, or several bits. Some or all of these values may also be signaled using the "flags" field. Some or all of the options signaled using these fields may also be signaled using the "version" field.
[0247] Examples of hardware for performing steps of the method of the presently disclosed embodiments FIG. 7 is a schematic diagram of an example of a data processing device configured to implement, in whole or in part, some embodiments of the present disclosure.
[0248] The data processing device 700 may be a device such as a microcomputer, a workstation, or a lightly portable device. As shown, the data processing device 700 preferably comprises a communication bus 713 connected thereto. - a central processing unit 711, such as a microprocessor, designated CPU or GPU (for Graphical Processing Unit); - a read-only memory 707, designated ROM, for storing, in whole or in part, a computer program for implementing the present disclosure; a random access memory 712, designated RAM, for storing the executable code of the method according to the embodiment of the present disclosure, as well as registers adapted to record variables and parameters necessary for carrying out the method according to the embodiment of the present disclosure; - at least one communication interface 702 connected to a communication network for transmitting data to and / or receiving data from a remote device;
[0249] Optionally, the data processing device 700 may also include one or more of the following components: - a data storage means 704, such as a hard disk, for storing a computer program for implementing the method according to one or more embodiments of the present disclosure; a disk drive 705 for a disk 706, the disk drive being adapted to read data from the disk 706 or to write data to this disk, and A screen 709 for serving as a graphical interface with the user by means of a keyboard 710 or any other pointing means.
[0250] The data processing device 700 may optionally be connected to various peripheral devices, including sensors 708 such as a digital camera and / or LiDAR, each connected to an input / output card (not shown), to provide data to the data processing device 700.
[0251] Preferably, a communication bus provides communication and interoperability between the various elements included in or connected to the data processing device 700. The representation of a bus is not limiting, and in particular a central processing unit is operable to communicate instructions to any element of the data processing device 700 directly or by means of another element of the data processing device 700.
[0252] The disk 706 may optionally be replaced by any information carrier, such as for example a compact disk (CD-ROM), a rewritable or non-rewritable compact disk, a ZIP disk, a USB key or a memory card, or in general by any information storage means readable by a microcomputer or microprocessor, which may be integrated into the device, possibly removable, and which may be adapted to store one or more programs whose execution can carry out the method according to the invention.
[0253] The executable code may optionally be stored either in the read-only memory 707, on the hard disk 704 or on a removable digital medium as previously mentioned, such as for example the disk 706. According to an optional variant, the executable code of the program may be stored in one of the storage means of the data processing device 700, such as the hard disk 704, and then received by means of a communications network via the interface 702, in order to be executed.
[0254] The central processing unit 711 is preferably adapted to control and direct the execution of instructions or parts of the software code of the program according to the invention, the instructions being stored in one of the storage means mentioned above. On power-up, the program stored in a non-volatile memory, for example the hard disk 704 or the read-only memory 707, is transferred to the random access memory 712, which memory contains the executable code of the program as well as registers for storing variables and parameters necessary to implement the invention.
[0255] In a preferred embodiment, the device is a programmable device that uses software to implement the invention, although the invention may also be implemented in hardware (e.g., in the form of an application specific integrated circuit or ASIC).
[0256] Although the present invention has been described with reference to particular embodiments, the present invention is not limited to those embodiments and modifications will be apparent to those skilled in the art that are within the scope of the present invention.
[0257] Many further modifications and variations will be suggested to those skilled in the art by reference to the above exemplary embodiments, which are given by way of example only and are not intended to limit the scope of the invention, as determined solely by the appended claims. In particular, different features from the various embodiments may be interchanged where appropriate.
[0258] The specific embodiments of the invention described above may be implemented singly or as a combination of elements from multiple embodiments, and features from various embodiments may be combined where appropriate or beneficial in a single embodiment.
[0259] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
Claims
1. A method for generating an ISOBMFF-based media file containing three-dimensional data, comprising: acquiring a plurality of subsets of the three-dimensional data; generating a plurality of input items, each of which describes said subset of said three-dimensional data; generating a derived item that includes descriptive data of the spatial configuration of the plurality of input items, the derived item being related to the plurality of input items by item references; generating the media file including the plurality of input items, the item references, and the derived items; How to prepare for this.
2. The method of claim 1 , wherein the descriptive data further comprises an indicator of whether the three-dimensional data of each subset of three-dimensional data uses the same frame of reference.
3. The method of claim 1 , wherein the description data further comprises an indicator of whether three-dimensional bounding boxes corresponding to the three-dimensional data described by each of at least two input items of the plurality of input items overlap.
4. The method of claim 1 , wherein the descriptive data further comprises an indicator for describing a three-dimensional array of cells, each input item of the plurality of input items corresponding to one of the cells.
5. The method of claim 4 , wherein the descriptive data further comprises an indicator of the order of the plurality of input items in the three-dimensional array of cells.
6. The method of claim 4 , wherein the descriptive data further comprises an indicator for describing a size of at least one of the cells.
7. The method of claim 4 , wherein the size of the cell is determined as a function of a three-dimensional bounding box corresponding to the three-dimensional data described by each of the plurality of input items.
8. The method described in claim 4, wherein the three-dimensional data is point cloud data, and the description data further includes an indicator indicating that at least one of the cells does not contain a point defined in the point cloud data.
9. The method of claim 1, wherein the three-dimensional data is point cloud data, and at least one item property is associated with at least one of the input items and / or derived items to describe at least one spatial operation to be performed on points defined within the point cloud data.
10. The method of claim 1, wherein the three-dimensional data is point cloud data, and the derived item includes at least one indicator for describing at least one spatial operation to be performed on points defined in the point cloud data.
11. The method of claim 1 , further comprising acquiring the three-dimensional data and dividing the acquired three-dimensional data into the plurality of subsets of the three-dimensional data.
12. A method for analyzing an ISOBMFF-based media file containing three-dimensional data, comprising: obtaining, from the media file, a derived item that includes descriptive data of a spatial configuration of a plurality of input items, the derived item being associated with the plurality of input items by item references; obtaining input items from the media file, each input item describing a subset of three-dimensional data of a plurality of subsets, the input items as a function of the item references; obtaining the subsets of the three-dimensional data from the media file as a function of the input items; generating the three-dimensional data as a function of the descriptive data, the three-dimensional data comprising the three-dimensional data of the plurality of subsets.
13. The method of claim 12 , further comprising determining whether the three-dimensional data of each subset of three-dimensional data uses the same frame of reference as a function of the indicator in the descriptive data.
14. 13. The method of claim 12, further comprising determining whether three-dimensional bounding boxes corresponding to the three-dimensional data described by each of at least two input items of the plurality of input items overlap as a function of an indicator of the description data.
15. The method of claim 12 , further comprising determining a three-dimensional array of cells as a function of the indicator of the descriptive data, each input item of the plurality of input items corresponding to one of the cells.
16. 16. The method of claim 15, further comprising determining an order of the plurality of input items within the three-dimensional array of cells as a function of the indicator of the descriptive data.
17. The method of claim 15 , further comprising determining a size of at least one of the cells as a function of an indicator of the descriptive data.
18. The method of claim 15, wherein the three-dimensional data is point cloud data, and further comprising determining, as a function of the indicators of the descriptive data, that at least one of the cells does not contain any points defined in the point cloud data.
19. The method of claim 12, wherein the three-dimensional data is point cloud data, and further comprising applying a spatial operation to points defined within the point cloud data as a function of at least one item property associated with at least one of the input items and / or the derived items.
20. The method of claim 12, wherein the three-dimensional data is point cloud data, and further comprising applying at least one spatial operation to be performed on points defined within the point cloud data as a function of at least one indicator of the derived item.
21. A computer program for a programmable device, comprising instructions for carrying out the steps of the method according to any one of claims 1 to 20, when the program is loaded and executed by the programmable device.
22. A non-transitory computer-readable storage medium storing the computer program of claim 21.
23. A processing device configured to perform the steps of the method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method
JP2023509086A
Method and apparatus for encapsulating image data in a file for progressive rendering
JP2023552029A
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
US20210209807A1
Three-dimensional data coding method, three-dimensional data decoding method, three-dimensional data coding device, and three-dimensional data decoding device
WO2020071416A1
Method and apparatus for encapsulating image data in a file for progressive rendering
WO2022129235A1