Point cloud data transmission device and method performed by this transmission device, and point cloud data reception device and method performed by this reception device

The method addresses latency and encoding/decoding complexity in point cloud data processing by grouping samples based on temporal levels, enabling efficient access and manipulation of point cloud data for high-quality VR and autonomous driving services.

JP7746424B2Active Publication Date: 2025-09-30LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023580651
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-18
Filing Date
2022-06-30
Publication Date
2025-09-30
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing large amounts of point cloud data for applications like VR, AR, and autonomous driving due to latency and encoding/decoding complexity, and lack support for temporal scalability in carrying geometry-based point cloud compressed data.

Method used

A method and apparatus for processing point cloud data that includes grouping samples based on temporal levels, generating a G-PCC file with sample group and temporal level information, and enabling efficient access and manipulation of desired components consistent with network capabilities.

Benefits of technology

The solution allows for high-quality point cloud services with reduced latency and encoding/decoding complexity, supporting temporal scalability and efficient access to G-PCC components, while reducing bits for signaling frame rate information and increasing bit efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746424000014
    Figure 0007746424000014
  • Figure 0007746424000015
    Figure 0007746424000015
  • Figure 0007746424000016
    Figure 0007746424000016
Patent Text Reader

Abstract

A transmitting device for point cloud data, a method performed in the transmitting device, a receiving device, and a method performed in the receiving device are provided. The method performed in the receiving device for point cloud data according to the present disclosure may include the steps of obtaining a geometry-based point cloud compression (G-PCC) file including the point cloud data, and extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on information on sample groups in which samples in the G-PCC file are grouped based on one or more temporal levels and information on the temporal levels.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method and apparatus for processing point cloud content. [Background technology]

[0002] Point cloud content is content expressed as a point cloud, which is a collection of points belonging to a coordinate system that represents a three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide a variety of services such as VR (virtual reality), AR (augmented reality), MR (mixed reality), and autonomous driving services. Since tens of thousands to hundreds of thousands of point data are required to represent point cloud content, a method for efficiently processing a huge amount of point data is required. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure provides an apparatus and method for efficiently processing point cloud data. The present disclosure provides a method and apparatus for processing point cloud data to address latency and encoding / decoding complexity.

[0004] The present disclosure also provides an apparatus and method for supporting temporal scalability in the carriage of geometry-based point cloud compressed data.

[0005] The present disclosure also provides an apparatus and method for providing a point cloud content service that efficiently stores a G-PCC bitstream in a single track within a file or splits it into multiple tracks and provides signaling for this.

[0006] The present disclosure also proposes an apparatus and method for processing file storage techniques to support efficient access to stored G-PCC bitstreams.

[0007] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]

[0008] A method performed in a point cloud data receiving device according to one embodiment of the present disclosure may include (comprise; configure; construct; set; encompass; contain; contain) a step of obtaining a G-PCC (Geometry-based point cloud compression) file containing the point cloud data - the G-PCC file including information on sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels, and information on the temporal levels - and a step of extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on the information on the sample groups and the information on the temporal levels.

[0009] A point cloud data receiving device according to another embodiment of the present disclosure includes a memory and at least one processor, wherein the at least one processor acquires a G-PCC (geometry-based point cloud compression) file containing the point cloud data - the G-PCC file includes information on sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels, and information on the temporal levels - and can extract one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on the information on the sample groups and the information on the temporal levels.

[0010] A method performed in a point cloud data transmission device according to another embodiment of the present disclosure may include generating information for sample groups in which G-PCC (geometry-based point cloud compression) samples are grouped based on one or more temporal levels, and generating a G-PCC file including information for the sample groups, information for the temporal levels, and the point cloud data.

[0011] According to another embodiment of the present disclosure, a point cloud data transmission device includes a memory and at least one processor, wherein the at least one processor generates information for sample groups in which G-PCC (geometry-based point cloud compression) samples are grouped based on one or more temporal levels, and generates a G-PCC file including the information for the sample groups, the information for the temporal levels, and the point cloud data. [Effects of the Invention]

[0012] The apparatus and method according to the embodiments of the present disclosure can process point cloud data with high efficiency.

[0013] The apparatus and method according to the embodiments of the present disclosure can provide high-quality point cloud services.

[0014] The apparatus and method according to the embodiments of the present disclosure can provide point cloud content for providing general-purpose services such as VR services and autonomous driving services.

[0015] The apparatus and method according to the embodiments of the present disclosure can provide time scalability that allows for efficient access to desired components of the G-PCC components.

[0016] The apparatus and method according to the embodiments of the present disclosure can improve the performance of point cloud content provision systems by supporting temporal scalability, allowing data to be manipulated at a high level consistent with network capabilities, decoder capabilities, etc.

[0017] The apparatus and method according to the embodiments of the present disclosure can reduce bits for signaling frame rate information and increase bit efficiency.

[0018] Apparatus and methods according to embodiments of the present disclosure can enable smooth and gradual replay by reducing the increase in replay complexity. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a block diagram illustrating an example of a point cloud content providing system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an example of a point cloud content providing process according to an embodiment of the present disclosure. [Figure 3] 1 illustrates an example of a point cloud video acquisition process according to an embodiment of the present disclosure. [Figure 4] 1 illustrates an example of a point cloud encoding device according to an embodiment of the present disclosure. [Figure 5] 1 illustrates an example of a voxel according to an embodiment of the present disclosure. [Figure 6] 1 illustrates an example of an octree and occupancy code according to an embodiment of the present disclosure. [Figure 7] 1 illustrates an example of a neighboring node pattern according to an embodiment of the present disclosure. [Figure 8] 10 illustrates an example of arranging points according to LOD distance values ​​according to an embodiment of the present disclosure. [Figure 9] 10 illustrates an example of a point configuration by LOD according to an embodiment of the present disclosure. [Figure 10] FIG. 1 is a block diagram illustrating an example of a point cloud decoding device according to an embodiment of the present disclosure. [Figure 11] FIG. 10 is a block diagram illustrating another example of a point cloud decoding device according to an embodiment of the present disclosure. [Figure 12] FIG. 10 is a block diagram illustrating another example of a transmission device according to an embodiment of the present disclosure. [Figure 13] FIG. 10 is a block diagram illustrating another example of a receiving device according to an embodiment of the present disclosure. [Figure 14] 1 illustrates an example of a structure that can be coupled with a point cloud data transmission / reception method / apparatus according to an embodiment of the present disclosure. [Figure 15] FIG. 10 is a block diagram illustrating another example of a transmission device according to an embodiment of the present disclosure. [Figure 16] 1 illustrates an example of spatial division of a bounding box into three-dimensional blocks according to an embodiment of the present disclosure. [Figure 17] FIG. 10 is a block diagram illustrating another example of a receiving device according to an embodiment of the present disclosure. [Figure 18] 1 illustrates an example of a bitstream structure according to an embodiment of the present disclosure. [Figure 19] 10 illustrates an example for identifying relationships between structures within a bitstream according to an embodiment of the present disclosure. [Figure 20] 1 illustrates reference relationships between structures within a bitstream according to an embodiment of the present disclosure. [Figure 21]1 illustrates an example of an SPS syntax structure according to an embodiment of the present disclosure. [Figure 22] 10 illustrates an example of an indication of an attribute type and an indication of a correspondence relationship between position components according to an embodiment of the present disclosure. [Figure 23] 1 illustrates an example for a GPS syntax structure according to an embodiment of the present disclosure. [Figure 24] 1 illustrates an example of an APS syntax structure according to an embodiment of the present disclosure. [Figure 25] 10 illustrates an example of an attribute coding type table according to an embodiment of the present disclosure. [Figure 26] 10 illustrates an example for a tile inventory syntax structure according to an embodiment of the present disclosure. [Figure 27-28] 1 illustrates an example for a geometry slice syntax structure according to an embodiment of the present disclosure. [Figure 29-30] 10 illustrates an example of an attribute slice syntax structure according to an embodiment of the present disclosure. [Figure 31] 10 illustrates an example of a metadata slice syntax structure according to an embodiment of the present disclosure. [Figure 32] 1 illustrates an example of a TLV encapsulation structure according to an embodiment of the present disclosure. [Figure 33] 1 illustrates an example of a TLV encapsulation syntax structure and payload type according to an embodiment of the present disclosure. [Figure 34] 1 illustrates an example for a file containing a single track according to an embodiment of the present disclosure. [Figure 35] 1 shows an example for a file containing multiple tracks according to an embodiment of the present disclosure. [Figure 36-37] 1 is a flow chart for an embodiment that supports temporal scalability. [Figure 38-39] 10 is a flowchart for an example of determining whether to signal and obtain frame rate information. [Figure 40-41]10 is a flow chart for an embodiment that can prevent duplicate signaling problems. DETAILED DESCRIPTION OF THE INVENTION

[0020] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0021] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.

[0022] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.

[0023] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.

[0024] In this disclosure, components that are distinguished from one another are used to clearly describe the respective features and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0025] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.

[0026] The present disclosure relates to encoding and decoding of point cloud-related data, and terms used in the present disclosure may have their ordinary meanings in the technical field to which the present disclosure pertains, unless they are newly defined in the present disclosure.

[0027] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0028] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0029] The present disclosure relates to compression of point cloud-related data. Various methods or embodiments of the present disclosure may be applied to the MPEG (Moving Picture Experts Group) PCC (point cloud compression or point cloud coding) standard (e.g., G-PCC or V-PCC standard) or next-generation video / image coding standards.

[0030] In this disclosure, a "point cloud" may refer to a collection of points located in three-dimensional space. Also, in this disclosure, a "point cloud content" may refer to a "point cloud video / image" that is content represented by a point cloud. Hereinafter, a "point cloud video / image" will be referred to as a "point cloud video." A point cloud video may include one or more frames, and a frame may be a still image or a picture. Therefore, a point cloud video may include point cloud images / frames / pictures and may be referred to as any one of a "point cloud image," a "point cloud frame," and a "point cloud picture."

[0031] In this disclosure, "point cloud data" may refer to data or information associated with each point in a point cloud. Point cloud data may include geometry and / or attributes. Point cloud data may also include meta data. Point cloud data may also be referred to as "point cloud content data" or "point cloud video data." Point cloud data may also be referred to as "point cloud content," "point cloud video," "G-PCC data," etc.

[0032] In the present disclosure, a point cloud object corresponding to point cloud data may be represented in the form of a box based on a coordinate system, and this box based on the coordinate system may be referred to as a bounding box. That is, the bounding box may be a rectangular cuboid that can contain all of the points of the point cloud, or may be a rectangular cuboid that contains the source point cloud frame.

[0033] In this disclosure, geometry includes the position (or position information) of each point, and this position can be expressed by parameters (e.g., x-axis value, y-axis value, and z-axis value) that represent a three-dimensional coordinate system (e.g., a coordinate system consisting of x-axis, y-axis, and z-axis). Geometry is sometimes referred to as "geometry information."

[0034] In the present disclosure, the attributes may include attributes of each point, and the attributes may include one or more of texture information, hue (RGB or YCbCr), reflectance (r), transparency, etc. of each point. The attributes may be referred to as "attribute information." The metadata may include various data related to acquisition in the acquisition process described below.

[0035] Overview of the point cloud content provision system

[0036] 1 illustrates an example of a system for providing point cloud content (hereinafter referred to as a "point cloud content providing system") according to an embodiment of the present disclosure. FIG. 2 illustrates an example of a process in which the point cloud content providing system provides point cloud content.

[0037] As shown in Fig. 1, the point cloud content providing system may include a transmission device 10 and a reception device 20. The point cloud content providing system may perform an acquisition process (S20), an encoding process (S21), a transmission process (S22), a decoding process (S23), a rendering process (S24), and / or a feedback process (S25) shown in Fig. 2 through the operations of the transmission device 10 and the reception device 20.

[0038] To provide point cloud content, the transmitting device 10 may acquire point cloud data and output a bitstream through a series of processes (e.g., encoding process) on the acquired point cloud data (original point cloud data). Here, the point cloud data may be output in a bitstream format after the encoding process. According to an embodiment, the transmitting device 10 may transmit the output bitstream to the receiving device 20 in a file or streaming (streaming segment) format via a digital storage medium or a network. The digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The receiving device 20 may process (e.g., decode or restore) the received data (e.g., encoded point cloud data) back into the original point cloud data and render it. Through these processes, point cloud content can be provided to a user, and the present disclosure may provide various embodiments necessary to effectively perform these processes.

[0039] As shown in FIG. 1, the transmission device 10 may include an acquisition unit 11, an encoding unit 12, an encapsulation processing unit 13, and a transmission unit 14, and the receiving device 20 may include a receiving unit 21, a decapsulation processing unit 22, a decoding unit 23, and a rendering unit 24.

[0040] The acquisition unit 11 may perform a step (S20) of acquiring a point cloud video through a capture, synthesis, or generation process, etc. Therefore, the acquisition unit 11 may also be referred to as a "point cloud video acquisition unit."

[0041] The acquisition process (S20) can generate point cloud data (geometry and / or attributes, etc.) for a large number of points. The acquisition process (S20) can also generate metadata related to the acquisition of the point cloud video. The acquisition process (S20) can also generate mesh data (e.g., triangular shape data) that indicates connectivity information between point clouds.

[0042] The metadata can include initial viewing orientation metadata, which can indicate whether the point cloud data is front-facing or rear-facing. The metadata is sometimes referred to as "auxiliary data," which is metadata for the point cloud.

[0043] The captured point cloud video may contain PLY (polygon file format or the Stanford triangle format) files. Because a point cloud video has one or more frames, the captured point cloud video may contain one or more PLY files. A PLY file contains point cloud data for each point.

[0044] To acquire point cloud video (or point cloud data), the acquisition unit 11 may be configured as a combination of a camera device capable of acquiring depth (depth information) and an RGB camera capable of extracting color information corresponding to the depth information. Here, the camera device capable of acquiring depth information may be a combination of an infrared pattern projector and an infrared camera. The acquisition unit 11 may also be configured as a LiDAR, but a radar system that measures the position coordinates of a reflector by emitting a LiDAR laser pulse and measuring the time it takes for the laser pulse to be reflected and returned can also be used.

[0045] The acquisition unit 110 can extract the form of geometry consisting of points in a three-dimensional space from the depth information, and extract attributes that express the hue, reflection, etc. of each point from the RGB information.

[0046] Methods for extracting (or capturing, acquiring, etc.) point cloud video (or point cloud data) include an inward-facing method for capturing a central object and an outward-facing method for capturing the external environment. Examples of the inward-facing method and the outward-facing method are shown in FIG. 3. (a) of FIG. 3 is an example of the inward-facing method, and (b) of FIG. 3 is an example of the outward-facing method.

[0047] As shown in Figure 3(a), the inward-facing method can be used to create point cloud content of the current surroundings of a vehicle, such as in autonomous driving. As shown in Figure 3(b), the outward-facing method can be used to create point cloud content that allows users to freely view key objects such as characters, players, objects, and actors in a 360-degree VR / AR environment. When creating point cloud content using multiple cameras, a camera calibration process may be performed before capturing content to set a global coordinate system between the cameras. A method of synthesizing an arbitrary point cloud video based on the captured point cloud video may also be used.

[0048] On the other hand, when providing a point cloud video of a computer-generated virtual space, capture via an actual camera may not be performed. In this case, post-processing may be required to improve the quality of the captured point cloud content. For example, in the acquisition step (S20), the maximum / minimum depth values ​​may be adjusted within the range provided by the camera equipment, and post-processing may be performed to remove unwanted areas (e.g., background) or point data of unwanted areas, or to fill spatial holes by recognizing connected spaces. As another example, post-processing may be performed to integrate point cloud data extracted from cameras sharing a spatial coordinate system into a single content through a conversion process to a global coordinate system for each point based on the position coordinates of each camera. As a result, a single point cloud content covering a wide range may be generated, or point cloud content with a high point density may be obtained.

[0049] The encoder 12 may perform an encoding process (S21) of encoding data (such as geometry, attributes, and / or metadata and / or mesh data) generated from the acquirer 11 into one or more bitstreams. Therefore, the encoder 12 may also be referred to as a "point cloud video encoder." The encoder 12 may encode the data generated from the acquirer 11 serially or in parallel.

[0050] The encoding process (S21) performed by the encoder 12 may be geometry-based point cloud compression (G-PCC). The encoder 12 may perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency.

[0051] The encoded point cloud data can be output in the form of a bitstream. Based on the G-PCC procedure, the encoding unit 12 can encode the point cloud data by separating it into geometry and attributes, as described below. In this case, the output bitstream can include a geometry bitstream containing the encoded geometry and an attribute bitstream containing the encoded attributes. The output bitstream can also include one or more of a metadata bitstream containing metadata, an auxiliary bitstream containing auxiliary data, and a mesh data bitstream containing mesh data. The encoding step (S21) will be described in more detail below. The bitstream containing the encoded point cloud data may also be called a "point cloud bitstream" or a "point cloud video bitstream."

[0052] The encapsulation processor 13 may perform a process of encapsulating one or more bitstreams output from the decoder 12 into the form of a file, a segment, or the like. Therefore, the encapsulation processor 13 may also be referred to as a "file / segment encapsulation module." Although the drawings illustrate an example in which the encapsulation processor 13 is a separate component / module in relation to the transmitter 14, the encapsulation processor 13 may also be included in the transmitter 14 depending on the embodiment.

[0053] The encapsulation processor 13 may encapsulate the data in a file format such as ISO Base Media File Format (ISOBMFF) or process it in other forms such as DASH segments. Depending on the embodiment, the encapsulation processor 13 may include metadata in the file format. For example, the metadata may be included in boxes at various levels in the ISOBMFF file format or may be included as data in a separate track within the file. Depending on the embodiment, the encapsulation processor 130 may encapsulate the metadata itself in the file. The metadata processed by the encapsulation processor 13 may be transmitted from a metadata processor or the like (not shown in the drawings). The metadata processor may be included in the encoding unit 12 or may be configured as a separate component / module.

[0054] The transmitting unit 14 may perform a transmitting step (S22) of processing the encapsulated point cloud bitstream according to a file format (processing for transmission). The transmitting unit 14 may transmit the bitstream or a file / segment including the bitstream to the receiving unit 21 of the receiving device 20 via a digital storage medium or a network. Therefore, the transmitting unit 14 may also be referred to as a "transmitter" or a "communication module."

[0055] The transmitter 14 may process the point cloud data according to any transmission protocol. Here, "processing the point cloud data according to any transmission protocol" may be "processing for transmission." The processing for transmission may include processing for transmission via a broadcast network or processing for transmission via broadband. Depending on the embodiment, the transmitter 14 may receive not only the point cloud data but also metadata from a metadata processing unit and perform processing for transmission on the transmitted metadata. Depending on the embodiment, the processing for transmission may be performed in a transmission processing unit, which may be included in the transmitter 14 or configured as a component / module separate from the transmitter 14.

[0056] The receiving unit 21 can receive the bitstream or a file / segment containing the bitstream transmitted by the transmission device 10. Depending on the transmission channel, the receiving unit 21 can receive the bitstream or a file / segment containing the bitstream via a broadcast network, or via broadband. Alternatively, the receiving unit 21 can receive the bitstream or a file / segment containing the bitstream via a digital storage medium.

[0057] The receiver 21 may perform processing according to a transmission protocol on the received bitstream or a file / segment including the bitstream. The receiver 21 may perform a reverse process of the transmission process (processing for transmission) corresponding to the processing for transmission performed by the transmitter 10. The receiver 21 may transfer encoded point cloud data from the received data to the decapsulation processor 22 and transfer metadata to the metadata parsing unit. The metadata may be in the form of a signaling table. Depending on the embodiment, the reverse process of the processing for transmission may be performed in the reception processor. The reception processor, the decapsulation processor 22, and the metadata parsing unit may each be included in the receiver 21 or configured as a component / module separate from the receiver 21.

[0058] The decapsulation processing unit 22 can decapsulate the point cloud data in a file format (i.e., a bitstream in a file format) transmitted from the receiving unit 21 or the receiving processing unit. Therefore, the decapsulation processing unit 22 is also called a "file / segment encapsulation module."

[0059] The decapsulation processing unit 22 may obtain a point cloud bitstream or a metadata bitstream by decapsulating a file using ISOBMFF or the like. Depending on the embodiment, metadata (metadata bitstream) may be included in the point cloud bitstream. The obtained point cloud bitstream may be transmitted to the decoding unit 23, and the obtained metadata bitstream may be transmitted to the metadata processing unit. The metadata processing unit may be included in the decoding unit 23 or may be configured as a separate component / module. The metadata obtained by the decapsulation processing unit 23 may be in the form of a box or track in a file format. If necessary, the decapsulation processing unit 23 may receive metadata required for decapsulation from the metadata processing unit. The metadata may be transmitted to the decoding unit 23 and used in the decoding process (S23), or may be transmitted to the rendering unit 24 and used in the rendering process (S24).

[0060] The decoding unit 23 receives the bitstream and performs a decoding process (S23) of decoding the point cloud bitstream (encoded point cloud data) by performing an operation corresponding to the operation of the encoding unit 12. Therefore, the decoding unit 23 is also called a "point cloud video decoder."

[0061] The decoding unit 23 may separate the point cloud data into geometry and attributes and decode them. For example, the decoding unit 23 may restore (decode) geometry from a geometry bitstream included in the point cloud bitstream, and may restore (decode) attributes based on an attribute bitstream included in the point cloud bitstream and the restored geometry. A 3D point cloud video / image may be restored based on position information according to the restored geometry and attributes (such as color or texture) according to the decoded attributes. The decoding process (S23) will be described in more detail below.

[0062] The rendering unit 24 may perform a rendering process (S24) of rendering the restored point cloud video. Therefore, the rendering unit 24 may also be called a "renderer."

[0063] The rendering process (S24) may refer to a process of rendering and displaying point cloud content in a 3D space, and may be performed in a desired rendering manner based on position information and attribute information of points decoded through the decoding process.

[0064] The points of the point cloud content may be rendered as a vertex having a certain thickness, a cube having a specific minimum size centered at the vertex position, or a circle centered at the vertex position. A user may view all or a portion of the rendered result through a VR / AR display, a general display, or the like. The rendered video may be displayed through a display unit. A user may view all or a portion of the rendered result through a VR / AR display, a general display, or the like.

[0065] The feedback process (S25) may include a process of transmitting various feedback information obtained in the rendering process (S24) or the display process to the transmitting device 10 or to other components in the receiving device 20. The feedback process (S25) may be performed by one or more of the components included in the receiving device 20 of Fig. 1, or by one or more of the components depicted in Fig. 10 and Fig. 11. Depending on the embodiment, the feedback process (S25) may be performed by a 'feedback unit' or a 'sensing / tracking unit'.

[0066] Interactivity for point cloud content consumption can be provided through the feedback process (S25). Depending on the embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc. can be fed back in the feedback process (S25). Depending on the embodiment, the user can interact with what is realized in the VR / AR / MR / autonomous driving environment, and in this case, information related to the interaction can be transmitted to the transmission device 10 or the service provider side in the feedback process (S25). Depending on the embodiment, the feedback process (S25) may not be performed.

[0067] Head orientation information may refer to information regarding the position, angle, movement, etc. of the user's head. Based on this information, information regarding the area the user is currently viewing in the point cloud video, i.e., viewport information, may be calculated.

[0068] The viewport information may be information about the area the user is currently viewing in the point cloud video. The viewpoint refers to the location where the user is viewing the point cloud video, and may refer to the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape of the area can be determined by the field of view (FOV). Gaze analysis using the viewport information can determine how the user consumes the point cloud video and how much they gaze at which area of ​​the point cloud video. Gaze analysis can be performed on the receiving side (receiving device) and transmitted to the transmitting side (transmitting device) via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the position / direction of the user's head, the vertical or horizontal FOV supported by the device, etc.

[0069] According to an embodiment, the feedback information may be not only transmitted to the transmitting side (transmitting device) but also consumed by the receiving side (receiving device). That is, the feedback information may be used to perform a decoding process, a rendering process, etc. at the receiving side (receiving device).

[0070] For example, the receiving device 20 may use head orientation information and / or viewport information to preferentially decode and render only the point cloud video for the area the user is currently viewing. The receiving unit 21 may receive all point cloud data, or may receive point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The decapsulation processing unit 22 may decapsulate all point cloud data, or may decapsulate point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The decoding unit 23 may decode all point cloud data, or may decode point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information.

[0071] Point cloud encoding device overview

[0072] 4 illustrates an example of a point cloud encoding device 400 according to an embodiment of the present disclosure. The point cloud encoding device 400 of FIG. 4 may correspond in configuration and function to the encoding unit 12 of FIG.

[0073] As shown in FIG. 4, the point cloud encoding device 400 may include a coordinate system conversion unit 405, a geometry quantization unit 410, an octree analysis unit 415, an approximation unit 420, a geometry encoding unit 425, a restoration unit 430, an attribute conversion unit 440, a RAHT conversion unit 445, an LOD generation unit 450, a lift unit 455, an attribute quantization unit 460, an attribute encoding unit 465, and / or a color conversion unit 435.

[0074] The point cloud data acquired by the acquisition unit 11 may undergo a process for adjusting the quality of the point cloud content (e.g., lossless, lossy, near-lossless) depending on the network conditions or the application. While each point of the acquired point cloud content may be transmitted without loss, real-time streaming may not be possible in such a case due to the large size of the point cloud content. Therefore, in order to smoothly provide the point cloud content, a process for reconstructing the point cloud content in accordance with the maximum target bitrate is required.

[0075] The process for adjusting the quality of point cloud content may be a process of reconstructing and encoding point position information (position information included in geometry information) or color information (color information included in attribute information), etc. The process of reconstructing and encoding point position information may be referred to as geometry coding, and the process of reconstructing and encoding attribute information related to each point may be referred to as attribute coding.

[0076] The geometry coding may include a geometry quantization process, a voxelization process, an octree analysis process, an approximation process, a geometry encoding process, and / or a coordinate system transformation process. The geometry coding may further include a geometry restoration process. The attribute coding may include a color transformation process, an attribute transformation process, a prediction transformation process, a lift transformation process, a RAHT transformation process, an attribute quantization process, an attribute encoding process, etc.

[0077] Geometry Coding

[0078] The coordinate system transformation process may correspond to a process of transforming a coordinate system for a point position. Therefore, the coordinate system transformation process may be referred to as "transform coordinates." The coordinate system transformation process may be performed by the coordinate system transformation unit 405. For example, the coordinate system transformation unit 405 may transform the point position from a global space coordinate system into position information in a three-dimensional space (e.g., a three-dimensional space expressed by an X-axis, Y-axis, and Z-axis coordinate system). The position information in the three-dimensional space according to the embodiment may be referred to as "geometry information."

[0079] The geometry quantization process may correspond to a process of quantizing position information of points and may be performed by the geometry quantization unit 410. For example, the geometry quantization unit 410 may search for position information having the smallest (x, y, z) values ​​among the position information of points and subtract the position information having the smallest (x, y, z) values ​​from the position information of each point. In addition, the geometry quantization unit 410 may perform the quantization process by multiplying the subtracted value by a preset quantization scale value and adjusting (raising or lowering) the result to a nearby integer value.

[0080] The voxelization process may correspond to a process of matching geometry information quantized through the quantization process to specific voxels existing in a three-dimensional space. The voxelization process may also be performed by the geometry quantization unit 410. The geometry quantization unit 410 may perform octree-based voxelization based on position information of points to reconstruct each point to which the quantization process has been applied.

[0081] An example of a voxel according to an embodiment of the present disclosure is shown in FIG. 5. A voxel may refer to a space for storing information about a point existing in three dimensions, similar to a pixel, which is the smallest unit having information about a two-dimensional image / video. A voxel is a portmanteau of the words "volume" and "pixel." As shown in FIG. 5, a voxel may refer to a three-dimensional cubic space generated by dividing a three-dimensional space (2 depth, 2 depth, 2 depth) into units (unit=1.0) based on each axis (x axis, y axis, and z axis). A voxel can estimate spatial coordinates based on its positional relationship with a voxel group and, like a pixel, may have color or reflectance information.

[0082] Only one point may exist (match) in one voxel. That is, information related to multiple points may exist in one voxel. Alternatively, information related to multiple points included in one voxel may be integrated into one point information. This adjustment can be performed selectively. When integrating and expressing one point information in one voxel, the position value of the center point of the voxel may be set based on the position value of the point existing in the voxel, and an attribute conversion process related to this may be performed. For example, the attribute conversion process may adjust the position value of the point included in the voxel or the center point of the voxel to the average value of the hue or reflectance of adjacent points within a specific radius.

[0083] The octree analyzer 415 can use an octree to efficiently manage the regions / positions of voxels. An example of an octree according to an embodiment of the present disclosure is shown in FIG. 6(a). To efficiently manage the space of a two-dimensional image, the entire space is divided based on the x-axis and y-axis, resulting in four spaces. Each of the four spaces is further divided based on the x-axis and y-axis, resulting in four spaces for each smaller space. The regions are divided until leaf nodes become pixels, and a quadtree can be used as a data structure to efficiently manage the regions by size and position.

[0084] Similarly, the present disclosure can apply the same method to efficiently manage a three-dimensional space by its position and size. However, since the Z-axis is added as shown in the middle of Figure 6(a), dividing the three-dimensional space based on the x-axis, y-axis, and z-axis can result in eight spaces. Furthermore, as shown on the right side of Figure 6(a), dividing each of the eight spaces again based on the x-axis, y-axis, and z-axis can result in eight further spaces for each smaller space.

[0085] The octree analysis unit 415 divides regions until leaf nodes become voxels, and can use an octree data structure that can manage eight child node regions to efficiently manage them according to region size and position.

[0086] To manage voxels that reflect the position of a point using an octree, the entire volume of the octree must be set to (0, 0, 0) to (2d, 2d, 2d). 2d is set to the value that forms the smallest bounding box that encloses all points in the point cloud, and d is the depth of the octree. The formula for calculating the d value is as shown in Equation 1 below, where d is the position value of the point to which the quantization process has been applied.

[0087]

number

[0088] An occupancy code can be expressed as an occupancy code, and an example of an occupancy code according to an embodiment of the present disclosure is shown in Fig. 6(b). The occupancy code of each node can be expressed as 1 if the node contains a point, and as 0 if the node does not contain a point.

[0089] Each node may have an 8-bit bitmap indicating whether it is occupied by its eight child nodes. For example, since the occupancy code of the node corresponding to the second depth (1-depth) in FIG. 6(b) is 00100001, the spaces (voxels or regions) corresponding to the third and eighth nodes can contain at least one point. Also, since the occupancy code of the child node (leaf node) of the third node is 10000111, the spaces corresponding to the first, sixth, seventh, and eighth leaf nodes of the leaf node can contain at least one point. Also, since the occupancy code of the child node (leaf node) of the eighth node is 01001111, the spaces corresponding to the second, fifth, sixth, seventh, and eighth leaf nodes of the leaf node can contain at least one point.

[0090] The geometry encoding process may correspond to a process of performing entropy coding on the dedicated code. The geometry encoding process may be performed by the geometry encoding unit 425. The geometry encoding unit 425 may perform entropy coding on the dedicated code. The generated dedicated code may be encoded directly or may be encoded through intra- and inter-coding processes to improve compression efficiency. The receiving device 20 may reconstruct an octree through the dedicated code.

[0091] On the other hand, in the case of a specific region with no or very few points, voxelizing the entire region may be inefficient. That is, since there are few points in the specific region, it may not be necessary to construct an entire octree. In such cases, an early termination method may be necessary.

[0092] For a specific region (a specific region that does not correspond to a leaf node), instead of dividing the node (specific node) corresponding to this specific region into eight subnodes (child nodes), the point cloud encoding device 400 can transmit the position of points directly only for that specific region, or can use a surface model to reconstruct the position of points within the specific region based on voxels.

[0093] A mode in which the position of each point is directly transmitted to a specific node may be called a direct mode. The point cloud encoding device 400 may check whether the conditions for enabling the direct mode are met.

[0094] Conditions for enabling direct mode include: 1) the direct mode use option must be activated; 2) the specific node is not a leaf node; 3) there must be points within the specific node that are less than a threshold; and 4) the total number of points to be transmitted directly must not exceed the threshold.

[0095] If all of these conditions are met, the point cloud encoding device 400 can entropy-code and transmit the position value of the direct point for the specific node via the geometry encoding unit 425.

[0096] The mode of reconstructing the positions of points in a specific region based on voxels using a surface model may be a trisoup mode. The trisoup mode may be performed by the approximation unit 420. The approximation unit 420 determines a specific level of the octree, and from the determined specific level, the positions of points in a node region can be reconstructed based on voxels using a surface model.

[0097] The point cloud encoding device 400 can also selectively apply the trisoup mode. Specifically, when using the trisoup mode, the point cloud encoding device 400 can specify the level (specific level) at which the trisoup mode is applied. For example, if the specified specific level is the same as the depth (d) of the octree, the trisoup mode may not be applied. In other words, the specified specific level must be smaller than the depth value of the octree.

[0098] A 3D cubic region of a node at a specified level is called a block, and one block can include one or more voxels. A block or voxel can also correspond to a brick. Each block may have 12 edges, and the approximation unit 420 can check whether each edge is adjacent to a voxel having a point. Each edge can be adjacent to multiple occupied voxels. A specific position of an edge adjacent to a voxel is called a vertex, and when multiple occupied voxels are adjacent to one edge, the approximation unit 420 can determine the average position of the positions as the vertex.

[0099] If a vertex exists, the point cloud encoding device 400 can entropy code the start point (x, y, z) of the edge, the direction vector (△x, △y, △z) of the edge, and the position value of the vertex (relative position value within the edge) via the geometry encoding unit 425.

[0100] The geometry restoration process may correspond to a process of generating restored geometry by reconstructing an octree and / or an approximated octree. The geometry restoration process may be performed by the restoration unit 430. The restoration unit 430 may perform the geometry restoration process through triangle reconstruction, up-sampling, voxelization, etc.

[0101] [Table 1]

[0102]

number

[0103]

number

[0104]

number

[0105] Also, the restoration unit 430 may find the minimum value of the added values ​​and perform a projection process along the axis where the minimum value is located.

[0106] For example, if the x element is minimum, the restoration unit 430 may project each vertex onto the x-axis based on the center of the block, and then project it onto the (y, z) plane. Also, if the value derived from projecting onto the (y, z) plane is (ai, bi), the restoration unit 430 may calculate a θ value through atan2(bi, ai) and align the vertices based on the θ value.

[0107] The method of reconstructing triangles according to the number of vertices can be created by combining them according to the sorting order as shown in Table 1 below. For example, if there are four vertices (n=4), two triangles (1, 2, 3) and (3, 4, 1) can be constructed. The first triangle 1, 2, 3 can be constructed from the first, second, and third vertices of the sorted vertices, and the second triangle 3, 4, 1 can be constructed from the third, fourth, and first vertices.

[0108] [Table 2]

[0109] The restoration unit 430 may perform an upsampling process to add intermediate points along the edges of triangles and voxelize them. The restoration unit 430 may generate additional points based on an upsampling factor and a block width. These points may be called refined vertices. The restoration unit 430 may voxelize the refined vertices, and the point cloud encoding device 400 may perform attribute coding based on the voxelized position values.

[0110] According to an embodiment, the geometry encoding unit 425 may improve compression efficiency by applying context adaptive arithmetic coding. The geometry encoding unit 425 may directly entropy code the occupancy code using the arithmetic code. According to an embodiment, the geometry encoding unit 425 may adaptively encode based on the occupancy of neighboring nodes (intra-coding) or based on the occupancy code of a previous frame (inter-coding). Here, a frame may refer to a collection of point cloud data generated at the same time. Intra-coding and inter-coding are optional processes and may be omitted.

[0111] Compression efficiency varies depending on the number of neighboring nodes referenced, and as the number of bits increases, the encoding process becomes more complex, but compression efficiency can be improved by focusing on one side. For example, if you have a 3-bit context, you may need to code it into 23 = 8 types. The part where coding is performed separately can affect the complexity of implementation, so it is necessary to match the compression efficiency and complexity to an appropriate level.

[0112] In the case of intra-coding, the geometry encoding unit 425 may first determine a neighbor pattern value based on the occupancy of neighboring nodes. An example of the neighbor pattern is shown in FIG.

[0113] 7(a) shows a cube corresponding to a node (the cube located in the center) and six cubes (adjacent nodes) that share at least one face with the cube. The illustrated nodes are at the same depth. The illustrated numbers indicate the weights (1, 2, 4, 8, 16, 32, etc.) associated with each of the six nodes. Each weight is assigned sequentially according to the position of the adjacent node.

[0114] (b) of FIG. 7 shows the adjacent node pattern value. The adjacent node pattern value is the sum of values ​​multiplied by the weights of occupied adjacent nodes (adjacent nodes with points). Therefore, the adjacent node pattern value can have a value from 0 to 63. When the adjacent node pattern value is 0, it indicates that there are no nodes with points (occupied nodes) among the adjacent nodes of the node. When the adjacent node pattern value is 63, it indicates that all adjacent nodes are occupied nodes. In (b) of FIG. 7, the adjacent nodes assigned weights 1, 2, 4, and 8 are occupied nodes, so the adjacent node pattern value is 15, which is the sum of 1, 2, 4, and 8.

[0115] The geometry encoding unit 425 may perform coding according to the adjacent node pattern value. For example, if the adjacent node pattern value is 63, the geometry encoding unit 425 may perform 64 types of coding. According to an embodiment, the geometry encoding unit 425 may reduce coding complexity by changing the adjacent node pattern value. For example, the change in the adjacent node pattern value may be performed based on a table that changes 64 to 10 or 6.

[0116] Attribute Coding

[0117] Attribute coding can be a process of coding attribute information based on the restored (reconstructed) geometry and the geometry before coordinate system transformation (original geometry). Because attributes can be dependent on geometry, the restored geometry can be used for attribute coding.

[0118] As mentioned above, attributes can include hue, reflectance, etc. The same attribute coding method can be applied to the information or parameters included in the attributes. Hue has three elements, and reflectance has one element, and each element can be processed independently.

[0119] Attribute coding may include a hue conversion process, an attribute conversion process, a prediction conversion process, a lift conversion process, a RAHT conversion process, an attribute quantization process, an attribute encoding process, etc. The prediction conversion process, the lift conversion process, and the RAHT conversion process may be used selectively, or one or more of them may be used in combination.

[0120] The hue conversion process may correspond to a process of converting the format of the hue in an attribute into another format. The hue conversion process may be performed by the color conversion unit 435. That is, the color conversion unit 435 may convert the hue in the attribute. For example, the color conversion unit 435 may perform a coding operation to convert the hue in the attribute from RGB to YCbCr. Depending on the embodiment, the operation of the color conversion unit 435, i.e., the hue conversion process, may be selectively applied depending on the hue value included in the attribute.

[0121] As described above, when one or more points exist in a voxel, the position values ​​of the points in the voxel can be set to the center of the voxel in order to integrate and display them as one point information for the voxel. This may require a process of converting the attribute values ​​associated with the points. In addition, an attribute conversion process can be performed even when the trisoup mode is performed.

[0122] The attribute conversion process may correspond to a process of converting attributes based on a position where geometry coding has not been performed and / or a reconstructed geometry. For example, the attribute conversion process may correspond to a process of converting attributes of a point at a position based on the position of the point included in a voxel. The attribute conversion process may be performed by the attribute conversion unit 440.

[0123] The attribute conversion unit 440 can calculate the average value of the central position value of a voxel and the attribute values ​​of adjacent points (neighboring points) within a specific radius. Alternatively, the attribute conversion unit 440 can apply weights based on the distance from the central position to the attribute values ​​and calculate the average value of the weighted attribute values. In this case, each voxel has a position and a calculated attribute value.

[0124] When searching for neighboring points within a specific location or radius, KD trees or Moulton codes can be used. KD trees are binary search trees that support a data structure that allows for location-based management of points, enabling fast nearest neighbor (NNS) searches. Moulton codes can be generated by mixing the bits of the three-dimensional position information (x, y, z) for all points. For example, if the (x, y, z) values ​​are (5, 9, 1), the bits representing (5, 9, 1) become (0101, 1001, 0001). Mixing these values ​​in the order of z, y, and x according to the bit index results in 010001000111, which is 1095. In other words, 1095 is the Moulton code value for (5, 9, 1). Points are aligned based on the Moulton code, and nearest neighbor search (NNS) is possible through a depth-first travertex process.

[0125] After the attribute transformation process, there may be cases where a nearest neighbor search (NNS) is required in other transformation processes for attribute coding, and in such cases, a KD tree or a Moulton code can be used.

[0126] The predictive transformation process may correspond to a process of predicting an attribute value of a current point (a point corresponding to a prediction target) based on attribute values ​​of one or more points (neighboring points) adjacent to the current point. The predictive transformation process may be performed by a level of detail (LOD) generating unit 450.

[0127] The predictive conversion is a method to which an LOD conversion technique is applied, and the LOD generating unit 450 can calculate and set the LOD value of each point based on the LOD distance value of each point.

[0128] An example of point organization according to LOD distance values ​​is shown in FIG. 8. In FIG. 8, the first image shows the original point cloud content, the second image shows the distribution of points with the lowest LOD, and the seventh image shows the distribution of points with the highest LOD, based on the direction of the arrow. As shown in FIG. 8, points with the lowest LOD may be sparsely distributed, and points with the highest LOD may be densely distributed. That is, as the LOD increases, the spacing (or distance) between points may become shorter.

[0129] Each point in the point cloud can be separated by LOD, and the configuration of points by LOD can include points that belong to a lower LOD than the LOD value. For example, the configuration of points with LOD level 2 can include all points that belong to LOD levels 1 and 2.

[0130] An example of the configuration of points by LOD is shown in Figure 9. The upper diagram of Figure 9 shows an example of points (P0 to P9) in point cloud content distributed in three-dimensional space. The original order in Figure 9 indicates the order of points P0 to P9 before LOD generation, and the LOD-based order in Figure 9 indicates the order of points after LOD generation.

[0131] 9, points can be reordered by LOD, with higher LODs including points from lower LODs. For example, LOD0 can include P0, P5, P4, and P2, LOD1 can include points from LOD0, P1, P6, and P3, and LOD2 can include points from LOD0, LOD1, and P9, P8, and P7.

[0132] The LOD generator 450 may generate a predictor for each point for predictive conversion. Therefore, if there are N points, N predictors may be generated. The predictor may be set by calculating a weight value (=1 / distance) based on the LOD value for each point, indexing information for adjacent points, and a distance value between the adjacent points. Here, the adjacent points may be points located within a distance set for each LOD from the current point.

[0133] In addition, the predictor may multiply the attribute values ​​of adjacent points by a "set weight" and set the average of the weighted attribute values ​​as the predicted attribute value of the current point. An attribute quantization process may be performed on a residual attribute value obtained by subtracting the predicted attribute value of the current point from the attribute value of the current point.

[0134] The lift conversion process may correspond to a process of reconstructing points into a set of detail levels through an LOD generation process, similar to the prediction conversion process. The lift conversion process may be performed by the lift unit 455. The lift conversion process may also include a process of generating a predictor for each point, a process of setting the calculated LOD to the predictor, a process of registering neighboring points, and a process of setting weights according to the distance between the current point and the neighboring points.

[0135] The difference between the lift transformation process and the prediction transformation process is that the lift transformation process is a method of cumulatively applying weights to attribute values. The method of cumulatively applying weights to attribute values ​​is as follows.

[0136] 1) There can be a separate array QW (quantization weight) that stores the weight value for each point. The initial value of all elements of QW is 1.0. The QW value of the predictor index of the adjacent node (adjacent point) registered in the predictor is multiplied by the weight of the predictor of the current point and added.

[0137] 2) To calculate the predicted attribute value, the attribute value of the point is multiplied by the weight and subtracted from the existing attribute value. This process is sometimes called the lift prediction process.

[0138] 3) Create temporary arrays called "updateweight" and "update" and initialize the elements in the arrays to 0.

[0139] 4) For all predictors, the calculated weight is multiplied by the weight stored in the QW to derive new weights, and the new weights are accumulated as the index of the adjacent node in updateweight, and the value obtained by multiplying the new weight by the attribute value of the index of the adjacent node is accumulated in update.

[0140] 5) For all predictors, divide the attribute value of update by the weight value of update weight of the predictor index and add the result to the existing attribute value. This process is sometimes called lift update process.

[0141] 6) For all predictors, the attribute values ​​updated through the lift update process are multiplied by the weights (stored in the QW) updated through the lift prediction process, the result (the multiplied value) is quantized, and the quantized value is entropy encoded.

[0142] The RAHT conversion process may correspond to a method of predicting attribute information of a node at a higher level using attribute information associated with a node at a lower level of the octree. That is, the RAHT conversion process may correspond to an attribute information intra-coding method using an octree backward scan. The RAHT conversion process may be performed by the RAHT conversion unit 445.

[0143] The RAHT conversion unit 445 scans the entire region with voxels and can perform the RAHT conversion process up to the root node by combining (merging) the voxels into larger blocks at each step. The RAHT conversion unit 445 performs the RAHT conversion process only on occupied nodes, so in the case of an unoccupied empty node, the RAHT conversion process can be performed on the node at the immediately upper level.

[0144] [Table 3]

[0145]

number

[0146] [Table 4]

[0147]

number

[0148] In Equation 6, the gDC value can also be quantized and entropy coded like the high-pass coefficient.

[0149] The attribute quantization process may correspond to a process of quantizing attributes output from the RAHT conversion unit 445, the LOD generation unit 450, and / or the lifting unit 455. The attribute quantization process may be performed by the attribute quantization unit 460. The attribute encoding process may correspond to a process of encoding the quantized attributes and outputting an attribute bitstream. The attribute encoding process may be performed by the attribute encoding unit 465.

[0150] For example, when the LOD generator 450 calculates a predicted attribute value of the current point, the attribute quantizer 460 may quantize a residual attribute value obtained by subtracting the predicted attribute value of the current point from the attribute value of the current point. An example of the attribute quantization process of the present disclosure is shown in Table 2.

[0151] [Table 5]

[0152] If there are no adjacent points in the predictor of each point, the attribute encoding unit 465 can directly entropy encode the attribute value (unquantized attribute value) of the current point. In contrast, if there are adjacent points in the predictor of the current point, the attribute encoding unit 465 can entropy encode the quantized residual attribute value.

[0153] As another example, if the lift unit 460 outputs a value obtained by multiplying an attribute value updated through a lift update process by a weight (stored in the QW) updated through a lift prediction process, the attribute quantization unit 460 can quantize the result (the value obtained by multiplication), and the attribute encoding unit 465 can entropy encode the quantized value.

[0154] Point Cloud Decoder Overview

[0155] 10 illustrates an example of a point cloud decoding device 1000 according to an embodiment of the present disclosure. The point cloud decoding device 1000 of FIG. 10 may correspond in configuration and function to the decoding unit 23 of FIG.

[0156] The point cloud decoding device 1000 may perform a decoding process based on data (bitstream) transmitted from the transmission device 10. The decoding process may include a process of restoring (decoding) a point cloud video by performing an operation corresponding to the encoding operation described above on the bitstream.

[0157] 10, the decoding process may include a geometry decoding process and an attribute decoding process. The geometry decoding process may be performed by a geometry decoding unit 1010, and the attribute decoding process may be performed by an attribute decoding unit 1020. That is, the point cloud decoding apparatus 1000 may include the geometry decoding unit 1010 and the attribute decoding unit 1020.

[0158] The geometry decoding unit 1010 can restore geometry from the geometry bitstream, and the attribute decoding unit 1020 can restore attributes based on the restored geometry and attribute bitstream. In addition, the point cloud decoding device 1000 can restore a 3D point cloud video (point cloud data) based on position information according to the restored geometry and attribute information according to the restored attributes.

[0159] 11 illustrates a specific example of a point cloud decoding apparatus 1100 according to another embodiment of the present disclosure. As illustrated in FIG. 11, the point cloud decoding apparatus 1100 may include a geometry decoding unit 1105, an octree synthesis unit 1110, an approximation synthesis unit 1115, a geometry restoration unit 1120, a coordinate system inverse transformation unit 1125, an attribute decoding unit 1130, an attribute inverse quantization unit 1135, a RATH transformation unit 1150, an LOD generation unit 1140, an inverse lifting unit 1145, and / or a color inverse transformation unit 1155.

[0160] The geometry decoding unit 1105, the octree synthesis unit 1110, the approximation synthesis unit 1115, the geometry restoration unit 1120, and the coordinate system inverse transformation unit 1150 may perform geometry decoding. Geometry decoding may be performed by the reverse process of the geometry coding described with reference to FIGS. 1 to 9. Geometry decoding may include direct coding and trisoup geometry decoding. Direct coding and trisoup geometry decoding may be selectively applied.

[0161] The geometry decoding unit 1105 can decode the received geometry bitstream based on arithmetic coding. The operation of the geometry decoding unit 1105 can correspond to the reverse process of the operation performed by the geometry encoding unit 435.

[0162] The octree synthesis unit 1110 can generate an octree by obtaining an occupation code from the decoded geometry bitstream (or from geometry information obtained as a decoding result). The operation of the octree synthesis unit 1110 can correspond to the reverse process of the operation performed by the octree analysis unit 415.

[0163] The approximation synthesis unit 1115 can synthesize a surface based on the decoded geometry and / or the generated octree if trisou geometry encoding is applied.

[0164] The geometry restoration unit 1120 may restore geometry based on the surface and the decoded geometry. When direct coding is applied, the geometry restoration unit 1120 may directly add position information of points to which direct coding is applied. Also, when trisoup geometry encoding is applied, the geometry restoration unit 1120 may restore geometry by performing a reconstruction operation, such as triangulation, upsampling, or voxelization. The restored geometry may include a point cloud picture or frame without attributes.

[0165] The coordinate system inverse transformation unit 1150 may transform the coordinate system based on the reconstructed geometry to obtain the position of the point. For example, the coordinate system inverse transformation unit 1150 may inversely transform the position of the point from a three-dimensional space (e.g., a three-dimensional space expressed by an X-axis, a Y-axis, and a Z-axis coordinate system) to position information in a global space coordinate system.

[0166] The attribute decoding unit 1130, the attribute inverse quantization unit 1135, the RAHT transform unit 1230, the LOD generation unit 1140, the inverse lift unit 1145, and / or the color inverse transform unit 1250 may perform attribute decoding. Attribute decoding may include RAHT transform decoding, predictive transform decoding, and lift transform decoding. The above three types of decoding may be used selectively, or a combination of one or more of the above decoding types may be used.

[0167] The attribute decoding unit 1130 can decode the attribute bitstream based on arithmetic coding. For example, if the attribute value of the current point is directly entropy encoded because there are no adjacent points in the predictor of each point, the attribute decoding unit 1130 can decode the attribute value of the current point (unquantized attribute value). As another example, if the predictor of the current point has adjacent points and the quantized residual attribute value is entropy encoded, the attribute decoding unit 1130 can decode the quantized residual attribute value.

[0168] The attribute inverse quantization unit 1135 may inverse quantize the decoded attribute bitstream or information about the attribute obtained as a result of decoding, and output the inverse quantized attribute (or attribute value). For example, if the attribute decoding unit 1130 outputs a quantized residual attribute value, the attribute inverse quantization unit 1135 may inverse quantize the quantized residual attribute value and output the residual attribute value. The inverse quantization process may be selectively applied depending on the attribute encoding method of the point cloud encoding device 400. That is, if the attribute value of the current point is directly encoded because there are no neighboring points in the predictor of each point, the attribute decoding unit 1130 may output the unquantized attribute value of the current point, and the attribute encoding process may be skipped. An example of the attribute inverse quantization process of the present disclosure is shown in Table 3.

[0169] [Table 6]

[0170] The RATH transform unit 1150, the LOD generator 1140, and / or the inverse lifting unit 1145 may process the reconstructed geometry and dequantized attributes. The RATH transform unit 1150, the LOD generator 1140, and / or the inverse lifting unit 1145 may selectively perform a decoding operation corresponding to the encoding operation of the point cloud encoding device 400.

[0171] The color inverse transform unit 1155 may perform inverse transform coding to inversely transform the color values ​​(or textures) included in the decoded attributes. The operation of the color inverse transform unit 1155 may be selectively performed depending on the operation of the color transform unit 435.

[0172] 12 illustrates another example of a transmission device according to an embodiment of the present disclosure. As illustrated in FIG. 12, the transmission device may include a data input unit 1205, a quantization processing unit 1210, a voxelization processing unit 1215, an octree occupancy code generation unit 1220, a surface model processing unit 1225, an intra / inter coding processing unit 1230, an arithmetic coder 1235, a metadata processing unit 1240, a hue conversion processing unit 1245, an attribute conversion processing unit 1250, a prediction / lift / RAHT conversion processing unit 1255, an arithmetic coder 1260, and a transmission processing unit 1265.

[0173] The function of the data input unit 1205 may correspond to the acquisition process performed by the acquisition unit 11 in FIG. 1. That is, the data input unit 1205 may acquire a point cloud video and generate point cloud data for a plurality of points. Geometry information (position information) in the point cloud data may be generated in the form of a geometry bitstream through a quantization processing unit 1210, a voxelization processing unit 1215, an octree occupation code generation unit 1220, a surface model processing unit 1225, an intra / inter coding processing unit 1230, and an arithmetic coder 1235. Attribute information in the point cloud data may be generated in the form of an attribute bitstream through a color conversion processing unit 1245, an attribute conversion processing unit 1250, a prediction / lift / RAHT conversion processing unit 1255, and an arithmetic coder 1260. The geometry bitstream, attribute bitstream, and / or metadata bitstream may be transmitted to a receiving device through processing by a transmission processing unit 1265.

[0174] Specifically, the function of the quantization processing unit 1210 may correspond to the quantization process performed by the geometry quantization unit 410 of Figure 4 and / or the function of the coordinate system conversion unit 405. The function of the voxelization processing unit 1215 may correspond to the voxelization process performed by the geometry quantization unit 410 of Figure 4, and the function of the octree occupation code generation unit 1220 may correspond to the function performed by the octree analysis unit 415 of Figure 4. The function of the surface model processing unit 1225 may correspond to the function performed by the approximation unit 420 of Figure 4, and the function of the intra / inter coding processing unit 1230 and the function of the arithmetic coder 1235 may correspond to the function performed by the geometry encoding unit 425 of Figure 4. The function of the metadata processing unit 1240 may correspond to the function of the metadata processing unit described with reference to Figure 1.

[0175] 1. The function of the hue conversion processing unit 1245 may correspond to the function performed by the color conversion unit 435 in Fig. 4, and the function of the attribute conversion processing unit 1250 may correspond to the function performed by the attribute conversion unit 440 in Fig. 4. The function of the prediction / lift / RAHT conversion processing unit 1255 may correspond to the function performed by the RAHT conversion unit 4450, LOD generation unit 450, and lift unit 455 in Fig. 4, and the function of the arithmetic coder 1260 may correspond to the function of the attribute encoding unit 465 in Fig. 4. The function of the transmission processing unit 1265 may correspond to the function performed by the transmission unit 14 and / or the encapsulation processing unit 13 in Fig. 1.

[0176] 13 illustrates another example of a receiving device according to an embodiment of the present disclosure. As illustrated in FIG. 13, the receiving device may include a receiving unit 1305, a receiving processing unit 1310, an arithmetic decoder 1315, a metadata parser 1335, a proprietary code-based octree reconstruction processing unit 1320, a surface model processing unit 1325, an inverse quantization processing unit 1330, an arithmetic decoder 1340, an inverse quantization processing unit 1345, a prediction / lift / RAHT inverse transform processing unit 1350, a hue inverse transform processing unit 1355, and a renderer 1360.

[0177] The function of the receiving unit 1305 may correspond to the function performed by the receiving unit 21 in FIG. 1, and the function of the receiving processing unit 1310 may correspond to the function performed by the decapsulation processing unit 22 in FIG. 1. That is, the receiving unit 1305 receives a bitstream from the transmission processing unit 1265, and the receiving processing unit 1310 extracts a geometry bitstream, an attribute bitstream, and / or a metadata bitstream through a decapsulation process. The geometry bitstream may be generated as reconstructed (restored) position values ​​(position information) through an arithmetic decoder 1315, a proprietary code-based octree reconstruction processing unit 1320, a surface model processing unit 1325, and an inverse quantization processing unit 1330. The attribute bitstream may be generated as reconstructed attribute values ​​through an attribute decoder 1340, an inverse quantization processing unit 1345, a prediction / lift / RAHT inverse transform processing unit 1350, and a hue inverse transform processing unit 1355. The metadata bitstream can be generated as metadata (or metadata information) restored through the metadata parser 1335. The position values, attribute values, and / or metadata can be rendered by the renderer 1360 to provide the user with an experience such as VR / AR / MR / autonomous driving.

[0178] Specifically, the function of the arithmetic decoder 1315 may correspond to the function performed by the geometry decoding unit 1105 of Figure 11, and the function of the proprietary code-based octree reconstruction processing unit 1320 may correspond to the function performed by the octree synthesis unit 1110 of Figure 11. The function of the surface model processing unit 1325 may correspond to the function performed by the approximation synthesis unit of Figure 11, and the function of the inverse quantization processing unit 1330 may correspond to the function performed by the geometry restoration unit 1120 and / or the coordinate system inverse transformation unit 1125 of Figure 11. The function of the metadata parser 1335 may correspond to the function performed by the metadata parsing unit described in Figure 1.

[0179] 11. The function of the arithmetic decoder 1340 may correspond to the function performed by the attribute decoding unit 1130 in Fig. 11, and the function of the inverse quantization processing unit 1345 may correspond to the function performed by the attribute inverse quantization unit 1135 in Fig. 11. The function of the prediction / lift / RAHT inverse transform processing unit 1350 may correspond to the function performed by the RAHT transform unit 1150, LOD generation unit 1140, and inverse lift unit 1145 in Fig. 11, and the function of the hue inverse transform processing unit 1355 may correspond to the function performed by the color inverse transform unit 1155 in Fig. 11.

[0180] FIG. 14 illustrates an example of a structure that can be coupled with a point cloud data transmission / reception method / apparatus according to an embodiment of the present disclosure.

[0181] The structure of Figure 14 shows a configuration in which at least one of an AI server, a robot, a self-driving vehicle, an XR device, a smartphone, a home appliance, and / or an HMD is connected to a cloud network. The robot, the self-driving vehicle, the XR device, the smartphone, or the home appliance may be referred to as a device. The XR device may correspond to a point cloud data device (PCC) according to an embodiment or may be linked to a PCC device.

[0182] A cloud network may refer to a network that forms part of or exists within a cloud computing infrastructure, and may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, a 5G network, or the like.

[0183] The server is connected to at least one of a robot, an autonomous vehicle, an XR device, a smartphone, a home appliance, and / or an HMD via a cloud network and can assist with at least some of the processing of the connected devices.

[0184] The HMD can indicate one of the types that can be realized by the XR device and / or the PCC device according to the embodiment. The device of the HMD type according to the embodiment can include a communication unit, a control unit, a memory unit, an I / O unit, a sensor unit, a power supply unit, and the like.

[0185] <PCC+XR>

[0186] The XR / PCC device may be realized by an HMD, a HUD equipped in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a fixed robot or a mobile robot, etc. by applying the PCC and / or XR technology.

[0187] The XR / PCC device can obtain information about the surrounding space or real objects by analyzing three-dimensional point cloud data or image data acquired from various sensors or an external device to generate position (geometry) data and attribute data for three-dimensional points, and render and output an XR object that outputs the information. For example, the XR / PCC device can output an XR object including additional information for a recognized object corresponding to the recognized object.

[0188] <PCC+XR+mobile phone>

[0189] [[ID=​​​​​​​​Autonomous vehicles can be realized as mobile robots, vehicles, unmanned aerial vehicles, etc. by applying PCC technology and XR technology. An autonomous vehicle to which XR / PCC technology is applied can refer to an autonomous vehicle equipped with a means for providing XR images, or an autonomous vehicle that can be controlled / interacted with within an XR image. In particular, an autonomous vehicle that can be controlled / interacted with within an XR image can be distinguished from an XR device and can be linked with it.

[0192] An autonomous vehicle equipped with a means for providing an XR / PCC image can acquire sensor information from sensors including a camera and output an XR / PCC image generated based on the acquired sensor information. For example, an autonomous vehicle equipped with a HUD can output an XR / PCC image to provide a passenger with an XR / PCC object corresponding to a real-world object or an object on a screen.

[0193] In this case, when an XR / PCC object is output to a HUD, at least a portion of the XR / PCC object may be output to overlap with an actual object toward which the passenger's gaze is directed. In contrast, when an XR / PCC object is output to a display within an autonomous vehicle, at least a portion of the XR / PCC object may be output to overlap with an object within the screen. For example, an autonomous vehicle may output XR / PCC objects corresponding to objects such as roadways, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, and buildings.

[0194] The VR technology, AR technology, MR technology, and / or PCC technology according to the embodiments may be applied to various devices. That is, VR technology is a display technology that provides real-world objects and backgrounds only as CG images. In contrast, AR technology refers to a technology that displays virtual CG images on top of images of actual objects. MR technology is similar to the aforementioned AR technology in that it mixes and combines virtual objects with the real world. However, AR technology clearly distinguishes between real objects and virtual objects created from CG images and uses virtual objects to complement real objects, whereas MR technology is distinct from AR technology in that virtual objects are considered to be equivalent to real objects. More specifically, for example, a hologram service is an application of the aforementioned MR technology. VR, AR, and MR technologies are sometimes collectively referred to as XR technology.

[0195] space division

[0196] Point cloud data (i.e., G-PCC data) can represent a volumetric encoding of a point cloud consisting of a sequence of frames (point cloud frames). Each point cloud frame can include the number of points, the positions of the points, and the attributes of the points. The number of points, the positions of the points, and the attributes of the points can vary from frame to frame. Each point cloud frame can represent a set of 3D points at a particular time instance, specified by their Cartesian coordinates (x, y, z) and zero or more attributes. Here, the Cartesian coordinates (x, y, z) can be positions or geometries.

[0197] According to an embodiment, the present disclosure may further perform a spatial partitioning process of partitioning point cloud data into one or more 3D blocks before encoding the point cloud data. The 3D block may refer to all or a portion of a 3D space occupied by the point cloud data. The 3D block may refer to one or more of a tile group, a tile, a slice, a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

[0198] A tile, which corresponds to a 3D block, may refer to all or a portion of a 3D space occupied by point cloud data. A slice, which corresponds to a 3D block, may also refer to all or a portion of a 3D space occupied by point cloud data. A tile may be divided into one or more slices based on the number of points included in the tile. A tile may be a group of slices having bounding box information. The bounding box information of each tile may be specified in a tile inventory (or a tile parameter set (TPS)). A tile may overlap other tiles within the bounding box. A slice may be a unit of data that is independently encoded or independently decoded. That is, a slice may be a set of points that can be independently encoded or decoded. According to an embodiment, a slice may be a series of syntax elements representing a portion or an entire coded point cloud frame. Each slice may include an index for identifying the tile to which the slice belongs.

[0199] The spatially divided 3D blocks may be processed independently or independently. For example, the spatially divided 3D blocks may be encoded or decoded independently or independently, and may be transmitted or received independently or independently. The spatially divided 3D blocks may be quantized or dequantized independently or independently, and may be transformed or inversely transformed independently or independently. The spatially divided 3D blocks may be rendered independently or independently. For example, encoding or decoding may be performed on a slice-by-slice or tile-by-tile basis. Quantization or inverse quantization may be performed differently for each tile or slice, or may be performed differently for each transformed or inversely transformed tile or slice.

[0200] In this way, by spatially dividing point cloud data into one or more 3D blocks and processing the spatially divided 3D blocks independently or independently, the process of processing the 3D blocks can be performed in real time with low latency. Furthermore, random access and parallel encoding or parallel decoding in the 3D space occupied by the point cloud data are possible, which can prevent errors from accumulating during the encoding or decoding process.

[0201] 15 is a block diagram illustrating an example of a transmission device 1500 that performs a spatial division process according to an embodiment of the present disclosure. As shown in FIG. 15, the transmission device 1500 may include a spatial division unit 1505 that performs the spatial division process, a signaling processing unit 1510, a geometry encoder 1515, an attribute encoder 1520, an encapsulation processing unit 1525, and / or a transmission processing unit 1530.

[0202] The spatial division unit 1505 may perform a spatial division process of dividing point cloud data into one or more 3D blocks based on a bounding box and / or a sub-bounding box. Through the spatial division process, the point cloud data may be divided into one or more tiles and / or one or more slices. Depending on the embodiment, through the spatial division process, the point cloud data may be divided into one or more tiles, and each divided tile may be further divided into one or more slices.

[0203] FIG. 16 shows an example of spatially dividing a bounding box (i.e., point cloud data) into one or more three-dimensional blocks. As shown in FIG. 16, the overall bounding box of the point cloud data can be divided into three tiles, i.e., tile #0, tile #1, and tile #2. Tile #0 can be further divided into two slices, i.e., slice #0 and slice #1. Tile #1 can be further divided into two slices, i.e., slice #2 and slice #3. Tile #2 can be further divided into slice #4.

[0204] The signaling processor 1510 may generate and / or process (e.g., entropy encode) signaling information and output it in the form of a bitstream. Hereinafter, the bitstream (in which the signaling information is encoded) output from the signaling processor is referred to as a "signaling bitstream." The signaling information may include information for or about spatial division. That is, the signaling information may include information related to the spatial division process performed by the spatial divider 1505.

[0205] When point cloud data is divided into one or more 3D blocks, information for decoding a portion of point cloud data corresponding to a specific tile or slice of the point cloud data may be required. Furthermore, information related to a 3D spatial region may be required to support spatial access (or partial access) of the point cloud data. Here, spatial access may refer to extracting only a portion of point cloud data required from the entire point cloud data from a file. The signaling information may include information for decoding a portion of the point cloud data, information related to a 3D spatial region for supporting spatial access, etc. For example, the signaling information may include 3D bounding box information, 3D spatial region information, tile information, and / or tile inventory information, etc.

[0206] The signaling information can be provided from the spatial division unit 1505, the geometry encoder 1515, the attribute encoder 1520, the transmission processing unit 1525, and / or the encapsulation processing unit 1530. In addition, the signaling processing unit 1510 can provide feedback information fed back from the receiving device 1700 of FIG. 17 to the spatial division unit 1505, the geometry encoder 1515, the attribute encoder 1520, the transmission processing unit 1525, and / or the encapsulation processing unit 1530.

[0207] The signaling information may be stored and signaled in a sample, sample entry, sample group, track group, or separate metadata track within a track. According to an embodiment, the signaling information may be signaled in units of a sequence parameter set (SPS) for sequence-level signaling, a geometry parameter set (GPS) for signaling geometry coding information, an attribute parameter set (APS) for signaling attribute coding information, a tile parameter set (TPS) (or tile inventory) for tile-level signaling, etc. Furthermore, the signaling information may be signaled in units of a coding unit such as a slice or a tile.

[0208] Meanwhile, the position of the 3D block (position information) can be output to the geometry encoder 1515, and the attribute of the 3D block (attribute information) can be output to the attribute encoder 1520.

[0209] The geometry encoder 1515 may construct an octree based on the position information, encode the constructed octree, and output a geometry bitstream. The geometry encoder 1515 may also reconstruct (restore) the octree and / or an approximated octree and output the reconstructed octree to the attribute encoder 1520. The reconstructed octree may be reconstructed geometry. The geometry encoder 1515 may perform all or some of the operations performed by the coordinate system converter 405, geometry quantizer 410, octree analyzer 415, approximator 420, geometry encoding unit 425, and / or reconstruction unit 430 of FIG. 4. Depending on the embodiment, the geometry encoder 1515 may perform all or some of the operations performed by the quantizer 1210, voxelization processor 1215, octree occupation code generator 1220, surface model processor 1225, intra / inter-coding processor 1230, and / or arithmetic coder 1235 of FIG. 12.

[0210] The attribute encoder 1520 may output an attribute bitstream by encoding attributes based on the restored geometry. The attribute encoder 1520 may perform all or some of the operations performed by the attribute conversion unit 440, the RAHT conversion unit 445, the LOD generation unit 450, the lift unit 455, the attribute quantization unit 460, the attribute encoding unit 465, and / or the color conversion unit 435 of FIG. 4. Depending on the embodiment, the attribute encoder 1520 may perform all or some of the operations performed by the attribute conversion processing unit 1250, the prediction / lift / RAHT conversion processing unit 1255, the arithmetic coder 1260, and / or the hue conversion processing unit 1245 of FIG. 12.

[0211] The encapsulation processor 1525 may encapsulate one or more input bitstreams into a file or a segment. For example, the encapsulation processor 1525 may encapsulate a geometry bitstream, an attribute bitstream, and a signaling bitstream separately, or may multiplex and encapsulate a geometry bitstream, an attribute bitstream, and a signaling bitstream. According to an embodiment, the encapsulation processor 1525 may encapsulate a bitstream (G-PCC bitstream) composed of a sequence of TLV (type-length-value) structures into a file. The TLV (or TLV encapsulation) structure constituting the G-PCC bitstream may include a geometry bitstream, an attribute bitstream, a signaling bitstream, etc. Depending on the embodiment, the G-PCC bitstream may be generated by the encapsulation processor 1525 or the transmission processor 1530. The TLV structure or TLV encapsulation structure will be described in detail below. Depending on the embodiment, the encapsulation processor 1525 may perform all or part of the operations performed by the encapsulation processor 13 of FIG.

[0212] The transmission processing unit 1530 can process encapsulated bitstreams or files / segments according to any transmission protocol. The transmission processing unit 1530 can perform all or part of the operations performed by the transmission unit 14 and transmission processing unit described with reference to Fig. 1 or the transmission processing unit 1265 in Fig. 12.

[0213] 17 is a block diagram illustrating an example of a receiving device 1700 according to an embodiment of the present disclosure. The receiving device 1700 may perform operations corresponding to those of the transmitting device 1500 that performs spatial division. As shown in FIG. 17, the receiving device 1700 may include a receiving processing unit 1705, a decapsulation processing unit 1710, a signaling processing unit 1715, a geometry decoder 1720, an attribute encoder 1725, and / or a post-processing unit 1730.

[0214] The receiving processor 1705 can receive a file / segment in which a G-PCC bitstream is encapsulated, a G-PCC bitstream, or a bitstream, and process these in accordance with a transmission protocol. The receiving processor 1705 can perform all or part of the operations performed by the receiving unit 21 and receiving processor described with reference to Fig. 1, or the receiving unit 1305 or receiving processor 1310 in Fig. 13.

[0215] The decapsulation processing unit 1710 can obtain a G-PCC bitstream by performing the reverse process of the operation performed by the encapsulation processing unit 1525. The decapsulation processing unit 1710 can obtain a G-PCC bitstream by decapsulating a file / segment. For example, the decapsulation processing unit 1710 can obtain a signaling bitstream and output it to the signaling processing unit 1715, obtain a geometry bitstream and output it to the geometry decoder 1720, and obtain an attribute bitstream and output it to the attribute decoder 1725. The decapsulation processing unit 1710 can perform all or part of the operations performed by the decapsulation processing unit 22 in FIG. 1 or the reception processing unit 1410 in FIG. 13.

[0216] The signaling processing unit 1715 can parse and decode signaling information by performing the reverse process of the operations performed by the signaling processing unit 1510. The signaling processing unit 1715 can parse and decode signaling information from the signaling bitstream. The signaling processing unit 1715 can provide the decoded signaling information to the geometry decoder 1720, the attribute decoder 1720, and / or the post-processing unit 1730.

[0217] The geometry decoder 1720 can reconstruct the geometry from the geometry bitstream by performing the reverse process of the operations performed by the geometry encoder 1515. The geometry decoder 1720 can reconstruct the geometry based on the signaling information (parameters related to the geometry). The reconstructed geometry can be provided to the attribute decoder 1725.

[0218] The attribute decoder 1725 can recover attributes from the attribute bitstream by performing the reverse process of the operations performed by the attribute encoder 1520. The attribute decoder 1725 can recover attributes based on signaling information (parameters associated with the attributes) and the recovered geometry.

[0219] The post-processing unit 1730 may restore point cloud data based on the restored geometry and the restored attributes. The restoration of point cloud data may be performed by matching the restored geometry and the restored attributes. According to an embodiment, when the restored point cloud data is in units of tiles and / or slices, the post-processing unit 1730 may restore a bounding box of the point cloud data by performing a reverse process of the spatial division process of the transmitting device 1500 based on the signaling information. According to an embodiment, when the bounding box is divided into a plurality of tiles and / or a plurality of slices through the spatial division process, the post-processing unit 1730 may restore a portion of the bounding box by combining some slices and / or some tiles based on the signaling information. Here, some slices and / or some tiles used to restore the bounding box may be slices and / or some tiles associated with a 3D spatial region for which spatial proximity is desired.

[0220] Bitstream

[0221] FIG. 18 shows an example of the structure of a bitstream according to an embodiment of the present disclosure, FIG. 19 shows an example of the identification relationship between structures within a bitstream according to an embodiment of the present disclosure, and FIG. 20 shows the reference relationship between structures within a bitstream according to an embodiment of the present disclosure.

[0222] When the geometry bitstream, the attribute bitstream, and / or the signaling bitstream are configured as one bitstream (or a G-PCC bitstream), the bitstream may include one or more sub-bitstreams.

[0223] As shown in FIG. 18, a bitstream may include one or more SPSs, one or more GPSs, one or more APSs (APS0, APS1), one or more TPSs, and / or one or more slices (slice0, ..., slicen). A tile is a slice group including one or more slices, so a bitstream may include one or more tiles. The TPS may include information about each tile (e.g., information such as bounding box coordinate values, height, and / or size), and each slice may include a geometry bitstream (Geom0) and / or one or more attribute bitstreams (Attr0, Attr1). For example, slice0 (slice0) may include a geometry bitstream (Geom00) and / or one or more attribute bitstreams (Attr00, Attr10).

[0224] The geometry bitstream in each slice may consist of a geometry slice header (Geom_slice_header) and geometry slice data (Geom_slice_data). The geometry slice header may include information about the parameter set identification information (geom_parameter_set_id) included in the GPS, a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and / or information about the data included in the geometry slice data (geom_slice_data) (geomBoxOrigin, geom_box_log2_scale, geom_max_node_size_log2, geom_num_points). geomBoxOrigin is geometry box origin information indicating the box origin of the geometry slice data, geom_box_log2_scale is information indicating the log scale of the geometry slice data, geom_max_node_size_log2 is information indicating the size of the root geometry octree node, and geom_num_points is information related to the number of points in the geometry slice data. The geometry slice data may include geometry information (or geometry data) of the point cloud data in that slice.

[0225] Each attribute bitstream in each slice may include an attribute slice header (Attr_slice_header) and attribute slice data (Atrr_slice_data). The attribute slice header may include information about the attribute slice data, and the attribute slice data may include attribute information (or attribute data) of the point cloud data in the slice. When there are multiple attribute bitstreams in one slice, each may include different attribute information. For example, one attribute bitstream may include attribute information corresponding to hue, and another attribute bitstream may include attribute information corresponding to reflectance.

[0226] As shown in FIGS. 19 and 20 , an SPS may include an identifier (seq_parameter_set_id) for identifying the SPS, and a GPS may include an identifier (geom_parameter_set_id) for identifying the GPS and an identifier (seq_parameter_set_id) indicating the active SPS to which the GPS belongs (references). Furthermore, an APS may include an identifier (attr_parameter_set_id) for identifying the APS and an identifier (seq_parameter_set_id) indicating the active SPS referenced by the APS. The geometry data includes a geometry slice header and geometry slice data, and the geometry slice header may include an identifier (geom_parameter_set_id) of the active GPS referenced by the geometry slice. The geometry slice header may further include an identifier (geom_slice_id) for identifying the geometry slice and / or an identifier (geom_tile_id) for identifying the tile. The attribute data includes an attribute slice header and attribute slice data, and the attribute slice header can include an identifier (attr_parameter_set_id) of the active APS referenced by the attribute slice and an identifier (geom_slice_id) for identifying the geometry slice associated with the attribute slice.

[0227] Through this reference relationship, a geometry slice can reference a GPS, and a GPS can reference an SPS. The SPS can also list available attributes and assign identifiers to each listed attribute to identify the decoding method. Attribute slices can be mapped to output attributes by identifiers, and attribute slices themselves can have dependencies on previously decoded geometry slices and APSs.

[0228] Depending on the embodiment, parameters required for encoding point cloud data may be newly defined in the parameter set of the point cloud data and / or the slice header. For example, when encoding attributes, parameters required for the encoding may be newly defined (added) to the APS, and when encoding based on tiles, parameters required for the encoding may be newly defined (added) to the tile and / or slice header.

[0229] SPS Syntax Structure

[0230] 21 illustrates an example of a syntax structure of an SPS according to an embodiment of the present disclosure. In FIG. 21, the syntax elements (or fields) shown in the syntax structure of the SPS may be syntax elements included in the SPS or may be syntax elements signaled via the SPS.

[0231] The main_profile_compatibility_flag may indicate whether the bitstream conforms to the Main Profile. For example, if the value of the main_profile_compatibility_flag is equal to a first value (e.g., 1), this may indicate that the bitstream conforms to the Main Profile, and if the value of the main_profile_compatibility_flag is equal to a second value (e.g., 0), this may indicate that the bitstream conforms to a profile other than the Main Profile.

[0232] The unique_point_positions_constraint_flag indicates whether all output points can have unique positions in each point cloud frame referenced by the current SPS. For example, if the value of unique_point_positions_constraint_flag is equal to a first value (e.g., 1), this indicates that all output points can have unique positions in each point cloud frame referenced by the current SPS. If the value of unique_point_positions_constraint_flag is equal to a second value (e.g., 0), this indicates that two or more output points can have the same position in any point cloud frame referenced by the current SPS. Also, even if all points are unique in each slice, points from different slices within a frame can overlap. In this case, the value of unique_point_positions_constraint_flag can be set to 0.

[0233] The level_idc may indicate the level to which the bitstream conforms. The sps_seq_parameter_set_id may indicate an identifier for the SPS referenced by other syntax elements.

[0234] sps_bounding_box_present_flag may indicate whether a bounding box is present in the SPS. For example, if the value of sps_bounding_box_present_flag is equal to a first value (e.g., 1), this may indicate that a bounding box is present in the SPS, and if the value of sps_bounding_box_present_flag is equal to a second value (e.g., 0), this may indicate that the size of the bounding box is undefined. If the value of sps_bounding_box_present_flag is equal to the first value (e.g., 1), sps_bounding_box_offset_x, sps_bounding_box_offset_y, sps_bounding_box_offset_z, sps_bounding_box_offset_log2_scale, sps_bounding_box_size_width, sps_bounding_box_size_height, and / or sps_bounding_box_size_depth may be further signaled.

[0235] sps_bounding_box_offset_x may indicate the quantized x-offset of the source bounding box in the Cartesian coordinate system; if there is no x-offset of the source bounding box, the value of sps_bounding_box_offset_x can be inferred to be 0. sps_bounding_box_offset_y may indicate the quantized y-offset of the source bounding box in the Cartesian coordinate system; if there is no y-offset of the source bounding box, the value of sps_bounding_box_offset_y can be inferred to be 0. sps_bounding_box_offset_z may indicate the quantized z-offset of the source bounding box in the Cartesian coordinate system; if there is no z-offset of the source bounding box, the value of sps_bounding_box_offset_z can be inferred to be 0. sps_bounding_box_offset_log2_scale may indicate the scale factors for scaling the quantized x, y, and z source bounding box offsets. sps_bounding_box_size_width can indicate the width (or horizontal dimension) of the source bounding box in a Cartesian coordinate system; if there is no width of the source bounding box, the value of sps_bounding_box_size_width can be inferred to be 1. sps_bounding_box_size_height can indicate the height of the source bounding box in a Cartesian coordinate system; if there is no height of the source bounding box, the value of sps_bounding_box_size_height can be inferred to be 1. sps_bounding_box_size_depth can indicate the depth of the source bounding box in a Cartesian coordinate system; if there is no depth of the source bounding box, the value of sps_bounding_box_size_depth can be inferred to be 1.

[0236] sps_source_scale_factor_numerator_minus1 plus 1 can indicate the scale factor numerator of the source point cloud. sps_source_scale_factor_denominator_minus1 plus 1 can indicate the scale factor denominator of the source point cloud. sps_num_attribute_sets can indicate the number of coded attributes in the bitstream. sps_num_attribute_sets must have a value between 0 and 63.

[0237] The attribute_dimension_minus1[i] and attribute_instance_id[i] may be further signaled for the number of coded attributes in the bitstream indicated by sps_num_attribute_sets. i may increase by 1 from 0 to the number of coded attributes in the bitstream minus 1. The value of attribute_dimension_minus1[i] plus 1 may indicate the number of components of the i-th attribute, and attribute_instance_id may indicate the instance identifier of the i-th attribute.

[0238] attribute_bitdepth_minus1[i], attribute_secondary_bitdepth_minus1[i], attribute_cicp_colour_primaries[i], attribute_cicp_transfer_characteristics[i], attribute_cicp_matrix_coeffs[i], and / or attribute_cicp_video_full_range_flag[i] may be further signaled if the value of attribute_dimension_minus1[i] is greater than 1. The value of attribute_bitdepth_minus1[i] plus 1 may indicate the bit depth for the first component (or the first component) of the i-th attribute signal. The value of attribute_secondary_bitdepth_minus1[i] plus 1 may indicate the bit depth for the second component (or second component) of the i-th attribute signal, and attribute_cicp_colour_primaries[i] may indicate the chromaticity coordinates of the color attribute source primaries of the i-th attribute. Attribute_cicp_transfer_characteristics[i] may indicate the source input linear optical intensity having a nominal real-valued range between 0 and 1 for the i-th attribute, which is the reference opto-electronic transfer characteristic function, or may indicate the inverse of the reference opto-electronic transfer characteristic function, which is a function of the output linear optical intensity.attribute_cicp_matrix_coeffs[i] may describe the matrix coefficients used to derive luma and chroma signals from the green, blue, and red (or Y, Z, X primaries) of the i-th attribute. attribute_cicp_video_full_range_flag[i] may indicate the range of the black level and luma and chroma signals derived from the E'Y, E'PB, and E'PR or E'R, E'G, and E'B actual value component signals of the i-th attribute. known_attribute_label_flag[i] may indicate whether know_attribute_label[i] or attribute_label_four_bytes[i] is signaled for the i-th attribute. For example, if the value of known_attribute_label_flag[i] is equal to a first value (e.g., 1), this may indicate that known_attribute_label is signaled for the i-th attribute, and if the value of known_attribute_label_flag[i] is equal to a second value (e.g., 0), this may indicate that attribute_label_four_bytes[i] is signaled for the i-th attribute. known_attribute_label[i] may indicate the type of the i-th attribute. For example, if the value of known_attribute_label[i] is equal to a first value (e.g., 0), this can indicate that the attribute is color; if the value of known_attribute_label[i] is equal to a second value (e.g., 1), this can indicate that the attribute is reflectance; and if the value of known_attribute_label[i] is equal to a third value (e.g., 2), this can indicate that the attribute is a frame index.The attribute_label_four_bytes[i] can indicate a known attribute type using a four-byte code, as shown in Figure 22a.

[0239] log2_max_frame_idx may indicate the number of bits used to signal the frame_idx syntax variable. For example, adding 1 to the value of log2_max_frame_idx may indicate the number of bits used to signal the frame_idx syntax variable. As shown in Figure 22b, axis_coding_order may indicate the correspondence between the X, Y, and Z output axis labels and the three position components in the reconstructed point cloud RecPic[pointidx][axis] with axis=0,...,2.

[0240] The sps_bypass_stream_enabled_flag may indicate whether the bypass coding mode is used to read the bitstream. For example, if the value of the sps_bypass_stream_enabled_flag is equal to a first value (e.g., 1), this may indicate that the bypass coding mode is used to read the bitstream. As another example, if the value of the sps_bypass_stream_enabled_flag is equal to a second value (e.g., 0), this may indicate that the bypass coding mode is not used to read the bitstream. The sps_extension_flag may indicate whether the sps_extension_data syntax element is present in the SPS syntax structure. For example, if the value of sps_extension_flag is equal to a first value (e.g., 1), this may indicate that the sps_extension_data syntax element is present in the SPS syntax structure, and if the value of sps_extension_flag is equal to a second value (e.g., 0), this may indicate that the sps_extension_data syntax element is not present in the SPS syntax structure. sps_extension_flag may also have to be equal to 0 in the bitstream. When the value of sps_extension_flag is equal to the first value (e.g., 1), sps_extension_data_flag may be further signaled. sps_extension_data_flag may have any value, and the presence and value of sps_extension_data_flag may not affect decoder conformance to the profile.

[0241] GPS Syntax Structure

[0242] Figure 23 shows an example of a syntax structure of GPS. In Figure 23, the syntax elements (or fields) shown in the syntax structure of GPS may be syntax elements included in GPS or may be syntax elements signaled via GPS.

[0243] The gps_geom_parameter_set_id may indicate an identifier of a GPS referenced by other syntax elements, and the gps_seq_parameter_set_id may indicate the value of the seq_parameter_set_id for the active SPS. The gps_box_present_flag may indicate whether additional bounding box information is provided in the geometry slice header referencing the current GPS. For example, if the value of the gps_box_present_flag is equal to a first value (e.g., 1), this may indicate that additional bounding box information is provided in the geometry slice header referencing the current GPS. If the value of the gps_box_present_flag is equal to a second value (e.g., 0), this may indicate that additional bounding box information is not provided in the geometry slice header referencing the current GPS. If the value of the gps_box_present_flag is equal to a first value (e.g., 1), the gps_gsh_box_log2_scale_present_flag may be further signaled. The gps_gsh_box_log2_scale_present_flag may indicate whether the gps_gsh_box_log2_scale is signaled in each geometry slice header that currently references a GPS. For example, if the value of the gps_gsh_box_log2_scale_present_flag is equal to a first value (e.g., 1), this may indicate that the gps_gsh_box_log2_scale is signaled in each geometry slice header that currently references a GPS. As another example, if the value of the gps_gsh_box_log2_scale_present_flag is equal to a second value (e.g., 0), this may indicate that the gps_gsh_box_log2_scale is not signaled in each geometry slice header that currently references a GPS, and that a common scale for all slices is signaled in the gps_gsh_box_log2_scale of the current GPS.If the value of gps_gsh_box_log2_scale_present_flag is equal to a second value (e.g., 0), gps_gsh_box_log2_scale may be further signaled. gps_gsh_box_log2_scale may indicate a common scale factor of the bounding box origin for all slices that currently reference the GPS.

[0244] The unique_geometry_points_flag may indicate whether all output points have unique positions within a slice in all slices that currently refer to the GPS. For example, if the value of the unique_geometry_points_flag is the same as the first value (e.g., 1), this may indicate that all output points have unique positions within a slice in all slices that currently refer to the GPS. If the value of the unique_geometry_points_flag is the same as the second value (e.g., 0), this may indicate that two or more output points may have the same position within a slice in all slices that currently refer to the GPS.

[0245] The geometry_planar_mode_flag may indicate whether the planar coding mode is activated. For example, if the value of the geometry_planar_mode_flag is equal to a first value (e.g., 1), this may indicate that the planar coding mode is activated, and if the value of the geometry_planar_mode_flag is equal to a second value (e.g., 0), this may indicate that the planar coding mode is not activated. When the value of the geometry_planar_mode_flag is equal to the first value (e.g., 1), the geom_planar_mode_th_idcm, geom_planar_mode_th[1], and / or geom_planar_mode_th[2] may be further signaled. The geom_planar_mode_th_idcm may indicate an activation threshold for the direct coding mode. The value of the geom_planar_mode_th_idcm may be an integer ranging from 0 to 127. geom_planar_mode_th[i] may indicate the activation threshold for the planar coding mode along the ith direction where the planar coding mode is most likely to be efficient, for i ranging from 0 to 2. geom_planar_mode_th[i] may be an integer in the range from 0 to 127.

[0246] The geometry_angular_mode_flag may indicate whether angular coding mode is activated. For example, if the value of the geometry_angular_mode_flag is equal to a first value (e.g., 1), this may indicate that angular coding mode is activated, and if the value of the geometry_angular_mode_flag is equal to a second value (e.g., 0), this may indicate that angular coding mode is not activated. If the value of the geometry_angular_mode_flag is equal to the first value (e.g., 1), lidar_head_position[0], lidar_head_position[1], lidar_head_position[2], number_lasers, planar_buffer_disabled, implicit_qtbt_angular_max_node_min_dim_log2_to_split_z, and / or implicit_qtbt_angular_max_diff_to_split_z may be further signaled.

[0247] lidar_head_position[0], lidar_head_position[1], and / or lidar_head_position[2] may indicate the (X, Y, Z) coordinates of the lidar head in a coordinate system with internal axes. number_lasers may indicate the number of lasers used for angular coding mode. laser_angle[i] and laser_correction[i] may each be signaled the number indicated by number_lasers, where i may increase in increments of 1 from 0 to 'number_lasers value - 1'. laser_angle[i] may indicate the tangent of the elevation angle of the ith laser relative to the horizontal plane defined by the 0th and 1st internal axes, and laser_correction[i] may indicate the correction of the ith laser position relative to lidar_head_position[2] along the second internal axis. planar_buffer_disabled may indicate whether buffer-based closest node tracking is used in the process of coding the planar mode flag and planar position in planar mode. For example, if the value of planar_buffer_disabled is equal to the first value (e.g., 1), this may indicate that buffer-based closest node tracking is not used in the process of coding the planar mode flag and planar position in planar mode. If the value of planar_buffer_disabled is equal to the second value (e.g., 0), this may indicate that buffer-based closest node tracking is used in the process of coding the planar mode flag and planar position in planar mode. If planar_buffer_disabled is not present, the value of planar_buffer_disabled may be inferred to be the second value (e.g., 0).implicit_qtbt_angular_max_node_min_dim_log2_to_split_z can indicate the log2 value of the node size at which a horizontal split of a node is more preferred than a vertical split. implicit_qtbt_angular_max_diff_to_split_z can indicate the log2 value of the vertical to horizontal node size ratio allowed for a node. If implicit_qtbt_angular_max_diff_to_split_z is not present, implicit_qtbt_angular_max_diff_to_split_z can be inferred to be 0.

[0248] The neighbor_context_restriction_flag may indicate whether the geometry node occupancy of the current node is coded in a context determined from a neighbor node located inside the parent node of the current node. For example, if the value of the neighbor_context_restriction_flag is equal to a first value (e.g., 0), this may indicate that the geometry node occupancy of the current node is coded in a context determined from a neighbor node located inside the parent node of the current node. If the value of the neighbor_context_restriction_flag is equal to a second value (e.g., 1), this may indicate that the geometry node occupancy of the current node is not coded in a context determined from a neighbor node located inside the parent node of the current node. The inferred_direct_coding_mode_enabled_flag may indicate whether the direct_mode_flag is present in the geometry node syntax. For example, if the value of inferred_direct_coding_mode_enabled_flag is equal to a first value (e.g., 1), this can indicate that direct_mode_flag is present in the geometry node syntax, and if the value of inferred_direct_coding_mode_enabled_flag is equal to a second value (e.g., 0), this can indicate that direct_mode_flag is not present in the geometry node syntax.

[0249] The bitwise_occupancy_coding_flag may indicate whether the geometry node occupancy is encoded using the bitwise contextualization of the syntax element occupancy_map. For example, if the value of the bitwise_occupancy_coding_flag is equal to a first value (e.g., 1), this may indicate that the geometry node occupancy is encoded using the bitwise contextualization of the syntax element occupancy_map. If the value of the bitwise_occupancy_coding_flag is equal to a second value (e.g., 0), this may indicate that the geometry node occupancy is encoded using the directory-encoded syntax element occupancy_map. The adjacent_child_contextualization_enabled_flag may indicate whether adjacent children of neighboring octree nodes are used for bitwise occupancy contextualization. For example, if the value of adjacent_child_contextualization_enabled_flag is equal to a first value (e.g., 1), this may indicate that adjacent children of adjacent octree nodes are used for bitwise occupancy contextualization, and if the value of adjacent_child_contextualization_enabled_flag is equal to a second value (e.g., 0), this may indicate that children of adjacent octree nodes are not used for bitwise occupancy contextualization.

[0250] log2_neighbour_avail_boundary can indicate the value of the variable NeighbAvailBoundary used in the decoding process. For example, if the value of neighbour_context_restriction_flag is the same as the first value (e.g., 1), NeighbAvailabilityMask can be set to 1, and if the value of neighbour_context_restriction_flag is the same as the second value (e.g., 0), NeighbAvailabilityMask can be set to 1 << log2_neighbour_avail_boundary. log2_intra_pred_max_node_size can indicate the octree node size eligible for occupancy intra prediction. log2_trisoup_node_size can specify the variable TrisoupNodeSize to the size of a triangular node.

[0251] The geom_scaling_enabled_flag may indicate whether a scaling process for the geometry position is applied during the geometry slice decoding process. For example, if the value of the geom_scaling_enabled_flag is equal to a first value (e.g., 1), this may indicate that a scaling process for the geometry position is performed during the geometry slice decoding process. If the value of the geom_scaling_enabled_flag is equal to a second value (e.g., 0), this may indicate that scaling is not required for the geometry position. When the value of the geom_scaling_enabled_flag is equal to the first value (e.g., 1), the geom_base_qp may be further signaled. The geom_base_qp may indicate the base value of the geometry position quantization parameter. The gps_implicit_geom_partition_flag may indicate whether implicit geometry partitioning is enabled for the sequence or slice. For example, if the value of gps_implicit_geom_partition_flag is equal to a first value (e.g., 1), this may indicate that implicit geometry partitioning is enabled for the sequence or slice, and if the value of gps_implicit_geom_partition_flag is equal to a second value (e.g., 0), this may indicate that implicit geometry partitioning is disabled for the sequence or slice. When the value of gps_implicit_geom_partition_flag is equal to the first value (e.g., 1), gps_max_num_implicit_qtbt_before_ot and gps_min_size_implicit_qtbt may be signaled.gps_max_num_implicit_qtbt_before_ot can indicate the maximum number of implicit QT and BT partitions before an OT partition. gps_min_size_implicit_qtbt can indicate the minimum size of implicit QT and BT partitions.

[0252] The gps_extension_flag may indicate whether the gps_extension_data syntax element is present in the GPS syntax structure. For example, if the value of the gps_extension_flag is equal to a first value (e.g., 1), this may indicate that the gps_extension_data syntax element is present in the GPS syntax structure, and if the value of the gps_extension_flag is equal to a second value (e.g., 0), this may indicate that the gps_extension_data syntax element is not present in the GPS syntax structure. If the value of the gps_extension_flag is equal to the first value (e.g., 1), the gps_extension_data_flag may be further signaled. The gps_extension_data_flag may have any value, and the presence and value of the gps_extension_data_flag may not affect decoder conformance to the profile.

[0253] APS Syntax Structure

[0254] Figure 24 shows an example of an APS syntax structure. In Figure 24, the syntax elements (or fields) shown in the APS syntax structure may be syntax elements included in the APS or may be syntax elements signaled via the APS.

[0255] aps_attr_parameter_set_id may provide an identifier of the APS for reference by other syntax elements, and aps_seq_parameter_set_id may indicate the value of sps_seq_parameter_set_id for an active SPS. attr_coding_type may indicate the coding type for the attribute. A table of attr_coding_type values ​​and their respective assigned attribute coding types is shown in FIG. 25. As shown in FIG. 25, if the value of attr_coding_type is a first value (e.g., 0), the coding type indicates conducting weight lifting; if the value of attr_coding_type is a second value (e.g., 1), the coding type indicates RAHT; and if the value of attr_coding_type is a third value (e.g., 2), the coding type indicates fixed weight lifting.

[0256] aps_attr_initial_qp may indicate the initial value of the variable SliceQp for each slice that references an APS. The value of aps_attr_initial_qp may be in the range of 4 to 51. aps_attr_chroma_qp_offset may indicate the offset for the initial quantization parameter signaled by aps_attr_initial_qp. aps_slice_qp_delta_present_flag may indicate whether the ash_attr_qp_delta_luma syntax element and the ash_attr_qp_delta_chroma syntax element are present in the attribute slice header (ASH). For example, if the value of aps_slice_qp_delta_present_flag is equal to a first value (e.g., 1), this can indicate that ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma are present in the attribute slice header (ASH), and if the value of aps_slice_qp_delta_present_flag is equal to a second value (e.g., 0), this can indicate that ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma are not present in the attribute slice header.

[0257] If the value of attr_coding_type is the first value (e.g., 0) or the third value (e.g., 2), i.e., if the coding type is prediction weight lifting or fixed weight lifting, lifting_num_pred_nearest_neighbours_minus1, lifting_search_range_minus1, and lifting_neighbour_bias[k] may be further signaled. The value of lifting_num_pred_nearest_neighbours_minus1 plus 1 may indicate the maximum number of nearest neighbors used for prediction. The value of the variable NumPredNearestNeighbours may be set equal to lifting_num_pred_nearest_neighbours (lifting_num_pred_nearest_neighbours_minus1 plus 1). The value of lifting_search_range_minus1 plus 1 can indicate the search range used for "determining nearest neighbors used for prediction" and "building distance-based levels of detail (LOD)". The variable LiftingSearchRange, which specifies the search range, can be calculated by adding 1 to the value of the lifting_search_range_minus1 field (LiftingSearchRange = lifting_search_range_minus1 + 1). lifting_neighbour_bias[k] is part of the nearest neighbor guidance process and can indicate the bias used to weight the kth component in the calculation of the Euclidean distance between two points.

[0258] If the value of attr_coding_type is a third value (e.g., 2), that is, if the coding type indicates fixed-weight lifting, lifting_scalability_enabled_flag can be further signaled. lifting_scalability_enabled_flag can indicate whether the attribute decoding process accepts pruned octree decoding results for the input geometry points. For example, if the value of lifting_scalability_enabled_flag is equal to a first value (e.g., 1), this can indicate that the attribute decoding process accepts pruned octree decoding results for the input geometry points. If the value of lifting_scalability_enabled_flag is equal to a second value (e.g., 0), this can indicate that the attribute decoding process requires full octree decoding results for the input geometry points. If the value of lifting_scalability_enabled_flag is not the first value (e.g., 1), lifting_num_detail_level_minus1 may be further signaled. lifting_num_detail_level_minus1 may indicate the number of LODs for attribute coding. A variable LevelDetailCount for specifying the number of LODs may be derived by adding 1 to the value of lifting_num_detail_level_minus1 (LevelDetailCount = lifting_num_detail_level_minus1 + 1).

[0259] If the value of lifting_num_detail_level_minus1 is greater than 1, lifting_lod_regular_sampling_enabled_flag may be further signaled. lifting_lod_regular_sampling_enabled_flag may indicate whether the LOD is built using the regular sampling strategy. For example, if the value of lifting_lod_regular_sampling_enabled_flag is equal to a first value (e.g., 1), this may indicate that the LOD is built using the regular sampling strategy, and if the value of lifting_lod_regular_sampling_enabled_flag is equal to a second value (e.g., 0), this may indicate that the distance-based sampling strategy is used instead.

[0260] If the value of lifting_scalability_enabled_flag is not the first value (e.g., 1), lifting_sampling_period_minus2[idx] or lifting_sampling_distance_squared_scale_minus1[idx] may be further signaled depending on whether the LOD is generated by the regular sampling strategy (value of lifting_lod_regular_sampling_enabled_flag). For example, if the value of lifting_lod_regular_sampling_enabled_flag is equal to the first value (e.g., 1), lifting_sampling_period_minus2[idx] may be signaled, and if the value of lifting_lod_regular_sampling_enabled_flag is equal to the second value (e.g., 0), lifting_sampling_distance_squared_scale_minus1[idx] may be signaled. If the value of idx is not 0 (idx!=0), lifting_sampling_distance_squared_offset[idx] may be further signaled. idx may increase from 0 in increments of 1 up to num_detail_level_minus1 minus 1. The value of lifting_sampling_period_minus2[idx] plus 2 may indicate the sampling period for LODidx. The value of lifting_sampling_distance_squared_scale_minus1[idx] plus 1 may indicate the scaling factor for the derivation of the squared sampling distance for LODidx. lifting_sampling_distance_squared_offset[idx] may indicate the offset for the derivation of the squared sampling distance for LODidx.

[0261] If the value of attr_coding_type is the same as the first value (e.g., 0), i.e., if the coding type is prediction weight lifting, lifting_adaptive_prediction_threshold, lifting_intra_lod_prediction_num_layers, lifting_max_num_direct_predictors, and inter_component_prediction_enabled_flag may be signaled. lifting_adaptive_prediction_threshold may indicate a threshold for enabling adaptive prediction. To switch the adaptive predictor selection mode, a variable AdaptivePredictionThreshold specifying the threshold may be set to the same value as lifting_adaptive_prediction_threshold. lifting_intra_lod_prediction_num_layers may indicate the number of LOD layers that a decoded point in the same LOD layer can refer to to generate a predicted value of a target point. For example, if the value of lifting_intra_lod_prediction_num_layers is the same as the value of LevelDetailCount, this may indicate that the target point can refer to decoded points in the same LOD layer for all LOD layers. If the value of lifting_intra_lod_prediction_num_layers is 0, this may indicate that the target point cannot refer to decoded points in the same LOD layer for any LOD layer. The value of lifting_intra_lod_prediction_num_layers may range from 0 to LevelDetailCount. lifting_max_num_direct_predictors may indicate the maximum number of predictors that can be used for direct prediction.The inter_component_prediction_enabled_flag may indicate whether the primary component of a multi-component attribute is used to predict the reconstructed value of a non-primary component. For example, if the value of the inter_component_prediction_enabled_flag is equal to a first value (e.g., 1), this may indicate that the primary component of a multi-component attribute is used to predict the reconstructed value of a non-primary component. As another example, if the value of the inter_component_prediction_enabled_flag is equal to a second value (e.g., 0), this may indicate that all attribute components are independently reconstructed.

[0262] If the value of attr_coding_type is equal to a second value (e.g., 1), that is, if the attribute coding type is RAHT, raht_prediction_enabled_flag may be signaled. raht_prediction_enabled_flag may indicate whether transform weight prediction from adjacent points is enabled in the RAHT decoding process. For example, if the value of raht_prediction_enabled_flag is equal to a first value (e.g., 1), this may indicate that transform weight prediction from adjacent points is enabled in the RAHT decoding process, and if the value of raht_prediction_enabled_flag is equal to a second value (e.g., 0), this may indicate that transform weight prediction from adjacent points is disabled in the RAHT decoding process. When the value of raht_prediction_enabled_flag is equal to the first value (e.g., 1), raht_prediction_threshold0 and raht_prediction_threshold1 may be further signaled. raht_prediction_threshold0 may indicate a threshold for terminating transform weighted prediction from adjacent points. raht_prediction_threshold1 may indicate a threshold for skipping transform weighted prediction from adjacent points.

[0263] aps_extension_flag may indicate whether the aps_extension_data_flag syntax element is present in the APS syntax structure. For example, if the value of aps_extension_flag is equal to a first value (e.g., 1), this may indicate that the aps_extension_data_flag syntax element is present in the APS syntax structure, and if the value of aps_extension_flag is equal to a second value (e.g., 0), this may indicate that the aps_extension_data_flag syntax element is not present in the APS syntax structure. When the value of aps_extension_flag is equal to the first value (e.g., 1), aps_extension_data_flag may be signaled. aps_extension_data_flag may have any value, and the presence and value of aps_extension_data_flag may not affect decoder conformance to the profile.

[0264] Tile Inventory Syntax Structure

[0265] Figure 26 shows an example of a syntax structure of a tile inventory. The tile inventory is also called a tile parameter set (TPS). In Figure 26, the syntax elements (or fields) shown in the syntax structure of the TPS may be syntax elements included in the TPS or may be syntax elements signaled via the TPS.

[0266] The tile_frame_idx may include an identification number that can be used to identify the purpose of the tile inventory. The tile_seq_parameter_set_id may indicate the value of sps_seq_parameter_set_id for the active SPS. The tile_id_present_flag may indicate a parameter for identifying a tile. For example, if the value of the tile_id_present_flag is equal to a first value (e.g., 1), this may indicate that the tile is identified by the value of the tile_id syntax element, and if the value of the tile_id_present_flag is equal to a second value (e.g., 0), this may indicate that the tile is identified by its position in the tile inventory. The tile_cnt may indicate the number of tile bounding boxes present in the tile inventory. The tile_bounding_box_bits may indicate the bit depth for representing the bounding box information for the tile inventory. The tile_id, tile_bounding_box_offset_xyz[tile_id][k], and tile_bounding_box_size_xyz[tile_id][k] may be signaled while the loop variable tileIdx increases by 1 from 0 to (the number of tile bounding boxes minus 1). The tile_id may identify a specific tile within the tile_inventory. The tile_id may be signaled when the value of tile_id_present_flag is the first value (e.g., 1), and may be signaled as many times as the number of tile bounding boxes. If the tile_id does not exist (if not signaled), the value of tile_id may be inferred as the index of the tile within the tile inventory provided by the loop variable tileIdx. It may be a bitstream compatibility requirement that all values ​​of tile_id must be unique within the tile inventory.tile_bounding_box_offset_xyz[tileId][k] and tile_bounding_box_size_xyz[tileId][k], where tile_bounding_box_offset_xyz[tileId][k] can indicate a bounding box that includes a slice identified by the same gsh_tile_id as tileId. tile_bounding_box_offset_xyz[tileId][k] can be the k-th component of the (x, y, z) origin coordinates of the tile bounding box relative to TileOrigin[k]. tile_bounding_box_size_xyz[tileId][k] can be the k-th component of the width, height, and depth of the tile bounding box.

[0267] tile_origin_xyz[k] can be signaled while variable k increases by 1 from 0 to 2. tile_origin_xyz[k] can indicate the k-th component of the tile origin in cartesian coordinates. The value of tile_origin_xyz[k] can be forced in the same way as sps_bounding_box_offset[k]. tile_origin_log2_scale can indicate a scale factor for scaling the components of tile_origin_xyz. The value of tile_origin_log2_scale can be forced in the same way as sps_bounding_box_offset_log2_scale. For k = 0, ···, 2, an array TileOrigin with elements TileOrigin[k] can be derived as 'TileOrigin[k]=tile_origin_xyz[k]<<tile_origin_log2_scale'.

[0268] Geometry Slice Syntax Structure

[0269] 27 and 28 show an example of a syntax structure of a geometry slice. The syntax elements (or fields) shown in Fig. 27 and 28 may be syntax elements included in the geometry slice or may be syntax elements signaled via the geometry slice.

[0270] A bitstream transmitted from a transmitting device to a receiving device may include one or more slices. Each slice may include a geometry slice and an attribute slice. Here, a geometry slice may be a geometry slice bitstream, and an attribute slice may be an attribute slice bitstream. A geometry slice may include a geometry slice header (GSH), and an attribute slice may include an attribute slice header (ASH).

[0271] As shown in Figure 27a, the geometry slice bitstream (geometry_slice_bitstream()) can include a geometry slice header (geometry_slice_header()) and geometry slice data (geometry_slice_data()). The geometry slice data can include geometry or geometry-related data associated with a portion or the entire point cloud. As shown in Figure 27(b), the syntax elements signaled via the geometry slice header are as follows:

[0272] gsh_geometry_parameter_set_id may indicate the value of gps_geom_parameter_set_id for the active GPS. gsh-tile_id may indicate the identifier of the tile referenced by the geometry slice header (GSH). gsh_slice_id may indicate the identifier of the slice for reference by other syntax elements. frame_idx may indicate log2_max_frame_idx + 1 least significant bits of the notional frame number counter. Consecutive slices with different frame_idx values ​​may form part of different output point cloud frames. Consecutive slices with no intermediate frame boundary marker data units and the same frame_idx value may form part of the same output point cloud frame. gsh_num_point may indicate the maximum number of coded points in the slice. It may be a bitstream compatibility requirement that the value of gsh_num_points must be equal to or greater than the number of decoded points in the slice.

[0273] If the value of gps_box_present_flag is equal to a first value (e.g., 1), gsh_box_log2_scale, gsh_box_origin_x, gsh_box_origin_y, and gsh_box_origin_z may be signaled. According to an embodiment, gsh_box_log2_scale may also be signaled if the value of gps_box_present_flag is equal to a first value (e.g., 1) and the value of gps_gsh_box_log2_scale_present_flag is equal to the first value (e.g., 1). gsh_box_log2_scale may indicate the scale factor of the bounding box origin for the slice. gsh_box_origin_x may indicate the x value of the bounding box origin scaled by the value of gsh_box_log2_scale, gsh_box_origin_y may indicate the y value of the bounding box origin scaled by the value of gsh_box_log2_scale, and gsh_box_origin_z may indicate the z value of the bounding box origin scaled by the value of gsh_box_log2_scale. The variables slice_origin_x, slice_origin_y, and / or slice_origin_z may be derived as follows:

[0274] If the value of gps_gsh_box_log2_scale_present_flag is equal to a second value (e.g., 0), then originScale is set equal to gsh_box_log2_scale. Otherwise, if the value of gps_gsh_box_log2_scale_present_flag is equal to a first value (e.g., 1), then originScale is set equal to gps_gsh_box_log2_scale. If the value of gps_box_present_flag is equal to a second value (e.g., 0), then the values ​​of the variables slice_origin_x, slice_origin_y, and slice_origin_z can be inferred to be 0. Otherwise, if the value of gps_box_present_flag is equal to the first value (e.g., 1), then the following formulas can be applied to the variables slice_origin_x, slice_origin_y, and slice_origin_z:

[0275] slice_origin_x=gsh_box_origin_x< <originScale

[0276] slice_origin_y=gsh_box_origin y< <originScale

[0277] slice_origin_z=gsh_box_origin_z< <originScale

[0278] If the value of gps_implicit_geom_partition_flag is equal to a first value (e.g., 1), gsh_log2_max_nodesize_x, gsh_log2_max_nodesize_y_minus_x, and gsh_log2_max_nodesize_z_minus_y may be further signaled. If the value of gps_implicit_geom_partition_flag is equal to a second value (e.g., 0), gps_log2_max_nodesize may be signaled.

[0279] gsh_log2_max_nodesize_x can represent the bounding box size in the x dimension, i.e., MaxNodesizeXLog2 used in the decoding process, as follows:

[0280] MaxNodeSizeXLog2=gsh_log2_max_nodesize_x

[0281] MaxNodeSizeX=1< <MaxNodeSizeXLog2

[0282] gsh_log2_max_nodesize_y_minus_x indicates the bounding box size in the y dimension, i.e., MaxNodesizeYLog2 used in the decoding process, as follows:

[0283] MaxNodeSizeYLog2=gsh_log2_max_nodesize_y_minus_x+MaxNodeSizeXLog2

[0284] MaxNodeSizeY=1< <MaxNodeSizeYLog2

[0285] gsh_log2_max_nodesize_z_minus_y indicates the bounding box size in the z dimension, i.e., MaxNodesizeZLog2 used in the decoding process, as follows:

[0286] MaxNodeSizeZLog2=gsh_log2_max_nodesize_z_minus_y+MaxNodeSizeYLog2

[0287] MaxNodeSizeZ=1< <MaxNodeSizeZLog2

[0288] gsh_log2_max_nodesize may indicate the size of the root geometry octree node if the value of gps_implicit_geom_partition_flag is the first value (e.g., 1). The variables MaxNodeSize and MaxGeometryOctreeDepth may be derived as follows:

[0289] MaxNodeSize=1< <gsh_log2_max_nodesize

[0290] MaxGeometryOctreeDepth=gsh_log2_max_nodesize log2_trisoup_node_size

[0291] When the value of geom_scaling_enabled_flag is equal to the first value (e.g., 1), geom_slice_qp_offset and geom_octre_qp_offsets_enabled_flag can be signaled. geom_slice_qp_offset can indicate an offset to the base geometry quantization parameter (geom_base_qp). geom_octre_qp_offsets_enabled_flag can indicate whether geom_node_qp_offset_eq0_flag is present in the geometry node syntax. For example, if the value of geom_octree_qp_offsets_enabled_flag is equal to a first value (e.g., 1), this may indicate that geom_node_qp_offset_eq0_flag is present in the geometry node syntax, and if the value of geom_octree_qp_offsets_enabled_flag is equal to a second value (e.g., 0), this may indicate that geom_node_qp_offset_eq0_flag is not present in the geometry node syntax. When the value of geom_octree_qp_offsets_enabled_flag is equal to the first value (e.g., 1), geom_octree_qp_offsets_depth may be signaled. geom_octree_qp_offsets_depth may indicate the depth of the geometry octree when geom_node_qp_offset_eq0_flag is present in the geometry node syntax.

[0292] As shown in FIG. 28, the syntax elements signaled via the geometry slice data are as follows: The geometry slice data may include a repeat statement (first repeat statement) repeated the number of times equal to the value of MaxGeometryOctreeDepth. MaxGeometryOctreeDepth may indicate the maximum depth of the geometry octree. In the first repeat statement, depth may increase by 1 from 0 to (MaxGeometryOctreeDepth-1). The first repeat statement may include a repeat statement (second repeat statement) repeated the number of times equal to the value of NumNodesAtDepth. NumNodesAtDepth[depth] may indicate the number of nodes to be decoded at the corresponding depth. In the second repeat statement, nodeidx may increase by 1 from 0 to (NumNodesAtDepth-). Through the first and second iterations, xN = NodeX[depth][nodeIdx], yN = NodeY[depth][nodeIdx], zN = NodeZ[depth][nodeIdx], and geometry_node(depth, nodeIdx, xN, yN, zN) can be signaled. The variables NodeX[depth][nodeIdx], NodeY[depth][nodeIdx], and NodeZ[depth][nodeIdx] can indicate the x, y, and z coordinates of the Idx-th node in decoding order at a given depth. The geometry bitstream for the corresponding depth can be transmitted via geometry_node(depth, nodeIdx, xN, yN, zN).

[0293] If the value of log2_trisoup_node_size is greater than 0, geometry_trisoup_data() can be further signaled. That is, if the size of the triangle node is greater than 0, the trisoup geometry encoded geometry bitstream can be signaled via geometry_trisoup_data().

[0294] Attribute Slice Syntax Structure

[0295] 29 and 30 show an example of the syntax structure of an attribute slice. The syntax elements (or fields) shown in Figures 29 and 30 may be syntax elements included in the attribute slice or may be syntax elements signaled via the attribute slice.

[0296] As shown in Figure 29a, the attribute slice bitstream (attribute_slice_bitstream()) can include an attribute slice header (attribute_slice_header()) and attribute slice data (attribute_slice_data()). The attribute slice data (attribute_slice_data()) can include attributes or attribute-related data associated with a portion or the entire point cloud. As shown in Figure 29b, the syntax elements signaled via the attribute slice header are as follows:

[0297] ash_attr_parameter_set_id can indicate the value of aps_attr_parameter_set_id in the active APS. ash_attr_sps_attr_idx can indicate an attribute set in the active SPS. ash_attr_geom_slice_id can indicate the value of gsh_slice_id in the active geometry slice header. aps_slice_qp_delta_present_flag can indicate whether the ash_attr_layer_qp_delta_luma and ash_attr_layer_qp_delta_chroma syntax elements are currently present in ASH. For example, when the value of aps_slice_qp_delta_present_flag is equal to a first value (e.g., 1), this may indicate that ash_attr_layer_qp_delta_luma and ash_attr_layer_qp_delta_chroma are currently present in ASH, and when the value of aps_slice_qp_delta_present_flag is equal to a second value (e.g., 0), this may indicate that ash_attr_layer_qp_delta_luma and ash_attr_layer_qp_delta_chroma are not currently present in ASH. When the value of aps_slice_qp_delta_present_flag is equal to a first value (e.g., 1), ash_attr_qp_delta_luma may be signaled. ash_attr_qp_delta_luma may indicate the luma delta quantization parameter (qp) from the initial slice qp in the active attribute parameter set. ash_attr_qp_delta_chroma may be signaled if the value of attribute_dimension_minus1[ash_attr_sps_attr_idx] is greater than 0. ash_attr_qp_delta_chroma may indicate the chroma delta quantization parameter (qp) from the initial slice qp in the active attribute parameter set.The variables InitialSliceQpY and InitialSliceQpC can be derived as follows.

[0298] InitialSliceQpY=aps_attrattr_initial_qp+ash_attr_qp_delta_luma

[0299] InitialSliceQpC=aps_attrtr_initial_qp+aps_attr_chroma_qp_offset+ash_attr_qp_delta_chroma

[0300] If the value of ash_attr_layer_qp_delta_present_flag is the same as the first value (e.g., 1), ash_attr_num_layer_qp_minus1 can be signaled. The value of ash_attr_num_layer_qp_minus1 plus 1 can indicate the number of layers for which ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma are signaled. If ash_attr_num_layer_qp is not signaled, the value of ash_attr_num_layer_qp can be inferred to be 0. The variable NumLayerQp, which specifies the number of layers, can be derived by adding 1 to the value of ash_attr_num_layer_qp_minus1 as follows: (NumLayerQp = ash_attr_num_layer_qp_minus1 + 1).

[0301] If the value of ash_attr_layer_qp_delta_present_flag is equal to the first value (e.g., 1), ash_attr_layer_qp_delta_luma[i] may be repeatedly signaled by the value of NumLayerQp. i may increase by 1 from 0 to (NumLayerQp-1). Also, in the repeated process of increasing i by 1, if the value of attribute_dimension_minus1[ash_attr_sps_attr_idx] is greater than 0, ash_attr_layer_qp_delta_chroma[i] may be further signaled. ash_attr_layer_qp_delta_luma may indicate the luma delta quantization parameter (qp) from InitialSliceQpy in each layer. ash_attr_layer_qp_delta_chroma may indicate the chroma delta quantization parameter qp from InitialSliceQpC in each layer. The variables SliceQpY[i] and SliceQpC[i] can be derived as follows:

[0302] SliceQpY[i]=InitialSliceQpY+ash_attr_layer_qp_delta_luma[i]

[0303] SliceQpC[i]=InitialSliceQpC+ash_attr_layer_qp_delta_chroma[i]

[0304] ash_attr_region_qp_delta_present_flag can also be signaled. If the value of ash_attr_region_qp_delta_present_flag is equal to a first value (e.g., 1), this can indicate that ash_attr_region_qp_delta, the region bounding box origin, and the size are present in the current attribute slice header. If the value of ash_attr_region_qp_delta_present_flag is equal to a second value (e.g., 0), this can indicate that ash_attr_region_qp_delta, the region bounding box origin, and the size are not present in the current attribute slice header. If the value of ash_attr_layer_qp_delta_present_flag is the same as the first value (e.g., 1), ash_attr_qp_region_box_origin_x, ash_attr_qp_region_box_origin_y, ash_attr_qp_region_box_origin_z, ash_attr_qp_region_box_width, ash_attr_qp_region_box_height, ash_attr_qp_region_box_depth, and ash_attr_region_qp_delta may be further signaled. ash_attr_qp_region_box_origin_x may indicate the x offset of the region bounding box relative to slice_origin_x, ash_attr_qp_region_box_origin_y may indicate the y offset of the region bounding box relative to slice_origin_y, and ash_attr_qp_region_box_origin_z may indicate the z offset of the region bounding box relative to slice_origin_z.ash_attr_qp_region_box_size_width may indicate the width of the region bounding box, ash_attr_qp_region_box_size_height may indicate the height of the region bounding box, and ash_attr_qp_region_box_size_depth may indicate the depth of the region bounding box. ash_attr_region_qp_delta may indicate the delta qp from SliceQpY[i] and SliceQpC[i] of the region specified by ash_attr_qp_region_box. The variable RegionboxDeltaQp, which specifies the region box delta quantization parameter, may be set to the same value as ash_attr_region_qp_delta.

[0305] As shown in Figure 30, the syntax elements signaled via the attribute slice data are as follows: Zerorun may indicate the number of (preceding) zeros before predIndex or residual. predIndex[i] may indicate the predictor index for decoding the i-th point value of the attribute. The value of predIndex[i] may range from 0 to the value of max_num_conductors.

[0306] Metadata Slice Syntax Structure

[0307] Figure 31 shows an example of the syntax structure of a metadata slice. The syntax elements (or fields) shown in Figure 31 may be syntax elements included in an attribute slice or may be syntax elements signaled via an attribute slice.

[0308] As shown in (a) of Figure 31, the metadata slice bitstream (metadata_slice_bitstream()) can include a metadata slice header (metadata_slice_header()) and metadata slice data (metadata_slice_data()). (b) of Figure 31 shows an example of the metadata slice header, and (c) of Figure 31 shows an example of the metadata slice data.

[0309] As shown in (b) of FIG. 31, syntax elements signaled via the metadata slice header are as follows: msh_slice_id may indicate an identifier for identifying the metadata slice bitstream. msh_geom_slice_id may indicate an identifier for identifying a geometry slice associated with the metadata carried in the metadata slice. msh_attr_id may indicate an identifier for identifying an attribute associated with the metadata carried in the metadata slice. msh_attr_slice_id may indicate an identifier for identifying an attribute slice associated with the metadata carried in the metadata slice. As shown in (c) of FIG. 31, a metadata bitstream (metadata_bitstream()) may be signaled via the metadata slice data.

[0310] TLV structure

[0311] As mentioned above, a G-PCC bitstream can refer to a bitstream of point cloud data consisting of a sequence of TLV structures, which are also referred to as "TLV encapsulation structures," "G-PCC TLV encapsulation structures," or "G-PCC TLV structures."

[0312] An example of a TLV encapsulation structure is shown in Figure 32, an example of a syntax structure of the TLV encapsulation is shown in Figure 33a, and an example of a payload type of the TLV encapsulation structure is shown in Figure 33b. As shown in Figure 32, each TLV encapsulation structure can be composed of a TLV type (TLV TYPE), a TLV length (TLV LENGTH), and / or a TLV payload (TLV PAYLOAD). The TLV type can be type information of the TLV payload, the TLV length can be length information of the TLV payload, and the TLV payload can be payload (or payload bytes). Looking at the syntax structure (tlv_encapsulation()) of the TLV encapsulation illustrated in Figure 33a, tlv_type can indicate type information of the TLV payload, tlv_num_payload_bytes can indicate length information of the TLV payload, and tlv_payload_byte[i] can indicate the TLV payload. tlv_payload_byte[i] can be signaled by the value of tlv_num_payload_bytes, where i can be incremented by 1 from 0 to (tlv_num_payload_bytes-1).

[0313] The TLV payload may include an SPS, a GPS, one or more APSs, a tile inventory, a geometry slice, one or more attribute slices, and one or more metadata slices. According to an embodiment, the TLV payload of each TLV encapsulation structure may include one of an SPS, a GPS, one or more APSs, a tile inventory, a geometry slice, one or more attribute slices, and one or more metadata slices, depending on the type information of the TLV payload. The data included in the TLV payload may be classified based on the type information of the TLV payload. For example, as shown in FIG. 33b, a value of 0 in the tlv_type field indicates that the data included in the TLV payload is an SPS, and a value of 1 in the tlv_type field indicates that the data included in the TLV payload is a GPS. A value of 2 in the tlv_type field indicates that the data included in the TLV payload is a geometry slice, and a value of 3 in the tlv_type field indicates that the data included in the TLV payload is an APS. A value of tlv_type of 4 indicates that the data contained in the TLV payload is an attribute slice, a value of 5 indicates that the data contained in the TLV payload is a tile inventory (or tile parameter set), a value of tlv_type of 6 indicates that the data contained in the TLV payload is a frame boundary marker, and a value of tlv_type of 7 indicates that the data contained in the TLV payload is a metadata slice. The payload of the TLV encapsulation structure can follow the format of a High Efficiency Video Coding (HEVC) Network Abstraction Layer (NAL) unit.

[0314] The information included in the SPS in the TLV payload may include some or all of the information included in the SPS of Figure 21. The information included in the tile inventory in the TLV payload may include some or all of the information included in the tile inventory of Figure 26. The information included in the GPS in the TLV payload may include some or all of the information included in the GPS of Figure 23. The information included in the APS in the TLV payload may include some or all of the information included in the APS of Figure 24. The information included in the geometry slice in the TLV payload may include some or all of the information included in the geometry slices of Figures 27 and 28. The information included in the attribute slice in the TLV payload may include some or all of the information included in the attribute slices of Figures 29 and 30. The information included in the metadata slice in the TLV payload may include some or all of the information included in the metadata slice of Figure 31.

[0315] Encapsulation / Decapsulation

[0316] Such a TLV encapsulation structure can be generated by the transmission unit, transmission processing unit, and encapsulation unit mentioned in this specification. The G-PCC bitstream configured with the TLV encapsulation structure can be transmitted to the receiving device as is, or can be encapsulated and transmitted to the receiving device. For example, the encapsulation processing unit 1525 can encapsulate the G-PCC bitstream configured with the TLV encapsulation structure in a file / segment format and transmit it. The decapsulation processing unit 1710 can decapsulate the encapsulated file / segment to obtain the G-PCC bitstream.

[0317] According to an embodiment, a G-PCC bitstream can be encapsulated in an ISOBMFF-based file format. In this case, the G-PCC bitstream can be stored in a single track or multiple tracks within an ISOBMFF file. Here, a single track or multiple tracks within a file is also referred to as a "track" or a "G-PCC track." An ISOBMFF-based file is also referred to as a container, container file, media file, G-PCC file, etc. Specifically, a file can be composed of boxes and / or information that can be referred to as ftyp, moov, mdat, etc.

[0318] The ftyp box (file type box) may provide file type or file compatibility-related information for the file. A receiving device may classify the file by referring to the ftyp box. The mdat box, also called a media data box, may contain actual media data. Depending on the embodiment, a geometry slice (or a coded geometry bitstream) and zero or more attribute slices (or coded attribute bitstreams) may be included in a sample of the mdat box in the file. Here, a sample may be referred to as a G-PCC sample. The moov box, also called a movie box, may contain metadata for the media data of the file. For example, the moov box may contain information necessary for decoding and playing the media data, and may contain information about the track and sample of the file. The moov box may serve as a container for all metadata. The moov box may be the top-layer box among metadata-related boxes.

[0319] According to an embodiment, the MOOV box may include a track (trak) box that provides information related to the track of a file. The trak box may include a media (mdia) box that provides media information for the track and a track reference container (tref) box that references the track and a sample of the file corresponding to the track. The media box may include a media information container (minf) box that provides information about the media data and a handler (hdlr) box that indicates the stream type. The minf box may include a sample table (stbl) box that provides metadata related to the samples in the mdat box. The stbl box may include a sample description (stsd) box that provides information on the coding type used and initialization information required for the coding type. According to an embodiment, the sample description (stsd) box may include a sample entry for the track. Depending on the embodiment, signaling information (or metadata) such as SPS, GPS, APS, tile inventory, etc. may be included in the sample entries of the moov box or the samples of the mdat box in the file.

[0320] A G-PCC track can be defined as a volumetric visual track that carries geometry slices (or coded geometry bitstreams), attribute slices (or coded attribute bitstreams), or both geometry slices and attribute slices. Depending on the embodiment, a volumetric visual track can be identified by a volumetric visual media handler type 'volv' in a HandlerBox of a MediaBox and / or a volumetric visual media header (vvhd) in a minf box of a MediaBox. The minf box is also referred to as a media information container or media information box. The minf box is contained in a MediaBox, which is contained in a track box, which can be contained in a moov box of the file. A single volumetric visual track or multiple volumetric visual tracks can be present in a file.

[0321] VolumetricVisualMediaHeaderBox

[0322] A volumetric visual track can use a volumetric visual media header (vvhd) box in a media information box (MediaInformationBox). The volumetric visual media header box can be defined as follows:

[0323] Box Type: 'vvhd'

[0324] Container:MediaInformationBox

[0325] Mandatory: Yes

[0326] Quantity: Exactly one

[0327] The syntax for the Volumetric Visual Media Header box is as follows:

[0328] aligned(8) class VolumetricVisualMediaHeaderBox

[0329] extends FullBox('vvhd', version=0,1){

[0330] }

[0331] In the above syntax, version may be an integer value indicating the version of the volumetric visual media header box.

[0332] VolumetricVisualSampleEntry

[0333] By way of example, a volumetric visual track may use volumetric visual sample entries for transmitting signaling information as follows:

[0334] class VolumetricVisualSampleEntry(codingname) extends SampleEntry(codingname){

[0335] unsigned int(8)

[32] compressorname;

[0336] / / other boxes from derived specifications

[0337] }

[0338] In the above syntax, compressorName may indicate the name of a compressor for informational purposes. According to an embodiment, a sample entry (i.e., an ancestor class of VolumetricVisualSmapleEntry) from which a VolumetricVisualSmapleEntry inherits may include a GPCC decoder configuration box.

[0339] G-PCC Decoder Configuration Box (GPCCConfigurationBox)

[0340] Depending on the embodiment, the G-PCC decoder configuration box may include GPCCDecoderConfigurationRecod() as follows:

[0341] class GPCCConfigurationBox extends Box('gpcC'){

[0342] GPCCDecoderConfigurationRecord() GPCCConfig;

[0343] }

[0344] According to an embodiment, GPCCDecoderConfigurationRecord() may provide G-PCC decoder configuration information for geometry-based point cloud content. The syntax of GPCCDecoderConfigurationRecord() may be defined as follows:

[0345] aligned(8) class GPCCDecoderConfigurationRecord{

[0346] unsigned int(8) configurationVersion=1;

[0347] unsigned int(8) profile_idc;

[0348] unsigned int(24)profile_compatibility_flags

[0349] unsigned int(8) level_idc;

[0350] unsigned int(8) numOfSetupUnitArrays;

[0351] for(i=0;i<numOfSetupUnitArrays;i++){

[0352] unsigned int(7) SetupUnitType;

[0353] bit(1)SetupUnit completeness;

[0354] unsigned int(8) numOfSepupUnit;

[0355] for(i=0; numOfSepupUnit;i++){

[0356] tlv_encapsulation setupUnit;

[0357] }

[0358] }

[0359] / / additional fields

[0360] }

[0361] The configuration Version may be a version field. Incompatible changes to the record may be indicated by a change in the version number. The values ​​for profile_idc, profile_compatibility_flags, and level_idc may be valid for all parameter sets that are activated when the bitstream described by the record is decoded. The profile_idc may indicate the profile to which the bitstream associated with the configuration record conforms. The profile_idc may contain a profile code to indicate a specific profile of the G-PCC. A value of 1 in the profile_compatibility_flags may indicate that the bitstream conforms to the profile indicated by the profile_idc field. Each bit in the profile_compatibility_flags may be set only if all parameter sets set that bit. The level_idc may contain a profile level code. The level_idc may indicate a capability level equal to or higher than the highest level indicated for the highest tier in all parameter sets. The numOfSetupUnitArrays may indicate the number of G-PCC setup unit arrays of the type indicated by the setupUnitTye. That is, numOfSetupUnitArrays may indicate the number of G-PCC setup unit arrays included in GPCCDecoderConfigurationRecord(). setupUnitType, setupUnit_completeness, and numOfSetupUnits may further be included in GPCCDecoderConfigurationRecord().setupUnitType, setupUnit_completeness, and numOfSetupUnits are contained by a repeat statement that is repeated the number of times equal to the value of numOfSetupUnitArrays, and this repeat statement can be repeated by incrementing by 1 until i is between 0 and (numOfSetupUnitArrays-1). setupUnitType can indicate the type of G-PCC setupUnits. That is, the value of setupUnitType can be one of the values ​​indicating SPS, GPS, APS, or tile inventory. A value of setupUnit_completeness of 1 can indicate that all setup units of a given type are in the next array and that there are none in the stream. Alternatively, a value of the setupUnit_completeness field of 0 can indicate that there are additional setup units of the indicated type in the stream. numOfSetuUnits can indicate the number of G-PCC setup units of the type indicated by setupUnitType. setupUnit(tlv_encapsulation setupUnit) can also be included in GPCCDecoderConfigurationRecord(). A setupUnit is contained by a repeat statement that is repeated the number of times specified by numOfSetupUnits, where i can be incremented by 1 from 0 to (numOfSetupUnits-1). A setupUnit can be a setup unit of the type indicated by setupUnitType, for example, SPS, GPS, APS, or an instance of a TLV encapsulation structure carrying a tile inventory.

[0362] A volumetric visual track may use a volumetric visual sample for transmitting actual data. A volumetric visual sample entry may be referred to as a sample entry or a G-PCC sample entry, and a volumetric visual sample may be referred to as a sample or a G-PCC sample. A single volumetric visual track may be referred to as a single track or a G-PCC single track, and multiple volumetric visual tracks may be referred to as multiple tracks or multiple G-PCC tracks. Signaling information related to sample grouping, track grouping, single track encapsulation of a G-PCC bitstream, or multiple track encapsulation of a G-PCC bitstream, or signaling information for supporting spatial proximity, may be added to a sample entry in the form of a box or a full box. The signaling information may include at least one of a GPCC entry information box (GPCCEntryInfoBox), a GPCC component type box (GPCCComponentTypeBox), a cubic region information box (CubicRegionInfoBox), a 3D bounding box information box (3DBoundingBoxInfoBox), or a tile inventory (TileInventoryBox).

[0363] GPCC Entry Information Structure

[0364] The syntax structure of the G-PCC entry information box (GPCCEntryInfoBox) can be defined as follows:

[0365] class GPCCEntryInfoBox extends Box('gpsb') {

[0366] GPCCEntryInfoStruct();

[0367] }

[0368] In the above syntax structure, a GPCCEntryInfoBox with a sample entry type of 'gpsb' can include a GPCCEntryInfoStruct(). The syntax of GPCCEntryInfoStruct() can be defined as follows:

[0369] aligned(8) class GPCCEntryInfoStruct {

[0370] unsigned int(1) main_entry_flag;

[0371] unsigned int(1) dependent_on;

[0372] if(dependent_on){ / / non-entry

[0373] unsigned int(16) dependency_id;

[0374] }

[0375] }

[0376] GPCCEntryInfoStruct() may include main_entry_flag and dependent_on. main_entry_flag may indicate whether this is an entry point for decoding a G-PCC bitstream. dependent_on indicates whether its decoding is dependent on others. If dependent_on is present in a sample entry, dependent_on may indicate that the decoding of a sample within a track is dependent on another track. If the value of dependent_on is 1, GPCCEntryInfoStruct() may further include dependency_id. dependency_id may indicate the identifier of the track for decoding the associated data. If dependency_id is present in a sample entry, dependency_id may indicate the identifier of the track carrying the G-PCC sub-bitstream on which the decoding of a sample within the track depends. If dependency_id exists in a sample group, dependency_id may represent an identifier of a sample that carries a G-PCC sub-bitstream on which the decoding of the associated sample depends.

[0377] G-PCC component information structure

[0378] The syntax structure of the G-PCC component type box (GPCCComponentTypeBox) can be defined as follows:

[0379] aligned(8) class GPCCComponentTypeBox extends FullBox('gtyp', version=0,0){

[0380] GPCCComponentTypeStruct();

[0381] }

[0382] A GPCCComponentTypeBox with a sample entry type of 'gtyp' can contain a GPCCComponentTypeStruct(). The syntax of GPCCComponentTypeStruct() can be defined as follows:

[0383] aligned(8) class GPCCComponentTypeStruct {

[0384] unsigned int(8) numOfComponents;

[0385] for(i=0;i <numOfComponents;i++) {

[0386] unsigned int(8) gpcc_type;

[0387] if(gpcc_type==4)

[0388] unsigned int(8)AttrIdx;

[0389] }

[0390] / / additional fields

[0391] }

[0392] numOfComponents may indicate the number of G-PCC components signaled in the GPCCComponentTypeStruct. gpcc_type may be included in the GPCCComponentTypeStruct by a repeat statement that is repeated the same number of times as the value of numOfComponents. This repeat statement may be repeated with i increasing by 1 from 0 to (numOfComponents-1). gpcc_type may indicate the type of G-PCC component. For example, a value of 2 for gpcc_type may indicate a geometry component, and a value of 4 may indicate an attribute component. When the value of gpcc_type is 4, i.e., indicating an attribute component, the repeat statement may further include AttrIdx. AttrIdx may indicate the identifier of the attribute signaled in SPS(). A G-PCC component type box (GPCCComponentTypeBox) may be included in a sample entry for multiple tracks. When a G-PCC Component Type Box (GPCCComponentTypeBox) is present in the sample entry of a track that carries some or all of a G-PCC bitstream, a GPCCComponentTypeStruct() can indicate one or more G-PCC component types carried by each track. The GPCCComponentTypeBox or GPCCComponentTypeStruct() that contains the GPCCComponentTypeStruct() are sometimes referred to as G-PCC component information.

[0393] Sample Group

[0394] The encapsulation processor mentioned in this disclosure may group one or more samples to generate a sample group. The encapsulation processor, metadata processor, or signaling processor mentioned in this disclosure may signal signaling information associated with the sample group to a sample, sample group, or sample entry. That is, sample group information associated with the sample group may be added to the sample, sample group, or sample entry. The sample group information may be 3D bounding box sample group information, 3D region sample group information, 3D tile sample group information, 3D tile inventory sample group information, etc.

[0395] Track Groups

[0396] The encapsulation processor mentioned in this disclosure may group one or more tracks to generate a track group. The encapsulation processor, metadata processor, or signaling processor mentioned in this disclosure may signal signaling information associated with the track group to a sample, track group, or sample entry. That is, the track group information associated with the track group may be added to the sample, track group, or sample entry. The track group information may be 3D bounding box track group information, point cloud composition track group information, spatial region track group information, 3D tile track group information, 3D tile inventory track group information, etc.

[0397] Sample Entry

[0398] Figure 34 is a diagram illustrating an ISOBMFF-based file including a single track. Figure 34(a) shows an example of the layout of an ISOBMFF-based file including a single track, and Figure 34(b) shows an example of a sample structure of an mdat box when a G-PCC bitstream is stored in a single track of a file. Figure 35 is a diagram illustrating an ISOBMFF-based file including multiple tracks. Figure 35(a) shows an example of the layout of an ISOBMFF-based file including multiple tracks, and Figure 35(b) shows an example of a sample structure of an mdat box when a G-PCC bitstream is stored in a single track of a file.

[0399] The stsd box (SampleDescriptionBox) included in the moov box of a file can contain a sample entry for a single track that stores a G-PCC bitstream. The SPS, GPS, APS, and tile inventory can be included in the sample entry of the moov box or the sample of the mdat box in a file. In addition, geometry slices and zero or more attribute slices can be included in the sample of the mdat box in a file. When a G-PCC bitstream is stored in a single track of a file, each sample can contain multiple G-PCC components. That is, each sample can consist of one or more TLV encapsulation structures. A sample entry for a single track can be defined as follows:

[0400] Sample Entry Type: 'gpe1','gpeg'

[0401] Container: SampleDescriptionBox

[0402] Mandatory: A'gpe1' or 'gpeg' sample entry is mandatory

[0403] Quantity: One or more sample entries may be present

[0404] The sample entry type 'gpel' or 'gpeg' is mandatory, and one or more sample entries can exist. A G-PCC track can use a VolumetricVisualSampleEntry with a sample entry type of 'gpel' or 'gpeg'. A sample entry of a G-PCC track can include a G-PCC decoder configuration box (GPCCConfigurationBox), which can include a G-PCC decoder configuration record (GPCCDecoderConfigurationRecord()). GPCCDecoderConfigurationRecord() can include at least one of configurationVersion, profile_idc, profile_compatibility_flags, level_idc, numOfSetupUnitArrays, SetupUnitType, completeness, numOfSepupUnit, and setupUnit. The setupUnit array field included in GPCCDecoderConfigurationRecord() can include a TLV encapsulated structure containing one SPS.

[0405] If the sample entry type is 'gpe1', all parameter sets, e.g., SPS, GPS, APS, and tile inventory, can be included in the setupUnits array. If the sample entry type is 'gpeg', the above parameter sets can be included in the setupUnits array (i.e., sample entry) or in the stream (i.e., sample). An example of the syntax of a G-PCC sample entry (GPCCSampleEntry) with a sample entry type of 'gpe1' is as follows:

[0406] aligned(8) class GPCCSampleEntry()

[0407] extends VolumetricVisualSampleEntry('gpe1'){

[0408] GPCCConfigurationBox config; / / mandatory

[0409] 3DBoundingBoxInfoBox();

[0410] CubicRegionInfoBox();

[0411] TileInventoryBox();

[0412] }

[0413] A G-PCC sample entry (GPCCSampleEntry) with a sample entry type of 'gpe1' can include a GPCCConfigurationBox, a 3DBoundingBoxInfoBox(), a CubicRegionInfoBox(), and a TileInventoryBox(). The 3DBoundingBoxInfoBox() can indicate 3D bounding box information of the point cloud data associated with the sample carried in the track. The CubicRegionInfoBox() can indicate one or more spatial region information of the point cloud data carried in the sample in the track. The TileInventoryBox() can indicate 3D tile inventory information of the point cloud data carried in the sample in the track.

[0414] As shown in Figure 34(b), a sample may include a TLV encapsulation structure including a geometry slice, a TLV encapsulation structure including one or more parameter sets, and a TLV encapsulation structure including one or more attribute slices.

[0415] As shown in (a) of Figure 35, when a G-PCC bitstream is carried in multiple tracks of an ISOBMFF-based file, each geometry slice or attribute slice can be mapped to an individual track. For example, a geometry slice can be mapped to track 1, and an attribute slice can be mapped to track 2. The track carrying the geometry slice (track 1) can be referred to as a geometry track or a G-PCC geometry track, and the track carrying the attribute slice (track 2) can be referred to as an attribute track or a G-PCC attribute track. The geometry track can be defined as a volumetric visual track carrying geometry slices, and the attribute track can be defined as a volumetric visual track carrying attribute slices.

[0416] A track that carries a portion of a G-PCC bitstream that contains both geometry and attribute slices is also called a multiplexed track. When geometry slices and attribute slices are stored in separate tracks, each sample in the track can contain at least one TLV encapsulation structure carrying data for a single G-PCC component. In this case, each sample may not contain both geometry and attributes, and may not contain multiple attributes. Multi-track encapsulation of a G-PCC bitstream can enable a G-PCC player to effectively access one of the G-PCC components. When a G-PCC bitstream is carried on multiple tracks, the following conditions must be met for a G-PCC player to effectively access one of the G-PCC components:

[0417] a) When a G-PCC bitstream consisting of a TLV encapsulation structure is carried in multiple tracks, the track carrying the geometry bitstream (or geometry slice) becomes the entry point.

[0418] b) In the sample entry, a new box is added to indicate the role of the stream included in the track. The new box can be the G-PCC Component Type Box (GPCCComponentTypeBox) described above. That is, the GPCCComponentTypeBox can be included in the sample entry for multiple tracks.

[0419] c) In tracks that carry only G-PCC geometry bitstreams, a track reference is introduced as the track that carries the G-PCC attribute bitstream.

[0420] GPCCComponentTypeBox may include GPCCComponentTypeStruct(). If GPCCComponentTypeBox is present in the sample entry of a track that carries part or all of the G-PCC bitstream, GPCCComponentTypeStruct() may indicate the type of one or more G-PCC components carried by each track (e.g., geometry, attribute). For example, a value of 2 in the gpcc_type field included in GPCCComponentTypeStruct() may indicate a geometry component, and a value of 4 may indicate an attribute component. Also, if the value of the gpcc_type field is 4, i.e., indicating an attribute component, an AttrIdx field indicating an identifier of the attribute signaled to SPS() may be further included.

[0421] When a G-PCC bitstream is carried in multiple tracks, the syntax of the sample entry can be defined as follows:

[0422] Sample Entry Type:'gpe1', 'gpeg', 'gpc1' or 'gpcg'

[0423] Container:SampleDescriptionBox

[0424] Mandatory:'gpc1','gpcg' sample entry is mandatory

[0425] Quantity:One or more sample entries may be present

[0426] The sample entry type 'gpc1', 'gpcg', 'gpc1', or 'gpcg' is required, and one or more sample entries can be present. Multiple tracks (e.g., geometry or attribute tracks) can use VolumetricVisualSampleEntry with a sample entry type of 'gpc1', 'gpcg', 'gpc1', or 'gpcg'. In a 'gpe1' sample entry, all parameter sets can be present in the setupUnit array. In a 'gpeg' sample entry, parameter sets can be present in the array or stream. In a 'gpe1' or 'gpeg' sample entry, a GPCCComponentTypeBox may be present. In a 'gpc1' sample entry, the SPS, GPS, and tile inventory can be present in the SetupUnit array of a track carrying a G-PCC geometry bitstream. All associated APSs can be present in the SetupUnit array of a track carrying a G-PCC attribute bitstream. In a 'gpcg' sample entry, SPS, GPS, APS or tile inventory may be present in the array or stream. In a 'gpc1' or 'gpcg' sample array, a GPCCComponentTypeBox may have to be present.

[0427] An example for the syntax of a sample entry in G-PCC is as follows:

[0428] aligned(8) class GPCCSampleEntry()

[0429] extends VolumetricVisualSampleEntry(codingname) {

[0430] GPCCConfigurationBox config; / / mandatory

[0431] GPCCComponentTypeBox type; / / optional

[0432] }

[0433] The compressorname, i.e., codingname, of the base class VolumetricVisualSampleEntry can indicate the name of the compressor to be used with the recommended "\013GPCC coding" value. In "\013GPCC coding", the first byte (octal 13 or decimal 11 represented by \013) is the number of remaining bytes and can indicate the number of bytes in the remaining string. The congif can contain G-PCC decoder configuration information. The info can represent G-PCC component information carried in each track. The info can indicate the component tile carried in the track, and can also indicate the attribute name, index, and attribute type of the G-PCC component carried in the G-PCC attribute track.

[0434] Sample Format

[0435] When the G-PCC bitstream is stored in a single track, the syntax for the sample format is as follows:

[0436] aligned(8) class GPCCSample

[0437] {

[0438] unsigned int GPCCLength=sample_size; / / Size of Sample

[0439] for(i=0;i <GPCCLength;) / / to end of the sample

[0440] {

[0441] tlv_encapsulation gpcc_unit;

[0442] i+=(1+4)+gpcc_unit.tlv_num_payload_bytes;

[0443] }

[0444] }

[0445] In the above syntax, each sample (GPCCSample) corresponds to a single point cloud frame and can consist of one or more TLV encapsulation structures belonging to the same presentation time. Each TLV encapsulation structure can contain a single type of TLV payload. In addition, a sample can be independent (e.g., a sync sample). GPCCLength indicates the length of the sample, and gpcc_unit can contain an instance of a TLV encapsulation structure containing a single G-PCC component (e.g., a geometry slice).

[0446] When a G-PCC bitstream is stored in multiple tracks, each sample may correspond to a single point cloud frame, and samples contributing to the same point cloud frame in different tracks may have the same presentation time. Each sample must consist of one or more G-PCC units of the G-PCC component displayed in the GPCCComponentInfoBox of the sample entry and zero or more G-PCC units carrying either a parameter set or a tile inventory. If a G-PCC unit containing a parameter set or a tile inventory is present in a sample, the corresponding F-PCC sample must appear before the G-PCC unit of the G-PCC component. Each sample may contain one or more G-PCC units containing attribute data units and zero or more G-PCC units carrying parameter sets. When a G-PCC bitstream is stored in multiple tracks, the syntax and semantics for the sample format may be the same as those for when a G-PCC bitstream is stored in a single track.

[0447] Subsample

[0448] In a receiving device, the geometry slice must be decoded first, and then the attribute slice must be decoded based on the decoded geometry. Therefore, if each sample is composed of multiple TLV encapsulation structures, each TLV encapsulation structure must be accessed in the sample. Also, if a sample is composed of multiple TLV encapsulation structures, each of the multiple TLV encapsulation structures can be stored as a subsample. The subsamples may be referred to as G-PCC subsamples. For example, if a sample includes a parameter set TLV encapsulation structure containing a parameter set, a geometry TLV encapsulation structure containing a geometry slice, and an attribute TLV encapsulation structure containing an attribute slice, the parameter set TLV encapsulation structure, the geometry TLV encapsulation structure, and the attribute TLV encapsulation structure can each be stored as a subsample. In this case, the type of TLV encapsulation structure carried in the subsample may be required to enable access to each G-PCC component in the sample.

[0449] When a G-PCC bitstream is stored in a single track, a G-PCC subsample can contain only one TLV encapsulation structure. One SubSampleInformationBox can be present in the Sample Table Box (SampleTableBox, stbl) of the moov box, or in each Track Fragment Box (TrackFragmentBox, traf) of the Movie Fragment Box (Moof). If a SubSampleInformationBox is present, the 8-bit type value of the TLV encapsulation structure can be included in the 32-bit codec_specific_parameters field of the subsample entry in the SubSampleInformationBox. If the TLV encapsulation structure contains an attribute payload, the 6-bit value of the attribute index can be included in the 32-bit codec_specific_parameters field of the subsample entry in the SubSampleInformationBox. Depending on the embodiment, the type of each subsample can be included in the codec_secific_parameters field of the subsample entry in the SubSampleInformationBox. Depending on the embodiment, the type of each subsample can be identified by parsing the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox. The codec_specific_parameters of the SubSampleInformationBox can be defined as follows:

[0450] if(flags==0) ​​{

[0451] unsigned int(8) PayloadType;

[0452] if(PayloadType==4){ / / attribute payload

[0453] unsigned int(6) AttrIdx;

[0454] bit(18) reserved=0;

[0455] }

[0456] else

[0457] bit(24) reserved=0;

[0458] } else if(flags==1){

[0459] unsigned int(1) tile_data;

[0460] bit(7) reserved=0;

[0461] if (tile_data)

[0462] unsigned int(24) tile_id;

[0463] else

[0464] bit(24) reserved=0;

[0465] }

[0466] In the above subsample syntax, payloadType may indicate the tlv_type of the TLV encapsulation structure within the subsample. For example, a payloadType value of 4 may indicate an attribute slice (i.e., attribute slice). attrIdx may indicate the identifier of the attribute information of the TLV encapsulation structure containing the attribute payload within the subsample. attrIdx may be the same as the ash_attr_sps_attr_idx of the TLV encapsulation structure containing the attribute payload within the subsample. tile_data may indicate whether the subsample contains one tile or another. A value of 1 for tile_data may indicate that the subsample contains a TLV encapsulation structure containing a geometry data unit or attribute data unit corresponding to one G-PCC tile. A value of 0 for tile_data may indicate that the subsample contains a TLV encapsulation structure containing each parameter set, tile inventory, or frame boundary marker. The tile_id may indicate the index of the G-PCC version to which the subsample relates in the tile inventory.

[0467] If subsamples are present when a G-PCC bitstream is stored in multiple tracks (in the case of multiple track encapsulation of G-PCC data in ISOBMFF), only the SubSampleInformationBox with flag = 1 in the SampleTableBox or the TrackFragmentBox of each MovieFragmentBox may need to be present. When a G-PCC bitstream is stored in multiple tracks, the syntax elements and semantics may be the same as when the G-PCC bitstream is stored in a single track and flag = 1.

[0468] Inter-Track Referencing

[0469] When a G-PCC bitstream is carried in multiple tracks (i.e., when the G-PCC geometry bitstream and attribute bitstream are carried in different (separate) tracks), a track reference tool can be used to link the tracks. One TrackReferenceTypeBox can be added to a TrackReferenceBox within the TrackBox of a G-PCC track. The TrackReferenceTypeBox can contain an array of track_IDs that specify the tracks to which the G-PCC track refers.

[0470] According to an embodiment, the present disclosure may provide an apparatus and method for supporting temporal scalability in the carriage of G-PCC data (hereinafter, may be referred to as a G-PCC bitstream, an encapsulated G-PCC bitstream, or a G-PCC file). The present disclosure may also propose an apparatus and method for providing a point cloud content service that efficiently stores a G-PCC bitstream in a single track within a file or divides it into multiple tracks and stores it, and provides signaling therefor. The present disclosure also proposes an apparatus and method for processing a file storage technique that can support efficient access to the stored G-PCC bitstream.

[0471] Temporal scalability

[0472] Temporal scalability may refer to a function that allows the possibility of extracting one or more subsets of independently coded frames. Temporal scalability may also refer to a function that divides G-PCC data into multiple different temporal levels and processes G-PCC frames belonging to different temporal levels independently. Support for temporal scalability allows a G-PCC player (or a transmission device and / or a receiving device of the present disclosure) to effectively access a desired component (target component) among G-PCC components. Support for temporal scalability also allows G-PCC frames to be processed independently, allowing support for temporal scalability at the system level to be expressed as more flexible temporal sub-layering. Support for temporal scalability also allows a system that processes G-PCC data (point cloud content provision system) to manipulate data at a high level to match network capabilities, decoder capabilities, etc., thereby improving the performance of the point cloud content provision system.

[0473] Problems with the prior art

[0474] Prior art related to point cloud content providing systems or G-PCC data carriage does not support temporal scalability. That is, prior art processes G-PCC data at only one temporal level. Therefore, prior art cannot provide the benefits of supporting temporal scalability described above.

[0475] 1. First Example

[0476] Flowcharts for a first embodiment supporting time scalability are shown in FIGS.

[0477] Referring to FIG. 36, the transmission device 10, 1500 may generate information on a sample group (S3610). As described below, a sample group may be a grouping of samples (G-PCC samples) in a G-PCC file based on one or more temporal levels. The transmission device 10, 1500 may also generate a G-PCC file (S3620). Specifically, the transmission device 10, 1500 may generate a G-PCC file including point cloud data, information on the sample group, and / or information on the temporal level. Steps S3610 and S3620 may be performed by one or more of the encapsulation processor 13, the encapsulation processor 1525, and / or the signaling processor 1510. The transmission device 10, 1500 may signal the G-PCC file.

[0478] Referring to FIG. 37, the receiving device 20, 1700 may acquire a G-PCC file (S3710). The G-PCC file may include point cloud data, information on sample groups, and information on temporal levels. The sample groups may be samples (G-PCC samples) in the G-PCC file grouped based on one or more temporal levels. The receiving device 20, 1700 may extract one or more samples belonging to a target temporal level from the samples (G-PCC samples) in the G-PCC file (S3720). Specifically, the receiving device 20, 1700 may decapsulate or extract samples belonging to a target temporal level based on the information on the sample group and / or the information on the temporal level. Here, the target temporal level may be a desired temporal level. According to an embodiment, step S3720 may include a step of determining whether one or more temporal levels exist, a step of determining whether a desired temporal level and / or a desired frame rate is available if one or more temporal levels exist, and a step of extracting samples in a sample group belonging to the desired temporal level if the desired temporal level and / or the desired frame rate is available. Step S3710 (and / or the step of parsing the sample entry) and step S3720 may be performed by one or more of the decapsulation processing unit 22, the decapsulation processing unit 1710, and / or the signaling processing unit 1715. According to an embodiment, the receiving device 20, 1700 may perform a step of determining or judging the temporal level of a sample in a G-PCC file based on information on the sample group and / or information on the temporal level between steps S3710 and S3720.

[0479] Sample grouping

[0480] Schemes for supporting temporal scalability may include a sample grouping scheme and a track grouping scheme. The sample grouping scheme may be a scheme for grouping samples in a G-PCC file according to temporal levels, and the track grouping scheme may be a scheme for grouping tracks in a G-PCC file according to temporal levels. Hereinafter, the present disclosure will be described focusing on an apparatus and method for supporting temporal scalability based on the sample grouping scheme.

[0481] A sample group can be used to associate samples with their designated temporal levels. That is, a sample group can indicate which samples belong to which temporal level. A sample group can also be information about the results of grouping one or more samples into one or more temporal levels. A sample group is also called a 'tele' sample group or a temporal level sample group 'tele'.

[0482] Information for sample groups

[0483] The information about the sample group may include information about the result of sample grouping. Thus, the information about the sample group may be information used to associate samples with their assigned temporal levels. That is, the information about the sample group may indicate which samples belong to which temporal level, or may be information about the result of grouping one or more samples into one or more temporal levels.

[0484] Information about sample groups can be present in tracks containing geometry data units. When G-PCC data is carried in multiple tracks, information about sample groups can be present only in geometry tracks to group each sample in the track into a specified temporal level. Samples in attribute tracks can be inferred based on their relationship to their associated geometry tracks. For example, samples in an attribute track can belong to the same temporal level as samples in their associated geometry track.

[0485] If information about a sample group exists in a G-PCC tile track referenced by a G-PCC tile base track, information about the sample group may also have to exist in the rest tile tracks referenced by the G-PCC tile base track. Here, the G-PCC tile track may be a volumetric visual track that carries all G-PCC components or a single G-PCC component corresponding to one or more G-PCC tiles. Also, the G-PCC tile base track may be a volumetric visual track that carries all parameter sets and tile inventory corresponding to the G-PCC tile track.

[0486] Information on temporal levels

[0487] Information about the temporal level can be signaled to describe the temporal scalability supported in a G-PCC file. The information about the temporal level can be present in the sample entry of the track that includes the sample group (or information about the sample group). For example, the information about the temporal level can be present in the GPCCDecoderConfigurationRecord() or in the G-PCC Scalability Information Box (GCCScalabilityInfoBox) that signals scalability information for the G-PCC track.

[0488] As expressed in the following syntax structure, the information on the temporal levels may include one or more of number information (e.g., num_temporal_levels), temporal level identification information (e.g., level_idc), and / or frame rate information. In addition, the information on the frame rate may include frame rate information (e.g., avg_frame_rate) and / or frame rate presence information (e.g., avg_frame_rate_present_flag).

[0489] aligned(8) class GPCCDecoderConfigurationRecord {

[0490] unsigned int(8) configurationVersion=1;

[0491] unsigned int(8) profile_idc;

[0492] unsigned int(24) profile_compatibility_flags;

[0493] unsigned int(3) num_temporal_levels;

[0494] unsigned int(1) avg_frame_rate_present;

[0495] bit(4) reserved='1111'b;

[0496] for(i=0;i==0||i <num_temporal_levels;i++) {

[0497] unsigned int(8)level_idc[i];

[0498] if (avg_frame_rate_present)

[0499] unsigned int(16)avg_frame_rate[i];

[0500] }}

[0501] unsigned int ( 8 ) numOfSetupUnits ;

[0502] for ( i = 0 ; i <numOfSetupUnits;i++) {

[0503] tv_encapsulation setupUnit; / / as defined in ISO / IEC 23090-9

[0504] }}

[0505] / / additional fields

[0506] }}

[0507] (1)In the snowflake

[0508] The number information indicates the number of temporal levels. For example, if the value of num_temporal_levels is greater than a first value (e.g., 1), the value of num_temporal_levels may indicate the number of temporal levels. According to an embodiment, the number information may further indicate whether a sample in the G-PCC file is temporally scalable. For example, if the value of num_temporal_levels is equal to a first value (e.g., 1), this may indicate that the sample is not temporally scalable (temporal scalability is not supported). As another example, if the value of num_temporal_levels is less than the first value (e.g., 1) (e.g., 0), this may indicate that it is unknown whether the sample is temporally scalable. In summary, if the value of num_temporal_levels is equal to or less than a first value (e.g., 1), the value of num_temporal_levels may indicate whether the sample is temporally scalable.

[0509] (2) Temporal level identification information

[0510] The temporal level identification information indicates a temporal level identifier (or level code) for the corresponding track. That is, the temporal level identification information may indicate the temporal level identifier of a sample in the corresponding track. According to an embodiment, if the value of num_temporal_levels is equal to or less than a first value (e.g., 1), the temporal level identification information may indicate the level code of a sample in a G-PCC file. The temporal level identification information may be generated as many times as the number of temporal levels indicated by the number information in step S3620, and may be obtained as many times as the number of temporal levels indicated by the number information in step S3710. For example, the temporal level identification information may be generated / obtained as i increases by 1 from 0 to 'num_temporal_level-1'.

[0511] (3) Information about frame rate

[0512] As described above, the information on the frame rate may include frame rate information (eg, avg_frame_rate) and / or frame rate present information (eg, avg_frame_rate_present_flag).

[0513] The frame rate information may indicate a frame rate at a temporal level. For example, the frame rate information may indicate a frame rate at a temporal level in frame units, and the frame rate indicated by the frame rate information may be an average frame rate. According to an embodiment, if the value of avg_frame_rate is equal to a first value (e.g., 0), this may indicate an unspecified average frame rate. The frame rate information may be generated for the number of temporal levels indicated by the number information, and may be acquired for the number of temporal levels indicated by the number information. For example, the frame rate information may be generated / acquired as i increases by 1 from 0 to 'num_temporal_level-1'.

[0514] The frame rate present information may indicate whether frame rate information is present (i.e., whether frame rate information is signaled). For example, if the value of Avg_frame_rate_present_flag is a first value (e.g., 1), this may indicate that frame rate information is present, and if the value of avg_frame_rate_present_flag is a second value (e.g., 0), this may indicate that frame rate information is not present.

[0515] The frame rate information may be signaled regardless of the value of the frame rate presence information. Depending on the embodiment, whether the frame rate information is signaled may be determined depending on the value of the frame rate presence information. Figures 38 and 39 are flowcharts illustrating an example in which whether the frame rate information is signaled is determined depending on the value of the frame rate presence information.

[0516] Referring to FIG. 38, the transmission device 10, 1500 may determine whether frame rate information is present (S3810). If frame rate information is not present, the transmission device 10, 1500 may include only frame rate presence information in the information regarding the temporal level. That is, if frame rate information is not present, the transmission device 10, 1500 may signal only frame rate presence information. In this case, the value of avg_frame_rate_present_flag may be equal to a second value (e.g., 0). In contrast, if frame rate information is present, the transmission device 10, 1500 may include frame rate presence information and frame rate information in the information regarding the temporal level (S3830). That is, if frame rate information is present, the transmission device 10, 1500 may signal frame rate presence information and frame rate information. In this case, the value of avg_frame_rate_present_flag may be equal to a first value (e.g., 1). Steps S3810 to S3830 can be performed by one or more of the encapsulation processor 13, the encapsulation processor 1525, and / or the signaling processor 1510.

[0517] Referring to FIG. 39, the receiving device 20, 1700 may acquire frame rate presence information (S3910). The receiving device 20, 1700 may also determine a value indicated by the frame rate presence information (S3920). If the value of avg_frame_rate_present_flag is equal to a second value (e.g., 0), the receiving device 20, 1700 may not acquire frame rate information and may terminate the process of acquiring information about the frame rate. Alternatively, if the value of avg_frame_rate_present_flag is equal to a first value (e.g., 1), the receiving device 20, 1700 may acquire frame rate information (S3930). The receiving device 20, 1700 may extract samples corresponding to a target frame rate from among samples in the G-PCC file (S3940). Here, the 'samples corresponding to the target frame rate' may be a temporal level or track having a frame rate corresponding to the target frame rate. Furthermore, the 'frame rate corresponding to the target frame rate' may include not only a frame rate having the same value as the target frame rate, but also a frame rate having a value less than the target frame rate. Steps S3910 to S3940 may be performed by one or more of the decapsulation processor 22, the decapsulation processor 1710, and / or the signaling processor 1715. Depending on the embodiment, a step of determining the frame rate of a temporal level (or track) based on the frame rate information may be performed between steps S3930 and S3940.

[0518] In this way, when whether to signal frame rate information is determined according to the value of the frame rate presence information, the bits for signaling frame rate information can be reduced, thereby increasing bit efficiency. For example, a specific file parser or player (e.g., a receiving device) can use frame rate information, but other file parsers or players may not want to use frame rate information or may not use frame rate information. Furthermore, when a G-PCC file is played by a file parser or player that does not want to use frame rate information or does not use frame rate information, setting the value of the frame rate presence information to a second value (e.g., 0) can increase bit efficiency for signaling frame rate information.

[0519] 2. Second Example

[0520] The second embodiment is an embodiment in which a condition for 'smallest composition time difference between successive samples' is added to the first embodiment.

[0521] When temporal scalability is used or activated for a G-PCC file, coded frames of the G-PCC bitstream can be aligned to different temporal levels, and the different temporal levels can be stored in different tracks. The frame rate of each temporal level (or each track) can be determined based on the frame rate information.

[0522] For example, samples may be arranged into three temporal levels (temporal level 0, temporal level 1, and temporal level 2), and each temporal level may be stored in one track, thereby forming three tracks (track 0, track 1, and track 2). In this case, a file parser or player (e.g., a receiving device) can easily determine the frame rate when playing track 0 alone, or the frame rate when playing tracks 0 and 1 together. Given a target frame rate, the file parser or player (e.g., a receiving device) can select a track to play. For example, if temporal level 0 is associated with a frame rate of 30 fps, temporal level 1 is associated with a frame rate of 60 fps, and temporal level 2 is associated with a frame rate of 120 fps, and the target frame rate is 60 fps, the file parser or player can easily determine that both track 0 and track 1 should be played.

[0523] On the other hand, if a sample (or a track associated with a higher temporal level) having a higher temporal level is associated with a higher frame rate, the complexity and size of the track can also increase as file playback progresses to the higher temporal level. That is, if a sample (or a track associated with a given temporal level) has twice the frame rate (twice the number of frames) of the track having the immediately lower temporal level, the complexity for playback can increase several times as the temporal level of the track to be played increases.

[0524] Therefore, for smooth and gradual playback, a restriction on the difference in frame rates of each track may be necessary. The restriction that the distance between two different frames (i.e., the connection time difference) in each track must be the same regardless of which track is selected for playback may correspond to an ideal condition for smooth and gradual playback. However, because this condition can occur under ideal circumstances, realistic constraints for smooth and gradual playback may be necessary in practical situations where ideal circumstances may not occur.

[0525] This disclosure proposes a constraint that the minimum coupling time difference between consecutive samples (frames) of each track for playback may be the same. Specifically, assuming that the temporal levels include a 'first temporal level' and a 'second temporal level' having a level value greater than the first temporal level, the 'minimum coupling time difference between consecutive samples in the second temporal level' may be the same as the 'minimum coupling time difference between consecutive samples in the first temporal level' or greater than the 'minimum coupling time difference between consecutive samples in the first temporal level'. That is, the 'minimum coupling time difference between consecutive samples in the second temporal level' may be equal to or greater than the 'minimum coupling time difference between consecutive samples in the first temporal level'.

[0526] When such constraints are applied, it is possible to prevent a temporal level having a relatively high value from containing an excessively large number of frames compared to a temporal level having a relatively low value, thereby reducing the increase in complexity for playback and enabling smooth, gradual playback.

[0527] 3. Third Example

[0528] The third embodiment is based on the first and / or second embodiment and is an embodiment that prevents redundant signaling of information for the temporal level.

[0529] As described above, sample groups (or information about sample groups) can be present in tracks containing geometry data units, and information about temporal levels can be present in sample entries of tracks containing sample groups (or information about sample groups). That is, information about temporal levels can only be present in tracks containing geometry tracks. Furthermore, samples in attribute tracks can be inferred based on their relationship with their associated geometry tracks. Therefore, signaling information about temporal levels for tracks that do not contain geometry tracks can cause redundant signaling problems.

[0530] The following syntax structure is according to the third embodiment.

[0531] aligned(8) class GPCCDecoderConfigurationRecord {

[0532] unsigned int(8) configurationVersion=1;

[0533] unsigned int(8) profile_idc;

[0534] unsigned int(24) profile_compatibility_flags;

[0535] unsigned int(1) multiple_temporal_level_tracks_flag;

[0536] unsigned int(1) frame_rate_present_flag;

[0537] unsigned int(3) num_temporal_levels;

[0538] bit(3) reserved = 0;

[0539] for(i=0; i<num_temporal_levels; i++){

[0540] unsigned int(8) level_idc;

[0541] unsigned int(16) temporal_level_id;

[0542] if (frame_rate_present_flag)

[0543] unsigned int(16) frame_rate;

[0544] }

[0545] unsigned int(8) numOfSetupUnits;

[0546] for (i=0; i<numOfSetupUnits; i++) {

[0547] tlv_encapsulation setupUnit; / / as defined in ISO / IEC 23090-9

[0548] }

[0549] / / additional fields

[0550] }

[0551] In the above syntax structure, multiple_temporal_level_tracks_flag is a syntax element included in the information for the temporal level and may indicate whether multiple temporal level tracks exist in the G-PCC file. For example, if the value of multiple_temporal_level_tracks_flag is equal to a first value (e.g., 1), this may indicate that the G-PCC bitstream frames are grouped into multiple temporal level tracks, and if the value of multiple_temporal_level_tracks_flag is equal to a second value (e.g., 2), this may indicate that all temporal level samples exist in a single track. If the type of component (gpcc_data) carried by the current track is an attribute and / or if the current track is a track with a predetermined sample entry type (e.g., 'gpc1' and / or 'gpcg'),' multiple_temporal_level_tracks_flag may not be signaled. In this case, the value of multiple_temporal_level_tracks_flag may be the same as the corresponding syntax element (information for the temporal level) in the geometry track referenced by the current track.

[0552] The frame_rate_present_flag is frame rate present information, and its meaning is as described above. If the type of component (gpcc_data) carried by the current track is an attribute and / or the current track is a track with a predetermined sample entry type (e.g., 'gpc1' and / or 'gpcg'), the frame_rate_present_flag may not be signaled. In this case, the value of the frame_rate_present_flag may be the same as the value of the corresponding syntax element (information for the temporal level) in the geometry track referenced by the current track.

[0553] num_temporal_levels is number information and may indicate the maximum number of temporal levels at which G-PCC bitstream frames are grouped. If information on temporal levels is not available or all frames are signaled at a single temporal level, the value of num_temporal_levels may be set to 1. The minimum value of num_temporal_levels may be 1. If the type of component (gpcc_data) carried by the current track is an attribute and / or if the current track is a track with a specified sample entry type (e.g., 'gpcc1' and / or 'gpcg'), num_temporal_levels may not be signaled. In this case, the value of num_temporal_levels may be the same as the corresponding syntax element (information on temporal levels) in the geometry track referenced by the current track.

[0554] 40 and 41 are flow charts for methods that can prevent the problem of duplicate signaling.

[0555] 40, the transmission device 10, 1500 can determine the type of component currently transported by the truck (the type of component of the current truck) (S4010). The type of component can be determined based on the value of gpcc_type and Table 4 below.

[0556] [Table 7]

[0557] If the type of component carried by the current track is an attribute (or attribute data) (gpcc_type==4), the transmission device 10, 1500 may not include information on the temporal level (S4020). That is, if the type of component carried by the current track is an attribute, the transmission device 10, 1500 may not signal information on the temporal level. In this case, one or more values ​​of multiple_temporal_level_tracks_flag, frame_rate_present_flag, and / or frame_rate_present_flag may be the same as the corresponding syntax element (information on the temporal level) in the geometry track referenced by the current track. Alternatively, if the type of component carried by the current track is an attribute, the transmission device 10, 1500 may include information on the temporal level (S4030). That is, if the type of component currently being transported by the truck is an attribute, the transmission device 10, 1500 can signal information on the temporal level. Steps S4010 to S4030 can be performed by one or more of the encapsulation processor 13, the encapsulation processor 1525, and / or the signaling processor 1510.

[0558] According to an embodiment, in step S4010, the transmission device 10, 1500 may further determine the type of sample entry of the current track. If the type of component carried by the current track is an attribute and the current track is a track having a predetermined sample entry type (e.g., 'gpc1' and / or 'gpcg'), the transmission device 10, 1500 may not include (signal) information on the temporal level (S4020). In this case, one or more values ​​of multiple_temporal_level_tracks_flag, frame_rate_present_flag, and / or frame_rate_present_flag may be the same as the corresponding syntax element (information on the temporal level) in the geometry track referenced by the current track. Alternatively, if the type of component carried by the current track is not an attribute, or if the current track is not a track with a specified sample entry type (e.g., 'gpc1' and / or 'gpcg'), the transmission device 10, 1500 may include (signal) information regarding the temporal level (S4030).

[0559] 41, the receiving device 20, 1700 can determine the type of component currently being transported by the truck (S4110). The type of component can be determined based on the value of gpcc_type and Table 4 above.

[0560] If the type of component carried by the current track is an attribute, the receiving device 20, 1700 may not acquire information about the temporal level (S4120). In this case, one or more values ​​of multiple_temporal_level_tracks_flag, frame_rate_present_flag, and / or frame_rate_present_flag may be set to the same as the corresponding syntax element (information about the temporal level) in the geometry track referenced by the current track. Alternatively, if the type of component carried by the current track is not an attribute, the receiving device 20, 1700 may acquire information about the temporal level (S4130). Steps S4110 to S4130 may be performed by one or more of the decapsulation processing unit 22, the decapsulation processing unit 1710, and / or the signaling processing unit 1715.

[0561] According to an embodiment, in step S4110, the receiving device 20, 1700 may further determine the type of sample entry of the current track. If the type of component carried by the current track is an attribute and the current track is a track having a predetermined sample entry type (e.g., 'gpc1' and / or 'gpcg'), the receiving device 20, 1700 may not acquire information about the temporal level (S4120). In this case, one or more values ​​of multiple_temporal_level_tracks_flag, frame_rate_present_flag, and / or frame_rate_present_flag may be set to the same as the corresponding syntax element (information about the temporal level) in the geometry track referenced by the current track. Alternatively, if the type of component carried by the current track is not an attribute or the current track is not a track having a predetermined sample entry type (e.g., 'gpc1' and / or 'gpcg'), the receiving device 20, 1700 may acquire information about the temporal level (S4130).

[0562] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of the various embodiments to be performed on a device or computer, and non-transitory computer-readable media on which such software or instructions are stored and executable on a device or computer.

[0563] [Industrial Applicability] Embodiments of the present disclosure may be used to provide point cloud content, and may be used to encode / decode point cloud data.

[0564] [Claims at the time of international application] [1] 1. A method performed at a receiving device of point cloud data, comprising: obtaining a geometry-based point cloud compression (G-PCC) file containing the point cloud data, the G-PCC file including information on sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels, and information on the temporal levels; and extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on the information on the sample group and the information on the temporal level. [2] The method according to [1], wherein the information on the temporal levels includes number information indicating the number of the temporal levels and temporal level identification information for identifying the temporal levels. [3] The method of [2], wherein the number information further indicates whether the samples in the G-PCC file are temporally scalable. [4] The number information is indicating the number of the temporal levels based on the value of the number information being greater than a first value; The method described in [3], which indicates whether or not there is time scalability for samples in the G-PCC file based on the value of the number information being less than or equal to the first value. [5] The method described in [1], wherein the information for the sample group is present in a geometry track among the multiple tracks that carries geometry data of the point cloud data, based on the fact that the G-PCC file is carried in multiple tracks. [6] The temporal level of one or more samples in an attribute track is the same as the temporal level of a corresponding sample in the geometry track; The method described in [5], wherein the attribute track is a track among the multiple tracks that carries attribute data of the point cloud data. [7] the information on the temporal level includes frame rate information indicating a frame rate of the temporal level; The method according to [1], wherein the extracting step extracts one or more samples from the G-PCC file that correspond to a target frame rate. [8] the information on the temporal level includes frame rate presence information indicating whether the frame rate information is present; The method of claim 7, wherein the frame rate information is included in the information for the temporal level based on the frame rate presence information indicating the presence of the frame rate information. [9] The temporal level is a first temporal level and a second temporal level having a level value greater than the first temporal level; The method of claim 7, wherein the smallest composition time difference between successive samples in the second temporal level is equal to or greater than the smallest composition time difference between successive samples in the first temporal level.

[10] The method described in [1], wherein the information for the temporal level has the same value as the corresponding information in the geometry track referenced by the current track, based on the type of component carried by the current track being an attribute of the point cloud data.

[11] The method of claim 10, wherein the information for the temporal level has the same value as the corresponding information in the geometry track referenced by the current track, based on the type of the component being an attribute of the point cloud data and the type of the sample entry of the current track being a predetermined sample entry type.

[12] The method of claim 11, wherein the predetermined sample entry type includes one or more of a gpc1 sample entry type or a gpcg sample entry type.

[13] A point cloud data receiving device, memory and; at least one processor; The at least one processor: obtaining a geometry-based point cloud compression (G-PCC) file containing the point cloud data, the G-PCC file including information on sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels, and information on the temporal levels; A receiving device that includes a step of extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on information on the sample group and information on the temporal level.

[14] A method performed in a point cloud data transmission device, generating information for sample groups in which geometry-based point cloud compression (G-PCC) samples are grouped based on one or more temporal levels; generating a G-PCC file including the information for the sample groups, the information for the temporal levels, and the point cloud data.

[15] A point cloud data transmission device, memory and; at least one processor; The at least one processor: A transmission device that generates information for sample groups in which G-PCC (geometry-based point cloud compression) samples are grouped based on one or more temporal levels, and generates a G-PCC file including the information for the sample groups, the information for the temporal levels, and the point cloud data.

Claims

1. 1. A method performed at a receiving device of point cloud data, comprising: obtaining a geometry-based point cloud compression (G-PCC) file containing the point cloud data; The G-PCC file includes information about sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels; and The G-PCC file contains information for the temporal level, extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on the information on the sample group and the information on the temporal level; The method, wherein the information on the temporal levels includes number information indicating the number of the temporal levels and temporal level identification information for identifying the temporal levels.

2. The method of claim 1, wherein the number information further indicates whether temporal scalability of the samples in the G-PCC file is supported.

3. The number information is indicating the number of the temporal levels based on the value of the number information being greater than a first value; The method of claim 2 , further comprising indicating whether temporal scalability of samples in the G-PCC file is supported based on the value of the number information being equal to or less than the first value.

4. 2. The method of claim 1, wherein the information about the sample group is present in a geometry track that carries geometry data of the point cloud data among the multiple tracks, based on the fact that the G-PCC file is carried in multiple tracks.

5. The temporal level of one or more samples in an attribute track is the same as the temporal level of a corresponding sample in the geometry track; The method of claim 4 , wherein the attribute track is a track among the multiple tracks that carries attribute data of the point cloud data.

6. The information on the temporal level includes frame rate information indicating a frame rate of the temporal level, The method of claim 1 , wherein the extracting step extracts one or more samples corresponding to a target frame rate from among the samples in the G-PCC file.

7. the information on the temporal level includes frame rate presence information indicating whether the frame rate information is present; The method of claim 6 , wherein the frame rate information is included in the information for the temporal level based on the frame rate presence information indicating the presence of the frame rate information.

8. The temporal level is a first temporal level and a second temporal level having a level value greater than the first temporal level; 7. The method of claim 6, wherein the smallest composition time difference between consecutive samples in the second temporal level is equal to or greater than the smallest composition time difference between consecutive samples in the first temporal level.

9. 2. The method of claim 1, wherein the information for the temporal level has the same value as corresponding information in a geometry track referenced by the current track, based on the type of component carried by the current track being an attribute of the point cloud data.

10. 10. The method of claim 9, wherein the information for the temporal level has the same value as corresponding information in a geometry track referenced by the current track, based on the type of the component being an attribute of the point cloud data and the type of a sample entry of the current track being a predetermined sample entry type.

11. The method of claim 10 , wherein the predetermined sample entry types include one or more of a gpc1 sample entry type or a gpcg sample entry type.

12. A point cloud data receiving device, memory; at least one processor; The at least one processor: obtaining a geometry-based point cloud compression (G-PCC) file containing the point cloud data, the G-PCC file including information on sample groups into which samples in the G-PCC file are grouped based on one or more temporal levels, and information on the temporal levels; extracting one or more samples belonging to a target temporal level from among the samples in the G-PCC file based on the information on the sample group and the information on the temporal level; The information on the temporal levels includes number information indicating the number of the temporal levels and temporal level identification information for identifying the temporal levels.

13. A method performed in a point cloud data transmission device, generating information for sample groups in which geometry-based point cloud compression (G-PCC) samples are grouped based on one or more temporal levels; generating a G-PCC file including information for the sample groups, information for the temporal levels, and the point cloud data; The method, wherein the information on the temporal levels includes number information indicating the number of the temporal levels and temporal level identification information for identifying the temporal levels.

14. A point cloud data transmission device, memory; at least one processor; The at least one processor: generating information for sample groups in which geometry-based point cloud compression (G-PCC) samples are grouped based on one or more temporal levels; generating a G-PCC file including information for the sample groups, information for the temporal levels, and the point cloud data; The information on the temporal levels includes number information indicating the number of the temporal levels and temporal level identification information for identifying the temporal levels.

Citation Information

Patent Citations

  • Point cloud data transmission apparatus, point cloud data transmission method, point cloud data reception apparatus and point cloud data reception method

    US20210029187A1

  • An apparatus, a method and a computer program for volumetric video

    WO2019162564A1

  • An apparatus, a method and a computer program for video encoding and decoding

    WO2020254720A1