Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

WO2026206006A1PCT designated stage Publication Date: 2026-10-01LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004803
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-26
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004803_01102026_PF_FP_ABST
    Figure KR2026004803_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method and device for decoding point cloud data are disclosed. The decoding method according to embodiments comprises the steps of: decoding geometry data of point cloud data within a bitstream; and decoding attribute data of the point cloud data, wherein each of the geometry data and the attribute data is decoded in units of FGSs, the FGSs are respectively mapped to subgroups spatially partitioned within a layer group defined as a group of consecutive tree levels of an occupancy tree, each of the FGSs is identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying a subgroup within the layer group, and the attribute decoding step may comprise a step of generating a plurality of LoDs within a specific FGS.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud data transmission device, point cloud data transmission method, point cloud data reception device and point cloud data reception method

[0001] The embodiments relate to a method and apparatus for processing point cloud content.

[0002] Point cloud content is content represented as a point cloud, which is a set of points belonging to a coordinate system that represents a three-dimensional space or volume. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), XR (Extended Reality), and autonomous driving services. However, representing point cloud content requires tens of thousands to hundreds of thousands of point data points. Therefore, a method is required to efficiently process a vast amount of point data.

[0003] In other words, there is a problem in that a large amount of throughput is required to transmit and receive point cloud data. Therefore, encoding for compression and decoding for decompression are performed during the process of transmitting and receiving point cloud data; however, due to the large size of the point cloud data, the computations are complex and time-consuming.

[0004] The technical problem according to the embodiments is to provide an apparatus and method for efficiently transmitting and receiving point clouds in order to solve the aforementioned problems, etc.

[0005] The technical problem according to the embodiments is to provide an apparatus and method for solving latency and encoding / decoding complexity.

[0006] The technical problem according to the embodiments is to provide an apparatus and method for efficiently performing partial encoding and decoding.

[0007] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that a person skilled in the art can infer based on the entire content described.

[0008] To achieve the above-described purpose and other advantages, the decoding method according to the embodiments may include the step of decoding geometry data of point cloud data within a bitstream and the step of decoding attribute data of the point cloud data.

[0009] According to embodiments, the geometry data and the attribute data are each decoded in units of fine segmentation slices (FGS), and each FGS is mapped to each subgroup spatially segmented within a layer group defined as a group of consecutive tree levels of an occupancy tree, and each FGS can be identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying the subgroup within the layer group.

[0010] According to embodiments, the attribute decoding step may include the step of generating a plurality of Level of Detail (LoD) within a specific FGS.

[0011] According to the embodiments, the LoD generation step includes the step of setting one of the plurality of LoDs as a reference LoD for inter-level prediction when LoD scalability is enabled for the FGS, and the reference LoD setting step may determine the reference LoD as a minimum reference LoD for inter-level prediction when the number of refinement points of the reference LoD is less than the cumulative number of refinement points belonging to LoDs that are finer than the reference LoD.

[0012] According to embodiments, the attribute decoding step may further include the step of performing a prediction reference search by setting a reference point search range for the inter-level prediction based on the LoD with the maximum tree depth among the one or more skipped LoDs when one or more LoDs within the layer group are skipped.

[0013] According to the embodiments, if the minimum reference LoD is greater than a preset threshold, decoding may not be performed for the child subgroups of the subgroup determined by the minimum reference LoD.

[0014] According to embodiments, the attribute decoding step may further include a step of calculating the distance between a predicted target point selected by the predicted reference search and a reference candidate point.

[0015] According to embodiments, the distance calculation step may calculate the distance between the prediction target point and the reference candidate point by aligning the attribute coordinates of the prediction target point and the reference candidate point based on the reference LoD, and then using an L1 norm that applies axis-specific weights to the difference between the aligned attribute coordinates.

[0016] According to embodiments, the decoder includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to decode geometry data of point cloud data within a bitstream and to decode attribute data of the point cloud data.

[0017] According to embodiments, the geometry data and the attribute data are each decoded in units of fine segmentation slices (FGS), and each FGS is mapped to each subgroup spatially segmented within a layer group defined as a group of consecutive tree levels of an occupancy tree, and each FGS can be identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying the subgroup within the layer group.

[0018] According to embodiments, the at least one processor may include a LoD generation unit that generates a plurality of Level of Detail (LoD) within a specific FGS.

[0019] According to the embodiments, when LoD scalability is enabled for the FGS, the LoD generation unit can set one of the plurality of LoDs as a reference LoD for inter-level prediction.

[0020] According to the embodiments, the LoD generation unit may determine the reference LoD as the minimum reference LoD for inter-level prediction when the number of refinement points of the reference LoD is smaller than the cumulative number of refinement points belonging to LoDs that are finer than the reference LoD.

[0021] According to embodiments, the encoding method may include the step of encoding geometry data of point cloud data and the step of encoding attribute data of the point cloud data.

[0022] According to embodiments, the geometry data and the attribute data are each encoded in units of fine segmentation slices (FGS), and each FGS is mapped to each subgroup spatially segmented within a layer group defined as a group of consecutive tree levels of an occupancy tree, and each FGS can be identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying the subgroup within the layer group.

[0023] According to embodiments, the attribute encoding step includes the step of generating a plurality of Level of Detail (LoD) within a specific FGS, and the LoD generation step may include the step of setting one of the plurality of LoDs as a reference LoD for inter-level prediction when LoD scalability is enabled for the FGS.

[0024] According to the embodiments, the reference LoD setting step may determine the reference LoD as the minimum reference LoD for inter-level prediction when the number of refinement points of the reference LoD is smaller than the cumulative number of refinement points belonging to LoDs that are finer than the reference LoD.

[0025] According to embodiments, the encoding device includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to encode geometry data of point cloud data and attribute data of point cloud data.

[0026] According to embodiments, a computer-readable storage medium can store a bitstream generated by the encoding method.

[0027] According to embodiments, the transmission method may include the step of obtaining a bitstream for point cloud data, the bitstream being generated based on the step of encoding geometry data of the point cloud data and the step of encoding attribute data of the point cloud data; and the step of transmitting data including the bitstream.

[0028] The device and method according to the embodiments can provide a high-quality point cloud service.

[0029] The device and method according to the embodiments can achieve various video codec methods.

[0030] The device and method according to the embodiments can provide general-purpose point cloud content, such as autonomous driving services.

[0031] The apparatus and method according to the embodiments can provide improved parallel processing and scalability by performing spatial adaptive partitioning of point cloud data for independent encoding and decoding of point cloud data.

[0032] The apparatus and method according to the embodiments can improve the encoding and decoding performance of a point cloud by dividing the point cloud data into tile and / or slice units to perform encoding and decoding, and by signaling the data necessary for this purpose.

[0033] The device and method according to the embodiments can divide and transmit compressed data according to certain criteria for point cloud data. In addition, when using layered coding, the compressed data can be divided and sent according to the layer. Therefore, the storage and transmission efficiency of the transmitting device can be increased.

[0034] The apparatus and method according to the embodiments can increase the efficiency of scalable attribute coding by eliminating the dependency of attribute coding for a coded geometry layer when using scalable attribute coding.

[0035] The apparatus and method according to the embodiments apply the same LoD generation and NN search performed in the attribute encoder of the transmitting device to the attribute decoder of the receiving device when scalable transmission and / or scalable decoding is performed. In particular, by performing the same position correction and NN search in the attribute encoder of the transmitting device and the attribute decoder of the receiving device, the neighbor candidates considered in the attribute encoder of the transmitting device and the neighbor candidates considered in the attribute decoder of the receiving device do not differ, and the order of nodes does not change. Therefore, there is an effect that decoder prediction errors do not occur in prediction transformations, such as those performing attribute prediction from neighbor nodes.

[0036] The apparatus and method according to the embodiments can determine a minimum reference detail level for inter-level prediction based on the refinement point distribution of each level of detail (LoD) within a fine segmentation slice (FGS) when scalable transmission and / or scalable decoding is performed, and by performing a prediction search based on the minimum reference detail level even in LoD skip situations, it can ensure prediction stability and improve restoration accuracy in a partial decoding environment.

[0037] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments.

[0038] FIG. 1 shows an example of a point cloud content provision system according to embodiments.

[0039] FIG. 2 is a block diagram illustrating a point cloud content provision operation according to embodiments.

[0040] FIG. 3 shows an example of a point cloud encoder according to embodiments.

[0041] FIG. 4 shows examples of octree and occupancy codes according to embodiments.

[0042] Figure 5 shows an example of a point configuration by LoD according to embodiments.

[0043] Figure 6 shows an example of a point configuration by LoD according to embodiments.

[0044] FIG. 7 shows an example of a point cloud decoder according to embodiments.

[0045] FIG. 8 is an example of a transmission device according to embodiments.

[0046] FIG. 9 is an example of a receiving device according to embodiments.

[0047] FIG. 10 shows an example of a structure that can be linked with a point cloud data transmission / reception method / device according to embodiments.

[0048] FIGS. 11 and FIGS. 12 are diagrams illustrating the encoding, transmission, and decoding processes of point cloud data according to embodiments.

[0049] FIG. 13 is a diagram showing examples of a device and method for transmitting and receiving point cloud data according to embodiments.

[0050] FIG. 14 is a diagram showing an example of a layer-based point cloud data configuration according to embodiments.

[0051] FIG. 15(a) shows the bitstream structure of geometry data according to embodiments, and FIG. 15(b) shows the bitstream structure of attribute data according to embodiments.

[0052] FIG. 16 is a diagram showing an example of the configuration of a bitstream for dividing and transmitting the bitstream in layers (or LoDs) according to embodiments.

[0053] FIG. 17 illustrates an example of a bitstream alignment method when a geometry bitstream and an attribute bitstream are multiplexed into a single bitstream according to embodiments.

[0054] FIG. 18 shows another example of a bitstream alignment method when a geometry bitstream and an attribute bitstream are multiplexed into a single bitstream according to embodiments.

[0055] FIGS. 19(a) to 19(c) are drawings showing examples of symmetric geometry-attribute selection according to embodiments.

[0056] FIGS. 20(a) to 20(c) are drawings showing examples of asymmetric geometry-attribute selection according to embodiments.

[0057] FIGS. 21(a) to 21(c) illustrate examples of methods for configuring slices containing point cloud data according to embodiments.

[0058] FIG. 22 shows a geometry coding layer structure according to embodiments.

[0059] FIG. 23 is a diagram showing the layer group and subgroup structure according to embodiments.

[0060] FIGS. 24(a) to 24(c) are drawings illustrating the representation of layer group-based point cloud data according to embodiments.

[0061] FIG. 25 shows a point cloud data transmission / reception device / method according to embodiments.

[0062] FIG. 26 is a diagram showing an example of performing NN search using full layer geometry information according to embodiments.

[0063] FIG. 27 is a diagram showing an example of performing an NN search on a downsampled geometry grid according to embodiments.

[0064] FIG. 28 is a diagram showing an example of performing an NN search on a position-corrected geometry grid according to embodiments.

[0065] FIGS. 29(a) to 29(c) are drawings showing examples of performing NN search on a geometry grid according to embodiments.

[0066] FIG. 30 is a drawing showing another example of a point cloud transmitting device according to embodiments.

[0067] FIG. 31 is a detailed block diagram of the LoD generation section of an attribute encoder according to embodiments.

[0068] FIG. 32 is a drawing showing another example of a point cloud receiving device according to embodiments.

[0069] FIG. 33 shows an example of a bitstream structure of point cloud data for transmission / reception according to embodiments.

[0070] FIGS. 34a to 34c are drawings showing an example of the syntax structure of a sequence parameter set (seq_parameter_set()) (SPS) according to the embodiments.

[0071] FIGS. 35a to 35c are drawings showing an example of the syntax structure of an attribute parameter set (attr_parameter_set()) (APS) according to the embodiments.

[0072] FIGS. 36(a) and FIGS. 36(b) are diagrams showing the relationship between layer groups, subgroups, and FGS according to embodiments.

[0073] FIGS. 37(a) to 37(c) are drawings illustrated to compare the distribution of the number of attribute points by detail level in an environment where LoD scalability is applied.

[0074] FIG. 38 is a graph showing an enlarged distribution of the number of points in a specific detail level range when LoD scalability according to the embodiments is applied.

[0075] Figure 39 is a diagram showing an example of compressing and providing geometry and attributes of point cloud data.

[0076] FIG. 40 is a diagram showing another example of compressing and servicing the geometry and attributes of point cloud data according to embodiments.

[0077] FIG. 41 is a diagram showing another example of compressing and providing geometry and attributes of point cloud data according to embodiments.

[0078] FIG. 42 shows a flowchart of a point cloud data transmission method according to embodiments.

[0079] FIG. 43 shows a flowchart of a method for receiving point cloud data according to embodiments.

[0080] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments may be practiced without these details.

[0081] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.

[0082] FIG. 1 shows an example of a point cloud content provision system according to embodiments.

[0083] The point cloud content providing system illustrated in FIG. 1 may include a transmission device (10000) and a reception device (10004). The transmission device (10000) and the reception device (10004) can communicate via wired or wireless means to transmit and receive point cloud data.

[0084] A transmission device (10000) according to embodiments can acquire, process, and transmit point cloud video (or point cloud content). According to embodiments, the transmission device (10000) may include a fixed station, a base transceiver system (BTS), a network, an Artificial Intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or server, etc. Additionally, according to embodiments, the transmission device (10000) may include a device that communicates with a base station and / or other wireless devices using wireless access technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an Internet of Things (IoT) device, an AI device / server, etc.

[0085] A transmission device (10000) according to embodiments includes a point cloud video acquisition unit (10001), a point cloud video encoder (10002), and / or a transmitter (or communication module), 10003.

[0086] A point cloud video acquisition unit (10001) according to the embodiments acquires a point cloud video through processing steps such as capture, synthesis, or generation. The point cloud video is a point cloud content represented as a point cloud, which is a set of points located in a three-dimensional space, and may be referred to as point cloud video data, etc. The point cloud video according to the embodiments may include one or more frames. A frame represents a still image / picture. Accordingly, the point cloud video may include a point cloud image / frame / picture and may be referred to as any one of a point cloud image, a frame, and a picture.

[0087] A point cloud video encoder (10002) according to the embodiments encodes the obtained point cloud video data. The point cloud video encoder (10002) can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiments may include Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding or next-generation coding. Furthermore, the point cloud compression coding according to the embodiments is not limited to the embodiments described above. The point cloud video encoder (10002) can output a bitstream containing the encoded point cloud video data. The bitstream may include not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0088] A transmitter (10003) according to the embodiments transmits a bitstream containing encoded point cloud video data. The bitstream according to the embodiments is encapsulated into a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not illustrated in the drawings, the transmission device (10000) may include an encapsulation unit (or encapsulation module) that performs an encapsulation operation. Additionally, according to the embodiments, the encapsulation unit may be included in the transmitter (10003). According to the embodiments, the file or segment may be transmitted to a receiving device (10004) via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter (10003) according to the embodiments can communicate wired or wirelessly with the receiving device (10004) (or receiver (10005)) via a network such as 4G, 5G, or 6G. Additionally, the transmitter (10003) can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). Additionally, the transmission device (10000) can transmit encapsulated data according to an on-demand method.

[0089] A receiving device (10004) according to embodiments includes a receiver (10005), a point cloud video decoder (10006), and / or a renderer (10007). According to embodiments, the receiving device (10004) may include a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Thing) device, AI device / server, etc., that communicates with a base station and / or other wireless device using wireless access technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).

[0090] A receiver (10005) according to the embodiments receives a bitstream containing point cloud video data or a file / segment containing the bitstream from a network or a storage medium. The receiver (10005) can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver (10005) according to the embodiments can output a bitstream by decapsulating the received file / segment. Additionally, according to the embodiments, the receiver (10005) may include a decapsulation unit (or decapsulation module) for performing a decapsulation operation. Additionally, the decapsulation unit may be implemented as an element (or component) separate from the receiver (10005).

[0091] A point cloud video decoder (10006) decodes a bitstream containing point cloud video data. The point cloud video decoder (10006) can decode the point cloud video data according to the way the point cloud video data is encoded (e.g., the reverse process of the operation of a point cloud video encoder (10002)). Accordingly, the point cloud video decoder (10006) can decode the point cloud video data by performing point cloud decompression coding, which is the reverse process of point cloud compression. Point cloud decompression coding includes G-PCC coding.

[0092] The renderer (10007) renders the decoded point cloud video data. In one embodiment, the renderer (10007) can render the decoded point cloud video data according to a viewport, etc. The renderer (10007) can render not only the point cloud video data but also audio data to output point cloud content. According to embodiments, the renderer (10007) may include a display for displaying the point cloud content. According to embodiments, the display may not be included in the renderer (10007) but may be implemented as a separate device or component.

[0093] The arrows indicated by dotted lines in the drawing represent the transmission path of feedback information obtained from the receiving device (10004). The feedback information is information intended to reflect interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is content for a service requiring interaction with a user (e.g., autonomous driving service, etc.), the feedback information may be transmitted to the content transmitting side (e.g., the transmitting device (10000)) and / or the service provider. Depending on the embodiments, the feedback information may be used in the receiving device (10004) as well as the transmitting device (10000), or it may not be provided.

[0094] Head orientation information according to the embodiments may refer to information regarding the user's head position, direction, angle, movement, etc. The receiving device (10004) according to the embodiments may calculate viewport information based on head orientation information. Viewport information is information about the area of ​​the point cloud video that the user is looking at (i.e., the area the user is currently looking at). That is, viewport information is information about the area the user is currently looking at within the point cloud video. In other words, the viewport or viewport area may refer to the area the user is looking at in the point cloud video. And the viewpoint is the point the user is looking at in the point cloud video, and may refer to the exact center point of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape occupied by that area may be determined by the FOV (Field Of View). Therefore, the receiving device (10004) may extract viewport information based on the vertical or horizontal FOV supported by the device in addition to head orientation information. Additionally, the receiving device (10004) can perform gaze analysis, etc. based on head orientation information and / or viewport information to determine the user's point cloud video consumption method, the point cloud video area the user gazes at, the gaze time, etc. According to embodiments, the receiving device (10004) can transmit feedback information including the gaze analysis results to the transmitting device (10000). According to embodiments, a device such as a VR / XR / AR / MR display can extract a viewport area based on the user's head position / direction, the vertical or horizontal FOV supported by the device, etc. According to embodiments, the head orientation information and viewport information may be referred to as feedback information, signaling information, or metadata.

[0095] Feedback information according to the embodiments may be obtained during the rendering and / or display process. Feedback information according to the embodiments may be obtained by one or more sensors included in the receiving device (10004). Additionally, according to the embodiments, feedback information may be obtained by the renderer (10007) or a separate external element (or device, component, etc.). The dotted line in FIG. 1 indicates the process of transmitting feedback information obtained from the renderer (10007). The feedback information may be transmitted to the transmitting side, as well as consumed at the receiving side. That is, the point cloud content providing system may process point cloud data (encoding / decoding / rendering) based on the feedback information. For example, the point cloud video decoder (10006) and the renderer (10007) may use the feedback information, namely head orientation information and / or viewport information, to preferentially decode and render only the point cloud video for the area currently being viewed by the user.

[0096] Additionally, the receiving device (10004) can transmit feedback information to the transmitting device (10000). The transmitting device (10000) (or the point cloud video encoder (10002)) can perform an encoding operation based on the feedback information. Thus, the point cloud content providing system can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information without processing (encoding / decoding) all point cloud data, and provide point cloud content to the user.

[0097] According to embodiments, the transmission device (10000) may be referred to as an encoder, transmission device, transmitter, transmission system, etc., and the receiving device (10004) may be referred to as a decoder, receiving device, receiver, receiving system, etc.

[0098] Point cloud data processed in the point cloud content providing system of FIG. 1 according to embodiments (processed through a series of processes of acquisition / encoding / transmission / decoding / rendering) may be referred to as point cloud content data or point cloud video data. According to embodiments, point cloud content data may be used as a concept including metadata or signaling information related to point cloud data.

[0099] The elements of the point cloud content delivery system illustrated in FIG. 1 may be implemented as hardware, software, processors, and / or combinations thereof.

[0100] FIG. 2 is a block diagram illustrating a point cloud content provision operation according to embodiments.

[0101] The block diagram of FIG. 2 illustrates the operation of the point cloud content provision system described in FIG. 1. As described above, the point cloud content provision system can process point cloud data based on point cloud compression coding (e.g., G-PCC).

[0102] A point cloud content providing system according to the embodiments (e.g., a point cloud transmission device (10000) or a point cloud video acquisition unit (10001)) can acquire a point cloud video (20000). The point cloud video is represented as a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to the embodiments may include a Ply (Polygon File format or the Stanford Triangle format) file. If the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as the geometry and / or attributes of the points. The geometry includes the positions of the points. The position of each point may be represented by parameters (e.g., values ​​of the X-axis, Y-axis, and Z-axis, respectively) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of XYZ axes). Attributes include attributes of points (e.g., texture information, color (YCbCr or RGB), reflectance (r), transparency, etc. of each point). A point has one or more attributes (or properties). For example, a point may have one attribute which is color, or two attributes which are color and reflectance. According to embodiments, geometry may be referred to as positions, geometry information, geometry data, etc., and attributes may be referred to as attributes, attribute information, attribute data, etc. Additionally, a point cloud content providing system (e.g., a point cloud transmission device (10000) or a point cloud video acquisition unit (10001)) may obtain point cloud data from information related to the acquisition process of point cloud video (e.g., depth information, color information, etc.).

[0103] A point cloud content providing system according to embodiments (e.g., a transmission device (10000) or a point cloud video encoder (10002)) can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression coding. As described above, point cloud data may include geometry and attributes of points. Accordingly, the point cloud content providing system can output a geometry bitstream by performing geometry encoding to encode geometry. The point cloud content providing system can output an attribute bitstream by performing attribute encoding to encode attributes. According to embodiments, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and attribute bitstream according to embodiments can be multiplexed and output as a single bitstream. The bitstream according to the embodiments may further include signaling information related to geometry encoding and attribute encoding.

[0104] A point cloud content providing system according to embodiments (e.g., a transmission device (10000) or a transmitter (10003)) can transmit encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data can be represented as a geometry bitstream and an attribute bitstream. Additionally, the encoded point cloud data can be transmitted in the form of a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and attribute encoding). Additionally, the point cloud content providing system can encapsulate the bitstream transmitting the encoded point cloud data and transmit it in the form of a file or segment.

[0105] A point cloud content providing system according to embodiments (e.g., a receiving device (10004) or a receiver (10005)) can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., a receiving device (10004) or a receiver (10005)) can demultiplex the bitstream.

[0106] A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode encoded point cloud data (e.g., a geometry bitstream, an attribute bitstream) transmitted as a bitstream. A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode point cloud video data based on signaling information related to the encoding of point cloud video data included in the bitstream. A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode the geometry bitstream to restore the positions (geometry) of the points. A point cloud content providing system can decode the attribute bitstream based on the restored geometry to restore the attributes of the points. A point cloud content delivery system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can restore a point cloud video based on positions according to the restored geometry and decoded attributes.

[0107] A point cloud content providing system according to embodiments (e.g., a receiving device (10004) or a renderer (10007)) can render decoded point cloud data (20004). The point cloud content providing system (e.g., a receiving device (10004) or a renderer (10007)) can render geometry and attributes decoded through a decoding process according to various rendering methods. Points of the point cloud content may be rendered as vertices having a certain thickness, cubes having a specific minimum size with the vertex location as the center, or circles with the vertex location as the center, etc. All or part of the rendered point cloud content is provided to a user through a display (e.g., a VR / AR display, a general display, etc.).

[0108] A point cloud content providing system (e.g., a receiving device (10004)) according to the embodiments can obtain feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. Since the feedback information and the operation of the point cloud content providing system according to the embodiments are the same as the feedback information and operation described in FIG. 1, a detailed description is omitted.

[0109] FIG. 3 shows an example of a point cloud encoder according to embodiments.

[0110] FIG. 3 shows an example of the point cloud video encoder (10002) of FIG. 1. The point cloud encoder reconstructs point cloud data (e.g., positions and / or attributes of points) and performs encoding operations to adjust the quality of point cloud content (e.g., lossless, lossy, near-lossless) according to network conditions or applications. If the total size of the point cloud content is large (e.g., point cloud content of 60 Gbps in the case of 30 fps), the point cloud content delivery system may not be able to stream the content in real time. Therefore, the point cloud content delivery system may reconstruct the point cloud content based on a maximum target bitrate to provide it according to the network environment.

[0111] As described in FIGS. 1 and 2, the point cloud encoder can perform geometry encoding and attribute encoding. Geometry encoding is performed before attribute encoding.

[0112] The point cloud encoder according to the embodiments comprises a coordinate system transformation unit (Transformation Coordinates, 30000), a quantization unit (Quantize and Remove Points (Voxelize), 30001), an octree analysis unit (Analyze Octree, 30002), a surface approximation analysis unit (Analyze Surface Approximation, 30003), an arithmetic encoder (Arithmetic Encode, 30004), a geometry reconstruction unit (Reconstruct Geometry, 30005), a color transformation unit (Transform Colors, 30006), an attribute transformation unit (Transfer Attributes, 30007), a RAHT transformation unit (30008), a LoD generation unit (Generated LoD, 30009), a lifting transformation unit (Lifting) (30010), and a coefficient quantization unit (Quantize Coefficients, 30011). It includes an and / or arithmetic encoder (30012). In the point cloud encoder of FIG. 3, the coordinate system transformation unit (30000), quantization unit (30001), octree analysis unit (30002), surface approximation analysis unit (30003), arithmetic encoder (30004), and geometry reconstruction unit (30005) can be grouped and referred to as a geometry encoder. Also, the color conversion unit (30006), attribute conversion unit (30007), RAHT conversion unit (30008), LoD generation unit (30009), lifting conversion unit (30010), coefficient quantization unit (30011), and / or arithmetic encoder (30012) can be grouped and referred to as an attribute encoder.

[0113] The coordinate system transformation unit (30000), quantization unit (30001), octree analysis unit (30002), surface approximation analysis unit (30003), arismetic encoder (30004), and geometry reconstruction unit (30005) can perform geometry encoding. Geometry encoding according to the embodiments may include octree geometry coding, direct coding, trisoup geometry encoding, and entropy encoding. Direct coding and trisoup geometry encoding are applied optionally or in combination. Additionally, geometry encoding is not limited to the above examples.

[0114] As illustrated in the drawings, the coordinate system conversion unit (30000) according to the embodiments receives positions and converts them into a coordinate system. For example, the positions can be converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to the embodiments may be referred to as geometry information.

[0115] The quantization unit (30001) according to the embodiments quantizes the geometry. For example, the quantization unit (30001) can quantize points based on the minimum position values ​​of all points (e.g., minimum values ​​on each axis for the X-axis, Y-axis, and Z-axis). The quantization unit (30001) performs a quantization operation to find the nearest integer value by multiplying the difference between the minimum position value and the position value of each point by a preset quantization scale value and then performing rounding down or rounding up. Thus, one or more points may have the same quantized position (or position value). The quantization unit (30001) according to the embodiments performs voxelization based on the quantized positions to reconstruct the quantized points. Just as the minimum unit containing 2D image / video information is a pixel, the points of the point cloud content (or 3D point cloud video) according to the embodiments may be contained in one or more voxels. A voxel is a combination of volume and pixel, and refers to a three-dimensional cubic space that is generated when a three-dimensional space is divided into units (unit=1.0) based on axes representing the three-dimensional space (e.g., X-axis, Y-axis, Z-axis). The quantization unit (40001) can match groups of points in the three-dimensional space to voxels. According to embodiments, a single voxel may contain only one point. According to embodiments, a single voxel may contain one or more points. In addition, to represent a single voxel as a single point, the position of the center of the voxel can be set based on the positions of one or more points included in the voxel. In this case, the attributes of all positions included in the voxel can be combined and assigned to the voxel.

[0116] The octree analysis unit (30002) according to the embodiments performs octree geometry coding (or octree coding) to represent the voxels in an octree structure. The octree structure represents points matched to the voxels based on an octree structure.

[0117] The surface approximation analysis unit (30003) according to the embodiments can analyze and approximate an octree. The octree analysis and approximation according to the embodiments is a process of analyzing to voxelize an area containing multiple points in order to efficiently provide octree and voxelization.

[0118] An arithmetic encoder (30004) according to the embodiments entropy-encodes an octree and / or an approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.

[0119] The color conversion unit (30006), attribute conversion unit (30007), RAHT conversion unit (30008), LoD generation unit (30009), lifting conversion unit (30010), coefficient quantization unit (30011) and / or arismetic encoder (30012) perform attribute encoding. As described above, a point may have one or more attributes. The attribute encoding according to the embodiments is applied equally to the attributes of a point. However, if a single attribute (e.g., color) includes one or more elements, independent attribute encoding is applied to each element. The attribute encoding according to the embodiments may include color conversion coding, attribute conversion coding, Region Adaptive Hierarchial Transform (RAHT) coding, prediction transformation (Interpolaration-based hierarchical nearest-neighbour prediction-Prediction Transform) coding, and lifting transformation (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-described RAHT coding, prediction transformation coding, and lifting transformation coding may be used optionally, or a combination of one or more of the codings may be used. Furthermore, the attribute encoding according to the embodiments is not limited to the examples described above.

[0120] The color conversion unit (30006) according to the embodiments performs color conversion coding that converts color values ​​(or textures) included in attributes. For example, the color conversion unit (30006) can convert the format of color information (e.g., convert from RGB to YCbCr). The operation of the color conversion unit (30006) according to the embodiments may be applied optionally depending on the color values ​​included in attributes.

[0121] The geometry reconstruction unit (30005) according to the embodiments reconstructs (decompresses) an octree and / or an approximated octree. The geometry reconstruction unit (30005) reconstructs an octree / voxel based on the results of analyzing the distribution of points. The reconstructed octree / voxel may be referred to as the reconstructed geometry (or restored geometry).

[0122] The attribute transformation unit (30007) according to the embodiments performs attribute transformation that transforms attributes based on positions where geometry encoding has not been performed and / or reconstructed geometry. As described above, since attributes are dependent on geometry, the attribute transformation unit (30007) can transform attributes based on reconstructed geometry information. For example, the attribute transformation unit (30007) can transform the attributes of a point at a position based on the position value of a point included in a voxel. As described above, when the position of the center point of a voxel is set based on the positions of one or more points included in a voxel, the attribute transformation unit (30007) transforms the attributes of one or more points. When trisoop geometry encoding is performed, the attribute conversion unit (30007) can convert attributes based on the trisoop geometry encoding.

[0123] The attribute transformation unit (30007) can perform attribute transformation by calculating the average value of attributes or attribute values ​​(e.g., the color or reflectance of each point) of neighboring points within a specific location / radius from the position (or position value) of the center point of each voxel. The attribute transformation unit (30007) can apply a weight based on the distance from the center point to each point when calculating the average value. Thus, each voxel has a position and a calculated attribute (or attribute value).

[0124] The attribute conversion unit (30007) can search for neighboring points within a specific location / radius from the position of the center point of each voxel based on a KD tree or a Molton code. A KD tree is a binary search tree that supports a data structure capable of managing points based on their positions to enable rapid Nearest Neighbor Search (NNS). A Molton code is generated by representing the coordinate values ​​(e.g., (x, y, z)) representing the 3D positions of all points as bit values ​​and mixing the bits. For example, if the coordinate values ​​representing the position of a point are (5, 9, 1), the bit values ​​of the coordinate values ​​are (0101, 1001, 0001). When the bit values ​​are mixed according to the bit indices in the order of z, y, and x, it becomes 010001000111. When this value is represented in decimal, it becomes 1095. That is, the Molton code value of the point with coordinates (5, 9, 1) is 1095. The attribute transformation unit (30007) sorts the points based on the Molton code value and can perform shortest neighbor point search (NNS) through a depth-first traversal process. After the attribute transformation operation, if shortest neighbor point search (NNS) is required in other transformation processes for attribute coding, a KD tree or Molton code is utilized.

[0125] As illustrated in the drawing, the converted attributes are input to the RAHT conversion unit (30008) and / or the LoD generation unit (30009).

[0126] The RAHT transformation unit (30008) according to the embodiments performs RAHT coding to predict attribute information based on reconstructed geometry information. For example, the RAHT transformation unit (30008) can predict attribute information of a node at an upper level of the octree based on attribute information associated with a node at a lower level of the octree.

[0127] The LoD generation unit (30009) according to the embodiments generates a Level of Detail (LoD) to perform predictive transformation coding. The LoD according to the embodiments represents the degree of detail of the point cloud content, and indicates that the smaller the LoD value, the lower the detail of the point cloud content, and the larger the LoD value, the higher the detail of the point cloud content. Points can be classified according to the LoD.

[0128] The lifting transformation unit (30010) according to the embodiments performs lifting transformation coding that transforms the attributes of the point cloud based on weights. As described above, the lifting transformation coding may be applied optionally.

[0129] The coefficient quantization unit (30011) according to the embodiments quantizes attribute-coded attributes based on coefficients.

[0130] An arismetic encoder (30012) according to the embodiments encodes quantized attributes based on arismetic coding.

[0131] The elements of the point cloud encoder of FIG. 3 may be implemented in hardware, software, firmware, or a combination thereof, comprising one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud encoder of FIG. 3 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud encoder of FIG. 3. One or more memories according to the embodiments may include high-speed random access memory and may include non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0132] FIG. 4 shows examples of octree and occupancy codes according to embodiments.

[0133] As described in FIGS. 1 to 3, a point cloud content providing system (point cloud video encoder (10002)) or a point cloud encoder (e.g., an octree analysis unit (30002)) performs octree geometry coding based on an octree structure (or octree coding) to efficiently manage the area and / or position of a voxel.

[0134] The top of FIG. 4 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiments is represented by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). The octree structure has two poles (0,0,0) and (2 d , 2 d , 2 d It is generated by recursively subdividing the bounding box (cubical axis-aligned bounding box) defined by ). 2d can be set to the value that constitutes the smallest bounding box enclosing all points of the point cloud content (or point cloud video). d represents the depth of the octree. The value of d is determined according to the following equation. In the equation below, (x int n , y int n , z int n ) represents the positions (or position values) of quantized points.

[0135]

[0136] As illustrated in the middle of the top of Fig. 4, the entire three-dimensional space can be divided into eight spaces according to the division. Each divided space is represented as a cube having six faces. As illustrated in the right of the top of Fig. 4, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Thus, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented as cubes having six faces. This division method is applied until the leaf nodes of the octree become voxels.

[0137] The bottom of Fig. 4 shows the occupancy code of an octree. The occupancy code of an octree is generated to indicate whether each of the eight partitioned spaces resulting from the partitioning of a single space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node represents the occupancy of the partitioned space, and the child node has a value of 1 bit. Thus, the occupancy code is represented as an 8-bit code. That is, if the space corresponding to the child node contains at least one point, the node has a value of 1. If the space corresponding to the child node does not contain a point (empty), the node has a value of 0. Since the occupancy code shown in Fig. 4 is 00100001, it indicates that the spaces corresponding to the 3rd and 8th child nodes among the eight child nodes each contain at least one point. As illustrated in the drawing, the 3rd child node and the 8th child node each have 8 child nodes, and each child node is represented by an 8-bit Occupancy code. The drawing indicates that the Occupancy code of the 3rd child node is 10000111 and the Occupancy code of the 8th child node is 01001111. A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30004)) can entropy-encode the Occupancy code. Additionally, to increase compression efficiency, the point cloud encoder can intra- / inter-encode the Occupancy code. A receiving device according to the embodiments (e.g., a receiving device (10004) or a point cloud video decoder (10006)) reconstructs the octree based on the Occupancy code.

[0138] A point cloud encoder according to the embodiments (e.g., the point cloud encoder of FIG. 3, or the octree analysis unit (30002)) can perform voxelization and octree coding to store the positions of the points. However, since points in a three-dimensional space are not always evenly distributed, there may be specific areas where few points exist. Therefore, performing voxelization on the entire three-dimensional space is inefficient. For example, if there are almost no points in a specific area, there is no need to perform voxelization up to that area.

[0139] Accordingly, the point cloud encoder according to the embodiments can perform direct coding, which directly codes the positions of points included in a specific region (or nodes excluding leaf nodes of an octree) without performing voxelization on the aforementioned specific region. The coordinates of the points directly coded according to the embodiments are referred to as the Direct Coding Mode (DCM). Additionally, the point cloud encoder according to the embodiments can perform trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) based on voxels using a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct coding and trisoup geometry encoding according to the embodiments may be performed optionally. In addition, direct coding and trisoop geometry encoding according to the embodiments can be performed in combination with octree geometry coding (or octree coding).

[0140] To perform direct coding, the option to use direct mode for applying direct coding must be enabled, the node to which direct coding is to be applied must not be a leaf node, and there must be points within a specific node that are below a threshold. In addition, the total number of points subject to direct coding must not exceed a preset threshold. If the above conditions are satisfied, the point cloud encoder (or arismetic encoder (30004)) according to the embodiments can entropy-code the positions (or position values) of the points.

[0141] A point cloud encoder according to the embodiments (e.g., a surface approximation analysis unit (30003)) can determine a specific level of an octree (where the level is smaller than the depth d of the octree) and, starting from that level, perform trisoop geometry encoding to reconstruct the position of points within a node area based on voxels using a surface model (trisoop mode). The point cloud encoder according to the embodiments can specify the level to which trisoop geometry encoding is applied. For example, if the specified level is equal to the depth of the octree, the point cloud encoder does not operate in trisoop mode. That is, the point cloud encoder according to the embodiments can operate in trisoop mode only when the specified level is smaller than the depth value of the octree. A three-dimensional cubic area of ​​nodes at the specified level according to the embodiments is referred to as a block. A block may include one or more voxels. A block or a voxel may correspond to a brick. Within each block, geometry is represented as a surface. A surface according to the embodiments may intersect each edge of the block at most once.

[0142] Since one block has 12 edges, there are at least 12 intersection points within one block. Each intersection point is referred to as a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to that edge among all blocks sharing that edge. An occupied voxel according to the embodiments means a voxel containing a point. The position of a vertex detected along an edge is the average position along the edge of all voxels adjacent to that edge among all blocks sharing that edge.

[0143] When a vertex is detected, the point cloud encoder according to the embodiments includes the starting point (x, y, z) of the edge and the direction vector of the edge ( x, y, z), vertex position values ​​(relative position values ​​within the edge) can be entropy-coded. When trisoop geometry encoding is applied, a point cloud encoder according to the embodiments (e.g., geometry reconstruction unit (30005)) can generate restored geometry (reconstructed geometry) by performing triangle reconstruction, up-sampling, and voxelization processes.

[0144] The vertices located on the edges of the block determine the surface passing through the block. The surface according to the embodiments is a non-planar polygon. The triangle reconstruction process reconstructs the surface represented by triangles based on the edge start point, the edge direction vector, and the vertex position value. The triangle reconstruction process is as follows: ① calculate the centroid value of each vertex, ② subtract the centroid value from each vertex value, ③ square the result, and add all the result together.

[0145]

[0146] Then, the minimum value of the sum is calculated, and a projection process is performed along the axis where the minimum value is located. For example, if the x-element is at its minimum, each vertex is projected along the x-axis relative to the center of the block and projected onto the (y, z) plane. If the value obtained by projecting onto the (y, z) plane is (ai, bi), the θ value is calculated using atan2(bi, ai), and the vertices are aligned based on the θ value. The table below shows the combinations of vertices to generate triangles depending on the number of vertices. The vertices are sorted in order from 1 to n. Table 1 below indicates that for four vertices, two triangles can be formed depending on the combination of vertices. The first triangle is composed of the 1st, 2nd, and 3rd vertices among the aligned vertices, and the second triangle can be composed of the 3rd, 4th, and 1st vertices among the aligned vertices.

[0147] Table 1. Triangles formed from vertices ordered 1,… , n

[0148] nTriangles3(1,2,3)4(1,2,3), (3,4,1)5(1,2,3), (3,4,5), (5,1,3)6(1,2,3), (3,4,5), (5,6,1), (1,3,5)7(1,2,3), (3,4,5), (5,6,7), (7,1,3), (3,5,7)8(1,2,3), (3,4,5), (5,6,7), (7,8,1), (1,3,5), (5,7,1)9(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,1,3), (3,5,7), (7,9,3)10(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,1), (1,3,5), (5,7,9), (9,1,5)11(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,1,3), (3,5,7), (7,9,11), (11,3,7)12(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,12,1), (1,3,5), (5,7,9), (9,11,1), (1,5,9)

[0149] The upsampling process is performed to voxelize by adding intermediate points along the edges of the triangle. Additional points are generated based on the upsampling factor value and the width of the block. The additional points are referred to as refined vertices. A point cloud encoder according to the embodiments can voxelize the refined vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized positions (or position values).

[0150] Figure 5 shows an example of a point configuration by LoD according to embodiments.

[0151] As described in FIGS. 1 to 4, the encoded geometry is reconstructed (decompressed) before attribute encoding is performed. When direct coding is applied, the geometry reconstruction operation may include changing the arrangement of the direct-coded points (e.g., placing the direct-coded points at the front of the point cloud data). When trisoop geometry encoding is applied, the geometry reconstruction process involves triangle reconstruction, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.

[0152] A point cloud encoder (e.g., an LoD generation unit (30009)) can reorganize points by LoD. The drawing shows point cloud content corresponding to the LoD. The left side of the drawing shows the original point cloud content. The second figure from the left of the drawing shows the distribution of points in the lowest LoD, and the rightmost figure of the drawing shows the distribution of points in the highest LoD. That is, points in the lowest (or coarse) LoD are distributed sparsely, and points in the highest (or finest) LoD are distributed finely. That is, according to the direction of the arrow indicated at the bottom of the drawing, as the LoD increases, the spacing (or distance) between points becomes shorter.

[0153] Figure 6 shows an example of a point configuration by LoD according to embodiments.

[0154] As described in FIGS. 1 to 5, a point cloud content providing system or a point cloud encoder (e.g., a point cloud video encoder (10002), the point cloud encoder of FIG. 3, or an LoD generation unit (30009)) can generate a LoD. The LoD is generated by reorganizing points into a set of refinement levels according to a set LoD distance value (or a set of Euclidean distances). The LoD generation process is performed in a point cloud decoder as well as a point cloud encoder.

[0155] The top of Fig. 6 shows examples of points (P0 to P9) of point cloud content distributed in three-dimensional space. The Original Order in Fig. 6 represents the order of points P0 to P9 prior to LoD generation. The LoD-based Order in Fig. 6 represents the order of points following LoD generation. Points are rearranged by LoD. Additionally, higher LoDs include points belonging to lower LoDs. As illustrated in Fig. 6, LoD0 includes P0, P5, P4, and P2. LoD1 includes the points of LoD0 and P1, P6, and P3. LoD2 includes the points of LoD0, the points of LoD1, and P9, P8, and P7.

[0156] As described in FIG. 3, the point cloud encoder according to the embodiments can perform predictive transform coding, lifting transform coding, and RAHT transform coding selectively or in combination.

[0157] The point cloud encoder according to the embodiments can generate predictors for points and perform predictive transformation coding to set the predicted attribute (or predicted attribute value) of each point. That is, N predictors can be generated for N points. The predictor according to the embodiments can calculate a weight (=1 / distance) value based on the LoD value of each point, indexing information for neighboring points within a set distance for each LoD, and the distance value to the neighboring points.

[0158] According to the embodiments, the predicted attribute (or attribute value) is set as the average value of the values ​​obtained by multiplying the attributes (or attribute values, e.g., color, reflectance, etc.) of neighboring points set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder according to the embodiments (e.g., coefficient quantization unit (30011)) can quantize and inverse quantize the residual values ​​(which may be referred to as residual attributes, residual attribute values, attribute prediction residual values, etc.) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. The quantization process is as shown in Tables 2 and 3 below.

[0159] int PCCQuantization(int value, int quantStep) {if( value >=0) {return floor(value / quantStep + 1.0 / 3.0);} else {return -floor(-value / quantStep + 1.0 / 3.0);}}

[0160] int PCCInverseQuantization(int value, int quantStep) {if( quantStep ==0) {return value;} else {return value * quantStep;}}

[0161] A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30012)) can entropy-code the quantized and inversely quantized residual values ​​as described above when there are neighboring points in the predictor of each point. A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30012)) can entropy-code the attributes of the corresponding point without performing the process described above when there are no neighboring points in the predictor of each point.

[0162] A point cloud encoder according to the embodiments (e.g., a lifting transformation unit (30010)) can perform lifting transformation coding by generating a predictor for each point, setting the calculated LoD in the predictor, registering neighboring points, and setting weights based on the distance to the neighboring points. The lifting transformation coding according to the embodiments is similar to the prediction transformation coding described above, but differs in that weights are cumulatively applied to attribute values. The process of cumulatively applying weights to attribute values ​​according to the embodiments is as follows.

[0163] 1) Create an array QW (QuantizationWight) to store the weight values ​​of each point. The initial value of all elements in QW is 1.0. Add the value obtained by multiplying the current point's predictor weight by the QW value of the predictor index of the neighboring node registered in the predictor.

[0164] 2) Lift prediction process: To calculate the predicted attribute value, the value obtained by multiplying the point's attribute value by a weight is subtracted from the existing attribute value.

[0165] 3) Create temporary arrays named updateweight and update, and initialize the temporary arrays to 0.

[0166] 4) For all predictors, the calculated weight is additionally multiplied by the weight stored in the QW corresponding to the predictor index, and the resulting weight is accumulated in the update weight array with the neighbor node index. In the update array, the value obtained by multiplying the attribute value of the neighbor node index by the calculated weight is accumulated.

[0167] 5) Lift update process: For all predictors, the attribute value of the update array is divided by the weight value of the update weight array at the predictor index, and the original attribute value is added back to the divided value.

[0168] 6) For all predictors, the predicted attribute value is calculated by additionally multiplying the attribute value updated through the lift update process by the weight (stored in QW) updated through the lift prediction process. A point cloud encoder according to the embodiments (e.g., coefficient quantizer (30011)) quantizes the predicted attribute value. Additionally, a point cloud encoder (e.g., arismetic encoder (30012)) entropies the quantized attribute value.

[0169] A point cloud encoder according to the embodiments (e.g., a RAHT transform unit (30008)) can perform RAHT transform coding to predict attributes of upper-level nodes using attributes associated with nodes at lower levels of the octree. RAHT transform coding is an example of attribute intra-coding through octree backward scanning. A point cloud encoder according to the embodiments scans from voxels to the entire region and repeats the merging process up to the root node, merging voxels into larger blocks at each step. The merging process according to the embodiments is performed only on occupied nodes. The merging process is not performed on empty nodes, and the merging process is performed on the node immediately above the empty node.

[0170] The following equation represents the RAHT transformation matrix. g lx,y,z represents the average attribute value of the voxels at level l. g lx,y,z can be calculated from gl+1 2x,y,z and gl+1 2x+1,y,z. g l 2x,y,z and the weights of gl 2x+1,y,z are w1=w l 2x,y,z and w2=wl 2x+1,y,z.

[0171]

[0172] g l-1 x,y,z is a low-pass value used in the merging process at the next higher level. h l-1 x,y,z ε₀ are high-pass coefficients, and the high-pass coefficients at each step are quantized and entropy-coded (e.g., encoding of an arismetic encoder (30012)). The weights are w l-1 x,y,z = w l 2x,y,z + wl is calculated as 2x+1,y,z. The root node is the last g 1 0,0,0 and g 1 0,0,1 It is generated as follows through.

[0173]

[0174] The gDC value is also quantized and entropy-coded, just like the high-pass coefficient.

[0175] FIG. 7 shows an example of a point cloud decoder according to embodiments.

[0176] The point cloud decoder illustrated in FIG. 7 is an example of a point cloud decoder and can perform a decoding operation, which is the reverse process of the encoding operation of the point cloud encoder described in FIG. 1 to 6.

[0177] As described in Fig. 1, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.

[0178] A point cloud decoder according to the embodiments comprises an arithmetic decoder (7000), a synthesize octree (7001), a synthesize surface approximation (7002), a reconstruct geometry (7003), an inverse transform coordinates (7004), an arithmetic decoder (7005), an inverse quantize (7006), a RAHT transform (7007), a generate LoD (7008), an inverse lifting (7009), and / or an inverse transform colors (7010).

[0179] An arismetic decoder (7000), an octree composite unit (7001), a surface offset composite unit (7002), a geometry reconstruction unit (7003), and a coordinate system inverse transformation unit (7004) can perform geometry decoding. Geometry decoding according to the embodiments may include direct coding and trisoup geometry decoding. Direct coding and trisoup geometry decoding are applied optionally. Additionally, geometry decoding is not limited to the above examples and is performed as the reverse process of geometry encoding described in FIGS. 1 through 6.

[0180] The arismetic decoder (7000) according to the embodiments decodes the received geometry bitstream based on arismetic coding. The operation of the arismetic decoder (7000) corresponds to the reverse process of the arismetic encoder (30004).

[0181] The octree synthesis unit (7001) according to the embodiments can generate an octree by obtaining an Occupancy code from a decoded geometry bitstream (or information regarding the geometry obtained as a result of decoding). A specific description of the Occupancy code is as described in FIGS. 1 to 6.

[0182] The surface off-relation synthesis unit (7002) according to the embodiments can synthesize a surface based on the decoded geometry and / or the generated octree when trisoop geometry encoding is applied.

[0183] The geometry reconstruction unit (7003) according to the embodiments can regenerate geometry based on a surface and / or decoded geometry. As described in FIGS. 1 through 6, direct coding and trisoop geometry encoding are applied optionally. Accordingly, the geometry reconstruction unit (7003) directly retrieves and adds position information of points to which direct coding has been applied. In addition, when trisoop geometry encoding is applied, the geometry reconstruction unit (7003) can restore geometry by performing reconstruction operations of the geometry reconstruction unit (30005), such as triangle reconstruction, up-sampling, and voxelization operations. Specific details are omitted as they are the same as those described in FIG. 4. The restored geometry may include a point cloud picture or frame that does not contain attributes.

[0184] The coordinate system inverse transformation unit (7004) according to the embodiments can obtain the positions of the points by transforming the coordinate system based on the restored geometry.

[0185] The arismetic decoder (7005), inverse quantization unit (7006), RAHT transformation unit (7007), LoD generation unit (7008), inverse lifting unit (7009), and / or color inverse transformation unit (7010) may perform attribute decoding. Attribute decoding according to the embodiments may include RAHT (Region Adaptive Hierarchial Transform) decoding, prediction transformation (Interpolaration-based hierarchical nearest-neighbour prediction-Prediction Transform) decoding, and lifting transformation (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) decoding. The three decodings described above may be used optionally, or a combination of one or more decodings may be used. Furthermore, attribute decoding according to the embodiments is not limited to the examples described above.

[0186] The arismetic decoder (7005) according to the embodiments decodes the attribute bitstream into arismetic coding.

[0187] The inverse quantization unit (7006) according to the embodiments inverse quantizes information about the decoded attribute bitstream or the attribute obtained as a result of decoding and outputs the inverse quantized attributes (or attribute values). Inverse quantization may be optionally applied based on the attribute encoding of the point cloud encoder.

[0188] According to embodiments, the RAHT transform unit (7007), LoD generate unit (7008), and / or inverse lifting unit (7009) may process the reconstructed geometry and inverse quantized attributes. As described above, the RAHT transform unit (7007), LoD generate unit (7008), and / or inverse lifting unit (7009) may optionally perform a corresponding decoding operation according to the encoding of the point cloud encoder.

[0189] The color inverse conversion unit (7010) according to the embodiments performs inverse conversion coding to inversely convert the color value (or texture) included in the decoded attributes. The operation of the color inverse conversion unit (7010) may be selectively performed based on the operation of the color conversion unit (30006) of the point cloud encoder.

[0190] The elements of the point cloud decoder of FIG. 7 may be implemented in hardware, software, firmware, or a combination thereof, comprising one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud decoder of FIG. 7 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud decoder of FIG. 7.

[0191] FIG. 8 is an example of a transmission device according to embodiments.

[0192] The transmission device illustrated in FIG. 8 is an example of the transmission device (10000) of FIG. 1 (or the point cloud encoder of FIG. 3). The transmission device illustrated in FIG. 8 can perform at least one of the same or similar operations and methods as the operations and encoding methods of the point cloud encoder described in FIG. 1 to 6. A transmission device according to embodiments may include a data input unit (8000), a quantization processing unit (8001), a voxelization processing unit (8002), an octree occupancy code generation unit (8003), a surface model processing unit (8004), an intra / inter coding processing unit (8005), an arithmetic coder (8006), a metadata processing unit (8007), a color conversion processing unit (8008), an attribute conversion processing unit (or attribute conversion processing unit) (8009), a prediction / lifting / RAHT conversion processing unit (8010), an arithmetic coder (8011) and / or a transmission processing unit (8012).

[0193] The data input unit (8000) according to the embodiments receives or acquires point cloud data. The data input unit (8000) may perform an operation and / or acquisition method identical or similar to the operation and / or acquisition method of the point cloud video acquisition unit (10001) (or the acquisition process (20000) described in FIG. 2).

[0194] The data input unit (8000), quantization processing unit (8001), voxelization processing unit (8002), octree occupancy code generation unit (8003), surface model processing unit (8004), intra / inter coding processing unit (8005), and arithmetic coder (8006) perform geometry encoding. Since the geometry encoding according to the embodiments is identical or similar to the geometry encoding described in FIGS. 1 to 6, a detailed description is omitted.

[0195] The quantization processing unit (8001) according to the embodiments quantizes geometry (e.g., location values ​​of points, or position values). The operation and / or quantization of the quantization processing unit (8001) is the same or similar to the operation and / or quantization of the quantization unit (30001) described in FIG. 3. The specific description is the same as that described in FIG. 1 through 6.

[0196] The voxelization processing unit (8002) according to the embodiments voxelizes the position values ​​of the quantized points. The voxelization processing unit (80002) may perform the same or similar operation and / or process as the operation and / or voxelization process of the quantization unit (30001) described in FIG. 3. The specific description is the same as that described in FIG. 1 to 6.

[0197] The octree occupancy code generation unit (8003) according to the embodiments performs octree coding on the positions of voxelized points based on an octree structure. The octree occupancy code generation unit (8003) can generate an occupancy code. The octree occupancy code generation unit (8003) can perform operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (or octree analysis unit (30002)) described in FIGS. 3 and 4. The specific description is the same as that described in FIGS. 1 through 6.

[0198] The surface model processing unit (8004) according to the embodiments can perform trisup geometry encoding that reconstructs the positions of points within a specific region (or node) based on a voxel based on a surface model. The surface model processing unit (8004) can perform operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (e.g., surface approximation analysis unit (30003)) described in FIG. 3. The specific description is the same as that described in FIG. 1 through 6.

[0199] According to the embodiments, the intra / inter coding processing unit (8005) can intra / inter code point cloud data. The intra / inter coding processing unit (8005) can perform coding identical or similar to intra / inter coding. According to the embodiments, the intra / inter coding processing unit (8005) may be included in an arismetic coder (8006).

[0200] An arithmetic coder (8006) according to the embodiments entropy-encodes an octree and / or an approximated octree of point cloud data. For example, the encoding method includes an arithmetic encoding method. The arithmetic coder (8006) performs the same or similar operation and / or method as the arithmetic encoder (30004).

[0201] A metadata processing unit (8007) according to the embodiments processes metadata regarding point cloud data, such as setting values, and provides it to necessary processing processes such as geometry encoding and / or attribute encoding. Additionally, a metadata processing unit (8007) according to the embodiments may generate and / or process signaling information related to geometry encoding and / or attribute encoding. The signaling information according to the embodiments may be encoded separately from geometry encoding and / or attribute encoding. Additionally, the signaling information according to the embodiments may be interleaved.

[0202] The color conversion processing unit (8008), attribute conversion processing unit (8009), prediction / lifting / RAHT conversion processing unit (8010), and arithmetic coder (8011) perform attribute encoding. Since the attribute encoding according to the embodiments is identical or similar to the attribute encoding described in FIGS. 1 to 6, a detailed description is omitted.

[0203] The color conversion processing unit (8008) according to the embodiments performs color conversion coding that converts color values ​​included in attributes. The color conversion processing unit (8008) may perform color conversion coding based on reconstructed geometry. The description of the reconstructed geometry is the same as that described in FIGS. 1 through 6. In addition, it performs the same or similar operation and / or method as the operation and / or method of the color conversion unit (30006) described in FIG. 3. A detailed description is omitted.

[0204] The attribute transformation processing unit (8009) according to the embodiments performs attribute transformation that transforms attributes based on positions where geometry encoding has not been performed and / or reconstructed geometry. The attribute transformation processing unit (8009) performs operations and / or methods identical or similar to the operations and / or methods of the attribute transformation unit (30007) described in FIG. 3. A detailed description is omitted. The prediction / lifting / RAHT transformation processing unit (8010) according to the embodiments may code the transformed attributes by RAHT coding, prediction transformation coding, and lifting transformation coding, or a combination thereof. The prediction / lifting / RAHT transformation processing unit (8010) performs at least one of operations identical or similar to the operations of the RAHT transformation unit (30008), LoD generation unit (30009), and lifting transformation unit (30010) described in FIG. 3. In addition, the descriptions of predictive transformation coding, lifting transformation coding, and RAHT transformation coding are the same as those described in Figures 1 to 6, so a detailed description is omitted.

[0205] The arismetic coder (8011) according to the embodiments can encode coded attributes based on arismetic coding. The arismetic coder (8011) performs the same or similar operation and / or method as the operation and / or method of the arismetic encoder (300012).

[0206] A transmission processing unit (8012) according to embodiments may transmit each bitstream containing encoded geometry and / or encoded attributes and metadata information, or may transmit the encoded geometry and / or encoded attributes and metadata information by configuring them into a single bitstream. When the encoded geometry and / or encoded attributes and metadata information according to embodiments is configured into a single bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to embodiments may include signaling information and slice data, including SPS (Sequence Parameter Set) for sequence-level signaling, GPS (Geometry Parameter Set) for signaling of geometry information coding, APS (Attribute Parameter Set) for signaling of attribute information coding, and TPS (Tile Parameter Set) for tile-level signaling. The slice data may include information about one or more slices. One slice according to embodiments is one geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0 It may include ).

[0207] A slice refers to a series of syntax elements representing all or part of a coded point cloud frame.

[0208] According to the embodiments, the TPS may include information regarding each tile (e.g., coordinate value information of a bounding box and height / size information, etc.) for one or more tiles. The geometry bitstream may include a header and a payload. The header of the geometry bitstream according to the embodiments may include identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information regarding data included in the payload, etc. As described above, the metadata processing unit (8007) according to the embodiments may generate and / or process signaling information and transmit it to the transmission processing unit (8012). According to the embodiments, the elements performing geometry encoding and the elements performing attribute encoding may share data / information with each other as indicated by the dotted lines. The transmission processing unit (8012) according to the embodiments may perform an operation and / or transmission method identical or similar to the operation and / or transmission method of the transmitter (10003). A detailed explanation is omitted as it is the same as that described in FIGS. 1 and 2.

[0209] FIG. 9 is an example of a receiving device according to embodiments.

[0210] The receiving device illustrated in FIG. 9 is an example of the receiving device (10004) of FIG. 1. The receiving device illustrated in FIG. 9 can perform at least one of the same or similar operations and methods as the operations and decoding methods of the point cloud decoder described in FIG. 1 to FIG. 8.

[0211] A receiving device according to the embodiments may include a receiving unit (9000), a receiving processing unit (9001), an arithmetic decoder (9002), an occupancy code-based octree reconstruction processing unit (9003), a surface model processing unit (triangle reconstruction, up-sampling, voxelization) (9004), an inverse quantization processing unit (9005), a metadata parser (9006), an arithmetic decoder (9007), an inverse quantization processing unit (9008), a prediction / lifting / RAHT inverse transformation processing unit (9009), a color inverse transformation processing unit (9010), and / or a renderer (9011). Each component of the decoding according to the embodiments may perform the inverse process of the components of the encoding according to the embodiments.

[0212] A receiver (9000) according to the embodiments receives point cloud data. The receiver (9000) may perform an operation and / or a receiving method identical or similar to the operation and / or receiving method of the receiver (10005) of FIG. 1. A detailed description is omitted.

[0213] A receiving processing unit (9001) according to the embodiments can obtain a geometry bitstream and / or an attribute bitstream from the received data. The receiving processing unit (9001) may be included in the receiving unit (9000).

[0214] The arismetic decoder (9002), the occupancy code-based octree reconstruction processing unit (9003), the surface model processing unit (9004), and the inverse quantization processing unit (9005) can perform geometry decoding. Since the geometry decoding according to the embodiments is identical or similar to the geometry decoding described in at least one of FIGS. 1 to 8, a detailed description is omitted.

[0215] The arismetic decoder (9002) according to the embodiments can decode a geometry bitstream based on arismetic coding. The arismetic decoder (9002) performs the same or similar operation and / or coding as the operation and / or coding of the arismetic decoder (7000).

[0216] According to the embodiments, the Occupancy code-based octree reconstruction processing unit (9003) can reconstruct an octree by obtaining an Occupancy code from a decoded geometry bitstream (or information regarding geometry obtained as a result of decoding). The Occupancy code-based octree reconstruction processing unit (9003) performs the same or similar operations and / or methods as the octree synthesis unit (7001) and / or octree generation method. According to the embodiments, the surface model processing unit (9004) can perform trisup geometry decoding and related geometry reconstruction (e.g., triangle reconstruction, up-sampling, voxelization) based on the surface model method when trisup geometry encoding is applied. The surface model processing unit (9004) performs the same or similar operations as the surface offset synthesis unit (7002) and / or geometry reconstruction unit (7003).

[0217] The inverse quantization processing unit (9005) according to the embodiments can inverse quantize the decoded geometry.

[0218] A metadata parser (9006) according to the embodiments can parse metadata included in the received point cloud data, such as setting values, etc. The metadata parser (9006) can pass the metadata to geometry decoding and / or attribute decoding. A specific description of the metadata is omitted as it is the same as described in FIG. 8.

[0219] The arismetic decoder (9007), inverse quantization processing unit (9008), prediction / lifting / RAHT inverse transformation processing unit (9009), and color inverse transformation processing unit (9010) perform attribute decoding. Since attribute decoding is identical or similar to the attribute decoding described in at least one of FIGS. 1 to 8, a detailed description is omitted.

[0220] The arismetic decoder (9007) according to the embodiments can decode an attribute bitstream into arismetic coding. The arismetic decoder (9007) can perform decoding of the attribute bitstream based on reconstructed geometry. The arismetic decoder (9007) performs the same or similar operation and / or coding as the operation and / or coding of the arismetic decoder (7005).

[0221] The inverse quantization processing unit (9008) according to the embodiments can inverse quantize the decoded attribute bitstream. The inverse quantization processing unit (9008) performs the same or similar operation and / or method as the operation and / or inverse quantization method of the inverse quantization unit (7006).

[0222] The prediction / lifting / RAHT inverse transformation processing unit (9009) according to the embodiments can process the reconstructed geometry and inverse quantized attributes. The prediction / lifting / RAHT inverse transformation processing unit (9009) performs at least one of the same or similar operations and / or decodings as the operations and / or decodings of the RAHT transformation unit (7007), LoD generation unit (7008), and / or inverse lifting unit (7009) of FIG. 7. The color inverse transformation processing unit (9010) according to the embodiments performs inverse transformation coding to inversely transform the color values ​​(or textures) included in the decoded attributes. The color inverse transformation processing unit (9010) performs the same or similar operations and / or inverse transformation coding as the operations and / or inverse transformation coding of the color inverse transformation unit (7010) of FIG. 7. A renderer (9011) according to the embodiments can render point cloud data.

[0223] FIG. 10 shows an example of a structure capable of interoperability with a point cloud data transmission / reception method / device according to embodiments.

[0224] The structure of FIG. 10 represents a configuration in which at least one of a server (1060), a robot (1010), an autonomous vehicle (1020), an XR device (1030), a smartphone (1040), a home appliance (1050) and / or an HMD (1070) is connected to a cloud network (1010). The robot (1010), the autonomous vehicle (1020), the XR device (1030), the smartphone (1040), or the home appliance (1050) are referred to as devices. Additionally, the XR device (1030) may correspond to a point cloud data (PCC) device according to the embodiments or may be linked with a PCC device.

[0225] The cloud network (1000) may refer to a network that constitutes part of the cloud computing infrastructure or exists within the cloud computing infrastructure. Here, the cloud network (1000) may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, or a 5G network, etc.

[0226] The server (1060) is connected to at least one of a robot (1010), an autonomous vehicle (1020), an XR device (1030), a smartphone (1040), a home appliance (1050) and / or an HMD (1070) via a cloud network (1000) and can assist in at least some of the processing of the connected devices (1010 to 1070).

[0227] The HMD (Head-Mount Display) (1070) represents one of the types in which an XR device and / or PCC device according to the embodiments may be implemented. A device of the HMD type according to the embodiments includes a communication unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit, etc.

[0228] Hereinafter, various embodiments of the device (1010 to 1050) to which the above-described technology is applied will be described. Here, the device (1010 to 1050) illustrated in FIG. 10 may be linked / coupled with a point cloud data transmission / reception device according to the above-described embodiments.

[0229] <PCC+XR>

[0230] The XR / PCC device (1030) may be implemented as a Head-Mount Display (HMD), a Head-Up Display (HUD) equipped in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, digital signage, a vehicle, a stationary robot, or a mobile robot by applying PCC and / or XR (AR+VR) technology.

[0231] The XR / PCC device (1030) can obtain information about surrounding space or real objects by analyzing 3D point cloud data or image data obtained through various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output an XR object to be output. For example, the XR / PCC device (1030) can output an XR object containing additional information about a recognized object by associating it with the recognized object.

[0232] <PCC+XR+모바일폰>

[0233] The XR / PCC device (1030) can be implemented as a mobile phone (1040) or the like by applying PCC technology.

[0234] The mobile phone (1040) can decode and display point cloud content based on PCC technology.

[0235] <PCC+자율주행+XR>

[0236] The autonomous vehicle (1020) can be implemented as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0237] An autonomous vehicle (1020) equipped with XR / PCC technology may refer to an autonomous vehicle equipped with means for providing XR images, or an autonomous vehicle that is the subject of control / interaction within the XR images. In particular, the autonomous vehicle (1020) that is the subject of control / interaction within the XR images is distinguished from the XR device (1030) and can be interconnected with it.

[0238] An autonomous vehicle (1020) equipped with means for providing XR / PCC images can acquire sensor information from sensors including cameras and output XR / PCC images generated based on the acquired sensor information. For example, the autonomous vehicle (1020) can provide an XR / PCC object corresponding to a real object or an object in the screen to the occupant by providing an XR / PCC object by outputting an XR / PCC image with a HUD.

[0239] At this time, when the XR / PCC object is displayed on the HUD, at least a portion of the XR / PCC object may be displayed so as to overlap with the actual object to which the occupant's gaze is directed. On the other hand, when the XR / PCC object is displayed on a display provided inside the autonomous vehicle, at least a portion of the XR / PCC object may be displayed so as to overlap with an object on the screen. For example, the autonomous vehicle (1220) may display XR / PCC objects corresponding to objects such as lanes, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, etc.

[0240] VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.

[0241] In other words, VR technology is a display technology that provides real-world objects or backgrounds solely as CG images. On the other hand, AR technology refers to a technology that displays virtual CG images alongside images of real objects. Furthermore, MR technology is similar to the aforementioned AR technology in that it mixes and combines virtual objects with the real world. However, it is distinguished from AR technology in that while AR technology maintains a clear distinction between real-world objects and virtual objects created from CG images, using virtual objects to complement real-world objects, MR technology regards virtual objects as having the same nature as real-world objects. To give a more specific example, the aforementioned MR technology is applied in hologram services.

[0242] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended Reality) technology. Accordingly, the embodiments of the present disclosure are applicable to all VR, AR, MR, and XR technologies. These technologies may utilize encoding / decoding based on PCC, V-PCC, and G-PCC technologies.

[0243] The PCC method / device according to the embodiments can be applied to a vehicle providing autonomous driving services.

[0244] Vehicles providing autonomous driving services are connected to PCC devices to enable wired / wireless communication.

[0245] When a point cloud data (PCC) transmitting / receiving device according to the embodiments is connected to a vehicle for wired / wireless communication, it can receive and process content data related to AR / VR / PCC services that can be provided along with an autonomous driving service, and transmit it to the vehicle. Additionally, when the point cloud data transmitting / receiving device is mounted on a vehicle, the point cloud transmitting / receiving device can receive and process content data related to AR / VR / PCC services according to a user input signal received through a user interface device, and provide it to the user. A vehicle or a user interface device according to the embodiments can receive a user input signal. The user input signal according to the embodiments may include a signal indicating an autonomous driving service.

[0246] The point cloud data transmission method / device according to the embodiments is interpreted as a term referring to the transmission device (10000) of FIG. 1, the point cloud video encoder (10002), the transmitter (10003), the acquisition-encoding-transmission (20000-20001-20002) of FIG. 2, the point cloud video encoder of FIG. 3, the transmission device of FIG. 8, the device of FIG. 10, the transmission device of FIG. 30, etc.

[0247] The method / device for receiving point cloud data according to the embodiments is interpreted as a term referring to the receiving device (10004) of FIG. 1, the receiver (10005), the point cloud video decoder (10006), the transmission-decoding-rendering (20002-20003-20004) of FIG. 2, the point cloud video decoder of FIG. 7, the receiving device of FIG. 9, the device of FIG. 10, the receiving device of FIG. 32, etc.

[0248] In addition, the point cloud data transmission / reception method / device according to the embodiments may be referred to simply as the method / device according to the embodiments.

[0249] According to the embodiments, geometry data, geometry information, location information, etc. constituting the point cloud data are interpreted as having the same meaning. Attribute data, attribute information, attribute information, etc. constituting the point cloud data are interpreted as having the same meaning.

[0250] The method / device according to the embodiments can process point cloud data with consideration of scalable transmission.

[0251] The method / device according to the embodiments describes a method for efficiently supporting selective decoding of a portion of data when such decoding is required due to receiver performance or transmission speed when transmitting / receiving point cloud data. In particular, this document proposes a technique to increase the efficiency of scalable coding, wherein the encoder on the transmitting side can selectively transmit information required by the decoder on the receiving side regarding already compressed data, and the decoder can decode it. In this case, the coding unit may be a tree level, LoD, layer group, subgroup, data unit, slice, fine granularity slice (FGS), etc.

[0252] In particular, this document proposes a method to enhance the efficiency of scalable coding among point cloud data compression methods. Here, scalable coding is a technology that can gradually change the resolution of data according to conditions such as the request / processing speed / performance / transmission bandwidth of the receiver, and is a technology that enables the transmitter to efficiently transmit compressed data and the receiver to decode the compressed data. To this end, in addition to the technology of this disclosure, split packing may be applied to effectively transmit point cloud data configured based on layers, and can be used as a compression method for efficient storage and transmission of large-capacity point cloud data with a wide distribution and high point density. Referring to a point cloud data transmission / reception device (or may be abbreviated as encoder / decoder) according to the embodiments illustrated in FIGS. 3 and 7, point cloud data is composed of a set of points, and each point is composed of geometry information (or referred to as geometry or geometry data) and attribute information (or referred to as attribute or attribute data). Geometry information refers to the 3D position information (xyz) of each point. That is, the position of each point is expressed by parameters in a coordinate system representing 3D space (for example, parameters of the three axes representing space: X, Y, and Z (x,y,z)). Furthermore, attribute information refers to the point's color (RGB, YUV, etc.), reflectance, normal vectors, transparency, etc. Point Cloud Compression (PCC) uses octree-based compression to efficiently compress distribution characteristics that are non-uniformly distributed in 3D space, and attribute information is compressed based on this.The point cloud video encoder and point cloud video decoder illustrated in FIGS. 3 and 7 can process operation(s) according to the embodiments through each component.

[0253] According to the embodiments, the transmitting device compresses the geometry information (e.g., location) and attribute information (e.g., color / brightness / reflectivity, etc.) of the point cloud data and transmits them to the receiving device. At this time, the point cloud data can be configured according to an octree structure with layers or according to the Level of Detail (LoD) based on the level of detail, and based on this, scalable point cloud data coding and representation are possible. At this time, it is possible to decode or represent only a part of the point cloud data depending on the performance or transmission speed of the receiving device, but currently there is no method to remove unnecessary data in advance.

[0254] That is, in cases where only a portion of the scalable point cloud compressed bitstream needs to be transmitted (e.g., decoding only a portion of the layers of scalable decoding), it is not possible to select and send only the necessary parts. Therefore, as shown in Fig. 11, the necessary parts must be re-encoded after decoding at the transmitting device, or as shown in Fig. 12, the entire thing must be transmitted to the receiving device and then the necessary data must be selectively applied after decoding at the receiving device.

[0255] However, in the case of Fig. 11, a delay may occur due to the time required for decoding and re-encoding, and in the case of Fig. 12, bandwidth efficiency is reduced because even unnecessary data is transmitted to the receiving device, and there is a disadvantage that data quality must be lowered when using a fixed bandwidth.

[0256] Accordingly, the method / device according to the embodiments provides slices so that the point cloud can be processed by dividing it into regions.

[0257] In particular, for octree-based position compression, entropy-based compression methods and direct coding can be used together, and in this case, we propose a slice configuration to efficiently utilize scalability.

[0258] In addition, the method / device according to the embodiments can define a slice segmentation structure of point cloud data and signal a scalable layer and a slice structure for scalable transmission.

[0259] The method / device according to the embodiments can process the bitstream by dividing it into specific units for efficient bitstream transmission and decoding.

[0260] The method / device according to the embodiments enables selective transmission and decoding of layered point cloud data at the bitstream level.

[0261] The units according to the embodiments may be referred to as LoD, layer, slice, etc. LoD is a term synonymous with LoD in attribute data coding, but in another sense, it may refer to a data unit for a layer structure of a bitstream. The LoD according to the embodiments may be a concept corresponding to a single depth or grouping two or more depths based on the depth (level) of the layer structure of point cloud data, for example, an octree or multiple trees. Similarly, a layer is intended to generate a unit of a sub-bitstream and is a concept corresponding to a single depth or grouping two or more depths, and may correspond to a single LoD or two or more LoDs. Furthermore, a slice is a unit for constituting a unit of a sub-bitstream and may correspond to a single depth, a part of a single depth, or two or more depths. Additionally, a slice may correspond to a single LoD, a part of a single LoD, or two or more LoDs. According to the embodiments, LoD, layer, and slice may correspond to or be included with one another. Additionally, the unit according to the embodiments may include LoD, layer, slice, layer group, subgroup, etc., and may be referred to as mutually complementary. According to the embodiments, in an octree structure, layer, depth, level, depth level, and tree level may be used with the same meaning.

[0262] In particular, in the case of scalable attribute coding, as shown on the left side of FIG. 13, full layer geometry is required during the decoding process, so there was a problem of having to unnecessarily transmit geometry detail information even when the receiver does not use some detail levels. This resulted in problems such as requiring an unnecessarily large amount of transmission bandwidth, or delay or complexity issues caused by the receiver having to process geometry details. However, as disclosed in the present disclosure, by configuring the coding unit into independent slices according to the meaning of tree level, LoD, layer group, subgroup unit, etc., and transmitting / receiving, full layer geometry is not required during the decoding process in the case of scalable attribute coding, as shown on the right side of FIG. 13. In other words, scalable attribute coding becomes possible even with only the geometry of the corresponding layer, rather than the full layer.

[0263] FIG. 14 is a diagram showing an example of a layer-based point cloud data configuration according to embodiments. FIG. 14 is an example of an octree structure in which the depth level of the root node is set to 0 and the depth level of the leaf node is set to 7.

[0264] The method / device according to the embodiments can encode and decode point cloud data by configuring layer-based point cloud data as shown in FIG. 14.

[0265] The layering of point cloud data according to the embodiments may have a layer structure in various aspects such as SNR, spatial resolution, color, temporal frequency, and bit depth depending on the application field, and may form layers in a direction in which the density of data increases based on an octree structure or an LoD structure.

[0266] In other words, when generating LoDs based on an octree structure, LoDs can be defined to increase in the direction of increasing detail, that is, in the direction of increasing octree depth levels. In this document, the term "layer" may be used interchangeably with "level," "depth," "depth level," and "tree level."

[0267] For example, in an octree structure having 7 depth levels excluding the root node level (or root level), LoD 0 is configured to include from the root node level to octree depth level 4, LoD 1 is configured to include from the root node level to octree depth level 5, and LoD 2 is configured to include from the root node level to octree depth level 7.

[0268] FIG. 15(a) shows the bitstream structure of geometry data according to embodiments, and FIG. 15(b) shows the bitstream structure of attribute data according to embodiments.

[0269] The method / device according to the embodiments can generate LoDs based on the layering of an octree structure as in FIG. 14, and can configure geometry bitstreams and attribute bitstreams as in FIG. 15 (a) and FIG. 15 (b).

[0270] A transmitting device according to the embodiments can divide a bitstream obtained through point cloud compression into a geometry bitstream and an attribute bitstream depending on the type of data and transmit them.

[0271] At this time, each bitstream can be configured as a slice and transmitted. According to the embodiments, a geometry bitstream (e.g., FIG. 15 (a)) and an attribute bitstream (e.g., FIG. 15 (b)) can each be configured as a single slice and transmitted regardless of layer information or LoD information. In this case, to use only a part of the layer or LoD, the process of decoding the bitstream, selecting only the part to be used and removing unnecessary parts, and re-encoding based only on the necessary information must be performed.

[0272] This document proposes a method to divide and transmit bitstreams in layers (or LoDs) to avoid these unnecessary intermediate processes.

[0273] FIG. 16 is a diagram showing an example of the configuration of a bitstream for dividing and transmitting the bitstream in layers (or LoDs) according to embodiments.

[0274] For example, considering the case of LoD-based PCC technology, it has a structure where a lower LoD is included in a higher LoD. That is, the upper LoD includes all the points of the lower LoD. And, if we define the information of points that are included in the current LoD but not in the previous LoD—that is, points newly added for each LoD—as R (rest or retained), the transmitting device can divide the initial LoD information and the information R newly included in each LoD into independent units (e.g., slices) and transmit them as shown in Fig. 16.

[0275] In other words, a set of points newly added compared to the previous LoD to constitute each LoD can be defined as Information R. FIG. 16 is an example in which LoD1 includes LoD0 + Information R1 and LoD2 includes LoD1 + Information R2. In this disclosure, Information R, Information R1, Information R2, etc., may be referred to as refinement levels. For example, Information R2 is a set of points that is not included in LoD1 but is included in LoD2.

[0276] According to one embodiment, points sampled for one or more octree depth levels can be determined as data belonging to information R. That is, a set of points sampled for one or more octree depth levels (i.e., matching to an occupied node) can be defined as information R. According to another embodiment, points sampled for one octree depth level may be divided into multiple pieces of information R according to certain criteria. In this case, various criteria for dividing one octree depth level into multiple pieces of information R may be considered as follows. For example, when dividing one octree depth level into M pieces of information R, the data within the information R may be grouped into M pieces of information R having consecutive Molton codes, or pieces with the same remainder when the Molton code sequence index is divided by M may be grouped into M pieces of information R, or pieces located at the same position when grouped as sibling nodes may be grouped into M pieces of information R. According to yet another embodiment, if necessary, a portion of the sampled points of multiple octree depth levels may be determined as information R.

[0277] FIG. 16 shows an example in which a geometry bitstream and an attribute bitstream are each divided into three slices according to embodiments. Each slice includes a header and a payload (or data unit) containing actual data (e.g., geometry data, attribute data). The header may include information about the slice. Additionally, the header may further include reference information related to the previous slice, the previous LoD, or the previous layer for LoD configuration.

[0278] For example, in Fig. 16, the geometry bitstream is divided into a slice that transmits geometry data belonging to LoD0, a slice that transmits geometry data belonging to information R1, and a slice that transmits geometry data belonging to information R2. And the attribute bitstream is divided into a slice that transmits attribute data belonging to LoD0, a slice that transmits attribute data belonging to information R1, and a slice that transmits attribute data belonging to information R2.

[0279] A receiving method / device according to the embodiments receives a bitstream divided into LoD or layers and can efficiently decode only the data to be used without complex intermediate processes.

[0280] At this time, various embodiments of the method for transmitting the bitstream can be applied.

[0281] For example, the geometry bitstream and the attribute bitstream may be transmitted separately, or the geometry bitstream and the attribute bitstream may be multiplexed into a single bitstream and transmitted.

[0282] Additionally, when each bitstream contains LoD0 and one or more pieces of information R, the delivery order of LoD0 and one or more pieces of information R may vary.

[0283] FIG. 16 is an example in which a geometry bitstream and an attribute bitstream are transmitted, respectively, and LoD0 containing the geometry bitstream and two pieces of information R(R1, R2) are transmitted sequentially, and LoD0 containing the attribute bitstream and two pieces of information R(R1, R2) are transmitted sequentially.

[0284] FIG. 17 illustrates an example of a bitstream alignment method when a geometry bitstream and an attribute bitstream are multiplexed into a single bitstream according to embodiments.

[0285] The transmission method / device according to the embodiments can transmit geometry data and attribute data serially as shown in FIG. 17 when transmitting a bitstream. In this case, depending on the type of data, the entire geometry data (or geometry information) can be transmitted first, and then the attribute data (or attribute information) can be transmitted. In this case, there is an advantage that the geometry data can be quickly restored based on the information of the transmitted bitstream.

[0286] FIG. 17 illustrates, for example, that layers (LoDs) containing geometry data may be positioned first within the bitstream, and layers (LoDs) containing attribute data may be positioned after the geometry layers. Since attribute data is dependent on geometry data, layers (LoDs) containing geometry data may be positioned before layers (LoDs) containing attribute data. FIG. 17 illustrates an example where LoD0 containing geometry data and two pieces of information R (R1, R2) are transmitted sequentially, followed by LoD0 containing attribute data and two pieces of information R (R1, R2) being transmitted sequentially. In this case, the positions can be varied according to the embodiments. Additionally, references between geometry headers are possible, and references between attribute headers and geometry headers are also possible.

[0287] FIG. 18 shows another example of a bitstream alignment method when a geometry bitstream and an attribute bitstream are multiplexed into a single bitstream according to embodiments.

[0288] When transmitting a bitstream, the transmission method / device according to the embodiments can transmit geometry data and attribute data serially as shown in FIG. 18. At this time, bitstreams constituting the same layer containing geometry data and attribute data may be collected and transmitted. In this case, if a compression technique capable of parallel decoding of geometry and attributes is used, the decoding time can be reduced. At this time, information that needs to be processed first (smaller LoD, geometry must precede attribute) can be placed first.

[0289] FIG. 18 is an example in which data is transmitted in the order of LoD0 containing geometry data, LoD0 containing attribute data, information R1 containing geometry data, information R1 containing attribute data, information R2 containing geometry data, and information R2 containing attribute data. At this time, the positions can be varied according to the embodiments. In addition, reference between geometry headers is possible, and reference between attribute headers and geometry headers is also possible.

[0290] The transmitting / receiving method / device according to the embodiments can efficiently select a desired layer (or LoD) at the bitstream level when transmitting and receiving a bitstream. In the bitstream alignment method according to the embodiments, when geometry information is collected and sent as in FIG. 17, a gap may occur in the middle after selecting a specific bitstream level, and in this case, the bitstream may need to be rearranged.

[0291] Meanwhile, when geometry data and attribute data are bundled and transmitted according to layers as in FIG. 18, necessary information may be selectively transmitted or / or unnecessary information may be selectively removed as in FIG. 19(a) to FIG. 19(c) or FIG. 20(a) to FIG. 20(c), depending on the application field.

[0292] FIGS. 19(a) to 19(c) are drawings showing examples of symmetric geometry-attribute selection according to embodiments.

[0293] In cases where a portion of the bitstream must be selected according to the embodiments, taking FIGS. 19(a) through 19(c) as examples, the transmitting device selects and transmits only up to LoD1 (i.e., LoD0+R1), and removes information R2 corresponding to the upper layer (i.e., the new portion of LoD2) from the bitstream and does not transmit it. In the case of symmetric geometry-attribute selection, geometry data and attribute data of the same layer are selected and transmitted simultaneously or are selected and removed simultaneously.

[0294] FIGS. 20(a) to 20(c) illustrate examples of asymmetric geometry-attribute selection according to embodiments. In the case of asymmetric geometry-attribute selection, only one of the geometry data and attribute data of the same layer is selected and transmitted or removed.

[0295] In cases where a portion of the bitstream needs to be selected according to the embodiments, taking FIGS. 20(a) to 20(c) as examples, the transmitting device selects and transmits LoD1 (LoD0 + R1) containing geometry data, LoD1 (LoD0 + R1) containing attribute data, and R2 containing geometry data, and removes R2 containing attribute data from the bitstream and does not transmit it. That is, attribute data is an example of selecting and transmitting the remaining data excluding the data of the upper layer (R2), and geometry data is an example of transmitting data from all layers (from Level 0 (root level) to Level 7 (leaf level) of the octree structure).

[0296] If a portion of the bitstream needs to be selected according to the embodiments, the portion of the bitstream can be selected by the symmetric geometry-attribute selection method of FIGS. 19(a) to 19(c) or the asymmetric geometry-attribute selection method of FIGS. 20(a) to 20(c) or by combining the symmetric geometry-attribute selection method and the asymmetric geometry-attribute selection method.

[0297] The bitstream splitting and selection of some bitstreams described above are intended to support the scalability of point cloud data.

[0298] Referring to Fig. 14, this document can support scalable encoding / decoding (scalability) when point cloud data is represented in an octree structure and separated by LoD (or layer).

[0299] The scalability function according to the embodiments may include slice level scalability and / or octree level scalability.

[0300] According to the embodiments, the LoD can be used as a unit to represent a set of one or more octree layers. Additionally, the LoD may have the meaning of a bundle of octree layers to be organized in slice units.

[0301] According to the embodiments, the LoD can be used in a broad sense, such as a unit for subdividing data in detail, by extending the meaning of LoD during attribute encoding / decoding.

[0302] That is, spatial scalability by the actual octree layer (or scalable attribute layer) can be provided for each octree layer, but if scalability is configured at the slice level before bitstream parsing, it can be selected at the LoD level.

[0303] For example, in the octree structure, levels 1 through 4 correspond to LoD0, levels 1 through 5 correspond to LoD1, and levels 1 through 7 (i.e., leaf levels) correspond to LoD2.

[0304] That is, taking Fig. 14 as an example, when utilizing scalability at the slice level, the provided scalable stages are three stages: LoD0, LoD1, and LoD2, and the scalable stages that can be provided in the decoding stage by the octree structure are eight stages ranging from the root level to the leaf level.

[0305] According to embodiments, when LoD0 to LoD2 are each composed of slices, the transcoder (see FIG. 11) of the receiver or transmitter may select only LoD0, or only LoD1, or LoD2 for scalable processing. In FIG. 14, LoD1 includes LoD0, and LoD2 includes LoD1 and LoD2.

[0306] For example, if only LoD0 is selected, the maximum octree level becomes 4, and one scalable layer among the octree layers from 0 to 4 can be selected during the decoding process. In this case, the receiving device can consider the node size obtained through the maximum octree level (or depth) as a leaf node, and the node size can be transmitted as signaling information.

[0307] For example, if LoD1 is selected, Layer 5 is added so that the maximum octree level becomes 5, and one scalable layer among octree layers 0 to 5 can be selected during the decoding process. In this case, the receiving device may consider the node size obtainable through the maximum octree level (or depth) as a leaf node, and the node size at this time can be transmitted as signaling information. According to embodiments, octree depth, octree layer, octree level, etc. refer to units for subdividing data in detail.

[0308] For example, if LoD2 is selected, layers 6 and 7 are added, making the maximum octree level 7, and one scalable layer among the octree layers 0 to 7 can be selected during the decoding process. In this case, the receiving device can consider the node size obtainable through the maximum octree level (or depth) as a leaf node, and the node size at this time can be transmitted as signaling information.

[0309] FIGS. 21(a) to 21(c) illustrate examples of methods for configuring slices containing point cloud data according to embodiments.

[0310] The transmission method / device / encoder according to the embodiments may be configured by dividing the G-PCC bit stream into a slice structure. The data unit for detailed data representation may be a slice.

[0311] For example, one or more octree layers (or depths) can be matched to a single slice.

[0312] A transmission method / device according to embodiments, e.g., an encoder, can construct a slice (41001)-based bitstream by scanning nodes (points) included in an octree in the direction of a scan order (41000). A slice may include nodes of one or more levels in an octree structure, may include only nodes of a specific level, or may include only some nodes of a specific level. Alternatively, it may include only some nodes of one or more levels.

[0313] FIG. 21(a) shows an example in which an octree structure is composed of seven slices, where slice (41002) can be composed of nodes from level 0 to level 4. Slice (41003) can be composed of some nodes of level 5, slice (41004) can be composed of other some nodes of level 5, and slice (41005) can be composed of yet another some nodes of level 5. That is, in FIG. 21(a), level 5 is divided into three slices. Similarly, in FIG. 21(a), level 6 (i.e., leaf level) is also divided into three slices. In other words, a slice can be composed of some nodes of a specific level.

[0314] FIG. 21(b) shows an example in which an octree structure is composed of four slices, wherein one slice is composed of nodes from level 0 to level 3 and some nodes of level 4, and one slice is composed of the remaining nodes of level 4 and some nodes of level 5. Additionally, one slice is composed of the remaining nodes of level 5 and some nodes of level 6, and one slice is composed of the remaining nodes of level 6.

[0315] FIG. 21(c) shows an example in which an octree structure is composed of five slices, with one slice composed of nodes from level 0 to level 3 and four slices composed of nodes from level 4 to level 6. That is, one slice is composed of some nodes of level 4, some nodes of level 5, and some nodes of level 6. In other words, one slice from level 4 to level 6 may include some data of level 4 and data of level 5 or level 6 corresponding to the child nodes of that data.

[0316] In other words, as shown in FIG. 21(b) and FIG. 21(c), when multiple octree layers are matched to a single slice, only some nodes of each layer may be included. When multiple slices constitute a geometry / attribute frame in this way, information necessary to configure the layers can be transmitted to the receiving device through signaling information. For example, the signaling information may include layer information included in each slice, node information included in each layer, etc.

[0317] An encoder and a device corresponding to the encoder according to the embodiments can encode point cloud data and generate and transmit a bitstream further comprising the encoded data and signaling information (or parameter information) regarding the point cloud data.

[0318] Furthermore, when generating a bitstream, the bitstream can be generated based on a bitstream structure according to the embodiments (e.g., see FIGS. 15-21, etc.). Accordingly, a receiving device, a decoder, a corresponding device, etc. according to the embodiments can receive and parse a bitstream configured to be suitable for a selective decoding structure of some data, and efficiently provide only a portion of the point cloud data by decoding it (see the right side of FIG. 13).

[0319] The following describes the scalable transmission of point cloud data.

[0320] A point cloud data transmission method / device according to the embodiments can scalably transmit a bitstream containing point cloud data, and a point cloud data reception method / device according to the embodiments can scalably receive and decode the bitstream.

[0321] When a bitstream having the structure described in FIGS. 11 to 21 is used for scalable transmission, signaling information for selecting the slice required by the receiving device can be transmitted to the receiving device. Scalable transmission may mean transmitting or decoding only a part of the bitstream, rather than transmitting or decoding the entire bitstream, and the result may be low-resolution point cloud data.

[0322] According to the embodiments, when applying scalable transmission to an octree-based geometry bitstream, for the bitstream of each octree layer (Fig. 14) from the root node to the leaf node, it must be possible to construct point cloud data using only information up to a specific octree layer.

[0323] To achieve this, the target octree layer must not have any dependency on information from its sub-octree layers. This can be a common constraint applied to geometry / attribute coding.

[0324] In addition, when performing scalable transmission, it is necessary to transmit a scalable structure to the receiving device for selecting scalable layers from the transmitting / receiving device. When considering the octree structure according to the embodiments, all octree layers may support scalable transmission, but scalable transmission may be enabled only for specific octree layers or lower. For example, if some of the octree layers are included, the receiving device is informed through signaling information which scalable layer the slice belongs to, so that the receiving device can determine whether the slice is necessary or unnecessary at the bitstream stage. In the example of FIG. 21(a), levels 0 (i.e., root level) through 4 (41002) do not support scalable transmission and constitute a single scalable layer, and the octree layers below can be configured to have a one-to-one match with the scalable layer. Generally, scalability can be supported for parts corresponding to leaf nodes, and as shown in FIG. 21(c), when multiple octree layers are included within a single slice, it can be defined to form a single scalable layer for those layers.

[0325] At this time, scalable transmission and scalable decoding can be distinguished and used depending on the purpose. According to the embodiments, scalable transmission can be used for the purpose of selecting information up to a specific layer without passing through a decoder at the transmitting / receiving device. According to the embodiments, scalable decoding can be used for the purpose of selecting a specific layer during coding. That is, scalable transmission supports the selection of necessary information without passing through a decoder in a compressed state (i.e., at the bitstream stage), thereby enabling the identification of a specific layer at the transmitting or receiving device. On the other hand, scalable decoding can be used in cases such as scalable representation by supporting encoding / decoding only up to the parts required during the encoding / decoding process.

[0326] In this case, the layer configuration for scalable transmission and the layer configuration for scalable decoding may differ. For example, the lower three octree layers including the leaf node may constitute a single layer from the perspective of scalable transmission, but from the perspective of scalable decoding, scalable decoding may be possible for the leaf node layer, leaf node layer-1, and leaf node layer-2 respectively, if they include all layer information.

[0327] FIG. 22 illustrates a geometry coding layer structure according to embodiments. In particular, the left side of FIG. 22 shows an example of three slices generated by a layer group structure in an encoder on the transmitting side, and the right side of FIG. 22 shows an example of a partially decoded output using two slices in a decoder on the receiving side.

[0328] According to the embodiments, when fine granularity slicing is enabled, the G-PCC bitstream can be divided into several sub-bitstreams. Here, fine granularity slicing may be referred to as layer group-based slicing. To effectively utilize the layering structure of G-PCC, each slice may contain coded data of a partial coding layer or a partial region. Through segmentation or partitioning paired with the coding layer structure, scalable transmission or spatial random access use cases can be supported in an efficient manner.

[0329] layer group based slice segmentation

[0330] In fine particle size slicing, each slice segment may contain data coded from a layer group defined as follows.

[0331] A layer group can be defined as a group of consecutive tree layers, and the starting and ending depths of the group of tree layers can be any number in the tree depth, and the starting depth can be smaller than the ending depth. The order of data coded in a slice segment can be the same as the order of data coded in a single slice.

[0332] For example, considering a geometry coding layer structure with eight coding layers as shown on the left side of FIG. 22, there are three layer groups, and each layer group is matched to a different slice. More specifically, Layer Group 1 for coding layers 0 through 4 is matched to slice 1, Layer Group 2 for coding layer 5 is matched to slice 2, and Layer Group 3 for coding layers 6 through 7 is matched to slice 3. When the first two slices (i.e., slices 1 and 2) are transmitted or selected, the decoded output becomes partial layers 0 through 5, as shown on the right side of FIG. 22. Using slices in the layer group structure allows for partial decoding of coding layers without accessing the entire bitstream.

[0333] Bitstream and point cloud data according to the embodiments can be generated based on coding layer-based slice partitioning. By slicing the bitstream at the end of the coding layer of the encoding process, the method / device according to the embodiments can select the relevant slice, thereby supporting scalable transmission or partial decoding.

[0334] The left side of FIG. 22 shows a geometry coding layer structure with eight layers, where each slice corresponds to a layer group. Layer group 1 includes coding layers 0 through 4. Layer group 2 includes coding layer 5. Layer group 3 is a group for coding layers 6 and 7. When geometry (or attributes) have a tree structure with eight levels (depths), bitstreams can be hierarchically organized by grouping data corresponding to one or more levels (depths). Each group can be included in a single slice.

[0335] The right side of Fig. 22 shows the decoded output when two of the three slices are selected. When the decoder selects Group 1 and Group 2, partial layers from level (depth) 0 to 5 of the tree are selected. That is, by using slices of the layer group structure, partial decoding of the coding layer can be supported without accessing the entire bitstream.

[0336] For a partial decoding process according to the embodiments, the encoder may generate three slices based on a layer group structure. The decoder according to the embodiments may select two of the three slices to perform partial decoding.

[0337] A bitstream according to the embodiments may include slices based on layer groups. Each slice may include a header containing signaling information regarding point cloud data (i.e., geometry data and / or attribute data) included in the slice. A receiving method / device according to the embodiments may select slices and decode point cloud data included in the payload of a slice based on a header included in a selected slice.

[0338] In addition to the layer group structure, the method / device according to the embodiments may further divide the layer group into multiple subgroups to account for spatial random access use cases. The subgroups according to the embodiments are mutually exclusive, and the set of subgroups may be identical to the layer group. Since the points of each subgroup form boundaries in a spatial region, the subgroups can be represented by subgroup bounding box information. Using spatial information, the layer group and subgroup structure can support access to the region of interest (ROI) by selecting slices that cover the ROI. Spatial random access within a frame or tile can be supported by efficiently comparing the bounding box information of the region of interest (ROI) with that of each slice.

[0339] The method / device according to the embodiments may configure a slice for transmitting point cloud data as shown on the left side of FIG. 22.

[0340] According to the embodiments, the entire coded bitstream may be contained in a single slice. Furthermore, for multiple slices, each slice may contain a sub-bitstream. The order of the slices may be the same as the order of the sub-bitstreams. And, each slice may be matched to a group of layers in a tree structure.

[0341] Also, just as the upper layer of a geometry tree does not affect the lower layer, slices may not affect previous slices.

[0342] The segmented slices according to the embodiments are efficient in terms of error robustness, effective transmission, and supporting region of interest.

[0343] 1) Error resilience

[0344] Compared to a single-slice structure, divided slices can be more robust to errors. That is, if a single slice contains the entire bitstream of a frame, data loss can affect the entire frame data. In contrast, when the bitstream is divided into multiple slices, even if at least one slice is lost, at least one slice unaffected by the loss can still be decoded.

[0345] 2) Scalable transmission

[0346] This document considers cases where multiple decoders with different capabilities can be supported.

[0347] If coded point cloud data (i.e., Point Cloud Compression (PCC) bitstream) is contained in a single slice, the Level of Depth (LoD) of the coded point cloud data can be determined prior to encoding. Therefore, multiple pre-encoded bitstreams with different resolutions of point cloud data can be transmitted independently. This can be inefficient in terms of large bandwidth or storage space.

[0348] If coded point cloud data (i.e., a Point Cloud Compression (PCC) bitstream) is contained in partitioned slices, a single bitstream can support different levels of decoders. From the decoder's perspective, the receiving device can select target layers and pass a partially selected bitstream to the decoder. Similarly, by using a single bitstream without partitioning the entire bitstream, a partial bitstream can be efficiently generated at the transmitting device.

[0349] 3) Region-based spatial scalability

[0350] In the context of the G-PCC requirements according to the embodiments, region-based spatial scalability can be defined as follows: that is, the compressed bitstream consists of one or more layers, so that a specific region of interest can have a higher density with additional layers, and the layers can be predicted from the lower layers.

[0351] To support this requirement, it is necessary to support region-wise different detailed representations. For example, in VR / AR applications, it is desirable to represent far objects with lower precision and nearer objects with higher precision. Additionally, the decoder can increase the resolution of the region of interest upon request. This can be implemented by using scalable structures of G-PCC, such as geometry octrees and scalable attribute coding schemes.

[0352] According to the embodiments, decoders must access the entire bitstream based on the current slice structure containing the entire geometry or attributes. This can cause bandwidth, memory, and decoder inefficiencies. Meanwhile, if the bitstream is segmented into multiple slices and each slice contains sub-bitstreams according to scalable layers, the decoder according to the embodiments can select a slice as needed before efficiently parsing the bitstream.

[0353] The method / device according to the embodiments can generate layer groups using the tree structure (or layer structure) of point cloud data.

[0354] Taking the left side of Fig. 22 as an example, there are eight layers within the geometry coding layer structure (e.g., octree structure), and three slices can be used to contain one or more layers. A group represents a group of layers. When scalable attribute coding is used, the tree structure is identical to the geometry tree structure. The same octree-slice mapping can be used to create attribute slice segments.

[0355] A layer group according to the embodiments represents a group of layer structure units that occur in G-PCC coding, such as an octree layer, a LoD layer, etc.

[0356] A subgroup can be represented as a set of adjacent nodes within a single layer group. For example, it can be composed of groups of adjacent nodes based on Morton code order, distance-based adjacency, or coding order. Nodes in a parent-child relationship may also exist within a single subgroup.

[0357] When a subgroup is defined, a boundary occurs in the middle of the layer, and parameters such as entropy_continuation_enabled_flag can be signaled regarding whether entropy continuity is maintained at the boundary. Additionally, continuity can be maintained by referencing the previous slice through ref_slice_id.

[0358] The tree structure according to the embodiments may be an octree structure, and the attribute layer structure or attribute coding tree according to the embodiments may include a Level of Detail (LoD) structure. That is, the tree structure for point cloud data includes layers corresponding to depth or level, and the layers may be grouped.

[0359] The method / device according to the embodiments (e.g., the octree analysis unit (30002) or LoD generation unit (30009) of FIG. 3, the octree synthesis unit (7002) or LoD generation unit (7008) of FIG. 7) can generate an octree structure of geometry or generate a LoD tree structure of attributes. Additionally, point cloud data can be grouped based on layers of the tree structure.

[0360] Referring to the left side of FIG. 22, multiple layers are grouped to form the first to third groups. One group can be further divided to form subgroups.

[0361] According to the embodiments, a slice may contain data coded from a layer group. Here, the layer group is defined as a group of consecutive tree layers, and the start and end depths of the tree layers may be specific numbers within the tree depth, and the start is smaller than the end.

[0362] The left side of Fig. 22 shows a geometry coding layer structure as an example of a tree structure, but a coding layer structure for attributes can be generated in the same way.

[0363] FIG. 23 shows a layer group and subgroup structure according to embodiments.

[0364] Referring to Fig. 23, point cloud data and bitstreams can be represented separated by bounding boxes.

[0365] Referring to FIG. 23, a subgroup structure and a bounding box corresponding to the subgroup are illustrated. Layer group 2 is divided into two subgroups (group2-1, group2-2) and included in different slices, and layer group 3 is divided into four subgroups (group3-1, group3-2, group3-3, group3-4) and included in different slices. Given the slices of the layer groups and subgroups along with bounding box information, spatial access can be performed by 1) comparing the bounding box of each slice with the ROI, 2) selecting the slice in which the subgroup bounding box overlaps with the ROI, and 3) decoding the selected slice.

[0366] When the ROI is considered in region 3-3, slices 1, 3, and 6 are selected as the subgroup bounding boxes of layer group 1, subgroup 2-2, and 3-3 that cover the ROI area. For effective spatial access, it is assumed that there is no dependency between subgroups of the same layer group. For live streaming or low-latency use cases, time efficiency can be improved by performing selection and decoding when each slice segment is received.

[0367] The method / device according to the embodiments may represent data as a tree (45000) composed of layers (which may be referred to as depths, levels, etc.) when encoding geometry and / or attributes. Point cloud data corresponding to each layer (depth / level) may be grouped into a layer group (or group, 45001). Layer group 2 may be further divided (segmented) into two subgroups (45002), and layer group 3 may be further divided (segmented) into four subgroups (45003). Each subgroup may be composed of each slice to generate a bitstream.

[0368] A receiving device according to the embodiments receives a bitstream, selects a specific slice from the bitstream, and can decode a bounding box corresponding to a subgroup included in the selected slice. For example, if slice 1 is selected, a bounding box (45004) corresponding to layer group 1 can be decoded. Layer group 1 may be data corresponding to the largest area. When additionally displaying a detailed area for layer group 1, the method / device according to the embodiments may select slice 3 and / or slice 6 to hierarchically access the bounding boxes (point cloud data) of subgroup 2-2 and / or subgroup 3-3 for the detailed area included in the area of ​​layer group 1.

[0369] Encoding and decoding of point cloud data using the layer group and subgroup of FIG. 23 can be performed by the transmitting / receiving device of FIG. 1, the encoding and decoding of FIG. 2, the transmitting device / method of FIG. 3, the receiving device / method of FIG. 7, the transmitting / receiving device / method of FIG. 8 and FIG. 9, the devices of FIG. 10, the transmitting / receiving device of FIG. 30 and FIG. 32, and the transmitting / receiving method of FIG. 42 and FIG. 43.

[0370] FIGS. 24(a) to 24(c) show representations of layer group-based point cloud data according to embodiments.

[0371] The device / method according to the embodiments can provide efficient access to large-scale point cloud data or high-density point cloud data through layer group slicing based on scalability and spatial access capabilities. Because point cloud data has a large number of points and a large data size, rendering or displaying content may take a significant amount of time. Therefore, as an alternative approach, the level of detail can be adjusted based on the viewer's interest. For example, when the viewer is far from a scene or object, structural or whole-area information is more important than local details, whereas as the viewer approaches a specific area or object, detailed information about the area of ​​interest is required. By using an adaptive method, the renderer according to the embodiments can efficiently provide data of sufficient quality to the viewer. FIGS. 24(a) through 24(c) illustrate an example of increasing detail for three levels of viewing distance that change based on the ROI.

[0372] The high-level view in Fig. 24(a) shows coarse detail, the mid-level view in Fig. 24(b) shows medium-level detail, and the low-level view in Fig. 24(c) shows fine-grained detail.

[0373] FIG. 25 shows a point cloud data transmission / reception device / method according to embodiments.

[0374] Multi-resolution ROIs can be supported when layer group slicing is used for G-PCC bitstream generation.

[0375] Referring to FIG. 25, multi-resolution ROIs can be supported by the scalability and spatial accessibility of hierarchical slicing. In FIG. 25, the encoder (47001) at the transmitting end can generate bitstream slices of spatial subgroups of each layer group or octree layer groups. Upon request, a slice matching an ROI of each resolution is selected and transmitted to the receiving end. The total bitstream size is reduced compared to a tile-based approach because it does not include details other than the requested ROI. At the receiver, the decoder (47004) can combine slices to produce three outputs, for example: 1) a high-level view output from a layer group, 2) a mid-level view output from a selected subgroup of layer group 1 and layer group 2, and 3) a low-level view output of good quality detail from layer group 1 and selected subgroups of layer groups 2 and 3. Since the outputs can be generated progressively, the receiver can provide a viewing experience such as a zooming function in which the resolution gradually increases from a high-level view to a low-level view.

[0376] The encoder (47001) above is a point cloud encoder according to the embodiments and may correspond to a geometry encoder and / or an attribute encoder. The encoder (47001) may slice point cloud data based on layer groups (or groups). Layers may be referred to as tree depths, LoD levels, etc. As in 47002, the depths of the geometry octree and / or levels of attribute layers, etc., may be divided into layer groups (or subgroups).

[0377] The slice selector (47003) can be combined with the encoder (47001) to select a divided slice (or sub-slice) and selectively transmit it partially, such as to layer group 1 to layer group 3.

[0378] The decoder (47004) can decode point cloud data transmitted optionally and partially. For example, for a high-level view, it can decode layer group 1 (high depth / layer / level or index 0, close to the root). And, for a mid-level view, it can decode based on layer group 1 and layer group 2 with a higher depth / level index than layer group 1 alone. Additionally, for a low-level view, it can decode based on layer group 1 to layer group 3.

[0379] Referring to FIG. 25, an encoder (47001) according to the embodiments can receive point cloud data as input and slice it into layer groups. That is, the point cloud data can be structured hierarchically and divided into layer groups. In this case, the hierarchical structure may refer to an octree structure or a Level of Detail (LoD). 47002 represents the point cloud data divided into layer groups. A slice selector (47003) can select a layer group (or a corresponding slice), and the selected slices are transmitted to a decoder (47004) at the receiving end. The decoder (47004) can combine the received slices according to the user's request to restore only layer group 1, restore layer groups 1 and 2, or restore all received layer groups. The layer groups are mutually hierarchical and differ in the degree of detail. When only Layer Group 1 is restored, the restored range is wide but cannot be expressed in detail, whereas when Layer Groups 1 through 3 are all restored, the restored range is narrow but can be expressed in detail.

[0380] So far, we have explained scalable coding and decoding of geometry data.

[0381] According to the embodiments, scalable coding and decoding can also be applied to attribute data.

[0382] The present disclosure describes a layer-based attribute compression method for improving the transmission efficiency of scalable attribute compression.

[0383] That is, the present disclosure can perform LoD generation and nearest neighbor (NN) search based on scalable locations.

[0384] In particular, the present disclosure proposes a method to improve the efficiency of scalable attribute coding by eliminating the dependency of attribute coding on coded geometry layers when using scalable attribute coding. The method proposed in the present disclosure may be related to LoD generation and / or NN search during the scalable attribute coding or decoding process. More generally, the present disclosure may be used in scalability-based applications that decode only some of the coding layers.

[0385] To further enhance the effect proposed in this disclosure, it can be used in conjunction with a slicing technique that divides and transmits geometry / attribute bitstreams (e.g., layer group slicing).

[0386] In the transmitting device of the present disclosure, LoD generation and NN search are performed in the LoD generation unit (30009) of FIG. 3 or the prediction / lifting / RAHT conversion processing unit (8009) of FIG. 8 as an example. This is to aid the understanding of those skilled in the art, and LoD generation and NN search may be performed in each block or module. Additionally, in the receiving device of the present disclosure, LoD generation and NN search are performed in the LoD generation unit (7008) of FIG. 7 or the prediction / lifting / RAHT inverse conversion processing unit (9009) of FIG. 9 as an example. This is to aid the understanding of those skilled in the art, and LoD generation and NN search may be performed in each block or module.

[0387] As described in the present disclosure, by configuring the coding unit into independent slices according to the meaning of tree level, LoD, layer group unit, etc., and transmitting / receiving, full layer geometry is not required during the decoding process in the case of scalable attribute coding. That is, scalable attribute coding becomes possible even with only the geometry of the corresponding layer, rather than the full layer.

[0388] For convenience of explanation, the present disclosure assumes that a tree structure is generated in a geometry encoder as shown in FIG. 14, and that a LoD is generated in an attribute encoder based on this tree structure. That is, FIG. 14 is an example of an octree structure in which the depth level of the root node is set to 0 and the depth level of the leaf node is set to 7.

[0389] In the octree structure of Fig. 14, LoD 0 is configured to include from the root node level to octree depth level 4, LoD 1 is configured to include from the root node level to octree depth level 5, and LoD 2 is configured to include from the root node level to octree depth level 7.

[0390] In other words, the attribute encoder generates one or more LoDs based on the octree structure. Here, LoD0 is a set consisting of points with the largest distances between them, and as l increases, LoD l The distance between points belonging to becomes smaller. The present disclosure describes a group having different LoDs LoDs l It can be referred to as a set.

[0391] And, the attribute encoder is LoD l When a set is created, LoD lBased on the set, X (>0) nearest neighbor (NN) points can be found in the group where the LoD is equal to or smaller (i.e., the distance between nodes is large) and registered as a set of neighbor points in the predictor. X is the maximum number of neighbor points that can be set, and in this disclosure, 3 is described as an example.

[0392] For example, in Fig. 6, neighbor points of P3 belonging to LoD1 are found in LoD0 and LoD1. For instance, if the maximum number (X) that can be set as neighbor points is 3, the 3 neighbor nodes closest to P3 can be P2, P4, and P6. These 3 nodes are registered as a set of neighbor points in the predictor of P3.

[0393] As described above, each point in the point cloud data may have a predictor. The attribute encoder according to the embodiments may compress and transmit attribute information based on neighboring points registered in the predictor.

[0394] In this way, the attribute encoder of the transmitting device generates LoDs based on points of the reconstructed geometry for attribute compression and performs the process of finding the nearest neighbor points of the point to be encoded based on the generated LoDs. Similarly, the attribute decoder of the receiving device also generates LoDs and performs the process of finding the nearest neighbor points of the point to be decoded based on the generated LoDs.

[0395] Meanwhile, assuming that layer group-based slicing is performed on the octree structure of FIG. 14 as shown on the left side of FIG. 22, LoD0 may include layer group 1, LoD1 may include layer groups 1 and 2, and LoD2 may include layer groups 1 to 3. In other words, the attribute encoder may generate LoD0 based on layer group 1, generate LoD1 based on layer groups 1 and 2, and generate LoD2 based on layer groups 1 to 3.

[0396] According to the embodiments, in an octree structure, layer, depth, level, depth level, and tree level may be used interchangeably.

[0397] For example, if only layer groups 1 and 2 are transmitted at the transmitting end, the layers transmitted in the octree structure as shown on the left side of Fig. 22 are from 0 to 5 (i.e., 5 layers), and the layers skipped are 6-7 (i.e., 2 layers). That is, when utilizing attribute scalability at the slice level, the scalable stages provided are 3 stages of LoD0, LoD1, and LoD2, and the scalable stages that can be provided in the decoding stage by the octree structure are 8 stages ranging from the root level (or depth or layer) to the leaf level (or depth or layer).

[0398] According to embodiments, when LoD0 to LoD2 are each composed of slices for attributes, the transmitting device and / or the receiving device may select only LoD0, or only LoD1, or LoD2 for attribute scalable processing. For example, if there are three LoDs, LoD0 can be used to represent low-resolution details, LoD1 can be used to represent medium-resolution details, and LoD2 can be used to represent high-resolution details.

[0399] At this time, LoD generation and neighbor point selection are based on the locations of points in the point cloud data provided by the geometry encoder.

[0400] That is, LoD generation and neighbor point selection are performed using full layer geometry information.

[0401] FIGS. 26 to 29 show point cloud data in 2D space for convenience of explanation, and can be expanded to 3D space by adding a z-axis.

[0402] FIG. 26 is a diagram showing an example of performing NN search using full-layer geometry information according to embodiments. That is, FIG. 26 represents point cloud data in a 2D space and is an example of performing NN search using full-layer geometry information in this 2D space. In the geometry grid of FIG. 26, the grid represents the positions of points (or geometry positions), and the circles inside the grid represent occupied nodes.

[0403] Assuming that 48010 in Fig. 26 is the neighbor search target node, ①, ②, and ③ may be the NNs of the 48010 node found by distance-based neighbor search. Here, ①, ②, and ③ may be the order of the nearest neighbor nodes. In this case, depending on the NN search method, adjacent nodes may be selected in Morton code order or adjacent nodes may be found based on Euclidean distance.

[0404] Meanwhile, as described above, when LoDs are composed of individual slices, the attribute encoder of the transmitting device and / or the attribute decoder of the receiving device may select some of the LoDs. For example, when LoD0 to LoD2 are composed of individual slices as in FIG. 14, the attribute encoder of the transmitting device and / or the attribute decoder of the receiving device may select only LoD0, or only LoD1, or LoD2 for scalable processing. Even in this case, the NN search in FIG. 26 is performed using full layer (or full depth) geometry information.

[0405] For example, let’s assume that the decoder processes only a portion of the geometry (i.e., does not decode some high-resolution details). That is, we could assume that the decoder selects LoD1 (i.e., medium-resolution details) to decode, or that it skips the bottom 1 layer in the octree structure to decode.

[0406] In this case, it can be determined that the neighbor search target node (48020) is located on the downsampled geometry grid as in FIG. 27, and NN nodes (or points) can be obtained based on Euclidean distance. That is, in FIG. 27 as well, the grid represents the positions of the points (or geometry positions), and the circles inside the grid represent the occupied nodes.

[0407] FIG. 27 is a diagram showing an example of performing an NN search on a geometry grid that is downsampled by the difference between the full depth LoD or layer group according to the embodiments and the LoD or layer group to which the node to which the current NN search is to be performed belongs.

[0408] That is, in the attribute encoder of the transmitting device, NN search is performed on a geometry grid such as Fig. 26, whereas in the attribute decoder of the receiving device, NN search is performed on a downsampled geometry grid such as Fig. 27.

[0409] This is because the attribute decoder of the receiving device does not use high-resolution details for geometry positions, which results in geometry repositioning and changes in the relative distance between nodes.

[0410] For example, it can be seen that the location of the neighbor search target node (48010) in FIG. 26 is located below the neighbor search target node (48020) in FIG. 27. Additionally, it can be seen that the order and location of the NN points are different.

[0411] That is, the nearest neighbor node in Fig. 26 may be determined to be the farthest neighbor node in Fig. 27. In this case, the neighbor candidate considered by the attribute encoder of the transmitting device and the neighbor candidate considered by the attribute decoder of the receiving device may differ, and even if the neighbor candidates are the same, the order of the NN nodes may change. In this case, a decoder prediction error may occur in the predicting transform, which performs attribute prediction from one or more neighbor nodes. That is, since the attribute decoder of the receiving device constructs the LoD and performs NN search using the remaining layers excluding the skipped layer, this problem may occur as the geometry grid of the transmitting device and the geometry grid of the receiving device differ.

[0412] In order to solve this, the present disclosure provides an embodiment in which an attribute encoder of a transmitting device performs an NN search (i.e., NN distance) by considering a skip layer (or missing layer), and an attribute decoder of a receiving device also performs an NN search (i.e., NN distance) by considering a skip layer (or missing layer) in the same way as the NN search performed in the attribute encoder of the transmitting device.

[0413] In other words, for scalable transmission, the attribute encoder of the transmitting device performs downsampling on the remaining layers excluding the skipped layer to create a downsampled geometry grid (i.e., tree depth), and performs upsampling using information related to the skip layer to correct the positions of occupied nodes, including neighbor search target nodes. That is, the positions of the nodes are corrected by performing upsampling to a full-depth geometry grid based on information related to the skip layer. Then, NN search is performed based on the corrected positions to calculate the distance between nodes.

[0414] FIG. 28 is a diagram showing an example of performing NN search on a geometry grid with full depth position correction according to embodiments. That is, FIG. 28 is an example of correcting the position of a node (i.e., dotted line) through upsampling from a down-sampled geometry grid (i.e., solid line) assuming the case where there is no lower layer, and calculating the distance between nodes based on the corrected position.

[0415] In FIG. 28, 48030 is a neighbor search target node on a position-corrected geometry grid, and ①, ②, and ③ may be the order of NN points of the neighbor search target node on the position-corrected geometry grid.

[0416] More specifically, downsampling is performed based on LoDs or layer groups within the tree structure (e.g., octree structure) to which the node to which NN search is to be performed belongs. Upsampling can then be performed based on signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information). In this case, the signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information) may be transmitted or received by being included in a sequence parameter set and / or an attribute parameter set. This is just one example; it may also be transmitted or received by being included in a tile inventory (or tile parameter set) and / or an attribute slice header. For the sake of convenience of explanation, the signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information) is referred to as skip layer related information. Furthermore, the LoD generation and NN search described above can be applied identically to the attribute encoder of the transmitting device and the attribute decoder of the receiving device. That is, the resolution details of the NN search performed in the attribute encoder of the transmitting device are made identical to the resolution details of the NN search performed in the attribute decoder of the receiving device.

[0417] In other words, when scalable transmission and / or scalable decoding is performed, the LoD generation and NN search performed in the attribute encoder of the transmitting device are applied identically in the attribute decoder of the receiving device, so that the neighbor candidates considered by the attribute encoder of the transmitting device and the neighbor candidates considered by the attribute decoder of the receiving device do not differ, and the order of the nodes also does not change. Therefore, decoder prediction errors do not occur in predictive transformations, such as those performing attribute prediction from neighbor nodes.

[0418] In the present disclosure, skip layer related information (root_bbox_origin / root_bbox_size and / or max_geometry_level_log2 information) represents the origin (origin position) information of the bounding box and the size information of the bounding box when it is full geometry.

[0419] That is, information related to skip layers is necessary to convert a downsampled geometry grid (i.e., tree depth) into a full geometry grid when performing NN search in an encoder or decoder. In other words, it is information for creating the dotted grid in FIG. 28. That is, in the present disclosure, the encoder and decoder create a downsampled geometry grid by performing downsampling on the remaining layers excluding the skipped layer, and correct the positions of occupied nodes including neighbor search target nodes by performing upsampling using information related to skip layers. Then, an NN search is performed based on the corrected positions to obtain the NN of the neighbor search target nodes.

[0420] FIGS. 29(a) to 29(c) illustrate examples of performing NN search on a geometry grid according to embodiments. LoD represents the degree of detail of point cloud content; since a smaller LoD value indicates lower detail of point cloud content and a larger LoD value indicates higher detail of point cloud content, the detail in FIG. 29(a) is the highest (e.g., high-resolution detail), and the detail in FIG. 29(c) is the lowest (e.g., low-resolution detail).

[0421] Assuming that FIG. 29(a) is a diagram showing an example of NN search results in a geometry grid subsampled based on LoD N (e.g., full layer geometry), FIG. 29(b) is a diagram showing an example of NN search results after performing position correction considering LoD N (i.e., full layer geometry) in a geometry grid downsampled based on LoD N-1, and FIG. 29(c) is a diagram showing an example of NN search results after performing position correction considering LoD N in a geometry grid downsampled based on LoD N-2.

[0422] That is, when calculating the distance between nodes using the method described above, the same node location can be corrected to a position on a sub-sampled geometry grid of full depth (or layer) as in FIG. 29(a) according to the LoD.

[0423] Below is code showing an example of implementing NN search using the proposed method.

[0424] The code below can be explained in two main steps (STEP 1 / 2).

[0425] STEP 1: The code below is an example of a case where layer group slicing is applied, assuming that skip layers are applied at the layer group level. In this case, since it is guaranteed that data exists up to the maximum Level of Detail (max LoD) within the corresponding layer group, it is possible to maximize the utilization of available detail information by using geometry resolution up to the tree level corresponding to the max LoD. If layer group slicing is not used and slices are divided and delivered according to each LoD, details below the geometry tree level that match the LoD containing the node performing the NN search (48020 in the example above) may be omitted (i.e., the >> shiftLayerGroup part in the code below).

[0426] STEP 2: At this point, a process of upsampling the downsampled geometry grid back to the original geometry grid (i.e., the geometry grid (or location) corresponding to the full layer (or full depth)) may be included. This serves to provide a common ground that allows NN distances calculated at different LoDs to be compared on the same geometry grid (i.e., corresponding to << (shiftLayerGroup + treeLvlGap) in the code below).

[0427] That is, the code below includes a process of downsampling a geometry grid by the difference between the full depth LoD or layer group and the LoD or layer group to which the node to which the current NN search is to be performed belongs, back to the original geometry grid (i.e., the geometry grid (or location) corresponding to the full layer (or full depth)).

[0428] According to the embodiments, the code below uses a method of correcting the position to the left, front, and bottom in a downsampled geometry grid, but depending on the application, a method of correcting to the centroid (Fig. 28 and / or Fig. 29 above) may also be used.

[0429] According to the embodiments, in the code below, reeLvlGap is a parameter for geometry grid upsampling, and shiftLayerGroup is a parameter for geometry grid downsampling. Also, rootNodeSizeLog2.max() represents the maximum tree level transmitted from the encoder, and rootNodeSizeLog2_coded.max() represents the maximum tree level decoded from the decoder. Additionally, maxNumLayers_curLayerGroup represents the maximum number of layers for the current layer group, and maxNumLayers represents the maximum number of coded layers.

[0430] In other words, the parameter treeLvlGap for geometry grid upsampling can be obtained by subtracting the decoder maximum tree level (rootNodeSizeLog2_coded.max()) from the encoder maximum tree level (rootNodeSizeLog2.max()). For example, if rootNodeSizeLog2.max() is 8 and rootNodeSizeLog2_coded.max() is 6, treeLvlGap can be 2. That is, there can be 2 skip layers.

[0431] The following is code showing an example of a function implementation for finding the nearest neighbor (NN) according to the embodiments.

[0432] inline void

[0433] computeNearestNeighbors(

[0434] const AttributeParameterSet&aps,

[0435] const AttributeBrickHeader& abh,

[0436] const std::vector <mortoncodewithindex>& packedVoxel,

[0437] const std::vector<uint32_t>& retained,

[0438] int32_t startIndex,

[0439] int32_t endIndex,

[0440] int32_t lodIndex,

[0441] std::vector<uint32_t>& indexes,

[0442] std::vector <pccpredictor>& predictors,

[0443] std::vector<uint32_t>& pointIndexToPredictorIndex,

[0444] int32_t& predIndex,

[0445] MortonIndexMap3d& atlas

[0446] , const LayerGroupSlicingParams& layerGroupParams

[0447] , const int curLayerGroup)

[0448] {

[0449] constexpr auto searchRangeNear = 2;

[0450] constexpr auto bucketSizeLog2 = 5;

[0451] constexpr auto bucketSize = 1 << bucketSizeLog2;

[0452] constexpr auto bucketSizeMinus1 = bucketSize - 1;

[0453] constexpr auto levelCount = 3;

[0454]

[0455] const int32_t shiftBits = aps.scalable_lifting_enabled_flag

[0456] ? 1 + lodIndex

[0457] : 1 + aps.dist2 + abh.attr_dist2_delta + lodIndex;

[0458] const int32_t shiftBits3 = 3 * shiftBits;

[0459] const int32_t log2CubeSize = atlas.cubeSizeLog2();

[0460] const int32_t atlasBits = 3 * log2CubeSize;

[0461] / NB: when the atlas boundary is greater than 2^63, all points belong

[0462] / to a single atlas. The clipping is necessary to avoid undefined

[0463] / behaviour of shifts greater than or equal to the word size.

[0464] const int32_t atlasBoundaryBit = std::min(63, shiftBits3 + atlasBits);

[0465]

[0466] const int32_t retainedSize = retained.size();

[0467] const int32_t indexesSize = endIndex - startIndex;

[0468] const auto rangeInterLod = aps.inter_lod_search_range;

[0469] const auto rangeIntraLod = aps.intra_lod_search_range;

[0470] static const uint8_t kNeighOffset

[0027] = {

[0471] 7, / { 0, 0, 0} 0

[0472] 3, / {-1, 0, 0} 1

[0473] 5, / { 0, -1, 0} 2

[0474] 6, / { 0, 0, -1} 3

[0475] 35, / { 1, 0, 0} 4

[0476] 21, / { 0, 1, 0} 5

[0477] 14, / { 0, 0, 1} 6

[0478] 28, / { 0, 1, 1} 7

[0479] 42, / { 1, 0, 1} 8

[0480] 49, / { 1, 1, 0} 9

[0481] 12, / { 0, -1, 1} 10

[0482] 10, / {-1, 0, 1} 11

[0483] 17, / {-1, 1, 0} 12

[0484] 20, / { 0, 1, -1} 13

[0485] 34, / { 1, 0, -1} 14

[0486] 33, / { 1, -1, 0} 15

[0487] 4, / { 0, -1, -1} 16

[0488] 2, / {-1, 0, -1} 17

[0489] 1, / {-1, -1, 0} 18

[0490] 56, / { 1, 1, 1} 19

[0491] 24, / {-1, 1, 1} 20

[0492] 40, / { 1, -1, 1} 21

[0493] 48, / { 1, 1, -1} 22

[0494] 32, / { 1, -1, -1} 23

[0495] 16, / {-1, 1, -1} 24

[0496] 8, / {-1, -1, 1} 25

[0497] 0 / {-1, -1, -1} 26

[0498] };

[0499] / Part that calculates the point position based on the current layer group

[0500] int treeLvlGap = 0; / Initialize geometry grid up-sampling parameter to 0

[0501] int shiftLayerGroup = 0; / Initialize geometry grid down-sampling parameter to 0

[0502] if (layerGroupParams.layerGroupEnabledFlag) {

[0503] treeLvlGap = layerGroupParams.rootNodeSizeLog2.max() - layerGroupParams.rootNodeSizeLog2_coded.max();

[0504] / rootNodeSizeLog2.max(): Encoder maximum tree level

[0505] / rootNodeSizeLog2_coded.max(): Decoder maximum tree level

[0506]

[0507] int maxNumLayers_curLayerGroup = 0; / Maximum number of layers for the current layer-group

[0508] int maxNumLayers = 0; / Maximum number of coded layers

[0509] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[0510] maxNumLayers += layerGroupParams.numLayersPerLayerGroup[i];

[0511] if(i <= curLayerGroup)

[0512] maxNumLayers_curLayerGroup += layerGroupParams.numLayersPerLayerGroup[i];

[0513] }

[0514]

[0515] shiftLayerGroup = maxNumLayers - maxNumLayers_curLayerGroup;

[0516] }

[0517]

[0518] / The point positions biased by lodNieghBias

[0519] / todo(df): preserve this

[0520] std::vector<point_t> biasedPos;

[0521] biasedPos.reserve(packedVoxel.size());

[0522] for (const auto& src : packedVoxel) {

[0523] auto point = clacIntermediatePosition(

[0524] aps.scalable_lifting_enabled_flag, lodIndex, src.position);

[0525] biasedPos.push_back(times((point >> shiftLayerGroup) << (shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[0526] }

[0527]

[0528] atlas.reserve(retainedSize);

[0529] std::vector<int32_t> neighborIndexes;

[0530] neighborIndexes.reserve(64);

[0531]

[0532] BoxHierarchy<bucketSizeLog2, levelCount> hBBoxes;

[0533] hBBoxes.resize(retainedSize);

[0534] for (int32_t i = 0, b = 0; i < retainedSize; ++b) {

[0535] hBBoxes.insert(biasedPos[retained[i]], i);

[0536] ++i;

[0537] for (int32_t k = 1; k < bucketSize && i < retainedSize; ++k, ++i) {

[0538] hBBoxes.insert(biasedPos[retained[i]], i);

[0539] }

[0540] }

[0541] hBBoxes.update();

[0542]

[0543] BoxHierarchy<bucketSizeLog2, levelCount> hIntraBBoxes;

[0544] if (lodIndex >= aps.intra_lod_prediction_skip_layers) {

[0545] hIntraBBoxes.resize(indexesSize);

[0546] for (int32_t i = startIndex, b = 0; i < endIndex; ++b) {

[0547] hIntraBBoxes.insert(biasedPos[indexes[i]], i - startIndex);

[0548] ++i;

[0549] for (int32_t k = 1; k < bucketSize && i < endIndex; ++k, ++i) {

[0550] hIntraBBoxes.insert(biasedPos[indexes[i]], i - startIndex);

[0551] }

[0552] }

[0553] hIntraBBoxes.update();

[0554] }

[0555]

[0556] const auto bucketSize0Log2 = hBBoxes.bucketSizeLog2(0);

[0557] const auto bucketSize1Log2 = hBBoxes.bucketSizeLog2(1);

[0558] const auto bucketSize2Log2 = hBBoxes.bucketSizeLog2(2);

[0559]

[0560] int64_t curAtlasId = -1;

[0561] int64_t lastMortonCodeShift3 = -1;

[0562] int64_t cubeIndex = 0;

[0563] for (int32_t i = startIndex, j = 0; i < endIndex; ++i) {

[0564] int32_t localIndexes[3] = { -1, -1, -1};

[0565] int64_t minDistances[3] = { std::numeric_limits<int64_t>::max(),

[0566] std::numeric_limits<int64_t>::max(),

[0567] std::numeric_limits<int64_t>::max()};

[0568]

[0569] const int32_t index = indexes[i];

[0570] const auto& pv = packedVoxel[index];

[0571] const int64_t mortonCode = pv.mortonCode;

[0572] const int64_t pointAtlasId = mortonCode >> atlasBoundaryBit;

[0573] const int64_t mortonCodeShiftBits3 = mortonCode >> shiftBits3;

[0574] const int32_t pointIndex = pv.index;

[0575] const auto bpoint = biasedPos[index];

[0576] indexes[i] = pointIndex;

[0577] auto& predictor = predictors[--predIndex];

[0578] pointIndexToPredictorIndex[pointIndex] = predIndex;

[0579]

[0580] if (retainedSize) {

[0581] while (j < retainedSize - 1

[0582] && mortonCode >= packedVoxel[retained[j]].mortonCode) {

[0583] ++j;

[0584] }

[0585] / 현재 포인트의 위치가 Atlas를 벗어나는 경우 update

[0586] if (curAtlasId != pointAtlasId) {

[0587] atlas.clearUpdates();

[0588] curAtlasId = pointAtlasId;

[0589] while (

[0590] cubeIndex < retainedSize

[0591] && (packedVoxel[retained[cubeIndex]].mortonCode >> atlasBoundaryBit)

[0592] == curAtlasId) {

[0593] atlas.set(

[0594] packedVoxel[retained[cubeIndex]].mortonCode >> shiftBits3,

[0595] cubeIndex);

[0596] ++cubeIndex;

[0597] }

[0598] }

[0599]

[0600] if (lastMortonCodeShift3 != mortonCodeShiftBits3) {

[0601] lastMortonCodeShift3 = mortonCodeShiftBits3;

[0602] const auto basePosition = morton3dAdd(mortonCodeShiftBits3, -1ll);

[0603] neighborIndexes.resize(0);

[0604] for (int32_t n = 0; n < 27; ++n) {

[0605] const auto neighbMortonCode =

[0606] morton3dAdd(basePosition, kNeighOffset[n]);

[0607] if ((neighbMortonCode >> atlasBits) != curAtlasId) {

[0608] continue;

[0609] }

[0610] const auto range = atlas.get(neighbMortonCode);

[0611] for (int32_t k = range.start; k < range.end; ++k) {

[0612] neighborIndexes.push_back(k);

[0613] }

[0614] }

[0615] }

[0616] / When performing NN search based on points belonging to the top LoD

[0617] Neighbor search targeting 27 adjacent points of a plane, line, or point

[0618] for (const auto k : neighborIndexes) {

[0619] updateNearestNeigh(

[0620] bpoint, biasedPos[retained[k]], k, localIndexes, minDistances);

[0621] }

[0622] / Morton Code Order neighbor search targeting points within a certain range

[0623] if (localIndexes[2] == -1) {

[0624] const auto center = localIndexes[0] == -1 ? j : localIndexes[0];

[0625] const auto k0 = std::max(0, center - rangeInterLod);

[0626] const auto k1 = std::min(retainedSize - 1, center + rangeInterLod);

[0627] / NN search focusing on nearby points : -searchRangeNear ~ searchRangeNear range

[0628] / center location

[0629] updateNearestNeighWithCheck(

[0630] bpoint, biasedPos[retained[center]], center, localIndexes,

[0631] minDistances);

[0632]

[0633] for (int32_t n = 1; n <= searchRangeNear; ++n) {

[0634] Positive direction from center in Morton code: 1 ~ searchRangeNear

[0635] const int32_t kp = center + n;

[0636] if (kp <= k1) {

[0637] updateNearestNeighWithCheck(

[0638] bpoint, biasedPos[retained[kp]], kp, localIndexes, minDistances);

[0639] }

[0640] Negative direction from center along Morton code: -1 ~ - searchRangeNear

[0641] const int32_t kn = center - n;

[0642] if (kn >= k0) {

[0643] updateNearestNeighWithCheck(

[0644] bpoint, biasedPos[retained[kn]], kn, localIndexes, minDistances);

[0645] }

[0646] }

[0647]

[0648] const int32_t p1 =

[0649] std::min(retainedSize - 1, center + searchRangeNear + 1);

[0650] const int32_t p0 = std::max(0, center - searchRangeNear - 1);

[0651]

[0652] / 먼 거리의 포인트를 대상으로 NN search

[0653] / 양의 방향 : searchRangeNear + 1 ~ k1

[0654] / search p1...k1

[0655] const int32_t b21 = k1 >> bucketSize2Log2;

[0656] const int32_t b20 = p1 >> bucketSize2Log2;

[0657] const int32_t b11 = k1 >> bucketSize1Log2;

[0658] const int32_t b10 = p1 >> bucketSize1Log2;

[0659] const int32_t b01 = k1 >> bucketSize0Log2;

[0660] const int32_t b00 = p1 >> bucketSize0Log2;

[0661] for (int32_t b2 = b20; b2 <= b21; ++b2) {

[0662] if (

[0663] localIndexes[2] != -1

[0664] && hBBoxes.bBox(b2, 2).getDist1(bpoint) >= minDistances[2])

[0665] continue;

[0666]

[0667] const auto alignedIndex1 = b2 << bucketSizeLog2;

[0668] const auto start1 = std::max(b10, alignedIndex1);

[0669] const auto end1 = std::min(b11, alignedIndex1 + bucketSizeMinus1);

[0670] for (int32_t b1 = start1; b1 <= end1; ++b1) {

[0671] if (

[0672] localIndexes[2] != -1

[0673] && hBBoxes.bBox(b1, 1).getDist1(bpoint) >= minDistances[2])

[0674] continue;

[0675]

[0676] const auto alignedIndex0 = b1 << bucketSizeLog2;

[0677] const auto start0 = std::max(b00, alignedIndex0);

[0678] const auto end0 = std::min(b01, alignedIndex0 + bucketSizeMinus1);

[0679] for (int32_t b0 = start0; b0 <= end0; ++b0) {

[0680] if (

[0681] localIndexes[2] != -1

[0682] && hBBoxes.bBox(b0, 0).getDist1(bpoint) >= minDistances[2])

[0683] continue;

[0684]

[0685] const int32_t alignedIndex = b0 << bucketSizeLog2;

[0686] const int32_t h0 = std::max(p1, alignedIndex);

[0687] const int32_t h1 = std::min(k1, alignedIndex + bucketSizeMinus1);

[0688] for (int32_t k = h0; k <= h1; ++k) {

[0689] updateNearestNeighWithCheck(

[0690] bpoint, biasedPos[retained[k]], k, localIndexes,

[0691] minDistances);

[0692] }

[0693] }

[0694] }

[0695] }

[0696]

[0697] / 음의 방향 : p1 ~ -searchRangeNear - 1

[0698] / search k0...p1

[0699] const int32_t c21 = p0 >> bucketSize2Log2;

[0700] const int32_t c20 = k0 >> bucketSize2Log2;

[0701] const int32_t c11 = p0 >> bucketSize1Log2;

[0702] const int32_t c10 = k0 >> bucketSize1Log2;

[0703] const int32_t c01 = p0 >> bucketSize0Log2;

[0704] const int32_t c00 = k0 >> bucketSize0Log2;

[0705] for (int32_t c2 = c21; c2 >= c20; --c2) {

[0706] if (

[0707] localIndexes[2] != -1

[0708] && hBBoxes.bBox(c2, 2).getDist1(bpoint) >= minDistances[2])

[0709] continue;

[0710]

[0711] const auto alignedIndex1 = c2 << bucketSizeLog2;

[0712] const auto start1 = std::max(c10, alignedIndex1);

[0713] const auto end1 = std::min(c11, alignedIndex1 + bucketSizeMinus1);

[0714] for (int32_t c1 = end1; c1 >= start1; --c1) {

[0715] if (

[0716] localIndexes[2] != -1

[0717] && hBBoxes.bBox(c1, 1).getDist1(bpoint) >= minDistances[2])

[0718] continue;

[0719]

[0720] const auto alignedIndex0 = c1 << bucketSizeLog2;

[0721] const auto start0 = std::max(c00, alignedIndex0);

[0722] const auto end0 = std::min(c01, alignedIndex0 + bucketSizeMinus1);

[0723] for (int32_t c0 = end0; c0 >= start0; --c0) {

[0724] if (

[0725] localIndexes[2] != -1

[0726] && hBBoxes.bBox(c0, 0).getDist1(bpoint) >= minDistances[2])

[0727] continue;

[0728]

[0729] const int32_t alignedIndex = c0 << bucketSizeLog2;

[0730] const int32_t h0 = std::max(k0, alignedIndex);

[0731] const int32_t h1 = std::min(p0, alignedIndex + bucketSizeMinus1);

[0732] for (int32_t k = h1; k >= h0; --k) {

[0733] updateNearestNeighWithCheck(

[0734] bpoint, biasedPos[retained[k]], k, localIndexes,

[0735] minDistances);

[0736] }

[0737] }

[0738] }

[0739] }

[0740] }

[0741]

[0742] predictor.neighborCount = (localIndexes[0] != -1)

[0743] + (localIndexes[1] != -1) + (localIndexes[2] != -1);

[0744]

[0745] for (int32_t h = 0; h < predictor.neighborCount; ++h)

[0746] localIndexes[h] = retained[localIndexes[h]];

[0747] }

[0748]

[0749] / When performing NN search on points belonging to the same LoD as the current point.

[0750] / Only points previously created from the current point can be used. (1 ~ searchRangeNear)

[0751] if (lodIndex >= aps.intra_lod_prediction_skip_layers) {

[0752] const int32_t k00 = i + 1;

[0753] const int32_t k01 = std::min(endIndex - 1, k00 + searchRangeNear);

[0754] for (int32_t k = k00; k <= k01; ++k) {

[0755] updateNearestNeigh(

[0756] bpoint, biasedPos[indexes[k]], indexes[k], localIndexes,

[0757] minDistances);

[0758] }

[0759] const int32_t k0 = k01 + 1 - startIndex;

[0760] const int32_t k1 =

[0761] std::min(endIndex - 1, k00 + rangeIntraLod) - startIndex;

[0762]

[0763]

[0764] / search k0...k1

[0765] const int32_t b21 = k1 >> bucketSize2Log2;

[0766] const int32_t b20 = k0 >> bucketSize2Log2;

[0767] const int32_t b11 = k1 >> bucketSize1Log2;

[0768] const int32_t b10 = k0 >> bucketSize1Log2;

[0769] const int32_t b01 = k1 >> bucketSize0Log2;

[0770] const int32_t b00 = k0 >> bucketSize0Log2;

[0771] for (int32_t b2 = b20; b2 <= b21; ++b2) {

[0772] if (

[0773] localIndexes[2] != -1

[0774] && hIntraBBoxes.bBox(b2, 2).getDist1(bpoint) >= minDistances[2])

[0775] continue;

[0776]

[0777] const auto alignedIndex1 = b2 << bucketSizeLog2;

[0778] const auto start1 = std::max(b10, alignedIndex1);

[0779] const auto end1 = std::min(b11, alignedIndex1 + bucketSizeMinus1);

[0780] for (int32_t b1 = start1; b1 <= end1; ++b1) {

[0781] if (

[0782] localIndexes[2] != -1

[0783] && hIntraBBoxes.bBox(b1, 1).getDist1(bpoint) >= minDistances[2])

[0784] continue;

[0785]

[0786] const auto alignedIndex0 = b1 << bucketSizeLog2;

[0787] const auto start0 = std::max(b00, alignedIndex0);

[0788] const auto end0 = std::min(b01, alignedIndex0 + bucketSizeMinus1);

[0789] for (int32_t b0 = start0; b0 <= end0; ++b0) {

[0790] if (

[0791] localIndexes[2] != -1

[0792] && hIntraBBoxes.bBox(b0, 0).getDist1(bpoint) >= minDistances[2])

[0793] continue;

[0794]

[0795] const int32_t alignedIndex = b0 << bucketSizeLog2;

[0796] const int32_t h0 = std::max(k0, alignedIndex);

[0797] const int32_t h1 = std::min(k1, alignedIndex + bucketSizeMinus1);

[0798] for (int32_t h = h0; h <= h1; ++h) {

[0799] const int32_t k = startIndex + h;

[0800] updateNearestNeigh(

[0801] bpoint, biasedPos[indexes[k]], indexes[k], localIndexes,

[0802] minDistances);

[0803] }

[0804] }

[0805] }

[0806] }

[0807] }

[0808]

[0809] predictor.neighborCount = std::min(

[0810] aps.num_pred_nearest_neighbours_minus1 + 1,

[0811] (localIndexes[0] != -1) + (localIndexes[1] != -1)

[0812] + (localIndexes[2] != -1));

[0813] for (int32_t h = 0; h < predictor.neighborCount; ++h) {

[0814] auto& neigh = predictor.neighbors[h];

[0815] neigh.predictorIndex = packedVoxel[localIndexes[h]].index;

[0816] neigh.weight = (biasedPos[localIndexes[h]] - bpoint).getNorm2<int64_t>();

[0817] }

[0818]

[0819] / Prune neighbours based upon max neigh range.

[0820] if (aps.scalable_lifting_enabled_flag) {

[0821] int maxNeighRange = aps.max_neigh_range_minus1 + 1;

[0822] int64_t maxDistance = 3ll * maxNeighRange << 2 * lodIndex;

[0823] if (aps.lodNeighBias == 1) {

[0824] predictor.pruneDistanceGt(maxDistance);

[0825] }

[0826] else {

[0827] auto curPt = clacIntermediatePosition(true, lodIndex, pv.position);

[0828]

[0829] for (int h = 1; h < predictor.neighborCount; h++) {

[0830] auto neighPt = clacIntermediatePosition(

[0831] true, lodIndex, packedVoxel[localIndexes[h]].position);

[0832]

[0833] / Discard this and subsequent points if distance limit exceeded

[0834] auto norm2 = (curPt - neighPt).getNorm2<int64_t>();

[0835] if (norm2 > maxDistance) {

[0836] predictor.neighborCount = h;

[0837] break;

[0838] }

[0839] }

[0840] }

[0841] }

[0842]

[0843] / distance를 기반으로 weight를 구해주는 부분

[0844] if (predictor.neighborCount > 1) {

[0845] if (predictor.neighbors[0].weight > predictor.neighbors[1].weight)

[0846] std::swap(predictor.neighbors[1], predictor.neighbors[0]);

[0847] if (predictor.neighborCount == 3) {

[0848] if (predictor.neighbors[1].weight > predictor.neighbors[2].weight) {

[0849] std::swap(predictor.neighbors[2], predictor.neighbors[1]);

[0850] if (predictor.neighbors[0].weight > predictor.neighbors[1].weight)

[0851] std::swap(predictor.neighbors[1], predictor.neighbors[0]);

[0852] }

[0853] }

[0854] }

[0855] }

[0856] }

[0857] FIG. 30 is a drawing showing another example of a point cloud transmitting device according to embodiments. The elements of the point cloud transmitting device illustrated in FIG. 30 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the point cloud transmitting device of FIG. 30 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud transmitting device of FIG. 30 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud transmitting device of FIG. 30. In FIG. 30, the execution order of each block may be changed, some blocks may be omitted, and some blocks may be newly added.

[0858] According to embodiments, the point cloud transmitting device may include a data input unit (51001), a signaling processing unit (51002), a geometry encoder (51003), an attribute encoder (51004), and a transmission processing unit (51005).

[0859] The geometry encoder (51003) and attribute encoder (51004) can perform some or all of the operations described in the point cloud video encoder (10002) of FIG. 1, the encoding (20001) of FIG. 2, the point cloud video encoder of FIG. 3, and the point cloud video encoder of FIG. 8.

[0860] The data input unit (51001) according to the embodiments receives or acquires point cloud data. The data input unit (51001) may perform some or all of the operations of the point cloud video acquisition unit (10001) of FIG. 1, or may perform some or all of the operations of the data input unit (8000) of FIG. 8.

[0861] The data input unit (51001) outputs the positions of the points of the point cloud data to the geometry encoder (51003) and the attributes of the points of the point cloud data to the attribute encoder (51004). Additionally, parameters are output to the signaling processing unit (51002). Depending on the embodiments, parameters may be provided to the geometry encoder (51003) and the attribute encoder (51004).

[0862] The geometry encoder (51003) constructs an octree using the positions of input points and performs geometry compression based on the octree. At this time, the prediction for geometry compression may be performed within a frame or between frames. In this document, the former is referred to as intra-frame prediction and the latter as inter-frame prediction. The geometry encoder (51003) performs entropy encoding on the compressed geometry information and outputs it to the transmission processing unit (51005) in the form of a geometry bitstream.

[0863] The geometry encoder (51003) reconstructs geometry information based on positions changed through compression and outputs the reconstructed (or decoded) geometry information to the attribute encoder (51004).

[0864] According to the embodiments, the geometry encoder (51003) constructs an octree using the positions of input points, performs layer group-based slicing on the octree, selects one or more slices, and then performs compression of the geometry information of the selected one or more slices. Since the layer group-based slicing and slice-unit geometry compression according to the embodiments have been described in detail in FIGS. 11 to 25, they are omitted here to avoid redundant description.

[0865] The attribute encoder (51004) compresses attribute information based on positions where geometry encoding has not been performed and / or reconstructed geometry information. In one embodiment, the attribute information may be coded by combining one or more of RAHT coding, LoD-based predictive transform coding, and lifting transform coding. The attribute encoder (51004) performs entropy encoding on the compressed attribute information and outputs it to the transmission processing unit (51005) in the form of an attribute bitstream.

[0866] According to the embodiments, the attribute encoder (51004) creates a downsampled geometry grid by performing sampling on the remaining layers excluding one or more skipped layers in the case of scalable transmission as described above, and corrects the positions of occupied nodes including neighbor search target nodes by performing upsampling using information related to the skipped layers. Then, it performs NN search based on the corrected positions to obtain the distance between nodes. Then, it performs attribute prediction and compression based on the NN nodes obtained in this way. Since the position correction and NN search according to the embodiments have been described in detail in FIGS. 26 to 29, they are omitted here to avoid redundant explanation.

[0867] The signaling processing unit (51002) may generate and / or process signaling information necessary for encoding / decoding / rendering geometry information and attribute information, and provide it to a geometry encoder (51003), an attribute encoder (51004), and / or a transmission processing unit (51005). Alternatively, the signaling processing unit (51002) may receive signaling information generated by the geometry encoder (51003), the attribute encoder (51004), and / or the transmission processing unit (51005). The signaling processing unit (51002) may also provide information fed back from a receiving device (e.g., head orientation information and / or viewport information) to the geometry encoder (51003), the attribute encoder (51004), and / or the transmission processing unit (51005).

[0868] In this specification, signaling information including skip layer-related information may be signaled and transmitted in units of parameter sets (SPS: sequence parameter set, GPS: geometry parameter set, APS: attribute parameter set, TPS: Tile Parameter Set (or tile inventory), etc.). In addition, it may be signaled and transmitted in units of coding units (or compression units or prediction units) of each image, such as slices or tiles.

[0869] The transmission processing unit (51005) may perform an operation and / or transmission method identical or similar to the operation and / or transmission method of the transmission processing unit (8012) of FIG. 8, and may perform an operation and / or transmission method identical or similar to the operation and / or transmission method of the transmitter (10003) of FIG. 1. A detailed description is omitted here and refers to the description of FIG. 1 or FIG. 8.

[0870] The transmission processing unit (51005) can multiplex the geometry bitstream output from the geometry encoder (51003), the attribute bitstream output from the attribute encoder (51004), and the signaling bitstream output from the signaling processing unit (51002) into a single bitstream and then transmit it as is or encapsulate it into a file or segment and transmit it. In this document, the file is in ISOBMFF file format as one example.

[0871] According to the embodiments, the file or segment may be transmitted to a receiving device or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmission processing unit (51005) according to the embodiments can communicate with the receiving device via wired / wireless communication through a network such as 4G, 5G, 6G, etc. Additionally, the transmission processing unit (51005) can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). Additionally, the transmission processing unit (51005) may transmit encapsulated data according to an on-demand method.

[0872] FIG. 31 is a detailed block diagram of a part of an attribute encoder according to embodiments. In particular, FIG. 31 is a detailed block diagram of a LoD generation unit that performs LoD generation and NN search among the attribute encoders. The LoD generation and NN search described in FIG. 31 may be performed in the LoD generation unit (30009) of FIG. 3 or in the prediction / lifting / RAHT transformation processing unit (8009) of FIG. 8.

[0873] The elements of the point cloud encoder of FIG. 31 may be implemented in hardware, software, firmware, or a combination thereof, comprising one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud encoder of FIG. 31 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud encoder of FIG. 31. One or more memories according to the embodiments may include high-speed random access memory and may include non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0874] The LoD generation unit of FIG. 31 may include a LoD subsampling unit (55001), a geometry grid adaptation unit (55003), and an NN search unit (55005).

[0875] According to the embodiments, the LoD subsampling unit (55001) can generate one or more LoDs by performing octree-based subsampling. For example, the LoD subsampling unit (55001) can generate LoDs by performing subsampling based on a tree structure (e.g., octree) of geometry reconstructed from the geometry encoder (51003).

[0876] According to the embodiments, the geometry grid adaptation unit (55003) performs level-based geometry grid adaptation as described above to correct the positions of the nodes of the geometry grid.

[0877] At this time, geometry grid adaptation corrects the positions of the nodes by performing downsampling based on the LoD or layer group to which the node to perform NN search belongs, and then performing upsampling on the downsampled geometry grid based on skip layer information. In the present disclosure, the LoD or layer group to which the node to perform NN search belongs may be a scalable transmitted LoD or layer group.

[0878] Additionally, upsampling can be performed based on signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information). In this case, the signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information) may be transmitted by being included in a sequence parameter set and / or an attribute parameter set. This is just one example, and it may also be transmitted by being included in a tile inventory (or tile parameter set) and / or an attribute slice header.

[0879] The NN search unit (55005) according to the embodiments can find up to three NNs by a distance-based NN method. At this time, depending on the NN search method, adjacent nodes in the Molton code order may be selected as NNs, or adjacent nodes may be selected as NNs based on Euclidean distance.

[0880] For example, in Fig. 28, 48030 is a neighbor search target node on a position-corrected geometry grid, and ①, ②, and ③ may be the order of NN points of the neighbor search target node on the position-corrected geometry grid.

[0881] FIG. 32 is a drawing showing another example of a point cloud receiving device according to embodiments. The elements of the point cloud receiving device illustrated in FIG. 32 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the point cloud receiving device of FIG. 32 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud receiving device of FIG. 32 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud receiving device of FIG. 32. The execution order of each block in FIG. 32 may be changed, some blocks may be omitted, and some blocks may be newly added.

[0882] According to embodiments, the point cloud receiving device may include a receiving processing unit (61001), a signaling processing unit (61002), a geometry decoder (61003), an attribute decoder (61004), and a post-processor (61005).

[0883] The receiving processing unit (61001) according to the embodiments may receive a single bitstream, or may receive a geometry bitstream, an attribute bitstream, and a signaling bitstream, respectively. When a file and / or segment is received, the receiving processing unit (61001) according to the embodiments may decapsulate the received file and / or segment and output it as a bitstream.

[0884] A receiving processing unit (61001) according to the embodiments can, when a bitstream is received (or decapsulated), demultiplex a geometry bitstream, an attribute bitstream, and / or a signaling bitstream from the bitstream, and output the demultiplexed signaling bitstream to a signaling processing unit (61002), the geometry bitstream to a geometry decoder (61003), and the attribute bitstream to an attribute decoder (61004).

[0885] According to the embodiments, when a geometry bitstream, an attribute bitstream, and / or a signaling bitstream are each received (or decapsulated), the receiving processing unit (61001) can transmit the signaling bitstream to the signaling processing unit (61002), the geometry bitstream to the geometry decoder (61003), and the attribute bitstream to the attribute decoder (61004).

[0886] The signaling processing unit (61002) can parse and process signaling information, such as SPS, GPS, APS, TPS, metadata, etc., from an input signaling bitstream and provide it to a geometry decoder (61003), an attribute decoder (61004), and a post-processing unit (61005). In another embodiment, signaling information included in a geometry data unit header and / or an attribute data unit header may also be parsed in advance by the signaling processing unit (61002) before decoding the corresponding slice data.

[0887] According to the embodiments, the signaling processing unit (61002) can also parse and process information signaled to the sequence parameter set and / or attribute parameter set (e.g., skip layer related information) and provide it to the attribute decoder (61004).

[0888] According to the embodiments, the geometry decoder (61003) can restore the geometry by performing the inverse process of the geometry encoder (51003) of FIG. 30 based on signaling information for the compressed geometry bitstream. The geometry information restored (or reconstructed) in the geometry decoder (61003) is provided to the attribute decoder (61004). The attribute decoder (61004) can restore the attributes by performing the inverse process of the attribute encoder (51004) of FIG. 30 based on signaling information and reconstructed geometry information for the compressed attribute bitstream.

[0889] According to the embodiments, the post-processing unit (61005) can reconstruct and display / render point cloud data by matching geometry information (i.e., positions) restored and output from the geometry decoder (61003) with attribute information restored and output from the attribute decoder (61004).

[0890] According to the embodiments, the attribute decoder (61004) may also include the LoD generation unit of FIG. 31.

[0891] That is, the LoD generation and NN search described above can be applied in the same way to the attribute encoder (51004) of the transmitting device and the attribute decoder (61004) of the receiving device.

[0892] According to the embodiments, the LoD subsampling unit (55001) can generate one or more LoDs by performing octree-based subsampling. For example, the LoD subsampling unit (55001) can generate LoDs by performing subsampling based on a tree structure (e.g., octree) of geometry reconstructed in the geometry decoder (61003).

[0893] According to the embodiments, the geometry grid adaptation unit (55003) performs level-based geometry grid adaptation as described above to correct the positions of the nodes of the geometry grid.

[0894] At this time, geometry grid adaptation corrects the positions of the nodes by performing downsampling based on the LoD or layer group to which the node to be performed NN search belongs, and then performing upsampling on the downsampled geometry grid based on skip layer information. In the present disclosure, the LoD or layer group to which the node to be performed NN search belongs may be a scalable decoding LoD or layer group.

[0895] Additionally, upsampling can be performed based on signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information). In this case, the signaling information (e.g., root_bbox_origin / root_bbox_size / max_geometry_level_log2 information) may be received included in a sequence parameter set and / or an attribute parameter set. This is just one example, and it may also be received included in a tile inventory (or tile parameter set) and / or an attribute slice header.

[0896] The NN search unit (55005) according to the embodiments can find up to three NNs by a distance-based NN method. At this time, depending on the NN search method, adjacent nodes in the Molton code order may be selected as NNs, or adjacent nodes may be selected as NNs based on Euclidean distance.

[0897] For example, in Fig. 28, 48030 is a neighbor search target node on a position-corrected geometry grid, and ①, ②, and ③ may be the order of NN points of the neighbor search target node on the position-corrected geometry grid.

[0898] In this way, when scalable transmission and / or scalable decoding is performed, the LoD generation and NN search performed in the attribute encoder (51004) of the transmitting device are applied identically in the attribute decoder (61004) of the receiving device, so that the neighbor candidates considered in the attribute encoder (51004) of the transmitting device and the neighbor candidates considered in the attribute decoder (61004) of the receiving device do not differ, and the order of the nodes does not differ. Therefore, decoder prediction errors do not occur in prediction transformations, such as those performing attribute prediction from neighbor nodes.

[0899] FIG. 33 illustrates an example of a bitstream structure of point cloud data for transmission / reception according to embodiments. According to embodiments, the bitstream output from any one of the point cloud video encoders of FIG. 1, FIG. 2, FIG. 3, FIG. 8, and FIG. 30 may be in the form of FIG. 33.

[0900] According to the embodiments, the bitstream of point cloud data provides tiles or slices so that the point cloud data can be divided and processed by region. Each region of the bitstream according to the embodiments may have different importance. Accordingly, when the point cloud data is divided into tiles, different filters (encoding methods) and different filter units can be applied to each tile. Additionally, when the point cloud data is divided into slices, different filters and different filter units can be applied to each slice.

[0901] A transmitting device according to the embodiments can transmit point cloud data according to the structure of a bitstream as shown in FIG. 33, thereby enabling the application of different encoding operations depending on importance and providing a method to use a high-quality encoding method in important areas. In addition, it can support efficient encoding and transmission according to the characteristics of point cloud data and provide attribute values ​​according to user requirements.

[0902] A receiving device according to the embodiments receives point cloud data according to the structure of a bitstream as shown in FIG. 33, thereby enabling different filtering (decoding methods) to be applied to different regions (regions divided into tiles or slices) instead of using a complex decoding (filtering) method for the entire point cloud data depending on the capacity of the receiving device. Accordingly, better image quality and appropriate system latency can be guaranteed in areas important to the user.

[0903] When a geometry bitstream, an attribute bitstream, and / or a signaling bitstream (or signaling information) according to the embodiments is composed of a single bitstream (or G-PCC bitstream) as shown in FIG. 33, the bitstream may include one or more sub-bitstreams. The bitstream according to the embodiments may include a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling of geometry information coding, one or more Attribute Parameter Sets (APS0, APS1) for signaling of attribute information coding, a Tile Inventory (or TPS) for tile-level signaling, and one or more slices (slice 0 to slice n). That is, the bitstream of point cloud data according to the embodiments may include one or more tiles, and each tile may be a group of slices including one or more slices (slice 0 to slice n). A tile inventory (i.e., TPS) according to the embodiments may include information regarding each tile (e.g., coordinate information of a tile bounding box and height / size information, etc.) for one or more tiles. Each slice may include one geometry bitstream (Geom0) and / or one or more attribute bitstreams (Attr0, Attr1). For example, slice 0 includes one geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0 It may include ).

[0904] The geometry bitstream within each slice may consist of a geometry slice header (geom_slice_header) and geometry slice data (geom_slice_data). According to embodiments, the geometry bitstream within each slice may be referred to as a geometry data unit, the geometry slice header as a geometry data unit header, and the geometry slice data as geometry data unit data. The geometry slice header (or geometry data unit header) according to embodiments may include identification information of a parameter set included in a geometry parameter set (GPS) (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information regarding data included in the geometry slice data (geom_slice_data) (geomBoxOrigin, geom_box_log2_scale, geom_max_node_size_log2, geom_num_points), etc. geomBoxOrigin is geometry box origin information representing the box origin of the geometry slice data, geom_box_log2_scale is information representing the log scale of the geometry slice data, geom_max_node_size_log2 is information representing the size of the root geometry octree node, and geom_num_points is information related to the number of points of the geometry slice data. The geometry slice data (or geometry data unit data) according to the embodiments may include geometry information (or geometry data) of the point cloud data within the slice.

[0905] Each attribute bitstream within each slice may consist of an attribute slice header (attr_slice_header) and attribute slice data (attr_slice_data). According to embodiments, the attribute bitstream within each slice may be referred to as an attribute data unit, the attribute slice header as an attribute data unit header, and the attribute slice data as attribute data unit data. The attribute slice header (or attribute data unit header) according to embodiments may include information regarding the corresponding attribute slice data (or the corresponding attribute data unit), and the attribute slice data may include attribute information of the point cloud data within the slice (or referred to as attribute data or attribute value). If there are multiple attribute bitstreams within a single slice, each may include different attribute information. For example, one attribute bitstream may contain attribute information corresponding to color, and another attribute stream may contain attribute information corresponding to reflectance.

[0906] According to the embodiments, parameters required for encoding and / or decoding point cloud data may be newly defined in parameter sets of point cloud data (e.g., SPS, GPS, APS, and TPS (or tile inventory), etc.) and / or in the header of the corresponding slice. For example, when performing encoding and / or decoding of geometry information, they may be added to the geometry parameter set (GPS), and when performing tile-based encoding and / or decoding, they may be added to the tile and / or slice header.

[0907] According to the embodiments, skip layer related information can be signaled to a sequence parameter set and / or an attribute parameter set.

[0908] According to the embodiments, skip layer related information may be signaled in the tile parameter set and / or attribute slice header.

[0909] According to the embodiments, if the syntax element defined below can be applied to multiple point cloud data streams as well as the current point cloud data stream, skip layer-related information can be transmitted through a parameter set of a higher concept.

[0910] According to the embodiments, information related to the skip layer may be defined in a corresponding location or a separate location depending on the application or system, allowing for different application scopes, application methods, etc. The term "field," as used in the syntax of this specification described below, may have the same meaning as "parameter" or "syntactic element."

[0911] According to the embodiments, parameters (which may be variously referred to as metadata, signaling information, etc.) containing skip layer-related information may be generated in a metadata processing unit (or metadata generator) or a signaling processing unit of a transmitting device and transmitted to a receiving device to be used in a decoding / reconstruction process. For example, parameters generated and transmitted by the transmitting device may be obtained from a metadata parser of the receiving device.

[0912] FIGS. 34a to 34c illustrate an example of the syntax structure of a sequence parameter set (seq_parameter_set()) (SPS) according to the present specification. The SPS may include sequence information of a point cloud data bitstream, and in particular, an example is shown that includes information related to a skip layer.

[0913] The syntax of FIGS. 34a to 34c is included in the bitstream of FIG. 33, is generated by a point cloud encoder according to the embodiments, and can be decoded by a point cloud decoder.

[0914] In FIGS. 34a to 34c, the simple_profile_compatibility_flag field may indicate whether the bitstream follows the main profile. For example, if the value of the simple_profile_compatibility_flag field is 1, it may indicate that the bitstream follows the simple profile. For example, if the value of the simple_profile_compatibility_flag field is 0, it may indicate that the bitstream follows a profile other than the simple profile.

[0915] If the value of the unique_point_positions_constraint_flag field is 1, all output points in each point cloud frame referenced by the current SPS can have unique positions. If the value of the unique_point_positions_constraint_flag field is 0, two or more output points in any point cloud frame referenced by the current SPS can have the same position. For example, even if all points are unique in each slice, points different from the slices within the frame may overlap. In that case, the value of the unique_point_positions_constraint_flag field is set to 0.

[0916] The level_idc field indicates the level that the bitstream follows.

[0917] The sps_seq_parameter_set_id field provides an identifier for the SPS for reference by other syntax elements.

[0918] The sps_num_attribute_sets field indicates the number of coded attributes in the bitstream.

[0919] The SPS according to the embodiments includes a loop that repeats as many times as the value of the sps_num_attribute_sets field. In one embodiment, i is initialized to 0 and increases by 1 each time the loop is executed, and the loop repeats until the value of i becomes the value of the sps_num_attribute_sets field. This loop may include the attribute_dimension_minus1[i] field and the attribute_instance_id[i] field. attribute_dimension_minus1[i] plus 1 represents the number of components of the i-th attribute.

[0920] The attribute_instance_id[i] field represents the instance identifier of the i-th attribute.

[0921] The known_attribute_label_flag[i] field indicates whether the known_attribute_label[i] field or the attribute_label_four_bytes[i] field is signaled for the i-th attribute. For example, if the value of the known_attribute_label_flag[i] field is 0, it indicates that the known_attribute_label[i] field is signaled for the i-th attribute, and if the value of the known_attribute_label_flag[i] field is 1, it indicates that the attribute_label_four_bytes[i] field is signaled for the i-th attribute.

[0922] The known_attribute_label[i] field indicates the type of the i-th attribute. For example, if the value of the known_attribute_label[i] field is 0, it indicates that the i-th attribute is color; if the value of the known_attribute_label[i] field is 1, it indicates that the i-th attribute is reflectance; and if the value of the known_attribute_label[i] field is 2, it indicates that the i-th attribute is frame index. Additionally, if the value of the known_attribute_label[i] field is 4, it indicates that the i-th attribute is transparency; and if the value of the known_attribute_label[i] field is 5, it indicates that the i-th attribute is normals.

[0923] The axis_coding_order field indicates the correspondence between the three position components in the reconstructed point cloud RecPic [pointidx] [axis] with X, Y, Z output axis labels and axis=0..2.

[0924] If the value of the bypass_stream_enabled_flag field is 1, it indicates that the bypass coding mode is used to read the bitstream. As another example, if the value of the bypass_stream_enabled_flag field is 0, it indicates that the bypass coding mode is not used to read the bitstream.

[0925] The sps_extension_flag field indicates whether the sps_extension_data syntax structure exists in the corresponding SPS syntax structure. For example, if the value of the sps_extension_present_flag field is 1, it indicates that the sps_extension_data syntax structure exists in this SPS syntax structure, and if it is 0, it indicates that it does not exist. The SPS according to the embodiments may further include the sps_extension_data_flag field if the value of the sps_extension_flag field is 1. The sps_extension_data_flag field may have any value.

[0926] If layer_group_enabled_flag is 1, it indicates that the geometry (or attribute) bitstream of a slice is contained in multiple slices that match the coding layer group or its subgroups. If layer_group_enabled_flag is 0, it indicates that the geometry (or attribute) bitstream is contained in a single slice.

[0927] num_layer_groups_minus1 plus 1 represents the number of layer groups, which are groups of consecutive tree layers that are part of the geometry coding tree structure. num_layer_groups_minus1 is in the range of 0 to the number of coding tree layers.

[0928] The layer_group_id field represents the identifier of the slice's layer group. The range of layer_group_id is from 0 to num_layer_groups_minus1.

[0929] num_layers_minus1 plus 1 represents the number of coding layers included in the i-th layer group. The total number of layer groups can be derived by adding all (num_layers_minus1[i] + 1) from i 0 to num_layer_groups_minus1.

[0930] If the subgroup_enabled_flag field is 1, it indicates that the i-th layer group is divided into two or more subgroups. In this case, the point sets in the subgroups are identical to the point sets in the layer group. When the subgroup_enabled_flag of the i-th layer group is 1, if j is greater than or equal to i, the subgroup_enabled_flag of the j-th layer group will be 1. If subgroup_enabled_flag is 0, it indicates that the current layer group is contained in a single slice without being divided into multiple subgroups.

[0931] subgroup_bbox_origin_bits_minus1 plus 1 indicates the bit length of the syntax element subgroup_bbox_origin.

[0932] subgroup_bbox_size_bits_minus1 plus 1 indicates the bit length of the syntax element subgroup_bbox_size.

[0933] According to the embodiments, if layer_group_enabled_flag is 1, the SPS may include at least one of the max_geometry_level_log2 field, root_bbox_origin_x field, root_bbox_origin_y field, root_bbox_origin_z field, root_bbox_size_x field, root_bbox_size_y field or root_bbox_size_z field, and this may be referred to as skip layer related information.

[0934] In the present disclosure, information related to skip layers may include information on the origin (origin position) of the bounding box when it is full geometry and information on the size of the bounding box. Information related to skip layers is information necessary to convert a downsampled geometry grid (i.e., tree depth) into a full geometry form when performing NN search in an encoder or decoder. That is, in the present disclosure, the encoder and decoder create a downsampled geometry grid by performing sampling on the remaining layers excluding the skipped layer, and correct the positions of occupied nodes including neighbor search target nodes by performing upsampling using information related to skip layers. Then, an NN search is performed based on the corrected positions to obtain the NN of the neighbor search target nodes.

[0935] In the present disclosure, all of the max_geometry_level_log2 field, root_bbox_origin_x field, root_bbox_origin_y field, root_bbox_origin_z field, root_bbox_size_x field, root_bbox_size_y field, or root_bbox_size_z field may be transmitted, but only the max_geometry_level_log2 field may be transmitted, or the root_bbox_origin_x field, root_bbox_origin_y field, root_bbox_origin_z field, root_bbox_size_x field, root_bbox_size_y field, and root_bbox_size_z field may be transmitted.

[0936] The above max_geometry_level_log2 field can represent the maximum geometry level in the case of full geometry. The value of the max_geometry_level_log2 field can be obtained as shown in the example below.

[0937] bound = quantInput.computeBoundingBox().max - quantInput.computeBoundingBox().min + 1;

[0938] Here, 'bound' represents a tight (i.e., accurate) bounding box calculated by finding the maximum and minimum values ​​along the x, y, and z axes using the actual point distribution. The fields root_bbox_origin_x, root_bbox_origin_y, and root_bbox_origin_z represent the origin of this bounding box, and the fields root_bbox_size_x, root_bbox_size_y, and root_bbox_size_z represent the size of this bounding box.

[0939] for (int k = 0; k < 3; k++)

[0940] rootNodeSizeLog2[k] = ceillog2(std::max(2, bound[k]));

[0941] max_geometry_level_log2 = rootNodeSizeLog2.max();

[0942] Here, rootNodeSizeLog2[k] represents the bounding box size of the corresponding axis (k), and max_geometry_level_log2 represents the maximum value among the bounding box sizes of the x, y, and z axes.

[0943] The above root_bbox_origin_x, root_bbox_origin_y, and root_bbox_origin_z fields can represent the origin (i.e., the location of the origin) at the left, front, and bottom positions of the minimum bounding box capable of containing the entire point cloud data in the case of full geometry.

[0944] The above root_bbox_size_x, root_bbox_size_y, and root_bbox_size_z fields may represent the length (i.e., size) in the x-axis, y-axis, and z-axis directions from the origin (i.e., the position of the origin) at the left, front, and bottom positions of the minimum bounding box that can contain the entire point cloud data in the case of full geometry.

[0945] In addition, if it is agreed in advance that point cloud data has been moved to a specific location (e.g., (0, 0, 0)), the origin value is not passed but can be inferred according to the agreement.

[0946] FIGS. 35a to 35c illustrate an embodiment of the syntax structure of an attribute parameter set (attr_parameter_set()) (APS) according to the present specification. The SPS may include sequence information of a point cloud data bitstream. The APS according to the embodiments may include information regarding a method for encoding attribute information of point cloud data contained in one or more slices, and in particular, an example is shown that includes information related to a skip layer.

[0947] The syntax of FIGS. 35a to 35c is included in the bitstream of FIG. 33, is generated by a point cloud encoder according to the embodiments, and can be decoded by a point cloud decoder.

[0948] The aps_attr_parameter_set_id field represents the identifier of the APS for reference by other syntax elements.

[0949] The aps_seq_parameter_set_id field represents the value of sps_seq_parameter_set_id for the active SPS.

[0950] The attr_coding_type field indicates the coding type for the attribute.

[0951] In one embodiment, if the value of the attr_coding_type field is 0, the coding type may indicate Region Adaptive Hierarchical Transform (RAHT); if 1, the coding type may indicate predicting weight lifting (or LoD with predicting transformation); if 2, the coding type may indicate fixed weight lifting (or LoD with lifting transform); and if 3, raw attribute data.

[0952] According to the embodiments, if the value of the attr_coding_type field is 2 or greater than 2, for example, if the coding type is LoD with lifting transform, the APS is

[0953] It may include at least one of the following fields: pred_set_size_minus1, pred_inter_lod_search_range, pred_dist_bias_minus1_xyz[k], last_comp_pred_enabled, lod_scalability_enabled, pred_max_range_minus1, max_geometry_level_log2, root_bbox_origin_x, root_bbox_origin_y, root_bbox_origin_z, root_bbox_size_x, root_bbox_size_y, and root_bbox_size_z.

[0954] The value obtained by adding 1 to the pred_set_size_minus1 field value specifies the maximum size of the per-point predictor set.

[0955] The pred_inter_lod_search_range field specifies the range of searchable indexes around a search center to find nearest neighbors to include in a point's predictor set in an extended inter-detail-level search.

[0956] The value obtained by adding 1 to the pred_dist_bias_minus1_xyz[k] field specifies the factor used to weight the k-th XYZ component of the distance vector between two point positions used to calculate inter-point distances in the predictor search for a single refinement point.

[0957] The expression PredBias[k] specifies the factor for the k-th STV component as follows.

[0958] PredBias[k] := pred_dist_bias_minus1_xyz[StvToXyz[k]] + 1

[0959] That is, the value obtained by adding 1 to the bias value of the XYZ component corresponding to the STV component k is used as the weight.

[0960] The last_comp_pred_enabled field indicates whether the second coefficient component among the three component attributes is used to predict the value of the third coefficient component. For example, if the value of the last_comp_pred_enabled field is 1, it is used, and if it is 0, it is not used. Also, if the last_comp_pred_enabled field does not exist, it is inferred as 0.

[0961] The lod_scalability_enabled field indicates whether attribute values ​​are coded using restricted LoD generation and predictor search. For example, if the value of the lod_scalability_enabled field is 1, it indicates that it is used, and if it is 0, it indicates that it is not used. That is, if the value of the lod_scalability_enabled field is 1, the attribute values ​​can be reconstructed for a partially decoded occupancy tree. If the lod_scalability_enabled field does not exist, it is inferred as 0.

[0962] The value of the pred_max_range_minus1 field plus 1 specifies, when present, the distance beyond which point predictor candidates shall be discarded during predictor set pruning for scalable attribute coding. This distance is expressed in units of the per-detail-level block size.

[0963] According to the embodiments, if the lod_scalability_enabled field is 1, the APS may further include at least one of the max_geometry_level_log2 field, the root_bbox_origin_x field, the root_bbox_origin_y field, the root_bbox_origin_z field, the root_bbox_size_x field, the root_bbox_size_y field, or the root_bbox_size_z field, which may be referred to as skip layer related information. In the present disclosure, the skip layer related information may include bounding box origin (origin position) information and bounding box size information when the geometry is full. The skip layer related information is information required to convert a downsampled geometry grid (i.e., tree depth) into a full geometry form when performing NN search in an encoder or decoder. That is, in the present disclosure, the encoder and decoder perform sampling with the remaining layers excluding the skipped layer to create a downsampled geometry grid, and perform upsampling using information related to the skipped layer to correct the positions of the occupied nodes including the neighbor search target nodes. Then, based on the corrected positions, an NN search is performed to obtain the NN of the neighbor search target nodes.

[0964] In the present disclosure, all of the max_geometry_level_log2 field, root_bbox_origin_x field, root_bbox_origin_y field, root_bbox_origin_z field, root_bbox_size_x field, root_bbox_size_y field, or root_bbox_size_z field may be transmitted, but only the max_geometry_level_log2 field may be transmitted, or the root_bbox_origin_x field, root_bbox_origin_y field, root_bbox_origin_z field, root_bbox_size_x field, root_bbox_size_y field, and root_bbox_size_z field may be transmitted.

[0965] The above max_geometry_level_log2 field can represent the maximum geometry level in the case of full geometry. The value of the max_geometry_level_log2 field can be obtained as shown in the example below.

[0966] bound = quantInput.computeBoundingBox().max - quantInput.computeBoundingBox().min + 1;

[0967] Here, bound represents a tight (i.e., accurate) bounding box obtained by finding the maximum and minimum values ​​along the x, y, and z axes using the actual point distribution (i.e., full geometry). The above-mentioned root_bbox_origin_x, root_bbox_origin_y, and root_bbox_origin_z fields represent the origin of this bounding box, and the above-mentioned root_bbox_size_x, root_bbox_size_y, and root_bbox_size_z fields represent the size of this bounding box.

[0968] for (int k = 0; k < 3; k++)

[0969] rootNodeSizeLog2[k] = ceillog2(std::max(2, bound[k]));

[0970] max_geometry_level_log2 = rootNodeSizeLog2.max();

[0971] Here, rootNodeSizeLog2[k] represents the bounding box size of the corresponding axis (k), and max_geometry_level_log2 represents the maximum value among the bounding box sizes of the x, y, and z axes.

[0972] The above root_bbox_origin_x, root_bbox_origin_y, and root_bbox_origin_z fields can represent the origin (i.e., the location of the origin) at the left, front, and bottom positions of the minimum bounding box capable of containing the entire point cloud data in the case of full geometry.

[0973] The above root_bbox_size_x, root_bbox_size_y, and root_bbox_size_z fields may represent the length (i.e., size) in the x-axis, y-axis, and z-axis directions from the origin (i.e., the position of the origin) at the left, front, and bottom positions of the minimum bounding box that can contain the entire point cloud data in the case of full geometry.

[0974] In addition, if it is agreed in advance that point cloud data has been moved to a specific location (e.g., (0, 0, 0)), the origin value is not passed but can be inferred according to the agreement.

[0975] The receiving method / device according to the embodiments has the following effects.

[0976] The present disclosure describes a method for dividing and transmitting compressed data according to certain criteria for point cloud data. In particular, as one application of the present disclosure, when layered coding (or hierarchical coding) is used, the compressed data can be divided and sent according to the layer, which increases the storage and transmission efficiency of the transmitting end. In particular, when applying scalable attribute coding, there is a disadvantage that the entire coded geometry data must be transmitted; however, through the proposal of the present disclosure, only the geometry layer that matches the level used in scalable attribute coding can be transmitted, so the transmission efficiency is further increased in the present disclosure.

[0977] Next, we will describe another embodiment of the aforementioned layer group slicing, that is, fine granularity slicing (FGS). In this case, some of the content may overlap with the aforementioned layer group slicing content.

[0978] Modifications and combinations between the embodiments in this disclosure are possible. The terms used in this document may be understood based on their intended meanings within the scope of their common usage in the art. The embodiments described below in this disclosure may be considered together with the disclosures described in the aforementioned embodiments.

[0979] First, for the sake of convenience of explanation, we will define some terms.

[0980] FGS (fine granularity slices) are a subset of slices that carry geometry or attributes of a subgroup within a layer group.

[0981] A layer group is a group of consecutive tree levels of an occupancy tree or an attribute-assigned occupancy tree.

[0982] A subgroup is a spatial subset of a layer-group, wherein the bounding box of a subgroup does not overlap with other subgroups in the same layer-group.

[0983] The root layer group is a layer group that includes the root node of the occupancy tree or the attribute-assigned occupancy tree.

[0984] A parent subgroup is a subgroup within a layer-group adjacent to the top layer minimum depth of the current subgroup, where the bounding box of the parent subgroup is a superset of the bounding box of the current subgroup.

[0985] A child subgroup is a subgroup within a layer-group adjacent to the bottom layer maximum depth of the current subgroup, where the bounding box of the child subgroup is a subset of the bounding box of the current subgroup.

[0986] An attribute-assigned occupancy tree is an occupancy tree where the attributes of a node in each tree level are assigned by the attributes of its child nodes.

[0987] At this time, one slice may consist of FGSs, and each FGS is mapped 1:1 to the FGS geometry or FGS attribute of a subgroup within the layer group.

[0988] FGSs or slices of FGSs within a slice are identified by a common slice identifier (slice_id).

[0989] In the present disclosure, FGS may be indicated by a pair of layer group indices (e.g., layer_group_id) and subgroup indices (e.g., subgroup_id) or by fgs_id. If FGSs have the same pair of indices, the FGSs may be one of the geometry or attributes of the subgroup indicated by the layer group index and the subgroup index.

[0990] All FGSs include geometry data units (GDUs) or dependent geometry data units (DGDUs) that code partial slice geometry, or attribute data units (ADUs) or dependent attribute data units (DADUs) that code partial slice attributes.

[0991] The first FGS in the slice is the FGS of the GDU. This FGS may be followed by the FGS of the DGDU, which vary depending on the previously decoded GDU and DGDU.

[0992] In the slice attribute, the first FGS is the ADU's FGS. This FGS may be followed by the DADU's FGS, which depend on the corresponding ADU or previously decoded ADUs and DADUs. The FGS of ADUs and DADUs occur after the FGS of GDUs and DGDUs.

[0993] If no attributes are present, that is, if only geometry is transmitted or only geometry FGS is decoded, the present disclosure may use predetermined attributes for geometry nodes. This may also apply to nodes generated in the case of partial region or depth decoding.

[0994] In the layer group structure of the present disclosure, nodes within an occupancy tree or an attribute-assigned occupancy tree are grouped into layer groups and subgroups.

[0995] A layer group is a group of consecutive tree levels, where every tree level belongs to only one layer group. The minimum depth of a subgroup is the maximum depth of the parent subgroup plus 1, or 0 in the case of the root layer group. The maximum depth of a subgroup is the minimum depth of the child subgroup minus 1, or the maximum depth of the occupancy tree in the case of the last layer group. In this case, a layer group is identified by the layer group index (layer_group_idx).

[0996] A subgroup is a spatial subset of a layer group, and a node within a tree level belongs to one of the subgroups of the layer group.

[0997] The range of locations of nodes in a subgroup is described by bounding boxes, and the range of a subgroup bounding box does not overlap with the bounding boxes of other subgroups within the same layer group.

[0998] The set of nodes of all subgroups within a single layer group is the same as the set of nodes within that layer group.

[0999] For the root layer group, there is only one subgroup. Subgroups within a layer group are identified by the subgroup index (subgroup_idx).

[1000] FIGS. 36(a) and FIGS. 36(b) are diagrams showing the relationship between layer groups, subgroups, and FGS according to embodiments. In particular, FIG. 36(a) shows the layer group structure of an Occupancy tree with a maximum depth of 8, and FIG. 36(b) shows the parent-child relationship between subgroups and their corresponding FGSs.

[1001] In FIG. 36(a), three layer groups are defined, with layer group 0 consisting of depths 0-3, layer group 1 consisting of depths 4-6, and layer group 2 consisting of depths 7 and 8. That is, layer group 0 consists of one subgroup, layer group 1 consists of two subgroups, and layer group 2 consists of three subgroups. In other words, excluding the root layer group, each layer group can be composed of subgroups, and each subgroup is represented by a pair consisting of a layer group index and a subgroup index. For example, the root layer group is represented as (0, 0).

[1002] In FIG. 36(b), the spatial region of the subgroups in FIG. 36(a) is illustrated by a rectangular bounding box in the xy plane. When the bounding box of a subgroup in a layer-group is a superset of the bounding box of one or more subgroups in the next layer-group, the subgroups in adjacent layer-groups are in a parent-child relationship. For example, subgroup (0,0) is the parent of subgroups (1,0) and (1,1). Similarly, subgroups (2,0) and (2,1) are children of subgroup (1,0). Each subgroup, represented by a pair of layer-group index and subgroup index, is located in different FGSs.

[1003] In FIG. 36(a) and FIG. 36(b), S l,n represents subgroup n associated with layer group l. And, d i represents the octree tree depth i.

[1004] According to the embodiments, when LoD scalability is applied to an FGS having a structure such as that shown in FIG. 36(a) and FIG. 36(b), partial density can be provided by not decoding fine details within a specific layer group. That is, in the case of an FGS, partial density decoding can be provided in units of LoD level bundles by skipping layer groups, and when LoD scalability is enabled, partial density decoding can be supported more finely by skipping in units of LoD levels within a layer group. In other words, LoD scalability can be applied even within an FGS, and in this case, partial density decoding can be supported by performing output without selectively decoding finer detail levels within a specific layer group.

[1005] In the present disclosure, a finer LoD refers to a LoD containing more points, and a coarser LoD refers to a LoD containing fewer points, and a refinement point refers to a point that exists in the corresponding LoD but not in the coarser LoD. The finer LoD may be used interchangeably with the finer detail level, and the coarser LoD may be used interchangeably with the coarser LoD.

[1006] Additionally, as shown in FIG. 36(a) and FIG. 36(b), each subgroup includes a part of the Occupancy Tree, and each subgroup includes one or more depths and LoD levels of the Occupancy Tree. That is, each subgroup includes a part of the Occupancy Tree and may include one or more Occupancy Tree depths and a plurality of corresponding LoDs.

[1007] Additionally, LoD scalability can be applied within FGS, in which case output is performed without decoding N levels from the subgroup bottom layer, corresponding to the number of skip layers N. That is, when LoD scalability is enabled, skipping can be performed by not decoding N detail levels from the subgroup bottom layer. In this case, if skip levels exist, the number of points in the lower LoD (i.e., finer LoD) may be lower than the number of points in the upper LoD (i.e., coarser LoD) during LoD generation due to a lack of geometry information. In other words, if skipped detail levels exist, the monotonic increase in the number of points between detail levels may not be maintained due to the lack of geometry information at those detail levels; consequently, during the LoD generation process, the number of points in the finer LoD may become smaller than the number of points in the coarser LoD. In such cases, the present disclosure may consider a finer LoD combined with a coarser LoD, and to this end, obtains a minimum inter-level prediction. That is, the present disclosure determines a minimum reference detail level for inter-level prediction through an encoder and / or decoder, and performs inter-level prediction based thereon.

[1008] In the present disclosure, inter-level prediction is performed using a set of predictors constructed through an inter-detail-level search. Here, an inter-detail-level search is a process of constructing a set of predictors by searching for spatially adjacent points among points belonging to a coarser detail level for a refinement point belonging to the current detail level. Meanwhile, in the present disclosure, the term "inter-level" refers to the relationship between detail levels and is used with the same meaning as "inter-detail-level."

[1009] In the present disclosure, the minimum reference detail level may be determined as a detail level where the number of refinement points corresponding to a specific detail level is less than the cumulative number of refinement points corresponding to detail levels that are finer than that detail level, and the minimum inter-level prediction may be performed by predicting the attribute values ​​of points belonging to finer detail levels using points included in the minimum reference detail level. That is, in the present disclosure, the minimum inter-level prediction refers to an inter-level prediction performed based on a detail level determined as the minimum reference detail level (MinInterRefLvl) among the detail levels that can be used as references.

[1010] FIGS. 37(a) to 37(c) are drawings illustrated to compare the distribution of the number of attribute points by detail level in an environment where LoD scalability is applied.

[1011] FIG. 38 is an enlarged graph showing the distribution of the number of points in a specific detail level range when LoD scalability according to the embodiments is applied. In particular, FIG. 38 is an enlarged view of the graph in FIG. 37(c).

[1012] More specifically, Fig. 37(a) shows example point cloud data (e.g., frog_00067_vox12). This data can be represented as an FGS-based structure containing multiple Levels of Detail (LoDs), where each LoD contains geometry and attribute information at different levels of precision. Fig. 37(b) is a graph illustrating the attribute number of LoDs per level of detail in the CTC Anchor method. That is, it can be seen that the number of attribute points increases as the LoD becomes finer. Figs. 37(c) and 38 show the attribute number of LoDs per level of detail when the Scalable Lifting Scheme (CE13.15) is applied. In Figs. 37(b), 37(c), and 38, the vertical axis represents the LoD level, and the horizontal axis represents the number of points contained in the corresponding LoD. Figures 37(c) and 38 show the result of skipping one LoD level for an MPEG test sequence, where it can be seen that the number of points in LoD 12 is fewer than in LoD 11. In other words, an inversion phenomenon is observed where the number of points in LoD 12 is lower than in LoD 11. This can occur when a specific LoD level is skipped, as in the example, or when some geometry depths are not generated due to partial decoding. That is, as illustrated in Figures 37(c) and 38, when a specific LoD level is skipped, the number of points in a finer detail level may be lower than that in a less fine detail level. This implies that the general hierarchical assumption that a finer LoD always contains more points is broken. In such situations, inter-level prediction using the most finer LoD as the reference level may degrade prediction efficiency. In other words, in such a distribution structure, it is not appropriate to select the reference level simply based on the layer order when performing inter-level prediction.

[1013] Accordingly, the present disclosure provides, as an embodiment, determining the minimum reference detail level (MinInterRefLvl) by comparing the number of refinement points corresponding to each LoD with the cumulative number of refinement points of finer LoDs to solve this problem.

[1014] For example, as shown in Fig. 37(c) and Fig. 38, if the number of points in LoD 12 is less than that in LoD 11, LoD 12 can be considered combined with 11, and to do this, a minimum inter-level prediction is performed.

[1015] In other words, the present disclosure enables stable inter-level prediction even when the number of refinement points of a specific LoD is less than the cumulative number of refinement points of more fine LoDs, by determining a minimum reference detail level through a refinement point-based comparison condition. That is to say, the present disclosure enables stable inter-level prediction even in environments where LoD skips exist by determining a minimum reference detail level by comparing the refinement point distribution of each detail level.

[1016] Specifically, the encoder and / or decoder of the present disclosure compares the number of refinement points corresponding to each detail level with the cumulative number of refinement points corresponding to detail levels that are finer than the corresponding detail level. Then, if the number of refinement points of the corresponding detail level is fewer than the cumulative number of refinement points of the finer detail levels, the detail level is determined as the minimum reference detail level for inter-level prediction.

[1017] This method provides a stable prediction reference standard even when the number of points between LoDs is reversed, and enables the construction of a more reliable set of neighbor points for the target point to be predicted.

[1018] The following is a detailed explanation of how to determine the minimum reference detail level (MinInterRefLvl) for inter-level predictor searches.

[1019] According to the embodiments, the variable MinInterRefLv identifies the finest detail level used as a reference for inter-detail level prediction. That is, the minimum reference detail level is identified by the variable MinInterRefLvl. The variable MinInterRefLvl represents the finest detail level among the detail levels that can be used as a reference for inter-detail level prediction.

[1020] In particular, when lod_scalability_enabled is set to 1, the variable MinInterRefLvl is determined as the finest detail level having fewer refinement points than the total number of refinement points associated with all finer detail levels. That is, the variable MinInterRefLvl is determined as the finest detail level among the detail levels for which the number of refinement points corresponding to that detail level is less than the total number of refinement points corresponding to the finer detail levels.

[1021] An exemplary code algorithm for implementing this is as follows.

[1022] MinInterRefLvl = 1

[1023] if (lod_scalability_enabled) {

[1024] for (lvl = 1; lvl < LodCnt 1; lvl++) {

[1025] if (LodRfmtPtCnt[lvl] < slice_num_points_minus1 LodPtCnt[lvl])

[1026] break

[1027] MinInterRefLvl++

[1028] }

[1029] }

[1030] In the above algorithm, MinInterRefLvl is set to an initial value of 1, and when the value of lod_scalability_enabled is true, that is, when LoD scalability is enabled, the minimum reference detail level is determined by repeatedly evaluating the condition for each detail level.

[1031] The above LodRfmtPtCnt[lvl] is the number of refinement points newly added to the detail level lvl, and slice_num_points_minus1 LodPtCnt[lvl] is the total number of refinement points existing in detail levels that are more fine than the corresponding detail level.

[1032] According to the above code algorithm, it is determined whether the number of refinement points at the current detail level (lvl) is less than the total number of refinement points (i.e., the sum) existing at finer detail levels.

[1033] If the number of refinement points at the current detail level (lvl) is less than the total number of refinement points (i.e., the sum) in finer detail levels, the loop terminates immediately, and the current detail level is determined as the minimum reference detail level (MinInterRefLvl) to be used for inter-level prediction. In other words, the first detail level to satisfy this condition becomes MinInterRefLvl. If the above condition is not satisfied, the current detail level is not yet the minimum reference detail level, so the value of MinInterRefLvl is increased to examine the next detail level as a candidate, and the above process is repeated.

[1034] In other words, it is determined whether the number of refinement points at the current detail level is less than the cumulative number of refinement points corresponding to finer detail levels, and the detail level that first satisfies this condition is determined as the minimum reference detail level (MinInterRefLvl) for inter-level prediction.

[1035] According to the embodiments, the above conditions can be derived as follows.

[1036] As previously mentioned, when lod_scalability_enabled is 1, the variable MinInterRefLvl is the finest detail level that has fewer refinement points than the total number of refinement points associated with all finer detail levels. That is, the variable MinInterRefLvl is the finest detail level among the detail levels for which the number of refinement points is less than the total number of refinement points corresponding to the finer detail levels.

[1037] According to the embodiments, the determination condition of MinInterRefLvl can be expressed by the following formula.

[1038]

[1039] In the above formula, LodRfmtPtCnt[lvl] represents the number of refinement points corresponding to detail level lvl, and LodRfmtPtCnt[k] represents the number of refinement points corresponding to detail level k. And, represents the cumulative number of refinement points corresponding to detail levels finer than the above detail level lvl. That is, the above formula represents a conditional expression for determining whether the number of refinement points of the current detail level is less than the total number of refinement points corresponding to detail levels finer than that. In other words, the above formula represents a comparison condition for determining the minimum reference detail level for inter-level prediction.

[1040] The following describes the relationship between the LoD and the refinement list. Depending on the embodiments, the relationship between the detail level and the refinement list is defined as follows.

[1041] LodRfmtPtCnt[lvl] = LodPtCnt[lvl] - LodPtCnt[lvl +1]

[1042] Here, LodPtCnt[lvl] represents the total number of points included in detail level lvl, LodPtCnt[lvl+1] represents the number of points included in a detail level coarser than detail level lvl, and LodRfmtPtCnt[lvl] represents the number of refinement points newly added to detail level lvl. As defined above, subtracting the number of points already included in a coarser detail level (lvl+1) from the total number of points included in detail level lvl yields the number of points newly included in the current detail level, i.e., the number of refinement points. In other words, the number of refinement points for a specific detail level is calculated as the difference between the total number of points in that detail level and the number of points in a coarser detail level. Thus, the number of refinement points for a specific detail level refers to the number of points that are included in that detail level but not in a coarser detail level.

[1043] The following is an expression that expands the cumulative number of refinement points corresponding to detail levels in terms of the difference in the number of points. That is, the right side of the above conditional expression can be expanded as follows.

[1044]

[1045] represents the sum of the number of refinement points from detail level 0 to detail level lvl-1. That is, it represents the cumulative number of refinement points corresponding to detail levels that are finer than lvl. In other words, by subtracting the number of points included in the coarser detail level k+1 from the number of points included in detail level k ((LodPtCnt[k] - LodPtCnt[k+1])), the number of newly added refinement points in detail level k can be calculated, and by performing this process from detail level 0 to detail level lvl-1 and summing the number of refinement points from detail level 0 to detail level lvl-1, the cumulative number of refinement points from detail level 0 to detail level lvl-1 is calculated.

[1046] Therefore, the conditions for determining the minimum reference detail level (MinInterRefLvl) can be summarized as follows.

[1047] LodRfmtPtCnt[lvl] < slice_num_points_minus1LodPtCnt[lvl]

[1048] That is, if the number of refinement points for a specific detail level is smaller than the total number of refinement points corresponding to finer detail levels, that detail level can be selected as the minimum reference detail level (MinInterRefLvl) for inter-level prediction.

[1049] The following is an explanation of the number of refinement points in FGS.

[1050] According to the embodiments, the present disclosure can determine a minimum reference detail level for an FGS attribute based on the number of refinement points within the FGS. That is, the MinInterRefLvl can be determined as the finest detail level among the detail levels in which the number of refinement points of that detail level within the FGS is less than the total number (i.e., cumulative number) of refinement points corresponding to finer detail levels.

[1051] However, since FGS is used to obtain partial LoDs, the slice_num_points_minus1 in the above conditional expression must be changed to the number of nodes / points generated in FGS. That is, the encoder and / or decoder of the present disclosure performs prediction only when the number of refinement points of a specific level (nodes / points that have not been subsampled to a higher LoD, i.e., a coarser detail level than the current detail level) for partial LoDs generated in FGS is smaller than the number of refinement points of the previous level, and in the opposite case, performs a process of accumulating and merging refinement points. The present disclosure can perform a predictor search from the maximum depth of the skipped LoD by considering cases where LoD skipping occurs.

[1052] That is, the encoder and / or decoder of the present disclosure can perform inter-level prediction for partial LoDs generated in FGS when the number of refinement points corresponding to a specific detail level (lvl) is smaller than the cumulative number of refinement points corresponding to finer detail levels. In addition, considering cases where LoD skipping occurs, a predictor search can be performed based on the detail level corresponding to the largest tree depth among the skipped detail levels.

[1053] An exemplary code algorithm for implementing this is as follows. That is, the minimum reference detail level for inter-level predictor searches in a specific FGS can be determined as follows.

[1054] MinInterRefLvl = LodMinLevel+1

[1055] if (lod_scalability_enabled) {

[1056] for (lvl = LodMinLevel+1; lvl < LodCnt - 1; lvl++) {

[1057] if (LodRfmtPtCnt[lvl] < num_nodes_minus1 LodPtCnt[lvl])

[1058] break

[1059] MinInterRefLvl++

[1060] }

[1061] }

[1062] In the above algorithm, if the value of lod_scalability_enabled is true, that is, if LoD scalability is enabled, the minimum reference detail level (MinInterRefLvl) for inter-level prediction can be calculated in the above FGS.

[1063] Specifically, the minimum reference detail level (MinInterRefLvl) can be initially set to the level immediately following the minimum detail level LodMinLevel, i.e., LodMinLevel + 1. This is to ensure that, in environments where partial decoding or constrained LoD generation is applied, the reference level for performing inter-level prediction starts at a level finer than LodMinLevel.

[1064] Then, if LoD scalability is enabled, the encoder and / or decoder may determine MinInterRefLvl by iteratively evaluating conditions for detail levels in the range from lvl = LodMinLevel + 1 to lvl < LodCnt - 1, where LodCnt represents the total number of detail levels generated in the corresponding FGS, and the last detail level may be excluded from the range used for prediction and reference.

[1065] In each iteration step, the encoder and / or decoder may calculate the number of refinement points LodRfmtPtCnt[lvl] corresponding to the detail level lvl and compare it with num_nodes_minus1 LodPtCnt[lvl]. Here, LodPtCnt[lvl] represents the number of points (or nodes) included in the detail level lvl in a specific FGS (or partial occupancy tree), and num_nodes_minus1 represents the total number of points (or nodes) considered in the FGS (or partial occupancy tree) minus 1. Thus, num_nodes_minus1 LodPtCnt[lvl] functions as a comparison criterion that reflects the scale of points (or nodes) included in detail levels finer than the corresponding detail level, i.e., the amount of accumulated information on the finer level side.

[1066] If LodRfmtPtCnt[lvl] is smaller than num_nodes_minus1 LodPtCnt[lvl], the encoder and / or decoder may determine that lvl as the minimum reference detail level (MinInterRefLvl) for inter-level prediction and terminate the iteration (break).

[1067] On the other hand, if the above condition is not satisfied (i.e., LodRfmtPtCnt[lvl] is not smaller than the comparison criterion), the encoder and / or decoder may determine that the minimum reference detail level has not yet been determined, increase the MinInterRefLvl value by 1, and then repeat the same comparison for the next detail level.

[1068] In other words, it is determined whether the number of refinement points at the current detail level in a specific FGS is less than the cumulative number of refinement points corresponding to finer detail levels, and the detail level that first satisfies this condition is determined as the minimum reference detail level (MinInterRefLvl) for inter-level prediction in that FGS.

[1069] Consequently, according to the present disclosure, even in an FGS environment where LoD scalability is enabled, the minimum detail level to be used as a reference for inter-level prediction can be reliably determined by reflecting the characteristics of the refinement point distribution, thereby improving the consistency of the prediction reference and the prediction efficiency.

[1070] Meanwhile, if MinInterRefLvl is greater than 1, the child subgroups of a parent node cannot be decoded because there is insufficient information about the parent nodes. Therefore, in the present disclosure, one embodiment is provided in that if MinInterRefLvl is greater than 1 for a parent subgroup, the child subgroups of that parent subgroup are not decoded. Additionally, in the present disclosure, for stricter constraints, one embodiment is provided in that the minimum reference detail level is determined for the subgroups of the last layer group of the current slice. That is, the minimum reference detail level can be determined for the subgroups belonging to the last layer group of the current slice.

[1071] However, FGS is characterized by performing progressive decoding based on the vertical continuity of depth / level. In this case, if skip LoD occurs for a specific subgroup, decoding of the child subgroup may not proceed.

[1072] Alternatively, skip LoD can be allowed only for the lowest layer group, as in the following code algorithm.

[1073] MinInterRefLvl = 1

[1074] if (lod_scalability_enabled) {

[1075] for (lvl = 1; lvl < num_layers_minus1[num_layer_groups_minus1]; lvl++) {

[1076] if (LodRfmtPtCnt[lvl] < num_nodes_minus1 LodPtCnt[lvl])

[1077] break

[1078] MinInterRefLvl++

[1079] }

[1080] }

[1081] In other words, the above code algorithm can be applied when LoD skipping is limited to the lowest layer group, and can be used to determine the minimum reference detail level (MinInterRefLvl) for inter-level prediction for subgroups belonging to the lowest layer group.

[1082] Specifically, the minimum reference detail level (MinInterRefLvl) can be set to 1 as an initial value. Subsequently, when LoD scalability is enabled, the encoder and / or decoder repeatedly evaluates conditions within the range lvl < num_layers_minus1[num_layer_groups_minus1], starting with the detail level index lvl from 1. Here, num_layer_groups_minus1 is information regarding the index of the current layer group, and num_layers_minus1[num_layer_groups_minus1] is a value indicating the range of detail levels considered in that layer group, and the iteration can be performed for detail levels within that range.

[1083] In each iteration step, the encoder and / or decoder checks LodRfmtPtCnt[lvl], the number of refinement points (or refinement nodes) corresponding to the detail level lvl, and compares it with num_nodes_minus1 LodPtCnt[lvl], where LodPtCnt[lvl] represents the number of points (or nodes) included in the detail level lvl, and num_nodes_minus1 represents the total number of points (or nodes) considered in the FGS (or partial occupancy tree) minus 1.

[1084] If LodRfmtPtCnt[lvl] is smaller than num_nodes_minus1 LodPtCnt[lvl], the encoder and / or decoder may determine that detail level lvl as the minimum reference detail level (MinInterRefLvl) for inter-level prediction and terminate the iteration (break). Conversely, if the above condition is not satisfied, the MinInterRefLvl value may be increased by 1 and the same comparison may be repeated for the next detail level.

[1085] Consequently, according to the present disclosure, a minimum reference detail level to be used for inter-level prediction can be determined by reflecting the distribution of refinement points by detail level in an environment where LoD scalability is enabled, thereby improving the stability of the prediction reference level selection.

[1086] In other words, when LoD scalability (lod_scalability_enabled) is enabled, the decoder can support partial density decoding by performing selective decoding at the layer group and / or detail level level. In this case, if a skip occurs at the detail level (LoD), sufficient parent node information for a specific subgroup may not be obtained; consequently, due to a lack of information regarding the parent subgroup, decoding of the child subgroup may be limited or decoding stability may be degraded.

[1087] Accordingly, in the present disclosure, instead of allowing LoD skips for all layer groups, LoD skips may be allowed, for example, only for the lowest layer group. That is, by restricting or prohibiting LoD skips for layer groups other than the lowest layer group, sufficient parent node information necessary for child subgroup decoding can be ensured.

[1088] In addition, even when LoD skipping is allowed only for the lowest layer group as described above, the selection of the reference detail level for inter-level prediction can affect prediction performance. Therefore, the decoder can determine the minimum reference detail level (MinInterRefLvl) by comparing the number of refinement points per detail level (LodRfmtPtCnt[lvl]) and the reference value (num_nodes_minus1 LodPtCnt[lvl]) for subgroups belonging to the lowest layer group. For example, the detail level that first satisfies the condition LodRfmtPtCnt[lvl] < num_nodes_minus1 LodPtCnt[lvl] is determined as MinInterRefLvl, and reference point search and prediction for inter-level prediction can be performed based on the determined MinInterRefLvl.

[1089] The following is an explanation of the distance calculation.

[1090] According to the embodiments, in the process of obtaining the biased L1 norm, it is necessary to perform quantization according to the detail level.

[1091] In the present disclosure, distance calculation using L1 norm (Norm1) can be performed as follows.

[1092] That is, in the process of inter-level prediction or prediction reference search, the distance between points can be calculated to determine the spatial similarity between a target point for prediction and a candidate point for reference. In the present disclosure, the distance between points can be calculated using a biased L1 norm, which can be defined as the Manhattan distance between two points with axis-specific weights applied.

[1093] In particular, when layer groups are enabled in FGS-based attribute decoding, point coordinates may have different resolutions depending on the Level of Detail (LoD); therefore, coordinates can be quantized in a manner corresponding to the corresponding level of detail before distance calculation. This allows for alignment of coordinate precision when calculating distances between points belonging to different levels of detail.

[1094] Specifically, the distance between two points A and B can be calculated by the L1 norm-based distance function Norm1 as in the following code algorithm.

[1095] Norm1(ptIdxA, ptIdxB, pcA, pcB) := dist[0] + dist[1] + dist[2]

[1096] where

[1097] dist[k] := Abs(posA[ptIdxA][k] posB[ptIdxB][k]) × PredBias[k]

[1098] posA[ptIdx][k] := pcA ?

[1099] RefAttrPos[ptIdxA][k]

[1100] : (fgs_layer_group_enabled

[1101] ? (AttrPos[ptIdxA][k] >> Lvl) << Lvl

[1102] : AttrPos[ptIdxA][k])

[1103] posB[ptIdx][k] := pcB ?

[1104] RefAttrPos[ptIdxB][k]

[1105] : (fgs_layer_group_enabled

[1106] ? (AttrPos[ptIdxB][k] >> Lvl) << Lvl

[1107] : AttrPos[ptIdxB][k]

[1108]

[1109] In the code algorithm above, Norm1(ptIdxA, ptIdxB, pcA, pcB) is a function representing the weighted Manhattan distance between two points, k represents a coordinate axis (e.g., X, Y, Z, or each component of the STV coordinate system), and PredBias[k] represents the distance calculation weight for that axis.

[1110] When the above parameter pcA is equal to 0, the above parameter ptIdxA represents the AttrPos index, and when the above parameter pcA is equal to 1, the above parameter ptIdxA represents the RefAttrPos index. That is, when point A is a reference point (pcA = 1), posA can be determined using the reference attribute coordinate RefAttrPos, and when point A is not a reference point (pcA = 0), posA can be determined using the general attribute coordinate AttrPos.

[1111] When the above parameter pcB is equal to 0, the above parameter ptIdxB represents the AttrPos index, and when the above parameter pcB is equal to 1, the above parameter ptIdxB represents the RefAttrPos index. That is, when point B is a reference point (pcB = 1), posB can be determined using the reference attribute coordinate RefAttrPos, and when point B is not a reference point (pcB = 0), posB can be determined using the general attribute coordinate AttrPos.

[1112] The result value of this function is specified by the expression Norm1.

[1113] The expressions posA[ptIdx][k] and posB[ptIdx][k] represent the attribute coordinates used for distance calculation.

[1114] That is, using the posA and posB coordinates determined as above, the distance component dist[k] for each axis is calculated, and by summing them, the final distance value Norm1 between two points can be calculated. The distance calculated in this way can be used as a criterion for determining the proximity between points during the inter-detail level search or predictor set construction process.

[1115] In this case, when fgs_layer_group_enabled is 1, the coordinate values ​​are quantized according to the detail level. That is, when the layer group function is enabled in FGS (fgs_layer_group_enabled = 1), the attribute coordinates of a point can be quantized according to the corresponding detail level Lvl before distance calculation. In the above code algorithm, the >> Lvl operation is an operation that reduces the coordinate values ​​to a precision corresponding to the detail level, and the << Lvl operation represents an operation that aligns the coordinate values ​​according to the reduced precision.

[1116] As such, in the present disclosure, by quantizing and aligning point coordinates according to the detail level and performing L1 norm-based distance calculation with axis-specific weights, stable distance calculation can be performed even between points belonging to different detail levels, thereby enabling more accurate reference point search in the inter-level prediction process.

[1117] The following is an explanation of transform quantization weights and LoD scalability.

[1118] According to the embodiments, when applying LoD scalability to FGS, it is necessary to apply transform coefficient weights—that is, weights applied to attribute transform coefficients—to LoD scalability. This is because, generally, FGS assumes that all LoD levels (i.e., detail levels) within a subgroup are decoded, whereas LoD scalability assumes that geometry depths do not exist due to skipping. In other words, when LoD scalability is applied, nodes corresponding to specific tree depths may be skipped during the geometry decoding process, and as a result, a situation may occur where geometry information corresponding to some LoD levels does not exist. In such cases, since the reference geometry structure for calculating transform coefficient weights may be inconsistent between the encoder and the decoder, it is necessary to perform geometry and attribute processing in a manner suitable for the LoD scalability environment.

[1119] To this end, the following methods can be considered.

[1120] As an example, a method of not skipping depth for FGS geometry can be considered. That is, a method of processing so that there is no skipped depth for FGS geometry can be considered.

[1121] According to the embodiments, when the partial occupancy tree of the FGS geometry is fully decoded, the transformation coefficient weights can be obtained as in the encoder. However, for partial density decoding, an additional process of quantizing the geometry nodes may be required for the LoD levels skipped in the FGS attributes. That is, if some LoD levels are skipped during the FGS attribute decoding process to support partial density decoding, an additional coordinate quantization process may be performed for the geometry nodes corresponding to the skipped LoD levels.

[1122] The following shows the process of generating partial slice geometry and partial slice attributes for general LoD scalability, and in this disclosure, PointPos can be applied by replacing it with SubgroupNodePos[layerGroupIdx][subgroupIdx].

[1123] Partial slice geometry

[1124] In an LoD scalability environment, the decoder may generate partial slice geometry with lower precision instead of full slice geometry. According to embodiments, the decoder may generate lower precision slice geometry at the point when geometry decoding is complete. That is, lower precision partial slice geometry may be generated as a post-processing step after geometry decoding is complete, which has substantially the same effect as stopping the decoding of the occupancy tree at a specific tree level and quantizing point locations to match that precision. Although such lower precision slice geometry is defined as a post-processing step, it is substantially the same as stopping the decoding of the occupancy tree at the beginning of the tree level where NodeSizeLog2 becomes equal to MinNodeSizeLog2, quantizing the locations of points encoded by direct nodes, and removing points existing at the same location.

[1125] In this case, low-precision slice geometry can be defined using the following variables.

[1126] The MinNodeSizeLog2 variable represents the application-specific minimum occupancy tree node size that specifies the precision of output points. In other words, MinNodeSizeLog2 is a variable representing the minimum node size that determines the precision of output points, indicating the minimum node size allowed in the occupancy tree.

[1127] The PartialPtIdx array maps the points included in the low-precision slice geometry to the point indices of the PointPos array. Here, PartialPtIdx[idx] represents the point index in the PointPos array. In other words, the PartialPtIdx array is a mapping array that indicates which index the points included in the low-precision slice geometry have in the original point array (PointPos).

[1128] The PartialPtCnt variable represents the cumulative number of points included in the low-precision slice geometry. In other words, the PartialPtCnt variable represents the total number of points included in the low-precision slice geometry.

[1129] The following is an explanation of the selection of partial point positions.

[1130] According to embodiments, the following process is performed to select points constituting low-precision slice geometry. That is, to generate low-precision slice geometry, the entire slice geometry may first be divided into spatial blocks of a certain size. According to embodiments, the entire slice geometry is spatially divided into a cubic block lattice with a side length of 2MinNodeSizeLog2. That is, the entire point space may be spatially divided into a cubic block lattice with a side length of 2MinNodeSizeLog2.

[1131] Subsequently, a point is selected from each occupied block, and the PointPos index of the selected point is recorded. That is, if one or more points exist for each block, that block is determined to be an occupied block, and a point within that occupied block can be selected and included in the partial slice geometry. In this process, an array storing the number of points per block may be used to check whether a point exists in each block, and if a point has already been selected in a specific block, additional points within the same block are not selected. The index of the selected point is recorded in the PartialPtIdx array, and the number of selected points can be accumulated and managed in PartialPtCnt.

[1132] The code algorithm below illustrates the point selection process for constructing the aforementioned low-precision slice geometry. More specifically, it illustrates the process of generating lower-precision slice geometry by reducing the number of points in the original high-precision slice geometry to leave only one point per spatial block. In the code algorithm below, the sparse array blkPtCnt is used to identify blocks of the partitioned geometry. A value greater than 0 in the array blkPtCnt[ps][pt][pv] indicates that the quantized point location (ps, pt, pv) is already included in the output partial slice geometry. Additionally, elements that are not set in the blkPtCnt array are assumed to be 0.

[1133] Specifically, in the code algorithm below, PartialPtCnt is first initialized to 0. PartialPtCnt is a variable that stores the accumulated number of points to be included in the low-precision slice geometry.

[1134] According to the embodiments, all points (PointPos array) included in the entire slice geometry are examined one by one. ptIdx represents the index of the point currently being examined, and PointCnt represents the total number of points. That is, for all points, the block position is calculated iteratively and the selection status is determined. At this time, the coordinates (x, y, z) of each point are stored in PointPos[ptIdx][0..2].

[1135] Here, a right bit shift (>>) of MinNodeSizeLog2 is performed to calculate the coordinates of the spatial block to which the point belongs. Ps represents the block index in the X-axis direction, pt represents the block index in the Y-axis direction, and pv represents the block index in the Z-axis direction. This operation has the same effect as dividing the space into cube blocks with a side length of 2MinNodeSizeLog2. The blkPtCnt array stores the number of points contained in each block.

[1136] According to embodiments, the number of points in the block (ps, pt, pv) to which the current point belongs is increased by 1, and if the number of points in the block is greater than 1, that is, if there is already a previously selected point in the same block, the current point is not selected and the process moves to the next point. As a result, at most one point is selected in each block.

[1137] If the current point is the first point to appear in the block, store the index ptIdx of that point in the PartialPtIdx array. At the same time, increment the PartialPtCnt value by 1 to update the number of selected points.

[1138] In this way, the above code algorithm divides the entire point space into a 3D block grid of size 2MinNodeSizeLog2, selects only the first point found for each block, and stores the indices of the selected points in the PartialPtIdx array. In addition, the number of selected points is recorded in PartialPtCnt.

[1139] PartialPtCnt = 0

[1140] for (ptIdx = 0; ptIdx < PointCnt; ptIdx++) {

[1141] ps = PointPos[ptIdx][0] >> MinNodeSizeLog2

[1142] pt = PointPos[ptIdx][1] >> MinNodeSizeLog2

[1143] pv = PointPos[ptIdx][2] >> MinNodeSizeLog2

[1144]

[1145] blkPtCnt[ps][pt][pv]++

[1146] if (blkPtCnt[ps][pt][pv] > 1)

[1147] continue

[1148]

[1149] PartialPtIdx[PartialPtCnt++] = ptIdx

[1150] }

[1151] The following is an explanation of the quantization of partial point positions.

[1152] According to embodiments, points selected for low-precision slice geometry can be quantized according to a minimum node size (MinNodeSizeLog2). That is, the positions of points constituting partial slice geometry can be quantized according to the minimum node size. According to embodiments, as in the code algorithm below, the coordinates of each point can be quantized using a bit shift operation so as to be aligned to a precision corresponding to the corresponding minimum node size. This process has the effect of aligning point coordinates to specific block boundaries and serves to reduce geometry precision in a partial decoding environment.

[1153] for (ptIdx = 0; ptIdx < PointCnt; ptIdx++)

[1154] for (k = 0; k < 3; k++)

[1155] PointPos[ptIdx][k] = (PointPos[ptIdx][k] >> MinNodeSizeLog2) << MinNodeSizeLog2

[1156] Additionally, if the MinNodeSizeLog2 value is greater than 1, the points may be at the center of their respective blocks. According to the embodiments, when the MinNodeSizeLog2 value is greater than 1, as in the code algorithm below, the coordinates may be further corrected so that each point is located at the center of its block. This allows the points included in the partial slice geometry to be placed at the center of each block.

[1157] Specifically, according to the code algorithm below, the above process moves the coordinates of all points to the center of the corresponding block based on the block size defined by MinNodeSizeLog2 for each coordinate axis (X, Y, Z). This is a process of adjusting the coordinates so that the points are located at the center of the block instead of at the corners of the block after the coordinate quantization performed in the previous step. In other words, it is a coordinate correction process to construct low-precision slice geometry by moving the quantized point coordinates to the center of the corresponding spatial block.

[1158] for (ptIdx = 0; ptIdx < PointCnt; ptIdx++)

[1159] for (k = 0; k < 3; k++)

[1160] PointPos[ptIdx][k] |= (MinNodeSizeLog2 > 1) << MinNodeSizeLog2 - 1

[1161] The following is an explanation of the points output.

[1162] According to the embodiments, in a partial decoding environment, not all points of the entire slice geometry are output, but only previously selected partial points are output. That is, only the points recorded in the PartialPtIdx array constitute the final set of output points. In this way, partial decoding is equivalent to outputting only selected points. The output points are the points corresponding to PartialPtIdx[i], i∈0 .. PartialPtCnt - 1. That is, the points identified by PartialPtIdx[i] (from i = 0 to PartialPtCnt 1) constitute the output points.

[1163] The following is an explanation of partial attribute decoding.

[1164] When LoD scalability is enabled (lod_scalability_enabled = 1), slice attributes can be restored according to the standard attribute decoding procedure, but with the following differences.

[1165] First, attribute decoding is performed based on the previously generated low-precision slice geometry, rather than the entire slice geometry.

[1166] Second, the finest detail level can be constructed using low-precision slice geometry rather than the entire slice geometry. This is practically equivalent to not reconstructing LoDs corresponding to detail levels smaller than MinNodeSizeLog2. In other words, it is equivalent to not reconstructing detail levels (LoDs) where lvl < MinNodeSizeLog2.

[1167] Specifically, as in the code algorithm below, the finest detail level may consist of points recorded in the PartialPtIdx array, wherein the point indices included in that detail level may be sorted according to the order of Morton-coded attribute coordinates. According to embodiments, the point indices included in the finest detail level may be sorted in ascending order based on the value of each Morton-coded attribute coordinate. That is, the indices of points included in the finest detail level are sorted in ascending order based on the Morton-coded attribute coordinate value corresponding to each point.

[1168] This alignment process can be used to efficiently determine spatial proximity during the subsequent inter-detail-level search and predictor set construction processes.

[1169] for (i = 0; i < PartialPtCnt; i++)

[1170] LodPtIdx[0][i] = PartialPtIdx[i]

[1171] LodPtCnt[0] = PartialPointCnt

[1172] As described above, by applying partial slice geometry and partial attribute decoding structures, the present disclosure can stably maintain the transformation coefficient weight calculation structure between the encoder and the decoder even in an LoD scalability environment, and attribute decoding and prediction processes can be consistently performed even when geometry depth skips occur.

[1173] In addition, the present disclosure may apply a fixed transform coefficient weight for each LoD level (i.e., detail level) to all FGSs, or, as a hybrid approach, apply a fixed transform coefficient weight only to FGSs where LoD scalability is allowed and an adaptive transform coefficient weight to FGSs where LoD scalability is not allowed. In this case, FGSs where LoD scalability is allowed may be determined by prior agreement (e.g., the last layer group including leaf nodes), or their availability may be signaled in the data unit header.

[1174] As described above, the present disclosure describes a method for dividing and transmitting compressed data according to a specific standard for point cloud data. As one application of the present disclosure, when using layered coding, the compressed data can be divided and sent according to the layer, which increases the storage and transmission efficiency of the transmitting end. In particular, when applying scalable attribute coding, there is a disadvantage that the entire coded geometry data must be sent; however, through the proposal of the present disclosure, only the geometry layer that matches the tree level used in scalable attribute coding can be transmitted, thereby further increasing transmission efficiency.

[1175] FIG. 39 illustrates an example of compressing and providing geometry and attributes of point cloud data. In other words, in a point cloud compression (PCC) based service, the compression rate or the number of data can be adjusted and sent depending on the receiver performance or transmission environment. However, when point cloud data is bundled into a single slice unit as shown in FIG. 39, if the receiver performance or transmission environment changes, it is necessary to either 1) convert the bitstream suitable for each environment in advance, store it separately, and select it during transmission, or 2) perform a conversion process (transcoding) prior to transmission. At this time, if the number of receiver environments to be supported increases or the transmission environment changes frequently, storage space issues or delays caused by conversion may become a problem.

[1176] FIG. 40 is a diagram showing another example of compressing and servicing the geometry and attributes of point cloud data according to embodiments.

[1177] As proposed in this disclosure, when compressed data is divided and transmitted according to layers, there is an advantage in that only the necessary parts of the pre-compressed data can be selectively transmitted through a bitstream selector at the bitstream stage without a separate conversion process. This is efficient in terms of storage space as only one storage space is required per stream, and efficient transmission is also possible in terms of bandwidth because only the necessary layers are selectively transmitted by the bitstream selector before transmission.

[1178] To describe the effects according to the features of the present disclosure from the perspective of the receiver, when using layered coding (or hierarchical coding) as one of the applications of the present disclosure, compressed data can be divided and sent according to the layer, and in this case, the efficiency of the receiver increases. In particular, when applying scalable attribute coding, there is a disadvantage that delay occurs and the receiver's computation is burdened as shown in FIG. 41 by receiving and decoding the entire coded geometry data. However, through the proposal of the present disclosure, by decoding only the geometry layer that matches the tree level used in scalable attribute coding, the delay factor is reduced, and the efficiency of the decoder can be increased by saving the computing power required for decoding.

[1179] FIG. 41 is a diagram illustrating the operation of the transmitting and receiving ends when transmitting point cloud data composed of layers. In this case, when transmitting information that allows the entire point cloud data to be restored regardless of the receiver's performance, the receiver requires a process (e.g., data selection or subsampling) to restore the point cloud data through decoding and then select only the point cloud data corresponding to the required layer. In this case, since the transmitted bitstream is already decoded, a delay may occur in a receiver aiming for low latency, or decoding may not be possible depending on the receiver's performance.

[1180] However, as proposed, if only the compressed data of the necessary layer is received depending on the layer, the receiver can selectively decode specific layers, thereby increasing decoder efficiency and having the advantage of supporting decoders of various performance levels.

[1181] FIG. 42 shows a flowchart of a point cloud data encoding method according to embodiments.

[1182] A point cloud data encoding method according to embodiments may include a step of encoding geometry data of point cloud data (S71001) and a step of encoding attribute data of the point cloud data based on input and / or reconstructed geometry data (S71002).

[1183] The step of encoding geometry data and attribute data of point cloud data (S71001, S71002) can perform part or all of the operation of the point cloud video encoder (10002) of FIG. 1, the encoding (20001) of FIG. 2, the point cloud video encoder of FIG. 3, the point cloud video encoder of FIG. 8, and the point cloud encoding of FIG. 30.

[1184] In one embodiment, the step of encoding attribute data (S71002) applies at least one LoD generation method among an octree-based LoD generation method, a distance-based LoD generation method, and a sampling-based LoD generation method to LoD l After creating the set, the above LoD l Based on the set, X (>0) nearest neighbor (NN) points can be found in the group with the same or smaller LoD (i.e., large distance between nodes) and registered as a set of neighbor points in the predictor.

[1185] According to the embodiments, the step of encoding attribute data (S71002) creates a downsampled geometry grid by performing sampling on the remaining layers excluding one or more skipped layers in the case of scalable transmission as described above, and corrects the positions of occupied nodes including neighbor search target nodes by performing upsampling using information related to the skipped layers. Then, an NN search is performed based on the corrected positions to obtain the distance between nodes. Then, attribute prediction and compression are performed based on the NN nodes obtained in this way. Since the position correction and NN search according to the embodiments have been described in detail in FIGS. 26 to 30 and FIG. 31, they are omitted here to avoid redundant explanation.

[1186] According to the embodiments, the step of encoding attribute data (S71002) determines the minimum reference detail level (MinInterRefLvl) by comparing the number of refinement points corresponding to each LoD and the cumulative number of refinement points of finer LoDs in order to perform inter-level prediction as described above. More specifically, the step of encoding attribute data (S71002) compares the number of refinement points corresponding to each detail level with the cumulative number of refinement points corresponding to detail levels that are finer than the corresponding detail level when lod_scalability_enabled is set to 1. Then, if the number of refinement points of the corresponding detail level is less than the cumulative number of refinement points of the finer detail levels, the detail level is determined as the minimum reference detail level for inter-level prediction. Additionally, the step of encoding attribute data (S71002) may proceed with reference point search and prediction for inter-level prediction based on the determined minimum reference detail level. As an example, the minimum reference detail level is similarly applied and determined within the FGS. That is, in a specific FGS, it is determined whether the number of refinement points of the current detail level is less than the cumulative number of refinement points corresponding to finer detail levels, and the detail level that first satisfies this condition is determined as the minimum reference detail level (MinInterRefLvl) for inter-level prediction in that FGS. Since the method for determining the minimum reference detail level has been sufficiently explained above, a detailed explanation is omitted here.

[1187] The encoding method described so far is performed by an encoding device (or referred to as an encoder). The encoding device includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: encode geometry data of point cloud data; and encode attribute data of point cloud data. The at least one processor may be configured to perform the operations of the method described above.

[1188] The embodiments further include a computer-readable storage medium that stores a bitstream generated by the method according to FIG. 42.

[1189] The embodiments further include a method comprising the steps of: acquiring a bitstream for point cloud data; generating the bitstream based on the steps of encoding geometry data of the point cloud data and encoding attribute data of the point cloud data; and transmitting data including the bitstream.

[1190] The bitstream according to the embodiments may further include signaling information.

[1191] The step of transmitting data including a bitstream according to the embodiments may be performed in the transmitter (10003) of FIG. 1, the transmission step (20002) of FIG. 2, the transmission processing unit (8012) of FIG. 8, or the transmission processing unit (51005) of FIG. 30.

[1192] FIG. 43 shows a flowchart of a point cloud data decoding method according to embodiments.

[1193] A point cloud data decoding method according to embodiments may include a step of decoding geometry data of point cloud data within a bitstream (S81001) and a step of decoding attribute data of point cloud data within a bitstream (S81002). The bitstream may further include signaling information.

[1194] The point cloud data decoding method according to the embodiments may further include the step of rendering the restored point cloud data based on the decoded geometry data and the decoded attribute data.

[1195] The step of decoding geometry data and attribute data according to the embodiments (S81001, S81002) can perform decoding in units of a slice based on a layer or layer group or a tile including one or more slices.

[1196] The step of decoding geometry data according to the embodiments (S81001) may perform part or all of the operation of the point cloud video decoder (10006) of FIG. 1, the decoding (20003) of FIG. 2, the point cloud video decoder of FIG. 7, the point cloud video decoder of FIG. 9, and the geometry decoder of FIG. 32.

[1197] The step (S81002) of decoding attribute data according to the embodiments may perform part or all of the operation of the point cloud video decoder (10006) of FIG. 1, the decoding (20003) of FIG. 2, the point cloud video decoder of FIG. 7, the point cloud video decoder of FIG. 9, and the attribute decoder of FIG. 32.

[1198] According to embodiments, at least one of the signaling information, for example, a sequence parameter set, an attribute parameter set, a tile parameter set, and an attribute slice header, may include skip layer-related information. According to embodiments, the skip layer-related information may include bounding box location information and size information. For example, the skip layer-related information may include root_bbox_origin and root_bbox_size information and / or max_geometry_level_log2 information. The skip layer-related information is information for making the resolution details of the NN search performed by the attribute encoder of the transmitting device and the resolution details of the NN search performed by the attribute decoder of the receiving device identical.

[1199] According to the embodiments, the attribute data decoding step (S81002) may generate LoDs by performing subsampling based on the tree structure (e.g., octree) of the geometry reconstructed in the geometry data decoding step (S8001). At this time, the attribute data decoding step (S81002) corrects the positions of nodes in the geometry grid by performing level-based geometry grid adaptation as described above. At this time, the geometry grid adaptation corrects the positions of n...

Claims

1. A step of decoding geometry data of point cloud data within a bitstream; and A step of decoding attribute data of the above point cloud data; comprising Decoding method.

2. In Paragraph 1, The geometry data and the attribute data are each decoded in units of fine segmentation slices (FGS), and Each of the above FGS is mapped to each spatially partitioned subgroup within a layer group defined as a group of consecutive tree levels of an occupancy tree, and Each of the above FGS is a decoding method identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying a subgroup within the layer group.

3. In Paragraph 2, The above attribute decoding step is, It includes the step of generating multiple Levels of Detail (LoDs) within a specific FGS, and The above LoD generation step is, When LoD scalability is enabled for the above FGS, the method includes the step of setting one of the plurality of LoDs as a reference LoD for inter-level prediction. A decoding method in which the above reference LoD setting step determines the above reference LoD as the minimum reference LoD for inter-level prediction when the number of refinement points of the above reference LoD is smaller than the cumulative number of refinement points belonging to LoDs that are finer than the above reference LoD.

4. In paragraph 3, the attribute decoding step is, A decoding method further comprising the step of performing a prediction reference search by setting a reference point search range for inter-level prediction based on the LoD with the maximum tree depth among the one or more skipped LoDs when one or more LoDs are skipped within the layer group.

5. In Paragraph 3, A decoding method that does not perform decoding on child subgroups of the subgroup determined by the minimum reference LoD when the minimum reference LoD is greater than a preset threshold.

6. In Paragraph 4, The above attribute decoding step is, The method further includes the step of calculating the distance between a predicted target point selected by the above-mentioned prediction reference search and a reference candidate point, The above distance calculation step is, A decoding method for calculating the distance between the predicted target point and the reference candidate point by aligning the attribute coordinates of the predicted target point and the reference candidate point based on the reference LoD, and then using an L1 norm that applies axis-specific weights to the difference between the aligned attribute coordinates.

7. Memory; and At least one processor connected to the memory; comprising, The above at least one processor is: Decoding geometry data of point cloud data within a bitstream; and Decoding attribute data of the above point cloud data; configured to do so, Decoding device.

8. In Paragraph 7, The geometry data and the attribute data are each decoded in units of fine segmentation slices (FGS), and Each of the above FGS is mapped to each spatially partitioned subgroup within a layer group defined as a group of consecutive tree levels of an occupancy tree, and Each of the above FGS is a decoding device identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying a subgroup within the layer group.

9. In paragraph 8, the above at least one processor, It includes a LoD generation unit that generates multiple Level of Detail (LoD) units within a specific FGS, and When LoD scalability is enabled for the FGS, the above LoD generation unit sets one of the plurality of LoDs as a reference LoD for inter-level prediction, and The above LoD generation unit is a decoding device that determines the reference LoD as the minimum reference LoD for inter-level prediction when the number of refinement points of the reference LoD is smaller than the cumulative number of refinement points belonging to LoDs that are finer than the reference LoD.

10. A step of encoding the geometry data of the point cloud data; and A step of encoding attribute data of the above point cloud data; comprising Encoding method.

11. In Paragraph 10, The geometry data and the attribute data are each encoded in units of fine segmentation slices (FGS), and Each of the above FGS is mapped to each spatially partitioned subgroup within a layer group defined as a group of consecutive tree levels of an occupancy tree, and Each of the above FGS is an encoding method identified by a combination of a layer group index for identifying the layer group and a subgroup index for identifying the subgroup within the layer group.

12. In Paragraph 11, The above attribute encoding step is, It includes the step of generating multiple Levels of Detail (LoDs) within a specific FGS, and The above LoD generation step is, When LoD scalability is enabled for the above FGS, the method includes the step of setting one of the plurality of LoDs as a reference LoD for inter-level prediction. The above reference LoD setting step is an encoding method for determining the reference LoD as the minimum reference LoD for inter-level prediction when the number of refinement points of the reference LoD is smaller than the cumulative number of refinement points belonging to LoDs that are finer than the reference LoD.

13. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Encoding the geometry data of the point cloud data; and Configured to encode the attribute data of the above point cloud data; Encoding device.

14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 10.

15. Step for acquiring a bitstream for point cloud data, The bitstream is generated based on the step of encoding geometry data of the point cloud data; and the step of encoding attribute data of the point cloud data; and A method comprising the step of transmitting data including the bitstream above.