Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

By encoding and decoding point cloud data and using the octree structure for geometric and attribute encoding, the problem of low point cloud data processing efficiency in the existing technology is solved, and efficient and high-quality point cloud services are achieved.

CN119999205APending Publication Date: 2025-05-13LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073204.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-18
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process large amounts of point cloud data, resulting in long wait times and high encoding/decoding complexity.

Method used

By encoding the point cloud data and sending it into a bitstream, the receiver then decoding it. The specific method includes using the octree structure for geometric encoding and attribute encoding to achieve effective point cloud data transmission.

Benefits of technology

It realizes efficient processing of point cloud data and provides high-quality point cloud services, suitable for providing general services such as VR and self-driving services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999205A_ABST
    Figure CN119999205A_ABST
Patent Text Reader

Abstract

A point cloud data transmission method according to an embodiment may comprise the steps of: encoding point cloud data; and transmitting the bitstream including the point cloud data. A point cloud data receiving method according to an embodiment may comprise the steps of: receiving a bitstream including point cloud data; and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments relate to a method and apparatus for processing point cloud content. Background Technology

[0002] Point cloud content is content represented by point clouds, which are collections of points belonging to a coordinate system representing three-dimensional space (or volume). Point cloud content can represent media configured in three dimensions and is used to provide various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), XR (Extended Reality), and autonomous driving services. However, tens of thousands to hundreds of thousands of point data points are required to represent point cloud content. Therefore, an efficient method for processing large amounts of point data is needed. Summary of the Invention

[0003] Technical issues

[0004] The embodiments provide an apparatus and method for efficiently processing point cloud data. The embodiments also provide a point cloud data processing method and apparatus for addressing latency and encoding / decoding complexity.

[0005] The technical scope of the embodiments is not limited to the above-described technical objectives, and can be extended to other technical objectives that can be inferred by those skilled in the art based on the entire content disclosed herein.

[0006] Technical solution

[0007] In one aspect of this disclosure, a method for transmitting point cloud data may include: encoding the point cloud data; and transmitting a bit stream containing the point cloud data. In another aspect of this disclosure, a method for receiving point cloud data may include: receiving a bit stream containing the point cloud data; and decoding the point cloud data.

[0008] Beneficial effects

[0009] The apparatus and method according to the embodiments can effectively process point cloud data.

[0010] The apparatus and method according to the embodiments can provide high-quality point cloud services.

[0011] The apparatus and method according to the embodiments can provide point cloud content for providing general services such as VR services and autonomous driving services. Attached Figure Description

[0012] The accompanying drawings, included to provide a further understanding of this disclosure and incorporated into and constituting a part of this application, illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar parts.

[0013] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown;

[0014] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an embodiment;

[0015] Figure 3 An exemplary point cloud encoder according to an embodiment is shown;

[0016] Figure 4 An example of an octree and occupancy code according to an embodiment is shown;

[0017] Figure 5 An example of point configuration in each LOD according to an embodiment is shown;

[0018] Figure 6 An example of point configuration in each LOD according to an embodiment is shown;

[0019] Figure 7 A point cloud decoder according to an embodiment is shown;

[0020] Figure 8 A transmitting apparatus according to an embodiment is shown;

[0021] Figure 9 A receiving device according to an embodiment is shown;

[0022] Figure 10 An exemplary structure that can be operated in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment is shown;

[0023] Figure 11 The process of encoding, transmitting, and decoding point cloud data according to an embodiment is illustrated;

[0024] Figure 12 The layer-based point cloud data configuration according to an embodiment and the geometry and attribute bitstream structure according to an embodiment are illustrated.

[0025] Figure 13 The bitstream configuration according to an embodiment is shown;

[0026] Figure 14 A bitstream sorting method according to an embodiment is shown;

[0027] Figure 15 A method for selecting geometric data and attribute data according to an embodiment is shown;

[0028] Figure 16 A method for configuring slices containing point cloud data according to an embodiment is shown;

[0029] Figure 17 The geometric compilation layer structure according to an embodiment is shown;

[0030] Figure 18 The layer group and subgroup structure according to an embodiment is shown;

[0031] Figure 19 Multi-resolution, multi-size ROIs are shown according to embodiments;

[0032] Figure 20 A layer assembly sheet according to an embodiment is shown;

[0033] Figure 21 The nearest neighbor search process according to an embodiment is illustrated;

[0034] Figure 22 The nearest neighbor search process according to an embodiment is illustrated;

[0035] Figure 23 The nearest neighbor search process according to an embodiment is illustrated;

[0036] Figure 24 An attribute encoding method according to an embodiment is shown;

[0037] Figure 25 An attribute decoding method according to an embodiment is shown;

[0038] Figure 26 A bitstream containing parameters and encoded point cloud data according to an embodiment is shown;

[0039] Figure 27 A set of sequence parameters according to an embodiment is shown;

[0040] Figure 28 The attribute parameter set and attribute data unit header according to an embodiment are shown;

[0041] Figure 29 The diagram illustrates the dependency attribute data unit header according to an embodiment;

[0042] Figure 30 The process of transmitting partial point cloud data according to an embodiment is illustrated;

[0043] Figure 31 A method for sending and receiving point cloud data according to an embodiment is shown;

[0044] Figure 32 A method for sending and receiving point cloud data according to an embodiment is shown;

[0045] Figure 33 A method for transmitting point cloud data according to an embodiment is shown; and

[0046] Figure 34A method for receiving point cloud data according to an embodiment is shown. Detailed Implementation

[0047] Preferred embodiments of the present disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the accompanying drawings is intended to illustrate exemplary embodiments of the present disclosure and not to show only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0048] Although most of the terms used in this disclosure are selected from commonly used terms in the art, some terms have been arbitrarily chosen by the applicant and their meanings are explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.

[0049] Figure 1 An exemplary point cloud content delivery system according to an embodiment is shown.

[0050] Figure 1 The point cloud content providing system shown may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 are capable of transmitting and receiving point cloud data via wired or wireless communication.

[0051] The point cloud data transmitting device 10000 according to an embodiment can acquire and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmitting device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device, and / or a server. According to an embodiment, the transmitting device 10000 may include devices configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers.

[0052] According to an embodiment, the transmitting device 10000 includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002, and / or a transmitter (or communication module) 10003.

[0053] The point cloud video acquisition unit 10001 according to an embodiment acquires point cloud video through processing procedures such as capture, synthesis, or generation. Point cloud video is point cloud content represented by a point cloud, which is a set of points located in 3D space, and may be referred to as point cloud video data, point cloud data, etc. The point cloud video according to an embodiment may include one or more frames. A frame represents a still image / scene. Therefore, point cloud video may include point cloud images / frames / scenes, and may be referred to as point cloud images, frames, or scenes.

[0054] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 may encode the point cloud video data based on point cloud compression compilation. Point cloud compression compilation according to an embodiment may include geometry-based point cloud compression (G-PCC) compilation and / or video-based point cloud compression (V-PCC) compilation or next-generation compilation. Point cloud compression compilation according to an embodiment is not limited to the above embodiments. The point cloud video encoder 10002 may output a bitstream containing encoded point cloud video data. The bitstream may contain not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0055] According to an embodiment, transmitter 10003 transmits a bitstream containing encoded point cloud video data. The bitstream, according to an embodiment, is encapsulated in a file or segment (e.g., a streaming segment) and transmitted via various networks such as broadcast networks and / or broadband networks. Although not shown in the figures, transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform encapsulation operations. According to an embodiment, the encapsulator may be included in transmitter 10003. According to an embodiment, the file or segment may be transmitted via a network to receiving device 10004 or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). Transmitter 10003 according to an embodiment is capable of wired / wireless communication with receiving device 10004 (or receiver 10005) via networks such as 4G, 5G, and 6G. Additionally, the transmitter may perform necessary data processing operations depending on the network system (e.g., a 4G, 5G, or 6G communication network system). Transmitting device 10000 may transmit encapsulated data on demand.

[0056] According to an embodiment, the receiving device 10004 includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include devices, robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0057] According to an embodiment, receiver 10005 receives a bitstream containing point cloud video data or a file / segment encapsulated with a bitstream from a network or storage medium. Receiver 10005 may perform necessary data processing according to the network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). According to an embodiment, receiver 10005 may decapsulate the received file / segment and output a bitstream. According to an embodiment, receiver 10005 may include a decapsulator (or decapsulator module) configured to perform a decapsulation operation. The decapsulator may be implemented as a separate element (or component) from receiver 10005.

[0058] The point cloud video decoder 10006 decodes the bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to the method in which the point cloud video data is encoded (e.g., the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression compilation (the reverse process of point cloud compression). Point cloud decompression compilation includes G-PCC compilation.

[0059] Renderer 10007 renders decoded point cloud video data. In one embodiment, renderer 10007 can render decoded point cloud video data according to a viewport, etc. Renderer 10007 can render not only point cloud video data but also audio data to output point cloud content. According to an embodiment, renderer 10007 may include a display configured to display point cloud content. According to an embodiment, the display may be implemented as a separate device or component rather than included in renderer 10007.

[0060] The arrows indicated by dashed lines in the diagram represent the transmission paths of the feedback information acquired by the receiving device 10004. The feedback information reflects the interactivity of the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is for a service requiring user interaction (e.g., autonomous driving services, etc.), the feedback information may be provided to the content sender (e.g., the sending device 10000) and / or the service provider. According to embodiments, the feedback information may be used in both the receiving device 10004 and the sending device 10000, or it may not be provided.

[0061] According to an embodiment, head orientation information is information about the position, orientation, angle, and movement of the user's head. The receiving device 10004 according to an embodiment can calculate viewport information based on the head orientation information. The viewport information can be information about the area of ​​the point cloud video that the user is viewing. The viewpoint is the point through which the user views the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of the area can be determined by the field of view (FOV). Therefore, in addition to head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area the user gazes at in the point cloud video, the gaze duration, etc. According to an embodiment, the receiving device 10004 can send feedback information including the gaze analysis results to the transmitting device 10000. The feedback information according to an embodiment can be acquired during rendering and / or display. The feedback information according to an embodiment can be acquired by one or more sensors included in the receiving device 10004. According to an embodiment, feedback information can be obtained by the renderer 10007 or by a separate external element (or device, component, etc.). Figure 1 The dashed lines in the diagram represent the process of sending feedback information obtained by the renderer 10007. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10006 can perform decoding operations based on the feedback information. The receiving device 10004 can send feedback information to the sending device 10000. The sending device 10000 (or the point cloud video data encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can effectively process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide the point cloud content to the user.

[0062] According to the embodiment, the transmitting device 10000 can be called an encoder, transmitting device, transmitter, etc., and the receiving device 10004 can be called a decoder, receiving device, receiver, etc.

[0063] According to the embodiments Figure 1 Point cloud data processed in a point cloud content provision system (through a series of processes including acquisition, encoding, transmission, decoding, and rendering) can be referred to as point cloud content data or point cloud video data. According to embodiments, point cloud content data can be used as a concept encompassing metadata or signaling information related to point cloud data.

[0064] Figure 1 The components of the point cloud content provided by the system can be implemented by hardware, software, processors, and / or combinations thereof.

[0065] Figure 2This is a block diagram illustrating the point cloud content provisioning operation according to an embodiment.

[0066] Figure 2 The block diagram shows Figure 1 The operation of the point cloud content providing system described herein. As mentioned above, the point cloud content providing system can process point cloud data based on point cloud compression compilation (e.g., G-PCC).

[0067] A point cloud content providing system (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) according to an embodiment can acquire point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system used to represent 3D space. The point cloud video according to an embodiment may include Ply (Polygon file format or Stanford Triangle format) files. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. A Ply file contains point cloud data such as point geometry and / or attributes. Geometry includes the position of the points. The position of each point may be represented by parameters (e.g., values ​​of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). Attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc., of each point). A point has one or more attributes. For example, a point may have a color attribute or two attributes: color and reflectivity. According to embodiments, geometry can be referred to as location, geometric information, geometric data, location information, location data, etc., and attributes can be referred to as attributes, attribute information, attribute data, etc. A point cloud content providing system (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) can obtain point cloud data from information related to the point cloud video acquisition process (e.g., depth information, color information, etc.).

[0068] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression compilation. As described above, point cloud data can include geometric information and attribute information about points. Therefore, the point cloud content providing system can perform geometric encoding to encode the geometry and output a geometric bitstream. The point cloud content providing system can perform attribute encoding to encode the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometric encoding. The geometric bitstream and attribute bitstream according to an embodiment can be multiplexed and output as a single bitstream. The bitstream according to an embodiment may also contain signaling information related to geometric encoding and attribute encoding.

[0069] A point cloud content providing system (e.g., transmitting device 10000 or transmitter 10003) according to an embodiment can transmit encoded point cloud data (20002). Figure 1 As shown, encoded point cloud data can be represented by geometric bitstreams and attribute bitstreams. Additionally, the encoded point cloud data can be transmitted as a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric encoding and attribute encoding). The point cloud content providing system can encapsulate the bitstream carrying the encoded point cloud data and transmit it as a file or fragment.

[0070] The point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to an embodiment can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bitstream.

[0071] A point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode encoded point cloud data (e.g., geometric bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode point cloud video data based on signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode the geometric bitstream to reconstruct the location (geometry) of the points. The point cloud content providing system can reconstruct the attributes of the points by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can reconstruct point cloud video based on location according to the reconstructed geometry and the decoded attributes.

[0072] A point cloud content providing system (e.g., receiving device 10004 or renderer 10007) according to an embodiment can render decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometry and attributes decoded through the decoding process. Points in the point cloud content can be rendered as vertices with a specific thickness, cubes with a specific minimum size centered at the corresponding vertex position, or circles centered at the corresponding vertex position. All or part of the rendered point cloud content is provided to a user through a display (e.g., a VR / AR display, a general display, etc.).

[0073] The point cloud content providing system (e.g., receiving device 10004) according to an embodiment can obtain feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. The feedback information and operation of the point cloud content providing system according to an embodiment are related to reference. Figure 1 The feedback information and operation described are the same, so their detailed description is omitted.

[0074] Figure 3 An exemplary point cloud encoder according to an embodiment is shown.

[0075] Figure 3 Show Figure 1 An example of a point cloud video encoder 10002. The point cloud encoder reconstructs and encodes point cloud data (e.g., point locations and / or attributes) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) based on network conditions or applications. When the total size of the point cloud content is large (e.g., providing 60 Gbps of point cloud content for 30 fps), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide the point cloud content according to network conditions, etc.

[0076] For reference Figure 1 and Figure 2 As described, the point cloud encoder can perform geometric encoding and attribute encoding. Geometric encoding is performed before attribute encoding.

[0077] The point cloud encoder according to the embodiment includes a coordinate transformer (transform coordinates) 30000, a quantizer (quantize and remove points (voxarization)) 30001, an octree analyzer (analyze octrees) 30002, a surface approximation analyzer (analyze surface approximations) 30003, an arithmetic encoder (arithmetic encoding) 30004, a geometry reconstructor (reconstruct geometry) 30005, a color transformer (transform colors) 30006, an attribute transformer (transform attributes) 30007, a RAHT transformer (RAHT) 30008, an LOD generator (generate LODs) 30009, a lift transformer (lift) 30010, a coefficient quantizer (quantize coefficients) 30011, and / or an arithmetic encoder (arithmetic encoding) 30012.

[0078] The coordinate transformer 30000, quantizer 30001, octree analyzer 30002, surface approximation analyzer 30003, arithmetic encoder 30004, and geometric reconstructor 30005 are capable of performing geometric coding. Geometric coding according to embodiments may include octree geometric compilation, predictive tree geometric compilation, direct compilation, triplet geometric coding, and entropy coding. Direct compilation and triplet geometric coding are applied selectively or in combination. Geometric coding is not limited to the examples described above.

[0079] As shown in the figure, the coordinate transformer 30000 according to an embodiment receives a position and transforms it into coordinates. For example, the position can be transformed into position information in three-dimensional space (e.g., three-dimensional space represented by the XYZ coordinate system). The position information in three-dimensional space according to an embodiment can be referred to as geometric information.

[0080] The quantizer 30001 according to the embodiment performs geometric quantization. For example, the quantizer 30001 may quantize points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 30001 performs a quantization operation: multiplying the difference between the minimum position value and the position value of each point by a preset quantization scaling value, and then finding the nearest integer value by rounding the value obtained by the multiplication. Thus, one or more points may have the same quantized position (or position value). The quantizer 30001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of pixels (the smallest unit containing 2D image / video information), the points of the point cloud content (or 3D point cloud video) according to the embodiment may be included in one or more voxels. As a combination of volume and pixel, the term voxel refers to the 3D cubic space generated when 3D space is divided into units (unit = 1.0) based on axes representing 3D space (e.g., X-axis, Y-axis, and Z-axis). The quantizer 30001 allows a group of points in 3D space to be matched with voxels. According to one embodiment, a voxel may include only one point. According to another embodiment, a voxel may include one or more points. To represent a voxel as a point, the location of the voxel's center can be set based on the locations of one or more points included in the voxel. In this case, attributes of all locations included in a voxel can be combined and assigned to the voxel.

[0081] According to the embodiment, the octree analyzer 30002 performs octree geometry compilation (or octree compilation) to present voxels in an octree structure. The octree structure represents points based on the matching of octree structures with voxels.

[0082] The surface approximation analyzer 30003 according to the embodiment can analyze and approximate an octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing multiple points to efficiently provide an octree and voxelization.

[0083] According to an embodiment, the arithmetic encoder 30004 performs entropy encoding on octrees and / or approximate octrees. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric bitstream is generated.

[0084] The attribute encoding is performed by a color transformer 30006, an attribute transformer 30007, a RAHT transformer 30008, a LOD generator 30009, a boosting transformer 30010, a coefficient quantizer 30011, and / or an arithmetic encoder 30012. As described above, a point may have one or more attributes. The attribute encoding according to the embodiment is also applied to the attributes that a point has. However, when an attribute (e.g., color) includes one or more elements, attribute encoding is applied independently to each element. The attribute encoding according to the embodiment includes color transformation compilation, attribute transformation compilation, region adaptive hierarchical transformation (RAHT) compilation, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) compilation, and interpolation-based hierarchical nearest neighbor prediction (boosting transformation) compilation with update / boosting steps. Depending on the point cloud content, the above-described RAHT compilation, prediction transformation compilation, and boosting transformation compilation may be used selectively, or a combination of one or more compilation schemes may be used. The attribute encoding according to the embodiment is not limited to the examples described above.

[0085] The color converter 30006 according to the embodiment performs color transformation compilation that transforms the color values ​​(or textures) included in the attributes. For example, the color converter 30006 can transform the format of color information (e.g., from RGB to YCbCr). Optionally, the operation of the color converter 30006 according to the embodiment can be applied based on the color values ​​included in the attributes.

[0086] According to the embodiment, the geometry reconstructor 30005 reconstructs (decompresses) octrees and / or approximate octrees. The geometry reconstructor 30005 reconstructs the octree / voxel based on the results of analyzing the point distribution. The reconstructed octree / voxel may be referred to as the reconstructed geometry (recovered geometry).

[0087] According to an embodiment, the attribute transformer 30007 performs attribute transformation to transform attributes based on reconstructed geometry and / or locations where geometric encoding is not performed. As described above, since attributes depend on geometry, the attribute transformer 30007 can transform attributes based on reconstructed geometric information. For example, based on the position value of a point included in a voxel, the attribute transformer 30007 can transform the attributes of the point at that location. As described above, when the center position of a voxel is set based on the positions of one or more points included in the voxel, the attribute transformer 30007 transforms the attributes of one or more points. When performing triadic geometric encoding, the attribute transformer 30007 can transform attributes based on the triadic geometric encoding.

[0088] The attribute transformer 30007 performs attribute transformation by calculating the average of the attributes or attribute values ​​(e.g., color or reflectivity of each point) of neighboring points within a specific location / radius from the center (or location value) of each voxel. The attribute transformer 30007 can apply weights based on the distance from the center to each point when calculating the average. Therefore, each voxel has a location and a calculated attribute (or attribute value).

[0089] The attribute transformer 30007 can search for nearest neighbors within a specific location / radius of the center of each voxel based on a KD-tree or Morton code. A KD-tree is a binary search tree and supports a data structure that allows points to be managed based on location, enabling fast nearest neighbor search (NNS). Morton codes are generated by representing the coordinates (e.g., (x, y, z)) of the 3D location of all points as bit values ​​and mixing the bits. For example, when the coordinates representing the point location are (5, 9, 1), the bit values ​​are (0101, 1001, 0001). Mixing the bit values ​​according to the bit index in the order of z, y, and x produces 010001000111. This value is represented as the decimal number 1095. That is, the Morton code value for the point with coordinates (5, 9, 1) is 1095. The attribute transformer 30007 can sort the points based on the Morton code values ​​and perform NNS using a depth-first traversal process. After an attribute transformation operation, use a KD tree or Morton code when NNS is needed in another transformation process used for attribute compilation.

[0090] As shown in the figure, the transformation properties are input to the RAHT transformer 40008 and / or the LOD generator 30009.

[0091] According to an embodiment, the RAHT transformer 30008 performs RAHT compilation for predicting attribute information based on reconstructed geometric information. For example, the RAHT transformer 30008 can predict the attribute information of higher-level nodes in an octree based on attribute information associated with lower-level nodes in the octree.

[0092] The LOD generator 30009 according to the embodiment generates a Level of Detail (LOD) to perform predictive transform compilation. The LOD according to the embodiment represents the level of detail of the point cloud content. As the LOD value decreases, it indicates a deterioration in the detail of the point cloud content. As the LOD value increases, it indicates an enhancement in the detail of the point cloud content. Points can be classified by LOD.

[0093] The lift transformer 30010 according to the embodiment performs a lift transform compilation that transforms point cloud attributes based on weights. As described above, the lift transform compilation can optionally be applied.

[0094] According to the embodiment, the coefficient quantizer 30011 quantizes the attribute encoded by the attribute based on the coefficient.

[0095] According to the embodiment, the arithmetic encoder 30012 encodes quantized attributes based on arithmetic compilation.

[0096] Although not shown in the figure, Figure 3 The elements of the point cloud encoder can be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. One or more processors can perform the above-described... Figure 3 At least one of the operation and / or functions of the elements of the point cloud encoder. Additionally, one or more processors are operable or perform operations for executing... Figure 3 The software program and / or instruction set for the operation and / or function of the elements of the point cloud encoder. One or more memories according to the embodiments may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices).

[0097] Figure 4 An example of an octree and occupancy code according to an embodiment is shown.

[0098] For reference Figures 1 to 3 As described, the point cloud content delivery system (point cloud video encoder 10002) or point cloud encoder (e.g., octree analyzer 30002) performs octree geometry compilation (or octree compilation) based on an octree structure to efficiently manage the regions and / or locations of voxels.

[0099] Figure 4 The upper part shows an octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., the X, Y, and Z axes). This is achieved by two poles (0, 0, 0) and (2... d , 2 d , 2 d An octree structure is created by recursively subdividing the bounding box aligned to the cubic axis. Here, 2d can be set as the value of the minimum bounding box that constitutes all points surrounding the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined in the following equation. In the following equation, (x int n , y int n , z int n ) indicates the position (or position value) of the quantized point.

[0100]

[0101] like Figure 4As shown in the upper center, the entire 3D space can be divided into eight spaces according to partitions. Each partitioned space is represented by a cube with six faces. For example... Figure 4 As shown in the upper right, each of the eight spaces is further subdivided based on a coordinate system axis (e.g., the X, Y, and Z axes). Thus, each space is divided into eight smaller spaces. These smaller spaces are also represented by cubes with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.

[0102] Figure 4 The lower part shows the octree occupancy code. The occupancy code generates the octree to indicate whether each of the eight partitions generated by dividing a space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of a partitioned space, and each child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when the space corresponding to a child node contains at least one point, the node is assigned a value of 1. When the space corresponding to a child node does not contain a point (the space is empty), the node is assigned a value of 0. Since... Figure 4 The occupancy code shown is 00100001, therefore, it indicates that the space corresponding to the third and eighth child nodes among the eight child nodes each contains at least one point. As shown, each of the third and eighth child nodes has eight child nodes, and the child nodes are represented by an 8-bit occupancy code. The figure shows that the occupancy code for the third child node is 10000111, and the occupancy code for the eighth child node is 01001111. A point cloud encoder (e.g., an arithmetic encoder 30004) according to an embodiment can perform entropy encoding on the occupancy code. To increase compression efficiency, the point cloud encoder can perform intra-frame / inter-frame compilation on the occupancy code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs the octree based on the occupancy code.

[0103] A point cloud encoder according to an embodiment (e.g., Figure 4 A point cloud encoder or octree analyzer (30002) can perform voxelization and octree compilation to store point locations. However, points are not always uniformly distributed in 3D space, so there may be specific regions with fewer points. Therefore, performing voxelization over the entire 3D space is inefficient. For example, when a specific region contains very few points, voxelization is not necessary in that specific region.

[0104] Therefore, for the aforementioned specific region (or nodes other than the leaf nodes of the octree), the point cloud encoder according to the embodiment can skip voxelization and perform direct compilation to directly encode the point positions included in the specific region. The coordinates of the directly compiled points according to the embodiment are referred to as Direct Compilation Mode (DCM). The point cloud encoder according to the embodiment can also perform triadic geometry encoding based on the surface model, which reconstructs the point positions in the specific region (or node) based on voxels. Triadic geometry encoding is a geometry encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct compilation and triadic geometry encoding according to the embodiment can be performed selectively. In addition, direct compilation and triadic geometry encoding according to the embodiment can be performed in combination with octree geometry encoding (or octree compilation).

[0105] To perform direct compilation, the option to apply direct compilation using direct mode should be enabled. The nodes to be directly compiled are not leaf nodes, and there should be fewer than a threshold number of points within a given node. Furthermore, the total number of points to be directly compiled should not exceed a preset threshold. When the above conditions are met, the point cloud encoder (or arithmetic encoder 30004) according to the embodiment can perform entropy compilation on the point locations (or location values).

[0106] A point cloud encoder (e.g., a surface approximation analyzer 30003) according to an embodiment can determine a specific level of an octree (a level less than the depth d of the octree) and can begin using a surface model at that level to perform triadic geometry encoding to reconstruct point locations in a node region based on voxels (triadic mode). The point cloud encoder according to an embodiment can specify the level at which triadic geometry encoding is applied. For example, the point cloud encoder does not operate in triadic mode when the specific level is equal to the depth of the octree. In other words, the point cloud encoder according to an embodiment can operate in triadic mode only when the specified level is less than the depth value of the octree. A 3D cubic region of a node at a specified level according to an embodiment is called a block. A block may include one or more voxels. A block or voxel may correspond to a cube. Geometry is represented as surfaces within each block. A surface according to an embodiment may intersect each edge of a block at most once.

[0107] A block has 12 edges, and therefore a block contains at least 12 intersections. Each intersection is called a vertex. Vertices along an edge are detected when there is at least one occupied voxel adjacent to the edge in all blocks sharing the edge. An occupied voxel, according to an embodiment, is a voxel containing a point. The vertex position detected along an edge is the average position of the edges of all voxels adjacent to the edge in all blocks sharing the edge.

[0108] Once a vertex is detected, the point cloud encoder according to the embodiment can perform entropy encoding on the edge's origin (x, y, z), the edge's direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When applying triad geometry encoding, the point cloud encoder according to the embodiment (e.g., geometry reconstructor 30005) can generate the restored geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization processes.

[0109] Vertices located at the edges of a block determine the surface passing through the block. According to the embodiment, the surface is a non-planar polygon. During triangle reconstruction, the surface represented by the triangles is reconstructed based on the origin of the edges, the direction vectors of the edges, and the position values ​​of the vertices. The triangle reconstruction process is performed as follows: i) calculating the centroid value of each vertex, ii) subtracting the centroid value from each vertex value, and iii) estimating the sum of squares of the values ​​obtained through the subtraction.

[0110]

[0111] The minimum value of the sum is estimated, and a projection process is performed based on the axis with the minimum value. For example, when element x is minimum, each vertex is projected onto the x-axis relative to the center of the block, and onto the (y, z) plane. When the value obtained by projection onto the (y, z) plane is (ai, bi), the value of θ is estimated by atan2(bi, ai), and the vertices are sorted based on the value of θ. The following shows the vertex combinations for creating triangles based on the number of vertices. Vertices are sorted from 1 to n. Table 1 below shows that for four vertices, two triangles can be constructed based on vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 from the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 from the sorted vertices.

[0112] A triangle formed from vertices sorted by 1, ..., n

[0113] n triangle

[0114] 3 (1,2,3)

[0115] 4 (1,2,3), (3,4,1)

[0116] 5 (1,2,3), (3,4,5), (5,1,3)

[0117] 6 (1,2,3), (3,4,5), (5,6,1), (1,3,5)

[0118] 7 (1,2,3), (3,4,5), (5,6,7), (7,1,3), (3,5,7)

[0119] 8 (1,2,3), (3,4,5), (5,6,7), (7,8,1), (1,3,5), (5,7,1)

[0120] 9 (1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,1,3), (3,5,7), (7,9,3)

[0121] 10 (1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,1), (1,3,5), (5,7,9),(9,1,5)

[0122] 11 (1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,1,3), (3,5,7),(7,9,11), (11,3,7)

[0123] 12 (1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,12,1), (1,3,5),(5,7,9), (9,11,1), (1,5,9)

[0124] An upsampling process is performed to add points along the edges of the triangle at the center, and voxelization is then performed. The added points are generated based on the upsampling factor and the width of the block. The added points are called thinned vertices. The point cloud encoder according to an embodiment can voxelize the thinned vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized positions (or position values).

[0125] Figure 5 An example of point configuration in each LOD according to an embodiment is shown.

[0126] For reference Figures 1 to 4 The described approach involves reconstructing (decompressing) the encoded geometry before performing attribute encoding. When applying direct compilation, the geometry reconstruction operation may include changing the placement of points in the direct compilation (e.g., placing points in front of the point cloud data). When applying triad geometry encoding, the geometry reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.

[0127] A point cloud encoder (e.g., LOD generator 30009) can classify (or reorganize) points according to LOD. The figure shows the point cloud content corresponding to LOD. The leftmost view in the figure represents the original point cloud content. The second view from the left in the figure represents the point distribution in the lowest LOD, and the rightmost view represents the point distribution in the highest LOD. That is, points are sparsely distributed in the lowest LOD and densely distributed in the highest LOD. In other words, as the LOD increases in the direction indicated by the arrow at the bottom of the figure, the space (or distance) between points narrows.

[0128] Figure 6 An example of point configuration for each LOD according to an embodiment is shown.

[0129] For reference Figures 1 to 5 As described, a point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, Figure 3 A point cloud encoder or LOD generator (30009) can generate LODs. LODs are generated by reorganizing points into a set of refined levels based on a set of LOD distance values ​​(or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.

[0130] Figure 6 The upper part shows examples of points (P0 to P9) of point cloud content distributed in 3D space. Figure 6 In this context, the original order represents the order of points P0 to P9 before LOD generation. Figure 6 In this context, LOD-based order represents the order in which points are generated according to their LOD. Points are reorganized by LOD. Additionally, higher LODs include points belonging to lower LODs. For example... Figure 6 As shown, LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.

[0131] For reference Figure 3 As described, the point cloud encoder according to the embodiments may selectively or in combination perform predictive transform compilation, boosting transform compilation, and RAHT transform compilation.

[0132] The point cloud encoder according to the embodiment can generate predictors for points to perform predictive transformations for setting the predictive attributes (or predictive attribute values) of each point. That is, N predictors can be generated for N points. The predictors according to the embodiment can calculate weights (=1 / distance) based on the LOD value of each point, index information of neighboring points existing within a set distance of each LOD, and the distance to the neighboring points.

[0133] According to the embodiment, the predicted attribute (or attribute value) is set as the average of the values ​​obtained by multiplying the attributes (or attribute values) of neighboring points (e.g., color, reflectivity, etc.) set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder (e.g., coefficient quantizer 30011) according to the embodiment can quantize and inverse quantize the residuals (which may be referred to as residual attributes, residual attribute values, or attribute prediction residuals, or attribute residuals) obtained by subtracting the predicted attribute (attribute value) from the attributes (attribute values) of each point. The quantization process is configured as shown in the table below.

[0134] Pseudocode for Residual Quantization of Table Attribute Prediction

[0135] int PCCQuantization(int value, int quantStep) {

[0136] if (value >= 0) {

[0137] return floor(value / quantStep + 1.0 / 3.0);

[0138] } else {

[0139] return -floor(-value / quantStep + 1.0 / 3.0);

[0140] }

[0141] }

[0142] Pseudocode for inverse quantization of residuals in table attribute prediction

[0143] int PCCInverseQuantization(int value, int quantStep) {

[0144] if (quantStep == 0) {

[0145] return value;

[0146] } else {

[0147] return value * quantStep;

[0148] }

[0149] }

[0150] When the predictors of each point have neighboring points, the point cloud encoder (e.g., arithmetic encoder 30012) according to the embodiment can perform entropy compilation on the residual values ​​of quantization and inverse quantization as described above.

[0151] When the predictor of each point has no neighboring points, the point cloud encoder (e.g., arithmetic encoder 30012) according to the embodiment can perform entropy compilation on the attributes of the corresponding point without performing the above operation.

[0152] The point cloud encoder (e.g., lift transformer 30010) according to an embodiment can generate predictors for each point, set the calculated LOD and register neighboring points in the predictors, and set weights based on the distance to the neighboring points to perform lift transform compilation. The lift transform compilation according to the embodiment is similar to the prediction transform compilation described above, but differs in that weights are applied cumulatively to attribute values. The process of cumulatively applying weights to attribute values ​​according to the embodiment is configured as follows.

[0153] 1) Create an array quantized weights (QW) to store the weight values ​​of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighboring nodes registered in the predictor by the weight of the current point's predictor, and add the values ​​obtained by multiplication.

[0154] 2) Improve the prediction process: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.

[0155] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to zero.

[0156] 4) The weights calculated by multiplying the weights computed for all predictors by the weights stored in the QW corresponding to the predictor index are summed with the updateweight array and used as the index of the neighbor node. The values ​​obtained by multiplying the attribute values ​​of the neighbor node indexes by the calculated weights are summed with the update array.

[0157] 5) Improve the update process: Divide the attribute values ​​of the update array of all predictors by the weight values ​​of the updateweight array of the predictor index, and add the existing attribute values ​​to the values ​​obtained by division.

[0158] 6) For all predictors, the predicted attribute is calculated by multiplying the attribute value updated through the boosting update process by the weight updated through the boosting prediction process (stored in QW). The predicted attribute value is quantized by a point cloud encoder (e.g., coefficient quantizer 30011) according to the embodiment. Additionally, the point cloud encoder (e.g., arithmetic encoder 30012) performs entropy encoding on the quantized attribute value.

[0159] A point cloud encoder according to an embodiment (e.g., RAHT transform 30008) can perform RAHT transform compilation, where attributes associated with lower-level nodes in an octree are used to predict attributes of higher-level nodes. RAHT transform compilation is an example of intra-frame attribute compilation by scanning backward through an octree. The point cloud encoder according to an embodiment scans the entire region starting from voxels and repeats a merging process at each step, merging voxels into larger blocks, until the root node is reached. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the node directly above an empty node.

[0160] The following equation represents the RAHT transformation matrix. In this equation, Indicates level The average attribute value of the voxels at that location. Based on and To calculate. and The weight is and .

[0161]

[0162] here, It is a low-pass value and is used in the merging process at the next higher level. This represents the high-pass coefficient. The high-pass coefficient at each step is quantized and subjected to entropy compilation (e.g., encoded by an arithmetic encoder 300012). Weights are calculated as follows: .pass and Create the root node as follows.

[0163]

[0164] The value of gDC is also quantized and subjected to entropy compilation, just like the high-pass coefficient.

[0165] Figure 7 A point cloud decoder according to an embodiment is shown.

[0166] Figure 7 The point cloud decoder shown is an example of a point cloud decoder and can perform decoding operations. Figures 1 to 6 The reverse process of the encoding operation of the point cloud encoder is shown.

[0167] For reference Figure 1 and Figure 6 As described, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.

[0168] The point cloud decoder according to the embodiment includes an arithmetic decoder (arithmetic decoding) 7000, an octree synthesizer (synthesized octree) 7001, a surface approximation synthesizer (synthesized surface approximation) 7002, a geometry reconstructor (reconstructed geometry) 7003, an inverse coordinate transformer (inverse coordinate transformation) 7004, an arithmetic decoder (arithmetic decoding) 7005, an inverse quantizer (inverse quantization) 7006, a RAHT transformer 7007, a LOD generator (generated LOD) 7008, an inverse lifter (inverse lift) 7009, and / or a color inverse transformer (inverse color transformation) 7010.

[0169] An arithmetic decoder 7000, an octree synthesizer 7001, a surface approximation synthesizer 7002, a geometry reconstructor 7003, and a coordinate inverse transformer 7004 can perform geometry decoding. Geometry decoding according to embodiments may include direct decoding and triplet geometry decoding. Direct decoding and triplet geometry decoding are selectively applied. Geometry decoding is not limited to the examples described above and is provided for reference only. Figures 1 to 6 The reverse process of the described geometric encoding is executed.

[0170] According to the embodiment, the arithmetic decoder 7000 decodes the received geometric bitstream based on arithmetic compilation. The operation of the arithmetic decoder 7000 corresponds to the inverse process of the arithmetic encoder 30004.

[0171] The octree synthesizer 7001 according to an embodiment can generate an octree by obtaining a octet code (or information about the geometry obtained as a decoding result) from the decoded geometry bitstream. The octet code is as follows: Figures 1 to 6 Please describe that configuration in detail.

[0172] When applying triplet geometry encoding, the surface approximation synthesizer 7002 according to the embodiment can synthesize the surface based on the decoded geometry and / or the generated octree.

[0173] According to an embodiment, the geometry reconstructor 7003 can regenerate geometry based on surface and / or decoded geometry. See also... Figures 1 to 9 As described, direct compilation and triadic geometry encoding are selectively applied. Therefore, geometry reconstructor 7003 directly imports and sums the positional information of points for which direct compilation has been applied. When triadic geometry encoding is applied, geometry reconstructor 7003 can reconstruct the geometry by performing reconstruction operations (e.g., triangle reconstruction, upsampling, and voxelization) of geometry reconstructor 30005. Details and references Figure 6 The descriptions are the same for all of them, so their descriptions are omitted. The reconstructed geometry may include point cloud images or frames that do not contain attributes.

[0174] According to the embodiment, the inverse coordinate transformer 7004 can obtain the point position based on the reconstructed geometric transformation coordinates.

[0175] An executable reference includes an arithmetic decoder 7005, an inverse quantizer 7006, a RAHT transformer 7007, a LOD generator 7008, an inverse booster 7009, and / or a color inverse transformer 7010. Figure 6 The attribute decoding described herein includes Region Adaptive Hierarchical Transformation (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) decoding, and interpolation-based hierarchical nearest neighbor prediction (lifting transformation) decoding with update / lifting steps. These three decoding schemes may be used selectively, or a combination of one or more decoding schemes may be used. The attribute decoding according to the embodiments is not limited to the examples described above.

[0176] According to the embodiment, the arithmetic decoder 7005 decodes the attribute bitstream through arithmetic compilation.

[0177] According to the embodiment, the inverse quantizer 7006 inversely quantizes information about the decoded attribute bitstream or the attributes obtained as a decoding result, and outputs the inversely quantized attributes (or attribute values). Inverse quantization can be selectively applied based on the attribute encoding of the point cloud encoder.

[0178] According to an embodiment, the RAHT transformer 7007, LOD generator 7008, and / or inverse lifter 7009 can handle the reconstructed geometry and inverse quantization attributes. As described above, the RAHT transformer 7007, LOD generator 7008, and / or inverse lifter 7009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.

[0179] According to the embodiment, the color inverse transformer 7010 performs inverse transformation compilation to inverse transform the color values ​​(or textures) included in the decoded attributes. The operation of the color inverse transformer 7010 can be selectively performed based on the operation of the color transformer 30006 of the point cloud encoder.

[0180] Although not shown in the figure, Figure 7 The elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. One or more processors can perform the above-described... Figure 7 The point cloud decoder's components include at least one or more operations and / or functions. Additionally, one or more processors are operable or perform operations for executing... Figure 7 The software program and / or instruction set for the operation and / or function of the elements of the point cloud decoder.

[0181] Figure 8 A transmitting apparatus according to an embodiment is shown.

[0182] Figure 8 The transmitting device shown is Figure 1The transmitting device 10000 (or Figure 3 Example of a point cloud encoder. Figure 8 The transmitting device shown can perform the same operation as the reference. Figures 1 to 6 The described point cloud encoder includes one or more of the same or similar operations and methods. The transmitting apparatus according to embodiments may include a data input unit 8000, a quantization processor 8001, a voxelization processor 8002, an octree occupancy code generator 8003, a surface model processor 8004, an intra / inter-frame compilation processor 8005, an arithmetic encoder 8006, a metadata processor 8007, a color transformation processor 8008, an attribute transformation processor 8009, a prediction / boosting / RAHT transformation processor 8010, an arithmetic encoder 8011, and / or a transmission processor 8012.

[0183] According to an embodiment, the data input unit 8000 receives or acquires point cloud data. The data input unit 8000 can perform operations and / or acquisition methods related to the point cloud video acquirer 10001 (or refer to...). Figure 2 The described acquisition process (20000) is the same or similar operation and / or acquisition method.

[0184] The data input unit 8000, quantization processor 8001, voxelization processor 8002, octree occupancy code generator 8003, surface model processor 8004, intra / inter-frame compilation processor 8005, and arithmetic encoder 8006 perform geometric encoding. Geometric encoding and reference according to the embodiment. Figures 1 to 9 The geometric codes described are the same or similar, so their detailed descriptions are omitted.

[0185] The quantization processor 8001 according to the embodiment quantizes geometry (e.g., point position values). The operation of the quantization processor 8001 and / or quantization with reference... Figure 3 The operation and / or quantization of the described quantizer 30001 are the same or similar. Details and references Figures 1 to 9 The descriptions are the same.

[0186] According to the embodiment, the voxelization processor 8002 voxels the quantized position values ​​of points. The voxelization processor 8002 can perform operations similar to those described above. Figure 3 The operation and / or voxelization process of the quantizer 30001 described are the same as or similar to the operation and / or process. Details and references Figures 1 to 6 The descriptions are the same.

[0187] According to the embodiment, the octree occupancy code generator 8003 performs octree compilation based on the voxelized positions of points in the octree structure. The octree occupancy code generator 8003 can generate occupancy codes. The octree occupancy code generator 8003 can execute and reference... Figure 3 and Figure 4The operations and / or methods described are the same as or similar to those of the point cloud encoder (or octree analyzer 30002). Details and references Figures 1 to 6 The descriptions are the same.

[0188] According to an embodiment, the surface model processor 8004 can perform triadic geometry encoding based on a surface model to reconstruct point positions in a specific region (or node) based on voxels. The surface model processor 8004 can perform operations with reference to... Figure 3 The operations and / or methods described are the same as or similar to those of the point cloud encoder (e.g., surface approximation analyzer 30003). Details and references are available. Figures 1 to 6 The descriptions are the same.

[0189] According to an embodiment, the intra / inter-frame compilation processor 8005 can perform intra / inter-frame compilation on point cloud data. The intra / inter-frame compilation processor 8005 can execute reference... Figure 7 The compilation described is the same as or similar to intra / inter-frame compilation. Details and references. Figure 7 Those described are the same. According to an embodiment, the intra / inter-frame coding processor 8005 may be included in the arithmetic encoder 8006.

[0190] According to an embodiment, the arithmetic encoder 8006 performs entropy encoding on octrees and / or approximate octrees of point cloud data. For example, the encoding scheme includes arithmetic encoding. The arithmetic encoder 8006 performs the same or similar operations and / or methods as the arithmetic encoder 30004.

[0191] According to an embodiment, the metadata processor 8007 processes metadata (e.g., set values) about point cloud data and provides it to necessary processing procedures such as geometric encoding and / or attribute encoding. Additionally, according to an embodiment, the metadata processor 8007 can generate and / or process signaling information related to geometric encoding and / or attribute encoding. The signaling information according to an embodiment can be encoded separately from the geometric encoding and / or attribute encoding. The signaling information according to an embodiment can be interleaved.

[0192] Color transformation processor 8008, attribute transformation processor 8009, prediction / boosting / RAHT transformation processor 8010, and arithmetic encoder 8011 perform attribute encoding. Attribute encoding and reference according to the embodiment. Figures 1 to 6 The attribute codes described are the same or similar, so their detailed descriptions are omitted.

[0193] According to an embodiment, a color transformation processor 8008 performs color transformation compilation to transform color values ​​included in attributes. The color transformation processor 8008 may perform color transformation compilation based on reconstructed geometry. The reconstructed geometry and reference... Figures 1 to 9 The description is the same. Furthermore, its execution is the same as the reference.Figure 3 The operation and / or methods of the described color converter 30006 are the same as or similar to those described. Detailed descriptions are omitted.

[0194] According to an embodiment, the attribute transformation processor 8009 performs attribute transformation to transform attributes based on reconstructed geometry and / or locations where geometric encoding is not performed. The attribute transformation processor 8009 performs transformations with reference to... Figure 3 The operations and / or methods of the described attribute transformer 30007 are the same as or similar to those described. Detailed descriptions thereof are omitted. The prediction / boosting / RAHT transformation processor 8010 according to the embodiment can encode the transformed attributes through any one or a combination of RAHT compilation, prediction transformation compilation, and boosting transformation compilation. The prediction / boosting / RAHT transformation processor 8010 performs operations and references... Figure 3 The RAHT transformer 30008, LOD generator 30009, and boost transformer 30010 described herein operate at least one of the same or similar operations. Additionally, the predictive transform compilation, boost transform compilation, and RAHT transform compilation are related to the reference... Figures 1 to 9 The descriptions are the same, so their detailed descriptions are omitted.

[0195] According to an embodiment, the arithmetic encoder 8011 can encode compiled attributes based on arithmetic compilation. The arithmetic encoder 8011 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 300012.

[0196] According to an embodiment, the transmission processor 8012 can transmit various bitstreams containing encoded geometric and / or encoded attribute and metadata information, or transmit a single bitstream configured with encoded geometric and / or encoded attribute and metadata information. When the encoded geometric and / or encoded attribute and metadata information according to an embodiment is configured as a single bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information and slice data. The signaling information includes a sequence parameter set (SPS) for sequence-level signaling, a geometric parameter set (GPS) for signaling for geometric information compilation, an attribute parameter set (APS) for signaling for attribute information compilation, and a tile parameter set (TPS) for tile-level signaling. The slice data may include information about one or more slices. A slice according to an embodiment may include a geometric bitstream Geom00 and one or more attribute bitstreams Attr00 and Attr10.

[0197] A slice is a series of syntactic elements that represent all or part of a compiled point cloud frame.

[0198] According to an embodiment, the TPS may include information about individual tiles in one or more tiles (e.g., coordinate information and height / size information about the bounding box). The geometric bitstream may include a header and a payload. The header of the geometric bitstream according to an embodiment may include a geom_parameter_set_id, a geom_tile_id, and a geom_slice_id included in the GPS, as well as information about the data contained in the payload. As described above, the metadata processor 8007 according to an embodiment may generate and / or process signaling information and transmit it to the transmission processor 8012. According to an embodiment, the element performing geometry encoding and the element performing attribute encoding may share data / information with each other, as indicated by the dashed lines. The transmission processor 8012 according to an embodiment may perform the same or similar operations and / or transmission methods as the transmitter 10003. Details and References Figure 1 and Figure 2 The descriptions are the same, so their descriptions are omitted.

[0199] Figure 9 An example of a receiving device according to an embodiment is shown.

[0200] Figure 9 The receiving device shown is Figure 1 The receiving device 10004 (or Figure 10 and Figure 11 An example of a point cloud decoder. Figure 9 The receiving device shown can perform the same operation as the reference. Figures 1 to 11 The same or similar one or more operations and methods described in the point cloud decoder.

[0201] The receiving apparatus according to an embodiment may include a receiver 9000, a receiving processor 9001, an arithmetic decoder 9002, an octree reconstruction processor based on occupancy codes 9003, a surface model processor (triangle reconstruction, upsampling, voxelization) 9004, an inverse quantization processor 9005, a metadata parser 9006, an arithmetic decoder 9007, an inverse quantization processor 9008, a prediction / boost / RAHT inverse transform processor 9009, a color inverse transform processor 9010, and / or a renderer 9011. Each decoding element according to an embodiment can perform the inverse process of the operation of the corresponding encoding element according to the embodiment.

[0202] Receiver 9000 according to an embodiment receives point cloud data. Receiver 9000 can perform operations related to... Figure 1 The operation and / or receiving method of the receiver 10005 are the same as or similar to those of the receiver. Detailed description omitted.

[0203] According to an embodiment, the receiving processor 9001 can acquire a geometric bitstream and / or an attribute bitstream from the received data. The receiving processor 9001 may be included in the receiver 9000.

[0204] The arithmetic decoder 9002, the octet-based octree reconstruction processor 9003, the surface model processor 9004, and the inverse quantization processor 9005 are capable of performing geometric decoding. Geometric decoding and reference according to the embodiment... Figures 1 to 10 The described geometric decodings are the same or similar, so their detailed descriptions are omitted.

[0205] The arithmetic decoder 9002 according to an embodiment can decode a geometric bitstream based on arithmetic compilation. The arithmetic decoder 9002 performs the same or similar operations and / or compilation as the arithmetic decoder 7000.

[0206] According to an embodiment, the octree reconstruction processor 9003 based on occupancy codes can reconstruct an octree by obtaining occupancy codes from the decoded geometric bitstream (or information about the geometry obtained as a decoding result). The octree reconstruction processor 9003 performs the same or similar operations and / or methods as the octree synthesizer 7001 and / or the octree generation method. When applying triad geometry encoding, the surface model processor 9004 according to an embodiment can perform triad geometry decoding and related geometric reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on surface modeling methods. The surface model processor 9004 performs the same or similar operations as the surface approximation synthesizer 7002 and / or the geometry reconstructor 7003.

[0207] The geometry of reversible quantization decoding according to the embodiment of the inverse quantization processor 9005.

[0208] Metadata parser 9006 according to an embodiment can parse metadata (e.g., set values) contained in received point cloud data. Metadata parser 9006 can pass the metadata to geometry decoder and / or attribute decoder. Metadata and reference Figure 8 The metadata described is the same, so its detailed description is omitted.

[0209] Arithmetic decoder 9007, inverse quantization processor 9008, prediction / boost / RAHT inverse transform processor 9009, and color inverse transform processor 9010 perform attribute decoding. Attribute decoding and reference Figures 1 to 10 The properties described are decoded the same or similarly, so their detailed descriptions are omitted.

[0210] The arithmetic decoder 9007 according to an embodiment can decode the attribute bitstream via arithmetic compilation. The arithmetic decoder 9007 can decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 9007 performs the same or similar operations and / or compilation as the arithmetic decoder 7005.

[0211] According to an embodiment, the inverse quantization processor 9008 can reversibly quantize and decode attribute bitstreams. The inverse quantization processor 9008 performs the same or similar operations and / or methods as the inverse quantizer 7006 and / or the inverse quantization method.

[0212] According to an embodiment, the predictive / boosting / RAHT inverse transform processor 9009 can process reconstructed geometry and inverse quantization attributes. The predictive / boosting / RAHT inverse transform processor 9009 performs one or more operations and / or decodings that are the same as or similar to those of the RAHT transformer 7007, LOD generator 7008, and / or inverse booster 7009. According to an embodiment, the color inverse transform processor 9010 performs inverse transform compilation to inverse transform color values ​​(or textures) included in the decoded attributes. The color inverse transform processor 9010 performs operations and / or inverse transform compilations that are the same as or similar to those of the color inverse transformer 7010. According to an embodiment, the renderer 9011 can render point cloud data.

[0213] Figure 10 An exemplary structure operable in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment is shown.

[0214] Figure 10 The structure represents a configuration in which at least one of the following components—server 1060, robot 1010, autonomous vehicle 1020, XR device 1030, smartphone 1040, home appliance 1050, and / or head-mounted display (HMD) 1070—is connected to cloud network 1000. Robot 1010, autonomous vehicle 1020, XR device 1030, smartphone 1040, or home appliance 1050 are referred to as devices. Furthermore, XR device 1030 may correspond to a point cloud data (PCC) device according to an embodiment or be operatively connected to a PCC device.

[0215] Cloud network 1000 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, cloud network 1000 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0216] Server 1060 can be connected via cloud network 1000 to at least one of robot 1010, self-driving vehicle 1020, XR device 1030, smartphone 1040, home appliance 1050 and / or HMD 1070, and can assist at least a portion of the processing of connected devices 1010 to 1070.

[0217] HMD 1070 represents one of the implementation types of the XR device and / or PCC device according to the embodiments. The HMD-type device according to the embodiments includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0218] Hereinafter, various embodiments of the apparatus 1010 to 1050 that apply the above-described technology will be described. Figure 10 The devices 1010 to 1050 shown are operable to be connected to / coupled to the point cloud data transmitting and receiving devices according to the above embodiments.

[0219] <PCC+XR>

[0220] The XR / PCC device 1030 may employ PCC technology and / or XR (AR+VR) technology, and may be implemented as an HMD, a head-up display (HUD) installed in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0221] The XR / PCC device 1030 can analyze 3D point cloud data or image data acquired through various sensors or from external devices and generate positional and attribute data about 3D points. Thus, the XR / PCC device 1030 can acquire information about the surrounding space or real-world objects and render and output XR objects. For example, the XR / PCC device 1030 can match an XR object, including auxiliary information about the identified object, with the identified object and output the matched XR object.

[0222] <PCC+XR+Mobile Phone>

[0223] The XR / PCC device 1030 can be implemented as a smartphone 1040 by applying PCC technology.

[0224] The 1040 smartphone can decode and display point cloud content based on PCC technology.

[0225] <PCC+Self-driving+XR>

[0226] The self-driving vehicle 1020 can be realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0227] The self-driving vehicle 1020 employing XR / PCC technology can refer to a self-driving vehicle equipped with means for providing XR images, or a self-driving vehicle serving as a control / interaction target in an XR image. Specifically, as a control / interaction target in an XR image, the self-driving vehicle 1020 can be distinguished from and operatively connected to the XR device 1030.

[0228] The autonomous vehicle 1020, equipped with means for providing XR / PCC images, can acquire sensor information from sensors including cameras and output generated XR / PCC images based on the acquired sensor information. For example, the autonomous vehicle 1020 may have a HUD and output XR / PCC images to it, thereby providing passengers with XR / PCC objects corresponding to real objects or objects presented on the screen.

[0229] When an XR / PCC object is output to a HUD, at least a portion of the XR / PCC object can be output to overlap with the actual object being pointed at by the passenger's eyes. Conversely, when an XR / PCC object is output to a display installed within the autonomous vehicle, at least a portion of the XR / PCC object can be output to overlap with objects on the screen. For example, the autonomous vehicle 1220 can output XR / PCC objects corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two-wheeled vehicles, pedestrians, and buildings.

[0230] Virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or point cloud compression (PCC) technologies according to the embodiments are applicable to various devices.

[0231] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. AR technology, on the other hand, refers to the technology of displaying virtually created CG images on top of images of real objects. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology in that AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.

[0232] Recently, VR, AR, and MR technologies have often been referred to as extended reality (XR) technologies rather than being explicitly distinguished from each other. Therefore, embodiments of this disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are suitable for such technologies.

[0233] The PCC method / apparatus according to the embodiments can be applied to vehicles that provide autonomous driving services.

[0234] Vehicles providing autonomous driving services connect to the PCC device for wired / wireless communication.

[0235] When the point cloud data (PCC) transmitting / receiving device according to an embodiment is connected to a vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services (which may be provided together with autonomous driving services) and transmit it to the vehicle. If the PCC transmitting / receiving device is installed in the vehicle, it can receive / process content data related to AR / VR / PCC services based on user input signals input through a user interface device and provide it to the user. The vehicle or user interface device according to an embodiment can receive user input signals. User input signals according to an embodiment may include signals indicating autonomous driving services.

[0236] The point cloud data transmission method / device according to the embodiments can be interpreted as a reference. Figure 1 10000 transmission devices Figure 1 Point cloud video encoder 10002 Figure 1 Transmitter 10003 Figure 2 Acquisition 20000 / Encoding 20001 / Transmission 20002 Figure 3 encoder, Figure 8 Transmission equipment Figure 10 equipment Figure 11 , 20 Encoders of 24 and 30 to 32 Figure 33 Terms related to transmission methods, etc.

[0237] The point cloud data transmission method / device according to the embodiments can be interpreted as a reference. Figure 1 10000 transmission devices Figure 1 Point cloud video encoder 10002 Figure 1 Transmitter 10003 Figure 2 Acquisition 20000 / Encoding 20001 / Transmission 20002 Figure 3 encoder, Figure 8 Transmission equipment Figure 10 equipment Figure 11 , 20 Encoders of sizes 25 and 30 to 32 Figure 34 Terms related to transmission methods, etc.

[0238] Furthermore, the point cloud data transmission / reception method / device according to the embodiments can be simply referred to as the method / device.

[0239] According to the embodiment, the geometric data, geometric information, location information, and geometry constituting the point cloud data will be interpreted as having the same meaning. Attribute data and the attribute information constituting the point cloud data will also be interpreted as having the same meaning.

[0240] The method / apparatus according to the embodiments may include and perform a boost transformation for partial compilation of G-PCC.

[0241] The embodiments include a method for effectively supporting selective decoding of a portion of data when necessary, due to receiver performance or transmission speed during the transmission and reception of point cloud data. The embodiments also include a method for selecting necessary information or removing unnecessary information from a bitstream unit by partitioning the geometric and attribute data delivered for each data unit into syntax units such as geometric octrees and Level of Detail (LoD).

[0242] The embodiments include a method for configuring a data structure composed of point clouds. More specifically, the embodiments disclose an encapsulation and signaling method for efficiently transmitting layer-configured point cloud compression (PCC) data, and include a method for applying the encapsulation and signaling method to a scalable PCC-based service. In particular, the embodiments include a method for configuring and sending and receiving segments to be more suitable for scalable PCC services when a direct compression mode is used for location compression. In particular, the embodiments can provide a compressed structure for efficiently storing and transmitting large point cloud data with wide distribution and high point density.

[0243] refer to Figure 3 Point cloud data consists of the location (geometry) (e.g., XYZ coordinates) and attributes (e.g., color, reflectivity, intensity, grayscale, opacity, etc.) of each data point. In point cloud compression (PCC), octree-based compression is performed to effectively compress the non-uniform distribution characteristics in 3D space, and attribute information is compressed based on this. Figure 3 This is a flowchart of the sending and receiving sides of PCC.

[0244] Figure 11 The process of encoding, transmitting, and decoding point cloud data according to an embodiment is illustrated.

[0245] The point cloud encoder 15000 is a transmission device that performs the transmission method according to the embodiment and can scalably encode and transmit point cloud data.

[0246] The point cloud decoder 15010 is a receiving device that performs the receiving method according to the embodiment and can scalably decode point cloud data.

[0247] The source data received by encoder 15000 may include geometric data and / or attribute data.

[0248] The encoder 15000 scalably encodes point cloud data but does not immediately generate a partial PCC bitstream. Instead, when it receives complete geometry and attribute data, it stores the data in a storage device connected to the encoder. The encoder then performs transcoding for partial encoding and generates and transmits a partial PCC bitstream. The decoder 15010 receives and decodes the partial PCC bitstream to reconstruct partial geometry and / or partial attributes.

[0249] Upon receiving the complete geometry and attributes, encoder 15000 can store the data in a storage device connected to the encoder and transcode the point cloud data using a low quantization parameter (QP) to generate and transmit a complete PCC bitstream. Decoder 15010 can receive and decode the complete PCC bitstream to reconstruct the complete geometry and / or complete attributes. Decoder 15010 can select portions of the geometry and / or attributes from the complete PCC bitstream via data selection.

[0250] The method / apparatus according to the embodiment compresses and transmits point cloud data by dividing the location information of data points and feature information such as color / brightness / reflectivity into geometric information and attribute information. In this case, an octree structure with layers can be configured according to the level of detail, or PCC data can be configured according to the level of detail (LoD). Scalable point cloud data compilation and representation can then be performed based on the configured structure or data. In this case, due to the performance of the receiver or the transmission rate, only a portion of the point cloud data can be decoded or represented.

[0251] In this process, the method / apparatus according to the embodiment can remove unnecessary data in advance. In other words, when only a portion of the scalable PCC bitstream needs to be sent (i.e., only some layers are decoded in scalable decoding), it is not possible to select and send only the necessary portion. Therefore, 1) the necessary portion needs to be re-encoded after decoding (15020), or 2) the receiver must selectively apply the operation after the entire data has been transmitted to the receiver (15030). However, in case 1), a delay may occur due to the time spent on decoding and re-encoding (15020). In case 2), bandwidth efficiency may be reduced due to the transmission of unnecessary data. Furthermore, when using fixed bandwidth, it may be necessary to reduce the data quality for transmission (15030).

[0252] Therefore, the method / apparatus according to the embodiment can define a slice segmentation structure for point cloud data and signal the scalable layer and slice structure for scalable transmission.

[0253] In an embodiment, to ensure efficient bitstream delivery and decoding, the bitstream can be divided into specific units to be processed.

[0254] For octree-based geometric compression, the methods / devices according to the embodiments can be used together with entropy-based compilation and direct compilation. In this case, a slice configuration is needed to effectively utilize scalability.

[0255] According to embodiments, units can be referred to as LODs, layers, slices, etc. LOD is the same term as LOD in attribute data compilation, but can refer to a data unit of a hierarchical structure of a bitstream. It can be a concept corresponding to a bundle of one or two or more depths of a hierarchical structure based on point cloud data, such as the depth (level) of an octree or multiple trees. Similarly, a layer is provided as a unit to generate a sub-bitstream and is a concept corresponding to a bundle of one or two or more depths, and can correspond to one LOD or two or more LODs. Furthermore, a slice is a unit for configuring units of a sub-bitstream and can correspond to one depth, a portion of a depth, or two or more depths. Additionally, it can correspond to one LOD, a portion of an LOD, or two or more LODs. According to embodiments, LODs, layers, and slices can correspond to each other, or one of LODs, layers, and slices can be included in another. Furthermore, units according to embodiments can include LODs, layers, slices, layer groups, or subgroups, and can be referred to as complementary to each other.

[0256] Furthermore, when using attribute compilation (such as predictive transformations), it is necessary to maintain compilation constraints and efficiency to support partial decoding. Implementations can address these challenges by leveraging lifting transformation methods for partial compilation, providing an update process that considers subgroup boundaries and / or methods such as subgroup incremental QP and hierarchical subgroup incremental QP.

[0257] Figure 12 The layer-based point cloud data configuration and the structure of the geometry and attribute bitstreams according to an embodiment are shown.

[0258] The transmission method / device according to the embodiment can be configured as follows: Figure 12 The layer-based point cloud data shown is used for encoding and decoding of point cloud data.

[0259] The embodiments relate to the efficient transmission and decoding of data in bitstream units of point cloud data configured in layers by selectively sending and decoding.

[0260] Point cloud data can be layered according to the application domain, with layered structures in terms of SNR, spatial resolution, color, temporal frequency, bit depth, etc., and layers can be constructed in the direction of increasing data density based on octree or LoD structure.

[0261] The method / apparatus according to the embodiments can be based on, for example... Figure 16 The layers shown are used to configure, encode, and decode geometric bitstreams and attribute bitstreams.

[0262] The bitstream obtained by the transmitting device / encoder according to the embodiment through point cloud compression can be divided into geometric data bitstream and attribute data bitstream according to the data type and then transmitted.

[0263] Each bitstream according to the embodiment can be composed of slices. Regardless of the layer information or LoD information, the geometric data bitstream and the attribute data bitstream can each be configured as a slice and delivered. In this case, when only a portion of the layer or LoD is used, the following should be performed: 1) decode the bitstream, 2) select only the desired portion and remove unnecessary portions, and 3) re-encode only based on the necessary information.

[0264] [[ID= A bitstream configuration according to an embodiment is shown.

[0265] The sending method / device according to the embodiment can generate, as shown in the example. ​ The bitstream shown, and the receiving method / device according to the embodiment, can decode the point cloud data included in the bitstream, such as... ​ As shown.

[0266] Bitstream configuration according to the embodiment

[0267] In this embodiment, to avoid unnecessary intermediate processes, the bitstream can be divided into layers (or LoD) and sent.

[0268] For example, in a LoD-based PCC structure, lower LoDs are included in higher LoDs. Information included in the current LoD but not in previous LoDs—that is, information newly included in each LoD—can be referred to as R (Rest). ​ As shown, the initial LoD information and the new information R included in each LoD can be partitioned and transmitted in each independent unit.

[0269] The transmission method / apparatus according to the embodiments can encode geometric data and generate a geometric bitstream. The geometric bitstream can be configured for each LOD or layer. The geometric bitstream may contain a header (geometric header) for each LOD or layer. The header may include reference information for the next LOD or layer. The current LOD (layer) may also include information R (geometric data) not included in previous LODs (layers).

[0270] The receiving method / device according to the embodiment can encode attribute data and generate an attribute bitstream. The attribute bitstream can be configured for each LOD or layer, and the attribute bitstream can contain a header (attribute header) for each LOD or layer. The header can include reference information for the next LOD or layer. The current LOD (layer) can also include information R (attribute data) not included in previous LODs (layers).

[0271] The receiving method / device according to the embodiments can receive a bitstream consisting of LODs or layers and efficiently decode only the necessary data without complex intermediate processes.

[0272] ​ A bitstream sorting method according to an embodiment is shown.

[0273] The method / apparatus according to the embodiments can be used for ​ Sort the bitstream, such as ​ As shown.

[0274] Bitstream sorting method according to the embodiment

[0275] When transmitting bit streams, the transmission method / device according to the embodiments can serially transmit geometry and attributes, such as... ​ As shown. In this case, depending on the type of data, the entire geometric information (geometric data) can be sent first, followed by the attribute information (attribute data). In this case, the geometric information can be quickly reconstructed based on the sent bitstream information.

[0276] exist ​ In (a), for example, the layer (LOD) containing geometric data may be located first in the bitstream, and the layer (LOD) containing attribute data may be located after the geometric layer. Since the attribute data depends on the geometric data, the geometric layer can be placed first. Furthermore, the positions can be varied depending on the embodiment. References can also be made between geometric headers and between attribute headers and geometric headers.

[0277] refer to ​-(b) allows for the collection and transmission of bitstreams comprising geometric and attribute data that constitute the same layer. In this case, the decoding execution time can be reduced by using compression techniques capable of decoding geometry and attributes in parallel. In this scenario, information requiring immediate processing (lower LoD, where geometry must precede attributes) can be placed first.

[0278] The first layer 1800 includes the geometric and attribute data corresponding to the lowest LOD 0 (layer 0) and each header. The second layer 1810 includes LOD 0 (layer 0) and also includes the geometric and attribute data of points in a new, more detailed layer 1 (LOD 1) that are not included in LOD 0 (layer 0) as information R1. The third layer 1820 can then be placed in a similar manner.

[0279] When sending and receiving bitstreams, the sending / receiving method / device according to the embodiments can effectively select the desired layer (or LoD) in the application field at the bitstream level. In the bitstream ordering method according to the embodiments, geometric information is collected and sent ( ​ Empty portions can occur in the middle after bitstream level selection. In this case, bitstream rearrangement may be necessary. This is in accordance with the geometry and properties of each layer's bundling and delivery. ​ Unnecessary information can be selectively removed based on the application fields as follows.

[0280] ​ A method for selecting geometric data and attribute data according to an embodiment is shown.

[0281] The bitstream selection is performed as follows, according to an embodiment of the method / device.

[0282] When it is necessary to select a bitstream as described above, the method / device according to the embodiment can select data at the bitstream level: 1) symmetric selection of geometry and attributes; 2) asymmetric selection of geometry and attributes; or 3) a combination of the above two methods.

[0283] 1) Symmetrical selection of geometry and properties

[0284] refer to ​ It shows the following situation: only when the LoD up to LOD 1 (LOD 0+R1) is selected (19000) and sent or decoded, the information corresponding to R2 (the new part in LOD 2) is removed for transmission / decoding, where R2 corresponds to the upper layer.

[0285] 2) Asymmetric selection of geometry and properties

[0286] The method / apparatus according to the embodiment can transmit geometry and attributes asymmetrically. Only the attributes of the upper layer (attribute R2) are removed (19001), and the complete geometry (from level 0 (root level) to level 7 (leaf level) in a triangular octree structure) can be selected and sent / decoded (19011).

[0287] refer to ​ When point cloud data is represented in an octree structure and hierarchically divided into LODs (or layers), scalable encoding / decoding (scalability) can be supported.

[0288] The scalability features according to the embodiments may include slice-level scalability and / or octree-level scalability.

[0289] According to the embodiment, the LoD (Level of Detail) can be used as a unit to represent a collection of one or more octree layers. Alternatively, it can represent a bundle of octree layers to be configured as slices.

[0290] In attribute encoding / decoding, the LOD according to the embodiment can be extended and used as a unit for detailed partitioning of data in a broader sense.

[0291] In other words, spatial scalability of the actual octree layer (or scalable attribute layer) can be provided for each octree layer. However, when configuring scalability in the slice prior to bitstream parsing, the choice can be made in the LoD depending on the implementation.

[0292] In an octree structure, LOD 0 can correspond to the root level of level 4, LOD 1 can correspond to the root level of level 5, and LOD 2 can correspond to the root level of level 7, which is the leaf level.

[0293] In other words, such as ​ As shown, when scalability is utilized in a slice, such as in the case of scalable transmission, the provided scalable steps can correspond to three steps of LOD 0, LOD 1 and LOD 2, and in the decoding operation, the scalable steps provided by the octree structure can correspond to eight steps from root to leaf.

[0294] According to an embodiment, for example, in ​ In the process, when LOD 0 to LOD 2 are configured as the corresponding slices, the transcoder of the receiver or transmitter ( ​ The transcoder (15040) can be configured to: 1) LOD 0 only, 2) LOD 0 and LOD 1, or 3) LOD 0, LOD 1 and LOD 2 for scalable processing.

[0295] Example 1: When only LOD 0 is selected, the maximum octree level can be 4, and a scalable layer can be selected from octree levels 0 to 4 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0296] Example 2: When LOD 0 and LOD 1 are selected, layer 5 can be added. Therefore, the maximum octree level can be 5, and a scalable layer can be selected from octree layers 0 to 5 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0297] According to an embodiment, the octree depth, octree level, and octree hierarchy can be units for detailed data partitioning.

[0298] Example 3: When LOD 0, LOD 1, and LOD 2 are selected, layers 6 and 7 can be added. Therefore, the maximum octree level can be 7, and a scalable layer can be selected from octree layers 0 to 7 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0299] ​ A method for slicing point cloud data according to an embodiment is shown.

[0300] Slicing configuration according to the embodiment

[0301] According to the embodiments, the transmission method / device / encoder can configure the G-PCC bitstream by dividing the bitstream in a slice structure. The data unit used for detailed data representation can be a slice.

[0302] A slice, according to an embodiment, can represent a data unit used to segment point cloud data. That is, a slice represents a portion of point cloud data. A slice can be referred to as a term representing a part or unit.

[0303] For example, one or more octree layers can be matched with a slice.

[0304] According to the embodiments, the transmission method / device (e.g., an encoder) can configure a bit stream based on slice 2001 by scanning the nodes (points) included in the octree in the direction of scan sequence 2000.

[0305] exist ​ In (a), some nodes in an octree can be included in a slice.

[0306] An octree layer (e.g., level 0 to level 4) can form a slice 2002.

[0307] Partial data from an octree level (e.g., level 5) can form each slice 2003, 2004, 2005.

[0308] Partial data from an octree level (e.g., level 6) can form each slice.

[0309] exist ​ -(b) and ​ In (c), when multiple octree layers match a slice, only some nodes from each layer may be included. In this way, when multiple slices constitute a geometry / attribute frame, the information needed to configure the layers can be delivered to the receiver. This information may include information about the layers included in each slice and information about the nodes included in each layer.

[0310] exist ​ In -(b), partial data from the octree levels (e.g., levels 0 to 3) and level 4 can be configured as a slice.

[0311] An octree level (e.g., partial data at level 4 and partial data at level 5) can be configured as a slice.

[0312] An octree level (e.g., partial data at level 5 and partial data at level 6) can be configured as a slice.

[0313] An octree level (e.g., a portion of the data at level 6) can be configured as a slice.

[0314] exist ​ In -(c), an octree level (e.g., data from level 0 to level 4) can be configured as a slice.

[0315] Partial data from each of the octree levels 5, 6, and 7 can be configured as a slice.

[0316] The encoder and the corresponding device according to the embodiment can encode point cloud data and can generate and send a bit stream including encoded data and parameter information related to the point cloud data.

[0317] Furthermore, when generating the bitstream, the bitstream can be generated based on the bitstream structure according to the embodiment (for example, see...). ​ Therefore, the receiving device, decoder, and corresponding device according to the embodiments can receive and parse bitstreams configured for selective partial data decoding, and partially decode and efficiently provide point cloud data (see [link]). ​ ).

[0318] Scalable transmission according to the embodiment

[0319] The point cloud data transmission method / device according to the embodiments can scalably transmit bit streams including point cloud data, and the point cloud data receiving method / device according to the embodiments can scalably receive and decode bit streams.

[0320] When according to ​ When the bitstream of the illustrated embodiment is used for scalable transmission, information needed to select the desired slice for the receiver can be sent to the receiver. Scalable transmission can mean transmitting or decoding only a portion of the bitstream, rather than decoding the entire bitstream, and the result can be low-resolution point cloud data.

[0321] When scalable transmission is applied to octree-based geometric bitstreams, for each octree level from the root node to the leaf node ( ​ For bitstreams, point cloud data may need to be configured with information that only extends to a specific octree layer.

[0322] Therefore, the target octree layer should not depend on information about the lower octree layers. This can be a constraint that applies to both geometry compilation and attribute compilation.

[0323] Additionally, in scalable transmission, a scalable structure is required for the transmitter / receiver to select scalable layers. Considering an octree structure according to an embodiment, all octree layers can support scalable transmission, or scalable transmission can be allowed only for specific octree layers or lower. When a slice includes some octree layers, the scalable layers that include the slice can be indicated. Therefore, it can be determined whether the slice is necessary / unnecessary at the bitstream stage. ​ In the example of -(a), the yellow portion starting from the root node constitutes a scalable layer, but does not support scalable transport. Subsequent octree layers can be matched with scalable layers in a one-to-one correspondence. Typically, the portions corresponding to leaf nodes can support scalability. When a slice includes multiple octree layers, a scalable layer can be defined to be configured for these layers.

[0324] In this context, scalable transmission and scalable decoding can be used separately, depending on the purpose. Scalable transmission can be used on the sending / receiving side for selecting information up to a specific layer, without involving the decoder. Scalable decoding is used to select a specific layer during compilation. That is, scalable transmission can support the selection of necessary information without involving the decoder in a compressed state (at the bitstream stage), allowing information to be sent or determined by the receiver. On the other hand, scalable decoding can support encoding / decoding only the data up to the required portion of the encoding / decoding process, and can therefore be used as a scalable representation in this case.

[0325] In this context, the layer configuration for scalable transport can differ from the layer configuration for scalable decoding. For example, three bottom octree layers including leaf nodes can constitute a single layer in terms of scalable transport. However, when all layer information is included in terms of scalable decoding, scalable decoding can be performed for each of leaf node layer n, leaf node layer n-1, and leaf node layer n-2.

[0326] The method / apparatus according to the embodiments can perform fine-grained slicing.

[0327] When fine-grained slicing is enabled, the G-PCC bitstream (see example) can be included in the encoding operation of the point cloud transmission method according to the embodiment. ​ (etc.) is split into multiple sub-bitstreams. To effectively utilize the hierarchical structure of G-PCC, each slice can include data for compiling a portion of the compilation layer or a portion of the region. By using slice partitioning paired with the compilation layer structure, scalable transport or spatial random access use cases can be supported in an efficient manner.

[0328] ​ The geometric compilation layer structure according to an embodiment is shown.

[0329] The point cloud data transmission method / device according to the embodiments, namely ​ 10000 transmission devices ​ Point cloud video encoder 10002 ​ Transmitter 10003 ​ Acquisition 20000 / Encoding 20001 / Transmission 20002 ​ encoder, ​ Transmission equipment ​ equipment ​ , ​ , ​ and ​ encoder, ​ Transmission methods, etc., can encode point clouds in a layered structure to generate layer-based bitstreams, such as... ​ As shown.

[0330] The point cloud data receiving method / device according to the embodiment, namely ​ Receiving device 10004 ​ Receiver 10005 ​ Point cloud video decoder 10006 ​ Sending 20002 / Decoding 20003 / Rendering 20004 ​ decoder ​ Receiving equipment ​ equipment ​ , ​ , ​ and​ decoder ​ The receiving methods, etc., can receive layer-based point cloud data and bitstreams, and selectively decode the data, such as... ​ As shown.

[0331] Bitstream and point cloud data, according to the embodiments, can be generated based on slice segmentation at the compilation layer. By slicing the bitstream at the end of the compilation layer of the encoding process, the method / apparatus according to the embodiments can select relevant slices to support scalable transmission or partial decoding.

[0332] ​ -(a) shows a geometric compilation layer structure with eight layers, where each slice corresponds to a layer group. Layer group 13900 contains compilation layers 0 through 4. Layer group 23901 contains compilation layer 5. Layer group 33902 is a group used for compiling layers 6 and 7. When a geometry (or attribute) has a tree structure with eight levels (depths), the bitstream can be configured hierarchically by grouping the data corresponding to one or more levels (depths). Each group can be included in a slice.

[0333] ​ -(b) shows the decoded output when two slices are selected from the three groups. When the decoder selects group 1 and group 2, partial layers of the tree with a depth of 0 to 5 can be selected. That is, partial decoding of the compiled layer can be supported by using slices of the layer group structure without accessing the entire bitstream.

[0334] For the partial decoding process according to the embodiment, the encoder can generate three slices based on the layer group structure. The decoder can perform partial decoding by selecting two slices from the three slices.

[0335] According to the embodiment, the bitstream ( ​ The receiving method / apparatus may include layer-based groups / slices 3903. Each slice may include a header containing signaling information related to the point cloud data (geometric data and / or attribute data) included in the slice. According to embodiments, the receiving method / apparatus may select a subset of slices and decode the point cloud data included in the slice's payload based on the headers included in the slices.

[0336] In addition to the layer group structure, considering the use case of spatial random access, the method / apparatus according to the embodiment can also divide the layer group into several subgroups (subgroups). Subgroups are mutually exclusive, and the set of subgroups can be the same as the layer group. Since the points of each subgroup form the boundary in the spatial domain, the subgroup can be represented by subgroup bounding box information. Based on spatial information, the layer group and subgroup structure can support spatial access. By effectively comparing the region of interest (ROI) with the bounding box information about each slice, spatial random access within a frame or tile can be supported.

[0337] The method / apparatus according to the embodiments can divide layers 1 to 3 (3900, 3901 and 3902) into one or more subgroups.

[0338] although ​ The geometric codec layer structure is shown as an example, but the attribute codec layer structure can be generated in a similar way.

[0339] The method / apparatus according to the embodiments can perform layer-group-based slice segmentation. In fine-grained slices, each slice may contain compiled data from layer groups defined below.

[0340] A layer group is defined as a set of consecutive tree layers whose start and end depths can be any number of tree depths, and whose start depth is less than its end depth. The order of compiled data in a slice fragment can be the same as the order of compiled data in a single slice.

[0341] For example, suppose a geometric compilation layer structure with eight layers is configured, such as ​ As shown in (a). In this example, there are three layer groups, each matching a different slice: layer group 1 for compiling layers 0 through 4, layer group 2 for compiling layer 5, and layer group 3 for compiling layers. When the first two slices are sent or selected, the decoded output will be a portion of layers 0 through 5, as shown in (a). ​ As shown in (b). By using slices in the layer group structure, partial decoding of the compiled layer can be supported even without accessing the entire bitstream.

[0342] A subgroup is a subset of a layer group whose points are adjacent to each other. Subgroups of a layer group are mutually exclusive, and the set of points in a subgroup of a layer group can be the same as the set of points in the layer group. The points in each subgroup are defined within a spatial region, so the boundaries of a subgroup can be described using subgroup bounding box information. Using spatial information, layer group and subgroup structures can support efficient access to the Region of Interest (ROI) by selecting slices that cover the ROI.

[0343] ​ The layer group structure and subgroup structure according to the embodiment are shown.

[0344] ​ The layer-based point cloud data and bitstream shown can be represented as follows: ​ The bounding box shown.

[0345] The subgroup structure and corresponding bounding boxes are shown. Layer group 2 is divided into two subgroups (group 2-1 and group 2-2), and layer group 3 is divided into four subgroups (group 3-1, group 3-2, group 3-3, and group 3-4). Subgroups of layer group 2 and subgroup 3 are contained in different slices. Given slices of layer groups and subgroups with bounding box information, 1) the bounding box of each slice can be compared with the ROI, and 2) slices whose subgroup bounding boxes are associated with the ROI can be selected, and spatial access can be performed. Then, 3) the selected slices are selected. When considering the ROI in region 3-3, slices 1, 3, and 6 are selected as subgroup bounding boxes of layer group 1 and subgroups 2-2 and 3-3 to cover the ROI region. For efficient spatial access, it is assumed that there are no dependencies between subgroups within the same layer group. In real-time streaming or low-latency use cases, selection and decoding operations can be performed as each slice fragment is received to improve time efficiency.

[0346] The method / apparatus according to the embodiments can represent data as layers (which may be referred to as depths or levels) as a layer tree 4000 during geometric and / or attribute encoding. Point cloud data corresponding to layers (depths / levels) can be grouped into layer groups (or groups) 4001 as shown in Figure 39. Each layer group can be further divided (segmented) into subgroups 4002. A bitstream can be generated by configuring each subgroup as a slice. The receiving device according to the embodiments can receive the bitstream, select a specific slice, decode the subgroups included in the slice, and decode the bounding boxes corresponding to the subgroups. For example, when slice 1 is selected, the bounding boxes 4003 corresponding to group 1 can be decoded. Group 1 may be data corresponding to the largest region. When the user wants to view the detailed region of group 1, the method / apparatus according to the embodiments can select slice 3 and / or slice 6, and can partially and hierarchically access the bounding boxes (point cloud data) of groups 2-2 and / or groups 3-3 for the detailed regions included in the region of group 1.

[0347] ​ Multi-resolution, multi-size ROIs are illustrated according to an embodiment.

[0348] Layer slicing provides efficient access to large-scale or dense point cloud data based on scalability and spatial access capabilities. Due to the large number of points and the large data size, rendering or displaying content can take a significant amount of time. As an alternative, the level of detail can be adjusted based on the viewer's interest. For example, structural or global region information is more important than local details when the viewer is far from the scene or object. Conversely, detailed information about the ROI is needed when the viewer is close to a specific region or object. Using adaptive methods, the renderer can effectively provide the viewer with data of sufficient quality. ​Example incremental detail changes at three viewing distance levels are shown, where the viewing distance varies based on the ROI. For example, the viewing distance could be 1) high-level view (coarse detail), 2) mid-level view (medium detail), and 3) low-level view (fine detail).

[0349] ​ A slice of a layer group according to an embodiment is shown.

[0350] The point cloud data transmission method / device according to the embodiment (i.e., ​ 10000 transmission devices ​ Point cloud video encoder 10002 ​ Transmitter 10003 ​ Acquisition 20000 / Encoding 20001 / Transmission 20002 ​ encoder, ​ Transmission equipment ​ equipment ​ , ​ , ​ and ​ encoder, ​ Transmission methods, etc., can be used to slice point cloud data based on layer groups, such as... ​ As shown.

[0351] The point cloud data receiving method / device according to the embodiment, namely ​ Receiving device 10004 ​ Receiver 10005 ​ Point cloud video decoder 10006 ​ Transmission 20002 / Decoding 20003 / Rendering 20004 ​ decoder ​ Receiving equipment ​ equipment ​ , ​ , ​ and ​ decoder ​ The receiving methods, etc., can receive and decode data based on layer group slices, such as... ​ As shown.

[0352] The method / apparatus according to the embodiments can support high-resolution ROI based on the scalability and spatial accessibility of layered slicing.

[0353] refer to ​The encoder can generate bitstream slices of octree layer groups or spatial subgroups of each layer group. Based on the request, slices matching the ROI for each resolution can be selected and sent. Since details other than the requested ROI are not included in the bitstream, the overall bitstream size can be smaller than that in a tile-based approach. The receiver's decoder can combine slices to generate three outputs. For example, 1) a high-level view output can be generated from slice 1 of layer group, 2) a mid-level view output can be generated from slice 1 of layer group and a selected subgroup of layer group 2, and 3) a low-level view with detailed output can be generated from selected subgroups of layer groups 1, 2, and 3. Because the output can be generated progressively, the receiver can provide a viewing experience such as zooming in / out. The resolution can gradually increase from the high-level view to the low-level view.

[0354] According to an embodiment, encoder 60000 can correspond to a geometric encoder and an attribute encoder as a point cloud encoder. The encoder can slice the point cloud data based on layer groups (or subgroups). This layer can be referred to as the depth of the tree, the level of the layer, etc. As shown in part 60000-1, the depth of the octree of the geometry and / or the level of the attribute layer can be divided into layer groups (or subgroups).

[0355] The slice selector 60001, in conjunction with the encoder 60000, can select segmented slices (or sub-slices) to selectively and partially send data, such as layer group 1, to layer group 3.

[0356] Decoder 6002 can decode point cloud data that is selectively and partially transmitted. For example, it can decode a high-level view of layer group 1 (which has a high depth / layer / level or index 0, or is close to the root). Subsequently, a mid-level view can be decoded by individually incrementing the depth / level index on layer group 1 based on layer group 1 and layer group 2. Low-level views can be decoded based on layer groups 1 through 3.

[0357] ​ The nearest neighbor search process according to an embodiment is illustrated.

[0358] ​ The process of finding the nearest neighbor to predict the current point when encoding and decoding a point based on a prediction method using a sending / receiving method / device, according to an embodiment, is illustrated.

[0359] The compilation (encoding / decoding) process according to the embodiment can be configured as follows.

[0360] To encode and decode points, the encoding and decoding method according to the embodiment may perform a nearest neighbor (NN) search to find the nearest neighbors similar to the current point and perform a prediction transformation. The nearest neighbor search may be referred to as an NN search, etc.

[0361] For example, the proposed G-PCC v1-based NN search method also finds prediction candidates from neighboring nodes with higher LoD and nodes compiled with the current LoD. For fast search, the search range is limited by the cubic boundary and the number of points in the Merton compilation order. Considering subgroup boundaries, when neighboring candidates are in the same layer group as the current node, the implementation can use an additional constraint called the intra-layer group search boundary.

[0362] Within a layer group, when a neighbor candidate is in the same layer group as the current node, the search is restricted to neighbor nodes in the same subgroup boundary.

[0363] In other words, neighboring nodes outside the subgroup bounding box cannot be considered as neighbor candidates for the current node, such as... ​ As shown. For example, when the current encoding / compilation target is the first subgroup 2100, the first subgroup 2100 includes nodes (points) belonging to LoD N-1 and LoD N. To predict the current node 2101, a prediction can be generated by referring to the current node's parent node or previous node. Nodes belonging to the second subgroup 2102, which is different from the first subgroup 2100 to which the current node 2101 belongs, are excluded from the candidate neighbors of the current node 2101.

[0364] Additionally, the implementation can use layer group adaptation locations in subgroup estimation and neighbor distance calculation. Therefore, considering the case of missing subgroups, discrepancies between the encoder and decoder can be prevented.

[0365] Layer group adaptation position: Indicates the number of LoDs of the compiled descendant subgroups to the right, and then the number of LoDs of the compiled descendant subgroups and skipped descendant subgroups to the left.

[0366] ​ The NN search process according to an embodiment is illustrated.

[0367] ​ It shows something similar to ​ The NN search process in [the context of the NN search].

[0368] When the current LoD is at the coarsest level of the current layer group, neighbor candidates can be in different layer groups. In this case, when a node is at the upper subgroup boundary, it can search for neighbors outside the subgroup boundary.

[0369] Inter-layer group search boundary: such as ​ As shown, when a neighbor candidate is in a higher-level group, the use of neighbor nodes at the boundary of the upper subgroup can be restricted.

[0370] ​ The NN search process according to an embodiment is illustrated.

[0371] ​ It shows something similar to ​ and ​ The NN search process in [the context of the NN search].

[0372] When using layer group adaptive positioning, the compiled positioning of nodes from the parent node can be handled differently when a node is within the bounding box of a subgroup of the current node.

[0373] The layer group adaptation position according to the embodiment can be configured as follows.

[0374] For nodes within the subgroup bounding box: the compilation position is shifted to the right by the number of LoDs of the current subgroup's compiled subgroup, then shifted to the left by the number of LoDs of the current subgroup's compiled subgroup and the number of skipped subgroups.

[0375] Nodes outside the subgroup bounding box: The number of LoDs in the compiled subgroup of the upper subgroup is shifted to the right by the number of LoDs in the compiled subgroup of the upper subgroup, and then shifted to the left by the number of LoDs in the compiled subgroup of the upper subgroup and the number of skipped subgroups.

[0376] In other words, when a neighbor candidate is within the subgroup bounding box, the details of its geometric location are taken into account at the highest level of the subgroup. On the other hand, when a neighbor candidate is outside the subgroup bounding box, the details of its geometric location are considered up to the highest level of the parent-child group.

[0377] refer to ​ In this embodiment, to generate a predicted value for the current node of the first subgroup 2100, reference can be made to nodes with higher LoD levels included in the first subgroup 2100 to which the current node belongs and / or previous nodes with the current LoD level in the first subgroup 2100 to which the current node belongs. Nodes in the second subgroup 2102 that are not in the first subgroup are excluded from the candidates used to predict the current node.

[0378] refer to ​ In this embodiment, a second subgroup to which the current node does not belong cannot be a neighbor candidate for predicting the current node. The first and second subgroups can be demarcated by a boundary. However, a parent-child group including the first and second subgroups can be a neighbor candidate for predicting the current node. Nodes belonging to a parent-child group with a higher LoD level than the current node in the first subgroup can be neighbor candidates for the current node.

[0379] refer to ​More specifically, the geometric location (geometric information, geometric shape, etc.) used to predict the current node can be considered. The geometric location values ​​of nodes belonging to the first subgroup 2100, which includes the current node 2101, and located at a lower LoD level than the current node, can be candidates for the geometric location of the current node's neighbors. That is, geometric locations can be selected from the range of nodes included in the subgroup bounding box. Nodes belonging to a second subgroup different from the first subgroup to which the current node belongs cannot be candidate neighbors of the current node. However, when the second subgroup is a sibling subgroup of a parent-child group that is the same as the first subgroup to which the current node belongs, nodes belonging to the parent-child group can be candidates for the current node's neighbors. When selecting nodes belonging to the parent-child group, the geometric location of nodes in the parent-child group does not need to be defined in detail up to the range of the subgroup, and only the LoD level to which the parent-child group belongs can be determined.

[0380] For example, the NN search according to the embodiment can be represented in pseudocode as follows. In the following text, the NN search method is applied to both the encoder and the decoder.

[0381] inline void

[0382] buildPredictorsFast(

[0383] const AttributeParameterSet& aps,

[0384] const AttributeBrickHeader& abh,

[0385] const PCCPointSet3& pointCloud,

[0386] int32_t minGeomNodeSizeLog2,

[0387] int geom_num_points_minus1,

[0388] std::vector <pccpredictor>& predictors,

[0389] std::vector<uint32_t>& numberOfPointsPerLevelOfDetail,

[0390] std::vector<uint32_t>& indexes,

[0391] LayerGroupSlicingParams& layerGroupParams)

[0392] {

[0393] const int32_t pointCount = int32_t(pointCloud.getPointCount());

[0394] assert(pointCount);

[0395] std::vector <mortoncodewithindex>packedVoxel;

[0396] computeMortonCodesUnsorted(pointCloud, aps.lodNeighBias,packedVoxel);

[0397] if (!aps.canonical_point_order_flag)

[0398] std::sort(packedVoxel.begin(), packedVoxel.end());

[0399] std::vector<uint32_t> retained, input, pointIndexToPredictorIndex;

[0400] pointIndexToPredictorIndex.resize(pointCount);

[0401] retained.reserve(pointCount);

[0402] std::vector<uint32_t> pointIdxToPackedVoxelIdx;

[0403] std::vector<uint32_t> dcmNodesList;

[0404] std::vector<std::vector <int>> pointIdxToSubgroupIdx;

[0405] int treeLvlGap = 0;

[0406] if (layerGroupParams.layerGroupEnabledFlag) {

[0407] pointIdxToPackedVoxelIdx.resize(pointCount);

[0408] for (uint32_t i = 0; i < pointCount; i++) {

[0409] auto pointIdx = packedVoxel[i].index;

[0410] pointIdxToPackedVoxelIdx[pointIdx] = i;

[0411] }

[0412] int layerIdx = 0;

[0413] int dcmNodesCount = 0;

[0414] int maxNumLayer = 0;

[0415] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++)

[0416] maxNumLayer += layerGroupParams.numLayersPerLayerGroup[i];

[0417] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[0418] for (int j = 0; j < layerGroupParams.numLayersPerLayerGroup[i]; j++,layerIdx++) {

[0419] for (int sbgrIdx = 0; sbgrIdx <= layerGroupParams.numSubgroupsMinus1[i]; sbgrIdx++) {

[0420] if (layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx].size() &&layerIdx < maxNumLayer - 1) {

[0421] for (int m = 0; m < layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx].size(); m++) {

[0422] auto pointIdx = layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx][m];

[0423] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointIdx];

[0424] dcmNodesList.push_back(packedVoxelIndex);

[0425] }

[0426] }

[0427] }

[0428] }

[0429] }

[0430] if (dcmNodesList.size()) {

[0431] std::sort(dcmNodesList.begin(), dcmNodesList.end());

[0432] input.resize(pointCount - dcmNodesList.size());

[0433] int dcmIdx = 0;

[0434] int nonIdcmIdx = 0;

[0435] for (uint32_t i = 0; i < pointCount; ++i) {

[0436] uint32_t packedVoxelIndex;

[0437] if (dcmIdx < dcmNodesList.size()) {

[0438] packedVoxelIndex = dcmNodesList[dcmIdx];

[0439] if (packedVoxel[packedVoxelIndex].mortonCode != packedVoxel[i].mortonCode)

[0440] dcmIdx++;

[0441] else if(nonIdcmIdx < input.size())

[0442] input[nonIdcmIdx++] = i;

[0443] }

[0444] else if (nonIdcmIdx < input.size())

[0445] input[nonIdcmIdx++] = i;

[0446] }

[0447] }

[0448] else {

[0449] input.resize(pointCount);

[0450] for (uint32_t i = 0; i < pointCount; ++i)

[0451] input[i] = i;

[0452] }

[0453] treeLvlGap = layerGroupParams.rootNodeSizeLog2.max() - layerGroupParams.rootNodeSizeLog2_coded.max();

[0454] std::vector <int>accNumLayers;

[0455] accNumLayers.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[0456] accNumLayers[0] = layerGroupParams.numLayersPerLayerGroup[0];

[0457] if (layerGroupParams.numLayerGroupsMinus1 > 0)

[0458] for (int i = 1; i <= layerGroupParams.numLayerGroupsMinus1; i++)

[0459] accNumLayers[i] = accNumLayers[i - 1] + layerGroupParams.numLayersPerLayerGroup[i];

[0460] (Here, a mapping function between the point index and the subgroup isused to improve speed.)

[0461] pointIdxToSubgroupIdx.resize(pointCount);

[0462] for (int pointIdx = 0; pointIdx < pointCount; pointIdx++) {pointIdxToSubgroupIdx[pointIdx].resize(layerGroupParams.numLayerGroupsMinus1+ 1);

[0463] for (int lyrGrpIdx = 1; lyrGrpIdx <= layerGroupParams.numLayerGroupsMinus1; lyrGrpIdx++) {

[0464] int shift = accNumLayers[layerGroupParams.numLayerGroupsMinus1] -accNumLayers[lyrGrpIdx];

[0465] auto pos = (pointCloud[pointIdx] >> shift) << (shift + treeLvlGap);

[0466] for (int subGrpIdx = 0; subGrpIdx <= layerGroupParams.numSubgroupsMinus1[lyrGrpIdx]; subGrpIdx++) {

[0467] if (layerGroupParams.sliceSelectionIndicationFlag[lyrGrpIdx][subGrpIdx]) {

[0468] auto bbox_min = layerGroupParams.subgrpBboxOrigin[lyrGrpIdx][subGrpIdx];

[0469] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[lyrGrpIdx][subGrpIdx];

[0470] if (pos.x() >= bbox_min[0] && pos.x() < bbox_max[0]

[0471] && pos.y() >= bbox_min[1] && pos.y() < bbox_max[1]

[0472] && pos.z() >= bbox_min[2] && pos.z() < bbox_max[2]) {pointIdxToSubgroupIdx[pointIdx][lyrGrpIdx] = subGrpIdx;

[0473] break;

[0474] }

[0475] }

[0476] }

[0477] }

[0478] }

[0479] }

[0480] Here, `pointIdxToSubgroupIdx` indicates which subgroup a point belongs to. By generating this array, the encoder and decoder can quickly identify the inclusion (or mapping) relationship between points and subgroups. According to the embodiment, the method / apparatus can generate an index of the subgroup to which the current point belongs while generating the subgroup, and can also generate array information indicating the mapping relationship between points and subgroup indices. During the encoding / decoding process, the method / apparatus can quickly identify the subgroup to which the current point belongs based on the array information.

[0481] else {

[0482] input.resize(pointCount);

[0483] for (uint32_t i = 0; i < pointCount; ++i)

[0484] input[i] = i;

[0485] }

[0486] }

[0487] / prepare output buffers

[0488] predictors.resize(pointCount);

[0489] numberOfPointsPerLevelOfDetail.resize(0);

[0490] indexes.resize(0);

[0491] indexes.reserve(pointCount);

[0492] numberOfPointsPerLevelOfDetail.reserve(21);

[0493] numberOfPointsPerLevelOfDetail.push_back(pointCount);

[0494] bool concatenateLayers = aps.scalable_lifting_enabled_flag;

[0495] if (layerGroupParams.layerGroupEnabledFlag)

[0496] concatenateLayers = false;

[0497] std::vector<uint32_t> indexesOfSubsample;

[0498] if (concatenateLayers)

[0499] indexesOfSubsample.reserve(pointCount);

[0500] std::vector<Box3<int32_t>> bBoxes;

[0501] const int32_t log2CubeSize = 7;

[0502] MortonIndexMap3d atlas;

[0503] atlas.resize(log2CubeSize);

[0504] atlas.init();

[0505] auto maxNumDetailLevels = aps.maxNumDetailLevels();

[0506] int32_t predIndex = int32_t(pointCount);

[0507] int maxNumLayers = 0;

[0508] std::vector <int>predIdxToSubgroupMap;

[0509] std::vector <int>predIdxToLayerGroupMap;

[0510] if (layerGroupParams.layerGroupEnabledFlag) {

[0511] predIdxToLayerGroupMap.resize(pointCount, -1);

[0512] predIdxToSubgroupMap.resize(pointCount, -1);

[0513] treeLvlGap = layerGroupParams.rootNodeSizeLog2.max() - layerGroupParams.rootNodeSizeLog2_coded.max();

[0514] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++)

[0515] maxNumLayers += layerGroupParams.numLayersPerLayerGroup[m];

[0516] maxNumDetailLevels = maxNumLayers + 1;

[0517] }

[0518] int maxLoD = 0;

[0519] for (auto lodIndex = minGeomNodeSizeLog2;

[0520] !input.empty() && lodIndex < maxNumDetailLevels; ++lodIndex) {

[0521] const int32_t startIndex = indexes.size();

[0522] if (lodIndex == maxNumDetailLevels - 1) {

[0523] for (const auto index : input) {

[0524] indexes.push_back(index);

[0525] }

[0526] }

[0527] else {

[0528] subsample(

[0529] aps, abh, pointCloud, packedVoxel, input, lodIndex, retained,indexes,

[0530] atlas, layerGroupParams );

[0531] }

[0532] const int32_t endIndex = indexes.size();

[0533] int shiftLayerGroup = 0;

[0534] int shiftPrtLayerGroup = 0;

[0535] int curLayerGroup = 0;

[0536] int curSubgroup = 0;

[0537] int prtLayerGroup = 0;

[0538] int prtSubgroup = 0;

[0539] std::vector<std::vector<uint32_t>> retained_subgroup, indexes_subgroup;

[0540] if (layerGroupParams.layerGroupEnabledFlag) {

[0541] int layerIdx = maxNumLayers - lodIndex;

[0542] int accNumLayers = 0;

[0543] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++) {

[0544] accNumLayers += layerGroupParams.numLayersPerLayerGroup[m];

[0545] if (accNumLayers >= layerIdx) {

[0546] curLayerGroup = m;

[0547] break;

[0548] }

[0549] }

[0550] if (accNumLayers - layerGroupParams.numLayersPerLayerGroup[curLayerGroup] + 1 == layerIdx && curLayerGroup > 0)

[0551] prtLayerGroup = curLayerGroup - 1;

[0552] else

[0553] prtLayerGroup = curLayerGroup;

[0554] retained_subgroup.clear();

[0555] indexes_subgroup.clear();

[0556] retained_subgroup.resize(layerGroupParams.numSubgroupsMinus1[prtLayerGroup] + 1);

[0557] indexes_subgroup.resize(layerGroupParams.numSubgroupsMinus1[curLayerGroup] + 1);

[0558] std::vector<std::vector<uint32_t>> dcmNodeList_ParentSubgroup;

[0559] int layerIdx_minus1 = layerIdx - 1;

[0560] if (dcmNodesList.size()) {

[0561] dcmNodeList_ParentSubgroup.resize(layerGroupParams.numSubgroupsMinus1[prtLayerGroup] + 1);

[0562] int lyrGrpIdx = 0;

[0563] if (layerIdx_minus1 > 0) {

[0564] for (int curLodIndex = 0; curLodIndex < layerIdx_minus1; curLodIndex++) {

[0565] lyrGrpIdx = 0;

[0566] int accNumLayers = 0;

[0567] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++) {

[0568] accNumLayers += layerGroupParams.numLayersPerLayerGroup[m];

[0569] if (accNumLayers > curLodIndex) {

[0570] lyrGrpIdx = m;

[0571] break;

[0572] }

[0573] }

[0574] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[lyrGrpIdx];m++) {

[0575] if (layerGroupParams.sliceSelectionIndicationFlag[lyrGrpIdx][m]) {

[0576] for (int k = 0; k < layerGroupParams.dcmNodesIdx[curLodIndex][m].size(); k++) {

[0577] auto pointCloudIndex = layerGroupParams.dcmNodesIdx[curLodIndex][m][k];

[0578] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointCloudIndex];

[0579] int subgrpIdx = pointIdxToSubgroupIdx[pointCloudIndex][prtLayerGroup];

[0580] dcmNodeList_ParentSubgroup[subgrpIdx].push_back(packedVoxelIndex);

[0581] }

[0582] }

[0583] }

[0584] }

[0585] }

[0586] }

[0587] for (int i = 0; i < retained.size(); i++) {

[0588] prtSubgroup = pointIdxToSubgroupIdx[packedVoxel[retained[i]].index][prtLayerGroup];

[0589] retained_subgroup[prtSubgroup].push_back(retained[i]);

[0590] }

[0591] if (dcmNodesList.size()) {

[0592] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; i++) {

[0593] if (layerGroupParams.sliceSelectionIndicationFlag[prtLayerGroup][i]

[0594] && dcmNodeList_ParentSubgroup[i].size()) {

[0595] for (int j = 0; j < dcmNodeList_ParentSubgroup[i].size(); j++)retained_subgroup[i].push_back(dcmNodeList_ParentSubgroup[i][j]);

[0596] std::sort(retained_subgroup[i].begin(), retained_subgroup[i].end());

[0597] }

[0598] }

[0599] if (layerIdx >= 2) {

[0600] auto prev = retained.size();

[0601] int parentLayerIndex = layerIdx - 2;

[0602] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; i++) {

[0603] if (layerGroupParams.sliceSelectionIndicationFlag[prtLayerGroup][i]){

[0604] for (int k = 0; k < layerGroupParams.dcmNodesIdx[parentLayerIndex][i].size(); k++) {

[0605] auto pointCloudIndex = layerGroupParams.dcmNodesIdx[parentLayerIndex][i][k];

[0606] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointCloudIndex];

[0607] retained.push_back(packedVoxelIndex);

[0608] }

[0609] }

[0610] }

[0611] std::sort(retained.begin(), retained.end());

[0612] }

[0613] }

[0614] for (int i = startIndex; i < endIndex; i++) {

[0615] curSubgroup = pointIdxToSubgroupIdx[packedVoxel[indexes[i]].index][curLayerGroup];

[0616] if (layerGroupParams.sliceSelectionIndicationFlag[curLayerGroup][curSubgroup])

[0617] indexes_subgroup[curSubgroup].push_back(indexes[i]);

[0618] }

[0619] }

[0620] std::vector<point_t> biasedPos_indexes, biasedPos_retained;

[0621] if (concatenateLayers && !layerGroupParams.layerGroupEnabledFlag) {

[0622] indexesOfSubsample.resize(endIndex);

[0623] if (startIndex != endIndex) {

[0624] for (int32_t i = startIndex; i < endIndex; i++)

[0625] indexesOfSubsample[i] = indexes[i];

[0626] int32_t numOfPointInSkipped = geom_num_points_minus1 + 1 -pointCount;

[0627] if (endIndex - startIndex <= startIndex + numOfPointInSkipped) {

[0628] concatenateLayers = false;

[0629] } else {

[0630] for (int32_t i = 0; i < startIndex; i++)

[0631] indexes[i] = indexesOfSubsample[i];

[0632] / reset predIndex

[0633] predIndex = pointCount;

[0634] for (int lod = 0; lod < lodIndex - minGeomNodeSizeLog2; lod++) {

[0635] int divided_startIndex =

[0636] pointCount - numberOfPointsPerLevelOfDetail[lod];

[0637] int divided_endIndex =

[0638] pointCount - numberOfPointsPerLevelOfDetail[lod + 1];

[0639] computeNearestNeighbors(

[0640] aps, abh, packedVoxel, retained, divided_startIndex,

[0641] divided_endIndex, lod + minGeomNodeSizeLog2, indexes, predictors,

[0642] pointIndexToPredictorIndex, predIndex, atlas, layerGroupParams,

[0643] curLayerGroup, curSubgroup, biasedPos_indexes, biasedPos_retained);

[0644] }

[0645] }

[0646] }

[0647] }

[0648] int numIndexedNodes = 0;

[0649] if(layerGroupParams.layerGroupEnabledFlag){

[0650] indexes.resize(startIndex);

[0651] int count = 0;

[0652] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[curLayerGroup]; i++) {

[0653] if (layerGroupParams.sliceSelectionIndicationFlag[curLayerGroup][i]){

[0654] if (curLayerGroup == prtLayerGroup)

[0655] prtSubgroup = i;

[0656] else {

[0657] auto cur_bbox_min = layerGroupParams.subgrpBboxOrigin[curLayerGroup][i];

[0658] auto cur_bbox_max = cur_bbox_min + layerGroupParams.subgrpBboxSize[curLayerGroup][i];

[0659] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; m++) {

[0660] auto bbox_min = layerGroupParams.subgrpBboxOrigin[prtLayerGroup][m];

[0661] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[prtLayerGroup][m];

[0662] if (cur_bbox_min.x() >= bbox_min[0] && cur_bbox_max.x() <= bbox_max[0]

[0663] && cur_bbox_min.y() >= bbox_min[1] && cur_bbox_max.y() <= bbox_max[1]

[0664] && cur_bbox_min.z() >= bbox_min[2] && cur_bbox_max.z() <= bbox_max[2]) {

[0665] prtSubgroup = m;

[0666] break;

[0667] }

[0668] }

[0669] }

[0670] int startIndex_subgroup = 0;

[0671] int endIndex_subgroup = indexes_subgroup[i].size();

[0672] int curSubgroup = i;

[0673] biasedPos_indexes.clear(); biasedPos_indexes.reserve(indexes_subgroup[curSubgroup].size());

[0674] for (int k = 0; k < indexes_subgroup[curSubgroup].size(); k++) {

[0675] auto pos = pointCloud[packedVoxel[indexes_subgroup[curSubgroup][k]].index];

[0676] auto point = pos;

[0677] biasedPos_indexes.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[0678] biasedPos_retained.clear(); biasedPos_retained.reserve(retained_subgroup[prtSubgroup].size());

[0679] auto cur_bbox_min = layerGroupParams.subgrpBboxOrigin[curLayerGroup][i];

[0680] auto cur_bbox_max = cur_bbox_min + layerGroupParams.subgrpBboxSize[curLayerGroup][i];

[0681] for (int k = 0; k < retained_subgroup[prtSubgroup].size(); k++) {

[0682] auto pos = pointCloud[packedVoxel[retained_subgroup[prtSubgroup][k]].index];

[0683] auto point = pos;

[0684] if (prtLayerGroup != curLayerGroup) {

[0685] if (pos.x() >= cur_bbox_min.x() && pos.x() < cur_bbox_max.x()

[0686] && pos.y() >= cur_bbox_min.y() && pos.y() < cur_bbox_max.y()

[0687] && pos.z() >= cur_bbox_min.z() && pos.z() < cur_bbox_max.z())biasedPos_retained.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[0688] else biasedPos_retained.push_back(times((point >> shiftPrtLayerGroup)<< (shiftPrtLayerGroup + treeLvlGap), aps.lodNeighBias));

[0689] }

[0690] else

[0691] biasedPos_retained.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[0692] }

[0693] computeNearestNeighbors(

[0694] aps, abh, packedVoxel, retained_subgroup[prtSubgroup], startIndex_subgroup, endIndex_subgroup, lodIndex, indexes_subgroup[curSubgroup],

[0695] predictors, pointIndexToPredictorIndex, predIndex, atlas,layerGroupParams, curLayerGroup, curSubgroup, biasedPos_indexes, biasedPos_retained);

[0696] for (int m = startIndex_subgroup; m < endIndex_subgroup; m++)

[0697] {

[0698] int cur_subgroup = i;

[0699] predIdxToLayerGroupMap.push_back(curLayerGroup);

[0700] predIdxToSubgroupMap.push_back(cur_subgroup); indexes.push_back(indexes_subgroup[cur_subgroup][m]);

[0701] }

[0702] numIndexedNodes += indexes_subgroup[i].size(); layerGroupParams.numberOfPointsPerLodPerSubgroups.push_back(indexes_subgroup[i].size());

[0703] }

[0704] else layerGroupParams.numberOfPointsPerLodPerSubgroups.push_back(0);

[0705] }

[0706] }

[0707] else

[0708] computeNearestNeighbors(

[0709] aps, abh, packedVoxel, retained, startIndex, endIndex, lodIndex,indexes,

[0710] predictors, pointIndexToPredictorIndex, predIndex, atlas,layerGroupParams,

[0711] curLayerGroup, curSubgroup, biasedPos_indexes, biasedPos_retained);

[0712] if (layerGroupParams.layerGroupEnabledFlag) {

[0713] if(pointCount > indexes.size())

[0714] numberOfPointsPerLevelOfDetail.push_back(pointCount - indexes.size());

[0715] }

[0716] else {

[0717] if (!retained.empty()) {

[0718] numberOfPointsPerLevelOfDetail.push_back(retained.size());

[0719] }

[0720] }

[0721] input.resize(0);

[0722] std::swap(retained, input);

[0723] }

[0724] std::reverse(indexes.begin(), indexes.end());

[0725] updatePredictors(pointIndexToPredictorIndex, predictors);

[0726] std::reverse(

[0727] numberOfPointsPerLevelOfDetail.begin(),

[0728] numberOfPointsPerLevelOfDetail.end());

[0729] Here, information about whether the neighbor is in the same subgroupis updated.

[0730] layerGroupParams.numRefNodesInTheSameSubgroup.resize(pointCount, 0);

[0731] if (layerGroupParams.layerGroupEnabledFlag) {

[0732] std::reverse(predIdxToLayerGroupMap.begin(),predIdxToLayerGroupMap.end());

[0733] std::reverse(predIdxToSubgroupMap.begin(), predIdxToSubgroupMap.end());

[0734] for (int i = 0; i < pointCount; i++) {

[0735] auto& predictor = predictors[i];

[0736] int curLayerGroup = predIdxToLayerGroupMap[i];

[0737] int curSubgroup = predIdxToSubgroupMap[i];

[0738] for (int k = 0; k < predictor.neighborCount; k++) {

[0739] int idx = predictor.neighbors[k].predictorIndex;

[0740] int neighborLayerGroup = predIdxToLayerGroupMap[idx];

[0741] int neighborSubgroup = predIdxToSubgroupMap[idx];

[0742] if(neighborLayerGroup == curLayerGroup && neighborSubgroup ==curSubgroup) {

[0743] predictor.neighbors[k].inTheSameSubgroupFlag = true;

[0744] layerGroupParams.numRefNodesInTheSameSubgroup[idx]++;

[0745] }

[0746] else

[0747] predictor.neighbors[k].inTheSameSubgroupFlag = false;

[0748] }

[0749] }

[0750] }

[0751] std::reverse(

[0752] layerGroupParams.numberOfPointsPerLodPerSubgroups.begin(),

[0753] layerGroupParams.numberOfPointsPerLodPerSubgroups.end());

[0754] }

[0755] inline void

[0756] buildPredictorsFast(

[0757] const AttributeParameterSet& aps,

[0758] const AttributeBrickHeader& abh,

[0759] const PCCPointSet3& pointCloud,

[0760] int32_t minGeomNodeSizeLog2,

[0761] int geom_num_points_minus1,

[0762] std::vector <pccpredictor>& predictors,

[0763] std::vector<uint32_t>& numberOfPointsPerLevelOfDetail,

[0764] std::vector<uint32_t>& indexes,

[0765] LayerGroupSlicingParams& layerGroupParams){

[0766] const int32_t pointCount = int32_t(pointCloud.getPointCount());

[0767] assert(pointCount);

[0768] std::vector <mortoncodewithindex>packedVoxel;

[0769] computeMortonCodesUnsorted(pointCloud, aps.lodNeighBias,packedVoxel);

[0770] if (!aps.canonical_point_order_flag)

[0771] std::sort(packedVoxel.begin(), packedVoxel.end());

[0772] std::vector<uint32_t> retained, input, pointIndexToPredictorIndex;

[0773] pointIndexToPredictorIndex.resize(pointCount);

[0774] retained.reserve(pointCount);

[0775] std::vector<uint32_t> pointIdxToPackedVoxelIdx;

[0776] std::vector<uint32_t> dcmNodesList;

[0777] if (layerGroupParams.layerGroupEnabledFlag) {

[0778] pointIdxToPackedVoxelIdx.resize(pointCount);

[0779] for (uint32_t i = 0; i < pointCount; i++) {

[0780] auto pointIdx = packedVoxel[i].index;

[0781] pointIdxToPackedVoxelIdx[pointIdx] = i;

[0782] }

[0783] int layerIdx = 0;

[0784] int dcmNodesCount = 0;

[0785] int maxNumLayer = 0;

[0786] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++)

[0787] maxNumLayer += layerGroupParams.numLayersPerLayerGroup[i];

[0788] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[0789] for (int j = 0; j < layerGroupParams.numLayersPerLayerGroup[i]; j++,layerIdx++) {

[0790] for (int sbgrIdx = 0; sbgrIdx <= layerGroupParams.numSubgroupsMinus1[i]; sbgrIdx++) {

[0791] if (layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx].size() &&layerIdx < maxNumLayer - 1) {

[0792] for (int m = 0; m < layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx].size(); m++) {

[0793] auto pointIdx = layerGroupParams.dcmNodesIdx[layerIdx][sbgrIdx][m];

[0794] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointIdx];dcmNodesList.push_back(packedVoxelIndex);

[0795] }

[0796] }

[0797] }

[0798] }

[0799] }

[0800] if (dcmNodesList.size()) {

[0801] input.resize(pointCount - dcmNodesList.size());

[0802] int dcmIdx = 0;

[0803] int nonIdcmIdx = 0;

[0804] for (uint32_t i = 0; i < pointCount; ++i) {

[0805] uint32_t packedVoxelIndex;

[0806] if (dcmIdx < dcmNodesList.size()) {

[0807] packedVoxelIndex = dcmNodesList[dcmIdx];

[0808] if (packedVoxel[packedVoxelIndex].mortonCode != packedVoxel[i].mortonCode)

[0809] dcmIdx++;

[0810] else if(nonIdcmIdx < input.size())

[0811] input[nonIdcmIdx++] = i;

[0812] }

[0813] else if (nonIdcmIdx < input.size())

[0814] input[nonIdcmIdx++] = i;

[0815] }

[0816] }

[0817] else {

[0818] input.resize(pointCount);

[0819] for (uint32_t i = 0; i < pointCount; ++i)

[0820] input[i] = i;

[0821] }

[0822] }

[0823] else {

[0824] input.resize(pointCount);

[0825] for (uint32_t i = 0; i < pointCount; ++i) {

[0826] input[i] = i;

[0827] }

[0828] }

[0829] Set up buffers for output values (prepare output buffers).

[0830] predictors.resize(pointCount);

[0831] numberOfPointsPerLevelOfDetail.resize(0);

[0832] indexes.resize(0);

[0833] indexes.reserve(pointCount);

[0834] numberOfPointsPerLevelOfDetail.reserve(21);

[0835] numberOfPointsPerLevelOfDetail.push_back(pointCount);

[0836] bool concatenateLayers = aps.scalable_lifting_enabled_flag;

[0837] if (layerGroupParams.layerGroupEnabledFlag)

[0838] concatenateLayers = false;

[0839] std::vector<uint32_t> indexesOfSubsample;

[0840] if (concatenateLayers)

[0841] indexesOfSubsample.reserve(pointCount);

[0842] std::vector<Box3<int32_t>> bBoxes;

[0843] const int32_t log2CubeSize = 7;

[0844] MortonIndexMap3d atlas;

[0845] atlas.resize(log2CubeSize);

[0846] atlas.init();

[0847] auto maxNumDetailLevels = aps.maxNumDetailLevels();

[0848] int32_t predIndex = int32_t(pointCount);

[0849] int treeLvlGap = 0;

[0850] int maxNumLayers = 0;

[0851] std::vector <int>predIdxToSubgroupMap;

[0852] std::vector <int>predIdxToLayerGroupMap;

[0853] if (layerGroupParams.layerGroupEnabledFlag) {

[0854] predIdxToLayerGroupMap.resize(pointCount, -1);

[0855] predIdxToSubgroupMap.resize(pointCount, -1);

[0856] treeLvlGap = layerGroupParams.rootNodeSizeLog2.max() - layerGroupParams.rootNodeSizeLog2_coded.max();

[0857] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++)

[0858] maxNumLayers += layerGroupParams.numLayersPerLayerGroup[m];

[0859] maxNumDetailLevels = maxNumLayers + 1;

[0860] }

[0861] int maxLoD = 0;

[0862] for (auto lodIndex = minGeomNodeSizeLog2;

[0863] !input.empty() && lodIndex < maxNumDetailLevels; ++lodIndex) {

[0864] const int32_t startIndex = indexes.size();

[0865] if (lodIndex == maxNumDetailLevels - 1) {

[0866] for (const auto index : input) {

[0867] indexes.push_back(index);

[0868] }

[0869] } else {

[0870] subsample(

[0871] aps, abh, pointCloud, packedVoxel, input, lodIndex, retained,indexes,

[0872] atlas, layerGroupParams );

[0873] }

[0874] const int32_t endIndex = indexes.size();

[0875] int shiftLayerGroup = 0;

[0876] int shiftPrtLayerGroup = 0;

[0877] int curLayerGroup = 0;

[0878] int curSubgroup = 0;

[0879] int prtLayerGroup = 0;

[0880] int prtSubgroup = 0;

[0881] std::vector<std::vector<uint32_t>> retained_subgroup, indexes_subgroup;

[0882] if (layerGroupParams.layerGroupEnabledFlag) {

[0883] int layerIdx = maxNumLayers - lodIndex;

[0884] int accNumLayers = 0;

[0885] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++) {

[0886] accNumLayers += layerGroupParams.numLayersPerLayerGroup[m];

[0887] if (accNumLayers >= layerIdx) {

[0888] curLayerGroup = m;

[0889] break;

[0890] }

[0891] }

[0892] if (accNumLayers - layerGroupParams.numLayersPerLayerGroup[curLayerGroup] + 1 == layerIdx && curLayerGroup > 0)

[0893] prtLayerGroup = curLayerGroup - 1;

[0894] else

[0895] prtLayerGroup = curLayerGroup;

[0896] int maxNumLayers_curLayerGroup = 0;

[0897] int maxNumLayers_prtLayerGroup = 0;

[0898] int maxNumLayers_total = 0;

[0899] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[0900] maxNumLayers_total += layerGroupParams.numLayersPerLayerGroup[i];

[0901] if (i <= curLayerGroup)

[0902] maxNumLayers_curLayerGroup += layerGroupParams.numLayersPerLayerGroup[i];

[0903] if (i <= prtLayerGroup)

[0904] maxNumLayers_prtLayerGroup += layerGroupParams.numLayersPerLayerGroup[i];

[0905] }

[0906] shiftLayerGroup = maxNumLayers_total - maxNumLayers_curLayerGroup;

[0907] shiftPrtLayerGroup = maxNumLayers_total - maxNumLayers_prtLayerGroup;

[0908] retained_subgroup.clear();

[0909] indexes_subgroup.clear();

[0910] retained_subgroup.resize(layerGroupParams.numSubgroupsMinus1[prtLayerGroup] + 1);

[0911] indexes_subgroup.resize(layerGroupParams.numSubgroupsMinus1[curLayerGroup] + 1);

[0912] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[curLayerGroup]; i++)

[0913] std::vector<std::vector<uint32_t>> dcmNodeList_ParentSubgroup;

[0914] int layerIdx_minus1 = layerIdx - 1;

[0915] if (dcmNodesList.size()) {dcmNodeList_ParentSubgroup.resize(layerGroupParams.numSubgroupsMinus1[prtLayerGroup] + 1);

[0916] int lyrGrpIdx = 0;

[0917] if (layerIdx_minus1 > 0) {

[0918] for (int curLodIndex = 0; curLodIndex < layerIdx_minus1; curLodIndex++) {

[0919] lyrGrpIdx = 0;

[0920] int accNumLayers = 0;

[0921] for (int m = 0; m <= layerGroupParams.numLayerGroupsMinus1; m++) {

[0922] accNumLayers += layerGroupParams.numLayersPerLayerGroup[m];if(accNumLayers > curLodIndex) {

[0923] lyrGrpIdx = m;

[0924] break;

[0925] }

[0926] }

[0927] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[lyrGrpIdx];m++) {

[0928] if (layerGroupParams.sliceSelectionIndicationFlag[lyrGrpIdx][m]) {

[0929] auto bbox_min = layerGroupParams.subgrpBboxOrigin[lyrGrpIdx][m];

[0930] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[lyrGrpIdx][m];

[0931] for (int subgrpIdx = 0; subgrpIdx <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; subgrpIdx++) {

[0932] auto retained_subgrp_bbox_min = layerGroupParams.subgrpBboxOrigin[prtLayerGroup][subgrpIdx];

[0933] auto retained_subgrp_bbox_max = retained_subgrp_bbox_min +layerGroupParams.subgrpBboxSize[prtLayerGroup][subgrpIdx];

[0934] if (retained_subgrp_bbox_min[0] >= bbox_min[0] && retained_subgrp_bbox_max[0] <= bbox_max[0]

[0935] && retained_subgrp_bbox_min[1] >= bbox_min[1] && retained_subgrp_bbox_max[1] <= bbox_max[1]

[0936] && retained_subgrp_bbox_min[2] >= bbox_min[2] && retained_subgrp_bbox_max[2] <= bbox_max[2]) {for (int k = 0; k < layerGroupParams.dcmNodesIdx[curLodIndex][m].size(); k++) {

[0937] auto pointCloudIndex = layerGroupParams.dcmNodesIdx[curLodIndex][m][k];

[0938] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointCloudIndex];dcmNodeList_ParentSubgroup[subgrpIdx].push_back(packedVoxelIndex);

[0939] }

[0940] }

[0941] }

[0942] }

[0943] }

[0944] }

[0945] for (int subgrpIdx = 0; subgrpIdx <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; subgrpIdx++)

[0946] }

[0947] }

[0948] for (int i = 0; i < retained.size(); i++) {

[0949] auto pos = (pointCloud[packedVoxel[retained[i]].index] >>shiftLayerGroup) << (shiftLayerGroup + treeLvlGap);

[0950] prtSubgroup = 0;

[0951] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; m++) {

[0952] if (layerGroupParams.sliceSelectionIndicationFlag[prtLayerGroup][m]){

[0953] auto bbox_min = layerGroupParams.subgrpBboxOrigin[prtLayerGroup][m];

[0954] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[prtLayerGroup][m];

[0955] if (pos.x() >= bbox_min[0] && pos.x() < bbox_max[0]

[0956] && pos.y() >= bbox_min[1] && pos.y() < bbox_max[1]

[0957] && pos.z() >= bbox_min[2] && pos.z() < bbox_max[2]) {

[0958] prtSubgroup = m;

[0959] break;

[0960] }

[0961] else if (m == layerGroupParams.numSubgroupsMinus1[prtLayerGroup] &&prtSubgroup == 0)

[0962] }

[0963] }

[0964] retained_subgroup[prtSubgroup].push_back(retained[i]);

[0965] }

[0966] if (dcmNodesList.size()) {

[0967] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; i++) {

[0968] if (layerGroupParams.sliceSelectionIndicationFlag[prtLayerGroup][i]

[0969] && dcmNodeList_ParentSubgroup[i].size()) {

[0970] for (int j = 0; j < dcmNodeList_ParentSubgroup[i].size(); j++)retained_subgroup[i].push_back(dcmNodeList_ParentSubgroup[i][j]);

[0971] std::sort(retained_subgroup[i].begin(), retained_subgroup[i].end());

[0972] }

[0973] }

[0974] if (layerIdx >= 2) {

[0975] std::cout << "(retained.size() = " << retained.size();

[0976] auto prev = retained.size();

[0977] int parentLayerIndex = layerIdx - 2;

[0978] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; i++) {

[0979] if (layerGroupParams.sliceSelectionIndicationFlag[prtLayerGroup][i]){

[0980] for (int k = 0; k < layerGroupParams.dcmNodesIdx[parentLayerIndex][i].size(); k++) {

[0981] auto pointCloudIndex = layerGroupParams.dcmNodesIdx[parentLayerIndex][i][k];

[0982] auto packedVoxelIndex = pointIdxToPackedVoxelIdx[pointCloudIndex];retained.push_back(packedVoxelIndex);

[0983] }

[0984] }

[0985] }

[0986] }

[0987] }

[0988] for (int i = startIndex; i < endIndex; i++) {

[0989] auto pos = (pointCloud[packedVoxel[indexes[i]].index] >>shiftLayerGroup) << (shiftLayerGroup + treeLvlGap);

[0990] curSubgroup = 0;

[0991] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[curLayerGroup]; m++) {

[0992] if (layerGroupParams.sliceSelectionIndicationFlag[curLayerGroup][m]){

[0993] auto bbox_min = layerGroupParams.subgrpBboxOrigin[curLayerGroup][m];

[0994] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[curLayerGroup][m];

[0995] if (pos.x() >= bbox_min[0] && pos.x() < bbox_max[0]

[0996] && pos.y() >= bbox_min[1] && pos.y() < bbox_max[1]

[0997] && pos.z() >= bbox_min[2] && pos.z() < bbox_max[2]) {

[0998] curSubgroup = m;

[0999] break;

[1000] }

[1001] }

[1002] }

[1003] std::vector<point_t> biasedPos_indexes, biasedPos_retained;

[1004] if (concatenateLayers && !layerGroupParams.layerGroupEnabledFlag) {

[1005] indexesOfSubsample.resize(endIndex);

[1006] if (startIndex != endIndex) {

[1007] for (int32_t i = startIndex; i < endIndex; i++)

[1008] indexesOfSubsample[i] = indexes[i];

[1009] int32_t numOfPointInSkipped = geom_num_points_minus1 + 1 -pointCount;

[1010] if (endIndex - startIndex <= startIndex + numOfPointInSkipped) {

[1011] concatenateLayers = false;

[1012] } else {

[1013] for (int32_t i = 0; i < startIndex; i++)

[1014] indexes[i] = indexesOfSubsample[i];

[1015] Reset predIndex.

[1016] predIndex = pointCount;

[1017] for (int lod = 0; lod < lodIndex - minGeomNodeSizeLog2; lod++) {

[1018] int divided_startIndex =

[1019] pointCount - numberOfPointsPerLevelOfDetail[lod];

[1020] int divided_endIndex =

[1021] pointCount - numberOfPointsPerLevelOfDetail[lod + 1];

[1022] computeNearestNeighbors(

[1023] aps, abh, packedVoxel, retained, divided_startIndex,

[1024] divided_endIndex, lod + minGeomNodeSizeLog2, indexes, predictors,

[1025] pointIndexToPredictorIndex, predIndex, atlas

[1026] , layerGroupParams, curLayerGroup, curSubgroup

[1027] , biasedPos_indexes, biasedPos_retained );

[1029] }

[1030] }

[1031] }

[1032] }

[1033] int numIndexedNodes = 0;

[1034] if(layerGroupParams.layerGroupEnabledFlag){

[1035] indexes.resize(startIndex);

[1036] int count = 0;

[1037] for (int i = 0; i <= layerGroupParams.numSubgroupsMinus1[curLayerGroup]; i++) {

[1038] if (layerGroupParams.sliceSelectionIndicationFlag[curLayerGroup][i]){

[1039] if (curLayerGroup == prtLayerGroup)

[1040] prtSubgroup = i; / prtSubgroup = curSubgroup

[1041] else {

[1042] auto cur_bbox_min = layerGroupParams.subgrpBboxOrigin[curLayerGroup][i];

[1043] auto cur_bbox_max = cur_bbox_min + layerGroupParams.subgrpBboxSize[curLayerGroup][i];

[1044] for (int m = 0; m <= layerGroupParams.numSubgroupsMinus1[prtLayerGroup]; m++) {

[1045] auto bbox_min = layerGroupParams.subgrpBboxOrigin[prtLayerGroup][m];

[1046] auto bbox_max = bbox_min + layerGroupParams.subgrpBboxSize[prtLayerGroup][m];

[1047] if (cur_bbox_min.x() >= bbox_min[0] && cur_bbox_max.x() <= bbox_max[0]

[1048] && cur_bbox_min.y() >= bbox_min[1] && cur_bbox_max.y() <= bbox_max[1]

[1049] && cur_bbox_min.z() >= bbox_min[2] && cur_bbox_max.z() <= bbox_max[2]) {

[1050] prtSubgroup = m;

[1051] break;

[1052] }

[1053] }

[1054] }

[1055] int startIndex_subgroup = 0;

[1056] int endIndex_subgroup = indexes_subgroup[i].size();

[1057] int curSubgroup = i;

[1058] biasedPos_indexes.clear();

[1059] biasedPos_indexes.reserve(indexes_subgroup[curSubgroup].size());

[1060] for (int k = 0; k < indexes_subgroup[curSubgroup].size(); k++) {

[1061] auto pos = pointCloud[packedVoxel[indexes_subgroup[curSubgroup][k]].index];

[1062] auto point = pos;

[1063] biasedPos_indexes.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[1064] }

[1065] biasedPos_retained.clear();

[1066] biasedPos_retained.reserve(retained_subgroup[prtSubgroup].size());

[1067] auto cur_bbox_min = layerGroupParams.subgrpBboxOrigin[curLayerGroup][i];

[1068] auto cur_bbox_max = cur_bbox_min + layerGroupParams.subgrpBboxSize[curLayerGroup][i];

[1069] for (int k = 0; k < retained_subgroup[prtSubgroup].size(); k++) {

[1070] auto pos = pointCloud[packedVoxel[retained_subgroup[prtSubgroup][k]].index];

[1071] auto point = pos;

[1072] if (prtLayerGroup != curLayerGroup) {

[1073] if (pos.x() >= cur_bbox_min.x() && pos.x() < cur_bbox_max.x() &&pos.y() >= cur_bbox_min.y() && pos.y() < cur_bbox_max.y() && pos.z() >=cur_bbox_min.z() && pos.z() < cur_bbox_max.z())

[1074] biasedPos_retained.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[1075] else biasedPos_retained.push_back(times((point >> shiftPrtLayerGroup)<< (shiftPrtLayerGroup + treeLvlGap), aps.lodNeighBias));

[1076] }

[1077] else

[1078] biasedPos_retained.push_back(times((point >> shiftLayerGroup) <<(shiftLayerGroup + treeLvlGap), aps.lodNeighBias));

[1079] }

[1080] computeNearestNeighbors(

[1081] aps, abh, packedVoxel, retained_subgroup[prtSubgroup], startIndex_subgroup, endIndex_subgroup, lodIndex, indexes_subgroup[i],

[1082] predictors, pointIndexToPredictorIndex, predIndex, atlas

[1083] , layerGroupParams, curLayerGroup, curSubgroup

[1084] , biasedPos_indexes, biasedPos_retained);

[1085] for (int m = startIndex_subgroup; m < endIndex_subgroup; m++)

[1086] Here, information about whether the neighbor is in the same subgroupis updated.

[1087] {

[1088] int cur_subgroup = i;

[1089] predIdxToLayerGroupMap.push_back(curLayerGroup);

[1090] predIdxToSubgroupMap.push_back(cur_subgroup);

[1091] indexes.push_back(indexes_subgroup[cur_subgroup][m]);

[1092] }

[1093] numIndexedNodes += indexes_subgroup[i].size();layerGroupParams.numberOfPointsPerLodPerSubgroups.push_back(indexes_subgroup[i].size());

[1094] }

[1095] else layerGroupParams.numberOfPointsPerLodPerSubgroups.push_back(0);

[1096] }

[1097] }

[1098] else

[1099] computeNearestNeighbors(

[1100] aps, abh, packedVoxel, retained, startIndex, endIndex, lodIndex,indexes,

[1101] predictors, pointIndexToPredictorIndex, predIndex, atlas

[1102] , layerGroupParams, curLayerGroup, curSubgroup

[1103] , biasedPos_indexes, biasedPos_retained);

[1104] if (layerGroupParams.layerGroupEnabledFlag) {

[1105] if(pointCount > indexes.size())

[1106] numberOfPointsPerLevelOfDetail.push_back(pointCount - indexes.size());

[1107] }

[1108] else {

[1109] if (!retained.empty()) {

[1110] numberOfPointsPerLevelOfDetail.push_back(retained.size());

[1111] }

[1112] }

[1113] input.resize(0);

[1114] std::swap(retained, input);

[1115] }

[1116] std::reverse(indexes.begin(), indexes.end());

[1117] updatePredictors(pointIndexToPredictorIndex, predictors);

[1118] std::reverse(

[1119] numberOfPointsPerLevelOfDetail.begin(),

[1120] numberOfPointsPerLevelOfDetail.end());

[1121] layerGroupParams.numRefNodesInTheSameSubgroup.resize(pointCount, 0);

[1122] Here, information about whether neighbors are in the same subgroup is updated.

[1123] / Update information about whether neighbors are in the same subgroup.

[1124] if (layerGroupParams.layerGroupEnabledFlag) {

[1125] std::reverse(predIdxToLayerGroupMap.begin(),predIdxToLayerGroupMap.end());

[1126] std::reverse(predIdxToSubgroupMap.begin(), predIdxToSubgroupMap.end());

[1127] for (int i = 0; i < pointCount; i++) {

[1128] auto& predictor = predictors[i];

[1129] int curLayerGroup = predIdxToLayerGroupMap[i];

[1130] int curSubgroup = predIdxToSubgroupMap[i];

[1131] for (int k = 0; k < predictor.neighborCount; k++) {

[1132] int idx = predictor.neighbors[k].predictorIndex;

[1133] int neighborLayerGroup = predIdxToLayerGroupMap[idx];

[1134] int neighborSubgroup = predIdxToSubgroupMap[idx];

[1135] if(neighborLayerGroup == curLayerGroup && neighborSubgroup ==curSubgroup) {

[1136] predictor.neighbors[k].inTheSameSubgroupFlag = true;

[1137] layerGroupParams.numRefNodesInTheSameSubgroup[idx]++;

[1138] }

[1139] else

[1140] predictor.neighbors[k].inTheSameSubgroupFlag = false;

[1141] }

[1142] }

[1143] }

[1144] }

[1145] The point cloud data sending / receiving method / device according to the embodiment can search for nearest neighbors and perform map shifting.

[1146] When using upper (parent) child groups with additional missing nodes, the differences in node positions between the parent and child child groups can be significant, and therefore neighbors may not be found in the current attribute graph. If the current point does not exist in the current graph, a graph shifting process can be performed to align the starting positions of nodes in the parent and child LoDs.

[1147] For example, a graph shift can be represented in pseudocode as follows.

[1148] if (curAtlasId != pointAtlasId) {

[1149] atlas.clearUpdates();

[1150] curAtlasId = pointAtlasId;

[1151] while(cubeIndex < retainedSize

[1152] && (packedVoxel[retained[cubeIndex]].mortonCode >> atlasBoundaryBit)< curAtlasId)

[1153] ++cubeIndex;

[1154] while (cubeIndex < retainedSize

[1155] && (packedVoxel[retained[cubeIndex]].mortonCode >> atlasBoundaryBit)==curAtlasId) {atlas.set(packedVoxel[retained[cubeIndex]].mortonCode >>shiftBits3, cubeIndex);

[1156] ++cubeIndex;

[1157] }

[1158] }

[1159] ​ The starting node included in a subgroup can be different from the starting node included in the parent group. In this case, the nodes in the parent group may not match the nodes in the subgroup, resulting in incorrect attribute values ​​being compiled. To avoid this, a procedure can be added to shift the graph of the nodes (points) included in the parent group so that it matches the graph of the starting node included in the subgroup.

[1160] Compare the current point's map ID, which is the target of the current attribute encoding, with the current map ID.

[1161] If the current map ID is different from the map ID of the current point, update the current map ID to the map ID of the current point.

[1162] When the graph ID of a node belonging to the upper part of the reserved graph is different from the current graph ID (it can be smaller than the current graph ID because the graph IDs are sorted in ascending order), a point index called the cube index can be added to select only the points belonging to the current graph and exclude the points that do not belong to the current graph.

[1163] For nodes (points) in a subgroup that belong to the current graph, the current graph information can be updated based on a combination of point index (called cube index) and graph index.

[1164] The process involves constructing a graph by selecting points belonging to the child group's bounding box (bbox) from the points belonging to the parent group when the child group's bounding box (bbox) is included in the parent group's bounding box (bbox).

[1165] Therefore, in "while(packedVoxel[retained[cubeIndex]].mortonCode third;glt;atlasBoundaryBit)third;curAtlasId)++cubeIndex", this process finds the initial starting point of the nodes in the parent-child group included in the bounding box of the sub-subgroup by excluding points that do not belong to the current atlas ID (curAtlasId) from the points included in LoD N-1. Then, through the second while loop, the retained points belonging to the current atlas ID (curAtlasId) are registered with the atlas.

[1166] Regarding the relationship between graph shifting and NN search, when selecting nodes in the parent-child group as neighbor candidates to predict the attributes of nodes in the child-child group, the graph of the parent-child group can be shifted based on the graph ID of the child-child group.

[1167] The point cloud data sending / receiving method / device according to the embodiment can search for nearest neighbors, shift maps (when necessary), and derive subgroup weights.

[1168] During prediction transformation, quantized weights are used to assign credit to nodes used in the prediction. In G-PCC v1, the weights are derived as the sum of the weights of relevant nodes that reference the current node as a neighbor. Based on this concept, the weights of nodes in a subgroup can be calculated using the following constraints.

[1169] Subgroup weight derivation: When the current node and its neighbors belong to the same subgroup, generate the sum of the weights of the relevant nodes that reference the current node as a neighbor.

[1170] refer to ​ For each point (node) in a subgroup, weights used to compensate for losses in neighbor selection can be allocated to nodes that are frequently selected as neighbors in order to reduce losses.

[1171] Weights are used for quantization.

[1172] The process of generating weights is configured as follows:

[1173] Weight(m)+=α*numRefNodes(m) / maxNumRefNodes+β

[1174] numRefNodes: Indicates the number of times this node is used as a neighbor at a specific node m.

[1175] maxNumRefNodes: Indicates the maximum number of nodes in the parent group that can use the current node as a neighbor.

[1176] α and β are variables.

[1177] The process of calculating the quantization weights can be represented by the following pseudocode.

[1178] inline void

[1179] computeQuantizationWeights(

[1180] const std::vector <pccpredictor>& predictors,

[1181] std::vector<int64_t>& quantizationWeights,

[1182] Vec3<int32_t> neighWeight,

[1183] LayerGroupSlicingParams& layerGroupParams,

[1184] bool flag = true) {

[1185] const size_t pointCount = predictors.size();

[1186] quantizationWeights.resize(pointCount);

[1187] for (size_t i = 0; i < pointCount; ++i) {

[1188] quantizationWeights[i] = (1 << kFixedPointWeightShift);

[1189] }

[1190] for (size_t i = 0; i < pointCount; ++i) {

[1191] const size_t predictorIndex = pointCount - i - 1;

[1192] const auto& predictor = predictors[predictorIndex];

[1193] const auto currentQuantWeight = quantizationWeights[predictorIndex];

[1194] for (size_t j = 0; j < predictor.neighborCount; ++j) {

[1195] if (predictor.neighbors[j].inTheSameSubgroupFlag) {

[1196] const size_t neighborPredIndex = predictor.neighbors[j].predictorIndex;

[1197] auto& neighborQuantWeight = quantizationWeights[neighborPredIndex];

[1198] neighborQuantWeight += divExp2RoundHalfInf(

[1199] neighWeight[j] * currentQuantWeight, kFixedPointWeightShift);

[1200] }

[1201] else if (flag) {

[1202] const size_t neighborPredIndex = predictor.neighbors[j].predictorIndex;

[1203] layerGroupParams.quantWeights_outOfSubgroup[neighborPredIndex] +=divExp2RoundHalfInf(

[1204] neighWeight[j] * currentQuantWeight, kFixedPointWeightShift);

[1205] }

[1206] }

[1207] }

[1208] }

[1209] This method ensures that the decoder is targeted for the entire fine-grained slice (FGS) (refer to ​ The layer-based slices (i.e., items at the reference slice granularity unit) and the missing FGS both produce the same subgroup weights as the encoder. However, the weights may be significantly smaller than expected, and may exacerbate the loss at higher LoDs. Implementations may include performing subgroup weight adjustments for each subgroup and each node to mitigate the loss.

[1210] Adjust the subgroup weights as follows.

[1211] Weight(m)+=α*numRefNodes(m) / maxNumRefNodes+β

[1212] The weight (m) is the subgroup weight of the m-th node.

[1213] numRefNodes(m) is the number of related nodes of the m-th node.

[1214] maxNumRefNodes is the maximum number of related nodes in the current subgroup.

[1215] α and β are weight adjustment coefficients that can be derived from the encoder.

[1216] As a method for adaptively correcting the weight of each node during the subgroup weight adjustment process, the degree of weight used for correction can be determined based on the number of times neighboring nodes reference each node within the subgroup, as described above. Besides the methods described above, another approach can be used for adaptive node correction weight calculation. Indicators indicating the importance of each node, such as proportional weights based on weight magnitude, weights outside the subgroup, and the number of nodes outside the subgroup, can be used to improve the correction accuracy of important nodes (i.e., nodes frequently referenced by other nodes).

[1217] The process of weight utilization and correction according to the embodiment can be represented by the following pseudocode.

[1218] template<typename T>

[1219] void SubgroupQuantizationWeightAdjustment(T& quantWeights,LayerGroupSlicingParams& layerGroupParams) {

[1220] int predCnt = 1; / starting from lod 1

[1221] int arrayIdx = 1; / starting from lod 1

[1222] layerGroupParams.quantWeights_top.clear();

[1223] layerGroupParams.quantWeights_top.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[1224] layerGroupParams.quantWeights_bottom.clear();

[1225] layerGroupParams.quantWeights_bottom.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[1226] std::vector<std::vector <int>> numPointsInSubgroups;

[1227] numPointsInSubgroups.resize(layerGroupParams.numLayerGroupsMinus1 +1);

[1228] std::vector<std::vector <int>> idxSelectedNodes;

[1229] idxSelectedNodes.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[1230] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[1231] int numSubgroups = layerGroupParams.numSubgroupsMinus1[i] + 1;

[1232] layerGroupParams.quantWeights_top[i].resize(numSubgroups);

[1233] layerGroupParams.quantWeights_bottom[i].resize(numSubgroups);

[1234] numPointsInSubgroups[i].resize(numSubgroups);

[1235] idxSelectedNodes[i].resize(numSubgroups);

[1236] std::vector <int>maxValue;

[1237] std::vector <int>maxNumRefNodes;

[1238] maxValue.resize(numSubgroups);

[1239] maxNumRefNodes.resize(numSubgroups);

[1240] if (i == 0) { / lod 0

[1241] idxSelectedNodes[0][0] = 0;

[1242] maxNumRefNodes[0] = layerGroupParams.numRefNodesInTheSameSubgroup[0];

[1243] numPointsInSubgroups[0][0] = layerGroupParams.numberOfPointsPerLodPerSubgroups[0];

[1244] }

[1245] for (int j = 0; j < layerGroupParams.numLayersPerLayerGroup[i]; j++){

[1246] for (int k = 0; k < numSubgroups; k++) {

[1247] int lyrgrpIdx = i;

[1248] int subgrpIdx = numSubgroups - 1 - k;

[1249] int numPoints = layerGroupParams.numberOfPointsPerLodPerSubgroups[arrayIdx++];

[1250] for (int m = 0; m < numPoints; m++) {

[1251] if (maxValue[subgrpIdx] < layerGroupParams.quantWeights_outOfSubgroup[predCnt])

[1252] maxValue[subgrpIdx] = layerGroupParams.quantWeights_outOfSubgroup[predCnt];

[1253] if (maxNumRefNodes[subgrpIdx] < layerGroupParams.numRefNodesInTheSameSubgroup[predCnt]) {

[1254] idxSelectedNodes[lyrgrpIdx][subgrpIdx] = predCnt;

[1255] maxNumRefNodes[subgrpIdx] = layerGroupParams.numRefNodesInTheSameSubgroup[predCnt];

[1256] }

[1257] if (j == layerGroupParams.numLayersPerLayerGroup[lyrgrpIdx] - 1) {

[1258] layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx] +=layerGroupParams.quantWeights_outOfSubgroup[predCnt];

[1259] }

[1260] predCnt++;

[1261] }

[1262] numPointsInSubgroups[lyrgrpIdx][subgrpIdx] += numPoints;

[1263] if (j == layerGroupParams.numLayersPerLayerGroup[lyrgrpIdx] - 1) {

[1264] layerGroupParams.quantWeights_top[lyrgrpIdx][subgrpIdx] = maxValue[subgrpIdx];

[1265] if (numPoints)

[1266] layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx] / =numPoints;

[1267] else

[1268] layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx] = 0;

[1269] }

[1270] }

[1271] }

[1272] }

[1273] predCnt = 1;

[1274] arrayIdx = 1;

[1275] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++) {

[1276] int numSubgroups = layerGroupParams.numSubgroupsMinus1[i] + 1;

[1277] if (i == 0) / for lod 0

[1278] quantWeights[0] += (layerGroupParams.quantWeights_top[0][0] -layerGroupParams.quantWeights_bottom[0][0])

[1279] * layerGroupParams.numRefNodesInTheSameSubgroup[0] / layerGroupParams.numRefNodesInTheSameSubgroup[idxSelectedNodes[0][0]]

[1280] + layerGroupParams.quantWeights_bottom[0][0];

[1281] for (int j = 0; j < layerGroupParams.numLayersPerLayerGroup[i]; j++){

[1282] for (int k = 0; k < numSubgroups; k++) {

[1283] int lyrgrpIdx = i;

[1284] int subgrpIdx = numSubgroups - 1 - k;

[1285] int numPoints = layerGroupParams.numberOfPointsPerLodPerSubgroups[arrayIdx++];

[1286] for (int m = 0; m < numPoints; m++) {

[1287] if (layerGroupParams.numRefNodesInTheSameSubgroup[idxSelectedNodes[lyrgrpIdx][subgrpIdx]])

[1288] quantWeights[predCnt] += (layerGroupParams.quantWeights_top[lyrgrpIdx][subgrpIdx] - layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx])

[1289] * layerGroupParams.numRefNodesInTheSameSubgroup[predCnt] / layerGroupParams.numRefNodesInTheSameSubgroup[idxSelectedNodes[lyrgrpIdx][subgrpIdx]]

[1290] + layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx];

[1291] else

[1292] quantWeights[predCnt] += layerGroupParams.quantWeights_bottom[lyrgrpIdx][subgrpIdx];

[1293] predCnt++;

[1294] }

[1295] }

[1296] }

[1297] }

[1298] }

[1299] The method / apparatus according to the embodiments can perform a boost transformation from the encoder's perspective as follows.

[1300] From the encoder's perspective, the sum of α and β can be calculated as the maximum value among the subgroup weights. β can be set as the average of the subgroup weights within the subgroup leaf nodes.

[1301] During the encoding process, each node is determined to be included in a subgroup based on subgroup bounding box information. When a subgroup changes, the encoder, context, zero-run information, arithmetic encoder, etc., for each subgroup are updated. Zero-run is a method to improve compression efficiency by transmitting consecutive zeros. However, dividing into subgroups can lead to disconnections within a LoD, thereby reducing compression efficiency. Implementations can improve compression efficiency by ensuring continuity between the end of the previous LoD and the beginning of the next LoD within the same subgroup.

[1302] Each time a subgroup containing points changes, the zero values ​​of the previous subgroup can be retained. When processing points belonging to the current subgroup, zeros can be accumulated.

[1303] For example, when the last three bits of LOD 1 are zero and the first five bits of LOD 2 are also zero, compression efficiency can be improved by considering eight consecutive zero bits for compression instead of compressing each LOD individually.

[1304] In cases where a layer group includes multiple subgroups, nodes belonging to different subgroups can be mixed (because they are sorted based on Morton codes). If a probabilistic update is performed on the context, it determines which context to use at the boundaries of the subgroups, such as the context for subgroup 1 and the context for subgroup 2; this change affects compression / construction performance.

[1305] For example, refer to ​ Subgroup 3-1 can have LoD N-1 and LoD N, and subgroup 3-2 can also have LoDN-1 and LoD N. Arithmetic encoding and zero-run information can be updated during attribute compilation within each subgroup. That is, the continuity of arithmetic compilation-related information within LoDs in a subgroup can increase compression / reconstruction efficiency.

[1306] As a method for correcting quantization weights, similar to Subgroup Quantization Weight Adjustment described above, quantization parameter (QP) adjustment can be applied to each subgroup, each layer within a subgroup, or each region within a subgroup. This allows the degree of quantization to be adjusted based on the characteristics of the subgroup or layers within a subgroup. The characteristics of the subgroup or layers within a subgroup can include the number of distribution points or a density representing the number of points per unit space. Values ​​calculated by the encoder or arbitrarily specified values ​​can be transmitted via the header as the QP difference for each subgroup, each layer, or each region. Subgroup quantization weight adjustment and QP adjustment can be used selectively or simultaneously for each subgroup, each layer within a subgroup, or each region within a subgroup.

[1307] The encoder boost transformation process can be represented by the following code:

[1308] void

[1309] AttributeEncoder::encodeColorsLift(

[1310] const AttributeDescription& desc,

[1311] const AttributeParameterSet& aps,

[1312] const QpSet& qpSet,

[1313] PCCPointSet3 & pointCloud

[1314] PCCResidualsEncoder& encoder___s,

[1315] std::vector<std::unique_ptr <entropyencoder>>& arithmeticEncoders,

[1316] LayerGroupSlicingParams& layerGroupParams,

[1317] const AttributeBrickHeader& abh,

[1318] const AttributeContexts& ctxtMem)

[1319] {

[1320] const size_t pointCount = pointCloud.getPointCount();

[1321] std::vector<uint64_t> weights;

[1322] if (layerGroupParams.layerGroupEnabledFlag) {

[1323] layerGroupParams.quantWeights_outOfSubgroup.clear();

[1324] layerGroupParams.quantWeights_outOfSubgroup.resize(_lods.predictors.size());

[1325] }

[1326] if (!aps.scalable_lifting_enabled_flag) {

[1327] PCCComputeQuantizationWeights(_lods.predictors, weights

[1328] , layerGroupParams, true);

[1329] } else {

[1330] computeQuantizationWeightsScalable(

[1331] _lods.predictors, _lods.numPointsInLod, pointCount, 0, weights);

[1332] }

[1333] if (layerGroupParams.layerGroupEnabledFlag && layerGroupParams.subgroupQuantWeightAdjEnabledFlag)

[1334] SubgroupQuantizationWeightAdjustment(quantWeights,layerGroupParams);

[1335] const size_t lodCount = _lods.numPointsInLod.size();

[1336] std::vector<Vec3<int64_t>> colors;

[1337] colors.resize(pointCount);

[1338] for (size_t index = 0; index < pointCount; ++index) {

[1339] const auto& color = pointCloud.getColor(_lods.indexes[index]);

[1340] for (size_t d = 0; d < 3; ++d) {

[1341] colors[index][d] = int32_t(color[d]) << kFixedPointAttributeShift;

[1342] }

[1343] }

[1344] for (size_t i = 0; (i + 1) < lodCount; ++i) {

[1345] const size_t lodIndex = lodCount - i - 1;

[1346] const size_t startIndex = _lods.numPointsInLod[lodIndex - 1];

[1347] const size_t endIndex = _lods.numPointsInLod[lodIndex];

[1348] PCCLiftPredict(_lods.predictors, startIndex, endIndex, true, colors);

[1349] PCCLiftUpdate(

[1350] _lods.predictors, weights, startIndex, endIndex, true, colors);

[1351] }

[1352] The coefficients used for predicting the final component at each level of detail (LoD) are calculated as follows.

[1353] / Coefficients used for each level of detail in the final component prediction

[1354] int8_t lastCompPredCoeff = 0;

[1355] if (aps.last_component_prediction_enabled_flag) {

[1356] _abh->attrLcpCoeffs = computeLastComponentPredictionCoeff(aps,colors);

[1357] lastCompPredCoeff = _abh->attrLcpCoeffs[0];

[1358] }

[1359] auto arithmeticEncoderIt = arithmeticEncoders.begin();

[1360] PCCResidualsEncoder encoder(aps, abh, ctxtMem, arithmeticEncoderIt->get());

[1361] Vec3 <int>bbox_min = { 0,0,0};

[1362] Vec3 <int>bbox_max = layerGroupParams.subgrpBboxOrigin[0][0];

[1363] std::vector<std::vector<std::unique_ptr <pccresidualsencoder>>>savedStateVector;

[1364] std::vector <int>zeroRunAccVector;

[1365] int curLayerGroupId = 0;

[1366] int curSubgroupId = 0;

[1367] int prevLayerGroupId = 0;

[1368] int prevSubgroupId = 0;

[1369] int numPrevSlices = 0;

[1370] int sum_layers = 1;

[1371] int startIdx, endIdx;

[1372] if (layerGroupParams.layerGroupEnabledFlag) {

[1373] savedStateVector.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[1374] zeroRunAccVector.resize(layerGroupParams.numLayerGroupsMinus1 + 1);

[1375] for (int layerGroupIdx = 0; layerGroupIdx <= layerGroupParams.numLayerGroupsMinus1; layerGroupIdx++) {

[1376] savedStateVector[layerGroupIdx].resize(layerGroupParams.numSubgroupsMinus1[layerGroupIdx] + 1);

[1377] }

[1378] }

[1379] else {

[1380] zeroRunAccVector.resize(1);

[1381] }

[1382] int numRemainingNodesInCurSubgroup = 0;

[1383] int cnt = 0;

[1384] int numSubgroups = 0;

[1385] for (int i = 0; i <= layerGroupParams.numLayerGroupsMinus1; i++)

[1386] for (int j = 0; j <= layerGroupParams.numSubgroupsMinus1[i]; j++)

[1387] numSubgroups++;

[1388] layerGroupParams.numPointsPerAttrSubgroups.resize(numSubgroups);

[1389] int zeroRun = 0;

[1390] int quantLayer = 0;

[1391] int lod = 0;

[1392] for (size_t predictorIndex = 0; predictorIndex < pointCount;

[1393] ++predictorIndex) {

[1394] if (predictorIndex == _lods.numPointsInLod[quantLayer]) {

[1395] quantLayer = std::min(int(qpSet.layers.size()) - 1, quantLayer + 1);

[1396] }

[1397] When a subgroup changes, the process of updating the encoder, context, zero-run information, and arithmetic encoder for each subgroup is represented by the following code.

[1398] if (layerGroupParams.layerGroupEnabledFlag) {

[1399] if (predictorIndex == 0) {

[1400] prevLayerGroupId = 0;

[1401] prevSubgroupId = 0;

[1402] curLayerGroupId = 0;

[1403] curSubgroupId = 0;

[1404] numRemainingNodesInCurSubgroup = layerGroupParams.numberOfPointsPerLodPerSubgroups[cnt++];

[1405] sum_layers += layerGroupParams.numLayersPerLayerGroup[curLayerGroupId];

[1406] zeroRunAccVector.resize(layerGroupParams.numSubgroupsMinus1[curLayerGroupId] + 1, 0);

[1407] layerGroupParams.numPointsPerAttrSubgroups[0] =numRemainingNodesInCurSubgroup;

[1408] startIdx = 0;

[1409] endIdx = numRemainingNodesInCurSubgroup;

[1410] }

[1411] else if (!numRemainingNodesInCurSubgroup) {

[1412] numRemainingNodesInCurSubgroup = layerGroupParams.numberOfPointsPerLodPerSubgroups[cnt++];

[1413] startIdx = endIdx;

[1414] endIdx += numRemainingNodesInCurSubgroup;

[1415] if (predictorIndex == _lods.numPointsInLod[lod] && curSubgroupId ==0) {

[1416] lod++;

[1417] prevLayerGroupId = curLayerGroupId;

[1418] prevSubgroupId = curSubgroupId;

[1419] if (lod == sum_layers) {

[1420] curLayerGroupId++;

[1421] curSubgroupId = layerGroupParams.numSubgroupsMinus1[curLayerGroupId];

[1422] int refLayerGroupId = layerGroupParams.refLayerGroupId[curLayerGroupId][curSubgroupId];

[1423] int refSubgroupId = layerGroupParams.refSubgroupId[curLayerGroupId][curSubgroupId];savedStateVector[prevLayerGroupId][prevSubgroupId].reset(newPCCResidualsEncoder(encoder));

[1424] for (int i = 0; i < layerGroupParams.numSubgroupsMinus1[prevLayerGroupId] + 1; i++) {

[1425] if (zeroRunAccVector[i]) {

[1426] encoder = *savedStateVector[prevLayerGroupId][i];encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices + i))->get();encoder.encodeRunLength(zeroRunAccVector[i]);savedStateVector[prevLayerGroupId][i].reset(new PCCResidualsEncoder(encoder));

[1427] }

[1428] }

[1429] zeroRunAccVector.clear();zeroRunAccVector.resize(layerGroupParams.numSubgroupsMinus1[curLayerGroupId] + 1, 0);

[1430] numPrevSlices += layerGroupParams.numSubgroupsMinus1[prevLayerGroupId] + 1;

[1431] encoder = *savedStateVector[refLayerGroupId][refSubgroupId];

[1432] encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices +curSubgroupId))->get();if (aps.last_component_prediction_enabled_flag)

[1433] lastCompPredCoeff = _abh->attrLcpCoeffs[lod];

[1434] }

[1435] else {

[1436] curSubgroupId = layerGroupParams.numSubgroupsMinus1[curLayerGroupId];savedStateVector[prevLayerGroupId][prevSubgroupId].reset(newPCCResidualsEncoder(encoder));

[1437] encoder = *savedStateVector[curLayerGroupId][curSubgroupId];

[1438] encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices +curSubgroupId))->get();

[1439] if (aps.last_component_prediction_enabled_flag)

[1440] lastCompPredCoeff = _abh->attrLcpCoeffs[lod];

[1441] }

[1442] }

[1443] else {

[1444] prevLayerGroupId = curLayerGroupId;

[1445] prevSubgroupId = curSubgroupId--;

[1446] if (lod == sum_layers) {

[1447] int refLayerGroupId = layerGroupParams.refLayerGroupId[curLayerGroupId][curSubgroupId];

[1448] int refSubgroupId = layerGroupParams.refSubgroupId[curLayerGroupId][curSubgroupId];savedStateVector[prevLayerGroupId][prevSubgroupId].reset(newPCCResidualsEncoder(encoder));

[1449] encoder = *savedStateVector[refLayerGroupId][refSubgroupId];

[1450] encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices +curSubgroupId))->get();

[1451] if (!curSubgroupId)

[1452] sum_layers += layerGroupParams.numLayersPerLayerGroup[curLayerGroupId];

[1453] }

[1454] else {savedStateVector[prevLayerGroupId][prevSubgroupId].reset(newPCCResidualsEncoder(encoder));

[1455] encoder = *savedStateVector[curLayerGroupId][curSubgroupId];

[1456] encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices +curSubgroupId))->get();

[1457] }

[1458] }

[1459] int idx = curSubgroupId;

[1460] for (int i = 0; i < curLayerGroupId; i++)

[1461] idx += (layerGroupParams.numSubgroupsMinus1[i] + 1);

[1462] layerGroupParams.numPointsPerAttrSubgroups[idx] +=numRemainingNodesInCurSubgroup;

[1463] }

[1464] if (numRemainingNodesInCurSubgroup)

[1465] numRemainingNodesInCurSubgroup--;

[1466] else {

[1467] predictorIndex--;

[1468] continue;

[1469] }

[1470] }

[1471] else {

[1472] if (predictorIndex == _lods.numPointsInLod[lod]) {

[1473] lod++;

[1474] if (aps.last_component_prediction_enabled_flag)

[1475] lastCompPredCoeff = _abh->attrLcpCoeffs[lod];

[1476] }

[1477] }

[1478] const auto pointIndex = _lods.indexes[predictorIndex];

[1479] auto quant = qpSet.quantizers(pointCloud[pointIndex], quantLayer);

[1480] if (layerGroupParams.layerGroupEnabledFlag && aps.aps_slice_qp_deltas_present_flag) {

[1481] auto qp_delta = layerGroupParams.qp_delta[curLayerGroupId][curSubgroupId];

[1482] qp_delta[0] -= layerGroupParams.qp_delta[0][0][0];

[1483] qp_delta[1] -= layerGroupParams.qp_delta[0][0][1];

[1484] quant = qpSet.quantizers(quantLayer, qp_delta);

[1485] }

[1486] const int64_t iQuantWeight = irsqrt(weights[predictorIndex]);

[1487] const int64_t quantWeight =

[1488] (weights[predictorIndex] * iQuantWeight + (1ull << 39)) >> 40;

[1489] auto& color = colors[predictorIndex];

[1490] int values[3];

[1491] values[0] = quant[0].quantize(color[0] * quantWeight);

[1492] int64_t scaled = quant[0].scale(values[0]);

[1493] color[0] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1494] values[1] = quant[1].quantize(color[1] * quantWeight);

[1495] scaled = quant[1].scale(values[1]);

[1496] color[1] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1497] color[2] -= (lastCompPredCoeff * color[1]) >> 2;

[1498] scaled *= lastCompPredCoeff;

[1499] scaled >>= 2;

[1500] values[2] = quant[1].quantize(color[2] * quantWeight);

[1501] scaled += quant[1].scale(values[2]);

[1502] color[2] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1503] if (!values[0] && !values[1] && !values[2])

[1504] ++zeroRunAccVector[curSubgroupId];

[1505] else {

[1506] encoder.encodeRunLength(zeroRunAccVector[curSubgroupId]);

[1507] encoder.encode(values[0], values[1], values[2]);

[1508] zeroRunAccVector[curSubgroupId] = 0;

[1509] }

[1510] }

[1511] if (layerGroupParams.layerGroupEnabledFlag) {

[1512] savedStateVector[curLayerGroupId][curSubgroupId].reset(newPCCResidualsEncoder(encoder));

[1513] for (int i = 0; i < layerGroupParams.numSubgroupsMinus1[curLayerGroupId] + 1; i++) {

[1514] if (zeroRunAccVector[i]) {

[1515] encoder = *savedStateVector[curLayerGroupId][i];

[1516] encoder.arithmeticEncoder = (arithmeticEncoderIt + (numPrevSlices +i))->get();

[1517] encoder.encodeRunLength(zeroRunAccVector[i]);

[1518] savedStateVector[curLayerGroupId][i].reset(new PCCResidualsEncoder(encoder));

[1519] }

[1520] }

[1521] }

[1522] else {

[1523] if (zeroRun)

[1524] encoder.encodeRunLength(zeroRun);

[1525] }

[1526] The following is the reconstruction process.

[1527] for (size_t lodIndex = 1; lodIndex < lodCount; ++lodIndex) {

[1528] const size_t startIndex = _lods.numPointsInLod[lodIndex - 1];

[1529] const size_t endIndex = _lods.numPointsInLod[lodIndex];

[1530] PCCLiftUpdate(

[1531] _lods.predictors, weights, startIndex, endIndex, false, colors);

[1532] PCCLiftPredict(_lods.predictors, startIndex, endIndex, false,colors);

[1533] }

[1534] int64_t clipMax = (1 << desc.bitdepth) - 1;

[1535] for (size_t f = 0; f < pointCount; ++f) {

[1536] const auto color0 =

[1537] divExp2RoundHalfInf(colors[f], kFixedPointAttributeShift);

[1538] Vec3<attr_t> color;

[1539] for (size_t d = 0; d < 3; ++d) {

[1540] color[d] = attr_t(PCCClip(color0[d], 0, clipMax));

[1541] }

[1542] pointCloud.setColor(_lods.indexes[f], color);

[1543] }

[1544] }

[1545] template<typename T>

[1546] void

[1547] PCCLiftPredict(

[1548] const std::vector <pccpredictor>& predictors,

[1549] const size_t startIndex,

[1550] const size_t endIndex,

[1551] const bool direct,

[1552] std::vector <t>& attributes)

[1553] {

[1554] const size_t predictorCount = endIndex - startIndex;

[1555] for (size_t index = 0; index < predictorCount; ++index) {

[1556] const size_t predictorIndex = predictorCount - index - 1 +startIndex;

[1557] const auto& predictor = predictors[predictorIndex];

[1558] auto& attribute = attributes[predictorIndex];

[1559] T predicted(T(0));

[1560] for (size_t i = 0; i < predictor.neighborCount; ++i) {

[1561] const size_t neighborPredIndex = predictor.neighbors[i].predictorIndex;

[1562] const uint32_t weight = predictor.neighbors[i].weight;

[1563] assert(neighborPredIndex < startIndex);

[1564] predicted += weight * attributes[neighborPredIndex];

[1565] }

[1566] predicted = divExp2RoundHalfInf(predicted, kFixedPointWeightShift);

[1567] if (direct) {

[1568] attribute -= predicted;

[1569] } else {

[1570] attribute += predicted;

[1571] }

[1572] }

[1573] } / -----------------------------------------------------------------

[1574] Regarding lifting transform, by restricting the update process to use only points belonging to a specific subgroup, it can be ensured that both the encoder and decoder obtain the same result, even if a portion of the decoding function is utilized.

[1575] template<typename T>

[1576] void

[1577] PCCLiftUpdate(

[1578] const std::vector <pccpredictor>& predictors,

[1579] const std::vector<uint64_t>& quantizationWeights,

[1580] const size_t startIndex,

[1581] const size_t endIndex,

[1582] const bool direct,

[1583] std::vector <t>& attributes)

[1584] {

[1585] std::vector<uint64_t> updateWeights;

[1586] updateWeights.resize(startIndex, uint64_t(0));

[1587] std::vector <t>updates

[1588] updates.resize(startIndex);

[1589] for (size_t index = 0; index < startIndex; ++index) {

[1590] updates[index] = int64_t(0);

[1591] }

[1592] const size_t predictorCount = endIndex - startIndex;

[1593] for (size_t index = 0; index < predictorCount; ++index) {

[1594] const size_t predictorIndex = predictorCount - index - 1 +startIndex;

[1595] const auto& predictor = predictors[predictorIndex];

[1596] const auto currentQuantWeight = quantizationWeights[predictorIndex];

[1597] for (size_t i = 0; i < predictor.neighborCount; ++i) {

[1598] Here, the following procedure is performed to check whether the current node's neighbors belong to the same subgroup as the current node.

[1599] if (predictor.neighbors[i].inTheSameSubgroupFlag) {

[1600] const size_t neighborPredIndex = predictor.neighbors[i].predictorIndex;

[1601] const auto weight = divExp2RoundHalfInf(

[1602] predictor.neighbors[i].weight * currentQuantWeight,

[1603] kFixedPointWeightShift);

[1604] assert(neighborPredIndex < startIndex);

[1605] updateWeights[neighborPredIndex] += weight;

[1606] updates[neighborPredIndex] += weight * attributes[predictorIndex];

[1607] }

[1608] }

[1609] }

[1610] for (size_t predictorIndex = 0; predictorIndex < startIndex;

[1611] ++predictorIndex) {

[1612] const uint32_t sumWeights = updateWeights[predictorIndex];

[1613] if (sumWeights) {

[1614] auto& update = updates[predictorIndex];

[1615] update = divApprox(update, sumWeights, 0);

[1616] auto& attribute = attributes[predictorIndex];

[1617] if (direct) {

[1618] attribute += update;

[1619] } else {

[1620] attribute -= update;

[1621] }

[1622] }

[1623] }

[1624] }

[1625] The lifting transform of the decoder, which corresponds to the encoder and performs the inverse process, is represented as follows.

[1626] void

[1627] AttributeDecoder::decodeColorsLift(

[1628] const AttributeDescription& desc,

[1629] const AttributeParameterSet& aps,

[1630] const AttributeBrickHeader& abh,

[1631] const QpSet& qpSet,

[1632] int geom_num_points_minus1,

[1633] int minGeomNodeSizeLog2,

[1634] PCCResidualsDecoder & decoder

[1635] PCCPointSet3 & pointCloud

[1636] LayerGroupSlicingParams& layerGroupParams,

[1637] const int curLayerGroup,

[1638] const int curSubgroup)

[1639] {

[1640] const size_t pointCount = pointCloud.getPointCount();

[1641] int8_t lastCompPredCoeff = 0;

[1642] if (layerGroupParams.layerGroupEnabledFlag){

[1643] if (curLayerGroup == 0) {

[1644] _weights.clear();

[1645] PCCComputeQuantizationWeights(_lods.predictors, _weights,

[1646] layerGroupParams, false);

[1647] _colors.resize(pointCount);

[1648] }

[1649] }

[1650] else if (!aps.scalable_lifting_enabled_flag) {

[1651] PCCComputeQuantizationWeights(_lods.predictors, _weights,

[1652] layerGroupParams, false);

[1653] }

[1654] else {

[1655] computeQuantizationWeightsScalable(

[1656] _lods.predictors, _lods.numPointsInLod, geom_num_points_minus1 + 1,

[1657] minGeomNodeSizeLog2, _weights);

[1658] }

[1659] if (layerGroupParams.layerGroupEnabledFlag && layerGroupParams.subgroupQuantWeightAdjEnabledFlag)

[1660] SubgroupQuantizationWeightAdjustment(_weights, layerGroupParams);

[1661] int zeroRunRem = 0;

[1662] int quantLayer = 0;

[1663] int idxStart = 0;

[1664] int idxEnd = pointCount;

[1665] int lodStart, lodEnd;

[1666] int subgroupIdx;

[1667] int64_t clipMax = (1 << desc.bitdepth) - 1;

[1668] if (layerGroupParams.layerGroupEnabledFlag) {

[1669] The point range according to the subgroup and LOD is determined as follows.

[1670] lodEnd = 0;

[1671] int prevLodEnd = 0;

[1672] for (int i = 0; i <= curLayerGroup; i++) {

[1673] lodEnd += layerGroupParams.numLayersPerLayerGroup[i];

[1674] if (i != curLayerGroup)

[1675] prevLodEnd += layerGroupParams.numLayersPerLayerGroup[i];

[1676] }

[1677] if (!curLayerGroup)

[1678] lodStart = 0;

[1679] else

[1680] lodStart = prevLodEnd + 1;

[1681] quantLayer = std::min(int(qpSet.layers.size()) - 1, prevLodEnd);

[1682] }

[1683] int subgroupIdxInit = layerGroupParams.numSubgroupsMinus1[curLayerGroup] - curSubgroup;

[1684] for (int i = 0; i < curLayerGroup; i++)

[1685] subgroupIdxInit += (layerGroupParams.numSubgroupsMinus1[i] + 1) * (layerGroupParams.numLayersPerLayerGroup[i]);

[1686] if (curLayerGroup > 0)

[1687] subgroupIdxInit++;

[1688] subgroupIdx = subgroupIdxInit;

[1689] int idxStartInit = 0;

[1690] for (int i = 0; i < subgroupIdx; i++)

[1691] idxStartInit += layerGroupParams.numberOfPointsPerLodPerSubgroups[i];

[1692] int idxEndInit = idxStartInit + layerGroupParams.numberOfPointsPerLodPerSubgroups[subgroupIdx];

[1693] idxStart = idxStartInit;

[1694] idxEnd = idxEndInit;

[1695] for (int lod = lodStart; lod <= lodEnd; lod++) {

[1696] if (aps.last_component_prediction_enabled_flag)

[1697] lastCompPredCoeff = abh.attrLcpCoeffs[lod];

[1698] for (size_t predictorIndex = idxStart; predictorIndex < idxEnd;

[1699] ++predictorIndex) {

[1700] const uint32_t pointIndex = _lods.indexes[predictorIndex];

[1701] auto quant = qpSet.quantizers(pointCloud[pointIndex], quantLayer);

[1702] if (layerGroupParams.layerGroupEnabledFlag && aps.aps_slice_qp_deltas_present_flag) {

[1703] auto qp_delta = layerGroupParams.qp_delta[curLayerGroup][curSubgroup];

[1704] qp_delta[0] -= layerGroupParams.qp_delta[0][0][0];

[1705] qp_delta[1] -= layerGroupParams.qp_delta[0][0][1];

[1706] quant = qpSet.quantizers(quantLayer, qp_delta);

[1707] }

[1708] if (--zeroRunRem < 0)

[1709] zeroRunRem = decoder.decodeRunLength();

[1710] int32_t values[3] = {};

[1711] if (!zeroRunRem)

[1712] decoder.decode(values);

[1713] const int64_t iQuantWeight = irsqrt(_weights[predictorIndex]);

[1714] auto& color = _colors[predictorIndex];

[1715] int64_t scaled = quant[0].scale(values[0]);

[1716] color[0] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1717] scaled = quant[1].scale(values[1]);

[1718] color[1] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1719] scaled *= lastCompPredCoeff;

[1720] scaled >>= 2;

[1721] scaled += quant[1].scale(values[2]);

[1722] color[2] = divExp2RoundHalfInf(scaled * iQuantWeight, 40);

[1723] }

[1724] if (lod != lodEnd) {

[1725] int prevIdx = subgroupIdx;

[1726] subgroupIdx += layerGroupParams.numSubgroupsMinus1[curLayerGroup] +1;

[1727] for (int i = prevIdx; i < subgroupIdx; i++)

[1728] idxStart += layerGroupParams.numberOfPointsPerLodPerSubgroups[i];

[1729] idxEnd = idxStart + layerGroupParams.numberOfPointsPerLodPerSubgroups[subgroupIdx];

[1730] quantLayer = std::min(int(qpSet.layers.size()) - 1, quantLayer + 1);

[1731] }

[1732] }

[1733] Reconstruct:

[1734] subgroupIdx = subgroupIdxInit;

[1735] idxStart = idxStartInit;

[1736] idxEnd = idxEndInit;

[1737] for (size_t lodIndex = lodStart; lodIndex <= lodEnd; lodIndex++) {

[1738] if (lodIndex > 0) {

[1739] PCCLiftUpdate(

[1740] _lods.predictors, _weights, idxStart, idxEnd, false, _colors);

[1741] PCCLiftPredict(_lods.predictors, idxStart, idxEnd, false, _colors);

[1742] }

[1743] if (lodIndex != lodEnd) {

[1744] int prevIdx = subgroupIdx;

[1745] subgroupIdx += layerGroupParams.numSubgroupsMinus1[curLayerGroup] +1;

[1746] for (int i = prevIdx; i < subgroupIdx; i++)

[1747] idxStart += layerGroupParams.numberOfPointsPerLodPerSubgroups[i];

[1748] idxEnd = idxStart + layerGroupParams.numberOfPointsPerLodPerSubgroups[subgroupIdx];

[1749] }

[1750] }

[1751] subgroupIdx = subgroupIdxInit;

[1752] idxStart = idxStartInit;

[1753] idxEnd = idxEndInit;

[1754] for (size_t lodIndex = lodStart; lodIndex <= lodEnd; lodIndex++) {

[1755] for (size_t f = idxStart; f < idxEnd; f++) {

[1756] const auto color0 =

[1757] divExp2RoundHalfInf(_colors[f], kFixedPointAttributeShift);

[1758] Vec3<attr_t> color;

[1759] for (size_t d = 0; d < 3; ++d) {

[1760] color[d] = attr_t(PCCClip(color0[d], int64_t(0), clipMax));

[1761] }

[1762] pointCloud.setColor(_lods.indexes[f], color);

[1763] }

[1764] if (lodIndex != lodEnd) {

[1765] int prevIdx = subgroupIdx;

[1766] subgroupIdx += layerGroupParams.numSubgroupsMinus1[curLayerGroup] +1;

[1767] for (int i = prevIdx; i < subgroupIdx; i++)

[1768] idxStart += layerGroupParams.numberOfPointsPerLodPerSubgroups[i];

[1769] idxEnd = idxStart + layerGroupParams.numberOfPointsPerLodPerSubgroups[subgroupIdx];

[1770] }

[1771] }

[1772] }

[1773] ​ An attribute encoding method according to an embodiment is shown.

[1774] according to ​ The point cloud data transmission method / device of the embodiment corresponds to ​ 10000 transmission devices ​ Point cloud video encoder 10002 ​ Transmitter 10003 ​ Acquisition 20000 / Encoding 20001 / Transmission 20002 ​ encoder, ​ Transmission equipment ​ equipment ​ , ​ and ​ encoder, ​ Transmission methods, etc.

[1775] ​ This is an exemplary flowchart for a layer group slice encoder for LoD-based attribute compilation. First, parameters for SPS, GPS, APS, LGSI, etc., are generated, and the layer group structure is configured. The layer group structure can be configured during or before generating the Layer Group Slice Inventory (LGSI). After performing geometry compilation, LoD generation is performed based on the geometry nodes. Attribute compilation layers (e.g., LoD, RAHT compilation layers) are matched with the layer group, and quantization weights are obtained based on the LoD. Then, for each attribute compilation layer, it is checked whether the layer group has changed, and for each node in the attribute compilation layer, it is checked whether the subgroup has changed. When a subgroup or layer group changes, the attribute encoder of the previous subgroup can be stored, and the attribute encoder of the current subgroup can be used to ensure continuity of context, zero runs, etc., within the subgroup, and to ensure independence from neighboring subgroups. After compressing nodes belonging to all attribute compilation layers, the compiled bitstream of each subgroup is encapsulated into each slice.

[1776] Although the above example is compiled based on LoD-based properties, it can also be applied to compile other properties included in the embodiments, such as RAHT.

[1777] refer to ​ The encoder can be based on ​ The flowchart describes the encoding of point cloud data. Depending on the encoder's encoding configuration, parameter values ​​can be generated. Here, parameters such as SPS, GPS, APS, LGSI, etc., can be generated and transmitted in the bitstream along with the encoded point cloud data. The encoder can first encode the geometric data of the point cloud data using a geometric encoder. After geometric encoding, the encoder can generate a Level of Detail (LOD) using an attribute encoder. During LOD generation, according to the above embodiment, attributes can be represented in a layer group structure (which can be referred to as a layer group including subgroups, including the same FGS, etc.). The LOD can be mapped to a layer group. One or more indices of the LOD (which can be referred to as layers, levels, depths, etc.) can be grouped (or mapped) into layer groups. Each layer group can be assigned a weight. For example, when the highest-level layer group is important, it can be assigned a larger weight. On the other hand, when the lowest-level layer group is important, it can be assigned a larger weight. The weight varies depending on the desired configuration, such as random access to and partial compilation of data at the level of detail. Layer groups can be determined. Here, the determined layer group can be a layer group to be partially encoded or partially transmitted. Subgroups can be determined within the subgroups belonging to the layer group. The determined subgroup can be a subgroup to be partially encoded or partially transmitted. Check if the subgroup has changed. If the subgroup has changed, context information can be saved during subgroup encoding, and necessary context information can be loaded. Node attributes are encoded. If the subgroup has not changed, node attributes can be encoded immediately. Attribute data unit headers can be generated when the end of a node and the end of the LOD are reached. Attribute bitstreams can be generated and transmitted along with geometry bitstreams. Additionally, predictive coding can be performed based on subgroups, such as... ​ As shown in the diagram.

[1778] ​ An attribute decoding method according to an embodiment is shown.

[1779] according to ​ The point cloud data receiving method / device of the embodiment corresponds to ​ Receiving device 10004 ​ Receiver 10005 ​ Point cloud video decoder 10006 ​ Transmission 20002 / Decoding 20003 / Rendering 20004 ​ decoder ​ Receiving equipment ​ equipment ​ , ​ and ​ decoder ​ The receiving method, etc.

[1780] The attribute decoder takes as input a fine-grained slice bitstream of attributes, decoded geometry, and a layer group structure. After geometry decoding, LoD generation is performed based on the provided decoded geometry, and weight derivation is performed according to the node-based relationships of the LoD. Information generated through these operations is used in subsequent attribute decoding operations. The first attribute slice can be operated on in the same way as previous attribute slices. Then, when layer group slicing is enabled, the dependent attribute data cell header can be parsed, and information from `layer_group_id` and `subgroup_id` can be retrieved. If necessary, information about the LoD can be signaled. Based on this information, a layer group matching the current LoD can be searched, and parent and child groups can be selected. When the layer group structure is applied to the geometry and attributes in the same manner, the reference and parent information used in geometry decoding can be used. Based on this information, information in the dependent attribute data cells can be decoded, and the point cloud data can be reconstructed.

[1781] During decoding, the LoD generation and weight derivation methods used by the encoder can be employed. Therefore, LoDs can be mapped to layer groups.

[1782] refer to ​ The decoder can be based on ​ The flowchart shown decodes point cloud data. Parameters contained in the bitstream, such as SPS, GPS, APS, and LGSI, can be parsed. The decoder's geometry decoder can decode the geometry data. The decoder's attribute decoder can generate a Level of Detail (LOD). When generating the LOD, a layer group structure (FGS) according to an embodiment can be used. LOD-specific weights can be derived. Weights can be applied to attributes in the reverse process of the encoder. Attribute data unit headers in the bitstream can be parsed. Attribute data units in the bitstream can be decoded. It is checked whether layer group slicing is enabled. When layer group slicing is enabled, dependent attribute data unit headers are parsed. LODs can be mapped to layer groups. A higher subgroup of the current subgroup of the current layer group can be selected. A higher subgroup can refer to the parent-child group of the current subgroup, the parent-parent-child group, etc. Dependent attribute data units can be decoded. It is checked whether the end of the geometry bitstream has been reached. The attribute bitstream corresponding to the geometry bitstream can be decoded. Both geometry and attributes are decoded, and the point cloud output is generated. Then, the decoding process is terminated. When layer group slicing is not enabled, decoding can be terminated without any layer group-based operations. Alternatively, predictive decoding can be performed based on subgroups, such as... ​ As shown in the diagram.

[1783] ​ A bitstream containing parameters and encoded point cloud data according to an embodiment is shown.

[1784] The point cloud data transmission method / device according to the embodiment (which corresponds to) ​ 10000 transmission devices ​ Point cloud video encoder 10002 ​ Transmitter 10003 ​ Acquisition 20000 / Encoding 20001 / Transmission 20002 ​ encoder, ​ Transmission equipment ​ equipment ​ , 20 Encoders of 24 and 30 to 32 ​ Transmission methods, etc., can generate bit streams, such as... ​ As shown.

[1785] according to ​ The point cloud data receiving method / device of the embodiment (which corresponds to) ​ Receiving device 10004 ​ Receiver 10005 ​ Point cloud video decoder 10006 ​ Sending 20002 / Decoding 20003 / Rendering 20004 ​ decoder ​ Receiving equipment ​ equipment ​ , 20 Decoders 25 and 30 to 32 ​ (Methods for receiving data, etc.) can parse bitstreams, such as... ​ As shown.

[1786] Information regarding separate slices (referring to layer groups, subgroups, FSGs, etc.) can be defined in the parameter set and SEI message according to the embodiment. It can be defined in the sequence parameter set, geometric parameter set, attribute parameter set, geometric slice header, and attribute slice header, and can be defined at corresponding locations or individual locations depending on the application or system to use different application scopes, different application methods, etc. In other words, it can have different meanings depending on where the signal is carried. When information is defined in SPS, it can be applied equally throughout the entire sequence. When defined in GPS, it can indicate that the information is used for location reconstruction. When defined in APS, it can indicate that the information is applied to attribute reconstruction. When defined in TPS, it can indicate that the signaling is applied only to points within a tile. When delivered on a slice basis, it can indicate that the signaling is applied only to the corresponding slice. Furthermore, depending on the application and system, different application scopes, different application methods, etc., can be defined at corresponding locations or individual locations. Additionally, when the syntax elements defined below apply to multiple point cloud data streams and the current point cloud data stream, information can be carried in the parent parameter set, etc.

[1787] Each abbreviation has the following meaning. Each abbreviation can be referred to by another term within the scope of its equivalent meaning: SPS: Sequence Parameter Set; GPS: Geometric Parameter Set; APS: Attribute Parameter Set; TPS: Tile Parameter Set; Geom: Geometric Bit Stream = Geometric Tile Header + Geometric Tile Data; Attr: Attribute Bit Stream = Attribute Brick Header + Attribute Brick Data.

[1788] Although defining information independently of compilation techniques has been described, it is possible to define information in conjunction with compilation methods. It can be defined within a tile parameter set to support scalability across different regions. Furthermore, when the syntax elements defined below apply to multiple point cloud data streams as well as the current point cloud data stream, information can be carried in higher-level parameter sets, etc.

[1789] Alternatively, a Network Abstraction Layer (NAL) unit can be defined, and relevant information for selecting the layer (such as layer_id) can be delivered, allowing the bitstream to be selected at the system level.

[1790] The parameters (which may be referred to as metadata, signaling information, etc.) according to the embodiments can be generated during the process of the transmitter according to the embodiments described below, and are delivered to the receiver according to the embodiments for the reconstruction process.

[1791] For example, the parameters according to the embodiments can be generated in the metadata processor (or metadata generator) of the transmitting device according to the embodiments described below, and obtained by the metadata parser of the receiving device according to the embodiments.

[1792] According to the embodiment, the transmission method / device generates, as shown in the example. ​ The information shown, and the transmission containing that information ​ The bitstream.

[1793] ​ A set of sequence parameters according to an embodiment is shown.

[1794] ​ It shows that it includes ​ The set of sequence parameters in the bit stream.

[1795] sps_extension_flag: Indicates whether the sequence parameter set is extended. Depending on whether the set is extended, it contains the following elements:

[1796] layer_group_enabled_flag: Indicates whether the bitstream included in the sequence is encoded based on layer groups (e.g., ​ ).

[1797] num_layer_groups_minus1: Indicates the number of layer groups contained in the bitstream.

[1798] layer_group_id[i]: Identifies each layer group.

[1799] num_layers_minus1[i]: Indicates the number of layers contained in the layer group.

[1800] subgroup_enabled_flag[i]: Indicates whether the subgroup includes subgroups, such as ​ As shown.

[1801] subgroup_bbox_origin_bits_minus1: Indicates the origin position of the bounding box of the subgroup.

[1802] subgroup_bbox_size_bits_minus1: Indicates the size (width, height, and depth) of the bounding box of the subgroup.

[1803] root_subgroup_bbox_origin: Indicates the origin of the bounding box that is the root of the subgroup's bounding box.

[1804] root_subgroup_bbox_size: Indicates the size (width, height, depth) of the bounding box that is the root of the subgroup's bounding box.

[1805] root_subgroup_bbox_origin indicates the origin of the bounding box of the subgroup of the root subgroup.

[1806] root_subgroup_bbox_size indicates the size of the bounding box of the subgroup of the root subgroup.

[1807] ​ The attribute parameter set and attribute data unit header according to an embodiment are shown.

[1808] ​ It shows that it includes ​ The attribute parameter set and attribute data unit header in the bit stream.

[1809] `aps_attr_ref_id_present_flag`: When equal to 1, it indicates the context reference of the attribute slice based on `attr_ref_layer_group_id` and `attr_ref_subgroup_id`. It indicates the context state of subsequent attribute slices, as indicated by `attr_context_reference_indication_flag`. When `aps_attr_ref_id_present_flag` equals 0, the context reference and context state are inherited from the geometry slice corresponding to the `layer_group_id` and `subgroup_id` of the currently dependent attribute slice.

[1810] The increments of subgroup_weight_adj_coeff_a_bits_minus1 and subgroup_weight_adj_coeff_b_bits_minus1 by 1 indicate the sizes of subgroup_weight_adj_coeff_a and subgroup_weight_adj_coeff_b, respectively.

[1811] subgroup_weight_adj_coeff_a and subgroup_weight_adj_coeff_b indicate the coefficients used for weight adjustment in the current subgroup.

[1812] ​ The diagram illustrates the dependency attribute data unit header according to an embodiment.

[1813] ​ It shows that it includes ​ Dependency attribute data unit header in the bitstream.

[1814] The `dadu_attribute_parameter_set_id` of the dependent attribute data cell indicates the active attribute parameter set (APS) indicated by `aps_attr_parameter_set_id`. The value of `dgdu_attribute_parameter_set_id` of the dependent geometry data cell is equal to the value of `adu_geometry_parameter_set_id` of the attribute data cell of the corresponding slice.

[1815] dadu_sps_attr_idx identifies the attribute that is compiled into the list of active SPS attributes.

[1816] dadu_slice_id indicates the attribute fragment to which the current dependent attribute data unit belongs.

[1817] `dadou_layer_group_id` is an indicator of the layer group of the slice. `dadou_layer_group_id` is in the range of 0 to `num_layer_groups_minus1`. If it does not exist, it is inferred to be 0.

[1818] `dadou_subgroup_id` indicates the subgroup identifier of the layer group referenced by `dadu_layer_group_id`. Here, `dadu_subgroup_id` indicates the order of slices within the same `dadu_layer_group_id`. If it does not exist, `dadu_subgroup_id` is inferred to be 0.

[1819] `attr_ref_layer_group_id` is an indicator that specifies the layer group identifier of the context reference of the current dependent attribute data cell. `attr_ref_layer_group_id` is in the range of 0 to `attr_layer_group_id` for the current dependent attribute data cell.

[1820] attr_ref_subgroup_id indicates the reference subgroup of the layer group indicated by attr_ref_layer_group_id.

[1821] When attr_context_reference_indication_flag equals 1, it indicates that the context state of the current dependent attribute slice is inherited by one or more subsequent dependent attribute slices. An attr_context_reference_indication_flag value of 0 indicates that the context state of the current dependent slice is not inherited by subsequent dependent slices.

[1822] The decoder can use `attr_context_reference_indication_flag` to manage the context buffer. When `attr_context_reference_indication_flag` equals 1, the context state of the current dependent slice is stored in the context buffer at the end of decoding. When `attr_context_reference_indication_flag` equals 0, the context state of the current dependent slice is not stored in the context buffer.

[1823] When `subgroup_weight_adj_coeff_present_flag` equals 1, it indicates that `subgroup_weight_adj_coeff_a` and `subgroup_weight_adj_coeff_b` exist for the current subgroup corresponding to the current slice. When `subgroup_weight_adj_coeff_present_flag` equals 0, it indicates that the subgroup weight adjustment coefficients do not exist. Therefore, it is inferred that the subgroup weight adjustment coefficients (`subgroup_weight_adj_coeff_a` and `subgroup_weight_adj_coeff_b`) are 0.

[1824] subgroup_delta_qp_coeff_present_flag: Indicates whether the delta QP coefficients used for the subgroup exist.

[1825] delta_qp_luma[i]: Indicates the incremental QP coefficients applied to each layer in the current layer group for brightness-related parameters.

[1826] delta_qp_chroma[i]: Indicates the incremental QP coefficients for chroma correlation applied to each layer in the current layer group.

[1827] ​ The process of sending partial point cloud data according to an embodiment is illustrated.

[1828] The point cloud data transmission method / device according to the embodiment (which corresponds to) ​ 10000 transmission devices ​ Point cloud video encoder 10002 ​ Transmitter 10003 ​ Acquisition 20000 / Encoding 20001 / Transmission 20002 ​ encoder, ​ Transmission equipment ​ equipment ​ , ​ , ​ and Figures 30 to 32 encoder, Figure 33 Transmission methods, etc., can scalably encode the entire point cloud data, store it in a storage device, perform transcoding for partial encoding, and then send the partial bit stream, such as... Figure 30 As shown.

[1829] Corresponding to Figure 1 Receiving device 10004 Figure 1 Receiver 10005 Figure 1 Point cloud video decoder 10006 Figure 2 Sending 20002 / Decoding 20003 / Rendering 20004 Figure 7 decoder Figure 9 Receiving equipment Figure 10 equipment Figure 11 , Figure 20 , Figure 25 and Figures 30 to 32 decoder Figure 34 The point cloud data receiving method / device according to the embodiments can receive and decode bit streams of partial point cloud data to reconstruct partial geometry and / or partial attributes.

[1830] The embodiments include a method for partitioning and transmitting compressed data based on specific standards for point cloud data. When using layered compilation, compressed data can be partitioned and sent according to layers, which can improve the efficiency on the receiving side.

[1831] The geometry and attributes of point cloud data can be compressed to provide services. In PCC-based services, the compression rate or amount of data used for transmission can be adjusted based on receiver performance or transmission environment.

[1832] When configuring point cloud data on a slice-by-slice basis, receiver performance or transmission environments may change. In this case, it is necessary to 1) pre-transcode the bitstream into a form suitable for each environment, store the bitstream separately, and select a portion for transmission, or 2) perform transcoding before transmission. In this case, if the number of receiver environments to be supported increases or the transmission environment changes frequently, storage space-related issues or transcoding-induced latency may occur.

[1833] Figure 31 A method for sending and receiving point cloud data according to an embodiment is shown.

[1834] and Figure 1 Transmitting device 10000 Figure 1 Point cloud video encoder 10002 Figure 1 Transmitter 10003 Figure 2 Get 20000 / Encode 20001 / Send 20002 Figure 3 encoder, Figure 8 The transmitting device Figure 10 equipment Figure 11 , Figure 20 , Figure 30 and Figure 32 encoder, Figure 33 The point cloud data transmission method / device according to the embodiments, corresponding to the transmission method, can scalably encode the entire point cloud data and store it in a storage device (storage can be skipped), and can select a bit stream to immediately transmit a portion of the point cloud, and... Figure 30 The results are different.

[1835] The point cloud data receiving method / device according to the embodiment (which corresponds to) Figure 1 Receiving device 10004 Figure 1 Receiver 10005 Figure 1 Point cloud video decoder 10006 Figure 2 Sending 20002 / Decoding 20003 / Rendering 20004 Figure 7 decoder Figure 9 Receiving equipment Figure 10 equipment Figure 11 , 20 30 and 32 decoders Figure 34 (The receiving method, etc.) can immediately receive and decode the bit stream of partial point cloud data to reconstruct partial geometry and / or partial attributes.

[1836] According to an embodiment, when compressed data is divided and transmitted according to layers, only the necessary portions of the pre-compressed data can be selectively transmitted in the bitstream step without a separate transcoding operation. This is efficient even in terms of storage space, as each stream requires only one storage space. Furthermore, efficient transmission can be performed even in terms of bandwidth, as only the necessary layers are selectively transmitted.

[1837] The point cloud data receiving method / device according to the embodiments can provide the following effects.

[1838] Figure 32 A method for sending and receiving point cloud data according to an embodiment is shown.

[1839] The point cloud data transmission method / device according to the embodiment (which corresponds to) Figure 1 10000 transmission devices Figure 1 Point cloud video encoder 10002 Figure 1 Transmitter 10003 Figure 2 Acquisition 20000 / Encoding 20001 / Transmission 20002 Figure 3 encoder, Figure 8 Transmission equipment Figure 10 equipment Figure 11 , Figure 20 , Figure 24 and Figures 30 to 32 encoder, Figure 33 Transcoding methods (such as those used for transmission) can scalably encode the entire point cloud data and store it in a storage device (storage can be skipped), perform transcoding, and then send the entire point cloud, such as... Figure 31 As shown.

[1840] The point cloud data receiving method / device according to the embodiment (which corresponds to) Figure 1 Receiving device 10004 Figure 1 Receiver 10005 Figure 1 Point cloud video decoder 10006 Figure 2 Sending 20002 / Decoding 20003 / Rendering 20004 Figure 7 decoder Figure 9 Receiving equipment Figure 10 equipment Figure 11 , Figure 20 , Figure 25 and Figures 30 to 32 decoder Figure 34 (The receiving method, etc.) can receive the entire point cloud, perform scalable decoding on it, and select data through subsampling to perform methods to reconstruct the entire geometry / attributes or reconstruct a portion of the geometry / attributes.

[1841] In other words, when transmitting layered point cloud data, advantageous effects can be achieved on both the transmitting and receiving sides. In this case, if information enabling the reconstruction of the entire PCC data is transmitted regardless of the receiver's performance, the receiver needs to select the data corresponding to the desired layer after reconstructing the point cloud data (i.e., perform data selection or subsampling). In this scenario, since the delivered bitstream has already been decoded, delays may occur in receivers aiming for low latency, or decoding may not be performed at all, depending on the receiver's performance.

[1842] According to an embodiment, when a bitstream is divided into slices and transmitted, the receiver can selectively send the bitstream to the decoder based on the decoder's performance or the density of the point cloud data represented by the application field. In this case, by performing the selection before decoding, decoder efficiency can be improved, and decoders with various performance characteristics can be supported.

[1843] Figure 33 A method for transmitting point cloud data according to an embodiment is shown.

[1844] S3300: The point cloud data transmission method according to the embodiment may include encoding the point cloud data.

[1845] The encoding operation according to the embodiment may include Figure 1 The transmitting device 10000, the point cloud video acquirer 10001, and the point cloud video encoder 10002, Figure 2 Obtaining 20000 / encoding 20001 Figure 3 encoder, Figure 8 encoder, Figure 10 XR equipment 1030, Figure 11 Encoder 15000, Figure 20 Encoder 60000 and slice selector 60001, Figure 24 encoding, Figures 26 to 29 Bitstream and parameter generation, and Figures 30 to 32 Scalable encoders and bitstream selectors.

[1846] S3301: The point cloud data transmission method according to the embodiment may further include transmitting a bit stream containing point cloud data.

[1847] The transmission operation according to the embodiment may include a transmission device 10000, Figure 1 Transmitter 10003 Figure 2 Sending 20002 Figure 3 Bit stream transmission Figure 8 Bit stream transmission Figure 11 Part or all of the PCC bit stream is sent. Figure 20 Layer-based bitstream transmission and Figures 30 to 32 Sending part or all of the bit stream.

[1848] Figure 34 A method for receiving point cloud data according to an embodiment is shown.

[1849] S3400: The point cloud data receiving method according to the embodiment may include receiving a bit stream containing point cloud data.

[1850] The receiving operation according to the embodiment may include Figure 1 Receiving device 10004 and receiver 10005 Figure 2 Sending 20002 Figure 7 and Figure 9 Bit stream reception Figure 10 XR equipment 1030, Figure 11 Reception of all or part of the bit stream Figure 20 and Figure 25 Layer-based bitstream reception Figures 26 to 29 Parameters and bitstream reception and Figures 30 to 32 Receive all or part of the bit stream.

[1851] S3401: The point cloud data receiving method according to the embodiment may further include decoding the point cloud data.

[1852] The decoding operation according to the embodiment may include Figure 1 The receiving device 10004, the point cloud video decoder 10006, and the receiver 10007, Figure 2 Decoding 20003 / Rendering 20004 Figure 7 and Figure 9 decoder Figure 10 XR equipment 1030, Figure 11 Decoder, selectable decoder, data selection, Figure 20 Decoder 60002, Figure 25 Parameter decoding and geometry / attribute decoding, Figures 26 to 29 Parameters and bitstream decoding, and Figures 30 to 32 The decoder, selectable decoder, and data selector / subsampler.

[1853] refer to Figure 1 The point cloud data transmission method according to the embodiments may include: encoding point cloud data and transmitting a bit stream containing point cloud data.

[1854] refer to Figure 18 , 25 and Figure 21 Regarding attribute / LOD / layer groups and nearest neighbor search, the encoding of point cloud data can include encoding the attributes of the point cloud data. Attribute encoding can include: generating layer groups based on the attribute's level of detail (LOD), and searching for neighbor candidates for the attribute based on neighbor nodes of subgroups within the layer group.

[1855] Regarding the pointIdxToSubgroupIdx information, the encoding of point cloud data may include encoding the attributes of the point cloud data. Attribute encoding may include: generating a layer group comprising at least one subgroup based on the level of detail (LOD) used for the attribute, and generating information indicating the relationship between at least one point in the point cloud data and at least one subgroup.

[1856] Regarding the subgroup boundary-based update described in connection with the lifting transformation, the encoding of attributes may further include performing a lifting transformation on the attributes. Contextual information, zero-run information, and arithmetic encoding information for encoding each of at least one subgroup may be updated, wherein the attributes in the layer group are arithmetically encoded by accumulating LOD-related zero-run information at the boundary between the first and second subgroups included in the layer group.

[1857] Regarding the derivation of subgroup weights, the encoding of attributes may also include performing predictive transformations on attributes based on quantized weights, where the quantized weights can be generated based on whether the current node is called a neighbor node.

[1858] refer to Figure 29 Regarding subgroup QP, subgroup increment QP, and layered subgroup increment QP, the encoding of attributes may also include corrected quantization weights. The quantization weights can be corrected based on the number of points included in the subgroups within the layer group, and the bitstream may contain luma-dependent quantization parameters and chroma-dependent quantization parameters for the layer group.

[1859] A point cloud data device for performing a point cloud data transmission method may include: an encoder configured to encode point cloud data; and a transmitter configured to transmit a bit stream containing point cloud data.

[1860] A point cloud data receiving method that performs the reverse process of a point cloud data transmission method may include: receiving a bit stream containing point cloud data, and decoding the point cloud data.

[1861] Decoding point cloud data can include decoding the attributes of the point cloud data. Decoding attributes can include: generating layer groups based on the level of detail (LOD) for the attribute, and searching for neighbor candidates for the attribute based on the neighbor nodes of the subgroups in the layer group.

[1862] Decoding point cloud data can include decoding the attributes of the point cloud data. Attribute decoding can include: generating a layer group comprising at least one subgroup based on the level of detail (LOD) used for the attribute, and generating information indicating the relationship between at least one point in the point cloud data and at least one subgroup.

[1863] Decoding an attribute may also include performing a promotion transformation on the attribute. Context information, zero-run information, and arithmetic decoding information for decoding each of at least one subgroup can be updated, and arithmetic decoding of attributes in a layer group can be performed by accumulating LOD-related zero-run information at the boundary between the first and second subgroups included in the layer group.

[1864] The decoding of an attribute may also include performing a predictive transformation on the attribute based on quantization weights, where the quantization weights may be generated based on whether the current node is called a neighbor node.

[1865] Decoding attributes can also include correcting quantization weights. Quantization weights can be corrected based on the number of points included in the subgroups within a layer group, and the bitstream can contain luma-dependent and chroma-dependent quantization parameters for the layer group.

[1866] refer to Figures 27 to 29 Regarding subgroup_weight_adj_coeff_a and subgroup_weight_adj_coeff_b, the bitstream can contain coefficient information related to the weights, which are associated with the subgroups that include point cloud data.

[1867] The receiving device for performing the point cloud data receiving method may include a receiver configured to receive a bit stream containing point cloud data, and a decoder configured to decode the point cloud data.

[1868] According to the embodiments, point cloud data can be compressed and reconstructed with high compression efficiency. Furthermore, spatial random access and scalable compilation can be achieved.

[1869] Embodiments have been described in accordance with the method and / or apparatus, and the descriptions of the method and apparatus may be applied complementaryly to each other.

[1870] Although the accompanying drawings have been described separately for simplicity, new embodiments can be designed by combining the embodiments shown in the corresponding drawings. The design of computer-readable recording media also falls within the scope of the appended claims and their equivalents, on which a person skilled in the art records programs for performing the above embodiments as needed. The apparatus and methods according to the embodiments are not limited to the configurations and methods of the above embodiments. Various modifications can be made to the embodiments by selectively combining all or some of the embodiments. Although preferred embodiments have been described with reference to the accompanying drawings, a person skilled in the art will understand that various modifications and variations can be made to the embodiments without departing from the spirit or scope of this disclosure as described in the appended claims. Such modifications should not be interpreted in isolation from the technical concept or perspective of the embodiments.

[1871] Various elements of the device according to the embodiments can be implemented by hardware, software, firmware, or a combination thereof. Various elements of the embodiments can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiments, the components according to the embodiments can be implemented as separate chips. According to the embodiments, at least one or more of the components of the device according to the embodiments can include one or more processors capable of executing one or more programs. The one or more programs can perform any one or more of the operations / methods according to the embodiments, or include instructions for performing them. Executable instructions for performing the methods / operations of the device according to the embodiments can be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a transient CRM or other computer program product configured to be executed by one or more processors. Additionally, the memory according to the embodiments can be used as a concept that covers not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Furthermore, processor-readable recording media can be distributed across computer systems connected via a network, allowing processor-readable code to be stored and executed in a distributed manner.

[1872] In this disclosure, " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Furthermore, "A, B" can mean "A and / or B". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, in this specification, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can mean 1) only A, 2) only B, or 3) both A and B. In other words, the term "or" as used in this document should be interpreted as indicating "additionally or alternatively".

[1873] Terms such as "first" and "second" can be used to describe various elements of the embodiments. However, the various components according to the embodiments should not be limited by the terms used above. These terms are only used to distinguish one element from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Both a first user input signal and a second user input signal are user input signals, but do not mean the same user input signal unless the context clearly specifies otherwise.

[1874] The terminology used to describe embodiments is for the purpose of describing particular embodiments and is not intended to limit the embodiments. As used in the description of embodiments and claims, the singular forms "a," "an," and "the" include plural indicators unless the context clearly specifies otherwise. The expression "and / or" is used to include all possible combinations of terms. Terms such as "comprising" or "having" are intended to indicate the presence of drawings, numbers, steps, elements, and / or components, and should be understood not to exclude the possibility of the additional presence of drawings, numbers, steps, elements, and / or components. As used herein, conditional expressions such as "if" and "when" are not limited to optional cases and are intended to perform related operations or interpret related definitions based on specific conditions when those conditions are met.

[1875] Operations according to the embodiments described herein can be performed by a transmitting / receiving device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control the various operations described herein. The processor may be referred to as a controller, etc. In the embodiments, operations may be performed by firmware, software, and / or combinations thereof. Firmware, software, and / or combinations thereof may be stored in a processor or memory.

[1876] The operations according to the above embodiments can be performed by the transmitting device and / or receiving device according to the embodiments. The transmitting / receiving device may include a transmitter / receiver configured to transmit and receive media data, a memory configured to store instructions (program code, algorithms, flowcharts and / or data) for the process according to the embodiments, and a processor configured to control the operation of the transmitting / receiving device.

[1877] The processor may be referred to as a controller, etc., and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above embodiments can be performed by the processor. Alternatively, the processor may be implemented as an encoder / decoder for the operations of the above embodiments.

[1878] [Public Mode]

[1879] As described above, the relevant details have already been described in the best mode for performing the embodiments.

[1880] [Industrial Applicability]

[1881] As described above, the embodiments are applicable in whole or in part to point cloud data sending / receiving devices and systems.

[1882] Those skilled in the art can change or modify the embodiments in various ways within the scope of the embodiments.

[1883] The embodiments may include variations / modifications within the scope of the claims and their equivalents.< / t> < / t> < / pccpredictor> < / t> < / pccpredictor> < / int> < / pccresidualsencoder> < / int> < / int> < / entropyencoder> < / int> < / int> < / int> < / int> < / pccpredictor> < / int> < / int> < / mortoncodewithindex> < / pccpredictor> < / int> < / int> < / int> < / int> < / mortoncodewithindex> < / pccpredictor>

Claims

1. A method for sending point cloud data, comprising: Encode point cloud data; as well as A bitstream containing the point cloud data is sent.

2. The method according to claim 1, wherein: The encoding of the point cloud data includes: Encoding attributes of the point cloud data; and The encoding of the attribute includes: generating a layer group based on a level of detail (LOD) for the attribute; and Neighbor candidates for the attribute are searched based on neighbor nodes of the subgroup in the layer group.

3. The method according to claim 1, wherein: The encoding of the point cloud data includes: Encoding the attributes of the point cloud data, The encoding of the attribute includes: generating a layer group including at least one sub-group based on a level of detail (LOD) for the attribute; and Information indicating a relationship between at least one point of the point cloud data and the at least one subgroup is generated.

4. The method according to claim 3, wherein: The encoding of the attribute also includes: performs a lifting transformation on said attribute, wherein the context information, the zero-run information and the arithmetic coding information for the encoding for each of the at least one subgroup are updated, The attributes in the layer group are arithmetically encoded by accumulating the zero-run information related to the LOD at a boundary between a first subgroup and a second subgroup included in the layer group.

5. The method according to claim 2, wherein: The encoding of the attribute also includes: performing a prediction transformation on the attributes based on the quantized weights, The quantization weight is generated based on whether the current node is referenced as a neighbor node.

6. The method according to claim 5, wherein: The encoding of the attribute also includes: Correcting the quantization weights; wherein the quantization weight is corrected based on the number of points included in the subgroup in the layer group, The bitstream contains luminance-related quantization parameters and chrominance-related quantization parameters for the layer group.

7. A device for sending a point cloud data transmission, comprising: an encoder configured to encode point cloud data; as well as A transmitter is configured to transmit a bit stream containing the point cloud data.

8. A method for receiving point cloud data, the method comprising: receiving a bitstream containing point cloud data; as well as The point cloud data is decoded.

9. The method according to claim 8, wherein: The decoding of the point cloud data includes: Decoding the attributes of the point cloud data, The decoding of the attribute includes: generating a layer group based on a level of detail (LOD) for the attribute; and Neighbor candidates for the attribute are searched based on neighbor nodes of the subgroup in the layer group.

10. The method according to claim 8, wherein: The decoding of the point cloud data includes: Decoding the attributes of the point cloud data, The decoding of the attribute includes: generating a layer group including at least one sub-group based on a level of detail (LOD) for the attribute; and Information indicating a relationship between at least one point of the point cloud data and the at least one subgroup is generated.

11. The method according to claim 10, wherein: The decoding of the attribute also includes: performs a lifting transformation on said attribute, wherein context information, zero-run information and arithmetic decoding information for the decoding for each of the at least one subgroup are updated, The attributes in the layer group are arithmetically decoded by accumulating the zero-run information related to the LOD at a boundary between a first subgroup and a second subgroup included in the layer group.

12. The method according to claim 9, wherein: The decoding of the attribute also includes: performing a prediction transformation on the attributes based on the quantized weights, The quantization weight is generated based on whether the current node is referenced as a neighbor node.

13. The method according to claim 12, wherein: The decoding of the attribute also includes: Correcting the quantization weights; wherein the quantization weight is corrected based on the number of points included in the subgroup in the layer group, The bitstream contains luminance-related quantization parameters and chrominance-related quantization parameters for the layer group.

14. The method according to claim 8, wherein: The bitstream includes coefficient information associated with weights for a subgroup including the point cloud data.

15. A device for receiving point cloud data, comprising: a receiver configured to receive a bitstream containing point cloud data; as well as A decoder is configured to decode the point cloud data.