Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method

By using point cloud compression encoding technology and feedback information processing, the efficiency and quality issues of point cloud data processing have been solved, enabling efficient point cloud services to support virtual reality, augmented reality, and autonomous driving.

CN115918092BActive Publication Date: 2025-11-14LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180044044.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-22
Filing Date
2021-06-17
Publication Date
2025-11-14
Estimated Expiration
2041-06-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently process large amounts of point cloud data, leading to issues such as waiting time and encoding/decoding complexity, which impact the quality and efficiency of point cloud services.

Method used

Point cloud compression coding technology is employed, including geometry-based point cloud compression (G-PCC) and video-based point cloud compression (V-PCC), combined with feedback information processing, to achieve efficient point cloud data transmission and rendering through a transmitting device and a receiving device.

Benefits of technology

It achieves highly efficient point cloud data processing, provides high-quality point cloud services, and supports applications such as virtual reality, augmented reality, and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115918092B_ABST
    Figure CN115918092B_ABST
Patent Text Reader

Abstract

To achieve the technical objective, the point cloud data transmission method according to the embodiments may include the following steps: encoding the point cloud data; and transmitting a bit stream including the point cloud data. The point cloud data reception method according to the embodiments may include the following steps: receiving a bit stream including the point cloud data; and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments relate to methods and apparatus for processing point cloud content. Background Technology

[0002] Point cloud content is content represented by point clouds, which are collections of points belonging to a coordinate system representing three-dimensional space. Point cloud content can represent media configured in three dimensions and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. However, tens of thousands to hundreds of thousands of points are needed to represent point cloud content. Therefore, methods for efficiently processing large amounts of point data are required. Summary of the Invention

[0003] Technical issues

[0004] The embodiments provide apparatus and methods for efficiently processing point cloud data. The embodiments also provide point cloud data processing methods and apparatuses to address latency and encoding / decoding complexity.

[0005] The technical scope of the implementation is not limited to the technical objectives mentioned above, and can be extended to other technical objectives that can be inferred by those skilled in the art based on all the contents disclosed herein.

[0006] Technical solution

[0007] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art upon examination of the following, or may be learned from practice of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, the claims, and the drawings.

[0008] Beneficial effects

[0009] The apparatus and method according to the embodiments can process point cloud data efficiently.

[0010] The apparatus and method according to the embodiments can provide high-quality point cloud services.

[0011] The apparatus and method according to the embodiments can provide point cloud content for providing general services such as VR services and autonomous driving services. Attached Figure Description

[0012] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and constitute a part of this application. They illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. For a better understanding of the various embodiments described below, reference should be made to the description of the following embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar parts. In the drawings:

[0013] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown;

[0014] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an implementation method;

[0015] Figure 3 An exemplary process for capturing point cloud video according to an embodiment is illustrated;

[0016] Figure 4 An exemplary point cloud encoder according to an implementation method is illustrated;

[0017] Figure 5 An example of a voxel according to an embodiment is shown;

[0018] Figure 6 An example of an octree and occupancy code according to an implementation is shown;

[0019] Figure 7 An example of a neighboring node pattern according to an implementation method is shown;

[0020] Figure 8 An example of point configuration in each LOD according to the implementation method is illustrated;

[0021] Figure 9 An example of point configuration in each LOD according to the implementation method is illustrated;

[0022] Figure 10 An example of a point cloud decoder according to an implementation method is shown;

[0023] Figure 11 An example of a point cloud decoder according to an implementation method is shown;

[0024] Figure 12 An example of a transmitting device according to an embodiment is shown;

[0025] Figure 13 An example of a receiving device according to an embodiment is shown;

[0026] Figure 14 An exemplary structure operable in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment is illustrated;

[0027] Figure 15 An example is provided of a method for transmitting point cloud data according to an implementation method;

[0028] Figure 16 An example of a point cloud data transmission device according to an embodiment is shown;

[0029] Figure 17 An example of a point cloud data receiving device according to an embodiment is shown;

[0030] Figure 18 The structure of a bitstream including point cloud data according to an embodiment is shown;

[0031] Figure 19 The set of sequence parameters according to the implementation method is shown;

[0032] Figure 20 The set of geometric parameters according to the implementation method is shown;

[0033] Figure 21 The block parameter set according to the implementation method is shown;

[0034] Figure 22 A geometric slice head according to an embodiment is shown;

[0035] Figure 23 A method for transmitting point cloud data according to an embodiment is illustrated; and

[0036] Figure 24 A method for receiving point cloud data according to an embodiment is illustrated. Detailed Implementation

[0037] Now, reference will be made in detail to preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description, given with reference to the accompanying drawings, is intended to explain exemplary embodiments of the present disclosure and not to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0038] While most of the terms used in this disclosure are selected from commonly used terms in the art, some terms have been arbitrarily chosen by the applicant, and their meanings will be explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the literal meaning of the terms rather than their simple names or connotations.

[0039] Figure 1 An exemplary point cloud content delivery system according to an implementation is shown.

[0040] Figure 1The point cloud content providing system illustrated herein may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 are capable of wired or wireless communication to transmit and receive point cloud data.

[0041] The point cloud data transmission device 10000 according to an embodiment can acquire and process point cloud video (or point cloud content) and transmit the point cloud video (or point cloud content). According to an embodiment, the transmission device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device, and / or a server. According to an embodiment, the transmission device 10000 may include devices configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers.

[0042] The transmitting device 10000 according to the embodiment includes a point cloud video acquirer 10001, a point cloud video encoder 10002 and / or a transmitter (or communication module) 10003.

[0043] The point cloud video acquirer 10001 according to an embodiment acquires point cloud video through processing procedures such as capture, synthesis, or generation. Point cloud video is point cloud content represented by a point cloud as a set of points in 3D space, and may be referred to as point cloud video data. The point cloud video according to an embodiment may include one or more frames. A frame represents a still image / picture. Therefore, point cloud video may include point cloud images / frames / pictures, and may be referred to as point cloud images, frames, or pictures.

[0044] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiments. The point cloud video encoder 10002 can output a bitstream containing the encoded point cloud video data. The bitstream may contain not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0045] According to an embodiment, transmitter 10003 transmits a bitstream containing encoded point cloud video data. The bitstream, according to an embodiment, is encapsulated in a file or segment (e.g., a streaming segment) and transmitted over various networks such as broadcast networks and / or broadband networks. Although not shown in the figures, transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform encapsulation operations. According to an embodiment, the encapsulator may be included in transmitter 10003. According to an embodiment, the file or segment can be transmitted over a network to receiving device 10004 or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). Transmitter 10003 according to an embodiment is capable of wired / wireless communication with receiving device 10004 (or receiver 10005) via networks such as 4G, 5G, and 6G. Additionally, the transmitter can perform necessary data processing operations depending on the network system (e.g., a 4G, 5G, or 6G communication network system). Transmitting device 10000 can transmit encapsulated data on demand.

[0046] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include devices, robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0047] According to an embodiment, receiver 10005 receives a bitstream containing point cloud video data or a file / segment encapsulating the bitstream from a network or storage medium. Receiver 10005 can perform necessary data processing according to the network system (e.g., 4G, 5G, 6G, etc. communication network systems). According to an embodiment, receiver 10005 can decapsulate the received file / segment and output the bitstream. According to an embodiment, receiver 10005 may include a decapsulator (or decapsulator module) configured to perform a decapsulation operation. The decapsulator may be implemented as a separate element (or component) from receiver 10005.

[0048] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to a method used to encode the point cloud video data (e.g., in the inverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression encoding, which is the inverse process of point cloud compression. Point cloud decompression encoding includes G-PCC encoding.

[0049] Renderer 10007 renders the decoded point cloud video data. Renderer 10007 can output point cloud content by rendering not only the point cloud video data but also the audio data. According to one embodiment, renderer 10007 may include a display configured to display the point cloud content. According to another embodiment, the display may be implemented as a separate device or component, rather than being included in renderer 10007.

[0050] The arrows indicated by the dashed lines in the diagram represent the transmission paths of the feedback information acquired by the receiving device 10004. The feedback information reflects the interactivity of the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is the content of a service requiring user interaction (e.g., autonomous driving service, etc.), the feedback information can be provided to the content sender (e.g., the sending device 10000) and / or the service provider. Depending on the implementation, the feedback information may be used in both the receiving device 10004 and the sending device 10000, or it may not be provided.

[0051] The head orientation information according to the embodiment is information about the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information can be information related to the area of ​​the point cloud video that the user is viewing. The viewpoint is the point through which the user is viewing the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of this area can be determined by the field of view (FOV). Therefore, in addition to the head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area the user gazes at in the point cloud video, the gaze duration, etc. According to the embodiment, the receiving device 10004 can send feedback information including the gaze analysis results to the transmitting device 10000. The feedback information according to the embodiment can be obtained during rendering and / or display processing. According to the embodiment, the feedback information can be obtained by one or more sensors included in the receiving device 10004. According to the embodiment, the feedback information can be obtained by the renderer 10007 or a separate external component (or device, part, etc.). Figure 1The dashed lines in the diagram represent the processing of feedback information received by renderer 10007. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, point cloud video decoder 10006 can perform decoding operations based on the feedback information. Receiving device 10004 can send feedback information to transmitting device 10000. Transmitting device 10000 (or point cloud video encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide point cloud content to the user.

[0052] According to the implementation, the transmitting device 10000 may be referred to as an encoder, transmitting device, transmitter, etc., and the receiving device 10004 may be referred to as a decoder, receiving device, receiver, etc.

[0053] (Through a series of processes including acquisition / encoding / sending / decoding / rendering) according to the implementation method Figure 1 The point cloud data processed in the point cloud content provision system can be referred to as point cloud content data or point cloud video data. According to implementation methods, point cloud content data can be used as a concept encompassing metadata or signaling information related to point cloud data.

[0054] Figure 1 The components of the point cloud content providing system illustrated herein can be implemented by hardware, software, processors, and / or combinations thereof.

[0055] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an implementation method.

[0056] Figure 2 The block diagram shows Figure 1 The operation of the point cloud content providing system described herein. As mentioned above, the point cloud content providing system can process point cloud data based on point cloud compression encoding (e.g., G-PCC).

[0057] A point cloud content providing system according to an embodiment (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) can acquire point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system used to represent 3D space. The point cloud video according to an embodiment may include a Ply (polygon file format or Stanford triangle format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. Geometry includes the position of the points. The position of each point may be represented by parameters (e.g., values ​​of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). Attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc., of each point). A point has one or more attributes. For example, a point may have an attribute as color or two attributes as color and reflectivity. According to the implementation, geometry can be referred to as location, geometric information, geometric data, etc., and attributes can be referred to as attributes, attribute information, attribute data, etc. A point cloud content providing system (e.g., a point cloud sending device 10000 or a point cloud video acquirer 10001) can obtain point cloud data from information related to the acquisition and processing of point cloud video (e.g., depth information, color information, etc.).

[0058] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression encoding. As described above, point cloud data can include the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry encoding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding to encode the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and attribute bitstream according to an embodiment can be multiplexed and output as a single bitstream. The bitstream according to an embodiment may also contain signaling information related to geometry encoding and attribute encoding.

[0059] A point cloud content providing system according to an embodiment (e.g., transmitting device 10000 or transmitter 10003) can transmit encoded point cloud data (20002). For example... Figure 1As illustrated, the encoded point cloud data can be represented by a geometric bitstream and an attribute bitstream. Additionally, the encoded point cloud data can be transmitted as a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric and attribute encoding). The point cloud content providing system can encapsulate the bitstream carrying the encoded point cloud data and transmit the bitstream as a file or segment.

[0060] The point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to the embodiment can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bitstream.

[0061] A point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode encoded point cloud data (e.g., geometric bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode point cloud video data based on signaling information related to the encoding of point cloud video data contained in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode the geometric bitstream to reconstruct the location (geometry) of points. The point cloud content providing system can reconstruct the attributes of points by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can reconstruct point cloud video based on the location of the reconstructed geometry and the decoded attributes.

[0062] A point cloud content providing system (e.g., receiving device 10004 or renderer 10007) according to an embodiment can render decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometry and attributes decoded through the decoding process. Points in the point cloud content can be rendered as vertices with a certain thickness, cubes with a specific minimum size centered at the corresponding vertex position, or circles centered at the corresponding vertex position. All or part of the rendered point cloud content is provided to a user through a display (e.g., a VR / AR display, a conventional display, etc.).

[0063] The point cloud content providing system (e.g., receiving device 10004) according to the embodiment can obtain feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. The feedback information and operation and reference of the point cloud content providing system according to the embodiment... Figure 1The feedback information and operations described are the same, so a detailed description of them is omitted.

[0064] Figure 3 An exemplary process for capturing point cloud video according to an implementation method is illustrated.

[0065] Figure 3 Examples of references are provided. Figures 1 to 2 The described point cloud content provides an example of point cloud video capture processing for the system.

[0066] Point cloud content includes point cloud videos (images and / or videos) representing objects and / or environments located in various 3D spaces (e.g., 3D spaces representing real environments, 3D spaces representing virtual environments, etc.). Therefore, the point cloud content providing system according to embodiments can use one or more cameras (e.g., infrared cameras capable of acquiring depth information, RGB cameras capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., infrared pattern projectors for acquiring depth information), LiDRA, etc., to capture point cloud videos. The point cloud content providing system according to embodiments can extract the shape of the geometry composed of points in 3D space from the depth information and extract the attributes of each point from the color information to obtain point cloud data. Images and / or videos according to embodiments can be captured based on at least one of inward-oriented and outward-oriented techniques.

[0067] Figure 3 The left side illustrates inward-facing technology. Inward-facing technology refers to the technique of capturing images of a central object using one or more cameras (or camera sensors) positioned around it. Inward-facing technology can be used to generate point cloud content that provides users with 360-degree images of key objects (e.g., VR / AR content that provides users with 360-degree images of objects such as characters, players, objects, or actors).

[0068] Figure 3 The right side illustrates outward-facing techniques. Outward-facing techniques refer to techniques that capture images of the environment of a central object, rather than the central object itself, using one or more cameras (or camera sensors) positioned around it. Outward-facing techniques can be used to generate point cloud content that provides the surrounding environment from a user's perspective (e.g., content representing the external environment that can be provided to users of autonomous vehicles).

[0069] As shown in the figure, point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate system is different in each camera; therefore, the point cloud content providing system can calibrate one or more cameras to set the global coordinate system before the capture operation. Alternatively, the point cloud content providing system can generate point cloud content by compositing arbitrary images and / or videos with images and / or videos captured using the aforementioned capture techniques. The point cloud content providing system may not perform this step when generating point cloud content representing virtual space. Figure 3 The capture operations described herein. The point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unwanted areas (e.g., background), identify spaces to which the captured images and / or videos are connected, and perform a space-hole filling operation when space holes exist.

[0070] A point cloud content delivery system can generate point cloud content by performing coordinate transformations on points in a point cloud video obtained from each camera. The system can perform coordinate transformations on points based on the coordinates of each camera's location. Therefore, the system can generate content representing a wide range or point cloud content with high-density points.

[0071] Figure 4 An exemplary point cloud encoder according to an implementation method is illustrated.

[0072] Figure 4 It shows Figure 1 An example of a point cloud video encoder 10002. The point cloud encoder reconstructs and encodes point cloud data (e.g., point locations and / or attributes) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) based on network conditions or applications. When the total size of the point cloud content is large (e.g., 60Gbps of point cloud content for 30fps), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide the point cloud content according to network conditions, etc.

[0073] For reference Figures 1 to 2 The point cloud encoder described herein can perform geometric encoding and attribute encoding. Geometric encoding is performed before attribute encoding.

[0074] The point cloud encoder according to the implementation includes a coordinate transformer (transform coordinates) 40000, a quantizer (quantize and remove points (voxarization)) 40001, an octree analyzer (analyze octrees) 40002, a surface approximation analyzer (analyze surface approximations) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometry reconstructor (reconstruct geometry) 40005, a color transformer (transform colors) 40006, an attribute transformer (transform attributes) 40007, a RAHT transformer (RAHT) 40008, an LOD generator (generate LODs) 40009, a lift transformer (lift) 40010, a coefficient quantizer (quantize coefficients) 40011, and / or an arithmetic encoder (arithmetic coding) 40012.

[0075] Coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometric reconstructor 40005 can perform geometric encoding. Geometric encoding according to the implementation may include octree geometric encoding, direct encoding, trisoup geometric encoding, and entropy encoding. Direct encoding and trisoup geometric encoding are applied selectively or in combination. Geometric encoding is not limited to the examples described above.

[0076] As shown in the figure, the coordinate transformer 40000 according to the embodiment receives the position and transforms it into coordinates. For example, the position can be transformed into position information in three-dimensional space (e.g., three-dimensional space represented by the XYZ coordinate system). The position information in three-dimensional space according to the embodiment can be referred to as geometric information.

[0077] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 can quantize points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 40001 performs the following quantization operation: multiplying the difference between the position value of each point and the minimum position value by a preset quantization scaling value, and then finding the nearest integer value by rounding the value obtained by multiplication. Thus, one or more points can have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of a pixel, which is the smallest unit containing 2D image / video information, the points of the point cloud content (or 3D point cloud video) according to the embodiment can be included in one or more voxels. The term voxel, a compound word of volume and pixel, refers to the 3D cubic space generated when 3D space is divided into units (unit = 1.0) based on axes representing 3D space (e.g., X-axis, Y-axis, and Z-axis). The quantizer 40001 can match groups of points in 3D space to voxels. According to one implementation, a voxel may include only one point. According to another implementation, a voxel may include one or more points. To represent a voxel as a point, the center position of the voxel can be set based on the positions of the one or more points included in the voxel. In this case, attributes of all positions included in a voxel can be combined and assigned to that voxel.

[0078] The octree analyzer 40002 according to the implementation performs octree geometric encoding (or octree encoding) to represent voxels in an octree structure. The octree structure represents points based on the octree structure and voxel matching.

[0079] The surface approximation analyzer 40003 according to the embodiment can analyze and approximate octrees. The octree analysis and approximation according to the embodiment analyzes regions containing multiple points to efficiently provide octree and voxelization processing.

[0080] The arithmetic encoder 40004 according to the embodiment performs entropy encoding on an octree and / or an approximate octree. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric bitstream is generated.

[0081] The attribute encoding is performed by a color transformer 40006, an attribute transformer 40007, a RAHT transformer 40008, an LOD generator 40009, a boosting transformer 40010, a coefficient quantizer 40011, and / or an arithmetic encoder 40012. As described above, a point can have one or more attributes. The attribute encoding according to the embodiment is also applied to the attributes a point possesses. However, when an attribute (e.g., color) comprises one or more elements, the attribute encoding is applied independently to each element. The attribute encoding according to the embodiment includes color transformation encoding, attribute transformation encoding, region adaptive hierarchical transformation (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) encoding, and interpolation-based hierarchical nearest neighbor prediction encoding with an update / boosting step (boosting transformation). Depending on the point cloud content, the RAHT encoding, prediction transformation encoding, and boosting transformation encoding described above can be used selectively, or a combination of one or more encoding schemes can be used. The attribute encoding according to the embodiment is not limited to the examples described above.

[0082] The color converter 40006 according to the embodiment performs color transformation encoding on the color values ​​(or textures) included in the transformation attributes. For example, the color converter 40006 can transform the format of color information (e.g., from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment can be optionally applied according to the color values ​​included in the attributes.

[0083] The geometry reconstructor 40005, according to the implementation method, reconstructs (decompresses) octrees and / or approximate octrees. The geometry reconstructor 40005 reconstructs the octree / voxel based on the distribution of analysis points. The reconstructed octree / voxel can be referred to as the reconstructed geometry (recovered geometry).

[0084] The attribute transformer 40007 according to the embodiment performs attribute transformation to transform attributes based on the location and / or reconstructed geometry that has not undergone geometric encoding. As described above, since attributes depend on geometry, the attribute transformer 40007 can transform attributes based on reconstructed geometric information. For example, based on the position values ​​of points included in a voxel, the attribute transformer 40007 can transform the attributes of points at that location. As described above, when the position of the voxel center is set based on the positions of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of said one or more points. When performing trigonometric Thomson geometric encoding, the attribute transformer 40007 can transform attributes based on the trigonometric Thomson geometric encoding.

[0085] The attribute transformer 40007 performs attribute transformation by calculating the average of the attributes or attribute values ​​(e.g., color or reflectivity of each point) of neighboring points within a specific position / radius from the center (or position value) of each voxel. The attribute transformer 40007 can apply weights based on the distance from the center to each point when calculating the average. Therefore, each voxel has a position and a calculated attribute (or attribute value).

[0086] The attribute transformer 40007 can search for nearest neighbors within a specific location / radius of the center of each voxel based on a KD-tree or Morton code. A KD-tree is a binary search tree and supports a data structure that allows points to be managed based on location so that nearest neighbor search (NNS) can be performed quickly. Morton codes are generated by representing the coordinates (e.g., (x, y, z)) of the 3D location of all points as bit values ​​and mixing those bits. For example, when the coordinates representing the location of a point are (5, 9, 1), the bit values ​​for the coordinates are (0101, 1001, 0001). The bit values ​​are mixed according to the bit index in the order of z, y, and x to produce 010001000111. This value is represented as the decimal number 1095. That is, the Morton code value for the point with coordinates (5, 9, 1) is 1095. The attribute transformer 40007 can sort the points based on their Morton code values ​​and perform NNS using a depth-first traversal process. After an attribute transformation operation, if an NNS is required in another transformation process used for attribute encoding, use a KD tree or Morton code.

[0087] As shown in the figure, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.

[0088] According to the implementation, the RAHT transformer 40008 performs RAHT encoding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT transformer 40008 can predict the attribute information of higher-level nodes in an octree based on the attribute information associated with lower-level nodes in the octree.

[0089] The LOD generator 40009 according to the embodiment generates a Level of Detail (LOD) to perform predictive transform coding. The LOD according to the embodiment represents the level of detail of the point cloud content. A decreasing LOD value indicates a decrease in the level of detail of the point cloud content. An increasing LOD value indicates an increase in the level of detail of the point cloud content. Points can be classified according to their LOD.

[0090] The lift transformer 40010 according to the embodiment performs lift transform coding to transform the attributes of the point cloud based on weights. As described above, lift transform coding may optionally be applied.

[0091] According to the implementation method, the coefficient quantizer 40011 quantizes the attribute after attribute encoding based on the coefficient.

[0092] According to the implementation method, the arithmetic encoder 40012 encodes the quantized attributes based on arithmetic encoding.

[0093] Although not shown in the figure, Figure 4 The elements of the point cloud encoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits, which are configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform the above-described... Figure 4 At least one of the operations and / or functions of the elements of the point cloud encoder. Additionally, one or more processors can operate on or execute a set of software programs and / or instructions to perform... Figure 4 The operation and / or function of the elements of the point cloud encoder. One or more memories according to the embodiments may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash storage devices or other non-volatile solid-state storage devices).

[0094] Figure 5 An example of a voxel according to an embodiment is shown.

[0095] Figure 5 This illustrates voxels in 3D space represented by a coordinate system consisting of three axes: the X-axis, Y-axis, and Z-axis. (See reference...) Figure 4 As described, a point cloud encoder (e.g., quantizer 40001) can perform voxelization. A voxel refers to the 3D cubic space generated when the 3D space is divided into cells (unit = 1.0) based on axes representing the 3D space (e.g., X-axis, Y-axis, and Z-axis). Figure 5 An example of voxels generated via an octree structure is shown, in which a bounding box aligned to the cubic axis, defined by two poles (0, 0, 0) and (2d, 2d, 2d), is recursively subdivided. A voxel comprises at least one point. The spatial coordinates of a voxel can be estimated based on its positional relationship to a group of voxels. As mentioned above, voxels possess properties similar to pixels in a 2D image / video (such as color or reflectivity). Details and references of voxels are provided. Figure 4 The details described are the same, so the description of it is omitted.

[0096] Figure 6 An example of an octree and occupancy code according to an implementation is shown.

[0097] For reference Figures 1 to 4The described point cloud content delivery system (point cloud video encoder 10002) or point cloud encoder (e.g., octree analyzer 40002) performs octree geometric encoding (or octree encoding) based on an octree structure to efficiently manage the regions and / or locations of voxels.

[0098] Figure 6 The upper part shows the octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes of the coordinate system (e.g., the X, Y, and Z axes). The octree structure is created by recursively subdividing bounding boxes aligned to the cubic axes defined by two poles (0, 0, 0) and (2d, 2d, 2d). Here, 2d can be set as the value of the minimum bounding box that constitutes all points around the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined by the following formula. In the following formula, (x int n ,y int n ,z int n ) indicates the position (or position value) of the quantization point.

[0099]

[0100] like Figure 6 As shown in the upper center, the entire 3D space can be divided into eight spaces according to partitions. Each partitioned space is represented by a cube with six faces. (See diagram below.) Figure 6 As shown in the upper right corner, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., the X, Y, and Z axes). Thus, each space is divided into eight smaller spaces. These smaller spaces are also represented by cubes with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.

[0101] Figure 6 The lower part shows the octree occupancy code. The occupancy code of the octree is generated to indicate whether each of the eight partitioned spaces resulting from partitioning a space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the partitioned space, and the child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when the space corresponding to the child node contains at least one point, the node is assigned a value of 1. When the space corresponding to the child node does not contain a point (the space is empty), the node is assigned a value of 0. Since... Figure 6The occupancy code shown is 00100001, indicating that the spaces corresponding to the third and eighth child nodes out of eight each contain at least one point. As shown, each of the third and eighth child nodes has eight child nodes, and each child node is represented by an 8-bit occupancy code. The figure shows the occupancy code for the third child node as 10000111, and the occupancy code for the eighth child node as 01001111. A point cloud encoder (e.g., an arithmetic encoder 40004) according to an embodiment can perform entropy coding on the occupancy code. To improve compression efficiency, the point cloud encoder can perform intra / inter-frame coding on the occupancy code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs the octree based on the occupancy code.

[0102] A point cloud encoder according to an implementation method (e.g., Figure 4 A point cloud encoder or octree analyzer (40002) can perform voxelization and octree encoding to store the locations of points. However, points are not always uniformly distributed in 3D space, so there will be specific regions with fewer points. Therefore, performing voxelization over the entire 3D space is inefficient. For example, when a specific region contains fewer points, it is not necessary to perform voxelization in that specific region.

[0103] Therefore, for the specific region mentioned above (or nodes other than the leaf nodes of the octree), the point cloud encoder according to the embodiment can skip voxelization and perform direct encoding to directly encode the positions of points included in the specific region. The coordinates of the directly encoded points according to the embodiment are called the Direct Encoding Mode (DCM). The point cloud encoder according to the embodiment can also perform trigonometric encoding based on the surface model to reconstruct the positions of points in the specific region (or node) based on voxels. Trigonometric encoding is a geometric encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Trigonometric encoding and direct encoding according to the embodiment can be selectively performed. In addition, trigonometric encoding and direct encoding according to the embodiment can be performed in combination with octree geometric encoding (or octree coding).

[0104] To perform direct encoding, the option to apply direct encoding using direct mode should be enabled. The node to be directly encoded is not a leaf node, and there should be fewer than a threshold number of points within that node. Furthermore, the total number of points to be directly encoded should not exceed a preset threshold. When the above conditions are met, the point cloud encoder (or arithmetic encoder 40004) according to the implementation can perform entropy encoding on the point positions (or position values).

[0105] A point cloud encoder according to an embodiment (e.g., a surface approximation analyzer 40003) can determine a specific level of the octree (a level less than the depth d of the octree) and can perform trigonometric tangent coding using a surface model starting from that level to reconstruct the location of points in the region of a node based on voxels (trigonometric tangent mode). The point cloud encoder according to an embodiment can specify the level at which trigonometric tangent coding will be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in trigonometric tangent mode. In other words, the point cloud encoder according to an embodiment can operate in trigonometric tangent mode only when the specified level is less than the depth value of the octree. The 3D cubic region of a node at the specified level according to an embodiment is called a block. A block may include one or more voxels. A block or voxel may correspond to a brick. Geometry is represented by a surface within each block. A surface according to an embodiment may intersect each edge of a block at most once.

[0106] A block has 12 edges, therefore there are at least 12 intersections within a block. Each intersection is called a vertex (or apex point). A vertex is detected along an edge when there is at least one occupied voxel adjacent to that edge in all blocks sharing that edge. An occupied voxel, according to the implementation, refers to a voxel containing a point. The position of a vertex detected along an edge is the average position of the edges of all voxels adjacent to that edge in all blocks sharing that edge.

[0107] Once a vertex is detected, the point cloud encoder according to the embodiment can perform entropy encoding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the vertex position value (relative position value within the edge). When applying trigonometric Tang geometry encoding, the point cloud encoder according to the embodiment (e.g., geometry reconstructor 40005) can generate the restored geometry (reconstructed geometry) by performing trigonometric reconstruction, upsampling, and voxelization.

[0108] Vertices at the edges of a block define the surface passing through the block. The surface, according to the implementation, is a non-planar polygon. In the triangulation process, the surface represented by triangles is reconstructed based on the starting point of the edge, the direction vector of the edge, and the position values ​​of the vertices. The triangulation process is performed by: ① calculating the centroid value of each vertex, ② subtracting the centroid value from each vertex value, and ③ estimating the sum of squares of the values ​​obtained through subtraction.

[0109]

[0110] Estimate the minimum value of the sum and perform projection processing based on the axis with the minimum value. For example, when element x is minimum, each vertex is projected onto the x-axis relative to the center of the block and onto the (y,z) plane. When the value obtained by projecting onto the (y,z) plane is (ai,bi), the value of θ is estimated by atan2(bi,ai), and the vertices are sorted according to the value of θ. Table 1 below shows the vertex combinations for creating triangles based on the number of vertices. The vertices are sorted from 1 to n. Table 1 below shows that for four vertices, two triangles can be constructed based on the combinations of vertices. The first triangle can be composed of vertices 1, 2, and 3 from the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 from the sorted vertices.

[0111] Table 2-1. Triangles formed by vertices sorted by 1, ..., n

[0112]

[0113] 12(1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,11),(11,12,1),(1,3,5),(5,7,9),(9,11,1),(1,5,9)

[0114] Upsampling is performed to add points along the edges of the triangle at the center and voxelization is then performed. The added points are generated based on the upsampling factor and the width of the block. The added points are called thinned vertices. A point cloud encoder according to an implementation can voxelize the thinned vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized locations (or location values).

[0115] Figure 7 An example of a neighbor node pattern according to an implementation method is shown.

[0116] To improve the compression efficiency of point cloud videos, the point cloud encoder according to the implementation method can perform entropy coding based on context-adaptive arithmetic coding.

[0117] For reference Figures 1 to 6 The described point cloud content delivery system or point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or...) Figure 4The arithmetic encoder 40004 can immediately perform entropy coding on the occupancy code. Alternatively, the point cloud content providing system or point cloud encoder can perform entropy coding (intra-frame coding) based on the occupancy code of the current node and the occupancy of neighboring nodes, or entropy coding (inter-frame coding) based on the occupancy code of the previous frame. According to the embodiment, a frame represents a collection of simultaneously generated point cloud videos. The compression efficiency of the intra-frame coding / inter-frame coding according to the embodiment can depend on the number of referenced neighboring nodes. As the number of bits increases, the operation becomes more complex, but the coding can be biased to one side, thereby increasing the compression efficiency. For example, when given a 3-bit context, 2... 3 = There are 8 methods to perform the encoding. The division of space for encoding affects the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be achieved.

[0118] Figure 7 This illustrates a process for obtaining occupancy patterns based on the occupancy of neighboring nodes. A point cloud encoder, according to an implementation, determines the occupancy of the neighboring nodes of each node in an octree and obtains the value of the neighboring node pattern. This neighboring node pattern is then used to infer the occupancy pattern of the node. Figure 7 The left side of the diagram shows the cube corresponding to the node (the cube in the middle) and six cubes sharing at least one face with the cube (neighboring nodes). The nodes shown in the diagram are nodes at the same depth. The numbers shown in the diagram represent the weights associated with the six nodes (1, 2, 4, 8, 16, and 32). Weights are assigned sequentially based on the position of the neighboring nodes.

[0119] Figure 7 The right side of the diagram shows the neighbor node pattern values. The neighbor node pattern value is the sum of the values ​​multiplied by the weights of occupied neighbor nodes (neighbor nodes with points). Therefore, the neighbor node pattern values ​​range from 0 to 63. When the neighbor node pattern value is 0, it indicates that none of the node's neighbors have a point (unoccupied node). When the neighbor node pattern value is 63, it indicates that all neighbor nodes are occupied nodes. As shown in the diagram, since the neighbor nodes assigned weights 1, 2, 4, and 8 are occupied nodes, the neighbor node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform encoding based on the neighbor node pattern values ​​(e.g., 64 encodings can be performed when the neighbor node pattern value is 63). According to implementations, the point cloud encoder can reduce encoding complexity by changing the neighbor node pattern values ​​(e.g., based on a table that changes 64 to 10 or 6).

[0120] Figure 8 An example of point configuration in each LOD according to the implementation method is shown.

[0121] For reference Figures 1 to 7The description describes the reconstruction (decompression) of encoded geometry before performing attribute encoding. When direct encoding is applied, geometry reconstruction operations may include changing the placement of directly encoded points (e.g., placing directly encoded points in front of the point cloud data). When triangulation geometry encoding is applied, geometry reconstruction processing is performed through triangulation, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.

[0122] Point cloud encoders (e.g., LOD generator 40009) can classify (reorganize) points using LOD. This figure illustrates the point cloud content corresponding to LOD. The leftmost image in the figure represents the original point cloud content. The second image from the left shows the distribution of points in the lowest LOD, and the rightmost image shows the distribution of points in the highest LOD. That is, points are sparsely distributed in the lowest LOD and densely distributed in the highest LOD. In other words, as the LOD increases in the direction indicated by the arrow at the bottom of the figure, the space (or distance) between points narrows.

[0123] Figure 9 An example of point configuration for each LOD according to the implementation method is shown.

[0124] For reference Figures 1 to 8 The described point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, ...) Figure 4 A point cloud encoder or LOD generator (40009) can generate LODs. LODs are generated by reorganizing points into a set of refinement levels based on a set of LOD distance values ​​(or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.

[0125] Figure 9 The upper part shows examples of points (P0 to P9) of point cloud content distributed in 3D space. Figure 9 In this context, the original order represents the order of points P0 to P9 before LOD generation. Figure 9 In this context, LOD-based ordering represents the order in which points are generated according to their LOD values. Points are reorganized using LOD. Furthermore, higher LOD values ​​include points belonging to lower LOD values. For example... Figure 9 As shown, LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.

[0126] For reference Figure 4 The point cloud encoder described herein, according to the embodiments, can selectively or in combination perform predictive transform coding, lifting transform coding, and RAHT transform coding.

[0127] The point cloud encoder according to the implementation can generate predictors for points by performing predictive transform coding to set the predictive attributes (or predictive attribute values) for each point. That is, N predictors can be generated for N points. The predictors according to the implementation can calculate weights (= 1 / distance) based on the LOD value of each point, indexed information related to neighboring points existing within a set distance for each LOD, and the distance to the neighboring points.

[0128] According to the implementation, the predicted attribute (or attribute value) is set as the average of the values ​​obtained by multiplying the attributes (or attribute values) of neighboring points (e.g., color, reflectivity, etc.) set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation can quantize and inverse quantize the residual (which may be referred to as residual attribute, residual attribute value, or attribute prediction residual) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. This quantization process is configured as shown in the table below.

[0129] Table. Pseudocode for Attribute Prediction Residual Quantization

[0130] int PCCQuantization(int value,int quantStep){

[0131] if(value>=0){

[0132] return floor(value / quantStep+1.0 / 3.0);

[0133] }else{

[0134] return-floor(-value / quantStep+1.0 / 3.0);

[0135] }

[0136] }

[0137] Table. Pseudocode for Inverse Quantization of Attribute Prediction Residuals

[0138] int PCCInverseQuantization(int value,int quantStep){

[0139] if(quantStep==0){

[0140] return value;

[0141] }else{

[0142] return value * quantStep;

[0143] }

[0144] }

[0145] When the predictor for each point has neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the residual values ​​after quantization and inverse quantization as described above. When the predictor for each point has no neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the attributes of the corresponding point without performing the above operations.

[0146] The point cloud encoder (e.g., lift transformer 40010) according to the embodiment can generate a predictor for each point, set the calculated LOD and register neighboring points in the predictor, and set weights according to the distance to the neighboring points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the prediction transform coding described above, but the difference is that the weights are applied cumulatively to the attribute values. The process of applying weights cumulatively to attribute values ​​according to the embodiment is configured as follows.

[0147] 1) Create an array QuantizationWeight(QW) to store the weight value for each point. All elements of QW are initialized to 1.0. The QW value of the predictor index of the neighboring nodes registered in the predictor is multiplied by the predictor weight of the current point, and the resulting values ​​are summed.

[0148] 2) Improved prediction processing: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.

[0149] 3) Create a temporary array called updateweight, and update and initialize the temporary array to zero.

[0150] 4) The weights calculated by multiplying the weights computed for all predictors by the weights stored in QW corresponding to the predictor indices are cumulatively added to the update weight array as the indices of neighboring nodes. The values ​​obtained by multiplying the attribute values ​​of the neighboring node indices by the calculated weights are cumulatively added to the update array.

[0151] 5) Improved update processing: Divide the attribute values ​​of the update array for all predictors by the weight values ​​of the update weight array of the predictor index, and add the existing attribute values ​​to the values ​​obtained by division.

[0152] 6) The predicted attribute is calculated by multiplying the attribute value updated by the boost update process for all predictors by the weight updated by the boost prediction process (stored in the QW). The predicted attribute value is quantized by a point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation. In addition, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute value.

[0153] A point cloud encoder according to an embodiment (e.g., RAHT transform 40008) can perform RAHT transform coding, where attributes associated with lower-level nodes in an octree are used to predict attributes of higher-level nodes. RAHT transform coding is an example of intra-frame attribute coding performed via a backward scan of an octree. The point cloud encoder according to an embodiment scans the entire region from voxels and repeats a merging process in each step, merging voxels into larger blocks, until the root node is reached. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed in a higher mode directly above an empty node.

[0154] The following equation represents the RAHT transformation matrix. In this equation, g l x,y,z This represents the average attribute value of the voxel at level l. It can be based on g. l+1 2x,y,z and g l+1 2x+1,y,z To calculate g l x,y,z g l 2x,y,z and g l 2x+1,y,z The weights are w1 = w l 2x,y,z and w2 = w l 2x+1,y,z .

[0155]

[0156] here, It is a low-pass value and is used in the next higher level of merge processing. This represents the high-pass coefficient. The high-pass coefficient in each step is quantized and undergoes entropy encoding (e.g., via an arithmetic encoder 40012). Weights are calculated as follows: pass and The root node is calculated as follows.

[0157]

[0158] Figure 10 An example of a point cloud decoder according to an implementation method is shown.

[0159] Figure 10 The point cloud decoder shown in the example is Figure 1 The example of the point cloud video decoder 10006 described in [the document], and can perform [operations] with [other functions]. Figure 1 The point cloud video decoder 10006 illustrated in the figure operates in the same or similar manner. As shown in the figure, the point cloud decoder can receive a geometry bitstream and an attribute bitstream contained in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs the decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and the attribute bitstream and outputs the decoded attributes. The decoded geometry and the decoded attributes are used to reconstruct the point cloud content (the decoded point cloud).

[0160] Figure 11 An example of a point cloud decoder according to an implementation method is shown.

[0161] Figure 11 The point cloud decoder shown in the example is Figure 10 The example shown is a point cloud decoder, which can be executed as... Figures 1 to 9 The example above illustrates the decoding operation, which is the inverse of the encoding operation of a point cloud encoder.

[0162] For reference Figure 1 and Figure 10 As described, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.

[0163] The point cloud decoder according to the implementation includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesized octree) 11001, a surface approximation synthesizer (synthesized surface approximation) 11002, a geometry reconstructor (reconstructed geometry) 11003, an inverse coordinate transformer (inverse coordinate transformation) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantizer (inverse quantization) 11006, a RAHT transformer 11007, an LOD generator (generated LOD) 11008, an inverse lifter (inverse lift) 11009, and / or an inverse color transformer (inverse color transformation) 11010.

[0164] Arithmetic decoder 11000, octree synthesizer 11001, surface approximation synthesizer 11002, geometric reconstructor 11003, and coordinate inverse transformer 11004 can perform geometric decoding. Geometric decoding according to embodiments may include direct encoding and trigonometric Thomson geometric decoding. Direct encoding and trigonometric Thomson geometric decoding are selectively applied. Geometric decoding is not limited to the examples described above and is provided for reference only. Figures 1 to 9 The inverse processing of the described geometric encoding is performed.

[0165] The arithmetic decoder 11000 according to the embodiment decodes the received geometric bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse processing of the arithmetic encoder 40004.

[0166] The octree synthesizer 11001 according to the embodiment can generate an octree by obtaining occupancy codes from the decoded geometric bitstream (or geometric information obtained as a decoding result). See reference... Figures 1 to 9 Configure the occupancy code in detail.

[0167] When applying trisoup geometry encoding, the surface approximation synthesizer 11002 according to the implementation can synthesize the surface based on the decoded geometry and / or the generated octree.

[0168] The geometry reconstructor 11003 according to the embodiment can regenerate geometry based on surfaces and / or decoded geometry. See reference... Figures 1 to 9 As described, direct encoding and trigonometric Tangle geometry encoding are selectively applied. Therefore, geometry reconstructor 11003 directly imports and adds positional information about points where direct encoding has been applied. When trigonometric Tangle geometry encoding is applied, geometry reconstructor 11003 can reconstruct the geometry by performing reconstruction operations (e.g., triangulation, upsampling, and voxelization) of geometry reconstructor 40005. Details and References Figure 6 The details described are the same, so their description is omitted. The reconstructed geometry may include point cloud images or frames that do not contain attributes.

[0169] According to the implementation, the inverse coordinate transformer 11004 can obtain the position of a point based on the reconstructed geometric transformation coordinates.

[0170] Arithmetic decoder 11005, inverse quantizer 11006, RAHT transformer 11007, LOD generator 11008, inverse booster 11009, and / or inverse color transformer 11010 can perform reference... Figure 10 The attribute decoding described herein includes Region Adaptive Hierarchical Transformation (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) decoding, and interpolation-based hierarchical nearest neighbor prediction decoding with an update / lifting step (lifting transformation). The above three decoding schemes may be used selectively, or a combination of one or more decoding schemes may be used. The attribute decoding according to the embodiments is not limited to the examples described above.

[0171] According to the embodiment, the arithmetic decoder 11005 decodes the attribute bitstream by arithmetic encoding.

[0172] The inverse quantizer 11006 according to the implementation performs inverse quantization on the information about the decoded attribute bitstream or attribute obtained as a decoding result, and outputs the inverse quantized attribute (or attribute value). Inverse quantization can be selectively applied based on the attribute encoding of the point cloud encoder.

[0173] According to the implementation, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can process the reconstructed geometry and inversely quantized attributes. As described above, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.

[0174] The inverse color transformer 11010 according to the embodiment performs inverse transformation encoding to inversely transform the color values ​​(or textures) included in the decoded attributes. The operation of the inverse color transformer 11010 can be selectively performed based on the operation of the color transformer 40006 of the point cloud video encoder.

[0175] Although not shown in the figure, Figure 11 The elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits, which are configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform the above-described... Figure 11 The point cloud decoder's components include at least one or more of their operations and / or functions. Additionally, one or more processors can operate on or execute software programs and / or sets of instructions to perform... Figure 11 The operation and / or functions of the components of the point cloud decoder.

[0176] Figure 12 An example of a transmitting device according to an embodiment is shown.

[0177] Figure 12 The transmitting device shown is Figure 1 The transmitting device 10000 (or Figure 4 Example of a point cloud encoder. Figure 12 The transmitting device illustrated in the example can perform and reference Figures 1 to 9The described point cloud encoder operation and method are one or more of the same or similar operations and methods. The transmitting apparatus according to the embodiment may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra / inter-frame coding processor 12005, an arithmetic encoder 12006, a metadata processor 12007, a color transformation processor 12008, an attribute transformation processor 12009, a prediction / boosting / RAHT transformation processor 12010, an arithmetic encoder 12011, and / or a transmitting processor 12012.

[0178] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 can perform operations and / or acquisition methods similar to those of the point cloud video acquirer 10001 (or refer to...). Figure 2 The described acquisition process (20000) is the same as or similar to the operation and / or acquisition method.

[0179] The data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame coding processor 12005, and arithmetic encoder 12006 perform geometric coding. Geometric coding according to the embodiment and reference... Figures 1 to 9 The geometric codes described are the same or similar, so a detailed description of them is omitted.

[0180] The quantization processor 12001 according to the embodiment quantizes geometry (e.g., point position values). The operation of the quantization processor 12001 and / or quantization with reference... Figure 4 The operation and / or quantization of the described quantizer 40001 are the same as or similar. Details and references Figures 1 to 9 The details described are the same.

[0181] The voxelization processor 12002 according to the embodiment performs voxelization on the quantized position values ​​of points. The voxelization processor 12002 can perform operations similar to those described above. Figure 4 The operation and / or voxelization process of the described quantizer 40001 are the same as or similar to the operation and / or processing. Details and references Figures 1 to 9 The details described are the same.

[0182] The octree occupancy code generator 12003 according to the embodiment performs octree encoding based on the voxelized positions of points in the octree structure. The octree occupancy code generator 12003 can generate occupancy codes. The octree occupancy code generator 12003 can perform and reference... Figure 4 and Figure 6The operations and / or methods described are the same as or similar to those of the point cloud video encoder (or octree analyzer 40002). Details and references Figures 1 to 9 The details described are the same.

[0183] According to the implementation, the surface model processor 12004 can perform trigonometric geometry encoding based on a surface model to reconstruct the positions of points in a specific region (or node) based on voxels. The surface model processor 12004 can perform operations related to reference... Figure 4 The operation and / or methods described are the same as or similar to those of the point cloud video encoder (e.g., surface approximation analyzer 40003). Details and references Figures 1 to 9 The details described are the same.

[0184] The intra / inter-frame coding processor 12005 according to the embodiment can perform intra / inter-frame coding on point cloud data. The intra / inter-frame coding processor 12005 can perform operations similar to those described above. Figure 7 The described intra / inter-frame coding is the same or similar. Details and references Figure 7 The details described are the same. According to the implementation, the intra / inter-frame coding processor 12005 may be included in the arithmetic encoder 12006.

[0185] The arithmetic encoder 12006 according to the embodiment performs entropy encoding on an octree and / or an approximate octree of point cloud data. For example, the encoding scheme includes arithmetic encoding. The arithmetic encoder 12006 performs the same or similar operations and / or methods as the arithmetic encoder 40004.

[0186] The metadata processor 12007 according to the embodiment processes metadata (e.g., set values) about point cloud data and provides it to necessary processing procedures such as geometric encoding and / or attribute encoding. Additionally, the metadata processor 12007 according to the embodiment can generate and / or process signaling information related to geometric encoding and / or attribute encoding. The signaling information according to the embodiment can be encoded separately from geometric encoding and / or attribute encoding. The signaling information according to the embodiment can be interleaved.

[0187] Color transformation processor 12008, attribute transformation processor 12009, prediction / boosting / RAHT transformation processor 12010, and arithmetic encoder 12011 perform attribute encoding. Attribute encoding and reference according to the implementation method. Figures 1 to 9 The attribute codes described are the same or similar, so detailed descriptions of them are omitted.

[0188] According to the embodiment, the color transformation processor 12008 performs color transformation encoding to transform color values ​​included in the attributes. The color transformation processor 12008 can perform color transformation encoding based on reconstructed geometry. Reconstructed geometry and reference... Figures 1 to 9 The description is the same. Additionally, it performs the same as the reference. Figure 4 The operation and / or methods of the described color converter 40006 are the same as or similar to those described. Detailed description of it is omitted.

[0189] The attribute transformation processor 12009 according to the implementation performs attribute transformation to transform attributes based on the reconstructed geometry and / or locations where geometric encoding has not been performed. The attribute transformation processor 12009 performs and references... Figure 4 The operation and / or method of the described attribute transformer 40007 are the same as or similar to those operations and / or methods. Detailed descriptions thereof are omitted. The prediction / boosting / RAHT transformation processor 12010 according to the embodiment can encode the transformed attribute by any one or a combination of RAHT encoding, prediction transformation encoding, and boosting transformation encoding. The prediction / boosting / RAHT transformation processor 12010 performs and references... Figure 4 The RAHT transformer 40008, LOD generator 40009, and lift transformer 40010 described herein operate in at least one of the same or similar manner. Furthermore, the predictive transform coding, lift transform coding, and RAHT transform coding are similar to those described in the reference... Figures 1 to 9 The descriptions are the same, so detailed descriptions of them are omitted.

[0190] The arithmetic encoder 12011 according to the embodiment can encode the encoded attributes based on arithmetic encoding. The arithmetic encoder 12011 performs the same or similar operations and / or methods as the arithmetic encoder 40012.

[0191] According to an embodiment, the transmitting processor 12012 can transmit each bitstream containing encoded geometric and / or encoded attribute and metadata information, or transmit a bitstream configured with encoded geometric and / or encoded attribute and metadata information. When the encoded geometric and / or encoded attribute and metadata information according to an embodiment is configured as a bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information, including a sequence parameter set (SPS) for sequence-level signaling, a geometric parameter set (GPS) for signaling for geometric information encoding, an attribute parameter set (APS) for signaling for attribute information encoding, and a tile parameter set (TPS) for tile-level signaling, and tile data. The tile data may include information about one or more tiles. A tile according to an embodiment may include a geometric bitstream Geom0.0 and one or more attribute bitstreams Attr0 0 and Attr1 0 .

[0192] A slice is a series of syntax elements that represent all or part of an encoded point cloud frame.

[0193] The TPS according to the embodiment may include information about each of one or more tiles (e.g., height / size information and coordinate information about the bounding box). The geometric bitstream may include a header and a payload. The header of the geometric bitstream according to the embodiment may include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, as well as information about the data contained in the payload. As described above, the metadata processor 12007 according to the embodiment may generate and / or process signaling information and send it to the transmit processor 12012. According to the embodiment, the elements for performing geometry encoding and the elements for performing attribute encoding may share data / information with each other, as indicated by the dashed lines. The transmit processor 12012 according to the embodiment may perform the same or similar operations and / or transmission methods as the transmitter 10003. Details and References Figure 1 and Figure 2 The details described are the same, so the description of it is omitted.

[0194] Figure 13 An example of a receiving device according to an embodiment is shown.

[0195] Figure 13 The receiving device illustrated in the example is Figure 1 The receiving device 10004 (or Figure 10 and Figure 11 Example of a point cloud decoder. Figure 13 The receiving device illustrated in the example can perform the same operation as the reference. Figures 1 to 11 The operations and methods described in the point cloud decoder are one or more of the same or similar operations and methods.

[0196] The receiving apparatus according to the embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an octree reconstruction processor based on occupancy code 13003, a surface model processor (triangulation, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / boosting / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each element for decoding according to the embodiment can perform the inverse processing of the operation of the corresponding element for encoding according to the embodiment.

[0197] Receiver 13000, according to an embodiment, receives point cloud data. Receiver 13000 can perform operations related to... Figure 1 The operation and / or receiving method of the receiver 10005 are the same as or similar to those of the receiver. Detailed description of it is omitted.

[0198] According to the embodiment, the receiving processor 13001 can acquire geometric bitstreams and / or attribute bitstreams from received data. The receiving processor 13001 may be included in the receiver 13000.

[0199] The arithmetic decoder 13002, the octet-based octree reconstruction processor 13003, the surface model processor 13004, and the inverse quantization processor 13005 can perform geometric decoding. Geometric decoding and reference according to the implementation method. Figures 1 to 10 The described geometric decodings are the same or similar, so a detailed description of them is omitted.

[0200] The arithmetic decoder 13002 according to the embodiment can decode geometric bitstreams based on arithmetic coding. The arithmetic decoder 13002 performs the same or similar operations and / or coding as the arithmetic decoder 11000.

[0201] According to the embodiment, the octree reconstruction processor 13003 based on occupancy codes can reconstruct an octree by obtaining occupancy codes from the decoded geometric bitstream (or geometric information obtained as a decoding result). The octree reconstruction processor 13003 performs operations and / or methods identical or similar to those of the octree synthesizer 11001 and / or the octree generation method. When applying trigonometric Tang geometry encoding, the surface model processor 13004 according to the embodiment can perform trigonometric Tang geometry decoding and related geometric reconstruction (e.g., triangulation, upsampling, voxelization) based on surface model methods. The surface model processor 13004 performs operations identical or similar to those of the surface approximation synthesizer 11002 and / or the geometry reconstructor 11003.

[0202] The inverse quantization processor 13005 according to the embodiment can perform inverse quantization on the decoded geometry.

[0203] The metadata parser 13006 according to the implementation can parse metadata contained in received point cloud data, such as set values. The metadata parser 13006 can transmit metadata for geometry decoding and / or attribute decoding. Metadata and reference Figure 12 The metadata described is the same, so a detailed description of it is omitted.

[0204] The arithmetic decoder 13007, inverse quantization processor 13008, prediction / boost / RAHT inverse transform processor 13009, and color inverse transform processor 13010 perform attribute decoding. Attribute decoding and reference Figures 1 to 10 The properties described are decoded in the same or similar ways, so detailed descriptions of them are omitted.

[0205] The arithmetic decoder 13007 according to the embodiment can decode the attribute bitstream via arithmetic coding. The arithmetic decoder 13007 can decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 13007 performs the same or similar operations and / or coding as the arithmetic decoder 11005.

[0206] The inverse quantization processor 13008 according to the embodiment can perform inverse quantization on the decoded attribute bitstream. The inverse quantization processor 13008 performs the same or similar operations and / or methods as the inverse quantizer 11006 and / or the inverse quantization method.

[0207] The prediction / boosting / RAHT inverse transform processor 13009 according to the embodiment can process the reconstructed geometry and inversely quantized attributes. The prediction / boosting / RAHT inverse transform processor 13009 performs one or more operations and / or decodings that are the same as or similar to those of the RAHT transformer 11007, LOD generator 11008, and / or inverse booster 11009. The color inverse transform processor 13010 according to the embodiment performs inverse transform encoding to inversely transform the color values ​​(or textures) included in the decoded attributes. The color inverse transform processor 13010 performs operations and / or inverse transform encodings that are the same as or similar to those of the inverse color transformer 11010. The renderer 13011 according to the embodiment can render point cloud data.

[0208] Figure 14 An exemplary structure of a combined point cloud data transmission / reception method / apparatus according to an embodiment is illustrated.

[0209] Figure 14The structure represents a configuration in which at least one of server 1460, robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, home appliance 1450, and / or head-mounted display (HMD) 1470 is connected to cloud network 1400. Robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, or home appliance 1450 are referred to as devices. Additionally, XR device 1430 may correspond to a point cloud data (PCC) device according to an embodiment, or may be operatively connected to a PCC device.

[0210] Cloud network 1400 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, cloud network 1400 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0211] Server 1460 can be connected via cloud network 1400 to at least one of robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, home appliance 14500 / or HMD 1470, and can assist at least a portion of the processing of connected devices 1410 to 1470.

[0212] HMD 1470 represents one type of implementation of an XR device and / or PCC device according to an embodiment. An HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0213] Various embodiments of the apparatus 1410 to 1450 applying the above-described technology will be described below. According to the above embodiments, Figure 14 The devices 1410 to 1450 illustrated herein can be operatively connected to / coupled to point cloud data transmitting and receiving devices.

[0214] <PCC+XR>

[0215] The XR / PCC device 1430 can employ PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) installed in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0216] The XR / PCC device 1430 can analyze 3D point cloud data or image data obtained through various sensors or from external devices, and generate position data and attribute data about the 3D points. Thereby, the XR / PCC device 1430 can obtain information about the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC device 1430 can match an XR object including auxiliary information about the recognized object with the recognized object, and output the matched XR object.

[0217] <PCC + XR + Mobile Phone>

[0218] The XR / PCC device 1430 can be implemented as a mobile phone by applying PCC technology.

[0219] The mobile phone can decode and display point cloud content based on PCC technology.

[0220] <PCC + Autopilot + XR>

[0221] The autonomous vehicle 1420 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0222] The autonomous vehicle 1420 applying XR / PCC technology can represent an autonomous vehicle provided with a device for providing XR images or an autonomous vehicle that is a control / interaction target in an XR image. Specifically, the autonomous vehicle 1420 that is a control / interaction target in an XR image can be separated from the XR device 1430, and can be operably connected to the XR device 1430.

[0223] The autonomous vehicle 1420 having a device for providing XR / PCC images can obtain sensor information from sensors including cameras, and output the generated XR / PCC images based on the obtained sensor information. For example, the autonomous vehicle 1420 can have a HUD and output XR / PCC images thereto, thereby providing an XR / PCC object corresponding to a real object or an object existing on the screen to the occupants.

[0224] When the XR / PCC object is output to the HUD, at least a part of the XR / PCC object can be output to overlap with the real object pointed at by the occupants' eyes. On the other hand, when the XR / PCC object is output to a display provided inside the autonomous vehicle, at least a part of the XR / PCC object can be output to overlap with the object on the screen. For example, the autonomous vehicle 1220 can output XR / PCC objects corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two - wheeled vehicles, pedestrians, and buildings.

[0225] Virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or point cloud compression (PCC) technologies according to the implementation methods are applicable to various devices.

[0226] In other words, VR technology is a technology that only provides CG images of real-world objects, backgrounds, etc. AR technology, on the other hand, refers to the technology of displaying virtually created CG images on top of images of real objects. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology in that AR technology explicitly distinguishes between real objects and virtual objects created as CG images, using virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.

[0227] Recently, VR, AR, and MR technologies have sometimes been referred to as Extended Reality (XR) technologies without being clearly distinguished from each other. Therefore, embodiments of this disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are suitable for such technologies.

[0228] The PCC method / apparatus according to the implementation method can be applied to vehicles that provide autonomous driving services.

[0229] Vehicles providing autonomous driving services connect to the PCC device for wired / wireless communication.

[0230] When the point cloud data (PCC) transmitting / receiving device according to the embodiment is connected to a vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services that can be provided along with autonomous driving services and transmit it to the vehicle. When the PCC transmitting / receiving device is installed in the vehicle, it can receive / process content data related to AR / VR / PCC services based on user input signals input through a user interface device and provide it to the user. The vehicle or user interface device according to the embodiment can receive user input signals. User input signals according to the embodiment may include signals indicating autonomous driving services.

[0231] The method / apparatus for transmitting point cloud data according to the embodiments is configured to refer to Figure 1 Transmitting device 10000 Figure 1 Point cloud video encoder 10002 Figure 1 Transmitter 10003 and Figure 2 Get 20000 / Encode 20001 / Send 20002 Figure 4 encoder,Figure 1 The transmitting device ​ The device ​ and ​ Terms related to the encoding process, etc.

[0232] The method / apparatus for receiving point cloud data according to the embodiments is configured to refer to ​ The receiving device 10004 ​ Receiver 10005 ​ Point cloud video decoder 10006 ​ Sending 20002 / Decoding 20003 / Rendering 20004 ​ and ​ decoder ​ The receiving device ​ The device ​ and ​ Terms related to the decoding process, etc.

[0233] Furthermore, the method / apparatus for transmitting and receiving point cloud data according to the embodiments can be simply referred to as the method / apparatus according to the embodiments.

[0234] According to the implementation method, geometric data, geometric information, and location information constituting point cloud data are interpreted as having the same meaning. Attribute data, attribute information, and attribute information constituting point cloud data are also interpreted as having the same meaning.

[0235] ​ A method for sending point cloud data according to an implementation method is illustrated.

[0236] ​ Transmitting device 10000, point cloud video encoder 10002, ​ encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmission devices, etc., can be based on ​ The method shown compresses (encodes) point cloud data. ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device ​ Point cloud data receiving devices, etc., can be based on ​ The method shown decodes point cloud data.

[0237] The method / apparatus according to the implementation can compress / reconstruct point cloud geometric information with low latency. To ensure efficient processing in this operation, a prediction tree can be used. When generating the prediction tree, points can be sorted based on groups, and the prediction tree can be generated quickly based on the sorting order.

[0238] There may be scenarios where point cloud content needs to be encoded with low latency. For example, situations where point cloud data should be captured and transmitted from LiDAR in real time or 3D map data should be received and processed in real time could correspond to the aforementioned scenarios.

[0239] Therefore, the method / apparatus according to the embodiments relates to: a group-based point ranking method for enhancing the compression / reconstruction efficiency of geometric information encoding based on prediction trees, which can be applied to geometric compression / reconstruction of geometry-based point cloud compression (G-PCC) for cloud content requiring low latency encoding; and a method for rapidly generating prediction trees based on ranking order.

[0240] For example, implementation methods may provide a point sorting method, a method for quickly generating a prediction tree based on the sorting order, and a signaling method that supports the above methods to reduce the bitstream size when generating the prediction tree.

[0241] The implementation relates to a method for selecting an attribute predictor to increase the attribute compression efficiency of G-PCC for 3D point cloud data compression. Hereinafter, the encoder and encoding device are referred to as encoders, and the decoder and decoding device are referred to as decoders.

[0242] A point cloud consists of a set of points, and each point can include geometric and attribute information. Geometric information is three-dimensional position (XYZ) information, and attribute information is the value of color (RGB, YUV, etc.) and / or reflectance. G-PCC encoding operations can include compressed geometry and compressed attribute information based on the geometry reconstructed by reconstructing the position information changed by compression (reconstructed geometry = decoded geometry). ​ , ​ and ​ ).

[0243] G-PCC decoding operations (which correspond to G-PCC encoding operations) may include receiving the encoded geometric bitstream and attribute bitstream (see...). ​ ), decoded geometry, and decoded attribute information based on the geometry reconstructed through decoding operations (see...). ​ , ​ , ​ and ​ ).

[0244] For geometric information compression, compression techniques based on octrees, triangular soups, or prediction trees can be used.

[0245] Typical examples of point cloud services requiring low latency may include real-time navigation using 3D map point clouds, or real-time capture, compression, and transmission of point clouds via LiDAR devices.

[0246] As a key feature for improving encoders and decoders for low-latency services, it may be necessary to begin with the functionality of compressing a portion of the point cloud data. In octree-based geometry encoding, points are scanned and encoded in a width-first manner. On the other hand, in prediction tree-based geometry compression, which aims for low-latency geometry compression, the same operation can be performed in a depth-first manner to minimize stepwise point scanning. Predictions can be generated from the geometric information between parent and child nodes of the tree, and the residuals can be entropy-encoded to configure the geometric bitstream. Since the depth-first scheme does not require stepwise scanning of all points, geometry encoding can be performed progressively on the captured point cloud data without waiting to capture all the data.

[0247] However, because it performs the operation in a depth-first manner, the depth-first scheme can have a larger residual than octree-based geometric coding, and thus can increase the size of the geometric bitstream, in which all points are analyzed and efficiently encoded.

[0248] In prediction tree-based compression, points are sorted, and tree generation is performed based on the sorted point order. Therefore, the order of points can have a significant impact on tree generation. That is, for nodes in the prediction tree, points that are close to each other may not be set as parent / child nodes. Conversely, points in adjacent positions in the order of the point array are very likely to be set as parent / child nodes.

[0249] The method / apparatus according to the embodiments has the effect of reducing the size of the geometric bitstream by changing the sorting points.

[0250] When generating a prediction tree based on a KD-tree, it can take a long time to search for points that are geographically close to each other. This characteristic can be an obstacle to low-latency, real-time services. The method / apparatus according to the implementation can reduce the time required to generate the prediction tree.

[0251] According to the implementation method, the prediction tree-based geometry compression can be performed by the PCC geometry encoder of the PCC encoder, and the geometry can be reconstructed by the PCC geometry decoder in the PCC decoder.

[0252] 1500: The method for sending point cloud data according to the implementation method may include the following steps: sorting the points to generate a prediction tree.

[0253] Sort the points to generate a prediction tree.

[0254] The method / apparatus according to the implementation can sort the points in a specific manner to effectively generate a prediction tree.

[0255] Before generating the prediction tree, points are ordered sequentially based on Morton code, radius, azimuth, elevation, sensor ID, or by applying the captured time in sequence. The ordering method can be applied in various ways depending on the characteristics of the content.

[0256] Sorting can be applied differently depending on the multiple partitioning stages. For example, content in the form of spinning data captured by a LiDAR device, and points (geometric data of point cloud data), can be sorted based on azimuth angles suitable for the content. Then, points with the same azimuth angle can be sorted based on radius. Then, points with the same radius can be sorted based on elevation angle.

[0257] Depending on the implementation, the sorting direction can be specified. Sorting can be performed in ascending, descending, or both. For example, orientation can be applied in descending order, and radius can be applied in ascending order.

[0258] The method / apparatus according to the implementation can apply grouping at sorting points.

[0259] Grouping is a method of sorting groups by specifying a particular range and performing sorting within that range.

[0260] For example, when grouping points based on azimuth, points within a specific azimuth range can be included in a group, and points within a group can be grouped by radius, and points with the same radius can be sorted by elevation angle.

[0261] Depending on the implementation method, various grouping criteria, such as sorting criteria, can exist. Grouping can be applied based on Morton code, radius, azimuth, elevation, sensor ID, acquisition time, etc., and a range of values ​​can be defined for each group.

[0262] For example, when using Morton codes, x, y, and z values ​​can be shifted to perform grouping. When using azimuth, the decimal point of the radian value can be rounded. Ranges can be set for elevation, radius, sensor ID, or acquisition time.

[0263] After grouping, further sorting can be performed in the group at time n.

[0264] When a prediction tree is generated by sorting points according to similar characteristics in a distribution based on the fact that points within a specific range can have similar characteristics in the distribution, the residual can be reduced due to the similarity pattern, thereby reducing the size of the bit stream.

[0265] 1501: The method for sending point cloud data according to the embodiment may further include the step of generating a prediction tree.

[0266] The method for generating a prediction tree according to the implementation may include generating a prediction tree using a KD tree and a group-based prediction tree.

[0267] 1. Using KD-tree

[0268] In geometric information compression coding based on prediction trees, the operation of generating the tree based on the sorting order used to generate the prediction tree can be configured as follows pseudocode.

[0269] Points[]: An array of all points;

[0270] pointCount: The total number of points;

[0271] second_sorted_idexes[]: An array of indices for the final sorted points;

[0272] KDTree: A KD-tree used for neighbor node search.

[0273] for(i=0; I <pointCount;i++){

[0274] 1)P=Points[second_sorted_idexes[i]]

[0275] 2) Search for neighbors close to P in the KD-tree.

[0276] 3) If no adjacent neighbor is found as a search result, register P is used as a node in the KD-tree.

[0277] When there are no nodes in the KD tree initially, P can be registered in the KD tree.

[0278] 4) If a node exists as a search result, check the number of child nodes of the original node associated with the corresponding node in the KD-tree. If the number is less than or equal to 3, register the child nodes of the original node. If the number is greater than or equal to 3, check the next nearest node. If the number of nodes checked is less than or equal to 3, register the checked node.

[0279] Depending on the implementation method, the number of child nodes can be set differently.

[0280] 5) Register the prediction result of P (3, excluding itself) as a node of the KD tree.

[0281] }

[0282] A prediction tree can be generated from sorted points, containing the neighbors most similar to the current point as a child node. The prediction tree can represent the parent-child relationship of each point (node), and a predicted value for that point can be generated from the prediction tree.

[0283] The prediction result for P can be obtained from the prediction results from the parent node, the prediction results from the grandparent node and the parent node, and the prediction results from the great-grandparent node, the grandparent node and the parent node. The prediction value can be obtained efficiently from the prediction tree, which includes parent / child nodes generated by searching nearby neighbors.

[0284] 2. Group-based prediction tree

[0285] According to the implementation method, in order to quickly generate a prediction tree, the prediction tree can be constructed rapidly by configuring points in groups according to parent-child relationships based on the sorting order without a KD tree and configuring the nearest points between consecutive groups according to parent-child relationships. This operation can be configured using the following pseudocode.

[0286] idx = 0

[0287] group__sorted_indxes[][]: Indices of the points belonging to the group

[0288] for(i=0; I <groupCount;i++){

[0289] for(j=0;j <groupCount[i].size();j++){

[0290] 1)P=Points[group_sorted_indxes[i][j]]

[0291] The generation of the prediction tree begins with the number of groups and the points included in those groups.

[0292] For example, generation can start from point j in group i.

[0293] 2) If j! = 0, set Points[group_sorted_indxes[i][j-1]] as the parent of P.

[0294] Parent-child relationships can be quickly established by setting the point corresponding to j-1 (which is the previous value of j) as the parent of P.

[0295] 3) If j = 0 and i! = 0, then

[0296] (1) Find the point in group_sorted_indxes[i-1] that is close to P (e.g., based on the difference in radius, x / y / z distance, etc.) and set that point as the parent of P.

[0297] Since the parent item of group i is searched in the previously processed group (i-1), the parent-child relationship can be established quickly.

[0298] }

[0299] }

[0300] Each operation according to the embodiments can be performed by a point cloud data transmitting / receiving device according to the embodiments, which can be implemented by hardware, software, firmware or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories.

[0301] ​ An example of a point cloud data transmission device according to an embodiment is shown.

[0302] ​ Transmitting device 10000 ​ Point cloud video encoder 10002 ​ encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmission devices, etc., may include, for example, ​ The structure is shown. According to an embodiment, each element may correspond to a point cloud data transmission device that can be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories.

[0303] The data input unit 16000 can receive point cloud data. The data input unit can receive geometric data, attribute data, and / or parameter information sets of the point cloud data.

[0304] The coordinate transformer 16001 can transform the coordinate information of point cloud data to achieve encoding.

[0305] The geometric information transformation quantization processor 16002 can quantize geometric data. For example, quantization can be applied to geometric data based on quantization parameters.

[0306] The spatial divider 16003 can divide geometric data based on spatial units in order to encode the geometric data.

[0307] The geometric information encoder 16004 can encode geometric data. The geometric information encoder can voxelize the geometric data. According to the implementation, voxelization can be optional. That is, when performing a prediction tree for geometric encoding, the geometric data may or may not be voxelized.

[0308] When the geometry encoding type is prediction-based encoding, the geometry information encoder can generate a prediction tree through a prediction tree generator, generate a prediction tree through a prediction determiner, and perform rate-distortion optimization (RDO) to select the best prediction mode. Predicted geometry values ​​based on the best prediction mode can then be generated.

[0309] A geometric information entropy encoder can entropy encode the residuals relative to the predicted values ​​and configure the geometric information bitstream.

[0310] The prediction tree generator can generate prediction trees based on the following settings: point sorting method (e.g., no sorting, sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, sorting by captured time order, etc.), grouping application method (e.g., no sorting, sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, sorting by captured time order, and combinations thereof). (etc.), grouping range (e.g., applying Morton code grouping to round numbers, indicating values ​​in other ranges when shift values, radii, azimuth, elevation, or acquisition time are decimal values), sorting methods within groups (e.g., no sorting, sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, sorting by acquisition time order, etc.), and rapid application of prediction trees (generating prediction trees based on parent-child relationships within groups, and generating prediction trees based on parent-child relationships between adjacent groups).

[0311] The prediction tree generator can generate information about the point sorting method (pred_geom_tree_sorting_type), the grouping application method (pred_geom_tree_sorting_type), the grouping range (pred_geom_tree_grouping_n_digit), the sorting method within the group (pred_geom_tree_sorting_type), and the fast application of the prediction tree (pred_geom_tree_build_method), and send the information to the receiving side.

[0312] The prediction tree generator can receive input from a point sorting method based on the settings of the transmitting device, and can sort these points according to the sorting method. Point sorting methods can include sorting by Morton code order, sorting by radius order, sorting by azimuth angle, sorting by elevation angle, sorting by sensor ID order, sorting by acquisition time order, or combinations thereof. The applied sorting method can be included as parameter information in the bitstream and delivered to the decoder. The sorting operation can be divided into multiple operations, and the sorting can be applied differently in each operation. Multiple point cloud data can be sorted differently in corresponding operations.

[0313] The prediction tree generator can receive input regarding whether grouping should be applied to the point sorting method, as well as the method for grouping according to the settings of the transmitting device. Grouping and sorting methods can include sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, sorting by the time of acquisition, or combinations thereof. The applied grouping and sorting method can be included as parameter information in the bitstream and delivered to the decoder. The sorting operation can be divided into multiple operations, and grouping can be applied in each operation. Multiple point cloud data can be grouped differently in the corresponding operations.

[0314] The prediction tree generator can receive input of a range of packets based on the settings of the sending device.

[0315] For example, in Morton code grouping, the shift value can be a grouping range setting value. When Morton codes for multiple points are expressed bitwise, the Morton codes can be shifted by a shift value indicating the grouping range, so that multiple points can be grouped (or sorted) into the same group.

[0316] When azimuth, elevation, time, etc., have values ​​including decimal points, the number of digits to be rounded can be a grouping range setting value. When multiple points are expressed as data with decimal points (such as azimuth), the multiple points can be grouped or sorted into the same group by rounding the data at specific decimal places.

[0317] In the case of sensor IDs or other data, the range value can be a grouping range setting value. Points included within a specific range can be grouped or sorted into the same group.

[0318] According to the implementation method, the grouping range can be included as parameter information in the bitstream and delivered to the decoder.

[0319] The prediction tree generator can receive input for a method of sorting points in a group according to settings of the transmitting device. Points belonging to a group can be sorted according to a sorting method within the group. Point sorting methods can include sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, sorting by captured time order, or combinations thereof. The applied sorting method can be included as parameter information in the bitstream and delivered to the decoder.

[0320] The prediction tree generator can receive input indicating whether to apply a fast prediction tree generation method based on the sorting order according to the settings of the transmitting device. In fast generation, a prediction tree can be generated according to the sorting order within a group, and parent-child relationships can be established by finding points that are close to the first point of consecutive groups within the group. The indication of whether to apply fast prediction tree generation can be included as parameter information in the bitstream and delivered to the decoder.

[0321] The geometric information encoder can check the geometric encoding type. According to an implementation, the geometric encoding type can be set by the point cloud data transmitting device. For example, the optimal encoding type can be determined based on the point cloud data. The encoding type according to an implementation may include octree-based encoding, prediction-based encoding, and / or triangle soup-based encoding.

[0322] In octree-based encoding, octrees can be generated for geometric data, and geometric data can be encoded based on octrees.

[0323] In triangular soup-based encoding, triangular soups can be generated for geometric data, and geometric data can be encoded based on triangular soups.

[0324] In prediction-based coding, prediction trees can be generated according to the implementation method, and prediction data can be determined based on RDO.

[0325] For the prediction-based encoding according to the implementation method, prediction tree construction type information, point sorting method, maximum distance information, etc., can be input. Furthermore, information about the applied prediction tree construction type and information about the applied point sorting method can be sent to the receiving decoder device within the prediction tree.

[0326] The geometry information encoder can reconstruct geometric data, that is, location data. The reconstructed (recovered) geometry data can then be sent to the attribute information encoder for attribute encoding.

[0327] When the geometry encoding type is prediction-based encoding, the geometry information encoder can generate a prediction tree through a prediction tree generator, and perform RDO based on the prediction tree generated by a rate-distortion optimization (RDO)-based prediction determiner to select the best prediction mode. Therefore, predicted values ​​for the geometry based on the best prediction mode can be generated.

[0328] The 16005 geometric information entropy encoder can encode point cloud data based on an entropy scheme. The geometric information encoder can generate an encoded geometric information bitstream.

[0329] The geometric information entropy encoder 16005 can construct a geometric information bit stream by entropy encoding the residual between geometric data and predicted values.

[0330] The attribute information encoder 16006 can receive and encode attribute data from the data input unit. Since the attribute data (or attributes) depends on the geometric data (or location), the attribute data can be encoded based on the geometric data reconstructed by the geometric location reconstructor. The data can be transformed for color encoding, where color is attribute information. The attribute information encoder can encode the attribute data based on an entropy scheme.

[0331] Geometric information encoders can generate geometric information bitstreams by encoding geometric data based on octrees, prediction trees, trigonometric soups, etc., and attribute information encoders can encode attribute data to generate attribute information bitstreams.

[0332] In addition, encoders that include geometry information encoders and attribute information encoders can include information about geometry encoding and attribute encoding as parameter information in the bitstream and deliver the bitstream to the decoder.

[0333] ​ An example of a point cloud data receiving device according to an embodiment is shown.

[0334] ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device ​ Point cloud data receiving devices, etc., can be based on ​ The method shown decodes point cloud data. According to an embodiment, each element may correspond to a point cloud data receiving device, which may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories.

[0335] The point cloud data receiving device according to the embodiment can receive a bit stream including geometric data and attribute data, decode the geometric data, and decode the attribute data.

[0336] ​ The operation can be followed ​ The opposite of the corresponding operation.

[0337] The geometry information decoder 17000 can receive geometry information bit streams.

[0338] The geometric information decoder can decode geometric data according to the entropy scheme by the geometric information entropy decoder.

[0339] The geometric information decoder can check the geometric coding type. The encoder can check whether octree-based coding or prediction-based coding is applied, and perform the corresponding decoding.

[0340] When the encoder applies octree-based encoding, the geometric information decoder can perform octree-based decoding through the octree reconstructor.

[0341] When the encoder applies prediction-based encoding, the geometric information decoder can perform prediction tree-based decoding through the prediction tree reconstructor.

[0342] The prediction tree reconstructor can reconstruct the prediction tree based on the checkpoint sorting method, grouping application, grouping application method, grouping application range, and fast prediction tree application, and decode the predicted values ​​of the geometry, based on the parameter information included in the bitstream.

[0343] The geometry information decoder can reconstruct geometry data (or location), generate reconstructed geometry data, and deliver the reconstructed geometry data to the attribute information decoder.

[0344] The geometric information decoder can predict geometric data through the geometric information predictor, transform the predicted data through the geometric information transform inverse quantization processor, and perform inverse quantization on the transformed data.

[0345] The geometric information decoder can perform inverse transformations on the coordinates of geometric data.

[0346] A geometric information decoder can decode a geometric information bitstream and reconstruct the geometric information.

[0347] The attribute information decoder 17010 can decode the attribute information bitstream and reconstruct the attribute information. The attribute information decoder can then decode the attribute information based on the reconstructed geometric information.

[0348] The attribute information decoder can decode the residuals of attribute data delivered from the encoder using an entropy scheme and an attribute residual information entropy decoder.

[0349] The attribute information decoder can inverse quantize the residual attribute information using a residual attribute information inverse quantization processor.

[0350] The inverse color transformation processor can perform inverse transformation on attribute data (or color) and reconstruct attribute information.

[0351] ​ The structure of a bitstream including point cloud data according to an embodiment is shown.

[0352] ​ Transmitting device 10000 ​ Point cloud video encoder 10002 ​ encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmission devices, etc., can generate data such as... ​ The configuration shown includes a bitstream of point cloud data. ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device ​ Point cloud data receiving devices, etc., can be used for, for example ​ The bitstream containing point cloud data in the configuration shown is decoded.

[0353] For example, the method / apparatus according to the embodiment can use signals to notify PCC encoding-related information. Signaling information according to the embodiment can be used on either the transmitting or receiving side. The signaling information according to the embodiment can be generated and transmitted by the transmitting / receiving apparatus according to the embodiment (e.g., the metadata processor of the transmitting apparatus (which may be referred to as a metadata generator, etc.)), and can be received and obtained by the metadata parser of the receiving apparatus. Each operation of the receiving apparatus according to the embodiment can be performed based on the signaling information. The encoded point cloud configuration is described below.

[0354] The abbreviations for information included in point cloud data are: SPS (Sequence Parameter Set); GPS (Geometry Parameter Set); APS (Attribute Parameter Set); TPS (Tile Parameter Set); Geom (Geometry Bitstream), where the Geometry Bitstream may include a geometry slice header and / or geometry slice data; and Attr (Attribute Bitstream), where the Attribute Bitstream may include an attribute tile header and attribute tile data. According to an implementation, a slice can represent a coding unit. Slice-based geometry can be encoded in geometry encoding, and slice-based attributes can be encoded in attribute encoding. According to an implementation, data units can be added to the bitstream structure. Geometry / attribute data can be delivered on a per-data-unit basis.

[0355] The method / apparatus according to the implementation method can be used with ​ The relevant point sorting / prediction tree generation option information is added to the SPS or GPS of the bitstream of the point cloud data to be delivered, such as... ​ and ​ As shown.

[0356] The method / apparatus according to the implementation method can be used with ​ The operation-related point sorting / prediction tree generation option information is added to the TPS or geometry head of each slice to be delivered, such as... ​ and ​ As shown.

[0357] The method / apparatus according to the implementation can process point cloud data on the basis of tiles or slices, so that the point cloud can be divided into regions to be processed.

[0358] When point cloud data is divided into multiple regions, options can be set to generate different sets of neighbor points for each region, providing various options such as low-complexity and low-reliability results or high-complexity and high-reliability results. These options can be configured differently based on the receiver's processing capabilities.

[0359] Therefore, when a point cloud is divided into tiles, different options can be applied to the corresponding tiles. The settings information for each tile can be carried in the TPS information.

[0360] When a point cloud is divided into slices, different options can be applied to the corresponding slices. The settings information for each slice can be carried in the slice header information.

[0361] ​ The set of sequence parameters according to the implementation method is shown.

[0362] ​ Transmitting device 10000 ​ Point cloud video encoder 10002 ​encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmission devices, etc. ​ As shown, it can generate point cloud data and such as... ​ The bit stream of sequence parameter information related to point cloud data is shown in the configuration. ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device ​ Point cloud data receiving devices, etc., can process point cloud data and other data such as... ​ The bitstream of sequence parameter information related to the point cloud data as configured is decoded.

[0363] Optional information related to point sorting and prediction tree generation according to the implementation method can be included in the sequence parameter set.

[0364] `pred_geom_tree_sorting_num_steps`: This indicates the number of point sorting steps that will be applied to the corresponding sequence. The point sorting method used to predict the geometric tree can be applied to the corresponding sequence in multiple steps. The sequence can represent frames that include point cloud data.

[0365] pred_geom_tree_sorting_type: can indicate the stepwise sorting method to be applied when generating the predicted geometric tree from the corresponding sequence.

[0366] 0 = No sorting;

[0367] 1 = Sort according to Morton code order;

[0368] 2 = Sort by radius;

[0369] 3 = Sort by azimuth angle;

[0370] 4 = Sort by elevation angle;

[0371] 5 = Sort by sensor ID;

[0372] 6 = Sort by the time of capture.

[0373] The value can be changed to another value based on each sorting method.

[0374] pred_geom_tree_sorting_ascending_flag: When sorting points according to each step in generating a predictive geometric tree with the corresponding sequence, this flag indicates information about whether the points are sorted in ascending (true) or descending (false) order for the corresponding step.

[0375] pred_geom_tree_group_sorting_flag: Indicates whether group-based sorting is performed for each sorting step when generating the predictive geometric tree for the corresponding sequence.

[0376] `pred_geom_tree_grouping_n_digit`: Indicates the range of groups applied when performing progressive sorting groupings while generating the predicted geometric tree from the corresponding sequence. Although the syntax for applying the group range is the same, the application of the group range can vary depending on `pred_geom_tree_sorting_type`.

[0377] For example, the range of Morton code blocks can be expressed as shift values.

[0378] When data such as radius, azimuth, or elevation have decimal points, the grouping range can be expressed as the number of digits to be rounded. Alternatively, the grouping range can be expressed as a value indicating the range.

[0379] `pred_geom_tree_build_method`: This can indicate the method used to generate the predicted geometry tree from the corresponding sequence.

[0380] 0 = Prediction tree generation based on sorting order;

[0381] 1 = Distance-based KD-tree prediction tree generation;

[0382] 3 = Prediction tree generation based on the order of sorted groups.

[0383] Each integer value can be set differently.

[0384] profile_idc: This can indicate profile information about the bitstream according to the implementation method. Available candidate values ​​can be reserved through ISO / IEC.

[0385] profile_compatibility_flags: When equal to 1, this indicates that the bitstream conforms to the profile indicated by profile_idc.

[0386] sps_num_attribute_sets: Indicates the number of encoded attributes in the bitstream. It can have values ​​ranging from 0 to 63.

[0387] attribute_dimension[i]: Indicates the number of components of the i-th attribute. Attribute components can include color and reflectivity.

[0388] attribute_instance_id[i]: Indicates the instance ID of the i-th attribute.

[0389] The method / apparatus according to the embodiments can be based on the following... ​ The predictive geometric coding related parameters are used to deliver signaling information.

[0390] For example, it can be done through ​ The `pred_geom_tree_sorting_type` information indicates the sorting method applied to the points (sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, or sorting by the time of capture).

[0391] It can be done ​ The `pred_geom_tree_sorting_type` information indicates the sorting method applied to the grouping (sorting by Morton code order, sorting by radius order, sorting by azimuth order, sorting by elevation order, sorting by sensor ID order, or sorting by the time of capture).

[0392] The grouping range (which can receive shift values ​​for Morton code grouping methods, receive numbers to be rounded for data with decimal points such as azimuth angles, and allows input of the number of digits to be rounded and the range for other data) can be determined by... ​ The pred_geom_tree_grouping_n_digit information is used to indicate this.

[0393] Within a group, sorting methods (sorting by Morton code order, by radius order, by azimuth order, by elevation order, by sensor ID order, by captured time order, or a combination thereof) can be derived from... ​ The pred_geom_tree_sorting_type information is used to indicate this.

[0394] Whether to apply a group-based fast prediction tree generation method can be determined by... ​ The information is indicated by the pred_geom_tree_group_sorting_flag.

[0395] Sorting methods applied to points, grouping sorting methods applied to grouping, intra-group sorting methods, etc. can be distinguished by the pred_geom_tree_sorting_num_steps information. When there is one or more pred_geom_tree_sorting_num_steps, there can be one sorting step and two sorting steps (sorting_type[0], sorting_type[1]). In the case of multiple sorting types, the type can be applied first in a large range and then in the grouping range.

[0396] For example, in ​ , when there are multiple steps, if i = 0 in "for(i = 0; i < pred_geom_tree_sorting_num_steps; i++)", the pred_geom_tree_sorting_type can indicate the sorting method applied to these points. If i = 1, the pred_geom_tree_sorting_type can indicate the grouping sorting method applied to grouping. If i = 2, the pred_geom_tree_sorting_type can indicate the intra-group sorting method. The order of the information indicated by i can vary according to the implementation.

[0397] ​ Shows a set of geometric parameters according to an embodiment.

[0398] ​ The transmitting device 10000 of ​ The point cloud video encoder 10002 of ​ The encoding of ​ The encoder of ​ The transmitting device of ​ The device of ​ The point cloud data transmitting device, etc. can generate a bitstream including point cloud data and geometric parameter information related to the point cloud data configured as shown in ​ . ​ The receiving device 10004 of ​ The point cloud video decoder 10006 of ​ The decoding of ​ And ​ The decoder of ​ The receiving device of ​ The device of ​ The point cloud data receiving device, etc. can decode a bitstream including point cloud data and sequence parameter information related to the point cloud data configured as shown in ​ .

[0399] The method / apparatus according to the implementation can add relevant option information for point sorting / prediction tree generation functions to the geometric parameter set for delivery.

[0400] `pred_geom_tree_sorting_num_steps`: This indicates the number of point sorting steps that will be applied to the corresponding sequence. The point sorting method used to predict the geometric tree can be applied to the corresponding sequence in multiple steps. The sequence can represent frames that include point cloud data.

[0401] pred_geom_tree_sorting_type: can indicate the stepwise sorting method to be applied when generating the predicted geometric tree from the corresponding sequence.

[0402] 0 = No sorting;

[0403] 1 = Sort according to Morton code order;

[0404] 2 = Sort by radius;

[0405] 3 = Sort by azimuth angle;

[0406] 4 = Sort by elevation angle;

[0407] 5 = Sort by sensor ID;

[0408] 6 = Sort by the time of capture.

[0409] The value can be changed to another value based on each sorting method.

[0410] pred_geom_tree_sorting_ascending_flag: When sorting points according to each step in generating a predictive geometric tree with the corresponding sequence, this flag indicates information about whether the points are sorted in ascending (true) or descending (false) order for the corresponding step.

[0411] pred_geom_tree_group_sorting_flag: Indicates whether group-based sorting is performed for each sorting step when generating the predictive geometric tree for the corresponding sequence.

[0412] `pred_geom_tree_grouping_n_digit`: Indicates the range of groups applied when performing progressive sorting groupings while generating the predicted geometric tree from the corresponding sequence. Although the syntax for applying the group range is the same, the application of the group range can vary depending on `pred_geom_tree_sorting_type`.

[0413] For example, the range of Morton code blocks can be expressed as shift values.

[0414] When data such as radius, azimuth, or elevation have decimal points, the grouping range can be expressed as the number of digits to be rounded. Alternatively, the grouping range can be expressed as a value indicating the range.

[0415] `pred_geom_tree_build_method`: This can indicate the method used to generate the predicted geometry tree from the corresponding sequence.

[0416] 0 = Prediction tree generation based on sorting order;

[0417] 1 = Distance-based KD-tree prediction tree generation;

[0418] 3 = Prediction tree generation based on the order of sorted groups.

[0419] Each integer value can be set differently.

[0420] gps_geom_parameter_set_id: Indicates the identifier for the GPS used by other syntax elements. The value of gps_geom_parameter_set_id can be in the range of 0 to 15 (inclusive).

[0421] gps_seq_parameter_set_id: This indicates the value of sps_seq_parameter_set_id used for the active SPS. The value of gps_seq_parameter_set_id will be in the range of 0 to 15 (inclusive).

[0422] ​ The set of block parameters according to the implementation method is shown.

[0423] ​ Transmitting device 10000 ​ Point cloud video encoder 10002 ​ encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmitting devices, etc., can generate point cloud data and such ​ The bitstream of tile parameter information related to point cloud data is shown in the configuration. ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device​ Point cloud data receiving devices, etc., can process point cloud data and other data such as... ​ The bitstream of tile parameter information related to the point cloud data configured as shown is decoded.

[0424] The method / apparatus according to the implementation can add relevant option information for point sorting / prediction tree generation functions to the tile parameter set for delivery.

[0425] `pred_geom_tree_sorting_num_steps`: This indicates the number of point sorting steps that will be applied to the corresponding tile. The point sorting method used to predict the geometric tree can be applied to the corresponding tile in multiple steps. A tile can represent a frame unit used to process point cloud data.

[0426] pred_geom_tree_sorting_type: can indicate the stepwise sorting method to be applied when generating the predicted geometry tree from the corresponding tile.

[0427] 0 = No sorting;

[0428] 1 = Sort according to Morton code order;

[0429] 2 = Sort by radius;

[0430] 3 = Sort by azimuth angle;

[0431] 4 = Sort by elevation angle;

[0432] 5 = Sort by sensor ID;

[0433] 6 = Sort by the time of capture.

[0434] The value can be changed to another value based on each sorting method.

[0435] pred_geom_tree_sorting_ascending_flag: When sorting points according to each step when generating a predictive geometry tree with corresponding tiles, this flag indicates information about whether the points are sorted in ascending (true) or descending (false) order for the corresponding step.

[0436] pred_geom_tree_group_sorting_flag: Indicates whether group-based sorting is performed for each sorting step when generating the predictive geometry tree for the corresponding tile.

[0437] `pred_geom_tree_grouping_n_digit`: Indicates the range of groups applied when performing progressive sorting groupings while generating the predicted geometry tree from the corresponding tiles. Although the syntax for applying the group range is the same, the application of the group range can vary depending on `pred_geom_tree_sorting_type`.

[0438] For example, the range of Morton code blocks can be expressed as shift values.

[0439] When data such as radius, azimuth, or elevation have decimal points, the grouping range can be expressed as the number of digits to be rounded. Alternatively, the grouping range can be expressed as a value indicating the range.

[0440] pred_geom_tree_build_method: can indicate the method used to generate the predicted geometry tree from the corresponding tiles.

[0441] 0 = Prediction tree generation based on sorting order;

[0442] 1 = Distance-based KD-tree prediction tree generation;

[0443] 3 = Prediction tree generation based on the order of sorted groups.

[0444] Each integer value can be set differently.

[0445] gps_geom_parameter_set_id: Indicates the identifier for the GPS used by other syntax elements. The value of gps_geom_parameter_set_id can be in the range of 0 to 15 (inclusive).

[0446] gps_seq_parameter_set_id: This indicates the value of sps_seq_parameter_set_id for the active SPS. The value of gps_seq_parameter_set_id will be in the range of 0 to 15 (inclusive).

[0447] num_tiles: Indicates the number of tiles signaled for the bitstream. If it does not exist, num_tiles can be assumed to be 0.

[0448] tile_bounding_box_offset_x[i]: Indicates the x-offset of the i-th tile in Cartesian coordinates. If it does not exist, it can be inferred to be the x-offset of tile_bounding_box_offset_x[0].

[0449] tile_bounding_box_offset_y[i]: Indicates the y-offset of the i-th tile in Cartesian coordinates. If it does not exist, it can be inferred to be the y-offset of tile_bounding_box_offset_y[0].

[0450] tile_bounding_box_offset_z[i]: Indicates the z-offset of the i-th tile in Cartesian coordinates. If it does not exist, it can be inferred to be the z-offset of tile_bounding_box_offset_z[0].

[0451] ​ A geometric slice head according to an embodiment is shown.

[0452] ​ Transmitting device 10000 ​ Point cloud video encoder 10002 ​ encoding, ​ encoder, ​ The transmitting device ​ The device ​ Point cloud data transmitting devices, etc., can generate point cloud data and such ​ The bitstream of slice header information associated with point cloud data is shown in the configuration. ​ The receiving device 10004 ​ Point cloud video decoder 10006 ​ Decoding ​ and ​ decoder ​ The receiving device ​ The device ​ Point cloud data receiving devices, etc., can process point cloud data and other data such as... ​ The bitstream of slice header information related to point cloud data configured as shown is decoded.

[0453] The method / apparatus according to the implementation can add relevant option information for point sorting / prediction tree generation functions to the geometric slice header for delivery. A slice can be a unit for encoding and decoding point cloud data.

[0454] `pred_geom_tree_sorting_num_steps`: This indicates the number of point sorting steps that will be applied to the corresponding slice. The point sorting method used to predict the geometry tree can be applied to the corresponding slice in multiple steps.

[0455] pred_geom_tree_sorting_type: can indicate the stepwise sorting method to be applied when generating the predicted geometry tree from the corresponding slice.

[0456] 0 = No sorting;

[0457] 1 = Sort according to Morton code order;

[0458] 2 = Sort by radius;

[0459] 3 = Sort by azimuth angle;

[0460] 4 = Sort by elevation angle;

[0461] 5 = Sort by sensor ID;

[0462] 6 = Sort by the time of capture.

[0463] The value can be changed to another value based on each sorting method.

[0464] pred_geom_tree_sorting_ascending_flag: When sorting points according to each step when generating a predictive geometry tree with corresponding slices, this flag indicates information about whether the points are sorted in ascending (true) or descending (false) order for the corresponding step.

[0465] pred_geom_tree_group_sorting_flag: Indicates whether group-based sorting is performed for each sorting step when generating the predictive geometry tree for the corresponding slice.

[0466] `pred_geom_tree_grouping_n_digit`: Indicates the range of groups applied when performing progressive sorting groupings as the predictive geometry tree is generated from the corresponding slice. Although the syntax for applying the group range is the same, the application of the group range can vary depending on `pred_geom_tree_sorting_type`.

[0467] For example, the range of Morton code blocks can be expressed as shift values.

[0468] When data such as radius, azimuth, or elevation have decimal points, the grouping range can be expressed as the number of digits to be rounded. Alternatively, the grouping range can be expressed as a value indicating the range.

[0469] pred_geom_tree_build_method: can indicate the method used to generate the predicted geometry tree from the corresponding slice:

[0470] 0 = Prediction tree generation based on sorting order;

[0471] 1 = Distance-based KD-tree prediction tree generation;

[0472] 3 = Prediction tree generation based on the order of sorted groups.

[0473] Each integer value can be set differently.

[0474] gsh_geometry_parameter_set_id: Indicates the value of gps_geom_parameter_set_id for the active GPS.

[0475] gsh_tile_id: Indicates the tile ID referenced by GSH. The value of gsh_tile_id can be in the range of 0 to XX (inclusive).

[0476] gsh_slice_id: Identifies the ID of the slice header used for reference by other syntax elements. The value of gsh_slice_id can be in the range of 0 to XX (inclusive).

[0477] ​ A method for sending point cloud data according to an implementation method is illustrated.

[0478] 2300: The point cloud data transmission method according to the embodiment may include encoding the point cloud data. According to the embodiment, the encoding operation may include: ​ Point cloud video acquisition 10001, ​ Point cloud video encoder 10002 ​ To obtain 20000 ​ The code 20001 ​ Operation of point cloud data encoder ​ Operation of point cloud data transmission device ​ The processing of each device, ​ Point cloud data encoding methods ​ Operation of point cloud data transmission device and ​ Bitstream generation of point cloud data.

[0479] 2301: The point cloud data transmission method according to the embodiment may further include transmitting a bit stream including point cloud data. According to the embodiment, the transmission operation includes... ​ The operation of transmitter 10003 ​ Sending 20002 ​ Sending of geometric bitstream and attribute bitstream ​ Sending encoded point cloud data ​ Data transmission from each device ​ The transmission of geometric information bitstreams and / or attribute information bitstreams and ​ The bitstream used for sending point cloud data.

[0480] ​ A method for receiving point cloud data according to an embodiment is illustrated.

[0481] 2400: The point cloud data receiving method according to an embodiment may include receiving a bit stream comprising point cloud data. According to an embodiment, the receiving operation may include... ​ The operation of receiver 10005, according to ​ Sending 20002 and receiving ​ and ​ The reception of bitstreams including geometric bitstreams and attribute bitstreams. ​ Operation of receiver 13000 ​ Data reception by each device ​ The reception of geometric information bitstream and attribute information bitstream and ​ The reception of bit streams.

[0482] 2401: The point cloud data receiving method according to the embodiment may further include decoding the point cloud data. According to the embodiment, the decoding operation may include... ​ Operation of point cloud video decoder 10006 ​ Decoding 20003 ​ and ​ Decoding point cloud data, ​ Operation of point cloud data receiving device ​ Data processing performed by each device ​ Point sorting and prediction tree generation, ​ The operation of the point cloud data receiving device and based on the included ​ The parameter information in the bitstream is used for decoding the geometric data and / or attribute data included in the bitstream.

[0483] The method for transmitting point cloud data according to the embodiments may include the following steps: encoding the point cloud data; and transmitting a bit stream including the point cloud data.

[0484] In the encoding of point cloud data according to the embodiments, for example, geometric encoding may include predictive geometric encoding of geometric data based on a prediction tree.

[0485] According to the implementation method, the points in the geometric data can be sorted based on a sorting type, which may include at least one of Morton order, azimuth order, or radial distance order. The points can be sorted by rounding over at least one of the azimuth, radius, or elevation angle.

[0486] The encoding of geometric data according to the implementation method may include a point-generating prediction tree based on the sorting of the geometric data.

[0487] The encoding of geometric data according to the implementation method may include predictive data based on the generation of geometric data from a prediction tree.

[0488] The encoding of geometric data according to the implementation method may include residuals of the geometric data generated based on the predicted data.

[0489] According to the implementation method, the bitstream may include parameter information for predicting geometric coding.

[0490] The device for receiving point cloud data according to the embodiments may include: a receiver configured to receive a bit stream including point cloud data; and a decoder configured to decode the point cloud data.

[0491] According to an implementation, for example, a decoder configured to decode point cloud data can perform predictive geometry decoding on geometric data based on a prediction tree.

[0492] According to the implementation, the bitstream may include parameter information for predicting geometric coding. The points in the geometric data may be sorted based on a sorting type, which may include at least one of Morton order, azimuth order, or radial distance order. The points may be sorted by rounding over at least one of azimuth, radius, or elevation angle.

[0493] The method / apparatus according to the embodiments can provide the following effects.

[0494] In scenarios requiring low latency services, such as when point cloud data should be captured and transmitted in real-time from LiDAR, or when providing services for real-time reception and processing of 3D map data, it may be necessary to reduce the time required for encoding and decoding. As a method to reduce this time, encoding can begin using partial point cloud information even without providing complete point cloud information, and streaming and decoding can be performed. This reduces latency.

[0495] To reduce the latency of geometry-based point cloud compression (G-PCC) techniques for low-latency 3D point cloud data compression services while maintaining the bitstream size, implementations provide a group-based sorting method that orders points according to distribution characteristics. The sorting order can significantly influence the bitstream size by impacting prediction tree generation, particularly the settings of parent and child nodes. Additionally, a method for rapidly generating prediction trees is provided.

[0496] The implementation provides a sorting method based on the content characteristics of an efficient prediction tree generation method for G-PCC encoders / decoders used for 3D point cloud data compression. Therefore, the efficiency of geometric compression encoding / decoding can be enhanced.

[0497] Therefore, the transmitting method / apparatus according to the embodiments can efficiently compress point cloud data and transmit the compressed data, and also deliver signaling information for the data. Therefore, the receiving method / apparatus according to the embodiments can also efficiently decode / reconstruct point cloud data.

[0498] The implementation has been described from the perspective of methods and / or apparatus, and the descriptions of methods and apparatus can be applied to complement each other.

[0499] Although the accompanying drawings have been described separately for simplicity, new embodiments can be designed by combining the embodiments illustrated in the various figures. The design of a computer-readable recording medium on which a program for performing the above embodiments is recorded, as required by those skilled in the art, also falls within the scope of the appended claims and their equivalents. The apparatus and methods according to the embodiments are not limited to the configurations and methods of the above embodiments. Various modifications can be made to the embodiments by selectively combining all or some of the embodiments. Although preferred embodiments have been described with reference to the accompanying drawings, those skilled in the art will appreciate that various modifications and variations can be made to the embodiments without departing from the spirit or scope of this disclosure as described in the appended claims. Such modifications should not be understood as independent of the technical concept or perspective of the embodiments.

[0500] Various elements of the apparatus according to the embodiments can be implemented by hardware, software, firmware, or a combination thereof. Various elements in the embodiments can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiments, the components according to the embodiments can be implemented as separate chips. According to the embodiments, at least one or more components of the apparatus according to the embodiments can include one or more processors capable of executing one or more programs. The one or more programs can execute any one or more of the operations / methods according to the embodiments, or include instructions for executing them. Executable instructions for performing the methods / operations of the apparatus according to the embodiments can be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a transient CRM or other computer program product configured to be executed by one or more processors. Furthermore, the memory according to the embodiments can be used as a concept encompassing not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it can be implemented in a carrier-based manner, such as by transmission over the Internet. Furthermore, the processor-readable recording medium can be distributed across a network-connected computer system, allowing processor-readable code to be stored and executed in a distributed manner.

[0501] In this specification, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Similarly, “A, B” can mean “A and / or B”. Additionally, “A / B / C” can mean “at least one of A, B, and / or C”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Additionally, in this specification, the term “or” should be interpreted as indicating “and / or”. For example, the expression “A or B” can mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” as used in this document should be interpreted as indicating “additionally or alternatively”.

[0502] Terms such as "first" and "second" can be used to describe various elements of the embodiments. However, the various components according to the embodiments should not be limited to the above terms. These terms are only used to distinguish one element from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of these terms should be interpreted without departing from the scope of the various embodiments. Both a first user input signal and a second user input signal are user input signals, but they do not mean the same user input signal unless the context clearly indicates otherwise.

[0503] The terminology used to describe embodiments is for the purpose of describing particular embodiments and is not intended to limit the embodiments. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly specifies otherwise. The expression “and / or” is used to include all possible combinations of terms. Terms such as “comprising” or “having” are intended to indicate the presence of figures, numbers, steps, elements, and / or components and should be understood not to exclude the possibility of additional figures, numbers, steps, elements, and / or components. As used herein, conditional expressions such as “if” and “when” are not limited to optional cases but are intended to be interpreted as performing a related operation or interpreting a related definition when a specific condition is met.

[0504] Operations according to the embodiments described in this specification can be performed by a transmitting / receiving device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control the various operations described in this specification. The processor may be referred to as a controller, etc. In the embodiments, operations may be performed by firmware, software, and / or combinations thereof. Firmware, software, and / or combinations thereof may be stored in a processor or memory.

[0505] The operations according to the above embodiments can be performed by the transmitting and / or receiving devices according to the embodiments. The transmitting / receiving device includes a transmitter / receiver configured to transmit and receive media data, a memory configured to store instructions (program code, algorithms, flowcharts, and / or data) for processing according to the embodiments, and a processor configured to control the operation of the transmitting / receiving device.

[0506] The processor may be referred to as a controller, etc., and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above embodiments can be performed by the processor. Alternatively, the processor may be implemented as an encoder / decoder for the operations of the above embodiments.

[0507] The mode of the present invention

[0508] As described above, the relevant content has been described in the best mode for implementing the method.

[0509] Industrial applicability

[0510] As described above, the implementation methods can be applied in whole or in part to point cloud data transmission / reception devices and systems.

[0511] Those skilled in the art will understand that various changes or modifications can be made to the implementation methods within the scope of the implementation methods.

[0512] Therefore, the embodiments are intended to cover modifications and variations of this disclosure, provided that they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for transmitting point cloud data, the method comprising the following steps: Geometric data of point cloud data is encoded based on prediction trees. The prediction tree is generated based on the points in the point cloud data. The points are sorted based on information related to their radius; Encode the attribute data of the point cloud data; and Send a bitstream including the point cloud data. The bit stream includes information representing the number of sorting steps associated with the prediction tree, and information representing the grouping range applied to the radius for the point in the prediction tree.

2. The method according to claim 1, in, The step of encoding the geometric data generates predicted data based on the prediction tree.

3. The method according to claim 2, in, The step of encoding the geometric data generates residuals of the geometric data based on the predicted data.

4. A device for transmitting point cloud data, the device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Geometric data of point cloud data is encoded based on prediction trees. The prediction tree is generated based on the points in the point cloud data. The points are sorted based on information related to their radius; The attribute data of the point cloud data is encoded; and Send a bitstream including the point cloud data. The bit stream includes information representing the number of sorting steps associated with the prediction tree, and information representing the grouping range applied to the radius for the point in the prediction tree.

5. The device according to claim 4, in, The at least one processor is configured to generate prediction data of the geometric data based on the prediction tree.

6. The device according to claim 5, in, The at least one processor is configured to generate residuals of the geometric data based on the predicted data.

7. A method for receiving point cloud data, the method comprising the following steps: Receive a bitstream including point cloud data; The geometric data of the point cloud data is decoded based on the prediction tree. The prediction tree is generated based on the points in the point cloud data. Wherein, the points are sorted based on information related to radius; and Decode the attribute data of the point cloud data. The bit stream includes information representing the number of sorting steps associated with the prediction tree, and information representing the grouping range applied to the radius of the point in the prediction tree.

8. A device for receiving point cloud data, the device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Receive a bitstream including point cloud data; The geometric data of the point cloud data is decoded based on the prediction tree. The prediction tree is generated based on the points in the point cloud data. The points are sorted based on information related to their radius; and Decode the attribute data of the point cloud data. The bit stream includes information representing the number of sorting steps associated with the prediction tree, and information representing the grouping range applied to the radius for the point in the prediction tree.