Point cloud data encoding / decoding method and point cloud data sending method

By encoding and decoding the geometric structure and attributes of point cloud data, using the octree structure and attribute encoder, the problem of inefficient point cloud data processing is solved and efficient point cloud services are achieved.

CN120302059APending Publication Date: 2025-07-11LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510718060.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-03-23
Filing Date
2021-01-05
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process large amounts of point cloud data, resulting in latency and encoding/decoding complexity problems.

Method used

The method of encoding and decoding point cloud data is adopted, including encoding geometric structures and attributes, encoding using an octree and attribute encoder, sending a bitstream through a transmitter, and decoding using a decoder.

Benefits of technology

It realizes efficient processing of point cloud data, provides high-quality point cloud services, and supports VR, autonomous driving and other services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302059A_ABST
    Figure CN120302059A_ABST
Patent Text Reader

Abstract

The invention relates to a point cloud data encoding / decoding method and a point cloud data transmitting method. According to the point cloud data transmission method of the embodiment, the point cloud data can be encoded and transmitted. The point cloud data processing method according to an embodiment enables receiving point cloud data and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the original application number 202180008152.4 (International Application No.: PCT / KR2021 / 000063, application date: January 5, 2021, invention title: Point cloud data sending device, sending method, processing device and processing method). Technical Field

[0002] Embodiments provide for providing point cloud content to provide various services such as VR (Virtual Reality), augmented reality (AR), mixed reality (MR), and autonomous driving services to users. Background Art

[0003] Point cloud content is content represented by a point cloud, which is a set of points belonging to a coordinate system representing a three-dimensional space. Point cloud content can represent three-dimensionally configured media and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. However, tens of thousands to hundreds of thousands of point data are required to represent point cloud content. Therefore, a method for efficiently processing a large amount of point data is needed. Summary of the Invention

[0004] Technical Problem

[0005] Embodiments provide an apparatus and method for efficiently processing point cloud data. Embodiments provide a point cloud data processing method and apparatus for solving latency and encoding / decoding complexity.

[0006] The technical scope of the embodiments is not limited to the above-mentioned technical objectives and may be extended to other technical objectives that can be inferred by those skilled in the art based on all the content disclosed herein.

[0007] Technical Solution

[0008] Therefore, in order to efficiently process point cloud data, a method for sending point cloud data according to some embodiments may include the following steps: encoding point cloud data including geometric structures and attributes; and sending a bitstream including the encoded point cloud data. According to some embodiments, the geometric structure represents the positions of the points of the point cloud data, and according to some embodiments, the attributes include at least one of the color and reflectivity of the points. The step of encoding the point cloud data includes encoding the geometric structure and encoding the attributes based on a complete or partial octree of the encoded geometric structure.

[0009] A device for transmitting point cloud data according to some embodiments may include: an encoder configured to encode point cloud data including geometric structures and attributes; and a transmitter configured to transmit a bitstream including the encoded point cloud data. The geometric structure according to some embodiments represents the positions of the points of the point cloud data, and the attributes according to some embodiments include at least one of the color and reflectivity of the points. The encoder according to some embodiments includes a geometric structure encoder that encodes the geometric structure and an attribute encoder that encodes the attributes based on the complete or partial octree of the encoded geometric structure.

[0010] A method for processing point cloud data according to some embodiments may include the steps of: receiving a bitstream including point cloud data; and decoding the point cloud data based on signaling information included in the bitstream. The step of decoding the point cloud data includes decoding the geometric structure included in the point cloud data and decoding the attributes including at least one of the color and reflectivity of the points based on the complete or partial octree of the decoded geometric structure. The geometric structure according to some embodiments represents the positions of the points of the point cloud data.

[0011] A device for processing point cloud data according to some embodiments may include: a receiver that receives a bitstream including point cloud data; and a decoder that decodes the point cloud data based on signaling information included in the bitstream. The decoder includes: a geometric structure decoder that decodes the geometric structure included in the point cloud data; and an attribute decoder that decodes the attributes including at least one of the color and reflectivity of the points based on the complete or partial octree of the decoded geometric structure. The geometric structure according to some embodiments represents the positions of the points of the point cloud data.

[0012] Advantageous Effects

[0013] The device / method according to some embodiments processes point cloud data efficiently.

[0014] The device / method according to some embodiments provides high-quality point cloud services.

[0015] The device / method according to some embodiments provides point cloud content for providing general services such as VR services, autonomous driving services, etc. Brief Description of the Drawings

[0016] The drawings, which are incorporated in and constitute a part of this application, are included to provide a further understanding of the present disclosure and illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. To better understand the various embodiments described below, reference should be made to the following embodiments described in conjunction with the drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar components.

[0017] Figure 1 Shows an exemplary point cloud content providing system according to an embodiment;

[0018] Figure 2 Is a block diagram illustrating an operation of providing point cloud content according to an embodiment;

[0019] Figure 3 Illustrates an exemplary process of capturing a point cloud video according to an embodiment;

[0020] Figure 4 Illustrates an exemplary point cloud encoder according to an embodiment;

[0021] Figure 5 Shows an example of a voxel according to an embodiment;

[0022] Figure 6 Shows an example of an octree and occupancy code according to an embodiment;

[0023] Figure 7 Shows an example of a neighboring node pattern according to an embodiment;

[0024] Figure 8 Illustrates an example of a point configuration in each LOD according to an embodiment;

[0025] Figure 9 Illustrates an example of a point configuration in each LOD according to an embodiment;

[0026] Figure 10 Illustrates a point cloud decoder according to an embodiment;

[0027] Figure 11 Illustrates a point cloud decoder according to an embodiment;

[0028] Figure 12 Illustrates a transmitting device according to an embodiment;

[0029] Figure 13 Illustrates a receiving device according to an embodiment;

[0030] Figure 14 Illustrates an exemplary structure operable in combination with a point cloud data transmission / reception method / device according to an embodiment;

[0031] Figure 15 Is a flowchart illustrating a point cloud encoding example;

[0032] Figure 16 Illustrates an example of a point and its neighboring points;

[0033] Figure 17 Shows an exemplary bitstream structure diagram;

[0034] Figure 18 illustrates an example of signaling information according to an embodiment;

[0035] Figure 19 exemplifies an example of signaling information according to an embodiment;

[0036] Figure 20 exemplifies a method for encoding relevant weights according to an embodiment;

[0037] Figure 21 exemplifies an example of spatial scalability decoding;

[0038] Figure 22 illustrates an exemplary improved quantization weight derivation process;

[0039] Figure 23 is a flowchart exemplifying point cloud encoding according to an embodiment;

[0040] Figure 24 is a flowchart exemplifying a method for transmitting point cloud data according to an embodiment;

[0041] Figure 25 is a flowchart exemplifying a method for processing point cloud data according to an embodiment. Detailed Embodiments

[0042] Now, reference will be made in detail to the preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to explain the exemplary embodiments of the present disclosure, and not to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0043] Although most of the terms used in the present disclosure have been selected from the commonly used general terms in the art, some terms have been arbitrarily selected by the applicant, and their meanings will be explained in detail as needed in the following description. Therefore, the present disclosure should be understood based on the original meaning of the terms rather than their simple names or meanings.

[0044] Figure 1 illustrates an exemplary point cloud content providing system according to an embodiment.

[0045] Figure 1 The point cloud content providing system exemplified in [reference] may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 can perform wired or wireless communication to transmit and receive point cloud data.

[0046] The point cloud data transmission device 10000 according to an embodiment can obtain and process point cloud video (or point cloud content), and transmit the point cloud video (or point cloud content). According to an embodiment, the transmission device 10000 can include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or server. According to an embodiment, the transmission device 10000 can include a device, a robot, a vehicle, an AR / VR / XR device, a portable device, a household appliance, an Internet of Things (IoT) device, and an AI device / server configured to perform communication with a base station and / or other wireless devices using a radio access technology (e.g., 5G new radio access technology (NR), Long Term Evolution (LTE)).

[0047] The transmission device 10000 according to an embodiment includes a point cloud video acquirer 10001, a point cloud video encoder 10002, and / or a transmitter (or communication module) 10003.

[0048] The point cloud video acquirer 10001 according to an embodiment acquires a point cloud video through a processing process such as capturing, synthesizing, or generating. The point cloud video is point cloud content represented by a point cloud that is a set of points in a 3D space, and can be referred to as point cloud video data. The point cloud video according to an embodiment can include one or more frames. A frame represents a still image / picture. Therefore, the point cloud video can include point cloud images / frames / pictures, and can be referred to as a point cloud image, frame, or picture.

[0049] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment can include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiment. The point cloud video encoder 10002 can output a bitstream containing the encoded point cloud video data. The bitstream can contain not only the encoded point cloud video data, but also signaling information related to the encoding of the point cloud video data.

[0050] The transmitter 10003 according to an embodiment transmits a bitstream including encoded point cloud video data. The bitstream according to an embodiment is encapsulated in a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not shown in the figure, the transmitting device 10000 may include an encapsulator (or an encapsulation module) configured to perform an encapsulation operation. According to an embodiment, the encapsulator may be included in the transmitter 10003. According to an embodiment, the file or segment may be transmitted to the receiving device 10004 through a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to an embodiment is capable of performing wired / wireless communication with the receiving device 10004 (or the receiver 10007) through networks such as 4G, 5G, 6G, etc. Additionally, the transmitter may perform necessary data processing operations according to a network system (e.g., a 4G, 5G, or 6G communication network system). The transmitting device 10000 may transmit the encapsulated data in an on-demand manner.

[0051] The receiving device 10004 according to an embodiment includes a receiver 10007, a point cloud video decoder 10006, and / or a renderer 10005. According to an embodiment, the receiving device 10004 may include a device, a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server configured to perform communication with a base station and / or other wireless devices using radio access technologies (e.g., 5G New Radio (NR), Long Term Evolution (LTE)).

[0052] The receiver 10007 according to an embodiment receives a bitstream including point cloud video data or a file / segment in which the bitstream is encapsulated from a network or a storage medium. The receiver 10007 may perform necessary data processing according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver 10007 according to an embodiment may perform de-encapsulation on the received file / segment and output the bitstream. According to an embodiment, the receiver 10007 may include a de-encapsulator (or a de-encapsulation module) configured to perform a de-encapsulation operation. The de-encapsulator may be implemented as an element (or a component) separate from the receiver 10007.

[0053] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to the method of encoding the point cloud video data (e.g., in the reverse process of the operation of the point cloud video encoder 10002). Thus, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression encoding, which is the reverse process of point cloud compression. The point cloud decompression encoding includes G-PCC encoding.

[0054] The renderer 10005 renders the decoded point cloud video data. The renderer 10005 can output point cloud content by rendering not only the point cloud video data but also audio data. According to an embodiment, the renderer 10005 may include a display configured to display the point cloud content. According to an embodiment, the display may be implemented as a separate device or component rather than being included in the renderer 10005.

[0055] The arrow indicated by the dashed line in the figure represents the transmission path of the feedback information acquired by the receiving device 10004. The feedback information is information for reflecting the interactivity with the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). In particular, when the point cloud content is content of a service that requires interaction with the user (e.g., an autonomous driving service, etc.), the feedback information may be provided to the content sender (e.g., the transmitting device 10000) and / or the service provider. According to an embodiment, the feedback information may be used in the receiving device 10004 and the transmitting device 10000, or may not be provided.

[0056] The head orientation information according to the embodiment is information regarding the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information can be information regarding the area of the point cloud video that the user is viewing. The viewing point is the point through which the user is viewing the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is an area centered on the viewing point, and the size and shape of this area can be determined by the field of view (FOV). Therefore, in addition to the head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. In addition, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area in the point cloud video that the user gazes at, the gaze time, etc. According to the embodiment, the receiving device 10004 can send feedback information including the gaze analysis result to the sending device 10000. The feedback information according to the embodiment can be obtained during the rendering and / or display processing. The feedback information according to the embodiment can be obtained by one or more sensors included in the receiving device 10004. According to the embodiment, the feedback information can be obtained by the renderer 10005 or a separate external component (or device, part, etc.). Figure 1 The dashed line in [Figure] indicates the process of sending the feedback information obtained by the renderer 10005. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can send the feedback information to the sending device 10000. The sending device 10000 (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process the necessary data (e.g., the point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide the point cloud content to the user.

[0057] According to the embodiment, the sending device 10000 can be referred to as an encoder, sending device, transmitter, etc., and the receiving device 10004 can be referred to as a decoder, receiving device, receiver, etc.

[0058] In the Figure 1 point cloud content providing system according to the embodiment, the point cloud data processed (through a series of processes of acquisition / encoding / sending / decoding / rendering) can be referred to as point cloud content data or point cloud video data. According to the embodiment, the point cloud content data can be used as a concept covering metadata or signaling information related to the point cloud data.

[0059] Figure 1 The elements of the point cloud content providing system illustrated in [Figure] can be implemented by hardware, software, a processor, and / or a combination thereof.

[0060] Figure 2 is a block diagram illustrating an operation of providing point cloud content according to an embodiment.

[0061] Figure 2 The block diagram of Figure 1 illustrates the operation of the point cloud content providing system described in. As described above, the point cloud content providing system can process point cloud data based on point cloud compression coding (e.g., G-PCC).

[0062] A point cloud content providing system according to an embodiment (e.g., the point cloud transmitting device 10000 or the point cloud video acquirer 10001) can acquire a point cloud video (20000). The point cloud video is represented by point clouds belonging to a coordinate system for representing a 3D space. A point cloud video according to an embodiment may include a Ply (Polygon File Format or Stanford Triangle Format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. The geometry includes the position of the points. The position of each point can be represented by parameters (e.g., values of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system composed of the X, Y, and Z axes). The attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc. for each point). A point has one or more attributes. For example, a point may have one attribute, i.e., color, or two attributes, i.e., color and reflectivity. According to an embodiment, the geometry may be referred to as position, geometry information, geometry data, etc., and the attributes may be referred to as attributes, attribute information, attribute data, etc. The point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquirer 10001) can obtain point cloud data from information related to the acquisition process of the point cloud video (e.g., depth information, color information, etc.).

[0063] A point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the point cloud video encoder 10002) can encode the point cloud data (20001). The point cloud content providing system can encode the point cloud data based on point cloud compression coding. As described above, the point cloud data may include the geometry and attributes of the points. Therefore, the point cloud content providing system can perform geometry encoding for encoding the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding for encoding the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and the attribute bitstream according to an embodiment can be multiplexed and output as one bitstream. The bitstream according to an embodiment may also contain signaling information related to geometry encoding and attribute encoding.

[0064] A point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the transmitter 10003) may transmit encoded point cloud data (20002). As Figure 1 illustrated, the encoded point cloud data may be represented by a geometric structure bitstream and an attribute bitstream. Additionally, the encoded point cloud data may be transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric structure encoding and attribute encoding). The point cloud content providing system may encapsulate the bitstream carrying the encoded point cloud data and transmit the bitstream in the form of a file or a segment.

[0065] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the receiver 10007) may receive a bitstream containing the encoded point cloud data. Additionally, the point cloud content providing system (e.g., the receiving device 10004 or the receiver 10007) may demultiplex the bitstream.

[0066] A point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the encoded point cloud data (e.g., geometric structure bitstream, attribute bitstream) transmitted in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the point cloud video data based on the signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the geometric structure bitstream to reconstruct the positions of the points (geometric structure). The point cloud content providing system may reconstruct the attributes of the points by decoding the attribute bitstream based on the reconstructed geometric structure. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may reconstruct the point cloud video based on the reconstructed geometric structure and the decoded attributes based on the positions.

[0067] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the renderer 10005) may render the decoded point cloud data (20004). The point cloud content providing system (e.g., the receiving device 10004 or the renderer 10005) may use various rendering methods to render the geometric structure and attributes decoded through the decoding process. The points in the point cloud content may be rendered as vertices with a certain thickness, cubes with a specific minimum size centered at the corresponding vertex positions, or circles centered at the corresponding vertex positions. All or part of the rendered point cloud content is provided to the user through a display (e.g., a VR / AR display, a common display, etc.).

[0068] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004) may obtain feedback information (20005). The point cloud content providing system may encode / decode the point cloud data based on the feedback information. The feedback information and operations of the point cloud content providing system according to an embodiment are the same as those described in the reference Figure 1 and thus the detailed description thereof is omitted.

[0069] Figure 3 Illustrates an exemplary process of capturing a point cloud video according to an embodiment.

[0070] Figure 3 Illustrates the exemplary point cloud video capture process of the point cloud content providing system described in the reference Figures 1 to 2 and thus the detailed description thereof is omitted.

[0071] The point cloud content includes a point cloud video (image and / or video) representing objects and / or environments located in various 3D spaces (e.g., a 3D space representing a real environment, a 3D space representing a virtual environment, etc.). Therefore, a point cloud content providing system according to an embodiment may use one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), a projector (e.g., an infrared mode projector for obtaining depth information), LiDAR, etc. to capture the point cloud video. A point cloud content providing system according to an embodiment may extract the shape of the geometric structure formed by points in the 3D space from the depth information, and extract the attributes of each point from the color information to obtain the point cloud data. The image and / or video according to an embodiment may be captured based on at least one of an inward-facing technique and an outward-facing technique.

[0072] Figure 3 The left part of illustrates the inward-facing technique. The inward-facing technique refers to a technique of capturing an image of a central object using one or more cameras (or camera sensors) arranged around the central object. The inward-facing technique may be used to generate point cloud content that provides a 360-degree image of a key object to the user (e.g., VR / AR content that provides a 360-degree image of an object (e.g., a key object such as a character, player, object, or actor) to the user).

[0073] Figure 3 The right part of illustrates the outward-facing technique. The outward-facing technique refers to a technique of capturing an image of the environment around the central object rather than the central object using one or more cameras (or camera sensors) arranged around the central object. The outward-facing technique may be used to generate point cloud content for providing the surrounding environment as seen from the user's perspective (e.g., content representing the external environment of a user that can be provided to an autonomous vehicle).

[0074] As shown in the figure, point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate systems are different for each camera. Therefore, the point cloud content providing system can calibrate one or more cameras before the capture operations to set a global coordinate system. Additionally, the point cloud content providing system can generate point cloud content by synthesizing any image and / or video with the images and / or videos captured by the above capture techniques. The point cloud content providing system may not perform the capture operations described in Figure 3 when generating point cloud content representing a virtual space. The point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unwanted regions (e.g., the background), identify the space to which the captured images and / or videos are connected, and perform an operation to fill in spatial holes when there are spatial holes.

[0075] The point cloud content providing system can generate a piece of point cloud content by performing coordinate transformation on the points of the point cloud videos obtained from each camera. The point cloud content providing system can perform coordinate transformation on the points based on the coordinates of each camera position. Therefore, the point cloud content providing system can generate point cloud content representing a wide range of content or can generate point cloud content with a high density of points.

[0076] Figure 4 An exemplary point cloud encoder according to an embodiment is illustrated.

[0077] Figure 4 Shows Figure 1 an example of the point cloud video encoder 10002. The point cloud encoder reconstructs and encodes the point cloud data (e.g., the positions and / or attributes of the points) to adjust the quality of the point cloud content (e.g., losslessly, lossily, or near losslessly) according to the network conditions or the application. When the overall size of the point cloud content is large (e.g., for 30fps, a point cloud content of 60Gbps is given), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bit rate according to the network environment, etc., to provide the point cloud content.

[0078] As described with reference to Figures 1 to 2 the point cloud encoder can perform geometric structure encoding and attribute encoding. The geometric structure encoding is performed before the attribute encoding.

[0079] The point cloud encoder according to an embodiment includes a coordinate transformer (transform coordinates) 40000, a quantizer (quantize and remove points (voxelization)) 40001, an octree analyzer (analyze octree) 40002, a surface approximation analyzer (analyze surface approximation) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometric structure reconstructor (reconstruct geometric structure) 40005, a color transformer (transform color) 40006, an attribute transformer (transform attribute) 40007, a RAHT transformer (RAHT) 40008, a LOD generator (generate LOD) 40009, a lifting transformer (lifting) 40010, a coefficient quantizer (quantize coefficients) 40011, and / or an arithmetic encoder (arithmetic coding) 40012.

[0080] The coordinate transformer 40000, the quantizer 40001, the octree analyzer 40002, the surface approximation analyzer 40003, the arithmetic encoder 40004, and the geometric structure reconstructor 40005 may perform geometric structure encoding. The geometric structure encoding according to an embodiment may include octree geometric structure encoding, direct encoding, trisoup geometric structure encoding (trisoup geometry encoding), and entropy encoding. The direct encoding and the trisoup geometric structure encoding are selectively or combinatorially applied. The geometric structure encoding is not limited to the above examples.

[0081] As shown in the figure, the coordinate transformer 40000 according to an embodiment receives a position and transforms it into coordinates. For example, the position may be transformed into position information in a three-dimensional space (e.g., a three-dimensional space represented by an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment may be referred to as geometric structure information.

[0082] The quantizer 40001 according to an embodiment quantizes geometric structures. For example, the quantizer 40001 may quantize points based on the minimum position values of all points (e.g., the minimum values on each of the X, Y, and Z axes). The quantizer 40001 performs the following quantization operation: multiplying the difference between the minimum position value and the position value of each point by a preset quantization scaling value, and then finding the closest integer value by rounding the value obtained by the multiplication. Thus, one or more points may have the same quantized position (or position value). The quantizer 40001 according to an embodiment performs voxelization based on the quantized positions to reconstruct the quantized points. As in the case of pixels which are the smallest units containing 2D image / video information, the points of the point cloud content (or 3D point cloud video) according to an embodiment may be included in one or more voxels. The term voxel, which is a composite of volume and pixel, refers to a 3D cubic space generated when a 3D space is divided into units (unit = 1.0) based on axes representing the 3D space (e.g., the X axis, the Y axis, and the Z axis). The quantizer 40001 may match multiple sets of points in the 3D space with voxels. According to an embodiment, one voxel may include only one point. According to an embodiment, one voxel may include one or more points. To represent a voxel as one point, the position of the center of the voxel may be set based on the positions of one or more points included in the voxel. In this case, the attributes of all positions included in one voxel may be combined and assigned to the voxel.

[0083] The octree analyzer 40002 according to an embodiment performs octree geometric structure encoding (or octree encoding) to represent voxels in an octree structure. The octree structure represents points that match the octree structure with voxels.

[0084] The surface approximation analyzer 40003 according to an embodiment may analyze and approximate the octree. The octree analysis and approximation according to an embodiment is a process of analyzing a region containing multiple points to efficiently provide an octree and voxelization.

[0085] The arithmetic encoder 40004 according to an embodiment performs entropy encoding on the octree and / or the approximated octree. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric structure bitstream is generated.

[0086] The color transformer 40006, the attribute transformer 40007, the RAHT transformer 40008, the LOD generator 40009, the lifting transformer 40010, the coefficient quantizer 40011, and / or the arithmetic coder 40012 perform attribute encoding. As described above, a point can have one or more attributes. The attribute encoding according to an embodiment is equally applied to the attributes that a point has. However, when an attribute (e.g., color) includes one or more elements, the attribute encoding is independently applied to each element. The attribute encoding according to an embodiment includes color transform encoding, attribute transform encoding, region adaptive hierarchical transform (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) encoding, and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform) encoding. Depending on the point cloud content, the above-mentioned RAHT encoding, prediction transform encoding, and lifting transform encoding can be selectively used, or a combination of one or more encoding schemes can be used. The attribute encoding according to an embodiment is not limited to the above examples.

[0087] The color transformer 40006 according to an embodiment performs color transform encoding of the color value (or texture) included in the transform attributes. For example, the color transformer 40006 can transform the format of the color information (e.g., from RGB to YCbCr). The operation of the color transformer 40006 according to an embodiment can be optionally applied according to the color value included in the attributes.

[0088] The geometry reconstructor 40005 according to an embodiment reconstructs (decompresses) an octree and / or an approximate octree. The geometry reconstructor 40005 reconstructs the octree / voxel based on the result of analyzing the point distribution. The reconstructed octree / voxel can be referred to as the reconstructed geometry (restored geometry).

[0089] The attribute transformer 40007 according to an embodiment performs an attribute transform to transform an attribute based on the reconstructed geometry and / or the position where geometry encoding has not been performed. As described above, since an attribute depends on the geometry, the attribute transformer 40007 can transform an attribute based on the reconstructed geometry information. For example, based on the position value of the points included in the voxel, the attribute transformer 40007 can transform the attributes of the points at that position. As described above, when the position of the voxel center is set based on the position of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of the one or more points. When performing trisoup geometry encoding, the attribute transformer 40007 can transform an attribute based on the trisoup geometry encoding.

[0090] The attribute transformer 40007 can perform an attribute transformation by calculating the average value of the attributes or attribute values (e.g., the color or reflectivity of each point) of the neighboring points within a specific position / radius from the center position (or position value) of each voxel. The attribute transformer 40007 can apply weights according to the distance from the center to each point when calculating the average value. Thus, each voxel has a position and a calculated attribute (or attribute value).

[0091] The attribute transformer 40007 can search for neighboring points existing within a specific position / radius from the center position of each voxel based on a K-D tree or a Morton code. A K-D tree is a binary search tree and supports a data structure that can manage points based on position such that a nearest neighbor search (NNS) can be performed quickly. The Morton code is generated by representing the coordinates (e.g., (x, y, z)) representing the 3D positions of all points as bit values and mixing the bits. For example, when the coordinates representing the position of a point are (5, 9, 1), the bit values of the coordinates are (0101, 1001, 0001). Mixing the bit values in the order of z, y, and x according to the bit index produces 010001000111. This value is represented as a decimal number 1095. That is, the Morton code value of the point with coordinates (5, 9, 1) is 1095. The attribute transformer 40007 can sort the points based on the Morton code value and perform NNS through depth-first traversal processing. When NNS is required in another transformation process for attribute encoding after the attribute transformation operation, a K-D tree or a Morton code is used.

[0092] As shown in the figure, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.

[0093] The RAHT transformer 40008 according to an embodiment performs RAHT encoding for predicting attribute information based on the reconstructed geometric structure information. For example, the RAHT transformer 40008 can predict the attribute information of the nodes in the higher layer of the octree based on the attribute information associated with the nodes in the lower layer of the octree.

[0094] The LOD generator 40009 according to an embodiment generates a level of detail (LOD) to perform predictive transform encoding. The LOD according to an embodiment is the level of detail of the point cloud content. As the LOD value decreases, it indicates a decrease in the level of detail of the point cloud content. As the LOD value increases, it indicates an increase in the level of detail of the point cloud content. The points can be classified according to the LOD.

[0095] The lifting transformer 40010 according to an embodiment performs lifting transform encoding for transforming the point cloud attributes based on weights. As described above, lifting transform encoding can be optionally applied.

[0096] The coefficient quantizer 40011 according to the embodiment quantizes the attribute after the attribute encoding based on the coefficient.

[0097] The arithmetic encoder 40012 according to the embodiment encodes the quantized attributes based on arithmetic coding.

[0098] Although not shown in this figure, Figure 4 The elements of the point cloud encoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may perform the above Figure 4 At least one of the operations and / or functions of the elements of the point cloud encoder. In addition, one or more processors can operate or execute a set of software programs and / or instructions to perform Figure 4 The operation and / or functionality of the elements of the point cloud encoder. One or more memories according to an embodiment may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices).

[0099] Figure 5 An example of a voxel according to an embodiment is shown.

[0100] Figure 5 1 shows a voxel in a 3D space represented by a coordinate system consisting of three axes, namely, the X axis, the Y axis, and the Z axis. Figure 4 As described, the point cloud encoder (eg, quantizer 40001) may perform voxelization. A voxel refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on axes representing the 3D space (eg, X-axis, Y-axis, and Z-axis). Figure 5 An example of a voxel generated by an octree structure is shown, in which the octree consists of two poles (0, 0, 0) and (2 d ,2 d ,2 d ) is recursively subdivided. A voxel consists of at least one point. The spatial coordinates of a voxel can be estimated based on the positional relationship with the voxel group. As mentioned above, a voxel has properties like pixels of a 2D image / video (such as color or reflectivity). The details of the voxel are similar to those of the reference Figure 4 The details described are the same, so their description is omitted.

[0101] Figure 6 Examples of octrees and occupancy codes are shown according to an embodiment.

[0102] As reference Figures 1 to 4As described, a point cloud content providing system (point cloud video encoder 10002) or a point cloud encoder (e.g., octree analyzer 40002) performs octree geometry encoding (or octree encoding) based on an octree structure to efficiently manage regions and / or positions of voxels.

[0103] Figure 6 The upper part shows an octree structure. The 3D space of the point cloud content according to an embodiment is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, and Z-axis). The octree structure is created by recursively subdividing a cube axis-aligned bounding box defined by two extreme points (0, 0, 0) and (2 d , 2 d , 2 d ). Here, 2 d can be set to a value that constitutes the minimum bounding box surrounding all points of the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined by the following formula. In the following formula, (x int n , y int n , z int n ) represents the position (or position value) of a quantized point.

[0104]

[0105] As Figure 6 shown in the middle of the upper part of, the entire 3D space can be divided into eight spaces according to partitioning. Each divided space is represented by a cube having six faces. As Figure 6 shown in the upper right of, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). Thus, each space is divided into eight smaller spaces. The divided smaller spaces are also represented by cubes having six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.

[0106] Figure 6 The lower part of shows an octree occupancy code. An octree occupancy code is generated to indicate whether each of the eight divided spaces resulting from dividing a space contains at least one point. Thus, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the divided space, and the child nodes have values in units of 1 bit. Thus, the occupancy code is represented as an 8-bit code. That is, when at least one point is included in the space corresponding to the child node, the node is assigned the value 1. When no point is included in the space corresponding to the child node (the space is empty), the node is assigned the value 0. Since Figure 6The occupancy code shown is 00100001, so it indicates that the spaces corresponding to the third and eighth child nodes among the eight child nodes each contain at least one point. As shown in the figure, each of the third and eighth child nodes has 8 child nodes, and the child nodes are represented by 8-bit occupancy codes. The occupancy code of the third child node shown in the figure is 10000111, and the occupancy code of the eighth child node is 01001111. A point cloud encoder according to an embodiment (e.g., arithmetic encoder 40004) may perform entropy encoding on the occupancy code. To improve compression efficiency, the point cloud encoder may perform intra / inter-frame encoding on the occupancy code. A receiving device according to an embodiment (e.g., receiving device 10004 or point cloud video decoder 10006) reconstructs an octree based on the occupancy code.

[0107] A point cloud encoder according to an embodiment (e.g., Figure 4 the point cloud encoder or octree analyzer 40002) may perform voxelization and octree encoding to store the positions of points. However, points are not always evenly distributed in 3D space, so there will be specific regions where there are fewer points. Therefore, performing voxelization on the entire 3D space is inefficient. For example, when a specific region contains fewer points, voxelization does not need to be performed in the specific region.

[0108] Therefore, for the above specific regions (or nodes other than the leaf nodes of the octree), a point cloud encoder according to an embodiment may skip voxelization and perform direct encoding to directly encode the positions of the points included in the specific regions. The coordinates of the directly encoded points according to an embodiment are referred to as the direct coding mode (DCM). A point cloud encoder according to an embodiment may also perform trisoup geometry encoding based on a surface model, thereby reconstructing the positions of points in a specific region (or node) based on voxels. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. The direct encoding and trisoup geometry encoding according to an embodiment may be selectively performed. In addition, the direct encoding and trisoup geometry encoding according to an embodiment may be performed in combination with octree geometry encoding (or octree encoding).

[0109] To perform direct encoding, the option to use the direct mode to apply direct encoding should be enabled. The node to which direct encoding will be applied is not a leaf node, and there should be fewer than a threshold number of points within the specific node. In addition, the total number of points to which direct encoding will be applied should not exceed a preset threshold. When the above conditions are met, a point cloud encoder according to an embodiment (or arithmetic encoder 40004) may perform entropy encoding on the position (or position value) of the points.

[0110] A point cloud encoder according to an embodiment (e.g., the surface approximation analyzer 40003) may determine a specific level of the octree (a level less than the depth d of the octree), and the surface model may be used starting from that level to perform trisoup geometry encoding to reconstruct the positions of points in the region of the node based on voxels (trisoup mode). The point cloud encoder according to an embodiment may specify the level at which trisoup geometry encoding will be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in the trisoup mode. In other words, the point cloud encoder according to an embodiment may operate in the trisoup mode only when the specified level is less than the depth value of the octree. The 3D cubic region of the node at the specified level according to an embodiment is called a block. A block may include one or more voxels. The block or voxel may correspond to a brick. The geometry is represented as a surface within each block. The surface according to an embodiment may intersect each edge of the block at most once.

[0111] A block has 12 edges, so there are at least 12 intersection points in a block. Each intersection point is called a vertex (or apex point). A vertex existing along an edge is detected when there is at least one occupied voxel adjacent to that edge among all the blocks sharing the edge. An occupied voxel according to an embodiment refers to a voxel containing a point. The position of the vertex detected along the edge is the average position of the edges of all the voxels adjacent to that edge among all the blocks sharing the edge.

[0112] Once the vertex is detected, the point cloud encoder according to an embodiment may perform entropy encoding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the voxel position value (the relative position value within the edge). When applying trisoup geometry encoding, the point cloud encoder according to an embodiment (e.g., the geometry reconstructor 40005) may generate a restored geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization processing.

[0113] The vertices at the edges of the block determine the surface passing through the block. The surface according to an embodiment is a non-planar polygon. In the triangle reconstruction process, the surface represented by a triangle is reconstructed based on the starting point of the edge, the direction vector of the edge, and the position value of the vertex. The triangle reconstruction process is performed through the following steps: i) calculating the centroid value of each vertex, ii) subtracting the centroid value from each vertex value, and iii) estimating the sum of the squares of the values obtained by the subtraction.

[0114]

[0115] Estimate the minimum value of the sum and perform projection processing based on the axis with the minimum value. For example, when the element x is the minimum value, each vertex is projected onto the x-axis relative to the center of the block and projected onto the (y, z) plane. When the value obtained by projecting onto the (y, z) plane is (ai, bi), estimate the value of θ by atan2(bi, ai) and sort the vertices according to the value of θ. The following table shows the vertex combinations for creating triangles according to the number of vertices. The vertices are sorted from 1 to n. The following table shows that for four vertices, two triangles can be constructed according to the vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 among the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 among the sorted vertices.

[0116] Table 2-1 Triangles formed by vertices sorted as 1, …, n n triangles

[0117]

[0118] Perform upsampling processing to add points in the middle along the sides of the triangles and perform voxelization. The added points are generated based on the upsampling factor and the width of the block. The added points are called refined vertices. The point cloud encoder according to the embodiment can voxelize the refined vertices. In addition, the point cloud encoder can perform attribute encoding based on the voxelization position (or position value).

[0119] Figure 7 An example of the neighboring node pattern according to the embodiment is shown.

[0120] In order to improve the compression efficiency of the point cloud video, the point cloud encoder according to the embodiment can perform entropy encoding based on context-adaptive arithmetic coding.

[0121] As described in the reference Figures 1 to 6 described, the point cloud content providing system or the point cloud encoder (e.g., the point cloud video encoder 10002, the point cloud encoder, or Figure 4 the arithmetic encoder 40004) of can immediately perform entropy encoding on the occupancy code. In addition, the point cloud content providing system or the point cloud encoder can perform entropy encoding (intra-frame encoding) based on the occupancy code of the current node and the occupancy of the neighboring nodes, or perform entropy encoding (inter-frame encoding) based on the occupancy code of the previous frame. The frame according to the embodiment represents a set of point cloud videos generated simultaneously. The compression efficiency of the intra-frame encoding / inter-frame encoding according to the embodiment can depend on the number of neighboring nodes being referenced. When the number of bits increases, the operation becomes more complex, but the encoding may be biased to one side, thereby increasing the compression efficiency. For example, when given 3-bit context, 2 3 = 8 methods need to be used to perform encoding. The parts divided for encoding affect the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be satisfied.

[0122] Figure 7 Illustrates the process of obtaining an occupancy pattern based on the occupancy of neighboring nodes. The point cloud encoder according to an embodiment determines the occupancy of neighboring nodes of each node of the octree and obtains the value of the neighboring pattern. The occupancy pattern of the node is inferred using the neighboring node pattern. Figure 7 The left part of shows a cube corresponding to a node (the cube in the middle) and six cubes (neighboring nodes) sharing at least one face with the cube. The nodes shown in the figure are nodes at the same depth. The numbers shown in the figure respectively represent the weights (1, 2, 4, 8, 16, and 32) associated with the six nodes. The weights are assigned sequentially according to the positions of the neighboring nodes.

[0123] Figure 7 The right part of shows the neighboring node pattern values. The neighboring node pattern value is the sum of the values obtained by multiplying the weights of the occupied neighboring nodes (neighboring nodes with points). Therefore, the neighboring node pattern value is from 0 to 63. When the neighboring node pattern value is 0, this indicates that there is no node without points (unoccupied node) among the neighboring nodes of this node. When the neighboring node pattern value is 63, this indicates that all neighboring nodes are occupied nodes. As shown in the figure, since the neighboring nodes assigned weights 1, 2, 4, and 8 are occupied nodes, the neighboring node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform encoding according to the neighboring node pattern value (for example, when the neighboring node pattern value is 63, 64 kinds of encodings can be performed). According to an embodiment, the point cloud encoder can reduce the encoding complexity by changing the neighboring node pattern value (for example, based on a table through which 64 is changed to 10 or 6).

[0124] Figure 8 Illustrates an example of the point configuration in each LOD according to an embodiment.

[0125] As described with reference to Figures 1 to 7 Before performing attribute encoding, the encoded geometry is reconstructed (decompressed). When direct encoding is applied, the geometry reconstruction operation may include changing the placement of the points after direct encoding (for example, placing the points after direct encoding at the front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since the attributes depend on the geometry, the attribute encoding is performed based on the reconstructed geometry.

[0126] A point cloud encoder (e.g., the LOD generator 40009) can classify (reorganize) points by LOD. This figure shows the point cloud content corresponding to the LOD. The leftmost picture in this figure represents the original point cloud content. The second picture from the left in this figure represents the distribution of points in the lowest LOD, and the rightmost picture in this figure represents the distribution of points in the highest LOD. That is, the points in the lowest LOD are sparsely distributed, and the points in the highest LOD are densely distributed. That is, as the LOD rises in the direction indicated by the arrow at the bottom of this figure, the space (or distance) between points becomes narrower.

[0127] Figure 9 An example of the point configuration for each LOD according to an embodiment is illustrated.

[0128] As referred to Figures 1 to 8 described, a point cloud content providing system or a point cloud encoder (e.g., the point cloud video encoder 10002, Figure 4 the point cloud encoder or the LOD generator 40009) can generate LOD. The LOD is generated by reorganizing points into a set of refinement levels according to a set of LOD distance values (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.

[0129] Figure 9 The upper part of Figure 9 shows examples of points (P0 to P9) in the point cloud content distributed in 3D space. In Figure 9 the original order represents the order of points P0 to P9 before generating the LOD. In Figure 9 the LOD-based order represents the order of points generated according to the LOD. The points are reorganized by the LOD. Additionally, the high LOD contains the points belonging to the lower LOD. As shown in Figure 9 LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.

[0130] As referred to Figure 4 described, the point cloud encoder according to an embodiment can selectively or combinatorially perform predictive transform coding, lifting transform coding, and RAHT transform coding.

[0131] The point cloud encoder according to an embodiment can generate a predictor for points to perform predictive transform coding to set the prediction attribute (or prediction attribute value) of each point. That is, N predictors can be generated for N points. The predictor according to an embodiment can calculate a weight (= 1 / distance) based on the LOD value of each point, the indexed information about neighboring points existing within the distance set for each LOD, and the distance to the neighboring points.

[0132] The predicted attribute (or attribute value) according to the embodiment is set to the average value obtained by multiplying the attributes (or attribute values) of neighboring points (e.g., color, reflectivity, etc.) set in the predictor of each point by weights (or weight values) calculated based on the distances from the neighboring points. A point cloud encoder according to the embodiment (e.g., coefficient quantizer 40011) may quantize and inverse-quantize the residuals (which may be referred to as residual attributes, residual attribute values, or attribute prediction residuals) obtained by subtracting the predicted attributes (attribute values) from the attributes (attribute values) of each point. The quantization process is configured as shown in the following table.

[0133] Table: Attribute Prediction Residual Quantization Pseudocode

[0134]

[0135]

[0136] Table: Attribute Prediction Residual Inverse Quantization Pseudocode

[0137]

[0138] When the predictor of each point has neighboring points, a point cloud encoder according to the embodiment (e.g., arithmetic encoder 40012) may perform entropy coding on the residual attribute values after quantization and inverse quantization as described above. When the predictor of each point has no neighboring points, a point cloud encoder according to the embodiment (e.g., arithmetic encoder 40012) may perform entropy coding on the attributes of the corresponding points without performing the above operations.

[0139] A point cloud encoder according to the embodiment (e.g., lifting transformer 40010) may generate a predictor for each point, set the calculated LOD, register neighboring points in the predictor, and set weights according to the distances from the neighboring points to perform lifting transform coding. The lifting transform coding according to the embodiment is similar to the above-mentioned prediction transform coding, but the difference is that the weights are applied to the attribute values accumulatively. The process of applying weights to the attribute values accumulatively according to the embodiment is configured as follows.

[0140] 1) Create an array quantization weight (QW) for storing the weight values of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighboring nodes registered in the predictor by the weight of the predictor of the current point, and add the values obtained by the multiplication.

[0141] 2) Lifting prediction process: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.

[0142] 3) Create a temporary array called updateweight and update and initialize this temporary array to zero.

[0143] 4) Add the weights calculated by multiplying the weights for all predictors by the weights corresponding to the predictor indices stored in QW to the update weight array cumulatively as indices of neighboring nodes. Add the values obtained by multiplying the attribute values of the indices of the neighboring nodes by the calculated weights to the update array cumulatively.

[0144] 5) Boost the update process: Divide the attribute values of the update array for all predictors by the weight values of the update weight array of the predictor indices, and add the existing attribute values to the values obtained by the division.

[0145] 6) Calculate the predicted attributes by multiplying the attribute values updated by the boost update process for all predictors by the weights (stored in QW) updated by the boost prediction process. Quantize the predicted attribute values according to the point cloud encoder of the embodiment (e.g., coefficient quantizer 40011). Additionally, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute values.

[0146] The point cloud encoder according to the embodiment (e.g., RAHT transformer 40008) may perform RAHT transform encoding in which the attributes of higher-level nodes are predicted using the attributes associated with the nodes at lower levels in the octree. RAHT transform encoding is an example of intra-frame encoding of attributes performed by reverse scanning of the octree. The point cloud encoder according to the embodiment scans the entire region of the voxels and repeats the merging process of merging the voxels into larger blocks at each step until reaching the root node. The merging process according to the embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the upper nodes directly above the empty nodes.

[0147] The following equation represents the RAHT transform matrix. In this equation, represents the average attribute value of the voxels at level l. can be calculated based on and . and The weights of and

[0148]

[0149] Here, is the low-pass value and is used in the merging process at the next higher level. Represents a high-pass coefficient. The high-pass coefficient in each step is quantized and subjected to entropy coding (e.g., encoded by an arithmetic coder 400012). The weights are calculated as as follows by and to create a root node.

[0150]

[0151] Figure 10 Illustrates a point cloud decoder according to an embodiment.

[0152] Figure 10 The point cloud decoder illustrated in Figure 1 is an example of the point cloud video decoder 10006 described in Figure 1 and can perform operations that are the same as or similar to those of the point cloud video decoder 10006 illustrated in Figure 1 . As shown in the figure, the point cloud decoder can receive a geometry bitstream and an attribute bitstream included in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs the decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and the attribute bitstream and outputs the decoded attributes. The decoded geometry and the decoded attributes are used to reconstruct the point cloud content (decoded point cloud).

[0153] Figure 11 Illustrates a point cloud decoder according to an embodiment.

[0154] Figure 11 The point cloud decoder illustrated in Figure 10 is an example of the point cloud decoder illustrated in Figures 1 to 9 and can perform a decoding operation that is the inverse process of the encoding operation of the point cloud encoder illustrated in Figures 1 to 9 .

[0155] As referred to in Figure 1 and Figure 10 described, the point cloud decoder can perform geometry decoding and attribute decoding. The geometry decoding is performed before the attribute decoding.

[0156] The point cloud decoder according to an embodiment includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesize octree) 11001, a surface approximation synthesizer (synthesize surface approximation) 11002, a geometric structure reconstructor (reconstruct geometric structure) 11003, a coordinate inverse transformer (inverse transform coordinates) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantization unit (inverse quantization) 11006, a RAHT transformer 11007, a LOD generator (generate LOD) 11008, an inverse lifting unit (inverse lifting) 11009, and / or a color inverse transformer (inverse transform color) 11010.

[0157] The arithmetic decoder 110000, the octree synthesizer 11001, the surface approximation synthesizer 11002, the geometric structure reconstructor 11003, and the coordinate inverse transformer 11004 may perform geometric structure decoding. The geometric structure decoding according to an embodiment may include direct decoding and trisoup geometric structure decoding. The direct decoding and trisoup geometric structure decoding are selectively applied. The geometric structure decoding is not limited to the above examples and is performed as the inverse process of the geometric structure encoding described for reference. Figures 1 to 9 The inverse process of the geometric structure encoding described above is performed.

[0158] The arithmetic decoder 11000 according to an embodiment decodes the received geometric structure bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.

[0159] The octree synthesizer 11001 according to an embodiment may generate an octree by obtaining occupancy codes from the decoded geometric structure bitstream (or the information about the geometric structure obtained as the decoding result). The occupancy codes are configured as described in the reference. Figures 1 to 9 configured as described in detail.

[0160] When trisoup geometric structure coding is applied, the surface approximation synthesizer 11002 according to an embodiment may synthesize a surface based on the decoded geometric structure and / or the generated octree.

[0161] The geometric structure reconstructor 11003 according to an embodiment may regenerate a geometric structure based on the surface and / or the decoded geometric structure. As described in the reference, direct coding and trisoup geometric structure coding are selectively applied. Therefore, the geometric structure reconstructor 11003 directly imports and adds the position information of the points to which direct coding is applied. When trisoup geometric structure coding is applied, the geometric structure reconstructor 11003 may reconstruct the geometric structure by performing the reconstruction operations of the geometric structure reconstructor 40005 (e.g., triangle reconstruction, upsampling, and voxelization). The details are as described in the reference. Figures 1 to 9 When trisoup geometric structure coding is applied, the geometric structure reconstructor 11003 may reconstruct the geometric structure by performing the reconstruction operations of the geometric structure reconstructor 40005 (e.g., triangle reconstruction, upsampling, and voxelization). The details are as described in the reference. Figure 6The details of the description are the same, so the description thereof is omitted. The reconstructed geometric structure may include a point cloud picture or frame that does not contain attributes.

[0162] According to an embodiment, the coordinate inverse converter 11004 can obtain the position of a point by transforming coordinates based on the reconstructed geometric structure.

[0163] The arithmetic decoder 11005, inverse quantizer 11006, RAHT converter 11007, LOD generator 11008, inverse lifter 11009, and / or color inverse converter 11010 can perform the attribute decoding described in the reference Figure 10 According to an embodiment, the attribute decoding includes region adaptive hierarchical transform (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) decoding, and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform) decoding. The above three decoding schemes can be selectively used, or a combination of one or more decoding schemes can be used. The attribute decoding according to an embodiment is not limited to the above examples.

[0164] According to an embodiment, the arithmetic decoder 11005 decodes the attribute bitstream by arithmetic decoding.

[0165] According to an embodiment, the inverse quantizer 11006 inverse quantizes the information about the decoded attribute bitstream or attribute obtained as a decoding result, and outputs the inverse quantized attribute (or attribute value). The inverse quantization can be selectively applied based on the attribute encoding of the point cloud encoder.

[0166] According to an embodiment, the RAHT converter 11007, LOD generator 11008, and / or inverse lifter 11009 can process the reconstructed geometric structure and the inverse quantized attribute. As described above, the RAHT converter 11007, LOD generator 11008, and / or inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.

[0167] According to an embodiment, the color inverse converter 11010 performs inverse transform decoding to inverse transform the color value (or texture) included in the decoded attribute. The operation of the color inverse converter 11010 can be selectively performed based on the operation of the color converter 40006 of the point cloud encoder.

[0168] Although not shown in this figure, Figure 11 the elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can execute the above Figure 11at least one or more of the operations and / or functions of the components of the point cloud video decoder. Additionally, one or more processors may operate or execute a set of software programs and / or instructions to perform Figure 11 the operations and / or functions of the components of the point cloud decoder.

[0169] Figure 12 An exemplary transmitting device according to an embodiment is illustrated.

[0170] Figure 12 The transmitting device shown in Figure 1 is an example of the transmitting device 10000 (or Figure 4 the point cloud encoder) of Figure 12 The transmitting device illustrated in Figures 1 to 9 may perform one or more of the same or similar operations and methods as the operations and methods of the point cloud encoder described with reference to

[0171] According to an embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 may perform the same or similar operations and / or acquisition methods as the operations and / or acquisition method of the point cloud video acquirer 10001 (or the acquisition process 20000 described with reference to Figure 2 ).

[0172] The data input unit 12000, the quantization processor 12001, the voxelization processor 12002, the octree occupancy code generator 12003, the surface model processor 12004, the intra / inter-frame encoding processor 12005, and the arithmetic encoder 12006 perform geometric structure encoding. The geometric structure encoding according to an embodiment is the same or similar to the geometric structure encoding described with reference to Figures 1 to 9 and thus a detailed description thereof is omitted.

[0173] According to an embodiment, the quantization processor 12001 quantizes geometric structures (e.g., the position values of points). The operations and / or quantization of the quantization processor 12001 are the same or similar to the operations and / or quantization of the quantizer 40001 described with reference to Figure 4 . The details are the same as the details described with reference to Figures 1 to 9 .

[0174] According to an embodiment, the voxelization processor 12002 voxelizes the quantized position values of the points. The voxelization processor 120002 may perform operations and / or processes that are the same as or similar to the operations of the quantizer 40001 and / or the voxelization process described with reference to Figure 4 The details are the same as those described with reference to Figures 1 to 9 The details are the same as those described with reference to

[0175] According to an embodiment, the octree occupancy code generator 12003 performs octree encoding on the voxelized positions of the points based on the octree structure. The octree occupancy code generator 12003 may generate occupancy codes. The octree occupancy code generator 12003 may perform operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (or octree analyzer 40002) described with reference to Figure 4 and Figure 6 The details are the same as those described with reference to Figures 1 to 9 The details are the same as those described with reference to

[0176] According to an embodiment, the surface model processor 12004 may perform trisoup geometry encoding based on the surface model to reconstruct the positions of the points in a specific region (or node) based on voxels. The surface model processor 12004 may perform operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (e.g., surface approximation analyzer 40003) described with reference to Figure 4 The details are the same as those described with reference to Figures 1 to 9 The details are the same as those described with reference to

[0177] According to an embodiment, the intra / inter-frame encoding processor 12005 may perform intra / inter-frame encoding on the point cloud data. The intra / inter-frame encoding processor 12005 may perform encoding that is the same as or similar to the intra / inter-frame encoding described with reference to Figure 7 The details are the same as those described with reference to Figure 7 The details are the same as those described with reference to. According to an embodiment, the intra / inter-frame encoding processor 12005 may be included in the arithmetic encoder 12006.

[0178] According to an embodiment, the arithmetic encoder 12006 performs entropy encoding on the octree and / or approximate octree of the point cloud data. For example, the encoding scheme includes arithmetic coding. The arithmetic encoder 12006 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 40004.

[0179] The metadata processor 12007 according to an embodiment processes metadata regarding point cloud data (e.g., set values) and provides it to necessary processing procedures such as geometric structure encoding and / or attribute encoding. Additionally, the metadata processor 12007 according to an embodiment may generate and / or process signaling information related to geometric structure encoding and / or attribute encoding. The signaling information according to an embodiment may be encoded separately from geometric structure encoding and / or attribute encoding. The signaling information according to an embodiment may be interleaved.

[0180] The color transform processor 12008, the attribute transform processor 12009, the prediction / lifting / RAHT transform processor 12010, and the arithmetic coder 12011 perform attribute encoding. The attribute encoding according to an embodiment is the same as or similar to the attribute encoding described in the reference Figures 1 to 9 and thus a detailed description thereof is omitted.

[0181] The color transform processor 12008 according to an embodiment performs color transform encoding to transform the color values included in the attributes. The color transform processor 12008 may perform color transform encoding based on the reconstructed geometric structure. The reconstructed geometric structure is the same as the one described in the reference Figures 1 to 9 . Additionally, it performs operations and / or methods that are the same as or similar to the operations and / or methods of the color transformer 40006 described in the reference Figure 4 . A detailed description thereof is omitted.

[0182] The attribute transform processor 12009 according to an embodiment performs attribute transformation to transform the attributes based on the reconstructed geometric structure and / or positions where geometric structure encoding has not been performed. The attribute transform processor 12009 performs operations and / or methods that are the same as or similar to the operations and / or methods of the attribute transformer 40007 described in the reference Figure 4 . A detailed description thereof is omitted. The prediction / lifting / RAHT transform processor 12010 according to an embodiment may encode the transformed attributes by any one or a combination of RAHT encoding, prediction transform encoding, and lifting transform encoding. The prediction / lifting / RAHT transform processor 12010 performs at least one of the operations that are the same as or similar to the operations of the RAHT transformer 40008, the LOD generator 40009, and the lifting transformer 40010 described in the reference Figure 4 . Additionally, the prediction transform encoding, the lifting transform encoding, and the RAHT transform encoding are the same as those described in the reference Figures 1 to 9 and thus a detailed description thereof is omitted.

[0183] According to an embodiment, the arithmetic encoder 12011 can encode the encoded attributes based on arithmetic coding. The arithmetic encoder 12011 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 400012.

[0184] According to an embodiment, the transmission processor 12012 can transmit each bitstream including the encoded geometry and / or the encoded attributes and metadata information, or transmit a single bitstream configured with the encoded geometry and / or the encoded attributes and metadata information. When the encoded geometry and / or the encoded attributes and metadata information according to an embodiment are configured as a single bitstream, the bitstream can include one or more sub-bitstreams. The bitstream according to an embodiment can include signaling information, which includes a sequence parameter set (SPS) for signaling at the sequence level, a geometry parameter set (GPS) for signaling for geometry information encoding, an attribute parameter set (APS) for signaling for attribute information encoding, and a tile parameter set (TPS) for signaling for tile level and slice data. The slice data can include information about one or more slices. A slice according to an embodiment can include a geometry bitstream Geom0 0 and one or more attribute bitstreams Attr0 0 and Attr1 0 .

[0185] A slice refers to a series of syntax elements that represent all or part of an encoded point cloud frame.

[0186] The TPS according to an embodiment can include information about each tile in one or more tiles (e.g., coordinate information and height / size information about the border). The geometry bitstream can include a header and a payload. The header of the geometry bitstream according to an embodiment can include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, as well as information about the data included in the payload. As described above, the metadata processor 12007 according to an embodiment can generate and / or process the signaling information and send it to the transmission processor 12012. According to an embodiment, the elements for performing geometry encoding and the elements for performing attribute encoding can share data / information with each other, as indicated by the dashed line. The transmission processor 12012 according to an embodiment can perform operations and / or a transmission method that are the same as or similar to those of the transmitter 10003. The details are the same as those described with reference to Figure 1 and Figure 2 and are thus omitted from the description.

[0187] Figure 13 Illustrates an exemplary receiving device according to an embodiment.

[0188] Figure 13 The receiving device illustrated in Figure 1 is the receiving device 10004 (or Figure 10 and Figure 11 an example of the point cloud decoder). Figure 13 The receiving device illustrated in Figures 1 to 11 may perform one or more of the operations and methods that are the same as or similar to the operations and methods of the point cloud decoder described with reference to

[0189] The receiving device according to an embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an occupancy-code-based octree reconstruction processor 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / lifting / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each element for decoding according to an embodiment may perform an inverse process of the operation of the corresponding element for encoding according to an embodiment.

[0190] The receiver 13000 according to an embodiment receives point cloud data. The receiver 13000 may perform operations and / or a receiving method that are the same as or similar to the operations and / or the receiving method of the receiver 10007 in Figure 1 . A detailed description thereof is omitted.

[0191] The receiving processor 13001 according to an embodiment may obtain a geometry bitstream and / or an attribute bitstream from the received data. The receiving processor 13001 may be included in the receiver 13000.

[0192] The arithmetic decoder 13002, the occupancy-code-based octree reconstruction processor 13003, the surface model processor 13004, and the inverse quantization processor 1305 may perform geometry decoding. The geometry decoding according to an embodiment is the same as or similar to the geometry decoding described with reference to Figures 1 to 10 , and thus a detailed description thereof is omitted.

[0193] The arithmetic decoder 13002 according to an embodiment may decode the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.

[0194] The occupancy - code - based octree reconstruction processor 13003 according to an embodiment may reconstruct an octree by obtaining occupancy codes from a decoded geometric structure bitstream (or information about the geometric structure obtained as a decoding result). The occupancy - code - based octree reconstruction processor 13003 performs operations and / or methods that are the same as or similar to those of the octree synthesizer 11001 and / or the octree generation method. When applying trisoup geometric structure encoding, the surface model processor 13004 according to an embodiment may perform trisoup geometric structure decoding and related geometric structure reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on the surface model method. The surface model processor 13004 performs operations that are the same as or similar to those of the surface approximation synthesizer 11002 and / or the geometric structure reconstructor 11003.

[0195] The inverse quantization processor 13005 according to an embodiment may perform inverse quantization on the decoded geometric structure.

[0196] The metadata parser 13006 according to an embodiment may parse metadata (e.g., set values) included in the received point cloud data. The metadata parser 13006 may transmit the metadata for geometric structure decoding and / or attribute decoding. The metadata is the same as the metadata described in the reference Figure 12 and thus a detailed description thereof is omitted.

[0197] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lifting / RAHT inverse transform processor 13009, and the color inverse transform processor 13010 perform attribute decoding. The attribute decoding is the same as or similar to the attribute decoding described in the reference Figures 1 to 10 and thus a detailed description thereof is omitted.

[0198] The arithmetic decoder 13007 according to an embodiment may decode an attribute bitstream by arithmetic coding. The arithmetic decoder 13007 may decode the attribute bitstream based on the reconstructed geometric structure. The arithmetic decoder 13007 performs operations and / or coding that are the same as or similar to those of the arithmetic decoder 11005.

[0199] The inverse quantization processor 13008 according to an embodiment may perform inverse quantization on the decoded attribute bitstream. The inverse quantization processor 13008 performs operations and / or methods that are the same as or similar to those of the inverse quantizer 11006 and / or the inverse quantization method.

[0200] According to an embodiment, the prediction / lifting / RAHT inverse transform processor 13009 may process the reconstructed geometry and the inverse-quantized attributes. The prediction / lifting / RAHT inverse transform processor 13009 performs one or more of the operations and / or decoding that are the same as or similar to the operations and / or decoding of the RAHT transform 11007, the LOD generator 11008, and / or the inverse lifting 11009. According to an embodiment, the color inverse transform processor 13010 performs inverse transform encoding to inverse-transform color values (or textures) included in the decoded attributes. The color inverse transform processor 13010 performs the operations and / or inverse transform encoding that are the same as or similar to the operations and / or inverse transform encoding of the color inverse transform 11010. According to an embodiment, the renderer 13011 may render point cloud data.

[0201] Figure 14 An exemplary structure operable in conjunction with a point cloud data sending / receiving method / apparatus according to an embodiment is illustrated.

[0202] Figure 14 The structure represents a configuration in which at least one of the server 1460, the robot 17100, the autonomous vehicle 1420, the XR device 1430, the smart phone 1440, the home appliance 1450, and / or the head-mounted display (HMD) 1470 is connected to the cloud network 1400. The robot 1410, the autonomous vehicle 1420, the XR device 1430, the smart phone 1440, or the home appliance 1450 is referred to as a device. Additionally, the XR device 1430 may correspond to a point cloud data (PCC) device according to an embodiment or may be operatively connected to a PCC device.

[0203] The cloud network 1400 may represent a part of or exist in a cloud computing infrastructure. Here, the cloud network 1400 may be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0204] The server 1460 may be connected to at least one of the robot 1410, the autonomous vehicle 1420, the XR device 1430, the smart phone 1440, the home appliance 1450, and / or the HMD 1470 through the cloud network 1400 and may assist in at least part of the processing of the connected devices 1410 to 1470.

[0205] The HMD 1470 represents one of the implementation types of the XR device and / or the PCC device according to an embodiment. The HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power unit.

[0206] In the following, various embodiments of apparatuses 1410 to 1450 to which the above technologies are applied will be described. According to the above embodiments, Figure 14 The apparatuses 1410 to 1450 illustrated in [reference number] can be operably connected / linked to a point cloud data transmitting device and a receiver.

[0207] <PCC+XR>

[0208] The XR / PCC apparatus 1430 can adopt PCC technology and / or XR (AR+VR) technology and can be implemented as a head-mounted display (HMD), a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a household appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0209] The XR / PCC apparatus 1430 can analyze 3D point cloud data or image data obtained through various sensors or from an external device and generate position data and attribute data regarding 3D points. Accordingly, the XR / PCC apparatus 1430 can obtain information regarding the surrounding space or real objects and render and output XR objects. For example, the XR / PCC apparatus 1430 can match an XR object including auxiliary information regarding a recognized object with the recognized object and output the matched XR object.

[0210] <PCC + XR + Mobile Phone>

[0211] The XR / PCC apparatus 1430 can be implemented as a mobile phone (smartphone) 1440 by applying PCC technology.

[0212] The mobile phone 1440 can decode and display point cloud content based on PCC technology.

[0213] <PCC+Autopilot+XR>

[0214] The autonomous vehicle 1420 can be implemented as a mobile robot, a vehicle, a drone, etc. by applying PCC technology and XR technology.

[0215] The autonomous vehicle 1420 applying XR / PCC technology can represent an autonomous vehicle provided with a device for providing XR images or an autonomous vehicle that is a control / interaction target in an XR image. Specifically, the autonomous vehicle 1420 that is a control / interaction target in an XR image can be distinguished from the XR apparatus 1430 and can be operably connected to the XR apparatus 1730.

[0216] An autonomous vehicle 1420 having an apparatus for providing XR / PCC images can acquire sensor information from sensors including cameras and output XR / PCC images generated based on the acquired sensor information. For example, the autonomous vehicle 1420 can have a HUD and output XR / PCC images thereto, thereby providing an XR / PCC object corresponding to a real object or an object existing on a screen to an occupant.

[0217] When the XR / PCC object is output to the HUD, at least a part of the XR object can be output to overlap with a real object that the occupant's eyes are gazing at. On the other hand, when the XR / PCC object is output to a display provided inside the autonomous vehicle, at least a part of the XR / PCC object can be output to overlap with an object on the screen. For example, the autonomous vehicle 1220 can output an XR / PCC object corresponding to an object such as a road, another vehicle, a traffic signal, a traffic sign, a two-wheeled vehicle, a pedestrian, and a building.

[0218] Virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology according to an embodiment are applicable to various apparatuses.

[0219] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to a technology of showing a virtual created CG image on an image of a real object. The MR technology is similar to the above AR technology in that the virtual object to be shown is mixed and combined with the real world. However, the MR technology is different from the AR technology in that the AR technology clearly distinguishes between a real object and a virtual object created as a CG image and uses the virtual object as a supplementary object for the real object, while the MR technology regards the virtual object as an object having the same characteristics as the real object. More specifically, an example of the application of the MR technology is a hologram service.

[0220] Recently, VR, AR, and MR technologies are sometimes referred to as extended reality (XR) technologies without being clearly distinguished from each other. Therefore, the embodiments of the present disclosure are applicable to any one of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, G-PCC technologies is applicable to such technologies.

[0221] The PCC method / apparatus according to an embodiment can be applied to a vehicle providing an autonomous driving service.

[0222] A vehicle providing an autonomous driving service is connected to a PCC apparatus to perform wired / wireless communication.

[0223] When the point cloud data (PCC) transmitting / receiving device according to an embodiment is connected to a vehicle for wired / wireless communication, the device may receive / process content data related to AR / VR / PCC services that can be provided together with an autonomous driving service, and transmit it to the vehicle. In the case where the PCC transmitting / receiving device is installed on the vehicle, the PCC transmitting / receiving device may receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device, and provide it to the user. A vehicle or a user interface device according to an embodiment may receive a user input signal. The user input signal according to an embodiment may include a signal indicating an autonomous driving service.

[0224] As described in the reference Figures 1 to 14 described, the point cloud processing device according to an embodiment (e.g., Figure 1 , Figure 12 and Figure 14 the transmitting device or the point cloud encoder described therein) selectively uses RAHT coding, predictive transform coding, and lifting transform coding, or a combination of one or more of the coding techniques according to the point cloud content to perform attribute coding. For example, RAHT coding and lifting transform coding can be used for lossy coding, which greatly compresses the point cloud content data. Predictive transform coding can be used for lossless coding.

[0225] As described above, the point cloud encoder according to an embodiment may generate a predictor for a point and perform predictive transform coding to set the prediction attribute (or prediction attribute value) for each point. According to an embodiment, predictive transform coding and lifting transform coding calculate the distance (or position) of each neighboring point based on the positions of the points stored within the proximity range of the point cloud (hereinafter referred to as the target point). The calculated distance is used as a reference or reference weight to predict the attribute (e.g., color, reflectivity, etc.) of the target point, or to update the attribute (or prediction attribute) of the target point when the distance to the neighboring point changes. The following formula represents the attribute of the target point predicted based on the neighboring points.

[0226] [Equation 1]

[0227]

[0228] In the above formula, P represents the prediction attribute of the target point, and d1, d2, and d3 represent the distances to each of the three neighboring points of the target point. The corresponding distances are combined and used as a reference weight. C1, C2, and C3 (not shown in the formula) represent the attributes of the corresponding neighboring points. Shift is a parameter for adjusting the average energy or power level of the point. The value of Shift is controlled by the hardware voltage operating range of the encoder or decoder.

[0229] According to an embodiment, considering the correlation between points (e.g., neighboring points), the point cloud encoder may change the above reference weights. Specifically, the point cloud encoder uses the correlation weights calculated by changing the above reference weights according to a correlation combination method to highlight the eigenvalue at the distance between the target point and the neighboring points. Therefore, the point cloud encoder can ensure a performance gain proportional to the correlation between points. Additionally, the point cloud encoder can generate a predictor considering the degree of correlation without changing the structures of the RAHT coding, predictive transform coding, and lifting transform coding described Figures 1 to 14 while considering the degree of correlation without changing the structures of the RAHT coding, predictive transform coding, and lifting transform coding described

[0230] Figure 15 is a flowchart illustrating an example of point cloud coding.

[0231] As described in the reference Figures 1 to 14 a point cloud transmitting device or a point cloud encoder (e.g., Figure 4 the point cloud encoder described in

[0232] receives an attribute (15100). The point cloud encoder (e.g., the LOD generator 4009) generates an LOD to perform a predictive transform (15200). The point cloud encoder reorganizes the points into levels of detail to generate an LOD. Therefore, the larger the level value of the LOD, the more detailed the point cloud content. According to an embodiment, the LOD may include points grouped based on the distance between points. The point cloud encoder reorganizes the points based on an octree structure. An iterative generation algorithm applicable to octree decoding may be applied to the grouped points according to the position or order of the points (e.g., Morton code order, etc.). In each iteration sequence, one or more levels of detail R0, R1,..., Ri belonging to one LOD (e.g., LODi) are generated. That is, the level of the LOD is a combination of levels of detail.

[0233] Additionally, the point cloud encoder described in the reference Figures 1 to 14 supports spatially adaptive decoding. According to an embodiment, spatially adaptive decoding is performed on some or all of the geometric structure and / or attributes according to the decoding performance of a point cloud receiving device (e.g., Figure 1 the receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device of Figures 10 to 11 to provide decoding of point cloud content with various resolutions. Spatially adaptive decoding includes at least one of adaptive geometric structure decoding for the geometric structure and adaptive attribute decoding for the attributes. Adaptive attribute decoding includes at least one of the above RAHT coding, predictive transform coding, and lifting transform coding. Therefore, the point cloud encoder can perform attribute encoding to allow the point cloud receiving device (the point cloud decoder described in the reference Figure 13The described receiving device, etc.) performs adaptable attribute decoding. The LOD for supporting spatial adaptable decoding can be generated by searching for neighboring points through an approximate nearest neighbor search method for points from the lowest point to the highest point in the octree structure. The nearest neighbor points of the corresponding points in the current LOD (e.g., LODl) are searched from the LOD (e.g., LODl - 1) at a level lower than the current LOD. The LODs at levels lower than the current LOD are combinations of the refinement levels of R0, R1, ..., and R1 - 1.

[0234] Configure a specific LOD generation algorithm as follows. According to an embodiment, (P i ) i=1...N is called a set of positions associated with the points of the point cloud. According to an embodiment, (M i ) i=1...N is the Morton code associated with the set of positions. The parameters D0 and ρ are defined as the initial sampling distance and the distance ratio between LODs, respectively. The distance ratio is always greater than 1 (ρ > 1).

[0235] According to an embodiment, the points are sorted in ascending order according to the Morton code values of the points. According to an embodiment, the parameter I represents an array of point indices sorted according to the above - mentioned processing. The LOD generation algorithm is executed iteratively. In each iteration k, the points belonging to LODk are extracted, and predictors for the extracted points are generated starting from k equal to 0 until all points are assigned to an LOD. Below, a more detailed process is described.

[0236] The sampling distance D is initialized to the initial sampling distance D0. In the iterations k ranging from 0 to the number of LODs, L(k) is the set of indices of the points belonging to the k - th LOD, and O(k) is the set of points belonging to the LODs corresponding to levels higher than k. After L(k) and O(k) are initialized, the LOD assignment and residuals of the points are calculated iteratively and input sequentially. This process is repeated for all indices in the array I. Here, L(k) and O(k) can be calculated and used in the process of generating predictors associated with the points of L(k). According to an embodiment, R(k) is the set of points that need to be added to LOD(k - 1) to obtain LOD(k) and is represented as follows.

[0237] R(k) = L(k)\L(k - 1), where "\" is the difference operator.

[0238] For each point i in R(k), an algorithm for finding h neighboring points of point i in O(k) and calculating the normal distance and the associated linear distance associated with point i is configured as follows. According to an embodiment, h, as a user - defined parameter, represents a constant for adjusting the maximum number of neighboring points used for predicting point i.

[0239] The counter j is initialized to zero (j = 0).

[0240] For a point i in R(k), Mi represents the Morton code associated with the point i. Mj represents the Morton code associated with the j-th element in O(k).

[0241] When Mi is greater than or equal to Mj and j is less than the size of O(k) (M i ≥M j and j < SizeOf(O(k))), the counter j is incremented by 1 (j←j+1) and the distance between Mi and the point associated with the index in O(k) is measured. The point is within a specific search range [j - SR2, j + SR2], and h nearest neighbor points ((n1, n2,..., n h )) and the normal distance between each neighbor point and i are tracked

[0242] In addition, based on the normal distance, the correlation squared distance between two nearest squared distances is calculated. The correlation squared distance can be expressed as follows.

[0243]

[0244] The calculation method is not limited to the above examples.

[0245] When the correlation squared distance between the target point and the last processed point is less than the threshold, the neighbor points of the last processed point are used for initial estimation and search. According to an embodiment, the threshold can be defined by the user. Points with a correlation squared distance greater than the threshold are excluded.

[0246] The LOD generation algorithm is also applied to a point cloud receiving device (e.g., referring to Figure 10 , Figure 11 and Figure 13 the described point cloud decoder and receiving device). Therefore, the point cloud receiving device generates an LOD based on the above LOD generation algorithm.

[0247] The point cloud encoder performs transform coding (15300). As described in Figures 1 to 14 , according to an embodiment, the point cloud encoder selectively uses RAHT coding, predictive transform coding, and lifting transform coding or a combination of one or more of the coding techniques according to the point cloud content.

[0248] According to an embodiment, the predictive transform coding includes prediction based on interpolation. Attributes associated with the point cloud are encoded and decoded in the order defined by the processing according to the LOD. In each operation, only the points that have been encoded or decoded are considered for prediction. The attribute of a point is predicted based on a weighted average of the attributes (or attribute values) of the neighboring points of the point. However, the points in the neighboring point group may be distributed near or far from the point. Therefore, when a larger weight is assigned to the densely distributed points compared to the points that are less densely distributed or far from the point, the actual correlation between these points can be reflected, and thus the predicted attribute can be calculated more accurately. Therefore, the attribute (or attribute value) of a point according to an embodiment can be predicted based on the distance to the nearest neighboring points of the point and the interpolation-based prediction (or interpolation prediction transform processing) using weights.

[0249] According to an embodiment, (a i ) i∈0...k-1 represents an attribute (or an attribute value). represents a set of the k nearest neighboring points of the point. is the j-th decoded and reconstructed attribute. The following equation represents the process of calculating the correlation distance using a cyclic correlation shift matrix.

[0250] [Equation 2]

[0251]

[0252] In this equation, represents the distance between the point and its neighboring points. is the correlation distance (or referred to as the correlation value) calculated by applying a matrix. The weighted average predicted attribute (correlation weight) calculated based on the correlation distance calculated in the above equation is represented as follows

[0253] [Equation 3]

[0254]

[0255] ​According to an embodiment, the lifting transform coding uses an update operator to calculate the predicted attributes of each point. The lifting transform coding calculates the predicted values and residual values of the points belonging to the highest level of LOD (e.g., LODn). The lifting transform coding can calculate the weights of the points. The update operator of the points can calculate the updated attribute values based on the calculated weights and residual values. The calculated updated attribute values are used to calculate the predicted attributes of the points in the next LOD (e.g., LODn-1 which is one level lower than LODn). One level of LOD includes the points included in other LODs of higher levels. That is, since the points included in the lower level LODs are more frequently used for prediction, the lifting transform coding based on LOD has a greater impact on the points belonging to the lower level LODs. Therefore, the update operator can perform the update operation based on the weights updated by adding the weights of neighboring points to the weights of the corresponding points.

[0256] According to an embodiment, the lifting transform coding can use the updated weights that reflect the correlation between neighboring points. The update operator performs the update operation based on the updated weights. Hereinafter, the process of updating the weights based on the correlation between neighboring points will be described. The updated weights that reflect the correlation between neighboring points can be referred to as correlation weights.

[0257] w(P) is the weight associated with point p. The following recursive operation is used to calculate w(P).

[0258] For all points, the value of w(P) is defined as 1.

[0259] Traverse the points in the reverse order of the order defined in the LOD structure.

[0260] For each point Q(i,j) belonging to LOD(j), the weights of the neighboring points of the point are updated. The following represents the update process.

[0261] w(P)←w(P)+w(C[Q(i,j),j])α(p,Q(i,j))

[0262] Here, C[(Q(i,j),j] represents the correlation squared distance between the point Q(i,j) in the j-th set of the nearest neighboring points. Therefore, the correlation weight w(P) according to the embodiment is updated considering the correlation squared distance between neighboring points. The update operator updates the attribute values based on the correlation weights and the prediction residuals.

[0263] According to an embodiment, the update process can be performed by program instructions stored in one or more memories included in the point cloud sending device and the receiving device. According to an embodiment, the program instructions can be executed by the point cloud encoder and / or decoder (or processor), and cause the point cloud encoder and / or decoder to update the attribute values.

[0264] According to an embodiment, the point cloud encoder performs quantization (15400). Since the quantization is the same as the quantization described in the reference, a detailed description thereof will be skipped. The weights according to the correlation between points as described above can also be applied to the quantization. Figures 1 to 14 According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights according to the correlation between points as described above can also be applied to the arithmetic coding.

[0265] According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights according to the correlation between points as described above can also be applied to the arithmetic coding. Figures 1 to 14 According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights according to the correlation between points as described above can also be applied to the arithmetic coding.

[0266] Figure 16 An example of a point and its neighboring points is illustrated.

[0267] Figure 16 A point P and four neighboring points C1, C2, C3, and C4, which are the targets of the prediction transform coding described as a reference, are shown. As described in the reference, the point cloud encoder performs prediction transform coding (e.g., interpolation-based prediction). The attribute (or attribute value) of point p can be a weighted average of the attributes of C1, C2, C3, and C4, which are the nearest neighboring points of point p. In this figure, δ Figure 15 represents the distance between point p and neighboring point C1, and δ Figure 15 represents the distance between point p and neighboring point C2. δ j represents the distance between point p and neighboring point C1, and δ j+1 represents the distance between point p and neighboring point C2. δ j+2 represents the distance between point p and neighboring point C3, and δ j+3 represents the distance between point p and neighboring point C4. As indicated by the dashed line in Figure 16 , neighboring points C2, C3, and C4 are relatively densely arranged compared to neighboring point C1.

[0268] Therefore, the prediction transform coding according to the embodiment multiplies the distance between point p and the neighboring points illustrated in Figure 16 by the circular correlation shift matrix shown in Equation 2 to calculate a correlation distance (e.g., ) that reflects the correlation between the neighboring points. Therefore, the prediction transform coding calculates the weighted average attribute (Equation 3) in consideration of the correlation according to the density of the neighboring points.

[0269] Figure 17 An exemplary bitstream structure diagram is shown.

[0270] A point cloud processing device (e.g., the transmitting device described in the references Figure 1 , Figure 12 and Figure 14 ) can transmit the encoded point cloud data in the form of a bitstream. The bitstream is a series of bits that form a representation of the point cloud data (or point cloud frame).

[0271] Point cloud data (or a point cloud frame) can be segmented into tiles and slices.

[0272] The point cloud data can be segmented into a plurality of slices and encoded in a bitstream. A slice is a set of points and is represented as a series of syntax elements representing all or part of the encoded point cloud data. A slice may or may not be dependent on other slices. Additionally, a slice may include a geometric structure data unit and may include one or more attribute data units, or may not include an attribute data unit. As described above, attribute encoding is performed based on geometric structure encoding. Thus, an attribute data unit is based on the geometric structure data unit in the same slice. That is, a point cloud data receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) can process attribute data based on the decoded geometric structure data. Thus, the geometric structure data unit must come before the associated attribute data unit in a slice. The data units in a slice must be consecutive, and no order of the slices is specified.

[0273] A tile is a (three-dimensional) rectangle in the shape of a parallelepiped within a bounding box (e.g., the bounding box described with reference to Figure 5 ). The bounding box may contain one or more tiles. A tile may overlap completely or partially with another tile. A tile may include one or more slices.

[0274] Thus, a point cloud data transmitting device can process data corresponding to a tile according to importance and provide high-quality point cloud content. That is, a point cloud data transmitting device according to an embodiment can perform point cloud compression encoding with better compression efficiency and appropriate latency on data corresponding to an area important to a user.

[0275] According to an embodiment, the bitstream contains signaling information and a plurality of slices (slice 0,..., slice n). As shown in the figure, before the signaling information slice in the bitstream. Thus, a point cloud data receiving device can first obtain the signaling information and sequentially or selectively process the plurality of slices based on the signaling information. As shown in the figure, slice 0 includes a geometric structure data unit (Geom0 0 ) and two attribute data units (Attr0 0 and Attr1 0 ). Additionally, the geometric structure data unit comes before the attribute data units in the same slice. Thus, a point cloud data receiving device processes (decodes) the geometric structure data unit (or geometric structure data), and then processes the attribute data unit (or attribute data) based on the processed geometric structure data. According to an embodiment, the signaling information may be referred to as signaling data, metadata, etc., and is not limited to this example.

[0276] According to an embodiment, the signaling information includes a sequence parameter set (SPS), a geometry parameter set (GPS), and one or more attribute parameter sets (APS). The SPS encodes information about the entire sequence such as profiles and levels, and may include comprehensive information about the entire sequence (sequence level) such as picture resolution and video format. The GPS is information about geometry encoding applied to the geometry included in the sequence (bitstream). The GPS may include information about an octree (e.g., the octree described in Figure 6 the octree described in Figure 6 ) and information about the octree depth. The APS is information about attribute encoding applied to the attributes included in the sequence (bitstream). As shown in the figure, the bitstream includes one or more APSs (e.g., APS0, APS1,...) according to the identifier for identifying the attribute.

[0277] According to an embodiment, the signaling information may further include TPS. The TPS is information about tiles and may include information such as identifiers and tile sizes. According to an embodiment, the signaling information is information at the sequence level (i.e., bitstream level) and is applied to the corresponding bitstream. In addition, the signaling information has a syntax structure including a syntax element and a descriptor for describing the syntax element. Pseudo-code for describing the syntax may be used. In addition, the point cloud receiving device may sequentially parse and process the syntax elements in the syntax.

[0278] Although not shown in the figure, according to an embodiment, the geometry data unit and the attribute data unit respectively include a geometry header and an attribute header. According to an embodiment, the geometry header and the attribute header are signaling information applied at the corresponding slice level and have the above syntax structure.

[0279] According to an embodiment, the geometry header contains information (or signaling information) for processing the corresponding geometry data unit. Therefore, the geometry header first appears in the geometry data unit. The point cloud receiving device may first parse the geometry header to process the geometry data unit. The geometry header is related to the GPS that contains information about the entire geometry. Therefore, the geometry header contains information specifying the gps_geom_parameter_set_id included in the GPS. The geometry header also contains tile information (e.g., tile_id) and a slice identifier related to the slice to which the geometry data unit belongs.

[0280] According to an embodiment, the attribute header contains information (or signaling information) for processing the corresponding attribute data unit. Thus, the attribute header first appears in the attribute data unit. The point cloud receiving device may first parse the attribute header to process the attribute data unit. The attribute header is associated with the APS that contains information about all attributes. Thus, the attribute header contains information specifying the aps_attr_parameter_set_id included in the APS. As described above, attribute decoding is based on geometric structure decoding. Thus, to determine the geometric structure data unit associated with the attribute data unit, the attribute header contains information specifying the slice identifier included in the geometric structure header.

[0281] When the point cloud data processing device performs attribute encoding based on the relevant weights described in the reference Figures 15 to 16 the signaling information in the bitstream may include information about the relevant weights. According to an embodiment, the information about the relevant weights may be included in the signaling information at the sequence level (e.g., SPS, APS, etc.) or included in the slice level (e.g., the attribute header).

[0282] Figure 18 An example of signaling information according to an embodiment is shown.

[0283] Figure 18 Shown is the reference Figure 17 described syntax structure of the SPS, and an example is illustrated in which the information about the relevant weights described in the reference Figure 17 is included in the SPS at the sequence level.

[0284] The syntax of the SPS includes the following syntax elements.

[0285] The profile_compatibility_flags indicate whether the bitstream conforms to a specific profile for decoding or conforms to another profile. The profile specifies the constraints imposed on the bitstream to specify the ability to decode the bitstream. Each profile is a subset of algorithm features and constraints and is supported by all decoders that follow the profile. This is used for decoding and can be defined according to the standard.

[0286] The level_idc indicates the level applied to the bitstream. This level is used in all profiles. Generally, the level corresponds to a specific decoder processing load and memory capacity.

[0287] The sps_bounding_box_present_flag indicates whether there is information about the bounding box in the SPS. A sps_bounding_box_present_flag equal to 1 indicates the existence of information about the bounding box. A sps_bounding_box_present_flag equal to 0 indicates that the information about the bounding box is undefined.

[0288] The following is the information about the bounding box included in the SPS when the sps_bounding_box_present_flag is equal to 1.

[0289] The sps_bounding_box_offset_x indicates the quantized x-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0290] The sps_bounding_box_offset_y indicates the quantized y-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0291] The sps_bounding_box_offset_z indicates the quantized z-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0292] The sps_bounding_box_scale_factor specifies the scale factor used to indicate the size of the source bounding box.

[0293] The sps_bounding_box_size_width indicates the width of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0294] The sps_bounding_box_size_height indicates the height of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0295] The sps_bounding_box_size_depth indicates the depth of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.

[0296] The SPS syntax also includes the following elements.

[0297] The sps_source_scale_factor indicates the scale factor of the source point cloud data.

[0298] The sps_seq_parameter_set_id is the identifier of the SPS for reference by other syntax elements (e.g., the seq_parameter_set_id in GPS).

[0299] The sps_num_attribute_sets indicates the number of encoded attributes in the bitstream. The value of sps_num_attribute_sets is in the range of 0 to 64.

[0300] The following for statement includes as many elements indicating information about each attribute as the number indicated by sps_num_attribute_sets. In the figure, i represents each attribute (or attribute set), and the value of i is greater than or equal to 0 and less than the number indicated by sps_num_attribute_sets.

[0301] The attribute_dimension_minus1[i] indicates a value that is 1 less than the number of components of the i-th attribute. When the attribute is color, the attribute corresponds to a three-dimensional signal representing the light characteristics of the target point. For example, the attribute can be signaled as three components of RGB (red, green, blue). The attribute can be signaled as three components of YUV, namely luminance (illuminance) and two chrominances (saturation). When the attribute is reflectance, the attribute corresponds to a one-dimensional signal representing the intensity ratio of the light reflectance of the target point.

[0302] The attribute_instance_id[i] indicates the instance id of the i-th attribute. The attribute_instance_id is used to distinguish the same attribute label and attribute.

[0303] The attribute_bitdepth_minus1[i] indicates a value that is 1 less than the bit depth of the first component of the i-th attribute signal. Adding 1 to the value specifies the bit depth of the first component.

[0304] The attribute_cicp_colour_primaries[i] indicates the chromaticity coordinates of the color attribute source primaries of the i-th attribute.

[0305] The attribute_cicp_transfer_characteristics[i] either indicates the reference photoelectric transfer characteristic function of the color attribute as a function of the source input linear optical intensity Lc with a nominal true value range of 0 to 1, or indicates the reciprocal of the reference electro-optical transfer characteristic function as a function of the output linear light intensity Lo with a nominal true value range of 0 to 1.

[0306] The attribute_cicp_matrix_coeffs[i] indicates the matrix coefficients used to derive the luminance and chrominance signals from the RBG or YXZ primaries.

[0307] The attribute_cicp_video_full_range_flag[i] indicates the range and black level of the luminance and chrominance signals derived from the true value component signals of E′Y, E′PB, and E′PR or E′R, E′G, and E′B.

[0308] As described above, the SPS syntax contains information about the relevant weights. The following elements indicate information about the relevant weights with respect to the reference Figure 15 and Figure 16 information about the relevant weights described.

[0309] The attribute_correlated_weight_flag indicates whether the calculated distance should be correlated. The attribute_correlated_weight_flag equal to 1 indicates that the calculated distance should be correlated (i.e., using the correlated weights). The attribute_correlated_weight_flag equal to 0 indicates that the calculated distance is not correlated. The default value is inferred to be 0.

[0310] The attribute_correlated_weight_method indicates the method for calculating the correlated weights when the attribute_correlated_weight_flag is equal to 1. For example, the methods for calculating the correlated weights include those described by Equations 2 to 4 with respect to the reference Figure 15 described. Therefore, the point cloud receiving device calculates the correlated weights according to the method indicated by the attribute_correlated_weight_method and performs interpolation-based prediction and lifting transform coding.

[0311] The information about the correlated weights according to the embodiment is not limited to the above examples. Therefore, the information about the correlated weights may also include information about the correlated weight variables and information about the number of neighboring nodes in the correlated point set.

[0312] According to an embodiment, the syntax of the SPS includes the following syntax elements.

[0313] The known_attribute_label_flag[i], known_attribute_label[i], and attribute_label_fourbytes[i] are used together to identify the data type carried in the i-th attribute. The known_attribute_label_flag[i] indicates whether the attribute is identified by the value of the known_attibute_label[i] or by another object identifier attribute_label_fourbytes[i].

[0314] The sps_extension_flag indicates whether there is a sps_extension_data_flag in the SPS. A sps_extension_flag equal to 0 indicates that the sps_extension_data_flag syntax element does not exist in the SPS syntax structure. The value 1 of the sps_extension_flag is reserved for future use. The decoder may ignore all sps_extension_data_flag syntax elements following a sps_extension_flag equal to 1.

[0315] The sps_extension_data_flag indicates whether there is data for future use and can have any value.

[0316] The SPS syntax is not limited to the above examples. For signaling efficiency, it may also include additional elements, or some of the elements shown in the figure may be excluded. Some elements may be signaled by signaling information other than the SPS (e.g., APS, attribute headers, etc.) or by an attribute data unit.

[0317] Figure 19 An example of signaling information according to an embodiment is illustrated.

[0318] Figure 19 is the syntax structure of the APS described with reference to Figure 17 and illustrates an example in which information about the relevant weights described with reference to Figure 17 is included in the APS at the sequence level.

[0319] The syntax of the APS includes the following syntax elements.

[0320] The aps_attr_parameter_set_id indicates the identifier of the APS for reference by other syntax elements. The value of the aps_attr_parameter_set_id is in the range of 0 to 15. One or more attribute data units are included in the bitstream (e.g., the bitstream described with reference to Figure 17 ), and each attribute data unit includes an attribute header. The attribute header includes a field having the same value as the aps_attr_parameter_set_id (e.g., ash_attr_parameter_set_id). The point cloud receiving device according to the embodiment parses the APS and processes the attribute data units referring to the same aps_attr_parameter_set_id based on the parsed APS and the attribute header.

[0321] The aps_seq_parameter_set_id specifies the value of the sps_seq_parameter_set_id that activates the SPS. The value of the aps_seq_parameter_set_id is in the range of 0 to 15.

[0322] The attr_coding_type indicates the attribute coding type for a given value of attr_coding_type. Attribute coding means attribute encoding. As described above, attribute encoding uses at least one of RAHT coding, predictive transform coding, and lifting transform coding, and the attr_coding_type indicates any one of the three coding types mentioned above. Therefore, the value of the attr_coding_type is equal to any one of 0, 1, or 2 in the bitstream. Other values of the attr_coding_type may be used by ISO / IEC subsequently. Thus, the point cloud receiving device according to the embodiment ignores the attr_coding_type having a value other than 0, 1, and 2. When the attr_coding_type is equal to 0, the attribute coding type is predictive transform coding. When the attr_coding_type is equal to 1, the attribute coding type is RAHT coding. When the attr_coding_type is equal to 2, the attribute coding type is lifting transform coding. The value of the attr_coding_type may change and is not limited to this example. For example, an attr_coding_type equal to 0 indicates that the attribute coding type is RAHT coding, an attr_coding_type equal to 1 indicates that the attribute coding type is LOD with predictive transform coding, and an attr_coding_type equal to 2 indicates that the attribute coding type is LOD with lifting transform coding.

[0323] The aps_attr_initial_qp indicates the initial value of the variable SliceQp for each slice with reference to the current APS.

[0324] The aps_attr_chroma_qp_offset specifies the offset applied to the initial quantization parameter signaled by the aps_attr_initial_qp.

[0325] The aps_slice_qp_delta_present_flag indicates whether there is a component QP offset indicated by the ash_attr_qp_offset in the header of the attribute data unit.

[0326] As described above, the APS syntax contains information about the reference Figure 15and Figure 16 information on the relevant weights described. As shown in the figure, the APS syntax includes the attribute_correlated_weight_flag and the attribute_correlated_weight method. Each element is the same as the element described with reference to Figure 18 and thus the description thereof is skipped.

[0327] The information on the relevant weights according to the embodiment is not limited to the above examples. Thus, the information on the relevant weights may also include information on the relevant weight variables and information on the number of neighbouring nodes in the relevant point set.

[0328] When the value of attr_coding_type indicates lifting transform coding, the following syntax elements exist in the APS.

[0329] lifting_num_pred_nearest_neighbours specifies the maximum number of nearest neighbours to be used for prediction.

[0330] lifting_max_num_direct_predictors indicates the maximum number of predictors to be used for direct prediction.

[0331] lifting_search_range specifies the search range for determining the nearest neighbours to be used for prediction and for establishing distance-based LOD.

[0332] lifting_lod_regular_sampling_enabled_flag indicates the sampling strategy for establishing LOD. A lifting_lod_regular_sampling_enabled_flag equal to 1 indicates using a regular sampling strategy to establish LOD. A lifting_lod_regular_sampling_enabled_flag equal to 0 indicates using a distance-based sampling strategy to establish LOD.

[0333] lifting_num_detail_levels_minus1 indicates the number of LODs for attribute coding. The value of lifting_num_detail_levels_minus1 is greater than or equal to 0.

[0334] The following for loop includes as many elements indicating information about each LOD as the number indicated by lifting_num_detail_levels_minus1. In the figure, idx indicates each LOD. The value of idx is greater than or equal to 0 and less than the number indicated by lifting_num_detail_levels_minus1.

[0335] When the value of lifting_lod_regular_sampling_enabled_flag is 1, lifting_sampling_period[idx] is included. When the value of lifting_lod_regular_sampling_enabled_flag is 0, lifting_sampling_distance_squared[idx] is included.

[0336] lifting_sampling_period[idx] specifies the sampling period for LOD idx.

[0337] lifting_sampling_distance_squared[idx] specifies the scaling factor used to derive the square of the sampling distance for LOD idx.

[0338] When attr_coding_type indicates that the attribute coding is prediction transform coding, the APS includes the following syntax elements.

[0339] lifting_adaptive_prediction_threshold indicates the threshold for enabling adaptive prediction.

[0340] lifting_intra_lod_prediction_num_layers specifies the number of LOD layers that can refer to decoded points in the same LoD layer to generate the predicted value of the target point.

[0341] According to an embodiment, the syntax of the APS includes the following syntax elements.

[0342] The aps_extension_flag indicates whether the aps_extension_data_flag exists in APS. An aps_extension_flag equal to 0 indicates that the aps_extension_data_flag syntax element does not exist in the APS syntax structure. The value 1 of the aps_extension_flag is reserved for future use. The decoder may ignore all aps_extension_data_flag syntax elements that follow an aps_extension_flag equal to 1.

[0343] The aps_extension_data_flag indicates whether there is data for future use and can have any value.

[0344] The APS syntax is not limited to the above examples. For signaling efficiency, it may also include additional elements or may exclude some of the elements shown in the figure. Some elements may be signaled by signaling information other than APS (e.g., attribute headers, etc.) or by attribute data units.

[0345] As described in the reference Figures 15 to 19 When there are two or more neighboring points, the correlation weight can be calculated based on the correlation between these points. As a method for calculating the correlation weight, the correlation weight can be calculated without changing the existing attribute coding algorithm, thus ensuring the flexibility of the system design. Additionally, by using the correlation weight to replace the weighting constant used in the prediction / lifting transform, the coding and decoding performance is improved. The method for calculating the correlation weight according to the embodiment and the correlation weight can be applied to all functions that require prediction algorithms such as quantization.

[0346] The method for calculating the correlation degree (e.g., Equation 2) may or may not include the distance between corresponding points. Additionally, the method for calculating the correlation degree may include operations of addition, multiplication, and division using the matrix described in the reference Figure 2 The method for calculating the correlation degree may also include operations of removing the correlation value or using the correlation value to assign and multiply each constant. Each constant may include not only integers but also complex numbers and may have a fixed value or a variable value.

[0347] The correlation weight is the sum of the weights based on the correlation degree and corresponds to the average value or variance (e.g., Equation 3). Weights not combined with the correlation can be used as the average value or variance.

[0348] Figure 20 An example of a method for encoding the correlation weight according to an embodiment is illustrated.

[0349] Figure 20Instructions are shown that represent methods of calculating correlation weights (e.g., Equation 2 and Equation 3) in various ways when there are three points.

[0350] In this figure, weighted_sum represents the approximate sum of each point and the degree of relevance, and the method of calculating the degree of relevance can vary according to the implementation. w0, w1, and w2 represent the correlation weights that reflect the degree of relevance calculated for the corresponding points.

[0351] The first box represents the process of calculating the degree of relevance by multiplying the correlation values by any constants α, β, and γ. The second box represents the process of calculating the degree of relevance based solely on the distance between the points. The third box represents the process of calculating the degree of relevance based on the square of the point distance. The fourth box represents the process of calculating the degree of relevance based on the sum of the point distances. Each box represents the correlation weights w0, w1, and w2 of the corresponding points, which reflect the degree of relevance calculated according to the process of calculating the degree of relevance. The correlation weights can vary according to the calculation method and the type of relevance. Additionally, the method of calculating the degree of relevance is not limited to the above examples.

[0352] Reference Figures 1 to 20 The described point cloud processing device supports spatial adaptive decoding. Spatial adaptive decoding is performed on all or part of the geometry and / or attributes according to the decoding performance of the point cloud receiving device (e.g., Figure 1 receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device) to provide decoding of point cloud content at various resolutions. According to an embodiment, the part of the geometry and attributes is referred to as partial geometry and partial attributes. The adaptive decoding applied to the geometry according to an embodiment is referred to as adaptive geometry decoding or geometry adaptive decoding. The adaptive decoding applied to the attributes according to an embodiment is referred to as adaptive attribute decoding or attribute adaptive decoding. As described in reference Figures 1 to 17 The points of the point cloud content are distributed in 3D space and are represented in an octree structure (e.g., the octree described in reference Figure 6 ). The octree structure is an octal tree structure in which the depth increases from the upper node to the lower node. According to an embodiment, the depth is referred to as the level and / or layer.

[0353] The point cloud processing device (or geometric structure encoder) performs geometric structure encoding based on the octree structure. Additionally, the point cloud processing device (or attribute encoder) generates LOD and performs attribute encoding (e.g., RAHT transform, prediction transform, lifting transform, etc.) based on the octree structure. Since the LOD is generated based on the octree structure, the octree structure is regarded as dividing the grouping of points and organizing the number of points of geometric structures and attributes. The levels of the LOD can correspond to the depth of the octree. Since the LOD (or octree depth) must be large enough to represent the original quality, spatial adaptability is very useful when the source point cloud is densely arranged even in a local area. Through spatial adaptability, the point cloud receiving device (or decoder) can provide low-resolution point cloud content such as a thumbnail with low decoder complexity and / or small bandwidth. When spatial adaptability decoding is supported, the point cloud processing device sends information for spatial adaptability decoding by referring to Figure 17 the signaling information (e.g., SPS, APS, attribute headers, etc.) included in the described bitstream.

[0354] The point cloud receiving device obtains the information for spatial adaptability decoding through the signaling information included in the bitstream. The point cloud receiving device performs geometric structure decoding on all or part of the geometric structures corresponding to a specific depth (or level) from the upper nodes to the lower nodes of the octree structure. As described above, attribute decoding is based on geometric structure decoding. Therefore, the point cloud receiving device can generate LOD based on the decoded geometric structure (or decoded octree structure) and perform attribute decoding on all and / or part of the attributes (e.g., RAHT transform, prediction transform, lifting transform, etc.).

[0355] Figure 21 An example of spatial adaptability decoding is illustrated.

[0356] The arrow 1800 shown in the figure indicates the direction in which the level of the LOD increases.

[0357] As referred to Figures 1 to 14 described, the point cloud processing device generates LOD based on the octree structure. The LOD is designed to manage the attributes of points with an octree structure, and an increase in the LOD value indicates an increase in the detail of the point content. The LOD can correspond to one or more depths of the octree structure. The highest node of the octree structure corresponds to the lowest depth or the first depth and is called the root. The lowest node of the octree structure corresponds to the highest depth or the last depth and is called the leaf. The depth of the octree structure increases in the direction from the root to the leaf, which is the same as the direction indicated by the arrow.

[0358] According to an embodiment, the point cloud decoder performs decoding 1811 for providing full-resolution point cloud content or decoding 1812 for providing low-resolution point cloud content according to its performance. The point cloud decoder provides full-resolution point cloud content by decoding 1811 the geometry bitstream 1811-1 and the attribute bitstream 1812-1 corresponding to the entire octree structure. The point cloud decoder provides low-resolution point cloud content by decoding 1812 the partial geometry bitstream 1812-1 and the partial attribute bitstream 1812-2 corresponding to a specific depth of the octree structure. Figure 21 An example of the lifting transform as the attribute decoding is illustrated, but the embodiment is not limited to this example.

[0359] As described above, the signaling information (e.g., SPS, APS, attribute header, etc.) in the bitstream (e.g., Figure 17 the bitstream) may include adaptability information (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag) related to the spatial adaptability decoding (or lifting transform) at the sequence level or slice level. As described above, the attribute decoding is performed based on the decoded geometry octree structure. The information related to the spatial adaptability decoding (or lifting transform) indicates whether the entire octree structure or a partial octree structure is required to decode the partial attributes.

[0360] The point cloud receiving device obtains the signaling information of the bitstream and performs adaptable attribute decoding based on the entire octree structure or the partial octree structure as the result of the geometry decoding according to the information related to the spatial adaptability decoding.

[0361] As described above, the LOD is generated based on the octree structure (e.g., Figure 15 operation 1520 in). Therefore, the levels of the LOD are generated based on the depth of the octree. When the octree structure changes, the LOD structure also changes. The point cloud encoder (e.g., the point cloud encoder described with reference to Figure 15 performs the lifting transform based on the LOD (e.g., Figure 15 operation 1530 in). As described above, the point cloud processing device performs the lifting transform coding from the highest level of the LOD to the lowest level of the LOD. According to an embodiment, the lifting transform coding uses an update operator to calculate the predicted attributes of each point. The update operator for the corresponding point can calculate the updated attribute value based on the calculated weight and residual value. According to an embodiment, the point cloud processing device determines (or calculates or derives) the quantization weight according to the quantization weight derivation process and performs quantization based on the determined quantization weight.

[0362] The point cloud receiving device (or point cloud decoder) according to the embodiment can perform inverse quantization by using an update operator to restore the attribute value and calculating the quantization weight in the same manner as the point cloud processing device.

[0363] The density of points belonging to one level of LOD is different from the density of points belonging to another level of LOD. For example, the density of points belonging to one level of LOD is lower than the density of points belonging to a lower level of LOD. Therefore, according to the embodiment, the quantization weight is derived from the sum of distances in a higher level of LOD. However, when performing spatially adaptable decoding, the point cloud receiving device cannot accurately calculate the quantization weight because it does not have information about the low LOD. Therefore, the quantization weight is fixed by the number of points in the LOD. The following shows the process of calculating the quantization weight for each LOD to support attribute decoding (e.g., lifting transform decoding) performed based on a partial octree structure.

[0364] For i = 0 to

[0365]

[0366] Here, i is a parameter indicating the level of each LOD, and the value of i is greater than or equal to 0 and less than the number of LODs (LODcount). pointCount is the number of points belonging to the corresponding LOD, and predictorCount represents the number of predictors of points belonging to LODs lower than the corresponding LOD. predictorCount[i] represents the number of predictors of points belonging to the corresponding LOD. As shown in this formula, the weight is calculated based on the number of attributes and a fixed constant (e.g., kFixedPointweightShift).

[0367] The above formula is used to calculate the quantization weight based on the LOD with densely arranged points rather than the information of points having predicted values for each LOD. Therefore, when points are evenly distributed in one or more LODs, or when the gap of LOD indices is large and the distribution of points is irregular, the performance of the decoder deteriorates according to the quantization weight.

[0368] Accordingly, the point cloud transmitting device and the point cloud receiving device according to the embodiment perform an improved quantization weight derivation process to obtain a mathematical optimization when performing a lifting transform coding (e.g., a lifting transform coding performed based on a partial octree structure). The improved quantization weight derivation process can calculate an improved quantization weight that can be changed without applying a fixed constant. Accordingly, the improved quantization value derivation process minimizes the change in the quantization weight according to the point cloud system without changing the fixed constant for each system. In addition, since the quantization weight corresponds to a mathematically optimized value, a higher performance gain is ensured compared to the existing quantization weight. In addition, since the improved quantization weight derivation process does not require operations on each point, the complexity of the point cloud receiver can be reduced.

[0369] According to an embodiment, the quantization weight derivation process may be performed by program instructions stored in one or more memories included in the point cloud transmitting device and the receiving device. The program instructions are executed by the point cloud encoder and / or decoder (or processor) and cause the point cloud encoder and / or decoder to calculate / deduce the quantization weight.

[0370] According to an embodiment, the lifting transform is represented as a linear function representing the sum of the quantization weight and the attribute value. The following equation is a linear function representing the lifting transform and is determined to satisfy the maximum benefit with the total resource allocation value. The resource according to the embodiment refers to the sum of the products of the weight of each point and the attribute value (voltage representing the attribute value).

[0371] [Equation 4]

[0372]

[0373] The parameter j is the index of each point, greater than or equal to 0 and less than or equal to N. The parameter N represents the total number of points in the expected point cloud. That is, N corresponds to the total number of points to be transmitted by the point cloud transmitting device. The parameter w j is the quantization weight (or improved quantization weight) and is determined (or calculated or deduced) by the improved quantization weight derivation process. The parameter a j represents the attribute value of each point. SumAttribute represents the sum of the quantization weight and the attribute value and takes the form of multiplication and addition of N linear functions.

[0374] The point cloud receiving device can obtain the luminance and color values of the corresponding point through the value of SumAttribute. The condition for optimizing the function representing the above lifting transform is represented by the following equation.

[0375] [Equation 5]

[0376]

[0377] Experience

[0378] w j ≥ 1, where j = 1, ..., N

[0379] As shown in the above formula, the function representing the lifting transformation can be optimized by applying constraints to the parameter w j of. According to an embodiment, w j is greater than or equal to 1. Additionally, w j (where j is greater than or equal to 0 and less than or equal to N) of the sum (or total weight) of N values is less than or equal to the total number of prediction points (TotalPredictedCount). This is intended to prevent the quantization weights from increasing infinitely and to prevent overflow problems in the memory required for the calculation. Additionally, the formula for driving the lifting transformation consists of a linear form and an integer space, maintaining the convexity of the function.

[0380] Any defined function fj constituting the above formula consists of multiplication and addition of integers and decimals. Therefore, the function fj is represented as a convex function, and the function fj is a concave function according to the necessary and sufficient conditions. When the values of w1, w2, .., and w N are greater than or equal to 1, the functions f1 of w1, f2 of w1, ... and the function fN are also convex functions, so the combination of two or more of the convex functions takes the form of a convex function. According to an embodiment, the weight (or weight constant) w j can be modified or changed and has an optimized value. Additionally, when the function fj is a combination or clustering of convex functions, g, which is the clustering function in fi, i.e., the clustering function, is also configured as a convex function. Therefore, an optimized value of the clustering of the weights grouped by LOD can be obtained. According to an embodiment, the function g has a maximum value and a minimum value.

[0381] As described above, the quantization weights are determined (calculated or derived) by TotalPredictedCount. When the point cloud transmitting device and the point cloud receiving device according to an embodiment transmit / receive information about the total number of points (e.g., TotalCount), the quantization weights have a global optimized value determined based on the total number of points. When the point cloud receiving device predicts TotalPredictedCount, the quantization weights have a local optimized value determined based on TotalPredictedCount. The above formula uses various methods such as Karush - Kuhn - Tucker (KKT) conditions, geometric / non - geometric structure techniques, and group - based power constraints. Since a Lagrange multiplier (λ ∈ R) in the real space, inequality constraints, and w representing the optimized value are defined *≥0, so the point cloud transmitting device and receiving device according to the embodiment can perform an improved quantization weight derivation process by generating conditions for calculating quantization weights (e.g., constraint changes).

[0382] Figure 22 An exemplary improved quantization weight derivation process is shown.

[0383] Figure 22 It is a flowchart illustrating an improved quantization weight derivation process. This flowchart includes one or more operations. These operations can be performed simultaneously or sequentially.

[0384] The improved quantization weight derivation process defines an initial weight (or initial quantization weight) (22100). According to the embodiment, the initial weight is determined based on the average transceiver power. That is, considering the data memory size, computational complexity, etc., the initial weight can be determined based on the average energy value or total energy value at the 32-bit and 64-bit levels. Additionally, when all point cloud data is normalized, the initial weight is determined to be 1.

[0385] The improved quantization weight derivation process reorders / counts the number of predicted points for each LOD (e.g., TotalPredictedCount as described above) (22200). This point represents a geometric structure point or an attribute point. The number of predicted points for each LOD can be obtained from the decoded geometric structure. The number of predicted points for each LOD is stored as a parameter. The improved quantization weight derivation process can store the number of actual points (e.g., TotalCount as described above) as a parameter without counting the number of predicted points.

[0386] The improved quantization weight derivation process calculates the total constraint (22300). Apply constraints to the weight or the sum of weights (e.g., the sum of w in Equation 8). According to the embodiment, the constraints can include but are not limited to the number of points in the total cumulative LOD for each LOD, the number of points in some LODs, the number of points belonging to a subset, and the number of points grouped according to the LOD. j That is, the sum of the weights of the attributes multiplied by each point. Therefore, the resources are proportionally allocated according to the LOD level and the points in the LOD.

[0387] The improved quantization weight derivation process calculates the weight assignment determined based on the calculated constraints (22400). That is, the improved quantization weight derivation process allocates resources based on the constraints according to each LoD level and the points in the LoD. As described above, the resources are the sum of the weights of the attributes multiplied by each point.

[0388] That is, the total number of points is allocated to each LoD level according to the weight based on the constraint.

[0389] Improve the quantization weight derivation process to update the weights (22500). The updated weights (or updated quantization weights) can be defined as the assigned weights, or can be generated by accumulating the assigned weights to the initial weights, or by modifying and combining some existing updated weights. The finally updated weights have optimized values.

[0390] In the following, the process of calculating the weight assignment determined based on the constraints calculated based on the reference Figure 22 will be described.

[0391] As described above, the improved quantization weight derivation process is based on the total number of prediction points (e.g., TotalPredictedCount above). Since the total number of prediction points is limited, the improved quantization weight derivation process can derive the same optimized value or a similar optimized value. According to an embodiment, the optimized value is calculated based on the ratio of the resources (the number of accumulated points of the LOD) to the total resources (the total number of points).

[0392] The process of calculating the number of accumulated points for each LOD is performed by program instructions stored in one or more memories included in the point cloud transmitting device and the receiving device. According to an embodiment, the program instructions are executed by the point cloud encoder and / or decoder (or processor), and cause the point cloud encoder and / or decoder to calculate the number of accumulated points for each LOD. The process of calculating the i-th optimized value (improved quantization weight) is represented as follows.

[0393]

[0394] OptimalWeight[i] represents the optimized value of the i-th LOD. numberOfPointsPerLOD[LodCount - 1] is the number of all points up to the LOD level corresponding to the value that is 1 less than the value of LoDCount. That is, numberOfPointsPerLOD[LodCount - 1] represents the sum of the number of points belonging to each LOD from the LOD at the level where i is 0 to the LOD at the level where i is LodCount - 1 (e.g., the number of points of LOD 0 equal to 1, the number of points of LOD 1 equal to 7,..., and the sum of the number of points of LOD LodCount - 1 equal to XX). According to an embodiment, numberOfPointsPerLOD[LodCount - 1] can be stored in the form of an index or the like, and can be obtained from the maximum index stored.

[0395] As described above, the i-th optimized value is generated based on the constraints and Ratio[i]. According to an embodiment, Ratio[i] may change according to the irregularity of the total energy that occurs in an actual environment such as the point cloud noise error. For example, numberOfpointsperLOD[i-1] may be used instead of numberOfpointsperLOD[i], or a part of numberOfpointsperLOD[i] may be grouped and regarded as a variable. According to an embodiment, the grouped variables may include sequential variables such as i+1 and i+2 or non-sequential variables such as i, i+4, and i+6. The constraint numberOfPointsPerLOD[LodCount-1] may be expressed as numberOfPointsPerLOD[k], where k is i-1, i-2,..., 0 or i+1, i+2. Additionally, a part of numberOfPointsPerLOD[k] may be grouped and regarded as a variable. For example, according to an embodiment, the grouped variables include sequential variables such as i+1 and i+2 or non-sequential variables such as k, k+4, and k+6. A specific constant may be subtracted from, added to, or combined with NumberOfPointsPerLOD[i] and NumberOfPointsPerLOD[k] represented by the variables i and k. For example, when two constants α and β are given for i and k, numberOfPointsPerLOD[i](+ or -)α or numberOfPointsPerLOD[i](* or / )α may be obtained, and numberOfPointsPerLOD[k](+ or may be -)β or numberOfPointsPerLOD[k](* or / )β may be obtained.

[0396] Different improved quantization weight derivation processes may be applied to the points of the corresponding LOD levels with the same attribute. For example, when i = 1, the optimized value may be inferred by numberOfPointPerLOD[LoDCount-1] / numberOfPointsPerLoD[i]. When i is greater than 1, the optimized value may be inferred by (numberOfPointPerLOD[LoDCount-1]-numberOfPointsPerLoD[i]) / numberOfPointsPerLoD[i].

[0397] Different improved quantization weight derivation processes may be applied to the points according to the LOD levels of different attributes. Therefore, the optimized values of the quantization weights for the LOD levels given when the attribute is reflectance and the quantization weights for the LOD levels given when the attribute is color are different.

[0398] Based on the LoD levels of corresponding sub-components (e.g., luminance, chrominance, etc.) in the same attribute, different improved quantization weight derivation processes can be applied to points. Thus, when the attribute is color (YCbCr), the quantization weights of the sub-component Y (luminance) and the sub-components Cb and Cr (chrominance) have different optimized values.

[0399] Figure 23 is a flowchart illustrating point cloud encoding according to an embodiment.

[0400] Reference Figures 1 to 22 The described point cloud transmitting device or point cloud encoder can perform attribute encoding to support the spatial adaptability decoding described in Reference Figure 21 According to an embodiment, the attribute encoding includes at least one of RAHT encoding, predictive transform encoding, or lifting transform encoding. The lifting transform encoding performed to support spatial adaptability decoding performs the improved quantization weight calculation process described in Reference Figures 21 to 22

[0401] As Figure 23 shown, the point cloud encoder receives an input (23100) of an attribute. According to an embodiment, the attribute includes color and reflectivity.

[0402] The point cloud encoder according to an embodiment performs attribute encoding (23200, 23210) according to the attribute type. As shown in the figure, the point cloud encoder independently performs encoding in the case where the attribute is color and in the case where the attribute is reflectivity. The two attribute encoding operations can be performed simultaneously or sequentially. The attribute encoding performs a lifting transform based on LOD (e.g., Figure 15 operation 1530). As described above, the point cloud processing device performs lifting transform encoding from the highest LOD level to the lowest LOD level.

[0403] As described above, the point cloud encoding can support spatial adaptability decoding (23300, 23310). According to an embodiment, the spatial adaptability decoding can perform decoding on some or all of the geometric structure and / or attributes according to the decoding performance of the point cloud receiving device (e.g., Figure 1 receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device) to provide point cloud content of various resolutions. Through spatial adaptability, the point cloud receiving device (or decoder) can provide low-resolution point cloud content such as a thumbnail with low decoder complexity and / or small bandwidth. When supporting spatial adaptability decoding, the point cloud processing device, by referring to Figure 17 ​The signaling information (e.g., SPS, APS, attribute headers, etc.) included in the described bitstream is used to send information for spatially adaptive decoding. As described above, the bitstream (e.g., Figure 17 the bitstream) may include adaptability information related to spatially adaptive decoding (or lifting transform) at the sequence level or slice level (adaptability information indicating whether decoding of attributes can be performed based on a spatial octree structure) (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag). As described above, attribute decoding is performed based on the decoded geometric octree structure. The information related to spatially adaptive decoding (or lifting transform) indicates whether the entire octree structure or a partial octree structure is required for partial attribute decoding. Therefore, the point cloud receiving device obtains the signaling information of the bitstream and performs adaptable attribute decoding based on the entire octree structure or partial octree structure as the result of geometric structure decoding according to the information related to spatially adaptive decoding.

[0404] When spatially adaptive decoding (e.g., lifting transform coding) based on a partial octree structure is supported, the point cloud encoder according to an embodiment uses a reference Figures 21 to 23 described improved quantization weight derivation process to calculate quantization weights (23400, 23410). The improved quantization weight derivation process is the same as the Figures 21 to 22 process, so the detailed description will be skipped. The point cloud receiving device (or point cloud decoder) according to an embodiment can perform inverse quantization by calculating quantization weights in the same manner as the point cloud processing device.

[0405] The point cloud encoder compresses color attributes and reflectivity attributes (23500, 23510).

[0406] The point cloud encoding shown in Figure 23 can be performed by program instructions stored in one or more memories included in the point cloud sending device and the receiving device. According to an embodiment, the program instructions are executed by the point cloud encoder and / or decoder (or processor) and cause the point cloud encoder and / or decoder to perform point cloud encoding and / or decoding.

[0407] Figure 24 is a flowchart illustrating a method of sending point cloud data according to an embodiment.

[0408] Figure 24 The flowchart 2400 of Figures 1 to 23 illustrates a point cloud data sending device by referring to Figure 1 、 Figure 12 and Figure 14A method of transmitting point cloud data by the described transmitting device or point cloud encoder). The point cloud data transmitting device encodes the point cloud data including geometric structures and attributes (2410). The geometric structure is information indicating the positions of points in the point cloud data, and the attributes include at least one of the color and reflectance of the points. The point cloud data transmitting device encodes the geometric structure. As described above with reference to Figures 1 to 23 described, the attribute encoding depends on the geometric structure encoding. Therefore, the point cloud data transmitting device encodes the attributes based on the whole or part of the octree structure of the encoded geometric structure, as described in reference Figure 21 described. The attributes are encoded based on the quantization weights of the points included in each level of detail (LOD) among one or more LODs. The quantization weights are determined based on the number of points and the number of points belonging to level l represented by the LOD. The details are the same as those described in reference Figures 20 to 23 described, and the description thereof will be skipped.

[0409] The quantization weights according to an embodiment are represented as follows.

[0410]

[0411] Here, represents the quantization weight of the LOD with level l, "total number of points" represents the number of points, "number of points in LOD l " represents the cumulative number of points belonging to level l represented by the LOD, and "number of points in R i " represents the number of points belonging only to the LOD with level i. The value of the number of points in LOD l is equal to the sum of the values of the number of points in R i where i ranges from 0 to l. The quantization weights according to an embodiment are the same as the process of calculating the i-th optimized value (improved quantization weight) or the calculated quantization weights described in reference Figures 21 to 22 described, so the detailed description thereof will be skipped.

[0412] As described in reference Figures 15 to 23 described, the point cloud data transmitting device generates one or more LODs by reordering the points, performs lifting transform coding on the attributes based on the one or more LODs, and quantizes the attributes after lifting transform coding based on the quantization weights. The quantization weights according to an embodiment are used to decode the attributes encoded based on the partial octree structure of the geometric structure. That is, as described in reference Figures 1 to 23 described, the point cloud transmitting device supports spatial scalability decoding. Spatial scalability decoding is based on the point cloud receiving device (e.g., Figure 1 the receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13The decoding performance of the receiving device) performs all or part of the geometric structure and / or attributes to provide decoded point cloud content at various resolutions.

[0413] The point cloud data transmitting device transmits a bitstream (e.g., the bitstream described in reference Figure 17 ) containing the encoded point cloud data (2420).

[0414] Therefore, the bitstream according to an embodiment (e.g., Figure 17 ) contains adaptability information (e.g., scalable_lifting_enabled_flag, lifting_scalability_enabled_flag) indicating whether the attributes encoded based on the partial octree structure can be decoded. As described in reference Figures 20 to 23 , the information for spatial adaptability decoding is sent through the signaling information (e.g., SPS, APS, attribute header, etc.) contained in the bitstream described in reference Figure 17 . As described above, the signaling information (e.g., SPS, APS, attribute header, etc.) in the bitstream (e.g., the bitstream of Figure 17 ) may include signaling information indicating whether the attributes decoded based on the spatial octree structure at the sequence level or slice level can be decoded. The receiver can obtain such information and perform spatial adaptability decoding. Since the operation of the point cloud data transmitting device is the same as that described in reference Figures 1 to 23 , a detailed description thereof will be skipped.

[0415] Figure 25 is a flowchart illustrating a method for processing point cloud data according to an embodiment.

[0416] Figure 25 The flowchart 2500 of Figures 1 to 23 illustrates a method for processing point cloud data by the point cloud data receiving device (e.g., the receiving device 10004 or the point cloud video decoder 10006) described in reference

[0417] The point cloud data receiving device (e.g., the receiving device 10004, Figure 13 the receiver of Figures 20 to 23 , etc.) receives a bitstream (2510) containing the point cloud data. The bitstream according to an embodiment contains the signaling information (e.g., SPS, APS, attribute header, etc.) necessary for decoding the point cloud data. As described in reference Figure 17 , the information for spatial adaptability decoding is sent through the signaling information (e.g., SPS, APS, attribute header, etc.) contained in the bitstream described in reference

[0418] The point cloud data receiving device (e.g., Figure 10 the decoder) decodes the point cloud data (2520) based on signaling information. The point cloud data receiving device (e.g., Figure 10 the geometry decoder) decodes the geometry included in the point cloud data. The geometry according to an embodiment is information indicating the positions of the points of the point cloud data. The point cloud data receiving device (e.g., Figure 10 the attribute decoder) decodes the attributes including at least one of the color and reflectance of the points based on the entire or partial octree structure of the decoded geometry. The attributes are decoded based on the quantization weights (or refined quantization weights) of the points included in each level of detail (LOD) among one or more LODs. The quantization weights are determined based on the number of points and the number of points belonging to level l represented by the LOD.

[0419] The signaling information necessary for decoding the point cloud data further includes adaptability information (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag) indicating whether the attributes can be decoded based on the partial octree structure. When the adaptability information indicates that the attributes can be decoded based on the partial octree structure, the quantization weights are determined for each point included in each LOD from the 0th LOD to the last LOD. The quantization weight calculation process is the same as the quantization weight calculation process described in reference Figures 21 to 23 and thus the detailed description thereof is skipped.

[0420] The point cloud data receiving device according to an embodiment generates one or more LODs by reordering the points, performs lifting transform decoding on the attributes based on the one or more LODs, and inverse quantizes the attributes after the lifting transform decoding based on the quantization weights. As described above in reference Figure 22 the quantization weights are determined by executing program instructions for calculating the quantization weights stored in the memory included in the point cloud receiving device. The point cloud data processing operations of the point cloud data receiving device are the same as those described in reference Figures 1 to 23 and thus the detailed description thereof is skipped.

[0421] According to Figures 1 to 25The components of the point cloud data processing apparatus described in can be implemented by hardware, software, firmware, or a combination thereof including one or more processors in combination with a memory. The components in the embodiments can be implemented by a single chip, for example, a single hardware circuit. According to an embodiment, the components according to the embodiment can be implemented separately as individual chips. Additionally, at least one or more components of the apparatus according to the embodiment can include one or more processors capable of executing one or more programs. One or more programs can execute Figures 1 to 25 any one or more of the operations / methods described in the operation / method of the point cloud data processing apparatus, or include instructions for performing the same operation / method.

[0422] Although the drawings have been described separately for simplicity, new embodiments can be designed by combining the embodiments illustrated in the corresponding figures. The design of a computer-readable recording medium on which a program for executing the above-described embodiments is recorded, which is required by those skilled in the art, also falls within the scope of the appended claims and their equivalents. The apparatus and method according to the embodiment can be unrestricted by the configurations and methods of the above-described embodiments. By selectively combining all or part of the embodiments, various modifications can be made to the embodiments. Although the preferred embodiments have been described with reference to the drawings, those skilled in the art will appreciate that various modifications and variations can be made to the embodiments without departing from the spirit or scope of the present disclosure described in the appended claims. Such modifications will not be understood independently of the technical ideas or viewpoints of the embodiments.

[0423] The descriptions of the apparatus and method according to the embodiment can be applied to complement each other. For example, the method for sending point cloud data according to the embodiment can be executed by the apparatus for sending point cloud data according to the embodiment or a component included in the apparatus for sending point cloud data. Additionally, the method for receiving point cloud data according to the embodiment can be executed by the apparatus for receiving point cloud data according to the embodiment or a component included in the apparatus for receiving point cloud data according to the embodiment.

[0424] The various elements of the device according to the embodiment can be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiment can be implemented by a single chip (e.g., a single hardware circuit). According to an embodiment, the components according to the embodiment can be implemented separately as separate chips. According to an embodiment, at least one or more components of the device according to the embodiment can include one or more processors capable of executing one or more programs. The one or more programs can execute any one or more of the operations / methods according to the embodiment, or include instructions for executing them. The executable instructions for executing the method / operation of the device according to the embodiment can be stored in a non-transitory CRM or other computer program products configured to be executed by one or more processors, or can be stored in a transient CRM or other computer program products configured to be executed by one or more processors. Additionally, the memory according to the embodiment can be used to cover not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it can also be implemented in the form of a carrier wave such as being transmitted via the Internet. Additionally, the processor-readable recording medium can be distributed among computer systems connected via a network such that the processor-readable code can be stored and executed in a distributed manner.

[0425] In this specification, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, in this specification, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" used in this document should be interpreted as indicating "additionally or alternatively".

[0426] Terms such as first and second can be used to describe the various elements of the embodiment. However, the various components according to the embodiment should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, the first user input signal can be referred to as the second user input signal. Similarly, the second user input signal can be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Both the first user input signal and the second user input signal are user input signals, but do not mean the same user input signal unless the context clearly indicates otherwise.

[0427] The terms used to describe the embodiments are used for the purpose of describing particular embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and the claims, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. The expression "and / or" is used to include all possible combinations of terms. Terms such as "including" or "having" are intended to indicate the components of diagrams, numbers, steps, elements and / or parts, and should be understood as not precluding the possibility of the additional presence of diagrams, numbers, steps, elements and / or components. As used herein, conditional expressions such as "if" and "when" are not limited to alternative cases and are intended to be interpreted as performing the relevant operations or interpreting the relevant definitions according to the specific conditions when the specific conditions are met.

[0428] The mode of the present disclosure

[0429] As described above, the relevant content has been described in the best mode of implementing the embodiments.

[0430] Industrial applicability

[0431] Those skilled in the art will appreciate that various changes or modifications can be made to the embodiments within the scope of the embodiments. Therefore, the embodiments are intended to cover the modified forms and variations of the present disclosure, provided that they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for decoding point cloud data, the method comprising the following steps: Decoding the geometric structure of the point cloud data based on an octree, wherein the geometric structure represents the positions of the points of the point cloud data, and wherein the octree includes nodes containing the points; and Decoding the attributes by generating levels of detail based on a partial octree of the geometric structure, wherein the attributes include at least one of the color and reflectivity of the points, wherein, for the levels of detail, quantization weights are applied to the attributes, wherein each quantization weight for each level of detail is derived based on the ratio of the total number of points associated with the level of detail to the number of points in each level of detail.

2. The method according to claim 1, Among them, wherein the geometric structure and the attributes are included in a bitstream, wherein the bitstream includes signaling information indicating whether to decode the encoded attributes based on the partial octree.

3. A method for encoding point cloud data, the method comprising the following steps: Encoding the geometric structure of the point cloud data based on an octree, wherein the geometric structure represents the positions of the points of the point cloud data, and wherein the octree includes nodes containing the points; and Encoding the attributes by generating levels of detail based on a partial octree of the geometric structure, wherein the attributes include at least one of the color and reflectivity of the points, wherein, for the levels of detail, quantization weights are applied to the attributes, wherein each quantization weight for each level of detail is derived based on the ratio of the total number of points associated with the level of detail to the number of points in each level of detail.

4. The method according to claim 3, Among them, wherein the geometric structure and the attributes are included in a bitstream, wherein the bitstream includes signaling information indicating whether to decode the encoded attributes based on the partial octree.

5. A method for transmitting point cloud data, the method comprising the following steps: Encoding the geometric structure of the point cloud data based on an octree, wherein the geometric structure represents the positions of the points of the point cloud data, and wherein the octree includes nodes containing the points; Encoding the attributes by generating levels of detail based on a partial octree of the geometric structure, wherein the attributes include at least one of the color and reflectivity of the points, wherein, for the levels of detail, quantization weights are applied to the attributes, wherein each quantization weight for each level of detail is derived based on the ratio of the total number of points associated with the level of detail to the number of points in each level of detail; and Transmitting data including a bitstream, the bitstream including the encoded geometric structure and the encoded attributes.