Encoding and decoding method and bit stream transmission method
By improving G-PCC technology, using the maximum proximity point distance and attribute correlation to select proximity points, the problem of low point cloud data transmission and processing efficiency is solved, and efficient point cloud data compression and encoding/decoding is achieved.
Patent Information
- Application Number
- CN202510223894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-13
- Filing Date
- 2021-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to efficiently process and transmit large amounts of point cloud data, resulting in long wait times and high encoding/decoding complexity.
By improving geometry-based point cloud compression (G-PCC) technology, the adjacent points predicted by attributes are selected using the maximum proximity distance, and the attribute correlation between points is taken into account when encoding, improving the compression performance and encoding/decoding efficiency of point cloud data.
It realizes efficient point cloud data transmission and processing, reduces the size of attribute bitstream, and improves compression efficiency and encoding/decoding performance.
Smart Images

Figure CN120075466A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202180028349.4 (International Application No.: PCT / KR2021 / 000588, filing date: January 15, 2021, invention title: Device for transmitting point cloud data, method for transmitting point cloud data, device for receiving point cloud data, and method for receiving point cloud data). Technical Field
[0002] Embodiments relate to a method and a device for processing point cloud content. Background Art
[0003] Point cloud content is content represented by a point cloud, which is a set of points belonging to a coordinate system representing a three-dimensional space. Point cloud content can represent media configured in three dimensions and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), XR (extended reality), and autonomous driving. However, tens of thousands to hundreds of thousands of point data are required to represent point cloud content. Therefore, methods for efficiently processing a large amount of point data are needed. Summary of the Invention
[0004] Technical Problem
[0005] An object of the present disclosure designed to solve the above problems is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for efficiently transmitting and receiving point clouds.
[0006] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for coping with latency and encoding / decoding complexity.
[0007] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for improving the compression performance of point clouds by improving the encoding technology of attribute information based on geometry-based point cloud compression (G-PCC).
[0008] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for enhancing compression efficiency while supporting parallel processing of attribute information of G-PCC.
[0009] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for reducing the size of an attribute bitstream and enhancing the compression efficiency of attributes by selecting neighboring points for attribute prediction in consideration of the attribute correlation between points in content when encoding the attribute information of G-PCC.
[0010] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method that reduce the size of an attribute bitstream and enhance the compression efficiency of an attribute by applying a maximum neighboring point distance when encoding attribute information of G-PCC and selecting neighboring points for attribute prediction.
[0011] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method that perform efficient encoding / decoding of various types of point cloud data by automatically calculating a maximum neighboring point range according to density to obtain a maximum neighboring point distance when encoding attribute information of G-PCC.
[0012] The object of the present disclosure is not limited to the above-mentioned objects, and other objects of the present disclosure not mentioned above will become clear to those of ordinary skill in the art after reviewing the following description.
[0013] Technical Solution
[0014] To achieve these objects and other advantages and in accordance with the object of the present disclosure, as implemented and broadly described herein, a point cloud data transmission method may include the following steps: obtaining point cloud data; encoding geometric information including the positions of points of the point cloud data; generating one or more levels of detail (LOD) based on the geometric information, and selecting one or more neighboring points of each point to be attribute-encoded based on the one or more LOD; encoding the attribute information of each point based on the one or more neighboring points of each selected point; and transmitting the encoded geometric information, the encoded attribute information, and signaling information.
[0015] According to an embodiment, one or more neighboring points of each selected point may be within a maximum neighboring point distance.
[0016] According to an embodiment, the maximum neighboring point distance may be determined based on a basic neighboring point distance and a maximum neighboring point range.
[0017] According to an embodiment, when generating the one or more LOD based on an octree, the basic neighboring point distance may be determined based on the diagonal distance of a node in a specific LOD.
[0018] According to an embodiment, selecting the one or more neighboring points may include the following steps: estimating the density of the point cloud data based on the basic neighboring point distance and the diagonal length of the bounding box of the point cloud data; and automatically calculating the maximum neighboring point range according to the estimated density.
[0019] According to an embodiment, information related to the maximum neighbor point range may be signaled in the signaling information.
[0020] In another aspect of the present disclosure, a point cloud data transmitting apparatus may include: an acquirer configured to acquire point cloud data; a geometry encoder configured to encode geometric information including positions of points of the point cloud data; and an attribute encoder configured to generate one or more levels of detail (LOD) based on the geometric information, select one or more neighbor points of each point to be attribute-encoded based on the one or more LOD, and encode attribute information of each point based on the one or more selected neighbor points of each point; and a transmitter configured to transmit the encoded geometric information, the encoded attribute information, and the signaling information.
[0021] According to an embodiment, the one or more neighbor points of each selected point may be within a maximum neighbor point distance.
[0022] According to an embodiment, the maximum neighbor point distance may be determined based on a basic neighbor point distance and a maximum neighbor point range.
[0023] According to an embodiment, when generating the one or more LOD based on an octree, the basic neighbor point distance may be determined based on a diagonal distance of a node in a specific LOD.
[0024] The attribute encoder may estimate a density of the point cloud data based on the basic neighbor point distance and a diagonal length of a bounding box of the point cloud data, and automatically calculate the maximum neighbor point range according to the estimated density.
[0025] According to an embodiment, information related to the maximum neighbor point range may be signaled in the signaling information.
[0026] In another aspect of the present disclosure, a point cloud data receiving apparatus may include: a receiver configured to receive geometric information, attribute information, and signaling information; a geometry decoder configured to decode the geometric information based on the signaling information; an attribute decoder configured to generate one or more levels of detail (LOD) based on the geometric information, select one or more neighbor points of each point to be attribute-encoded based on the one or more LOD, and decode attribute information of each point based on the one or more selected neighbor points of each point and the signaling information; and a renderer configured to render point cloud data reconstructed based on the decoded geometric information and the decoded attribute information.
[0027] According to an embodiment, one or more neighboring points of each selected point may be within a maximum neighboring point distance.
[0028] According to an embodiment, the maximum neighboring point distance may be determined based on a basic neighboring point distance and a maximum neighboring point range.
[0029] According to an embodiment, when generating the one or more LODs based on an octree, the basic neighboring point distance may be determined based on the diagonal distance of a node in a specific LOD.
[0030] The attribute decoder may estimate the density of the point cloud data based on the basic neighboring point distance and the diagonal length of the bounding box of the point cloud data, and automatically calculate the maximum neighboring point range according to the estimated density.
[0031] According to an embodiment, the attribute decoder may obtain the maximum neighboring point range from the signaling information.
[0032] Advantageous Effects
[0033] The point cloud data transmission method, the point cloud data transmission device, the point cloud data reception method, and the point cloud reception device according to an embodiment may provide a high-quality point cloud service.
[0034] The point cloud data transmission method, the point cloud data transmission device, the point cloud data reception method, and the point cloud reception device according to an embodiment may implement various video encoding and decoding methods.
[0035] The point cloud data transmission method, the point cloud data transmission device, the point cloud data reception method, and the point cloud reception device according to an embodiment may provide general point cloud content such as an autonomous driving service (or a self-driving service).
[0036] The point cloud data transmission method, the point cloud data transmission device, the point cloud data reception method, and the point cloud data reception device according to an embodiment may perform spatial adaptive segmentation of the point cloud data to independently encode and decode the point cloud data, thereby improving parallel processing and providing scalability.
[0037] The point cloud data transmission method, the point cloud data transmission device, the point cloud data reception method, and the point cloud data reception device according to an embodiment may perform encoding and decoding and thus signal necessary data by spatially dividing the point cloud data in units of tiles and / or slices, thereby improving the encoding and decoding performance of the point cloud.
[0038] The point cloud data sending method and device according to the embodiment, and the point cloud data receiving method and device reduce the size of the attribute bitstream and enhance the compression efficiency of the attribute by selecting neighboring points for attribute prediction while taking into account the attribute correlation between the points of the content when the attribute information of G-PCC is encoded.
[0039] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device can cause neighboring points for attribute prediction to be selected based on the maximum neighboring point distance when encoding the attribute information of G-PCC, thereby reducing the size of the attribute bitstream and improving the attribute compression efficiency.
[0040] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device can cause neighboring points for attribute prediction to be selected based on the maximum neighboring point distance when encoding the attribute information of G-PCC, thereby removing points outside the maximum neighboring point distance, that is, points that interfere with prediction in the neighboring point set. Therefore, when the correlation of the attribute is low due to the attribute characteristics of the content, the average value of the prediction can be increased, the residual value from the point attribute value can be prevented from increasing, and weights can be not assigned to unimportant information. Accordingly, the size of the attribute bitstream can be reduced and the PSNR accuracy can be increased.
[0041] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device can configure the maximum neighboring point range according to the LOD configuration method, service characteristics, or content characteristics. Accordingly, a maximum neighboring point range suitable for the attribute characteristics can be provided.
[0042] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device can configure the maximum range of neighboring points according to the LOD configuration method, service characteristics, or content characteristics. Accordingly, various types of point cloud data can be efficiently encoded / decoded.
[0043] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device can automatically calculate the maximum neighboring point range according to the density, thereby providing the best compression efficiency.
[0044] According to an embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method, and the point cloud data receiving device may estimate the maximum neighboring point distance using at least one of an octree-based calculation of the maximum neighboring point distance, a distance-based calculation of the maximum neighboring point distance, a sampling-based calculation of the maximum neighboring point distance, a calculation of the maximum neighboring point distance by calculating the average difference of Morton codes between LODs, or a calculation of the maximum neighboring point distance by calculating the average distance difference between LODs, and send a signal notifying relevant information. Accordingly, the attribute compression / decompression efficiency may be improved through an optimal set of neighboring points. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings are included to provide a further understanding of the present disclosure and are incorporated into and constitute a part of this application. The drawings illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. In the drawings:
[0046] Figure 1 An exemplary point cloud content providing system according to an embodiment is illustrated.
[0047] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.
[0048] Figure 3 An exemplary process of capturing a point cloud video according to an embodiment is illustrated.
[0049] Figure 4 An exemplary block diagram of a point cloud video encoder according to an embodiment is illustrated.
[0050] Figure 5 An example of a voxel in a 3D space according to an embodiment is illustrated.
[0051] Figure 6 An example of an octree and an occupancy code according to an embodiment is illustrated.
[0052] Figure 7 An example of a neighboring node pattern according to an embodiment is illustrated.
[0053] Figure 8 An example of a point configuration of point cloud content for each LOD according to an embodiment is illustrated.
[0054] Figure 9 An example of a point configuration of point cloud content for each LOD according to an embodiment is illustrated.
[0055] Figure 10 An example of a block diagram of a point cloud video decoder according to an embodiment is illustrated.
[0056] Figure 11Illustrates an example of a point cloud video decoder according to an embodiment.
[0057] Figure 12 Illustrates the configuration of point cloud video encoding of a transmitting device according to an embodiment.
[0058] Figure 13 Illustrates the configuration of point cloud video decoding of a receiving device according to an embodiment.
[0059] Figure 14 Illustrates an exemplary structure that can be connected during operation to a method / device for transmitting and receiving point cloud data according to an embodiment.
[0060] Figure 15 Illustrates an example of a point cloud transmitting device according to an embodiment.
[0061] Figure 16 (a) to Figure 16 (c) of illustrates an embodiment of dividing a bounding box into one or more tiles.
[0062] Figure 17 Illustrates examples of a geometry encoder and an attribute encoder according to an embodiment.
[0063] Figure 18 Is a diagram illustrating an example of generating LOD based on an octree according to an embodiment.
[0064] Figure 19 Is a diagram illustrating an example of arranging points in a point cloud in Morton code order according to an embodiment.
[0065] Figure 20 Is a diagram illustrating an example of searching for neighboring points based on LOD according to an embodiment.
[0066] Figure 21 (a) and Figure 21 (b) of illustrates an example of point cloud content according to an embodiment.
[0067] Figure 22 (a) and Figure 22 (b) of illustrates examples of the average distance, minimum distance, and maximum distance of each point belonging to a neighboring point set according to an embodiment.
[0068] Figure 23 (a) of illustrates an example of obtaining the diagonal distance of an octree node of LOD0 according to an embodiment.
[0069] Figure 23 (b) of illustrates an example of obtaining the diagonal distance of an octree node of LOD1 according to an embodiment.
[0070] Figure 24 of (a) and Figure 24 Example (b) illustrates the range of neighboring points that can be selected at each LOD.
[0071] Figure 25 Example shows the basic neighboring point distance belonging to each LOD according to an embodiment.
[0072] Figure 26 is a flowchart illustrating the automatic calculation of the maximum neighboring point range based on density according to an embodiment of the present disclosure.
[0073] Figure 27 is a flowchart illustrating the automatic calculation of the maximum neighboring point range based on density according to another embodiment of the present disclosure.
[0074] Figure 28 Example illustrates another example of searching for neighboring points based on LOD according to an embodiment.
[0075] Figure 29 Example illustrates another example of searching for neighboring points based on LOD according to an embodiment.
[0076] Figure 30 Example illustrates an example of a point cloud receiving device according to an embodiment.
[0077] Figure 31 Example illustrates examples of a geometry decoder and an attribute decoder according to an embodiment.
[0078] Figure 32 Example illustrates an exemplary bitstream structure of point cloud data for transmission / reception according to an embodiment.
[0079] Figure 33 Example illustrates an exemplary bitstream structure of point cloud data according to an embodiment.
[0080] Figure 34 Example illustrates the connection relationship between components in the bitstream of point cloud data according to an embodiment.
[0081] Figure 35 Example illustrates an embodiment of the syntax structure of a sequence parameter set according to an embodiment.
[0082] Figure 36 Example illustrates an embodiment of the syntax structure of a geometry parameter set according to an embodiment.
[0083] Figure 37 Example illustrates an embodiment of the syntax structure of an attribute parameter set according to an embodiment.
[0084] Figure 38 Example illustrates an embodiment of the syntax structure of a tile parameter set according to an embodiment.
[0085] Figure 39 Illustrates an embodiment of the syntax structure of a geometric slice bitstream () according to an embodiment.
[0086] Figure 40 Illustrates an embodiment of the syntax structure of a geometric slice header according to an embodiment.
[0087] Figure 41 Illustrates an embodiment of the syntax structure of geometric slice data according to an embodiment.
[0088] Figure 42 Illustrates an embodiment of the syntax structure of an attribute slice bitstream () according to an embodiment.
[0089] Figure 43 Illustrates an embodiment of the syntax structure of an attribute slice header according to an embodiment.
[0090] Figure 44 Illustrates an embodiment of the syntax structure of attribute slice data according to an embodiment.
[0091] Figure 45 Is a flowchart of a method for transmitting point cloud data according to an embodiment.
[0092] Figure 46 Is a flowchart of a method for receiving point cloud data according to an embodiment. Detailed embodiments
[0093] Now, a description will be given in detail with reference to the accompanying drawings according to the exemplary embodiments disclosed herein. For a brief description with reference to the accompanying drawings, the same or equivalent components may be provided with the same reference numerals, and their descriptions will not be repeated. It should be noted that the following examples are only for illustrating the present disclosure and do not limit the scope of the present disclosure. What can be easily inferred by those skilled in the technical field to which the present invention pertains from the detailed description and examples of the present disclosure will be construed as being within the scope of the present disclosure.
[0094] The detailed description in this specification should be construed as illustrative in all aspects and not restrictive. The scope of the present disclosure should be determined by the appended claims and their legal equivalents, and all changes falling within the meaning and equivalents of the appended claims are intended to be covered herein.
[0095] Now, preferred embodiments of the present disclosure will be described in detail, and examples of these embodiments are illustrated in the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to explain the exemplary embodiments of the present disclosure, rather than showing the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. Although most of the terms used in this specification have been selected from commonly used general terms in the art, the applicant has arbitrarily selected some terms, and their meanings will be explained in detail as needed in the following description. Therefore, the present disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings. Additionally, the following drawings and detailed description should not be construed as limited to the specifically described embodiments, but should be construed as including equivalents or alternatives to the embodiments described in the drawings and detailed description.
[0096] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown.
[0097] Figure 1 The point cloud content providing system illustrated in may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 can perform wired or wireless communication to transmit and receive point cloud data.
[0098] The point cloud data transmitting device 10000 according to an embodiment can protect and process a point cloud video (or point cloud content) and transmit the point cloud video (or point cloud content). According to an embodiment, the transmitting device 10000 can include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or server. According to an embodiment, the transmitting device 10000 can include a device, a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server configured to perform communication with a base station and / or other wireless devices using a radio access technology (e.g., 5G new radio access technology (NR), long term evolution (LTE)).
[0099] The transmitting device 10000 according to an embodiment includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002, and / or a transmitter (or communication module) 10003.
[0100] The point cloud video acquisition unit 10001 according to an embodiment acquires a point cloud video through a process such as capture, synthesis, or generation. The point cloud video is point cloud content represented by a point cloud that is a set of points in a 3D space, and may be referred to as point cloud video data. The point cloud video according to an embodiment may include one or more frames. A frame represents a still image / picture. Thus, the point cloud video may include point cloud images / frames / pictures, and may be referred to as point cloud images, frames, or pictures.
[0101] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 may encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiment. The point cloud video encoder 10002 may output a bitstream containing the encoded point cloud video data. The bitstream may contain not only the encoded point cloud video data, but also signaling information related to the encoding of the point cloud video data.
[0102] The transmitter 10003 according to an embodiment transmits the bitstream containing the encoded point cloud video data. The bitstream according to an embodiment is encapsulated in a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not shown in the figure, the transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform an encapsulation operation. According to an embodiment, the encapsulator may be included in the transmitter 10003. According to an embodiment, the file or segment may be sent to the receiving device 10004 through a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to an embodiment is capable of wired / wireless communication with the receiving device 10004 (or the receiver 10005) through networks such as 4G, 5G, 6G, etc. Additionally, the transmitter may perform necessary data processing operations according to the network system (e.g., a 4G, 5G, or 6G communication network system). The transmitting device 10000 may send the encapsulated data in an on-demand manner.
[0103] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include a device, a robot, a vehicle, an AR / VR / XR device, a portable device, a household appliance, an Internet of Things (IoT) device, and an AI device / server configured to perform communication with a base station and / or other wireless devices using a radio access technology (e.g., 5G New Radio (NR), Long-Term Evolution (LTE)).
[0104] The receiver 10005 according to an embodiment receives a bitstream containing point cloud video data or a file / segment encapsulating the bitstream from a network or a storage medium. The receiver 10005 may perform necessary data processing according to a network system (e.g., a 4G, 5G, 6G, etc. communication network system). The receiver 10005 according to an embodiment may unpack the received file / segment and output the bitstream. According to an embodiment, the receiver 10005 may include a de-packager (or de-packaging module) configured to perform a de-packaging operation. The de-packager may be implemented as an element (or component or module) separate from the receiver 10005.
[0105] The point cloud video decoder 10006 decodes the bitstream containing the point cloud video data. The point cloud video decoder 10006 may decode the point cloud video data according to a method of encoding the point cloud video data (e.g., in the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 may decode the point cloud video data by performing point cloud decompression encoding, which is the reverse process of point cloud compression. The point cloud decompression encoding includes G-PCC encoding.
[0106] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 may output point cloud content by rendering not only the point cloud video data but also audio data. According to an embodiment, the renderer 10007 may include a display configured to display the point cloud content. According to an embodiment, the display may be implemented as a separate device or component rather than being included in the renderer 10007.
[0107] The arrow indicated by the dashed line in the figure represents the transmission path of the feedback information acquired by the receiving device 10004. The feedback information is information used to reflect the interactivity of the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). In particular, when the point cloud content is the content of a service that requires interaction with the user (e.g., an autonomous driving service, etc.), the feedback information can be provided to the content sender (e.g., the transmitting device 10000) and / or the service provider. According to an embodiment, the feedback information can be used in the receiving device 10004 and the transmitting device 10000, or may not be provided.
[0108] The head orientation information according to an embodiment is information about the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to an embodiment can calculate the viewport information based on the head orientation information. The viewport information can be information related to the area of the point cloud video that the user is watching. The viewing point is the point through which the user is watching the point cloud video and can refer to the center point of the viewport area. That is, the viewport is an area centered on the viewing point, and the size and shape of this area can be determined by the field of view (FOV). Therefore, in addition to the head orientation information, the receiving device 10004 can also extract the viewport information based on the vertical or horizontal FOV supported by the device. In addition, the receiving device 10004 performs gaze analysis, etc., to check the way the user consumes the point cloud, the area where the user gazes in the point cloud video, the gaze time, etc. According to an embodiment, the receiving device 10004 can send the feedback information including the gaze analysis result to the transmitting device 10000. The feedback information according to an embodiment can be obtained in the rendering and / or display processing. The feedback information according to an embodiment can be protected by one or more sensors included in the receiving device 10004. According to an embodiment, the feedback information can be protected by the renderer 10007 or a separate external element (or device, component, etc.).
[0109] Figure 1 The dashed line in represents the process of transmitting the feedback information protected by the renderer 10007. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can send the feedback information to the transmitting device 10000. The transmitting device 10000 (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process the necessary data (e.g., the point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data and provide the point cloud content to the user.
[0110] According to an embodiment, the transmitting device 10000 may be referred to as an encoder, a transmitting device, a transmitter, a transmitting system, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, a receiving system, etc.
[0111] (Through a series of processes of acquisition / encoding / transmission / decoding / rendering) In the point cloud content providing system according to an embodiment Figure 1 The point cloud data processed can be referred to as point cloud content data or point cloud video data. According to an embodiment, the point cloud content data can be used as a concept encompassing metadata or signaling information related to the point cloud data.
[0112] Figure 1 The elements of the point cloud content providing system illustrated in can be implemented by hardware, software, a processor, and / or a combination thereof.
[0113] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.
[0114] Figure 2 The block diagram of shows Figure 1 the operations of the point cloud content providing system described in. As described above, the point cloud content providing system can process point cloud data based on point cloud compression coding (e.g., G-PCC).
[0115] The point cloud content providing system according to an embodiment (e.g., the point cloud transmitting device 10000 or the point cloud video acquisition unit 10001) can acquire a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system for representing a 3D space. The point cloud video according to an embodiment may include a Ply (Polygon File Format or Stanford Triangle Format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. The geometry includes the position of the points. The position of each point can be represented by parameters (e.g., the values of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system composed of the X, Y, and Z axes). The attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc. for each point). A point has one or more attributes. For example, a point can have an attribute as color or two attributes as color and reflectivity.
[0116] According to an embodiment, the geometry can be referred to as position, geometric information, geometric data, etc., and the attributes can be referred to as attributes, attribute information, attribute data, etc.
[0117] A point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquisition unit 10001) may obtain point cloud data according to information related to the acquisition process of the point cloud video (e.g., depth information, color information, etc.).
[0118] The point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the point cloud video encoder 10002) may perform an encoding process (20001) on the point cloud data. The point cloud content providing system may encode the point cloud data based on point cloud compression encoding. As described above, the point cloud data may include the geometric structure and attributes of points. Therefore, the point cloud content providing system may perform geometric encoding for encoding the geometric structure and output a geometric bitstream. The point cloud content providing system may perform attribute encoding for encoding the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system may perform attribute encoding based on geometric encoding. The geometric bitstream and the attribute bitstream according to an embodiment may be multiplexed and output as one bitstream. The bitstream according to an embodiment may further include signaling information related to geometric encoding and attribute encoding.
[0119] The point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the transmitter 10003) may transmit the encoded point cloud data. As Figure 1 illustrated, the encoded point cloud data may be represented by a geometric bitstream and an attribute bitstream. In addition, the encoded point cloud data may be transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric encoding and attribute encoding). The point cloud content providing system may encapsulate the bitstream carrying the encoded point cloud data and transmit the bitstream in the form of a file or a segment.
[0120] The point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the receiver 10005) may receive a bitstream containing the encoded point cloud data. In addition, the point cloud content providing system (e.g., the receiving device 10004 or the receiver 10005) may demultiplex the bitstream.
[0121] A point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) can decode the encoded point cloud data (e.g., the geometry bitstream, the attribute bitstream) transmitted in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) can decode the point cloud video data based on the signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) can decode the geometry bitstream to reconstruct the positions of the points (geometry structure). The point cloud content providing system can reconstruct the attributes of the points by decoding the attribute bitstream based on the reconstructed geometry structure. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) can reconstruct the point cloud video based on the positions according to the reconstructed geometry structure and the decoded attributes.
[0122] The point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the renderer 10007) can render the decoded point cloud data. The point cloud content providing system (e.g., the receiving device 10004 or the renderer 10007) can use various rendering methods to render the geometry structure and attributes decoded by the decoding process. The points in the point cloud can be rendered as vertices with a certain thickness, cubes with a specific minimum size centered at the corresponding vertex positions, or circles centered at the corresponding vertex positions. All or part of the rendered point cloud content is provided to the user through a display (e.g., a VR / AR display, a common display, etc.).
[0123] The point cloud content providing system according to an embodiment (e.g., the receiving device 10004) can protect the feedback information (20005). The point cloud content providing system can encode and / or decode the point cloud data based on the feedback information. The feedback information and operations of the point cloud content providing system according to an embodiment are the same as the feedback information and operations described in the reference Figure 1 and thus the detailed description thereof is omitted.
[0124] Figure 3 An exemplary process of capturing a point cloud video according to an embodiment is illustrated.
[0125] Figure 3 Illustrated in the reference Figures 1 to 2 is the exemplary point cloud video capturing process of the point cloud content providing system described.
[0126] The point cloud content includes a point cloud video (image and / or video) representing objects and / or environments located in various 3D spaces (e.g., a 3D space representing a real environment, a 3D space representing a virtual environment, etc.). Therefore, a point cloud content providing system according to an embodiment can use one or more cameras (e.g., an infrared camera capable of protecting depth information, an RGB camera capable of extracting color information corresponding to depth information, etc.), a projector (e.g., an infrared pattern projector for protecting depth information), LiDAR, etc. to capture a point cloud video. A point cloud content providing system according to an embodiment can extract the shape of a geometric structure composed of points in a 3D space from the depth information, and extract the attributes of each point from the color information to protect the point cloud data. Images and / or videos according to an embodiment can be captured based on at least one of an in-facing technique and an out-facing technique.
[0127] Figure 3 The left part of [Figure] illustrates the in-facing technique. The in-facing technique refers to a technique of capturing an image of a central object using one or more cameras (or camera sensors) arranged around the central object. Point cloud content (e.g., VR / AR content that provides a 360-degree image of an object (e.g., a key object such as a character, player, object, or actor) to a user) can be generated using the in-facing technique.
[0128] Figure 3 The right part of [Figure] illustrates the out-facing technique. The out-facing technique refers to a technique of capturing an image of the environment around a central object rather than the central object using one or more cameras (or camera sensors) arranged around the central object. Point cloud content for providing the surrounding environment as seen from the user's perspective (e.g., content representing the external environment of a user that can be provided to an autonomous vehicle) can be generated using the out-facing technique.
[0129] As Figure 3 shown in [Figure], point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate system is different among each camera. Therefore, the point cloud content providing system can calibrate one or more cameras before the capture operation to set a global coordinate system. In addition, the point cloud content providing system can generate point cloud content by synthesizing an arbitrary image and / or video with the images and / or videos captured by the above capture techniques. The point cloud content providing system cannot perform the capture operations described in Figure 3 when generating point cloud content representing a virtual space. A point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unnecessary regions (e.g., the background), identify the space to which the captured images and / or videos are connected, and perform an operation to fill in spatial holes when there are spatial holes.
[0130] The point cloud content providing system can generate a piece of point cloud content by performing coordinate transformation on the points of the point cloud video protected by each camera. The point cloud content providing system can perform coordinate transformation on the points based on the coordinates of each camera position. Therefore, the point cloud content providing system can generate point cloud content representing a wide range of content, or can generate point cloud content with a high density of points.
[0131] Figure 4 An exemplary point cloud video encoder according to an embodiment is illustrated.
[0132] Figure 4 Shows Figure 1 An example of the point cloud video encoder 10002. The point cloud video encoder reconstructs and encodes point cloud data (e.g., the position and / or attributes of points) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) according to network conditions or applications. When the total size of the point cloud content is large (e.g., for 30fps, a point cloud content of 60Gbps is given), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bit rate to provide the point cloud content according to the network environment, etc.
[0133] As described in reference Figures 1 to 2 The point cloud video encoder can perform geometric coding and attribute coding. Geometric coding is performed before attribute coding.
[0134] The point cloud video encoder according to an embodiment includes a coordinate transformation unit 40000, a quantization unit 40001, an octree analysis unit 40002, a surface approximation analysis unit 40003, an arithmetic encoder 40004, a geometric reconstruction unit 40005, a color transformation unit 40006, an attribute transformation unit 40007, a RAHT unit 40008, a LOD generation unit 40009, a lifting transformation unit 40010, a coefficient quantization unit 40011, and / or an arithmetic encoder 40012.
[0135] The coordinate transformation unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometric reconstruction unit 40005 can perform geometric coding. Geometric coding according to an embodiment can include octree geometric coding, direct coding, trisoup (trisoup) geometric coding, and entropy coding. Direct coding and trisoup geometric coding are selectively or combinatorially applied. Geometric coding is not limited to the above examples.
[0136] As shown in the figure, the coordinate transformation unit 40000 according to an embodiment receives a position and transforms it into coordinates. For example, the position may be transformed into position information in a three-dimensional space (e.g., a three-dimensional space represented by an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment may be referred to as geometric information.
[0137] The quantization unit 40001 according to an embodiment quantizes the geometric information. For example, the quantization unit 40001 may quantize points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantization unit 40001 performs the following quantization operation: multiplying the difference between the position value of each point and the minimum position value by a preset quantization scaling value, and then finding the closest integer value by rounding the value obtained by the multiplication. Thus, one or more points may have the same quantized position (or position value). The quantization unit 40001 according to an embodiment performs voxelization based on the quantized position to reconstruct the quantized points. Voxelization means the minimum unit for representing position information in a 3D space. The points of the point cloud content (or 3D point cloud video) according to an embodiment may be included in one or more voxels. The term voxel, which is a compound word of volume and pixel, refers to a 3D cubic space generated when a 3D space is divided into units (unit = 1.0) based on axes representing the 3D space (e.g., the X axis, the Y axis, and the Z axis). The quantization unit 40001 may match a group of points in the 3D space with voxels. According to an embodiment, one voxel may include only one point. According to an embodiment, one voxel may include one or more points. To represent a voxel as a point, the position of the center point of the voxel may be set based on the position of one or more points included in the voxel. In this case, the attributes of all positions included in one voxel may be combined and assigned to the voxel.
[0138] The octree analysis unit 40002 according to an embodiment performs octree geometry encoding (or octree encoding) to present the voxels in an octree structure. The octree structure represents points that match the voxels based on the octree structure.
[0139] The surface approximation analysis unit 40003 according to an embodiment may analyze and approximate the octree. The octree analysis and approximation according to an embodiment is a process of analyzing a region containing multiple points to efficiently provide the octree and voxelization.
[0140] The arithmetic encoder 40004 according to an embodiment performs entropy encoding on the octree and / or the approximated octree. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric bitstream is generated.
[0141] The color transformation unit 40006, the attribute transformation unit 40007, the RAHT unit 40008, the LOD generation unit 40009, the lifting transformation unit 40010, the coefficient quantization unit 40011, and / or the arithmetic encoder 40012 perform attribute encoding. As described above, a point may have one or more attributes. The attribute encoding according to the embodiment is also applied to the attributes that a point has. However, when an attribute (e.g., color) includes one or more elements, the attribute encoding is independently applied to each element. The attribute encoding according to the embodiment includes color transformation encoding, attribute transformation encoding, region adaptive hierarchical transformation (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) encoding, and interpolation-based hierarchical nearest neighbor prediction encoding with an update / lifting step (lifting transformation). According to the point cloud content, the above RAHT encoding, prediction transformation encoding, and lifting transformation encoding can be selectively used, or a combination of one or more encoding schemes can be used. The attribute encoding according to the embodiment is not limited to the above examples.
[0142] The color transformation unit 40006 according to the embodiment performs color transformation encoding of the color value (or texture) included in the transformation attribute. For example, the color transformation unit 40006 may transform the format of the color information (e.g., from RGB to YCbCr). The operation of the color transformation unit 40006 according to the embodiment can be optionally applied according to the color value included in the attribute.
[0143] The geometry reconstruction unit 40005 according to the embodiment reconstructs (decompresses) an octree and / or an approximate octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of points. The reconstructed octree / voxel may be referred to as a reconstructed geometric structure (restored geometric structure).
[0144] The attribute transformation unit 40007 according to the embodiment performs an attribute transformation to transform an attribute based on the position where geometric encoding has not been performed and / or the reconstructed geometric structure. As described above, since the attribute depends on the geometric structure, the attribute transformation unit 40007 can transform the attribute based on the reconstructed geometric information. For example, based on the position value of the points included in the voxel, the attribute transformation unit 40007 can transform the attributes of the points at that position. As described above, when the position of the voxel center is set based on the position of one or more points included in the voxel, the attribute transformation unit 40007 transforms the attributes of the one or more points. When performing trisoup geometric encoding, the attribute transformation unit 40007 can transform the attribute based on the trisoup geometric encoding.
[0145] The attribute transformation unit 40007 can perform attribute transformation by calculating the average of the attributes or attribute values (e.g., the color or reflectance of each point) of the neighbor points within a specific position / radius from the position (or position value) of the center of each voxel. The attribute transformation unit 40007 can apply weights according to the distance from the center to each point when calculating the average. Thus, each voxel has a position and a calculated attribute (or attribute value).
[0146] The attribute transformation unit 40007 can search for neighbor points existing within a specific position / radius from the position of the center of each voxel based on a K-D tree or a Morton code. A K-D tree is a binary search tree and supports a data structure for managing points based on position so that a nearest neighbor search (NNS) can be performed quickly. The Morton code is generated by presenting the coordinates (e.g., (x, y, z)) representing the 3D positions of all points as bit values and mixing the bits. For example, when the coordinates representing the position of a point are (5, 9, 1), the bit values of the coordinates are (0101, 1001, 0001). Mixing the bit values in the order of z, y, and x according to the bit indices produces 010001000111. This value is represented as a decimal number 1095. That is, the Morton code value of the point with coordinates (5, 9, 1) is 1095. The attribute transformation unit 40007 can sort the points based on the Morton code values and perform NNS through depth-first traversal processing. After the attribute transformation operation, when NNS is required in another transformation process for attribute encoding, a K-D tree or a Morton code is used.
[0147] As shown in the figure, the transformed attributes are input to the RAHT unit 40008 and / or the LOD generation unit 40009.
[0148] The RAHT unit 40008 according to an embodiment performs RAHT encoding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT unit 40008 can predict the attribute information of the higher-level nodes in the octree based on the attribute information associated with the lower-level nodes in the octree.
[0149] The LOD generation unit 40009 according to an embodiment generates a level of detail (LOD). The LOD according to an embodiment is the level of detail of the point cloud content. As the LOD value decreases, it indicates a decrease in the level of detail of the point cloud content. As the LOD value increases, it indicates an enhancement in the level of detail of the point cloud content. The points can be classified according to the LOD.
[0150] The lifting transformation unit 40010 according to an embodiment performs lifting transformation encoding for transforming the attributes of the point cloud based on weights. As described above, the lifting transformation encoding can be optionally applied.
[0151] The coefficient quantization unit 40011 according to the embodiment quantizes the attribute after the attribute encoding based on the coefficient.
[0152] The arithmetic encoder 40012 according to the embodiment encodes the quantized attributes based on arithmetic coding.
[0153] Although not shown in this figure, Figure 4 The elements of the point cloud video encoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud content providing device. The one or more processors may perform the above Figure 4 At least one of the operations and / or functions of the elements of the point cloud video encoder. In addition, one or more processors can operate or execute a set of software programs and / or instructions to perform Figure 4 The operation and / or functionality of the elements of the point cloud video encoder. One or more memories according to an embodiment may include a high-speed random access memory, or include a non-volatile memory (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices).
[0154] Figure 5 An example of a voxel according to an embodiment is shown.
[0155] Figure 5 2 shows voxels located in a 3D space represented by a coordinate system consisting of three axes, namely, an X-axis, a Y-axis, and a Z-axis. Figure 4 As described, the point cloud video encoder (eg, quantization unit 40001) may perform voxelization. A voxel refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on axes representing the 3D space (eg, X-axis, Y-axis, and Z-axis). Figure 5 An example of a voxel generated by an octree structure is shown, in which the octree consists of two poles (0, 0, 0) and (2 d ,2 d ,2 d ) is recursively subdivided. A voxel consists of at least one point. The spatial coordinates of the voxel can be estimated based on the positional relationship with the voxel group. As mentioned above, the voxel has properties like the pixels of a 2D image / video (such as color or reflectivity). The details of the voxel are similar to those of the reference Figure 4 The details described are the same, so their description is omitted.
[0156] Figure 6 Examples of octrees and occupancy codes are shown according to an embodiment.
[0157] As reference Figures 1 to 4As described, the point cloud content providing system (point cloud video encoder 10002) or the octree analysis unit 40002 of the point cloud video encoder performs octree geometry coding (or octree coding) based on the octree structure to efficiently manage the regions and / or positions of voxels.
[0158] Figure 6 The upper part of shows the octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). The octree structure is created by recursively subdividing the cube axis-aligned bounding box defined by two extreme points (0, 0, 0) and (2 d , 2 d , 2 d ). Here, 2 d can be set to the value of the minimum bounding box that encloses all the points of the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined in Equation 1. In Equation 1, (x int n , y int n , z int n ) represents the position (or position value) of the quantized point.
[0159] [Equation 1]
[0160]
[0161] As Figure 6 shown in the middle of the upper part of, the entire 3D space can be divided into eight spaces according to the partition. Each divided space is represented by a cube with six faces. As Figure 6 shown in the upper right side of, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). Thus, each space is divided into eight smaller spaces. The divided smaller spaces are also represented by cubes with six faces. This division scheme is applied until the leaf nodes of the octree become voxels.
[0162] Figure 6 The lower part of shows the octree occupancy code. The octree occupancy code is generated to indicate whether each of the eight divided spaces resulting from dividing a space contains at least one point. Thus, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the divided space, and the child node has a value of 1 bit. Thus, the occupancy code is represented as an 8-bit code. That is, when at least one point is included in the space corresponding to the child node, the node is given the value 1. When no point is included in the space corresponding to the child node (the space is empty), the node is given the value 0. Since Figure 6The occupancy code shown is 00100001, so it indicates that the spaces corresponding to the third and eighth child nodes among the eight child nodes each contain at least one point. As shown in the figure, each of the third and eighth child nodes has 8 child nodes, and the child nodes are represented by 8-bit occupancy codes. The figure shows that the occupancy code of the third child node is 10000111, and the occupancy code of the eighth child node is 01001111. A point cloud video encoder according to an embodiment (e.g., arithmetic encoder 40004) may perform entropy encoding on the occupancy code. To improve compression efficiency, the point cloud video encoder may perform intra / inter-frame encoding on the occupancy code. A receiving device according to an embodiment (e.g., receiving device 10004 or point cloud video decoder 10006) reconstructs the octree based on the occupancy code.
[0163] A point cloud video encoder according to an embodiment (e.g., octree analysis unit 40002) may perform voxelization and octree encoding to store the positions of points. However, points are not always evenly distributed in 3D space, so there will be specific regions where there are fewer points. Therefore, it is inefficient to perform voxelization on the entire 3D space. For example, when a specific region contains fewer points, voxelization does not need to be performed in the specific region.
[0164] Therefore, for the above specific region (or a node other than the leaf node of the octree), a point cloud video encoder according to an embodiment may skip voxelization and perform direct encoding to directly encode the positions of the points included in the specific region. The coordinates of the points directly encoded according to an embodiment are referred to as the direct coding mode (DCM). A point cloud video encoder according to an embodiment may also perform trisoup geometry encoding based on a surface model to reconstruct the positions of points in the specific region (or node) based on voxels. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangular meshes. Therefore, the point cloud video decoder can generate a point cloud from the mesh surface. The trisoup geometry encoding according to an embodiment and the direct encoding may be selectively performed. In addition, the trisoup geometry encoding and the direct encoding according to an embodiment may be performed in combination with octree geometry encoding (or octree encoding).
[0165] To perform direct encoding, the option to use the direct mode to apply direct encoding should be enabled. The node to which direct encoding will be applied is not a leaf node, and there should be fewer than a threshold number of points within the specific node. In addition, the total number of points to which direct encoding will be applied should not exceed a preset threshold. When the above conditions are met, a point cloud video encoder according to an embodiment (or arithmetic encoder 40004) may perform entropy encoding on the positions (or position values) of the points.
[0166] A point cloud video encoder according to an embodiment (e.g., the surface approximation analysis unit 40003) may determine a specific level of the octree (a level less than the depth d of the octree), and may start from that level to perform trisoup geometry encoding using a surface model to reconstruct the positions of points in the region of the node based on voxels (trisoup mode). The point cloud video encoder according to an embodiment may specify the level at which trisoup geometry encoding will be applied. For example, when the specific level is equal to the depth of the octree, the point cloud video encoder does not operate in the trisoup mode. In other words, the point cloud video encoder according to an embodiment may operate in the trisoup mode only when the specified level is less than the depth value of the octree. The 3D cubic region of the node at the specified level according to an embodiment is called a block. A block may include one or more voxels. A block or a voxel may correspond to a brick. The geometric structure is represented as a surface within each block. The surface according to an embodiment may intersect each edge of the block at most once.
[0167] A block has 12 edges, so there are at least 12 intersection points in a block. Each intersection point is called a vertex (or apex point). A vertex existing along an edge is detected when there is at least one occupied voxel adjacent to that edge among all the blocks sharing the edge. An occupied voxel according to an embodiment refers to a voxel containing a point. The position of the vertex detected along the edge is the average position of the edges of all the voxels adjacent to that edge among all the blocks sharing the edge.
[0168] Once the vertex is detected, the point cloud video encoder according to an embodiment may perform entropy encoding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the vertex position value (the relative position value within the edge). When applying trisoup geometry encoding, the point cloud video encoder according to an embodiment (e.g., the geometry reconstruction unit 40005) may generate a restored geometric structure (reconstructed geometric structure) by performing triangle reconstruction, upsampling, and voxelization processing.
[0169] The vertex at the edge of the block determines the surface passing through the block. The surface according to an embodiment is a non-planar polygon. In the triangle reconstruction process, the surface represented by a triangle is reconstructed based on the starting point of the edge, the direction vector of the edge, and the vertex position value. According to Equation 2, the triangle reconstruction process is performed by the following operations: ① calculating the centroid value of each vertex, ② subtracting the center value from each vertex value, and ③ estimating the sum of the squares of the values obtained by the subtraction.
[0170] [Equation 2]
[0171]
[0172] Then, estimate the minimum value of the sum, and perform projection processing based on the axis with the minimum value. For example, when the element x is the smallest, each vertex is projected onto the x-axis relative to the center of the block and projected onto the (y,z) plane. When the values obtained by projecting onto the (y,z) plane are (ai,bi), estimate the value of θ by atan2(bi,ai), and sort the vertices according to the value of θ. Table 1 below shows the vertex combinations for creating triangles according to the number of vertices. The vertices are sorted from 1 to n. Table 1 below shows that for four vertices, two triangles can be constructed according to the vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 among the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 among the sorted vertices.
[0173] Table 1. Triangles formed by vertices sorted as 1,…,n [Table 1]
[0174]
[0175] Perform upsampling processing to add points in the middle along the sides of the triangles and perform voxelization. The added points are generated based on the upsampling factor and the width of the block. The added points are called refined vertices. The point cloud video encoder according to the embodiment can perform voxelization on the refined vertices. In addition, the point cloud video encoder can perform attribute encoding based on the voxelization position (or position value). Figure 7 An example of the neighbor node pattern according to the embodiment is illustrated. To improve the compression efficiency of the point cloud video, the point cloud video encoder according to the embodiment can perform entropy encoding based on context adaptive arithmetic coding.
[0176] As described in reference Figures 1 to 6 described, Figure 1 the point cloud content providing system or the point cloud video encoder 10002 or Figure 4 the point cloud video encoder or the arithmetic encoder 40004 of 3 can immediately perform entropy encoding on the occupancy code. In addition, the point cloud content providing system or the point cloud video encoder can perform entropy encoding (intra-frame encoding) based on the occupancy code of the current node and the occupancy of the neighbor nodes, or perform entropy encoding (inter-frame encoding) based on the occupancy code of the previous frame. The frame according to the embodiment represents a set of point cloud videos generated simultaneously. The compression efficiency of the intra-frame encoding / inter-frame encoding according to the embodiment can depend on the number of neighbor nodes referred to. When the number of bits increases, the operation becomes complex, but the encoding can be biased to one side, thereby increasing the compression efficiency. For example, when given 3-bit context, 2 3 = 8 methods need to be used to perform encoding. The parts divided for encoding affect the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be satisfied.
[0177] Figure 7 Illustrates a process of obtaining an occupancy pattern based on the occupancy of neighboring nodes. A point cloud video encoder according to an embodiment determines the occupancy of neighboring nodes of each node of an octree and obtains a value of a neighbor pattern. The neighbor node pattern is used to infer the occupancy pattern of the node. Figure 7 The upper part of shows a cube corresponding to a node (the cube in the middle) and six cubes (neighboring nodes) that share at least one face with the cube. The nodes shown in the figure are nodes at the same depth. The numbers shown in the figure respectively represent the weights (1, 2, 4, 8, 16, and 32) associated with the six nodes. The weights are assigned in sequence according to the positions of the neighboring nodes.
[0178] Figure 7 The lower part of shows the neighbor node pattern values. The neighbor node pattern value is the sum of the values obtained by multiplying the weights of the occupied neighboring nodes (neighboring nodes with points). Therefore, the neighbor node pattern value is from 0 to 63. When the neighbor node pattern value is 0, it indicates that there are no nodes with points among the neighboring nodes of the node (unoccupied node). When the neighbor node pattern value is 63, it indicates that all neighboring nodes are occupied nodes. As shown in the figure, since the neighboring nodes assigned weights 1, 2, 4, and 8 are occupied nodes, the neighbor node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud video encoder can perform encoding according to the neighbor node pattern value (for example, when the neighbor node pattern value is 63, 64 kinds of encoding can be performed). According to an embodiment, the point cloud video encoder can reduce the encoding complexity by changing the neighbor node pattern value (for example, based on a table through which 64 is changed to 10 or 6).
[0179] Figure 8 Illustrates an example of the point configuration in each LOD according to an embodiment.
[0180] As referred to Figures 1 to 7 described, before performing attribute encoding, the encoded geometry is reconstructed (decompressed). When direct encoding is applied, the geometry reconstruction operation may include changing the placement of the directly encoded points (for example, placing the directly encoded points at the front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since the attributes depend on the geometry, the attribute encoding is performed based on the reconstructed geometry.
[0181] A point cloud video encoder (for example, the LOD generation unit 40009) can classify (reorganize or group) points by LOD. Figure 8 Shows the point cloud content corresponding to the LOD. Figure 8 The leftmost picture in represents the original point cloud content. Figure 8 The second picture from the left represents the distribution of points in the lowest LOD, andFigure 8 The rightmost picture in Figure 8 shows the distribution of points in the highest LOD. That is, the points in the lowest LOD are sparsely distributed, and the points in the highest LOD are densely distributed. That is, as the LOD increases in the direction indicated by the arrow pointed at the bottom by
[0182] Figure 9 An example of the point configuration for each LOD according to an embodiment is illustrated.
[0183] As referred to Figures 1 to 8 described, a point cloud content providing system or a point cloud video encoder (e.g., Figure 1 the point cloud video encoder 10002 of Figure 4 the point cloud video encoder or the LOD generation unit 40009) can generate LODs. LODs are generated by reorganizing points into a set of refinement levels according to a set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud video encoder but also by the point cloud video decoder.
[0184] Figure 9 The upper part of Figure 9 shows examples of points (P0 to P9) of point cloud content distributed in a 3D space. In Figure 9 the original order represents the order of points P0 to P9 before LOD generation. In Figure 9 the LOD-based order represents the order of points generated according to LOD. Points are reorganized by LOD. Additionally, the high LOD contains points belonging to the lower LOD. As shown in
[0185] As referred to Figure 4 described, the point cloud video encoder according to an embodiment can selectively or combinatorially perform LOD-based predictive transform coding, LOD-based lifting transform coding, and RAHT transform coding.
[0186] The point cloud video encoder according to an embodiment can generate predictors for points to perform LOD-based predictive transform coding to set the prediction attributes (or prediction attribute values) of each point. That is, N predictors can be generated for N points. The predictor according to an embodiment can calculate weights (= 1 / distance) based on the LOD value of each point, the indexed information related to neighbor points existing within a set distance for each LOD, and the distance to neighbor points.
[0187] The predicted attribute (or attribute value) according to the embodiment is set to the average value of the values obtained by multiplying the attributes (or attribute values) of the neighboring points (e.g., color, reflectance, etc.) set in the predictor of each point by the weights (or weight values) calculated based on the distances to each neighboring point. A point cloud video encoder according to an embodiment (e.g., the coefficient quantization unit 40011) may quantize and inverse-quantize the residual of each point (which may be referred to as residual attribute, residual attribute value, attribute prediction residual value, or prediction error attribute value, etc.) obtained by subtracting the predicted attribute (or attribute value) of each point from the attribute of each point (i.e., the original attribute value). Configure the quantization process performed on the residual attribute values in the transmitting device as shown in Table 2. Configure the inverse quantization process performed on the residual attribute values in the receiving device as shown in Table 3.
[0188] [Table 2]
[0189] int PCCQuantization(int value, int quantStep) { if (value >= 0) { return floor(value / quantStep + 1.0 / 3.0); } else { return -floor(-value / quantStep + 1.0 / 3.0); } }
[0190] [Table 3]
[0191] int PCCInverseQuantization(int value, int quantStep) { if (quantStep == 0) { return value; } else { return value * quantStep; } }
[0192] When the predictor of each point has neighboring points, a point cloud video encoder according to an embodiment (e.g., the arithmetic encoder 40012) may perform entropy encoding on the quantized and inverse-quantized residual attribute values as described above. When the predictor of each point has no neighboring points, a point cloud video encoder according to an embodiment (e.g., the arithmetic encoder 40012) may perform entropy encoding on the attribute of the corresponding point without performing the above operations. A point cloud video encoder according to an embodiment (e.g., the lifting transform unit 40010) may generate the predictor of each point, set the calculated LOD, register the neighboring points in the predictor, and set weights according to the distances to the neighboring points to perform lifting transform encoding. The lifting transform encoding according to the embodiment is similar to the above-described prediction transform encoding, but the difference is that the weights are applied to the attribute values accumulatively. Configure the process of applying weights to the attribute values accumulatively according to the embodiment as follows.
[0193] 1) Create an array Quantization Weight (QW) (quantization weight) for storing the weight values of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighboring node registered in the predictor by the weight of the predictor of the current point, and add the values obtained by the multiplication.
[0194] 2) Lifting prediction process: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.
[0195] 3) Create a temporary array called updateweight, and update and initialize this temporary array to zero.
[0196] 4) Add the weights calculated by multiplying the weights calculated for all predictors by the weights corresponding to the predictor indices stored in QW to the update weight array accumulatively as the indices of neighbor nodes. Add the values obtained by multiplying the attribute values of the indices of neighbor nodes by the calculated weights to the update array accumulatively.
[0197] 5) Boost the update process: Divide the attribute values of the update array for all predictors by the weight values of the update weight array of the predictor indices, and add the existing attribute values to the values obtained by the division.
[0198] 6) Calculate the predicted attributes by multiplying the attribute values updated by the boosting update process for all predictors by the weights updated by the boosting prediction process (stored in QW). Quantize the predicted attribute values according to the point cloud video encoder according to the embodiment (e.g., coefficient quantization unit 40011). Additionally, the point cloud video encoder (e.g., arithmetic encoder 40012) performs entropy coding on the quantized attribute values.
[0199] The point cloud video encoder according to the embodiment (e.g., RAHT unit 40008) may perform RAHT transform coding, where the attributes of higher-level nodes are predicted using the attributes associated with lower-level nodes in the octree. RAHT transform coding is an example of intra-frame coding of attributes through octree backward scanning. The point cloud video encoder according to the embodiment scans the entire region from the voxels and repeats the merging process of merging voxels into larger blocks at each step until reaching the root node. The merging process according to the embodiment is only performed on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed in a higher mode directly above the empty node.
[0200] The following Equation 3 represents the RAHT transform matrix. In Equation 3, represents the average attribute value of the voxels at level l. It can be based on and to calculate and The weights of and
[0201] [Equation 3]
[0202]
[0203] Here, is the low-pass value and is used in the merging process at the next higher level. represents a high-pass coefficient. The high-pass coefficient in each step is quantized and undergoes entropy coding (e.g., encoded by an arithmetic coder 40012). The weights are calculated as as in Equation 4 by and to calculate the root node.
[0204] [Equation 4]
[0205]
[0206] The value of gDC is also quantized and undergoes entropy coding like the high-pass coefficient.
[0207] Figure 10 illustrates a point cloud video decoder according to an embodiment.
[0208] Figure 10 The point cloud video decoder illustrated in Figure 1 is an example of the point cloud video decoder 10006 described in Figure 1 and can perform operations the same as or similar to those of the point cloud video decoder 10006 illustrated in
[0209] Figure 11 illustrates a point cloud video decoder according to an embodiment.
[0210] Figure 11 The point cloud video decoder illustrated in Figure 10 is an example of the point cloud video decoder illustrated in Figures 1 to 9 and can perform decoding operations that are the inverse processing of the encoding operations of the point cloud video encoder illustrated in
[0211] As described with reference to Figure 1 and Figure 10 the point cloud video decoder can perform geometric decoding and attribute decoding. Geometric decoding is performed before attribute decoding.
[0212] The point cloud video decoder according to the embodiment includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesize octree) 11001, a surface approximation synthesizer (synthesize surface approximation) 11002, a geometry reconstruction unit (reconstruct geometric structure) 11003, a coordinate inverse transformer (inverse transform coordinates) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantization unit (inverse quantization) 11006, a RAHT transformer 11007, a LOD generator (generate LOD) 11008, an inverse lifting unit (inverse lifting) 11009, and / or an inverse color transformation unit (inverse transform color) 11010.
[0213] The arithmetic decoder 11000, the octree synthesizer 11001, the surface approximation synthesizer 11002, the geometry reconstruction unit 11003, and the coordinate inverse transformer 11004 may perform geometry decoding. The geometry decoding according to the embodiment may include direct decoding and trisoup geometry decoding. The direct decoding and trisoup geometry decoding are selectively applied. The geometry decoding is not limited to the above examples and is performed as the inverse process of the geometry encoding described for reference. Figures 1 to 9 The inverse process of the described geometry encoding is performed.
[0214] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic encoding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.
[0215] The octree synthesizer 11001 according to the embodiment may generate an octree by obtaining occupancy codes from the decoded geometry bitstream (or the information about the geometric structure protected as the decoding result). The occupancy codes are configured as described in the reference. Figures 1 to 9 configured in detail as described in the reference.
[0216] When trisoup geometry encoding is applied, the surface approximation synthesizer 11002 according to the embodiment may synthesize a surface based on the decoded geometric structure and / or the generated octree.
[0217] The geometry reconstruction unit 11003 according to the embodiment may regenerate a geometric structure based on the surface and / or the decoded geometric structure. As described in the reference, the direct encoding and trisoup geometry encoding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly imports and adds the position information of the points to which the direct encoding is applied. When trisoup geometry encoding is applied, the geometry reconstruction unit 11003 may reconstruct the geometric structure by performing the reconstruction operations (e.g., triangle reconstruction, upsampling, and voxelization) of the geometry reconstruction unit 40005. The details are the same as those described in the reference, so the description thereof is omitted. The reconstructed geometric structure may include a point cloud picture or frame that does not contain attributes. Figures 1 to 9 described in the reference, so the description thereof is omitted. The reconstructed geometric structure may include a point cloud picture or frame that does not contain attributes. Figure 6 The details are the same as those described in the reference, so the description thereof is omitted. The reconstructed geometric structure may include a point cloud picture or frame that does not contain attributes.
[0218] The coordinate inverse transformer 11004 according to the embodiment can obtain the position of a point by transforming coordinates based on the reconstructed geometric structure.
[0219] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transformer 11007, the LOD generation unit 11008, the inverse lifter 11009, and / or the inverse color transformation unit 11010 can perform the reference Figure 10 described attribute decoding. The attribute decoding according to the embodiment includes region adaptive hierarchical transform (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) decoding, and interpolation-based hierarchical nearest neighbor prediction decoding with an update / lifting step (lifting transform). The above three decoding schemes can be selectively used, or a combination of one or more decoding schemes can be used. The attribute decoding according to the embodiment is not limited to the above examples.
[0220] The arithmetic decoder 11005 according to the embodiment decodes the attribute bitstream by arithmetic coding.
[0221] The inverse quantization unit 11006 according to the embodiment inverse quantizes the information about the decoded attribute bitstream or attribute protected as a decoding result, and outputs the inverse quantized attribute (or attribute value). The inverse quantization can be selectively applied based on the attribute coding of the point cloud video encoder.
[0222] According to the embodiment, the RAHT transformer 11007, the LOD generation unit 11008, and / or the inverse lifter 11009 can process the reconstructed geometric structure and the inverse quantized attribute. As described above, the RAHT transformer 11007, the LOD generation unit 11008, and / or the inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud video encoder.
[0223] The inverse color transformation unit 11010 according to the embodiment performs inverse transform coding to inverse transform the color value (or texture) included in the decoded attribute. The operation of the inverse color transformation unit 11010 can be selectively performed based on the operation of the color transformation unit 40006 of the point cloud video encoder.
[0224] Although not shown in this figure, Figure 11 the elements of the point cloud video decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud content providing device. The one or more processors can execute the above Figure 11at least one or more of the operations and / or functions of the components of the point cloud video decoder. Additionally, one or more processors may operate or execute a set of software programs and / or instructions to perform Figure 11 the operations and / or functions of the components of the point cloud video decoder.
[0225] Figure 12 illustrates a transmitting device according to an embodiment.
[0226] Figure 12 The transmitting device shown in Figure 1 is an example of the transmitting device 10000 (or Figure 4 the point cloud video encoder) of Figure 12 The transmitting device illustrated in Figures 1 to 9 may perform one or more of the operations and methods that are the same as or similar to the operations and methods of the point cloud video encoder described with reference to
[0227] According to an embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 may perform operations and / or acquisition methods that are the same as or similar to the operations and / or acquisition method of the point cloud video acquisition unit 10001 (or the acquisition process 20000 described with reference to Figure 2 ).
[0228] The data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame encoding processor 12005, and arithmetic encoder 12006 perform geometric encoding. The geometric encoding according to an embodiment is the same as or similar to the geometric encoding described with reference to Figures 1 to 9 and thus a detailed description thereof is omitted.
[0229] According to an embodiment, the quantization processor 12001 quantizes geometric structures (e.g., position values of points). The operations and / or quantization of the quantization processor 12001 are the same as or similar to the operations and / or quantization of the quantization unit 40001 described with reference to Figure 4 . The details are the same as the details described with reference to Figures 1 to 9 .
[0230] According to an embodiment, the voxelization processor 12002 voxelizes the quantized position values of the points. The voxelization processor 12002 may perform operations and / or processes that are the same as or similar to the operations and / or processes of the quantization unit 40001 described with reference to Figure 4 and / or the voxelization process. The details are the same as the details described with reference to Figures 1 to 9 the description.
[0231] According to an embodiment, the octree occupancy code generator 12003 performs octree encoding on the voxelized positions of the points based on the octree structure. The octree occupancy code generator 12003 may generate occupancy codes. The octree occupancy code generator 12003 may perform operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud video encoder (or octree analysis unit 40002) described with reference to Figure 4 and Figure 6 the description. The details are the same as the details described with reference to Figures 1 to 9 the description.
[0232] According to an embodiment, the surface model processor 12004 may perform trisoup geometry encoding based on the surface model to reconstruct the positions of the points in a specific region (or node) based on voxels. The surface model processor 12004 may perform operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud video encoder (e.g., surface approximation analysis unit 40003) described with reference to Figure 4 the description. The details are the same as the details described with reference to Figures 1 to 9 the description.
[0233] According to an embodiment, the intra / inter-frame encoding processor 12005 may perform intra / inter-frame encoding on the point cloud data. The intra / inter-frame encoding processor 12005 may perform encoding that is the same as or similar to the intra / inter-frame encoding described with reference to Figure 7 the description. The details are the same as the details described with reference to Figure 7 the description. According to an embodiment, the intra / inter-frame encoding processor 12005 may be included in the arithmetic encoder 12006.
[0234] According to an embodiment, the arithmetic encoder 12006 performs entropy encoding on the octree and / or approximate octree of the point cloud data. For example, the encoding scheme includes arithmetic encoding. The arithmetic encoder 12006 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 40004.
[0235] The metadata processor 12007 according to the embodiment processes metadata (e.g., set values) regarding point cloud data and provides it to necessary processing procedures such as geometric coding and / or attribute coding. Additionally, the metadata processor 12007 according to the embodiment may generate and / or process signaling information related to geometric coding and / or attribute coding. The signaling information according to the embodiment may be encoded separately from geometric coding and / or attribute coding. The signaling information according to the embodiment may be interleaved.
[0236] The color transformation processor 12008, the attribute transformation processor 12009, the LOD / upsampling / RAHT transformation processor 12010, and the arithmetic coder 12011 perform attribute coding. The attribute coding according to the embodiment is the same as or similar to the attribute coding described in the reference Figures 1 to 9 and thus a detailed description thereof is omitted.
[0237] The color transformation processor 12008 according to the embodiment performs color transformation coding to transform the color values included in the attributes. The color transformation processor 12008 may perform color transformation coding based on the reconstructed geometry. The reconstructed geometry is the same as that described in the reference Figures 1 to 9 . Additionally, it performs operations and / or methods that are the same as or similar to the operations and / or methods of the color transformation unit 40006 described in the reference Figure 4 . A detailed description thereof is omitted.
[0238] The attribute transformation processor 12009 according to the embodiment performs attribute transformation to transform the attributes based on the reconstructed geometry and / or positions where geometric coding has not been performed. The attribute transformation processor 12009 performs operations and / or methods that are the same as or similar to the operations and / or methods of the attribute transformation unit 40007 described in the reference Figure 4 . A detailed description thereof is omitted. The LOD / upsampling / RAHT transformation processor 12010 according to the embodiment may encode the transformed attributes by any one or a combination of RAHT coding, predictive transformation coding, and upsampling transformation coding. The LOD / upsampling / RAHT transformation processor 12010 performs at least one of the operations that are the same as or similar to the operations of the RAHT unit 40008, the LOD generation unit 40009, and the upsampling transformation unit 40010 described in the reference Figure 4 . Additionally, the predictive transformation coding, the upsampling transformation coding, and the RAHT transformation coding are the same as those described in the reference Figures 1 to 9 and thus a detailed description thereof is omitted.
[0239] According to an embodiment, the arithmetic encoder 12011 can encode the encoded attributes based on arithmetic coding. The arithmetic encoder 12011 performs the same or similar operations and / or methods as those of the arithmetic encoder 40012.
[0240] According to an embodiment, the transmission processor 12012 can transmit each bitstream including the encoded geometry and / or the encoded attributes and / or the metadata (or metadata information), or transmit a single bitstream configured with the encoded geometry and / or the encoded attributes and / or the metadata. When the encoded geometry and / or the encoded attributes and / or the metadata according to an embodiment are configured as a single bitstream, the bitstream can include one or more sub-bitstreams. The bitstream according to an embodiment can include signaling information, which includes a sequence parameter set (SPS) for sequence-level signaling, a geometry parameter set (GPS) for signaling for geometry information encoding, an attribute parameter set (APS) for signaling for attribute information encoding, and a tile parameter set (TPS or tile list) and slice data for tile-level signaling. The slice data can include information about one or more slices. A slice according to an embodiment can include a geometry bitstream Geom0 0 and one or more attribute bitstreams Attr0 0 and Attr1 0 .
[0241] A slice is a series of syntax elements representing all or part of an encoded point cloud frame.
[0242] According to an embodiment, the TPS can include information about each tile of one or more tiles (e.g., height / size information and coordinate information about a bounding box). The geometry bitstream can include a header and a payload. The header of the geometry bitstream according to an embodiment can include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, as well as information about the data included in the payload. As described above, according to an embodiment, the metadata processor 12007 can generate and / or process the signaling information and send it to the transmission processor 12012. According to an embodiment, the elements for performing geometry encoding and the elements for performing attribute encoding can share data / information with each other, as indicated by the dashed line. The transmission processor 12012 according to an embodiment can perform the same or similar operations and / or transmission methods as those of the transmitter 10003. The details are the same as those described with reference to Figure 1 and Figure 2 and are thus omitted from the description.
[0243] Figure 13 An example of a receiving device according to an embodiment is illustrated.
[0244] Figure 13 The receiving device illustrated in Figure 1 is an example of the receiving device 10004 (or Figure 10 and Figure 11 the point cloud video decoder). Figure 13 The receiving device illustrated in Figures 1 to 11 can perform one or more of the operations and methods that are the same as or similar to the operations and methods of the point cloud video decoder described with reference to
[0245] The receiving device according to an embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processor 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a LOD / lifting / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each element for decoding according to an embodiment can perform an inverse process of the operation of the corresponding element for encoding according to an embodiment.
[0246] The receiver 13000 according to an embodiment receives point cloud data. The receiver 13000 can perform operations and / or a receiving method that are the same as or similar to the operations and / or the receiving method of the receiver 10005 of Figure 1 . A detailed description thereof is omitted.
[0247] The receiving processor 13001 according to an embodiment can obtain a geometry bitstream and / or an attribute bitstream from the received data. The receiving processor 13001 can be included in the receiver 13000.
[0248] The arithmetic decoder 13002, the occupancy code-based octree reconstruction processor 13003, the surface model processor 13004, and the inverse quantization processor 13005 can perform geometry decoding. The geometry decoding according to an embodiment is the same as or similar to the geometry decoding described with reference to Figures 1 to 10 . A detailed description thereof is omitted.
[0249] The arithmetic decoder 13002 according to an embodiment can decode the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.
[0250] The occupancy code - based octree reconstruction processor 13003 according to an embodiment may reconstruct an octree by obtaining an occupancy code from a decoded geometry bitstream (or information on a geometry structure protected as a decoding result). The occupancy code - based octree reconstruction processor 13003 performs operations and / or methods that are the same as or similar to those of the octree synthesizer 11001 and / or the octree generation method. When applying trisoup geometry coding, the surface model processor 13004 according to an embodiment may perform trisoup geometry decoding and related geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on a surface model method. The surface model processor 13004 performs operations that are the same as or similar to those of the surface approximation synthesizer 11002 and / or the geometry reconstruction unit 11003.
[0251] The inverse quantization processor 13005 according to an embodiment may perform inverse quantization on the decoded geometry structure.
[0252] The metadata parser 13006 according to an embodiment may parse metadata (e.g., set values) included in the received point cloud data. The metadata parser 13006 may pass the metadata for geometry decoding and / or attribute decoding. The metadata is the same as the metadata described in the reference Figure 12 and thus a detailed description thereof is omitted.
[0253] The arithmetic decoder 13007, the inverse quantization processor 13008, the LOD / RAHT inverse transform processor 13009, and the color inverse transform processor 13010 perform attribute decoding. The attribute decoding is the same as or similar to the attribute decoding described in the reference Figures 1 to 10 and thus a detailed description thereof is omitted.
[0254] The arithmetic decoder 13007 according to an embodiment may decode an attribute bitstream by arithmetic coding. The arithmetic decoder 13007 may decode the attribute bitstream based on the reconstructed geometry structure. The arithmetic decoder 13007 performs operations and / or coding that are the same as or similar to those of the arithmetic decoder 11005.
[0255] The inverse quantization processor 13008 according to an embodiment may perform inverse quantization on the decoded attribute bitstream. The inverse quantization processor 13008 performs operations and / or methods that are the same as or similar to those of the inverse quantization unit 11006 and / or the inverse quantization method.
[0256] The LOD / lifting / RAHT inverse transform processor 13009 according to an embodiment may process the reconstructed geometry and the attributes after inverse quantization. The prediction / lifting / RAHT inverse transform processor 1301 performs one or more of the operations and / or decoding that are the same as or similar to the operations and / or decoding of the RAHT transformer 11007, the LOD generation unit 11008, and / or the inverse lifter 11009. The color inverse transform processor 13010 according to an embodiment performs inverse transform coding to inverse-transform the color values (or textures) included in the decoded attributes. The color inverse transform processor 13010 performs the operations and / or inverse transform coding that are the same as or similar to the operations and / or inverse transform coding of the inverse color transform unit 11010. The renderer 13011 according to an embodiment may render point cloud data.
[0257] Figure 14 An exemplary structure operably connectable to a method / apparatus for transmitting and receiving point cloud data according to an embodiment is shown.
[0258] Figure 14 The structure represents a configuration in which at least one of the server 17600, the robot 17100, the autonomous vehicle 17200, the XR device 17300, the smart phone 17400, the home appliance 17500, and / or the head-mounted display (HMD) 17700 is connected to the cloud network 17000. The robot 17100, the autonomous vehicle 17200, the XR device 17300, the smart phone 17400, or the home appliance 17500 is referred to as a device. Additionally, the XR device 17300 may correspond to a point cloud compression data (PCC) device according to an embodiment or may be operably connected to a PCC device.
[0259] The cloud network 17000 may represent a network that forms part of or exists in a cloud computing infrastructure. Here, the cloud network 17000 may be configured using a 3G network, a 4G or Long-Term Evolution (LTE) network, or a 5G network.
[0260] The server 17600 may be connected to at least one of the robot 17100, the autonomous vehicle 17200, the XR device 17300, the smart phone 17400, the home appliance 17500, and / or the HMD 17700 via the cloud network 17000 and may assist with at least a part of the processing of the connected devices 17100 to 17700.
[0261] The HMD 17700 represents one of the implementation types of the XR device and / or the PCC device according to an embodiment. The HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power unit.
[0262] In the following, various embodiments of apparatuses 17100 to 17500 to which the above technologies are applied will be described. According to the above embodiments, Figure 14 the apparatuses 17100 to 17500 illustrated in [the above] can be operably connected / coupled to a point cloud data sending device and a receiver.
[0263] <PCC+XR>
[0264] An XR / PCC apparatus can adopt PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a household appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.
[0265] The XR / PCC apparatus can analyze 3D point cloud data or image data obtained through various sensors or from an external device, and generate position data and attribute data regarding 3D points. Thereby, the XR / PCC apparatus can obtain information about the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC apparatus can match an XR object including auxiliary information about the recognized object with the recognized object, and output the matched XR object.
[0266] <PCC+Autopilot+XR>
[0267] An autonomous driving vehicle 17200 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0268] An autonomous driving vehicle 17200 applying XR / PCC technology can refer to an autonomous driving vehicle provided with a device for providing an XR image or an autonomous driving vehicle that is a control / interaction target in an XR image. Specifically, the autonomous driving vehicle 17200 that is a control / interaction target in an XR image can be distinguished from an XR device 17300, and can be operably connected to the XR device 17300.
[0269] An autonomous driving vehicle 17200 having a device for providing an XR / PCC image can obtain sensor information from a sensor including a camera, and output the generated XR / PCC image based on the obtained sensor information. For example, the autonomous driving vehicle 17200 can have an HUD and output the XR / PCC image thereto, thereby providing an XR / PCC object corresponding to a real object or an object existing on a screen to an occupant.
[0270] When an XR / PCC object is output to the HUD, at least a part of the XR / PCC object can be output to overlap with a real object that the occupant's eyes are directed at. On the other hand, when an XR / PCC object is output to a display provided inside an autonomous vehicle, at least a part of the XR / PCC object can be output to overlap with an object on the screen. For example, the autonomous vehicle 17200 can output an XR / PCC object corresponding to an object such as a road, another vehicle, a traffic signal, a traffic sign, a two-wheeler, a pedestrian, and a building.
[0271] According to an embodiment, virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology are applicable to various devices.
[0272] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to a technology that shows virtual-created CG images on an image of a real object. The MR technology is similar to the above AR technology in that the virtual object to be shown is mixed and combined with the real world. However, the MR technology is different from the AR technology in that the AR technology clearly distinguishes a real object from a virtual object created as a CG image and uses the virtual object as a supplementary object for the real object, while the MR technology regards the virtual object as an object having the same characteristics as the real object. More specifically, an example of the application of the MR technology is a hologram service.
[0273] Recently, VR, AR, and MR technologies are sometimes referred to as extended reality (XR) technologies without being clearly distinguished from each other. Therefore, the embodiments of the present disclosure are applicable to any one of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies is applicable to such technologies.
[0274] The PCC method / apparatus according to an embodiment can be applied to a vehicle that provides an autonomous driving service.
[0275] A vehicle that provides an autonomous driving service is connected to a PCC device for wired / wireless communication.
[0276] When the point cloud compression data (PCC) transmission / reception device according to an embodiment is connected to a vehicle for wired / wireless communication, the device may receive / process content data related to AR / VR / PCC services that can be provided together with an autonomous driving service and transmit it to the vehicle. When the PCC transmission / reception device is installed on the vehicle, the PCC transmission / reception device may receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device and provide it to the user. A vehicle or a user interface device according to an embodiment may receive a user input signal. The user input signal according to an embodiment may include a signal indicating an autonomous driving service.
[0277] In addition, the point cloud video encoder of the sender may also perform a spatial segmentation process of spatially dividing (or partitioning) the point cloud data into one or more 3D blocks before encoding the point cloud data. That is, in order to perform and process the encoding and transmission operations of the transmission device and the decoding and rendering operations of the reception device in real time with low latency, the transmission device may spatially divide the point cloud data into multiple regions. Additionally, the transmission device may encode the spatially segmented regions (or blocks) independently or non-independently, thereby achieving random access and parallel encoding in the three-dimensional space occupied by the point cloud data. Further, the transmission device and the reception device may independently or non-independently perform encoding and decoding for each spatially segmented region (or block), thereby preventing the accumulation of errors in the encoding process and the decoding process.
[0278] Figure 15 It is a diagram illustrating another example of a point cloud transmission device according to an embodiment including a spatial divider.
[0279] A point cloud transmission device according to an embodiment may include a data input unit 51001, a coordinate transformation unit 51002, a quantization processor 51003, a spatial divider 51004, a signaling processor 51005, a geometry encoder 51006, an attribute encoder 51007, and a transmission processor 51008. According to an embodiment, the coordinate transformation unit 51002, the quantization processor 51003, the spatial divider 51004, the geometry encoder 51006, and the attribute encoder 51007 may be referred to as a point cloud video encoder.
[0280] The data input unit 51001 may perform Figure 1 some or all of the operations of the point cloud video acquisition unit 10001, or may perform Figure 12 some or all of the operations of the data input unit 12000. The coordinate transformation unit 51002 may perform Figure 4 some or all of the operations of the coordinate transformation unit 40000. Additionally, the quantization processor 51003 may perform Figure 4Some or all operations of the quantization unit 40001, or may be performed Figure 12 Some or all operations of the quantization processor 12001.
[0281] The spatial splitter 51004 may split the space of the point cloud data that has been quantized and output from the quantization processor 51003 into one or more 3D blocks based on the bounding box and / or sub-bounding boxes. Here, the 3D block may refer to a tile group, a tile, a slice, a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In one embodiment, the signaling information for spatial splitting is entropy encoded by the signaling processor 51005 and then sent in the form of a bitstream through the transmission processor 51008.
[0282] Figure 16 of (a) to Figure 16 of (c) illustrate an embodiment of splitting the bounding box into one or more tiles. As Figure 16 shown in (a), the point cloud object corresponding to the point cloud data may be represented in the form of a coordinate system-based box, which is called a bounding box. In other words, the bounding box represents a cube that can contain all the points in the point cloud.
[0283] Figure 16 of (b) and Figure 16 of (c) illustrate an example in which Figure 16 the bounding box of (a) is split into tile 1# and tile 2#, and tile 2# is further split into slice 1# and slice 2#.
[0284] In one embodiment, the point cloud content may be one person such as an actor, multiple persons, one object, or multiple objects. On a larger scale, it may be a map for autonomous driving or a map for indoor navigation of a robot. In this case, the point cloud content may be a large amount of locally connected data. In this case, the point cloud content cannot be encoded / decoded all at once, so tile splitting may be performed before compressing the point cloud content. For example, room #101 in a building may be split into one tile, and room #102 in the building may be split into another tile. To support fast encoding / decoding by applying parallelization to the split tiles, the tiles may be further split (or divided) into slices. This operation may be referred to as slice splitting (or dividing).
[0285] That is, according to the embodiment, the tile may represent a partial region (e.g., a rectangular cube) of the 3D space occupied by the point cloud data. According to the embodiment, the tile may include one or more slices. The tile according to the embodiment may be split into one or more slices, so that the point cloud video encoder can encode the point cloud data in parallel.
[0286] A slice may represent a unit of data (or bitstream) that can be independently encoded by a point cloud video encoder according to an embodiment and / or a unit of data (or bitstream) that can be independently decoded by a point cloud video decoder. A slice may be a collection of data in the 3D space occupied by the point cloud data or a collection of some of the point cloud data. A slice according to an embodiment may represent a collection or region of points included in a tile according to an embodiment. According to an embodiment, a tile may be divided into one or more slices based on the number of points included in one tile. For example, one tile may be a collection of points divided by the number of points. According to an embodiment, a tile may be divided into one or more slices based on the number of points, and some data may be split or merged in the division process. That is, a slice may be a unit that can be independently encoded within the corresponding tile. In this way, the tiles obtained by spatial division are divided into one or more slices for fast and efficient processing.
[0287] A point cloud video encoder according to an embodiment may encode point cloud data slice by slice or tile by tile, where a tile includes one or more slices. Additionally, a point cloud video encoder according to an embodiment may perform different quantization and / or transformation on each tile or each slice.
[0288] The positions of one or more 3D blocks (e.g., slices) spatially divided by the spatial divider 51004 are output to the geometry encoder 51006, and the attribute information (or attributes) is output to the attribute encoder 51007. The position may be position information about the points included in the division unit (box, block, tile, tile group, or slice) and is referred to as geometric information.
[0289] The geometry encoder 51006 constructs and encodes (i.e., compresses) an octree based on the positions output from the spatial divider 51004 to output a geometry bitstream. The geometry encoder 51006 may reconstruct and / or approximate the octree and output it to the attribute encoder 51007. The reconstructed octree may be referred to as the reconstructed geometry (restored geometry).
[0290] The attribute encoder 51007 encodes (i.e., compresses) the attributes output from the spatial divider 51004 based on the reconstructed geometry output from the geometry encoder 51006 and outputs an attribute bitstream.
[0291] Figure 17 is a detailed block diagram illustrating another example of the geometry encoder 51006 and the attribute encoder 51007 according to an embodiment.
[0292] Figure 17The voxelization processor 53001, octree generator 53002, geometric information predictor 53003, and arithmetic encoder 53004 of the geometric encoder 51006 can perform Figure 4 Some or all of the operations of the octree analysis unit 40002, surface approximation analysis unit 40003, arithmetic encoder 40004, and geometric reconstruction unit 40005 of Figure 12 Some or all of the operations of the voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame encoding processor 12005, and arithmetic encoder 12006 of
[0293] Figure 17 The attribute encoder 51007 of
[0294] In an embodiment, a quantization processor may also be provided between the space divider 51004 and the voxelization processor 53001. The quantization processor quantizes the positions of one or more 3D blocks (e.g., slices) spatially divided by the space divider 51004. In this case, the quantization processor can perform Figure 4 Some or all of the operations of the quantization unit 40001 of Figure 12 Some or all of the operations of the quantization processor 12001 of Figure 15 The quantization processor 51003 of
[0295] The voxelization processor 53001 according to an embodiment performs voxelization based on the positions of one or more spatially divided 3D blocks (e.g., slices) or their quantized positions. Voxelization refers to the smallest unit for representing position information in 3D space. The points of the point cloud content (or 3D point cloud video) according to an embodiment can be included in one or more voxels. According to an embodiment, one voxel can include one or more points. In an embodiment, when quantization is performed before voxelization, multiple points can belong to one voxel.
[0296] In this specification, when two or more points are included in one voxel, these two or more points are referred to as duplicate points. That is, in geometric encoding processing, overlapping points can be generated through geometric quantization and voxelization.
[0297] The voxelization processor 53001 according to an embodiment may output duplicate points belonging to one voxel to the octree generator 53002 without merging them, or may merge multiple points into one point and output the merged point to the octree generator 53002.
[0298] The octree generator 53002 according to an embodiment generates an octree based on the voxels output from the voxelization processor 53001.
[0299] The geometry information predictor 53003 according to an embodiment predicts and compresses geometry information based on the octree generated by the octree generator 53002, and outputs the predicted and compressed information to the arithmetic encoder 53004. In addition, the geometry information predictor 53003 reconstructs the geometry based on the positions changed by compression, and outputs the reconstructed (or decoded) geometry to the LOD configuration unit 53007 of the attribute encoder 51007. The reconstruction of the geometry information may be performed in a device or component separate from the geometry information predictor 53003. In another embodiment, the reconstructed geometry may also be provided to the attribute transformation processor 53006 of the attribute encoder 51007.
[0300] The color transformation processor 53005 of the attribute encoder 51007 corresponds to Figure 4 the color transformation unit 40006 of Figure 12 or the color transformation processor 12008 of Figures 1 to 9 . The color transformation processor 53005 according to an embodiment performs color transformation encoding for transforming the color values (or textures) included in the attributes provided from the data input unit 51001 and / or the spatial divider 51004. For example, the color transformation processor 53005 may transform the format of the color information (e.g., from RGB to YCbCr). The operation of the color transformation processor 53005 according to an embodiment may be selectively applied according to the color values included in the attributes. In another embodiment, the color transformation processor 53005 may perform color transformation encoding based on the reconstructed geometry. For details of the geometry reconstruction, refer to the description of
[0301] The attribute transformation processor 53006 according to an embodiment may perform attribute transformation for transforming attributes based on the positions where geometry encoding has not been performed and / or the reconstructed geometry.
[0302] The attribute transformation processor 53006 may be referred to as a recoloring unit.
[0303] The operation of the attribute transformation processor 53006 according to an embodiment may be selectively applied according to whether duplicate points are merged. According to an embodiment, the merging of duplicate points may be performed by the voxelization processor 53001 or the octree generator 53002 of the geometry encoder 51006.
[0304] In this specification, when points belonging to a voxel are merged into one point in the voxelization processor 53001 or the octree generator 53002, the attribute transformation processor 53006 performs an attribute transformation. For example.
[0305] The attribute transformation processor 53006 performs operations and / or methods that are the same as or similar to the operations and / or methods of Figure 4 the attribute transformation unit 40007 of Figure 12 or the operations and / or methods of the attribute transformation processor 12009 of
[0306] According to an embodiment, the geometric information reconstructed by the geometric information predictor 53003 and the attribute information output from the attribute transformation processor 53006 are provided to the LOD configuration unit 53007 for attribute compression.
[0307] According to an embodiment, the attribute information output from the attribute transformation processor 53006 can be compressed based on the reconstructed geometric information by one or a combination of two or more of RAHT coding, LOD-based prediction transform coding, and lifting transform coding.
[0308] Hereinafter, as an embodiment, it is assumed that attribute compression is performed by one or a combination of LOD-based prediction transform coding and lifting transform coding. Therefore, the description of RAHT coding will be omitted. For details of RAHT transform coding, refer to the description of Figures 1 to 9
[0309] The LOD configuration unit 53007 according to an embodiment generates a level of detail (LOD).
[0310] The LOD is a degree representing the level of detail of the point cloud content. As the LOD value decreases, it indicates that the level of detail of the point cloud content decreases. As the LOD value increases, it indicates that the level of detail of the point cloud content increases. The points of the reconstructed geometry (i.e., the reconstructed positions) can be classified according to the LOD.
[0311] In an embodiment, in prediction transform coding and lifting transform coding, points can be divided into LODs and grouped.
[0312] This operation can be referred to as LOD generation processing, and groups with different LODs can be referred to as LOD l sets. Here, l represents the LOD and is an integer starting from 0. LOD 0 is a set composed of points with the largest distance therebetween. As l increases, the distance between points belonging to LOD l decreases.
[0313] When the LOD value configuration unit 53007 generates LODl When aggregating, the neighboring point aggregation configuration unit 53008 according to the embodiment can be based on the LOD l Find X (>0) nearest neighbor (NN) points in the group with the same or lower LOD (i.e., large distance between nodes), and register them as the neighboring set in the predictor. X is the maximum number of points that can be set as neighboring points. X can be input as a user parameter (also known as an encoder option), or signaled in the signaling information by the signaling processor 51005 (e.g., the lifting_num_pred_nearest_neighbours field signaled in the APS).
[0314] Reference Figure 9 As an example, at LOD 0 and LOD 1 Find the neighboring points of P3 belonging to LOD 1 For example, when the maximum number of points (X) that can be set as neighboring points is 3, the three nearest neighbor nodes of P3 can be P2, P4, and P6. These three nodes are registered as the neighboring set in the predictor of P3. In the embodiment, among the registered nodes, the neighboring node P4 is the closest to P3 in terms of distance, then P6 is the next closest node, and P2 is the farthest among the three nodes. Here, X = 3 is only an embodiment configured to provide an understanding of the present disclosure. The value of X can be different.
[0315] As described above, all points of the point cloud data can each have a predictor.
[0316] The attribute information prediction unit 53009 according to the embodiment predicts the attribute value from the neighboring points registered in the predictor, and obtains the residual attribute value of the corresponding point based on the predicted attribute value. The residual attribute value is output to the residual attribute information quantization processor 53010.
[0317] Next, the LOD generation and neighboring point search will be described in detail.
[0318] As described above, for attribute compression, the attribute encoder performs the process of generating the LOD based on the reconstructed geometry points and searching for the nearest neighbor points of the points to be encoded based on the generated LOD. According to the embodiment, even the attribute decoder of the receiving device performs the process of generating the LOD and searching for the nearest neighbor points of the points to be decoded based on the generated LOD.
[0319] According to the embodiment, the LOD configuration unit 53007 can configure one or more LODs using one or more LOD generation methods (or LOD configuration methods).
[0320] According to an embodiment, the LOD generation method used in the LOD configuration unit 53007 may be input as a user parameter or may be signaled in the signaling information by the signaling processor 51005. For example, the LOD generation method may be signaled in the APS.
[0321] According to an embodiment, the LOD generation method may be classified into an octree-based LOD generation method, a distance-based LOD generation method, and a sampling-based LOD generation method.
[0322] Figure 18 is a diagram illustrating an example of generating an LOD based on an octree according to an embodiment.
[0323] According to an embodiment, when generating an LOD based on an octree, each depth level of the octree may be matched with each LOD, as Figure 18 illustrated. That is, the octree-based LOD generation method is a method of generating an LOD using the following feature: the details of the point cloud data are gradually increased as the depth level in the octree structure increases (i.e., in the direction from the root to the leaf). According to an embodiment, the octree-based LOD configuration may be performed from the root node to the leaf node or, conversely, from the leaf node to the root node.
[0324] The distance-based LOD generation method according to an embodiment is a method of arranging points in Morton codes and generating an LOD based on the distance between the points.
[0325] The sampling-based LOD generation method according to an embodiment is a method of arranging points in Morton codes and classifying every k-th point as a candidate for a lower LOD that can be a set of neighboring points. That is, during sampling, every k-th point may be selected and classified as belonging to the LOD 0 to LOD l-1 lower candidate set, and the remaining points may be registered in the LOD l set. That is, the selected points become a candidate group, and the candidate group may be selected as a set of neighboring points for the points registered in the LOD l set. k may vary according to the point cloud content. Here, the points may be points of the captured point cloud or points of the reconstructed geometry.
[0326] In the present disclosure, the set of LOD l and the lower candidate set corresponding to LOD 0 to LOD l-1 will be referred to as a reserved set or an LOD reserved set. LOD 0 is a set composed of points with the largest distance therebetween. As l increases, the distance between the points belonging to LOD l decreases.
[0327] According to an embodiment, the LOD l points in the set are also arranged based on Morton code order. Additionally, the LOD 0 to the LOD l-1 points included in each of the sets are also arranged based on Morton code order.
[0328] Figure 19 is a diagram illustrating an example of arranging points of a point cloud in Morton code order according to an embodiment.
[0329] That is, a Morton code for each point is generated based on the x, y, and z position values of each point of the point cloud. When the Morton codes of the points of the point cloud are generated by this process, the points of the point cloud can be arranged in Morton code order. According to an embodiment, the points of the point cloud can be arranged in ascending order of the Morton code. The order of the points arranged in ascending order of the Morton code can be referred to as the Morton order.
[0330] The LOD configuration unit 53007 according to an embodiment can generate an LOD by sampling the points arranged in Morton order.
[0331] The LOD configuration unit 53007 generates an LOD by applying at least one of an octree-based LOD generation method, a distance-based LOD generation method, or a sampling-based LOD generation method l set, the neighboring point set configuration unit 53008 can be based on the LOD l set to search for X (>0) nearest neighbor (NN) points in a group with the same or lower LOD (i.e., a large distance between nodes), and register the X NN points as a neighboring point set in the predictor.
[0332] In this case, since it takes a lot of time to search all points to configure the neighboring point set, an embodiment of searching for neighboring points within a neighboring point search range that includes only a part of the points is provided. The search range refers to the number of points and can be 128, 256, or another value. According to an embodiment, information about the search range can be set in the neighboring point set configuration unit 53008, or can be input as a user parameter. Alternatively, information about the search range can be signaled in signaling information by the signaling processor 51005 (e.g., the lifting_search_range field signaled in the APS).
[0333] Figure 20 is a diagram illustrating an example of searching for neighboring points in a search range based on the LOD according to an embodiment. The arrows illustrated in the figure indicate the Morton order according to an embodiment.
[0334] In Figure 20 , in an embodiment, the index list includes the LOD set to which the point to be encoded belongs (i.e., the LODl ) And the retention list includes at least one lower LOD set based on the LOD l set (e.g., LOD 0 to LOD l-1 ).
[0335] According to an embodiment, the points of the index list are arranged in ascending order based on the size of the Morton code, and the points of the retention list are also arranged in ascending order based on the size of the Morton code. Therefore, the foremost points among the points arranged in Morton order in the index list and the retention list have the smallest Morton code size.
[0336] The neighboring point set configuration unit 53008 according to an embodiment may search among the points belonging to the LOD 0 to LOD l-1 set and / or among the points belonging to the LOD l set for the point that is sequentially located before the point Px (i.e., the point having a Morton code less than or equal to the Morton code of the point Px) and having the Morton code closest to the Morton code of the point Px, so as to generate a neighboring point set of the point Px (i.e., the point to be encoded or the current point) belonging to the LOD l set. In the present disclosure, the point to be searched will be referred to as Pi or the center point.
[0337] According to an embodiment, when searching for the center point Pi, the neighboring point set configuration unit 53008 may search for the point Pi having the Morton code closest to the Morton code of the point Px among all the points located before the point Px, or search for the point Pi having the Morton code closest to the Morton code of the point Px among the points within the search range. In the present disclosure, the search range may be configured by the neighboring point set configuration unit 53008, or may be input as a user parameter. Additionally, information about the search range may be signaled in the signaling information by the signaling processor 51005. The present disclosure provides embodiments for searching for the center point Pi in the retention list when the number of LODs is multiple and for searching for the center point Pi in the index list when the number of LODs is 1. For example, when the number of LODs is 1, the search range may be determined based on the position of the current point in the list arranged by Morton code.
[0338] The neighboring point set configuration unit 53008 according to an embodiment compares the point Px with the points located before the searched (or selected) center point Pi (i.e., to the left of the center point in Figure 20 ), and after the searched (or selected) center point Pi (i.e., to the right of the center point in Figure 20The distance values between points belonging to the neighboring point search range (to the right of the center point in []) are provided. The neighboring point set configuration unit 53008 may register X (e.g., 3) NN points as the neighboring point set. In an embodiment, the neighboring point search range is the number of points. The neighboring point search range according to an embodiment may include one or more points located before (i.e., in front of) and / or after (i.e., behind) the center point Pi in Morton order. X is the maximum number of points that can be registered as neighboring points.
[0339] According to an embodiment, information about the neighboring point search range and information about the maximum number X of points that can be registered as neighboring points may be configured by the neighboring point set configuration unit 53008, or may be input as user parameters, or may be signaled in signaling information by the signaling processor 51005 (e.g., the lifting_search_range field and the lifting_num_pred_nearest_neighbours field signaled in APS). According to an embodiment, the actual search range may be a value obtained by multiplying the value of the lifting_search_range field by 2 and then adding the resulting value to the center point (i.e., (value of the lifting_search_range field × 2) + center point), a value obtained by adding the center value to the value of the lifting_search_range field (i.e., value of the lifting_search_range field + center point), or the value of the lifting_search_range field. The present disclosure provides embodiments for searching for neighboring points of point Px in the reservation list when the number of LODs is multiple and for searching for neighboring points of point Px in the index list when the number of LODs is 1.
[0340] For example, if the number of LODs is 2 or more and the information about the search range (e.g., the lifting_search_range field) is 128, the actual search range includes the center point Pi in the reservation list arranged in Morton code, 128 points before the center point Pi, and 128 points after the center point Pi. As another example, if the number of LODs is 1 and the information about the search range is 128, the actual search range includes the center point Pi in the index list arranged in Morton code and 128 points before the center point Pi. As another example, if the number of LODs is 1 and the information about the search range is 128, the actual search range includes 128 points before the current point Px in the index list arranged in Morton code.
[0341] According to an embodiment, the neighboring point set configuration unit 53008 may compare distance values between points within an actual search range and a point Px, and register X NN points as a neighboring point set of the point Px.
[0342] On the other hand, depending on the object from which the content is captured, the capture scheme, or the device used to capture the content, the attribute characteristics of the content may appear different.
[0343] Figure 21 of (a) and Figure 21 (b) of FIG. is a diagram illustrating an example of point cloud content according to an embodiment.
[0344] Figure 21 (a) of FIG. is an example of dense point cloud data obtained by capturing an object with a 3D scanner, and the correlation between the attributes of the object may be high. However, as shown in (a) of FIG., even when the point cloud data consists of one object, the density of the points in the point cloud data can affect the similarity of the prediction performed based on neighboring points. Figure 21 (a) of FIG. shows that even when the point cloud data consists of one object, the density of the points in the point cloud data can affect the similarity of the prediction performed based on neighboring points.
[0345] Figure 21 (b) of FIG. illustrates an example of sparse point cloud data obtained by capturing a wide area with a LiDAR device, and the attribute correlation between points may be low. An example of dense point cloud data may be a still image, and an example of sparse point cloud data may be a drone or an autonomous vehicle. As shown in (b) of FIG., when the point cloud data consists of multiple objects, especially when a large area is captured with real-time LiDAR, the density of the points in the captured point cloud data is low. Accordingly, the similarity of neighboring points may be low, which may reduce the compression efficiency. Figure 21 (b) of FIG. shows that when the point cloud data consists of multiple objects, especially when a large area is captured with real-time LiDAR, the density of the points in the captured point cloud data is low. Accordingly, the similarity of neighboring points may be low, which may reduce the compression efficiency.
[0346] In this way, depending on the object from which the content is captured, the capture scheme, or the device used to capture the content, the attribute characteristics of the content may appear different. If neighboring point search is performed to configure a neighboring point set when encoding attributes without considering different attribute characteristics, the distances between corresponding points and the registered neighboring points may not all be adjacent. For example, when X (e.g., 3) NN points are registered as a neighboring point set after comparing distance values between points belonging to the neighboring point search range before and after comparing a point Px with a center point Pi, at least one of the X registered neighboring points may be far from the point Px. In other words, at least one of the X registered neighboring points may not be an actual neighboring point of the point Px.
[0347] This phenomenon is more likely to occur in sparse point cloud data than in dense point cloud data. For example, as shown in FIG. Figure 21As exemplified in (a) of, if there is an object in the content that has color continuity / similarity and is densely captured, and the captured area is small, then even when the points are somewhat far apart from neighboring points, there may be content with high attribute correlation between the points. Conversely, as Figure 21 exemplified in (b) of, if a relatively wide area is sparsely captured by a LiDAR device, then content with color correlation may exist only when the distance to neighboring points is within a specific range. Therefore, when encoding the attributes of sparse point cloud data as exemplified in (b) of Figure 21 , some neighboring points registered through neighboring point search may not be the actual neighboring points of the point to be encoded.
[0348] When generating LOD, the probability of this phenomenon occurring is high when the points are arranged in Morton order. That is, although Morton order quickly arranges neighboring points, due to the zigzag scanning feature of the Morton code, jump segments (i.e., Figure 19 the discontinuous long lines in ) occur, and then the position of the points changes greatly. In this case, some neighboring points registered through neighboring point search within the search range may not be the actual NN points of the point to be encoded.
[0349] Figure 22 (a) of and Figure 22 (b) of are diagrams exemplifying examples of the average distance, minimum distance, and maximum distance of points belonging to a neighboring point set according to an embodiment.
[0350] Figure 22 (a) of exemplifies an example of the 100th frame of content called FORD, and Figure 22 (b) of exemplifies an example of the first frame of content called QNX.
[0351] Figure 22 (a) of and Figure 22 (b) of are exemplary embodiments that help those skilled in the art understand. Figure 22 (a) of and Figure 22 (b) of may be different frames of the same content or specific frames of different content.
[0352] For ease of explanation, Figure 22 the frame exemplified in (a) of is called the first frame, and Figure 22 the frame exemplified in (b) of is called the second frame.
[0353] In Figure 22In the first frame of (a), the total number of points is 80265. Among them, only one NN point is registered in 4 points, 2 NN points are registered in 60 points, and 3 NN points are registered in 80201 points. That is, the maximum number of points that can be registered as NN points is 3, but it can be 1 or 2 depending on the position of the points in the frame.
[0354] In Figure 22 In the second frame of (b), the total number of points is 31279. Among them, only one NN point is registered in 3496 points, 2 NN points are registered in 219 points, and 3 NN points are registered in 27564 points.
[0355] In Figure 22 In the first frame of (a) and Figure 22 In the second frame of (b), NN1 is an NN point, NN2 is the next NN point, and NN3 is the third NN point. Figure 22 In the first frame of (a), the maximum distance of NN2 is 1407616 and the maximum distance of NN3 is 1892992, while Figure 22 In the second frame of (b), the maximum distance of NN2 is 116007040 and the maximum distance of NN3 is 310887696.
[0356] That is, it can be understood that there are quite large differences in distance according to the content or frame. Therefore, some of the registered neighboring points may not be actual neighboring points.
[0357] Thus, if the distance difference between the point to be encoded and the registered neighboring points becomes large, the attribute difference between these two points may be large. In addition, when prediction is performed based on the neighboring point set including these neighboring points and a residual attribute value is obtained, the residual attribute value may increase, resulting in an increase in the bitstream size.
[0358] In other words, the characteristics of the content or frame may affect the configuration of the neighboring point set used for attribute prediction, and this may affect the compression efficiency of the attributes.
[0359] Therefore, in the present disclosure, when configuring the neighboring point set for encoding the attributes of the point cloud content, the neighboring points are selected by considering the attribute correlation between the points of the content so that meaningless points are not selected as the neighboring point set. Then, since the residual attribute value decreases and thus the bitstream size decreases, the compression efficiency of the attributes is improved. In other words, the present disclosure uses a method of considering the attribute correlation between the points of the content to limit the points that can be selected as the neighboring point set to improve the compression efficiency of the attributes.
[0360] The present disclosure provides embodiments of configuring a set of neighboring points in the attribute encoding process of point cloud content by applying a maximum distance of neighboring points (also referred to as a maximum NN distance) to consider the attribute correlation between points of the content.
[0361] The present disclosure provides embodiments of configuring a set of neighboring points by applying a search range and / or a maximum NN distance when encoding the attributes of point cloud content.
[0362] According to an embodiment, the distance between points for selecting the nearest neighbors is less than or equal to a maximum neighboring point distance (also referred to as a maximum nearest neighbor distance).
[0363] For example, as a result of comparing the distance values between a comparison point Px and points belonging to the neighboring point search range before and after a central point Pi, points among X (e.g., 3) NN points that are farther than the maximum NN distance are not registered in the set of neighboring points of the point Px. In other words, as a result of comparing the distance values between a comparison point Px and points belonging to the neighboring point search range before and after a central point Pi, only points within the maximum NN distance among X (e.g., 3) NN points are registered in the set of neighboring points of the point Px. For example, if one of the three points with the closest distance is farther than the maximum NN distance, the farther point can be excluded, and only the remaining two points can be registered in the set of neighboring points of the point Px.
[0364] The neighboring point set configuration unit 53008 according to an embodiment can apply the maximum NN distance according to the attribute characteristics of the content to achieve the best compression efficiency of the attributes regardless of the attribute characteristics of the content. Here, the maximum NN distance can be calculated differently according to the LOD generation method (octree-based, distance-based, or sampling-based LOD generation method). That is, the maximum NN distance can be applied differently for each LOD.
[0365] In the present disclosure, the maximum NN distance can be used interchangeably with the maximum distance of neighboring points or the maximum neighboring point distance. In other words, in the present disclosure, the maximum NN distance, the maximum distance of neighboring points, and the maximum neighboring point distance have the same meaning.
[0366] In the present disclosure, the maximum NN distance can be obtained by multiplying a basic neighboring point distance (referred to as a basic distance or a reference distance) by NN_range, as exemplified in Equation 5 below.
[0367] [Equation 5]
[0368] Maximum neighboring point distance = basic neighboring point distance * NN_range
[0369] In Equation 5, NN_range is the range within which neighboring points can be selected, and can be referred to as the maximum range within which neighboring points can be selected, the maximum neighboring point range, the neighboring point range, or the nearest neighbor range.
[0370] According to an embodiment, the neighboring point set configuration unit 53008 can automatically or manually set the NN_range according to the characteristics of the content, or set the NN_range through an input as a user parameter (also referred to as an encoder option). Additionally, information related to the NN_range can be signaled in the signaling information by the signaling processor 51005. The signaling information including information related to the NN_range can be at least one of SPS, APS, a set of tile parameters, or an attribute slice header.
[0371] When calculating the NN_range automatically, the attribute decoder of the receiving device can also automatically calculate the NN_range and apply it to the neighboring point search.
[0372] When information related to the NN_range is signaled in the signaling information, the attribute decoder of the receiving device can calculate the maximum neighboring point distance based on the signaling information. In this case, the neighboring points used in the attribute prediction in the attribute coding can be reconstructed and applied to the attribute decoding.
[0373] According to an embodiment, the neighboring point set configuration unit 53008 can calculate / configure the basic neighboring point distance by combining one or more of an octree-based method, a distance-based method, a sampling-based method, a mean-difference-based method for Morton codes for each LOD, and a mean-distance-difference-based scheme for each LOD. According to an embodiment, the method for calculating / configuring the basic neighboring point distance and / or the basic neighboring point distance information can be signaled in the signaling information (e.g., at least one of SPS, APS, TPS, or an attribute slice header) by the signaling processor 51005.
[0374] Next, embodiments of obtaining the basic neighboring point distance and the maximum neighboring point distance when generating LOD based on an octree will be described.
[0375] As Figure 18 illustrated, when the LOD configuration unit 53007 generates an octree-based LOD, each depth level of the octree can be matched with each LOD.
[0376] According to an embodiment, the neighboring point set configuration unit 53008 can obtain the maximum neighboring point distance based on the octree-based LOD.
[0377] According to an embodiment, when the LOD configuration unit 53007 generates an octree-based LOD, spatial scalability can be supported. With spatial scalability, when the source point cloud is dense, a lower-resolution point cloud can be accessed, like a thumbnail with lower decoder complexity and / or smaller bandwidth. For example, the spatial scalability function of geometry can be provided by a process of encoding or decoding only the occupied bits up to a selected depth level by adjusting the depth level of the octree during the encoding / decoding of geometry. Additionally, even during the encoding / decoding of attributes, the spatial scalability function of attributes can be provided by generating an LOD from a selected depth level of the octree and configuring the points for which the encoding / decoding of the attributes will be performed.
[0378] According to an embodiment, for spatial scalability, the LOD configuration unit 53007 can generate an octree-based LOD, and the neighboring point set configuration unit 53008 can perform a neighboring point search based on the generated octree-based LOD.
[0379] When the neighboring point set configuration unit 53008 according to an embodiment searches for neighboring points in a previous LOD based on points belonging to the currently octree-based generated LOD, the maximum neighboring point distance can be obtained by multiplying a basic neighboring point distance (referred to as the basic distance or reference distance) by NN_range (i.e., basic neighboring point distance × NN_range).
[0380] According to an embodiment, the maximum distance (i.e., diagonal distance) of a node in a specific octree-based generated LOD can be defined as the basic distance (or basic neighboring point distance).
[0381] According to an embodiment, when searching for neighboring points in a previous LOD (e.g., LODs in the reserved list l to LOD 0 to LOD l-1 ) based on points belonging to the current LOD (e.g., the LODs in the index list), the diagonal distance of a higher node (parent node) of the octree node of the current LOD can be the basic distance for obtaining the maximum neighboring point distance.
[0382] According to an embodiment, the basic distance, which is the diagonal distance (i.e., the maximum distance) of a node at a specific LOD, can be obtained as shown in Equation 6 below.
[0383] [Equation 6]
[0384]
[0385] According to another embodiment, the basic distance, which is the diagonal distance (i.e., the maximum distance) of a node at a specific LOD, can be obtained based on L2 as shown in Equation 7 below.
[0386] [Formula 7]
[0387]
[0388] Generally, when calculating the distance between two points, the Manhattan distance calculation method or the Euclidean distance calculation method is used. The Manhattan distance will be referred to as the L1 distance and the Euclidean distance will be referred to as the L2 distance. The Euclidean space can be defined using the Euclidean distance, and the norm corresponding to this distance will be referred to as the Euclidean norm or the L2 norm. A norm is a method (function) for measuring the length or magnitude of a vector.
[0389] Figure 23 Example (a) of [object] shows obtaining the diagonal distance of an octree node of LOD 0 According to an embodiment, when Figure 23 the example of (a) of [object] is applied to Formula 6, the basic distance becomes and when Figure 23 the example of (a) of [object] is applied to Formula 7, the basic distance becomes 3.
[0390] Figure 23 Example (b) of [object] shows calculating the diagonal distance of an octree node of LOD 1 According to an embodiment, when Figure 23 the example of (b) of [object] is applied to Formula 6, the basic distance becomes and when Figure 23 the example of (b) of [object] is applied to Formula 7, the basic distance becomes 12.
[0391] Figure 24 Examples (a) of [object] and Figure 24 (b) of [object] show examples of the maximum range (NN_range) of neighboring points that can be selected at each LOD. More specifically, Figure 24 example (a) of [object] shows an example when NN_range is 1 at LOD 0 and Figure 24 example (b) of [object] shows an example when NN_range is 3 at LOD 0 According to an embodiment, NN_range can be set automatically or manually according to the characteristics of the content, or can be set by being input as a user parameter. Information related to NN_range can be signaled in signaling information. The signaling information including information related to NN_range can be at least one of SPS, APS, tile parameter set, or attribute slice header. According to an embodiment, the range NN_range that can be set as neighboring points can be set to any value (for example, a multiple of the basic distance) regardless of the octree node range.
[0392] According to an embodiment, when Figure 24 the example of (a) is applied to Equation 5 and Equation 6, the maximum neighboring point distance becomes and when Figure 24 the example of (b) is applied to Equation 5 and Equation 6, the maximum neighboring point distance becomes
[0393] According to an embodiment, when Figure 24 the example of (a) is applied to Equation 5 and Equation 7, the maximum neighboring point distance becomes 3 (= 3 * 1), and when Figure 24 the example of (b) is applied to Equation 5 and Equation 7, the maximum neighboring point distance becomes 9 (= 3 * 3).
[0394] Next, an embodiment of obtaining the basic neighboring point distance and the maximum neighboring point distance when generating distance-based LOD will be described.
[0395] According to an embodiment, when configuring distance-based LOD, the basic neighboring point distance can be set to the distance for configuring the points at the current LOD (i.e., LOD l ). That is, the basic distance for each LOD can be set to the distance (dist2 L ) applied to the LOD generation at the L level of the LOD.
[0396] Figure 25 Examples of the basic neighboring point distance belonging to each LOD according to an embodiment are illustrated.
[0397] Therefore, when configuring distance-based LOD, the maximum neighboring point distance can be obtained by multiplying the distance-based basic neighboring point distance dist2 L by NN_range.
[0398] According to an embodiment, NN_range can be automatically or manually set by the neighboring point set configuration unit 53008 according to the characteristics of the content, or can be set by an input as a user parameter. Information about NN_range can be signaled in signaling information. The signaling information including information about NN_range can be at least one of SPS, APS, tile parameter set, or attribute slice header. According to an embodiment, NN_range that can be set as a neighboring point can be set to any value (e.g., a multiple of the basic distance). According to an embodiment, NN_Range can be adjusted according to the method of calculating the distance to neighboring points.
[0399] Next, an embodiment of obtaining the basic neighboring point distance and the maximum neighboring point distance when generating sampling-based LOD will be described.
[0400] According to an embodiment, when generating a sampling (i.e., based on selection)-based LOD, a base distance can be determined according to a sampling rate. According to an embodiment, in a sampling-based LOD configuration method, a Morton code is generated based on the position value of a point, points are arranged according to the Morton code, and then points not corresponding to the k-th point according to the arrangement order are registered in the current LOD. Thus, each LOD can have a different k and can be represented as k L 。
[0401] According to an embodiment, when generating an LOD by applying a sampling method to points arranged based on a Morton code, the average distance of consecutive k L points at each LOD can be set as the base neighboring point distance at each LOD. That is, the average distance of consecutive k L points at the current LOD arranged based on the Morton code can be set as the base neighboring point distance of the current LOD. According to another embodiment, when generating an LOD by applying a sampling method to points arranged based on a Morton code, the depth level of an octree can be estimated from the current LOD level, and the diagonal length of a node of the estimated octree depth level can be set as the base neighboring point distance. That is, the base neighboring point distance can be obtained by applying the above formula 6 or formula 7.
[0402] Therefore, when configuring a sampling-based LOD, the maximum neighboring point distance can be obtained by multiplying the sampling-based base neighboring point distance by NN_range.
[0403] According to an embodiment, NN_range can be automatically or manually set by a neighboring point set configuration unit 53008 according to the characteristics of the content, or can be set by an input as a user parameter. Information about NN_range can be signaled in signaling information. The signaling information including information about NN_range can be at least one of SPS, APS, a tile parameter set, or an attribute slice header. According to an embodiment, NN_range that can be set as a neighboring point can be set to any value (e.g., a multiple of the base distance). According to an embodiment, NN_Range can be adjusted according to the method of calculating the distance to a neighboring point.
[0404] In the present disclosure, a method other than the above method of obtaining the base neighboring point distance can be used to obtain the base neighboring point distance.
[0405] For example, when configuring an LOD, the base neighboring point distance can be configured by calculating the average difference of the Morton codes of sampling points. Additionally, the maximum neighboring point distance can be obtained by multiplying the configured base neighboring point distance by NN_range.
[0406] As another example, whether using a distance-based method or a sampling-based method, the basic proximity point distance can be configured by calculating the average distance difference of the LOD of the current configuration. Additionally, the maximum proximity point distance can be obtained by multiplying the configured basic proximity point distance by NN_range.
[0407] As another example, the basic proximity point distance for each LOD can be received as a user parameter and then applied when obtaining the maximum proximity point distance. Thereafter, information about the received basic proximity point distance can be signaled in the signaling information.
[0408] As described above, the basic proximity point distance can be obtained by applying an optimal basic proximity point distance calculation method according to the LOD configuration method. In another embodiment, the method for calculating the basic proximity point distance can be applied by selecting any desired method and signaling the selected method in the signaling information so that the points can be restored in the decoder of the receiving device using this method. According to the embodiment, a method for calculating the maximum proximity point distance that is suitable for the attributes of the point cloud content and can maximize the compression performance can be used.
[0409] Next, an embodiment for automatically calculating the maximum proximity point range (NN_range) will be described.
[0410] According to the embodiment, the method for obtaining the maximum proximity point range can vary according to the LOD generation method.
[0411] According to the embodiment, the proximity point set configuration unit 53008 can automatically calculate the maximum proximity point range according to the density, or can automatically calculate the maximum proximity point range by inferring the density based on the length of the axis and the number of points.
[0412] Figure 26 is a flowchart illustrating automatically calculating the maximum proximity point range according to an embodiment of the present disclosure.
[0413] First, some of the scanned points are used to calculate the basic proximity point distance for each LOD (operation 55001). Here, according to the embodiment, it is assumed that the basic proximity point distance for each LOD is greater than 0 (>0). In the present disclosure, for simplicity, the basic proximity point distance for each LOD is referred to as dist2 0 .
[0414] According to the embodiment, the basic proximity point distance can be obtained differently according to the LOD generation method.
[0415] For example, when generating LOD based on an octree, the diagonal distance (i.e., the maximum distance) of a node in a specific LOD can be set to the base proximity point distance (e.g., see Equation 6 or Equation 7). According to an embodiment, when there are multiple LODs, the diagonal distance of the upper node (parent node) of the octree node of the current LOD can be set to the base proximity point distance.
[0416] For example, when generating LOD based on distance, the base proximity point distance can be set to the distance used to construct points at the level of the current LOD (i.e., LOD l ). That is, the base proximity point distance for each LOD can be set to the distance dist2 applied to LOD generation at level L of the LOD L (see Figure 25 ).
[0417] For example, when generating LOD by applying a sampling method to points arranged based on Morton codes, the average distance of every k L consecutive points in each LOD can be set to the base proximity point distance at each LOD. That is, the average distance of every k L consecutive points in the current LOD arranged based on Morton codes can be set to the base proximity point distance of the current LOD. As another example, when generating LOD by applying a sampling method to points arranged based on Morton codes, the depth level of the octree can be estimated from the current LOD level, and the diagonal length of the nodes at the estimated octree depth level can be set to the base proximity point distance. In this case, the base proximity point distance can be obtained by applying Equation 6 or Equation 7.
[0418] According to an embodiment, the number of points selected for calculating the base proximity point distance for each LOD can be one quarter, or other intervals can be considered for selection.
[0419] Once the base proximity point distance is calculated in operation 55001, the diagonal length of the bounding box of the points is calculated (operation 55002). In the present disclosure, for simplicity, the diagonal length of the bounding box will be referred to as BBox_diagonal_dist or BBox_diagonal_length.
[0420] According to an embodiment, the bounding box can be set by scanning the point cloud data. According to another embodiment, the bounding box can be inferred and calculated with reference to the octree depth for octree construction.
[0421] The density rate (or density_estimation) can be estimated based on the basic proximity point distance calculated in operation 55001 and the diagonal length of the bounding box calculated in operation 55002 (operation 55003).
[0422] In an embodiment, the density rate can be obtained by dividing the basic proximity point distance by the diagonal length of the bounding box (density_rate = dist2 0 / BBox_diagonal_dist). In this case, as the density rate increases, the density decreases.
[0423] In another embodiment, the density rate can be obtained by dividing the diagonal length of the bounding box by the basic proximity point distance (density_rate = BBox_diagonal_dist / dist2 0 ). In this case, when setting the maximum proximity point range according to the density rate, the Kn value and sign can change. In this case, as the density rate decreases, the density decreases. According to the embodiment, the Kn value can be a table value for setting the maximum proximity point range according to the calculated value obtained by automatically calculating the maximum proximity point range.
[0424] Once the density rate is obtained in operation 55003, the maximum proximity point range is set according to the density rate (operation 55004).
[0425] According to the embodiment, when the density rate is greater than Kn (where n >= 0 and Kn < 1), the value of n or n + 1 can be set as the maximum proximity point range (n: density_rate > Kn (n >= 0, Kn < 1)).
[0426] According to the embodiment, the Kn value can be a table value for setting the maximum proximity point range according to the density rate calculated when automatically calculating the maximum proximity point range.
[0427] According to the embodiment, the Kn value can be set to a default value according to the LOD generation method, or can be set differently according to user input. That is, the information related to the Kn value can be variably changed and signaled in the signaling information, or a fixed value can be preset by default in the attribute encoder / decoder.
[0428] For example, assume K1 = 0.000001 and the density rate is greater than 0.000001, then the maximum proximity point range (NN_range) is 1.
[0429] According to an embodiment, the maximum neighboring point range information automatically calculated according to the density rate can be signaled in signaling information. For example, the signaling information including the maximum neighboring point range information can be at least one of SPS, APS, a tile parameter set, or an attribute slice header.
[0430] Figure 27 is a flowchart illustrating automatically calculating a maximum neighboring point range based on density according to another embodiment of the present disclosure. In Figure 27 the embodiment, the maximum neighboring point range is calculated by inferring density based on the axis length and the number of points.
[0431] According to an embodiment, density can be inferred based on at least one of the minimum axis length of the bounding box, the maximum axis length of the bounding box, and the intermediate axis length of the bounding box.
[0432] Figure 27 An embodiment of obtaining the maximum neighboring point range based on the minimum axis length of the bounding box and the number of points is illustrated.
[0433] That is, the minimum axis length of the bounding box of the points is calculated (operation 57001). In the present disclosure, for simplicity, the minimum axis length of the bounding box is referred to as BBox_min_axis_length.
[0434] The density rate (density_rate or density_estimation) can be estimated based on the number of points and the minimum axis length of the bounding box calculated in operation 57001 (operation 57002).
[0435] In an embodiment, the density rate can be obtained by dividing the minimum axis length of the bounding box by the number of points (density_rate = BBox_min_axis_length / number of points). In this case, as the density rate increases, the density decreases.
[0436] Once the density rate is obtained in operation 57002, the maximum neighboring point range is set according to the density rate (operation 57003).
[0437] According to an embodiment, when the density rate is greater than Kn (where n >= 0 and Kn < 1), the value of n or n + 1 can be set as the maximum neighboring point range (n: density_rate > Kn (n >= 0, Kn < 1)).
[0438] According to an embodiment, the Kn value can be a table value for setting the maximum neighboring point range according to the density rate calculated when automatically calculating the maximum neighboring point range.
[0439] According to an embodiment, the Kn value can be set to a default value according to the LOD generation method, or can be set differently according to user input. That is, the information related to the Kn value can be variably changed and signaled in the signaling information, or a fixed value can be preset by default in the attribute encoder / decoder.
[0440] For example, assume K1 = 0.05 and the density rate is greater than 0.05, then the maximum nearest neighbor range (NN_range) is 1.
[0441] According to another embodiment, the density rate can be estimated based on the maximum axis length of the bounding box and the number of points, or can be estimated based on the intermediate axis length of the bounding box and the number of points. Then, the maximum nearest neighbor range is calculated based on the density rate. In this case, the method of calculating the maximum nearest neighbor range is the same as the method of calculating the maximum nearest neighbor range based on the minimum axis length of the bounding box described above.
[0442] According to an embodiment, a method of calculating the maximum nearest neighbor range that can match the attributes of the point cloud content and maximize the compression performance can be used.
[0443] According to an embodiment, any desired method of calculating the maximum nearest neighbor range can be selected and applied, or a combination of one or more methods can be applied. Alternatively, information about whether to perform automatic calculation can be signaled and sent in the signaling information. According to an embodiment, when the information about whether to perform automatic calculation indicates automatic calculation, the attribute decoder of the receiving device can automatically calculate the maximum nearest neighbor range.
[0444] According to an embodiment, information regarding a basic neighboring point distance calculation method (e.g., nn_base_distance_calculation_method_type), basic neighboring point distance information (e.g., nn_base_distance), maximum neighboring point range information (e.g., nearest_neighbore_max_range), minimum neighboring point range information (e.g., nearest_neighbor_min_range), information on how to apply the maximum neighboring point distance (e.g., nn_range_filtering_location_type), information on whether to automatically calculate the maximum neighboring point range (e.g., automatic_nn_range_calculation_flag), information on a method for calculating the maximum neighboring point range (e.g., automatic_nn_range_method_type), and information related to the Kn value of the automatic maximum neighboring point range (e.g., automatic_max_nn_range_in_table, automatic_nn_range_table_k) may be referred to as neighboring point selection related option information. The neighboring point selection related option information may be signaled and transmitted in signaling information. In an embodiment, the signaling information may be at least one of a sequence parameter set, an attribute parameter set, a tile parameter set, or an attribute slice header.
[0445] According to an embodiment, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may set a basic neighboring point distance calculation method (e.g., nn_base_distance_calculation_method_type) in the signaling information according to a LOD generation method (or a LOD configuration method). According to an embodiment, the basic neighboring point distance calculation method may include basic neighboring distance calculation based on an octree, basic neighboring point distance calculation based on distance, basic neighboring point distance calculation based on sampling, basic neighboring point distance calculation based on an average difference of Morton codes between LODs, basic neighboring point distance calculation based on an average distance difference between LODs, and basic neighboring point distance calculation according to an input basic neighboring point distance.
[0446] According to an embodiment, the basic neighboring point distance calculated according to the basic neighboring point distance calculation method may be applied to the calculation of the maximum neighboring point range. Additionally, the basic neighboring point distance calculation method may be signaled and transmitted in the signaling information. In this case, the attribute decoder may calculate the maximum neighboring point range based on the signaling information and apply it to the maximum neighboring point distance.
[0447] According to an embodiment, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may receive base neighboring point distance information (base_distance) as a user parameter (or an encoder option), and apply it to the calculation of the maximum neighboring point distance and / or the maximum neighboring point range. Additionally, when the base neighboring point distance information is signaled in the signaling information, the attribute decoder of the receiving device may calculate the maximum neighboring point distance and / or the maximum neighboring point range based on the signaling information.
[0448] According to an embodiment, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may signal the maximum neighboring point range information in the signaling information, and the attribute decoder of the receiving device may calculate the maximum neighboring point distance based on the signaling information. According to an embodiment, the maximum neighboring point range may be set differently for each LOD.
[0449] According to an embodiment, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may signal the minimum neighboring point range information in the signaling information, and the attribute decoder of the receiving device may calculate the maximum neighboring point distance based on the signaling information. According to an embodiment, the minimum neighboring point range may be set differently for each LOD.
[0450] According to an embodiment, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may receive information indicating whether to automatically calculate the maximum neighboring point range or the minimum neighboring point range or both the maximum neighboring point range and the minimum neighboring point range as a user parameter (or an encoder option), and signal in the signaling information the information indicating whether to perform the automatic calculation.
[0451] According to an embodiment, when the neighboring point set configuration unit 53008 and / or the signaling processor 51005 automatically calculates the maximum neighboring point range, it may receive the Kn value as a user parameter (or an encoder option), and signal in the signaling information the information related to the Kn value. According to an embodiment, the Kn value may be a table value for setting the maximum neighboring point range according to the calculated value obtained by automatically calculating the maximum neighboring point range. According to an embodiment, when the value of the information indicating whether to perform the automatic calculation is TRUE, the attribute decoder of the receiving device may use the information related to the Kn value signaled in the signaling information to determine the maximum neighboring point range.
[0452] According to an embodiment, when automatically calculating the maximum neighboring point range, the neighboring point set configuration unit 53008 and / or the signaling processor 51005 may receive information about the method of calculating the maximum neighboring point range as a user parameter (or encoder option). Additionally, information about the maximum neighboring point range calculation method (e.g., automatic_nn_range_method_type) may be signaled in the signaling information. According to an embodiment, the method of automatically calculating the maximum neighboring point range may include a method of calculating the maximum neighboring point range according to density (based on the denominator bounding box diagonal length or the numerator of the bounding box diagonal length) and a method of inferring density based on the axis length and the number of points and calculating the maximum neighboring point range (based on the maximum, minimum, or intermediate axis of the bounding box). According to an embodiment, the neighboring point set configuration unit 53008 may use the method of calculating the maximum neighboring point range according to density and / or the method of inferring density based on the axis length and the number of points and calculating the maximum neighboring point range as an automatic calculation method to automatically calculate the maximum neighboring point range. According to an embodiment, when the value of the information about whether to perform automatic calculation is TRUE, the attribute decoder of the receiving device may use the information about the maximum neighboring point range calculation method signaled in the signaling information to determine the maximum neighboring point range.
[0453] According to an embodiment, the neighboring point set configuration unit 53008 may calculate the maximum neighboring point distance based on the base neighboring point distance and the maximum neighboring point range.
[0454] According to an embodiment, the neighboring point set configuration unit 53008 may register only one or more points having a distance shorter than the calculated maximum neighboring point distance as the neighboring point set of the corresponding point.
[0455] According to an embodiment, the neighboring point set configuration unit 53008 may receive the method of applying the maximum neighboring point distance (e.g., nn_range_filtering_location_type) as a user parameter (or encoder option) and signal the method of applying the maximum neighboring point distance in the signaling information. According to an embodiment, the method of applying the maximum neighboring point distance may include a method of calculating the distance between points when selecting neighboring points and not selecting points outside the maximum neighboring point distance and a method of registering all neighboring points by calculating the distance between points and deleting points exceeding the maximum neighboring point distance among the registered points. According to an embodiment, the attribute decoder of the receiving device may configure the neighboring point set by applying the information about the method of applying the maximum neighboring point distance signaled in the signaling information (e.g., nn_range_filtering_location_type) to the neighboring point search.
[0456] The above LOD configuration method, basic proximity point distance acquisition method, maximum proximity point range (NN_range) acquisition method, and maximum proximity point distance acquisition method can be applied in the same or a similar manner as when the property decoder of the receiving device configures the proximity point set and LOD to predict the transform or lifting transform.
[0457] According to an embodiment, if the LOD is generated based on the basic proximity point distance and the maximum proximity point range (NN_range) as described above 1 and the maximum proximity point distance is determined, the proximity point set configuration unit 53008 searches for X (e.g., 3) NN points among the points within the search range in the group with the same or lower LOD (i.e., large distance between nodes) based on the LOD 1 set. Then, the proximity point set configuration unit 53008 may register only the NN points within the maximum proximity point distance among the X (e.g., 3) NN points as the proximity point set. Therefore, the number of NN points registered as the proximity point set is equal to or less than X. In other words, the NN points among the X NN points that are not within the maximum proximity point distance are not registered as the proximity point set and are excluded from the proximity point set. For example, if two of the three NN points are not within the maximum proximity point distance, the two NN points are excluded and only the other NN point (i.e., one NN point) is registered as the proximity point set.
[0458] According to another embodiment, if the LOD is generated based on the basic proximity point distance and the maximum proximity point range (NN_range) as described above 1 and the maximum proximity point distance is determined, the proximity point set configuration unit 53008 may search for X (e.g., 3) NN points within the maximum proximity point distance among the points within the search range in the group with the same or lower LOD (i.e., large distance between nodes) based on the LOD 1 set and register the X NN points as the proximity point set.
[0459] That is, the number of NN points that can be registered as the proximity point set may vary according to the timing (or position) at which the maximum proximity point distance is applied. In other words, the number of NN points registered as the proximity point set may be different depending on whether the maximum proximity point distance is applied after searching for X NN points through proximity point search or when calculating the distance between points. X is the maximum number of points that can be set as proximity points and can be input as a user parameter or signaled in the signaling information by the signaling processor 51005 (e.g., the lifting_num_pred_nearest_neighbours field signaled in the APS).
[0460] Figure 28A diagram illustrating an example of searching for neighboring points by applying a search range and a maximum neighboring point distance according to an embodiment. The arrows illustrated in the figure indicate the Morton order according to the embodiment.
[0461] In Figure 28 , in an embodiment, the index list includes the LOD set to which the point to be encoded belongs (i.e., LOD l ) and the reservation list includes at least one lower LOD set based on the LOD l set (e.g., LOD 0 to LOD l-1 ).
[0462] According to an embodiment, the points in the index list and the points in the reservation list are arranged in ascending order based on the size of the Morton code. Therefore, the foremost point among the points arranged in Morton order in the index list and the reservation list has the smallest Morton code size.
[0463] The neighboring point set configuration unit 53008 according to an embodiment may search, among the points belonging to the LOD 0 to LOD l-1 sets and / or the points belonging to the LOD 1 set, for the central point Pi having the Morton code closest to the Morton code of the point Px among the points that are sequentially located before the point Px (i.e., the points having a Morton code less than or equal to the Morton code of the point Px), in order to register the neighboring point set of the point Px (i.e., the point to be encoded or the current point) belonging to the LOD l set.
[0464] According to an embodiment, when searching for the central point Pi, the neighboring point set configuration unit 53008 may search for the point Pi having the Morton code closest to the Morton code of the point Px among all the points located before the point Px, or search for the point Pi having the Morton code closest to the Morton code of the point Px among the points within the search range. In the present disclosure, the search range may be configured by the neighboring point set configuration unit 53008, or may be input as a user parameter (also referred to as an encoder option). Additionally, information related to the search range may be signaled in the signaling information by the signaling processor 51005. The present disclosure provides embodiments of searching for the central point Pi in the reservation list when the number of LODs is multiple and searching for the central point Pi in the index list when the number of LODs is 1. For example, when the number of LODs is 1, the search range may be determined based on the position of the current point in the list arranged by the Morton code.
[0465] The neighboring point set configuration unit 53008 according to an embodiment performs a comparison between the point Px and the points before the searched (or selected) central point Pi (i.e., in Figure 28to the left of the center point in) and after the searched (or selected) center point Pi (i.e., to the right of the center point in Figure 28 a comparison of the distance values between the points belonging to the neighboring point search range. The neighboring point set configuration unit 53008 may select X (e.g., 3) NN points and register only the points within the maximum neighboring point distance at the LOD to which Px among the selected X points belongs as the neighboring point set of point Px. In an embodiment, the neighboring point search range is the number of points. The neighboring point search range according to an embodiment may include one or more points located before (i.e., in front of) and / or after (i.e., behind) the center point Pi in Morton order. X is the maximum number of points that can be registered as neighboring points.
[0466] According to an embodiment, the information related to the neighboring point search range and the information about the maximum number X of points that can be registered as neighboring points may be configured by the neighboring point set configuration unit 53008, or may be input as user parameters, or signaled in the signaling information by the signaling processor 51005 (e.g., the lifting_search_range field and the lifting_num_pred_nearest_neighbours field signaled in APS). According to an embodiment, the actual search range may be a value obtained by multiplying the value of the lifting_search_range field by 2 and then adding the center point to the obtained value (i.e., (value of the lifting_search_range field × 2) + center point), a value obtained by adding the center value to the value of the lifting_search_range field (i.e., value of the lifting_search_range field + center point), or the value of the lifting_search_range field. The present disclosure provides embodiments of searching for neighboring points of point Px in the reserved list when the number of LODs is multiple and searching for neighboring points of point Px in the index list when the number of LODs is 1.
[0467] For example, if the number of LODs is two or more and the information related to the search range (e.g., the lifting_search_range field) is 128, the actual search range includes the center point Pi in the reserved list arranged by Morton code, 128 points before the center point Pi, and 128 points after the center point Pi. As another example, if the number of LODs is one and the information related to the search range is 128, the actual search range includes the center point Pi in the index list arranged by Morton code and 128 points before the center point Pi. As another example, if the number of LODs is one and the information related to the search range is 128, the actual search range includes 128 points before the current point Px in the index list arranged by Morton code.
[0468] According to an embodiment, the neighboring point set configuration unit 53008 may compare the distance values between the points in the actual search range and the point Px to search for X NN points, and register only the points within the maximum neighboring point distance at the LOD to which Px belongs among the X points as the neighboring point set of the point Px. That is, the neighboring points registered as the neighboring point set of the point Px are limited to the points within the maximum neighboring point distance at the LOD to which Px belongs among the X points.
[0469] As described above, the present disclosure can provide an effect of enhancing the attribute compression efficiency by restricting the points that can be selected as the neighboring point set in consideration of the attribute characteristics (correlation) between the points of the point cloud content.
[0470] Figure 29 FIG. is a diagram illustrating another example of searching for neighboring points by applying a search range and a maximum neighboring point distance according to an embodiment. The arrows illustrated in the figure indicate the Morton order according to the embodiment.
[0471] Since Figure 29 the example of Figure 28 is similar to the example of Figure 29 except for the position where the maximum neighboring point distance is applied, the repeated description thereof will be omitted herein. Therefore, for the parts omitted or not described in Figure 28 refer to the description of
[0472] According to an embodiment, the points in the index list and the points in the reserved list are arranged in ascending order based on the size of the Morton code.
[0473] The neighboring point set configuration unit 53008 according to an embodiment may be in the LOD 0 to the LOD l-1 points of the set and / or belonging to the LOD lAmong the points in the set, search for the center point Pi with the Morton code closest to the Morton code of point Px among the points that are sequentially before point Px (i.e., points with Morton codes less than or equal to the Morton code of point Px) in order to register the ones belonging to the LOD l The set of neighboring points of the point Px (i.e., the point to be encoded or the current point) in the set.
[0474] According to the neighboring point set configuration unit 53008 of the embodiment, perform a comparison of the distance values between point Px and the points belonging to the neighboring point search range before the searched (or selected) center point Pi (i.e., to the left of the center point in Figure 29 and after the searched (or selected) center point Pi (i.e., to the right of the center point in Figure 29 ). The neighboring point set configuration unit 53008 may select X (e.g., 3) NN points within the maximum neighboring point distance at the LOD to which point Px belongs and register these X points as the set of neighboring points of point Px.
[0475] The present disclosure provides embodiments for searching for the neighboring points of point Px in the actual search range of the reservation list when the number of LODs is multiple and for searching for the neighboring points of point Px in the actual search range of the index list when the number of LODs is 1.
[0476] For example, if the number of LODs is 2 or more and the information about the search range (e.g., the lifting_search_range field) is 128, the actual search range includes the center point Pi in the reservation list arranged by Morton code, 128 points before the center point Pi, and 128 points after the center point Pi. As another example, if the number of LODs is 1 and the information about the search range is 128, the actual search range includes the center point Pi in the index list arranged by Morton code and 128 points before the center point Pi. As another example, if the number of LODs is 1 and the information about the search range is 128, the actual search range includes 128 points before the current point Px in the index list arranged by Morton code.
[0477] When comparing the distance values between the points in the actual search range and point Px, the neighboring point set configuration unit 53008 of the embodiment selects X (e.g., 3) NN points within the maximum neighboring point distance at the LOD to which point Px belongs from among the points in the actual search range and registers these X points as the set of neighboring points of point Px. That is, the neighboring points registered as the set of neighboring points of point Px are limited to the points within the maximum neighboring point distance at the LOD to which point Px belongs among the X points.
[0478] As described above, the present disclosure can provide an effect of enhancing the compression efficiency of attributes by selecting a set of neighboring points in consideration of the attribute characteristics (correlations) between points in the point cloud content.
[0479] According to the above-described embodiment, when a set of neighboring points is registered in each predictor of a point to be encoded in the neighboring point set configuration unit 53008, the attribute information prediction unit 53009 predicts the attribute value of the corresponding point from one or more neighboring points registered in the predictor. As described above, when configuring the set of neighboring points, by applying the maximum neighboring point distance, the number of neighboring points included in the set of neighboring points registered in each predictor is equal to or less than X (e.g., 3). According to an embodiment, the predictor of a point can register a 1 / 2 distance (= weight) based on the distance values to each neighboring point with the registered set of neighboring points. For example, taking (P2 P4 P6) as the set of neighboring points, the predictor of the P3 node calculates the weight based on the distance values to each neighboring point. According to an embodiment, the weight of a neighboring point can be
[0480] According to an embodiment, when the set of neighboring points is registered in the predictor, the neighboring point set configuration unit 53008 or the attribute information prediction unit 53009 can normalize the weight of each neighboring point with the total weight of the neighboring points included in the set of neighboring points.
[0481] For example, the weights of all neighboring points in the neighboring set of the P3 node are summed and the weight of each neighboring point is divided by the sum of the weights ( / total_weight, / total_weight, / total_weight). Thus, the weight of each neighboring point is normalized.
[0482] Then, the attribute information prediction unit 5301 can predict the attribute value through the predictor.
[0483] According to an embodiment, the average of the values obtained by multiplying the attributes (e.g., color, reflectance, etc.) of the neighboring points registered in the predictor by the weight (or the normalized weight) can be set as the prediction result (i.e., the predicted attribute value). Alternatively, the attribute of a specific point can be set as the prediction result (i.e., the predicted attribute value). According to an embodiment, the predicted attribute value can be referred to as predicted attribute information. Additionally, the residual attribute value (or residual attribute information or residual) can be obtained by subtracting the predicted attribute value (or predicted attribute information) of a point from the attribute value (i.e., the original attribute value) of the point.
[0484] According to an embodiment, the compression result value can be pre-calculated by applying various prediction patterns (or predictor indices), and then the prediction pattern (i.e., predictor index) that generates the smallest bitstream can be selected from among the prediction patterns.
[0485] Next, the process of selecting the prediction pattern will be described in detail.
[0486] In this specification, the prediction pattern has the same meaning as the predictor index (Preindex) and can be broadly referred to as a prediction method.
[0487] In an embodiment, the attribute information prediction unit 53009 can perform the process of finding the most suitable prediction pattern for each point and setting the found prediction pattern in the predictor corresponding to the point.
[0488] According to an embodiment, the prediction pattern in which the predicted attribute value is calculated by weighted average (i.e., the average obtained by multiplying the attributes of neighboring points set in the predictor of each point by the weights calculated based on the distances to each neighboring point) will be referred to as prediction pattern 0. Additionally, the prediction pattern in which the attribute of the first neighboring point is set as the predicted attribute value will be referred to as prediction pattern 1, the prediction pattern in which the attribute of the second neighboring point is set as the predicted attribute value will be referred to as prediction pattern 2, and the prediction pattern in which the attribute of the third neighboring point is set as the predicted attribute value will be referred to as prediction pattern 3. In other words, a value of the prediction pattern (or predictor index) equal to 0 can indicate predicting the attribute value by weighted average, and a value equal to 1 can indicate predicting the attribute value by the first neighboring node (i.e., neighboring point). A value equal to 2 can indicate predicting the attribute value by the second neighboring node, and a value equal to 3 can indicate predicting the attribute value by the third neighboring node.
[0489] According to an embodiment, the residual attribute value in prediction pattern 0, the residual attribute value in prediction pattern 1, the residual attribute value in prediction pattern 2, and the residual attribute value in prediction pattern 3 can be calculated, and a score or a double score can be calculated based on each residual attribute value. Then, the prediction pattern with the lowest calculated score can be selected and set as the prediction pattern for the corresponding point.
[0490] According to an embodiment, when a preset condition is satisfied, the process of finding the most suitable prediction pattern among multiple prediction patterns and setting it as the prediction pattern for the corresponding point can be executed. Therefore, when the preset condition is not satisfied, a fixed (or predefined) prediction pattern (e.g., prediction pattern 0) in which the predicted attribute value is calculated by weighted average can be set as the prediction pattern for the point without performing the process of finding the most suitable prediction pattern. In an embodiment, this process is performed for each point.
[0491] According to an embodiment, when the difference between attribute elements (e.g., R, G, B) between neighboring points registered in the predictor of a point is greater than or equal to a preset threshold (e.g., lifting_adaptive_prediction_threshold), or when the difference between attribute elements (e.g., R, G, B) between neighboring points registered in the predictor of a point is calculated and the sum of the maximum differences of the elements is greater than or equal to the preset threshold, a preset condition can be satisfied for a specific point. For example, assume that point P3 is a specific point, and points P2, P4, and P6 are registered as neighboring points of point P3. Additionally, assume that when the differences in R, G, and B values between points P2 and P4, between points P2 and P6, and between points P4 and P6 are calculated, the maximum difference in the R value is obtained between points P2 and P4, the maximum difference in the G value is obtained between points P4 and P6, and the maximum difference in the B value is obtained between points P2 and P6. Additionally, assume that among the maximum difference in the R value (between P2 and P4), the maximum difference in the G value (between P4 and P6), and the maximum difference in the B value (between P2 and P6), the difference in the R value between points P2 and P4 is the largest.
[0492] Under these assumptions, when the difference in the R value between points P2 and P4 is greater than or equal to the preset threshold, or when the sum of the difference in the R value between points P2 and P4, the difference in the G value between points P4 and P6, and the difference in the B value between points P2 and P6 is greater than or equal to the preset threshold, the process of finding the most suitable prediction mode among multiple candidate prediction modes can be performed. Additionally, the prediction mode (e.g., predIndex) can be signaled only when the difference in the R value between points P2 and P4 is greater than or equal to the preset threshold or the sum of the difference in the R value between points P2 and P4, the difference in the G value between points P4 and P6, and the difference in the B value between points P2 and P6 is greater than or equal to the preset threshold.
[0493] According to another embodiment, when the maximum difference between the values of the attributes (e.g., reflectance) of neighboring points registered in the predictor of a specific point is greater than or equal to a preset threshold (e.g., lifting_adaptive_prediction_threshold), a preset condition can be satisfied for that point. Additionally, assume that among the reflectance differences between points P2 and P4, between points P2 and P6, and between points P4 and P6, the reflectance difference between points P2 and P4 is the largest.
[0494] Under this assumption, when the reflectance difference between point P2 and P4 is greater than or equal to a preset threshold, the process of finding the most suitable prediction mode among multiple candidate prediction modes can be performed. Additionally, the prediction mode (e.g., predIndex) can be signaled only when the reflectance difference between point P2 and P4 is greater than or equal to the preset threshold.
[0495] According to an embodiment, the selected prediction mode (e.g., predIndex) of the corresponding point can be signaled in the attribute slice data. In this case, the sender transmits the residual attribute value obtained based on the selected prediction mode, and the receiver obtains the predicted attribute value of the corresponding point based on the signaled prediction mode and adds the predicted attribute value and the received residual attribute value to restore the attribute value of the corresponding point.
[0496] In another embodiment, when the prediction mode is not signaled, the sender can calculate the predicted attribute value based on the prediction mode set as the default mode (e.g., prediction mode 0) and calculate and transmit the residual attribute value based on the difference between the original attribute value and the predicted attribute value. The receiver can calculate the predicted attribute value based on the prediction mode set as the default mode (e.g., prediction mode 0) and restore the attribute value by adding this value to the received residual attribute value.
[0497] According to an embodiment, the threshold (e.g., the lifting_adaptive_prediction_threshold field signaled in APS) can be signaled in the signaling information or directly input through the signaling processor 51005.
[0498] According to an embodiment, when the preset condition is satisfied for a specific point as described above, a predictor candidate can be generated. The predictor candidate is referred to as a prediction mode or a predictor index.
[0499] According to an embodiment, prediction modes 1 to 3 can be included in the predictor candidate. According to an embodiment, prediction mode 0 may or may not be included in the predictor candidate. According to an embodiment, at least one prediction mode not mentioned above can also be included in the predictor candidate.
[0500] The prediction mode set for each point through the above process and the residual attribute value in the set prediction mode are output to the residual attribute information quantization processor 53010.
[0501] According to an embodiment, the residual attribute information quantization processor 53010 can apply zero run length coding to the input residual attribute value.
[0502] According to an embodiment of the present disclosure, quantization and zero run length coding can be performed on the residual attribute value.
[0503] According to an embodiment, the arithmetic encoder 53011 applies arithmetic coding to the residual attribute values output from the residual attribute information quantizer 53010 and the prediction mode, and outputs the result as an attribute bitstream.
[0504] The geometric bitstream compressed and output by the geometric encoder 51006 and the attribute bitstream compressed and output by the attribute encoder 51007 are output to the transmission processor 51008.
[0505] The transmission processor 51008 according to an embodiment may perform operations and / or a transmission method that are the same as or similar to the operations and / or the transmission method of the Figure 12 transmission processor 12012, and perform operations and / or a transmission method that are the same as or similar to the operations and / or the transmission method of the Figure 1 transmitter 10003. More specifically, reference will be made to the description of Figure 1 or Figure 12 .
[0506] The transmission processor 51008 according to an embodiment may transmit the geometric bitstream output from the geometric encoder 51006, the attribute bitstream output from the attribute encoder 51007, and the signaling bitstream output from the signaling processor 51005 separately, or may multiplex the bitstreams into one bitstream to be transmitted.
[0507] The transmission processor 51008 according to an embodiment may encapsulate the bitstream in a file or a segment (e.g., a streaming segment), and then transmit the encapsulated bitstream through various networks such as a broadcast network and / or a broadband network.
[0508] The signaling processor 51005 according to an embodiment may generate and / or process signaling information, and output it as a bitstream to the transmission processor 51008. The signaling information generated and / or processed by the signaling processor 51005 will be provided to the geometric encoder 51006, the attribute encoder 51007, and / or the transmission processor 51008 for geometric coding, attribute coding, and transmission processing. Alternatively, the signaling processor 51005 may receive the signaling information generated by the geometric encoder 51006, the attribute encoder 51007, and / or the transmission processor 51008.
[0509] In the present disclosure, signaling information may be signaled and transmitted in units of parameter sets (such as Sequence Parameter Set (SPS), Geometry Parameter Set (GPS), Attribute Parameter Set (APS), Tile Parameter Set (TPS), etc.). Additionally, it may be signaled and transmitted based on coding units of each image such as slices or tiles. In the present disclosure, signaling information may include metadata (e.g., set values) related to point cloud data and may be provided to the geometry encoder 51006, the attribute encoder 51007, and / or the transmission processor 51008 for geometry coding, attribute coding, and transmission processing. Depending on the application, signaling information may also be defined on the system side such as in file formats, Dynamic Adaptive Streaming over HTTP (DASH), or MPEG Media Transport (MMT), or on the wired interface side such as in High-Definition Multimedia Interface (HDMI), DisplayPort, Video Electronics Standards Association (VESA), or CTA.
[0510] The method / apparatus according to an embodiment may signal relevant information to add / perform the operations of the embodiment. The signaling information according to an embodiment may be used in a transmitting device and / or a receiving device.
[0511] According to an embodiment of the present disclosure, information about the maximum number of predictors (lifting_max_num_direct_predictors) for attribute prediction, threshold information (lifting_adaptive_prediction_threshold) for enabling adaptive prediction of attributes, option information related to neighboring point selection, etc. may be signaled in at least one of the sequence parameter set, the attribute parameter set, the tile parameter set, or the attribute slice header. Additionally, according to an embodiment, predictor index information (predIndex) indicating a prediction mode corresponding to a predictor candidate selected from among a plurality of predictor candidates may be signaled in the attribute slice data.
[0512] According to an embodiment, the neighboring point selection related option information may include at least one of information about a basic neighboring point distance calculation method (e.g., nn_base_distance_calculation_method_type), basic neighboring point distance information (e.g., nn_base_distance), maximum neighboring point range information (e.g., nearest_neighbore_max_range), minimum neighboring point range information (e.g., nearest_neighbor_min_range), information about a method of applying the maximum neighboring point distance (e.g., nn_range_filtering_location_type), information about whether to automatically calculate the maximum neighboring point range (e.g., automatic_nn_range_calculation_flag), information about a method of calculating the maximum neighboring point range (e.g., automatic_nn_range_method_type), and information related to the Kn value (e.g., automatic_max_nn_range_in_table, automatic_nn_range_table_k). According to an embodiment, the neighboring point selection related option information may further include at least one of information about a maximum number of points that can be set as neighboring points (e.g., lifting_num_pred_nearest_neighbours), search range related information (e.g., lifting_search_range), and / or LOD configuration method.
[0513] Similar to the point cloud video encoder of the above-described transmitting device, the point cloud video decoder of the receiving device performs processing similar to or the same as generating an LOD l set, finding the nearest neighboring points based on the LOD l set and registering them as a neighboring point set in the predictor, calculating weights based on the distance to each neighboring point, and performing normalization using the weights. Then, the decoder decodes the received prediction mode and predicts the attribute value of the point according to the decoded prediction mode. Additionally, after the received residual attribute value is decoded, the decoded residual attribute value may be added to the predicted attribute value to restore the attribute value of the point.
[0514] Figure 30 is a diagram showing another exemplary point cloud receiving device according to an embodiment.
[0515] The point cloud receiving apparatus according to an embodiment may include a receiving processor 61001, a signaling processor 61002, a geometry decoder 61003, an attribute decoder 61004, and a post-processor 61005. According to an embodiment, the geometry decoder 61003 and the attribute decoder 61004 may be referred to as a point cloud video decoder. According to an embodiment, the point cloud video decoder may be referred to as a PCC decoder, a PCC decoding unit, a point cloud decoder, a point cloud decoding unit, etc.
[0516] The receiving processor 61001 according to an embodiment may receive a single bitstream, or may separately receive a geometry bitstream, an attribute bitstream, and a signaling bitstream. When receiving a file and / or a segment, the receiving processor 61001 according to an embodiment may unpack the received file and / or segment, and output the unpacked file and / or segment as a bitstream.
[0517] When a single bitstream is received (or unpacked), the receiving processor 61001 according to an embodiment may demultiplex a geometry bitstream, an attribute bitstream, and / or a signaling bitstream from the single bitstream. The receiving processor 61001 may output the demultiplexed signaling bitstream to the signaling processor 61002, output the geometry bitstream to the geometry decoder 61003, and output the attribute bitstream to the attribute decoder 61004.
[0518] When the geometry bitstream, the attribute bitstream, and / or the signaling bitstream are separately received (or unpacked), the receiving processor 61001 according to an embodiment may transmit the signaling bitstream to the signaling processor 61002, transmit the geometry bitstream to the geometry decoder 61003, and transmit the attribute bitstream to the attribute decoder 61004.
[0519] The signaling processor 61002 may parse signaling information from the input signaling bitstream, for example, information included in SPS, GPS, APS, TPS, metadata, etc., process the parsed information, and provide the processed information to the geometry decoder 61003, the attribute decoder 61004, and the post-processor 61005. In another embodiment, the signaling information included in the geometry slice header and / or the attribute slice header may also be parsed by the signaling processor 61002 before decoding the corresponding slice data. That is, when the point cloud data is segmented into tiles and / or slices at the sending side as Figure 16 shown, the TPS includes the number of slices included in each tile, and accordingly, the point cloud video decoder according to an embodiment may check the number of slices and quickly parse the information for parallel decoding.
[0520] Thus, the point cloud video decoder according to the present disclosure can quickly parse a bitstream containing point cloud data when it receives an SPS with a reduced amount of data. The receiving device can decode a tile after receiving it and can decode each slice based on the GPS and APS included in each tile. Thereby, the decoding efficiency can be maximized.
[0521] That is, the geometry decoder 61003 can perform an inverse process of the operation of the geometry encoder 51006 that Figure 15 compresses the geometry bitstream based on signaling information (e.g., geometry-related parameters) to reconstruct the geometry. The geometry restored (or reconstructed) by the geometry decoder 61003 is provided to the attribute decoder 61004. The attribute decoder 61004 can perform an inverse process of the operation of the attribute encoder 51007 that Figure 15 compresses the attribute bitstream based on signaling information (e.g., attribute-related parameters) and the reconstructed geometry to restore the attributes. According to an embodiment, when the point cloud data is segmented into tiles and / or slices on the transmission side as shown in Figure 16 , the geometry decoder 61003 and the attribute decoder 61004 perform geometry decoding and attribute decoding tile by tile and / or slice by slice.
[0522] Figure 31 is a detailed block diagram illustrating another example of the geometry decoder 61003 and the attribute decoder 61004 according to an embodiment.
[0523] Figure 31 The arithmetic decoder 63001, octree reconstruction unit 63002, geometry information prediction unit 63003, inverse quantization processor 63004, and coordinate inverse transformation unit 63005 included in the geometry decoder 61003 in Figure 11 can perform some or all of the operations of the arithmetic decoder 11000, octree synthesis unit (octree synthesizer) 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate inverse transformation unit (coordinate inverse transformer) 11004 in Figure 13 , or can perform some or all of the operations of the arithmetic decoder 13002, occupancy code-based octree reconstruction processor 13003, surface model processor 13004, and inverse quantization processor 13005 in
[0524] According to an embodiment, when information about the maximum number of predictors (lifting_max_num_direct_predictors) to be used for attribute prediction, threshold information (lifting_adaptive_prediction_threshold) for enabling adaptive prediction of attributes, neighboring point selection related information, etc. is signaled in at least one of a sequence parameter set (SPS), an attribute parameter set (APS), a tile parameter set (TPS), or an attribute slice header, they may be obtained by the signaling processor 61002 and provided to the attribute decoder 61004, or may be directly obtained by the attribute decoder 61004.
[0525] The attribute decoder 61004 according to an embodiment may include an arithmetic decoder 63006, a LOD configuration unit 63007, a neighboring point set configuration unit 63008, an attribute information prediction unit (attribute information predictor) 63009, a residual attribute information inverse quantization processor 63010, and an inverse color transformation processor 63011.
[0526] The arithmetic decoder 63006 according to an embodiment may perform arithmetic decoding on an input attribute bitstream. The arithmetic decoder 63006 may decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 63006 performs operations and / or decoding that are the same as or similar to the operations and / or decoding of the arithmetic decoder 11005 of Figure 11 or the arithmetic decoder 13007 of Figure 13 and / or decoding.
[0527] According to an embodiment, the attribute information output from the arithmetic decoder 63006 may be decoded by one or a combination of two or more of RAHT decoding, LOD-based prediction transform decoding, and lifting transform decoding based on the reconstructed geometry information.
[0528] This is described as an embodiment in which the transmitting device performs attribute compression by one or a combination of LOD-based prediction transform coding and lifting transform coding. Accordingly, an embodiment in which the receiving device performs attribute decoding by one or a combination of LOD-based prediction transform decoding and lifting transform decoding will be described. The description of RAHT decoding of the receiving device will be omitted.
[0529] According to an embodiment, the attribute bitstream arithmetically decoded by the arithmetic decoder 63006 is provided to the LOD configuration unit 63007. According to an embodiment, the attribute bitstream provided from the arithmetic decoder 63006 to the LOD configuration unit 63007 may include a prediction mode and a residual attribute value.
[0530] According to an embodiment, the LOD configuration unit 63007 generates one or more LODs in the same or a similar manner as the LOD configuration unit 53007 of the transmitting device, and outputs the generated one or more LODs to the neighboring point set configuration unit 63008.
[0531] According to an embodiment, the LOD configuration unit 63007 may configure one or more LODs using one or more LOD generation methods (or LOD configuration methods). According to an embodiment, the LOD generation method used in the LOD configuration unit 63007 may be provided by the signaling processor 61002. For example, the LOD generation method may be signaled in the APS of the signaling information. According to an embodiment, the LOD generation method may be classified into an octree-based LOD generation method, a distance-based LOD generation method, and a sampling-based LOD generation method.
[0532] According to an embodiment, a group having different LODs is referred to as an LOD l set. Here, l represents the LOD and is an integer starting from 0. The LOD 0 is a set composed of points with the largest distance therebetween. As l increases, the distance between points belonging to the LOD l decreases.
[0533] According to an embodiment, the prediction mode and residual attribute value encoded by the transmitting device may be provided for each LOD or only for leaf nodes.
[0534] In one embodiment, when the LOD configuration unit 63007 generates the LOD l set, the neighboring point set configuration unit 63008 may search for neighboring points equal to or less than X (e.g., 3) points in a group having the same or a lower LOD (i.e., a large distance between nodes) based on the LOD l set, and register the searched neighboring points as a neighboring point set in the predictor.
[0535] According to an embodiment, the neighboring point set configuration unit 63008 configures the neighboring point set by obtaining a search range and / or a maximum neighboring point distance based on signaling information and / or applying the search range and / or the maximum neighboring point distance to the neighboring point search process.
[0536] According to an embodiment, the neighboring point set configuration unit 63008 can obtain the maximum neighboring point distance by multiplying the basic neighboring point distance by the NN_range. The NN_range is the range within which neighboring points can be selected, and will be referred to as the maximum range within which neighboring points can be selected, the maximum neighboring point range, the neighboring point range, or the nearest neighbor range. According to an embodiment, the neighboring point set configuration unit 63008 can calculate / configure the basic neighboring point distance by combining one or more of an octree-based method, a distance-based method, a sampling-based method, a method based on the average difference of Morton codes for each LOD, and a method based on the average distance difference for each LOD based on signaling information.
[0537] Since a detailed description of calculating / configuring the basic neighboring point distance by combining one or more of an octree-based method, a distance-based method, a sampling-based method, a method based on the average difference of Morton codes for each LOD, and a method based on the average distance difference for each LOD in the neighboring point set configuration unit 63008 according to an embodiment has been given in detail in the process of encoding attributes of the above transmission device, the details thereof will be omitted herein.
[0538] According to an embodiment, the neighboring point set configuration unit 63008 can obtain the basic neighboring point distance from the neighboring point selection related option information included in the signaling information.
[0539] According to an embodiment, the neighboring point set configuration unit 63008 can automatically or manually set the NN_range according to the characteristics of the content. The neighboring point set configuration unit 63008 can automatically or manually set the NN_range based on the neighboring point selection related option information included in the signaling information according to the characteristics of the content.
[0540] According to an embodiment, the neighboring point set configuration unit 63008 can directly obtain the maximum neighboring point range (NN_range) from the neighboring point selection related option information, or can calculate the NN_range based on the neighboring point selection related option information.
[0541] Details of automatically calculating the NN_range by the neighboring point set configuration unit 63008 according to an embodiment have been described above in the description of the attribute encoding process of the transmission device, so the description thereof will be skipped.
[0542] According to an embodiment, the neighbor point selection related option information may include at least one of information about a basic neighbor point distance calculation method (e.g., nn_base_distance_calculation_method_type), basic neighbor point distance information (e.g., nn_base_distance), maximum neighbor point range information (e.g., nearest_neighbore_max_range), minimum neighbor point range information (e.g., nearest_neighbor_min_range), information about a method of applying the maximum neighbor point distance (e.g., nn_range_filtering_location_type), information about whether to automatically calculate the maximum neighbor point range (e.g., automatic_nn_range_calculation_flag), information about a method of calculating the maximum neighbor point range (e.g., automatic_nn_range_method_type), and information related to the Kn value (e.g., automatic_max_nn_range_in_table, automatic_nn_range_table_k). According to an embodiment, the neighbor point selection related option information may further include information about a maximum number of points that can be set as neighbor points (e.g., lifting_num_pred_nearest_neighbours), search range related information (e.g., lifting_search_range), and / or LOD configuration method. In the present disclosure, for simplicity, at least one of information about a basic neighbor point distance calculation method (e.g., nn_base_distance_calculation_method_type) and basic neighbor point distance information (e.g., nn_base_distance) may be referred to as basic neighbor point distance related information.
[0543] According to an embodiment, the information about a basic neighbor point distance calculation method (e.g., nn_base_distance_calculation_method_type) may indicate the use of a basic distance input as a user parameter (or encoder option), octree-based maximum neighbor point distance calculation, distance-based maximum neighbor point distance calculation, sampling-based maximum neighbor point distance calculation, maximum neighbor point distance calculation by calculating an average difference of Morton codes between LODs, or maximum neighbor point distance calculation by calculating an average distance difference between LODs.
[0544] According to an embodiment, information about a method of applying a maximum neighbor point distance (e.g., nn_range_filtering_location_type) may indicate a method of calculating a distance between points when selecting neighbor points and not selecting points outside the maximum neighbor point distance, or a method of registering all neighbor points by calculating a distance between points and removing points exceeding the maximum neighbor point distance among the registered points.
[0545] According to an embodiment, information about whether to automatically calculate a maximum neighbor point range (e.g., automatic_nn_range_calculation_flag) may indicate whether to automatically calculate the maximum neighbor point range.
[0546] According to an embodiment, information about a method of calculating a maximum neighbor point range (e.g., automatic_nn_range_method_type) may indicate a method of calculating a maximum neighbor point range by estimating density based on a diagonal length of a bounding box or a method of calculating a maximum neighbor point range by estimating density based on an axis length of a bounding box and the number of points.
[0547] The method of calculating a maximum neighbor point range by estimating density based on a diagonal length of a bounding box may include estimating density using the diagonal length of the bounding box as a denominator and estimating density using the diagonal length of the bounding box as a numerator.
[0548] The method of calculating a maximum neighbor point range by estimating density based on an axis length of a bounding box and the number of points may include a method of estimating density based on a maximum axis length of a bounding box and the number of points, a method of estimating density based on a minimum axis length of a bounding box and the number of points, and a method of estimating density based on an intermediate axis length and the number of points.
[0549] According to an embodiment, the neighbor point set configuration unit 63008 estimates density for automatic calculation using one or a combination of two or more of the following methods: a method of estimating density using the diagonal length of a bounding box as a denominator, a method of estimating density using the diagonal length of a bounding box as a numerator, a method of estimating density based on a maximum axis length of a bounding box and the number of points, a method of estimating density based on a minimum axis length of a bounding box and the number of points, and a method of estimating density based on an intermediate axis length and the number of points.
[0550] Information related to the Kn value may include the number of entries in a table of Kn values (e.g., automatic_max_nn_range_in_table) and the Kn value (e.g., automatic_nn_range_table_k).
[0551] According to an embodiment, the neighboring point set configuration unit 63008 may automatically calculate a maximum neighboring point range (NN_range) based on information about whether to automatically calculate the maximum neighboring point range (e.g., automatic_nn_range_calculation_flag). For example, when the information about whether to automatically calculate the maximum range (e.g., automatic_nn_range_calculation_flag) indicates automatic calculation, the maximum neighboring point range (NN_range) may be automatically calculated based on information related to the basic neighboring point distance, information about the maximum neighboring point range calculation method, and information related to the Kn value.
[0552] According to an embodiment, if the maximum neighboring point distance is determined as described above, the neighboring point set configuration unit 63008 searches for X (e.g., 3) NN points among the points within the search range in a group having the same or lower LOD (i.e., large distance between nodes) based on the LOD l set, as in Figure 28 Then, the neighboring point set configuration unit 63008 may register only the NN points within the maximum neighboring point distance among the X (e.g., 3) NN points as the neighboring point set. Therefore, the number of NN points registered as the neighboring point set is equal to or less than X. In other words, the NN points among the X NN points that are not within the maximum neighboring point distance are not registered as the neighboring point set and are excluded from the neighboring point set. For example, if two of the three NN points are not within the maximum neighboring point distance, the two NN points are excluded and only the other NN point (i.e., one NN point) is registered as the neighboring point set.
[0553] According to another embodiment, if the maximum neighboring point distance is determined as described above, as in Figure 29 the neighboring point set configuration unit 53008 may search for X (e.g., 3) NN points within the maximum neighboring point distance among the points within the search range in a group having the same or lower LOD (i.e., large distance between nodes) based on the LOD 1 set and register the X NN points as the neighboring point set.
[0554] That is, the number of NN points that can be registered as the neighboring point set may vary according to the timing (or position) at which the maximum neighboring point distance is applied. In other words, as in Figure 29 the number of NN points registered as the neighboring point set may be different depending on whether the maximum neighboring point distance is applied after searching for X NN points through neighboring point search or when calculating the distance between points.
[0555] Referring to Figure 28, the neighboring point set configuration unit 63008 can compare the distance values between the points within the actual search range and the point Px to search for X NN points, and register only the points within the maximum neighboring point distance at the LOD to which Px belongs among the X points. That is, the neighboring points registered as the neighboring point set of point Px are limited to the points within the maximum neighboring point distance at the LOD to which point Px belongs among the X points.
[0556] Refer to the following as an example Figure 29 , when comparing the distance values between the points within the actual search range and the point Px, the neighboring point set configuration unit 63008 selects X (e.g., 3) NN points within the maximum neighboring point distance at the LOD to which point Px belongs from among the points within the actual search range, and registers the selected points as the neighboring point set of point Px. That is, the neighboring points registered as the neighboring point set of point Px are limited to the points within the maximum neighboring point distance at the LOD to which point Px belongs.
[0557] In Figure 28 and Figure 29 , the actual search range can be a value obtained by multiplying the value of the lifting_search_range field by 2 and then adding the center point to the resulting value (i.e., (lifting_search_range field value × 2) + center point), a value obtained by adding the center value to the lifting_search_range field value (i.e., lifting_search_range field value + center point), or the value of the lifting_search_range field. In an embodiment, the neighboring point set configuration unit 63008 searches for the neighboring points of point Px in the reservation list when the number of LODs is multiple, and searches for the neighboring points of point Px in the index list when the number of LODs is 1.
[0558] For example, if the number of LODs is 2 or more and the information about the search range (e.g., the lifting_search_range field) is 128, the actual search range includes the center point Pi in the reservation list arranged in Morton code, 128 points before the center point Pi, and 128 points after the center point Pi. As another example, if the number of LODs is 1 and the information about the search range is 128, the actual search range includes the center point Pi in the index list arranged in Morton code and 128 points before the center point Pi. As another example, if the number of LODs is 1 and the information related to the search range is 128, the actual search range includes 128 points before the current point Px in the index list arranged in Morton code.
[0559] For example, assume that the neighboring point set configuration unit 63008 selects points P2, P4, and P6 as neighboring points of point P3 (i.e., the node) belonging to the LOD 1 and registers the selected points as a neighboring point set in the predictor of P3 (see Figure 9 ).
[0560] According to an embodiment, when the neighboring point set is registered in each predictor of the point to be decoded in the neighboring point set configuration unit 63008, the attribute information prediction unit (attribute information predictor) 63009 predicts the attribute value of the corresponding point from one or more neighboring points registered in each predictor.
[0561] According to an embodiment, the attribute information predictor 63009 performs a process of predicting the attribute value of a point based on a prediction mode of a specific point. The attribute prediction process is performed on all points or at least some points of the reconstructed geometry.
[0562] The prediction mode of the specific point according to the embodiment may be one of prediction mode 0 to prediction mode 3.
[0563] According to an embodiment, prediction mode 0 is a mode of calculating a predicted attribute value by weighted average, prediction mode 1 is a mode of determining the attribute of the first neighboring point as the predicted attribute value, prediction mode 2 is a mode of determining the attribute of the second neighboring point as the predicted attribute value, and prediction mode 3 is a mode of determining the attribute of the third neighboring point as the predicted attribute value.
[0564] According to an embodiment, when the maximum difference between the attribute values of the neighboring points registered in the predictor of a point is less than a preset threshold, the attribute encoder on the transmitting side sets prediction mode 0 as the prediction mode of the point. When the maximum difference is greater than or equal to the preset threshold, the attribute encoder applies the RDO method to multiple candidate prediction modes and sets one of the candidate prediction modes as the prediction mode of the point. In the embodiment, this process is performed for each point.
[0565] According to an embodiment, the prediction mode (predIndex) of a point selected by applying the RDO method may be signaled in the attribute slice data. Accordingly, the prediction mode of a point can be obtained from the attribute slice data.
[0566] According to an embodiment, the attribute information prediction unit 63009 may predict the attribute value of each point based on the prediction mode of each point set as described above.
[0567] For example, when it is assumed that the prediction mode of point P3 is prediction mode 0, the average value of the values obtained by multiplying the attributes of points P2, P4, and P6, which are neighboring points registered in the predictor of point P3, by weights (or normalized weights) may be calculated. The calculated average value may be determined as the predicted attribute value of the point.
[0568] As another example, when the prediction mode of point P3 is prediction mode 1, the attribute value of point P4, which is a neighboring point registered in the predictor at point P3, can be determined as the predicted attribute value of point P3.
[0569] As another example, when the prediction mode of point P3 is prediction mode 2, the attribute value of point P6, which is a neighboring point registered in the predictor at point P3, can be determined as the predicted attribute value of point P3.
[0570] As another example, when the prediction mode of point P3 is prediction mode 3, the attribute value of point P2, which is a neighboring point registered in the predictor at point P3, can be determined as the predicted attribute value of point P3.
[0571] Once the attribute information prediction unit 63009 obtains the predicted attribute value of a point based on the prediction mode of the point, the residual attribute information inverse quantization processor 63010 restores the attribute value of the point by adding the predicted attribute value of the point predicted by the predictor 63009 to the received residual attribute value of the point, and then performs inverse quantization as an inverse process of the quantization process of the transmitting device.
[0572] In an embodiment, when zero run - length coding is applied to the residual attribute value of a point on the transmission side, the residual attribute information inverse quantization processor 63010 performs zero run - length decoding on the residual attribute value of the point and then performs inverse quantization.
[0573] The attribute value restored by the residual attribute information inverse quantization processor 63010 is output to the inverse color transform processor 63011.
[0574] The inverse color transform processor 63011 performs inverse transform coding of the inverse transform of the color value (or texture) included in the restored attribute value, and then outputs the attribute to the post - processor 61005. The inverse color transform processor 63011 performs operations and / or inverse transform coding that are the same as or similar to those of Figure 11 the inverse color transform unit 11010 of Figure 13 or the color inverse transform processor 13010 of
[0575] The post - processor 61005 can reconstruct the point cloud data by matching the position restored and output by the geometry decoder 61003 with the attribute restored and output by the attribute decoder 61004. Additionally, when the reconstructed point cloud data is in units of tiles and / or slices, the post - processor 61005 can perform the inverse process of the spatial segmentation on the transmission side based on signaling information. For example, when as Figure 16 shown in (b) of Figure 16 and Figure 16When the bounding box shown in (a) is divided into tiles and slices, the tiles and / or slices can be combined based on signaling information to restore the bounding box as shown in Figure 16 (a).
[0576] Figure 32 An example of a bitstream structure of point cloud data for transmission / reception according to an embodiment is illustrated.
[0577] Relevant information can be signaled to add / perform the above embodiments. Signaling information according to an embodiment can be used in a point cloud video encoder at a transmitting end or a point cloud video decoder at a receiving end.
[0578] A point cloud video encoder according to an embodiment can generate a bitstream as illustrated in Figure 32 by encoding geometric information and attribute information as described above. In addition, signaling information regarding point cloud data can be generated and processed by at least one of a geometric encoder, an attribute encoder, or a signaling processor of the point cloud video encoder and can be included in the bitstream.
[0579] Signaling information according to an embodiment can be received / obtained by at least one of a geometric decoder, an attribute decoder, and a signaling processor of the point cloud video decoder.
[0580] A bitstream according to an embodiment can be divided into a geometric bitstream, an attribute bitstream, and a signaling bitstream to be transmitted / received, or can be combined into one bitstream and transmitted / received.
[0581] When a geometric bitstream, an attribute bitstream, and a signaling bitstream according to an embodiment are configured as one bitstream, the bitstream can include one or more sub-bitstreams. A bitstream according to an embodiment can include a sequence parameter set (SPS) for sequence-level signaling, a geometric parameter set (GPS) for signaling geometric information encoding, one or more attribute parameter sets (APS) (APS 0 、APS 1 ) for signaling attribute information encoding, a tile parameter set (TPS) for tile-level signaling, and one or more slices (slice 0 to slice n). That is, a bitstream of point cloud data according to an embodiment can include one or more tiles, and each tile can be a set of slices including one or more slices (slice 0 to slice n). A TPS according to an embodiment can contain information about each of one or more tiles (e.g., coordinate value information and height / size information about a bounding box). Each slice can include a geometric bitstream (Geom0) and one or more attribute bitstreams (Attr0 and Attr1). For example, the first slice (slice 0) can include a geometric bitstream (Geom00 ) and one or more attribute bitstreams (Attr 0 , Attr1 0 ).
[0582] The geometric bitstream (or geometric slice) in each slice may be composed of a geometric slice header (geom_slice_header) and geometric slice data (geom_slice_data). According to an embodiment, the geom_slice_header may include identification information (geom_parameter_set_id) of a parameter set included in the GPS, a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id), as well as information about the data contained in the geometric slice data (geom_slice_data) (geomBoxOrigin, geom_box_log2_scale, geom_max_node_size_log2, geom_num_points). geomBoxOrigin is geometric box origin information indicating the origin of the box of the geometric slice data, geom_box_log2_scale is information indicating the logarithmic scale (log scale) of the geometric slice data, geom_max_node_size_log2 is information indicating the size of the root geometric octree node, and geom_num_points is information related to the number of points in the geometric slice data. According to an embodiment, the geom_slice_data may include geometric information (or geometric data) about the point cloud data in the corresponding slice.
[0583] Each attribute bitstream (or attribute slice) in each slice may be composed of an attribute slice header (attr_slice_header) and attribute slice data (attr_slice_data). According to an embodiment, the attr_slice_header may include information about the corresponding attribute slice data. The attribute slice data may contain attribute information (or attribute data or attribute values) about the point cloud data in the corresponding slice. When there are multiple attribute bitstreams in a slice, each bitstream may contain different attribute information. For example, one attribute bitstream may contain attribute information corresponding to color, while another attribute stream may contain attribute information corresponding to reflectivity.
[0584] Figure 33 An exemplary bitstream structure of point cloud data according to an embodiment is shown.
[0585] Figure 34 The connection relationship between components in the bitstream of point cloud data according to an embodiment is illustrated.
[0586] Figure 33 and Figure 34 the bitstream structure of the point cloud data illustrated in Figure 32 can represent the bitstream structure of the point cloud data shown in
[0587] According to an embodiment, the SPS may include an identifier (seq_parameter_set_id) for identifying the SPS, and the GPS may include an identifier (geom_parameter_set_id) for identifying the GPS and an identifier (seq_parameter_set_id) indicating the active SPS to which the GPS belongs. The APS may include an identifier (attr_parameter_set_id) for identifying the APS and an identifier (seq_parameter_set_id) indicating the active SPS to which the APS belongs.
[0588] According to an embodiment, the geometry bitstream (or geometry slice) may include a geometry slice header and geometry slice data. The geometry slice header may include an identifier (geom_parameter_set_id) of the active GPS referenced by the corresponding geometry slice. In addition, the geometry slice header may further include an identifier (geom_slice_id) for identifying the corresponding geometry slice and / or an identifier (geom_tile_id) for identifying the corresponding tile. The geometry slice data may include geometry information belonging to the corresponding slice.
[0589] According to an embodiment, the attribute bitstream (or attribute slice) may include an attribute slice header and attribute slice data. The attribute slice header may include an identifier (attr_parameter_set_id) of the active APS to be referenced by the corresponding attribute slice and an identifier (geom_slice_id) for identifying the geometry slice related to the attribute slice. The attribute slice data may include attribute information belonging to the corresponding slice.
[0590] That is, the geometry slice references the GPS, and the GPS references the SPS. Additionally, the SPS lists the available attributes, assigns identifiers to each attribute, and identifies the decoding method. According to the identifiers, the attribute slices are mapped to the output attributes. The attribute slices depend on the previous (decoded) geometry slices and the APS. The APS references the SPS.
[0591] According to an embodiment, parameters necessary for encoding the point cloud data may be newly defined in the parameter set of the point cloud data and / or the corresponding slice header. For example, when performing the encoding of the attribute information, parameters may be added to the APS. When performing tile-based encoding, parameters may be added to the tile and / or the slice header.
[0592] AsFigure 32 , Figure 33 and Figure 34 As shown in Figure 34 , the present disclosure provides tiles or slices such that point cloud data can be segmented and processed by region. According to an embodiment, corresponding regions of the bitstream may have different importance. Thus, when the point cloud data is segmented into tiles, different filters (encoding methods) and different filter units can be applied to each tile. When the point cloud data is segmented into slices, different filters and different filter units can be applied to each slice.
[0593] When the point cloud data is segmented and compressed, a transmitting device and a receiving device according to an embodiment can transmit and receive a bitstream in a high-level syntax structure to selectively transmit attribute information in the segmented regions.
[0594] A transmitting device according to an embodiment can transmit point cloud data according to a bitstream structure as shown in Figure 32 , Figure 33 and Figure 34 . Thus, a method of applying different encoding operations to important regions and using a good-quality encoding method can be provided. Additionally, efficient encoding and transmission can be supported according to the characteristics of the point cloud data, and attribute values can be provided according to user requirements.
[0595] A receiving device according to an embodiment can receive point cloud data according to a bitstream structure as shown in Figure 32 , Figure 33 and Figure 34 . Thus, different filtering (decoding) methods can be applied to corresponding regions (regions segmented into tiles or slices), rather than applying a complex decoding (filtering) method to the entire point cloud data. Therefore, better image quality in important regions is provided to the user, and appropriate latency of the system can be ensured.
[0596] As described above, tiles or slices are provided to process point cloud data by segmenting the point cloud data by region. When segmenting the point cloud data by region, an option of generating different sets of neighboring points for each region can be set. Thereby, a selection method with low complexity and slightly lower reliability or a selection method with high complexity and high reliability can be provided.
[0597] Option information related to neighboring point selection of a sequence required in the processing of encoding / decoding attribute information according to an embodiment can be signaled in the SPS and / or APS.
[0598] According to an embodiment, if there are tiles or slices with different attribute characteristics in the same sequence, option information related to neighboring point selection of the sequence can be signaled in the TPS of each slice and / or the attribute slice header.
[0599] According to an embodiment, when point cloud data is divided by region, the attribute characteristics of a specific region may be different from those of the sequence, and thus different configurations can be made through different maximum neighboring point range configuration functions.
[0600] Therefore, when point cloud data is divided into tiles, different maximum neighboring point ranges can be applied to each tile. Additionally, when point cloud data is divided into slices, different maximum neighboring point ranges can be applied to each slice.
[0601] According to an embodiment, at least one of SPS, APS, TPS, and the attribute slice header of each slice may include option information related to neighboring point selection. According to an embodiment, the option information related to neighboring point selection may include information related to the maximum neighboring point range.
[0602] According to an embodiment, the option information related to neighboring point selection may include information about the basic neighboring point distance calculation method (e.g., nn_base_distance_calculation_method_type), basic neighboring point distance information (e.g., nn_base_distance), maximum neighboring point range information (e.g., nearest_neighbore_max_range), minimum neighboring point range information (e.g., nearest_neighbor_min_range), information about the method of applying the maximum neighboring point distance (e.g., nn_range_filtering_location_type), information about whether to automatically calculate the maximum neighboring point range (e.g., automatic_nn_range_calculation_flag), information about the method of calculating the maximum neighboring point range (e.g., automatic_nn_range_method_type), and information related to the Kn value (e.g., automatic_max_nn_range_in_table, automatic_nn_range_table_k), among others.
[0603] According to an embodiment, the option information related to neighboring point selection may further include information about the maximum number of points that can be set as neighboring points (e.g., lifting_num_pred_nearest_neighbours), search range related information (e.g., lifting_search_range), and / or LOD configuration method.
[0604] A field as a term used in the syntax of the present disclosure described later may have the same meaning as a parameter or an element.
[0605] Figure 35 An embodiment of the syntax structure of a Sequence Parameter Set (SPS) (seq_parameter_set_rbsp()) according to the present disclosure is shown. The SPS may include sequence information about the point cloud data bitstream. In particular, in this example, the SPS includes neighboring point selection related option information.
[0606] The SPS according to an embodiment may include a profile_idc field, a profile_compatibility_flags field, a level_idc field, a sps_bound_box_present_flag field, a sps_source_scale_factor field, a sps_seq_parameter_set_id field, a sps_num_attribute_sets field, and a sps_extension_present_flag field.
[0607] The profile_idc field indicates the profile that the bitstream conforms to.
[0608] The profile_compatibility_flags field equal to 1 may indicate that the bitstream conforms to the profile indicated by profile_idc.
[0609] The level_idc field indicates the level that the bitstream conforms to.
[0610] The sps_bounding_box_present_flag field indicates whether source bounding box information is signaled in the SPS. The source bounding box information may include offset and size information about the source bounding box. For example, the sps_bounding_box_present_flag field equal to 1 indicates that source bounding box information is signaled in the SPS. The sps_bounding_box_present_flag field equal to 0 indicates that source bounding box information is not signaled. The sps_source_scale_factor field indicates the scaling factor of the source point cloud.
[0611] The sps_seq_parameter_set_id field provides an identifier for the SPS for reference by other syntax elements.
[0612] The sps_num_attribute_sets field indicates the number of encoded attributes in the bitstream.
[0613] The sps_extension_present_flag field specifies whether the sps_extension_data syntax structure exists in the SPS syntax structure. For example, an sps_extension_present_flag field equal to 1 specifies that the sps_extension_data syntax structure exists in the SPS syntax structure. An sps_extension_present_flag field equal to 0 specifies that the syntax structure does not exist. When it does not exist, it is inferred that the value of the sps_extension_present_flag field is equal to 0.
[0614] When the sps_bounding_box_present_flag field is equal to 1, the SPS according to the embodiment may further include an sps_bounding_box_offset_x field, an sps_bounding_box_offset_y field, an sps_bounding_box_offset_z field, an sps_bounding_box_scale_factor field, an sps_bounding_box_size_width field, an sps_bounding_box_size_height field, and an sps_bounding_box_size_depth field.
[0615] The sps_bounding_box_offset_x field indicates the x offset of the source bounding box in Cartesian coordinates. When the x offset of the source bounding box does not exist, the value of the sps_bounding_box_offset_x is 0.
[0616] The sps_bounding_box_offset_y field indicates the y offset of the source bounding box in Cartesian coordinates. When the y offset of the source bounding box does not exist, the value of the sps_bounding_box_offset_y is 0.
[0617] The sps_bounding_box_offset_z field indicates the z offset of the source bounding box in Cartesian coordinates. When the z offset of the source bounding box does not exist, the value of the sps_bounding_box_offset_z is 0.
[0618] The sps_bounding_box_scale_factor field indicates the scaling factor of the source bounding box in Cartesian coordinates. When the scaling factor of the source bounding box does not exist, the value of the sps_bounding_box_scale_factor can be 1.
[0619] The sps_bounding_box_size_width field indicates the width of the source bounding box in Cartesian coordinates. When the width of the source bounding box does not exist, the value of the sps_bounding_box_size_width field can be 1.
[0620] The sps_bounding_box_size_height field indicates the height of the source bounding box in Cartesian coordinates. When the height of the source bounding box does not exist, the value of the sps_bounding_box_size_height field can be 1.
[0621] The sps_bounding_box_size_depth field indicates the depth of the source bounding box in Cartesian coordinates. When the depth of the source bounding box does not exist, the value of the sps_bounding_box_size_depth field can be 1.
[0622] The SPS according to an embodiment includes an iterative statement repeated as many times as the value of the sps_num_attribute_sets field. In an embodiment, i is initialized to 0 and incremented by 1 each time the iterative statement is executed. The iterative statement is repeated until the value of i becomes equal to the value of the sps_num_attribute_sets field. The iterative statement may include the attribute_dimension[i] field, the attribute_instance_id[i] field, the attribute_bitdepth[i] field, the attribute_cicp_colour_primaries[i] field, the attribute_cicp_transfer_characteristics[i] field, the attribute_cicp_matrix_coeffs[i] field, the attribute_cicp_video_full_range_flag[i] field, and the known_attribute_label_flag[i] field.
[0623] The attribute_dimension[i] field specifies the number of components of the i-th attribute.
[0624] The attribute_instance_id[i] field specifies the instance ID of the i-th attribute.
[0625] The attribute_bitdepth[i] field specifies the bit depth of the i-th attribute signal.
[0626] The attribute_cicp_colour_primaries[i] field indicates the chromaticity coordinates of the colour attribute source primaries for the i-th attribute.
[0627] The attribute_cicp_transfer_characteristics[i] field either indicates the reference photoelectric transfer characteristic function of the colour attribute as a function of the source input linear optical intensity with a nominal true value range of 0 to 1, or indicates the inverse of the reference electro-optical transfer characteristic function as a function of the output linear light intensity.
[0628] The attribute_cicp_matrix_coeffs[i] field describes the matrix coefficients used when deriving the luminance and chrominance signals from the green, blue, and red or Y, Z, and X primaries.
[0629] The attribute_cicp_video_full_range_flag[i] field indicates the black level and range of the luminance and chrominance signals derived from the E′Y, E′PB, and E′PR or E′R, E′G, and E′B true value component signals.
[0630] The known_attribute_label_flag[i] field specifies whether the known_attribute_label field or the attribute_label_four_bytes field is signaled for the i-th attribute. For example, a value of the known_attribute_label_flag[i] field equal to 0 specifies that the known_attribute_label field is signaled for the i-th attribute. A value of the known_attribute_label_flag[i] field equal to 1 specifies that the attribute_label_four_bytes field is signaled for the i-th attribute.
[0631] The known_attribute_label[i] field can specify the attribute type. For example, a known_attribute_label[i] field equal to 0 can specify that the i-th attribute is colour. A known_attribute_label[i] field equal to 1 specifies that the i-th attribute is reflectance. A known_attribute_label[i] field equal to 2 can specify that the i-th attribute is a frame index.
[0632] The attribute_label_four_bytes field indicates a known attribute type with a 4-byte code.
[0633] In this example, the attribute_label_four_bytes field indicates color when equal to 0 and reflectance when equal to 1.
[0634] According to an embodiment, when the sps_extension_present_flag field is equal to 1, the SPS may further include a sps_extension_data_flag field.
[0635] The sps_extension_data_flag field may have any value.
[0636] According to an embodiment, the SPS may further include the following neighboring point selection related option information.
[0637] According to an embodiment, the neighboring point selection related option information may be included in iteration statements having the same number of iterations as the value of the above sps_num_attribute_sets field.
[0638] That is, the iteration statements may further include the nn_base_distance_calculation_method_type[i] field, the nearest_neighbour_max_range[i] field, the nearest_neighbour_min_range[i] field, the nn_range_filtering_location_type[i] field, and the automatic_nnn_range_calculation_flag[i] field.
[0639] The nn_base_distance_calculation_method_type[i] field may indicate the base nearest neighbor distance calculation method applied when compressing the i-th attribute of the corresponding sequence. For example, an nn_base_distance_calculation_method_type[i] field equal to 0 may indicate the use of the base distance. An nn_base_distance_calculation_method_type[i] field equal to 1 may indicate an octree-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type[i] field equal to 2 may indicate a distance-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type[i] field equal to 3 may indicate a sampling-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type[i] field equal to 4 may indicate a maximum nearest neighbor distance calculation by calculating the average difference of Morton codes between LODs. An nn_base_distance_calculation_method_type[i] field equal to 5 may indicate a maximum nearest neighbor distance calculation by calculating the average distance difference between LODs.
[0640] According to an embodiment, when the value of the nn_base_distance_calculation_method_type[i] field is 0, that is, when this field indicates the use of the input base distance, the iterative statement may further include the nn_base_distance[i] field.
[0641] The nn_base_distance[i] field may indicate the base nearest neighbor distance applied when compressing the i-th attribute of the corresponding sequence.
[0642] The nearest_neighbour_max_range[i] field may indicate the maximum nearest neighbor range applied when compressing the i-th attribute of the sequence. According to an embodiment, the value of the nearest_neighbour_max_range[i] field may be used as the value of NN_range in Equation 5. According to an embodiment, the nearest_neighbour_max_range[i] field may be used to limit the distance of the points registered as nearest neighbors. For example, when the LOD is generated based on an octree, the value of the nearest_neighbour_max_range[i] field may be the number of octree nodes around the point.
[0643] The nearest_neighbour_min_range[i] field may indicate the minimum nearest neighbour range applied when compressing the i-th attribute of the compressed sequence.
[0644] The nn_range_filtering_location_type[i] field may indicate the method of applying the maximum nearest neighbour distance when compressing the i-th attribute of the compressed sequence. For example, when the value of the nn_range_filtering_location_type[i] field is 0, it may indicate a method of calculating the distance between points when selecting nearest neighbours and not selecting points outside the maximum nearest neighbour distance. As another example, when the value of the nn_range_filtering_location_type[i] field is 1, it may indicate a method of registering all nearest neighbours by calculating the distance between points and removing points that exceed the maximum nearest neighbour distance among the registered points.
[0645] The automatic_nn_range_calculation_flag[i] field may indicate whether to automatically calculate the maximum nearest neighbour range when compressing the i-th attribute of the compressed sequence. For example, when the value of the automatic_nn_range_calculation_flag[i] field is 0 (i.e., TRUE), it may indicate the automatic calculation of the maximum nearest neighbour range.
[0646] According to an embodiment, when the value of the automatic_nn_range_calculation_flag[i] field is 0, that is, when this field indicates the automatic calculation of the maximum nearest neighbour range, the iterative statement may further include the automatic_nn_range_method_type[i] field and the automatic_max_nn_range_in_table[i] field.
[0647] The automatic_nn_range_method_type[i] field may indicate the method of calculating the maximum nearest neighbour range when compressing the i-th attribute of the compressed sequence.
[0648] For example, the automatic_nn_range_method_type[i] field equal to 1 can indicate a method of estimating density using the diagonal length of the bounding box as the denominator. The automatic_nn_range_method_type[i] field equal to 2 can indicate a method of estimating density using the diagonal length of the bounding box as the numerator. The automatic_nn_range_method_type[i] field equal to 3 can indicate a method of estimating density based on the maximum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 4 can indicate a method of estimating density based on the minimum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 5 can indicate a method of estimating density based on the intermediate axis length and the number of points. As another example, the automatic_nn_range_method_type[i] field equal to 6 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the maximum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 7 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the minimum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 8 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the intermediate axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 9 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the numerator with the method of estimating density based on the maximum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 10 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the numerator with the method of estimating density based on the minimum axis length of the bounding box and the number of points. The automatic_nn_range_method_type[i] field equal to 11 can indicate the method of estimating density using the diagonal length of the bounding box as the numerator and the method of estimating density based on the intermediate axis length of the bounding box and the number of points.
[0649] The automatic_max_nn_range_in_table[i] field may indicate the number of entries in the table of Kn values when compressing the i-th attribute of the compression sequence.
[0650] The SPS according to an embodiment may include an iterative statement with the same number of iterations as the value of the automatic_max_nn_range_in_table[i] field. In this case, according to the embodiment, j is initialized to 0 and incremented by 1 each time the iterative statement is executed. The iterative statement is iterated until the value of j becomes the value of the automatic_max_nn_range_in_table[i] field. The iterative statement may include the automatic_nn_range_table_k[i][j] field.
[0651] The automatic_nn_range_table_k[i][j] field may indicate the Kn value when compressing the i-th attribute of the corresponding sequence.
[0652] Figure 36 An embodiment of the syntax structure of a geometric parameter set (GPS) (seq_parameter_set_rbsp()) according to the present disclosure is shown. The GPS according to an embodiment may contain information about a method for encoding geometric information in point cloud data included in one or more slices.
[0653] According to an embodiment, the GPS may include a gps_geom_parameter_set_id field, a gps_seq_parameter_set_id field, a gps_box_present_flag field, a unique_geometry_points_flag field, a neighbour_context_restriction_flag field, an inferred_direct_coding_mode_enabled_flag field, a bitwise_occupancy_coding_flag field, an adjacent_child_contextualization_enabled_flag field, a log2_neighbour_avail_boundary field, a log2_intra_pred_max_node_size field, a log2_trisoup_node_size field, and a gps_extension_present_flag field.
[0654] The gps_geom_parameter_set_id field provides an identifier for the GPS for reference by other syntax elements.
[0655] The gps_seq_parameter_set_id field specifies the value of sps_seq_parameter_set_id for the active SPS.
[0656] The gps_box_present_flag field specifies whether additional bounding box information is provided in the geometry slice header that references the current GPS. For example, a gps_box_present_flag field equal to 1 can specify that additional bounding box information is provided in the geometry header that references the current GPS. Thus, when the gps_box_present_flag field is equal to 1, the GPS can also include the gps_gsh_box_log2_scale_present_flag field.
[0657] The gps_gsh_box_log2_scale_present_flag field specifies whether the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. For example, a gps_gsh_box_log2_scale_present_flag field equal to 1 can specify that the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. As another example, a gps_gsh_box_log2_scale_present_flag field equal to 0 can specify that the gps_gsh_box_log2_scale field is not signaled in each geometry slice header and that the common scaling for all slices is signaled in the gps_gsh_box_log2_scale field of the current GPS.
[0658] When the gps_gsh_box_log2_scale_present_flag field is equal to 0, the GPS can also include the gps_gsh_box_log2_scale field.
[0659] The gps_gsh_box_log2_scale field indicates the common scaling factor of the bounding box origin for all slices that reference the current GPS.
[0660] The unique_geometry_points_flag field indicates whether all output points have unique positions. For example, a unique_geometry_points_flag field equal to 1 indicates that all output points have unique positions. A unique_geometry_points_flag field equal to 0 indicates that two or more of the output points can have the same position in all slices referencing the current GPS.
[0661] The neighbor_context_restriction_flag field indicates the context used for octree occupancy coding. For example, a neighbor_context_restriction_flag field equal to 0 indicates that octree occupancy coding uses the context determined from six neighboring parent nodes. A neighbor_context_restriction_flag field equal to 1 indicates that octree occupancy coding uses only the context determined from sibling nodes.
[0662] The inferred_direct_coding_mode_enabled_flag field indicates whether there is a direct_mode_flag field in the geometry node syntax. For example, an inferred_direct_coding_mode_enabled_flag field equal to 1 indicates that there can be a direct_mode_flag field in the geometry node syntax. For example, an inferred_direct_coding_mode_enabled_flag field equal to 0 indicates that there is no direct_mode_flag field in the geometry node syntax.
[0663] The bitwise_occupancy_coding_flag field indicates whether bitwise contextualization of the syntax element occupancy map is used to code geometry node occupancy. For example, a bitwise_occupancy_coding_flag field equal to 1 indicates that bitwise contextualization of the syntax element occupancy_map is used to code geometry node occupancy. For example, a bitwise_occupancy_coding_flag field equal to 0 indicates that the dictionary-coded syntax element occupancy_byte is used to code geometry node occupancy.
[0664] The adjacent_child_contextualization_enabled_flag field indicates whether the adjacent children of an adjacent octree node are used for bitwise occupancy contextualization. For example, an adjacent_child_contextualization_enabled_flag field equal to 1 indicates that the adjacent children of an adjacent octree node are used for bitwise occupancy contextualization. For example, an adjacent_child_contextualization_enabled_flag equal to 0 indicates that the children of an adjacent octree node are not used for occupancy contextualization.
[0665] The log2_neighbour_avail_boundary field specifies the value of the variable NeighbAvailBoundary used in the decoding process as follows:
[0666] NeighbAvailBoundary = 2 log2 _ neighbour _ avail _ boundary .
[0667] For example, when the neighbour_context_restriction_flag field is equal to 1, NeighbAvailabilityMask can be set to be equal to 1. For example, when the neighbour_context_restriction_flag field is equal to 0, NeighbAvailabilityMask can be set to be equal to 1 << log2_neighbour_avail_boundary.
[0668] The log2_intra_pred_max_node_size field specifies the octree node size eligible for intra prediction occupancy.
[0669] The log2_trisoup_node_size field specifies the variable TrisoupNodeSize as the size of the triangular node as follows:
[0670] TrisoupNodeSize = 1 << log2_trisoup_node_size.
[0671] The gps_extension_present_flag field specifies whether the gps_extension_data syntax structure exists in the GPS syntax structure. For example, a gps_extension_present_flag equal to 1 specifies that the gps_extension_data syntax structure exists in the GPS syntax. For example, a gps_extension_present_flag equal to 0 specifies that the syntax structure does not exist in the GPS syntax.
[0672] When the value of the gps_extension_present_flag field is equal to 1, the GPS according to the embodiment may further include a gps_extension_data_flag field.
[0673] The gps_extension_data_flag field may have any value. Its presence and value do not affect the consistency of the decoder with the profile.
[0674] Figure 37 An embodiment of the syntax structure of an attribute parameter set (APS) (attribute_parameter_set()) according to the present disclosure is shown. The APS according to the embodiment may contain information about a method for encoding attribute information in point cloud data included in one or more slices. According to the embodiment, the APS may include adjacent point selection related option information.
[0675] The APS according to the embodiment may include an aps_attr_parameter_set_id field, an aps_seq_parameter_set_id field, an attr_coding_type field, an aps_attr_initial_qp field, an aps_attr_chroma_qp_offset field, an aps_slice_qp_delta_present_flag field, and an aps_flag field.
[0676] The aps_attr_parameter_set_id field provides an identifier for the APS for reference by other syntax elements.
[0677] The aps_seq_parameter_set_id field indicates the value of sps_seq_parameter_set_id of the active SPS.
[0678] The attr_coding_type field indicates the coding type of the attribute.
[0679] In this example, the attr_coding_type field equal to 0 indicates predictive weight lifting (or LOD with predictive transform) as the coding type. The attr_coding_type field equal to 1 indicates RAHT as the coding type. The attr_coding_type field equal to 2 indicates fixed weight lifting (or LOD with lifting transform).
[0680] The aps_attr_initial_qp field indicates the initial value of the variable SliceQp for each slice of the reference APS. When decoding a non-zero value of slice_qp_delta_luma or slice_qp_delta_luma, the initial value of SliceQp is modified at the attribute slice segment layer.
[0681] The aps_attr_chroma_qp_offset field indicates the offset to the initial quantization parameter signaled by the syntax aps_attr_initial_qp.
[0682] The aps_slice_qp_delta_present_flag field indicates whether the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements are present in the attribute slice header (ASH). For example, the aps_slice_qp_delta_present_flag field equal to 1 indicates the presence of the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements in the ASH. For example, the aps_slice_qp_delta_present_flag field equal to 0 indicates the absence of the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements in the ASH.
[0683] When the value of the attr_coding_type field is 0 or 2, i.e., when the coding type is predictive weight lifting (or LOD with predictive transform) or fixed weight lifting (or LOD with lifting transform), the APS according to the embodiment may further include the lifting_num_pred_nearest_neighbours field, the lifting_max_num_direct_predictors field, the lifting_search_range field, the lifting_lod_regular_sampling_enabled_flag field, the lifting_num_detail_levels_minus1 field.
[0684] The lifting_num_pred_nearest_neighbours field indicates the maximum number of nearest neighbor points (i.e., X values) to be used for prediction.
[0685] The lifting_max_num_direct_predictors field indicates the maximum number of predictors to be used for direct prediction. The value of the variable MaxNumPredictors used in the decoding process is as follows:
[0686] MaxNumPredictors = lifting_max_num_direct_predictors field + 1
[0687] The lifting_search_range field indicates the search range used to determine the nearest neighbor points.
[0688] The lifting_num_detail_levels_minus1 field indicates the number of detail levels of the attribute coding.
[0689] The lifting_lod_regular_sampling_enabled_flag field indicates whether to construct the level of detail (LOD) by using a regular sampling strategy. For example, a lifting_lod_regular_sampling_enabled_flag equal to 1 indicates that the level of detail (LOD) is constructed by using a regular sampling strategy. A lifting_lod_regular_sampling_enabled_flag equal to 0 indicates that an alternative distance-based sampling strategy is used.
[0690] The APS according to the embodiment includes an iterative statement that repeats as many times as the value of the lifting_num_detail_levels_minus1 field. In the embodiment, an index (idx) is initialized to 0 and incremented by 1 each time the iterative statement is executed, and the iterative statement is repeated until the index (idx) is greater than the value of the lifting_num_detail_levels_minus1 field. When the value of the lifting_lod_decimation_enabled_flag field is true (e.g., 1), the iterative statement may include the lifting_sampling_period[idx] field, and when the value of the lifting_lod_decimation_enabled_flag field is false (e.g., 0), the iterative statement may include the lifting_sampling_distance_squared[idx] field.
[0691] The lifting_sampling_period[idx] field indicates the sampling period of the level of detail idx.
[0692] The lifting_sampling_distance_squared[idx] field indicates the square of the sampling distance of the level of detail idx.
[0693] When the value of the attr_coding_type field is 0, i.e., when the coding type is prediction weight lifting (or LOD with prediction transform), the APS according to the embodiment may further include the lifting_adaptive_prediction_threshold field and the lifting_intra_lod_prediction_num_layers field.
[0694] The lifting_adaptive_prediction_threshold field indicates the threshold for enabling adaptive prediction.
[0695] The lifting_intra_lod_prediction_num_layers field indicates the number of LOD levels for which decoded points in the same LOD level can be referred to generate the prediction value of the target point. For example, the lifting_intra_lod_prediction_num_layers field equal to num_detail_levels_minus1 plus 1 indicates that for all LOD levels, the target point can refer to the decoded points in the same LOD level. For example, the lifting_intra_lod_prediction_num_layers field equal to 0 indicates that for any LoD level, the target point cannot refer to the decoded points in the same LOD level.
[0696] The aps_extension_present_flag field indicates whether the aps_extension_data syntax structure exists in the APS syntax structure. For example, the aps_extension_present_flag field equal to 1 indicates that the aps_extension_data syntax structure exists in the APS syntax structure. For example, the aps_extension_present_flag field equal to 0 indicates that the syntax structure does not exist in the APS syntax structure.
[0697] When the value of the aps_extension_present_flag field is 1, the APS according to the embodiment may further include an aps_extension_data_flag field.
[0698] The aps_extension_data_flag field can have any value. Its presence and value do not affect the consistency of the decoder with the profile.
[0699] The APS according to the embodiment may further include the following adjacent point selection related option information.
[0700] When the value of the attr_coding_type field is 0 or 2, that is, when the coding type is "prediction weight lifting (or LOD with prediction transform)" or "fixed weight lifting (or LOD with lifting transform)", the APS according to the embodiment may further include a different_nn_range_in_tile_flag field and a different_nn_range_per_lod_flag field.
[0701] The different_nn_range_in_tile_flag field may indicate whether the sequence is using different maximum / minimum neighboring point ranges for tiles segmented from the sequence.
[0702] The different_nn_range_per_lod_flag field may indicate whether different maximum / minimum neighboring point ranges are used for each LOD.
[0703] For example, when the value of the different_nn_range_per_lod_flag field is false (FALSE), the APS may also include a nearest_neighbour_max_range field, a nearest_neighbour_min_range field, a nn_base_distance_calculation_method_type field, a nn_range_filtering_location_type field, and an automatic_nn_range_calculation_flag field.
[0704] The nearest_neighbour_max_range field may indicate the maximum neighboring point range when compressing the corresponding attribute. According to an embodiment, the nearest_neighbour_max_range field may be used to limit the distance of points registered as neighboring points. According to an embodiment, the value of the nearest_neighbour_max_range field may be used as the value of NN_range in Equation 5.
[0705] The nearest_neighbour_min_range field may indicate the minimum neighboring point range when compressing the attribute.
[0706] The nn_base_distance_calculation_method_type field may indicate the basic neighboring point distance calculation method applied when compressing the attribute.
[0707] For example, an nn_base_distance_calculation_method_type field equal to 0 may indicate the use of a base distance. An nn_base_distance_calculation_method_type field equal to 1 may indicate an octree-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type field equal to 2 may indicate a distance-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type field equal to 3 may indicate a sampling-based maximum nearest neighbor distance calculation. An nn_base_distance_calculation_method_type field equal to 4 may indicate a maximum nearest neighbor distance calculation by computing the average difference of Morton codes between LODs. An nn_base_distance_calculation_method_type field equal to 5 may indicate a maximum nearest neighbor distance calculation by computing the average distance difference between LODs.
[0708] According to an embodiment, when the value of the nn_base_distance_calculation_method_type field is 0, i.e., when the field indicates the use of an input base distance, the iterative statement may further include an nn_base_distance field.
[0709] The nn_base_distance field may indicate the base nearest neighbor distance applied when compressing attributes.
[0710] The nn_range_filtering_location_type field may indicate a method of applying a maximum nearest neighbor distance when compressing a corresponding attribute. For example, when the value of the nn_range_filtering_location_type field is 0, it may indicate a method of calculating the distance between points when selecting nearest neighbors and not selecting points outside the maximum nearest neighbor distance. As another example, when the value of the nn_range_filtering_location_type field is 1, it may indicate a method of registering all nearest neighbors by calculating the distance between points and removing points beyond the maximum nearest neighbor distance among the registered points.
[0711] The automatic_nn_range_calculation_flag field can indicate whether to automatically calculate the maximum neighboring point range when compressing the corresponding attribute. For example, when the value of the automatic_nn_range_calculation_flag field is 0 (i.e., TRUE), it can indicate the automatic calculation of the maximum neighboring point range.
[0712] According to an embodiment, when the value of the automatic_nn_range_calculation_flag field is 0, that is, when this field indicates the automatic calculation of the maximum neighboring point range, the iterative statement may further include an automatic_nn_range_method_type field and an automatic_max_nn_range_in_table field.
[0713] The automatic_nn_range_method_type field can indicate the method for calculating the maximum neighboring point range when compressing the corresponding attribute.
[0714] For example, an automatic_nn_range_method_type field equal to 1 can indicate a method of estimating density using the diagonal length of the bounding box as the denominator. An automatic_nn_range_method_type field equal to 2 can indicate a method of estimating density using the diagonal length of the bounding box as the numerator. An automatic_nn_range_method_type field equal to 3 can indicate a method of estimating density based on the maximum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 4 can indicate a method of estimating density based on the minimum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 5 can indicate a method of estimating density based on the intermediate axis length and the number of points. As another example, an automatic_nn_range_method_type field equal to 6 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the maximum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 7 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the minimum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 8 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the denominator with the method of estimating density based on the intermediate axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 9 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the numerator with the method of estimating density based on the maximum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 10 can indicate a calculation method that combines the method of estimating density using the diagonal length of the bounding box as the numerator with the method of estimating density based on the minimum axis length of the bounding box and the number of points. An automatic_nn_range_method_type field equal to 11 can indicate a method of estimating density using the diagonal length of the bounding box as the numerator and a method of estimating density based on the intermediate axis length of the bounding box and the number of points.
[0715] The automatic_max_nn_range_in_table field can indicate the number of entries in the table of Kn values when compressing the i-th attribute of the sequence.
[0716] The APS according to the embodiment may include iteration statements with the same number of iterations as the value of the automatic_max_nn_range_in_table field. In this case, according to the embodiment, j is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of j becomes the value of the automatic_max_nn_range_in_table field. The iteration statement may include the automatic_nn_range_table_k[j] field.
[0717] The automatic_nn_range_table_k[j] field may indicate the Kn value when compressing the i-th attribute of the corresponding sequence.
[0718] According to the embodiment, when the value of the different_nn_range_per_lod_flag field is TRUE, the APS further includes iteration statements with the same number of iterations as the value of the lifting_num_detail_levels_minus1 field. In this case, according to the embodiment, the index (idx) is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the index (idx) is greater than the value of the lifting_num_detail_levels_minus1 field. The iteration statement may further include the nearest_neighbour_max_range[idx] field, the nearest_neighbour_min_range[idx] field, the nn_base_distance_calculation_method_type[idx] field, the nn_range_filtering_location_type[idx] field, and the automatic_nn_range_calculation_flag[idx] field.
[0719] The nearest_neighbour_max_range[idx] field may indicate the maximum nearest neighbour range for LOD idx. According to the embodiment, the nearest_neighbour_max_range[idx] field may be used to limit the distance of the points registered as the nearest neighbours of LOD idx. According to the embodiment, the value of the nearest_neighbour_max_range[idx] field may be used as the NN_range value of LOD idx.
[0720] The nearest_neighbour_min_range[idx] field may indicate the minimum nearest neighbour range for LOD idx.
[0721] The nn_base_distance_calculation_method_type[idx] field may indicate the basic nearest neighbour distance calculation method for LOD idx. As for the value assigned to the nn_base_distance_calculation_method_type[idx] field and the definition (or meaning) of the value, refer to the description of the nn_base_distance_calculation_method_type field.
[0722] According to an embodiment, when the value of the nn_base_distance_calculation_method_type[idx] field is 0, that is, when this field indicates the use of the input basic distance, the iterative statement may further include the nn_base_distance[idx] field.
[0723] The nn_base_distance[idx] field may indicate the basic nearest neighbour distance for LOD idx.
[0724] The nn_range_filtering_location_type[idx] field may indicate the method of applying the maximum nearest neighbour distance for LOD idx. For example, when the value of the nn_range_filtering_location_type[idx] field is 0, it may indicate a method of calculating the distance between points when selecting nearest neighbours and not selecting points outside the maximum nearest neighbour distance. As another example, when the value of the nn_range_filtering_location_type[idx] field is 1, it may indicate a method of registering all nearest neighbours by calculating the distance between points and removing points that exceed the maximum nearest neighbour distance among the registered points.
[0725] The automatic_nn_range_calculation_flag[idx] field may indicate whether to automatically calculate the maximum nearest neighbour range for LOD idx. For example, when the value of the automatic_nn_range_calculation_flag[idx] field is 0 (i.e., TRUE), it may indicate the automatic calculation of the maximum nearest neighbour range.
[0726] According to an embodiment, when the value of the automatic_nn_range_calculation_flag[i] field is 0, that is, when this field indicates the automatic calculation of the maximum neighboring point range, the iterative statement may further include the automatic_nn_range_method_type[idx] field and the automatic_max_nn_range_in_table[idx] field.
[0727] The automatic_nn_range_method_type[idx] field may indicate the method for calculating the maximum neighboring point range when compressing the corresponding attribute. Regarding the value assigned to the automatic_nn_range_method_type[idx] field and the definition (or meaning) of the value, refer to the description of the value, and refer to the description of the automatic_nn_range_method_type field.
[0728] The automatic_max_nn_range_in_table[idx] field may indicate the number of entries in the table of the Kn value for LOD idx.
[0729] The APS according to an embodiment may include...
Claims
1. A coding method, the coding method comprises the following steps: encoding geometric information including the positions of points of point cloud data; generating one or more levels of detail (LOD) based on the geometric information and selecting one or more neighboring points of each point to be attribute-encoded based on the one or more LOD, wherein the selected one or more neighboring points of each point are within a maximum neighboring point distance, and wherein the maximum neighboring point distance is obtained by multiplying a base distance by a maximum neighboring point range; encoding the attribute information of each point based on the selected one or more neighboring points of each point; and generating a bitstream including the encoded geometric information, the encoded attribute information, and signaling information, wherein the step of encoding the geometric information includes: quantizing the geometric information, and arithmetically encoding the quantized geometric information, and wherein the signaling information includes information related to the maximum neighboring point range.
2. The coding method according to claim 1, wherein the maximum neighboring point range is calculated based on the density value of the point cloud data, and wherein the density value of the point cloud data is estimated based on the distance for attribute encoding and the bounding box of the point cloud data.
3. The coding method according to claim 2, wherein the density value of the point cloud data is estimated by dividing the distance by the diagonal length of the bounding box of the point cloud data.
4. The coding method according to claim 1, wherein the number of the selected one or more neighboring points of each point is limited to the maximum number of neighboring points.
5. A decoding method, the decoding method comprises the following steps: receiving geometric information, attribute information, and signaling information; decoding the geometric information based on the signaling information, wherein the positions of points of point cloud data are included in the decoded geometric information; and decoding the attribute information of each point based on the one or more neighboring points of each point and the signaling information, wherein one or more neighboring points of each point to be attribute-decoded are selected based on one or more levels of detail (LOD), wherein the selected one or more neighboring points of each point are within a maximum neighboring point distance, wherein the maximum neighboring point distance is obtained by multiplying a base distance by a maximum neighboring point range, wherein the signaling information includes information related to the maximum neighboring point range, and wherein the step of decoding the geometric information includes: arithmetically decoding the geometric information, and performing inverse quantization on the arithmetically decoded geometric information.
6. The decoding method according to claim 5, wherein the maximum neighboring point range is calculated according to the density value of the point cloud data, wherein the density value of the point cloud data is estimated based on the distance for attribute decoding and the bounding box of the point cloud data.
7. The decoding method according to claim 6, wherein the density value of the point cloud data is estimated by dividing the distance by the diagonal length of the bounding box of the point cloud data.
8. The decoding method according to claim 5, wherein, the number of one or more neighboring points selected for each point is limited to the maximum number of neighboring points.
9. A method for transmitting a bitstream, the transmitting method comprising the steps of: encoding geometric information including the positions of points of point cloud data; generating one or more levels of detail (LOD) based on the geometric information and selecting one or more neighboring points for each point to be attribute-encoded based on the one or more LODs, wherein the one or more neighboring points selected for each point are within a maximum neighboring point distance, and wherein the maximum neighboring point distance is obtained by multiplying a base distance by a maximum neighboring point range; encoding the attribute information for each point based on the one or more neighboring points selected for each point; generating the bitstream including the encoded geometric information, the encoded attribute information, and signaling information; and transmitting the bitstream, wherein the step of encoding the geometric information comprises: quantizing the geometric information, and arithmetic-encoding the quantized geometric information, and wherein the signaling information includes information related to the maximum neighboring point range.
10. The transmitting method according to claim 9, wherein, the maximum neighboring point range is calculated based on the density value of the point cloud data, and wherein the density value of the point cloud data is estimated based on the distance for attribute encoding and the bounding box of the point cloud data.
11. The transmitting method according to claim 10, wherein, the density value of the point cloud data is estimated by dividing the distance by the diagonal length of the bounding box of the point cloud data.
12. The transmitting method according to claim 9, wherein, the number of one or more neighboring points selected for each point is limited to the maximum number of neighboring points.