Decoding attribute values in geometry-based point cloud compression

By generating a predicted value candidate set based on the position comparison of the second and third points in the G-PCC decoding device, the problem of degradation of prediction quality in the prior art is solved, and the compression effect of point cloud is improved.

CN120035843APending Publication Date: 2025-05-23QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072484.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-17
Filing Date
2023-10-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art may lead to a reduction in prediction quality in certain point cloud decoding scenarios, which in turn affects the compression effect.

Method used

The predicted value candidate list is adjusted to improve the prediction quality by generating a predicted value candidate set based on the position comparison of the second and third points in the G-PCC decoding device, rather than based solely on the closest point.

Benefits of technology

The replacement process is implemented in scenarios that are possible to improve the predictive quality, while avoiding replacement in scenarios that may reduce the predictive quality, thereby improving the compression effect of the point cloud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035843A_ABST
    Figure CN120035843A_ABST
Patent Text Reader

Abstract

An apparatus for processing point cloud data is configured to: determine a first attribute value for a first point of a point cloud, the first point being a decoded point closest to a current point of the point cloud; determining a second attribute value and a third attribute value for a second point and a third point of the point cloud, the second point and the third point being a second proximate decoded point and a third proximate decoded point; determining a fourth attribute value for a fourth point of the point cloud, the fourth point being a decoded point that is farther or the same distance from the current point than the third point; a predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value is generated based on a comparison of the location of the second point and the location of the third point.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Patent Application No. 18 / 488,816, filed on October 17, 2023, and U.S. Provisional Patent Application No. 63 / 380,222, filed on October 19, 2022, the entire contents of which are incorporated herein by reference. U.S. Patent Application No. 18 / 488,816, filed on October 17, 2023, claims the benefit of U.S. Provisional Patent Application No. 63 / 380,222, filed on October 19, 2022. Technical Field

[0003] The present disclosure relates to point cloud encoding and decoding. Background Art

[0004] A point cloud is a collection of points in a 3D space. A point may correspond to a point on an object in the 3D space. Therefore, a point cloud can be used to represent the physical content of a 3D space. Point clouds may have utility in a wide variety of situations. For example, a point cloud may be used in the context of an autonomous vehicle to represent the location of an object on the road. In another example, in order to locate virtual objects in augmented reality (AR) or mixed reality (MR) applications, a point cloud may be used in the context of representing the physical content of the environment. Point cloud compression is the process of encoding and decoding a point cloud. Encoding a point cloud can reduce the amount of data required for storage and transmission of the point cloud. Summary of the invention

[0005] In order to predict the value of the attribute, the G-PCC encoder and the G-PCC decoder can be configured to follow the same list construction process so that each generates the same prediction value candidate list. The G-PCC encoder can then signal the G-PCC decoder which candidate in the list will be used as the prediction value. The G-PCC decoding device can generate an initial list with M prediction value candidates, and the M prediction value candidates correspond to the M points closest to the current point (e.g., closest in terms of distance). For example, M can be equal to 3 or some other integer value. Based on the position of the M prediction value candidates relative to the current point and relative to each other, the candidates in the M prediction value candidates can be replaced with other candidates farther away from the current point, but due to the position, better predictions can be provided. The replacement process generally provides better predictions by generating a candidate list that is more likely to include prediction value candidates with values ​​close to the actual attribute value, and therefore provides better compression.

[0006] However, for certain specific decoding scenarios, the replacement process may result in poor prediction and, therefore, poor compression. The present disclosure describes techniques for preventing the replacement process from being called in scenarios where replacement is more likely to reduce prediction quality while still performing replacement in scenarios where replacement is likely to improve prediction quality. For example, according to the techniques of the present disclosure, a G-PCC decoding device may be configured to generate a candidate set of prediction values ​​based on a comparison of the position of a second point with the position of a third point, wherein the second point of the point cloud is the second closest decoded point to the current point of the point cloud, and the third point of the point cloud is the third closest decoded point to the current point of the point cloud. By performing replacement or preventing replacement based on a comparison of the position of the second point with the position of the third point rather than solely based on the relative position of the closest decoded point to the current point, a G-PCC decoding device configured to perform the techniques of the present disclosure can achieve better prediction, and therefore better compression.

[0007] According to an example of the present disclosure, a device for processing point cloud data includes: a memory configured to store point cloud data; and one or more processors implemented in a circuit and configured to: determine a first attribute value for a first point of the point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determine a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determine a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determine a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is either a decoded point that is third closest to the current point of the point cloud, or a decoded point that is third closest to the current point of the point cloud. Compared to a decoded point that is farther away from the current point, or compared to a decoded point that is the same distance from the current point as the third point; determining a candidate set of prediction values ​​for the attribute value of the current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine the candidate set of prediction values ​​for the current point of the point cloud, one or more processors are also configured to generate a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of the second point with a position of the third point; and decoding the attribute value of the current point based on the candidate set of prediction values.

[0008] According to an example of the present disclosure, a method for processing point cloud data includes: determining a first attribute value for a first point of the point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is a point that is farther away from the current point than the third point. an already decoded point; determining a candidate set of prediction values ​​for the attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine a candidate set of prediction values ​​for the current point of the point cloud, determining the candidate set of prediction values ​​comprises generating a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of a second point with a position of a third point; and decoding the attribute value of the current point based on the candidate set of prediction values.

[0009] A computer-readable storage medium stores instructions, which, when executed by one or more processors, cause the one or more processors to: determine a first attribute value for a first point of a point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determine a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determine a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determine a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is a decoded point that is further away from the current point than the third point; A decoded point, or a decoded point that is the same distance from the current point as the third point; determining a candidate set of prediction values ​​for the attribute value of the current point of the point cloud, wherein the current point defines the intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine the candidate set of prediction values ​​for the current point of the point cloud, one or more processors are also configured to generate a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of the second point with a position of the third point; and decoding the attribute value of the current point based on the candidate set of prediction values.

[0010] According to an example of the present disclosure, a device for processing point cloud data includes: a component for determining a first attribute value for a first point of a point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; a component for determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; a component for determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; a component for determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is a point that is farther away from the current point than the third point a decoded point; a component for determining a candidate set of prediction values ​​for an attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine a candidate set of prediction values ​​for a current point of the point cloud, the component for determining the candidate set of prediction values ​​includes a component for generating a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of a second point with a position of a third point; and a component for decoding the attribute value of the current point based on the candidate set of prediction values.

[0011] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram illustrating an example encoding and decoding system that may perform the techniques of this disclosure.

[0013] Figure 2 is a block diagram illustrating an example Geometry Point Cloud Compression (G-PCC) encoder.

[0014] Figure 3 is a block diagram illustrating an example G-PCC decoder.

[0015] Figure 4 is a conceptual diagram illustrating an example octree partitioning for geometry coding.

[0016] Figure 5 A more detailed diagram Figure 2 Block diagram of an example geometry encoding unit.

[0017] Figure 6 A more detailed diagram Figure 2 Block diagram of an example attribute encoding unit.

[0018] Figure 7 A more detailed diagram Figure 3 Block diagram of an example geometry decoding unit.

[0019] Figure 8 A more detailed diagram Figure 3 Block diagram of an example attribute decoding unit.

[0020] Fig. 9 is a conceptual diagram illustrating an example of a prediction tree.

[0021] Fig. 10A and Fig. 10B is a conceptual diagram illustrating an example of a rotating Light Detection and Ranging (LIDAR) acquisition model.

[0022] Fig.11 An example of inter-frame prediction of a current point based on a point in a reference frame is shown.

[0023] Fig.12 is a flow diagram illustrating an example decoding flow associated with a syntax element indicating whether a node is coded in an inter-prediction mode or an intra-prediction mode.

[0024] Fig.13 Additional inter-frame prediction value points are shown that are obtained from the first point having an azimuth angle greater than the inter-frame prediction value point.

[0025] Fig.14 A flow chart illustrating the motion compensation process when the reference frame is stored in the spherical domain and motion compensation is performed in the Cartesian domain is shown.

[0026] Fig.15 A reference structure for intra-frame level-of-detail (LOD) prediction is shown.

[0027] Fig.16 An example of processing neighboring points in subsequent LODs that are at the same distance as the current point is shown.

[0028] Fig.17 An example of processing neighboring points in the same LOD that are at the same distance as the current point is shown.

[0029] Fig.18 An example of actual nearest neighbors with significant jumps in Morton order is shown.

[0030] Fig.19 An example of voxel neighbors is shown.

[0031] Fig. 20 An example of spatial segmentation is shown.

[0032] Fig.21 An example of neighbor point distribution is shown.

[0033] Fig. 22 An example of generating List1 and List2 for performing a neighbor search process for attribute LOD prediction is shown.

[0034] Fig.23 The diagram illustrates the concept of relative direction.

[0035] Fig.24 is a flow diagram illustrating example operation of a G-PCC decoder in accordance with one or more techniques of this disclosure.

[0036] Fig.25 is a conceptual diagram illustrating an example ranging system that may be used with one or more techniques of this disclosure.

[0037] Fig.26 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of this disclosure may be used.

[0038] Fig. 27 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of this disclosure may be used.

[0039] Fig.28 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure may be employed. DETAILED DESCRIPTION

[0040] "Geometry-based Point Cloud Compression" (G-PCC) directly compresses 3D geometry, i.e., the locations of a set of points in 3D space. G-PCC also compresses associated attribute values. For example, the attributes can be color information (such as R / G / B, Y / Cb / Cr), reflection information, or other attributes (such as reflectivity, temperature values, humidity values, latitude coordinates, longitude coordinates, etc.). Point clouds can be captured by various cameras or sensors (such as light detection and ranging (LIDAR) scanners or 3D scanners) and can also be generated by computers. Point cloud data can be used for various applications, including but not limited to construction (e.g., modeling), graphics (e.g., 3D models for visualization and animation), and the automotive industry (e.g., LIDAR sensors used to aid navigation).

[0041] According to a technique for compressing attribute values, a decoding device (e.g., a G-PCC encoder or a G-PCC decoder) may predict an attribute value for a current point based on attribute values ​​for already decoded points. A residual value, i.e., the difference between the predicted attribute value and the actual attribute value for the current point, is then signaled in the bitstream so that the G-PCC decoder may determine the decoded attribute value as the sum of the predicted value and the residual value. If lossless compression is used, the decoded attribute value will be equal to the actual attribute value before compression. If lossy compression is used (e.g., due to quantized residual values), the decoded attribute value may be different from, but generally relatively close to, the actual attribute value before compression.

[0042] In order to determine the predicted value of the attribute, the G-PCC encoder and the G-PCC decoder can be configured to follow the same list construction process so that each generates the same predicted value candidate list. The G-PCC encoder can then signal the G-PCC decoder which candidate in the list will be used as the predicted value. The G-PCC decoding device can generate an initial list with M predicted value candidates, and the M predicted value candidates correspond to the M points closest to the current point. For example, M can be equal to 3 or some other integer value. Based on the position of the M predicted value candidates relative to the current point and relative to each other, the candidates in the M predicted value candidates can be replaced with other candidates farther away from the current point, but due to the position, better predictions can be provided. The replacement process generally provides better predictions by generating a candidate list that is more likely to include predicted value candidates with values ​​close to the actual attribute value, and therefore provides better compression.

[0043] However, for certain specific decoding scenarios, the replacement process may result in poor prediction and, therefore, poor compression. The present disclosure describes techniques for preventing the replacement process from being called in scenarios where replacement is more likely to reduce prediction quality while still performing replacement in scenarios where replacement is likely to improve prediction quality. For example, according to the techniques of the present disclosure, a G-PCC decoding device may be configured to generate a candidate set of prediction values ​​based on a comparison of the position of a second point with the position of a third point, wherein the second point of the point cloud is the second closest decoded point to the current point of the point cloud, and the third point of the point cloud is the third closest decoded point to the current point of the point cloud. By performing replacement or preventing replacement based on a comparison of the position of the second point with the position of the third point rather than solely based on the relative position of the closest decoded point to the current point, a G-PCC decoding device configured to perform the techniques of the present disclosure can achieve better prediction, and therefore better compression.

[0044] Figure 11 is a block diagram illustrating an example encoding and decoding system 100 that can perform the techniques of the present disclosure. The techniques of the present disclosure are generally directed to decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. In general, point cloud data includes any data used to process point clouds. Decoding can effectively compress and / or decompress point cloud data.

[0045] like Figure 1 As shown, system 100 includes a source device 102 and a destination device 116. Source device 102 provides encoded point cloud data to be decoded by destination device 116. Figure 1 In the example of , source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may include any of a wide variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets (such as smart phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.

[0046] exist Figure 1 In the example of , the source device 102 includes a data source 104, a memory 106, a G-PCC encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a G-PCC decoder 300, a memory 120, and a data consumer 118. According to the present disclosure, the G-PCC encoder 200 of the source device 102 and the G-PCC decoder 300 of the destination device 116 can be configured to apply the techniques related to predictive geometric decoding of the present disclosure. Therefore, the source device 102 represents an example of an encoding device, and the destination device 116 represents an example of a decoding device. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, the destination device 116 may interface with an external data consumer instead of including a data consumer in the same device.

[0047] like Figure 1The system 100 shown is only an example. In general, other digital encoding and / or decoding devices can perform the techniques related to predictive geometry decoding disclosed in the present invention. The source device 102 and the destination device 116 are only examples of such devices, wherein the source device 102 generates decoded data to be transmitted to the destination device 116. The present disclosure refers to a "decoding" device as a device that performs the decoding (encoding and / or decoding) of data. Therefore, the G-PCC encoder 200 and the G-PCC decoder 300 represent examples of decoding devices, specifically encoders and decoders, respectively. In some examples, the source device 102 and the destination device 116 can operate in a substantially symmetrical manner, so that each of the source device 102 and the destination device 116 includes encoding and decoding components. Therefore, the system 100 can support one-way or two-way transmission between the source device 102 and the destination device 116, for example, for streaming, playback, broadcasting, telephone, navigation and other applications.

[0048] In general, the data source 104 represents the source of data (i.e., the original unencoded point cloud data), and can provide a continuous series of "frames" of data to the G-PCC encoder 200, which encodes the data of the frames. The data source 104 of the source device 102 can include a point cloud capture device, such as any of a variety of cameras or sensors, for example, a 3D scanner or a light detection and ranging (LIDAR) device, one or more cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or additionally, the point cloud data can be generated by a computer based on a scanner, camera, sensor or other data. For example, the data source 104 can generate computer graphics-based data as source data, or generate a combination of live data, archived data and computer-generated data. In each case, the G-PCC encoder 200 encodes captured, pre-captured or computer-generated data. The G-PCC encoder 200 can rearrange the frames from the order in which they are received (sometimes referred to as "display order") into a decoding order for decoding. G-PCC encoder 200 may generate one or more bitstreams including the encoded data. Source device 102 may then output the encoded data onto computer-readable medium 110 via output interface 108 for receipt and / or retrieval by input interface 122 of destination device 116, for example.

[0049] The memory 106 of the source device 102 and the memory 120 of the destination device 116 can represent general purpose memory. In some examples, the memory 106 and the memory 120 can store raw data, for example, raw data from the data source 104 and raw decoded point cloud data from the G-PCC decoder 300. Additionally or alternatively, the memory 106 and the memory 120 can store software instructions that can be executed by, for example, the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although in this example, the memory 106 and the memory 120 are shown as being separated from the G-PCC encoder 200 and the G-PCC decoder 300, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 can also include internal memory for functionally similar or equivalent purposes. In addition, the memory 106 and the memory 120 can store encoded data, for example, data output from the G-PCC encoder 200 and input to the G-PCC decoder 300. In some examples, portions of memory 106 and memory 120 may be allocated as one or more buffers, for example, to store raw decoded and / or encoded data.For example, memory 106 and memory 120 may store data representing a point cloud.

[0050] The computer-readable medium 110 may represent any type of medium or device capable of transmitting the encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to transmit the encoded data directly to the destination device 116 in real time (e.g., via a radio frequency network or a computer-based network). According to a communication standard such as a wireless communication protocol, the output interface 108 can modulate a transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 102 to the destination device 116.

[0051] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.

[0052] In some examples, source device 102 may output the encoded data to file server 114 or another intermediate storage device that may store the encoded data generated by source device 102. Destination device 116 may access the stored data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing the encoded data and transmitting the encoded data to destination device 116. File server 114 may represent a network server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded data from file server 114 via any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing the encoded data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a downloading transmission protocol, or a combination thereof.

[0053] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to communicate data (such as encoded data) according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to communicate data (such as encoded data) according to other wireless standards (such as IEEE 802.11 specifications, IEEE802.15 specifications (e.g., ZigBee TM ),BluetoothTM Standards, etc.) to transfer data (such as encoded data). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-a-chip (SoC) devices. For example, source device 102 may include a SoC device to perform functions attributed to G-PCC encoder 200 and / or output interface 108, and destination device 116 may include a SoC device to perform functions attributed to G-PCC decoder 300 and / or input interface 122.

[0054] The techniques of this disclosure may be applied to encoding and decoding to support any of a variety of applications, such as communications between autonomous vehicles, communications between scanners, cameras, sensors, and processing devices (such as local or remote servers), geographic mapping, or other applications.

[0055] The input interface 122 of the destination device 116 receives the encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded bitstream may include signaling information defined by the G-PCC encoder 200, such as syntax elements with values ​​describing characteristics and / or processing of coding units (e.g., slices, pictures, groups of pictures, sequences, etc.), and the signaling information is also used by the G-PCC decoder 300. The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the location of a physical object. In some examples, the data consumer 118 may include a display to present imagery based on a point cloud.

[0056] The G-PCC encoder 200 and the G-PCC decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (digital signal processor, DSP), application specific integrated circuits (application specific integrated circuit, ASIC), field programmable gate arrays (field programmable gate array, FPGA), discrete logic, software, hardware, firmware or any combination thereof. When these techniques are partially implemented in software, the device can store the instructions of the software in a suitable non-temporary computer-readable medium, and use one or more processors to execute these instructions in hardware to perform the technology disclosed in the present invention. Each of the G-PCC encoder 200 and the G-PCC decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (encoder / decoder, CODEC) in the corresponding device. The device including the G-PCC encoder 200 and / or the G-PCC decoder 300 may include one or more integrated circuits, microprocessors and / or other types of devices.

[0057] The G-PCC encoder 200 and the G-PCC decoder 300 may operate according to a coding standard such as a video point cloud compression (V-PCC) standard or a geometric point cloud compression (G-PCC) standard. The present disclosure may generally relate to coding (e.g., encoding and decoding) of pictures to include the process of encoding or decoding data. The encoded bitstream typically includes a series of values ​​for syntax elements that represent coding decisions (e.g., coding modes).

[0058] The present disclosure may generally relate to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of values ​​of syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 may signal the values ​​of syntax elements in a bitstream. Generally speaking, signaling refers to generating values ​​in a bitstream. As described above, the source device 102 may transmit the bitstream to the destination device 116 in substantially real time or non-real time, such as may occur when the syntax elements are stored to the storage device 112 for later retrieval by the destination device 116.

[0059] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is studying the potential need for standardization of point cloud coding techniques with compression capabilities significantly exceeding current methods, with the goal of creating a standard. The group is working together on a collaborative exploration activity called the "3-Dimensional Graphics Team (3DG)" to evaluate compression technology designs proposed by their experts in this field.

[0060] Point cloud compression activities are divided into two different approaches. The first approach is "Video Point Cloud Compression" (V-PCC), which segments 3D objects and projects the segments into multiple 2D planes (represented as "patches" in 2D frames), and the segments are further decoded by traditional 2D video codecs (such as High Efficiency Video Coding (HEVC) (ITU-T H.265) codecs). The second approach is "Geometry-based Point Cloud Compression" (G-PCC), which directly compresses 3D geometry (i.e., the positions of a set of points in 3D space) and associated attribute values ​​(for each point associated with the 3D geometry). G-PCC addresses the compression of point clouds in category 1 (static point clouds) and category 3 (dynamically acquired point clouds). The latest draft of the G-PCC standard is available in "G-PCC DIS" (ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, January 2020), and a description of the codec is available in "G-PCC Codec Description v6" (ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, January 2020).

[0061] A point cloud contains a collection of points in 3D space and may have attributes associated with the points. The attributes may be color information (such as R, G, B or Y, Cb, Cr), or reflectance information, or other attributes. Point clouds may be captured by various cameras or sensors (such as LIDAR sensors and 3D scanners), and may also be computer generated. Point cloud data is used in a variety of applications, including but not limited to construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to aid navigation).

[0062] The 3D space occupied by the point cloud data can be surrounded by a virtual bounding box. The positions of the points in the bounding box can be represented by a certain precision; therefore, the positions of one or more points can be quantized based on the precision. At the smallest level, the bounding box is divided into voxels, which are the smallest spatial units represented by a unit cube. A voxel in a bounding box can be associated with zero, one, or more than one point. The bounding box can be divided into multiple cubic / cuboid regions (which can be referred to as tiles). Each tile can be decoded into one or more slices. The partitioning of the bounding box into slices and slices can be based on the number of points in each partition, or based on other considerations (for example, a particular area can be decoded as a slice). The tile area can be further partitioned using partitioning decisions similar to those in a video codec.

[0063] Figure 2 An overview of the G-PCC encoder 200 is provided. Figure 3 An overview of the G-PCC decoder 300 is provided. The modules shown are logical and do not necessarily correspond one-to-one with implemented code. Figure 2 In the example of , the G-PCC encoder 200 may include a geometry encoding unit 250 and an attribute encoding unit 260. In general, the geometry encoding unit 250 is configured to encode the positions of points in the point cloud frame to produce a geometry bitstream 203. The attribute encoding unit 260 is configured to encode the attributes of the points of the point cloud frame to produce an attribute bitstream 205. As will be explained below, the attribute encoding unit 260 may also encode the attributes using the positions as well as the encoded geometry (e.g., reconstruction) from the geometry encoding unit 250.

[0064] exist Figure 3 In the example of , the G-PCC decoder 300 may include a geometry decoding unit 350 and an attribute decoding unit 360. In general, the geometry encoding unit 350 is configured to decode the geometry bitstream 203 to recover the positions of the points in the point cloud frame. The attribute decoding unit 360 is configured to decode the attribute bitstream 205 to recover the attributes of the points of the point cloud frame. As will be explained below, the attribute decoding unit 360 can also encode attributes using the positions of the decoded geometry (e.g., reconstruction) from the geometry decoding unit 350.

[0065] In both the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are first decoded. The attribute decoding depends on the decoded geometry. Figure 5-Figure 8 In , coding units with vertical hashing are an option commonly used for category 1 data. Diagonal cross-hatching coding units are an option commonly used for category 3 data. All other modules are common between category 1 and category 3.

[0066] For category 3 data, the compressed geometry is typically represented as an octree of individual voxels from the root down to the leaf level. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to blocks larger than the leaf level of voxels) plus a model that approximates the surfaces within each leaf of the pruned octree. In this way, both category 1 and category 3 data share the octree decoding mechanism, while category 1 data can additionally approximate the voxels within each leaf with a surface model. The surface model used is a triangulation of 1-10 triangles per block to produce a triangle soup. Therefore, category 1 geometry codecs are called Trisoup geometry codecs, while category 3 geometry codecs are called Octree geometry codecs.

[0067] At each node of the octree, occupancy is signaled (when not inferred) for one or more of its children (up to eight nodes). Multiple neighborhoods are specified, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, an edge, or a vertex with the current octree node, and so on. Within each neighborhood, the occupancy of the node and / or its children can be used to predict the occupancy of the current node or its children. For points that are sparsely populated in some nodes of the octree, the codec also supports a direct decoding mode in which the 3D position of the point is directly encoded. A signaling flag can be sent to indicate the signaling direct mode. At the lowest level, the number of points associated with the octree node / leaf node can also be decoded.

[0068] Once the geometry is decoded, the attributes corresponding to the geometry points are decoded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute value representing the reconstructed point can be derived.

[0069] There are three attribute coding methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (prediction transform), and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). RAHT and lifting are typically used for category 1 data, while prediction is typically used for category 3 data. However, either method can be used for any data, and just like the geometry codec in G-PCC, the attribute coding method used to decode the point cloud is specified in the bitstream.

[0070] The decoding of attributes can be performed in levels of detail (LoD), where with each LoD, a finer representation of the point cloud attributes can be obtained. Each LoD can be specified based on a distance metric to neighboring nodes or based on a sampling distance.

[0071] At the G-PCC encoder 200, the residual obtained as an output of the coding method for the attribute is quantized. The residual can be obtained by subtracting the attribute value from a prediction derived based on points in the neighborhood of the current point and based on the attribute value of previously encoded points. The quantized residual can be decoded using context adaptive arithmetic coding.

[0072] G-PCC encoder 200 and G-PCC decoder 300 can be configured to use predictive geometry decoding as an alternative to octree geometry decoding to decode point cloud data. In predictive tree decoding, the nodes of the point cloud are arranged in a tree structure (which defines a prediction structure), and various prediction strategies are used to predict the coordinates of each node in the tree relative to its predicted value. The node as the root vertex has no predicted value. Other nodes can have 1, 2, 3 or more child nodes. Other nodes can be leaf nodes without child nodes. In an example, each node of the prediction has only one parent node.

[0073] The G-PCC encoder 200 may employ any algorithm to generate the prediction tree; the algorithm used may be determined based on the application / use case, and several strategies may be used. For each node, the residual coordinate values ​​are decoded in the bitstream starting from the root node in a depth-first manner. Predictive geometry decoding may be particularly useful for category 3 (LIDAR acquired) point cloud data (e.g., for low latency applications).

[0074] Figure 4 4 is a conceptual diagram illustrating an example octree partition for geometric decoding. At each node of the octree 400, the G-PCC encoder 200 can signal the occupancy of one or more child nodes (e.g., up to eight nodes) of the node to the G-PCC decoder 300 (when the G-PCC decoder 300 does not infer the occupancy). Specify multiple neighborhoods, including (a) nodes that share a side with the current octree node, (b) nodes that share a side, a side or a vertex with the current octree node, and so on. In each neighborhood, the occupancy of the node and / or the child nodes of the node can be used to predict the occupancy of the current node or the child nodes of the node. For points sparsely populated in certain nodes of the octree, the codec also supports a direct decoding mode in which the 3D position of the point is directly encoded. The G-PCC encoder 200 can send a signaling notification flag to indicate the direct mode of signaling notification. At the lowest level, the number of points associated with the octree node / leaf node can also be decoded.

[0075] Once the geometry is decoded, the attributes corresponding to the geometry points are decoded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute value representing the reconstructed point can be derived.

[0076] There are three attribute coding processes in G-PCC: Region Adaptive Hierarchical Transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (prediction transform), and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). RAHT and lifting are typically used for category 1 data, while prediction is typically used for category 3 data. However, either process can be used for any data, and as with the geometry codec in G-PCC, the attribute coding process used to decode point clouds is specified in the bitstream.

[0077] The decoding of attributes can be performed in LODs, where with each level of detail, a finer representation of the point cloud attributes can be obtained. Each level of detail can be specified based on a distance metric to neighboring nodes or based on a sampling distance.

[0078] At the G-PCC encoder 200, the residual obtained as an output of the coding process for the attribute is quantized. The residual can be obtained by subtracting the attribute value from a prediction derived based on points in the neighborhood of the current point and based on the attribute value of previously encoded points. The quantized residual can be decoded using context adaptive arithmetic coding.

[0079] Figure 5 A more detailed diagram Figure 2 2 is a block diagram of an example of a geometry encoding unit 250. The geometry encoding unit 250 may include a coordinate transformation unit 202, a voxelization unit 206, a prediction tree construction unit 207, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, and a geometry reconstruction unit 216.

[0080] like Figure 5 As shown in the example of , the geometry encoding unit 250 can obtain a set of positions of points in the point cloud. In one example, the geometry encoding unit 250 can obtain the position of the points in the point cloud from the data source 104 ( Figure 1 ) to obtain a set of positions and a set of attributes of points in the point cloud. The positions may include coordinates of the points in the point cloud. The geometry encoding unit 250 may generate a geometry bitstream 203 including an encoded representation of the positions of the points in the point cloud.

[0081] The coordinate transformation unit 202 may apply a transformation to the coordinates of the point to transform the coordinates from the initial domain to the transformed domain. The present disclosure may refer to the transformed coordinates as transformed coordinates. The voxelization unit 206 may voxelize the transformed coordinates. Voxelization of the transformed coordinates may include quantizing and removing some points of the point cloud. In other words, multiple points of the point cloud may be grouped into a single "voxel", which may be considered a point in some respects thereafter.

[0082] The prediction tree construction unit 207 may be configured to generate a prediction tree based on the voxelized transform coordinates. The prediction tree construction unit 207 may be configured to perform any of the prediction tree decoding techniques described above in an intra prediction mode or an inter prediction mode. To perform prediction tree decoding using inter prediction, the prediction tree construction unit 207 may access points in a previously encoded frame from the geometry reconstruction unit 216. The dashed line from the geometry reconstruction unit 216 shows the data path when inter prediction is performed. The arithmetic coding unit 214 may entropy code the syntax elements representing the encoded prediction tree.

[0083] The geometry coding unit 250 may perform octree-based decoding instead of prediction tree-based decoding. The octree analysis unit 210 may generate an octree based on the voxelized transform coordinates. The surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may entropy code the syntax elements representing the octree and / or surface information determined by the surface approximation analysis unit 212. The geometry coding unit 250 may output these syntax elements in the geometry bitstream 203. The geometry bitstream 203 may also include other syntax elements, including non-arithmetic-coded syntax elements.

[0084] Octree-based decoding may be performed as an intra prediction technique or an inter prediction technique. To perform octree decoding using inter prediction, the octree analysis unit 210 and the surface approximation analysis unit 212 may access points in a previously encoded frame from the geometry reconstruction unit 216. The dashed line from the geometry reconstruction unit 216 shows the data path when inter prediction is performed.

[0085] The geometry reconstruction unit 216 may reconstruct the transform coordinates of the points in the point cloud based on the octree, the prediction tree, data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transform coordinates reconstructed by the geometry reconstruction unit 216 may be different from the original number of points in the point cloud. The present disclosure may refer to the resulting points as reconstructed points.

[0086] Figure 6 is a block diagram more specifically illustrating Figure 2 an example of the attribute coding unit 260. The attribute coding unit 250 may include a color transform unit 204, an attribute transfer unit 208, a RAHT unit 218, a prediction coding unit 219, a LoD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, an arithmetic coding unit 226, and an attribute reconstruction unit 228. The attribute coding unit 260 may encode the attributes of the points of the point cloud to generate an attribute bitstream 205 including an encoded representation of a set of attributes. The attributes may include information about the points in the point cloud, such as the color associated with the points in the point cloud.

[0087] The color transform unit 204 may apply a transform to transform the color information of the attribute to a different domain. For example, the color transform unit 204 may transform the color information from the RGB color space to the YCbCr color space. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud. The attribute transfer unit 208 may use the original position of the point and the position generated from the attribute encoding unit 250 (e.g., from the geometry reconstruction unit 216) for the transfer.

[0088] The RAHT unit 218 may apply RAHT decoding to the attributes of the reconstructed points. In some examples, under RAHT, the attributes of a block of 2×2×2 point positions are obtained and transformed along one direction to obtain four low-frequency nodes (L) and four high-frequency nodes (H). Subsequently, the four low-frequency nodes (L) are transformed in a second direction to obtain two low-frequency nodes (LL) and two high-frequency nodes (LH). The two low-frequency nodes (LL) are transformed along a third direction to obtain one low-frequency node (LLL) and one high-frequency node (LLH). The low-frequency node LLL corresponds to the DC coefficient, and the high-frequency nodes H, LH, and LLH correspond to the AC coefficient. The transform in each direction may be a 1-D transform with two coefficient weights. The low-frequency coefficients may be regarded as coefficients of a 2×2×2 block for the next higher level RAHT transform, and the AC coefficients may be encoded without change; this transform continues until the top root node. The tree traversal for encoding is used from top to bottom to calculate the weights to be used for the coefficients; the transform order is from bottom to top. The coefficients may then be quantized and encoded.

[0089] Alternatively or additionally, the LoD generation unit 220 and the lifting unit 222 may apply LoD processing and lifting to the attributes of the reconstructed points, respectively. LoD generation is used to divide the attributes into different refinement levels. Each refinement level provides a refinement of the attributes of the point cloud. The first refinement level provides a rough approximation and contains few points; subsequent refinement levels typically contain more points, and so on. The refinement level can be constructed using a distance-based metric, or one or more other classification criteria (e.g., subsampling according to a specific order) can also be used. Therefore, all reconstructed points can be included in a refinement level. Each detail level is generated by taking the union of all points to a specific refinement level: for example, LoD1 is obtained based on the refinement level RL1, LoD2 is obtained based on RL1 and RL2,..., LoDN is obtained by the union of RL1, RL2, RLN. In some cases, LoD generation can be followed by a prediction scheme (e.g., a predictive transform), in which the attributes associated with each point in the LoD are predicted based on a weighted average of previous points, and the residual is quantized and entropy decoded. The lifting scheme is built on top of the predictive transform mechanism, where an update operator is used to update the coefficients, and adaptive quantization of the coefficients is performed.

[0090] The prediction coding unit 219 can be configured to determine the attribute value of the current point based on the attribute values ​​of the points that have been encoded. For example, the prediction coding unit 219 can be configured to determine a list of prediction candidates and select a candidate from the list as a prediction of the attribute value of the current point. The prediction coding unit 219 can then determine a residual value representing the difference between the predicted attribute value of the current point and the actual attribute value of the current point. The present disclosure describes techniques for generating a candidate list for predictive geometric decoding. Predictive geometric decoding will be described in more detail below. The prediction coding unit 219 includes a neighbor replacement unit (NRU) 221, which can be configured to perform the techniques of the present disclosure, including the neighbor replacement techniques described in more detail below.

[0091] The RAHT unit 218, the prediction coding unit 219, and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218, the prediction geometry coding unit 219, or the lifting unit 222. In the case of the prediction coding unit 219, the coefficient quantization unit 224 may quantize the determined residual value, or may skip quantization. The arithmetic coding unit 226 may apply arithmetic coding to the syntax elements representing the quantized coefficients. The G-PCC encoder 200 may output these syntax elements in the attribute bitstream 205. The attribute bitstream 205 may also include other syntax elements, including non-arithmetic coded syntax elements.

[0092] As with the geometry coding unit 250, the attribute coding unit 260 may use intra-prediction or inter-prediction techniques to encode the attributes. The above description of the attribute coding unit 260 generally describes intra-prediction techniques. In other examples, the RAHT unit 215, the LoD generation unit 220, and / or the lifting unit 222 may also use attributes from previously encoded frames to further encode the attributes of the current frame. In this regard, the attribute reconstruction unit 228 may be configured to reconstruct the encoded attributes and store them for possible future use in inter-prediction encoding.

[0093] Figure 7 A more detailed diagram Figure 3 The geometry decoding unit 350 may be configured to perform the Figure 5 The geometry decoding unit 350 receives the geometry bitstream 203 and generates the positions of the points of the point cloud frame. The geometry decoding unit 350 may include a geometry arithmetic decoding unit 302, an octree synthesis unit 306, a prediction tree synthesis unit 307, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, and an inverse coordinate transformation unit 320.

[0094] Geometry decoding unit 350 may receive geometry bitstream 203. Geometry arithmetic decoding unit 302 may apply arithmetic decoding (eg, Context-Adaptive Binary Arithmetic Coding (CABAC) or other types of arithmetic decoding) to syntax elements in geometry bitstream 203.

[0095] The octree synthesis unit 306 may synthesize an octree based on the syntax elements parsed from the geometry bitstream 203. Starting from the root node of the octree, the occupancy of each of the eight child nodes at each octree level is signaled in the bitstream. When the signaling indicates that a child node at a particular octree level is occupied, the occupancy of the child nodes of the child node is signaled. The signaling of the nodes at each octree level is signaled before proceeding to the next octree level.

[0096] At the final level of the octree, each node corresponds to a voxel position; when a leaf node is occupied, one or more points may be specified as being occupied at the voxel position. In some cases, due to quantization, some branches of the octree may terminate earlier than the final level. In this case, a leaf node is considered to be an occupied node with no child nodes. In the case of using surface approximation in the geometry bitstream 203, the surface approximation synthesis unit 310 may determine the surface model based on syntax elements parsed from the geometry bitstream 203 and based on the octree.

[0097] Octree-based coding may be performed as an intra prediction technique or an inter prediction technique. To perform octree coding using inter prediction, the octree synthesis unit 306 and the surface approximation synthesis unit 310 may access points in previously decoded frames from the geometry reconstruction unit 312. The dashed lines from the geometry reconstruction unit 312 show the data path when performing inter prediction.

[0098] The prediction tree synthesis unit may synthesize a prediction tree based on syntax elements parsed from the geometry bitstream 203. The prediction tree synthesis unit 307 may be configured to synthesize the prediction tree using any of the techniques described above, including using both intra prediction techniques or inter prediction techniques. In order to perform prediction tree coding using inter prediction, the prediction tree synthesis unit 307 may access points in previously decoded frames from the geometry reconstruction unit 312. The dashed line from the geometry reconstruction unit 312 shows the data path when performing inter prediction.

[0099] The geometric reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. For each position at a leaf node of the octree, the geometric reconstruction unit 312 may reconstruct the node position by using the binary representation of the leaf node in the octree. At each respective leaf node, the number of points at the respective leaf node is signaled; this indicates the number of repeated points at the same voxel position. When geometric quantization is used, the point position is scaled for determining the reconstructed point position value.

[0100] The inverse transform coordinate unit 320 can apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transform domain back to the original domain. The positions of the points in the point cloud can be in the floating point domain, but the point positions in the G-PCC codec are decoded in the integer domain. The inverse transform can be used to convert the positions back to the original domain.

[0101] Figure 8 A more detailed diagram Figure 3 The attribute decoding unit 360 may be configured to perform the Figure 6 The attribute decoding unit 360 receives the attribute bitstream 205 and generates the attributes of the points of the point cloud frame. The attribute decoding unit 360 may include an attribute arithmetic decoding unit 304, an inverse quantization unit 308, an inverse RAHT unit 314, a prediction decoding unit 315, a LoD generation unit 316, an inverse lifting unit 318, an inverse transform color unit 322, and an attribute reconstruction unit 328.

[0102] The attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream 205. The inverse quantization unit 308 may inverse quantize the attribute value. The attribute value may be based on the syntax elements obtained from the attribute bitstream 205 (eg, including the syntax elements decoded by the attribute arithmetic decoding unit 304).

[0103] Depending on how the attribute value is encoded, the inverse RAHT unit 314 can perform RAHT decoding to determine the color value of the point of the point cloud based on the inverse quantized attribute value. RAHT decoding is performed from the top to the bottom of the tree. At each level, the low-frequency coefficients and high-frequency coefficients derived from the inverse quantization process are used to derive the component value. At the leaf node, the derived value corresponds to the attribute value of the coefficient. The weight derivation process of the point is similar to the process used at the G-PCC encoder 200. Alternatively, the LoD generation unit 316 and the inverse lifting unit 318 can use a detail level-based technique to determine the color value of the point of the point cloud. The LoD generation unit 316 decodes each LoD that gives a gradually finer representation of the point attribute. Using the prediction transform, the LoD generation unit 316 derives the prediction of the point from the weighted sum of the points that were in the previous LoD or previously reconstructed in the same LoD. The LoD generation unit 316 can add the prediction to the residual (which is obtained after inverse quantization) to obtain the reconstructed value of the attribute. When a lifting scheme is used, the LoD generation unit 316 may also include an update operator to update the coefficients used to derive the property value. In this case, the LoD generation unit 316 may also apply inverse adaptive quantization.

[0104] The predictive decoding unit 315 can be configured to determine the attribute value of the current point based on the attribute values ​​of the points that have been encoded. For example, the predictive decoding unit 315 can be configured to determine a list of prediction candidates (i.e., the same list determined by the predictive geometry encoding unit 219), and select a candidate from the list as a prediction of the attribute value for the current point. The predictive decoding unit 315 can then determine a residual value representing the difference between the predicted attribute value of the current point and the actual attribute value of the current point, and add the residual value to the predicted value to determine the final attribute value of the current point. The present disclosure describes techniques for generating a candidate list for predictive geometry decoding. Predictive geometry decoding will be described in more detail below. The predictive decoding unit 315 includes an NRU 317, which can be configured to perform the techniques of the present disclosure, including the neighbor point replacement techniques described in more detail below.

[0105] In addition, Figure 8In the example of , the inverse color transform unit 322 can apply an inverse color transform to the color value. The inverse color transform can be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 can transform the color information from the RGB color space to the YCbCr color space. Accordingly, the inverse color transform unit 322 can transform the color information from the YCbCr color space to the RGB color space.

[0106] The attribute reconstruction unit 328 may be configured to store attributes from previously decoded frames. Attribute decoding may be performed as an intra prediction technique or an inter prediction technique. To perform attribute decoding using inter prediction, the inverse RAHT unit 314, the predictive decoding unit 315, and / or the LoD generation unit 316 may access attributes from previously decoded frames from the attribute reconstruction unit 328.

[0107] Shows Figure 5-Figure 8 The various units of the G-PCC encoder 200 and the G-PCC decoder 300 are described to help understand the operations performed by the G-PCC encoder 200 and the G-PCC decoder 300. The unit can be implemented as a fixed-function circuit, a programmable circuit, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset on an operation that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functions in the operations that can be performed. For example, a programmable circuit can execute software or firmware that enables the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (for example, to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.

[0108] Predictive geometry coding (e.g., see the G-PCC codec description) was introduced as an alternative to octree geometry coding, where nodes are arranged in a tree structure (which defines the prediction structure) and various prediction strategies are used to predict the coordinates of each node in the tree relative to the predicted value of that node. Fig. 9 An example of a prediction tree (i.e., a directed graph where arrows point in the direction of prediction) is shown. Horizontal hash nodes are root vertices and have no prediction values; cross-hatched nodes have two children; diagonal hash nodes have 3 children; non-hashed nodes have one child, and vertical hash nodes are leaf nodes and they have no children. Except for the root node, each node has only one parent node.

[0109] Fig. 9900 is a conceptual diagram illustrating an example of a prediction tree. Node 900 is the root vertex and has no predicted value. Nodes 902 and 904 have two child nodes. Node 906 has three child nodes. Nodes 908, 910, 912, 914, and 916 are leaf nodes and they have no child nodes. The remaining nodes each have one child node. Except for the root node 900, each node has only one parent node.

[0110] Four prediction strategies are specified for each node based on its parent node (p0), grandparent node (p1), and great-grandparent node (p2):

[0111] No forecast / Zero forecast(0)

[0112] Incremental prediction (p0)

[0113] Linear prediction (2*p0–p1)

[0114] Parallelogram prediction (2*p0+p1–p2)

[0115] The G-PCC encoder 200 may employ any algorithm to generate the prediction tree; the algorithm used may be determined based on the application / use case, and several strategies may be used. Some strategies are described in the G-PCC codec description.

[0116] For each node, the residual coordinate values ​​are decoded in the bitstream starting from the root node in a depth-first manner. For example, the G-PCC encoder 200 may decode the residual coordinate values ​​in the bitstream.

[0117] Predictive geometry decoding is mainly useful for category 3 (LIDAR acquired) point cloud data (eg, for low latency applications).

[0118] Fig. 10A and Fig. 10B is a conceptual diagram showing an example of a rotating LIDAR acquisition model. An angular mode for predictive geometry decoding is now described. The angular mode can be used for predictive geometry decoding, where the characteristics of the LIDAR sensor can be used to more efficiently decode the prediction tree. The coordinates of the position are converted to the (r, φ, i) (radius, azimuth, and laser index) domain 600, and the prediction is performed in this domain 600 (e.g., the residuals are decoded in the (r, φ, i) domain). Due to rounding errors, decoding in (r, φ, i) is not lossless, so a second set of residuals corresponding to Cartesian coordinates can be decoded. A description of the encoding and decoding strategies for the angular mode for predictive geometry decoding is reproduced below from the G-PCC codec description.

[0119] This technique focuses on point clouds acquired using a rotating LIDAR model. Here, the LIDAR 602 has N lasers (e.g., N=16, 32, 64) rotating around the Z axis according to an azimuth angle φ. Each laser can have a different elevation angle θ(i) i=1…N and height Assume that the laser i hits according to Figure 10A-10B A point M with Cartesian integer coordinates (x, y, z) defined in the coordinate system shown.

[0120] This technique uses three parameters (r, φ, i) to model the position of M, which are calculated as follows:

[0121]

[0122] φ=atan2(y,x)

[0123]

[0124] More precisely, this technique uses a quantized version of (r,φ,i) denoted as ), where three integers and i are calculated as follows:

[0125]

[0126] Among them (q r ,o r ) and (q φ ,o φ ) are control and is the quantization parameter of the precision of t. sign(t) is a function that returns 1 if t is positive and -1 otherwise. |t| is the absolute value of t.

[0127] To avoid reconstruction mismatches due to the use of floating point operations, andtan(θ(i)) i=1…N The values ​​are precomputed and quantized as follows:

[0128]

[0129] in and (q θ ,o θ ) are control and The quantization parameter of the accuracy.

[0130] The reconstructed Cartesian coordinates are obtained as follows:

[0131]

[0132] where app_cos(.) and app_sin(.) are approximations of cos(.) and sin(.). The calculations may be performed using fixed-point representation, lookup tables, and / or linear interpolation.

[0133] Notice, may be different from (x,y,z) due to various reasons (such as quantization, approximation, model inaccuracy, model parameter inaccuracy, etc.).

[0134] Assume (r x ,r y ,r z ) is the reconstructed residual defined as follows:

[0135]

[0136] Using this technique, the G-PCC encoder 200 can proceed as follows:

[0137] 1) Model parameters and And the quantization parameter q r , q θ and q φ Encoding

[0138] 2) Apply the geometry prediction scheme described in the text of “ISO / IEC FDIS23090-9 Geometry-based point cloud compression” (ISO / IEC JTC 1 / SC29 / WG7m55637, conference call, October 2020)

[0139] 3) For the expression

[0140] New prediction values ​​that take advantage of LIDAR properties can be introduced. For example, the rotation speed of a LIDAR scanner around the z-axis is usually constant. Therefore, the current The following can be predicted:

[0141]

[0142] in

[0143] (δ φ (k)) k=1…Kis a set of potential speeds that the G-PCC encoder 200 can use. The index k can be written explicitly to the bitstream, or can be inferred from the context based on a deterministic strategy applied by both the G-PCC encoder 200 and the G-PCC decoder 300, and n(j) is the number of skip points that can be written explicitly to the bitstream, or can be inferred from the context based on a deterministic strategy applied by both the G-PCC encoder 200 and the G-PCC decoder 300. n(j) is also referred to herein as a "phi multiplier". Note that the phi multiplier is currently only used with incremental prediction values.

[0144] 4) Reconstruct the residual (r) using each node pair x ,r y ,r z ) to encode

[0145] The G-PCC decoder 300 may proceed as follows:

[0146] 1) Model parameters and Decode and set the parameter q r , q θ and q φ Quantify

[0147] 2) The geometry prediction scheme associated with the node is described in the text of "ISO / IEC FDIS23090-9 Geometry-based point cloud compression" (ISO / IEC JTC 1 / SC29 / WG7m55637, conference call, October 2020) Decode the parameters.

[0148] 3) Calculate the reconstructed coordinates as described above

[0149] 4) For the residual (r x ,r y ,r z ) to decode

[0150] As discussed in more detail below, lossy compression can be supported by quantizing the reconstructed residual (r x ,r y ,r z )

[0151] 5) Calculate the original coordinates (x, y, z) as follows:

[0152]

[0153] Lossy compression can be achieved by reconstructing the residual (r x ,r y,r z ) is achieved by applying quantization or by discarding points.

[0154] The quantized reconstruction residual can be calculated as follows:

[0155]

[0156]

[0157] Among them, (q x ,o x ), (q y ,o y ) and (q z ,o z ) are control and For example, the G-PCC encoder 200 or the G-PCC decoder 300 may calculate the quantized residual.

[0158] The G-PCC encoder 200 or the G-PCC decoder 300 may use trellis quantization to further improve RD (rate-distortion) performance results.

[0159] The quantization parameter can be changed at sequence / frame / slice / block level to achieve region-adaptive quality and / or for rate control purposes.

[0160] Inter-frame prediction in G-PCC predictive geometric decoding is described in "G-PCC Second Edition Codec Description" (ISO / IEC JTC1 / SC29 / WG 7MDS21558, conference call, April 2022) (hereinafter referred to as "MDS21558").

[0161] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform inter-frame prediction for predictive geometric decoding, as described in MDS21558, "[G-PCC] EE13.2 Report on Inter-frame Prediction - Test 2" by AK Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz (ISO / IEC JTC1 / SC29 / WG7m56839, April 2021) (hereinafter referred to as "m56839") and "[G-PCC] [EE13.2 related] Additional results for predictive geometric inter-frame prediction" by AK Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz (ISO / IEC JTC1 / SC29 / WG7 m56841, April 2021) (hereinafter referred to as "m56841").

[0162] Predictive geometry coding uses a prediction tree structure to predict the position of a point. When angle coding is enabled, the x, y, z coordinates are transformed into radius, azimuth, and laserID, and residuals are signaled in these three coordinates and in the x, y, z dimensions. Intra-frame prediction for radius, azimuth, and laserID can be one of four modes, and the predicted value is a node in the prediction tree that is classified as a parent node, a grandfather node, and a great-grandfather node relative to the current node. Predictive geometry coding, currently designed in G-PCC version 1, is an intra-frame decoding tool that uses only points in the same frame for prediction. In addition, using points from previously decoded frames can provide better predictions, and therefore provide better compression performance.

[0163] Inter prediction was originally proposed in MDS21558 and m56839 to predict the radius of a point based on a reference frame. For each point in the prediction tree, determine whether the point is inter-predicted or intra-predicted (indicated by a flag). When it is intra-predicted, the intra-prediction mode of predictive geometry decoding is used. When inter prediction is used, the azimuth and laserID are still predicted using intra prediction, and the radius is predicted based on the point in the reference frame that has the same laserID as the current point and the azimuth closest to the current azimuth. In addition to radius prediction, further improvements to the process in m56841 also implement inter prediction of azimuth and laserID. When inter decoding is applied, the radius, azimuth and laserID of the current point are predicted based on points near the azimuth position of previously decoded points in the reference frame. In addition, separate context sets are used for inter and intra prediction.

[0164] exist Fig.11 The process in m56841 is illustrated in FIG. Fig.11 is a conceptual diagram illustrating an example of inter-frame prediction of a current point (curPoint) 1100 in a current frame based on a point (interPredPt) 1102 in a reference frame. Extending inter-frame prediction to azimuth, radius, and laserID may include the following steps:

[0165] • For a given point, select the previous decode point (prevDecP0) 1104.

[0166] • Select a location point (refFrameP0) 1106 in the reference frame that has the same scaled azimuth angle and laserID as prevDecP0 1104 .

[0167] • In the reference frame, find the first point (interPredPt) 1102 whose azimuth angle is greater than refFrameP0 1106. The point interPredPt 1102 may also be referred to as the "next" inter-frame predictor.

[0168] Fig.12 1 is a flowchart illustrating an example decoding process associated with an "inter-frame flag" signaled for each point. The inter-frame flag signaled for a point indicates whether inter-frame prediction is applied to the point. The flowcharts of the present disclosure are provided as examples. Other examples may include more, fewer, or different steps, or the steps may be performed in a different order.

[0169] exist Fig.12 In the example, the G-PCC decoder 300 can determine whether the inter-frame flag of the next point to be decoded (i.e., the current point of the current frame of the point cloud data) indicates that the current point is inter-frame predicted (800). If the inter-frame flag of the current point does not indicate that the current point is inter-frame predicted (the "no" branch of 1200), the G-PCC decoder 300 can identify an intra-frame prediction candidate (812). For example, the G-PCC decoder 300 can determine the intra-frame prediction strategy (e.g., no prediction, incremental prediction, linear prediction, parallelogram prediction, etc.) to determine the prediction value of the current point. The syntax element (pred_mode) signaled in the geometry bitstream 203 can indicate the intra-frame prediction strategy to be used to determine the prediction value of the current point.

[0170] On the other hand, if the inter-frame flag of the current point indicates that the current point is inter-frame predicted (the "yes" branch of 1200), the G-PCC decoder 300 can identify the previous point in the decoding order (e.g., the previous point 1108) (1202). The previous point can have coordinates (r, phi, and laserID). The G-PCC decoder 300 can then derive the quantized phi coordinates (i.e., azimuth coordinates) of the previous point (804). The quantized phi coordinates can be represented as Q(phi). The G-PCC decoder 300 can then check the reference frame (e.g., reference frame 1106) for a point (i.e., an inter-frame prediction point (e.g., interPredPt 1104)) having a quantized phi coordinate greater than the quantized phi coordinate of the previous point (1206). The G-PCC decoder 300 can use the inter-frame prediction point as a prediction value for the current point (1208).

[0171] Regardless of whether the G-PCC decoder 300 uses intra-frame prediction (e.g., as described with respect to step 1212) or inter-frame prediction (e.g., as described with respect to steps 1202-1208) to determine the prediction value for the current point, the G-PCC decoder 300 can add an incremental phi multiplier (1210).

[0172] Fig.13 is a conceptual diagram illustrating an example additional inter-frame prediction value point 1300 obtained from a first point having an azimuth angle greater than the inter-frame prediction value point 1314. In the inter-frame prediction process of the prediction geometry described above regarding Fig.11 when applying inter-frame decoding, for the current point (current point 1100), the radius, azimuth angle, and laser ID are predicted based on points (inter-frame prediction points 1104) near the co-located azimuth angle position (reference position 1110) in the reference frame (reference frame 1106). In the example of Fig.13 the G-PCC encoder 200 and the G-PCC decoder 300 can use the following steps to determine the additional inter-frame prediction value point 1300:

[0173] a) For a given point (current point 1300 of the current frame 1304), determine the previous point 1302 in the current frame 1304 ( Fig.13 "previous decoded point" in

[0174] b) Determine the reference position 1306 in the reference frame 1308 having the same scaled azimuth angle and laser ID as the previous point 1302 determined in step a) ( Fig.13 "reference point having the same scaled azimuth angle and laser ID" in

[0175] c) Determine the position in the reference frame 1308 as the first point having an azimuth angle (e.g., scaled azimuth angle) greater than the reference position 1306 determined in step b) to be used as the inter-frame prediction value point ( Fig.13 inter-frame prediction point 1310 in

[0176] As Fig.13 shown ( Fig.13 "additional inter-frame prediction point 1312" in

[0177] Fig.13 ), the additional inter-frame prediction value point can be obtained by finding the first point having an azimuth angle (e.g., scaled azimuth angle) greater than the inter-frame prediction point 1310 determined in step c). If inter-frame decoding has been applied, additional signaling can be used to indicate which prediction value to select. The additional inter-frame prediction value point can also be referred to as the "NextNext" inter-frame prediction point.

[0177] In some examples, the G-PCC encoder 200 (e.g., the arithmetic coding unit 214 of the G-PCC encoder 200) and the G-PCC decoder 300 (e.g., the geometric arithmetic decoding unit 302 of the G-PCC decoder 300) can apply a context selection algorithm for coding the inter prediction flag. The inter prediction flag values ​​of five previously decoded points can be used to select the context of the inter prediction flag in predictive geometry decoding.

[0178] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to perform global motion compensation. When global motion (GM) parameters are available, inter-frame prediction may be applied using reference frames that are motion compensated using GM parameters, as described in "[G-PCC] [New Proposal] Results on Inter-frame Prediction for Predictive Geometry Coding" by AK Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz (ISO / IEC JTC1 / SC29 / WG7 m59650, April 2022). The GM parameters may include rotation parameters and / or translation parameters.

[0179] Fig.14 A flow chart illustrating the motion compensation process when the reference frame is stored in the spherical domain and the motion compensation is performed in the Cartesian domain is shown. Fig.14 In the example of , the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to convert the reference frame 1402 in the spherical domain from the spherical domain to the Cartesian domain (1404) to generate a reference frame 1406 in the Cartesian domain. The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to perform motion compensation (1408) on the reference frame 1406 to generate a compensated reference frame 1410 in the Cartesian domain. The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to convert the compensated reference frame 1410 from the Cartesian domain to the spherical domain (1412) to generate a compensated reference frame 1414 in the spherical domain.

[0180] Typically, global motion compensation is applied in the Cartesian domain, but in some cases, global motion compensation can also be performed in the spherical domain. Depending on which domain stores the reference frame and which domain compensates the reference frame, one or more of a Cartesian domain → spherical domain conversion or a spherical domain → Cartesian domain conversion can be applied. For example, when the reference frame is stored in the spherical domain and motion compensation is performed in the Cartesian domain, the motion compensation process may include Fig.14 One or more steps as shown.

[0181] In this case, the compensated reference frame can be used for inter prediction. Given a position (x, y, z) in Cartesian coordinates, the corresponding radius and azimuth are calculated as follows (floating point implementation) (as in the CartesianToSpherical conversion function):

[0182] int64_t r0=int64_t(std::round(hypot(xyz[0],xyz[1])));

[0183] auto phi0=std::round((atan2(xyz[1],xyz[0]) / (2.0*M_PI))*scalePhi);

[0184] Among them, scalePhi is modified for different rate points in the lossy configuration; when the geometry is decoded losslessly, the maximum value of 24 bits is used for the azimuth. A fixed-point implementation of the azimuth is available in the convertXyZToRpl function.

[0185] radius:

[0186]

[0187]

[0188] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform attribute prediction, including a default attribute prediction with three neighbors. For example, attribute prediction can be performed by the prediction encoding unit 219 and the prediction decoding unit 315. The simplified prediction structure in the case where LOD is equal to 1 is described in "Improved G-PCC lossless and near lossless decoding" (ISO / IEC JTC1 / SC29 / WG11 input document m44899, Macau, China, October 2018).

[0189] Assume (P i ) i=1…N is the set of locations associated with the points of the point cloud, and let (M i ) i=1…N is the same as (P i ) i=1…N First, the points are sorted in ascending order according to their associated Morton codes. Let I be the array of point indices sorted according to this process. The encoder / decoder compresses / decompresses the points respectively according to the order defined by I. In each iteration i, a point P is selected i In the same way as the current version of G-PCC, P i The distance to s (e.g., s=64) previous points, and select k (e.g., k=3) Pi The nearest neighbor points of are used for prediction.

[0190] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform intra LOD prediction for attribute prediction transforms, as described in "Modifications to the reference structure in TMC13 for attribute prediction transforms" (ISO / IEC JTC1 / SC29 WG11 input document m46107, Marrakesh, Morocco, January 2019) and "[G-PCC] [New] Modifications to intra LOD prediction for attribute prediction transform decoding" (ISO / IEC JTC1 / SC29 / WG11 input document m54633, online, July 2022).

[0191] Fig.15 The reference structure for intra LOD prediction is shown. The EnableReferringSameLoD flag is introduced to control the reference structure used for prediction transforms in order to maintain a tradeoff between decoding efficiency and parallel processing. If the EnableReferringSameLoD flag is set to 1, 3D points in the same LoD can be used for prediction. Arrow 1502 represents an example of intra LOD prediction, where P6 and P1 are both in LOD1, and P6 is predicted based on P1. Arrow 1504 represents an example of non-intra LOD prediction, where P6 and P4 are in different LODs. Note that when the number of LODs is equal to 1, intra LOD prediction is always used.

[0192] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform neighbor searches at the same distance, as described in "[G-PCC] [New Proposal] Improved Implementation of Prediction and Lifting Schemes" (ISO / IEC JTC1 / SC29 / WG11 input document m51010, Geneva, Switzerland, October 2019). The lifting and prediction schemes make extensive use of nearest neighbor searches during the LOD generation and prediction value construction phases. Document m51010 describes how to handle neighbors at the same distance.

[0193] Fig.16 An example of processing neighboring points in subsequent LODs that are at the same distance as the current point is shown. According to document m51010, according to Fig.16 The priority shown is used to process the neighboring points in the subsequent LODs that are at the same distance as the current point. Fig.16In the example of , smaller values ​​of the Morton-based index correspond to smaller values ​​of the Morton code, and smaller values ​​of the priority index correspond to higher priorities. Basically, points are included in the neighbor list based on the concept of distance. When two points have the same distance, there needs to be a well-defined way to determine which point to select. For the purpose of example, assume that two points X1 and X2 are at the same distance. If X1 is searched first, then X1 can be inserted into the list. If X2 is searched, then X2 can only be inserted if there is room or there is another entry in the list with a distance greater than X2. That is, X2 will not replace X1 - so for this, X1 has a higher priority than X2 because X1 comes first. On the other hand, if X2 is searched first before X1, then X2 will be inserted instead of X1. Therefore, it can be said that the priority index corresponds to the order of search.

[0194] Fig.17 An example of processing neighboring points with the same distance as the current point in the same LOD is shown. According to document m51010, according to Fig.17 The priority described in the above is used to process neighboring points in the same LOD that are at the same distance as the current point. Fig.17 In the example of , smaller values ​​of the Morton-based index correspond to smaller values ​​of the Morton code, and smaller values ​​of the priority index correspond to higher priorities.

[0195] Neighbors in subsequent LODs have higher priority than neighbors in the same LOD.

[0196] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform an optimized nearest neighbor search for boosting / prediction as described in "CE13.6 Report on Attribute LOD Construction and Neighbor Search" (ISO / IEC JTC1 / SC29 / WG11 Input Document m54668, online, July 2020).

[0197] The approximate k-NN solution actually gives a good approximation of k-NN, but fails to capture the actual nearest neighbors when significant jumps in the Morton order are observed between neighboring points. An example of this scenario is Fig.18 This is illustrated by points P 1802 and Q 1804.

[0198] To improve the nearest neighbor search, the first optimization uses a lookup table to speed up the k-NN search, and more precisely, determines whether the neighbors of a voxel are occupied, and uses the occupancy to determine the k-NN of the current point. More precisely, let N(i,1), N(i,2), …, N(i,H) (see Fig.19 As an example of a voxel neighbor) is R dThe set of neighbors of P(i) in . For example, the 6 / 18 / 26-connectivity union of P(i) is {P(i)}, and C = [0,…,2 c -1]×[0,…,2 c -1]×[0,…,2 c -1] is the bounding cube of B (i.e., ).

[0199] Fig.19 An example of voxel neighbors is shown. Fig.19 In the example of , voxel neighbor 1902 has 6-connectivity. That is, there are 6 edge-to-edge connections. Voxel neighbor 1904 has 18-connectivity, and voxel neighbor 1906 has 26-connectivity.

[0200] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to precompute a LUT that maps any point X of C to a range of indices describing points in B that have the same position (or Morton code) as point X. The set of points S(i) of the k-NN to be searched to determine a point P(i)∈A is constructed as follows:

[0201] ○S(i)←{}

[0202] ○ For each neighbor point N(i,1), use the LUT to check if the neighbor point belongs to B. If N(i,1)∈B, then add all points in the range ρ(N(i,1)) to S(i).

[0203] ○ If S(i) has fewer than k elements

[0204] ○If S(i) is empty, then S(i)←{Q(j),Q(j+1),Q(j-1),…,Q(j+Δ),Q(j-Δ)}

[0205] ○ If S(i) has at least one element

[0206] ■Set * is the index of the nearest neighbor point P(i) in S(i)

[0207] ■ Apply additional refinement search to S(j * )={Q(j * ),Q(j * +1),Q(j * -1),…,Q(j * +Δ),Q(j * -Δ)}

[0208] The LUT and linear search may also be applied in the reverse order (eg, linear search first, then LUT based search).

[0209] Fig. 20 An example of spatial partitioning is shown. Allocating a LUT to hold C may be expensive in terms of memory. To reduce this requirement, the bounding cube C2002 can be partitioned into 2 e The LUT is initialized each time P(i) enters a new subcube. Only the points in B that are in that subcube will be added to the LUT. For points on the border of a subcube, the neighborhood relationship between the subcube borders is ignored.

[0210] This approach can take advantage of the sparsity and LOD structure of the point cloud to reduce the size of the k-NN search and improve its efficiency. More precisely, the distance constraint used for the boosting / prediction scheme is as follows:

[0211] It constrains the distance sequence {d(1), d(2), …d(L)} as follows:

[0212] ○

[0213] ○d(l+1)=2×d(l)

[0214] n0 is a parameter calculated by the encoder and explicitly signaled in the bitstream.

[0215] If n0 is to be determined, the G-PCC encoder 200 and the G-PCC decoder 300 may first select a random subset of the points of the point cloud and, for each point, calculate the distance between the point and its nearest neighbor. Let δ0 be the pth percentile of these distances (e.g., the 75th percentile). n0 is chosen to prove The smallest integer.

[0216] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform a neighbor search process for attribute LOD prediction as described in "[G-PCC] [EE13.49] Report on attribute LOD prediction" by W. Zhang, T. Tian, ​​L. Yang, F. Yang, M.-L. Champel, and S. Gao (ISO / IEC JTC1 / SC29 / WG7 m58236, October 2021) and "[G-PCC] [New] Improvements to neighbor search for attribute LOD prediction" by W. Zhang, L. Yang, F. Yang, T. Tian, ​​ML. Champel, and S. Gao (ISO / IEC JTC1 / SC29 / WG07 m57324, July 2021).

[0217] In the G-PCC version 1 design, attribute prediction is performed between the point to be decoded and its N nearest neighbors. Considering the complex distribution of 3D points, choosing the nearest point as the prediction value (i.e., using distance as the only criterion) may not always be optimal.

[0218] Fig.21 An example of neighboring point distribution is shown. The proposed process takes point distribution into account when selecting potential predictors. More specifically, the relative positions of potential predictors of the current point P 2102 are considered together with their distances to P2102. Fig.21 In the example of , although P2 2104 is closer to the current point P 2102 than P3 2106, P3 2106 may be a better prediction value of P 2102. The following steps detail the process. According to the first step, the G-PCC encoder 200 and the G-PCC decoder 300 generate two neighbor lists. List1 includes the nearest 3 neighbor points obtained using the existing process in G-PCC. List2 includes N (e.g., 3) points that were missed when updating List1. The final prediction value list is generated by updating List1 using the points in List2, as described in steps 2 and 3.

[0219] Fig. 22Shows the generation of List1 and List2. For purposes of explanation, consider eight neighboring points N0, …, N7, which are sequentially (in this order) considered candidates for the prediction list of the current point P. Chart 2202 shows the distance of each of the eight neighboring points from the current point P (e.g., the distance between N0 and P is 3, the distance between N1 and P is 4, and so on). List1 2204 and List2 2206 show the list generation process for the first list and the second list. In the initial state, moving from left to right, the first three neighboring points — N0, N1, and N2 — have been processed and included in List1 based on increasing distance from P; i.e., dist(P,N0) <= dist(P,N2) <= dist(P,N1), where dist(P,X) represents the distance between point P and point X; List2 is initially empty, and k is set to be equal to 3, where (k–3) indicates the index in List2 where the next candidate should be added. The next neighboring point to be processed is N3, and dist(P,N3) = dist(P,P2) = 4, where when n <= 3, Pn indicates the nth candidate in the current List1, or when n is greater than 3 and less than or equal to 6, Pn indicates the (n-3)th candidate in List2. When dist(P,N3) = dist(P,P2), if List2 is not full (considering the size of List2 to be 3), then N3 is added to List2. When List2 is empty, N3 is added as Pk (k = 3), and k is incremented by 1. The next candidate is N4, and dist(P,N4) = 1 < dist(P,P0) = 3. In this case, the current candidate in P2 (which is N1) is added to List2 as Pk (k = 4), and k is incremented by 1; the remaining candidates in List1 are pushed to the right, and N4 is added as a new candidate to P0. The next candidate is N5, and dist(P,N5) = 2 < dist(P,P1) = 3. In this case, the current candidate in P2 (which is N0) is added to List2 as Pk (k = 5), and k is incremented by 1; and N5 is added as a new candidate to P1. Note that the counter k has now reached 6, and when there can only be 3 entries in List2, k is reset to 3. The next candidate is N6, and dist(P,N6) = 2 < dist(P,P2) = 3. In this case, the current candidate in P2 (which is N0) is added to List2 as Pk (k = 3), and k is incremented by 1; and N6 is added as a new candidate to P2. The final candidate is N7, and dist(N7) = 3. When dist(N7) > dist(P2), it is not added to any list. At the end of the neighboring point list generation process, List1 has neighboring points N4, N5, and N6, and List2 has neighboring points N0, N2, and N1.

[0220] According to the second step, the G-PCC encoder 200 and the G-PCC decoder 300 check the distribution of the points in List1. If P1 or P2 is already in a direction strictly opposite to P0, such as Fig.23 As shown, the distribution of neighboring points is considered to be well dispersed, and therefore there is no need to check the points in List2. Otherwise, check point Pn in List2, if dist(Pn,P)<=T1 and Pn is in a direction strictly opposite to P0, then use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction strictly opposite to P1, then use Pn to replace P2. In the current embodiment, T1=w*dist(P2,P), T2=w*dist(P1,P), and w<<5=54.

[0221] According to the third step, the G-PCC encoder 200 and the G-PCC decoder 30 perform a check of the point distribution (e.g., a roughly relative check). If P2 is not replaced after step 2, and P2 or P1 is in the same direction as P0 (P0, P1 and P2 are in the same area), the distribution check is softened to some extent. Specifically, the distribution check is based on a predefined roughly relative direction, such as Fig.23 As shown. Similar to step 2, check point Pn in List1. If dist(Pn,P)<=T1 and Pn is in a direction roughly opposite to P0, use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction roughly opposite to P1, use Pn to replace P2.

[0222] Fig.23 The concept of relative direction is illustrated. The current point P defines the intersection of the x-axis, y-axis, and z-axis, where the x-axis, y-axis, and z-axis form the xy plane, the xz plane, and the yz plane. Two points located on opposite sides of the xy plane, the xz plane, and the yz plane are strictly opposite points. Points located on opposite sides of two of the xy plane, the xz plane, and the yz plane and located on the same side as one of the xy plane, the xz plane, and the yz plane are substantially opposite points. For example, referring to Fig.23 , two points are strictly opposite if they are located in quadrants 0 and 7, 1 and 6, 2 and 5, or 3 and 4.

[0223] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to signal (eg, send and receive, respectively) a flag that enables a new neighbor search process.

[0224] The modified neighbor search process is enabled by a flag signaled in the APS; when the flag is 1, the modified neighbor search process is enabled, otherwise the default neighbor search process in G-PCC version 1 is used. The current signaling in G-PCC is as follows (the enable flag is called predictionWithDistributionEnabled):

[0225]

[0226]

[0227] The code for replacing neighboring points is provided as follows: numend1 and numend2 are indices corresponding to thresholds T1 and T2. In this case, T1 and T2 are derived based on weights w = distCoefficient / 32 = 54 / 32.

[0228]

[0229]

[0230]

[0231]

[0232] The problems and solutions of the first aspect (aspect 1) of the present disclosure will now be described. A local partitioning unit (LPU) is specified as a 3D spatial region of a point cloud frame; global motion compensation may or may not be applied to each LPU. In G-PCC, there are two types of LPUs:

[0233] - Type 0: Type defined using thresholds. In this type, two thresholds are signaled, specifying a lower threshold T1 and an upper threshold T2, respectively. Points with z-coordinate values ​​between T1 and T2 are considered to be one type 0 LPU (also called ground points), and the remaining points are considered to be another type 0 LPU (also called object points).

[0234] - Type 1: defined by a cuboid with specified width, height and depth. Points within the cuboid with the specified width, height and depth are considered to be part of the corresponding LPU.

[0235] For type 1 LPUs, only width, height and depth need to be signaled, and for type 0 LPUs, only these two thresholds need to be signaled. However, for type 1 LPUs, the threshold is currently also signaled. This results in signaling of unnecessary bits in the bitstream. The current signaling may be as follows:

[0236]

[0237]

[0238] As a proposed solution to the above problem, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine whether to signal a threshold based on the LPU type, and avoid signaling the threshold based on a determination that the threshold does not need to be signaled for a particular LPU type.

[0239] In some cases, the determination may also depend on whether the coding mode is prediction geometry or octree geometry.

[0240] An example implementation of the solution of aspect 1 is as follows, wherein the delimiter <add> and< / add> The text between represents the text to be added, and the separator <del> and< / del> The text between the and is the deleted text.

[0241] Signaling application conditions for the threshold are proposed; the conditions depend on the LPU type and whether octree decoding is applied.

[0242]

[0243]

[0244] In some examples, the signaling may be modified as follows:

[0245]

[0246]

[0247] The problem and solution of the second aspect (Aspect 2) of the present disclosure will now be described. In the above-mentioned neighbor search process for attribute LOD prediction, the G-PCC encoder 200 and the G-PCC decoder 300 are configured to derive six nearest neighbors. Depending on certain conditions, the third neighbor is replaced by one of the fourth neighbor, the fifth neighbor, or the sixth neighbor. Regardless of the maximum number of neighbors that can be used for attribute inter-frame prediction, a flag (the separator below) that enables the modified neighbor search process is signaled. <section 1>and< / section 1> This results in unnecessary signaling of bits, making the coding inefficient.

[0248]

[0249] As a proposed solution to the above problem, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine whether to signal an enable flag for a modified neighbor search based on the maximum number of attributed neighbors. When the enable flag is determined to not be signaled / not present, a default neighbor search process is applied (effectively inferring that the enable flag is 0).

[0250] In one example, the determination may be based on a condition that the maximum number of attribute neighbors is greater than or equal to three.

[0251] In another example, the determination may be based on a condition that a maximum number of attribute neighbors is equal to three.

[0252] An example implementation of the solution of aspect 2 is as follows: The signaling of the flag enabling neighbor search is modified by adding a condition on the number of neighbors (&&aps.num_pred_nearest_neighbors_minus1>=2) as follows:

[0253]

[0254] When predictionWithDistributionEnabled is not signaled, the value of predictionWithDistributionEnabled is inferred to be 0.

[0255] The problems and solutions of the third aspect (Aspect 3) of the present disclosure will now be described. The following examples describe several processes for replacing the third prediction value with one of the fourth prediction value candidate, the fifth prediction value candidate, or the sixth prediction value candidate. However, some of these techniques may also be applied to replace other prediction values ​​(e.g., the zeroth prediction value or the first prediction value). For convenience, replacing the prediction value may also be referred to as a replacement direction.

[0256] The first problem (problem 1) of aspect 3 relates to lacking strict relative check to P1 and P2. According to the first problem, the neighbor point search process for attribute LOD prediction above utilizes the derivation of six nearest neighbor points. Neighbor points are arranged in a way that the distance from the current point increases. Depending on certain conditions, the third neighbor point in the list is replaced by one of the fourth neighbor point, the fifth neighbor point or the sixth neighbor point. The basic idea of ​​the improved search process is to include prediction candidates that are in "relative direction" or "roughly relative direction" to provide a better distribution of predicted values. However, the currently defined rule is not optimal, because in some cases, even in the first three candidates, there is a relative or roughly relative direction, replacement still occurs. This causes a farther neighbor point (fourth, fifth or sixth) to replace a closer neighbor point (second) that causes suboptimal performance to be included in the list.

[0257] The proposed solution to problem 1 is that when P1 and P2 are in strictly opposite directions to each other, the neighboring points are considered to be well dispersed and the list is not updated. This may be implemented as follows. For step 2 as described above, the G-PCC encoder 200 and the G-PCC decoder 300 may perform a modified check on the point distribution (e.g., a strictly relative check). The modification described with respect to step 2 is to use the separator <add> and< / add> The text between the characters represents the text to be added and the separator <del> and< / del> The text between the and are the text to be deleted.

[0258] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to check the distribution of the points in List1. If P1 or P2 is already in a direction strictly opposite to P0 (the definition of opposite is as follows Fig.23 As shown), we believe that the distribution of neighboring points is already well dispersed, so there is no need to check the points in List2. <add> Otherwise, if P1 is in a strictly opposite direction to P2, there is no need to check the points in List2 anymore.< / add> Otherwise, check point Pn in List2, if dist(Pn,P)<=T1 and Pn is in a direction strictly opposite to P0, then use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction strictly opposite to P1, then use Pn to replace P2. In the current implementation, T1=w*dist(P2,P), T2=w*dist(P1,P), and w<<5=54.

[0259] The second problem (Problem 2) of Aspect 3 relates to the unnecessary second threshold T2 in the strict relative check. According to Problem 2, two thresholds are specified when checking the strictly relative direction. When comparing the strictly relative direction to P0, threshold T1 is used, and when comparing the strictly relative direction to P1, threshold T2 is used. Because the list is originally composed based on distances, dist(P1,P) may be less than or equal to dist(P2,P); therefore T2<=T1. When comparing the distances of various points to P, points that satisfy dist(Pn,P)<=T2 always satisfy dist(Pn,P)<=T1. Having two thresholds and two comparisons leads to unnecessary complexity. If the purpose is to check whether any point is strictly relative to P0 or P1, it may be easier to check directly with one threshold.

[0260] The proposed solution to Problem 2 may be implemented as follows:

[0261] Step 2: Check of point distribution (strict relative check)

[0262] Check the distribution of points in List1. If P1 or P2 is already in a direction strictly opposite to P0 (the definition of relative is as follows Fig.23 If the neighboring points are well dispersed, then we believe that the distribution of neighboring points is well dispersed, so we no longer need to check the points in List2. Otherwise, check the point Pn in List2. If dist(Pn,P)<=T1 and Pn is at the same level as P0, <add> or P1< / add> Strictly relative direction, use Pn instead of P2 <del> ; If not, then if dist(Pn,P)<=T2 and Pn is in the strict opposite direction to P1, then replace P2 with Pn< / del> In the current implementation, T1 = w*dist(P2,P), <del> T2=w*dist(P1,P),< / del> And w<<5=54.

[0263] The third problem of aspect 3 (problem 3) relates to the lack of a roughly relative check for P1 and P2. According to the third problem, when the algorithm reaches step 3, it has been established that P0, P1 and P2 are in the same semi-3D plane, which then triggers a check for roughly relative directions. Similar to problem 1, there is no test whether P2 and P1 are roughly opposite to each other. If P1 and P2 are in roughly opposite directions, replacing P2 with another point farther away from the current point P may be suboptimal. An implementation of a solution to problem 3 may be as follows:

[0264] Step 3: Perform a check of point distribution (roughly relative check)

[0265] If P2 is not replaced after step 2, and P2 <del> or P1< / del> In the same direction as P0 (P0, P1 and P2 are in the same area), the distribution check is softened to some extent. Specifically, the distribution check is based on predefined approximate relative directions, such as Fig.23 As shown. Similar to step 2, check point Pn in List1. If dist(Pn,P)<=T1 and Pn is in a direction roughly opposite to P0, use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction roughly opposite to P1, use Pn to replace P2.

[0266] If p2 has not been replaced after step 2, and <del> P2 or< / del> P1 is in the same direction as P0 (P0, P1 and P2 are in the same region), <add> And P0 is not roughly opposite to P2,< / add> The distribution check is softened to some extent. Specifically, the distribution check is based on predefined approximate relative directions, such as Fig.23 As shown. Similar to step 2, check point Pn in List1. If dist(Pn,P)<=T1 and Pn is in a direction roughly opposite to P0, use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction roughly opposite to P1, use Pn to replace P2.

[0267] In one example, P2 may be replaced only if P2 is not substantially opposite to P1 or P0.

[0268] Step 3: Perform a check of point distribution (roughly relative check)

[0269] If p2 is not replaced after step 2, and P2 <del> or P1< / del> is in the same direction as P0 (P0, P1 and P2 are in the same region), <add> and P1 is not roughly opposite to P2,< / add> The distribution check is softened to some extent. Specifically, the distribution check is based on predefined approximate relative directions, such as Fig.23 As shown. Similar to step 2, check point Pn in List1. If dist(Pn,P)<=T1 and Pn is in a direction roughly opposite to P0, use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction roughly opposite to P1, use Pn to replace P2.

[0270] If P2 has not been replaced after step 2, and <del> P2 or< / del> P1 is at the same level as P0 <add>same direction (P0, P1 and P2 are in the same region), <add> And P0 is not roughly opposite to P2,< / add> The distribution check is softened to some extent. Specifically, the distribution check is based on predefined approximate relative directions, such as Fig.23 As shown. Similar to step 2, check point Pn in List1. If dist(Pn,P)<=T1 and Pn is in a direction roughly opposite to P0, use Pn to replace P2; if not, then if dist(Pn,P)<=T2 and Pn is in a direction roughly opposite to P1, use Pn to replace P2.

[0271] The fourth problem (problem 4) of aspect 3 relates to the lack of a check for equality of the two directions. According to problem 4, when the algorithm reaches step 3, it has been established that P0, P1, and P2 are in the same semi-3D plane, which then triggers a check for roughly opposite directions. If P0 is equal to P2 or P0 is equal to P1, step 3 is triggered; however, if P1 is equal to P2, no check is performed. If P1 is equal to P2, and P2 is not replaced, less "distributed" prediction values ​​may result, which in turn is not optimal for prediction.

[0272] The proposed solution is to check all equality cases when checking for roughly relative direction predictions.

[0273] - When all three directions are equal (the separator below <section 2>and< / section 2> ), if any of directions 3 to numend1 is roughly opposite to direction 0 / 1, direction 2 is substituted.

[0274] - When direction 2 is equal to direction 0 or direction 1, and direction 0 is not equal to direction 1 (the separator below <section 3>and< / section 3> The text between ):

[0275] ○ If direction 1 is roughly opposite to direction 0, do not replace.

[0276] ○ Otherwise, direction2 is replaced by direction3 to numend1 with the smallest index that is not equal to direction0 or direction1.

[0277] (Because the three directions are on the same half-plane, if a direction is not equal to 0 or 1, then the direction not equal to 0 or 1 is automatically approximately opposite to 0 or 1).

[0278] - When direction 2 is not equal to direction 0 or 1, and direction 0 is equal to direction 1 (the separator below <section 4>and< / section 4> The text between ):

[0279] ○ If direction 2 is roughly opposite to direction 0, do not replace.

[0280] ○ Otherwise, direction 2 is replaced by direction 3 to numend1 with the smallest index approximately opposite direction 0.

[0281] Another situation is that no two directions are equal among directions 0, 1, and 2. Since there are only four directions in a 3D half-plane, at least two directions among 0, 1, and 2 are approximately opposite.

[0282]

[0283]

[0284]

[0285] The following examples represent sample solutions to the problems described above. In these examples, the delimiter <add> and< / add> The text between represents the text to be added, and the separator <del> and< / del> The text between the delimiters represents the deleted text.<Aspect X,Solution Y> and< / Aspect X,Solution Y> Intended to identify text corresponding to the aspects and solutions described above.

[0286] A first example including solutions to Problem 1, Problem 2, and Problem 3 is provided below:

[0287]

[0288]

[0289]

[0290]

[0291]

[0292]

[0293] A second example including solutions to Problem 1, Problem 2, and Problem 4 is provided below:

[0294]

[0295]

[0296]

[0297]

[0298] A third example corresponding to another implementation of Example 2 is as follows:

[0299]

[0300]

[0301]

[0302] The examples in various aspects of this disclosure may be used alone or in any combination.

[0303] Fig.24 is a flow chart illustrating example operation of G-PCC decoder 300 according to one or more techniques of this disclosure. Fig.24 The technology can be performed by other types of G-PCC decoding devices. In addition, Fig.24 The techniques may be performed in whole or in part by a decoding function of a G-PCC encoding device (such as the G-PCC encoder 200).

[0304] exist Fig.24 In the example of , G-PCC decoder 300 determines a first property value for a point closest to a current point of the point cloud ( 2402 ). The closest point corresponds to an already decoded point that is less than a threshold distance from the current point.

[0305] G-PCC decoder 300 determines a second property value for a point that is second closest to the current point of the point cloud (2404). The second closest point corresponds to a second already decoded point that is less than a threshold distance from the current point.

[0306] G-PCC decoder 300 determines a third property value for a point that is third closest to the current point of the point cloud (2406). The third closest point corresponds to a third already decoded point that is less than a threshold distance from the current point.

[0307] G-PCC decoder 300 determines a fourth attribute value for a fourth point of the point cloud (2408). The fourth point corresponds to a fourth decoded point that is less than a threshold distance from the current point. The fourth point may be a fourth closest point to the current point, or may be, for example, a fifth or sixth closest point.

[0308] The G-PCC decoder 300 determines a candidate set of prediction values ​​for the attribute value of the current point of the point cloud based on a comparison of the position of the second point with the position of the third point (2410). As described in more detail above, the comparison of the position of the second point with the position of the third point may, for example, include determining whether the second point and the third point are strictly relative, approximately relative, or not relative. For example, if the second point and the third point are strictly relative or approximately relative, the G-PCC decoder 300 may include the attribute values ​​of the first point, the second point, and the third point in the candidate set of prediction values. In an example scenario, if the first point, the second point, and the third point are not strictly relative, and if the first point, the second point, and the third point are not approximately relative, only the attribute value of the fourth point is included in the candidate set of prediction values.

[0309] In some cases, determining whether the second point and the third point are strictly relative, approximately relative, or not relative can be performed only in response to determining that the first point and the second point are not strictly relative or approximately relative and / or in response to determining that the first point and the third point are not strictly relative or approximately relative.

[0310] The G-PCC decoder 300 decodes the attribute value of the current point based on the candidate set of prediction values ​​(2412). The G-PCC decoder 300 can, for example, reconstruct a point cloud based on the decoded attribute values. The reconstructed point cloud can then be used for various applications, including Figure 25-28 The application described in .

[0311] Fig.25 is a conceptual diagram illustrating an example ranging system 2500 that can be used with one or more techniques of this disclosure. Fig.25 In the example of , ranging system 2500 includes illuminator 2502 and sensor 2504. Illuminator 2502 can emit light 2506. In some examples, illuminator 2502 can emit light 2506 as one or more laser beams. Light 2506 can be at one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, light 2506 is not a coherent laser. When light 2506 encounters an object (such as object 2508), light 2506 produces return light 2510. Return light 2510 may include backscattered light and / or reflected light. Return light 2510 may pass through lens 2511, which guides return light 2510 to produce an image 2512 of object 2508 on sensor 2504. Sensor 2504 generates signal 2518 based on image 2512. Image 2512 may include a set of points (e.g., such as Fig.25 as shown by the points in image 2512).

[0312] In some examples, illuminator 2502 and sensor 2504 can be mounted on a rotating structure so that illuminator 2502 and sensor 2504 capture a 360-degree view of the environment. In other examples, ranging system 2500 can include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable illuminator 2502 and sensor 2504 to detect objects within a certain range (e.g., up to 360 degrees). Fig.25 The example shows only a single illuminator 2502 and sensor 2504, but the ranging system 2500 can include multiple sets of illuminators and sensors.

[0313] In some examples, the illuminator 2502 generates a structured light pattern. In such an example, the ranging system 2500 may include a plurality of sensors 2504 on which respective images of the structured light pattern are formed. The ranging system 2500 may use the differences between the images of the structured light pattern to determine the distance to an object 2508 from which the structured light pattern is backscattered. When the object 2508 is relatively close to the sensor 2504 (e.g., 0.2 meters to 2 meters), the ranging system based on structured light can have a high level of accuracy (e.g., accuracy in the sub-millimeter range). This high level of accuracy can be used for facial recognition applications, such as unlocking mobile devices (e.g., mobile phones, tablet computers, etc.) and security applications.

[0314] In some examples, the ranging system 2500 is a system based on time of flight (ToF). In some examples where the ranging system 2500 is a ToF-based system, the illuminator 2502 generates pulses of light. In other words, the illuminator 2502 can modulate the amplitude of the emitted light 2506. In such an example, the sensor 2504 detects the return light 2510 from the pulse of light 2506 generated by the illuminator 2502. The ranging system 2500 can then determine the distance to the object 2508 from which the light 2506 is backscattered based on the delay between the time when the light 2506 is emitted and detected and the known speed of light in air. In some examples, instead of modulating the amplitude of the emitted light 2506 (or in addition), the illuminator 2502 can modulate the phase of the emitted light 2506. In such an example, sensor 2504 can detect the phase of return light 2110 from object 2508 and determine the distance to a point on object 2508 using the speed of light and based on the time difference between the time when illuminator 2502 generates light 2506 with a particular phase and the time when sensor 2504 detects return light 2510 with the particular phase.

[0315] In other examples, a point cloud can be generated without using the illuminator 2502. For example, in some examples, the sensor 2504 of the ranging system 2500 can include two or more optical cameras. In such examples, the ranging system 2500 can use the optical cameras to capture a stereoscopic image of the environment, including the object 2508. The ranging system 2500 (e.g., the point cloud generator 2520) can then calculate the differences between the positions in the stereoscopic image. The ranging system 2500 can then use the differences to determine the distances to the shown positions in the stereoscopic image. Based on these distances, the point cloud generator 2520 can generate a point cloud.

[0316] The sensor 2504 can also detect other attributes of the object 2508, such as color and reflection information. In Fig.25 examples, the point cloud generator 2520 can generate a point cloud based on the signal 2518 generated by the sensor 2504. The ranging system 2500 and / or the point cloud generator 2520 can form part of the data source 104 ( Figure 1 ).

[0317] Fig.26 is a conceptual diagram showing an example of a vehicle-based scenario in which one or more techniques of the present disclosure can be used. In Fig.26 examples, the vehicle 2600 includes a laser assembly 2602, such as a LIDAR system. Although not shown in Fig.26 examples, the vehicle 2600 can also include a data source and a G-PCC encoder, such as the G-PCC encoder 200 ( Figure 1 ). In Fig.26 examples, the laser assembly 2602 emits a laser beam 2604, and the laser beam 2604 is reflected from a pedestrian 2606 or other object in the road. The data source of the vehicle 2600 can generate a point cloud based on the signal generated by the laser assembly 2602. The G-PCC encoder of the vehicle 2600 can encode the point cloud to generate a bitstream 2608. The bitstream 2608 can include far fewer bits than the unencoded point cloud obtained by the G-PCC encoder. The output interface of the vehicle 2600 (e.g., the output interface 108 ( Figure 1 )) can transmit the bitstream 2608 to one or more other devices. Thus, the vehicle 2600 can transmit the bitstream 2608 to other devices faster than the unencoded point cloud data. Additionally, the bitstream 2608 may require less data storage capacity.

[0318] In Fig.26 examples, the vehicle 2600 can transmit the bitstream 2608 to another vehicle 2610. The vehicle 2610 can include a G-PCC decoder, such as the G-PCC decoder 300 ( Figure 1 ). The G-PCC decoder of vehicle 2610 can decode bitstream 2608 to reconstruct the point cloud. Vehicle 2610 can use the reconstructed point cloud for various purposes. For example, vehicle 2610 can determine that pedestrian 2606 is in the road ahead of vehicle 2600 based on the reconstructed point cloud, and therefore (e.g., even before the driver of vehicle 2610 realizes that pedestrian 2606 is in the road) begin to slow down. Thus, in some examples, vehicle 2610 can perform autonomous navigation operations, generate notifications or warnings, or perform another action based on the reconstructed point cloud.

[0319] Additionally or alternatively, vehicle 2600 may transmit bitstream 2608 to server system 2612. Server system 2612 may use bitstream 2608 for various purposes. For example, server system 2612 may store bitstream 2608 for subsequent reconstruction of a point cloud. In this example, server system 2612 may use the point cloud along with other data (e.g., vehicle telemetry data generated by vehicle 2600) to train an autonomous driving system. In other examples, server system 2612 may store bitstream 2608 for subsequent reconstruction for use in a forensic crash investigation (e.g., if vehicle 2600 collides with pedestrian 2606), or may transmit notifications or instructions for navigation to vehicle 2600 or vehicle 2610.

[0320] Fig. 27 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of the present disclosure may be used. Extended reality (XR) is a term used to cover a range of technologies, including augmented reality (AR), mixed reality (MR), and virtual reality (VR). Fig. 27 In the example of , a first user 2700 is located at a first location 2702. User 2700 wears an XR headset 2704. As an alternative to the XR headset 2704, user 2700 may use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 2704 includes a depth detection sensor, such as a LIDAR system, that detects the location of points on an object 2706 at the first location 2702. The data source of the XR headset 2704 may use a signal generated by the depth detection sensor to generate a point cloud representation of the object 2706 at the location 2702. The XR headset 2704 may include a G-PCC encoder (e.g., Figure 1 G-PCC encoder 200).

[0321] The XR headset 2704 may transmit the bitstream 2708 (e.g., via a network such as the Internet) to an XR headset 2710 worn by a user 2712 at a second location 2714. The XR headset 2710 may decode the bitstream 2708 to reconstruct the point cloud. The XR headset 2710 may use the point cloud to generate an XR visualization (e.g., AR, MR, VR visualization) representing an object 2706 at the location 2702. Thus, in some examples, such as when the XR headset 2710 generates a VR visualization, the user 2712 at the location 2714 may have a 3D immersive experience of the location 2702. In some examples, the XR headset 2710 may determine the location of a virtual object based on the reconstructed point cloud. For example, the XR headset 2710 may determine that the environment (e.g., location 2702) includes a flat surface based on the reconstructed point cloud, and then determine that the virtual object (e.g., a cartoon character) will be positioned on the flat surface. The XR headset 2710 may generate an XR visualization in which the virtual objects are in the determined positions. For example, the XR headset 2710 may display a cartoon character sitting on a flat surface.

[0322] Fig.28 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure may be used. Fig.28 In an example of FIG. 2 , a mobile device 2800 (such as a mobile phone or tablet computer) includes a depth detection sensor, such as a LIDAR system, that detects the location of points on an object 2802 in the environment of the mobile device 2800. A data source of the mobile device 2800 can generate a point cloud representation of the object 2802 using a signal generated by the depth detection sensor. The mobile device 2800 can include a G-PCC encoder (e.g., Figure 1 G-PCC encoder 200). Fig.28 In an example, mobile device 2800 can transmit a bitstream to remote device 2806 (such as a server system or other mobile device). Remote device 2806 can decode bitstream 2804 to reconstruct a point cloud. Remote device 2806 can use point cloud for various purposes. For example, remote device 2806 can use point cloud to generate a map of the environment of mobile device 2800. For example, remote device 2806 can generate a map of the interior of a building based on the reconstructed point cloud. In another example, remote device 2806 can generate an image (e.g., computer graphics) based on point cloud. For example, remote device 2806 can use the points of point cloud as vertices of polygons, and use the color attributes of the points as the basis for shading the polygons. In some examples, remote device 2806 can use point cloud to perform facial recognition.

[0323] The following numbered clauses illustrate one or more aspects of the devices and techniques described in this disclosure.

[0324] Clause 1: A device for processing point cloud data, the device comprising: a memory configured to store point cloud data; and one or more processors implemented in circuitry and configured to: determine a first attribute value for a first point of the point cloud, wherein the first point of the point cloud is the decoded point closest to the current point of the point cloud; determine a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is the decoded point second closest to the current point of the point cloud; determine a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is the decoded point third closest to the current point of the point cloud; determine a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is either a decoded point farther from the current point than the third point or a decoded point at the same distance from the current point as the third point; determine a set of predicted value candidates for the attribute value of the current point of the point cloud, wherein the current point defines the intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, y-axis, and z-axis form an x-y plane, an x-z plane, and a y-z plane, and wherein, to determine the set of predicted value candidates for the current point of the point cloud, the one or more processors are further configured to generate a set of predicted value candidates having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point; and decode the attribute value of the current point based on the set of predicted value candidates.

[0325] Clause 2: The device according to Clause 1, wherein, to generate a set of predicted value candidates having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determine whether the second point is strictly opposite the third point, wherein two points are strictly opposite if they are on opposite sides of the x-y plane, opposite sides of the x-z plane, and opposite sides of the y-z plane; and in response to determining that the second point is strictly opposite the third point, include the first attribute value, the second attribute value, and the third attribute value in the set of predicted value candidates.

[0326] Clause 3: The device according to Clause 2, wherein, to generate a set of predicted value candidates having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determine whether the second point is strictly opposite the first point; and in response to determining that the second point is not strictly opposite the first point, determine whether the second point is strictly opposite the third point.

[0327] Article 4: A device according to Article 2, wherein in order to generate a candidate set of prediction values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point, one or more processors are also configured to: determine whether the third point is strictly opposite to the first point; and in response to determining that the third point is not strictly opposite to the first point, determine whether the second point is strictly opposite to the third point.

[0328] Article 5: A device according to Article 2, wherein in order to generate a candidate set of prediction values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point, one or more processors are also configured to: determine whether the second point is strictly opposite to the first point; determine whether the third point is strictly opposite to the first point; and in response to determining that the second point is not strictly opposite to the first point and the third point is not strictly opposite to the first point, determine whether the second point is strictly opposite to the third point.

[0329] Article 6: A device according to Article 1, wherein in order to generate a prediction value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: in response to determining that the second point is on the opposite side of the xy plane, the opposite side of the xz plane and the opposite side of the yz plane relative to the third point, include the first attribute value, the second attribute value and the third attribute value in the prediction value candidate set, and not include the fourth attribute value in the prediction value candidate set.

[0330] Article 7: A device according to Article 1, wherein in order to generate a candidate set of prediction values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point, one or more processors are also configured to: determine whether the third point and the second point are approximately opposite, wherein the two points are approximately opposite if they are on opposite sides of two planes among the xy plane, the xz plane and the yz plane, and on the same side as one of the xy plane, the xz plane and the yz plane; and in response to determining that the second point is approximately opposite to the third point, include the first attribute value, the second attribute value and the third attribute value in the candidate set of prediction values.

[0331] Article 8: A device according to any one of Articles 1-7, wherein one or more processors are further configured to: determine whether the maximum number of neighboring points used for prediction is at least three; and in response to determining that the maximum number of neighboring points used for prediction is at least three, receive a grammatical element indicating that a fourth attribute value of a fourth point is eligible to be included in the set of prediction value candidates.

[0332] Article 9: A device according to any one of Articles 1-8, wherein in order to decode the attribute value of the current point based on the prediction value candidate set, one or more processors are also configured to: determine a candidate value from the prediction value candidate set; receive a residual value; and determine the attribute value of the current point based on the candidate value and the residual value.

[0333] Article 10: A device according to any one of Articles 1-9, wherein the first attribute value, the second attribute value, the third attribute value, the fourth attribute value and the attribute value of the current point include color values.

[0334] Article 11: A device according to any one of Articles 1-10, wherein the one or more processors are further configured to generate a candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, and to decode the attribute value of the current point based on the candidate set of prediction values ​​as part of a process of encoding point cloud data.

[0335] Clause 12: The device of any one of clauses 1-12, wherein the one or more processors are further configured to reconstruct a point cloud based on the attribute values ​​of the current point. Clause 13: The device of clause 12, wherein the one or more processors are further configured to generate a map of the interior of the building based on the reconstructed point cloud.

[0336] Clause 14: The apparatus of clause 12, wherein the one or more processors are further configured to perform autonomous navigation operations based on the reconstructed point cloud.

[0337] Clause 15: The apparatus of clause 12, wherein the one or more processors are further configured to generate computer graphics based on the reconstructed point cloud.

[0338] Clause 16: The apparatus of clause 12, wherein the one or more processors are configured to: determine a position of a virtual object based on the reconstructed point cloud; and generate an extended reality (XR) visualization in which the virtual object is at the determined position.

[0339] Clause 17: The apparatus of clause 12, further comprising a display for presenting an image based on the reconstructed point cloud.

[0340] Clause 18: A device according to any one of clauses 1 to 17, wherein the device is one of a mobile phone or a tablet computer.

[0341] Article 19: The device of any one of Articles 1-17, wherein the device is a vehicle.

[0342] Clause 20: A device as described in any of Clauses 1-17, wherein the device is an extended reality device.

[0343] Article 21: A method for processing point cloud data, the method comprising: determining a first attribute value for a first point of the point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is farther away from the current point than the third point an already decoded point; determining a candidate set of prediction values ​​for the attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine a candidate set of prediction values ​​for the current point of the point cloud, determining the candidate set of prediction values ​​comprises generating a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of a second point with a position of a third point; and decoding the attribute value of the current point based on the candidate set of prediction values.

[0344] Article 22: A device according to Article 21, wherein generating a candidate set of predicted values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point includes: determining whether the second point is strictly opposite to the third point, wherein the two points are strictly opposite if the two points are on opposite sides of the xy plane, opposite sides of the xz plane and opposite sides of the yz plane; and in response to determining that the second point is strictly opposite to the third point, including the first attribute value, the second attribute value and the third attribute value in the candidate set of predicted values.

[0345] Article 23: An apparatus according to Article 22, wherein generating a candidate set of predicted values ​​having a subset of first attribute values, second attribute values, third attribute values, and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point includes: determining whether the second point is strictly opposite to the first point; and in response to determining that the second point is not strictly opposite to the first point, determining whether the second point is strictly opposite to the third point.

[0346] Article 24: A device according to Article 22, wherein generating a candidate set of predicted values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point includes: determining whether the third point is strictly opposite to the first point; and in response to determining that the third point is not strictly opposite to the first point, determining whether the second point is strictly opposite to the third point.

[0347] Article 25: According to the device described in Article 22, generating a candidate set of predicted values ​​having a subset of first attribute values, second attribute values, third attribute values ​​and fourth attribute values ​​based on a comparison of the position of the second point with the position of the third point includes: determining whether the second point is strictly opposite to the first point; determining whether the third point is strictly opposite to the first point; and in response to determining that the second point is not strictly opposite to the first point and the third point is not strictly opposite to the first point, determining whether the second point is strictly opposite to the third point.

[0348] Article 26: A device according to Article 21, wherein generating a prediction value candidate set having a subset of first attribute value, second attribute value, third attribute value and fourth attribute value based on comparison of the position of the second point with the position of the third point includes: in response to determining that the second point is on the opposite side of the xy plane, the opposite side of the xz plane and the opposite side of the yz plane relative to the third point, including the first attribute value, the second attribute value and the third attribute value in the prediction value candidate set, and not including the fourth attribute value in the prediction value candidate set.

[0349] Article 27: A device according to Article 21, wherein generating a prediction value candidate set having a subset of first attribute value, second attribute value, third attribute value and fourth attribute value based on a comparison of the position of the second point with the position of the third point includes: determining whether the third point and the second point are approximately opposite, wherein the two points are approximately opposite if they are on opposite sides of two planes among the xy plane, the xz plane and the yz plane, and on the same side as one of the xy plane, the xz plane and the yz plane; and in response to determining that the second point is approximately opposite to the third point, including the first attribute value, the second attribute value and the third attribute value in the prediction value candidate set.

[0350] Article 28: The apparatus according to any one of Articles 21-27 further includes: determining whether the maximum number of neighboring points used for prediction is at least three; and in response to determining that the maximum number of neighboring points used for prediction is at least three, receiving a grammatical element indicating that a fourth attribute value of a fourth point is eligible to be included in the set of prediction value candidates.

[0351] Article 29: A device according to any one of Articles 21-28, wherein decoding the attribute value of the current point based on the prediction value candidate set includes: determining a candidate value from the prediction value candidate set; receiving a residual value; and determining the attribute value of the current point based on the candidate value and the residual value.

[0352] Article 30: A device according to any one of Articles 21-29, wherein the first attribute value, the second attribute value, the third attribute value, the fourth attribute value and the attribute value of the current point include color values.

[0353] Article 31: The device according to any one of Articles 21-31 also includes: generating a candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, and decoding the attribute value of the current point based on the candidate set of prediction values ​​as part of the process of encoding point cloud data.

[0354] Article 32: A computer-readable storage medium storing instructions, which, when executed by one or more processors, causes the one or more processors to perform any of the methods of Articles 21-31.

[0355] Article 33: A device for processing point cloud data, the method comprising: a component for determining a first attribute value for a first point of the point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; a component for determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; a component for determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; a component for determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is farther away from the current point than the third point a decoded point; a component for determining a candidate set of prediction values ​​for an attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein in order to determine a candidate set of prediction values ​​for a current point of the point cloud, the component for determining the candidate set of prediction values ​​includes a component for generating a candidate set of prediction values ​​having a subset of a first attribute value, a second attribute value, a third attribute value, and a fourth attribute value based on a comparison of a position of a second point with a position of a third point; and a component for decoding the attribute value of the current point based on the candidate set of prediction values.

[0356] It should be appreciated that, depending on the example, certain actions or events of any technique described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). In addition, in some examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiple processors) rather than sequentially.

[0357] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted on a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media (which corresponds to tangible media such as data storage media) or communication media (including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol). In this manner, a computer-readable medium may generally correspond to (1) a non-temporary tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0358] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the required program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is properly referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave) is used to transmit instructions from a website, server or other remote source, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other temporary media, but point to non-temporary tangible storage media. Disks and optical disks used herein include compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks and blue-ray discs, wherein disks usually reproduce data magnetically, and optical discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0359] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the terms "processor" and "processing circuitry" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Furthermore, these techniques may be implemented entirely in one or more circuits or logic elements.

[0360] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or collections of ICs (e.g., chipsets). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units, including one or more processors as described above in combination with appropriate software and / or firmware.

[0361] Various examples have been described. These and other examples are within the scope of the following claims.< / section> < / section> < / section> < / add> < / section>

Claims

1. A device for processing point cloud data, the device include: A memory configured to store the point cloud data; as well as One or more processors implemented in circuitry and configured to: determining a first attribute value for a first point of a point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is either a decoded point that is farther from the current point than the third point or a decoded point that is the same distance from the current point as the third point; determining a candidate set of prediction values ​​for an attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein to determine the candidate set of prediction values ​​for the current point of the point cloud, the one or more processors are further configured to generate the candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of a position of the second point with a position of the third point; as well as The attribute value of the current point is decoded based on the candidate set of prediction values.

2. The device according to claim 1, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determining whether the second point is strictly opposite to the third point, wherein the two points are strictly opposite if they are on opposite sides of the xy plane, on opposite sides of the xz plane, and on opposite sides of the yz plane; as well as In response to determining that the second point is strictly opposite to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the prediction value candidate set.

3. The device according to claim 2, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determining whether the second point is strictly opposite to the first point; as well as In response to determining that the second point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

4. The device according to claim 2, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determining whether the third point is strictly opposite to the first point; as well as In response to determining that the third point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

5. The device according to claim 2, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determining whether the second point is strictly opposite to the first point; determining whether the third point is strictly opposite to the first point; as well as In response to determining that the second point is not strictly opposite to the first point and the third point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

6. The device according to claim 1, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: In response to determining that the second point is on an opposite side of the xy plane, an opposite side of the xz plane, and an opposite side of the yz plane relative to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the prediction value candidate set, while the fourth attribute value is not included in the prediction value candidate set.

7. The device according to claim 1, in, To generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, the one or more processors are further configured to: determining whether the third point and the second point are substantially opposite, wherein the two points are substantially opposite if the two points are on opposite sides of two of the xy plane, the xz plane, and the yz plane and on the same side as one of the xy plane, the xz plane, and the yz plane; as well as In response to determining that the second point is substantially opposite to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the set of predicted value candidates.

8. The device according to claim 1, in, The one or more processors are further configured to: determining whether the maximum number of neighbors used for prediction is at least three; and In response to determining that the maximum number of neighboring points for prediction is at least three, a syntax element is received indicating that a fourth property value of the fourth point qualifies for inclusion in the set of predictor candidates.

9. The device according to claim 1, in, In order to decode the attribute value of the current point based on the candidate set of prediction values, the one or more processors are further configured to: Determining a candidate value from the set of predicted value candidates; receiving residual values; and The attribute value of the current point is determined based on the candidate value and the residual value.

10. The device according to claim 1, in, The first attribute value, the second attribute value, the third attribute value, the fourth attribute value, and the attribute value of the current point include color values.

11. The device according to claim 1, in, The one or more processors are also configured to generate the candidate set of predicted values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, and to decode the attribute value of the current point based on the candidate set of predicted values ​​as part of a process of encoding the point cloud data.

12. The device according to claim 1, in, The one or more processors are further configured to reconstruct the point cloud based on the attribute value of the current point.

13. The device according to claim 12, in, The one or more processors are also configured to generate a map of the interior of the building based on the reconstructed point cloud.

14. The device according to claim 12, in, The one or more processors are further configured to perform autonomous navigation operations based on the reconstructed point cloud.

15. The device according to claim 12, in, The one or more processors are also configured to generate computer graphics based on the reconstructed point cloud.

16. The device according to claim 12, in, The one or more processors are configured to: determining a position of a virtual object based on the reconstructed point cloud; and An extended reality (XR) visualization is generated, wherein the virtual object is at the determined location.

17. The apparatus of claim 12, further comprising a display for presenting an image based on the reconstructed point cloud.

18. The device according to claim 1, in, The device is one of a mobile phone or a tablet computer.

19. The device according to claim 1, in, The device is a vehicle.

20. The device according to claim 1, in, The device is an extended reality device.

21. A method for processing point cloud data, the method include: determining a first attribute value for a first point of a point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is a decoded point that is farther from the current point than the third point; determining a candidate set of prediction values ​​for an attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein to determine the candidate set of prediction values ​​for the current point of the point cloud, wherein determining the candidate set of prediction values ​​comprises generating the candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of a position of the second point with a position of the third point; as well as The attribute value of the current point is decoded based on the candidate set of prediction values.

22. The device according to claim 21, in, Generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point comprises: determining whether the second point is strictly opposite to the third point, wherein the two points are strictly opposite if they are on opposite sides of the xy plane, on opposite sides of the xz plane, and on opposite sides of the yz plane; and In response to determining that the second point is strictly opposite to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the prediction value candidate set.

23. The device according to claim 22, in, Generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point comprises: determining whether the second point is strictly opposite to the first point; and In response to determining that the second point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

24. The device according to claim 22, in, Generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point comprises: determining whether the third point is strictly opposite to the first point; and In response to determining that the third point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

25. The apparatus according to claim 22, generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point include: determining whether the second point is strictly opposite to the first point; determining whether the third point is strictly opposite to the first point; as well as In response to determining that the second point is not strictly opposite to the first point and the third point is not strictly opposite to the first point, it is determined whether the second point is strictly opposite to the third point.

26. The apparatus according to claim 21, in, Generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point comprises: In response to determining that the second point is on an opposite side of the xy plane, an opposite side of the xz plane, and an opposite side of the yz plane relative to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the prediction value candidate set, while the fourth attribute value is not included in the prediction value candidate set.

27. The apparatus according to claim 21, in, Generating the predicted value candidate set having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point comprises: determining whether the third point and the second point are substantially opposite, wherein the two points are substantially opposite if the two points are on opposite sides of two of the xy plane, the xz plane, and the yz plane and on the same side as one of the xy plane, the xz plane, and the yz plane; and In response to determining that the second point is substantially opposite to the third point, the first attribute value, the second attribute value, and the third attribute value are included in the set of predicted value candidates.

28. The device according to claim 21, further comprising: include: Determine whether the maximum number of neighbors used for prediction is at least three; as well as In response to determining that the maximum number of neighboring points for prediction is at least three, a syntax element is received indicating that a fourth property value of the fourth point qualifies for inclusion in the set of predictor candidates.

29. The apparatus according to claim 21, in, Decoding the attribute value of the current point based on the predicted value candidate set includes: Determining a candidate value from the set of predicted value candidates; receiving residual values; and The attribute value of the current point is determined based on the candidate value and the residual value.

30. The apparatus according to claim 21, in, The first attribute value, the second attribute value, the third attribute value, the fourth attribute value, and the attribute value of the current point include color values.

31. The device according to claim 21, further comprising: include: Generate a candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of the position of the second point with the position of the third point, and decode the attribute value of the current point based on the candidate set of prediction values ​​as part of a process of encoding the point cloud data.

32. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to: determining a first attribute value for a first point of a point cloud, wherein the first point of the point cloud is a decoded point that is closest to a current point of the point cloud; determining a second attribute value for a second point of the point cloud, wherein the second point of the point cloud is a decoded point that is second closest to the current point of the point cloud; determining a third attribute value for a third point of the point cloud, wherein the third point of the point cloud is a decoded point that is third closest to the current point of the point cloud; determining a fourth attribute value for a fourth point of the point cloud, wherein the fourth point of the point cloud is either a decoded point that is farther from the current point than the third point or a decoded point that is the same distance from the current point as the third point; determining a candidate set of prediction values ​​for an attribute value of a current point of the point cloud, wherein the current point defines an intersection of an x-axis, a y-axis, and a z-axis, wherein the x-axis, the y-axis, and the z-axis form an xy plane, an xz plane, and a yz plane, wherein to determine the candidate set of prediction values ​​for the current point of the point cloud, the one or more processors are further configured to generate the candidate set of prediction values ​​having a subset of the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value based on a comparison of a position of the second point with a position of the third point; as well as The attribute value of the current point is decoded based on the candidate set of prediction values.