Predictor index signaling for prediction transform in geometry-based point cloud compression
By decoupling the predictor index and the reconstruction of the neighborhood point attribute values, a joint decoding technique is adopted to solve the problem of high computational complexity in point cloud compression and improve the decoding efficiency.
Patent Information
- Application Number
- CN202180017635.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-06
- Filing Date
- 2021-04-07
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-04-07
AI Technical Summary
In the prior art, the point cloud compression process couples the analysis of the predictor index of the current point with the reconstruction of the attribute values of the neighboring points, resulting in high computational complexity and affecting decoding efficiency.
By decoupling the parsing of predictor index and the reconstruction of attribute values of neighborhood points, a joint decoding technique is adopted to independently parse the predictor index and residual value, and compare them based on the maximum difference and threshold to reduce the computational complexity.
The computational complexity of the point cloud compression process is reduced, and the decoding efficiency and performance are improved.
Smart Images

Figure CN115211117B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application 17 / 223,780, filed April 6, 2021, U.S. Provisional Patent Application 63 / 006,634, filed April 7, 2020, U.S. Provisional Patent Application 63 / 010,519, filed April 15, 2020, U.S. Provisional Patent Application 63 / 012,557, filed April 20, 2020, U.S. Provisional Patent Application 63 / 043,596, filed June 24, 2020, U.S. Provisional Patent Application 63 / 087,774, filed October 5, 2020, and U.S. Provisional Patent Application 63 / 090,567, filed October 12, 2020, each of which is incorporated by reference in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to point cloud encoding and decoding. BACKGROUND
[0003] A point cloud is a collection of points in a three-dimensional space. The points can correspond to points on objects in the three-dimensional space. Thus, a point cloud can be used to represent the physical content of a three-dimensional space. Point clouds can have utility in various scenarios. For example, a point cloud can be used in the context of an autonomous vehicle to represent the locations of objects on a road. In another example, a point cloud can be used in the context of representing the physical content of an environment for the purpose of positioning virtual objects in an augmented reality (AR) or mixed reality (MR) application. Point cloud compression is a process for encoding and decoding point clouds. Encoding a point cloud can reduce the amount of data needed to store and transmit the point cloud. SUMMARY
[0004] Generally, this disclosure describes techniques for geometry-based point cloud compression (G-PCC), including techniques for signaling predictor indices for predicting transforms in G-PCC. As described herein, the techniques of this disclosure can decouple the signaling of a predictor index for a current point of a point cloud from the reconstruction of attribute values for points in a neighborhood of the current point. This decoupling can reduce computational complexity in the parsing process of determining whether a bitstream includes a syntax element indicating a predictor index for the current point.
[0005] In one example, the present disclosure describes a method for decoding point cloud data, the method comprising: applying an inverse function to a set of one or more jointly decoded values based on a comparison of a maximum difference value and a threshold to recover: (i) a residual value of an attribute value for a current point of the point cloud data, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on attribute values of one or more neighboring points; determining a predicted attribute value based on the predictor index; and reconstructing the attribute value of the current point based on the residual value and the predicted attribute value.
[0006] In another example, the present disclosure describes a method for encoding point cloud data, the method comprising: determining a predictor index for a current point of the point cloud based on a maximum difference being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determining a set of residual values for the attribute values of the current point; applying a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and signaling the jointly decoded values.
[0007] In another example, the present disclosure describes a device for decoding a point cloud, the device comprising: a memory for storing data representing the point cloud; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: apply an inverse function to a set of one or more jointly decoded values based on a comparison of a maximum difference value and a threshold to recover: (i) a residual value for an attribute value of a current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on attribute values of one or more neighboring points; determine a predicted attribute value based on the predictor index; and reconstruct the attribute value of the current point based on the residual value and the predicted attribute value.
[0008] In another example, the present disclosure describes a device for encoding a point cloud, the device comprising: a memory for storing data representing the point cloud; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: determine a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determine a set of residual values for the attribute values of the current point; apply a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and signal the jointly decoded values.
[0009] In another example, the present disclosure describes a device for decoding a point cloud, the device comprising: a device for decoding a point cloud, the device comprising: a unit for applying an inverse function to a set of one or more jointly decoded values based on a comparison of a maximum difference value and a threshold to recover: (i) a residual value of an attribute value for a current point of the point cloud data, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on attribute values of one or more neighboring points; a unit for determining a predicted attribute value based on the predictor index; and a unit for reconstructing the attribute value of the current point based on the residual value and the predicted attribute value.
[0010] In another example, the present disclosure describes a device for encoding a point cloud, the device comprising: a unit for determining a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; a unit for determining a set of residual values for the attribute values of the current point; a unit for applying a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and a unit for signaling the jointly decoded values.
[0011] In another example, the present disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: apply an inverse function to a set of one or more jointly decoded values based on a comparison of a maximum difference value and a threshold to recover: (i) a residual value for an attribute value of a current point in point cloud data, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on attribute values of one or more neighboring points; determine a predicted attribute value based on the predictor index; and reconstruct the attribute value of the current point based on the residual value and the predicted attribute value.
[0012] In another example, the present disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determine a set of residual values for the attribute values of the current point; apply a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and signal the jointly decoded values.
[0013] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram illustrating an example encoding and decoding system that may perform the techniques of this disclosure.
[0015] Figure 2 is a block diagram showing an example overview of a Geometric Point Cloud Compression (G-PCC) encoder.
[0016] Figure 3 is a block diagram showing an example overview of a G-PCC decoder.
[0017] Figure 4 is a conceptual diagram illustrating an example of applying prediction in level of detail (LoD) order.
[0018] Figure 5 is a flow chart illustrating an example process for determining an attribute predictor for a point of a point cloud.
[0019] Figure 6is a flow chart illustrating an example process for determining attribute prediction values for points of a point cloud, in accordance with one or more techniques of this disclosure.
[0020] Figure 7 is a flow chart illustrating an example process for determining attribute prediction values for points of a point cloud, in accordance with one or more techniques of this disclosure.
[0021] Figure 8 is a conceptual diagram illustrating example pseudo-code snippets for the parsing phase and the reconstruction phase, in accordance with one or more techniques of this disclosure.
[0022] Figure 9 is a conceptual diagram illustrating example pseudo-code snippets for the parsing phase and the reconstruction phase, in accordance with one or more techniques of this disclosure.
[0023] Figure 10 is a flow chart illustrating an example process for determining attribute prediction values for points of a point cloud, in accordance with one or more techniques of this disclosure.
[0024] Figure 11 is a flowchart illustrating an example process in accordance with one or more techniques of this disclosure.
[0025] Figure 12 is a flowchart illustrating an example encoding process in accordance with one or more techniques of this disclosure.
[0026] Figure 13 is a flowchart illustrating an example decoding process in accordance with one or more techniques of this disclosure.
[0027] Figure 14 is a flowchart illustrating an example process in accordance with one or more techniques of this disclosure.
[0028] Figure 15 is a flowchart illustrating an example process in accordance with one or more techniques of this disclosure.
[0029] Figure 16 is a conceptual diagram illustrating an example ranging system that may be used with one or more techniques of this disclosure.
[0030] Figure 17 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of this disclosure may be employed.
[0031] Figure 18 is a conceptual diagram illustrating an example augmented reality system in which one or more techniques of this disclosure may be used.
[0032] Figure 19is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure may be employed. DETAILED DESCRIPTION
[0033] If the maximum change in the attribute value of the points in the neighborhood of the current point in the point cloud is less than a predetermined threshold, the G-PCC encoder may signal the predictor index for the current point. The G-PCC decoder may use the predictor index for the current point to determine the predicted attribute value for the current point. The G-PCC decoder may reconstruct the attribute value of the current point based on the residual value for the current point and the predicted attribute value for the current point.
[0034] Traditionally, a G-PCC encoder signals a predictor index only when the maximum change in the property values of the points in the neighborhood of the current point is less than a threshold. Therefore, in order for a G-PCC decoder to determine whether the bitstream includes a predictor index for the current point, the property values of the points in the neighborhood of the current point must be fully reconstructed. In this way, the parsing of the predictor index for the current point and the reconstruction of the property values of the points in the neighborhood of the current point are coupled. This coupling can introduce computational complexity. In contrast, if the parsing of the predictor index for the current point and the reconstruction of the property values of the points in the neighborhood of the current point are decoupled, the G-PCC decoder may be able to parse the bitstream and later reconstruct the property values of the points.
[0035] This disclosure describes techniques that can decouple the parsing of a predictor index for a current point from the reconstruction of attribute values for points in the neighborhood of the current point. Thus, the techniques of this disclosure can reduce computational complexity. For example, the predictor index and residual value for the current point can be jointly decoded. The jointly decoded value for the current point is signaled independently of the attribute values for the points in the neighborhood of the current point.
[0036] Figure 1 is a block diagram illustrating an example encoding and decoding system 100 that can implement the techniques of this disclosure. In summary, the techniques of this disclosure relate to decoding (encoding and / or decoding) point cloud data, i.e., to support point cloud compression. Generally, point cloud data includes any data used to process a point cloud. The decoding can be effective in compressing and / or decompressing point cloud data.
[0037] like Figure 1 As shown in FIG, system 100 includes a source device 102 and a destination device 116. Source device 102 provides encoded point cloud data to be decoded by destination device 116. Figure 1In the example of FIG. 1, source device 102 provides point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can comprise any of a wide variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 can be equipped for wireless communication.
[0038] In Figure 1 In the example of FIG. 1, source device 102 includes data source 104, memory 106, G-PCC encoder 200, and output interface 108. Destination device 116 includes input interface 122, G-PCC decoder 300, memory 120, and data consumer 118. In accordance with this disclosure, G-PCC encoder 200 of source device 102 and G-PCC decoder 300 of destination device 116 can be configured to apply the techniques of this disclosure related to geometry-based point cloud compression. Thus, source device 102 represents an example of an encoding device, while destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 can include other components or arrangements. For example, source device 102 can receive data (e.g., point cloud data) from an internal or external source. Likewise, destination device 116 can interface with an external data consumer, rather than include a data consumer in the same device.
[0039] As Figure 1 System 100 as shown in FIG. 1 is merely one example. In general, other digital encoding and / or decoding devices can perform the techniques of this disclosure related to geometry point cloud compression. Source device 102 and destination device 116 are merely examples of such devices in which source device 102 generates coded data for transmission to destination device 116. This disclosure refers to “coding” devices as devices that perform coding (encoding and / or decoding) of data. Thus, G-PCC encoder 200 and G-PCC decoder 300 represent examples of coding devices (specifically, encoders and decoders, respectively). In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes encoding and decoding components. Hence, system 100 can support one-way or two-way transmission between source device 102 and destination device 116, e.g., for streaming, playback, broadcast, telephony, navigation, and other applications.
[0040] Generally, the data source 104 represents a source of data (i.e., raw, unencoded point cloud data) and can provide a sequential series of "frames" of data to the G-PCC encoder 200, which encodes the data for the frames. The data source 104 of the source device 102 can include a point cloud capture device, such as any of a variety of cameras or sensors, for example, a 3D scanner or a light detection and ranging (LIDAR) device, one or more cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or in addition, the point cloud data can be computer-generated or other data from a scanner, camera, or sensor. For example, the data source 104 can generate computer graphics-based data as source data, or produce a combination of real-time data, archived data, and computer-generated data. Thus, in some examples, the source device 102 can generate a point cloud. In each case, the G-PCC encoder 200 encodes captured, pre-captured, or computer-generated point cloud data. The G-PCC encoder 200 can generate one or more bitstreams including the encoded point cloud data. Source device 102 may then output the encoded data onto computer-readable medium 110 via output interface 108 to be received and / or retrieved by, for example, input interface 122 of destination device 116 .
[0041] The memory 106 of the source device 102 and the memory 120 of the destination device 116 can represent general purpose memory. In some examples, the memory 106 and the memory 120 can store original data, for example, original data from the data source 104 and original decoded data from the G-PCC decoder 300. In addition or alternatively, the memory 106 and the memory 120 can store software instructions executable by, for example, the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although the memory 106 and the memory 120 are shown as separate from the G-PCC encoder 200 and the G-PCC decoder 300 in this example, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 can also include internal memory for functionally similar or equivalent purposes. In addition, the memory 106 and the memory 120 can store, for example, encoded data output from the G-PCC encoder 200 and input to the G-PCC decoder 300. In some examples, portions of memory 106 and memory 120 may be allocated as one or more buffers, eg, to store raw decoded and / or encoded data. For example, memory 106 and memory 120 may store data representing a point cloud.
[0042] The computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to send the encoded data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 can modulate the transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission information according to a communication standard such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other device that can be useful for facilitating communication from the source device 102 to the destination device 116.
[0043] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.
[0044] In some examples, source device 102 may output the encoded data to a file server 114 or another intermediate storage device that may store the encoded data generated by source device 102. Destination device 116 may access the stored data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing the encoded data and transmitting the encoded data to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded data from file server 114 via any standard data connection, including an internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two, suitable for accessing the encoded data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a download transfer protocol, or a combination thereof.
[0045] The output interface 108 and the input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded data) according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded data) according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee TM ), Bluetooth TM Standards, etc.) to transmit data (such as encoded data). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include a SoC device for performing the functions attributed to G-PCC encoder 200 and / or output interface 108, and destination device 116 may include a SoC device for performing the functions attributed to G-PCC decoder 300 and / or input interface 122.
[0046] The techniques of this disclosure may be applied to encoding and decoding to support any of a variety of applications, such as communications between autonomous vehicles, communications between scanners, cameras, sensors, and processing devices (such as local or remote servers), geographic mapping, or other applications.
[0047] The input interface 122 of the destination device 116 receives the encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded bitstream may include signaling information defined by the G-PCC encoder 200 (which is also used by the G-PCC decoder 300), such as syntax elements with values describing the characteristics and / or processing of the decoded units. The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the position of a physical object. In some examples, the data consumer 118 may include a display that presents imagery based on the point cloud.
[0048] The G-PCC encoder 200 and the G-PCC decoder 300 can each be implemented as any of a variety of appropriate encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store instructions for the software in an appropriate non-transitory computer-readable medium, and use one or more processors to execute instructions in hardware to perform the technology of the present disclosure. Each of the G-PCC encoder 200 and the G-PCC decoder 300 can be included in one or more encoders or decoders, and any one of the encoders or decoders can be integrated as a part of the combined encoder / decoder (CODEC) in the corresponding device. The device including the G-PCC encoder 200 and / or the G-PCC decoder 300 can include one or more integrated circuits, microprocessors, and / or other types of devices.
[0049] The G-PCC encoder 200 and the G-PCC decoder 300 may operate in accordance with a coding standard such as the Video Point Cloud Compression (V-PCC) standard or the Geometric Point Cloud Compression (G-PCC) standard. In general, the present disclosure may relate to coding (e.g., encoding and decoding) of a picture to include the process of encoding or decoding data. The coded bitstream typically includes a series of values for syntax elements that represent coding decisions (e.g., coding modes).
[0050] In general, the present disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" may generally refer to the transmission of values for syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 may signal values for syntax elements in the bitstream. Generally, signaling refers to generating values in the bitstream. As described above, the source device 102 may transmit the bitstream to the destination device 116 in substantially real time or in non-real time (such as may occur when storing syntax elements to the storage device 112 for later retrieval by the destination device 116).
[0051] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is investigating the potential need for standardization of point cloud coding techniques with compression capabilities significantly exceeding those of current methods, with the goal of creating a standard. The groups are collaborating on this initiative within a collaborative effort known as the 3D Graphics Group (3DG), evaluating compression technology designs proposed by experts in the field.
[0052] Point cloud compression activities are categorized into two different approaches. The first approach is "Video Point Cloud Compression" (V-PCC), which segments 3D objects and projects the segments into multiple 2D planes (which are represented as "patches" in a 2D frame), which are further decoded by a traditional 2D video codec such as the High Efficiency Video Coding (HEVC) (ITU-T H.265) codec. The second approach is "Geometry-based Point Cloud Compression" (G-PCC), which directly compresses the 3D geometry (i.e., the position of a set of points in 3D space) and the associated attribute values for each point associated with the 3D geometry. G-PCC addresses the compression of point clouds under both category 1 (static point clouds) and category 3 (dynamically acquired point clouds). The latest draft of the G-PCC standard is available in the following documents: G-PCC DIS, ISO / IEC JTC1 / SC29 / WG11w19088, Brussels, Belgium, January 2019; and the description of the codec is available in the following document: G-PCC Codec Descriptionv6, ISO / IEC JTC1 / SC29 / WG11w19091, Brussels, Belgium, January 2019.
[0053] A point cloud contains a collection of points in 3D space and can have attributes associated with the points. Attributes can be color information (such as R, G, B or Y, Cb, Cr), reflectivity information, or other attributes. In this example, the set of R, G, and B values for a point is called the attribute of that point, and each individual R, G, and B value is called the "component" of the attribute. The number of components of an attribute is called the "dimension" of the attribute. For example, an attribute with R, G, and B components has a dimension of 3. In another example, an attribute of a point can have Y, Cb, and Cr components. In another example, an attribute of a point can have a reflectivity component. A single point can have multiple attributes. For example, a point can have a first attribute with R, G, and B components and a second attribute with a reflectivity component. The components of an attribute are associated with a predefined order. For example, in an attribute with R, G, and B components, the R component may come first, followed by the G component, and then the B component.
[0054] Point clouds can be captured by various cameras or sensors (such as LIDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in a variety of applications, including but not limited to: construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to aid navigation).
[0055] The 3D space occupied by the point cloud data can be surrounded by a virtual bounding box. The positions of the points in the bounding box can be represented with a certain precision; therefore, the positions of one or more points can be quantized based on the precision. At the smallest level, the bounding box is split into voxels, which are the smallest spatial units represented by a unit cube. A voxel in a bounding box can be associated with zero, one, or more than one points. The bounding box can be split into multiple cubes / cube regions, which can be referred to as tiles. Each tile can be decoded into one or more slices. The partitioning of the bounding box into slices and tiles can be based on the number of points in each partition, or based on other considerations (for example, a particular area can be decoded as a tile). The slice area can be further partitioned using split decisions similar to those in a video codec.
[0056] Figure 2 is a block diagram illustrating an example overview of a G-PCC encoder 200 . Figure 3 is a block diagram showing an example overview of a G-PCC decoder 300. The modules shown are logical and do not necessarily correspond one-to-one with the code implemented in the reference implementation of the G-PCC codec, i.e., the TMC13 test model software studied by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11).
[0057] In both the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud positions are first decoded. The attribute decoding depends on the decoded geometry. Figure 2 and Figure 3 In the , gray-shaded modules are options typically used for Category 1 data. Diagonal cross-hatched modules are options typically used for Category 3 data. All other modules are common between Category 1 and Category 3.
[0058] For category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level of a single voxel. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than a voxel) plus a model that approximates the surface within each leaf of the pruned octree. In this way, both category 1 and category 3 data share the octree decoding mechanism, while category 1 data can additionally utilize a surface model (called trisoup decoding) to approximate the voxels within each leaf. The surface model used is a triangulation that includes 1-10 triangles per block, resulting in a triangle soup. Therefore, a category 1 geometry codec can be called a trisoup geometry codec, while a category 3 geometry codec is called an octree geometry codec.
[0059] At each node of the octree, the occupancy is signaled (when not inferred) for one or more of its children (up to eight nodes). A number of neighborhoods are specified, including: (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge or vertex with the current octree node, etc. In each neighborhood, the occupancy of the current node or its children can be predicted using the occupancy of the nodes and / or their children. For points that are sparsely populated in certain nodes of the octree, the codec also supports a direct coding mode where the 3D position of the point is directly encoded. A flag can be signaled to indicate that the direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node can also be coded.
[0060] Once the geometry is coded, the attributes corresponding to the geometry points are coded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute values representing the reconstructed point can be derived.
[0061] There are three attribute coding methods in G-PCC: region adaptive hierarchical transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (predictive transform), and interpolation-based hierarchical nearest neighbor prediction with update / boosting step (boosting transform). RAHT and boosting are typically used for category 1 data, while prediction is typically used for category 3 data. However, any method can be used for any data, and as with the geometry codec in G-PCC, the attribute coding method used to code the point cloud is specified in the bitstream.
[0062] The coding of attributes can be done in levels of detail (LoD), where with each level of detail a finer representation of the point cloud attributes can be obtained. Each level of detail can be specified based on a distance metric to neighboring nodes or based on a sampling distance.
[0063] At the G-PCC encoder 200, the residual obtained as output of the coding method for attributes is quantized. The quantized residual can be coded using context adaptive arithmetic coding.
[0064] In the example of Figure 2 G-PCC encoder 200 can include a coordinate transform unit 202, a color transform unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic encoding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, a LoD generation unit 220, a boosting unit 222, a coefficient quantization unit 224, and an arithmetic encoding unit 226.
[0065] As in Figure 2As shown in the example of , the G-PCC encoder 200 can receive a set of positions and a set of attributes. The positions can include the coordinates of a point in the point cloud. The attributes can include information about the point in the point cloud, such as a color associated with the point in the point cloud.
[0066] The coordinate transformation unit 202 can apply a transformation to the coordinates of the point to transform the coordinates from the original domain to the transformed domain. This disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 can apply a transformation to transform the color information of the attribute to a different domain. For example, the color transformation unit 204 can transform the color information from the RGB color space to the YCbCr color space.
[0067] In addition, Figure 2 In the example of , the voxelization unit 206 may voxelize the transformed coordinates. Voxelizing the transformed coordinates may include quantizing and removing some points in the point cloud. In other words, multiple points in the point cloud may be grouped into a single "voxel", which in some aspects may be considered as a point thereafter. Furthermore, the octree analysis unit 210 may generate an octree based on the voxelized transformed coordinates. In addition, in Figure 2 In the example of , the surface approximation analysis unit 212 can analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 can entropy encode syntax elements representing the octree and / or information about the surface determined by the surface approximation analysis unit 212. The G-PCC encoder 200 can output these syntax elements in the geometry bitstream.
[0068] The geometry reconstruction unit 216 can reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometry reconstruction unit 216 may be different from the original number of points in the point cloud. This disclosure may refer to the resulting points as reconstructed points. The attribute transfer unit 208 can transfer attributes of the original points of the point cloud to the reconstructed points of the point cloud.
[0069] Furthermore, RAHT unit 218 may apply RAHT coding to the attributes of the reconstructed points. Alternatively or in addition, LoD generation unit 220 and lifting unit 222 may apply LoD processing and lifting, respectively, to the attributes of the reconstructed points. RAHT unit 218 and lifting unit 222 may generate coefficients based on the attributes. Coefficient quantization unit 224 may quantize the coefficients generated by RAHT unit 218 or lifting unit 222. Arithmetic coding unit 226 may apply arithmetic coding to syntax elements representing the quantized coefficients. G-PCC encoder 200 may output these syntax elements in the attribute bitstream.
[0070] exist Figure 3 In the example, the G-PCC decoder 300 may include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, a LoD generation unit 316, an inverse lifting unit 318, an inverse transform coordinate unit 320 and an inverse transform color unit 322.
[0071] The G-PCC decoder 300 may obtain a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., context-adaptive binary arithmetic coding (CABAC) or other types of arithmetic decoding) to the syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream.
[0072] The octree synthesis unit 306 may synthesize the octree based on syntax elements parsed from the geometry bitstream. In the case where surface approximation is used in the geometry bitstream, the surface approximation synthesis unit 310 may determine the surface model based on syntax elements parsed from the geometry bitstream and based on the octree.
[0073] Furthermore, the geometry reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. The inverse transform coordinate unit 320 may apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0074] In addition, Figure 3 In the example of , the inverse quantization unit 308 may inverse quantize the property value. The property value may be based on syntax elements obtained from the property bitstream (eg, including syntax elements decoded by the property arithmetic decoding unit 304).
[0075] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT decoding to determine color values for the points of the point cloud based on the inverse quantized attribute values. Alternatively, the LoD generation unit 316 and the inverse lifting unit 318 may use a level of detail based technique to determine color values for the points of the point cloud.
[0076] In addition, Figure 3 In the example of FIG, the inverse color transform unit 322 can apply an inverse color transform to the color values. The inverse color transform can be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 can transform the color information from the RGB color space to the YCbCr color space. Thus, the inverse color transform unit 322 can transform the color information from the YCbCr color space to the RGB color space.
[0077] Shown Figure 2 and Figure 3 The various units of the encoder 200 and decoder 300 are described to help understand the operations performed by the encoder 200 and decoder 300. The unit can be implemented as a fixed-function circuit, a programmable circuit, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset with respect to the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality with respect to the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.
[0078] In the attribute coding aspect of the Test Model Under Consideration 13 (TMC13) of G-PCC, the LoD (Level of Detail) of each 3D point is generated based on the distance of other points relative to that 3D point. Figure 4 As shown in the example of [ 15 ], the attribute value of a 3D point in each LoD can be encoded by applying predictions in an order based on the LoD. For example, the attribute value of point P2 is predicted by calculating the distance-based weighted average of points P0, P5, and P4 that were encoded or decoded before P2. The distance-based weighted average can be called a "default predictor."
[0079] in other words, Figure 4 A set of points 400 is shown, labeled P0 through P9. The G-PCC encoder 200 may group the points 400 into two or more LoDs. The G-PCC encoder 200 may signal information associated with the points 400 so that the G-PCC decoder 300 may decode the information associated with the points without reference to information associated with any points that are at a greater LoD than the point. Figure 4In the example of FIG, points P0, P5, P4, and P2 are in LoD0, points P1, P6, and P3 are in LoD1, and points P9, P8, and P7 are in LoD2. Therefore, the G-PCC decoder 300 can decode information associated with points P0, P5, P4, and P2 without referencing information associated with points P1, P6, P3, P9, P8, and P7. The G-PCC decoder 300 can decode information associated with points P1, P6, and P3 without referencing information associated with points P9, P8, and P7, but potentially with reference to points P0, P5, P4, and P2. Similarly, the G-PCC decoder 300 may need to use information associated with points P0, P5, P4, P2, P1, P6, and P3 to decode information associated with points P9, P8, and P7.
[0080] As an alternative to using a default predictor based on a distance-based weighted average of the attribute values of points in the neighborhood of the current point as the predictor for the attribute value of the current point, the G-PCC codec also allows a predictor to be selected from multiple predictors when the change in the attribute values of the points in the neighborhood of the current point is greater than or equal to a predetermined threshold.
[0081] For example, Figure 5 is a flow chart illustrating an example process for determining an attribute predictor for a point in a point cloud. Figure 5As shown in the example of , a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate the maximum change among the neighbors of a current point (i.e., points in the neighborhood of the current point) (500). For ease of explanation, the present disclosure may refer to the maximum change among the neighbors of the current point as "maxdiff" or "MaxPredDiff". The G-PCC decoder may determine maxdiff as the difference between the maximum property value of the neighboring points and the minimum property value of the neighboring points. In addition, the G-PCC decoder may determine whether maxdiff is less than a predetermined threshold (denoted as aps_prediction_threshold) (502). In the event that maxdiff is less than the predetermined threshold (the "yes" branch of 502), the G-PCC decoder may use a default predictor (i.e., a distance-based weighted average of the property values of the neighbors of the current point) as the predictor for the property value of the current point (504). When maxdiff is less than the predetermined threshold, no predictor index is signaled. In other words, if the change is less than a predetermined threshold, only the default predictor (weighted average) is used without signaling the index of the prediction candidate. When compared to the scenario where multiple predictors are used for all points (i.e., when predIndex[] is signaled for all points), evaluating whether maxdiff is less than a predetermined threshold to determine whether a single (default) predictor or multiple predictors are available can significantly reduce the overhead for signaling the predictor index (predIndex).
[0082] If maxdiff is not less than a predetermined threshold (i.e., maxdiff is greater than or equal to the predetermined threshold) ("No" branch of 502), the index of the selected prediction candidate (i.e., predictor index) is signaled (506). The predictor index indicates whether the default predictor or the attribute value of a specific neighbor among the neighbors is used as the predictor for the attribute value of the current point. The neighbors are sorted in an order based on the LoD in terms of distance, and neighbors in any higher LoD level than the current point are not included.
[0083] For example, about Figure 4For example, when the attribute value of P2 is encoded by using multiple predictor candidates (i.e., when the G-PCC encoder 200 encodes the attribute value of P2 using a weighted average of the attribute values of P0, P5 and P4), the G-PCC encoder 200 sets the predictor index to be equal to 0. When the G-PCC encoder 200 encodes the attribute value of P2 using the attribute value of the nearest neighbor point to P2 (i.e., P4), the G-PCC encoder 200 sets the predictor index to be equal to 1. Similarly, when the G-PCC encoder 200 encodes the attribute value of P2 using the attribute values of the next nearest neighbor points to P2 (i.e., P5 and P0), the G-PCC encoder 200 sets the predictor index to be equal to 2 or 3, respectively (e.g., as shown in Table 1 below).
[0084] Table 1
[0085] Predictor Index Predicted value 0 Weighted average 1 P4 (first closest point) 2 P5 (second closest point) 3 P0 (third closest point)
[0086] The syntax structure related to this example is shown in Table 2 below, where the emphasized parts are shown with the <! >... <! > tags.
[0087] Table 2
[0088]
[0089]
[0090] MaxPredDiff[ ] is computed for each point in the point cloud. The corresponding derivation process for TMC13 is shown below.
[0091] For single-component attributes:
[0092] The variable MaxPredDiff[ i ] is computed as follows.
[0093] Let be the set of k nearest neighbors of the current point i, and let be their decoded / reconstructed attribute values. The number of nearest neighbors k shall be in the range of 1 to lifting num pred nearest neighbors. The decoded / reconstructed attribute values of the neighbors are derived according to the prediction lifting decoding process (8.3.3).
[0094] For single-component attributes:
[0095]
[0096] For multi-component attributes:
[0097]
[0098] As shown in Table 2, the syntax element lifting_adaptive_prediction_threshold indicates a predetermined threshold. In Table 2, the syntax element indicating the predictor index for the current point (pred_index) is only signaled when MaxPredDiff[i] is greater than or equal to the predetermined threshold indicated by lifting_adaptive_prediction_threshold (and other conditions).
[0099] In TMC13, the computation of MaxPredDiff[] requires attribute reconstruction of previously decoded neighbor points. This also indicates that the parsing of predIndex[] depends on attribute reconstruction. For this reason, it is not possible to decouple the parsing and reconstruction processes for attribute coding. However, this can pose a technical problem because making the parsing process independent of the reconstruction process can be beneficial for the G-PCC decoder 300 because the G-PCC decoder 300 can parse the entire bitstream in a first step and the second step only involves the reconstruction process. For each point, the variation of the neighborhood attribute values needs to be computed in order to derive predIndex, and the computational complexity of this computation is significant.
[0100] As a second example technical problem associated with the process of TMC13 for determining whether to signal the predictor index, the maxPredDiff[] computation does not take into account the case when the primary component of an attribute has a different bit-depth compared to the secondary component of the same attribute.
[0101] As a third example technical problem associated with the process of TMC13 for determining whether to signal the predictor index, the lifting_adaptive_prediction_threshold syntax element is always signaled, but can not be used in some cases. For example, if the maximum number of predictors that can be used is equal to 0, then it can not be necessary to signal the lifting_adaptive_prediction_threshold syntax element.
[0102] In this disclosure, three different methods are described to remove the dependency of attribute reconstruction on the parsing of predlndex, which automatically allows the parsing and reconstruction process for attribute coding to be decoupled. This disclosure also describes techniques to handle the maxPredDiff computation for bit-depth difference in the primary and secondary components. This disclosure also describes techniques for signaling the lifting_adaptive_prediction_threshold syntax element. Examples of the various techniques of this disclosure can be used individually or in any combination.
[0103] According to a first method for removing the dependency of attribute reconstruction on the parsing of predlndex, the predlndex parsing depends on the attribute residual values. The basic assumption of the first method is that when the local variation in the neighborhood of the current point is large (in which case multiple predictor candidates can be beneficial), the resulting residual is likely to be high. Conversely, when the local variation is low (in which case a single default predictor can be sufficient), the resulting residual is likely to be low. Following this assumption, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) can make the determination between a single predictor or multiple predictors based on the quantized attribute residual values of the neighbors. Since this decision is based only on the residual values among the already parsed information, the parsing of the attribute does not depend on the reconstruction process.
[0104] An example of such a process is shown in Figure 6 Specifically, Figure 6 is a flowchart showing an example process for determining an attribute predictor value for a point of a point cloud according to one or more techniques of this disclosure. In Figure 6 In an example of Figure 6 In an example of
[0105] Thus, in some examples, the G-PCC encoder 200 may determine, based on the quantized attribute residual values for the points in the point cloud, whether to (i) determine that the predicted attribute value for the current point of the point cloud (i.e., a predictor of the attribute value for the current point of the point cloud) is a weighted average of a set of neighboring points, or (ii) signal in the bitstream a value of an index indicating the predicted attribute value for the current point. Additionally, the G-PCC encoder 200 may determine the attribute residual value for the current point based on the predicted attribute value for the current point.
[0106] Similarly, in some examples, the G-PCC decoder 300 can determine, based on the quantized attribute residual values for the points in the point cloud, whether to (i) determine the predicted attribute value of the current point of the point cloud to be a weighted average of a set of neighboring points, or (ii) determine the predicted attribute value of the current point based on a value signaled in the bitstream. In this example, the G-PCC decoder 300 can reconstruct the attribute value of the current point based on the predicted attribute value of the current point.
[0107] In the example where the attribute value of the current point is a color attribute, the residual values for the three color components signaled can be represented as res0, res1, and res2, respectively. If |res0|+|res1|+|res2|>=Thres0, the G-PCC decoder can parse predIndex from the bitstream. Otherwise, the G-PCC decoder can use the default weighted predictor. Thres0 can be signaled in a sequence parameter set (SPS) for sequence level control, an attribute parameter set (APS) for frame level control, or an attribute parameter set (APS) for area level control. In some examples, the G-PCC decoder derives Thres0 from the attribute quantization step size. In some examples, Thres0 is set to a predetermined value.
[0108] The G-PCC algorithm of TMC13 signals the quantized residual property value (i.e., signed exponential Golomb code) in se(v). The signed Golomb code couples the sign and amplitude into a single value, i.e., the signed value is mapped to an unsigned value before the actual encoding, which is also shown in Table 3 below.
[0109] Table 3
[0110]
[0111]
[0112] If the unsigned values of the res0, res1, and res2 values are "ures0," "ures1," and "ures2" (and therefore, ures0, ures1, ures2 are greater than or equal to 0), then if ures0+ures1+ures2>=Thres1, the G-PCC encoder 200 may signal predIndex and the G-PCC decoder 300 may parse predIndex. Otherwise, the G-PCC encoder 200 and the G-PCC decoder 300 may use a default predictor. Compared to the above condition |res0|+|res1|+|res2|>=Thres0, this example may be slightly simpler because, in this example, the G-PCC decoder directly sums the unsigned residuals from the parsing process, rather than converting the unsigned residual values to signed values and then summing the resulting amplitudes, which makes parsing predIndex less complicated. The syntactic changes are shown in Table 4 below. Throughout this disclosure, deleted text utilizes <d> ...< / d> Tags are used to mark the text, and the inserted text is marked with ... Label to mark.
[0113] Table 4
[0114]
[0115] In Table 4, all components of the residuals are used to calculate the sum of the residuals. For the color attribute (dimension 3), all three components are used to calculate the sum. Similarly, for the reflectance residual (dimension 1), only one component is used to calculate the sum. It is also possible that for multi-dimensional attributes, a subset of the components is used when calculating the sum.
[0116] In some examples, we can also define a monotonically increasing function f n and the condition for predIndex (in this example for the color attribute) resolution is:
[0117] If f0(res0)+f1(res1)+f2(res2)>=Thres2, then parse predIndex.
[0118] Otherwise, the G-PCC decoder uses the default predictor. A specific subset may be f0=f1=f 2..... =f n =f, that is, the functions are exactly the same.
[0119] When xl >= x0 and yl >= y0 and zl >= z0 (in this example, 3D space is used, and in general can be extended to any multi-dimensional space), a monotonic function f can also be defined in the attribute dimension space, e.g., f(xl,yl,zl) >= f(x0,y0,z0), and the predIndex resolution can be:
[0120] If f(res0, resl, res2) >= Thres3, resolve predIndex.
[0121] Otherwise, the G-PCC decoder uses the default predictor.
[0122] The functions described in the above examples can be predetermined and can be signaled in the SPS or APS or at the region level. In some examples, the functions described in the above examples depend on the attribute quantization parameter. In some examples, the residual values can be obtained from the residuals of one or more points in the neighborhood of the current point. For example, for each component, the maximum of the residuals of the points in the neighborhood can be used to determine the residual values resl, res2, and res3. In some examples, for each component, the mean or variance of the residuals of the points in the neighborhood can be used to determine the residual values resl, res2, and res3. The neighborhood can include the current point.
[0123] According to a second method for removing the dependency of predIndex resolution on attribute reconstruction, predIndex concealment is performed based on the parity of the attribute residual values. The second method completely avoids resolving predIndex. Instead, when multiple predictors are available, the G-PCC decoder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) can derive predIndex from the parity of the attribute residual values.
[0124] Figure 7 FIG. 1 is a flowchart illustrating an example process for determining an attribute prediction value for a point of a point cloud, in accordance with one or more techniques of the present disclosure. Figure 7 The process of FIG. 1 is consistent with the second method. In Figure 7In an example of , a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate a maximum change in attribute values of neighbors of a current point of a point cloud (maxdiff) (700). Additionally, the G-PCC decoder may determine whether maxdiff is less than a predetermined threshold (e.g., aps_prediction_threshold) (702). In response to determining that maxdiff is less than the predetermined threshold (the "yes" branch of 702), the G-PCC decoder may use a default predictor (i.e., a weighted average) as a predictor for the attribute value of the current point (704). However, if maxdiff is greater than or equal to the predetermined threshold (the "no" branch of 706), the G-PCC decoder may derive a predictor index (predIdx) for the attribute value of the current point from the parity of the attribute residual values of the neighbors of the current point (706). The G-PCC decoder may use the predictor indicated by the derived predictor index as the predictor for the attribute value of the current point.
[0125] Similar to the method of TMC13, Figure 7 The decision in [ ] about whether to use a single predictor or select a predictor from among multiple predictors can be based on maxPredDiff[ ], i.e., the variance in the neighborhood of the attribute reconstruction value. When maxPredDiff[ ] is greater than or equal to a threshold (meaning that multiple predictor candidates are being used), predIndex can be derived from the residual parity at the reconstruction stage rather than at the parsing stage. Thus, there is no attribute reconstruction dependency at the parsing stage.
[0126] Figure 8 is a conceptual diagram illustrating example pseudo-code snippets for the parsing phase 800 and the reconstruction phase 802 according to one or more techniques of this disclosure. Figure 8 As shown in the fragment 804 of FIG, during the parsing stage 800, the G-PCC decoder 300 may parse the residual for each point from the point with index 0 to the point with index npoint. Figure 8 As shown in segment 806 of , during the reconstruction stage 802, the G-PCC decoder 300 may calculate maxPredDiff for each point from the point with index 0 to the point with index npoint, and derive a prediction index (predIndex) for the point if maxPredDiff is greater than or equal to a predetermined threshold.
[0127] Table 5 below shows changes to the attribute slice data syntax structure according to the second method of the present disclosure.
[0128] Table 5
[0129]
[0130]
[0131] Thus, in some examples, the G-PCC encoder 200 may derive a predictor index for a current point in the point cloud based on the parity of attribute residual values for neighboring points in the point cloud. Additionally, the G-PCC encoder 200 may determine a predicted attribute value for the current point based on the predictor index for the current point. The G-PCC encoder 200 may determine an attribute residual value for the current point based on the predicted attribute value for the current point.
[0132] Similarly, the G-PCC decoder 300 can derive a predictor index for a current point in the point cloud based on the parity of the attribute residual values for neighboring points in the point cloud. The G-PCC decoder 300 can determine a predicted attribute value for the current point based on the predictor index for the current point. Furthermore, the G-PCC decoder 300 can reconstruct an attribute value for the current point based on the predicted attribute value for the current point.
[0133] In some examples, the G-PCC decoder can use a modular operation (shown as mod) to calculate parity. For example, for a color attribute with three residuals res0, res1, and res2, four different predictor candidates can be used (such as in the common test condition of G-PCC, with predidex 0, 1, 2, and 3). In this example, the G-PCC decoder can use one of the following methods to calculate parity:
[0134] Parity = (res0 + res1 + res2) mod 4.
[0135] Parity = (|res0| + |res1| + |res2|) mod 4.
[0136] • Or, for the unsigned residuals explained previously, i.e., ures0, ures1, ures2, parity = (ures0 + ures1 + ures2) mod 4.
[0137] In general, if the total number of candidates is N, mod N can be used instead of mod 4. For each attribute, the number N (ie, the total number of candidates) can be signaled / predetermined. Different attributes can use different N values.
[0138] In some examples, parity is equal to predIndex. In some examples, parity is another one-to-one mapping between parity and predIndex (substantially different permutations between parity and predIndex are possible), which can be predefined or explicitly signaled in the SPS, APS or regional level. The one-to-one mapping between parity and predIndex can also be adaptive in a single slice. For example, as the attribute value is decoded, the statistics used for different predIndex can be updated. Based on the statistics, the mapping can be updated after decoding each point attribute.
[0139] An adaptive mapping between predIndex and parity such that parity 0 is mapped to the most likely predIndex may provide better decoding performance. Typically, smaller parities may be mapped to the most likely predIndex, for example, as shown in Table 6 below.
[0140] Table 6
[0141] predIndex Parity (bits) Most likely predindex 0(00) The second most likely predindex 1(01) 3rd most likely predindex 2(10) Least likely predindex 3(11)
[0142] In some examples, the G-PCC decoder uses an exponential moving average based probability update with the following parameters: scale and log2_window. scale specifies the precision of the probability, and log2_window is the logarithm (base 2) of the exponential moving average window size.
[0143]
[0144] For example, when the accuracy of probability estimation is N bits, the value of scale is set to (1 < <N)。也可以使用其它概率估计方法(例如,简单移动平均等)。
[0145] After deriving the predIndex for each point, the G-PCC decoder can update the probability. The G-PCC decoder can then sort the predIndex based on the probability and update the mapping between predIndex and parity. The G-PCC decoder can then use the updated mapping to map the next symbol.
[0146] The G-PCC decoder may also calculate parity based on a subset of the residual values. For example, the G-PCC decoder may use only res1 and res2, such that parity = (res1 + res2) mod N, and so on. Other examples of the present disclosure may be modified such that the G-PCC decoder uses a subset of the residual values to derive parity.
[0147] A special case may occur when the G-PCC decoder calculates parity based on a subset of residual values when N=4 (i.e., when there are four candidates in total), so the 2-bit parity can be hidden into the even / odd parity of res1 and res2, i.e., parity=2*(res1%2)+(res2%2), as shown in Table 7 below.
[0148] Table 7
[0149] res1 is an even number, res2 is an even number Parity = 0 (00) res1 is an even number, res2 is an odd number Parity = 1(01) res1 is an odd number, res2 is an even number Parity = 2(10) res1 is an odd number, res2 is an odd number Parity = 3(11)
[0150] Another case may occur when N=2, so one bit can be hidden in the odd / even parity of the residual (res0), as shown in Table 8 below.
[0151] Table 8
[0152] res0 is an even number Parity = 0 res0 is an odd number Parity = 1
[0153] The example using parity can be combined with the example where the decision on whether to use a single predictor or multiple predictors is based on the sum of the residual values (rather than the maxPredDiff value). For example, Figure 9 is a conceptual diagram illustrating example pseudo-code snippets for the parsing phase 900 and the reconstruction phase 902 according to one or more techniques of this disclosure. Figure 9 As shown in the fragment 904 of FIG, during the parsing stage 800, the G-PCC decoder 300 may parse the residual for each point from the point with index 0 to the point with index npoint. Figure 9 As shown in segment 906 of , during the reconstruction stage 902, the G-PCC decoder 300 may calculate the sum of residual values (Sum_residual) for each point from the point with index 0 to the point with index npoint, and derive a prediction index (predIndex) for the point if the Sum_residual is greater than or equal to a predetermined threshold.
[0154] In order to satisfy the constraint of matching the parity with predIndex, the G-PCC encoder 200 may need to change or modify the residual value by a small amount.Thus, for lossless reconstruction, concealment may not be possible since concealment is generally a lossy technique.
[0155] In some examples, the sum of the residual calculations used for parity derivation involves a subset or over all components.Thus, for a single-dimensional attribute, parity is calculated as parity=(res0)mod N.
[0156] As described above, in some examples, in order to match the parity with the predIndex, the G-PCC encoder 200 cannot use all allowed combinations of residual values. This can be equivalently viewed as an increase in the step size in the quantization process. To counteract this increase, when concealment is required, the actual quantization parameter value used for the attribute can be reduced by using a negative QP offset (or reducing the QP value). This QP offset can be predetermined or signaled. When the attribute has multiple components (e.g., color) compared to a single or small number of components (e.g., reflectivity), the equivalent quantization step size increase is low. Therefore, depending on the number of components, the QP offset can be set when it needs to be predetermined (or not signaled).
[0157] A third approach for removing the dependency of attribute reconstruction on the parsing of predIndex involves multiple predictor signaling with candidate reordering. In other words, the third approach involves an example for decoupling the parsing and reconstruction processes using multiple predictor candidates for each point, so that the parsing process does not affect the reconstruction. However, during the reconstruction phase, candidates can be reordered when the change in the reconstructed attribute value is above a threshold.
[0158] Thus, in some examples consistent with the third approach, the G-PCC encoder 200 may determine the order of the candidates in the candidate list based on a comparison of changes in reconstructed attribute values for adjacent points of the point cloud. The candidate list includes candidates for weighted averages of reconstructed attribute values for adjacent points and candidates for individual reconstructed attribute values for adjacent points. The G-PCC encoder 200 may determine a predictor index for the current point. Additionally, the G-PCC encoder 200 may determine a predicted attribute value for the current point based on the indicated candidate in the candidate list. The predictor index for the current point indicates the indicated candidate. The G-PCC encoder 200 may determine an attribute residual value for the current point based on the predicted attribute value for the current point.
[0159] In some examples consistent with the third approach, the G-PCC decoder 300 may determine the order of the candidates in the candidate list based on a comparison of changes in reconstructed attribute values for adjacent points of the point cloud. The candidate list includes candidates for a weighted average of the reconstructed attribute values for the adjacent points and candidates for individual reconstructed attribute values for the adjacent points. The G-PCC decoder 300 may determine a predictor index for the current point. Additionally, the G-PCC decoder 300 may determine a predicted attribute value for the current point based on the indicated candidate in the candidate list. The predictor index for the current point indicates the indicated candidate. The G-PCC decoder 300 may reconstruct the attribute value of the current point based on the predicted attribute value of the current point.
[0160] Figure 10is a flow chart illustrating an example process for determining attribute prediction values for points of a point cloud, in accordance with one or more techniques of this disclosure. Figure 10 The process is the same as the third method. Figure 10 As shown in , if the change is greater than a threshold, the weighted average predictor is moved to the end of the list (predictor index 3). This may be particularly useful for sparse point cloud data. The reordering can be predetermined, explicitly signaled in the APS, or as a region-level flag, or implicitly derived based on neighbor attribute values.
[0161] More specifically, in Figure 10 In an example of , a G-PCC decoder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) may calculate the maximum change (maxdiff) among the neighbors of the current point (1000). The G-PCC decoder may then determine whether maxdiff is less than a predetermined threshold (aps_prediction_threshold) (1002). In response to determining that maxdiff is less than the predetermined threshold (the "yes" branch of 1002), the G-PCC decoder may determine a predictor for the property value of the current point based on a first candidate list (1004). In response to determining that maxdiff is not less than (i.e., greater than or equal to) the predetermined threshold (the "no" branch of 1002), the G-PCC decoder may determine a predictor for the property value of the current point based on a second candidate list (1006). In the first candidate list, predIndex=0 corresponds to a default predictor. In the second candidate list, predIndex=3 corresponds to a default predictor. Therefore, whether maxdiff is less than a predetermined threshold or greater than or equal to the predetermined threshold, predIndex is signaled.
[0162] In some examples, there may be more than two orders. In such examples, multiple reorderings may be applied based on the maxdiff value, e.g.
[0163]
[0164] maxDiff may indicate the difference between the maximum attribute value of the adjacent points and the minimum attribute value of the adjacent points.
[0165] A fourth approach for removing the dependency of attribute reconstruction on the parsed predIndex involves jointly decoding the residual value and predIndex. In other words, another approach to decoupling the parsed and predictor index (predIndex) is to jointly decode predIndex and the residual and signal the resulting code as the residual.
[0166] Figure 11 is a flowchart illustrating an example process in accordance with one or more techniques of this disclosure. Figure 11 The process is consistent with the fourth method of the present disclosure. Figure 11 As described in , when multiple predictor processes (right branch) are called, given the residuals (res0, res1, and res2, assuming three components of the residuals (e.g., color attributes)) and the predictor index, the residuals and the predictor index (predindex) are jointly decoded using a predetermined / fixed or signaled / adaptive function "f", and the jointly decoded values (i.e., res0', res1', and res2') are signaled as the residuals in the bitstream. Figure 11 On the left branch of , in case the predictor index is not applicable, the G-PCC encoder 200 signals the residual directly in the bitstream.
[0167] At the decoder side, the residual is parsed and decoded in the normal way. During attribute reconstruction, if the G-PCC decoder 300 calls multiple predictor branches ( Figure 11 ), the G-PCC decoder 300 can use the inverse function of "f" (denoted as "invf") to recover the residuals (res0, res1, and res2) and predIndex from the jointly decoded residuals (res0', res1', and res2'). It should be noted that "f" should be a reversible function, that is, invf should exist. It can also be observed that this process can be applied to both lossless and lossy decoding, because the original residuals (res0, res1, and res2) can always be recovered during this process.
[0168] Specifically, in Figure 11In an example, a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate a maximum change (maxdiff) among neighbors of a current point (1100). The G-PCC decoder may then determine whether maxdiff is less than a predetermined threshold (aps_prediction_threshold) (1102). In response to determining that maxdiff is less than the predetermined threshold (the "yes" branch of 1002), the G-PCC decoder may determine that the predictor for the property value of the current point is a default weighted average predictor (1104). In response to determining that maxdiff is not less than (i.e., greater than or equal to) the predetermined threshold (the "no" branch of 1102), the G-PCC decoder (in the case where the G-PCC decoder is G-PCC encoder 200) may use a function that determines a jointly decoded value based on the predictor index and the residual property value of the current point (1106). In the case where the G-PCC coder is the G-PCC decoder 300 , the G-PCC coder may use an inverse function to determine a predictor index and a residual property value of a current point based on the jointly coded value.
[0169] For the number of predictors N, the following functions can be used. The function and the inverse function can be implemented in various ways. For example, in one example, the function f and the inverse function invf can be defined in equations (1) and (2) as:
[0170] f(res0,res1,res2,predindex)={sign(res0)*(|res0|*N+predindex),res1,res2}={res0',res1res2'} (1)
[0171] invf(res0',res1',res2')={|res0'|mod N,sign(res0')*(|res0|-(|res0|modN)),res1',res2'}
[0172] ={predIndex,res0,res1,res2} (2) In equations (1) and (2) above, res0 is jointly decoded with predIndex, while res1 and res2 remain unchanged. Similar content can be applied to other arrangements without loss of generality, for example, res1 is jointly decoded, while res0 and res2 remain unchanged.
[0173] In a second example of how the function and inverse function can be implemented, when N is a power of 2, two residuals (in the text below, for res0 and res1) can be jointly decoded and one residual is unaffected, as shown in equations (3) and (4) below.
[0174] f(res0,res1,res2,predindex)={sign(res0)*(|res0|*(N / 2)+predindex / 2),sign(res1)*(|res1|*(N / 2)+predindex%2),res2}={res0',res1',res2'} (3)
[0175] invf(res0',res1',res2')={(|res0'|mod(N / 2))*2+(|res1'|mod(N / 2)),sign(res0')*(|res0|-(|res0|(mod N / 2)),sign(res1')*(|res1|-(|res1|(mod N / 2)),res2'}={predIndex,res0,res1,res2} (4)
[0176] Other arrangements may also be applied. For example, in a third example of how the function and inverse function may be implemented, only one residual (e.g., res0) is jointly decoded with predindex, and the resulting jointly decoded residual is res0' (res1'=res1, res2'=res2 - unaffected), and the derivation of res0 and predindex from res0' is as follows.
[0177]
[0178] Other arrangements may also be used. Combinations of such examples may also be used.
[0179] For the G-PCC v1 standard, the maximum number of direct predictors is 3. Therefore, there can be up to 4 predictors (considering the default weighted predictor and 3 direct predictors). In order to achieve better flexibility, it is also possible to allow only direct predictors, that is, excluding the default weighted predictor value. There is an APS level flag that indicates whether only direct predictors are allowed. The G-PCC encoder 200 can signal a value sigIndex corresponding to the predictor index (predIndex). The determination of sigIndex can be signaled in the modulo operation of the joint residual. As described elsewhere in the present disclosure, the G-PCC decoder 300 can determine the predictor index based on sigIndex.
[0180] The process for jointly decoding the residual value and the predictor index can be implemented in various ways. For example, in a first example process, if the attribute has a dimension greater than 1, the last two components of the residual of the attribute are used. Otherwise, for a single attribute dimension, only the components of the residual are used. Where such joint signaling is applicable, the maximum number of predictors can range from 2 to 4. Details of the determination of the joint residual for color (multi-component) and reflectance (single-component) attributes are described below.
[0181] The following example applies to the color (multi-component) attribute when the number of predictors is equal to 4. In this example, sigIndex may be mapped to the modulo 2 of the last two components of the joint residual (ie, res1' and res2'), as described in Table 9 below.
[0182] Table 9
[0183] |res1'|mod 2 |res2'|mod 2 sigIndex 0 0 0 0 1 1 1 0 2 1 1 3
[0184] Given the (res1, res2, sigIndex) tuple (ie, the original residual and the signaled index), the G-PCC encoder 200 may compute the joint residual (res1′, res2′) using equations (5) and (6) as follows:
[0185] res1'=sign(res1)*(|res1|*2+sigIndex / 2) (5)
[0186] res2'=sign(res2)*(|res2|*2+sigIndex%2) (6)
[0187] Alternatively, the G-PCC encoder 200 may calculate the joint residual (res1′, res2′) using equations (7) and (8) as follows:
[0188] res1'=sign(res1)*(|res1|<<1+sigIndex>>1) (7)
[0189] res2'=sign(res2)*(|res2|*<<1+sigIndex&1) (8)
[0190] Accordingly, in some examples, based on the number of predictors in the predictor list being equal to 4, the G-PCC encoder 200 can determine res1’ = sign(res1)*(|res1|<<1 + sigIndex>>1) and res2’ = sign(res2)*(|res2|*<<1 + sigIndex&1), where res1’ is a first jointly coded value in a set of jointly coded values, res2’ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in a set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is a predictor index.
[0191] At the decoder side, the G-PCC decoder 300 can perform the inverse operation using Equations (9), (10), and (11) as follows:
[0192] res1 = sign(res1’)*((|res1’|-(|res1’| % 2)) / 2) (9)
[0193] res2 = sign(res2’)*((|res2’|-(|res2’| % 2)) / 2) (10)
[0194] sigIndex = (|res1’| % 2))*2 + (|res2’| % (2)) (11)
[0195] Alternatively, the G-PCC decoder 300 can perform the inverse operation using Equations (12), (13), and (14) as follows:
[0196] res1 = sign(res1’)*(|res1’|>>1) (12)
[0197] res2 = sign(res2’)*(|res2’|>>1) (13)
[0198] sigIndex = (|res1’| & 1)<<1 + (|res2’| & 1) (14)
[0199] Thus, in some examples, based on the number of predictors in the predictor list being equal to 4, the G-PCC decoder 300 can reconstruct the first residual value by calculating the following: res1 = sign(res1')*(|res1'|>>1), where res1 is the first residual value and res1' is the first jointly decoded value. The G-PCC decoder 300 can reconstruct the second residual value by calculating the following: res2 = sign(res2')*(|res2'|>>1), where res2 is the second residual value and res2' is the second jointly decoded value. The G-PCC decoder 300 can determine the predictor index by calculating the following: sigIndex = (|res1'|&1)<<1+(|res2'|&1), where sigIndex is the predictor index, res1' is the first jointly decoded value, and res2' is the second jointly decoded value.
[0200] The following example may be applicable to the color attribute when the number of predictors is equal to 3. In this example, sigIndex may be mapped to the modulo 2 of the last two components of the joint residual (ie, res1' and res2'), as described in Table 10 below:
[0201] Table 10
[0202] |res1'|mod 2 |res2'|mod 2 sigIndex 0 - 0 1 0 1 1 1 2
[0203] Therefore, given the (res1, res2, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual (res1′, res2′) using equations (15) and (16) as follows:
[0204] res1'=sign(res1)*(|res1|*2+(sigIndex>0)) (15)
[0205] res2'=(sigIndex>0)? sign(res2)*(|res2|*2+(sigIndex-1)%2):res2 (16)
[0206] Alternatively, the G-PCC encoder 200 may calculate the joint residual using equations (17) and (18) as follows:
[0207] res1'=sign(res1)*(|res1|<<1+(sigIndex>0)) (17)
[0208] res2'=(sigIndex>0)? sign(res2)*(|res2|<<1+(sigIndex-1)):res2 (18)
[0209] Thus, in some examples, based on the number of predictors in the predictor list being equal to 3, the G-PCC encoder 200 may determine that res1′=sign(res1)*(|res1|<<1+(sigIndex>0)) and res2′=(sigIndex>0)?sign(res2)*(|res2|<<1+(sigIndex-1)):res2, where res1′ is a first jointly decoded value in the set of jointly decoded values, res2′ is a second jointly decoded value in the set of jointly decoded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
[0210] On the decoder side, the G-PCC decoder 300 may perform the inverse operation using equations (19), (20), and (21) as follows:
[0211] res1=sign(res1')*((|res1'|-(|res1'|%2)) / 2) (19)
[0212] res2=(|res1'|%2>0)? sign(res2')*((|res2'|-(|res2'|%2)) / 2):res2'(20)
[0213] sigIndex=(|res1'|%2))+((|res1'|%2>0)?(|res2'|%(2)):0) (21)
[0214] Alternatively, the G-PCC decoder 300 may perform the inverse operation using equations (22), (23), and (24) as follows:
[0215] res1=sign(res1')*(|res1'|>>1) (22)
[0216] res2=(|res1'|&1>0)? sign(res2')*(|res2'|>>1):res2' (23)
[0217] sigIndex=(|res1'|&1))+((|res1'|&1>0)?(|res2'|&1):0) (24)
[0218] Thus, in some examples, based on the number of predictors in the predictor list being equal to 3, the G-PCC decoder 300 can reconstruct the first residual value by calculating the following: res1 = sign(res1') * (|res1'|>>1), where res1 is the first residual value and res1' is the first jointly decoded value. The G-PCC decoder 300 can reconstruct the second residual value by calculating the following: res2 = (|res1'| & 1>0) ? sign(res2') * (|res2'|>>1): res2', where res2 is the second residual value and res2' is the second jointly decoded value. The G-PCC decoder 300 may determine the predictor index by calculating: sigIndex=(|res1′|&1))+((|res1′|&1>0)?(|res2′|&1):0), where sigIndex is the predictor index, res1′ is the first jointly coded value, and res2′ is the second jointly coded value.
[0219] The following example applies to the color attribute when the number of predictors is equal to 2. In this example, sigIndex can be mapped to the modulo 2 of the last two components of the joint residual (ie, res1' and res2'), as described in Table 11 below:
[0220] Table 11
[0221] |res1'|mod 2 |res2'|mod 2 sigIndex 0 - 0 1 - 1
[0222] Therefore, given the (res1, res2, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual (res1′, res2′) using equations (25) and (26) as follows:
[0223] res1'=sign(res1)*(|res1|*2+sigIndex%2) (25)
[0224] res2'=res2 (26)
[0225] Alternatively, the G-PCC encoder 200 may calculate the joint residual using equations (27) and (28) as follows:
[0226] res1'=sign(res1)*(|res1|<<1+(sigIndex&1)) (27)
[0227] res2'=res2 (28)
[0228] Thus, in some examples, based on the number of predictors being equal to 2, the G-PCC encoder 200 may determine res1′=sign(res1)*(|res1|<<1+(sigIndex&1)) and res2′=res2, where res1′ is a first jointly decoded value in the set of jointly decoded values, res2′ is a second jointly decoded value in the set of jointly decoded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
[0229] On the decoder side, the G-PCC decoder 300 may perform the inverse operation using equations (29), (30), and (31) as follows:
[0230] res1=sign(res1')*((|res1'|-(|res1'|%2)) / 2) (29)
[0231] res2=res2' (30)
[0232] sigIndex=(|res1'|%2)) (31)
[0233] Alternatively, the G-PCC decoder 300 may perform the inverse operation using equations (32), (33), and (35) as follows:
[0234] res1=sign(res1')*(|res1'|>>1) (32)
[0235] res2=res2' (33)
[0236] sigIndex=(|res1'|&1)) (34)
[0237] Thus, in some examples, based on the number of predictors in the predictor list being equal to 2, the G-PCC decoder 300 can reconstruct the first residual value by calculating the following: res1 = sign(res1')*(|res1'|>>1), where res1 is the first residual value and res1' is the first jointly decoded value. The G-PCC decoder 300 can reconstruct the second residual value by calculating the following: res2 = res2', where res2 is the second residual value and res2' is the second jointly decoded value. The G-PCC decoder 300 can determine the predictor index by calculating the following: sigIndex = (|res1'|&1)), where sigIndex is the predictor index and res1' is the first jointly decoded value.
[0238] The following example applies to the reflectivity (single component) property when the number of predictors is equal to 4. In this example, if the number of predictors is N, then given a (res, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual res' using the following equation (35):
[0239] res'=sign(res)*(|res|*N+res%N) (35)
[0240] Alternatively, the G-PCC encoder 200 may calculate the joint residual using equation (36) as follows:
[0241] res'=sign(res)*(|res|<<2+sigIndex) (36)
[0242] On the decoder side, the G-PCC decoder 300 may perform the inverse operation using equation (37) as follows:
[0243] res=sign(res')*((|res'|-(|res1'|%N)) / N) (37)
[0244] Alternatively, the G-PCC decoder 300 may perform the inverse operation using equation (38) as follows:
[0245] res=sign(res')*((|res'|>>2) (38)
[0246] The G-PCC decoder 300 may calculate sigIndex using the following equation (39):
[0247] sigIndex=res'&3 (39)
[0248] The following example applies to the reflectivity property when the number of predictors is equal to 3. In this example, given a (res, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual res' (where res1' is the intermediate residual) using equations (40) and (41) as follows:
[0249] res1'=(sigIndex>0)? (|res|<<1+(sigIndex-1)):|res| (40)
[0250] res'=sign(res)*(|res1'|<<1+(sigIndex>0)) (41)
[0251] At the decoder side, the G-PCC decoder 300 may perform the inverse operation using equations (42), (43), and (44) as follows (where res1 is the intermediate residual):
[0252] res1=|res'|>>1 (42)
[0253] res=((|res'|&1)>0)? (res1>>1):res1 (43)
[0254] sigIndex=(res'&1)+((res'&1)>0)? (res1&1) (44)
[0255] The following example applies to the reflectivity property when the number of predictors is equal to 2. In this example, given a (res, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual res' using the following equation (45):
[0256] res'=sign(res)*(|res|<<1+sigIndex) (45)
[0257] On the decoder side, the G-PCC decoder 300 may perform the inverse operation using equations (46) and (47) as follows:
[0258] res=sign(res')*((|res'|>>1) (46)
[0259] sigIndex=res'&1 (47)
[0260] exist Figure 11 In the example of FIG, when the neighborhood change (maxdiff) is greater than a predetermined threshold, the G-PCC decoder (e.g., the G-PCC encoder 200 or the G-PCC decoder 300) can use a default predictor plus up to three direct predictors. However, in the current G-PCC standard, there is no flexibility to remove the default weighted predictor.
[0261] In the present disclosure, a flag (e.g., lifting_only_direct_predictors) at the APS level (or alternatively, in the slice header) is introduced to indicate whether only direct predictors are used. If the flag is set to 1, the predictor list excludes the default predictor. Otherwise, if the flag is set to zero, the flag retains the default weighted predictor. For category 3 content, the default predictor may not be helpful when the variation is greater than a predetermined threshold. Table 12 below shows an example syntax structure that includes the lifting_only_direct_predictors syntax element.
[0262] Table 12
[0263]
[0264] As described above, the G-PCC encoder 200 may signal a value sigIndex corresponding to the predictor index (predIndex). The mapping between sigIndex and predIndex may be fixed or adaptive. In the present disclosure, the mapping and its nature (fixed or adaptive) may be predetermined and need not be signaled. For reflectivity, a fixed mapping is used in the following cases:
[0265] max_num_direct_predictors is set to 2 (or 1 or 3), and
[0266] only_direct_predictor is set to 1.
[0267] Therefore, only the first two direct neighbors can be used for prediction. Therefore, the G-PCC decoder can use the mapping shown in Table 13 below:
[0268] Table 13
[0269] predIndex sigIndex 1 0 2 1
[0270] For color attributes, adaptive mapping (or fixed mapping) is used in the following cases:
[0271] For Category 1A / 1B data:
[0272] o max_num_direct_predictors is set to 3, and
[0273] οonly_direct_predictor is set to 0.
[0274] For Category 3 data:
[0275] o max_num_direct_predictors is set to 3, and
[0276] οonly_direct_predictor is set to 1.
[0277] Category 1A and 1B data are typically dense point clouds, i.e., clouds in which points are arranged in a relatively dense manner. For example, Category 1A and Category 1B data may be point cloud data captured by an augmented reality (AR) system. In some examples, a 3D model of a person or structure may be Category 1A or Category 1B data. Category 3 data is less dense point cloud, such as point cloud captured using LIDAR, point cloud representing a 3D map, etc.
[0278] For adaptive mapping, the main motivation is to use a smaller sigIndex for the most likely predIndex (because the joint residual derived with a smaller sigIndex is more efficient for decoding), and vice versa. Therefore, in adaptive mapping, the G-PCC decoder can use the mapping shown in Table 14 below:
[0279] Table 14
[0280] predIndex sigIndex Most likely predindex 0 The second most likely predindex 1 3rd most likely predindex 2 Least likely predindex 3
[0281] To track the probability of each predIndex, the G-PCC decoder can update the probability of occurrence of different predIndex for each symbol. In some examples, the G-PCC decoder can use probability updates based on exponential moving averages, where G-PCC can use the following parameters: scale = 1024 (indicating the precision of the probability value) and log2_window = 6 (log2 value of the exponential moving average window size).
[0282]
[0283] Therefore, after deriving the predIndex for each point (where the maximum change is greater than a threshold), the G-PCC decoder can update the probability, sort the predIndex based on the probability, and update the mapping (between predIndex and parity). The G-PCC decoder can then use the updated mapping to derive the predIndex from the parity associated with the next symbol.
[0284] In general, multiple adaptive mappings can be implemented, where each mapping is associated with a specific range of maximum variation, i.e., adaptive mapping[j] can be used for class j where the maximum variation is such that threshold(j)<=maximumvariation <threshold(j+1))。使用多个类可以基于变化的程度来更好地跟踪predIndex的概率。
[0285] In the second example process for jointly decoding residual values and predictor indices, when the number of predictors is 4 and the attribute dimension is greater than 2, there is an alternative method to perform the joint decoding:
[0286] The following example of the second example process may apply when the number of predictors is equal to 4. In this example, sigIndex may be mapped to the modulo 2 of the last three components of the joint residual (ie, res1', res2', res0'), as described in Table 15 below.
[0287] Table 15
[0288] |res1'|mod 2 |res2'|mod 2 |res0'|mod 2 sigIndex 0 - - 0 1 0 - 1 1 1 0 2 1 1 1 3
[0289] Therefore, given the (res0, res1, res2, sigIndex) tuple (i.e., the original residual and the signaled index), the G-PCC encoder 200 can calculate the joint residual (res0, res1′, res2′) using equations (48), (49), and (50) as follows:
[0290] res1'=sign(res1)*(|res1|*2+(sigIndex>0)) (48)
[0291] res2'=(sigIndex>0)? sign(res2)*(|res2|*2+(sigIndex-1)%2):res2 (49)
[0292] res0'=(sigIndex>1)? sign(res0)*(|res0|*2+(sigIndex-2)%2):res0(50)
[0293] On the decoder side, the G-PCC decoder 300 may perform the inverse operation using equations (51), (52), (53), and (54) as follows:
[0294] res1=sign(res1')*((|res1'|-(|res1'|%2)) / 2) (51)
[0295] res2=(|res1'|%2>0)? sign(res2')*((|res2'|-(|res2'|%2)) / 2):res2'(52)
[0296] res0=(|res1'|%2>0&&|res2'|%2>0)? sign(res0')*((|res0'|-(|res0'|%2)) / 2):res0' (53)
[0298] sigIndex=(|res1'|%2))+((|res1'|%2>0)?(|res2'|%(2)):0)+((|res1'|%2>0&&|res2'|%2>0)?(|res0'|%(2)):0) (54)
[0299] In the third example process for jointly decoding residual values and predictor indices, for multi-component attributes, it may also be possible to perform joint decoding using only one component. This is similar to the single-component attribute decoding described in the first example process for joint decoding. The component to which joint decoding is applied can be fixed / predetermined, or signaled per slice / frame or at the sequence level.
[0300] Figure 12 is a flowchart illustrating an example encoding process according to one or more techniques of this disclosure. Figure 12 In an example, based on a maximum difference value (e.g., maxdiff) being less than a threshold (e.g., aps_prediction_threshold), the G-PCC encoder 200 may determine a predictor index for a current point in the point cloud, which indicates a predictor in a predictor list, where the predictors in the predictor list are based on attribute values of one or more neighboring points (1200). For example, the predictor index may indicate a method for determining how to predict an attribute value for the current point based on the attribute values of one or more neighboring points. The G-PCC encoder 200 may determine maxdiff as the difference between the maximum attribute value of the neighboring points and the minimum attribute value of the neighboring points. In some examples, the G-PCC encoder 200 may determine the predictors in the predictor list. Each predictor in the predictor list indicates a corresponding set of attribute values. The predictor index indicates a predictor in the predictor list that indicates a predicted attribute value for the current point. In some examples, the G-PCC encoder 200 signals a syntax element indicating whether the predictor list includes a default predictor. Furthermore, in some examples, based on the maximum number of direct predictors being greater than 0, the G-PCC encoder 200 may signal a boost adaptive prediction threshold syntax element in the bitstream indicating a boost adaptive prediction threshold.
[0301] In addition, Figure 12 In the example of , the G-PCC encoder 200 may determine a set of residual values for the property value of the current point ( 1202 ). The G-PCC encoder 200 may determine the residual value for the property value of the current point as the difference between the original value of the property value at the current point and the predicted property value for the current point.
[0302] The G-PCC encoder 200 may apply a function that generates a set of one or more jointly coded values based on: (i) a set of residual values for the property value of the current point, and (ii) a predictor index (1204). For example, in some examples where the number of predictors is equal to 4, the G-PCC encoder 200 may use equations (48), (49), and (50) to determine the jointly coded values. In other examples, the G-PCC encoder 200 may use other equations described in this disclosure to determine the jointly coded values.
[0303] The G-PCC encoder 200 may signal the jointly coded value 1206. For example, the G-PCC encoder 200 may include a syntax element in the bitstream indicating the jointly coded value.
[0304] As previously described, lifting_adaptive_prediction_threshold is a signaled threshold that the G-PCC decoder 300 may use to decide whether to use a single predictor (i.e., when only the default predictor can be used as the predictor for the property value of the current point) or to use multiple predictors (i.e., when one of the multiple predictors is used as the predictor for the property value of the current point). Table 16 below shows the syntax structure in which the lifting_adaptive_prediction_threshold syntax element is signaled.
[0305] Table 16
[0306]
[0307] The lifting_max_num_direct_predictors syntax element shown in Table 16 indicates how many direct predictors will be used with the weighted predictor when decoding using multiple predictors. A value of zero indicates that no direct predictor is used. Therefore, a value of zero indicates that a default single predictor is used for all points, or in other words, multiple predictors cannot be used.
[0308] Therefore, when lifting_max_num_direct_predictors > 0, only the lifting_adaptive_prediction_threshold syntax element should be signaled, e.g., as shown in Table 17 below.
[0309] Table 17
[0310]
[0311] Second, in the current version of the specification, there is no maximum value set for lifting_adaptive_prediction_threshold. It is proposed to have a semantic constraint on the maximum value, which is 1 < <max(attribute_bitdepth_minus1+1,attribute_bitdepth_secondary_minus1+1)。可以注意的是,当lifting_adaptive_prediction_threshold被设置为该最大值时,任何点都将不使用多个预测器。对该最大值的设置将使比特流一致性测试容易。
[0312] Therefore, the following changes are present in the following semantics for the lifting_adaptive_prediction_threshold syntax element (using … label to indicate).
[0313] lifting_adaptive_prediction_threshold specifies the threshold for enabling adaptive prediction. The value of lifting_adaptive_prediction_threshold[] should be between 0 and 1< <max(attribute_bitdepth_minus1[]+1,attribute_bitdepth_secondary_minus1[]+1) within the range.
[0314] Alternatively, the semantics for the lifting_adaptive_prediction_threshold syntax element may be defined as follows, where … Label indication changes:
[0315] lifting_adaptive_prediction_threshold specifies the threshold for enabling adaptive prediction. The value of lifting_adaptive_prediction_threshold[] should be between 0 and 1<<max(attribute_bitdepth_minus[]+1,(attribute_dimension_minus1> 0)? attribute_bitdepth_secondary_minus1[]+1):0) within the range.
[0316] Accordingly, in some examples, G-PCC encoder 200 can determine that the bitstream includes a lifting adaptive prediction threshold syntax element that indicates a lifting adaptive prediction threshold based on the maximum number of direct predictors being greater than 0. Based on a change in reconstructed attribute values of neighboring points of a current point of the point cloud being greater than the lifting adaptive prediction threshold, G-PCC encoder 200 can determine a predictor index for the current point. Further, G-PCC encoder 200 can determine a predicted attribute value for the current point based on the indicated candidate in the candidate list. The predictor index for the current point indicates the indicated candidate. G-PCC encoder 200 can determine an attribute residual value for the current point based on the predicted attribute value for the current point.
[0317] Similarly, in some examples, G-PCC decoder 300 can determine that the bitstream includes a lifting adaptive prediction threshold syntax element that indicates a lifting adaptive prediction threshold based on the first syntax element indicating that the maximum number of direct predictors is greater than 0. Based on a change in reconstructed attribute values of neighboring points of a current point of the point cloud being greater than the lifting adaptive prediction threshold, G-PCC decoder 300 can determine a candidate list that includes one or more direct predictors. Each of the one or more direct predictors is an attribute value of one of the neighboring points. G-PCC decoder 300 can determine a predictor index for the current point. Further, G-PCC decoder 300 can determine a predicted attribute value for the current point based on the indicated candidate in the candidate list. The predictor index for the current point indicates the indicated candidate. Further, G-PCC decoder 300 can reconstruct an attribute value for the current point based on the predicted attribute value for the current point.
[0318] Figure 13 is a flowchart illustrating an example decoding process in accordance with one or more techniques of this disclosure. In Figure 13 In examples of FIG. 13, based on a comparison that a maximum difference (e.g., maxdiff) is less than a threshold (e.g., aps_prediction_threshold), G-PCC decoder 300 can apply an inverse function to a set of one or more jointly coded values to recover (i) a residual value for an attribute value of a current point of point cloud data, and (ii) a predictor index that indicates a predictor in a predictor list, where the predictor in the predictor list is based on attribute values of one or more neighbor points (1300). In some examples, G-PCC decoder 300 can determine that the bitstream includes a lifting adaptive prediction threshold syntax element (e.g., lifting_adaptive_prediction_threshold) that indicates a lifting adaptive prediction threshold based on a first syntax element (e.g., lifting_max_num_direct_predictors) indicating that the maximum number of direct predictors is greater than 0.
[0319] In some examples (such as examples where the G-PCC decoder 300 is decoding a color attribute), the G-PCC decoder 300 may reconstruct a first residual value based on a first jointly decoded value in the set of jointly decoded values, reconstruct a second residual value based on a second jointly decoded value in the set of jointly decoded values, and determine a predictor index (e.g., predIndex) based on the first and second jointly decoded values in the set of jointly decoded values. For example, in some examples, the G-PCC decoder 300 may determine the predictor index based on a modulo 2 of the first jointly decoded value and the second jointly decoded value. In some examples where the number of predictors is equal to 4, the G-PCC decoder 300 may use equations (51), (52), (53), and (54) to determine the residual value and predictor index for the attribute value of the current point. In other examples, the G-PCC decoder 300 may use other equations described in the present disclosure to determine the residual value and predictor index for the attribute value of the current point. For example, in an example where the G-PCC decoder is decoding a single-dimensional attribute, the G-PCC decoder 300 may reconstruct the residual value based on the jointly decoded values in the set of jointly decoded values and determine the predictor index based on the jointly decoded values in the set of jointly decoded values. For example, the G-PCC decoder 300 may use equations (37), (38), (39), (42), (43), (44), (46), and (47) to reconstruct the residual value and determine the predictor index.
[0320] In addition, Figure 13 In an example of the present invention, the G-PCC decoder 300 may determine a predicted property value based on a predictor index (1302). For example, the G-PCC decoder 300 may generate a candidate list corresponding to different values of the predictor index. The candidate list may include candidates specifying property values for individual neighbors of the current point. In some examples, the candidate list may include a default predictor based on a weighted average of property values for two or more neighbors of the current point.
[0321] In some examples, the G-PCC decoder 300 can determine whether the predictor list includes a default predictor based on a syntax element (e.g., lifting_only_direct_predictors). In this example, each predictor in the predictor list indicates a corresponding set of property values. In this example, the G-PCC decoder 300 can determine a predictor in the predictor list based on a predictor index, where the determined predictor indicates a prediction property value.
[0322] In addition, the G-PCC decoder 300 may reconstruct the property value of the current point based on the residual value and the predicted property value (1304). For example, the G-PCC decoder 300 may add the residual value to the corresponding predicted property value to reconstruct the property value of the current point.
[0323] In some examples of the present disclosure, a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate maxPredDiff in different ways for different primary and secondary bit depths. As described elsewhere in this disclosure, maxPredDiff[] may be calculated by the maximum change across all components of the attribute. However, if the bit depths of the primary and secondary components of the attribute are different, the G-PCC decoder may need to compensate for the bit depth when calculating maxPredDiff. If the bit depth of the i-th component of the attribute is defined as BD i , then the maxPredDiff calculation can be modified as follows:
[0324]
[0325]
[0326] Currently, for prediction transforms, attributes of points in the point cloud are predicted based on the attribute values of neighboring points. Under common test conditions (such as the G-PCC reference software), four different predictors are available: a weighted predictor and its three direct neighbors (which is signaled when aps.max_num_direct_predictors = 3), and a predictor index using truncated unary (TU) binarization, where cMax = aps.max_num_direct_predictors. To reduce associated signaling, selection of the four different predictors is only made available when the change in attribute value is above a threshold signaled in the APS and the number of available neighbors is greater than 1. A neighboring point is considered "available" if its attribute was decoded before the current point and is available for use in predicting the attribute of the current point. The G-PCC decoder can use strategies (such as checking points within a radius centered on the current point) to identify neighboring points and then use the attribute values of one or more of the available neighboring points as predictors. In cases where the current point is isolated or is the first point to be decoded, the current point may have no neighbors.
[0327] In this scenario, when neighborCount=2, multiple predictors will be called. However, as shown in Table 18 below, the predictor index is still decoded using TU binarization, where cMax=aps.max_num_direct_predictors(=3), but predIndex=3 is not available. Table 18, <! >…< / !> Label indicates emphasis.
[0328] Table 18
[0329] predIndex Binary 0 0 1 10 2 110 <!>3< / !> <!>111< / !>
[0330] This redundancy case occurs rarely (and can therefore be considered a "corner" case) when there are not enough neighbors available to generate all direct predictors. Either of the following solutions (labeled Solution 1 and Solution 2) may have an impact on compression performance. However, the elimination of this redundancy (although very infrequent) may be beneficial.
[0331] Solution 1: Use truncated unary binarization, where maxval = min(neighCount, aps.max_num_direct_predictor), for example, Figure 14 as shown in the example. Figure 14 is a flowchart illustrating an example process according to one or more techniques of this disclosure. Figure 14 In an example of , a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate a maximum difference (maxdiff) among neighbors of a current point (i.e., points in a neighborhood of the current point) (1400). The G-PCC decoder may then determine whether a neighbor count (i.e., the number of available neighbors) is greater than 1 and whether maxdiff is greater than or equal to a predetermined threshold (aps_prediction_threshold) (1402). In response to determining that the neighbor count is less than or equal to 1 or maxdiff is less than the predetermined threshold (the "no" branch of 1402), the G-PCC decoder may use a default weighted average predictor to predict the property value of the current point (1404). However, if the neighbor count is greater than 1 and maxdiff is greater than or equal to the predetermined threshold (the "yes" branch of 1402), a predictor index may be signaled, and the G-PCC decoder may predict the property value of the current point based on the predictor indicated by the predictor index (1406).
[0332] The following gives Figure 14 The corresponding sample source code, where … The label indicates the added text, and <d> …< / d>Tags indicate deletion.
[0333] if(maxDiff>=aps.adaptive_prediction_threshold){
[0334] int maxMode=std::min((int)predictor.neighborCount,aps.max_num_direct_predictors);
[0335] predictor.predMode=decoder.decodePredMode(maxMode);
[0336] <d> predictor.predMode=decoder.decodePredMode(aps.max_num_direct_predictors);< / d>
[0337] }.
[0338] For neighCount=2, the binarization table is shown in Table 19 below:
[0339] Table 19
[0340] predIndex Binary 0 0 1 10 2 11 <d> 0< / d> <d> 3< / d> <d> 111< / d>
[0341] In another example, when NeighborCount>=(aps.max_num_direct_predictors) and Maxdiff>=Threshold (eg, aps_prediction_threshold), multiple predictor processes are called, as in Figure 15 as shown in the example. Figure 15 is a flowchart illustrating an example process according to one or more techniques of this disclosure. Figure 15 In an example of , a G-PCC decoder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may calculate a maximum difference (maxdiff) among neighbors of a current point (i.e., points in a neighborhood of the current point) (1500). The G-PCC decoder may then determine whether a neighbor count (i.e., the number of available neighbors) is greater than or equal to a maximum number of direct predictors (e.g., max_num_direct_predictors) and whether maxdiff is greater than or equal to a predetermined threshold (aps_prediction_threshold) (1502). In response to determining that the neighbor count is less than the maximum number of direct predictors or maxdiff is less than the predetermined threshold (the "no" branch of 1502), the G-PCC decoder may use a default weighted average predictor to predict a property value for the current point (1504). However, if the neighbor count is greater than or equal to the maximum number of direct predictors and maxdiff is greater than or equal to a predetermined threshold ("yes" branch of 1502), the predictor index may be signaled and the G-PCC decoder may predict the property value of the current point based on the predictor indicated by the predictor index (1506).
[0342] Figure 16is a conceptual diagram illustrating an example ranging system 1600 that can be used with one or more techniques of this disclosure. Figure 16 In the example of , ranging system 1600 includes illuminator 1602 and sensor 1604. Illuminator 1602 can emit light 1606. In some examples, illuminator 1602 can emit light 1606 as one or more laser beams. Light 1606 can have one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, light 1606 is not a coherent laser. When light 1606 encounters an object (such as object 1608), light 1606 creates return light 1610. Return light 1610 can include backscattered and / or reflected light. Return light 1610 can pass through lens 1611, which guides return light 1610 to create an image 1612 of object 1608 on sensor 1604. Sensor 1604 generates signal 1614 based on image 1612. Image 1612 can include a collection of points (e.g., as represented by Figure 16 (represented by the dots in image 1612).
[0343] In some examples, illuminator 1602 and sensor 1604 can be mounted on a rotating structure so that illuminator 1602 and sensor 1604 capture a 360-degree view of the environment. In other examples, ranging system 1600 can include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable illuminator 1602 and sensor 1604 to detect objects within a certain range (e.g., up to 360 degrees). Although Figure 16 The example shows only a single illuminator 1602 and sensor 1604, but the ranging system 1600 can include multiple groups of illuminators and sensors.
[0344] In some examples, illuminator 1602 generates a structured light pattern. In such examples, ranging system 1600 may include multiple sensors on which respective images of the structured light pattern are formed. Ranging system 1600 can use the differences between the images of the structured light pattern to determine the distance to object 1608 from which the structured light pattern was backscattered. When object 1608 is relatively close to sensor 1604 (e.g., 0.2 meters to 2 meters), the structured light-based ranging system can have a high level of accuracy (e.g., accuracy in the sub-millimeter range). This high level of accuracy can be useful in facial recognition applications, such as unlocking mobile devices (e.g., mobile phones, tablet computers, etc.) and for security applications.
[0345] In some examples, the ranging system 1600 includes a time-of-flight (ToF)-based system. In some examples in which the ranging system 1600 is a ToF-based system, the illuminator 1602 generates pulses of light. In other words, the illuminator 1602 can modulate the amplitude of the emitted light 1606. In such examples, the sensor 1604 detects the return light 1610 from the pulses of light 1606 generated by the illuminator 1602. The ranging system 1600 can then determine the distance to the object 1608 from which the light 1606 was backscattered based on the delay between when the light 1606 was emitted and when it was detected, and the known speed of light in air. In some examples, the illuminator 1602 can modulate the phase of the emitted light 1404 instead of (or in addition to) modulating the amplitude of the emitted light 1606. In such examples, the sensor 1604 can detect the phase of the return light 1610 from the object 1608, and determine the distance to a point on the object 1608 using the speed of light and based on the time difference between when the illuminator 1602 generated the light 1606 at a particular phase and when the sensor 1604 detected the return light 1610 at that particular phase.
[0346] In other examples, a point cloud can be generated without using an illuminator 1602. For example, in some examples, the sensor 1604 of the ranging system 1600 can include two or more optical cameras. In such examples, the ranging system 1600 can use the optical cameras to capture stereoscopic images of the environment including the object 1608. The ranging system 1600 (e.g., the point cloud generator 1620) can then calculate the differences between locations in the stereoscopic images. The ranging system 1600 can then use the differences to determine distances to the locations shown in the stereoscopic images. From these distances, the point cloud generator 1620 can generate a point cloud.
[0347] The sensor 904 can also detect other properties of the object 1608, such as color and reflectivity information. In Figure 16 In some examples, the point cloud generator 1620 can generate a point cloud based on the signals 918 generated by the sensor 1604. The ranging system 1600 and / or the point cloud generator 1620 can form part of the data source 104 Figure 1 The point cloud generator 1620 can use the color and reflectivity information to generate attributes for the points of the point cloud data.
[0348] Figure 17 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of the present disclosure can be used. In Figure 17 In some examples, the vehicle 1700 includes a laser package 1702, such as a LIDAR system. Although in Figure 17Not shown in the example of FIG, but the vehicle 1700 may also include a data source (such as data source 104 ( Figure 1 )) and a G-PCC encoder (such as G-PCC encoder 200 ( Figure 1 )).exist Figure 17 In the example of FIG, laser package 1702 emits a laser beam 1704 that reflects off a pedestrian 1706 or other object in the road. A data source of vehicle 1700 can generate a point cloud based on the signal generated by laser package 1702. A G-PCC encoder of vehicle 1700 can encode the point cloud to generate a bitstream 1708, such as geometry bitstream 203 ( Figure 2 ) and attribute bitstream 205 ( Figure 2 ). The bitstream 1708 may include significantly fewer bits than the unencoded point cloud obtained by the G-PCC encoder. The output interface of the vehicle 1700 (e.g., the output interface 108 ( Figure 1 )) can send the bitstream 1708 to one or more other devices. Therefore, the vehicle 1700 may be able to send the bitstream 1708 to other devices faster (compared to unencoded point cloud data). In addition, the bitstream 1708 may require less data storage capacity.
[0349] The techniques of this disclosure can further reduce the complexity associated with decoding the bitstream 1708. For example, jointly recovering the residual value and the predictor index from the jointly decoded values can enable the G-PCC decoder to parse the bitstream 1708 in a first step and then perform reconstruction in a second step. This can reduce the cost of implementing the G-PCC decoder.
[0350] exist Figure 17 In the example of FIG, vehicle 1700 may send a bitstream 1708 to another vehicle 1710. Vehicle 1710 may include a G-PCC decoder, such as G-PCC decoder 300 ( Figure 1 ). The G-PCC decoder of vehicle 1710 can decode bitstream 1708 to reconstruct the point cloud. Vehicle 1710 can use the reconstructed point cloud for various purposes. For example, vehicle 1710 can determine that pedestrian 1706 is in the road in front of vehicle 1700 based on the reconstructed point cloud and therefore begin to slow down (e.g., even before the driver of vehicle 1710 realizes that pedestrian 1706 is in the road). Therefore, in some examples, vehicle 1710 can perform autonomous navigation operations, generate notifications or warnings, or perform another action based on the reconstructed point cloud.
[0351] Additionally or alternatively, vehicle 1700 may send bitstream 1708 to server system 1712. Server system 1712 may use bitstream 1708 for various purposes. For example, server system 1712 may store bitstream 1708 for subsequent reconstruction of the point cloud. In this example, server system 1712 may use the point cloud along with other data (e.g., vehicle telemetry data generated by vehicle 1700) to train an autonomous driving system. In other examples, server system 1712 may store bitstream 1708 for subsequent reconstruction for a forensic accident investigation (e.g., if vehicle 1700 collides with pedestrian 1706).
[0352] Figure 18 is a conceptual diagram illustrating an example extended reality system in which one or more technologies of the present disclosure may be used. Extended reality (XR) is a term used to encompass a range of technologies including augmented reality (AR), mixed reality (MR), and virtual reality (VR). Figure 18 In the example of FIG1 , a first user 1800 is located in a first location 1802. The user 1800 wears an XR headset 1104. As an alternative to the XR headset 1804, the user 1800 can use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1804 can include a depth detection sensor (such as a LIDAR system) that detects the position of a point on an object 1806 at the location 1802. The data source of the XR headset 1804 can use the signal generated by the depth detection sensor to generate a point cloud representation of the object 1806 at the location 1802. The XR headset 1804 can include a G-PCC encoder (e.g., Figure 1 A G-PCC encoder 200 is configured to encode the point cloud to generate a bitstream 1108.
[0353] The techniques of this disclosure can further reduce the complexity associated with decoding the bitstream 1808. For example, jointly recovering the residual value and the predictor index from the jointly decoded values can enable the G-PCC decoder to parse the bitstream 1808 in a first step and then perform reconstruction in a second step. This can reduce the cost of implementing the G-PCC decoder.
[0354] The XR headset 1804 can transmit the bitstream 1808 to an XR headset 1810 worn by a user 1812 at a second location 1814 (e.g., via a network such as the Internet). The XR headset 1810 can decode the bitstream 1108 to reconstruct the point cloud. The XR headset 1810 can use the point cloud to generate an XR visualization (e.g., an AR, MR, VR visualization) representing the object 1806 at the location 1802. Thus, in some examples, such as when the XR headset 1810 generates a VR visualization, the user 1812 at the location 1814 can have a 3D immersive experience of the location 1802. In some examples, the XR headset 1810 can determine a location for a virtual object based on the reconstructed point cloud. For example, the XR headset 1810 can determine, based on the reconstructed point cloud, that the environment (e.g., the location 1802) includes a flat surface, and then determine that a virtual object (e.g., a cartoon character) is to be positioned on the flat surface. The XR headset 1810 can generate an XR visualization in which the virtual object is located at the determined location. For example, the XR headset 1810 can display the cartoon character on the flat surface.
[0355] Figure 19 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure can be used. In Figure 19 In an example, a mobile device 1900 (such as a mobile phone or tablet computer) includes a depth detection sensor (such as a LIDAR system) that detects locations of points on an object 1902 in an environment of the mobile device 1900. A data source of the mobile device 1900 can generate a point cloud representation of the object 1902 using signals generated by the depth detection sensor. The mobile device 1900 can include a G-PCC encoder (e.g., a G-PCC encoder 200 of Figure 1 configured to encode a point cloud to generate a bitstream 1904. In Figure 19 In an example, the mobile device 1900 can transmit the bitstream 1904 to a remote device 1906 (such as a server system or other mobile device). The remote device 1906 can decode the bitstream 1904 to reconstruct the point cloud. The remote device 1906 can use the point cloud for various purposes. For example, the remote device 1906 can use the point cloud to generate a map of the environment of the mobile device 1900. For example, the remote device 1906 can generate a map of an interior of a building based on the reconstructed point cloud. In another example, the remote device 1906 can generate imagery (e.g., computer graphics) based on the point cloud. For example, the remote device 1906 can use points of the point cloud as vertices of polygons and use color attributes of the points as a basis for shading the polygons. In some examples, the remote device 1906 can use the point cloud to perform facial recognition.
[0356] The techniques of this disclosure can further reduce the complexity associated with decoding the bitstream 1904. For example, jointly recovering the residual value and the predictor index from the jointly decoded values can enable the G-PCC decoder to parse the bitstream 1904 in a first step and then perform reconstruction in a second step. This can reduce the cost of implementing the G-PCC decoder.
[0357] The following is a non-limiting list of aspects of one or more techniques in accordance with this disclosure.
[0358] Aspect 1A: A method for decoding point cloud data, comprising: determining, based on quantized attribute residual values for points in a point cloud, whether to (i) determine that a predicted attribute value of a current point of the point cloud is a weighted average of a set of neighboring points, or (ii) determine the predicted attribute value of the current point based on a value signaled in a bitstream; and reconstructing the attribute value of the current point based on the predicted attribute value of the current point.
[0359] Aspect 2A: A method for encoding point cloud data, comprising: determining, based on a quantized attribute residual value for a point in the point cloud, whether to (i) determine that a predicted attribute value of a current point of the point cloud is a weighted average of a set of neighboring points, or (ii) signaling in a bitstream the value of an index indicating the predicted attribute value of the current point; and determining an attribute residual value for the current point based on the predicted attribute value of the current point.
[0360] Aspect 3A: A method for decoding point cloud data, comprising: deriving a predictor index for a current point of the point cloud based on the parity of attribute residual values for adjacent points in the point cloud; determining a predicted attribute value for the current point based on the predictor index for the current point; and reconstructing an attribute value of the current point based on the predicted attribute value of the current point.
[0361] Aspect 4A: A method for encoding point cloud data, comprising: deriving a predictor index for a current point of the point cloud based on the parity of attribute residual values for adjacent points in the point cloud; determining a predicted attribute value for the current point based on the predictor index for the current point; and determining an attribute residual value for the current point based on the predicted attribute value of the current point.
[0362] Aspect 5A: A method according to Aspect 3A or Aspect 4A, wherein deriving the predictor index for the current point includes: applying a one-to-one mapping between the parity and the predictor index for the current point to determine the predictor index for the current point based on the parity of the attribute residual values for the neighboring points in the point cloud.
[0363] Aspect 6A: The method of aspect 5A, wherein the mapping is based on a probability of a prediction index among a plurality of prediction indices.
[0364] Aspect 7A: The method according to aspect 6A further comprises: determining the probability of the prediction index based on an exponential moving average.
[0365] Aspect 8A: The method of aspect 6A, further comprising: updating the probabilities after deriving the predictor index for the current point; and reordering the mapping based on the updated probabilities.
[0366] Aspect 9A: A method for decoding point cloud data, comprising: determining an order of candidates in a candidate list based on a comparison of changes in reconstructed attribute values for adjacent points of a point cloud, wherein the candidate list includes candidates for a weighted average of the reconstructed attribute values for the adjacent points and candidates for individual reconstructed attribute values for the adjacent points; determining a predictor index for a current point; determining a predicted attribute value for the current point based on an indicated candidate in the candidate list, wherein the predictor index for the current point indicates the indicated candidate; and reconstructing the attribute value of the current point based on the predicted attribute value of the current point.
[0367] Aspect 10A: A method for encoding point cloud data, comprising: determining an order of candidates in a candidate list based on a comparison of changes in reconstructed attribute values for adjacent points of a point cloud, wherein the candidate list includes candidates for weighted averages of the reconstructed attribute values for the adjacent points and candidates for individual reconstructed attribute values for the adjacent points; determining a predictor index for a current point; determining a predicted attribute value for the current point based on an indicated candidate in the candidate list, wherein the predictor index for the current point indicates the indicated candidate; and determining an attribute residual value for the current point based on the predicted attribute value for the current point.
[0368] Aspect 11A: A method for decoding point cloud data, comprising: determining a candidate list of one or more direct predictors, each of the one or more direct predictors being an attribute value of a neighboring point among neighboring points; determining a predictor index for a current point; determining a predicted attribute value for the current point based on an indicated candidate in the candidate list, wherein the predictor index for the current point indicates the indicated candidate; and reconstructing an attribute value of the current point based on the predicted attribute value of the current point.
[0369] Aspect 12A: A method for encoding point cloud data, comprising: determining a predictor index for a current point; determining a predicted attribute value for the current point based on an indicated candidate in a candidate list, wherein the predictor index for the current point indicates the indicated candidate; and determining an attribute residual value for the current point based on the predicted attribute value of the current point.
[0370] Aspect 13A: The method according to aspect 11A or 12A, wherein the standard imposes a constraint on the boost adaptive prediction threshold syntax element, the constraint requiring the boost adaptive prediction threshold syntax element to be between 0 and 1 < <max(属性比特深度,辅attribute_bitdepth)的范围内。
[0371] Aspect 14A: The method according to Aspect 13A, further comprising: generating the point cloud.
[0372] Aspect 15A: A method for decoding point cloud data, comprising: applying an inverse function to a set of one or more jointly decoded values to recover the following based on a comparison of a maximum difference value and a threshold: (i) a residual value for an attribute value of a current point, and (ii) a predictor index; determining a predicted attribute value based on the predictor index; and reconstructing the attribute value of the current point based on the residual value and the predicted attribute value.
[0373] Aspect 16A: A method according to Aspect 15A, wherein applying the inverse function includes: reconstructing a first residual value based on a first jointly decoded value in the set of jointly decoded values; reconstructing a second residual value based on a second jointly decoded value in the set of jointly decoded values; and determining a predictor index based on the first jointly decoded value and the second jointly decoded value in the set of jointly decoded values.
[0374] Aspect 17A: The method of aspect 16A, wherein the predictor index is determined based on a modulo 2 of the first jointly coded value and the second jointly coded value.
[0375] Aspect 18A: A method according to aspect 16A or 17A, wherein reconstructing the first residual value includes calculating: res1=sign(res1')*((|res1'|-(|res1'|%2)) / 2), where res1 is the first residual value and res1' is the first jointly decoded value, wherein reconstructing the first residual value includes calculating res2=(|res1'|%2>0)?sign(res2')*((|res2'|-(|res2'|%2)) / 2):res2', where res2 is the second residual value and res2' is the second jointly decoded value, wherein applying the first inverse function further includes: reconstructing a third residual value by calculating: res0=(|res1'|%2>0 && |res2'|%2>0)? sign(res0')*((|res0'|-(|res0'|%2)) / 2):res0', wherein res0 is the third residual value, res1' is the first jointly decoded value, res2' is the second jointly decoded value, res0' is the third jointly decoded value in the set of jointly decoded values, and wherein determining the predictor index comprises calculating: sigIndex=(|res1'|%2))+((|res1'|%2>0)?(|res2'|%(2)):0)+((|res1'|%2>0&&|res2'|%2>0)?(|res0'|%(2)):0), wherein sigIndex is the predictor index, res1' is the first jointly decoded value, res2' is the second jointly decoded value, and res0' is the third jointly decoded value in the set of jointly decoded values.
[0376] Aspect 19A: A method according to any one of Aspects 15A-18A, wherein determining the prediction attribute value includes: determining whether a predictor list includes a default predictor based on a syntax element, each predictor in the predictor list indicating a corresponding set of attribute values; and determining a predictor in the predictor list based on the predictor index, wherein the determined predictor indicates the prediction attribute value.
[0377] Aspect 20A: A method for encoding point cloud data comprises: determining a predictor index for a current point of the point cloud based on a comparison of a maximum difference value and a threshold, the predictor index indicating a method for determining how to predict an attribute value of the current point based on attribute values of one or more neighboring points; determining a set of residual values for the attribute value of the current point; applying a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute value of the current point, and (ii) the predictor index; and signaling the jointly decoded values.
[0378] Aspect 21A: The method of aspect 20A, wherein applying the function comprises determining: res1′=sign(res1)*(|res1|*2+(sigIndex>0))res2′=(sigIndex>0)? sign(res2)*(|res2|*2+(sigIndex-1)%2):res2res0′=(sigIndex>1)? sign(res0)*(|res0|*2+(sigIndex-2)%2):res0, where res1′ is the first jointly coded value in the set of jointly coded values, res2′ is the second jointly coded value in the set of jointly coded values, res0′ is the third jointly coded value in the set of jointly coded values, res1 is the first residual value in the set of residual values, res2 is the second residual value in the set of residual values, res0 is the third residual value in the set of residual values, and sigIndex is the predictor index.
[0379] Aspect 22A: The method according to aspects 20A-21A further includes: determining a predictor in a predictor list, wherein each predictor in the predictor list indicates a corresponding set of property values, wherein the predictor index indicates a predictor in the predictor list, and the predictor indicates a predicted property value for the current point; and signaling a syntax element, wherein the syntax element indicates whether the predictor list includes a default predictor.
[0380] Aspect 23A: A method for decoding point cloud data, comprising: determining a predicted attribute value for a current point in the point cloud based on a predictor index, based on the number of available neighbors for the current point being greater than 1 and a maximum difference being greater than or equal to a threshold; and determining an attribute value for the current point based on the predicted attribute value.
[0381] Aspect 24A: A method for encoding point cloud data, comprising: determining a predictor index for a current point of the point cloud based on the number of available neighbors for the current point being greater than 1 and a maximum difference being greater than or equal to a threshold, the predictor index indicating a method for determining how to predict a predicted attribute value for the current point based on attribute values of one or more neighboring points; determining the predicted attribute value for the current point by applying the indicated method; determining a residual attribute value for the current point based on the predicted attribute value for the current point; and signaling the predictor index and the residual attribute value for the current point.
[0382] Aspect 25A: A method for decoding point cloud data, comprising: determining a predicted attribute value for a current point of the point cloud based on a predictor index based on that the number of available neighbors for the current point is greater than a maximum number of direct predictors and a maximum difference is greater than or equal to a threshold, the predictor index indicating a method for determining how to predict an attribute value of the current point based on attribute values of one or more neighbor points for the current point; and reconstructing the attribute value of the current point based on the predicted attribute value for the current point.
[0383] Aspect 26A: A method for encoding point cloud data, comprising: determining a predictor index for a current point of the point cloud based on the number of available neighbors for the current point being greater than 1 and a maximum difference being greater than or equal to a threshold, the predictor index indicating a method for determining how to predict a predicted attribute value for the current point based on attribute values of one or more neighbor points for the current point; determining the predicted attribute value for the current point by applying the indicated method; determining a residual attribute value for the current point based on the predicted attribute value for the current point; and signaling the predictor index and the residual attribute value for the current point.
[0384] Aspect 28A: The apparatus of aspect 27A, wherein the one or more units comprise one or more processors implemented in circuitry.
[0385] Aspect 29A: The apparatus according to Aspect 27A or 28A, further comprising: a memory for storing data representing the point cloud.
[0386] Aspect 30A: The apparatus of aspects 27A-29A, wherein the apparatus comprises a decoder.
[0387] Aspect 31A: The apparatus of aspects 27A-30A, wherein the apparatus comprises an encoder.
[0388] Aspect 32A: The device of aspects 27A-31A, further comprising: means for generating the point cloud data.
[0389] Aspect 33A: The device of aspects 27A-32A, further comprising: a display for presenting imagery based on the point cloud data.
[0390] Aspect 1B: A method for decoding point cloud data, comprising: applying, based on a comparison of a maximum difference and a threshold, an inverse function to a set of one or more jointly coded values to recover (i) a residual value for an attribute value of a current point of the point cloud data and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on attribute values of one or more neighbor points; determining a predicted attribute value based on the predictor index; and reconstructing the attribute value of the current point based on the residual value and the predicted attribute value.
[0391] Aspect 2B: The method of aspect 1B, wherein applying the inverse function comprises: reconstructing a first residual value based on a first jointly coded value in the set of jointly coded values; reconstructing a second residual value based on a second jointly coded value in the set of jointly coded values; and determining the predictor index based on the first jointly coded value and the second jointly coded value in the set of jointly coded values.
[0392] Aspect 3B: The method of aspect 2B, wherein determining the predictor index comprises: determining the predictor index based on a modulo 2 of the first jointly coded value and the second jointly coded value.
[0393] Aspect 4B: The method of aspect 2B, wherein based on a number of predictors in the predictor list being equal to 4: reconstructing the first residual value comprises computing: resl = sign(resl') * (|resl'| » 1), where resl is the first residual value and resl' is the first jointly coded value; reconstructing the second residual value comprises computing: res2 = sign(res2') * (|res2'| » 1), where res2 is the second residual value and res2' is the second jointly coded value; and determining the predictor index by computing: sigIndex = (|resl'| & 1) « 1 + (|res2'| & 1), where sigIndex is the predictor index, resl' is the first jointly coded value, and res2' is the second jointly coded value.
[0394] Aspect 5B: The method according to any one of aspects 2B or 4B, wherein, based on the number of predictors in the predictor list being equal to 3: reconstructing the first residual value comprises calculating: res1=sign(res1′)*(|res1′|>>1), where res1 is the first residual value and res1′ is the first jointly decoded value; reconstructing the second residual value comprises calculating: res2=(|res1′|&1>0)?sign(res2′)*(|res2′|>>1):res2′, where res2 is the second residual value and res2′ is the second jointly decoded value; and determining the predictor index comprises calculating: sigIndex=(|res1′|&1))+((|res1′|&1>0)?(|res2′|&1):0), where sigIndex is the predictor index, res1′ is the first jointly decoded value, and res2′ is the second jointly decoded value.
[0395] Aspect 6B: A method according to any one of Aspects 2B or 4B-5B, wherein, based on the number of predictors in the predictor list being equal to 2: reconstructing the first residual value includes calculating: res1=sign(res1')*(|res1'|>>1), where res1 is the first residual value, and res1' is the first jointly decoded value; reconstructing the second residual value includes calculating: res2=res2', where res2 is the second residual value, and res2' is the second jointly decoded value; and determining the predictor index includes calculating: sigIndex=(|res1'|&1)), where sigIndex is the predictor index, and res1' is the first jointly decoded value.
[0396] Aspect 7B: A method according to any one of aspects 1B-6B, wherein determining the prediction attribute value includes: determining whether a predictor list includes a default predictor based on a syntax element, each predictor in the predictor list indicating a corresponding set of attribute values; and determining a predictor in the predictor list based on the predictor index, wherein the determined predictor indicates the prediction attribute value.
[0397] Aspect 8B: A method according to any one of Aspects 1B-7B, wherein the threshold is a boost adaptive prediction threshold, and the method further includes: determining that the bitstream includes a boost adaptive prediction threshold syntax element indicating the boost adaptive prediction threshold based on the first syntax element indicating that the maximum number of direct predictors is greater than 0.
[0398] Aspect 9B: The method of any one of aspects 1B or 7B-8B, wherein applying the inverse function comprises reconstructing a residual value based on a jointly coded value of the set of jointly coded values, and determining the predictor index based on the jointly coded value of the set of jointly coded values.
[0399] Aspect 10B: A method of encoding point cloud data, comprising: determining a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a list of predictors, wherein the predictors in the list of predictors are based on attribute values of one or more neighbor points; determining a set of residual values for the attribute values of the current point; applying a function that generates a set of one or more jointly coded values based on (i) the set of residual values for attribute values of the current point, and (ii) the predictor index; and signaling the jointly coded values.
[0400] Aspect 11B: The method of aspect 10B, wherein, based on a number of predictors in the list of predictors being equal to 4, applying the function comprises determining: res1’ = sign(res1)*(|res1|«1 + sigIndex»1), res2’ = sign(res2)*(|res2|*«1 + sigIndex&1), where res1’ is a first jointly coded value of the set of jointly coded values, res2’ is a second jointly coded value of the set of jointly coded values, res1 is a first residual value of the set of residual values, res2 is a second residual value of the set of residual values, and sigIndex is the predictor index.
[0401] Aspect 12B: The method of any one of aspects 10B-11B, wherein, based on a number of predictors in the list of predictors being equal to 3, applying the function comprises determining: res1’ = sign(res1)*(|res1|«1 + (sigIndex>0)), res2’ = (sigIndex>0)? sign(res2)*(|res2|«1 + (sigIndex-1)): res2, where res1’ is a first jointly coded value of the set of jointly coded values, res2’ is a second jointly coded value of the set of jointly coded values, res1 is a first residual value of the set of residual values, res2 is a second residual value of the set of residual values, and sigIndex is the predictor index.
[0402] Aspect 13B: A method according to any one of Aspects 10B-12B, wherein, based on the number of predictors in the predictor list being equal to 2, applying the function includes determining: res1'=sign(res1)*(|res1|<<1+(sigIndex&1)), res2'=res2, wherein res1' is the first jointly decoded value in the set of jointly decoded values, res2' is the second jointly decoded value in the set of jointly decoded values, res1 is the first residual value in the set of residual values, res2 is the second residual value in the set of residual values, and sigIndex is the predictor index.
[0403] Aspect 14B: The method according to any one of Aspects 10B-13B further includes: determining a predictor in a predictor list, wherein each predictor in the predictor list indicates a corresponding set of property values, wherein the predictor index indicates a predictor in the predictor list that indicates a predicted property value for the current point; and signaling a syntax element, wherein the syntax element indicates whether the predictor list includes a default predictor.
[0404] Aspect 15B: A method according to any one of Aspects 10B-14B, wherein the threshold is a boost adaptive prediction threshold, and the method further comprises: based on the maximum number of direct predictors being greater than 0, signaling a boost adaptive prediction threshold syntax element indicating the boost adaptive prediction threshold in the bitstream.
[0405] Aspect 16B: An apparatus for decoding a point cloud, comprising: a memory for storing data representing the point cloud; and one or more processors coupled to the memory and implemented in a circuit, the one or more processors being configured to: apply an inverse function to a set of one or more jointly decoded values based on a comparison of a maximum difference value and a threshold to recover: (i) a residual value for an attribute value of a current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on the attribute values of one or more neighboring points; determine a predicted attribute value based on the predictor index; and reconstruct the attribute value of the current point based on the residual value and the predicted attribute value.
[0406] Aspect 17B: An apparatus according to Aspect 16B, wherein, as part of applying the inverse function, the one or more processors are configured to: reconstruct a first residual value based on a first jointly decoded value in the set of jointly decoded values; reconstruct a second residual value based on a second jointly decoded value in the set of jointly decoded values; and determine a predictor index based on the first jointly decoded value and the second jointly decoded value in the set of jointly decoded values.
[0407] Aspect 18B: The apparatus of Aspect 17B, wherein, as part of determining the predictor index, the one or more processors are configured to determine the predictor index based on a modulo 2 of the first jointly decoded value and the second jointly decoded value.
[0408] Aspect 19B: An apparatus according to Aspect 17B, wherein, based on the number of predictors in the predictor list being equal to 4: as part of reconstructing the first residual value, the one or more processors are configured to calculate: res1 = sign(res1')*(|res1'|>>1), where res1 is the first residual value, and res1' is the first jointly decoded value; as part of reconstructing the second residual value, the one or more processors are configured to calculate: res2 = sign(res2')*(|res2'|>>1), where res2 is the second residual value, and res2' is the second jointly decoded value; and as part of determining the predictor index, the one or more processors are configured to calculate: sigIndex = (|res1'|&1)<<1+(|res2'|&1), where sigIndex is the predictor index, res1' is the first jointly decoded value, and res2' is the second jointly decoded value.
[0409] Aspect 20B: An apparatus according to any one of Aspects 17B or 19B, wherein, based on the number of predictors in the predictor list being equal to 3: as part of reconstructing the first residual value, the one or more processors are configured to calculate: res1 = sign(res1')*(|res1'|>>1), where res1 is the first residual value, and res1' is the first jointly decoded value; as part of reconstructing the second residual value, the one or more processors are configured to calculate: res2 = (|res1'|&1>0)? sign(res2')*(|res2'|>>1):res2', wherein res2 is the second residual value and res2' is the second jointly decoded value; and as part of determining the predictor index, the one or more processors are configured to calculate: sigIndex=(|res1'|&1))+((|res1'|&1>0)?(|res2'|&1):0), wherein sigIndex is the predictor index, res1' is the first jointly decoded value, and res2' is the second jointly decoded value.
[0410] Aspect 21B: An apparatus according to any one of Aspects 17B or 19B-20B, wherein, based on the number of predictors in the predictor list being equal to 2: as part of reconstructing the first residual value, the one or more processors are configured to calculate: res1 = sign(res1')*(|res1'|>>1), where res1 is the first residual value, and res1' is the first jointly decoded value; as part of reconstructing the second residual value, the one or more processors are configured to calculate: res2 = res2', where res2 is the second residual value, and res2' is the second jointly decoded value; and as part of determining the predictor index, the one or more processors are configured to calculate: sigIndex = (|res1'|&1)), where sigIndex is the predictor index, and res1' is the first jointly decoded value.
[0411] Aspect 22B: An apparatus according to any one of Aspects 16B-21B, wherein, as part of determining the prediction attribute value, the one or more processors are configured to: determine whether a predictor list includes a default predictor based on a grammatical element, each predictor in the predictor list indicating a corresponding set of attribute values; and determine a predictor in the predictor list based on the predictor index, wherein the determined predictor indicates the prediction attribute value.
[0412] Aspect 23B: An apparatus according to any one of Aspects 16B-22B, wherein the threshold is a boost adaptive prediction threshold, and the one or more processors are further configured to: determine, based on the first syntax element indicating that the maximum number of direct predictors is greater than 0, that the bitstream includes a boost adaptive prediction threshold syntax element indicating the boost adaptive prediction threshold.
[0413] Aspect 24B: The apparatus according to any one of Aspects 16B-23B, further comprising: a display for presenting an image based on the point cloud data.
[0414] Aspect 25B: An apparatus according to any one of Aspects 16B or 22B-24B, wherein, as part of applying the inverse function, the one or more processors are configured to: reconstruct a residual value based on a jointly decoded value in the set of jointly decoded values; and determine a predictor index based on the jointly decoded value in the set of jointly decoded values.
[0415] Aspect 26B: A device for encoding a point cloud, comprising: a memory for storing data representing the point cloud; and one or more processors coupled to the memory and implemented in a circuit, the one or more processors being configured to: determine a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determine a set of residual values for the attribute values of the current point; apply a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and signal the jointly decoded values.
[0416] Aspect 27B: An apparatus according to Aspect 26B, wherein, based on the number of predictors in the predictor list being equal to 4, as part of applying the function, the one or more processors are configured to determine: res1'=sign(res1)*(|res1|<<1+sigIndex>>1), res2'=sign(res2)*(|res2|*<<1+sigIndex&1), wherein res1' is a first jointly decoded value in the set of jointly decoded values, res2' is a second jointly decoded value in the set of jointly decoded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
[0417] Aspect 28B: The device of any of aspects 26B-27B, wherein based on a number of predictors in the predictor list being equal to 3, as part of applying the function, the one or more processors are configured to determine: res1’ = sign(resl) * (|resl| « 1 + (sigIndex > 0)), res2’ = (sigIndex > 0)? sign(res2) * (|res2| « 1 + (sigIndex - 1)) : res2, where res1’ is a first jointly coded value of the set of jointly coded values, res2’ is a second jointly coded value of the set of jointly coded values, resl is a first residual value of the set of residual values, res2 is a second residual value of the set of residual values, and sigIndex is the predictor index.
[0418] Aspect 29B: The device of any of aspects 26B-28B, wherein based on a number of predictors in the predictor list being equal to 2, as part of applying the function, the one or more processors are configured to determine: res1’ = sign(resl) * (|resl| « 1 + (sigIndex & 1)), res2’ = res2, where res1’ is a first jointly coded value of the set of jointly coded values, res2’ is a second jointly coded value of the set of jointly coded values, resl is a first residual value of the set of residual values, res2 is a second residual value of the set of residual values, and sigIndex is the predictor index.
[0419] Aspect 30B: The device of any of aspects 26B-29B, wherein the one or more processors are further configured to: determine a predictor in a predictor list, wherein each predictor in the predictor list indicates a respective set of attribute values, wherein the predictor index indicates the predictor in the predictor list that indicates a predicted attribute value for the current point; and signal a syntax element that indicates whether the predictor list includes a default predictor.
[0420] Aspect 31B: The device of any of aspects 26B-30B, wherein the threshold is a lifting adaptive prediction threshold, and the one or more processors are further configured to: based on a maximum number of direct predictors being greater than 0, signal, in a bitstream, a lifting adaptive prediction threshold syntax element that indicates the lifting adaptive prediction threshold.
[0421] Aspect 32B: The device of any of aspects 26B-31B, further comprising: means for generating the point cloud data.
[0422] Aspect 33B: A device for decoding a point cloud comprising: means for applying, based on a comparison of a maximum difference value and a threshold, an inverse function to a set of one or more jointly coded values to recover (i) a residual value for an attribute value of a current point of point cloud data and (ii) a predictor index that indicates a predictor in a list of predictors, wherein the predictor in the list of predictors is based on attribute values of one or more neighbor points; means for determining, based on the predictor index, a predicted attribute value; and means for reconstructing the attribute value of the current point based on the residual value and the predicted attribute value.
[0423] Aspect 34B: A device for encoding a point cloud comprising: means for determining, based on a maximum difference value being less than a threshold, a predictor index for a current point of the point cloud, the predictor index indicating a predictor in a list of predictors, wherein the predictor in the list of predictors is based on attribute values of one or more neighbor points; means for determining a set of residual values for the attribute value of the current point; means for applying a function that generates a set of one or more jointly coded values based on (i) the set of residual values for the attribute value of the current point and (ii) the predictor index; and means for signaling the jointly coded values.
[0424] Aspect 35B: A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: apply, based on a comparison of a maximum difference value and a threshold, an inverse function to a set of one or more jointly coded values to recover (i) a residual value for an attribute value of a current point of point cloud data and (ii) a predictor index that indicates a predictor in a list of predictors, wherein the predictor in the list of predictors is based on attribute values of one or more neighbor points; determine, based on the predictor index, a predicted attribute value; and reconstruct the attribute value of the current point based on the residual value and the predicted attribute value.
[0425] Aspect 36B: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the following operations: determine a predictor index for a current point of the point cloud based on a maximum difference value being less than a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determine a set of residual values for the attribute values of the current point; apply a function that generates a set of one or more jointly decoded values based on: (i) the set of residual values for the attribute values of the current point, and (ii) the predictor index; and signal the jointly decoded values.
[0426] It is to be appreciated that, depending on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the techniques). Furthermore, in some examples, actions or events may be performed concurrently rather than sequentially, for example, through multithreading, interrupt handling, or multiple processors.
[0427] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media or communication media, including any media that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.
[0428] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0429] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein the term "processor" and "processing circuitry" can refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0430] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require
[0431] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for decoding point cloud data, the method comprising: Applying an inverse function to a set of one or more jointly decoded values based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of a current point of the point cloud data being greater than or equal to a threshold to recover: (i) a residual value for the attribute value of the current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on the attribute values of the one or more neighboring points; determining a predicted attribute value based on the predictor index; and The attribute value of the current point is reconstructed based on the residual value and the predicted attribute value.
2. The method according to claim 1, wherein Applying the inverse function includes: reconstructing a first residual value based on a first jointly coded value in the set of jointly coded values; reconstructing a second residual value based on a second jointly coded value in the set of jointly coded values; and The predictor index is determined based on the first jointly coded value and the second jointly coded value in the set of jointly coded values.
3. The method according to claim 2, wherein: Determining the predictor index includes determining the predictor index based on a modulo 2 of the first jointly coded value and the second jointly coded value.
4. The method according to claim 2, wherein: Based on the number of predictors in the predictor list being equal to 4: Reconstructing the first residual value includes calculating: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly decoded value; Reconstructing the second residual value includes calculating: res2=sign(res2')*(|res2'|>>1) wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and The predictor index is determined by calculating: sigIndex=(|res1'|&1)<<1+(|res2'|&1) where sigIndex is the predictor index, res1 ′ is the first jointly coded value, and res2 ′ is the second jointly coded value.
5. The method according to claim 2, wherein: Based on the number of predictors in the predictor list being equal to 3: Reconstructing the first residual value includes calculating: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly decoded value; Reconstructing the second residual value includes calculating: res2=(|res1'|&1>0)? sign(res2')*(|res2'|>>1):res2' wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and Determining the predictor index includes calculating: sigIndex=(|res1'|&1))+((|res1'|&1>0)?(|res2'|&1):0) where sigIndex is the predictor index, res1 ′ is the first jointly coded value, and res2 ′ is the second jointly coded value.
6. The method according to claim 2, wherein: Based on the number of predictors in the predictor list being equal to 2: Reconstructing the first residual value includes calculating: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly decoded value; Reconstructing the second residual value includes calculating: res2=res2' wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and Determining the predictor index includes calculating: sigIndex = (|res1'|&1)) where sigIndex is the predictor index, and res1 ′ is the first jointly coded value.
7. The method according to claim 1, wherein Determining the predicted attribute value includes: determining, based on the syntax element, whether a predictor list includes a default predictor, each predictor in the predictor list indicating a corresponding set of property values; and A predictor in the predictor list is determined based on the predictor index, wherein the determined predictor indicates the predicted property value.
8. The method according to claim 1, wherein The threshold is a lifting adaptive prediction threshold, and the method further comprises determining, based on the first syntax element indicating that the maximum number of direct predictors is greater than zero, that the bitstream includes a lifting adaptive prediction threshold syntax element indicating the lifting adaptive prediction threshold.
9. The method according to claim 1, wherein Applying the inverse function includes: reconstructing a residual value based on a jointly coded value in the set of jointly coded values; and The predictor index is determined based on the jointly coded value in the set of jointly coded values.
10. A method for encoding point cloud data, the method comprising: determining a predictor index for the current point based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of the current point in the point cloud data being greater than or equal to a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determining a set of residual values for the attribute value of the current point; applying a function that generates a set of one or more jointly coded values based on: (i) the set of residual values for the property value of the current point, and (ii) the predictor index; and The jointly coded value is signaled.
11. The method according to claim 10, wherein: Based on the number of predictors in the predictor list being equal to 4, applying the function includes determining: res1'=sign(res1)*(|res1|<<1+sigIndex>>1) res2'=sign(res2)*(|res2|*<<1+sigIndex&1) wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
12. The method according to claim 10, wherein: Based on the number of predictors in the predictor list being equal to 3, applying the function includes determining: res1'=sign(res1)*(|res1|<<1+(sigIndex>0)) res2'=(sigIndex>0)? sign(res2)*(|res2|<<1+(sigIndex-1)):res2 wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
13. The method according to claim 10, wherein: Based on the number of predictors in the predictor list being equal to 2, applying the function includes determining: res1'=sign(res1)*(|res1|<<1+(sigIndex&1)) res2'=res2 wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
14. The method according to claim 10, further comprising: determining a predictor in a predictor list, wherein each predictor in the predictor list indicates a corresponding set of property values, wherein the predictor index indicates a predictor in the predictor list that indicates a predicted property value for the current point; and A syntax element is signaled, the syntax element indicating whether the predictor list includes a default predictor.
15. The method according to claim 10, wherein The threshold is a boosted adaptive prediction threshold, and the method further comprises: A boost adaptive prediction threshold syntax element indicating the boost adaptive prediction threshold is signaled in the bitstream based on the maximum number of direct predictors being greater than zero.
16. A device for decoding point cloud data, the device comprising: A memory, configured to store the point cloud data; as well as one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: Applying an inverse function to a set of one or more jointly decoded values to recover the following based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of a current point of the point cloud data being greater than or equal to a threshold: (i) a residual value for the attribute value of the current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on the attribute values of the one or more neighboring points; determining a predicted attribute value based on the predictor index; and The attribute value of the current point is reconstructed based on the residual value and the predicted attribute value.
17. The apparatus according to claim 16, wherein As part of applying the inverse function, the one or more processors are configured to: reconstructing a first residual value based on a first jointly coded value in the set of jointly coded values; reconstructing a second residual value based on a second jointly coded value in the set of jointly coded values; as well as A predictor index is determined based on the first jointly coded value and the second jointly coded value in the set of jointly coded values.
18. The apparatus according to claim 17, wherein As part of determining the predictor index, the one or more processors are configured to determine the predictor index based on a modulo 2 of the first jointly coded value and the second jointly coded value.
19. The apparatus according to claim 17, wherein Based on the number of predictors in the predictor list being equal to 4: As part of reconstructing the first residual value, the one or more processors are configured to compute: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly decoded value; As part of reconstructing the second residual value, the one or more processors are configured to compute: res2=sign(res2')*(|res2'|>>1) wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and As part of determining the predictor index, the one or more processors are configured to calculate: sigIndex=(|res1'|&1)<<1+(|res2'|&1) where sigIndex is the predictor index, res1 ′ is the first jointly coded value, and res2 ′ is the second jointly coded value.
20. The apparatus of claim 17, wherein: Based on the number of predictors in the predictor list being equal to 3: As part of reconstructing the first residual value, the one or more processors are configured to compute: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly decoded value; As part of reconstructing the second residual value, the one or more processors are configured to compute: res2=(|res1'|&1>0)? sign(res2')*(|res2'|>>1):res2' wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and As part of determining the predictor index, the one or more processors are configured to calculate: sigIndex=(|res1'|&1))+((|res1'|&1>0)?(|res2'|&1):0) where sigIndex is the predictor index, res1 ′ is the first jointly coded value, and res2 ′ is the second jointly coded value.
21. The apparatus of claim 17, wherein: Based on the number of predictors in the predictor list being equal to 2: As part of reconstructing the first residual value, the one or more processors are configured to compute: res1=sign(res1')*(|res1'|>>1) wherein res1 is the first residual value, and res1′ is the first jointly coded value, As part of reconstructing the second residual value, the one or more processors are configured to compute: res2=res2' wherein res2 is the second residual value, and res2′ is the second jointly decoded value; and As part of determining the predictor index, the one or more processors are configured to calculate: sigIndex = (|res1'|&1)) where sigIndex is the predictor index, and res1 ′ is the first jointly coded value.
22. The apparatus of claim 16, wherein: As part of determining the predicted attribute value, the one or more processors are configured to: determining, based on the syntax element, whether a predictor list includes a default predictor, each predictor in the predictor list indicating a corresponding set of property values; as well as A predictor in the predictor list is determined based on the predictor index, wherein the determined predictor indicates the predicted property value.
23. The apparatus of claim 16, wherein: The threshold is a lifting adaptive prediction threshold, and the one or more processors are further configured to determine that the bitstream includes a lifting adaptive prediction threshold syntax element indicating the lifting adaptive prediction threshold based on the first syntax element indicating that the maximum number of direct predictors is greater than zero.
24. The apparatus of claim 16, further comprising: A display is used to present an image based on the point cloud data.
25. The apparatus of claim 16, wherein: As part of applying the inverse function, the one or more processors are configured to: reconstructing a residual value based on a jointly coded value in the set of jointly coded values; and A predictor index is determined based on the jointly coded value in the set of jointly coded values.
26. A device for encoding point cloud data, the device comprising: A memory, configured to store the point cloud data; as well as one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: determining a predictor index for the current point based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of the current point in the point cloud data being greater than or equal to a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; determining a set of residual values for the attribute value of the current point; applying a function that generates a set of one or more jointly coded values based on: (i) the set of residual values for the property value of the current point, and (ii) the predictor index; and The jointly coded value is signaled.
27. The apparatus of claim 26, wherein: Based on the number of predictors in the predictor list being equal to four, as part of applying the function, the one or more processors are configured to determine: res1'=sign(res1)*(|res1|<<1+sigIndex>>1) res2'=sign(res2)*(|res2|*<<1+sigIndex&1) wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
28. The apparatus of claim 26, wherein: Based on the number of predictors in the predictor list being equal to three, as part of applying the function, the one or more processors are configured to determine: res1'=sign(res1)*(|res1|<<1+(sigIndex>0)) res2'=(sigIndex>0)? sign(res2)*(|res2|<<1+(sigIndex-1)):res2 wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
29. The apparatus of claim 26, wherein: Based on the number of predictors in the predictor list being equal to two, as part of applying the function, the one or more processors are configured to determine: res1'=sign(res1)*(|res1|<<1+(sigIndex&1)) res2'=res2 wherein res1′ is a first jointly coded value in the set of jointly coded values, res2′ is a second jointly coded value in the set of jointly coded values, res1 is a first residual value in the set of residual values, res2 is a second residual value in the set of residual values, and sigIndex is the predictor index.
30. The apparatus of claim 26, wherein: The one or more processors are further configured to: determining a predictor in a predictor list, wherein each predictor in the predictor list indicates a corresponding set of property values, wherein the predictor index indicates a predictor in the predictor list that indicates a predicted property value for the current point; and A syntax element is signaled, the syntax element indicating whether the predictor list includes a default predictor.
31. The apparatus of claim 26, wherein The threshold is a lift adaptive prediction threshold, and the one or more processors are further configured to signal a lift adaptive prediction threshold syntax element in the bitstream indicating the lift adaptive prediction threshold based on a maximum number of direct predictors being greater than zero.
32. The apparatus of claim 26, further comprising: A device for generating the point cloud data.
33. A device for decoding point cloud data, the device comprising: A unit for applying an inverse function to a set of one or more jointly decoded values based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of a current point of the point cloud data being greater than or equal to a threshold to recover: (i) a residual value for the attribute value of the current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein the predictor in the predictor list is based on the attribute values of the one or more neighboring points; means for determining a predicted attribute value based on the predictor index; and A unit is configured to reconstruct the property value of the current point based on the residual value and the predicted property value.
34. A device for encoding point cloud data, the device comprising: means for determining a predictor index for a current point based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of the current point in the point cloud data being greater than or equal to a threshold, the predictor index indicating a predictor in a predictor list, wherein the predictors in the predictor list are based on attribute values of one or more neighboring points; means for determining a set of residual values for said attribute value of said current point; means for applying a function that generates a set of one or more jointly coded values based on: (i) the set of residual values for the property value of the current point, and (ii) the predictor index; and Means for signaling the jointly coded value.
35. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: Applying an inverse function to a set of one or more jointly decoded values based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of a current point of the point cloud data being greater than or equal to a threshold to recover: (i) a residual value of the attribute value for the current point, and (ii) a predictor index indicating a predictor in a predictor list, wherein The predictors in the predictor list are based on the attribute values of one or more neighboring points; determining a predicted attribute value based on the predictor index; as well as The attribute value of the current point is reconstructed based on the residual value and the predicted attribute value.
36. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: Determining a predictor index for the current point based on a difference between a maximum attribute value and a minimum attribute value of a plurality of neighboring points in a neighborhood of the current point in the point cloud data being greater than or equal to a threshold, the predictor index indicating a predictor in a predictor list, wherein The predictors in the predictor list are based on the attribute values of one or more neighboring points; determining a set of residual values for the attribute value of the current point; applying a function that generates a set of one or more jointly coded values based on: (i) the set of residual values for the property value of the current point, and (ii) the predictor index; as well as The jointly coded value is signaled.
Citation Information
Patent Citations
Point cloud attribute compression method based on intra-frame prediction
CN108322742A
Method and device for predictive picture encoding and decoding
CN110663254A