Inter-predictive coding using radial interpolation for predictive geometry based point cloud compression
Patent Information
- Application Number
- JP2024519710
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-21
- Filing Date
- 2022-09-22
- Publication Date
- 2025-09-17
AI Technical Summary
Inter-prediction for predictive geometry coding in point cloud compression is complex due to irregular sampling grids, leading to increased computational complexity and inefficiency.
The use of radial interpolation techniques to determine inter-predictors for point cloud data, reducing complexity by applying radial interpolation to reference points within a reference frame.
This approach improves coding accuracy and efficiency while reducing power consumption in compressing and decompressing geometric point clouds.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] This application claims priority to U.S. Patent Application No. 17 / 933,920, filed September 21, 2022, U.S. Provisional Patent Application No. 63 / 252,093, filed October 4, 2021, and U.S. Provisional Patent Application No. 63 / 254,472, filed October 11, 2021, the entire contents of each of which are incorporated by reference. U.S. Patent Application No. 17 / 933,920, filed September 21, 2022, claims the benefit of U.S. Provisional Patent Application No. 63 / 252,093, filed October 4, 2021, and U.S. Provisional Patent Application No. 63 / 254,472, filed October 11, 2021.
[0002] The present disclosure relates to encoding and decoding of point clouds. [Background technology]
[0003] A point cloud is a collection of points in three-dimensional space. The points may correspond to points on an object in the three-dimensional space. Therefore, a point cloud can be used to represent the physical contents of the three-dimensional space. A point cloud can have utility in a wide variety of situations. For example, a point cloud can be used to represent the position of an object on a road in the context of an autonomous vehicle. In another example, a point cloud can be used for the purpose of positioning a virtual object in an augmented reality (AR) or mixed reality (MR) application in the context of representing the physical contents of an environment. Point cloud compression is the process of encoding and decoding a point cloud. Encoding a point cloud can reduce the amount of data required for storage and transmission of the point cloud. Summary of the Invention
[0004] In general, this disclosure describes techniques for inter prediction for point cloud compression. In particular, this disclosure describes techniques for inter prediction for predictive geometry coding using radial interpolation.
[0005] In some embodiments, inter prediction for predictive geometry coding involves searching for inter prediction points in a reference frame. However, these points are not localized on a regular 2D array sampling grid. Interpolation for inter prediction for predictive geometry coding can be complicated due to variations in inter-sample spacing. The techniques of this disclosure can use radial interpolation to reduce the complexity of inter prediction for predictive geometry coding.
[0006] In one embodiment, the present disclosure describes a method for coding point cloud data, the method including: determining at least two reference points in a reference point cloud frame of the point cloud data; applying radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data; and coding the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0007] In another embodiment, the present disclosure describes a device for coding point cloud data, the device comprising: a memory configured to store the point cloud data; and one or more processors communicatively coupled to the memory, the one or more processors configured to determine at least two reference points in a reference point cloud frame of the point cloud data, apply radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data, and code the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0008] In another embodiment, the present disclosure describes a device for coding point cloud data, comprising: means for determining at least two reference points in a reference point cloud frame of the point cloud data; means for applying radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data; and means for coding the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0009] In another embodiment, the present disclosure describes a non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine at least two reference points in a reference point cloud frame of point cloud data, apply radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of point cloud data, and code the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0010] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram illustrating an example encoding and decoding system capable of implementing the techniques of this disclosure. [Diagram 2] FIG. 1 is a block diagram illustrating an example Geometry Point Cloud Compression (G-PCC) encoder in accordance with one or more techniques of this disclosure. [Diagram 3]FIG. 2 is a block diagram illustrating an example G-PCC decoder in accordance with one or more techniques of this disclosure. [Figure 4] 1 is a conceptual diagram illustrating an example octree partitioning for geometry coding in accordance with one or more techniques of this disclosure. [Diagram 5] FIG. 1 is a conceptual diagram illustrating an example of a prediction tree in accordance with one or more techniques of this disclosure. [Figure 6A] FIG. 1 is a conceptual diagram illustrating an example of a rotating Light Detection and Ranging (LIDAR) acquisition model in accordance with one or more techniques of the present disclosure. [Figure 6B] FIG. 1 is a conceptual diagram illustrating an example of a rotating Light Detection and Ranging (LIDAR) acquisition model in accordance with one or more techniques of the present disclosure. [Figure 7] 1 is a conceptual diagram illustrating an example of inter-prediction of a current point from a point in a reference frame, according to one or more techniques of this disclosure. [Figure 8] 1 is a flow diagram illustrating an example operation of a G-PCC decoder in accordance with one or more techniques of this disclosure. [Figure 9] 1 is a conceptual diagram illustrating an example of an additional inter predictor point obtained from an initial point having a larger azimuth angle than an inter predictor point, in accordance with one or more techniques of this disclosure. [Figure 10] FIG. 1 is a flow diagram illustrating an example radial interpolation technique in accordance with one or more aspects of the present disclosure. [Figure 11] FIG. 1 is a conceptual diagram illustrating an example distance measurement system that can be used with one or more techniques of the present disclosure. [Figure 12] FIG. 1 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of the present disclosure may be used. [Figure 13] FIG. 1 is a conceptual diagram illustrating an example extended reality system in which one or more techniques of the present disclosure can be used. [Figure 14]FIG. 1 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of the present disclosure may be used. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Inter prediction for predictive geometry coding relies on searching for inter prediction points in reference frames. In contrast to two-dimensional (2D) video coding, these points are not localized on a regular 2D array sampling grid. Furthermore, in 2D video coding, the 2D array sampling grid allows relatively simple interpolation filters to be applied on a block-by-block basis to obtain prediction blocks with sub-pixel accuracy (e.g., 1 / 2, 1 / 4, 1 / 8 sub-pixel, etc.), which provides significant coding efficiency for inter coding. However, interpolation on an irregular sampling grid (e.g., a three-dimensional point cloud) can be complicated due to variations in the spacing between samples. The techniques of this disclosure can reduce the complexity of inter prediction for predictive geometry coding by using radial interpolation. By reducing the complexity of inter prediction, the techniques of this disclosure can improve coding accuracy and efficiency and reduce the power required to code geometric point clouds.
[0013] 1 is a block diagram illustrating an example encoding and decoding system 100 capable of implementing the techniques of this disclosure. The techniques of this disclosure are generally directed to coding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. In general, point cloud data includes any data for processing a point cloud. Coding may be useful in compressing and / or decompressing the point cloud data.
[0014] As shown in Fig. 1, the system 100 includes a source device 102 and a destination device 116. The source device 102 provides encoded point cloud data to be decoded by the destination device 116. Specifically, in the embodiment of Fig. 1, the source device 102 provides the point cloud data to the destination device 116 via a computer-readable medium 110. The source device 102 and the destination device 116 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, ground or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, and the like. In some cases, the source device 102 and the destination device 116 may be capable of wireless communication.
[0015] In the embodiment of FIG. 1, the source device 102 includes a data source 104, a memory 106, a G-PCC encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a G-PCC decoder 300, a memory 120, and a data consumer 118. In accordance with the present disclosure, the G-PCC encoder 200 of the source device 102 and the G-PCC decoder 300 of the destination device 116 can be configured to apply techniques of the present disclosure related to inter-prediction in predictive geometry coding using radial interpolation. Thus, the source device 102 represents an example of an encoding device, while the destination device 116 represents an example of a decoding device. In other embodiments, the source device 102 and the destination device 116 may include other components or arrangements. For example, the source device 102 may receive data (e.g., point cloud data) from an internal or external source. Similarly, the destination device 116 may interface with an external data consumer rather than including the data consumer within the same device.
[0016] The system 100 as shown in FIG. 1 is merely an example. In general, other digital encoding and / or decoding devices may perform the techniques of this disclosure related to inter-prediction in predictive geometry coding using radial interpolation. The source device 102 and the destination device 116 are merely examples of such devices in which the source device 102 generates coded data for transmission to the destination device 116. In this disclosure, devices that perform coding (encoding and / or decoding) of data are referred to as "coding" devices. Thus, the G-PCC encoder 200 and the G-PCC decoder 300 represent examples of coding devices, specifically, encoders and decoders, respectively. In some embodiments, the source device 102 and the destination device 116 may operate in a substantially symmetrical manner such that each of the source device 102 and the destination device 116 includes an encoding component and a decoding component. Thus, the system 100 may support one-way or two-way transmission between the source device 102 and the destination device 116, for example, for streaming, playback, broadcast, telephony, navigation, and other applications.
[0017] In general, the data source 104 represents a source of data (i.e., raw, unencoded point cloud data) and can provide a sequential series of “frames” of those data to the G-PCC encoder 200, which encodes the data for those frames. The data source 104 of the source device 102 can include a point cloud capture device, such as any of a variety of cameras or sensors, e.g., a 3D scanner or a light detection and ranging (LIDAR) device, one or more video cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively, or in addition, the point cloud data can be computer-generated from a scanner, camera, sensor, or other data. For example, the data source 104 can generate computer graphics-based data as source data, or a combination of live data, archive data, and computer-generated data. In either case, the G-PCC encoder 200 encodes the captured, pre-captured, or computer-generated data. The G-PCC encoder 200 can reorder the frames from the order in which they were received (what may be referred to as a "display order") into a coding order for coding. The G-PCC encoder 200 can generate one or more bitstreams that include the encoded data. The source device 102 can then output the encoded data onto a computer-readable medium 110 via an output interface 108 for receipt and / or retrieval, for example, by an input interface 122 of a destination device 116.
[0018] The memory 106 of the source device 102 and the memory 120 of the destination device 116 may represent general purpose memories. In some embodiments, the memory 106 and the memory 120 may store raw data, e.g., raw data from the data source 104 and raw decoded data from the G-PCC decoder 300. Additionally or alternatively, the memory 106 and the memory 120 may store software instructions, e.g., executable by the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although the memory 106 and the memory 120 are shown in this embodiment as separate from the G-PCC encoder 200 and the G-PCC decoder 300, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 may also include internal memories for functionally similar or equivalent purposes. Additionally, memory 106 and memory 120 may store encoded data, for example, output from G-PCC encoder 200 and input to G-PCC decoder 300. In some embodiments, portions of memory 106 and memory 120 may be allocated as one or more buffers, for example, for storing raw decoded data and / or encoded data. For example, memory 106 and memory 120 may store data representing a point cloud.
[0019] The computer-readable medium 110 may represent any type of medium or device capable of transferring encoded data from the source device 102 to the destination device 116. In one embodiment, the computer-readable medium 110 represents a communication medium for enabling the source device 102 to transmit encoded data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 may modulate a transmission signal including the encoded data, and the input interface 122 may demodulate a received transmission signal according to a communication standard, such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum, or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device 102 to the destination device 116.
[0020] In some embodiments, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.
[0021] In some embodiments, the source device 102 can output the encoded data to a file server 114 or to another intermediate storage device capable of storing the encoded data generated by the source device 102. The destination device 116 can access the stored data from the file server 114 via streaming or download. The file server 114 can be any type of server device capable of storing the encoded data and transmitting the encoded data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. The destination device 116 can access the encoded data from the file server 114 through any standard data connection, including an Internet connection. The data connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded data stored on the file server 114. The file server 114 and the input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0022] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of the various IEEE 802.11 standards, or other physical components. In embodiments in which output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE-Advanced, 5G, etc. In some embodiments in which output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee™), the Bluetooth™ standard, etc. In some embodiments, the source device 102 and / or the destination device 116 may include corresponding system-on-a-chip (SoC) devices. For example, the source device 102 may include a SoC device for performing functions attributable to the G-PCC encoder 200 and / or the output interface 108, and the destination device 116 may include a SoC device for performing functions attributable to the G-PCC decoder 300 and / or the input interface 122.
[0023] The techniques of this disclosure may be applied to encoding and decoding in support of any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors and processing devices such as local or remote servers, geographic mapping, or other applications.
[0024] The input interface 122 of the destination device 116 receives the encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded bitstream may include signaling information defined by the G-PCC encoder 200, such as syntax elements having values that describe characteristics and / or processing of coding units (e.g., slices, pictures, groups of pictures, sequences, etc.), which is also used by the G-PCC decoder 300. The data consumer 118 uses the decoded data. For example, the data consumer 118 can use the decoded data to determine the location of a physical object. In some embodiments, the data consumer 118 may include a display for presenting an image based on the point cloud.
[0025] Each of the G-PCC encoder 200 and the G-PCC decoder 300 may be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Each of the G-PCC encoder 200 and the G-PCC decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. A device including G-PCC encoder 200 and / or G-PCC decoder 300 may comprise one or more integrated circuits, microprocessors, and / or other types of devices.
[0026] The G-PCC encoder 200 and the G-PCC decoder 300 may operate according to a coding standard, such as the video point cloud compression (V-PCC) standard or the geometry point cloud compression (G-PCC) standard. This disclosure may generally refer to coding (e.g., encoding and decoding) a picture to include the process of encoding or decoding data. An encoded bitstream generally includes a set of values for syntax elements that represent coding decisions (e.g., coding modes).
[0027] This disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to communication of values for syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 may signal values for syntax elements in the bitstream. In general, signaling refers to generating values in the bitstream. As discussed above, the source device 102 may transfer the bitstream to the destination device 116 in substantially real-time, or may transfer non-real-time, such as may occur when the syntax elements are stored in the storage device 112 for later retrieval by the destination device 116.
[0028] ISO / IEC MPEG (JTC1 / SC29 / WG11) and, more recently, ISO / IEC MPEG 3DG (JTC1 / SC29 / WG7) are examining the potential need for and seeking to develop a standard for a point cloud coding technique that would provide compression capabilities significantly beyond those of current methods. The groups, in a collaborative effort known as the 3-Dimensional Graphics Team (3DG), have aligned themselves in this quest to evaluate compression technology designs proposed by those with expertise in this field.
[0029] Point cloud compression activities are categorized into two different approaches. The first approach is "video point cloud compression" (V-PCC), which segments a 3D object and projects the segments (represented as "patches" in a 2D frame) into multiple 2D planes, which are further coded by a legacy 2D video codec, such as the High Efficiency Video Coding (HEVC) (ITU-T H.265) codec. The second approach is "geometry-based point cloud compression" (G-PCC), which directly compresses the 3D geometry, i.e., the positions of a set of points in 3D space, and the associated attribute values (for each point associated with that 3D geometry). G-PCC addresses the compression of point clouds in both category 1 (static point clouds) and category 3 (dynamically acquired point clouds). A recent draft of the G-PCC standard is available in ISO / IEC FDIS 23090-9 Geometry-based Point Cloud Compression (ISO / IEC JTC1 / SC29 / WG7 m55637, Teleconference, October 2020) and a description of the codec is available in G-PCC Codec Description (ISO / IEC JTC1 / SC29 / WG7 MDS20626, Teleconference, July 2021) (hereafter "G-PCC Codec Description").
[0030] A point cloud comprises a set of points in 3D space and may have attributes associated with the points. The attributes may be color information such as R, G, B, or Y, Cb, Cr, or reflectance information, or other attributes. Point clouds may be captured by various cameras or sensors, such as LIDAR sensors and 3D scanners, or may be computer generated. Point cloud data is used in a variety of applications, including but not limited to architecture (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors used to aid in navigation).
[0031] The 3D space occupied by the point cloud data can be enclosed by a virtual bounding box. The positions of points within the bounding box can be represented with a certain precision, and therefore the positions of one or more points can be quantized based on that precision. At the smallest level, the bounding box is divided into voxels, the smallest units of space represented by a unit cube. A voxel within a bounding box can be associated with zero points, one point, or more than one point. The bounding box can be divided into multiple cubic / rectangular regions, sometimes called tiles. Each tile can be coded into one or more slices. The partitioning of the bounding box into slices and tiles can be based on the number of points in each partition or other considerations (e.g., a particular region can be coded as a tile). The slice regions can be further partitioned using partitioning decisions similar to those in video codecs.
[0032] Figure 2 provides an overview of a G-PCC encoder 200. Figure 3 provides an overview of a G-PCC decoder 300. The illustrated modules are logical and do not necessarily correspond one-to-one to the implementation code in the reference implementation of the G-PCC codec, i.e., the TMC13 test model software considered by ISO / IEC MPEG (JTC1 / SC29 / WG11).
[0033] In both the G-PCC encoder 200 and the G-PCC decoder 300, the location of the point cloud is coded first. The attribute coding depends on the geometry to be decoded. In Fig. 2 and Fig. 3, the crosshatched module is an option typically used for category 1 data. The diagonal crosshatched module is an option typically used for category 3 data. All other modules are common to category 1 and category 3. See, for example, ISO / IEC FDIS 23090-9 Geometry-based Point Cloud Compression (ISO / IEC JTC1 / SC29 / WG7 m55637, Teleconference, October 2020).
[0034] For geometry point clouds, there are two different types of coding techniques: octree coding and predictive tree coding. Octree coding is discussed next. For category 3 data, the compressed geometry is typically represented as an octree from the root to the leaf level of individual voxels. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root to the leaf level of a block larger than a voxel) plus a model that approximates the surface within each leaf of the pruned octree. In this way, while both category 1 and category 3 data share the octree coding mechanism, category 1 data can further approximate the voxels within each leaf using a surface model. The surface model used is a triangulation involving 1-10 triangles per block, resulting in a triangle soup. Hence, category 1 geometry codecs are known as Trisoup geometry codecs, while category 3 geometry codecs are known as Octree geometry codecs.
[0035] FIG. 4 is an example octree partition for geometry coding according to the techniques of this disclosure. At each node of the octree 400, the G-PCC encoder 200 can signal the occupancy (if the occupancy is not inferred by the G-PCC decoder 300) for one or more of the node's child nodes (e.g., up to eight nodes) to the G-PCC decoder 300. Multiple neighborhoods are specified, including (a) nodes that share faces with the current octree node, (b) nodes that share faces, edges, or vertices with the current octree node, etc. Within each neighborhood, the occupancy of the node and / or its children can be used to predict the occupancy of the current node or the current node's children. For sparsely populated points at a particular node of the octree, the codec also supports a direct coding mode, where the 3D position of the point is directly coded. The G-PCC encoder 200 can signal a flag to indicate that the direct mode is signaled. At the lowest level, the number of points associated with the octree node / leaf node can also be coded.
[0036] When geometry is coded, attributes corresponding to those geometry points are coded. If there are multiple attribute points corresponding to one reconstructed / decoded geometry point, an attribute value representing the reconstructed point can be derived.
[0037] There are three attribute coding methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) coding, Interpolation-based Hierarchical Nearest Neighbor Prediction (Prediction Transform), and Interpolation-based Hierarchical Nearest Neighbor Prediction with Update / Lifting Steps (Lifting Transform). RAHT and lifting are typically used for category 1 data, while prediction is typically used for category 3 data. However, either method can be used for any data, and as with the geometry codec in G-PCC, the attribute coding method used to code the point cloud is specified in the bitstream.
[0038] The coding of attributes can be performed at a level-of-detail (LOD), where each level of detail can be used to obtain a more refined representation of the point cloud attributes. Each level of detail can be specified based on a distance metric from neighboring nodes or based on a sampling distance.
[0039] In the G-PCC encoder 200, the residual obtained as the output of the attribute coding method is quantized. The residual can be obtained by subtracting the attribute value from a prediction derived based on points in the neighborhood of the current point and based on the attribute values of previously coded points. The quantized residual can be coded using context-adaptive arithmetic coding.
[0040] In the embodiment of FIG. 2, the G-PCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, an LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0041] As shown in the example of FIG. 2, the G-PCC encoder 200 can obtain a set of positions of points in a point cloud and a set of attributes. The G-PCC encoder 200 can obtain the set of positions of points in a point cloud and a set of attributes from a data source 104 (FIG. 1). The positions can include coordinates of the points in the point cloud. The attributes can include information about the points in the point cloud, such as a color associated with the points in the point cloud. The G-PCC encoder 200 can generate a geometry bitstream 203 that includes an encoded representation of the positions of the points in the point cloud. The G-PCC encoder 200 can also generate an attribute bitstream 205 that includes an encoded representation of the set of attributes.
[0042] The coordinate transformation unit 202 can apply a transformation to the coordinates of the points to convert the coordinates from an initial domain to a transformation domain. In this disclosure, the transformed coordinates may be referred to as transformed coordinates. The color transformation unit 204 can apply a transformation to convert color information of the attributes to a different domain. For example, the color transformation unit 204 can convert color information from an RGB color space to a YCbCr color space.
[0043] Further, in the embodiment of FIG. 2, the voxelization unit 206 may voxelize the transformed coordinates. The voxelization of the transformed coordinates may include quantization and removing some points of the point cloud. In other words, multiple points of the point cloud may be contained within a single "voxel", which may be treated as one point in some respects hereafter. Further, the octree analysis unit 210 may generate an octree based on the voxelized transformed coordinates. Further, in the embodiment of FIG. 2, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may entropy code syntax elements representing the octree information and / or the surface information determined by the surface approximation analysis unit 212. The G-PCC encoder 200 may output these syntax elements in the geometry bitstream 203. The geometry bitstream 203 may also include other syntax elements, including syntax elements that are not arithmetically coded.
[0044] The geometry reconstruction unit 216 may reconstruct transformation coordinates of points in the point cloud based on the octree, the data indicative of the surface determined by the surface approximation analysis unit 212, and / or other information. The number of transformation coordinates reconstructed by the geometry reconstruction unit 216 may differ from the original number of points in the point cloud due to voxelization and surface approximation. In this disclosure, the resulting points may be referred to as reconstructed points. The attribute transfer unit 208 may transfer attributes of the original points of the point cloud to the reconstructed points of the point cloud.
[0045] The switch 228 may communicatively couple the attribute transfer unit 208 to the RAHT unit 218 via the terminal 230, where the RAHT unit 218 may apply RAHT coding to the attributes of the reconstructed points. In some embodiments, under RAHT, the attributes of a 2×2×2 block of point locations are obtained and transformed along one direction to obtain four low-frequency nodes (L) and four high-frequency nodes (H). The four low-frequency nodes (L) are then subsequently transformed in a second direction to obtain two low-frequency nodes (LL) and two high-frequency nodes (LH). The two low-frequency nodes (LL) are transformed in a third direction to obtain one low-frequency node (LLL) and one high-frequency node (LLH). The low-frequency node LLL corresponds to the DC coefficient, and the high-frequency nodes H, LH, and LLH correspond to the AC coefficients. The transformation in each direction may be a 1-D transformation using two coefficient weights. The low frequency coefficients can be obtained as 2x2x2 blocks of coefficients for the next higher level RAHT transform, and the AC coefficients are coded without modification, and such transformation continues up to the top root node. The tree traversal for coding is a top-to-bottom traversal, which is used to calculate the weights to be used for the coefficients, and the transformation order is bottom-to-top. The coefficients can then be quantized and coded.
[0046] Alternatively, or in addition, the switch 228 can communicatively couple the attribute transfer unit 208 to the LOD generation unit 220 via the terminal 232, in which case the lifting unit 222 can apply LOD processing and lifting, respectively, to the attributes of the reconstructed points. The LOD generation is used to split the attributes into different refinement levels. Each refinement level provides refinement to the attributes of the point cloud. The first refinement level provides a coarse approximation and includes a small number of points, subsequent refinement levels typically include more points, and so on. The refinement levels can be constructed using a distance-based metric, or one or more other classification criteria (e.g., subsampling from a particular rank) can also be used. Hence, all reconstructed points can be included in a refinement level. Each level of detail is generated by taking the union of all points up to a particular refinement level, e.g., LOD1 is obtained based on refinement level RL1, LOD2 is obtained based on RL1 and RL2, ...LODN is obtained by the union of RL1, RL2, ...RLN. In some cases, LOD generation can be followed by a prediction scheme (e.g., predictive transformation), where attributes associated with each point at that LOD are predicted from a weighted average of previous points, and the residual is quantized and entropy coded. A lifting scheme builds on top of the predictive transformation mechanism, updating the coefficients using an update operator, and adaptive quantization of the coefficients is performed.
[0047] The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 may output these syntax elements in an attribute bitstream 205. The attribute bitstream 205 may also include other syntax elements, including non-arithmetically coded syntax elements.
[0048] In the embodiment of FIG. 3, the G-PCC decoder 300 may include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, a LoD generation unit 316, an inverse lifting unit 318, an inverse coordinate transformation unit 320, and an inverse color transformation unit 322.
[0049] The G-PCC decoder 300 may obtain a geometry bitstream 203 and an attribute bitstream 205. A geometry arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., Context-Adaptive Binary Arithmetic Coding (CABAC) or other types of arithmetic decoding) to syntax elements in the geometry bitstream 203. Similarly, an attribute arithmetic decoding unit 304 may apply arithmetic decoding to syntax elements in the attribute bitstream 205.
[0050] The octree synthesis unit 306 can synthesize an octree based on syntax elements parsed from the geometry bitstream 203. Starting from the root node of the octree, the occupancy of each of the eight child nodes at each octree level is signaled in the bitstream. If the signaling indicates that a child node at a particular octree level is occupied, the occupancy of the children of this child node is signaled. The signaling of the nodes at each octree level is signaled before proceeding to a subsequent octree level. At the final level of the octree, each node corresponds to a voxel position, and if a leaf node is occupied, one or more points can be specified to be occupied at that voxel position. In some cases, due to quantization, some branches of the octree may terminate before the final level. In such cases, the leaf node is considered as an occupied node that does not have any child nodes. In cases where surface approximations are used in the geometry bitstream 203, the surface approximation synthesis unit 310 can determine a surface model based on syntax elements parsed from the geometry bitstream 203 and based on the octree.
[0051] Furthermore, the geometry reconstruction unit 312 may perform reconstruction to determine coordinates of points in the point cloud. For each position in the leaf nodes of the octree, the geometry reconstruction unit 312 may reconstruct the node position by using the binary representation of the leaf node in the octree. At each corresponding leaf node, the number of points in the corresponding leaf node is signaled, which indicates the number of overlapping points at the same voxel position. If geometry quantization is used, the point positions are scaled to determine the value of the reconstructed point position.
[0052] The inverse coordinate transformation unit 320 can apply an inverse transform to the reconstructed coordinates (positions) of the points in the point cloud to transform them from the transformed domain back to the original domain. The positions of the points in the point cloud may be in the floating point domain, but the point positions in the G-PCC codec are coded in the integer domain. An inverse transform can be used to transform those positions back to the original domain.
[0053] 3, the inverse quantization unit 308 may inverse quantize the attribute value. The attribute value may be based on a syntax element obtained from the attribute bitstream 205 (e.g., including a syntax element decoded by the attribute arithmetic decoding unit 304).
[0054] Depending on how the attribute values are encoded, the switch 328 can communicatively couple the inverse quantization unit 308 to the RAHT unit 314 via a terminal 330 or to the LOD generation unit 316 via a terminal 332. If the switch 328 communicatively couples the inverse quantization unit 308 to the RAHT unit 314, the RAHT unit 314 can determine color values for points of the point cloud based on the inverse quantized attribute values. RAHT decoding is performed from the top to the bottom of the tree. At each level, component values are derived using the low and high frequency coefficients derived from the inverse quantization process. At the leaf nodes, the derived values correspond to the attribute values of the coefficients. The weight derivation process for a point is similar to the process used in the G-PCC encoder 200. Alternatively, if the switch 328 communicatively couples the inverse quantization unit 308 to the LOD generation unit 316, the LOD generation unit 316 and the inverse lifting unit 318 can use a level of detail-based technique to determine color values for points of the point cloud. The LOD generation unit 316 decodes each LOD, which represents a progressive refinement of the point's attributes. In the case of predictive transformation, the LOD generation unit 316 can derive a predicted value for the point from a weighted sum of points present in the previous LOD or previously reconstructed points in the same LOD. The LOD generation unit 316 can obtain a reconstructed value for the attribute by adding the predicted value to the residual (obtained after inverse quantization). If a lifting scheme is used, the LOD generation unit 316 may also include an update operator to update the coefficients used to derive the attribute value. The LOD generation unit 316 can also apply inverse adaptive quantization in this case.
[0055] 3 embodiment, the inverse color transformation unit 322 may apply an inverse color transformation to the color values. The inverse color transformation may be the inverse of the color transformation applied by the color transformation unit 204 of the encoder 200. For example, the color transformation unit 204 may transform the color information from an RGB color space to a YCbCr color space. Thus, the inverse color transformation unit 322 may transform the color information from the YCbCr color space to the RGB color space.
[0056] The various units in FIG. 2 and FIG. 3 are shown to aid in understanding the operations performed by the encoder 200 and the decoder 300. These units may be implemented as fixed function circuits, programmable circuits, or a combination thereof. A fixed function circuit refers to a circuit that provides a specific function and is preconfigured for the operations that it can perform. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexibility in the operations that it can perform. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. Although a fixed function circuit may execute software instructions (e.g., to receive parameters or output parameters), the type of operations that the fixed function circuit performs is generally immutable. In some embodiments, one or more of the units may be individual circuit blocks (fixed function or programmable), and in some embodiments, one or more of the units may be an integrated circuit.
[0057] Predictive geometry coding (see, for example, G-PCC Codec Description) has been introduced as an alternative to octree geometry coding, where nodes are arranged in a tree structure (defining the prediction structure) and various prediction strategies are used to predict the coordinates of each node in the tree with respect to its predictor. Figure 5 shows an example of a predictive tree, a directed graph, with arrows pointing to the prediction direction. The horizontal crosshatched node is the root vertex and has no predictor, the crosshatched node has two children, the diagonal crosshatched node has three children, the non-crosshatched node has one child, and the vertical crosshatched node is a leaf node and has no children. Except for the root node, all nodes have only one parent node.
[0058] 5 is a conceptual diagram illustrating one embodiment of a prediction tree. Node 500 is the root vertex and has no predictors. Nodes 502 and 504 have two children. Node 506 has three children. Nodes 508, 510, 512, 514, and 516 are leaf nodes and have no children. The remaining nodes each have one child. Except for the root node 500, all nodes have only one parent node.
[0059] For each node, four prediction strategies are specified based on its parent (p0), grandparent (p1), and great-grandparent (p2): No Prediction / Zero Prediction(0) Delta prediction (p0) Linear prediction (2 * p0-p1) Parallelogram prediction (2 * p0+p1-p2)
[0060] The G-PCC encoder 200 can employ any algorithm to generate the prediction tree, the algorithm used can be determined based on the application / use case, and several strategies can be used, some of which are described in the G-PCC Codec Description.
[0061] For each node, the residual coordinate values are coded in the bitstream starting from the root node in a depth-first manner. For example, the G-PCC encoder 200 can code the residual coordinate values in the bitstream.
[0062] Predictive geometry coding is primarily useful for category 3 (LIDAR-acquired) point cloud data, e.g., for low-latency applications.
[0063] 6A and 6B are conceptual diagrams illustrating an embodiment of a rotational LIDAR acquisition model. The angular mode for predictive geometry coding is described next. In predictive geometry coding, the angular mode can be used, where the characteristics of the LIDAR sensor can be exploited in coding the prediction tree more efficiently. The coordinates of the position are transformed into the (r,φ,i) (radius, azimuth, and laser index) domain 600, and the prediction is performed in this domain 600 (e.g., the residuals are coded in the r,φ,i domain). Because coding in r,φ,i is not lossless due to errors in rounding, a second set of residuals, corresponding to Cartesian coordinates, can be coded. A description of the encoding and decoding strategy used for the angular mode for predictive geometry coding is reproduced below from the G-PCC Codec Description.
[0064] The angular mode technique can focus on point clouds acquired using a rotational LIDAR model. In this case, the LIDAR 602 has N lasers (e.g., N=16, 32, 64) that rotate around the Z axis according to an azimuth angle φ. Each laser is oriented at a different elevation angle θ(i). i=1...N and height
number
[0065] This technique models the position of M using three parameters (r, φ, i) calculated as follows:
[0066]
number
[0067] More precisely, this technology
[0068]
number
[0069]
number
[0070]
number
[0071]
number
[0072] To avoid reconstruction inconsistencies due to the use of floating-point arithmetic,
number
[0073]
number
number
[0074]
number
[0075] The reconstructed Cartesian coordinates are obtained as follows:
[0076]
number
[0077]
number
[0078] (r x ,r y ,r z ) may be the reconstruction residual defined as:
[0079]
number
[0080] For this technique, the G-PCC encoder 200 may proceed as follows: 1) Model parameters
[0081]
number
number
[0082]
number
[0083] A new predictor can be introduced that exploits the characteristics of the LIDAR. For example, the rotation speed of the LIDAR scanner around the z-axis is usually constant. Therefore, the G-PCC encoder 200 can use the current
[0084]
number
[0085]
number
[0086] The G-PCC decoder 300 may proceed as follows: 1) Model parameters
[0087]
number
number
[0088]
number
[0089]
number
[0090]
number
[0091] Lossy compression reduces the reconstruction residual (r x ,r y ,r z This can be achieved by applying quantization to , or by deleting points.
[0092] The quantized reconstruction residual can be calculated as follows:
[0093]
number
[0094]
number
[0095] The G-PCC encoder 200 or the G-PCC decoder 300 can use trellis quantization to further improve the rate-distortion (RD) performance results.
[0096] The quantization parameters may be varied at the sequence / frame / slice / block level to achieve region-adaptive quality and / or for rate control purposes.
[0097] Inter prediction in G-PCC predictive geometry coding is discussed next. For information on G-PCC predictive geometry coding, see Technologies under consideration in G-PCC (ISO / IEC JTC1 / SC29 / WG7 MDS20648, Teleconference, July 2021); A.K. Ramasubramonian, B. Ray, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC][New]Inter prediction with predictive geometry coding (ISO / IEC JTC1 / SC29 / WG7 m56117, January 2021); A.K. Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC]EE13.2 report on inter prediction, Test2 (ISO / IEC JTC1 / SC29 / WG7 m56839, April 2021); and A.K. Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz, [G-PCC][EE13.2-related] Additional results for inter prediction for predictive geometry (ISO / IEC JTC1 / SC29 / WG7 m56841, April 2021).
[0098] Predictive geometry coding uses a prediction tree structure to predict the location of a point. If angle coding is enabled, the x, y, z coordinates are converted to radius, azimuth, and laserID, and the residual can be signaled in these three coordinates as well as in the x, y, and z dimensions. The intra prediction used for radius, azimuth, and laserID can be one of four modes, and the predictors are the nodes classified as parent, grandparent, and great-grandparent in the prediction tree with respect to the current node. Predictive geometry coding as designed in G-PCC Ed.1 is an intra-coding tool since it only uses points within the same frame for prediction. Furthermore, using points from previously decoded frames may provide better prediction and therefore better compression performance.
[0099] For inter prediction, as first proposed in A.K. Ramasubramonian, B. Ray, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] [New] Inter prediction with predictive geometry coding (ISO / IEC JTC1 / SC29 / WG7 m56117, January 2021) and A.K. Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC] EE13.2 report on inter prediction, Test2 (ISO / IEC JTC1 / SC29 / WG7 m56839, April 2021), the proposal was to predict the radius of the points from the reference frame. For each point in the prediction tree, the G-PCC decoder 300 can determine whether the point is inter-predicted or intra-predicted (e.g., the G-PCC encoder 200 can indicate such inter-prediction or intra-prediction by a flag that the G-PCC encoder 200 can signal in the bitstream). If intra-predicted, the intra-prediction mode of predictive geometry coding is used. If inter-prediction is used, the azimuth angle and laserID are still predicted using intra-prediction, while the radius is predicted from the point in the reference frame that has the same laserID as the current point and the azimuth angle closest to the current azimuth angle. Further modifications to this technique by A. K. Ramasubramonian, G. Van der Auwera, L. Pham Van, M. Karczewicz, [G-PCC][EE13.2-related] Additional results for inter prediction for predictive geometry (ISO / IEC JTC1 / SC29 / WG7 m56841, April 2021) also allow inter prediction of azimuth and laserID in addition to radius prediction.When inter-coding is applied, the radius, azimuth and laserID of the current point are predicted based on points in the reference frame that are near the azimuth position of the previously decoded point. Furthermore, separate sets of contexts are used for inter-prediction and intra-prediction.
[0100] 7 is a conceptual diagram illustrating one embodiment of inter-prediction of a current point from a point in a reference frame. The extension of inter-prediction to azimuth, radius, and laserID includes the following steps, which can be performed, for example, by the G-PCC decoder 300: 1) For a given point (eg, the current point curPoint 700 in the current frame 704), select the previous decoded point (prevDecP0 702). 2) Select a location in the reference frame 70 (eg, refFrameP0 706) that has the same scaling azimuth and laserID as prevDecP0 702. 3) Find the first point (interPredPt 710) in the reference frame 708 that has an azimuth angle greater than the azimuth angle of refFrameP0 706.
[0101] Figure 8 is a flow diagram illustrating an example operation of a G-PCC decoder. Figure 8 shows the decoding flow associated with the "inter_flag" signaled pointwise. This technique is available in InterEM-v3.0.
[0102] For example, the G-PCC decoder 300 may determine whether the inter flag is true (e.g., equal to 1) (800). If the inter flag is true (the "Yes" path out of block 800), the G-PCC decoder 300 may select a previous decoded point in the decoding order using the radius, azimuth, and laserID (802). The G-PCC decoder 300 may derive a quantized phi, Q(phi) (e.g., a quantized azimuth value), of the selected previous decoded point (e.g., prevDecP0 702) (804). The G-PCC decoder 300 may check a reference frame (e.g., reference frame 708 of FIG. 7) for a point from which interPredPt 710 may be derived, where the quantized phi of such point is greater than Q(phi) (806). The G-PCC decoder 300 may then use 808 interPredPt 710 as an inter predictor for the current point curPoint 700. The G-PCC decoder 300 may then use a delta-phi multiplier, e.g., n(j)×δ as described above. φ (k) can be added to the first order residual (810).
[0103] If the inter flag is false (e.g., equal to 0) (the "No" path out of block 800), the G-PCC decoder 300 may select an intra prediction candidate (812) and apply intra prediction. The G-PCC decoder 300 may then add a delta-phi multiplier to generate a first-order residual (810).
[0104] FIG. 9 is a conceptual diagram illustrating an example of an additional inter predictor point obtained from a first point with a larger azimuth angle than the inter predictor point. The additional predictor candidates are discussed next. Information on the additional predictor candidates can be found in KL Loi, T. Nishi, T. Sugio, [G-PCC][New]Inter Prediction for Improved Quantization of Azimuthal Angle in Predictive Geometry Coding (ISO / IEC JTC1 / SC29 / WG7 m57351, July 2021). In the inter prediction technique for predicted geometry described above, the radius, azimuth angle, and laserID of the current point are predicted using the following steps based on points existing near the same azimuth angle position in the reference frame when inter coding is applied, for example by the G-PCC decoder 300: 1) For a given point (eg, the current point Curr Point 900), select a previous decoded point (eg, Prev Dec Point 902). 2) Select a location in the reference frame 908 (eg, Ref Point 906) that has the same scaling azimuth and laserID as the previous decoded point (eg, Prev Dec Point 902). 3) Select a location (Inter Pred Point 910) in the reference frame 908 to use as an inter predictor point from the first point with a larger azimuth angle than a location in the reference frame 908 that has the same scaling azimuth angle and laserID as the previous decoding point (e.g., Prev Dec Point 902).
[0105] This technique adds an additional inter predictor point 912, which is obtained by finding the first point that has a larger azimuth than an inter predictor point (e.g., Inter Pred Point 910), as shown in Figure 9. If inter coding is applied by the G-PCC encoder 200, additional signaling is used to indicate which of the predictors is selected. For example, the G-PCC encoder 200 can signal to the G-PCC decoder 300 which of the predictors is selected.
[0106] Improved inter prediction flag coding is discussed next. Information on improved inter prediction flag coding can be found in AK Ramasubramonian, L. Pham Van, G. Van der Auwera, M. Karczewicz, [G-PCC][New proposal]Improvement to inter prediction using predictive geometry coding (ISO / IEC JTC1 / SC29 / WG7 m57299, July 2021). An improved context selection algorithm for coding the inter prediction flag can be applied. The G-PCC encoder 200 can use the inter prediction flag values of the five previously coded points to select the context of the inter prediction flag in predictive geometry coding.
[0107] Inter prediction techniques for predictive geometry coding rely on searching for inter prediction points in reference frames, as described above. In contrast to two-dimensional (2D) video coding, these points are not localized on a regular 2D array sampling grid. Furthermore, in 2D video coding, the 2D array sampling grid allows relatively simple interpolation filters to be applied block-by-block to obtain prediction blocks with sub-pixel accuracy (e.g., 1 / 2, 1 / 4, 1 / 8 sub-pixel, etc.), which provides significant coding efficiency for inter coding. Similarly, interpolation for inter prediction for predictive geometry coding can be expected to have coding efficiency benefits. However, interpolation on an irregular sampling grid is complicated due to the variation in the spacing between samples.
[0108] According to the techniques of the present disclosure, any of the following techniques can be applied independently or in a combined manner (eg, in any combination).
[0109] Although the discussion in this disclosure is primarily directed to polar coordinate systems, the techniques disclosed in this application can also be applied to other coordinate systems, such as Cartesian, spherical, or any custom coordinate system that can be used to represent / code the location and attributes of a point cloud. In particular, G-PCC utilizes radius and azimuth from a spherical coordinate system in combination with a laser identifier or laserID (e.g., elevation angle) from the LIDAR sensor that captured the point.
[0110] At least one radial inter predictor for at least one current point in the current point cloud frame can be obtained by applying at least one radial interpolation technique to at least two points belonging to a reference point cloud frame. The reference point cloud frame can be a frame of point cloud data whose points can be used as references in predicting at least one current point in the current frame of point cloud data. The radial inter predictor can be a predictor derived from a point of the reference point cloud frame that can be used to predict a radius of at least one current point in the current point cloud frame.
[0111] The G-PCC encoder 200 or the G-PCC decoder 300 may select at least two reference points. The at least two reference points may be at least two points in a reference point cloud frame that are used to obtain at least one radial inter predictor. The at least two reference points may be selected in the reference point cloud frame according to the following technique: In general, the at least two reference points may be a grouping of reference points that belong to a reference point cloud frame (e.g., Reference Frame 908) with a certain range of azimuth, laserID, and / or radius. The purpose of the grouping is to apply at least one radial interpolation technique to the points that belong to the group to generate a radial inter predictor for at least one current point (e.g., Curr Point 900) in a current point cloud frame (e.g., Current Frame 904). The reference point cloud frame may be partially modified to compensate for motion, for example, by global motion, such as rotation (or affine, perspective, etc.) and / or translation, or local motion (e.g., region-based motion). For example, if a LIDAR system is on or in a car and the car is turning, the G-PCC encoder 200 or the G-PCC decoder 300 can apply global motion compensation. The purpose of the motion compensation is to obtain spatial alignment between the points of the reference point cloud frame and the points of the current point cloud frame, which results in a better inter-predictor.
[0112] The position coordinates of the reference point or the current point may be quantized, approximated, scaled, etc. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may quantize, approximate, scale, etc. the position coordinates of the reference point or the current point. For example, the G-PCC decoder 300 may determine the precision or representation bit depth of those positions by analyzing parameters signaled in the bitstream, or the G-PCC decoder 300 may build an array or map of reference point cloud positions with an array cell size or bin size for quantizing the point coordinates, which the G-PCC decoder 300 may determine by analyzing parameters signaled in the bitstream.
[0113] At least two reference points may have the same (or similar, e.g., adjacent) laserIDs and may be ordered according to their azimuth angles in the reference point cloud frame, either in ascending or descending order of azimuth angles. The adjacent laserID may be one laserID higher or lower than the other. If the azimuth angles are the same, they may be further ordered according to radius, either in ascending or descending order. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may order at least two reference points.
[0114] At least two reference points may have the same (or similar, e.g., adjacent) azimuth angles and may be ordered according to their laserID in the reference point cloud frame, either in ascending or descending order of laserID. Adjacent azimuth angles may be one azimuth angle above or below the other. If the laserIDs are identical, they may be further ordered according to radius, either in ascending or descending order.
[0115] In general, the position of a group of reference points belonging to a reference point cloud frame is related to the position of a current point in a current frame to be inter-predicted. The position of the reference point may be partially modified by motion compensation. For example, the G-PCC encoder 200 or the G-PCC decoder can partially modify the position of the reference point using motion compensation.
[0116] The positions of the at least two reference points can be determined by the azimuth angle and the laserID of the current point in the current point cloud frame that is to be inter-predicted. For example, the G-PCC encoder 200 can determine the positions of the at least two reference points using the azimuth angle and the laserID of the current point.
[0117] The location of a first of the at least two reference points having a laserID equal to the laserID of the current point can be determined by searching for a closest azimuth angle that is less than the azimuth angle of the current point, and the location of a second of the at least two reference points can be determined by searching for a closest azimuth angle that is greater than the azimuth angle of the current point.
[0118] The additional reference points may be determined by searching for the second, next closest point with an azimuth angle smaller or larger than the azimuth angle of the current point. The number of additional reference points may depend on the radial interpolation method applied to the reference points. For example, the radial interpolation method may be averaging, weighted averaging, linear interpolation, mathematical trigonometry, radial interpolation filtering using a filter with three or more coefficients, grouping of points on a segment, or any other interpolation method.
[0119] The position of a first of at least two reference points having an azimuth angle equal to the azimuth angle currAzim of the current point can be determined (e.g., by the G-PCC encoder 200 or the G-PCC decoder 300) by searching for the closest laserID that is smaller than the laserID angle of the current point currLaserId, and the position of a second of the at least two reference points can be determined by searching for the closest laserID that is larger than the laserID of the current point.
[0120] The G-PCC encoder 200 or the G-PCC decoder 300 can determine the additional reference points by searching for the second, next closest point with a laserID smaller or larger than the laserID of the current point. The number of additional reference points may depend on the interpolation method applied to the reference points.
[0121] In one embodiment, if there is no position in the reference frame with an azimuth angle equal to currAzim for a laserID being searched for by one of the techniques disclosed above (e.g., the closest laserID less than currLaserId, or the closest laserId greater than currLaserId, etc.), then the position with the azimuth angle closest to currAzim can be selected for the laserID being searched for.
[0122] In general, the G-PCC encoder 200 or the G-PCC decoder 300 can determine the location of a group of reference points by a range of azimuth angles and laserIDs, which determines the neighborhood of the azimuth angle and laserID of the current point to be predicted. This range can depend on the radial interpolation method to be applied to those reference points. For example, the radial interpolation method can be averaging, weighted averaging, linear interpolation, mathematical trigonometry, radial interpolation filtering using a filter with three or more coefficients, grouping of points on a segment, or any other interpolation method.
[0123] If one or more points in a group of reference points are unavailable or are duplicate predictors (e.g., the same radius as another reference point), the G-PCC encoder 200 or the G-PCC decoder 300 may select one or more of the following techniques to replace the unavailable / duplicate points: 1) Select a default predictor (e.g., select a default radius value); 2) use the search order for each location in the group of reference points (e.g. if an azimuth equal to currAzim is searched for with the closest laserID greater than currLaserId, if that point is unavailable / duplicate, select the next laserID and point with the same azimuth as currAzim); 3) Do not select an alternate point and do not use that point for the interpolation.
[0124] The two reference points used for radial interpolation can be selected to be equal to the first and second predictor candidates as described above for inter prediction for predictive geometry coding, and additional predictor candidates, for example, as shown in any of Figures 7-9. In general, additional nearby reference points (e.g., in addition to the two reference points) can be selected in a similar manner. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can select additional nearby reference points.
[0125] For example, as described above with respect to inter prediction and additional predictor candidates for predictive geometry coding as shown in Figures 7-9, the location of the reference point can be determined using a previously decoded point as described. Alternatively, the parent of the current point can be used to determine the reference point. The parent can be defined as the parent node of the current node in the tree structure that is being built for the current point cloud frame. Alternatively, a grandparent node or a great-grandparent node, etc. can be used.
[0126] If at least two reference points from a reference point cloud frame have the same laserID value, they are used for the radial interpolation technique as follows. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can use one of the following radial interpolation techniques. If they have the same azimuth angle, a similar interpolation can be applied.
[0127] The G-PCC encoder 200 or the G-PCC decoder 300 may apply a simple averaging where equal weights are applied to the radius values corresponding to each reference point. Radius_smaller and radius_greater may be defined as the radius values associated with the reference points having azimuth angles smaller and larger than that of the current point, respectively. For example, average_radius=(1 / 2) * radius_smaller+(1 / 2) * radius_greater.
[0128] The G-PCC encoder 200 or the G-PCC decoder 300 can apply linear interpolation using the azimuth angles of the reference point and the current point to obtain a prediction radius of the current point. Let current_azimuth be the azimuth angle associated with the current point, and let ref_smaller_azimuth and ref_greater_azimuth be the azimuth angles associated with the reference point that are smaller than the azimuth angle of the current point and the azimuth angles associated with the reference point that are larger than the azimuth angle of the current point, respectively. For example, as follows: weight1=(current_azimuth-ref_smaller_azimuth) / (ref_greater_azimuth-ref_smaller_azimuth weight2=(ref_greater_azimuth-current_azimuth) / (ref_greater_azimuth-ref_smaller_azimuth), i.e. (1.0-weight1). linear_interpolated_radius=weight2 * radius_smaller + weight 1 * radius_greater
[0129] The G-PCC encoder 200 or the G-PCC decoder 300 can apply mathematical trigonometry to calculate the interpolation radius as shown in the following pseudocode: Let predRight[1] and predLeft[1] be azimuth angles greater than and less than the azimuth angle of the current point, respectively, predRight[0] and predLeft[0] be radius values associated with these azimuth angles, and geom_angular_azimuth_scale be a parameter indicating the representation precision or bit depth of the azimuth angles.
[0130] [Table 1]
[0131] The G-PCC encoder 200 or the G-PCC decoder 300 may use more than two reference points, for example, four, six, or eight reference points, for a radial interpolation filter with more than two coefficients, for example, a cubic interpolation filter. The azimuth angle spacing between these reference points may be irregular, so the interpolation filter may be adaptively generated to fit these azimuth angle spacings.
[0132] A plurality of reference points may be grouped into a segment or block to generate an interpolated radius predictor for a plurality of current points grouped into a corresponding segment or block. For example, the G-PCC encoder 200 or the G-PCC decoder 300 may group a plurality of reference points to generate an interpolated radius predictor.
[0133] Before applying the radius interpolation technique, the conditions can be checked. For example, the G-PCC encoder 200 or the G-PCC decoder 300 can perform the condition check. For example, the difference between two radius values associated with two reference points may be required to be smaller than a threshold value, that is, ABS(radius1 - radius2) < threshold value (ABS is the absolute value). For example, if one potential reference point is on a road sign near the G-PCC encoder 200 and the other potential reference point is on a building 100 meters away, performing radius interpolation between the two potential reference points may not generate an accurate predicted radius. This threshold value may be signaled in the bitstream or can be a predetermined constant. Alternatively, multiple threshold ranges can be specified for abs(radius1 - radius2), and different interpolation schemes can be applied for each threshold range. For example, when ABS(radius1 - radius2) < threshold value 1, linear interpolation can be applied. However, when threshold value 1 < ABS(radius1 - radius2) < threshold value 2, cubic interpolation can be applied, and when ABS(radius1 - radius2) > threshold value 2, no interpolation can be applied.
[0134] The type of interpolation technique applied (e.g., average, linear, cubic, etc.) can be signaled in the bitstream, for example, for each current point, for each current segment / block, for each slice, for each frame, for each sequence, etc. Furthermore, the interpolation coefficient or weight can also be signaled in the bitstream. For example, the G-PCC encoder 200 can signal such syntax elements in the bitstream.
[0135] For example, as shown in any of Figures 7-9, if two reference points are selected, the G-PCC encoder 200 or the G-PCC decoder 300 can obtain the interpolation radius by simple averaging of the two radius values corresponding to the two reference points. Alternatively, a weighted average can be applied, or another interpolation technique can be used as described above. If more than two reference points are selected, an interpolation technique utilizing more than two points can be applied.
[0136] After the interpolated radius predictor is obtained, the G-PCC encoder 200 can calculate the radius residual between the radius of the current point and the predicted radius. For example, the G-PCC encoder 200 can calculate the radius residual by subtracting the radius of the current point from the predicted radius or subtracting the predicted radius from the radius of the current point. This residual is coded in the bitstream. The G-PCC decoder 300 can obtain the same interpolated radius predictor, decode the radius residual from the bitstream, and add the radius residual to the radius predictor to obtain the radius of the current point.
[0137] The radial residual, including for example the sign and magnitude, may be context coded using an arithmetic coder. The radial residual may be quantized.
[0138] The context for coding the radial residual obtained using inter prediction may be separate from the context for coding the radial residual obtained using intra prediction.
[0139] The azimuth of the current point to be predicted can be used to determine the location of the interpolation between the reference points, but it may be necessary for the azimuth of the current point to be decoded before the radius can be predicted.
[0140] Decoding the azimuth angle of the current point may include obtaining an azimuth angle predictor and decoding the azimuth angle residual. In some embodiments, decoding the azimuth angle residual depends on the decoding radius of the current point to determine the inverse quantizer, inverse scaling, or representation bit depth of the azimuth angle residual. This results in a causal decoding dependency between the azimuth angle and the radius of the current point. For example, the G-PCC decoder 300 can decode the azimuth angle of the current point.
[0141] To avoid this causal decoding dependency, the inverse quantization, inverse scaling, or representation bit depth of the current azimuth angle residual can instead depend on the decoding radius value of the previously reconstructed current point, e.g., in the decoding order.
[0142] Alternatively, the inverse quantization, inverse scaling, or representation bit depth of the current azimuth angle residual may instead depend on the radius value of a reference point in the neighborhood of the current point, e.g., the radius of a reference point corresponding to a previously decoded azimuth angle of the current point.
[0143] For example, if two reference points are selected, as shown in any of Figures 7 to 9, there is no need to first code the azimuth angle of the current point, and the above-mentioned dependency on radius can be avoided.
[0144] The following pseudo-code implementation of the G-PCC decoder 300 algorithm shows linear radial interpolation for inter prediction for predictive geometry coding. The code is based on TMC13v12, which corresponds to the first version of G-PCC.
[0145] [Table 2-1]
[0146] [Table 2-2]
[0147]
Table 2-3
[0148]
Table 2-4
[0149]
Table 2-5
[0150]
Table 2-6
[0151] FIG. 10 is a flow diagram illustrating an example radial interpolation technique according to one or more aspects of the present disclosure. The G-PCC encoder 200 or the G-PCC decoder 300 may determine at least two reference points in a reference point cloud frame of point cloud data (1000). For example, the G-PCC encoder 200 or the G-PCC decoder 300 selects at least two reference points (e.g., an inter pred point 910 and an additional inter pred point 912) in a reference frame 908. The G-PCC encoder 200 or the G-PCC decoder 300 may apply radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of point cloud data (1002). For example, the G-PCC encoder 200 or the G-PCC decoder 300 may apply any of the radial interpolation techniques described herein to the at least two reference points to obtain at least one radial inter predictor for the current point. The G-PCC encoder 200 or the G-PCC decoder 300 may code the current point cloud frame based on at least one radial inter predictor for at least one current point in the current point cloud frame (1004). For example, the G-PCC encoder 200 may encode the current point cloud frame using at least one radial inter predictor for at least one current point, and the G-PCC decoder 300 may decode the current point cloud frame using at least one radial inter predictor for at least one current point.
[0152] In some embodiments, the at least two reference points belong to a group of reference points within a range of at least one of an azimuth angle, a laser identifier, or a radius. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 can compensate the reference point cloud frame for motion by applying at least one of a rotation, a translation, or a local motion.
[0153] In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may at least one of quantize, approximate, or scale at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point. In some embodiments, the G-PCC encoder 200 may signal or the G-PCC decoder 300 may analyze the bit depth of at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may build an array or map of the position coordinates of the at least two reference points to quantize the position coordinates of the at least two reference points.
[0154] In some embodiments, the at least two reference points have the same or similar (e.g., within a predefined range) laser identifiers. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may order the at least two reference points according to corresponding azimuth angles associated with the at least two reference points. In some embodiments, the at least two reference points have the same azimuth angle. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may further order the at least two reference points according to corresponding radii associated with the at least two reference points.
[0155] In some embodiments, the at least two reference points have the same or similar (e.g., within a predefined range) azimuth angles. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may order the at least two reference points according to corresponding laser identifiers associated with the at least two reference points. In some embodiments, the at least two reference points have the same laser identifier. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may further order the at least two reference points according to corresponding radii associated with the at least two reference points.
[0156] In some embodiments, the positions of the at least two reference points are associated with the position of the at least one current point. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may determine the positions of the at least two reference points based on the azimuth angle and the laser identifier of a first current point of the at least one current point. In some embodiments, in determining the positions of the at least two reference points, the G-PCC encoder 200 or the G-PCC decoder 300 may search for a first closest azimuth angle that is smaller than the azimuth angle of the first current point to determine the position of the first reference point of the at least two reference points, and may search for a second closest azimuth angle that is larger than the azimuth angle of the first current point to determine the position of the second reference point of the at least two reference points.
[0157] In some embodiments, when determining the position of the at least two reference points, the G-PCC encoder 200 or the G-PCC decoder 300 may search for a first closest laser ID that is smaller than the laser ID of the first current point to determine the position of the first reference point of the at least two reference points, and may search for a second closest laser ID that is greater than the laser ID of the first current point to determine the position of the second reference point of the at least two reference points.
[0158] In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 can determine the positions of the at least two reference points based on an azimuth angle and a range of the laser identifier. In some embodiments, the azimuth angle and the range of the laser identifier are based on a radial interpolation method to be applied to the at least two reference points.
[0159] In some embodiments, one or more of the at least two reference points are unavailable or include a duplicate predictor. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 can replace the unavailable or duplicate predictor with a default predictor, or with the reference point having the next closest laser identifier, the next closest azimuth angle, or can discard the unavailable or duplicate predictor.
[0160] In some embodiments, the at least two reference points include the same laser identifier or the same azimuth angle. In some embodiments, when applying the radial interpolation to the at least two reference points, the G-PCC encoder 200 or the G-PCC decoder 300 can average the radius values of the at least two reference points with equal weights, apply linear interpolation to the radius values of the at least two reference points using the azimuth angles or laser identifiers of the at least two reference points and the current point to obtain a predicted radius of the current point, or apply mathematical trigonometry to the radius values of the at least two reference points.
[0161] In some embodiments, the at least two reference points include three or more reference points, and when applying radial interpolation to the at least two reference points, the G-PCC encoder 200 or the G-PCC decoder 300 may apply a radial interpolation filter having three or more coefficients to the three or more reference points.
[0162] In some embodiments, the current point is one of at least two current points, and the G-PCC encoder 200 or the G-PCC decoder 300 can group the at least two reference points into a first segment, group the at least two current points into a second segment, and generate multiple interpolated radius predictors for the second segment based on the first segment.
[0163] In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may determine that a difference between a radius value of a first reference point of the at least two reference points and a radius value of a second reference point of the at least two reference points is less than a threshold value before applying radial interpolation to the at least two reference points. In some embodiments, the threshold value is a first threshold value, and the G-PCC encoder 200 or the G-PCC decoder 300 may determine a type of interpolation to apply based on comparing the difference between the radius value of the first reference point and the radius value of the second reference point to the first threshold value and a second threshold value.
[0164] In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may at least one of: a) signal or parse syntax elements indicating the interpolation method to be applied, or b) signal or parse syntax elements indicating the coefficients or weights of the interpolation.
[0165] In some embodiments, the G-PCC encoder 200 and the G-PCC decoder 300 may determine the radial residual. In some embodiments, the G-PCC encoder 200 or the G-PCC decoder 300 may context code the radial residual, where the context for context coding the radial residual is separate from the context for context coding the radial residual using intra prediction.
[0166] In some embodiments, before determining the prediction radius, the G-PCC decoder can decode the azimuth angle residual of the current point. In some embodiments, the inverse quantization, inverse scaling, or representation bit depth depends on the decoded radius value of the previously reconstructed point or the radius value of one of the two or more reference points. In some embodiments, in determining the at least two reference points in the reference point cloud frame, the G-PCC encoder 200 or the G-PCC decoder 300 can use the previously decoded points to determine the at least two reference points in the reference point cloud frame.
[0167] In some embodiments, when determining at least two reference points in a reference point cloud frame, the G-PCC encoder 200 or the G-PCC decoder 300 may determine at least two reference points in the reference point cloud frame using a parent node, a grandparent node, or a great-grandparent node of at least one current point.
[0168] In some embodiments, the at least two reference points include two reference points, and in applying the radial interpolation to the at least two reference points, the G-PCC encoder 200 or the G-PCC decoder 300 averages two radius values corresponding to the two reference points. In some embodiments, in averaging the two radius values, the G-PCC encoder 200 or the G-PCC decoder 300 applies a weighted average to the two radius values.
[0169] FIG. 11 is a conceptual diagram illustrating an example distance measurement system 1100 that can be used with one or more techniques of the present disclosure. In the embodiment of FIG. 11, the distance measurement system 1100 includes an illuminator 1102 and a sensor 1104. The illuminator 1102 can emit light 1106. In some embodiments, the illuminator 1102 can emit the light 1106 as one or more laser beams. The light 1106 can be of one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other embodiments, the light 1106 is not a coherent laser light. When the light 1106 encounters an object, such as an object 1108, the light 1106 produces return light 1110. The return light 1110 can include backscattered light and / or reflected light. The returning light 1110 may pass through a lens 1111, which directs the returning light 1110 to produce an image 1112 of the object 1108 on the sensor 1104. The sensor 1104 generates a signal 1114 based on the image 1112. The image 1112 may include a set of points (e.g., as represented by the dots in the image 1112 of FIG. 11).
[0170] In some embodiments, the illuminator 1102 and the sensor 1104 can be mounted on a rotating structure such that the illuminator 1102 and the sensor 1104 capture a 360 degree view of the environment. In other embodiments, the distance measurement system 1100 can include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable the illuminator 1102 and the sensor 1104 to detect the distance of an object within a certain range (e.g., up to 360 degrees). Although the embodiment of FIG. 11 shows only a single illuminator 1102 and sensor 1104, the distance measurement system 1100 can include multiple sets of illuminators and sensors.
[0171] In some embodiments, the illuminator 1102 generates a structured light pattern. In such embodiments, the distance measurement system 1100 may include multiple sensors 1104 on which corresponding images of the structured light patterns are formed. The distance measurement system 1100 may use the difference between the images of the structured light patterns to determine the distance to an object 1108 from which the structured light pattern is backscattered. A structured light based distance measurement system may have a high level of accuracy (e.g., accuracy in the sub-millimeter range) when the object 1108 is relatively close (e.g., between 0.2 meters and 2 meters) to the sensor 1104. This high level of accuracy may be useful in facial recognition applications, such as unlocking a mobile device (e.g., a mobile phone, a tablet computer, etc.), and for security applications.
[0172] In some embodiments, the distance measurement system 1100 is a time of flight (ToF) based system. In some embodiments when the distance measurement system 1100 is a ToF based system, the illuminator 1102 generates a pulse of light. In other words, the illuminator 1102 can modulate the amplitude of the emitted light 1106. In such embodiments, the sensor 1104 detects the return light 1110 from the pulse of light 1106 generated by the illuminator 1102. The distance measurement system 1100 can then determine the distance to the object 1108 from which the light 1106 is backscattered based on the delay time between the emission and detection of the light 1106 and the known speed of light in air. In some embodiments, rather than (or in addition to) modulating the amplitude of the emitted light 1106, the illuminator 1102 can modulate the phase of the emitted light 1106. In such an embodiment, the sensor 1104 can detect the phase of the returning light 1110 from the object 1108 and can determine the distance to a point on the object 1108 using the speed of light and based on the time difference between when the illuminator 1102 generates the light 1106 at a particular phase and when the sensor 1104 detects the returning light 1110 at that particular phase.
[0173] In other examples, a point cloud can be generated without the use of an illuminator 1102. For example, in some examples, the sensor 1104 of the distance measurement system 1100 can include two or more optical cameras. In such examples, the distance measurement system 1100 can use the optical cameras to capture a stereo image of an environment, including the object 1108. The distance measurement system 1100 can include a point cloud generator 1116 that can calculate disparities between locations in the stereo image. The distance measurement system 1100 can then use those disparities to determine distances to locations shown in the stereo image. From these distances, the point cloud generator 1116 can generate a point cloud.
[0174] The sensor 1104 may also detect other attributes of the object 1108, such as color and reflectance information. In the example of Fig. 11, a point cloud generator 1116 may generate a point cloud based on the signal 1114 generated by the sensor 1104. The distance measurement system 1100 and / or the point cloud generator 1116 may form part of the data source 104 (Fig. 1). Thus, the point cloud generated by the distance measurement system 1100 may be encoded and / or decoded according to any of the techniques of this disclosure.
[0175] FIG. 12 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of the present disclosure may be used. In the example of FIG. 12, a vehicle 1200 includes a distance measurement system 1202. The distance measurement system 1202 may be implemented in a manner discussed with respect to FIG. 12. Although not shown in the example of FIG. 12, the vehicle 1200 may also include a data source, such as the data source 104 (FIG. 1), and a G-PCC encoder, such as the G-PCC encoder 200 (FIG. 1). In the example of FIG. 12, the distance measurement system 1202 emits a laser beam 1204 that reflects off a pedestrian 1206 or other object in a road. The data source of the vehicle 1200 may generate a point cloud based on the signal generated by the distance measurement system 1202. The G-PCC encoder of the vehicle 1200 may encode the point cloud to generate a bit stream 1208, such as a geometry bit stream (FIG. 2) and an attribute bit stream (FIG. 2). The bitstream 1208 may contain many fewer bits than the unencoded point cloud obtained by the G-PCC encoder. In some embodiments, the G-PCC encoder of the vehicle 1200 may encode the bitstream 1208 using radial interpolation as described above. In some embodiments, the G-PCC decoder of the vehicle 1210 may decode the bitstream 1208 using radial interpolation as described above.
[0176] An output interface of the vehicle 1200 (e.g., output interface 108 (FIG. 1)) can transmit the bitstream 1208 to one or more other devices. The bitstream 1208 may include much fewer bits than the unencoded point cloud obtained by the G-PCC encoder. Therefore, the vehicle 1200 may be able to transmit the bitstream 1208 to other devices more quickly than the unencoded point cloud data. Furthermore, the bitstream 1208 may require less data storage capacity.
[0177] In the example of FIG. 12, the vehicle 1200 can transmit the bitstream 1208 to another vehicle 1210. The vehicle 1210 can have a G-PCC decoder, such as the G-PCC decoder 300 (FIG. 1). The G-PCC decoder of the vehicle 1210 can decode the bitstream 1208 to reconstruct a point cloud. The vehicle 1210 can use the reconstructed point cloud for various purposes. For example, the vehicle 1210 can determine based on the reconstructed point cloud that a pedestrian 1206 is present in the road ahead of the vehicle 1200, and can therefore begin to decelerate, for example, even before the driver of the vehicle 1210 recognizes that a pedestrian 1006 is present in the road. Thus, in some examples, the vehicle 1210 can perform autonomous navigation operations based on the reconstructed point cloud.
[0178] Additionally or alternatively, the vehicle 1200 can transmit the bitstream 1208 to a server system 1212. The server system 1212 can use the bitstream 1208 for various purposes. For example, the server system 1212 can store the bitstream 1208 for subsequent reconstruction of a point cloud. In this example, the server system 1212 can use the point cloud along with other data (e.g., vehicle telemetry data generated by the vehicle 1200) to train an autonomous driving system. In other examples, the server system 1212 can store the bitstream 1208 for subsequent reconstruction for forensic crash investigation.
[0179] FIG. 13 is a conceptual diagram illustrating an example extended reality system that can use one or more techniques of the present disclosure. Extended reality (XR) is a term used to encompass a wide range of technologies, including augmented reality (AR), mixed reality (MR), and virtual reality (VR). In the example of FIG. 13, a user 1300 is located at a first location 1302. The user 1300 is wearing an XR headset 1304. As an alternative to the XR headset 1304, the user 1300 can also use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1304 includes a depth detection sensor, such as a distance measurement system, that detects the position of a point on an object 1306 at the location 1302. A data source of the XR headset 1304 can use the signal generated by the depth detection sensor to generate a point cloud representation of the object 1306 at the location 1302. The XR headset 1304 may include a G-PCC encoder (e.g., the G-PCC encoder 200 of FIG. 1) configured to encode the point cloud to generate a bitstream 1308. In some embodiments, the G-PCC encoder of the XR headset 1304 may use radial interpolation when encoding the point cloud, as described above.
[0180] The XR headset 1304 can transmit the bitstream 1308 (e.g., over a network such as the Internet) to an XR headset 1310 worn by a user 1312 at a second location 1314. The XR headset 1310 can decode the bitstream 1308 to reconstruct the point cloud. In some embodiments, the G-PCC decoder of the XR headset 1310 can use radial interpolation in decoding the point cloud, as described above.
[0181] The XR headset 1310 can use the point cloud to generate an XR visualization (e.g., AR, MR, VR visualization) that represents the object 1306 at the location 1302. Thus, in some examples, such as when the XR headset 1310 generates a VR visualization, the user 1312 can have a 3D immersive experience of the location 1302. In some examples, the XR headset 1310 can determine a location of a virtual object based on the reconstructed point cloud. For example, the XR headset 1310 can determine that the environment (e.g., the location 1302) includes a flat surface based on the reconstructed point cloud, and then determine that a virtual object (e.g., a cartoon character) is to be positioned on the flat surface. The XR headset 1310 can generate an XR visualization in which the virtual object is present at the determined location. For example, the XR headset 1310 can show a cartoon character sitting on the flat surface.
[0182] 14 is a conceptual diagram illustrating an example mobile device system that may use one or more techniques of this disclosure. In the example of FIG. 14, a mobile device 1400, such as a mobile phone or tablet computer, includes a distance measurement system, such as a LIDAR system, that detects the location of points on an object 1402 in the environment of the mobile device 1400. A data source of the mobile device 1400 may generate a point cloud representation of the object 1402 using signals generated by a depth detection sensor. The mobile device 1400 may include a G-PCC encoder (e.g., the G-PCC encoder 200 of FIG. 1) configured to encode the point cloud to generate a bitstream 1404.
[0183] In the example of Figure 14, the mobile device 1200 can transmit the bitstream to a remote device 1406, such as a server system or another mobile device. The remote device 1406 can decode the bitstream 1404 to reconstruct the point cloud. In some examples, the G-PCC decoder of the remote device 1406 can use radial interpolation in decoding the point cloud, as described above.
[0184] The remote device 1406 can use the point cloud for various purposes. For example, the remote device 1406 can use the point cloud to generate a map of the environment of the mobile device 1400. For example, the remote device 1406 can generate a map of the interior of a building based on the reconstructed point cloud. In another example, the remote device 1406 can generate imagery (e.g., computer graphics) based on the point cloud. For example, the remote device 1406 can use the points of the point cloud as vertices of a polygon and use the color attributes of the points as a basis for shading the polygon. In some examples, the remote device 1406 can use the reconstructed point cloud for facial recognition or other security applications.
[0185] The embodiments of the various aspects of the present disclosure may be used individually or in any combination.
[0186] This disclosure includes the following provisions:
[0187] Clause 1A. A method for coding point cloud data, comprising: 1. A method comprising: determining at least two reference points in a reference point cloud frame of point cloud data; applying radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of point cloud data; and coding the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0188] Clause 2A. The method of clause 1A, wherein at least two reference points belong to a group of reference points within a range of azimuth angle, laser identifier (ID), or radius.
[0189] Clause 3A. The method of clause 1A or clause 2A, further comprising compensating the reference point cloud frame for motion.
[0190] Clause 4A. The method of clause 3A, wherein compensating the reference point cloud frame for motion includes applying at least one of a rotation, a translation, or a local motion.
[0191] Clause 5A. The method of any of clauses 1A-4A, further comprising quantizing at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0192] Clause 6A. The method of any of clauses 1A-5A, further comprising approximating at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0193] Clause 7A. The method of any of clauses 1A-6A, further comprising scaling at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0194] Clause 8A. The method of any of clauses 5A-7A, further comprising signaling or parsing a bit depth of the position coordinates.
[0195] Clause 9A. The method of clause 5A, further comprising constructing an array or map of position coordinates of the at least two reference points to further quantize the position coordinates of the at least two reference points.
[0196] Clause 10A. The method of any of clauses 1A-8A, wherein at least two reference points have the same or similar laser identifier (ID), and the method further includes ordering the at least two reference points according to corresponding azimuth angles associated with the at least two reference points.
[0197] Clause 11A. The method of clause 10A, wherein the at least two reference points have the same azimuth angle, the method further including further ordering the at least two reference points according to corresponding radii associated with the at least two reference points.
[0198] Clause 12A. The method of any of clauses 1A-8A, wherein at least two reference points have the same or similar azimuth angles, and the method further includes ordering the at least two reference points according to corresponding laser identifiers (IDs) associated with the at least two reference points.
[0199] Clause 13A. The method of clause 12A, wherein at least two reference points have the same laser ID, the method further including further ordering the at least two reference points according to corresponding radii associated with the at least two reference points.
[0200] Clause 14A. The method of any of clauses 1A to 13A, wherein the positions of at least two reference points are associated with the position of at least one current point.
[0201] Clause 15A. The method of any of clauses 1A-14A, further comprising determining positions of at least two reference points based on an azimuth angle and a laser identifier (ID) of a first current point of the at least one current point.
[0202] Clause 16A. The method of clause 15A, wherein determining the position of the at least two reference points includes searching for a first closest azimuth angle that is smaller than the azimuth angle of the first current point to determine the position of a first reference point of the at least two reference points, and searching for a second closest azimuth angle that is greater than the azimuth angle of the first current point to determine the position of a second reference point of the at least two reference points.
[0203] Clause 17A. The method of clause 15A, wherein determining the location of the at least two reference points includes searching for a first closest laser ID that is smaller than a laser identifier (ID) of the first current point to determine the location of a first reference point of the at least two reference points, and searching for a second closest laser ID that is greater than the laser ID of the first current point to determine the location of a second reference point of the at least two reference points.
[0204] Clause 18A. The method of clause 16A or clause 17A, further comprising determining an additional reference point of the at least two reference points.
[0205] Clause 19A. The method of any of clauses 1A-18A, further comprising determining positions of at least two reference points based on an azimuth angle and a range of laser identifiers (IDs).
[0206] Clause 20A. The method of clause 19A, wherein the azimuth angle and laser identifier range are based on an interpolation method.
[0207] Clause 21A. The method of any of clauses 1A-20A, wherein one or more of the at least two reference points are unavailable or include a redundant predictor, and the method further includes replacing the unavailable predictor or the redundant predictor.
[0208] Clause 22A. The method of clause 21A, wherein replacing the unavailable or redundant predictor includes replacing the unavailable or redundant predictor with a default predictor, a reference point having the next closest laser identifier (ID), or discarding the unavailable or redundant predictor.
[0209] Clause 23A. The method of any of clauses 1A-22A, wherein at least two reference points include an identical laser identifier (ID).
[0210] Clause 24A. The method of clause 23A, further comprising averaging with equal weights a radius associated with a first reference point of the at least two reference points and a radius associated with a second reference point of the at least two reference points.
[0211] Clause 25A. The method of clause 23A, further comprising applying linear interpolation to azimuth angles of the at least two reference points and the current point to obtain a predicted radius of the current point.
[0212] Clause 26A. The method of clause 23A, further comprising applying mathematical trigonometry to determine the interpolation radius.
[0213] Clause 27A. The method of clause 23A, wherein the at least two reference points include three or more reference points, the method further comprising applying a radial interpolation filter having three or more coefficients to the three or more reference points.
[0214] Clause 28A. The method of any of clauses 23A-26A, further comprising grouping two or more reference points into a first segment, grouping two or more current points into a second segment, and generating a plurality of interpolated radius predictors for the second segment based on the first segment.
[0215] Clause 29A. The method of any of clauses 23A to 27A, further comprising determining that a difference between a radius value of a first reference point of the two or more reference points and a radius value of a second reference point of the two or more reference points is less than a threshold value.
[0216] Clause 30A. The method of clause 28A, wherein the threshold is predetermined or the threshold is signaled in the bitstream.
[0217] Clause 31A. The method of clause 28A or clause 29A, wherein the threshold is a first threshold, and the method further includes determining a type of interpolation to apply based on comparing a difference between the radius value of the first reference point and the radius value of the second reference point with the first threshold and a second threshold.
[0218] Clause 32A. The method of any of clauses 23A to 31A, further comprising signaling or parsing a syntax element indicating the interpolation method to be applied.
[0219] Clause 33A. The method of any of clauses 23A to 32A, further comprising signaling or parsing syntax elements indicating coefficients or weights of the interpolation.
[0220] Clause 34A. The method of any of clauses 23A-33A, further comprising determining a radius residual.
[0221] Clause 35A. The method of clause 34A, wherein the radius residual includes a sign and an absolute value.
[0222] Clause 36A. The method of clause 34A or 35A, further comprising context coding the radial residual, wherein a context for context coding the radial residual is separate from a context for context coding the radial residual using intra prediction.
[0223] Clause 37A. The method of any of clauses 34A-36A, further comprising decoding an azimuth angle of the current point prior to determining the predicted radius.
[0224] Clause 38A. The method of clause 37A, further comprising determining an azimuth angle predictor for the current point and decoding the azimuth angle residual.
[0225] Clause 39A. The method of clause 38A, wherein the inverse quantization, inverse scaling, or representation bit depth is dependent on a previously reconstructed decoding radius value of the current point.
[0226] Clause 40A. The method of clause 38A, wherein the inverse quantization, inverse scaling, or representation bit depth is dependent on a radius value of one of the two or more reference points.
[0227] Clause 41A. The method of any of clauses 1A to 40A, further comprising generating a point cloud.
[0228] Clause 42A. Any of the methods of clauses 1A to 41A, wherein determining at least two reference points in the reference point group frame includes determining at least two reference points in the reference point group frame using previously decoded points.
[0229] Clause 43A. Any of the methods of clauses 1A to 41A, wherein determining at least two reference points in the reference point cloud frame includes determining at least two reference points in the reference point cloud frame using a parent node, a grandparent node, or a great-grandparent node of at least one current point.
[0230] Clause 44A. The method of any of clauses 1A-43A, wherein the at least two reference points include two reference points, and applying radial interpolation to the at least two reference points includes averaging two radius values corresponding to the two reference points.
[0231] Clause 45A. The method of clause 44A, in which averaging the two radius values includes applying a weighted average to the two radius values.
[0232] Clause 46A. The method of any of clauses 1A-43A, wherein the at least two reference points include three or more reference points and applying radial interpolation utilizes the three or more reference points.
[0233] Clause 47A. The method of any of clauses 1A-45A, further comprising refraining from initially coding an azimuth angle of at least one current point.
[0234] Clause 48A. A device for processing point clouds, the device comprising one or more means for carrying out the methods of any of clauses 1A to 47A.
[0235] Clause 49A. A device of clause 48, wherein the one or more means include one or more processors implemented in circuitry.
[0236] Clause 50A. The device of any of clauses 48A or 49A, further comprising a memory for storing data representing the point cloud.
[0237] Clause 51A. A device of any of clauses 48A to 50A, wherein the device comprises a decoder.
[0238] Clause 52A. The device of any of clauses 48A to 51A, wherein the device comprises an encoder.
[0239] Clause 53A. The device of any of clauses 48A to 50A, further comprising a device for generating a point cloud.
[0240] Clause 54A. The device of any of clauses 48A to 53A, further comprising a display for presenting an image based on the point cloud.
[0241] Clause 55A. A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform any of the methods of clauses 1A to 47A.
[0242] Clause 1B. A method for coding point cloud data, comprising: determining at least two reference points in a reference point cloud frame of the point cloud data; applying radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data; and coding the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0243] Clause 2B. The method of clause 1B, wherein at least two of the reference points belong to a group of reference points within a range of at least one of an azimuth angle, a laser identifier (ID), or a radius.
[0244] Clause 3B. The method of clause 1B or clause 2B, further comprising compensating the reference point cloud frame for motion by applying at least one of a rotation, a translation, or a local motion.
[0245] Clause 4B. The method of clauses 1B-3B, further comprising at least one of quantizing, approximating, or scaling at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0246] Clause 5B. The method of clause 4B, further comprising signaling or analyzing a bit depth of at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0247] Clause 6B. The method of clause 4B, further comprising constructing an array or map of position coordinates of the at least two reference points to quantize the position coordinates of the at least two reference points.
[0248] Clause 7B. The method of clauses 1B-6B, wherein at least two reference points have the same or adjacent laser identifiers (IDs), and the method further includes ordering the at least two reference points according to corresponding azimuth angles associated with the at least two reference points.
[0249] Clause 8B. The method of clause 7B, wherein the at least two reference points have the same azimuth angle, the method further including further ordering the at least two reference points according to corresponding radii associated with the at least two reference points.
[0250] Clause 9B. The method of clauses 1B-6B, wherein at least two reference points have the same or adjacent azimuth angles, and the method further includes ordering the at least two reference points according to corresponding laser identifiers associated with the at least two reference points.
[0251] Clause 10B. The method of clause 9B, wherein at least two reference points have an identical laser identifier, the method further including further ordering the at least two reference points according to corresponding radii associated with the at least two reference points.
[0252] Clause 11B. The method of any of clauses 1B-10B, wherein the positions of at least two reference points are associated with the position of at least one current point.
[0253] Clause 12B. The method of any of clauses 1B-11B, further comprising determining positions of at least two reference points based on an azimuth angle and a laser identifier of a first current point of the at least one current point.
[0254] Clause 13B. The method of clause 12B, wherein determining a position of the at least two reference points includes searching for a first closest azimuth angle that is smaller than the azimuth angle of the first current point to determine a position of a first reference point of the at least two reference points, and searching for a second closest azimuth angle that is greater than the azimuth angle of the first current point to determine a position of a second reference point of the at least two reference points.
[0255] Clause 14B. The method of clause 12B, wherein determining the location of the at least two reference points includes searching for a first closest laser ID that is smaller than a laser identifier (ID) of the first current point to determine the location of a first reference point of the at least two reference points, and searching for a second closest laser ID that is greater than the laser ID of the first current point to determine the location of a second reference point of the at least two reference points.
[0256] Clause 15B. The method of any of clauses 1B-11B, further comprising determining positions of at least two reference points based on the azimuth angle and the range of the laser identifier.
[0257] Clause 16B. The method of clause 15B, based on a radial interpolation method in which the azimuth angle and the laser identifier range are applied to at least two reference points.
[0258] Clause 17B. The method of any of clauses 1B-16B, wherein one or more of the at least two reference points is unavailable or includes a duplicate predictor, and the method further includes replacing the unavailable or duplicate predictor with a default predictor or a reference point having a next closest laser identifier or azimuth angle, or discarding the unavailable or duplicate predictor.
[0259] Clause 18B. The method of any of clauses 1B-17B, wherein at least two reference points include the same laser identifier or the same azimuth angle.
[0260] Clause 19B. The method of clause 18B, wherein applying radial interpolation to the at least two reference points includes averaging radius values of the at least two reference points with equal weights, applying linear interpolation to the radius values of the at least two reference points using azimuth angles or laser identifiers of the at least two reference points and the current point to obtain a predicted radius of the current point, or applying mathematical trigonometry to the radius values of the at least two reference points.
[0261] Clause 20B. The method of clause 18B, wherein the at least two reference points include three or more reference points, and applying radial interpolation to the at least two reference points includes applying a radial interpolation filter having three or more coefficients to the three or more reference points.
[0262] Clause 21B. The method of clause 18B, wherein the current point is one of at least two current points, and the method further includes grouping the at least two reference points into a first segment, grouping the at least two current points into a second segment, and generating a plurality of interpolated radius predictors for the second segment based on the first segment.
[0263] Clause 22B. The method of clause 18B, wherein the at least two reference points include two reference points, and applying radial interpolation to the at least two reference points includes averaging two radius values corresponding to the two reference points.
[0264] Clause 23B. The method of clause 22B, wherein averaging the two radius values includes applying a weighted average to the two radius values.
[0265] Clause 24B. The method of any of clauses 18B-23B, further comprising determining that a difference between a radius value of a first reference point of the at least two reference points and a radius value of a second reference point of the at least two reference points is less than a threshold value before applying radial interpolation to the at least two reference points.
[0266] Clause 25B. The method of clause 24B, wherein the threshold is a first threshold, and the method further includes determining a type of interpolation to apply based on comparing a difference between the radius value of the first reference point and the radius value of the second reference point to the first threshold and a second threshold.
[0267] Clause 26B. Any of the methods of clauses 18B to 25B further comprising at least one of: a) signaling or parsing a syntax element indicating the interpolation method to be applied; or b) signaling or parsing a syntax element indicating the coefficients or weights of the interpolation.
[0268] Clause 27B. The method of any of clauses 1B-26B, further comprising determining a radius residual.
[0269] Clause 28B. The method of clause 27B, further comprising context coding the radius residual, wherein a context for context coding the radius residual is separate from a context for context coding the radius residual using intra prediction.
[0270] Clause 29B. The method of clause 27B or clause 28B, further comprising decoding an azimuth angle residual for the current point prior to determining the predicted radius.
[0271] Clause 30B. The method of clause 29B, wherein the inverse quantization, inverse scaling, or representation bit depth is dependent on a decoded radius value of a previously reconstructed point or a radius value of one of the two or more reference points.
[0272] Clause 31B. Any of the methods of clauses 1B to 30B, wherein determining at least two reference points in the reference point group frame includes determining at least two reference points in the reference point group frame using previously decoded points.
[0273] Clause 32B. Any of the methods of clauses 1B to 30B, wherein determining at least two reference points in the reference point cloud frame includes determining at least two reference points in the reference point cloud frame using a parent node, a grandparent node, or a great-grandparent node of at least one current point.
[0274] Clause 33B. A device for coding point cloud data, comprising: a memory configured to store the point cloud data; and one or more processors communicatively coupled to the memory, the one or more processors configured to determine at least two reference points in a reference point cloud frame of the point cloud data, apply radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data, and code the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0275] Clause 34B. The device of clause 33B, wherein at least two of the reference points belong to a group of reference points within a range of at least one of an azimuth angle, a laser identifier (ID), or a radius.
[0276] Clause 35B. The device of clause 33B or clause 34B, wherein the one or more processors are further configured to compensate the reference point cloud frame for motion by applying at least one of rotation, translation, or local motion.
[0277] Clause 36B. The device of clauses 33B to 35B, wherein the one or more processors are further configured to at least one of quantize, approximate, or scale at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
[0278] Clause 37B. The device of clause 36B, wherein the one or more processors are further configured to signal or analyze a bit depth of at least one of the position coordinates of the at least two reference points or the position coordinate of the at least one current point.
[0279] Clause 38B. The device of clause 36B, wherein the one or more processors are further configured to construct an array or map of position coordinates of the at least two reference points to quantize the position coordinates of the at least two reference points.
[0280] Clause 39B. The device of clauses 33B-38B, wherein at least two reference points have the same or adjacent laser identifiers (IDs), and the one or more processors are further configured to order the at least two reference points according to corresponding azimuth angles associated with the at least two reference points.
[0281] Clause 40B. The device of clause 39B, wherein the at least two reference points have the same azimuth angle, and the one or more processors are further configured to further order the at least two reference points according to corresponding radii associated with the at least two reference points.
[0282] Clause 41B. The device of clauses 33B-38B, wherein the at least two reference points have the same or adjacent azimuth angles, and the one or more processors are further configured to order the at least two reference points according to corresponding laser identifiers associated with the at least two reference points.
[0283] Clause 42B. The device of clause 41B, wherein at least two reference points have an identical laser identifier, and the one or more processors are further configured to further order the at least two reference points according to corresponding radii associated with the at least two reference points.
[0284] Clause 43B. A device according to any of clauses 33B to 42B, wherein the positions of at least two reference points are associated with the position of at least one current point.
[0285] Clause 44B. The device of any of clauses 33B to 43B, wherein the one or more processors are further configured to determine positions of at least two reference points based on an azimuth angle and a laser identifier of a first current point of the at least one current point.
[0286] Clause 45B. The device of clause 44B, wherein as part of determining the position of the at least two reference points, the one or more processors are configured to search for a first closest azimuth angle that is smaller than the azimuth angle of the first current point to determine the position of a first reference point of the at least two reference points, and to search for a second closest azimuth angle that is greater than the azimuth angle of the first current point to determine the position of a second reference point of the at least two reference points.
[0287] Clause 46B. The device of clause 44B, wherein as part of determining the position of the at least two reference points, the one or more processors are configured to search for a first closest laser ID that is smaller than a laser ID of the first current point to determine the position of the first reference point of the at least two reference points, and to search for a second closest laser ID that is greater than the laser ID of the first current point to determine the position of the second reference point of the at least two reference points.
[0288] Clause 47B. The device of any of clauses 33B to 43B, wherein the one or more processors are further configured to determine positions of at least two reference points based on an azimuth angle and a range of the laser identifier.
[0289] Clause 48B. The device of clause 47B, wherein the azimuth angle and the laser identifier range are based on a radial interpolation method to be applied to at least two reference points.
[0290] Clause 49B. The device of any of clauses 33B to 48B, wherein one or more of the at least two reference points are unavailable or include a duplicate predictor, and the one or more processors are further configured to replace the unavailable or duplicate predictor with a default predictor or a reference point having a next-closest laser identifier or azimuth angle, or to discard the unavailable or duplicate predictor.
[0291] Clause 50B. The device of any of clauses 33B to 49B, wherein at least two reference points include the same laser identifier or the same azimuth angle.
[0292] Clause 51B. The device of clause 50B, wherein as part of applying radial interpolation to the at least two reference points, the one or more processors are configured to average the radius values of the at least two reference points with equal weights, apply linear interpolation to the radius values of the at least two reference points using azimuth angles or laser identifiers of the at least two reference points and the current point to obtain a predicted radius of the current point.
[0293] Clause 52B. The device of clause 50B, wherein the at least two reference points include three or more reference points, and as part of applying radial interpolation to the at least two reference points, the one or more processors are configured to apply a radial interpolation filter having three or more coefficients to the three or more reference points.
[0294] Clause 53B. The device of clause 50B, wherein the current point is one of at least two current points, and the one or more processors are further configured to group the at least two reference points into a first segment, group the at least two current points into a second segment, and generate, for the second segment, a plurality of interpolated radius predictors based on the first segment.
[0295] Clause 54B. The device of any of clauses 50B, wherein the at least two reference points include two reference points, and as part of applying radial interpolation to the at least two reference points, the one or more processors are configured to average two radius values corresponding to the two reference points.
[0296] Clause 55B. The device of clause 54B, wherein averaging the two radius values includes applying a weighted average to the two radius values.
[0297] Clause 56B. The device of any of clauses 51B to 55B, wherein the one or more processors are further configured to determine that before applying radial interpolation to the at least two reference points, a difference between a radius value of a first reference point of the at least two reference points and a radius value of a second reference point of the at least two reference points is less than a threshold value.
[0298] Clause 57B. The device of clause 56B, wherein the threshold is a first threshold, and the one or more processors are further configured to determine a type of interpolation to apply based on comparing a difference between the radius value of the first reference point and the radius value of the second reference point to the first threshold and a second threshold.
[0299] Clause 58B. The device of any of clauses 50B to 57B, wherein the one or more processors are further configured to at least one of: a) signal or parse syntax elements indicating an interpolation method to be applied; or b) signal or parse syntax elements indicating interpolation coefficients or weights.
[0300] Clause 59B. The device of any of clauses 50B-58B, wherein the one or more processors are further configured to determine a radius residual.
[0301] Clause 60B. The device of clause 59B, wherein the one or more processors are further configured to context code the radial residual, and wherein a context for context coding the radial residual is separate from a context for context coding the radial residual using intra prediction.
[0302] Clause 61B. The device of clause 59B or clause 60B, wherein the one or more processors are further configured to decode an azimuth angle residual for the current point before determining the predicted radius.
[0303] Clause 62B. The device of clause 61B, wherein the inverse quantization, inverse scaling, or representation bit depth is dependent on a decoded radius value of a previously reconstructed point or a radius value of one of two or more reference points.
[0304] Clause 63B. The device of any of clauses 33B to 62B, wherein as part of determining at least two reference points in the reference point cloud frame, the one or more processors are configured to determine at least two reference points in the reference point cloud frame using previously decoded points.
[0305] Clause 64B. The device of any of clauses 33B to 62B, wherein as part of determining at least two reference points in the reference point cloud frame, the one or more processors are further configured to determine at least two reference points in the reference point cloud frame using a parent node, a grandparent node, or a great-grandparent node of at least one current point.
[0306] Clause 65B. The device is any device of clauses 33B to 64B, including a vehicle, a robot, an extended reality system, or a smartphone.
[0307] Clause 66B. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine at least two reference points in a reference point cloud frame of point cloud data, apply radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of point cloud data, and code the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0308] Clause 67B. A device for coding point cloud data, comprising: means for determining at least two reference points in a reference point cloud frame of the point cloud data; means for applying radial interpolation to the at least two reference points to obtain at least one radial inter predictor for at least one current point in a current point cloud frame of the point cloud data; and means for coding the current point cloud frame based on the at least one radial inter predictor for the at least one current point in the current point cloud frame.
[0309] It should be appreciated that, depending on the embodiment, certain acts or events of any of the techniques described herein may be performed in a different order, or may be added, combined, or omitted entirely (e.g., not all acts or events described are required to practice the techniques). Furthermore, in certain embodiments, acts or events may be performed simultaneously rather than sequentially, for example, through multithreading, interrupt processing, or multiple processors.
[0310] In one or more embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium, which includes any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0311] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where a disk typically reproduces data magnetically, while a disc reproduces data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0312] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the terms "processor" and "processing circuitry" as used herein may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a composite codec. These techniques may also be fully implemented in one or more circuits or logic elements.
[0313] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, the various units can be combined into a codec hardware unit or can be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0314] Various embodiments have been described. These and other embodiments are within the scope of the following claims.
Claims
1. 1. A method for coding point cloud data, comprising: determining at least two reference points within a reference point cloud frame of the point cloud data, the at least two reference points being expressed in a (radius, azimuth, laser identifier) coordinate domain; applying radial interpolation to the at least two reference points to obtain at least one radial inter-predictor for at least one current point in a current point cloud frame of the point cloud data; coding the current point cloud frame based on the at least one radial inter-predictor for the at least one current point in the current point cloud frame; wherein the at least two reference points include the same laser identifier.
2. the at least two reference points belong to a group of reference points within a range of at least one of an azimuth angle, a laser identifier (ID), or a radius; The method of claim 1 , wherein the method comprises applying the radial interpolation to the at least two reference points belonging to the group of reference points.
3. The method of claim 1 , further comprising compensating the reference point cloud frame for motion by applying at least one of a rotation, a translation, or a local motion.
4. 2. The method of claim 1, further comprising at least one of quantizing, approximating, or scaling at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
5. The method of claim 4 , further comprising signaling or analyzing a bit depth of at least one of the position coordinates of the at least two reference points or the position coordinates of the at least one current point.
6. The method of claim 4 , further comprising constructing an array or map of the position coordinates of the at least two reference points to quantize the position coordinates of the at least two reference points.
7. The method described in claim 1, further comprising ordering the at least two reference points according to their respective azimuth angles associated with the at least two reference points.
8. the at least two reference points have adjacent azimuth angles; The method further includes ordering the at least two reference points according to respective laser identifiers associated with the at least two reference points; the at least two reference points have the same laser identifier; The method of claim 1 , wherein the method further comprises further ordering the at least two reference points according to respective radii associated with the at least two reference points.
9. The method of claim 1 , further comprising determining positions of the at least two reference points based on an azimuth angle and a laser identifier of a first current point of the at least one current point.
10. Determining the locations of the at least two reference points searching for a first closest azimuth angle that is less than the azimuth angle of the first current point to determine a position of a first reference point of the at least two reference points; searching for a second closest azimuth angle greater than the azimuth angle of the first current point to determine a position of a second reference point of the at least two reference points; 10. The method of claim 9, comprising:
11. Determining the locations of the at least two reference points searching for a first closest laser identifier (ID) that is less than a laser ID of the first current point to determine a location of a first reference point of the at least two reference points; searching for a second closest laser ID greater than the laser ID of the first current point to determine a location of a second reference point of the at least two reference points; The method of claim 10, comprising:
12. determining positions of the at least two reference points based on an azimuth angle and a range of laser identifiers (IDs); The method of claim 1 , wherein the ranges of azimuth angles and laser identifiers are based on a radial interpolation method to be applied to the at least two reference points.
13. 1. A device for coding point cloud data, comprising: a memory configured to store the point cloud data; one or more processors communicatively coupled to the memory, determining at least two reference points within a reference point cloud frame of the point cloud data, the at least two reference points being expressed in a (radius, azimuth, laser identifier) coordinate domain; applying radial interpolation to the at least two reference points to obtain at least one radial inter-predictor for at least one current point in a current point cloud frame of the point cloud data; coding the current point cloud frame based on the at least one radial inter-predictor for the at least one current point in the current point cloud frame; one or more processors configured to wherein the at least two reference points include an identical laser identifier.
14. The device of claim 13, further comprising means for performing a method according to any one of claims 2 to 12.
15. 13. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any one of claims 1 to 12.