Angular modes and intra-tree quantization in geometry point cloud compression
By scaling quantized values without clipping to align coordinate positions with the origin, the technique addresses reduced efficiency in conventional point cloud coding, achieving enhanced coding gains from angular mode and in-tree quantization.
Patent Information
- Application Number
- JP2023520503
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-06
- Filing Date
- 2021-10-07
- Publication Date
- 2026-01-28
- Estimated Expiration
- 2041-10-07
AI Technical Summary
Conventional point cloud coding methods suffer reduced coding efficiency when angular mode and in-tree quantization are enabled together, as quantization bits are in a different scale space and not in the same domain as the original points, making simultaneous use of both techniques less beneficial.
Techniques for deriving scaled values from point/position coordinate values without clipping, allowing the G-PCC decoder to determine coordinate positions relative to the origin in the same scale space, thereby enhancing coding gains from angular mode with in-tree quantization.
Enables improved coding efficiency by aligning quantized values with the original point values' scale space, maintaining accuracy and enhancing coding gains when angular mode and in-tree quantization are used together.
Smart Images

Figure 0007808099000002 
Figure 0007808099000003 
Figure 0007808099000004
Abstract
Description
[Technical Field]
[0001] This application is U.S. Patent Application No. 17 / 495,621, filed October 6, 2021; U.S. Provisional Patent Application No. 63 / 088,938, filed October 7, 2020; U.S. Provisional Patent Application No. 63 / 090,629, filed October 12, 2020; and This application claims priority to U.S. Provisional Patent Application No. 63 / 091,821, filed October 14, 2020, the entire contents of which are incorporated herein by reference. U.S. Patent Application No. 17 / 495,621, filed October 6, 2021, U.S. Provisional Patent Application No. 63 / 088,938, filed October 7, 2020; U.S. Provisional Patent Application No. 63 / 090,629, filed October 12, 2020; and This application claims the benefit of U.S. Provisional Patent Application No. 63 / 091,821, filed October 14, 2020.
[0002] The present disclosure relates to point cloud encoding and decoding. [Background technology]
[0003] A point cloud is a collection of points in three-dimensional space. The points may correspond to points on an object in the three-dimensional space. Thus, a point cloud can be used to represent the physical content of a three-dimensional space. Point clouds can have utility in a wide variety of situations. For example, a point cloud can be used in the context of autonomous vehicles to represent the location of objects on a road. In another example, a point cloud can be used in the context of representing the physical content of an environment for purposes of positioning virtual objects in augmented reality (AR) or mixed reality (MR) applications. Point cloud compression is a process for encoding and decoding a point cloud. Encoding a point cloud can reduce the amount of data required for storage and transmission of the point cloud. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] G-PCC DIS, ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, January 2020 [Non-patent document 2] G-PCC Codec Description v6, ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, January 2020 [Non-patent document 3] Sebastien Lasserre, Jonathan Taquet, "[GPCC][CE 13.22 related] An improvement of the planar coding mode," ISO / IEC JTC1 / SC29 / WG11 MPEG / m50642, Geneva, Switzerland, October 2019. [Non-patent document 4] Sebastien Lasserre, David Flynn, "Planar mode in octree-based geometry coding," ISO / IEC JTC1 / SC29 / WG11 MPEG / m48906, Gothenburg, Sweden, July 2019 [Non-Patent Document 5] Sebastien Lasserre, Jonathan Taquet, "[GPCC] CE 13.22 report on angular mode," ISO / IEC JTC1 / SC29 / WG11 MPEG / m51594, Brussels, Belgium, January 2020. Summary of the Invention [Problem to be solved by the invention]
[0005] In conventional coding of point cloud frames, angular mode provides substantial gains in coding efficiency. However, when in-tree quantization is enabled, the gains achieved from angular mode become much smaller, and in some cases even produce losses. In angular mode, quantization bits are used for context derivation, and the quantization bits are in a different scale space and not in the same domain as the original points. This reduces the usefulness of both angular mode and in-tree quantization, and therefore, it may not be beneficial to enable both simultaneously. [Means for solving the problem]
[0006] This disclosure describes techniques for deriving scaled values xS from point / position coordinate values x to derive the node / point position relative to the lidar origin when angle mode and in-tree quantization are used together. More specifically, by scaling quantized values representing coordinate values without clipping, a G-PCC decoder can determine scaled values representing coordinate positions relative to the origin position in a way that puts the scaled values in the same scale space as the original point values with sufficient accuracy to achieve coding gain from angle mode in conjunction with in-tree quantization.
[0007] According to one example, a device for decoding a bitstream including point cloud data comprises: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in a circuit configuration, wherein the one or more processors are configured to: determine, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determine, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determine a quantized value for the node representing a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value representing the coordinate position relative to the origin position; and determine a context for context decoding a planar position syntax element for the angular mode based on the scaled value representing the coordinate position relative to the origin position.
[0008] According to another example, a method for decoding a bitstream including point cloud data, the method comprising: determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determining, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0009] According to another example, a device for encoding a bitstream including point cloud data comprises: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in a circuit configuration, wherein the one or more processors are configured to: determine that intra-tree quantization is enabled for a node; determine that an angular mode is activated for the node; in response to the intra-tree quantization being enabled for the node, determine a quantized value for the node representing a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value representing the coordinate position relative to the origin position; and determine a context for context encoding a planar position syntax element for the angular mode based on the scaled value representing the coordinate position relative to the origin position.
[0010] According to another example, a method for encoding a bitstream including point cloud data, the method comprising: determining that in-tree quantization is enabled for a node; determining that an angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context encoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0011] According to another example, a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to determine, based on syntax signaled in a bitstream, that in-tree quantization is enabled for a node; determine, based on syntax signaled in the bitstream, that angular mode is activated for the node; determine, in response to in-tree quantization being enabled for the node, a quantized value for the node representing a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value representing the coordinate position relative to the origin position; and determine, based on the scaled value representing the coordinate position relative to the origin position, a context for context decoding a planar position syntax element for the angular mode.
[0012] According to another example, a device for decoding a bitstream including point cloud data comprises: means for determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; means for determining, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; means for determining, in response to in-tree quantization being enabled for the node, a quantized value for the node that represents a coordinate position relative to an origin position; means for scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and means for determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0013] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example encoding and decoding system that may implement the techniques of this disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an example Geometry Point Cloud Compression (G-PCC) encoder. [Figure 3] FIG. 2 is a block diagram illustrating an exemplary G-PCC decoder. [Figure 4] FIG. 1 is a conceptual diagram illustrating an exemplary plane occupancy in the vertical direction. [Figure 5] FIG. 10 is a conceptual diagram of an example in which a context index is determined based on a laser beam position above or below a marker point of a node, in accordance with one or more techniques of the present disclosure. [Figure 6] FIG. 10 is a conceptual diagram illustrating an example three-context index determination. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example context index determination for coding vertical plane position of a planar mode based on laser beam position using sections separated by fine dotted lines. [Figure 8A] 10 is a flowchart illustrating an exemplary operation for encoding vertical plane position. [Figure 8B] 10 is a flowchart illustrating an exemplary operation for decoding vertical plane position. [Figure 9A] 10 is a flowchart illustrating example operations for coding vertical plane position, in accordance with one or more techniques of this disclosure. [Figure 9B] 10 is a flowchart illustrating example operations for coding vertical plane position, in accordance with one or more techniques of this disclosure. [Figure 10] FIG. 1 is a conceptual diagram illustrating an example distance measurement system that may be used with one or more techniques of the present disclosure. [Figure 11]FIG. 1 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of the present disclosure may be used. [Figure 12] FIG. 1 is a conceptual diagram illustrating an example augmented reality system in which one or more techniques of this disclosure may be used. [Figure 13] FIG. 1 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of the present disclosure may be used. DETAILED DESCRIPTION OF THE INVENTION
[0015] ISO / IEC MPEG (JTC 1 / SC29 / WG11), and more recently ISO / IEC MPEG 3DG (JTC 1 / SC29 / WG 7), are exploring standardization of point cloud coding techniques with compression capabilities potentially exceeding those of existing methods. The group formerly known as MPEG, which has now dissolved and split into individual working groups, is working together on this quest in a collaborative effort called the 3D Graphics Team (3DG) to evaluate compression technology designs proposed by experts in this area.
[0016] Point cloud compression activities are typically categorized into two different approaches. The first approach is "Video Point Cloud Compression" (V-PCC), which involves segmenting a 3D object and projecting the segments in multiple 2D planes (represented as "patches" in a 2D frame), which are further coded with a legacy 2D video codec such as HEVC. The second approach is "Geometry Point Cloud Compression" (G-PCC), which involves directly compressing the 3D geometry, i.e., the positions of a set of points in 3D space, and the associated attribute values (for each point related to the 3D geometry). G-PCC addresses the compression of point clouds in both Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds).
[0017] A point cloud includes a set of points in 3D space and may have attributes associated with the points. The attributes may be, for example, color information such as R / G / B, Y / Cb / Cr, reflectance information, or other such attributes. Point clouds may be captured by various cameras or sensors, such as light detection and ranging (LIDAR) scanners or 3D scanners, or may also be computer-generated. Point cloud data may be used in a variety of applications, including, but not limited to, architecture (e.g., modeling), graphics (e.g., 3D models for visualization and animation), and the automotive industry (e.g., LIDAR sensors used to aid navigation).
[0018] The 3D space occupied by the point cloud data may be covered by a virtual bounding box. The positions of points within the bounding box may be represented with some precision. Thus, the positions of one or more points may be quantized based on the precision. At the smallest level, the bounding box is divided into voxels, which are the smallest units of space represented by a unit cube. A voxel within the bounding box may be associated with zero, one, or more points. The bounding box may be divided into multiple cubes / cubic regions, sometimes called tiles, and each tile may be coded into one or more slices. The partitioning of the bounding box into slices and tiles may be based on the number of points in each partition or other considerations, such as coding a particular region as a tile. The slice regions may be further partitioned using partitioning decisions similar to those in video codecs.
[0019] G-PCC encoders and decoders may support planar coding mode and angular coding mode, sometimes referred to as planar mode and angular mode, respectively. Planar mode is a technique that may improve coding of node occupation. Planar mode may be used when all occupied child nodes of a node are adjacent to the plane and on sides of the plane associated with increasing coordinate values with respect to a dimension orthogonal to the plane. For example, planar mode may be used for a node when all occupied child nodes of the node are above or below a horizontal plane passing through the center point of the node, and planar mode may be used for a node when all occupied child nodes of the node are on the near side or far side of a vertical plane passing through the center point of the node. For a node, the G-PCC encoder may encode a syntax element for each of the x, y, and z dimensions to specify whether that dimension is coded using planar mode, and for each dimension coded using planar mode, a planar position syntax element (i.e., a syntax element indicating the planar position) may be signaled for the respective dimension. The plane position syntax element for a dimension indicates whether the plane orthogonal to that dimension is at the first or second position. When the plane is at the first position, it corresponds to the boundary of the node. When the plane is at the second position, it passes through the 3D center of the node. More generally, when the plane is at the first position, points within the node are on that side of the node at the first position but not on that side at the second position; when the plane is at the second position, points within the node are on that side of the node at the second position but not on that side at the second position. Thus, for the z dimension, a G-PCC coder may code the vertical plane position of the plane mode within the nodes of the octree that represent the 3D positions of points in the point cloud.
[0020] Point clouds can often be captured using LIDAR sensors or other laser-based sensors. Angle coding mode is optionally used in conjunction with planar mode to improve the coding of vertical (e.g., z) plane position syntax elements by employing knowledge of the position and elevation angle of the sensing laser beam in a typical LIDAR sensor. Additionally, angle coding mode can optionally be used to improve the coding of vertical z position bits in inferred direct coding mode (IDCM).
[0021] The G-PCC coder may also support intra-tree quantization. Intra-tree geometry scaling provides a means to quantize (encoder) and scale (decoder) geometry positions even while the coding tree is being constructed. Each point in a point cloud is located at a specific geometric position. In octree coding, information about the position may not be directly signaled, but rather, the octree occupancy across the hierarchy (from the root node to the leaf nodes) may be used to indicate the geometric position of the point. In the encoder, the points, or rather the occupied positions of the points, are placed in the leaf nodes of the octree. The size of the octree depends on the bit depth of the positions in each dimension. Starting from the root node, the occupancy of each of the eight octants is coded (differently). The occupancy at the root node effectively codes the most significant bits (MSBs) of the point position in three dimensions. This process continues down to the leaf nodes, which indicate the position of the point. The decoder follows a similar process of determining the occupancy of the octree nodes at each level up to the leaf nodes to determine the position of the point.
[0022] Geometry quantization is applied at a particular node depth in the octree signaled in the bitstream. Node depth generally refers to a particular level in the octree parsing. In the simple case where only the octree is considered (without QTBT), we consider the octree to have 12 levels. At the root node, each child node has 2 11 ×2 11 ×2 11 Each of these nodes may be considered to be at a node depth of 1. The children of each of the root node's child nodes have a size of 2 10 ×2 10 ×2 10 , which are considered to be nodes at node depth 2, and so on. In a simple example, if a node coordinate is 12 bits and the depth to which quantization is to be applied is 3, the first 3 MSBs of the node coordinate (called the MSB part of the position) are not quantized, and only the last 9 LSBs of the node coordinate (called the LSB part of the position) are quantized. Due to quantization, the 9 LSBs may be reduced to a smaller number of bits, such as N bits, where N is less than or equal to 9. This may result in some reduction in bit rate at the expense of reconstruction accuracy. The resulting node coordinate size becomes N+3 (which is ≦12). Similarly, at the decoder, the N LSBs are scaled and clipped to a maximum value of 1<<(9-1), which ensures that the scaled value does not exceed the 9 LSB bits of the original point. The final scaled position is calculated by splicing the 3 MSBs and the 9 scaled LSBs.
[0023] In conventional coding of point cloud frames, angular mode provides substantial gains in coding efficiency. However, when in-tree quantization is enabled, the gains of angular mode become significantly smaller and may even result in losses. For angular modes (IDCM angles and planar angles), quantization bits are used for context derivation, and the quantization bits are in a different scale space and not in the same domain as the original points. This reduces the usefulness of both angular mode and in-tree quantization, and therefore, it may not be beneficial to enable both simultaneously.
[0024] This disclosure describes techniques for deriving scaled values xS from point / position coordinate values x to derive the node / point position relative to the lidar origin when angle mode and intra-tree quantization are used together. More specifically, by scaling quantized values representing coordinate values without clipping, a G-PCC decoder can determine scaled values representing coordinate positions relative to the origin position in a way that puts the scaled values in the same scale space as the original point values with sufficient accuracy to achieve coding gain from angle mode.
[0025] This disclosure uses the term G-PCC coder to refer generically to a G-PCC encoder and / or a G-PCC decoder. Moreover, some techniques described in this disclosure with respect to decoding may also apply to encoding, and vice versa. For example, often, a G-PCC encoder and a G-PCC decoder are configured to perform the same or reciprocal processes. Also, a G-PCC encoder typically performs decoding as part of the process of determining how to encode.
[0026] 1 is a block diagram illustrating an example encoding and decoding system 100 that may perform the techniques of this disclosure. The techniques of this disclosure are generally directed to coding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. In general, point cloud data includes any data for processing a point cloud. Coding may be effective in compressing and / or decompressing point cloud data.
[0027] 1, system 100 includes a source device 102 and a destination device 116. Source device 102 provides encoded point cloud data to be decoded by destination device 116. In particular, in the example of FIG. 1, source device 102 provides the point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, terrestrial or marine vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication.
[0028] 1, source device 102 includes data source 104, memory 106, G-PCC encoder 200, and output interface 108. Destination device 116 includes input interface 122, G-PCC decoder 300, memory 120, and data consumer 118. According to this disclosure, G-PCC encoder 200 of source device 102 and G-PCC decoder 300 of destination device 116 may be configured to apply techniques of this disclosure related to angle modes and intra-tree quantization in G-PCC.
[0029] Accordingly, source device 102 represents an example of an encoding device, and destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 may include other components or configurations. For example, source device 102 may receive data (e.g., point cloud data) from an internal source or an external source. Similarly, destination device 116 may interface with an external data consumer rather than including the data consumer within the same device.
[0030] System 100 as shown in FIG. 1 is merely an example. In general, other digital encoding and / or decoding devices may perform the techniques of this disclosure related to angle modes and intra-tree quantization in G-PCC. Source device 102 and destination device 116 are merely examples of devices in which source device 102 generates coded data for transmission to destination device 116. This disclosure refers to devices that perform coding (encoding and / or decoding) of data as “coding” devices. Accordingly, G-PCC encoder 200 and G-PCC decoder 300 represent examples of coding devices, specifically, encoders and decoders, respectively. In some examples, source device 102 and destination device 116 may operate substantially symmetrically, such that each of source device 102 and destination device 116 includes encoding and decoding components. Thus, system 100 may support one-way or two-way transmission between source device 102 and destination device 116, for example, streaming, playback, broadcast, telephony, navigation, and other applications.
[0031] Generally, the data source 104 represents a source of data (i.e., raw, unencoded point cloud data) and may provide a sequential series of “frames” of data to the G-PCC encoder 200, which encodes the data for the frames. The data source 104 of the source device 102 may include a point cloud capture device, such as any of a variety of cameras or sensors, e.g., a 3D scanner or light detection and ranging (LIDAR) device, one or more video cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or additionally, the point cloud data may be computer-generated from scanners, cameras, sensors, or other data. For example, the data source 104 may generate computer-graphics-based data as source data, or may create a combination of live data, archived data, and computer-generated data. In each case, the G-PCC encoder 200 encodes the captured data, pre-captured data, or computer-generated data. The G-PCC encoder 200 may reorder the frames from the order in which they were received (sometimes called "display order") into a coding order for coding. The G-PCC encoder 200 may generate one or more bitstreams including the encoded data. The source device 102 may then output the encoded data onto a computer-readable medium 110 via an output interface 108, for receipt and / or retrieval by, for example, an input interface 122 of a destination device 116.
[0032] Memory 106 of source device 102 and memory 120 of destination device 116 may represent general-purpose memory. In some examples, memory 106 and memory 120 may store raw data, e.g., raw data from data source 104 and raw decoded data from G-PCC decoder 300. Additionally or alternatively, memory 106 and memory 120 may store software instructions executable by, e.g., G-PCC encoder 200 and G-PCC decoder 300, respectively. While memory 106 and memory 120 are shown separately from G-PCC encoder 200 and G-PCC decoder 300 in this example, it should be understood that G-PCC encoder 200 and G-PCC decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memory 106 and memory 120 may store encoded data, e.g., output from G-PCC encoder 200 and input to G-PCC decoder 300. In some examples, portions of memory 106 and memory 120 may be allocated as one or more buffers, e.g., for storing raw decoded and / or encoded data. For example, memory 106 and memory 120 may store data representing point clouds.
[0033] The computer-readable medium 110 may represent any type of medium or device capable of transporting encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to transmit encoded data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 may modulate a transmission signal containing the encoded data, and the input interface 122 may demodulate a received transmission signal, in accordance with a communication standard such as a wireless communication protocol. The communication medium may comprise any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device 102 to the destination device 116.
[0034] In some examples, source device 102 may output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded data.
[0035] In some examples, source device 102 may output the encoded data to file server 114 or another intermediate storage device, which may store the encoded data generated by source device 102. Destination device 116 may access the stored data from file server 114 via streaming or download. File server 114 may be any type of server device capable of storing encoded data and transmitting the encoded data to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network-attached storage (NAS) device. Destination device 116 may access the encoded data from file server 114 through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing encoded data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0036] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of the various IEEE 802.11 standards, or other physical components. In examples in which output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE-Advanced, 5G, etc. In some examples in which output interface 108 comprises a wireless transmitter, output interface 108 and input interface 122 may be configured to transfer data, such as encoded data, according to other wireless standards such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee™), the Bluetooth™ standard, etc. In some examples, source device 102 and / or destination device 116 may include respective system-on-chip (SoC) devices. For example, the source device 102 may include an SoC device for performing the functionality described as being in the G-PCC encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for performing the functionality described as being in the G-PCC decoder 300 and / or the input interface 122.
[0037] The techniques of this disclosure may be applied to encoding and decoding in support of any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors and processing devices such as local or remote servers, geographic mapping, or other applications.
[0038] The input interface 122 of the destination device 116 receives the encoded bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded bitstream may include signaling information defined by the G-PCC encoder 200, such as syntax elements having values that describe the characteristics and / or processing of a coded unit (e.g., a slice, a picture, a group of pictures, a sequence, etc.), which is also used by the G-PCC decoder 300. The data consumer 118 uses the decoded data. For example, the data consumer 118 may use the decoded data to determine the location of a physical object. In some examples, the data consumer 118 may include a display for presenting an image based on the point cloud.
[0039] The G-PCC encoder 200 and the G-PCC decoder 300 may each be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, a device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. The G-PCC encoder 200 and the G-PCC decoder 300 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) within the respective device. A device including the G-PCC encoder 200 and / or the G-PCC decoder 300 may comprise one or more integrated circuits, microprocessors, and / or other types of devices.
[0040] The G-PCC encoder 200 and the G-PCC decoder 300 may operate according to a coding standard, namely, the Geometry Point Cloud Compression (G-PCC) standard. Although the encoder 200 and the decoder 300 are described as a G-PCC encoder 200 and a G-PCC decoder 300, the encoder 200 and the decoder 300 should not be considered limited to operating according to the G-PCC standard. In some examples, the encoder 200 and the decoder 300 may operate according to the Video Point Cloud Compression (V-PCC) standard. This disclosure may generally refer to coding (e.g., encoding and decoding) of a picture to include the process of encoding or decoding data. An encoded bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes).
[0041] This disclosure may generally refer to “signaling” some information, such as syntax elements. The term “signaling” may generally refer to communicating values for syntax elements and / or other data used to decode encoded data. That is, G-PCC encoder 200 may signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As mentioned above, source device 102 may transport the bitstream to destination device 116 substantially in real time or not in real time, as may be done when storing syntax elements in storage device 112 for later retrieval by destination device 116.
[0042] As explained above, ISO / IEC MPEG (JTC1 / SC29 / WG11) is examining the potential need for and aims to develop a standard for point cloud coding techniques with compression capabilities significantly exceeding those of current methods. The group is working together on this exploration in a collaborative effort called the 3-Dimensional Graphics Team (3DG) to evaluate compression technology designs proposed by those experts in this field.
[0043] Point cloud compression activities are categorized into two different approaches. The first approach is "video point cloud compression" (V-PCC), which segments a 3D object and projects the segments in multiple 2D planes (represented as "patches" in a 2D frame), which are further coded by a legacy 2D video codec such as the High Efficiency Video Coding (HEVC) (ITU-T H.265) codec. The second approach is "geometry-based point cloud compression" (G-PCC), which directly compresses the 3D geometry, i.e., the location of a set of points in 3D space, and the associated attribute values (for each point related to the 3D geometry). G-PCC addresses the compression of point clouds in both Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds). A recent draft of the G-PCC standard is available at G-PCC DIS, ISO / IEC JTC1 / SC29 / WG11 w19088, Brussels, Belgium, January 2020, and the codec description is available at G-PCC Codec Description v6, ISO / IEC JTC1 / SC29 / WG11 w19091, Brussels, Belgium, January 2020.
[0044] A point cloud includes a set of points in 3D space and may have attributes associated with the points. The attributes may be color information such as R, G, B, or Y, Cb, Cr, or reflectance information, or other attributes. Point clouds may be captured by various cameras or sensors, such as LIDAR sensors and 3D scanners, or may also be computer-generated. Point cloud data is used in a variety of applications, including, but not limited to, architecture (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors used to aid navigation).
[0045] The 3D space occupied by the point cloud data may be covered by a virtual bounding box. The positions of points within the bounding box may be represented with some precision, and therefore, the positions of one or more points may be quantized based on the precision. At the smallest level, the bounding box is divided into voxels, which are the smallest units of space represented by a unit cube. A voxel within the bounding box may be associated with zero, one, or more points. The bounding box may be divided into multiple cubes / cubic regions, sometimes called tiles. Each tile may be coded into one or more slices. The partitioning of the bounding box into slices and tiles may be based on the number of points in each partition or other considerations (e.g., a particular region may be coded as a tile). The slice regions may be further partitioned using partitioning decisions similar to those in video codecs.
[0046] Figure 2 provides an overview of a G-PCC encoder 200. Figure 3 provides an overview of a G-PCC decoder 300. The illustrated modules are logical and do not necessarily correspond one-to-one to the implemented code in the reference implementation of the G-PCC codec, i.e., the TMC13 test model software considered by ISO / IEC MPEG (JTC1 / SC29 / WG11).
[0047] In both the G-PCC encoder 200 and the G-PCC decoder 300, the point cloud position is coded first. The attribute coding depends on the geometry being decoded. In Figures 2 and 3, the surface approximation analysis unit 212, the surface approximation synthesis unit 310, and the RAHT units 218 and 314 represent options commonly used for Category 1 data, while the LOD generation units 220 and 316, the lifting unit 222, and the inverse lifting unit 318 represent options commonly used for Category 3 data. All other units may be common between Category 1 and Category 3.
[0048] For Category 3 data, the compressed geometry is typically represented as an octree of individual voxels, from the root all the way down to the leaf level. For Category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree of blocks larger than a voxel, from the root down to the leaf level), plus a model that approximates the surface within each leaf of the pruned octree. In this way, both Category 1 and Category 3 data share the octree coding mechanism, and Category 1 data may additionally approximate the voxels within each leaf using a surface model. The surface model used is a triangulation with 1 to 10 triangles per block, resulting in a triangle soup. Therefore, Category 1 geometry codecs are called Trisoup geometry codecs, and Category 3 geometry codecs are called Octree geometry codecs.
[0049] At each node in the octree, occupancy is signaled (when not inferred) for one or more of its child nodes (up to 8 nodes). Multiple neighborhoods are specified, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge, or vertex with the current octree node, etc. Within each neighborhood, the occupancy of the node and / or its children can be used to predict the occupancy of the current node or its children. For sparse points within some nodes of the octree, the codec also supports a direct coding mode, in which the 3D position of the point is directly coded. A flag may be signaled to indicate that direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node may also be coded.
[0050] When geometry is coded, attributes corresponding to the geometry points are coded. When there are multiple attribute points corresponding to one geometry point to be reconstructed / decoded, an attribute value representing the reconstructed point can be derived.
[0051] G-PCC has three attribute coding methods: Region Adaptive Hierarchical Transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (prediction transform), and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform). RAHT and lifting are typically used for Category 1 data, and prediction is typically used for Category 3 data. However, either method may be used for any data; just like with geometry codecs in G-PCC, the attribute coding method used to code the point cloud is specified in the bitstream.
[0052] The coding of attributes may be done at a level-of-detail (LOD), where each level of detail may be used to obtain a finer representation of the point cloud attributes, and each level of detail may be specified based on a distance metric from neighboring nodes or based on a sampling distance.
[0053] In the G-PCC encoder 200, the residual obtained as the output of the coding method for the attribute is quantized. The quantized residual may be coded using context-adaptive arithmetic coding.
[0054] In the example of FIG. 2, the G-PCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, an LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0055] 2, the G-PCC encoder 200 may receive a set of locations and a set of attributes. The locations may include coordinates of points in the point cloud. The attributes may include information about the points in the point cloud, such as a color associated with the points in the point cloud.
[0056] The coordinate transformation unit 202 may apply a transform to the coordinates of the points to convert the coordinates from an initial domain to a transformation domain. This disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 may apply a transform to convert color information of the attributes to a different domain. For example, the color transformation unit 204 may convert color information from an RGB color space to a YCbCr color space.
[0057] Further, in the example of FIG. 2, the voxelization unit 206 may voxelize the transformed coordinates. Voxelizing the transformed coordinates may include quantization and removing some points of the point cloud. In other words, multiple points of the point cloud may be contained within a single "voxel," which may then be treated as one point in some respects. Further, the octree analysis unit 210 may generate an octree based on the voxelized transformed coordinates. Additionally, in the example of FIG. 2, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may entropy code syntax elements representing the octree and / or surface information determined by the surface approximation analysis unit 212. The G-PCC encoder 200 may output these syntax elements in a geometry bitstream.
[0058] The geometry reconstruction unit 216 may reconstruct transformation coordinates of points in the point cloud based on the octree, data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. The number of transformation coordinates reconstructed by the geometry reconstruction unit 216 may differ from the original number of points in the point cloud due to voxelization and surface approximation. This disclosure may refer to the obtained points as reconstructed points. The attribute transfer unit 208 may transfer attributes of the original points of the point cloud to the reconstructed points of the point cloud.
[0059] Furthermore, the RAHT unit 218 may apply RAHT coding to the attributes of the reconstruction points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting, respectively, to the attributes of the reconstruction points. The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 may output these syntax elements in an attribute bitstream.
[0060] In the example of FIG. 3, the G-PCC decoder 300 may include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, an inverse coordinate transformation unit 320, and an inverse color transformation unit 322.
[0061] The G-PCC decoder 300 may obtain a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 302 of the G-PCC decoder 300 may apply arithmetic decoding (e.g., context-adaptive binary arithmetic coding (CABAC) or other types of arithmetic decoding) to syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to syntax elements in the attribute bitstream.
[0062] The octree synthesis unit 306 may synthesize an octree based on syntax elements parsed from the geometry bitstream. In cases where surface approximation is used in the geometry bitstream, the surface approximation synthesis unit 310 may determine a surface model based on the syntax elements parsed from the geometry bitstream and based on the octree.
[0063] Further, the geometry reconstruction unit 312 may perform reconstruction to determine coordinates of points in the point cloud. The inverse coordinate transformation unit 320 may apply an inverse transform to the reconstructed coordinates to transform the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0064] 3, the inverse quantization unit 308 may inverse quantize the attribute values, which may be based on syntax elements obtained from the attribute bitstream (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0065] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT coding to determine color values for the points of the point cloud based on the dequantized attribute values. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 may use a level-of-detail-based technique to determine color values for the points of the point cloud.
[0066] 3, the inverse color transformation unit 322 may apply an inverse color transformation to the color values. The inverse color transformation may be the inverse of the color transformation applied by the color transformation unit 204 of the G-PCC encoder 200. For example, the color transformation unit 204 may transform the color information from the RGB color space to the YCbCr color space. Accordingly, the inverse color transformation unit 322 may transform the color information from the YCbCr color space to the RGB color space.
[0067] Various units in FIGS. 2 and 3 are illustrated to aid in understanding the operations performed by G-PCC encoder 200 and G-PCC decoder 300. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides specific functionality and is preconfigured for the operations that may be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that may be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. A fixed-function circuit may execute software instructions (e.g., to receive or output parameters), but the types of operations that the fixed-function circuit performs are generally invariant. In some examples, one or more of the units may be different circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be an integrated circuit.
[0068] The G-PCC encoder 200 and the G-PCC decoder 300 can be configured to code point cloud data using a planar coding mode, an angle coding mode, and an orientation coding mode. The planar coding mode was adopted at the 128th MPEG meeting in Geneva, Switzerland. A planar coding mode can be applied to each of the three dimensions x, y, and z at each node location. When a planar coding mode is specified for a particular dimension (e.g., z) at a node point, the planar coding mode indicates that all children of that node occupy only one of the z half-planes, as described below with respect to FIG. 4. Similar illustrations (e.g., techniques) apply to the planar coding mode in the x and y dimensions.
[0069] The angular coding mode was adopted at the 129th MPEG meeting in Brussels, Belgium. The following description is based on the original MPEG contribution, Sebastien Lasserre, Jonathan Taquet, "[GPCC][CE 13.22 related] An improvement of the planar coding mode," ISO / IEC JTC1 / SC29 / WG11 MPEG / m50642, Geneva, Switzerland, October 2019, and w19088. The angular coding mode is optionally used in conjunction with the planar mode (as described, for example, in Sebastien Lasserre, David Flynn, “[GPCC] Planar mode in octree-based geometry coding,” ISO / IEC JTC1 / SC29 / WG11 MPEG / m48906, Gothenburg, Sweden, July 2019) to improve the coding of vertical (z) plane position syntax elements by employing knowledge of the position and elevation angle of the sensing laser beam in a typical LIDAR sensor (see, for example, Sebastien Lasserre, Jonathan Taquet, “[GPCC] CE 13.22 report on angular mode,” ISO / IEC JTC1 / SC29 / WG11 MPEG / m51594, Brussels, Belgium, January 2020).
[0070] The orientation coding mode was adopted at the 130th MPEG Teleconferencing Meeting. The orientation coding mode is similar to the angle mode, extending the angle mode to the coding of the (x) and (y) plane position syntax elements of the planar mode and improving the coding of the x position bit or the y position bit in the IDCM. The orientation mode uses the sampling information of the orientation of each laser (e.g., the number of points that the laser can acquire in one rotation). In this disclosure, the term "angle mode" may also refer to the orientation mode, as described below.
[0071] Figure 4 is a conceptual diagram illustrating exemplary plane occupancy in the vertical direction. In the example of Figure 4, node 400 is partitioned into eight child nodes. Child nodes 402A-402H may be occupied or unoccupied. In the example of Figure 4, occupied child nodes are shaded. When one or more child nodes 402A-402D are occupied and none of child nodes 402E-402H are occupied, G-PCC encoder 200 may signal a plane_position syntax element with a value of 0 to indicate that all occupied child nodes are adjacent on the positive side of the plane of node 400's minimum z coordinate (i.e., the side with increasing z coordinate). When one or more child nodes 402E-402H are occupied and none of the child nodes 402A-402D are occupied, the G-PCC encoder 200 may signal a planar position syntax element with a value of 1 to indicate that all occupied child nodes are adjacent on the positive side of the plane of the midpoint z coordinate of node 400. In this way, the planar position syntax element may indicate the vertical plane position of the planar mode within node 400.
[0072] An angular coding mode may also optionally be used to improve the coding of vertical z-position bits in IDCM (Sebastien Lasserre, Jonathan Taquet, “[GPCC] CE 13.22 report on angular mode,” ISO / IEC JTC1 / SC29 / WG11 MPEG / m51594, Brussels, Belgium, January 2020). IDCM is a mode in which the positions of multiple points within a node are explicitly (e.g., directly) signaled relative to one point within the node. In angular coding mode, the positions of points may be signaled relative to the origin of the node, and the relationship of the points, captured using, for example, laser characteristics, is used to effectively compress the positions.
[0073] The angular coding mode may also be used when a point cloud is generated based on data generated by a distance measurement system, such as a LIDAR system. The LIDAR system may include a set of lasers arrayed in a vertical plane at different angles relative to an origin. The LIDAR system may rotate about a vertical axis. The LIDAR system may use the returned laser light to determine the distance and position of points in the point cloud. The laser beam emitted by the laser of the LIDAR system may be characterized by a set of parameters.
[0074] In the following description, laser, laser beam, laser sensor, or sensor, or other similar terms, may refer to any sensor capable of returning a distance measure and spatial orientation, potentially including an indication of time, for example, a typical LIDAR sensor.
[0075] The G-PCC encoder 200 or the G-PCC decoder 300 may code (e.g., encode or decode, respectively) the vertical plane position of a planar mode within a node by selecting a laser index from among a set of laser candidates signaled in a parameter set such as a geometry parameter set, where the selected laser index indicates the laser beam that intersects the node or is closest to the center of the node. The intersection or proximity of the laser beam with the node determines the context index (e.g., contextAngle, contextAnglePhiX, or contextAnglePhiY described below) used to arithmetically code the vertical plane position of the planar mode.
[0076] Thus, in some examples, a G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may code a vertical plane position of a planar mode within a node of an octree representing a three-dimensional position of a point in a point cloud. For ease of explanation, this disclosure may refer to the node that the G-PCC coder is coding as the current node. As part of coding the vertical plane position of a planar mode, the G-PCC coder may determine a laser index of a laser candidate within a set of laser candidates. The determined laser index indicates the laser beam that is closest to or intersects with the current node. The set of laser candidates may include each of the lasers in a LIDAR array. In some examples, the set of laser candidates may be indicated in a parameter set, such as a geometry parameter set. Additionally, as part of coding the vertical plane position, the G-PCC coder determines a context index based on the intersection or proximity of the laser beam with the current node. For example, the G-PCC coder may determine a context index based on whether the laser beam is above a first distance threshold, between the first and second distance thresholds, between the second and third distance thresholds, or below the third distance threshold. Further, as part of coding the vertical surface position, the G-PCC coder arithmetically codes the vertical surface position of the planar mode using the context indicated by the determined context index.
[0077] There may be an eligibility condition for determining whether a vertical plane position of the planar mode within the current node is eligible to be coded using the angle mode. If the vertical plane position is not eligible to be coded using the angle mode, the vertical plane position of the planar mode may be coded without employing sensor information. In some examples, the eligibility condition may determine whether only one laser beam intersects with the current node. In other words, if only one laser beam (i.e., no more than two laser beams) intersects with the current node, the vertical plane position of the current node may be eligible to be coded using the angle mode. In some examples, the eligibility condition may determine a minimum angle difference between lasers in a set of laser candidates. In other words, if the angle surrounding the current node is smaller than the minimum angle between the laser beams, the current node may be eligible to be coded using the angle mode. The angle surrounding the current node is the angle measured from the laser origin between a line passing through the far lower corner of the node and a line passing through the near upper corner of the node. When the angle surrounding the current node is less than the minimum angular difference between the laser beams, at most one laser beam intersects the node. In some examples, the eligibility condition is such that the vertical node dimension is less than (or equal to) the minimum angular difference. In other words, if the vertical dimension of the node is less than or equal to the vertical distance between the laser beams separated by the minimum angular difference at the vertical edge of the current node closest to the laser origin, the current node may be eligible to be coded using angular mode.
[0078] As described above, the G-PCC coder may select the laser index of the laser beam that intersects or is closest to the current node. In some examples, the G-PCC coder may determine the index of the laser that intersects or is closest to the current node by selecting the laser beam that is closest to a marker point in the current node. In some examples, the marker point in the current node may be a center point of the current node having coordinates in half of all three dimensions (e.g., cubic / cuboid dimensions) of the current node. In other examples, the marker point in the current node may be any other point that is part of the current node, such as any point within the node, or on a node side, or on a node edge, or at a node corner.
[0079] The G-PCC coder may determine whether a laser is near a marker point based on comparing the angle difference of candidate lasers. For example, the G-PCC coder may compare the difference between the angle of the laser beam and the angle of the marker point. The angle of the laser beam may be defined as being between the horizontal plane (z=0) and the direction of the laser beam. The angle of the marker point may be defined as being between the horizontal plane and the direction of a virtual beam to the marker point. The origin in this case may be co-located with the center of the sensor or laser. Alternatively, in some examples, a mathematical function or trigonometric function such as tangent may be applied to the angle before comparison.
[0080] In some examples, the G-PCC coder may determine whether the laser is near a marker point based on a comparison of vertical coordinate differences. For example, the G-PCC coder may compare the vertical coordinate of the marker point relative to the sensor origin (e.g., the z-coordinate of the marker point) with the vertical coordinate of the laser intersection with the node or its proximity to the node. The G-PCC coder may obtain the vertical coordinate of the laser intersection with the node by multiplying (trigonometrically) the tangent of the angle between the horizontal plane and the laser direction by a distance calculated by taking the Euclidean distance based on the (x, y) coordinates of the marker point.
[0081] In some examples, the G-PCC coder may determine a context index to use for coding a vertical plane position syntax element based on the relative positions of the laser beam and the marker point. For example, the G-PCC coder may determine that the context index is a first context index when the laser beam is above the marker point, and that the context index is a second context index when the laser beam is below the marker point. The G-PCC coder may determine whether the laser beam is above or below the marker point in a manner similar to determining the laser beam index that intersects or is closest to a node, for example, by comparing angle differences, tangent differences of angle values, or vertical coordinate differences.
[0082] In some examples, as part of determining the context index, the G-PCC coder may determine a distance threshold. The G-PCC coder may use the distance threshold to compare the distance between the laser beam and the marker point. The distance threshold may divide the distance range within the current node into intervals. The intervals may be equal or unequal in length. If the laser beam is within a distance range determined by the distance intervals, each interval of the multiple intervals may correspond to a context index. In some examples, there are two distance thresholds determined by equal distance offsets above and below the marker point, defining three distance intervals corresponding to three context indexes. The G-PCC coder may determine whether the laser beam belongs to an interval in a manner similar to determining the laser beam index of the laser that intersects or is closest to the node (e.g., by comparing angle differences, comparing tangent differences of angle values, or comparing vertical coordinate differences).
[0083] The above principles of employing sensor information are not limited to coding the vertical (Z) plane position syntax element of the planar mode within a node, but similar principles may also be applied to coding the X or Y plane position syntax element of the planar mode within a node. The X or Y plane position mode of the planar mode may be chosen by the encoder if it is more appropriate for coding the point distribution within the node. For example, if all occupied child nodes are on one side of a plane oriented in the X direction, an X plane position syntax element may be used to code the point distribution within the node. If all occupied child nodes are on one side of a plane oriented in the Y direction, a Y plane position syntax element may be used to code the point distribution within the node. Additionally, a combination of two or more planes oriented in the X, Y, or Z directions may be used to indicate the occupancy of a child node.
[0084] As an example, the G-PCC coder may determine the context from among the two contexts based on whether the laser beam is above or below the marker point (i.e., whether it is located above or below the marker point). In this example, the marker point is the center of the node. This is illustrated in FIG. 5. More specifically, FIG. 5 is a conceptual diagram of an example in which a context index is determined based on a laser beam position 500A, 500B above or below a marker point 502 of a node 504. Thus, in the example of FIG. 5, if the laser intersecting the node 504 is above the marker point 502, as shown with respect to laser beam position 500A, the G-PCC coder selects a first context index (e.g., Ctx=0). In the example of FIG. 5, the marker point 502 is located at the center of the node 504. If the laser intersecting the node 504 is below the marker point 502, as shown with respect to laser beam position 500B, the G-PCC coder selects a second context index (e.g., Ctx=1).
[0085] In some examples, the G-PCC coder determines three contexts based on whether the laser beam is located above or below two distance thresholds or halfway between the distance thresholds. In this example, the marker point is the center of the node. This is shown in FIG. 6. More specifically, FIG. 6 is a conceptual diagram illustrating an example three-context index determination for a node 600. In FIG. 6, the distance interval thresholds are indicated using thin dotted lines 602A, 602B. The laser beams are indicated using solid lines 604A, 604B. Each of these laser beams may be a laser candidate. The center point 606 (marker point) is indicated using a white circle. Thus, in the example of FIG. 6, if the laser (such as the laser corresponding to line 604A) is above line 602A, the G-PCC coder selects context index ctx1. If the laser is between lines 602A, 602B, the G-PCC coder may select context index ctx0. If a laser (such as the laser corresponding to line 604B) is below line 602B, the G-PCC coder may select context index ctx2.
[0086] In some examples, the G-PCC coder uses four contexts to code the vertical surface position in planar mode when angular mode is used. In such examples, the G-PCC coder may determine a context index based on the position of the laser beam within four intervals. An example of this is shown in FIG. 7. FIG. 7 is a conceptual diagram illustrating an example context index determination for coding the vertical surface position in planar mode (angular mode) based on the laser beam position (solid arrow) with the intervals separated by thin dotted lines. In the example of FIG. 7, lines 700A, 700B, and 700C correspond to distance interval thresholds. Further, in the example of FIG. 7, marker point 702 is located at the center of node 704. Line 706 corresponds to the laser beam intersecting node 704. Because line 706 is above line 700A, the G-PCC coder may select context index ctx2 for use in coding the vertical surface position.
[0087] 8A is a flowchart illustrating exemplary operations for encoding vertical plane positions. G-PCC encoder 200 may perform the operations of FIG. 8A as part of encoding a point cloud.
[0088] In the example of Figure 8A, the G-PCC encoder 200 (e.g., the arithmetic coding unit 214 of the G-PCC encoder 200 (Figure 2)) may encode 800 vertical plane positions of the planar mode in nodes of a tree (e.g., an octree) that represents the three-dimensional positions of points in a point cloud represented by the point cloud data. In other words, the G-PCC encoder 200 may encode the vertical plane positions.
[0089] As part of encoding the vertical plane position of the planar mode, the G-PCC encoder 200 (e.g., the arithmetic coding unit 214) may determine a laser index of a laser candidate in the set of laser candidates, the determined laser index indicating a laser beam that intersects with the node (802). The G-PCC encoder 200 may determine the laser index according to any of the examples provided elsewhere in this disclosure.
[0090] Additionally, the G-PCC encoder 200 (e.g., the arithmetic coding unit 214) may determine a context index based on whether the laser beam is above a first distance threshold, between the first and second distance thresholds, between the second and third distance thresholds, or below the third distance threshold (804). For example, in the example of FIG. 7, the G-PCC encoder 200 may determine a context index based on whether the laser beam is above a first distance threshold (corresponding to line 700A), between the first and second distance thresholds (corresponding to line 700B), between the second and third distance thresholds (corresponding to line 700C), or below the third distance threshold. In some examples, to determine the position of the laser beam relative to the first, second, and third distance thresholds, the G-PCC encoder 200 may determine a laser differential angle (e.g., thetaLaserDelta below) by subtracting the tangent of the angle of a line passing through the center of the node from the tangent of the angle of the laser beam, may determine a top angle differential (e.g., DeltaTop below) by subtracting a shift value from the laser differential angle, and may determine a bottom angle differential (e.g., DeltaBot below) by adding the shift value to the laser differential angle.
[0091] The G-PCC encoder 200 may perform a first comparison to determine whether the laser differential angle is greater than or equal to 0 (e.g., thetaLaserDelta≧0). The G-PCC encoder 200 may set a context index to 0 or 1 based on whether the laser differential angle is greater than or equal to 0 (e.g., contextAngular[Child] = thetaLaserDelta≧0 ? 0 : 1). Additionally, the G-PCC encoder 200 may perform a second comparison to determine whether the top angular difference is greater than or equal to 0 (e.g., DeltaTop≧0). When the top angular difference is greater than or equal to 0, the laser beam is above the first distance threshold. The G-PCC encoder 200 may also perform a third comparison to determine whether the bottom angular difference is less than 0 (e.g., DeltaBottom<0). When the bottom angular difference is less than 0, the laser beam is below the third distance threshold. The G-PCC encoder 200 may increment the context index by 2 based on the top angle difference being greater than or equal to 0 (e.g., if (DeltaTop ≧ 0) contextAngular[Child] += 2), or based on the bottom angle difference being less than 0 (e.g., else if (DeltaBottom < 0) contextAngular[Child] += 2).
[0092] The G-PCC encoder 200 (e.g., the arithmetic coding unit 214 of the G-PCC encoder 200) may arithmetically encode the vertical plane position in the planar mode using the context indicated by the determined context index (806). For example, the G-PCC encoder 200 may perform CABAC encoding on the syntax element indicating the vertical plane position.
[0093] 8B is a flowchart illustrating an example operation for decoding a vertical surface position. The G-PCC decoder 300 may perform the operations of FIG. 8B as part of reconstructing a point cloud represented by point cloud data. In the example of FIG. 8B, the G-PCC decoder 300 (e.g., the geometry arithmetic decoding unit 302 of FIG. 3) may decode 850 a vertical surface position of a planar mode in a node of a tree (e.g., an octree) representing the three-dimensional position of a point in the point cloud. In other words, the G-PCC decoder 300 may decode the vertical surface position.
[0094] As part of decoding the vertical plane position of the planar mode, the G-PCC decoder 300 (e.g., the geometry arithmetic decoding unit 302) may determine a laser index of a laser candidate in the set of laser candidates, the determined laser index indicating the laser beam that intersects with or is closest to the node (852). The G-PCC decoder 300 may determine the laser index according to any of the examples provided elsewhere in this disclosure.
[0095] Additionally, the G-PCC decoder 300 may determine a context index (contextAngular) based on whether the laser beam is above the first distance threshold, between the first and second distance thresholds, between the second and third distance thresholds, or below the third distance threshold (854). The G-PCC decoder 300 may determine the context index in the same manner as the G-PCC encoder 200, as described above.
[0096] The G-PCC decoder 300 (e.g., the geometry arithmetic decoding unit 302 of the G-PCC decoder 300) may arithmetically decode the vertical plane position for the planar mode using the context indicated by the determined context index (856). For example, the G-PCC decoder 300 may perform CABAC decoding on a syntax element indicating the vertical plane position. In some examples, the G-PCC decoder 300 may determine the position of one or more points in the point cloud based on the vertical plane position. For example, the G-PCC decoder 300 may determine the location of an occupied child node of a node based on the vertical plane position. The G-PCC decoder 300 may then process the occupied child node to determine the position of the point within the occupied child node and may not need to perform further processing on the unoccupied child node.
[0097] A G-PCC coder (e.g., G-PCC encoder 200 or G-PCC decoder 300) may code (i.e., encode or decode) the IDCM vertical point position offset within the node, in part by selecting a laser index from among a set of laser candidates. The set of laser candidates may be signaled in a parameter set, such as a geometry parameter set, and the selected laser index indicates a laser beam that intersects with the node. The set of laser candidates may correspond to lasers in a LIDAR array. The G-PCC coder may determine a context index for arithmetically coding bins (bits) from the IDCM vertical point position offset based on the intersection of the laser beam with the node.
[0098] As introduced above, the G-PCC encoder 200 and the G-PCC decoder 300 can be configured to perform context derivation for the angular mode. Planar mode context derivation for three coordinates x, y, and z is described below. Because the LIDAR laser rotates in the xy plane, the processing of the x and y coordinates can be similar to each other and different from the processing of the z coordinate. In other examples, two different planes can be similar to a third different plane. To determine the node position and node midpoint, the G-PCC encoder 200 and the G-PCC decoder 300 may derive the variables absPos and midNode as described below, where child.pos and childSizeLog2 refer to the position and size of the current node, respectively. { Vec3<int64_t> absPos = {child.pos[0] << childSizeLog2[0], child.pos[1] << childSizeLog2[1], child.pos[2] << childSizeLog2[2]}; / / Eligibility Vec3<int64_t> midNode = {1 << (childSizeLog2[0] ? childSizeLog2[0] - 1 : 0), 1 << (childSizeLog2[1] ? childSizeLog2[1] - 1 : 0), 1 << (childSizeLog2[2] ? childSizeLog2[2] - 1 : 0)};
[0099] When the node size is too large (when more than one laser can pass through the node), the G-PCC encoder 200, and therefore the G-PCC decoder 300, disables angle mode. The G-PCC encoder 200 calculates an estimate of the shortest distance between two adjacent lasers at the node's center distance (deltaAngle) and compares the estimate with the node size. When the current node is larger than a threshold, the G-PCC encoder 200 disables angle mode. The variable headPos refers to the position of the LIDAR head and can be used to find the coordinates xLidar / yLidar / zLidar relative to the LIDAR origin as follows: uint64_t xLidar = std::abs(((absPos[0] - headPos[0] + midNode[0]) << 8) - 128); uint64_t yLidar = std::abs(((absPos[1] - headPos[1] + midNode[1]) << 8) - 128); uint64_t rL1 = (xLidar + yLidar) >> 1; uint64_t deltaAngleR = deltaAngle * rL1; if (deltaAngleR <= (midNode[2] << 26)) return -1;
[0100] In the next step, the G-PCC encoder 200 calculates the angular altitude theta32 of the node center in the z direction (zLidar) and estimates the index of the laser closest to the node center, laserIndex, as follows: / / Determine the inverse of r (1 / sqrt(r2) = irsqrt(r2)) uint64_t r2 = xLidar * xLidar + yLidar * yLidar; uint64_t rInv = irsqrt(r2); / / Determine the uncorrected theta int64_t zLidar = ((absPos[2] - headPos[2] + midNode[2]) << 1) - 1; int64_t theta = zLidar * rInv; int theta32 = theta >= 0 ? theta >> 15 : -((-theta) >> 15); / / Determine the laser int laserIndex = int(child.laserIndex); if (laserIndex == 255 || deltaAngleR <= (midNode[2] << (26 + 2))) { auto end = thetaLaser + numLasers - 1; auto it = std::upper_bound(thetaLaser + 1, end, theta32); if (theta32 - *std::prev(it) <= *it - theta32) --it; laserIndex = std::distance(thetaLaser, it); child.laserIndex = uint8_t(laserIndex); }
[0101] Once the nearest laser is estimated, the G-PCC encoder 200 estimates the context for the planar modes of the x and y coordinates. In some examples, the G-PCC encoder 200 determines the context for only one of x and y, and when one coordinate is coded using the angle planar mode, the other coordinate is not planar coded. For each laser, the last orientation coded / predicted is stored in the phiBuffer. Based on the laser's speed (input parameter), the G-PCC encoder 200 updates the orientation prediction and determines the context by comparing the updated orientation prediction with the orientation of the node origin. Then, it checks which part of the current node the nearest laser passes through, which can be used to estimate one of six contexts for the x / y dimensions as follows: / / -- PHI -- / / angle int posx = absPos[0] - headPos[0]; int posy = absPos[1] - headPos[1]; int phiNode = iatan2(posy + midNode[1], posx + midNode[0]); int phiNode0 = iatan2(posy, posx); / / Discover the predictors int predPhi = phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Use the predictor if (predPhi != 0x80000000) { / / Basis shift predictor int Nshift = ((predPhi - phiNode) * phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= phiZi.delta(laserIndex) * Nshift; / / ctx orientation x or y int angleL = phiNode0 - predPhi; int angleR = phiNode - predPhi; int contextAnglePhi = (angleL >= 0 && angleR >= 0) || (angleL < 0 && angleR < 0) ? 2 : 0; angleL = std::abs(angleL); angleR = std::abs(angleR); if (angleL > angleR) { contextAnglePhi++; int temp = angleL; angleL = angleR; angleR = temp; } if (angleR > (angleL << 2)) contextAnglePhi += 4; if (std::abs(posx) <= std::abs(posy)) *contextAnglePhiX = contextAnglePhi; else *contextAnglePhiY = contextAnglePhi; }
[0102] For the z-dimensional context, three thresholds are chosen: theta32, theta32-zShift, and theta32+zShift. By comparing the elevation angle of the closest laser, thetaLaser[laserIndex], with the three thresholds, the G-PCC encoder 200 derives one of four contexts for the planar mode of the z coordinate as follows: / / -- THETA -- int thetaLaserDelta = thetaLaser[laserIndex] - theta32; int64_t hr = zLaser[laserIndex] * rInv; thetaLaserDelta += hr >= 0 ? -(hr >> 17) : ((-hr) >> 17); int64_t zShift = (rInv << childSizeLog2[2] )>> 20; int thetaLaserDeltaBot = thetaLaserDelta + zShift; int thetaLaserDeltaTop = thetaLaserDelta - zShift; int contextAngle = thetaLaserDelta >= 0 ? 0 : 1; if (thetaLaserDeltaTop >= 0) contextAngle += 2; else if (thetaLaserDeltaBot < 0) contextAngle += 2; return contextAngle; }
[0103] As introduced above, the G-PCC encoder 200 and the G-PCC decoder 300 can also be configured to determine an IDCM mode context. IDCM refers to a mode in G-PCC in which the position of a point within a node is coded without using the octree structure within that node. IDCM can be beneficial for point clouds with sparse nodes or outliers, for example, where coding the coordinates of a point in such a case can be cheaper than coding the point with a complete octree structure. In G-PCC, only one or two different positions can be coded as IDCM nodes within a node.
[0104] When the angular mode is used, the G-PCC encoder 200 may context code bits corresponding to the position based on laser parameters. An exemplary method for deriving context for bits coded in the angular IDCM mode is described below. The following description describes techniques that may be performed by the G-PCC encoder 200 and / or the G-PCC decoder 300.
[0105] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine the nodePosition and the laser index. As mentioned above, in the IDCM, up to two different positions within a node may be encoded. The first step may be to determine the number of points occupying up to two positions, which may be done as follows: int numPoints = 1; bool numPointsGt1 = _arithmeticDecoder->decode(_ctxNumIdcmPointsGt1); numPoints += numPointsGt1; int numDuplicatePoints = 0; if (!geom_unique_points_flag && !numPointsGt1) { numDuplicatePoints = !_arithmeticDecoder->decode(_ctxSinglePointPerBlock); if (numDuplicatePoints) { bool singleDup = _arithmeticDecoder->decode(_ctxSingleIdcmDupPoint); if (!singleDup) numDuplicatePoints += 1 + _arithmeticDecoder->decodeExpGolomb(0, _ctxPointCountPerBlock);} }
[0106] For the next step, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to check whether the planar mode is applied to the current node. If the planar mode is applied, the MSB of the corresponding coordinate may be signaled separately (planar), and effectively only the remaining bits are coded. Coding of the MSB using the planar mode may be performed before IDCM coding. The node size is updated from effectiveNodeSizeLog2 to nodeSizeLog2Rem as follows: / / Update the node size after the plane and determine the upper part of the position from the plane Vec3<int32_t> deltaPlanar{0, 0, 0}; Vec3 <int>nodeSizeLog2Rem = nodeSizeLog2; for (int k = 0; k < 3; k++) if (nodeSizeLog2Rem[k] > 0 && (planar.planarMode & (1 << k))) { deltaPlanar[k] |= (planar.planePosBits & (1 << k) ? 1 : 0); nodeSizeLog2Rem[k]--; }
[0107] For the next step, the G-PCC encoder 200 and the G-PCC decoder 300 calculate the node position posNodeLidar relative to the LIDAR origin headpost as follows: The x and y positions are used to determine whether the x or y coordinates of positions within this node should be context coded, while other positions are bypass coded. Vec3 <bool>directIdcm = !angularIdcm; point_t posNodeLidar; if (angularIdcm) { posNodeLidar = point_t( node.pos[0] << nodeSizeLog2[0], node.pos[1] << nodeSizeLog2[1], node.pos[2] << nodeSizeLog2[2]) - headPos; bool codeXorY = std::abs(posNodeLidar[0]) <= std::abs(posNodeLidar[1]); directIdcm.x() = !codeXorY; directIdcm.y() = codeXorY; }
[0108] When more than one point is coded, the G-PCC encoder 200 and the G-PCC decoder 300 order the points, which further reduces the number of bits required to code the point positions. This is achieved by coding one or more MSBs of the x- or y- or z-position, which may result in a further reduction in the number of bits remaining to be coded (nodeSizeLog2Rem). / / Decode two unordered points Vec3<int32_t> deltaPos[2]; deltaPos[0] = deltaPlanar; deltaPos[1] = deltaPlanar; if (numPoints == 2 && joint_2pt_idcm_enabled_flag) decodeOrdered2ptPrefix(directIdcm, nodeSizeLog2Rem, deltaPos);
[0109] Since some bits have already been coded (due to the planar coding and joint coding of the points), the G-PCC encoder 200 and the G-PCC decoder 300 may calculate the midpoint of the updated node position and may use this midpoint to determine the nearest laser passing through the node midpoint as follows: if (angularIdcm) { for (int idx = 0; idx < 3; ++idx) { int N = nodeSizeLog2[idx] - nodeSizeLog2Rem[idx]; for (int mask = N ? 1 << (N - 1) : 0; mask; mask >>= 1) { if (deltaPos[0][idx] & mask) posNodeLidar[idx] += mask << nodeSizeLog2Rem[idx]; } if (nodeSizeLog2Rem[idx]) posNodeLidar[idx] += 1 << (nodeSizeLog2Rem[idx] - 1); } node.laserIndex = findLaser(posNodeLidar, thetaLaser, numLasers); }
[0110] When angle mode is enabled, the G-PCC encoder 200 and the G-PCC decoder 300 may context code bits in positions that have not already been coded. For example, one bit of x and y may be context coded, and the other may be bypass coded, while the bits of z are context coded as described below. Otherwise, all remaining bits may be bypass coded. Vec3<int32_t> pos; for (int i = 0; i < numPoints; i++) { if (angularIdcm) { *(outputPoints++) = pos = decodePointPositionAngular( nodeSizeLog2, nodeSizeLog2Rem, node, planar, headPos, zLaser, thetaLaser, deltaPos[i]); } else *(outputPoints++) = pos = decodePointPosition(nodeSizeLog2Rem, deltaPos[i]); }
[0111] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine a context for the angle IDCM. Determining a context for the IDCM node position based on laser parameters will now be described.
[0112] The G-PCC encoder 200 and the G-PCC decoder 300 may calculate the node position posXyz relative to the LIDAR head position head post, which may again be used to determine which coordinate (x or y) to context code. If posXyz[1]≧posXyz[0] (i.e., the absolute y coordinate value is larger), the G-PCC encoder 200 and the G-PCC decoder 300 may bypass code the y coordinate bits and context code the x coordinate bits. Otherwise (i.e., the absolute y coordinate value is smaller than the absolute x coordinate), the G-PCC encoder 200 and the G-PCC decoder 300 may bypass code the x coordinate bits and context code the y coordinate bits. As each bit is coded, an updated node posXyz is calculated. { Vec3<int32_t> delta = deltaPlanar; Vec3 <int>posXyz = {(child.pos[0] << nodeSizeLog2[0]) - headPos[0], (child.pos[1] << nodeSizeLog2[1]) - headPos[1], (child.pos[2] << nodeSizeLog2[2]) - headPos[2]}; / / -- PHI -- / / Code x or y directly and calculate phi of the node bool codeXorY = std::abs(posXyz[0]) <= std::abs(posXyz[1]); if (codeXorY) { / / Direct coding of y if (nodeSizeLog2AfterPlanar[1]) for (int i = nodeSizeLog2AfterPlanar[1]; i > 0; i--) { delta[1] <<= 1; delta[1] |= _arithmeticDecoder->decode(); } posXyz[1] += delta[1]; posXyz[0] += delta[0] << nodeSizeLog2AfterPlanar[0]; } else { / / Direct code x if (nodeSizeLog2AfterPlanar[0]) for (int i = nodeSizeLog2AfterPlanar[0]; i > 0; i--) { delta[0] <<= 1; delta[0] |= _arithmeticDecoder->decode(); } posXyz[0] += delta[0]; posXyz[1] += delta[1] << nodeSizeLog2AfterPlanar[1]; }
[0113] In the next step, the G-PCC encoder 200 and the G-PCC decoder 300 calculate the laser index for the point using the laser index derived in the manner described above for the planar mode context and the laser index residual included in the bitstream. / / Discover the predictors int phiNode = iatan2(posXyz[1], posXyz[0]); int laserNode = int(child.laserIndex); / / Laser residual int laserIndex = laserNode + decodeThetaRes();
[0114] The G-PCC encoder 200 and the G-PCC decoder 300 use this updated laserIndex to determine a prediction for the orientation (phi), which is stored in the buffer _phiBuffer. Because the laser is approximated to uniformly sample points within the orientation range, the G-PCC encoder 200 and the G-PCC decoder 300 may use phiZi.delta to indicate the minimum difference in orientation value. Thus, the number of points that may have been skipped from the predicted phi value is calculated as nShift, and the predicted phi value is updated. int predPhi = _phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Basis shift predictor int nShift = ((predPhi - phiNode) * _phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= _phiZi.delta(laserIndex) * nShift;
[0115] For coordinates between x and y that are context coded, the G-PCC decoder 300 decodes the remaining nodeSizeLog2AfterPlanar[0] bits (which are context coded). For each bit, the G-PCC decoder 300 determines an updated context value by updating the node position by updating posXY as each bit is decoded. The G-PCC decoder 300 also recalculates the predicted phi value when the node position is updated. One of the six (or eight) contexts is chosen by comparing the orientation of the node position, the orientation of a position offset by half the node size in the context-coded direction (i.e., x or y), and the predicted orientation. Once the position is coded, the node's orientation phiNode is stored in a buffer. / / Choose x or y int* posXY = codeXorY ? &posXyz[0] : &posXyz[1]; int idx = codeXorY ? 0 : 1; / / Orientation coding for x or y int mask2 = codeXorY ? (nodeSizeLog2AfterPlanar[0] > 0 ? 1 << (nodeSizeLog2AfterPlanar[0] - 1) : 0) : (nodeSizeLog2AfterPlanar[1] > 0 ? 1 << (nodeSizeLog2AfterPlanar[1] - 1) : 0); for (; mask2; mask2 >>= 1) { / / Left and right angles int phiR = codeXorY? iatan2(posXyz[1], posXyz[0] + mask2) : iatan2(posXyz[1] + mask2, posXyz[0]); int phiL = phiNode; / / ctx orientation int angleL = phiL - predPhi; int angleR = phiR - predPhi; int contextAnglePhi = (angleL >= 0 && angleR >= 0) || (angleL < 0 && angleR < 0)? 2 : 0; angleL = std::abs(angleL); angleR = std::abs(angleR); if (angleL > angleR) { contextAnglePhi++; int temp = angleL; angleL = angleR; angleR = temp; } if (angleR > (angleL << 1)) contextAnglePhi += 4; / / Entropy coding bool bit = _arithmeticDecoder->decode( _ctxPlanarPlaneLastIndexAngularPhiIDCM[contextAnglePhi]); delta[idx] <<= 1; if (bit) { delta[idx] |= 1; *posXY += mask2; phiNode = phiR; predPhi = _phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Basis shift predictor int nShift = ((predPhi - phiNode) * _phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= _phiZi.delta(laserIndex) * nShift; } } / / Update the buffer phi _phiBuffer[laserIndex] = phiNode;
[0116] Once the x and y positions are coded, the z coordinate is coded similarly to the derivation of the context for the planar mode bit in the z direction (described above for the planar mode context), using the generalization to multiple bits as described above. As each z coordinate bit is coded, the G-PCC decoder 300 recalculates the context for the next bit using the updated node position and a z value that is offset from the node position by half the node size in the z direction (posXyz[2] + maskz). Again, the G-PCC decoder 300 chooses one of the four contexts to code the z coordinate bit. / / -- THETA -- int maskz = nodeSizeLog2AfterPlanar[2] > 0 ? 1 << (nodeSizeLog2AfterPlanar[2] - 1) : 0; if (!maskz) return delta; int posz0 = posXyz[2]; posXyz[2] += delta[2] << nodeSizeLog2AfterPlanar[2]; / / Since x and y are known, / / r is also known and bit-independent for z uint64_t xLidar = (int64_t(posXyz[0] ) << 8) - 128; uint64_t yLidar = (int64_t(posXyz[1] ) << 8) - 128; uint64_t r2 = xLidar * xLidar + yLidar * yLidar; int64_t rInv = irsqrt(r2); / / Use angle to code bit for z. Eligibility is implicit. Laser is known. int64_t hr = zLaser[laserIndex] * rInv; int fixedThetaLaser = thetaLaser[laserIndex] + int(hr >= 0 ? -(hr >> 17) : ((-hr) >> 17)); int zShift = (rInv << nodeSizeLog2AfterPlanar[2]) >> 18; for (int bitIdxZ = nodeSizeLog2AfterPlanar[2]; bitIdxZ > 0; bitIdxZ--, maskz >>= 1, zShift >>= 1) { / / Determine the uncorrected theta int64_t zLidar = ((posXyz[2] + maskz) << 1) - 1; int64_t theta = zLidar * rInv; int theta32 = theta >= 0 ? theta >> 15 : -((-theta) >> 15); int thetaLaserDelta = fixedThetaLaser - theta32; int thetaLaserDeltaBot = thetaLaserDelta + zShift; int thetaLaserDeltaTop = thetaLaserDelta - zShift; int contextAngle = thetaLaserDelta >= 0 ? 0 : 1; if (thetaLaserDeltaTop >= 0) contextAngle += 2; else if (thetaLaserDeltaBot < 0) contextAngle += 2; delta[2] <<= 1; delta[2] |= _arithmeticDecoder->decode( _ctxPlanarPlaneLastIndexAngularIdcm[contextAngle]); posXyz[2] = posz0 + (delta[2] << (bitIdxZ - 1)); } return delta; }
[0117] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to use in-tree quantization. In-tree geometry scaling provides a means to quantize (e.g., in the G-PCC encoder 200) and scale (e.g., in the G-PCC decoder 300) geometry positions even while the coding tree is being constructed. In the current draft text, the quantization step size and scaling of positions are applied as follows: The shift value (sh) is calculated as follows: qpScaled = qp << qpDivFactorLog2 sh = qpScaled >> 3 The scaling process is defined as follows: scaled_x = (x * ( 8 + qpScaled % 8 ) << sh + 4 ) >> 3
[0118] The G-PCC encoder 200 and the G-PCC decoder 300 apply geometry quantization at a specific node depth in the octree (signaled in the bitstream). In a simple example, if a node coordinate is 12 bits and the depth to which quantization should be applied is 3, the first three MSBs of the node coordinate, called the MSB part of the position, are not quantized. Only the last nine LSBs of the node coordinate, called the LSB part of the position, are quantized. Due to quantization, the nine LSBs may be reduced to a smaller number of bits, referred to herein as N, where N is less than 9. This quantization results in some reduction in bit rate at the expense of reconstruction accuracy. The resulting node coordinate size is N+3 (which will be ≦12).
[0119] Similarly, in the G-PCC decoder 300, the N LSBs are scaled and clipped to a maximum value of 1<<(9-1), which ensures that the scaled value does not exceed the 9 LSB bits of the original point. The final scaled position is calculated by splicing the 3 MSBs and the 9 scaled LSB bits.
[0120] In addition, a qp scale factor is also signaled, which determines the minimum number of QPs that can be defined per doubling of the step size. In G-PCC, this factor can take the values 0, 1, 2, and 3.
[0121] The number of QP points per doubling of the step size for various values of qpDivFactorLog2 is specified in the table below.
[0122] [Table 1]
[0123] For all qpDivFactorLog2 values, the step size derivation is kept the same as the scaling process, as 8 QP points for doubling the step size, and the step size derivation is done by adjusting the QP value before the shift bit calculation and scaling process.
[0124] Existing techniques have several potential problems. For example, in conventional coding of point cloud frames, angular mode provides substantial gains in coding efficiency. However, when in-tree quantization is enabled, the gains of angular mode become much smaller and in some cases actually produce losses. For angular modes (IDCM angles and planar angles), quantization bits are used for context derivation. The quantization bits are not in the same domain as the original points, but in a different scale space. This reduces their usefulness, as it may not be beneficial to enable both angular mode and in-tree quantization simultaneously.
[0125] This disclosure describes techniques that may address some of the problems introduced above. The various techniques described herein may be implemented either individually or in combination.
[0126] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to derive a scaled value xS for the node / point position coordinate, given a point / position coordinate value x, to derive the node / point position relative to the lidar origin. In some examples, the scaling operation may include scaling one or more bits of the position, and may additionally include a maximum number of bit values being scaled, where the maximum number is based on a signaled value, such as a value derived from the maximum depth to which quantization should be applied. In some examples, a similar scaled position may also be derived for the y or z coordinate, or possibly more than one coordinate.
[0127] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 use the scaled values to determine the relative position of the node with respect to the LIDAR head position, or more generally, the origin position. The G-PCC encoder 200 and the G-PCC decoder 300 may determine the scaled values to be used based, for example, on whether any quantization has been applied to the node position / coordinate. In some examples, the G-PCC encoder 200 and the G-PCC decoder 300 may use the scaled values when the node position / coordinate value is in a domain that is not the same as the head position. For example, if the node coordinate x_0 and the head position h_0 are described in the same domain (scale), and the value of x_0 is modified to be x by applying quantization, the scaled values are used to determine the relative position with respect to the LIDAR head.
[0128] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 use one or more scaled coordinate values to calculate laser characteristics. The laser characteristics may include, for example, an elevation angle (the angle made by the laser with respect to the x-y plane) or a laser head offset. The laser characteristics may include an azimuth or an azimuth prediction.
[0129] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to perform scaling operations without clipping. For example, the MSB bit of a position may be added with a scaled value of the LSB portion of that bit. In some scenarios, the G-PCC encoder 200 and the G-PCC decoder 300 may apply clipping, in which case the MSB bit of a position is added with a clipped version of the scaled value of the LSB portion of the bit. In some examples, some scaling operations may include clipping, while other scaling operations may not apply clipping. This may be determined by the coordinate being coded, the depth at which QP is applied, and the QP value.
[0130] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to restrict angle modes (in whole or in part) based on the value of qpDivFactorLog2. For example, when the value of qpDivFactorLog2 is 0, 1, or 2, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to disable angle modes (in whole or in part), i.e., only when a power-of-two step size is allowed for geometry quantization / scaling.
[0131] In some examples, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to enable the angular mode (e.g., as indicated in a parameter set, e.g., GPS) only when the QP values of all nodes referencing the parameter set correspond to a step size that is a power of 2. In another example, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to either fully or partially disable the angular mode when intra-tree geometry scaling is enabled (e.g., by setting geom_scaling_enabled_flag equal to 1). In the above example, the partial disabling of the angular mode may include one or more of disabling context derivation of planar mode-related parameters (or bits) using the angular mode and disabling context derivation of IDCM mode-related parameters (or bits) using the angular mode.
[0132] According to the techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to modify the context derivation for an angle mode based on a QP value for a particular node. In one example, one or more thresholds used in the context derivation may be specified based on a QP value (e.g., a function of the QP). The threshold may, for example, remain unchanged for a QP value of 0, but may be modified when the QP is greater than 0. In another example, the threshold may remain unchanged for a QP value that results in a step size that is a power of 2, but may be modified otherwise.
[0133] The modified QP value may be specified using a table of QP values and scale factors, and / or an offset for modifying the threshold. For example, when modified for QP, a fixed multiplier may be used to scale the threshold (for a non-zero QP, double the threshold). In another example, a threshold "z" may be modified as z_mod=func(z, QP), where func() is a predetermined function or a function indicated in the bitstream. For example, func(z, QP)=a(QP)*z+b(QP), where a(QP) and b(QP) are parameter sets based on the QP value (function a() or b() may be a linear or non-linear function).
[0134] The threshold modification may be applied to one or more of the three components x, y, or z. The modification may be applied to one or more of the theta / laser angle related context or the orientation related context. The threshold modification may be applied to one or more of the planar angle mode and the angle IDCM mode.
[0135] According to techniques of this disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to use one or more thresholds to determine whether to enable angle mode, and may modify the one or more thresholds based on the QP using one or more of the techniques used for context derivation. (For example, if angle mode is disabled when the node size is larger than a threshold, the threshold may be modified based on the QP, such as by doubling the threshold for some QP values.)
[0136] One exemplary implementation uses the scaled point values to determine relative positions for the LIDAR nodes. <add> and< / add> In the middle of the addition is shown, and the identifier <del> and< / del> The deletion is indicated between
[0137] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine a planar mode context.
[0138] A function invQuantPositionAngular is defined that applies a scaling operation to the three coordinates of the point based on the step size derived from qp and the bits indicated by quantMasks. This function may be defined as follows: In this example, the scaling operation is similar to the inverse scaling operation defined for in-tree quantization, but any general scaling operation may be applied. quantMasks is a set of 1s and 0s that indicates which bits of the node position should be scaled. For example, if quantMasks=00000111, only the last three bits are scaled. In this example, after the scaling operation, the unscaled portion of the position and the scaled portion of the position are added to form the "scaled" position. The unscaled portion may be shifted before being added to the scaled portion. <add> invQuantPositionAngular(int qp, Vec3<uint32_t> quantMasks, const Vec3<int32_t>& pos) { QuantizerGeom quantizer(qp); int shiftBits = QuantizerGeom::qpShift(qp); Vec3<int64_t> recon; for (int k = 0; k < 3; k++) { int lowPart = (pos[k] & quantMasks[k]) >> shiftBits; int highPart = pos[k] & ~(quantMasks[k]); int lowPartScaled = PCCClip(quantizer.scale(lowPart), 0, quantMasks[k]); recon[k] = highPart | lowPartScaled; } return recon; < / add> In some examples, this process is applied only when qp is not 0. In other examples, the PCCClip() operation may not be performed, and the recon[ ] variable is obtained by adding highPart and lowPartSCaled as follows: <add>int lowPartScaled = quantizer.scale(lowPart) recon[k] = highPart + lowPartScaled;< / add>
[0139] In determining the context for the plane variables, the G-PCC encoder 200 and the G-PCC decoder 300 apply scaling operations to the child node position absPos, midNode (which indicates half the size of the node), and childSize (which indicates the size of the node), and the scaled versions of these values are used in the context derivation process. { Vec3<int64_t> absPos = {child.pos[0] << childSizeLog2[0], child.pos[1] << childSizeLog2[1], child.pos[2] << childSizeLog2[2]}; / / Eligibility Vec3<int64_t> midNode = {1 << (childSizeLog2[0] ? childSizeLog2[0] - 1 : 0), 1 << (childSizeLog2[1] ? childSizeLog2[1] - 1 : 0), 1 << (childSizeLog2[2] ? childSizeLog2[2] - 1 : 0)}; <add> Vec3<int64_t> childSize = { 1 << childSizeLog2[0], 1 << childSizeLog2[1], 1 << childSizeLog2[2]}; if (child.qp) { absPos = invQuantPositionAngular(child.qp, quantMasks, absPos); midNode = invQuantPositionAngular(child.qp, quantMasks, midNode); childSize = invQuantPositionAngular(child.qp, quantMasks, childSize); } < / add> uint64_t xLidar = std::abs(((absPos[0] - headPos[0] + midNode[0] ) << 8) - 128); uint64_t yLidar = std::abs(((absPos[1] - headPos[1] + midNode[1]) << 8) - 128); uint64_t rL1 = (xLidar + yLidar) >> 1; uint64_t deltaAngleR = deltaAngle * rL1; if (deltaAngleR <= (midNode[2] << 26)) return -1; / / Determine the inverse of r (1 / sqrt(r2) = irsqrt(r2)) uint64_t r2 = xLidar * xLidar + yLidar * yLidar; uint64_t rInv = irsqrt(r2); / / Determine the uncorrected theta int64_t zLidar = ((absPos[2] - headPos[2] + midNode[2]) << 1) - 1; int64_t theta = zLidar * rInv; int theta32 = theta >= 0? theta >> 15 : -((-theta) >> 15); / / Determine the laser int laserIndex = int(child.laserIndex); if (laserIndex == 255 || deltaAngleR <= (midNode[2] << (26 + 2))) { auto end = thetaLaser + numLasers - 1; auto it = std::upper_bound(thetaLaser + 1, end, theta32); if (theta32 - *std::prev(it) <= *it - theta32) --it; laserIndex = std::distance(thetaLaser, it); child.laserIndex = uint8_t(laserIndex); } / / -- PHI -- / / Angle int posx = absPos[0] - headPos[0]; int posy = absPos[1] - headPos[1]; int phiNode = iatan2(posy + midNode[1], posx + midNode[0]); int phiNode0 = iatan2(posy, posx); / / Find the predictor int predPhi = phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Use the predictor if (predPhi != 0x80000000) { / / Basis shift predictor int Nshift = ((predPhi - phiNode) * phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= phiZi.delta(laserIndex) * Nshift; / / ctx orientation x or y int angleL = phiNode0 - predPhi; int angleR = phiNode - predPhi; int contextAnglePhi = (angleL >= 0 && angleR >= 0) || (angleL < 0 && angleR < 0) ? 2 : 0; angleL = std::abs(angleL); angleR = std::abs(angleR); if (angleL > angleR) { contextAnglePhi++; int temp = angleL; angleL = angleR; angleR = temp; } if (angleR > (angleL << 2)) contextAnglePhi += 4; if (std::abs(posx) <= std::abs(posy)) *contextAnglePhiX = contextAnglePhi; else *contextAnglePhiY = contextAnglePhi; } / / -- THETA -- int thetaLaserDelta = thetaLaser[laserIndex] - theta32; int64_t hr = zLaser[laserIndex] * rInv; thetaLaserDelta += hr >= 0 ? -(hr >> 17) : ((-hr) >> 17); <del> int64_t zShift = (rInv << childSizeLog2[2]) >> 20;< / del> <add> int64_t zShift = (rInv * childSize[2]) >> 20;< / add> int thetaLaserDeltaBot = thetaLaserDelta + zShift; int thetaLaserDeltaTop = thetaLaserDelta - zShift; int contextAngle = thetaLaserDelta >= 0 ? 0 : 1; if (thetaLaserDeltaTop >= 0) contextAngle += 2; else if (thetaLaserDeltaBot < 0) contextAngle += 2; return contextAngle; }
[0140] In some examples, the G-PCC encoder 200 and the G-PCC decoder 300 may be configured to perform zShift using multiplications only when the scaling step size (derived from the QP) is not a power of 2. Otherwise (the step size is a power of 2), zShift is implemented as a shift operation (as shown above).
[0141] This disclosure describes modifications for the IDCM angle context. The inverse scale function is defined as follows, where the scaling operation is similar to the previously defined function and the argument of the function (noClip) determines whether clipping should be applied or not: class InvQuantizer { int qp; Vec3<uint32_t> quantMasks; public: InvQuantizer(int qpVal, Vec3<uint32_t> quantMaskInp) : qp(qpVal), quantMasks(quantMaskInp) {} int32_t invQuantPositionComp( int idx, const int32_t& pos, bool noClip = false) { if (!qp) return pos; QuantizerGeom quantizer(qp); int shiftBits = QuantizerGeom::qpShift(qp); int32_t recon; int lowPart = pos & (quantMasks[idx] >> shiftBits); int highPart = pos ^ lowPart; int lowPartScaled = quantizer.scale(lowPart); if (!noClip) lowPartScaled = PCCClip(lowPartScaled, 0, quantMasks[idx]); recon = (highPart << shiftBits) + lowPartScaled; return recon; } };
[0142] In some examples, the recon variable may be derived as follows: recon = (highPart << shiftBits) | lowPartScaled;
[0143] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine the nodePosition and laser index, for example, in a manner similar to that described above.
[0144] The G-PCC decoder 300 may initialize the scaling operation (also called inverse quantization) using the QP of the node and the quantization mask applicable to the point location as follows: <add> InvQuantizer invQuantizerIDCM(child.qp, posQuantBitMasks);< / add>
[0145] In the code below, wherever the node's position relative to the LIDAR origin is to be estimated, a scaled value of the position is used. This ensures that the relative position relative to the LIDAR origin is correctly estimated, and that the laser index, altitude, and azimuth are correctly calculated. When a node position is updated by adding a mask value or node size, the added value is also dequantized before adding (or, in some cases, the value is added to the unquantized position and then scaled). int numPoints = 1; bool numPointsGt1 = _arithmeticDecoder->decode(_ctxNumIdcmPointsGt1); numPoints += numPointsGt1; int numDuplicatePoints = 0; if (!geom_unique_points_flag && !numPointsGt1) { numDuplicatePoints = !_arithmeticDecoder->decode(_ctxSinglePointPerBlock); if (numDuplicatePoints) { bool singleDup = _arithmeticDecoder->decode(_ctxSingleIdcmDupPoint); if (!singleDup) numDuplicatePoints += 1 + _arithmeticDecoder->decodeExpGolomb(0, _ctxPointCountPerBlock); } } / / Update the node size after the plane and determine the upper part of the position from the plane Vec3<int32_t> deltaPlanar{0, 0, 0}; Vec3 <int>nodeSizeLog2Rem = nodeSizeLog2; for (int k = 0; k < 3; k++) if (nodeSizeLog2Rem[k] > 0 && (planar.planarMode & (1 << k))) { deltaPlanar[k] |= (planar.planePosBits & (1 << k) ? 1 : 0); nodeSizeLog2Rem[k]--; } / / Indicates which components are coded directly or using angles / / Contextualization Vec3 <bool>directIdcm = !angularIdcm; point_t posNodeLidar; if (angularIdcm) { posNodeLidar = point_t( node.pos[0] << nodeSizeLog2[0], node.pos[1] << nodeSizeLog2[1], node.pos[2] << nodeSizeLog2[2]) <del> - headPos;< / del> <add>if (node.qp) posNodeLidar = invQuantPosition(node.qp, posQuantBitMasks, posNodeLidar); posNodeLidar -= headPos;< / add> bool codeXorY = std::abs(posNodeLidar[0]) <= std::abs(posNodeLidar[1]); directIdcm.x() = !codeXorY; directIdcm.y() = codeXorY; } / / Decode two unordered points Vec3<int32_t> deltaPos[2]; deltaPos[0] = deltaPlanar; deltaPos[1] = deltaPlanar; if (numPoints == 2 && joint_2pt_idcm_enabled_flag) decodeOrdered2ptPrefix(directIdcm, nodeSizeLog2Rem, deltaPos); if (angularIdcm) { <add> InvQuantizer invQuantizerIDCM(node.qp, posQuantBitMasks);< / add> for (int idx = 0; idx < 3; ++idx) { <add> int delta = 0;< / add> int N = nodeSizeLog2[idx] - nodeSizeLog2Rem[idx]; for (int mask = N ? 1 << (N - 1) : 0; mask; mask >>= 1) { if (deltaPos[0][idx] & mask) <add> delta< / add> <del> posNodeLida[idx]< / del> += mask << nodeSizeLog2Rem[idx]; } if (nodeSizeLog2Rem[idx]) <add> delta< / add> <del> posNodeLidar[idx]< / del> += 1 << (nodeSizeLog2Rem[idx] - 1); <add> posNodeLidar[idx] +=< / add> <add> invQuantizerIDCM.invQuantPositionComp(idx, delta, true);< / add> } node.laserIndex = findLaser(posNodeLidar, thetaLaser, numLasers); } Vec3<int32_t> pos; for (int i = 0; i < numPoints; i++) { if (angularIdcm) { *(outputPoints++) = pos = decodePointPositionAngular( nodeSizeLog2, nodeSizeLog2Rem, node, planar, headPos, zLaser, thetaLaser, deltaPos[i]); } else *(outputPoints++) = pos = decodePointPosition(nodeSizeLog2Rem, deltaPos[i]); } for (int i = 0; i < numDuplicatePoints; i++) *(outputPoints++) = pos; return numPoints + numDuplicatePoints;
[0146] The G-PCC encoder 200 and the G-PCC decoder 300 may be configured to determine a context for the angle IDCM.
[0147] As above, when comparing node positions relative to the LIDAR head position, scaled versions of the node size and node position are used. { Vec3<int32_t> delta = deltaPlanar; <del> Vec3 <int>posXyz = {(child.pos[0] << nodeSizeLog2[0]) - headPos[0], (child.pos[1] << nodeSizeLog2[1]) - headPos[1], (child.pos[2] << nodeSizeLog2[2]) - headPos[2]}; < / int> < / del> <add> Vec3 <int>posXyz = { (child.pos[0] << nodeSizeLog2[0]), (child.pos[1] << nodeSizeLog2[1]), (child.pos[2] << nodeSizeLog2[2])}; Vec3 <int>posXyzBeforeQuant = posXyz; if (child.qp) posXyz = invQuantPosition(child.qp, posQuantBitMasks, posXyz); posXyz -= headPos; < / int> < / int> < / add> / / -- PHI -- / / Code x or y directly and calculate phi of the node bool codeXorY = std::abs(posXyz[0]) <= std::abs(posXyz[1]); if (codeXorY) { / / Direct coding of y if (nodeSizeLog2AfterPlanar[1]) for (int i = nodeSizeLog2AfterPlanar[1]; i > 0; i--) { delta[1] <<= 1; delta[1] |= _arithmeticDecoder->decode(); } <del> posXyz[1] += delta[1]; posXyz[0] += delta[0] << nodeSizeLog2AfterPlanar[0]; < / del> <add> posXyz[1]= invQuantizerIDCM.invQuantPositionComp(1, posXyzBeforeQuant[1] + delta[1]) - headPos[1]; posXyz[0] += invQuantizerIDCM.invQuantPositionComp( 0, delta[0] << nodeSizeLog2AfterPlanar[0], true); < / add> } else { / / Direct coding of x if (nodeSizeLog2AfterPlanar[0]) for (int i = nodeSizeLog2AfterPlanar[0]; i > 0; i--) { delta[0] <<= 1; delta[0] |= _arithmeticDecoder->decode(); } <del> posXyz[0] += delta[0]; posXyz[1] += delta[1] << nodeSizeLog2AfterPlanar[1]; < / del> <add> posXyz[0] = invQuantizerIDCM.invQuantPositionComp(0, posXyzBeforeQuant[0] + delta[0]) - headPos[0]; posXyz[1] += invQuantizerIDCM.invQuantPositionComp( 1, delta[1] << nodeSizeLog2AfterPlanar[1], true); < / add> } / / Discover the predictors int phiNode = iatan2(posXyz[1], posXyz[0]); int laserNode = int(child.laserIndex); / / Laser residual int laserIndex = laserNode + decodeThetaRes(); int predPhi = _phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Basis shift predictor int nShift = ((predPhi - phiNode) * _phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= _phiZi.delta(laserIndex) * nShift; / / Choose x or y int* posXY = codeXorY ? &posXyz[0] : &posXyz[1]; int idx = codeXorY ? 0 : 1; / / Orientation coding for x or y int mask2 = codeXorY ? (nodeSizeLog2AfterPlanar[0] > 0 ? 1 << (nodeSizeLog2AfterPlanar[0] - 1) : (nodeSizeLog2AfterPlanar[1] > 0 ? 1 << (nodeSizeLog2AfterPlanar[1] - 1) for (; mask2; mask2 >>= 1) { / / Angle left and right <del> int phiR = codeXorY ? iatan2(posXyz[1], posXyz[0] + mask2) : iatan2(posXyz[1] + mask2, posXyz[0]); < / del> <add> int32_t scaledMask = invQuantizerIDCM.invQuantPositionComp(codeXorY ? 0 : 1, mask2, true); int phiR = codeXorY ? iatan2(posXyz[1], posXyz[0] + scaledMask) : iatan2(posXyz[1] + scaledMask, posXyz[0]); < / add> int phiL = phiNode; / / ctx orientation int angleL = phiL - predPhi; int angleR = phiR - predPhi; int contextAnglePhi = (angleL >= 0 && angleR >= 0) || (angleL < 0 && angleR < 0)? 2 : 0; angleL = std::abs(angleL); angleR = std::abs(angleR); if (angleL > angleR) { contextAnglePhi++; int temp = angleL; angleL = angleR; angleR = temp; } if (angleR > (angleL << 1)) contextAnglePhi += 4; / / Entropy coding bool bit = _arithmeticDecoder->decode( _ctxPlanarPlaneLastIndexAngularPhiIDCM[contextAnglePhi]); delta[idx] <<= 1; if (bit) { delta[idx] |= 1;<…>[[ID=…]]<…>phiNode = phiR; predPhi = _phiBuffer[laserIndex]; if (predPhi == 0x80000000) predPhi = phiNode; / / Basis shift predictor int nShift = ((predPhi - phiNode) * _phiZi.invDelta(laserIndex) + 536870912) >> 30; predPhi -= _phiZi.delta(laserIndex) * nShift; } } / / Update the buffer phi _phiBuffer[laserIndex] = phiNode; / / -- THETA -- int maskz = nodeSizeLog2AfterPlanar[2] > 0 ? 1 << (nodeSizeLog2AfterPlanar[2] - 1) : 0; if (!maskz) return delta; int posz0 = posXyz[2]; <del> posXyz[2] += delta[2] << nodeSizeLog2AfterPlanar[2];< / del> <add> posXyz[2] += invQuantizerIDCM.invQuantPositionComp( 2, delta[2] << nodeSizeLog2AfterPlanar[2], true); < / add> / / Since x and y are known, / / r is also known and bit-independent for z uint64_t xLidar = (int64_t(posXyz[0]) << 8) - 128; uint64_t yLidar = (int64_t(posXyz[1]) << 8) - 128; uint64_t r2 = xLidar * xLidar + yLidar * yLidar; int64_t rInv = irsqrt(r2); / / Use angle to code bit for z. Eligibility is implicit. Laser is known. int64_t hr = zLaser[laserIndex] * rInv; int fixedThetaLaser = thetaLaser[laserIndex] + int(hr >= 0 ? -(hr >> 17) : ((-hr) >> 17)); <del> int zShift = (rInv << nodeSizeLog2AfterPlanar[2]) >> 18;< / del> <add> int zShift = (child.qp ? ( rInv * invQuantizerIDCM.invQuantPositionComp( 2, 1 << nodeSizeLog2AfterPlanar[2], true)) : (rInv << nodeSizeLog2AfterPlanar[2])) >> 18; < / add> for (int bitIdxZ = nodeSizeLog2AfterPlanar[2]; bitIdxZ > 0; bitIdxZ--, maskz >>= 1, zShift >>= 1) { / / Determine the uncorrected theta <del> int64_t zLidar = ((posXyz[2] + maskz) << 1) - 1;< / del> <add> int scaledMaskZ = invQuantizerIDCM.invQuantPositionComp(2, maskz, true);< / add> <add> int64_t zLidar = ((posXyz[2] + scaledMaskZ) << 1) - 1;< / add> int64_t theta = zLidar * rInv; int theta32 = theta >= 0 ? theta >> 15 : -((-theta) >> 15); int thetaLaserDelta = fixedThetaLaser - theta32; int thetaLaserDeltaBot = thetaLaserDelta + zShift; int thetaLaserDeltaTop = thetaLaserDelta - zShift; int contextAngle = thetaLaserDelta >= 0 ? 0 : 1; if (thetaLaserDeltaTop >= 0) contextAngle += 2; else if (thetaLaserDeltaBot < 0) contextAngle += 2; delta[2] <<= 1; delta[2] |= _arithmeticDecoder->decode( _ctxPlanarPlaneLastIndexAngularIdcm[contextAngle]); <del> posXyz[2] = posz0 + (delta[2] << (bitIdxZ - 1));< / del> <add> if (delta[2] & 1)< / add> <add> posXyz[2] += scaledMaskZ;< / add> } return delta; }
[0148] The rest of the derivation is similar to that described above.
[0149] The following example shows an implementation where the thresholds for the z coordinate context in planar angle mode are updated as follows (the updated parts are identified between ** and **): Vec3<int64_t> childSize = { 1 << childSizeLog2[0], 1 << childSizeLog2[1], 1 << childSizeLog2[2]}; if (child.qp) { absPos = invQuantPositionAngular(child.qp, quantMasks, absPos); midNode = invQuantPositionAngular(child.qp, quantMasks, midNode); childSize = invQuantPositionAngular(child.qp, quantMasks, childSize); **childSize[2] = 2*childSize[2];** }
[0150] In this example, the remainder of the changes shown in the example above may also remain.
[0151] The examples in the various aspects of the present disclosure may be used individually or in any combination.
[0152] 9A is a flowchart illustrating an example operation of G-PCC encoder 200 in accordance with one or more techniques of this disclosure. G-PCC encoder 200 determines that in-tree quantization is enabled for a node (900). G-PCC encoder 200 determines that planar mode is activated for the node (902). In response to in-tree quantization being enabled for the node, G-PCC encoder 200 determines a quantized value for the node that represents a coordinate position relative to an origin position (904). G-PCC encoder 200 scales the quantized value without clipping to determine a scaled value that represents a coordinate position relative to the origin position (906).
[0153] To scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the G-PCC encoder 200 may be configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine a scaled value representing a coordinate position relative to the origin position. To scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the G-PCC encoder 200 may be configured to determine a shift amount for the MSBs based on a quantization parameter for the node, shift the MSBs based on the shift amount, and add the scaled LSBs to the shifted MSBs to determine a scaled value representing a coordinate position relative to the origin position. The G-PCC encoder 200 may receive an indication of the number of bits in the group of LSBs, for example, in syntax signaled in the bitstream. The indication may be either an explicit indication or an implicit indication. As an example, based on the received node depth, the G-PCC encoder 200 may be configured to derive the number of MSBs and LSBs.
[0154] The G-PCC encoder 200 determines a context for context encoding a planar position syntax element for the angle mode based on the scaled value representing a coordinate position relative to the origin position (908). The planar position syntax element may indicate, for example, a vertical plane position. To determine a context for context encoding a planar position syntax element for the angle mode based on the scaled value representing a coordinate position relative to the origin position, the G-PCC encoder 200 may be configured to determine one or more laser characteristics based on the scaled value and the origin position and encode the planar position syntax element for the angle mode based on the laser characteristics. To determine a context for context encoding a planar position syntax element for the angle mode based on the laser characteristics, the G-PCC encoder 200 may be configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0155] The above steps 904, 906, and 908 may be performed, for example, as part of a decoding operation performed by the G-PCC encoder 200. The G-PCC encoder 200 may perform decoding as part of encoding. For example, to determine whether a particular encoding scheme provides a desirable rate-distortion tradeoff, the G-PCC encoder 200 may encode point cloud data and then decode the encoded point cloud data so that the decoded point cloud data can be compared to the original point cloud data to determine the amount of distortion. When using various predictive coding tools, the G-PCC encoder 200 may also encode point cloud data and then decode the encoded point cloud data so that the G-PCC encoder 200 can perform predictions based on the same point cloud data available to the G-PCC decoder.
[0156] 9B is a flowchart illustrating an example operation of a G-PCC decoder 300 in accordance with one or more techniques of this disclosure. The G-PCC decoder 300 determines, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node (920). The G-PCC decoder 300 determines, based on syntax signaled in the bitstream, that planar mode is activated for the node (922). In response to in-tree quantization being enabled for the node, the G-PCC decoder 300 determines, for the node, a quantized value that represents a coordinate position relative to an origin position (924).
[0157] The G-PCC decoder 300 scales (926) the quantized values without clipping to determine scaled values that represent coordinate positions relative to the origin position.
[0158] To scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the G-PCC decoder 300 may be configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine a scaled value representing a coordinate position relative to the origin position. To scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the G-PCC decoder 300 may be configured to determine a shift amount for the MSBs based on a quantization parameter for the node, shift the MSBs based on the shift amount, and add the scaled LSBs to the shifted MSBs to determine a scaled value representing a coordinate position relative to the origin position. The G-PCC decoder 300 may receive an indication of the number of bits in the group of LSBs, for example, in syntax signaled in the bitstream. The indication may be either an explicit indication or an implicit indication. As an example, based on the received node depth, the G-PCC decoder 300 may be configured to derive the number of MSBs and LSBs.
[0159] The G-PCC decoder 300 determines a context for context decoding, e.g., arithmetically decoding, a planar position syntax element for the angular mode based on the scaled value representing a coordinate position relative to the origin position (928). The planar position syntax element may indicate, for example, a vertical plane position. To determine a context for context decoding a planar position syntax element for the angular mode based on the scaled value representing a coordinate position relative to the origin position, the G-PCC decoder 300 may be configured to determine one or more laser characteristics based on the scaled value and the origin position and decode the planar position syntax element for the angular mode based on the laser characteristics. To determine a context for context decoding a planar position syntax element for the angular mode based on the laser characteristics, the G-PCC decoder 300 may be configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0160] The G-PCC decoder 300 may be configured to reconstruct the point cloud, for example, by determining the positions of one or more points of the point cloud based on the plane positions.
[0161] FIG. 10 is a conceptual diagram illustrating an exemplary distance measurement system 1000 that may be used with one or more techniques of the present disclosure. In the example of FIG. 10, the distance measurement system 1000 includes an illuminator 1002 and a sensor 1004. The illuminator 1002 may emit a light beam 1006. In some examples, the illuminator 1002 may emit the light beam 1006 as one or more laser beams. The light beam 1006 may be at one or more wavelengths, such as infrared wavelengths or visible light wavelengths. In other examples, the light beam 1006 is not coherent laser light. When the light beam 1006 encounters an object, such as an object 1008, the light beam 1006 produces a return light beam 1010. The return light beam 1010 may include backscattered light and / or reflected light. The return light beam 1010 may pass through a lens 1011, which directs the return light beam 1010 to produce an image 1012 of the object 1008 on the sensor 1004. The sensor 1004 generates a signal 1018 based on the image 1012. The image 1012 may comprise a set of points (e.g., as represented by the dots in the image 1012 of Figure 10).
[0162] In some examples, the illuminator 1002 and sensor 1004 may be mounted on a spinning structure such that the illuminator 1002 and sensor 1004 capture a 360-degree view of the environment. In other examples, the distance measurement system 1000 may include one or more optical components (e.g., mirrors, collimators, diffraction gratings, etc.) that enable the illuminator 1002 and sensor 1004 to detect objects within a certain range (e.g., up to 360 degrees). Although the example of FIG. 10 shows only a single illuminator 1002 and sensor 1004, the distance measurement system 1000 may include multiple sets of illuminators and sensors.
[0163] In some examples, the illuminator 1002 generates a structured illumination pattern. In such examples, the distance measurement system 1000 may include multiple sensors 1004 on which respective images of the structured illumination pattern are formed. The distance measurement system 1000 may use the parallax between the images of the structured illumination pattern to determine the distance to an object 1008 from which the structured illumination pattern is backscattered. The structured illumination-based distance measurement system may have a high level of accuracy (e.g., accuracy in the sub-millimeter range) when the object 1008 is relatively close to the sensor 1004 (e.g., between 0.2 meters and 2 meters). This high level of accuracy may be useful in facial recognition applications, such as unlocking a mobile device (e.g., a mobile phone, a tablet computer, etc.), and for security applications.
[0164] In some examples, the distance measurement system 1000 is a time-of-flight (ToF)-based system. In some examples where the distance measurement system 1000 is a ToF-based system, the illuminator 1002 generates pulses of light. In other words, the illuminator 1002 may modulate the amplitude of the emitted light beam 1006. In such examples, the sensor 1004 detects a return light beam 1010 from the pulse of light beam 1006 generated by the illuminator 1002. The distance measurement system 1000 can then determine the distance to the object 1008 from which the light beam 1006 backscatters based on the delay between when the light beam 1006 is emitted and when it is detected and the known speed of light in air. In some examples, rather than (or in addition to) modulating the amplitude of the emitted light beam 1006, the illuminator 1002 may modulate the phase of the emitted light beam 1006. In such an example, the sensor 1004 may detect the phase of the return light ray 1010 from the object 1008 and may use the speed of light and determine the distance to a point on the object 1008 based on the time difference between when the illuminator 1002 generates the light ray 1006 at a particular phase and when the sensor 1004 detects the return light ray 1010 at that particular phase.
[0165] In other examples, the point cloud may be generated without the use of the illuminator 1002. For example, in some examples, the sensor 1004 of the distance measurement system 1000 may include two or more optical cameras. In such examples, the distance measurement system 1000 may use the optical cameras to capture stereo images of an environment including the object 1008. The distance measurement system 1000 (e.g., the point cloud generator 1020) may then calculate disparities between locations in the stereo images. The distance measurement system 1000 may then use the disparities to determine distances to locations shown in the stereo images. From these distances, the point cloud generator 1020 may generate a point cloud.
[0166] The sensor 1004 may also detect other attributes of the object 1008, such as color and reflectance information. In the example of Figure 10, the point cloud generator 1020 may generate a point cloud based on the signal 1018 generated by the sensor 1004. The distance measurement system 1000 and / or the point cloud generator 1020 may form part of the data source 104 (Figure 1).
[0167] FIG. 11 is a conceptual diagram illustrating an example vehicle-based scenario in which one or more techniques of the present disclosure may be used. In the example of FIG. 11, a vehicle 1100 includes a laser package 1102, such as a LIDAR system. Although not shown in the example of FIG. 11, the vehicle 1100 may also include a data source and a G-PCC encoder, such as G-PCC encoder 200 (FIG. 1). In the example of FIG. 11, the laser package 1102 emits a laser beam 1104 that reflects off a pedestrian 1106 or other object in the road. The data source of the vehicle 1100 may generate a point cloud based on a signal generated by the laser package 1102. The G-PCC encoder of the vehicle 1100 may encode the point cloud to generate a bit stream 1108. The bit stream 1108 may include significantly fewer bits than the unencoded point cloud obtained by the G-PCC encoder. An output interface of the vehicle 1100 (e.g., output interface 108 (FIG. 1)) may transmit the bit stream 1108 to one or more other devices. Therefore, the vehicle 1100 may be able to transmit the bitstream 1108 to other devices more quickly than unencoded point cloud data. Additionally, the bitstream 1108 may require less data storage capacity.
[0168] In the example of FIG. 11 , vehicle 1100 may transmit bitstream 1108 to another vehicle 1110. Vehicle 1110 may include a G-PCC decoder, such as G-PCC decoder 300 (FIG. 1). The G-PCC decoder of vehicle 1110 may decode bitstream 1108 and reconstruct a point cloud. Vehicle 1110 may use the reconstructed point cloud for various purposes. For example, vehicle 1110 may determine that pedestrian 1106 is in the road ahead of vehicle 1100 based on the reconstructed point cloud and thus may begin to slow down, for example, even before the driver of vehicle 1110 realizes that pedestrian 1106 is in the road. Thus, in some examples, vehicle 1110 may perform autonomous navigation operations, generate notifications or alerts, or take another action based on the reconstructed point cloud.
[0169] Additionally or alternatively, vehicle 1100 may transmit bitstream 1108 to server system 1112. Server system 1112 may use bitstream 1108 for various purposes. For example, server system 1112 may store bitstream 1108 for subsequent reconstruction of a point cloud. In this example, server system 1112 may use the point cloud along with other data (e.g., vehicle telemetry data generated by vehicle 1100) to train an autonomous driving system. In other examples, server system 1112 may store bitstream 1108 for subsequent reconstruction for forensic crash investigation (e.g., if vehicle 1100 collides with pedestrian 1106) or transmit notifications or instructions for navigation to vehicle 1100 or vehicle 1110.
[0170] FIG. 12 is a conceptual diagram illustrating an example augmented reality system in which one or more techniques of this disclosure may be used. Extended reality (XR) is a term used to cover a range of technologies, including augmented reality (AR), mixed reality (MR), and virtual reality (VR). In the example of FIG. 12, a first user 1200 is located at a first location 1202. The user 1200 is wearing an XR headset 1204. As an alternative to the XR headset 1204, the user 1200 may use a mobile device (e.g., a mobile phone, a tablet computer, etc.). The XR headset 1204 includes a depth-sensing sensor, such as a LIDAR system, that detects the position of a point on an object 1206 at the first location 1202. A data source of the XR headset 1204 may generate a point cloud representation of the object 1206 at the location 1202 using signals generated by the depth-sensing sensor. The XR headset 1204 may include a G-PCC encoder (e.g., G-PCC encoder 200 of FIG. 1) configured to encode the point cloud to generate a bitstream 1208.
[0171] The XR headset 1204 may transmit the bitstream 1208 (e.g., over a network such as the internet) to an XR headset 1210 worn by a user 1212 at a second location 1214. The XR headset 1210 may decode the bitstream 1208 and reconstruct a point cloud. The XR headset 1210 may use the point cloud to generate an XR visualization (e.g., an AR, MR, or VR visualization) representing the object 1206 at the location 1202. Thus, in some examples, such as when the XR headset 1210 generates a VR visualization, the user 1212 at the location 1214 may have a 3D immersive experience of the location 1202. In some examples, the XR headset 1210 may determine the position of a virtual object based on the reconstructed point cloud. For example, the XR headset 1210 may determine, based on the reconstructed point cloud, that the environment (e.g., location 1202) includes a flat surface and may then determine that a virtual object (e.g., an animated character) should be placed on the flat surface. The XR headset 1210 may generate an XR visualization with the virtual object at the determined location. For example, the XR headset 1210 may show the animated character sitting on the flat surface.
[0172] FIG. 13 is a conceptual diagram illustrating an example mobile device system in which one or more techniques of this disclosure may be used. In the example of FIG. 13, a mobile device 1300, such as a mobile phone or tablet computer, includes a depth-sensing sensor, such as a LIDAR system, that detects the positions of points on an object 1302 in the environment of the mobile device 1300. A data source of the mobile device 1300 may generate a point cloud representation of the object 1302 using signals generated by the depth-sensing sensor. The mobile device 1300 may include a G-PCC encoder (e.g., G-PCC encoder 200 of FIG. 1) configured to encode the point cloud to generate a bitstream 1304. In the example of FIG. 13, the mobile device 1300 may transmit the bitstream to a remote device 1306, such as a server system or another mobile device. The remote device 1306 may decode the bitstream 1304 to reconstruct the point cloud. The remote device 1306 may use the point cloud for various purposes. For example, the remote device 1306 may use the point cloud to generate a map of the environment of the mobile device 1300. For example, the remote device 1306 may generate a map of the interior of a building based on the reconstructed point cloud. In another example, the remote device 1306 may generate imagery (e.g., computer graphics) based on the point cloud. For example, the remote device 1306 may use the points of the point cloud as vertices of a polygon and may use the color attributes of the points as a basis for shading the polygon. In some examples, the remote device 1306 may perform facial recognition using the point cloud.
[0173] The following numbered clauses illustrate one or more aspects of the devices and techniques described in this disclosure.
[0174] Clause 1A. A method of processing point cloud data, the method comprising receiving data representing a point cloud and processing the data representing the point cloud according to any one or more techniques of the present disclosure for generating a point cloud.
[0175] Clause 2A. A device for processing a point cloud, the device comprising one or more means for receiving data representing a point cloud and processing the data representing the point cloud in accordance with any one or more techniques of the present disclosure for generating a point cloud.
[0176] Clause 3A. The method of clause 1A or the device of clause 2A, wherein processing data representing a point cloud according to any one or more techniques of the present disclosure for generating a point cloud comprises disabling angular mode either in whole or in part.
[0177] Clause 4A. The device of clause 3A, wherein the one or more means comprise one or more processors implemented in circuitry.
[0178] Clause 5A. The device of any one of clauses 2A to 4A, further comprising a memory for storing data representing a point cloud.
[0179] Clause 6A. The device of any of clauses 2A to 5A, wherein the device comprises a decoder.
[0180] Clause 7A. The device of any of clauses 2A to 5A, wherein the device comprises an encoder.
[0181] Clause 8A. The device of any one of clauses 2A to 7A, further comprising a device for generating a point cloud.
[0182] Clause 9A. The device of any of clauses 2A to 8A, further comprising a display for presenting an image based on the point cloud.
[0183] Clause 10A. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to receive data representing a point cloud and process the data representing the point cloud according to any one or more techniques of the present disclosure for generating a point cloud.
[0184] Clause 1B. A device for decoding a bitstream including point cloud data, the device comprising: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors being configured to: determine for a node based on syntax signaled in the bitstream that intra-tree quantization is enabled; determine for the node based on syntax signaled in the bitstream that angular mode is activated for the node; in response to intra-tree quantization being enabled for the node, determine a quantized value for the node that represents a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determine a context for context decoding of a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0185] Clause 2B. The device of clause 1B, wherein to determine a context for context decoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position, the one or more processors are configured to determine one or more laser characteristics based on the scaled value and the origin position, and decode the planar position syntax element for the angular mode based on the laser characteristics.
[0186] Clause 3B. The device of Clause 2B, wherein to determine a context for context decoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position, the one or more processors are further configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0187] Clause 4B. The device of clause 2B, wherein the planar position syntax element indicates a vertical plane position.
[0188] Clause 5B. The device of clause 2B, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0189] Clause 6B. The device of clause 1B, wherein the one or more processors are further configured to arithmetically decode the planar position in the angular mode using a context indicated by the determined context.
[0190] Clause 7B. The device of clause 1B, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine a scaled value representing a coordinate position relative to the origin position.
[0191] Clause 8B. The device of clause 7B, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a shift amount for the MSB based on a quantization parameter for the node, shift the MSB based on the shift amount, and add the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0192] Clause 9B. The device of clause 6B, wherein the one or more processors are further configured to receive, in syntax signaled in the bitstream, an indication of the number of bits in the group of LSBs.
[0193] Clause 10B. The device of clause 1B, wherein the one or more processors are further configured to reconstruct the point cloud.
[0194] Clause 11B. The device of clause 10B, wherein the one or more processors are configured to determine positions of one or more points of the point cloud based on the planar positions as part of reconstructing the point cloud.
[0195] Clause 12B. The device of clause 11B, wherein the one or more processors are further configured to generate a map of the interior of the building based on the reconstructed point cloud.
[0196] Clause 13B. The device of clause 11B, wherein the one or more processors are further configured to perform autonomous navigation operations based on the reconstructed point cloud.
[0197] Clause 14B. The device of clause 11B, wherein the one or more processors are further configured to generate computer graphics based on the reconstructed point cloud.
[0198] Clause 15B. The device of clause 11B, wherein the one or more processors are configured to determine a position of a virtual object based on the reconstructed point cloud and generate an extended reality (XR) visualization with the virtual object at the determined position.
[0199] Clause 16B. The device of clause 11B, further comprising a display for presenting an image based on the reconstructed point cloud.
[0200] Clause 17B. The device of clause 1B, wherein the device is one of a mobile phone or a tablet computer.
[0201] Clause 18B. A device according to clause 1B, where the device is a vehicle.
[0202] Clause 19B. The device of clause 1B, wherein the device is an augmented reality device.
[0203] Clause 20B. A method for decoding a bitstream including point cloud data, the method comprising: determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determining, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0204] Clause 21B. The method of clause 20B, wherein determining a context for context decoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position comprises determining one or more laser characteristics based on the scaled value and the origin position, and decoding the planar position syntax element for the angular mode based on the laser characteristics.
[0205] Clause 22B. The method of clause 21B, wherein determining a context for context decoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position further comprises determining a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0206] Clause 23B. The method of clause 21B, wherein the planar position syntax element indicates a vertical plane position.
[0207] Clause 24B. The method of clause 21B, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0208] Clause 25B. The method of clause 20B, further comprising arithmetically decoding the angular mode planar position using the context indicated by the determined context.
[0209] Clause 26B. The method of clause 20B, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scaling the LSBs without clipping to determine scaled LSBs, and adding the scaled LSBs to the MSBs to determine a scaled value representing a coordinate position relative to the origin position.
[0210] Clause 27B. The method of clause 26B, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a shift amount for an MSB based on a quantization parameter for the node, shifting the MSB based on the shift amount, and adding the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0211] Clause 28B. The method of clause 25B, further comprising receiving in syntax signaled in the bitstream an indication of the number of bits in the group of LSBs.
[0212] Clause 29B. The method of clause 20B, further comprising reconstructing the point cloud.
[0213] Clause 30B. The method of clause 29B, wherein reconstructing the point cloud comprises determining positions of one or more points of the point cloud based on the planar positions.
[0214] Clause 31B. The method of clause 29B, further comprising generating a map of the interior of the building based on the reconstructed point cloud.
[0215] Clause 32B. The method of clause 29B, further comprising performing an autonomous navigation operation based on the reconstructed point cloud.
[0216] Clause 33B. The method of clause 29B, further comprising generating computer graphics based on the reconstructed point cloud.
[0217] Clause 34B. The method of clause 29B, further comprising determining a position of a virtual object based on the reconstructed point cloud; and generating an extended reality (XR) visualization with the virtual object at the determined position.
[0218] Clause 35B. A device for encoding a bitstream including point cloud data, the device comprising: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors being configured to: determine that intra-tree quantization is enabled for a node; determine that an angular mode is activated for the node; in response to intra-tree quantization being enabled for the node, determine a quantized value for the node that represents a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determine a context for context encoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0219] Clause 36B. The device of clause 35B, wherein to determine a context for context encoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position, the one or more processors are configured to determine one or more laser characteristics based on the scaled value and the origin position, and decode the planar position syntax element for the angular mode based on the laser characteristics.
[0220] Clause 37B. The device of Clause 36B, wherein to determine a context for context encoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position, the one or more processors are further configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0221] Clause 38B. The device of clause 36B, wherein the planar position syntax element indicates a vertical plane position.
[0222] Clause 39B. The device of clause 36B, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0223] Clause 40B. The device of clause 35B, wherein the one or more processors are further configured to arithmetically encode the planar position of the angular mode using a context indicated by the determined context.
[0224] Clause 41B. The device of clause 35B, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine the scaled value representing the coordinate position relative to the origin position.
[0225] Clause 42B. The device of clause 41B, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a shift amount for the MSB based on a quantization parameter for the node, shift the MSB based on the shift amount, and add the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0226] Clause 43B. The device of clause 35B, wherein the one or more processors are further configured to reconstruct the point cloud.
[0227] Clause 44B. The device of clause 35B, wherein, to reconstruct the point cloud, the one or more processors are further configured to determine positions of one or more points of the point cloud based on the planar positions.
[0228] Clause 45B. The device of clause 35B, wherein the device is one of a mobile phone or a tablet computer.
[0229] Clause 46B. A device under clause 35B, where the device is a vehicle.
[0230] Clause 47B. The device of clause 35B, wherein the device is an augmented reality device.
[0231] Clause 48B. A method for encoding a bitstream including point cloud data, the method comprising: determining that in-tree quantization is enabled for a node; determining that an angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context encoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0232] Clause 49B. The method of clause 48B, wherein determining a context for context encoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position comprises determining one or more laser characteristics based on the scaled value and the origin position, and decoding the planar position syntax element for the angular mode based on the laser characteristics.
[0233] Clause 50B. The method of clause 49B, wherein determining a context for context encoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position further comprises determining a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0234] Clause 51B. The method of clause 49B, wherein the planar position syntax element indicates a vertical plane position.
[0235] Clause 52B. The method of clause 49B, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0236] Clause 53B. The method of clause 48B, further comprising arithmetically encoding the planar position of the angular mode using the context indicated by the determined context.
[0237] Clause 54B. The method of clause 48B, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scaling the LSBs without clipping to determine scaled LSBs, and adding the scaled LSBs to the MSBs to determine a scaled value representing a coordinate position relative to the origin position.
[0238] Clause 55B. The method of clause 54B, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a shift amount for an MSB based on a quantization parameter for the node, shifting the MSB based on the shift amount, and adding the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0239] Clause 56B. The method of clause 48B, further comprising reconstructing the point cloud.
[0240] Clause 57B. The method of clause 56B, wherein reconstructing the point cloud comprises determining positions of one or more points of the point cloud based on the planar positions.
[0241] Clause 58B. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to determine, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determine, based on syntax signaled in the bitstream, that angle mode is activated for the node; in response to in-tree quantization being enabled for the node, determine, for the node, a quantized value representing a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value representing the coordinate position relative to the origin position; and determine, based on the scaled value representing the coordinate position relative to the origin position, a context for context decoding a planar position syntax element for the angle mode.
[0242] Clause 59B. A device for decoding a bitstream including point cloud data, the device comprising: means for determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; means for determining, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; means for determining, in response to in-tree quantization being enabled for the node, a quantized value for the node that represents a coordinate position relative to an origin position; means for scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and means for determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0243] Clause 1C. A device for decoding a bitstream including point cloud data, the device comprising: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors being configured to: determine for a node based on syntax signaled in the bitstream that intra-tree quantization is enabled; determine for the node based on syntax signaled in the bitstream that angular mode is activated for the node; in response to intra-tree quantization being enabled for the node, determine a quantized value for the node that represents a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determine a context for context decoding of a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0244] Clause 2C. The device of Clause 1C, wherein to determine a context for context decoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position, the one or more processors are configured to determine one or more laser characteristics based on the scaled value and the origin position, and decode the planar position syntax element for the angular mode based on the laser characteristics.
[0245] Clause 3C. The device of Clause 2C, wherein to determine a context for context decoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position, the one or more processors are further configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0246] Clause 4C. The device of any of clauses 1C to 3C, wherein the planar position syntax element indicates a vertical plane position.
[0247] Clause 5C. The device of any of clauses 2C-4C, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0248] Clause 6C. The device of any of clauses 1C to 5C, wherein the one or more processors are further configured to arithmetically decode the planar position in the angular mode using a context indicated by the determined context.
[0249] Clause 7C. The device of any of clauses 1C to 6C, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine the scaled value representing the coordinate position relative to the origin position.
[0250] Clause 8C. The device of clause 7C, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a shift amount for the MSB based on a quantization parameter for the node, shift the MSB based on the shift amount, and add the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0251] Clause 9C. The device of any of clauses 6C to 8C, wherein the one or more processors are further configured to receive, in syntax signaled in the bitstream, an indication of the number of bits in the group of LSBs.
[0252] Clause 10C. The device of any of clauses 1C-9C, wherein the one or more processors are further configured to reconstruct the point cloud.
[0253] Clause 11C. The device of clause 10C, wherein the one or more processors are configured to determine positions of one or more points of the point cloud based on the planar positions as part of reconstructing the point cloud.
[0254] Clause 12C. The device of clause 11C, wherein the one or more processors are further configured to generate a map of the interior of the building based on the reconstructed point cloud.
[0255] Clause 13C. The device of clause 11C, wherein the one or more processors are further configured to perform autonomous navigation operations based on the reconstructed point cloud.
[0256] Clause 14C. The device of clause 11C, wherein the one or more processors are further configured to generate computer graphics based on the reconstructed point cloud.
[0257] Clause 15C. The device of clause 11C, wherein the one or more processors are configured to determine a position of a virtual object based on the reconstructed point cloud and generate an extended reality (XR) visualization with the virtual object at the determined position.
[0258] Clause 16C. The device of any of clauses 11C to 15C, further comprising a display for presenting an image based on the reconstructed point cloud.
[0259] Clause 17C. The device of any of clauses 1C to 16C, wherein the device is one of a mobile phone or a tablet computer.
[0260] Clause 18C. A device of any of clauses 1C to 16C, wherein the device is a vehicle.
[0261] Clause 19C. The device of any of clauses 1C to 16C, wherein the device is an augmented reality device.
[0262] Clause 20C. A method for decoding a bitstream including point cloud data, the method comprising: determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determining, for the node, based on syntax signaled in the bitstream, that angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0263] Clause 21C. The method of clause 20C, wherein determining a context for context-decoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position comprises determining one or more laser characteristics based on the scaled value and the origin position, and decoding the planar position syntax element for the angular mode based on the laser characteristics.
[0264] Clause 22C. The method of clause 21C, wherein determining a context for context decoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position further comprises determining a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0265] Clause 23C. The method of any of clauses 20C to 22C, wherein the planar position syntax element indicates a vertical plane position.
[0266] Clause 24C. The method of any of clauses 21C-23C, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0267] Clause 25C. The method of any of clauses 20C to 24C, further comprising arithmetically decoding the planar position of the angular mode using the context indicated by the determined context.
[0268] Clause 26C. The method of any of clauses 20C to 25C, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scaling the LSBs without clipping to determine scaled LSBs, and adding the scaled LSBs to the MSBs to determine the scaled value representing the coordinate position relative to the origin position.
[0269] Clause 27C. The method of clause 26C, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a shift amount for an MSB based on a quantization parameter for the node, shifting the MSB based on the shift amount, and adding the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0270] Clause 28C. The method of any of clauses 25C to 27C, further comprising receiving an indication of the number of bits in the group of LSBs in syntax signaled in the bitstream.
[0271] Clause 29C. The method of any of clauses 20C to 28C, further comprising reconstructing the point cloud.
[0272] Clause 30C. The method of clause 29C, wherein reconstructing the point cloud comprises determining positions of one or more points of the point cloud based on the planar positions.
[0273] Clause 31C. The method of clause 29C, further comprising generating a map of the interior of the building based on the reconstructed point cloud.
[0274] Clause 32C. The method of clause 29C, further comprising performing an autonomous navigation operation based on the reconstructed point cloud.
[0275] Clause 33C. The method of clause 29C, further comprising generating computer graphics based on the reconstructed point cloud.
[0276] Clause 34C. The method of clause 29C, further comprising determining a position of a virtual object based on the reconstructed point cloud; and generating an extended reality (XR) visualization with the virtual object at the determined position.
[0277] Clause 35C. A device for encoding a bitstream including point cloud data, the device comprising: a memory for storing the point cloud data; and one or more processors coupled to the memory and implemented in circuitry, the one or more processors configured to: determine that intra-tree quantization is enabled for a node; determine that an angular mode is activated for the node; in response to intra-tree quantization being enabled for the node, determine a quantized value for the node that represents a coordinate position relative to an origin position; scale the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determine a context for context encoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0278] Clause 36C. The device of clause 35C, wherein to determine a context for context encoding a planar position syntax element for an angular mode based on the scaled value representing a coordinate position relative to an origin position, the one or more processors are configured to determine one or more laser characteristics based on the scaled value and the origin position, and decode the planar position syntax element for the angular mode based on the laser characteristics.
[0279] Clause 37C. The device of Clause 36C, wherein to determine a context for context encoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position, the one or more processors are further configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0280] Clause 38C. The device of any one of clauses 35C to 37C, wherein the horizontal position syntax element indicates a vertical plane position.
[0281] Clause 39C. The device of any of clauses 36C-38C, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0282] Clause 40C. The device of any of clauses 35C to 39C, wherein the one or more processors are further configured to arithmetically encode the planar position of the angular mode using a context indicated by the determined context.
[0283] Clause 41C. The device of any of clauses 35C to 40C, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scale the LSBs without clipping to determine scaled LSBs, and add the scaled LSBs to the MSBs to determine the scaled value representing the coordinate position relative to the origin position.
[0284] Clause 42C. The device of clause 41C, wherein to scale the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position, the one or more processors are configured to determine a shift amount for the MSB based on a quantization parameter for the node, shift the MSB based on the shift amount, and add the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0285] Clause 43C. The device of any of clauses 35C to 42C, wherein the one or more processors are further configured to reconstruct the point cloud.
[0286] Clause 44C. The device of any of clauses 35C to 43C, wherein, to reconstruct the point cloud, the one or more processors are further configured to determine positions of one or more points of the point cloud based on the planar positions.
[0287] Clause 45C. The device of any of clauses 35C to 44C, wherein the device is one of a mobile phone or a tablet computer.
[0288] Clause 46C. A device of any of clauses 35C to 44C, wherein the device is a vehicle.
[0289] Clause 47C. The device of any of clauses 35C to 44C, wherein the device is an augmented reality device.
[0290] Clause 48C. A method for encoding a bitstream including point cloud data, the method comprising: determining that in-tree quantization is enabled for a node; determining that an angular mode is activated for the node; in response to in-tree quantization being enabled for the node, determining a quantized value for the node that represents a coordinate position relative to an origin position; scaling the quantized value without clipping to determine a scaled value that represents the coordinate position relative to the origin position; and determining a context for context encoding a planar position syntax element for the angular mode based on the scaled value that represents the coordinate position relative to the origin position.
[0291] Clause 49C. The method of clause 48C, wherein determining a context for context encoding a planar position syntax element for an angular mode based on a scaled value representing a coordinate position relative to an origin position comprises determining one or more laser characteristics based on the scaled value and the origin position, and decoding the planar position syntax element for the angular mode based on the laser characteristics.
[0292] Clause 50C. The method of clause 49C, wherein determining a context for context encoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position relative to the origin position further comprises determining a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
[0293] Clause 51C. The method of any of clauses 48C to 50C, wherein the planar position syntax element indicates a vertical plane position.
[0294] Clause 52C. The method of any of clauses 49C-51C, wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
[0295] Clause 53C. The method of any of clauses 48C to 52C, further comprising arithmetically encoding the planar position of the angular mode using the context indicated by the determined context.
[0296] Clause 54C. The method of any of clauses 48C to 53C, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a group of most significant bits (MSBs) and a group of least significant bits (LSBs), scaling the LSBs without clipping to determine scaled LSBs, and adding the scaled LSBs to the MSBs to determine the scaled value representing the coordinate position relative to the origin position.
[0297] Clause 55C. The method of clause 54C, wherein scaling the quantized values without clipping to determine a scaled value representing a coordinate position relative to the origin position comprises determining a shift amount for an MSB based on a quantization parameter for the node, shifting the MSB based on the shift amount, and adding the scaled LSB to the shifted MSB to determine a scaled value representing a coordinate position relative to the origin position.
[0298] Clause 56C. The method of any of clauses 48C to 55C, further comprising reconstructing the point cloud.
[0299] Clause 57C. The method of clause 56C, wherein reconstructing the point cloud comprises determining positions of one or more points of the point cloud based on the planar positions.
[0300] It should be appreciated that, depending on the example, some acts or events of any of the techniques described herein may be performed in a different sequence, added, combined, or entirely excluded (e.g., not all described acts or events may be necessary to practice the techniques). Moreover, in some examples, acts or events may be performed in parallel rather than sequentially, for example, through multithreaded processing, interrupt processing, or multiple processors.
[0301] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.
[0302] By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0303] The instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuitry. Accordingly, the terms "processor" and "processing circuitry," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuits or logic elements.
[0304] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0305] Various examples are described. These and other examples are within the scope of the following claims. [Explanation of symbols]
[0306] 100 Encoding and Decoding Systems 102 Source Devices 104 Data Sources 106 memory 108 Output Interface 110 Computer-Readable Medium 112 Storage Devices 114 File Server 116 Destination Device 118 Data Consumers 120 memory 122 input interface 200 G-PCC Encoder 202 Coordinate Transformation Unit 204 Color Conversion Unit 206 Voxelization Unit 208 Attribute Transfer Unit 210 Octree Analysis Unit 212 Surface Approximation Analysis Unit 214 Arithmetic Coding Unit 216 Geometry Reconstruction Unit 218 RAHT Unit 220 LOD generation units 222 Lifting Unit 224 Coefficient Quantization Unit 226 Arithmetic Coding Unit 300 G-PCC decoder 302 Geometry Arithmetic Decoding Unit 304 Attribute Arithmetic Decoding Unit 306 octree synthesis unit 308 Inverse Quantization Unit 310 Surface Approximation Synthesis Unit 312 Geometry Reconstruction Unit 314 RAHT unit 316 LOD generation units 318 Reverse Lifting Unit 320 Inverse Coordinate Transformation Unit 322 Inverse Color Conversion Unit 400 nodes 402 Child Node 500 laser beam positions 502 marker points 504 nodes 600 nodes 602 Distance Section Threshold 604 Laser Beam 606 Center point, Marker point 700 distance section threshold 702 marker points 704 nodes 706 Laser Beam 1000 Distance Measurement System 1002 Lighting equipment 1004 Sensor 1006 Ray of light 1008 Object 1010 Return Ray 1011 Lens 1012 images 1018 signal 1020 point cloud generator 1100 vehicles 1102 Laser Package 1104 Laser Beam 1106 Pedestrians 1108 Bitstream 1110 Another vehicle 1112 Server System 1200 First User 1202 First Location 1204 XR Headset 1206 Object 1208 bitstream 1210 XR Headset 1212 users 1214 Second Location 1300 mobile devices 1302 Object 1304 Bitstream 1306 Remote Device< / bool> < / int> < / int> < / bool> < / int>
Claims
1. 1. A device for decoding a bitstream containing point cloud data, comprising: a memory for storing the point cloud data; one or more processors implemented in circuitry coupled to the memory, the one or more processors: determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determining for the node based on the syntax signaled in the bitstream that an angular mode is activated for the node; determining, for the node, a quantized value representing a coordinate position in response to enabling intra-tree quantization for the node and activating the angular mode for the node; scaling the quantized values without clipping to determine scaled values representing the coordinate positions; and configured to determine a context for context decoding a planar position syntax element for the angle mode based on the scaled value representing the coordinate position. device.
2. determining the context for context decoding the planar position syntax element for the angular mode based on the scaled value representing the coordinate position; determining one or more laser characteristics based on the scaled values and the origin position; configured to decode the plane position syntax element for the angular mode based on the laser characteristic.
10. The device of claim 1.
3. 3. The device of claim 2, wherein to determine the context for context decoding the planar position syntax element for the angle mode based on the scaled value representing the coordinate position, the one or more processors are further configured to determine a context index based on whether a laser beam having the determined one or more laser characteristics is above a first distance threshold, between the first distance threshold and a second distance threshold, between the second distance threshold and a third distance threshold, or below the third distance threshold.
4. The device of claim 2 , wherein the one or more laser characteristics include laser elevation, laser head offset, or azimuth.
5. The device of claim 1 , wherein the planar position syntax element indicates a vertical plane position.
6. 10. The device of claim 1, wherein the one or more processors are further configured to arithmetically decode the planar position syntax element for the angular mode using a context indicated by the determined context.
7. the one or more processors scaling the quantized values without clipping to determine the scaled values representing the coordinate positions; determining a group of most significant bits (MSBs) and a group of least significant bits (LSBs); scaling the LSBs without clipping to determine scaled LSBs; configured to add the scaled LSB to the MSB to determine the scaled value representing the coordinate position.
10. The device of claim 1.
8. the one or more processors scaling the quantized values without clipping to determine the scaled values representing the coordinate positions; determining a shift amount for an MSB based on a quantization parameter for the node; shifting the MSB based on the shift amount; configured to add the scaled LSB to the shifted MSB to determine the scaled value representing the coordinate position.
10. The device of claim 1.
9. 8. The device of claim 7, wherein the one or more processors are further configured to receive in the syntax signaled in the bitstream an indication of the number of bits in the group of LSBs.
10. The device of claim 1 , wherein the one or more processors are further configured to reconstruct a point cloud from the point cloud data.
11. the one or more processors are configured to determine, as part of reconstructing the point cloud, a position of one or more points of the point cloud based on the plane position syntax element; 11. The device of claim 10, wherein the one or more processors are further configured to: generate a map of an interior of a building based on the point cloud; perform autonomous navigation operations based on the point cloud; and / or generate computer graphics based on the point cloud.
12. 1. A method for decoding a bitstream containing point cloud data, comprising: determining, based on syntax signaled in the bitstream, that in-tree quantization is enabled for a node; determining for the node based on the syntax signaled in the bitstream that an angular mode is activated for the node; determining, for the node, a quantized value representing a coordinate position in response to enabling intra-tree quantization for the node and activating the angular mode for the node; scaling the quantized values without clipping to determine scaled values representing the coordinate positions; determining a context for context decoding a planar position syntax element for the angular mode based on the scaled value representing the coordinate position; A method for providing the above.
13. 1. A device for encoding a bitstream comprising point cloud data, comprising: a memory for storing the point cloud data; one or more processors implemented in circuitry coupled to the memory, the one or more processors: determining that in-tree quantization is enabled for the node; determining that an angular mode is activated for the node; determining, for the node, a quantized value representing a coordinate position in response to enabling intra-tree quantization for the node and activating the angular mode for the node; scaling the quantized values without clipping to determine scaled values representing the coordinate positions; configured to determine a context for context encoding a planar position syntax element for the angular mode based on the scaled value representing the coordinate position. device.
14. 1. A method for encoding a bitstream containing point cloud data, comprising: determining that in-tree quantization is enabled for a node; determining that an angular mode is activated for the node; determining, for the node, a quantized value representing a coordinate position in response to enabling intra-tree quantization for the node and activating the angular mode for the node; scaling the quantized values without clipping to determine scaled values representing the coordinate positions; determining a context for context encoding a planar position syntax element for the angular mode based on the scaled value representing the coordinate position; A method for providing the above.
15. 15. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 12 or 14. A computer-readable storage medium.
Citation Information
Patent Citations
System and method for ordered representation and feature extraction for point clouds obtained by detection and ranging sensor
US20200302237A1
Method and apparatus for point cloud compression
US20200394822A1
Angular mode syntax for tree-based point cloud coding
US20220351423A1
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020101021A1
Angular priors for improved prediction in tree-based point cloud coding
WO2021084295A1