High level syntax for geometry-based point cloud compression
By optimizing the high-level syntax of G-PCC and utilizing the interaction relationships of syntax elements, the signaling notification overhead is reduced, thus solving the problem of high signaling notification overhead in the G-PCC standard and improving encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-08
- Publication Date
- 2026-03-20
AI Technical Summary
The existing geometric point cloud compression (G-PCC) standard results in high signaling overhead and may include unnecessary bit signaling notifications, affecting encoding and decoding efficiency.
By improving the high-level syntax of G-PCC and leveraging the interaction relationships between syntax elements, redundant signaling notifications can be reduced. For example, the values of certain syntax elements can be inferred rather than directly signaled, and other syntax elements can be represented using fewer bits, thus optimizing the signaling notification overhead.
It reduces signaling overhead during G-PCC encoding and decoding, lowers processing power consumption, and improves encoding and decoding efficiency.
Smart Images

Figure CN114930858B_ABST
Abstract
Description
[0001] This patent application claims priority to U.S. Patent Application Serial No. 17 / 143,607, filed January 7, 2021, which claims benefit of U.S. Provisional Patent Application Serial No. 62 / 958,399, filed January 8, 2020, U.S. Provisional Patent Application Serial No. 62 / 960,472, filed January 13, 2020, and U.S. Provisional Patent Application Serial No. 62 / 968,578, filed January 31, 2020, the entire contents of each of which are incorporated herein by reference. TECHNICAL FIELD
[0002] This disclosure relates to point cloud encoding and decoding. SUMMARY
[0003] In general, this disclosure describes techniques for point cloud encoding and decoding, including techniques related to geometry-based point cloud compression (G-PCC). The details of one or more examples are set forth in the accompanying drawings and the description below. One example draft of the G-PCC standard can result in point cloud coding techniques that include unnecessary signaling and / or signaling that contains more bits than can be necessary. This disclosure includes techniques for reducing the signaling overhead associated with G-PCC. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
[0004] In one example, a method of coding a point cloud includes determining a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scale factor for the point cloud, determining a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scale factor for the point cloud, and processing the point cloud based at least in part on the scale factor for the point cloud.
[0005] In another example, an apparatus for coding a point cloud includes a memory configured to store the point cloud, and one or more processors communicatively coupled with the memory, the one or more processors configured to determine a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scale factor for the point cloud, determine a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scale factor for the point cloud, and process the point cloud based at least in part on the scale factor for the point cloud.
[0006] In another example, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to determine a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scale factor for the point cloud, determine a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scale factor for the point cloud, and process the point cloud based at least in part on the scale factor for the point cloud.
[0007] In another example, a device includes means for determining a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scaling factor for a point cloud, means for determining a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scaling factor for the point cloud, and means for processing the point cloud based at least in part on the scaling factor for the point cloud.
[0008] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 is a block diagram illustrating an example encoding and decoding system that can perform the techniques of this disclosure.
[0010] Figure 2 is a block diagram illustrating an example geometry point cloud compression (G-PCC) encoder.
[0011] Figure 3 is a block diagram illustrating an example G-PCC decoder.
[0012] Figure 4 is a conceptual diagram illustrating relationships between a sequence parameter set, a geometry parameter set, a geometry slice header, an attribute parameter set, and an attribute slice header.
[0013] Figure 5 is a flow diagram illustrating example signaling techniques in accordance with this disclosure. DETAILED DESCRIPTION
[0014] One example draft of the G-PCC standard can result in point cloud coding techniques that include unnecessary signaling and / or signaling that contains more bits than can be necessary. As such, signaling in accordance with the draft standard can result in poorer functionality and greater signaling overhead than can be possible with other ways. In accordance with the techniques of this disclosure, certain syntax elements of the draft G-PCC standard can be represented with fewer bits than set forth in the draft standard, while other certain syntax elements can not be signaled. By representing certain syntax elements with fewer bits and not signaling other certain syntax elements, a G-PCC encoder can reduce signaling overhead associated with the syntax elements, and thus can reduce processing power drain of the G-PCC encoder and / or a G-PCC decoder.
[0015] Figure 1is a block diagram illustrating an example encoding and decoding system 100 that can perform the techniques of this disclosure. The techniques of this disclosure generally relate to signaling associated with coding (encoding and / or decoding) point cloud data. Generally, point cloud data includes any data for processing a point cloud. Coding can efficiently compress and / or decompress the point cloud data.
[0016] As Figure 1 shown, system 100 includes a source device 102 and a destination device 116. Source device 102 provides encoded point cloud data to be decoded by destination device 116. Specifically, in the example of Figure 1 , source device 102 provides the point cloud data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can comprise any of a wide variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, land or sea vehicles, spacecraft, aircraft, robots, LIDAR devices, satellites, and the like. In some cases, source device 102 and destination device 116 can be equipped for wireless communication.
[0017] In the example of Figure 1 , source device 102 includes data source 104, memory 106, G-PCC encoder 200, and output interface 108. Destination device 116 includes input interface 122, G-PCC decoder 300, memory 120, and data consumer 118. In accordance with this disclosure, G-PCC encoder 200 of source device 102 and G-PCC decoder 300 of destination device 116 can be configured to apply the techniques of this disclosure related to high-level syntax for geometry-based point cloud compression. Thus, source device 102 represents an example of an encoding device, while destination device 116 represents an example of a decoding device. In other examples, source device 102 and destination device 116 can include other components or arrangements. For example, source device 102 can receive data (e.g., point cloud data) from an internal or external source. Likewise, destination device 116 can interface with an external data consumer rather than include a data consumer in the same device.
[0018] As Figure 1The illustrated system 100 is merely one example. In general, other digital encoding and / or decoding devices can perform the techniques of this disclosure related to high-level syntax for geometry point cloud compression. Source device 102 and destination device 116 are merely examples of such devices that can generate encoded data for transmission to destination device 116. This disclosure refers to "coding" devices as devices that perform encoding (encoding and / or decoding) of data. Thus, G-PCC encoder 200 and G-PCC decoder 300 are representative of coding devices, specifically examples of encoders and decoders, respectively. In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes encoding and decoding components. Hence, system 100 can support one-way or two-way transmission between source device 102 and destination device 116, e.g., for streaming, playback, broadcast, telephony, navigation, and other applications.
[0019] In general, data source 104 represents a source of data (i.e., raw, unencoded point cloud data) and can provide successive series of "frames" of the data to G-PCC encoder 200, which encodes the data of the frames. Data source 104 of source device 102 can include a point cloud capture device such as any of a variety of cameras or sensors, e.g., a 3D scanner or a light detection and ranging (LIDAR) device, one or more video cameras, an archive containing previously captured data, and / or a data feed interface for receiving data from a data content provider. Alternatively or additionally, point cloud data can be computer-generated from scanner, camera, sensor, or other data. For example, data source 104 can generate computer graphics-based data as source data, or produce a combination of live data, archived data, and computer-generated data. In each case, G-PCC encoder 200 encodes the captured, pre-captured, or computer-generated data. G-PCC encoder 200 can rearrange the frames from the received order (sometimes referred to as "display order") into an encoding order for encoding. G-PCC encoder 200 can generate one or more bitstreams that include encoded data. Source device 102 can then output the encoded data via output interface 108 onto computer-readable medium 110 for reception and / or retrieval by an input interface 122 of, e.g., destination device 116.
[0020] The memories 106 of the source device 102 and 120 of the destination device 116 can be representative of general purpose memories. In some examples, the memories 106 and 120 can store raw data, e.g., raw data from the data source 104 and raw data from the G-PCC encoder 200, decoded data. Additionally or alternatively, the memories 106 and 120 can store software instructions executable by, e.g., the G-PCC encoder 200 and the G-PCC decoder 300, respectively. Although the memories 106 and 120 are shown separately from the G-PCC encoder 200 and the G-PCC decoder 300 in this example, it should be understood that the G-PCC encoder 200 and the G-PCC decoder 300 can also include internal memories for functionally similar or equivalent purposes. Moreover, the memories 106 and 120 can store encoded data, e.g., output from the G-PCC encoder 200 and input to the G-PCC decoder 300. In some examples, portions of the memories 106 and 120 can be allocated as one or more buffers, e.g., to store raw data, decoded data, and / or encoded data. For example, the memories 106 and 120 can store data representative of a point cloud.
[0021] The computer-readable medium 110 can represent any type of medium or device capable of transporting the encoded data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium to enable the source device 102 to transmit encoded data directly to the destination device 116 in real-time, e.g., via a radio frequency network or computer-based network. The output interface 108 can modulate a transmission signal including the encoded data, and the input interface 122 can demodulate the received transmission signal, according to a communication standard, such as a wireless communication protocol. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from the source device 102 to the destination device 116.
[0022] In some examples, the source device 102 can output encoded data from the output interface 108 to a storage device 112. Similarly, the destination device 116 can access encoded data from the storage device 112 via the input interface 122. The storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded data.
[0023] In some examples, source device 102 can output encoded data to a file server 114 or another intermediate storage device that can store encoded data generated by source device 102. Destination device 116 can access stored data from file server 114 via streaming or download. File server 114 can be any type of server device that is capable of storing encoded data and transmitting this encoded data to destination device 116. File server 114 can represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 can access encoded data from file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing encoded data stored on file server 114. File server 114 and input interface 122 can be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0024] Output interface 108 and input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 can be configured to transfer data such as encoded data according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to transfer data such as encoded data according to other wireless standards, such as IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee TM ), Bluetooth TM standards, and the like. In some examples, source device 102 and / or destination device 116 can include respective system on a chip (SoC) devices. For example, source device 102 can include an SoC device to perform the functionality attributed to G-PCC encoder 200 and / or output interface 108, while destination device 116 can include an SoC device to perform the functionality attributed to G-PCC decoder 300 and / or input interface 122.
[0025] The techniques of this disclosure can be applied in support of encoding and decoding any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices such as local or remote servers, geographic mapping, or other applications.
[0026] Input interface 122 of destination device 116 receives an encoded bitstream from computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded bitstream can include signaling information defined by G-PCC encoder 200 that is also used by G-PCC decoder 300, such as syntax elements having values that describe characteristics and / or processing of coded units (e.g., slices, pictures, groups of pictures, sequences, or the like). The signaling information can include syntax elements as defined according to the techniques of this disclosure, or can exclude certain syntax elements in certain conditions as set out in this disclosure. Data consumer 118 uses the decoded data. For example, data consumer 118 can use the decoded data to determine a position of a physical object. In some examples, data consumer 118 can include a display device to render images based on the point cloud.
[0027] G-PCC encoder 200 and G-PCC decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more processors, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware, or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of G-PCC encoder 200 and G-PCC decoder 300 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device. A device including G-PCC encoder 200 and / or G-PCC decoder 300 can comprise one or more integrated circuits (ICs), microprocessors, and / or other types of devices.
[0028] G-PCC encoder 200 and G-PCC decoder 300 can operate according to a coding standard, such as the Geometric Point Cloud Compression (G-PCC) standard, for video point cloud compression (V-PCC). This disclosure can generally relate to coding (e.g., encoding and decoding) of pictures, to include processing of data for encoding or decoding. An encoded bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes).
[0029] This disclosure can generally relate to "signaling" certain information, such as syntax elements. The term "signaling" can generally relate to a communication regarding values for syntax elements and / or other data used to decode encoded data. That is, the G-PCC encoder 200 can signal values for syntax elements in a bitstream. These syntax elements can include syntax elements as defined in this disclosure. In general, signaling involves generating values in a bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 in substantially real-time or non-real-time, such as can occur when syntax elements are stored to a storage device 112 for later retrieval by the destination device 116.
[0030] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) is studying the potential need for standardization of point cloud coding technology with compression capability significantly exceeding other approaches, and is working to create the standard. The group is studying this exploratory activity together in a collaborative task called the Three-Dimensional Graphics Team (3DG) to evaluate compression technology designs proposed by experts in this field.
[0031] Point cloud compression activities are categorized into two different approaches. The first approach is "Video Point Cloud Compression" (V-PCC), which segments 3D objects and projects these segments into multiple 2D planes (represented as "patches" in 2D frames), which are further coded by a conventional 2D video codec such as a High Efficiency Video Coding (HEVC) (ITU-T H.265) codec. The second approach is "Geometry-based Point Cloud Compression" (G-PCC), which directly compresses 3D geometry, i.e., positions of a set of points in 3D space, and associated attribute values (for each point associated with the 3D geometry). G-PCC addresses compression of point clouds in both Category 1 (static point clouds) and Category 3 (dynamically acquired point clouds).
[0032] Point clouds contain a set of points in 3D space and can have attributes associated with the points. The attributes can be color information such as R, G, B, or Y, Cb, Cr, or reflectance information, or other attributes. Point clouds can be captured by various capturing devices such as cameras or sensors such as LIDAR sensors, and 3D scanners, and can also be computer generated. Point cloud data is used in various applications including, but not limited to, construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to help with navigation).
[0033] The 3D space occupied by the point cloud data can be enclosed by a virtual bounding box. The positions of points in the bounding box can be represented by a certain precision; from which, the positions of one or more points can be quantized based on the precision. At the minimum level, the bounding box is divided into voxels, which are the smallest unit of space represented by a unit cube. Voxels in the bounding box can be associated with zero, one, or more than one point. The bounding box can be divided into multiple cubic / tessellated regions, which can be referred to as tiles. Each tile can be coded into one or more slices. The partitioning of the bounding box into slices and tiles can be based on the number of points in each partition, or based on other considerations (e.g., certain regions can be coded into tiles). Slice regions can be further partitioned by using similar partitioning decisions as in video codecs.
[0034] Figure 2 An overview of the G-PCC encoder 200 is provided. Figure 2 An overview of the G-PCC decoder 300 is provided. The illustrated modules are logical and do not necessarily one-to-one correspond to implementation code in a reference implementation of a G-PCC codec, i.e., the TMC13 test model software studied by ISO / IEC MPEG (JTC 1 / SC 29 / WG 11).
[0035] In both the G-PCC encoder 200 and the G-PCC decoder 300, point cloud positions are first coded. Attribute coding depends on the decoded geometry. In the G-PCC encoder 200, the decoded geometry is used to determine the attribute coding. Figure 2 and Figure 3 In the G-PCC encoder 200 and the G-PCC decoder 300, point cloud positions are first coded. Attribute coding depends on the decoded geometry. In the G-PCC encoder 200, the decoded geometry is used to determine the attribute coding.
[0036] For Category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level of individual voxels. For Category 1 data, the compressed geometry is typically represented as a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model of the surface within each leaf of the pruned octree. In this way, both Category 1 and Category 3 share the octree coding mechanism, while Category 1 can additionally utilize a surface model to approximate the voxels within each leaf. The surface model used is a triangulation that includes 1-10 triangles per block, resulting in a triangle soup. The Category 1 geometry coder is thus referred to as the Trisoup geometry coder, while the Category 3 geometry coder is referred to as the Octree geometry coder.
[0037] At each node of the octree, the occupancy (when not inferred) is signaled for one or more of its children (up to eight nodes). Multiple neighborhoods are specified, including (a) nodes that share a face with the current octree node, (b) nodes that share a face, edge, or vertex with the current octree node, etc. Within each neighborhood, the occupancy of the node and / or its children can be used to predict the occupancy of the current node or its children. For points that are sparsely distributed in certain nodes of the octree, the coder also supports a direct coding mode in which the 3D position of the point is directly coded. A flag can be signaled to indicate that the direct mode is signaled. At the lowest level, the number of points associated with an octree node / leaf node can also be coded.
[0038] Once the geometry is coded, the attributes corresponding to the geometry points are coded. When there are multiple attribute points corresponding to one reconstructed / decoded geometry point, the attribute values representing the reconstructed point can be derived.
[0039] There are three attribute coding methods in G-PCC: Region Adaptive Hierarchical Transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (predictive transform), and interpolation-based hierarchical nearest neighbor prediction with an update / boost step (boosting transform). RAHT and boosting are typically used for Category 1 data, while prediction is typically used for Category 3 data. However, any method can be used for any data, and, like with the geometry coders in G-PCC, the user (e.g., G-PCC encoder 200) has the option for choosing which of the 3 attribute coders to use.
[0040] The coding of attributes can be done with a level of detail (LOD), in which a finer representation of the point cloud attributes can be obtained with each level of detail. Each level of detail can be specified based on a distance metric to neighboring nodes or based on a sampling distance.
[0041] At the G-PCC encoder 200, the residual obtained as output of the coding method for attributes is quantized. The quantized residual can be coded using context adaptive arithmetic coding.
[0042] In examples of the G-PCC encoder 200, the G-PCC encoder 200 can include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit (GRU) 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226. Figure 2
[0043] As shown in examples of the G-PCC encoder 200, the G-PCC encoder 200 can accept a set of positions and a set of attributes. The positions can include coordinates of points in a point cloud. The attributes can include information about points in the point cloud, such as color associated with points in the point cloud. Figure 2 The coordinate transformation unit 202 can apply a transformation to the coordinates of the points to transform the coordinates from an original domain to a transformed domain. The disclosure can refer to the transformed coordinates as transformed coordinates. The color transformation unit 204 can apply a transformation for transforming color information of the attributes to a different domain. For example, the color transformation unit 204 can transform color information from an RGB color space to a YCbCr color space.
[0044] In addition, in examples of the G-PCC encoder 200, the voxelization unit 206 can voxelize the transformed coordinates. The voxelization of the transformed coordinates can include quantization and removal of certain points of the point cloud. In other words, multiple points of the point cloud can be grouped into a single "voxel," which can then be treated as a point in some aspects. In addition, the octree analysis unit 210 can generate an octree based on the voxelized transformed coordinates. Additionally, in examples of the G-PCC encoder 200, the surface approximation analysis unit 212 can analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 can entropy encode syntax elements representing information of the octree and / or the surface determined by the surface approximation analysis unit 212. The G-PCC 200 can output these syntax elements in a geometry bitstream.
[0045] Figure 2 Figure 2
[0046] The geometry reconstruction unit 216 can reconstruct the transformed coordinates of the points in the point cloud based on an octree, data indicative of the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometry reconstruction unit 216 can be different from the original number of points of the point cloud. The disclosure can refer to the resulting points as reconstructed points. The attribute transfer unit 208 can transfer attributes of the original points of the point cloud to the reconstructed points of the point cloud.
[0047] Further, the RAHT unit 218 can apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 can apply LOD processing and lifting, respectively, to the attributes of the reconstructed points. The RAHT unit 218 and the lifting unit 222 can generate coefficients based on the attributes. The coefficient quantization unit 224 can quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 can apply arithmetic coding to syntax elements representing the quantized coefficients. The G-PCC encoder 200 can output these syntax elements in an attribute bitstream.
[0048] In the example of FIG. 3, the G-PCC decoder 300 can include a geometry arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometry reconstruction unit 312, a RAHT unit 314, a LoD generation unit 316, an inverse lifting unit 318, an inverse transformed coordinate unit 320, and an inverse transformed color unit 322. Figure 3
[0049] The G-PCC decoder 300 can obtain a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 302 of the G-PCC decoder 300 can apply arithmetic decoding (e.g., context adaptive binary arithmetic coding (CABAC) or other types of arithmetic decoding) to syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 304 can apply arithmetic decoding to syntax elements in the attribute bitstream.
[0050] The octree synthesis unit 306 can synthesize an octree based on syntax elements parsed from the geometry bitstream. In instances in which a surface approximation is used for the geometry bitstream, the surface approximation synthesis unit 310 can determine a surface model based on syntax elements parsed from the geometry bitstream and based on the octree.
[0051] Further, the geometry reconstruction unit 312 can perform reconstruction to determine coordinates of the points in the point cloud. The inverse transformed coordinate unit 320 can apply an inverse transform to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0052] Additionally, in Figure 3 In the example, the inverse quantization unit 308 can inverse quantize the attribute value. The attribute value can be based on syntax elements obtained from the attribute bitstream (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0053] Depending on how the attribute values are encoded, RAHT unit 314 can perform RAHT to determine the color values of points in the point cloud based on the inversely quantized attribute values. Alternatively, LoD generation unit 316 and inverse boosting unit 318 can use level-of-detail (LOD) based techniques to determine the color values of points in the point cloud.
[0054] In addition, Figure 3 In the example, the inverse color transformation unit 322 can apply an inverse color transformation to color values. The inverse color transformation can be the inverse of the color transformation applied by the color transformation unit 204 of the G-PCC encoder 200. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space. Correspondingly, the inverse color transformation unit 322 can transform color information from the YCbCr color space to the RGB color space.
[0055] Figure 2 and Figure 3 Various units are shown to aid in understanding the operations performed by the G-PCC encoder 200 and the G-PCC decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is pre-programmed for the operations it can perform. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations it can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of that software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), while in other examples, one or more of the units may be integrated circuits.
[0056] The document w18887 mentioned above is a G-PCC draft document. This draft of the G-PCC standard can include non-essential signaling and / or signaling that contains more bits than can be necessary. The representation of the high level syntax (HLS) of G-PCC can be improved to achieve better functionality and reduce the associated signaling cost of different HLS related parameters. According to the techniques of this disclosure, certain syntax elements of document w18887 can be represented with fewer bits than prohibited in the draft standard, while other certain syntax elements can not be signaled. By representing certain syntax elements with fewer bits and not signaling other certain syntax elements, the G-PCC encoder 200 can reduce the signaling overhead associated with the syntax elements, and thus can reduce the processing power drain of the G-PCC encoder 200 and / or the G-PCC decoder 300.
[0057] This disclosure discusses various improvements to the high level syntax of G-PCC that can be made with respect to document w18887. The various techniques set forth herein can be applied independently, or one or more techniques can be applied in any combination.
[0058] The interaction of three particular syntax elements is now discussed. The geometry parameter set (GPS) syntax in document w18887 includes these three syntax elements: log2_trisoup_node_size (which indicates the size of the triangle node); inferred_direct_coding_mode_enabled_flag (which indicates whether direct_mode_flag can be present in the geometry node syntax); and unique_geometry_points_flag (which indicates whether all output points have unique positions within a given slice in all slices that refer to the current GPS). While the G-PCC encoder 200 can signal these syntax elements (e.g., log2_trisoup_node_size, inferred_direct_coding_mode_enabled_flag, and unique_geometry_points_flag) independently, they have inherent interaction with each other, which is evident in their semantics, as follows. Throughout this document, <e> …< / e> The text between the tags represents the emphasized portion of the original G-PCC syntax description of document w18887. Furthermore, throughout this disclosure, <m> …< / m> The text between the tags represents the modified text of the original G-PCC syntax description in document w18887 due to the techniques of this disclosure. <d> …< / d> The text between the tags represents the text deleted from the original G-PCC syntax description in document w18887 due to the techniques of this disclosure.
[0059]
[0060] log2_trisoup_node_size specifies the variable TrisoupNodeSize as the size of a triangle node as follows.
[0061] TrisoupNodeSize = 1 « log2_trisoup_node_size
[0062] When log2_trisoup_node_size is equal to 0, the geometry bitstream only includes octree coding syntax. <m>When log2_trisoup_node_size is greater than 0, the following are the bitstream conformance requirements:
[0063] - inferred_direct_coding_mode_enabled_flag must be equal to 0, and
[0064] - unique_geometry_points_flag must be equal to 1.< / m>
[0065] Accordingly, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine whether a value of a trisoup syntax element indicating a triangle node size is greater than 0, based on the value of the trisoup syntax element being greater than 0: infer a value of an inferred direct coding mode enabling syntax element indicating whether a direct mode syntax element is present in the bitstream is 0; and infer a value of a unique geometry point syntax element indicating whether all output points in all slices referring to a current geometry parameter set have unique positions within the respective slices is 1. Coding the point cloud can be based, at least in part (and in some examples, further) on the value of the trisoup syntax element. In this way, G-PCC encoder 200 can reduce signaling overhead when compared to the signaling described in document w18887.
[0066] In some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine that a bitstream is inconsistent with a coding standard based on a first syntax element not being equal to 0 or a second syntax element not being equal to 1, where: the first syntax element equal to 1 indicates whether a third syntax element can be present in a geometry node syntax, the second syntax element equal to 1 indicates that all output points in all slices referring to a current geometry parameter set have unique positions within the slices, and the third syntax element indicates whether a single child node of a current node is a leaf node and contains one or more delta point coordinates.
[0067] From the semantics of document w18887, it can be inferred that when log2_trisoup_node_size is greater than 0, the G-PCC decoder 300 can infer inferred_direct_coding_mode_enabled_flag and unique_geometry_points_flag without parsing the inferred_direct_coding_mode_enabled_flag and unique_geometry_points_flag syntax elements in the received bitstream. Therefore, according to one example of the disclosure, the G-PCC encoder 200 and the G-PCC decoder 300 can operate according to the following modification that removes redundant signaling and bypasses the conformance check when trisoup is enabled.
[0068]
[0069]
[0070] unique_geometry_points_flag equal to 1 indicates that in all slices referring to the current GPS, all output points have unique positions within the slice. unique_geometry_points_flag equal to 0 indicates that in all slices referring to the current GPS, two or more of the output points can have the same position within the slice. <m>When unique geometry points flag is not present in the bitstream, it is inferred to be equal to 1.< / m>
[0071] inferred_direct_coding_mode_enabled_flag equal to 1 indicates that direct_mode_flag can be present in the geometry node syntax. inferred_direct_coding_mode_enabled_flag equal to 0 indicates that direct_mode_flag is not present in the geometry node syntax. <m>When inferred direct coding mode enabled flag is not present in the bitstream, it is inferred to be equal to 0.< / m>
[0072] log2_trisoup_node_size specifies the variable TrisoupNodeSize as the size of a triangle node as follows.
[0073] TrisoupNodeSize = 1 « log2_trisoup_node_size
[0074] When log2_trisoup_node_size is equal to 0, the geometry bitstream includes only octree coding syntax. <d>When log2_trisoup_node_size is greater than 0, the following is the bitstream conformance requirement:
[0075] - inferred_direct_coding_mode_enabled_flag must be 0, and
[0076] - unique_geometry_points_flag must be equal to 1.< / d>
[0077] direct mode flag equal to 1 indicates that the single child node of the current node is a leaf node and contains one or more delta point coordinates. direct mode flag equal to 0 indicates that the single child node of the current node is an internal octree node. When not present, the value of direct mode flag is inferred to be 0.
[0078] Signaling of lifting num pred nearest neighbours (which indicates the maximum number of nearest neighbors used for prediction) and other similar syntax elements is now discussed. According to document w18887, a G-PCC encoder such as G-PCC encoder 200 can signal lifting num pred nearest neighbours in the attribute parameter set that specifies the maximum nearest neighbors used for prediction (e.g., for prediction transforms and lifting transforms) with a minimum value of 1.
[0079]
[0080] lifting num pred nearest neighbours specifies the maximum number of nearest neighbors to be used for prediction. The value of lifting num pred nearest neighbours will be in the range of 1 to <m>xx< / m> .
[0081] Since the minimum value of this syntax is 1, a G-PCC encoder can instead signal lifting num pred nearest neighbours minusl. The following shows the mentioned syntax and semantics.
[0082]
[0083] lifting num pred nearest neighbours <m>_minus1 plus 1< / m> specifies the maximum number of nearest neighbors to be used for prediction. The value of lifting num pred nearest neighbours will be in the range of 1 to <m>xx< / m> .
[0084] Accordingly, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a lifting syntax element, where a value of the lifting syntax element plus one specifies a maximum number of nearest neighbors to be used for prediction. The device can code the point cloud based at least in part (and in some examples, further) on the prediction. In this way, G-PCC encoder 200 can reduce a number of bits used to signal the lifting syntax element when compared to the signaling described in document w18887, thereby reducing signaling overhead.
[0085] In some examples, a device can determine a geometry slice header syntax element, where a value of the geometry slice header syntax element plus one specifies a number of points in a geometry slice. The device can determine an attribute bit depth syntax element, where a value of the attribute bit depth syntax element plus one specifies a bit depth of an attribute. The device can also determine a unique segment number syntax element, where a value of the unique segment number syntax element plus one specifies a number of unique segments.
[0086] In some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can code a syntax element (e.g., lifting num pred nearest neighbours minusl) where the syntax element plus one specifies a maximum number of nearest neighbors to be used for prediction. The coder can code the point cloud using the syntax element.
[0087] Similarly, coding of the following syntax elements is modified to zero value according to the techniques of this disclosure, as these syntax elements can be undesirable:
[0088] 1. gsh num points (which indicates a number of coded points in a point cloud in a slice) can be coded as gsh num points minusl.
[0089] 2. sps num attributes (which indicates a number of attributes in a sequence) can be coded as sps num attribute sets minusl (this change is not appropriate and can be ignored if G-PCC will support attribute-less point clouds).
[0090] 3. attribute dimension [] (which indicates an attribute associated with a point cloud) can be coded as attribute dimension minusl [].
[0091] 4. attribute_bitdepth[] (which indicates the attribute bit depth associated with the point cloud) can be coded as attribute_bitdepth_minusl[] or attribute_bitdepth_minusN[], where N is expected to be the smallest attribute bit depth supported. For example, the value of N can be set to 8.
[0092] 5. num_unique_segments (which indicates the number of unique segments in the trisoup mode) can be coded as num_unique_segments_minusl.
[0093] 6. num_vertices (which indicates the number of vertices in the trisoup mode) can be coded as num_vertices_minusl.
[0094]
[0095]
[0096] As mentioned above, attribute_bitdepth_minusl can be used instead of attribute_bitdepth_minus8 as shown in the table above.
[0097]
[0098] Similar to lifting_num_pred_nearest_neighbours, the semantics of the above syntax elements can be updated as appropriate.
[0099] In some scenarios, instead of G-PCC encoder 200 directly signaling the value of num_unique_segments, G-PCC encoder 200 can first scale down the value of num_unique_segments using a ceiling function that limits the scaled down value of num_unique_segments to a maximum value of 255:
[0100] num_unique_segments_downscaled = Ceil(num_unique_segments / K), where K is a positive integer, which can be fixed or variable. G-PCC encoder 200 can signal num_unique_segments_downscaled. If K is variable, G-PCC encoder 200 can also signal the value of K. For example, in G-PCC reference software, a fixed value of K = 8 is used. This methodology reduces the associated signaling cost, considering that num_unique_segments can have very high values, and variable length coding, such as exponential-Golomb coding, can lead to too much signaling. For example, if num_unique_segments = 10002, and K = 8, then num_unique_segments_dowsampled = Ceil(10002 / 8) = 1251. Fewer bits can be needed to signal "1251" (e.g., num_unique_segments_downscaled) compared to "10002" (e.g., num_unique_segments). Thus, in some examples, G-PCC encoder can signal num_unique_segments_downscaled (e.g., 1251) and G-PCC decoder can parse the signaled num_unique_segments_downscaled (e.g., 1251) and compute num_unique_segments (e.g., 10002) by multiplying num_unique_segments_downscaled (e.g., 1251) by K (e.g., 8). In this scenario, num_unique_segments_downscaled can always have a minimum value of 1, so num_unique_segments_downscaled_minusl can be signaled instead. In the above example, G-PCC encoder 200 can signal 1250 (instead of 1251).
[0101] Signaling of lifting_sampling_period[] (which indicates the sampling period for level of detail idx) is now discussed. For level of detail (LoD) generation, two ways of generation are outlined in document w18887: 1) based on distance, and 2) based on regular sampling. For regular sampling based generation, G-PCC encoder 200 can signal the sampling factor for each LoD. The following shows the syntax and semantics.
[0102]
[0103]
[0104] lifting_sampling_period[idx] specifies the sampling period for detail level idx. The value of lifting_sampling_period[] will be in the range of 0 to <m>xx< / m>
[0105] However, by definition, LoD idx is generated by subsampling LoD(idx+1), e.g. one point is excluded per lifting_sampling_period[idx]. Therefore, it can not be possible for lifting_sampling_period[idx] to be equal to 0, and even the value 1 does not indicate subsampling. Therefore, lifting_sampling_period should have a minimum value of 2, and thus, G-PCC encoder 200 can instead signal lifting_sampling_period_minus2[idx]. The following shows the mentioned syntax and semantics.
[0106]
[0107] lifting_sampling_period <m>_minus2[idx] plus 2< / m> specifies the sampling period for detail level idx. The value of lifting_sampling_period[idx] will be in the range of <m>2 to xx< / m>
[0108] Thus, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a value of a sampling period syntax element, the value of the sampling period syntax element indicating a sampling period for a detail level index. The device can code a point cloud based at least in part (in some examples, further) on the sampling period. In some examples, the value of the sampling period syntax element plus 2 specifies the sampling period for the detail level index. In this way, G-PCC encoder 200 can reduce the number of bits used to signal the sampling period syntax element when compared to the signaling described in document w18887, thereby reducing signaling overhead.
[0109] In some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can code a syntax element, wherein the syntax element plus 2 specifies a sampling period for a detail level index; and code a point cloud using the syntax element.
[0110] Also, in some cases, a constant sampling factor for all LoD layers can be sufficient. For example, according to document w18887, TMC13 setup uses a sampling period of 4 for all layers. Thus, according to the techniques of this disclosure, G-PCC encoder 200 can signal a flag that indicates whether the sampling factor is the same for all layers. The corresponding syntax and semantics are as follows.
[0111]
[0112]
[0113] <m>lifting_lod_constant_sampling_period_flag equal to 1 specifies that the sampling period is the same for all levels of detail.
[0114] lifting_lod_constant_sampling_period_flag equal to 0 indicates that the sampling period is different for different LoD levels. When lifting_lod_constant_sampling_period_flag is not present in the bitstream, it is inferred to be 0.< / m>
[0115] <m>lifting_constant_sampling_period_minus2 plus 2 specifies the constant sampling period of the regular sampling strategy that will be used for LoD generation when lifting lod constant sampling period flag is equal to 1.< / m>
[0116] lifting_sampling_period_minus2[ idx ] plus 2 specifies the sampling period for detail level idx. The value of lifting_sampling_period_minus2[ idx ] will be in the range of 0 to xx. <m>When lifting_sampling_period_minus2[idx] is not present in the bitstream, it is inferred to be equal to lifting_constant_sampling_period_minus2.< / m>
[0117] Thus, a device such as G-PCC encoder 200 or G-PCC decoder 300 can code a first syntax element (e.g., lifting_lod_constant_sampling_period_flag), where the first syntax element equal to 1 specifies that the sampling period is the same for all detail levels (LoD), where the first syntax element equal to 0 specifies that the sampling period is different for different LoD layers, and code a point cloud using the first syntax element. Further, G-PCC encoder 200 or G-PCC decoder 300 can code a second syntax element (e.g., lifting_constant_sampling_period_minus2), where the second syntax element plus 2 specifies a constant sampling period for a regular sampling strategy that will be used for detail level (LoD) generation when the first syntax element is equal to 1. G-PCC encoder 200 or G-PCC decoder 300 can determine a value of a third syntax (e.g., lifting_sample_period_minus2) element plus 2 specifies the sampling period for a LoD index, where when the third syntax element is not present in the bitstream, the decoder infers the value of the third syntax element to be equal to the value of the second syntax element.
[0118] Signaling of lifting_sampling_distance_squared[] (which indicates the square of the sampling distance for detail level idx) is now discussed. For distance-based LoD generation, G-PCC encoder 200 can signal the distance for each LoD layer in the bitstream with the corresponding syntax and semantics shown as follows.
[0119]
[0120] However, these distances often grow exponentially with the LoD level, so according to the techniques of this disclosure, G-PCC encoder 200 and / or G-PCC decoder 300 can determine:
[0121] lifting_sampling_distance_squared[idx] = (lifting_sampling_distance_squared_scale_minusl[idx - 1] + 1) * lifting_sampling_distance_squared[idx - 1] + lifting_sampling_distance_squared_offset[idx - 1]
[0122] Accordingly, according to the techniques of this disclosure, G-PCC encoder 200 can signal lifting_sampling_distance_squared_scale_minusl[idx - 1] and lifting_sampling_distance_squared_offset[idx - 1] as shown in the syntax and semantics mentioned below.
[0123]
[0124]
[0125] <d>lifting_sampling_distance_squared[idx] specifies the square of the sampling distance for detail level idx. The value of lifting_sampling_distance_squared[] will be in the range of 0 to xx.< / d>
[0126] <m>lifting_sampling_distance_squared_scale_minus1[idx] plus 1 specifies the scale factor used to derive the square of the sampling distance for detail level idx. The value of lifting_sampling_distance_squared_scale_minus1[idx] will be in the range of 0 to xx. When lifting_sampling_distance_squared_scale_minus1[idx] is not present in the bitstream, it is inferred to be equal to 0.< / m>
[0127] <m>lifting_sampling_distance_squared_offset[idx] specifies the offset used to derive the square of the sampling distance for detail level idx. The value of lifting_sampling_distance_squared_offset[idx] will be in the range of 0 to xx. When lifting_sampling_distance_squared_offset[idx] is not present in the bitstream, it is inferred to be equal to 0.< / m>
[0128] <m>For idx = 0... num_detail_level_minusl, inclusive num_detail_level_minusl, the sampling distance for detail level idx, lifting_sampling_distance_squared[idx] is derived as follows:
[0129]
[0130] Accordingly, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a value of a first syntax element (e.g., lifting_sampling_distance_squared_scale_minusl), where the value of the first syntax element plus one specifies a scale factor for deriving a squared sampling distance for a level of detail (LoD) index; determine a value of a second syntax element (e.g., lifting_sampling_distance_squared_offset), where the value of the second syntax element specifies an offset for deriving the squared sampling distance for the LoD index; determine the sampling distance for the LoD index based on the value of the first syntax element and the value of the second syntax element; and code a point cloud based on the sampling distance for the LoD index.
[0131] Signaling of sps_source_scale_factor, which indicates a scale factor of a source point cloud, is now discussed. The scale factor can be used to scale the point positions in physical dimensions before displaying the point cloud. According to document w18887, a G-PCC encoder such as G-PCC encoder 200 can signal sps_source_scale_factor as a floating-point number with 32-bit precision, e.g., u(32).
[0132]
[0133] sps_source_scale_factor indicates a scale factor of a source point cloud.
[0134] However, instead of a floating-point representation, a rational representation can be preferred, e.g., each of the numerator and the denominator can be coded as an integer. Accordingly, G-PCC encoder 200 or G-PCC decoder 300 can code the numerator and the denominator. However, for a valid rational representation (excluding 0 or undefined), neither the numerator nor the denominator can be zero. Accordingly, G-PCC encoder 200 can alternatively signal <m>sps_source_scale_factor_numerator_minus1< / m> and <m>sps_source_scale_factor_denominator_minus1< / m> , according to techniques of this disclosure.
[0135]
[0136] The advantage of this approach is that in scenarios where higher precision (of the fractional value) is needed, more bits can be spent on coding the numerator and the denominator. On the other hand, when the coarse precision is sufficient, less bits can be used. This approach is more scalable compared to the approach in document w18887, which uses a constant 32-bit representation regardless of the required precision.
[0137] Also, the concept of point cloud usability information (PCUI) can be useful, which can deal mainly with parameters that are not part of the decoding process, but can be useful for display purposes (e.g., similar to the video usability information (VUI) in HEVC or Versatile Video Coding (VVC) in the Joint Video Expert Team (JVET) of ITU-T SG 16 WP3 and ISO / IEC JTC 1 / SC 29 / WG 11), which can be signaled at the end of the SPS. In such scenarios, the above-mentioned syntax, e.g., sps_source_scale_factor or sps_source_scale_factor_numerator_minus1 and sps_source_scale_factor_denominator_minus1, can be moved to the PCUI.
[0138] Thus, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scale factor for a point cloud, determine a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scale factor for the point cloud, and code the point cloud based at least in part (in some examples, further) on the scale factor for the point cloud. In some examples, the value of the numerator syntax element plus one specifies the numerator of the scale factor for the point cloud. In some examples, the value of the denominator syntax element plus one specifies the denominator of the scale factor for the point cloud. In this way, G-PCC encoder 200 can reduce the number of bits used to signal the scale factor when compared to the signaling described in document w18887, thereby reducing signaling overhead.
[0139] In some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a value of a first syntax element, the value of the first syntax element indicating a numerator of a scale factor for a point cloud, determine a value of a second syntax element, the value of the second syntax element indicating a denominator of the scale factor for the point cloud, and code the point cloud using the scale factor for the point cloud.
[0140] Coding of various parameter set IDs will now be discussed. According to document w18887, a G-PCC encoder such as G-PCC encoder 200 can use variable length coding to signal various parameter set IDs such as SPS ID, GPS ID, and APS ID. However, G-PCC decoder 300 can need to parse these IDs to decode a slice.
[0141] Figure 4 is a conceptual diagram showing the relationship between sequence parameter set, geometry parameter set, geometry slice header, attribute parameter set, and attribute slice header. G-PCC decoder 300 can use geometry slice header (GSH) 402 to find the associated GPS for a geometry slice, e.g., GPS 404, by referring to the GPS ID. Using GPS 404, G-PCC decoder 300 can find the associated SPS, e.g., SPS 410, by referring to the SPS ID, as shown in the example of Figure 4 Similarly, G-PCC decoder 300 can use attribute slice header 406 to find the associated APS, e.g., APS 408, by referring to the APS ID. Using APS 408, G-PCC decoder 300 can find the associated SPS, SPS 410, by referring to the SPS ID.
[0142] However, due to the fact that each of SPS ID, GPS ID, and APS ID can have a value of 0 to 15 (inclusive) indicating a range that is a power of 2, and to ease parsing of such IDs, according to the techniques of this disclosure, G-PCC encoder 200 can use fixed length coding to encode these IDs, e.g., use u(4) (or 4 bits) for SPS ID, for GPS ID, and / or for APS ID. The syntax modification is shown below.
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149] Accordingly, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a parameter set identifier (ID), where the parameter set ID is coded using a fixed length. In some examples, the fixed length is 4 bits. In some examples, the parameter set ID identifies a parameter set. The parameter set can include one of a sequence parameter set, a geometry parameter set, or an attribute parameter set. In this way, G-PCC encoder 200 can reduce the number of bits used to signal the parameter set ID, thereby reducing signaling overhead when compared to the signaling described in document w18887.
[0150] In some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can use a fixed length value to indicate a syntax element, where the syntax element is one of a sequence parameter set (SPS) identifier, a geometry parameter set identifier, an attribute parameter set identifier, a geometry slice header identifier, or an attribute slice header identifier; and code a bitstream that includes an encoded representation of a point cloud, where the bitstream includes at least one of an SPS identified by the SPS identifier, a geometry parameter set identified by the geometry parameter set identifier, an attribute parameter set identified by the attribute parameter set identifier, a geometry slice header identified by the geometry slice header identifier, or an attribute slice header identified by the attribute slice header identifier.
[0151] Also, according to the techniques of this disclosure, G-PCC encoder 200 can signal sps_seq_parameter_set_id (which indicates an identifier for the SPS) at the very beginning of the SPS. Alternatively, G-PCC encoder 200 can signal sps_seq_parameter_set_id in the SPS at least before syntax elements that require variable length parsing (e.g., sps_bounding_box_present_flag (which indicates whether a bounding box is present in the sequence)). This can facilitate G-PCC decoder 300 to quickly parse sps_seq_parameter_set_id. Accordingly, one of four possible examples can be employed, as follows. Any of examples 1-4 can reduce the complexity of parsing sps_seq_parameter_set_id for G-PCC decoder 300.
[0152] Example 1:
[0153]
[0154] Example 2:
[0155]
[0156]
[0157] Example 3:
[0158]
[0159]
[0160] Example 4:
[0161]
[0162] Accordingly, in some examples, the parameter set (indicated by the parameter set ID) is a sequence parameter set, and the parameter set ID follows a syntax element indicating an upper bound level associated with the point cloud data (such as a number of allowed points in a point cloud frame or slice, a bounding box size, a maximum octree root node dimension, etc.), and precedes a syntax element indicating whether a bounding box is present in the sequence parameter set. In this way, the G-PCC encoder 200 can place the parameter set ID in a location in the sequence parameter set, which can reduce the decoding latency of the G-PCC decoder 300.
[0163] Now discuss coding syntax elements based on the number of attribute dimensions. Attributes can have different types, and each attribute can have one or more dimensions. For example, a color attribute can have three dimensions, while a frame_idx attribute can have only one dimension. The frame_idx attribute can indicate a frame number of a point cloud frame. However, certain syntax elements in the G-PCC syntax structure are signaled without considering the number of dimensions of the associated attribute, particularly those attributes associated with chroma components / dimensions: aps_attr_chroma_qp_offset, ash_attr_qp_delta_chroma, and ash_attr_layer_qp_delta_chroma[i]. According to the techniques of this disclosure, a G-PCC codec such as the G-PCC encoder 200 can signal these syntax elements with attribute_dimension[i] that specifies the number of dimensions associated with the ith attribute. The changes to the syntax and semantics in document w18887 are as follows:
[0164]
[0165]
[0166] ash_attr_qp_delta_luma specifies the luminance delta qp from the initial slice qp in the active attribute parameter set. When ash_attr_qp_delta_luma is not signaled, the value of ash_attr_qp_delta_luma is inferred to be 0.
[0167] ash_attr_qp_delta_chroma specifies the chroma delta qp from the initial slice qp in the active attribute parameter set. When ash_attr_qp_delta_chroma is not signaled, the value of ash_attr_qp_delta_chroma is inferred to be 0.
[0168] The variables InitialSliceQpY and InitialSliceQpC are derived as follows:
[0169] InitialSliceQpY = aps_attr attr_initial_qp + ash_attr_qp_delta_luma
[0170] InitialSliceQpC = aps_attr attr_initial_qp + aps_attr chroma_qp_offset + ash_attr_qp_delta_chroma
[0171] ash_attr_layer_qp_delta_luma specifies the luminance delta qp from InitialSliceQpY in each layer. When ash_attr_layer_qp_delta_luma is not signaled, the value of ash_attr_layer_qp_delta_luma is inferred to be 0 for all layers.
[0172] ash_attr_layer_qp_delta_chroma specifies the chroma delta qp from InitialSliceQpC in each layer. When ash_attr_layer_qp_delta_chroma is not signaled, the value of ash_attr_layer_qp_delta_chroma is inferred to be 0 for all layers.
[0173] The variables SliceQpY[i] and SliceQpC[i] are derived as follows, where i = 0...num_layer - 1:
[0174] for (i = 0; i < num_layer; i++) {
[0175] SliceQpY[i] = InitialSliceQpY + ash attr layer qp delta luma[i]
[0176] SliceQpC[i] = InitialSliceQpC + ash attr layer qp delta chroma[i]
[0177] }
[0178] The delta qp is the difference between the initial qp (signaled in the attribute parameter set (APS)) and the qp used for the current sample.
[0179] Thus, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a value of an index entry associated with a number of dimensions of a current attribute; determine whether the value of the index entry is greater than 1; and refrain from coding at least one of a first chroma delta qp syntax element indicating a chroma delta qp from an initial slice qp in an active attribute parameter set or a second chroma delta qp syntax element indicating a chroma delta qp from an initial slice qp chroma in each layer based on the value of the index entry not being greater than 1. Coding the point cloud can be based at least in part (in some examples, further) on the number of dimensions of the current attribute. In this way, G-PCC encoder 200 can reduce signaling overhead when compared to document w18887.
[0180] G-PCC encoder 200 can signal a syntax element aps attr chroma qp offset in the attribute parameter set where the number of dimensions associated with the attribute is unknown. Thus, the present disclosure does not address changes to coding of this syntax element.
[0181] Similarly, G-PCC encoder 200 can signal information about bit depths and sub-bit depths intended for coding of luma and chroma components in the SPS, respectively. According to the techniques of the present disclosure, G-PCC encoder 200 can signal the sub-bit depth information when the attribute dimension is greater than 1.
[0182]
[0183]
[0184] Accordingly, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine whether to code a syntax element based on an index of a current attribute dimension, where the syntax element indicates one of: a luma delta quantization parameter (QP) from an initial slice QP in an active attribute parameter set, a chroma delta QP from the initial slice QP in the active attribute parameter set, or an indication of a secondary bit depth; and code the point cloud based on the syntax element.
[0185] Signaling of the number of points in a point cloud is now discussed. In a recent review from a national agency, it was suggested that the syntax element gsh num points, which indicates the number of coded points in a point cloud in a slice, is redundant because the number of points in a point cloud will be computed by a decoder when decoding the geometry of the point cloud.
[0186] While the number of points is not necessary for decoding a point cloud, the syntax element is helpful in indicating the amount of complexity or resources required to decode the point cloud. A decoder such as G-PCC decoder 300 can be able to allocate resources using the number of coded points in a point cloud. A decoder such as G-PCC decoder 300 can also be able to negotiate a point cloud bitstream based on the number of points coded in the bitstream. For example, a decoder can negotiate multiple layers of point cloud data to be decoded based on available decoding resources such as processing power. Accordingly, it can be useful to indicate this information in the bitstream. It can not be feasible for a G-PCC decoder to derive the number of coded points in a point cloud because the derivation can require decoding the entire geometry of the point cloud.
[0187] Accordingly, in accordance with the techniques of this disclosure, G-PCC encoder 200 can signal the number of coded points in a point cloud in a slice header (as signaled as gsh num points in document w18887) or at a higher level, which can be useful to G-PCC decoder 300.
[0188] For information that is not explicitly required for decoding of a point cloud bitstream but can be useful to a decoder in typical application scenarios, a separate syntax structure or syntax substructure within the SPS can be defined (e.g., point cloud availability information similar to video availability information in video). Syntax elements such as gsh num points or sps source scale factor (or variants of these syntax elements mentioned in other parts of this disclosure) can be signaled in this new structure or substructure.
[0189] Furthermore, in some cases, point clouds can be coded with a scalable manner with levels of detail. The number of points in increasing levels of detail can increase, and a decoder can be interested in knowing the number of points associated with each level of detail when determining which level or levels to decode. Thus, G-PCC 200 can signal the number of points in each level of detail as follows.
[0190] Thus, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can determine a level of detail syntax element that indicates a number of points associated with a level of detail. The device can code a point cloud based, at least in part (in some examples, further) on the number of points associated with the level of detail. In this way, G-PCC decoder 300 can be able to determine the amount of complexity or resources required to decode a point cloud, and thereby allocate resources using the number of coded points in the point cloud.
[0191] Also, in some examples, the calculation of the exact number of points in a point cloud can not be necessary, and for G-PCC decoder 300 resource allocation, signaling an upper bound of the number of points can be sufficient, G-PCC encoder 200 can instead signal a maximum number of points. This can also reduce the burden on G-PCC encoder 200 to make an accurate estimate of the number of points, which can delay the generation of the syntax structure to the end of the encoding. Other aspects such as alignment of LoD generation and octree structure are contemplated for such use cases.
[0192]
[0193] [Ed. The prefix "aaa_" in the table above will be replaced by the syntax structure that will exist based on these syntax elements.]
[0194] aaa_max_num_points_minus1 plus 1 specifies the maximum number of points in the point cloud.
[0195] aaa_lod_points_present_flag equal to 1 specifies that the number of points in each level of detail is present in the syntax structure. aaa_lod_points_present_flag equal to 0 specifies that the number of points in each level of detail is not signaled.
[0196] aaa_num_lods_minus1 plus 1 specifies the number of levels of detail for which the number of points in each LoD is signaled.
[0197] aaa_max_num_points_in_lod_minus1[ i ] plus 1 specifies the maximum number of points in the i-th LoD.
[0198] The following consistency is required, when aaa_lod_points_present_flag is equal to 1, the sum of (aaa_max_num_points_in_lod_minus1[i] + 1) for the range from 0 to aaa_num_lods_minus1 (inclusive) should be less than or equal to (aaa_max_num_points_minus1 + 1).
[0199] The following consistency is also required, when present, aaa_max_num_points_in_lod_minus1[i] should be greater than or equal to the number of points decoded in the decoder for the level of detail.
[0200] Thus, in some examples, a device such as G-PCC encoder 200 or G-PCC decoder 300 can code a syntax structure including one or more of: a first syntax element, where the first syntax element plus 1 specifies a maximum number of points in a point cloud, a second syntax element, where the second syntax element specifies whether a number of points in each level of detail (LoD) is present in the syntax structure, a third syntax element, where the third syntax element plus 1 specifies a number of levels of detail for which a number of points in each LoD is signaled, or a fourth syntax element, where the fourth syntax element plus 1 specifies a maximum number of points in one of the levels of detail; and code the point cloud based on the syntax structure.
[0201] In some examples, the value of aaa_num_lods_minus1 plus 1 can be constrained to be the number of LoDs in a frame.
[0202] In some examples, G-PCC encoder 200 can signal the exact number of points in a point cloud or level of detail, rather than signaling a maximum number of points.
[0203] G-PCC encoder 200 can signal one or more of the syntax elements specified above in a bitstream that is part of a parameter set (geometry parameter set, attribute parameter set), or in a geometry or attribute payload, or in an SEI message.
[0204] When a point cloud contains multiple frames, one or more of the following can apply:
[0205] 1. In some examples, the G-PCC encoder 200 can signal the number of points signaled per LOD in an attribute parameter set; and the total number of points signaled (aaa num points minusl) can apply to the number of points in each frame, or the maximum number of points in each frame of the point cloud (e.g., by the G-PCC encoder 200).
[0206] 2. In some examples, the G-PCC encoder 200 can signal a syntax structure per frame (e.g., as part of a parameter set, a payload, or an SEI message).
[0207] 3. In some examples, the number of points in a point cloud / LoD applies to all frames; the number of LoD information (aaa num lods minusl) can not equal the number of LoDs in a frame. When the number of LoDs in a frame is more than the number of LoD information, the G-PCC decoder 300 can infer that there is no information indicated about the number of LoDs for those points (or infer a default value). When the number of LoDs in a frame is less than the number of LoD information, the G-PCC decoder 300 can ignore the number of points for additional LoD information that is not present in the frame.
[0208] 4. In some examples, the G-PCC encoder 200 can update this syntax structure when the number of points changes for a frame or each LoD, and signal the updated syntax structure.
[0209] 5. In some examples, the G-PCC encoder 200 can signal other characteristics of the point cloud (in addition to the number of points) per frame, or per layer, or per LoD.
[0210] Figure 5 is a flowchart illustrating an example signaling technique according to this disclosure. A device can determine a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scaling factor of a point cloud (502). For example, one or more processors of the G-PCC encoder 200 can determine a rational representation of a scaling factor of a point cloud to include a numerator, and can determine a value of a numerator syntax element based on the numerator. The G-PCC encoder 200 can signal the numerator syntax element in a bitstream, and one or more processors of the G-PCC decoder 300 can parse the numerator syntax element from the bitstream to determine the value of the numerator syntax element.
[0211] The device can determine a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of a scaling factor of the point cloud (504). For example, one or more processors of the G-PCC encoder 200 can determine a rational representation of the scaling factor of the point cloud to include a denominator, and can determine the value of the denominator syntax element based on the denominator. The G-PCC encoder 200 can signal the denominator syntax element in the bitstream, and one or more processors of the G-PCC decoder 300 can parse the denominator syntax element from the bitstream to determine the value of the denominator syntax element. The device can process the point cloud based at least in part on the scaling factor of the point cloud (506). For example, the G-PCC encoder 200 can process the point cloud based at least in part on the scaling factor of the point cloud. The G-PCC decoder 300 can process the point cloud based at least in part on the scaling factor of the point cloud. For example, the G-PCC decoder 300 can scale the point cloud prior to display based on the scaling factor.
[0212] In some examples, the value of the numerator syntax element plus one specifies a numerator of the scaling factor of the point cloud. In some examples, the value of the denominator syntax element plus one specifies a denominator of the scaling factor of the point cloud.
[0213] In some examples, a device such as the G-PCC encoder 200 or the G-PCC decoder 300 can determine a parameter set identifier (ID), where the parameter set ID is coded using a fixed length. In some examples, the fixed length is 4 bits, and the parameter set ID identifies a parameter set. In some examples, the parameter set can include one of a sequence parameter set, a geometry parameter set, or an attribute parameter set. In some examples, the parameter set is a sequence parameter set and the parameter set ID follows a syntax element indicating a level of image data compression and precedes a syntax element indicating whether a bounding box is present in the sequence parameter set.
[0214] In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 can determine a value of a sampling period syntax element, the value of the sampling period syntax element indicating a sampling period for a detail level index. In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 can further code the point cloud based on or based at least in part on the sampling period. In some examples, the value of the sampling period syntax element plus 2 specifies the sampling period for the detail level index.
[0215] In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 can determine a lift syntax element, where a value of the lift syntax element plus 1 specifies a maximum number of nearest neighbors to be used for prediction. In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 can further code the point cloud based on or based at least in part on the prediction.
[0216] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine a geometry slice header syntax element, where a value of the geometry slice header syntax element plus one specifies a number of points in a geometry slice. G-PCC encoder 200 or G-PCC decoder 300 can determine an attribute bit depth syntax element, where a value of the attribute bit depth syntax element plus one specifies a bit depth of an attribute. G-PCC encoder 200 or G-PCC decoder 300 determines a number of unique patches syntax element, where a value of the number of unique patches syntax element plus one specifies a number of unique patches.
[0217] In some examples, G-PCC encoder 200 or G-PCC decoder 300 can determine a value of an index entry associated with a number of dimensions of a current attribute. G-PCC encoder 200 or G-PCC decoder 300 can determine whether the value of the index entry is greater than one; and based on the value of the index entry not being greater than one, refrain from coding at least one of a first chroma delta qp syntax element (e.g., ash attr qp delta chroma) or a second chroma delta qp syntax element (e.g., ash attr layer qp delta chroma), the first chroma delta qp syntax element indicating a chroma delta qp from an initial slice qp in an active attribute parameter set, the second chroma delta qp syntax element indicating a second chroma delta qp from an initial slice qp chroma in each layer. In some examples, coding the point cloud is further based on the number of dimensions of the current attribute.
[0218] In some examples, G-PCC decoder 300 can determine whether a value of a trisoup syntax element indicating a size of a triangle node is greater than zero. Based on the value of the trisoup syntax element being greater than zero, G-PCC decoder 300 can infer a value of an inferred direct coding mode enabling syntax element is zero indicating whether a direct mode syntax element is present in a bitstream, and infer a value of a unique geometry point syntax element is one indicating whether all output points involved in all slices of a current geometry parameter set have unique positions within the respective slices. In some examples, G-PCC encoder 200 can determine whether a value of a trisoup syntax element indicating a size of a triangle node is greater than zero. Based on the value of the trisoup syntax element being greater than zero, G-PCC encoder 200 can refrain from signaling an inferred direct coding mode enabling syntax element in a bitstream, and refrain from signaling a unique geometry point syntax element in the bitstream. In some examples, G-PCC encoder 200 or G-PCC decoder 300 can further code the point cloud based on or based at least in part on the value of the trisoup syntax element.
[0219] In some examples, the G-PCC encoder 200 or the G-PCC decoder 300 can determine a detail level syntax element that indicates a number of points associated with a detail level. The G-PCC encoder 200 or the G-PCC decoder 300 can also code the point cloud based on, or based on at least in part on, the number of points associated with the detail level.
[0220] Examples in various aspects of the present disclosure can be used alone or in any combination. The present disclosure includes the following clauses.
[0221] Clause 1A. A method of coding point cloud data, the method comprising: determining that a bitstream is not inconsistent with a coding standard based on a first syntax element not being equal to 0 or a second syntax element not being equal to 1, wherein: the first syntax element equal to 1 indicates whether a third syntax element can be present in a geometry node syntax, the second syntax element equal to 1 indicates that all output points have unique positions within a slice in all slices referring to a current geometry parameter set, and the third syntax element indicates whether a single child node of a current node is a leaf node and contains one or more delta point coordinates.
[0222] Clause IB. A method of coding a point cloud, the method comprising: coding a syntax element, wherein the syntax element plus 1 specifies a maximum number of nearest neighbors to be used for prediction; and coding the point cloud using the syntax element.
[0223] Clause 1C. A method of coding a point cloud, the method comprising: coding a syntax element, wherein the syntax element plus 2 specifies a sampling period for a level of detail index; and coding the point cloud using the syntax element.
[0224] Clause ID. A method of coding a point cloud, the method comprising: coding a syntax element, wherein the syntax element equal to 1 specifies that a sampling period is the same for all levels of detail (LoDs), and wherein the syntax element equal to 0 specifies that the sampling period is different for different LoD levels; and coding the point cloud using the syntax element.
[0225] Clause 2D. The method of clause ID, wherein the syntax element is a first syntax element, and the method comprises: coding a second syntax element, wherein the second syntax element plus 2 specifies a constant sampling period of a regular sampling strategy to be used for level of detail (LoD) generation when the first syntax element is equal to 1.
[0226] Clause 3D. The method of clause 2D, further comprising: determining that a value of a third syntax element plus 2 specifies a sampling period for a LoD index, wherein the decoder infers the value of the third syntax element to be equal to the value of the second syntax element when the third syntax element is not present in the bitstream.
[0227] Clause 1E. A method of coding a point cloud, the method comprising: determining a value of a first syntax element, wherein the value of the first syntax element plus one specifies a scaling factor used to derive a square of a sampling distance for a level of detail (LoD) index; determining a value of a second syntax element, wherein the value of the second syntax element specifies an offset used to derive the square of the sampling distance for the LoD index; determining the sampling distance for the LoD index based on the value of the first syntax element and the value of the second syntax element; and coding the point cloud based on the sampling distance for the LoD index.
[0228] Clause 1F. A method of coding a point cloud, the method comprising: determining a value of a first syntax element, the value of the first syntax element indicating a numerator of a scaling factor for the point cloud; determining a value of a second syntax element, the value of the second syntax element indicating a denominator of the scaling factor for the point cloud; and coding the point cloud using the scaling factor for the point cloud.
[0229] Clause 1G. A method of coding a point cloud, the method comprising: using a fixed length value to indicate a syntax element, wherein the syntax element is one of a sequence parameter set (SPS) identifier, a geometry parameter set identifier, an attribute parameter set identifier, a geometry slice header identifier, or an attribute slice header identifier; and coding a bitstream comprising an encoded representation of the point cloud, wherein the bitstream comprises at least one of an SPS identified by the SPS identifier, a geometry parameter set identified by the geometry parameter set identifier, an attribute parameter set identified by the attribute parameter set identifier, a geometry slice header identified by the geometry slice header identifier, or an attribute slice header identified by the attribute slice header identifier.
[0230] Clause 2G. The method of clause 1G, wherein: the syntax element is the SPS identifier; and the syntax element is a first occurring syntax element in the SPS.
[0231] Clause 3G. The method of clause 1G, wherein: the syntax element is the SPS identifier; and the syntax element is a first occurring syntax element in the SPS after a set of reserved bits in the SPS.
[0232] Clause 4G. The method of clause 1G, wherein: the syntax element is the SPS identifier; and the syntax element is a first occurring syntax element in the SPS after a set of bits in which a profile is defined.
[0233] Clause 5G. The method of clause 1G, wherein: the syntax element is the first syntax element and the syntax element is the SPS identifier; and the syntax element is a first occurring syntax element in the SPS after a unique point position constraint flag.
[0234] Clause 6G. The method of clause 1G, wherein the syntax element is an SPS identifier; and the syntax element is the first syntax element that occurs after a layer indicator syntax element in the SPS.
[0235] Clause 1H. A method of coding a point cloud, the method comprising: determining, based on an index of a current attribute dimension, whether to code a syntax element, wherein the syntax element indicates one of: a luma delta quantization parameter (QP) from an initial slice QP in an active attribute parameter set, a chroma delta QP from the initial slice QP in the active attribute parameter set, or an indication of a secondary bit depth; and coding the point cloud based on the syntax element.
[0236] Clause 1I. A method of coding a point cloud, the method comprising: coding a syntax structure comprising one or more of: a first syntax element, wherein the first syntax element plus one specifies a maximum number of points in the point cloud, a second syntax element, wherein the second syntax element specifies whether a number of points in each level of detail (LoD) is present in the syntax structure, a third syntax element, wherein the third syntax element plus one specifies a number of levels of detail for which a number of points in each LoD is signaled, or a fourth syntax element, wherein the fourth syntax element plus one specifies a maximum number of points in one of the levels of detail; and coding the point cloud based on the syntax structure.
[0237] Clause 1J. A method of encoding a point cloud, the method comprising: segmenting a three-dimensional object of the point cloud into a number of unique patches; for each of the unique patches, forming a number of patches by projecting the unique patch in a number of two-dimensional planes; and signaling, in a bitstream, a reduced number of unique patches.
[0238] Clause 2J. A method of decoding a point cloud, the method comprising: obtaining, from a bitstream, a reduced number of unique patches; determining, based on the reduced number of unique patches, a number of unique patches; for each patch of a number of unique patches comprising the number of unique patches, reconstructing the patch based on a number of patches for the patch, the patches for the patch being projections of the patch into a number of two-dimensional planes.
[0239] Clause 1K. A method of coding a point cloud, the method comprising: determining a value of a numerator syntax element, the value of the numerator syntax element indicating a numerator of a scaling factor of the point cloud; determining a value of a denominator syntax element, the value of the denominator syntax element indicating a denominator of the scaling factor of the point cloud; and coding the point cloud based at least in part on the scaling factor of the point cloud.
[0240] Clause 2K. The method of clause 1K, wherein the value of the numerator syntax element plus one specifies the numerator of the scaling factor of the point cloud.
[0241] Clause 3K. The method of clause 1K or 2K, wherein a value of the denominator syntax element plus one specifies a denominator of a scaling factor of the point cloud.
[0242] Clause 1L. A method of coding a point cloud, the method comprising: determining a parameter set identifier (ID), wherein the parameter set ID is coded using a fixed length; and coding the point cloud based at least in part on a parameter set identified by the parameter set ID.
[0243] Clause 2L. The method of clause 1L, wherein the fixed length is 4 bits, wherein the parameter set ID identifies a parameter set, and wherein the parameter set comprises one of a sequence parameter set, a geometry parameter set, or an attribute parameter set.
[0244] Clause 3L. The method of clause 1L or 2L, wherein the parameter set comprises a sequence parameter set and wherein the parameter set ID follows a syntax element indicating a level of image data compression and precedes a syntax element indicating whether a bounding box is present in the sequence parameter set.
[0245] Clause 1M. A method of coding a point cloud, the method comprising: determining a value of a sampling period syntax element, the value of the sampling period syntax element indicating a sampling period for a detail level index; and coding the point cloud based at least in part on the sampling period.
[0246] Clause 2M. The method of clause 1M, wherein a value of the sampling period syntax element plus 2 specifies the sampling period for the detail level index.
[0247] Clause 1N. A method of coding a point cloud, the method comprising: determining a lifting syntax element, wherein a value of the lifting syntax element plus 1 specifies a maximum number of nearest neighbors to be used for prediction; and coding the point cloud based at least in part on the prediction.
[0248] Clause 2N, the method of clause 1N, further comprising: determining a geometry slice header syntax element, wherein a value of the geometry slice header syntax element plus 1 specifies a number of points in a geometry slice; determining an attribute bit depth syntax element, wherein a value of the attribute bit depth syntax element plus 1 specifies a bit depth of an attribute; and determining a unique segment number syntax element, wherein the unique segment number syntax element plus 1 specifies a number of unique segments.
[0249] Clause 1O. A method of coding a point cloud, the method comprising: determining a value of an index entry associated with a number of dimensions of a current attribute; determining whether the value of the index entry is greater than 1; based on the value of the index entry not being greater than 1, refraining from coding at least one of a first chroma delta qp syntax element or a second chroma delta qp syntax element, the first chroma delta qp syntax element indicating a chroma delta qp from an initial slice qp in an active attribute parameter set, the second chroma delta qp syntax element indicating a chroma delta qp from an initial slice qp chroma in each layer; and coding the point cloud based at least in part on the number of dimensions of the current attribute.
[0250] Clause 1P. A method of coding a point cloud, the method comprising: determining whether a value of a trisoup syntax element indicating a size of a triangle node is greater than 0, wherein the value of the trisoup syntax element being 0 indicates that a bitstream includes only octree coding syntax; based on the value of the trisoup syntax element being greater than 0: inferring a value of an inferred direct coding mode enabling syntax element indicating whether a direct mode syntax element is present in the bitstream is 0; and inferring a value of a unique geometry point syntax element indicating whether all output points involved in all slices of a current geometry parameter set have unique positions within the respective slices is 1; and coding the point cloud based at least in part on the value of the trisoup syntax element.
[0251] Clause 1Q. A method of coding a point cloud, the method comprising: determining a level of detail syntax element indicating a number of points associated with a level of detail; and coding the point cloud based at least in part on the number of points associated with the level of detail.
[0252] Clause 1R. The method of any of clauses 1A-1Q, wherein coding comprises decoding.
[0253] Clause 2R. The method of any of clauses 1A-1Q, wherein coding comprises encoding.
[0254] Clause 1S. A device for coding a point cloud, the device comprising one or more means for performing the method of any of clauses 1A-2R.
[0255] Clause 2S. The device of clause 1S, wherein the one or more means comprise one or more processors implemented in circuitry.
[0256] Clause 3S. The device of any of clauses 1S or 2S, further comprising a memory for storing data representing the point cloud.
[0257] Clause 4S. The device of any of clauses 1S-3S, wherein the device comprises a decoder.
[0258] Clause 5SL. The device of any of clauses IS - 3S, wherein the device comprises an encoder.
[0259] Clause 6S. The device of any of clauses IS - 5S, further comprising means for generating a point cloud.
[0260] Clause 7S. The device of any of clauses IS - 5S, further comprising a display for presenting an image based on the point cloud.
[0261] Clause 8S. A computer-readable medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any of clauses IA - 2R.
[0262] It is recognized that depending on the example, certain acts or events of any of the techniques described herein can be performed in a different sequence, can be added, modified or omitted (e.g., not all described acts or events are necessary for the practice of the techniques), and / or can be performed concurrently in any combination (e.g., thread execution, processing, or the like). Furthermore, certain acts or events can be performed at least partially concurrently with, in parallel with, or in some cases, in some manner prior to or after other acts or events involving the same or different processing threads or processing units.
[0263] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0264] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0265] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein the term "processor" and "processing circuitry" can refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0266] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described herein as being stored in or on memory, which can include one or more types of computer-readable storage media. In some examples, a component, module, or unit can be a software module or firmware module encoded in one or more types of computer-readable storage media (e.g., non-transitory computer-readable storage media). In some examples, a component, module, or unit can be a hardware module, such as one or more electronic circuits (e.g., one or more processors or one or more integrated circuits).
[0267] Various examples have been described. These and other examples are within the scope of the following claims.< / m>
Claims
1. A method for encoding and decoding point clouds, the method comprising: Determine the value of a molecular syntax element, the value of which indicates the molecule of the scaling factor of the point cloud; Determine the value of the denominator syntax element, the value of which indicates the denominator of the scaling factor of the point cloud; as well as The point cloud is processed at least in part based on the scaling factor of the point cloud.
2. The method of claim 1, wherein processing the point cloud includes scaling the point cloud.
3. The method of claim 1, wherein the value of the molecular syntax element is added to the molecule specifying the scaling factor of the point cloud.
4. The method of claim 1, wherein the value of the denominator syntax element is added to the denominator by a scaling factor specifying the point cloud.
5. The method of claim 1, further comprising: Determine the parameter set identifier (ID). The parameter set ID is encoded and decoded using a fixed length, and the point cloud is encoded and decoded based on a parameter set identified by the parameter set ID.
6. The method of claim 5, wherein the fixed length is 4 bits, wherein the parameter set ID identifies the parameter set, and wherein the parameter set includes one of a sequence parameter set, a geometric parameter set, or an attribute parameter set.
7. The method of claim 6, wherein the parameter set includes the sequence parameter set and wherein the parameter set ID follows a syntax element indicating an upper limit level associated with point cloud data and precedes a syntax element indicating the presence of a bounding box in the sequence parameter set.
8. The method of claim 1, further comprising: Determine the value of the sampling period syntax element, which indicates the sampling period used for level-of-detail indexing. The encoding and decoding of the point cloud is also based on the sampling period.
9. The method of claim 8, wherein the value of the sampling period syntax element plus 2 specifies the sampling period used for the level of detail index.
10. The method of claim 1, further comprising: Determine the lifting syntax element, where the value of the lifting syntax element plus 1 specifies the maximum number of nearest neighbors that will be used for prediction. The encoding and decoding of the point cloud is also based on the prediction.
11. The method of claim 10, further comprising: Determine the geometry slice header syntax element, wherein the value of the geometry slice header syntax element plus 1 specifies the number of points in the geometry slice; Determine the attribute bit depth syntax element, wherein the value of the attribute bit depth syntax element plus 1 specifies the bit depth of the attribute; as well as Determine the number of unique fragments syntax elements, where the value of the number of unique fragments syntax elements plus 1 specifies the number of unique fragments.
12. The method of claim 1, further comprising: Determine the value of the index entry associated with the number of dimensions of the current attribute; Determine whether the value of the index entry is greater than 1; and based on the fact that the value of the index entry is not greater than 1, suppress encoding / decoding of at least one of the first chroma increment qp syntax element or the second chroma increment qp syntax element, wherein the first chroma increment qp syntax element indicates the chroma increment qp starting from the initial slice qp in the active attribute parameter set, and the second chroma increment qp syntax element indicates the chroma increment qp starting from the chroma of the initial slice qp in each layer. The encoding and decoding of the point cloud is also based on the number of dimensions of the current attribute.
13. The method of claim 1, further comprising: Determine whether the value of the trisoup syntax element indicating the size of the triangle node is greater than 0, wherein a value of 0 for the trisoup syntax element indicates that the bitstream only includes octree encoding / decoding syntax; Based on the fact that the value of the trisoup syntax element is greater than 0: The value of the inferred direct encoding / decoding mode enabled syntax element is 0, indicating whether the direct mode syntax element exists in the bitstream. as well as The inference indicator determines whether all output points in all slices of the current geometry parameter set have a unique geometric point syntax element with a value of 1, indicating whether they have a unique position within the corresponding slice. The encoding and decoding of the point cloud is also based on the value of the trisoup syntax element.
14. The method of claim 1, further comprising: The detail level syntax element determines the number of points associated with the detail level. The encoding and decoding of the point cloud is also based on the number of points associated with the level of detail.
15. An apparatus for encoding and decoding point clouds, the apparatus comprising: A memory is configured to store the point cloud; as well as One or more processors, communicatively coupled to the memory, are configured to: Determine the value of a molecular syntax element, the value of which indicates the molecule of the scaling factor of the point cloud; Determine the value of the denominator syntax element, the value of which indicates the denominator of the scaling factor of the point cloud; as well as The point cloud is processed at least in part based on the scaling factor of the point cloud.
16. The device of claim 15, wherein, as part of processing the point cloud, the one or more processors are configured to scale the point cloud.
17. The device of claim 15, wherein the value of the molecular syntax element is added to the molecule specifying the scaling factor of the point cloud.
18. The device of claim 15, wherein the value of the denominator syntax element is added to the denominator by a scaling factor specifying the point cloud.
19. The device of claim 15, wherein the one or more processors are further configured to: Determine the parameter set identifier (ID). The parameter set ID is encoded and decoded using a fixed length, and the encoding and decoding of the point cloud is also based on the parameter set identified by the parameter set ID.
20. The device of claim 19, wherein the fixed length is 4 bits, wherein the parameter set ID identifies the parameter set, and wherein the parameter set includes one of a sequence parameter set, a geometric parameter set, or an attribute parameter set.
21. The device of claim 20, wherein the parameter set includes the sequence parameter set, and wherein the parameter set ID follows a syntax element indicating an upper limit level associated with point cloud data and precedes a syntax element indicating the presence of a bounding box in the sequence parameter set.
22. The device of claim 15, wherein the one or more processors are further configured to: Determine the value of the sampling period syntax element, which indicates the sampling period used for level-of-detail indexing. The encoding and decoding of the point cloud is also based on the sampling period.
23. The device of claim 22, wherein the value of the sampling period syntax element plus 2 specifies the sampling period used for the level of detail index.
24. The device of claim 15, wherein the one or more processors are further configured to: Determine the lifting syntax element, where the value of the lifting syntax element plus 1 specifies the maximum number of nearest neighbors that will be used for prediction. The encoding and decoding of the point cloud is also based on the prediction.
25. The device of claim 24, wherein the one or more processors are further configured to: Determine the geometry slice header syntax element, wherein the value of the geometry slice header syntax element plus 1 specifies the number of points in the geometry slice; Determine the attribute bit depth syntax element, wherein the value of the attribute bit depth syntax element plus 1 specifies the bit depth of the attribute; and Determine the number of unique fragments syntax elements, where the value of the number of unique fragments syntax elements plus 1 specifies the number of unique fragments.
26. The device of claim 15, wherein the one or more processors are further configured to: Determine the value of the index entry associated with the number of dimensions of the current attribute; Determine whether the value of the index entry is greater than 1; and based on the fact that the value of the index entry is not greater than 1, suppress encoding / decoding of at least one of the first chroma increment qp syntax element or the second chroma increment qp syntax element, wherein the first chroma increment qp syntax element indicates the chroma increment qp starting from the initial slice qp in the active attribute parameter set, and the second chroma increment qp syntax element indicates the chroma increment qp starting from the chroma of the initial slice qp in each layer. The encoding and decoding of the point cloud is also based on the number of dimensions of the current attribute.
27. The device of claim 15, wherein the one or more processors are further configured to: Determine whether the value of the trisoup syntax element indicating the size of the triangle node is greater than 0, wherein a value of 0 for the trisoup syntax element indicates that the bitstream only includes octree encoding / decoding syntax; Based on the fact that the value of the trisoup syntax element is greater than 0: The value of the inferred direct encoding / decoding mode enabled syntax element is 0, indicating whether the direct mode syntax element exists in the bitstream; and The inference indicator determines whether all output points in all slices of the current geometry parameter set have a unique geometric point syntax element with a value of 1, indicating whether they have a unique position within the corresponding slice. The encoding and decoding of the point cloud is also based on the value of the trisoup syntax element.
28. The device of claim 15, wherein the one or more processors are further configured to: The detail level syntax element determines the number of points associated with the detail level. The encoding and decoding of the point cloud is also based on the number of points associated with the level of detail.
29. The apparatus of claim 15, further comprising a display device configured to display the point cloud.
30. The apparatus of claim 15, further comprising a point cloud capturing device configured to capture the point cloud.
31. A non-transitory computer-readable storage medium storing instructions, which, when executed by one or more processors, cause the one or more processors to: Determine the value of a molecular syntax element, the value of which indicates the scaling factor of the point cloud; Determine the value of the denominator syntax element, the value of which indicates the denominator of the scaling factor of the point cloud; and The point cloud is processed at least in part based on the scaling factor of the point cloud.
32. An apparatus for encoding and decoding point clouds, the apparatus comprising: A component for determining the value of a molecular syntax element, the value of which indicates the scaling factor of the point cloud; A component for determining the value of a denominator syntax element, the value of which indicates the denominator of the scaling factor of the point cloud; as well as A component for processing the point cloud based at least in part on the scaling factor of the point cloud.
33. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-14.
Citation Information
Patent Citations
Transmission device, transmission method, reception device, and reception method
US20160234500A1