Encoding and decoding the geometry of a point cloud
Patent Information
- Application Number
- PCT/EP2026/057953
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-22
- Filing Date
- 2026-03-20
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026057953_01102026_PF_FP_ABST
Abstract
Description
[0001] ENCODING AND DECODING THE GEOMETRY OF A POINT CLOUD
[0002] FIELD OF THE INVENTION
[0003] The invention relates to coding and decoding of a point cloud representing the external surface of a 3D object. In particular, it relates to encoding / decoding of the geometry of such a point cloud.
[0004] BACKGROUND OF THE INVENTION
[0005] Traditional visual data describes an object or scene using a series of points that each comprise a position in two dimensions (x and y) and one or more optional attributes like color. Volumetric visual data adds another positional dimension to this traditional visual data. Volumetric visual data describes an object or scene using a series of points that each comprise a position in three dimensions (x, y, and z) and one or more optional attributes like color, reflectance, time stamp, etc. Compared to traditional visual data, volumetric visual data may provide a more immersive way to experience visual data.
[0006] For example, an object or scene described by volumetric visual data may be viewed from any (or multiple) angles, whereas traditional visual data may generally only be viewed from the angle in which it was captured or rendered. Volumetric visual data may be used in many applications, including Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR). Sparse volumetric visual data may be used in the automotive industry for the representation of 3D maps (cartography) or as input to assisted driving systems. In the latter use case, volumetric visual data is typically input to driving decision algorithms. In another example, volumetric visual data may be used to store valuable objects in digital form. In applications for preserving cultural heritage, the goal is to keep a representation of objects that may be threatened by natural disasters. For example, statues, vases, and temples may be entirely scanned and stored as volumetric visual data having several billions of samples. This use case for volumetric visual data may be particularly relevant for valuable objects in locations where earthquakes, tsunamis, and typhoons are frequent. Volumetric visual data may be in the form of a volumetric frame that describes an object or scene captured at a particular time instance or in the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video) that describes an object or scene captured at multiple different time instances.
[0007] One format for storing volumetric visual data is point clouds. A point cloud comprises a collection of points in three-dimensional (3D) space. Each point in a point cloud may comprise geometry information that indicates the point’s position in 3D space. For example, the geometry information may indicate the point’s position in 3D space using three Cartesian coordinates (x, y, and z) or using spherical coordinates (r, phi, theta) (e.g., when acquired by a rotating sensor). The positions of points in a point cloud may be quantized according to a space precision, which may be the same or different in each dimension. The quantization process may create a grid in 3D space. One or more points residing within each sub-grid volume may be mapped to the sub-grid center coordinates, referred to as voxels. A voxel (also referred to as a volumetric pixel) may be considered as a 3D extension of pixels corresponding tothe 2D image grid coordinates. For example, similar to a pixel being the smallest unit when dividing the 2D space (or 2D image) into discrete, uniform (e.g., equally sized) regions, a voxel may be the smallest unit of volume when dividing 3D space into discrete, uniform regions. The sub-grid center coordinates (which correspond to voxels) may be referred to as a voxelized grid. A point in a point cloud may further comprise one or more types of attribute information. Attribute information may indicate a property of a point’s visual appearance. For example, attribute information may indicate a texture (e.g., color) of the point, a material type of the point, transparency information of the point, reflectance information of the point, a normal vector to a surface of the point, a velocity at the point, an acceleration at the point, a time stamp indicating when the point was captured, or a modality indicating how the point was captured (e.g., running, walking, or flying). In another example, a point in a point cloud may comprise light field data in the form of multiple view-dependent texture information. Light field data may be another type of optional attribute information.
[0008] The points in a point cloud may describe an object or a scene. For example, the points in a point cloud may describe the external surface and / or the internal structure of an object or scene. The object or scene may be synthetically generated by a computer or may be generated from the capture of a real-world object or scene. The geometry information of a real-world object or scene may be obtained by 3D scanning and / or photogrammetry. 3D scanning may include laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning may obtain geometry information by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to an object or scene being scanned. Photogrammetry may obtain geometry information by triangulating the same feature or point in different spatially shifted 2D photographs. Point cloud data may be in the form of a point cloud frame that describes an object or scene captured at a particular time instance or in the form of a sequence of point cloud frames (referred to as a point cloud sequence or point cloud video) that describes an object or scene captured at multiple different time instances.
[0009] The data size of a point cloud frame or sequence may be too large for storage and / or transmission in many applications. For example, a single point cloud may comprise over a million points or even billions of points, where each point may comprise geometry information and one or more optional types of attribute information. The geometry information of each point may comprise three Cartesian coordinates (x, y, and z) or spherical coordinates (r, phi, theta) that are each represented, for example, using at least 10 bits per component or 30 bits in total. The attribute information of each point may comprise a texture corresponding to three color components (e.g., R, G, and B color components) that are each represented, for example, using 8-10 bits per component or 24-30 bits in total. A single point therefore comprises at least 54 bits of information in this example, with at least 30 bits of geometry information and at least 24 bits of texture. If a point cloud frame includes a million such points, each point cloud frame would require 54 million bits or 54 megabits to represent. In case of dynamic point clouds that change over time, at a frame rate of 30 frames per second, a data rate of 1.62 gigabits per second would be required to transmit the points of the point cloud sequence. Therefore, raw representations of point clouds may require a large amount of data and the practical deployment of point-cloud-based technologies may need compression technologies that enable the storage and distribution of point clouds with reasonable cost.
[0010] Encoding may be used to compress and / or reduce the data size of a point cloud frame or sequence to provide for more efficient storage and / or transmission. Decoding may be used to decompress a compressed point cloud frame or sequence for display and / or other forms of consumption (e.g., by a machine learning-based device, neural network-based device, artificial intelligence-based device, or other forms of consumption by other types of machine-based processing algorithms and / or devices).
[0011] Compression of point clouds may be lossy (introducing differences relative to the original data) for the distribution to and visualization by an end-user, for example, on AR or VR glasses or any other 3D-capable device. Lossy compression may allow for a high ratio of compression but may imply a trade-off between compression and visual quality perceived by an end-user. Other frameworks, like medical applications or autonomous driving, may require lossless compression to avoid altering the results of a decision obtained based on the analysis of the transmitted and decompressed point cloud frame.
[0012] SUMMARY OF THE INVENTION
[0013] To improve the encoding and decoding the geometry of a point cloud, the present embodiments set out to remedy at least one of the drawbacks of the prior art with a method of encoding a point cloud geometry, comprising: deriving, based on one or more parameters used for obtaining sets of residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of residual RAHT coefficients, a set of one or more values from sets of residual RAHT coefficients of the node; and quantizing each value of the set of one or more values based on a parametric quantizer.
[0014] Also disclosed is a method of decoding a point cloud geometry, comprising: deriving, based on one or more parameters used for obtaining sets of quantized residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of quantized residual RAHT coefficients, a set of one or more quantized values from sets of quantized residual RAHT coefficients of the node; and dequantizing each quantized value of the set of one or more quantized values based on a parametric dequantizer.
[0015] Also disclosed is an encoder comprising one or more processor configured to derive, based on one or more parameters used for obtaining sets of residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of residual RAHT coefficients, a set of one or more values from sets of residual RAHT coefficients of the node; and quantize each value of the set of one or more values based on a parametric quantizer.
[0016] Also disclosed is a decoder comprising one or more processor configured to derive, based on one or more parameters used for obtaining sets of quantized residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of quantized residual RAHT coefficients, a set of one or more quantized values from sets of quantized residual RAHT coefficients of the node; and dequantizing each quantized value of the set of one or more quantized values based on a parametric dequantizer.Also disclosed is a computer program product comprising computer program code means which, when executed on a computing device having a processing system, cause the processing system to perform all of the steps of the method described above.
[0017] The specific nature of the present embodiments as well as other objects, advantages, features and uses of the present embodiments will become evident from the following description of examples taken in conjunction with the accompanying drawings.
[0018] BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Examples of several of the various embodiments of the present disclosure are described herein with reference to the drawings.
[0020] FIG. 1 illustrates an exemplary point cloud coding / decoding system in which embodiments of the present disclosure may be implemented.
[0021] FIG. 2 illustrates the Morton order of eight sub-cuboids split from a cuboid.
[0022] FIG. 3 illustrates an example processing or scanning order for the first three levels of an occupancy tree.
[0023] FIG. 4 illustrates an example of already-coded occupancies of cuboids that may be used to code the occupancy of a current child cuboid.
[0024] FIG. 5 illustrates an example of a dynamic reduction function DR that may be used in dynamic OBUF.
[0025] FIG. 6 illustrates a flowchart of an example method for coding the occupancy (e.g., as indicated by a single bit) of a current child cuboid using dynamic OBUF.
[0026] FIG. 7 illustrates an example of an occupied cube of size NxNxN (where N > 1) that corresponds to a TriSoup node of an occupancy tree.
[0027] FIG. 8A illustrates an example cube corresponding to a TriSoup node with a number K of TriSoup vertices Vk.
[0028] FIG. 8B illustrates an example refinement to the TriSoup model by coding a centroid residual vector Cres into the bitstream such as to use C+Cres instead of C as pivoting vertex for the triangles.
[0029] FIG. 8C illustrates an example of coding a centroid residual vector Cres in / from the bitstream such that an adjusted centroid C+Cres is used instead of centroid C for generating TriSoup triangles of a cuboid corresponding to a portion of a point cloud.
[0030] FIG. 9A and FIG. 9B illustrate examples of voxelization.
[0031] FIG. 10 illustrates an example process for encoding geometry and attributes of a current point cloud.
[0032] FIG. 11 illustrates an example process for encoding attributes associated with a portion of geometry of the decoded geometry.
[0033] FIG. 12 illustrates an example process for decoding geometry and attributes of a current point cloud.FIG. 13 illustrates an example process for decoding attributes associated with a portion of geometry of the decoded geometry.
[0034] FIG. 13A illustrates an example of point-to-point projection distance between a point of a portion of geometry and its nearest neighbor points of a motion compensated geometry.
[0035] FIG. 13B illustrates an example of point-to-point projection distance between a point of the reference point cloud for attributes and its nearest neighbor point of a motion compensated geometry.
[0036] FIG. 14 illustrates an example process for determining attribute predictors of attributes associated with portion of geometry.
[0037] FIG. 15 illustrates an example process for encoding attributes based on attribute predictors associated with portion of geometry.
[0038] FIG. 16 illustrates another example process for encoding attributes based on attribute predictors associated with portion.
[0039] FIG. 17 illustrates an example process for decoding attributes based on attribute predictors associated with portion of geometry.
[0040] FIG. 18 illustrates another example process for decoding attributes based on attribute predictors associated with portion of geometry.
[0041] FIG.19 illustrates an example RAHT transformation applied on child nodes of an octree parent node along three successive directions.
[0042] FIG. 20 illustrates an example RAHT transformation being applied to all octree nodes at depth ‘d’ to determine DC coefficients at depth d-1 and AC coefficients.
[0043] FIG. 21A illustrates an example process for encoding attribute information for child RAHT nodes at depth d of a parent RAHT node at depth d-1 using top-down coding and inter-depth prediction, according to some embodiments.
[0044] FIG. 2 IB illustrates an example process for decoding attribute information for child RAHT nodes at depth d of a parent RAHT node at depth d-1 using top-down decoding and inter-depth prediction, according to some embodiments.
[0045] FIG. 22 illustrates an example process for up-sampling the mean sums of attributes values of RAHT nodes at depth d-1, such as including the parent RAHT node and the already -coded neighboring RAHT nodes, according to some embodiments.
[0046] FIG. 23 illustrates an example process for encoding attributes of a RAHT node using a prediction mode, according to some embodiments.
[0047] FIG. 24 illustrates an example process for performing block related to obtaining the prediction mode for encoding attributes of the RAHT node, according to embodiments.
[0048] FIG. 25 illustrates an example process for decoding attributes of a RAHT node using a prediction mode, according to some embodiments.
[0049] FIG. 26 illustrates an example process for performing block related to obtaining the prediction mode for decoding attributes of the RAHT node, according to embodiments.FIG. 27A illustrates an example of uniform quantizer / dequantizer with a dead zone process, according to embodiments.
[0050] FIG. 27B illustrates an example of quantizer / dequantizer with a dead zone process, according to embodiments.
[0051] FIG. 28 illustrates an example process for encoding attributes of a RAHT node using a prediction mode, according to some embodiments.
[0052] FIG. 29 illustrates an example process for quantizing sets of residual RAHT coefficients, according to some embodiments.
[0053] FIG. 30 illustrates an example process for decoding attributes of a RAHT node using a prediction mode, according to some embodiments.
[0054] FIG. 31 illustrates an example process for dequantizing sets of quantized residual AC coefficients, according to some embodiments.
[0055] FIG. 31 A illustrates an example of a process for determining a parametric quantizer / dequantizer based on one or more parameters of a parametric probability distribution of a set of already dequantized values, according to some embodiments.
[0056] FIG. 3 IB illustrates an example of rolling buffer to obtain already dequantized residual RAHT coefficients of RAHT node(s), according to some embodiments.
[0057] FIG. 32 illustrates a typical observed distribution of the magnitude of residual AC coefficients obtained by coding attributes of a point cloud frame.
[0058] FIG. 33 illustrates a flowchart of an example method for quantizing sets of residual RAHT coefficients, according to some embodiments.
[0059] FIG. 34 illustrates a flowchart of an example method for dequantizing sets of quantized residual RAHT coefficients, according to some embodiments.
[0060] FIG. 35 illustrates a block diagram of an example computer system in which embodiments of the present disclosure may be implemented.
[0061] DETAILED DESCRIPTION OF THE FIGURES
[0062] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure, including structures, systems, and methods, may be practiced without these specific details. The description and representation herein are the common means used by those experienced or skilled in the art to most effectively convey the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the disclosure.
[0063] References in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further,when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0064] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0065] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0066] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks.
[0067] FIG. 1 illustrates an exemplary point cloud coding system 100 in which embodiments of the present disclosure may be implemented. Point cloud coding system 100 comprises a source device 102, a transmission medium 104, and a destination device 106. Source device 102 encodes a point cloud sequence 108 into a bitstream 110 formore efficient storage and / or transmission. Source device 102 may store and / or transmit bitstream 110 to destination device 106 via transmission medium 104. Destination device 106 decodes bitstream 110 to display point cloud sequence 108 or for other forms of consumption. Destination device 106 may receive bitstream 110 from source device 102 via a storage medium or transmission medium 104. Source device 102 and destination device 106 may be any one of a number ofdifferent devices, including a cluster of interconnected computer systems acting as a pool of seamless resources (also referred to as a cloud of computers or cloud computer), a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, a wearable device, a television, a camera, a video gaming console, a set-top box, a video streaming device, an autonomous vehicle, or a head mounted display. A head mounted display may allow a user to view a VR, AR, or MR scene and adjust the view of the scene based on movement of the user’s head. A head mounted display may be tethered to a processing device (e.g., a server, desktop computer, set-top box, or video gaming counsel) or may be fully self-contained.
[0068] To encode point cloud sequence 108 into bitstream 110, source device 102 may comprise a point cloud source 112, an encoder 114, and an output interface 116. Point cloud source 112 may provide or generate point cloud sequence 108 from a capture of a natural scene and / or a synthetically generated scene. A synthetically generated scene may be a scene comprising computer generated graphics. Point cloud source 112 may comprise one or more point cloud capture devices (e.g., one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices), a point cloud archive comprising previously captured natural scenes and / or synthetically generated scenes, a point cloud feed interface to receive captured natural scenes and / or synthetically generated scenes from a point cloud content provider, and / or a processor to generate synthetic point cloud scenes.
[0069] As shown in FIG. 1, a point cloud sequence 108 may comprise a series of point cloud frames 124. A point cloud frame may describe an object or scene captured at a particular time instance. Point cloud sequence 108 may achieve the impression of motion when a constant or variable time is used to successively present point cloud frames 124 of point cloud sequence 108. A point cloud frame may comprise a collection of points 126 in 3D space. Each of points 126 may comprise geometry information that indicates the point’s position in 3D space. For example, the geometry information may indicate the point’s position in 3D space using three Cartesian coordinates (x, y, and z). One or more of points 126 may further comprise one or more types of attribute information. Attribute information may indicate a property of a point’s visual appearance. For example, attribute information may indicate a texture (e.g., color) of a point, a material type of a point, transparency information of a point, reflectance information of a point, a normal vector to a surface of a point, a velocity at a point, an acceleration at a point, a time stamp indicating when a point was captured, a modality indicating how a point was captured (e.g., running, walking, or flying). In another example, one or more of points 126 may comprise light field data in the form of multiple view-dependent texture information. Light field data may be another type of optional attribute information. Color attribute information of one or more of points 126 may comprise a luminance value and two chrominance values. The luminance value may represent the brightness (or luma component, Y) of the point. The chrominance values may respectively represent the blue and red components of the point (or chroma components, Cb and Cr) separate from the brightness. Other color attribute values are possible based on different color schemes (e.g., an RGB or monochrome color scheme).Encoder 114 may encode point cloud sequence 108 into bitstream 110. To encode point cloud sequence 108, encoder 114 may apply one or more lossy compression techniques and / or prediction techniques to reduce redundant information in point cloud sequence 108. Redundant information is information that may be predicted at a decoder and therefore may not be needed to be transmitted to the decoder for accurate decoding of point cloud sequence 108. For example, Motion Picture Expert Group (MPEG) introduced a geometry-based point cloud compression (G-PCC) standard (ISO / IEC standard 23090-9: Geometry-based point cloud compression). G-PCC specifies the encoded bitstream syntax and semantics for transmission and / or storage of a compressed point cloud frame and the decoder operation for reconstructing the compressed point cloud frame from the bitstream. During standardization of G-PCC, a reference software (ISO / IEC standard 23090-21: Reference Software for G-PCC) was developed to encode the geometry and attribute information of a point cloud frame. To encode geometry information of a point cloud frame, the G-PCC reference software encoder may perform voxelization by quantizing positions of points in a point cloud, which creates a grid in 3D space. The G-PCC reference software encoder may map the points to the center coordinates of the sub-grid volume (or voxel) that their quantized locations reside. The G-PCC reference software encoder may perform geometry analysis using an occupancy tree to compress the geometry information. The G-PCC reference software encoder may entropy encode the result of the geometry analysis to further compress the geometry information. To encode attribute information of a point cloud, the G-PCC reference software encoder may apply a transform tool, such as Region Adaptive Hierarchical Transform (RAHT), the Predicting Transform, and / or the Lifting Transform. The Lifting Transform may be built on top of the Predicting Transform but with an extra update / lifting step. Consequently, these two transforms may be referred to as Predicting / Lifting Transform or pred lift. Encoder 114 may operate in a same or similar manner to an encoder provided by the G-PCC reference software.
[0070] Output interface 116 may be configured to write and / or store bitstream 110 onto transmission medium 104 for transmission to destination device 106. In addition, or alternatively, output interface 116 may be configured to transmit, upload, and / or stream bitstream 110 to destination device 106 via transmission medium 104. Output interface 116 may comprise a wired and / or wireless transmitter configured to transmit, upload, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, 3rd Generation Partnership Project (3GPP) standards, Institute of Electrical and Electronics Engineers (IEEE) standards, Internet Protocol (IP) standards, and Wireless Application Protocol (WAP) standards.
[0071] Transmission medium 104 may comprise a wireless, wired, and / or computer readable medium. For example, transmission medium 104 may comprise one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic memory. In addition or alternatively, transmission medium 104 may comprise one more networks (e.g., the Internet) or file servers configured to store and / or transmit encoded video data.To decode bitstream 110 into point cloud sequence 108 for display or other forms of consumption, destination device 106 may comprise an input interface 118, a decoder 120, and a point cloud display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104 by source device 102. In addition, or alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may comprise a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as those mentioned above.
[0072] Decoder 120 may decode point cloud sequence 108 from encoded bitstream 110. For example, decoder 120 may operate in a same or similar manner to a decoder provided by G-PCC reference software. In some examples, decoder 120 may decode a point cloud sequence that approximates point cloud sequence 108 due to, for example, lossy compression of point cloud sequence 108 by encoder 114 and / or errors introduced into encoded bitstream 110 during transmission to destination device 106.
[0073] Point cloud display 122 may display point cloud sequence 108 to a user. Point cloud display 122 may comprise a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, a 3D display, a holographic display, a head mounted display, or any other display device suitable for displaying point cloud sequence 108.
[0074] It should be noted that point cloud coding / decoding system 100 is presented by way of example and not limitation. In the example of FIG. 1, point cloud coding / decoding system 100 may have other components and / or arrangements. For example, point cloud source 112 may be external to source device 102. Similarly, point cloud display 122 may be external to destination device 106 or omitted altogether where point cloud sequence is intended for consumption by a machine and / or storage device. In another example, source device 102 may further comprise a point cloud decoder and destination device 106 may comprise a point cloud encoder. In such an example, source device 102 may be configured to further receive an encoded bit stream from destination device 106 to support two-way point cloud transmission between the devices.
[0075] As mentioned above, an encoder may quantize the positions of points in a point cloud according to a space precision, which may be the same or different in each dimension of the points. The quantization process may create a grid in 3D space. The encoder may map any points residing within each sub-grid volume to the sub-grid center coordinates, referred to as a voxel (or a volumetric pixel). A voxel may be considered as a 3D extension of pixels corresponding to 2D image grid coordinates.
[0076] The encoder may represent or code the point cloud using an occupancy tree. For example, the encoder may split the initial volume or cuboid (also referred to as a bounding box) containing the point cloud into sub-cuboids. The encoder may then recursively split each sub-cuboid that contains at least one point of the point cloud. The encoder may not further split sub-cuboids that do not contain at least one point of the point cloud. A sub-cuboid that contains at least one point of the point cloud may be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud may be referred to as an unoccupied sub-cuboid. The encoder may split an occupied cuboid into, for example,two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub-cuboids (to form an octree). The encoder may split an occupied cuboid to obtain sub-cuboids all with the same size and shape at a given depth level of the occupancy tree by splitting following a plane passing through the middle of edges of the cuboid.
[0077] The initial volume or cuboid containing the point cloud may correspond to the root node of the occupancy tree. Each occupied sub-cuboid, split from the initial volume / cuboid, may correspond to a node (of the root node) in a second level of the occupancy tree. Each occupied sub-cuboid, split from an occupied sub-cuboid in the second level, may correspond to a node (off the occupied sub-cuboid in the second level from which it was split) in a third level of the occupancy tree. The occupancy tree structure may continue to form in this manner for each recursive split iteration until, for example, a maximum depth level of the occupancy tree is reached or each occupied sub-cuboid has a volume corresponding to one voxel.
[0078] Each non-leaf node of the occupancy tree may comprise or be associated with an occupancy word representing an occupancy state of the cuboid corresponding to the node. For example, a node of the occupancy tree corresponding to a cuboid that is split into 8 sub-cuboids may comprise or be associated with a 1-byte occupancy word. Each bit (referred to as an occupancy bit) of the 1-byte occupancy word may represent or indicate the occupancy of a different one of the eight sub-cuboids. Occupied subcuboids may be represented or indicated by a binary value of 1 in the 1 -byte occupancy word and unoccupied sub-cuboids may be represented or indicated by a binary value of 0 in the 1-byte occupancy word. In other examples, occupied and un-occupied sub-cuboids may be represented or indicated by opposite 1 -bit binary values in the 1-byte occupancy word.
[0079] Each bit of an occupancy word may represent or indicate the occupancy of a different one of the eight sub-cuboids following the so-called Morton order. For example, the least significant bit of an occupancy word may represent or indicate the occupancy of a first one of the eight sub-cuboids following the Morton order, the second least significant bit of an occupancy word may represent or indicate the occupancy of a second one of the eight sub-cuboids following the Morton order, etc.
[0080] FIG. 2 illustrates the Morton order of eight sub-cuboids 202-216 split from a cuboid 200. Subcuboids 202-216 are labeled based on their Morton order, with child node 202 being the first in Morton order and child node 216 being the last in Morton order. The Morton order for sub-cuboids 202-216 is a local lexicographic order in xyz.
[0081] The geometry of the point cloud is represented by, and therefore may be determined from, the initial volume and the occupancy words of the nodes in the occupancy tree. The encoder may therefore transmit the initial volume and the occupancy words of the nodes in the occupancy tree in a bitstream to a decoder for reconstructing the point cloud. Before transmitting the initial volume and the occupancy words of the nodes in the occupancy tree, the encoder may entropy encode the occupancy words. For example, the encoder may encode an occupancy bit of an occupancy word of a node corresponding to a cuboid, based on one or more occupancy bits of occupancy words of other nodes corresponding to cuboids that are adjacent or spatially close to the cuboid of the occupancy bit being encoded.An encoder and / or decoder may code occupancy bits of occupancy words in sequence of a scan order. For example, an encoder and / or decoder may scan an occupancy tree in breadth-first order: all the occupancy words of the nodes of a given depth (or level) within the occupancy tree may be scanned before scanning the occupancy words of the nodes of the next depth (or level). Within a depth, the encoder and / or decoder may scan the occupancy words of nodes in the Morton order. Within a node, the encoder and / or decoder may scan the occupancy bits of the occupancy word of the node further in the Morton order.
[0082] FIG. 3 illustrates an example of this scanning order for the first three levels of an occupancy tree 300. At each level of occupancy tree 300, a plurality of cuboids (e.g., cubes) are generated. In FIG. 3, a cube 302 corresponding to the root node of occupancy tree 300 is divided into eight sub-cubes. Two subcubes 304 and 306 of the eight sub-cubes are occupied, while the other six sub-cubes are unoccupied. Following the Morton order, a first eight-bit occupancy word occWi.i is constructed to represent the occupancy word of the root node. The least significant occupancy bit of the first eight-bit occupancy word occWi.i represents or indicates the occupancy of the first sub-cube of the eight sub-cubes in Morton order, the second least significant occupancy bit of the first eight-bit occupancy word occWi.i represents or indicates the occupancy of the second sub-cube of the eight sub-cubes in Morton order, etc.
[0083] Each of the two occupied sub-cubes 304 and 306 corresponds to a node off the root node in a second level of occupancy tree 300. The two occupied sub-cubes 304 and 306 are each further split into eight sub-cubes. One of the sub-cubes 308 of the eight sub-cubes split from sub-cube 304 is occupied, while the other seven sub-cubes are unoccupied. Three of the sub-cubes 310, 312, and 314 of the eight sub-cubes split from sub-cube 306 are occupied, while the other five sub-cubes of the eight sub-cubes split from sub-cube 306 are unoccupied. Two second eight-bit occupancy words occW2,i and occW2,2 are constructed in this order to respectively represent the occupancy word of the node corresponding to subcube 304 and the occupancy word of the node corresponding to sub-cube 306.
[0084] Each of the four occupied sub-cubes 308, 310, 312, and 314 corresponds to a node in a third level of occupancy tree 300. The four occupied sub-cubes 308, 310, 312, and 314 are each further split into eight sub-cubes or 32 sub-cubes in total. Four third eight-bit occupancy words occWs i. occW3,2, occWs s and occWs.4 are constructed in this order to respectively represent the occupancy word of the node corresponding to sub-cube 308, the occupancy word of the node corresponding to sub-cube 310, the occupancy word of the node corresponding to sub-cube 312, and the occupancy word of the node corresponding to sub-cube 314.
[0085] Following the scanning order discussed above, the occupancy words of this exemplary occupancy tree 300 may be entropy coded (e.g., entropy encoded by an encoder and entropy decoded by a decoder) as the succession of the seven occupancy words occWi.i to OCCWS AS a consequence of the breadth-first scanning order, when entropy coding the occupancy word of a current child node belonging to a current parent node, the occupancy words of all nodes having the same depth (or level) as the current parent node have already been entropy coded. In addition, the occupancy words of all nodes having the same depth (or level) as the current child node and having a lower Morton order than the current child node have alsoalready been entropy coded. Part of these already coded occupancy words may be used to entropy code the occupancy word of the current child node. For example, the already coded occupancy words of neighboring parent and child nodes may be used to entropy code the occupancy word of the current child node. When entropy coding a particular occupancy bit of the occupancy word of the current child node, the occupancy bits of the occupancy word having a lower Morton order than the particular occupancy bit have also already been entropy coded and may be used to code the occupancy bit of the occupancy word of the current child node.
[0086] FIG. 4 illustrates an example neighborhood of cuboids with already-coded (previously-coded) occupancy bits that may be used to entropy code the occupancy bit of a current child cuboid 400. The neighborhood of cuboids with already-coded occupancy bits may be determined based on the scanning order of an occupancy tree representing the geometry of the cuboids in FIG. 4 as discussed above. As illustrated in FIG. 4, current child cuboid 400 belongs to a current parent cuboid 402. Following the scanning order of the occupancy words and occupancy bits of nodes of the occupancy tree, the occupancy bits of four child cuboids 404, 406, 408, and 410, belonging to the same current parent cuboid 402, have already been coded. Also, the occupancy bit of child cuboids 412 of preceding parent cuboids have already been coded. Furthermore, the occupancy bits of parent cuboids 414, for which the occupancy bits of child cuboids have not already been coded, have already been coded. Therefore, the already-coded occupancy bits of cuboids 404, 406, 408, 410, 412, and 414 may be used to code the occupancy bit of the current child cuboid 400.
[0087] The number of possible occupancy configurations for a neighborhood of a current child cuboid may be 2N, where N is the number of cuboids in the neighborhood of the current child cuboid with already-coded occupancy bits. The neighborhood of the current child cuboid may comprise several dozens of cuboids, among them the 26 adjacent parent cuboids sharing a face, an, edge, or a vertex with the parent cuboid of the current child cuboid and also several adjacent child cuboids (with occupancy bits already coded) sharing a face, an edge, or a vertex with the current child cuboid. Even limited to a subset of the adjacent cuboids, the occupancy configuration for a neighborhood of the current child cuboid may have billions of possible occupancy configurations making its direct use impractical. The occupancy configuration for a neighborhood of the current child cuboid may be used by an encoder and / or decoder to select the context (or equivalently the probability model), among a set of contexts, of a binary entropy coder (e.g., binary arithmetic coder) that codes the occupancy bit of the current child cuboid. The contextbased binary entropy coding may be similar to the Context Adaptive Binary Arithmetic Coder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).
[0088] Several methods may be used by an encoder and / or decoder to reduce the occupancy configurations for a neighborhood of a current child cuboid being coded to a practical number of reduced occupancy configurations. Firstly, the 26or 64 occupancy configurations of the six adjacent parent cuboids sharing a face with the parent cuboid of the current child cuboid may be reduced to 9 occupancy configurations by using geometry invariance. Secondly, an occupancy score for the current child cuboid may be obtained from the 226occupancy configurations of the 26 adjacent parent cuboids. The score maybe further reduced into a ternary occupancy prediction (“predicted occupied”, “unsure”, “predicted unoccupied”) by applying score thresholds. Thirdly, the number of occupied and the number of unoccupied adjacent child cuboids may be used instead of the individual occupancies of these child cuboids.
[0089] An encoder and / or decoder employing one or more of the above methods may reduce the number of possible occupancy configurations for a neighborhood of a current child cuboid to a more manageable number (e.g., a few thousands). However, it has been observed that instead of associating a reduced number of contexts (or probability models) directly to the reduced occupancy configurations, another mechanism may be used, namely Optimal Binary Coders with Update on the Fly (OBUF). An encoder and / or decoder may implement OBUF to limit the number of contexts to a lower number (e.g., 32 contexts).
[0090] OBUF may use a limited number (e.g., 32) of contexts that may be fixed. These contexts may be ordered, referred to by a context index (e.g., a context index in the range of 0 to 31), and associated from a lowest virtual probability to a highest virtual probability to code a 1. A Uook-Up Table (UUT) of context indices may be initialized at the beginning of a point cloud coding process. For example, the UUT may initially point to a context (e.g., context with context index 15), among the limited number of contexts, with the median virtual probability to code a 1 for all input. This UUT may take an occupancy configuration for a neighborhood of current child cuboid as input and output the context index associated with the occupancy configuration. Consequently, the UUT may have as many entries as reduced occupancy configurations (e.g., around a few thousand). The coding of the occupancy bit of a current child cuboid may follow the steps of determining the reduced occupancy configuration of the current child node, obtaining a context index by applying the reduced occupancy configuration as an entry to the UUT, coding the occupancy bit of the current child cuboid by using the context pointed to (or indicated) by the context index, and finally updating the UUT entry corresponding to the reduced occupancy configuration depending on the value of the coded occupancy bit of the current child cuboid. If a binary 0 (e.g., indicating the current child cuboid is unoccupied) is coded, the UUT entry may be decreased to a lower context index value, and if a binary 1 (e.g., indicating the current child cuboid is occupied) is coded, the UUT entry may be increased to a higher context index value. The update process of the context index may be based on a theoretical model of optimal distribution for virtual probabilities associated with the limited number of contexts. This virtual probability for a context may be fixed by a model and may be different from the internal probability of the context that evolves during the coding of bits of data. The evolution of the internal context may follow a well-known process similar to the process in CABAC.
[0091] An encoder and / or decoder may implement a “dynamic OBUF” scheme that may handle a much larger number of occupancy configurations for a neighborhood of a current child cuboid than can be handled by general OBUF, while maintaining complexity within reasonable bounds. The use of a larger number of occupancy configurations for a neighborhood of a current child cuboid may lead to improved compression capabilities. By using an occupancy tree compressed by OBUF, an encoder and / or decoder may reach a lossless compression performance as good as 1 bit per point (bpp) for coding the geometry ofdense point clouds. An encoder and / or decoder may implement dynamic OBUF to potentially further reduce the bitrate by more than 25% to 0.7 bpp.
[0092] OBUF may not take as input a large variety of reduced occupancy configurations for a neighborhood of a current child cuboid, thus potentially leading to a loss of useful correlation. The size of the LUT of context indices may be increased to handle more various occupancy configurations for a neighborhood of a current child cuboid as input. However, by doing so, statistics may be diluted, and compression performance may be reduced. For example, if the LUT has millions of entries and the point cloud has a hundred thousand points, then most of the entries are never visited. Worse yet, many entries may be visited only a few times and their associated context indices may not be updated enough times to reflect any meaningful correlation between the occupancy configuration value and the probability of occupancy of the current child cuboid. Dynamic OBUF may be implemented to mitigate the dilution of statistics due to the increase in the number of occupancy configurations for a neighborhood of a current child cuboid. This mitigation is performed by a “dynamic reduction” of occupancy configurations in dynamic OBUF.
[0093] Dynamic OBUF may add an extra step of reduction of occupancy configurations for a neighborhood of a current child cuboid before applying the LUT of context indices. This step may be called a dynamic reduction because it evolves based on the progress of the coding of the point cloud or, more precisely, based on already visited occupancy configurations.
[0094] As discussed above, many possible occupancy configurations for a neighborhood of a current child cuboid are potentially involved but only a subset may be visited during the coding of a point cloud. This subset may characterize the type of the point cloud. For example, when coding AR or VR dense point clouds, most of the visited occupancy configurations may exhibit occupied adjacent cuboids of a current child cuboid. On the other hand, when coding sensor-acquired sparse point clouds, most of the visited occupancy configurations may exhibit only a few occupied adjacent cuboids of a current child cuboid. The role of the dynamic reduction may be to obtain a more precise correlation based on the most visited occupancy configuration while putting aside (or reducing aggressively) other occupancy configurations that are much less visited. The dynamic reduction may be updated on-the-fly, as detailed below, after each visit of an occupancy configuration during the coding of occupancy data.
[0095] FIG. 5 illustrates an example of a dynamic reduction function DR that may be used in dynamic OBUF. The dynamic reduction function DR may be obtained by masking bits PJ of occupancy configurations 500:
[0096] P = Pi ... PK
[0097] made of K bits. The size of the mask may decrease when occupancy configurations are visited a certain number of times. The initial dynamic reduction function DR0may mask all bits for all occupancy configurations such that it is a constant function DR°(P) = 0 for all occupancy configurations p. After each coding of an occupancy bit, the dynamic reduction function may evolve from a function DRnto an updated function DRn+1. The function may be defined by:
[0098] P’ =DRn(P) = Pi ... pkn(P)where kn(P) 510 is the number of non-masked bits. The initialization of DR0may correspond to ko(P)=O, and the natural evolution of the reduction function towards finer statistics may lead to an increasing number of non-masked bits kn(P) < kn+i (P). The dynamic reduction function may be entirely determined by the values of knfor all occupancy configurations p.
[0099] The visits to occupancy configurations may be tracked by a variable NV(P’) for all dynamically reduced occupancy configurations P’= DRn(P). After the coding of an occupancy bit based on an occupancy configuration pv, the corresponding number of visits NV(PV’) may be increased by one. If this number of visits NV(PV’) is greater than a threshold thv,
[0100] NV(PV’) > thv
[0101] then the number of unmasked bits kn(P) may be increased by one for all occupancy configurations P being dynamically reduced to pv’. Practically, this corresponds to replacing the dynamically reduced occupancy configuration pv’ by the two new dynamically reduced occupancy configurations P°’ and P” defined by P°’ = Pv’O = pvi ... pVkn(p)0 and p1’ = pv’ 1 = pV! ... p^l.
[0102] In other words, the number of unmasked bits has been increased by one kn+i (P) = kn(P) + 1 for all occupancy configurations P such that DRn(P) = pv’. The number of visits of the two new dynamically reduced occupancy configurations may then be initialized to zero:
[0103] NV(P°’) = NV(P1’) = O. (I)
[0104] At the start of the coding, the initial number of visits for the initial dynamic reduction function DR0may be set to
[0105] NV(DR°(P)) = NV(0) = 0,
[0106] and the evolution of NV on dynamically reduced occupancy configurations may now be entirely defined.
[0107] When a dynamically reduced occupancy configuration pv’ is replaced by the two new dynamically reduced occupancy configurations P°’ and P”, the corresponding LUT entry LUT[pv’] may be replaced by the two new entries LUT[P°’] and LUT[P” ] that are initialized by the context index associated with pv’,
[0108] LUT[P0’] = LUTfP1’] = LUT[PV’], (II)
[0109] and then evolve separately. The evolution of the LUT of context indices on dynamically reduced occupancy configurations may thus be entirely defined.
[0110] The reduction function DRnmay be modeled by a series of growing binary trees Tn520 whose leaf nodes 530 are the reduced occupancy configurations P’ = DRn(P). The initial tree may be the single root node associated with 0 = DR°(P). The replacement of the dynamically reduced to pv’ by p0’ and p1’ corresponds to growing the tree Tnfrom the leaf node associated with pv’ by attaching to it two new nodes associated with p0’ and P”. The tree Tn lmay be obtained by this growth. The number of visits NV and the LUT of context indices may be defined on the leaf nodes and evolve with the growth of the tree through equations (I) and (II).
[0111] In some examples, dynamic OBUF may be practically implemented by storage of the array NV[P’] and the LUT[P’] of context indices, as well as the trees Tn520. An alternative to the storage of the trees may be to store the array kn[P] 510 of the number of non-masked bits.A limitation for implementing dynamic OBUF may be its memory footprint. In some applications, a few million occupancy configurations may be practically handled, leading to about 20 bits Pi constituting an entry configuration to the reduction function DR. Each bit Pi may correspond to the occupancy status of a neighboring cuboid of a current child cuboid or a set of neighboring cuboids of a current child cuboid.
[0112] Higher bits Pi (e.g., Po, Pi, etc.) may be the first bits to be unmasked during the evolution of the dynamic reduction function DR. Therefore, the order of neighbor-based information put in the bits Pi may impact the compression performance. In some examples, neighboring information may be ordered from highest priority to lower priority and put in this order into the bits p from higher to lower weight. For example, the priority may be, from the most important to the least important, occupancy of sets of adjacent neighboring child cuboids, then occupancy of adjacent neighboring child cuboids, then occupancy of adjacent neighboring parent cuboids, then occupancy of non-adjacent neighboring child nodes, and finally occupancy of non-adjacent neighboring parent nodes. Adjacent nodes sharing a face with the current child node may also have higher priority than adjacent nodes sharing an edge or, worse, only a vertex with the current child node.
[0113] FIG. 6 illustrates a flowchart of an exemplary method for coding the occupancy bit of a current child cuboid using dynamic OBUF. The method of the flowchart begins at block 602. At block 602, an encoder and / or decoder may determine the occupancy configuration p of already-coded cuboids in a neighborhood of the current child cuboid. At block 604, the encoder and / or decoder may dynamically reduce the occupancy configuration P into a reduced occupancy configuration P’ = DRn(P). At block 606, the encoder and / or decoder may lookup context index LUT[P’] in the LUT of the dynamic OBUF. At block 608, the encoder and / or decoder may select the context (or probability model) pointed to by the context index. At block 610, the encoder and / or decoder may entropy code (e.g., arithmetic code) the occupancy bit of the current child cuboid based on the context. Thus, the occupancy bit of the current child cuboid may be coded based on occupancy bits of the already-coded cuboids neighboring the current child cuboid .
[0114] Although not shown in FIG. 6, the encoder and / or decoder may further update the reduction function DRninto DRn+1and update the context index EUT[P’] based on the occupancy bit of the current child cuboid. In addition, the method of FIG. 6 may be repeated for additional or all child cuboids of parent cuboids corresponding to nodes of the occupancy tree in a scan order, such as the scan order discussed above with respect to FIG. 3.
[0115] In general, the occupancy tree is a lossless compression technique. The occupancy tree may be adapted to provide lossy compression by modifying the point cloud on the encoder side (e.g., downsampling, removing points, moving points, etc.) but the lossy compression performance may be reduced / weak. However, the use of the occupancy tree as a lossless compression technique may be very useful for dense point clouds.
[0116] One approach to lossy compression for point cloud geometry may be to set the maximum depth of the occupancy tree to not reach the smallest volume size of one voxel but instead to stop at a biggervolume size (e.g., NxNxN cubes, where N > 1). The geometry of the points belonging to each occupied leaf node associated with the bigger volumes may then be modeled. This approach may be particularly suited for dense and smooth point clouds that may be locally modeled by smooth functions like planes or polynomials. The coding cost may become the cost of the occupancy tree plus the cost of the local model in each of the occupied leaf nodes.
[0117] A scheme for modeling the geometry of the points belonging to each occupied leaf node, associated with a volume size larger than one voxel, may use sets of triangles as local models. This scheme may be referred to as the “TriSoup” scheme. TriSoup is short for “Triangle Soup” because the connectivity between triangles may not be part of the models. An occupied leaf node, of an occupancy tree, that corresponds to a cuboid with a volume greater than one voxel may be referred to as a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node may be referred to as a TriSoup edge. A TriSoup node may comprise a presence flag (sk) for each TriSoup edge of its corresponding occupied cuboid. A presence flag (sk) of a TriSoup edge may indicate (a presence of or) whether a TriSoup vertex (Vk) is present or not on the TriSoup edge. At most one TriSoup vertex (Vk) may be present on a TriSoup edge. For each vertex (Vk) present on a TriSoup edge of an occupied cuboid, the Tri Soup node corresponding to the occupied cuboid may further comprise a position (pk) of the vertex (Vk) along the TriSoup edge.
[0118] In addition to the occupancy words of an occupancy tree, an encoder may entropy encode, for each TriSoup node of the occupancy tree, a TriSoup vertex presence flag (and a position of a TriSoup vertex, if present, along a TriSoup edge) of each TriSoup edge belonging to the TriSoup node. A decoder may similarly entropy decode the TriSoup vertex presence flags and positions of each TriSoup vertex along a respective TriSoup edge belonging to a TriSoup node of the occupancy tree, in addition to the occupancy words of the occupancy tree.
[0119] FIG. 7 illustrates an example of an occupied cube 700 of size NxNxN (where N > 1) that corresponds to a TriSoup node of an occupancy tree. Occupied cube 700 comprises TriSoup edges 710-721. The TriSoup node, corresponding to occupied cube 700, comprises a presence flag (sk) for each TriSoup edge of TriSoup edges 710-721. The presence flag of TriSoup edge 714 indicates that a TriSoup vertex Vi is present on TriSoup edge 714. The presence flag of TriSoup edge 715 indicates that a TriSoup vertex V2 is present on TriSoup edge 715. The presence flag of TriSoup edge 716 indicates that a TriSoup vertex V3 is present on TriSoup edge 716. The presence flag of TriSoup edge 717 indicates that a TriSoup vertex V4 is present on TriSoup edge 718. The presence flags of the remaining TriSoup edges each indicates that a TriSoup vertex is not present on their corresponding TriSoup edge. The TriSoup node, corresponding to occupied cube 700, further comprises a position (pk) for each TriSoup Vertex present along one of its TriSoup edges 710-721. More specifically, the TriSoup node (corresponding to occupied cube 700) further comprises a position pi for TriSoup vertex Vi, a position p2 for TriSoup vertex V2, a position ps for TriSoup vertex V3, and a position p4 for TriSoup vertex V4. The TriSoup vertices may be shared among TriSoup nodes along TriSoup edge(s) in common.In some examples, a presence flag (sk) and, if the presence flag (sk) indicates the presence of a vertex, a position (pk) (the presence flag (sk) and position (pk) individually or collectively referred to as vertex information) of the vertex along a current Tri Soup edge may be entropy coded based on already-coded presence flags and positions (of present TriSoup vertices) of TriSoup edges that neighbor the current TriSoup edge. A presence flag (sk) and, if the presence flag (sk) indicates the presence of a vertex, a position (pk) on (e.g., indicating a position of the vertex along) a current TriSoup edge may be additionally or alternatively entropy coded based on occupancies of cuboids that neighbor the current TriSoup edge. Similar to the entropy coding of the occupancy bits of the occupancy tree, a configuration PTS for a neighborhood (also referred to as a neighborhood configuration TS) of a current TriSoup edge may be obtained and dynamically reduced into a reduced configuration PTS’ = DRn(PTS) by using a dynamic OBUF scheme for TriSoup. A context index LUTfPTS’] may be obtained from the OBUF LUT and at least a part of the vertex information of the current Tri Soup edge may be entropy coded using the context (or probability model) pointed to by the context index.
[0120] In order to use a binary entropy coder to entropy code at least part of the vertex information of the current TriSoup edge, the TriSoup vertex position (pk) (if present) along its TriSoup edge may be binarized. A number of bits Nb may be set for the quantization of the TriSoup vertex position (pk) along the TriSoup edge of length N that is uniformly divided into 2Nbquantization intervals. By doing so, the TriSoup vertex position (pk) may be represented by Nb bits (pkj, j=l,...,Nb) that may be individually coded by the dynamic OBUF scheme as well as the bit corresponding to the presence flag (sk). The neighborhood configuration PTS, the OBUF reduction function DRn, and thus the context index may depend on the nature / characteristic / property of the coded bit (presence flag (sk), highest position bit (pki), second highest position bit (pk2), etc.). Therefore, there may be several dynamic OBUF schemes implemented, with each dedicated to a specific bit of information (presence flag (sk) or position bit (pkJ)) of the vertex information.
[0121] FIG.8A illustrates a cuboid 800 (e.g., a cube) corresponding to a TriSoup node with a number K of TriSoup vertices Vk. Within cuboid 800, TriSoup triangles may be constructed from the TriSoup vertices Vk if at least three (K>3) TriSoup vertices are present on the TriSoup edges of cuboid 800. In the example of FIG. 8A, 4 TriSoup vertices are present and therefore TriSoup triangles are constructed. The TriSoup triangles may be constructed around the centroid vertex C defined as the mean of the TriSoup vertices Vk. In some examples, to construct the TriSoup triangles, a dominant direction may first be determined, then vertices Vk may be ordered by turning around this direction, and finally the following K TriSoup triangles (listed as triples of vertices) are constructed: V1V2C, V2V3C, ..., VKVIC. The dominant direction may be chosen among the three directions parallel to the axis of the 3D space to increase or maximize the 2D surface of the triangles when projected along the dominant direction. By doing so, the dominant direction may be somewhat perpendicular to a local surface defined by the points of the point cloud belonging to the TriSoup node.
[0122] FIG. 8B illustrates a refinement to the TriSoup model by coding a centroid residual vector Cres into the bitstream such as to use C+Cres instead of C as a pivoting vertex for constructing / generating thetriangles. By doing so, the vertex C+Cresmay be closer to the points of the point cloud than the centroid C used to model the points, which reduces the reconstruction error and leads to lower distortion at the cost of a small increase in bitrate needed for coding Cres.
[0123] FIG. 8C illustrates a more detailed example of coding a centroid residual vector Cres in / from the bitstream such that an adjusted centroid C+Cres is used instead of centroid C for generating TriSoup triangles of a cuboid 800 (corresponding to a TriSoup node) corresponding to a portion of a point cloud, according to some embodiments. For example, the triangles may be generated based on adjusted centroid C+Cres and adjacent pairs of vertices of an ordering of the vertices V1-V4, determined as described above with respect to FIG. 8A. Further, as described above, the TriSoup triangles of the cuboid may be voxelized at the decoder to generate voxels representing (or modeling) the portion, of the point cloud, corresponding to the cuboid. A unit vector n (i.e., also referred to as a normalized vector) may be determined as a normalized mean vector of normal vectors to the triangles (V1V2C, V2V3C, ... , VKVIC) constructed by centroid C and pairs of the vertices of the cuboid by pivoting around the centroid C (e.g., as described in FIG. 8A). For example, the unit vector n may be determined as the normalized vector based on a mean of cross-products representing areas of the triangles (^j C x V2C + V2C x V3C + — I- VKC X V C ) / K. For example, the unit vector n may be determined by dividing the mean vector (n) by the norm (or length) of the mean vector (i.e., n = n / ||n||).
[0124] A value resulting from each cross product is equal to an area of a parallelogram formed by the two vectors in the cross product. Therefore, the value may be representative of an area of a triangle formed by the two vectors because the area of the triangle is equal to half of the value. Accordingly, since the vector n indicates a direction of the triangles (e.g., TriSoup triangles) representing (e.g., modeling) the portion of the point cloud, the vector n may be indicative of the direction normal to a local surface representative of the portion of the point cloud. In some examples, to maximize the effect of the centroid residual while minimizing its coding cost, a one-component residual aresalong the line (C, n) 810 may be coded instead of a 3D residual vector.
[0125] Cres OCresYl
[0126] The residual value aresmay be determined by the encoder as the intersection between the current point cloud and the line (C, n), which is along the same direction of the normalized vector n. For example, a set of points, of the portion of the point cloud, closest (e.g., within a threshold distance, a threshold number of points) to the line may be determined. The set of points may be projected on the line and the residual value ares may be determined as the mean component along the line of the projected points. In some examples, the mean may be determined as a weighted mean whose weights depend on the distance of the set of points from the line. For example, a point from the set closer to the line may have a higher weight than another point from the set farther from the line.
[0127] In some examples, the residual value areSmay be quantized. For example, it may be quantized by a uniform quantization function having quantization step similar to the quantization precision of the TriSoup vertices Vk. By doing so, the quantization error may be maintained to be uniform over all vertices Vk and C+Cres such that the local surface is uniformly approximated.In some examples, the residual value aresmay be binarized and entropy coded into the bitstream, e.g., by using a unary-based coding scheme. In some examples, the residual value aresmay be coded using a set of flags. For example, a flag fo may be coded to indicate if the residual value aresis equal to zero. If the flag fo indicates the residual value aresis zero, no further syntax elements may be needed. If the flag fo indicates the residual value aresis not zero, a sign bit indicating a sign may be coded and the residual magnitude |ares|-l may be coded using an entropy code. For example, the residual magnitude may be coded using a unary coding scheme that codes successive flags f (i>l) indicating if the residual value magnitude |ares| is equal to ‘i’. A binary entropy coder may binarize the residual value aresinto the flags f (i>0) and entropy code the binarized residual value as well as the sign bit.
[0128] In some examples, compression of the residual value aresmay be improved by determining bounds as shown in FIG. 8C. As shown, the line (C, n) 810 intersects the current cuboid 800 (corresponding to a TriSoup node) at two bounding points 820 and 821 and the encoder may impose that the adjusted centroid vertex C+Cres is located between the two bounding points 820 and 821. These bounding points 820 and 821 also bound the residual value ares(which may be quantized) as belonging to an integral interval [m, M] where m < 0 < M. By doing so, some bits of the binarized residual value aresmay be inferred. For example, if m=M=0, then residual value aresis necessarily equal to zero. In another example, if m=0<M, then the sign bit is necessarily positive. More generally, if the residual value aresis not equal to zero and its sign is known, its magnitude |ares| may be determined to be bounded by either |m| or M such that the magnitude may be coded by a truncated unary coding scheme that may infer the value of the last of successive flags f (i> 1) .
[0129] In some examples, the binary entropy coder used to code the binarized residual value aresmay be a context-adaptive binary arithmetic coder (CABAC) such that the probability model (also referred to as a context or an entropy coder) used to code at least one bit (e.g., f or sign bit) of the binarized residual value ares are updated depending on precedingly coded bits. In some examples, the probability model of the binary entropy coder may be determined based on contextual information such as the values of the bounds m and M, the position of vertices Vk, or the size of the cuboid. In some examples, the selection of the probability model (i.e., also referred equivalently as an entropy coder or context) may be performed by a dynamic OBUF scheme with the contextual information described above as inputs.
[0130] The reconstruction of a decoded point cloud from the set of TriSoup triangles may be referred to as “voxelization” and may be performed, e.g., by ray tracing or rasterization, for each triangle individually before duplicate voxels from the voxelized triangles are removed.
[0131] FIG. 9A illustrates an example of voxelization using ray tracing, according to some embodiments. For example, ray-triangle intersection algorithms, such as the Moller-Trumbore algorithm, rely on launching rays to determine whether rays intersect with TriSoup triangles and if so, at what points of the TriSoup triangles. Rays may be launched from integral coordinates that correspond to the centers of voxels. As illustrated by FIG. 9A, rays such as ray 900 may be launched parallel to one of the three coordinate axes of the 3D space, starting from integral coordinates (sometimes referred to as integer coordinates) such as an origin point 905 (shown as origin or starting point Pstart).An intersection point 904 (shown as Pint), if any, between ray 900 and a TriSoup triangle 901 belonging to a cube 902, corresponding to a TriSoup node, may be rounded (e.g., quantized) to obtain a decoded point corresponding to a voxel. For example, a ray, launched parallel to a coordinate axis in 3D space, may intersect a TriSoup triangle if and only if the projection, along the ray direction, of the center of a voxel belongs to the TriSoup triangle. In other words, the ray may be determined to intersect the TriSoup triangle if the point of intersection corresponds to the center of the voxel. In some examples, this intersection may be determined by applying a ray-triangle intersection algorithm (e.g., tracing or ray casting technique) such as the Moller-Trumbore algorithm to generate voxels representing the triangle.
[0132] Ray tracing techniques such as the Moller-Trumbore algorithm is based on generating, with respect to a triangle, barycentric coordinates of points of intersection between rays and a plane of the triangle. Then, points of the triangle may be determined from the barycentric coordinates.
[0133] FIG. 9B illustrates an example of voxelization using barycentric coordinates (u, v, w) of a point 912 (P) relative to a TriSoup triangle 910 having vertices labeled A, B, and C in the 3D space, according to some embodiments. In some examples, point 912 may be determined as an intersection between a ray and a plane of TriSoup triangle 910 (e.g., containing or passing through the three vertices A, B, and C of TriSoup triangle 910). For example, the ray may be launched parallel to one of the three coordinate axes in 3D space. In some examples, this intersection point 912 may be uniquely represented as a sum of the three vertices of TriSoup triangle 910:
[0134] P= uA + vB + wC
[0135] under the condition u + v + w = I . Therefore, any point P of the plane (containing Tri Soup triangle 910) has unique coordinates (u,v,w) in the barycentric coordinate system. A point with barycentric coordinates (u,v,w) includes an ordered triple of numbers u, v, and w. A point with barycentric coordinates (u,v,w) that sum to 1 (i.e., u + v + w = 1) is known as homogeneous barycentric coordinates or normalized barycentric coordinates. The barycentric coordinates of the intersection point with respect to TriSoup triangle 910 may be determined using, e.g., the well-known Moller-Trumbore algorithm.
[0136] By converting points with Cartesian coordinates in 3D space to homogeneous barycentric coordinates, the three vertices A, B, C of TriSoup triangle 910 have respective barycentric coordinates A(l,0,0), B(0,l,0) and C(0,0,l). In some examples, the convex hull (i.e., TriSoup triangle 910) of the three vertices A, B, and C is equal to the set of all points such that the barycentric coordinates u, v, and w is each greater than or equal to zero:
[0137] 0 < u, v, w
[0138] Therefore, in some examples, the intersection point may be determined to belong to TriSoup triangle 910 based on the intersection point having barycentric coordinates with an ordered triple of values that is each greater than or equal to zero. Relatedly, if at least one of barycentric coordinates (i.e., one of u, v, or w) is negative or less than 0, then the intersection point may be determined to not belong to TriSoup triangle because it will be on the plane, but not on an edge or within the TriSoup triangle. In some examples, a point determined to belong to TriSoup triangle 910 may be the ray intersecting TriSoup triangle 910 (e.g., within or at an edge of TriSoup triangle 910).Atribute coding is a process to code atributes of a current point cloud, e.g., atributes associated with the geometry of the current point cloud. Atributes coding may be performed globally on the decoded (e.g., reconstructed) geometry of a current point cloud, but such global coding induces high memory traffic and footprint as well as high computation complexity. A two-pass encoding / decoding on the geometry and then on the atributes, after completion of geometry encoding / decoding, is required and induces even higher memory traffic and footprint as well as overall latency before outputing geometry and atributes of a first point of the decoded point cloud.
[0139] Atribute Coding Units (ACU) has been introduced to enable local coding of atributes. ACU may be determined by segmenting an overall decoded geometry (e.g., decoded geometry 1013 of FIG. 10) of a current point cloud (e.g., current point cloud 1011 of FIG. 10) into a set of ACUs. Each ACU comprises (e.g., contains) a portion of geometry of the decoded geometry, which indicates 3D positions in the 3D space of a subset of points of the decoded geometry (e.g., as decoded by the decoder or encoded and then decoded by the encoder). As used herein, points of the decoded geometry may refer to voxels, as described above.
[0140] The atributes encoding / decoding is then localized to portions of the overall decoded geometry and the atribute coding of each ACU (associated with each portion of geometry of the current point) may be processed locally. Memory traffic and footprint as well as computation complexity are then reduced compared to global atribute coding.
[0141] For example, the geometry of the current point cloud may be restricted to a subset of nodes of the occupancy tree and the restricted geometry and the associated atributes may be encoded / decoded locally by segmenting the restricted geometry into ACUs.
[0142] For example, the 3D space encompassing a point cloud may be split into regions defined by subsets of nodes of the occupancy tree. The geometry of a first region of the 3D space may be encoded to obtain a first part of the decoded geometry of the current pint cloud that is segmented into a first set of ACUs that are atribute encoded. Then, the geometry of a second region of the 3D space may be encoded to obtain a second part of the decoded geometry of the current point cloud that is segmented into a second set of ACUs that are atribute encoded, etc. The memory footprint and traffic are thus limited within a few regions, due to some neighborhood prediction between regions, and the latency of the point cloud codec is reduced to the time needed for encoding geometry and atributes of a few regions. Smaller regions will lead to smaller memory footprint and traffic, and to shorter latency. In some examples, the regions may be slices (also referred to as bricks) of a volume encompassing / containing the point cloud.
[0143] FIG. 10 illustrates an example process 1000 for encoding geometry and atributes of a current point cloud (1011), according to some embodiments.
[0144] For example, process 1000 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 1010-1030 may represent components within the encoder.
[0145] The current point cloud (1011) may be a point cloud frame of a sequence of point cloud frames of a dynamic point cloud.At block 1010, an encoder may encode into a bitstream (1090) a geometry information (1012) representative of the geometry of the current point cloud (1011). The encoder may obtain a decoded (e.g., reconstructed) geometry (1013) of the point cloud (1011) as discussed above. At block 1020, the encoder may determine at least one portion of geometry (1022) of the decoded geometry (1013). Each portion of geometry (1022) comprises a subset of points of the decoded geometry (1013), e.g., comprising positions in the 3D space of the subset of points.
[0146] In some examples, the decoded geometry (1013) may be segmented into a set of ACUs, and each ACU of the set of ACUs comprises a respective portion of geometry (1022) of the decoded geometry (1013),
[0147] For example, the encoder may further encode, in the bitstream (1090), portion information (1021), for example, as part of the attribute information (1032).
[0148] For example, the portion information (1021) may indicate segmentation selections of the decoded geometry (1013) into the set of ACUs, e.g., the portion information (1021) indicates how the decoded geometry (1013) is segmented into the set of ACUs.
[0149] Attributes of the current point cloud (1011), e.g., attributes associated with the points of the current point cloud, are typically coded after the coding (e.g., including encoding and / or decoding) of the underlying geometry has been performed. If the geometry coding is a lossless coding (e.g., by using an octree scheme), the encoder has direct access to the attribute values associated with the decoded geometry. The attributes associated with each point of the current point cloud are the attributes associated with the corresponding point of the decoded geometry (1013).
[0150] In some examples, if the geometry coding is a lossy coding (e.g., by using a TriSoup scheme), the decoded geometry (1013) differs from the geometry of the current point cloud. In these examples, the attributes (10111) of the current point cloud may be mapped by the encoder from the geometry of the current point cloud (1011) to the decoded geometry (1013) such as to determine mapped attributes (1031) associated with the decoded geometry (1013). For example, the mapped attributes are assigned to each point of the decoded geometry (1013).
[0151] At block 1025, the encoder may determine mapped attributes (1031) of the decoded geometry (1013) by mapping the attributes (10111) of the current point cloud (1011) to the decoded geometry (1013).
[0152] As discussed above, attributes may indicate a property of a point’s visual appearance such as texture, color, material, transparency, reflectance, time stamp, velocity, etc. For attributes that are colors, this attribute mapping performed by the encoder is known as a recoloring process because the colors of the original geometry are used to color (e.g., recolor) the decoded geometry.
[0153] In some examples, attributes comprise colors and the mapped attributes may be determined based on recoloring the attributes.
[0154] In some examples, mapped attributes may be determined based on a k nearest neighbor (KNN) search algorithm (e.g., using a space partitioning algorithm such as a KD Tree search, a Ball / metric Tree search, brute force search, etc.) to determine nearest points from the geometry of the current point cloud(1011) to the decoded geometry (1013). For example, a mapped attribute of a point of the decoded geometry (1013) may be the average attribute values associated with the nearest points of the current point cloud (1011) relative to the point of the decoded geometry (1013).
[0155] In some examples, when the geometry compression is lossless, the decoded geometry (1013) is the same as the geometry of the current point cloud (1011). In these examples, the attribute mapping associates attributes of each point of the current point cloud (1011) to the same point of the decoded geometry (1013).
[0156] At block 1030, the encoder may encode, in the bitstream (1090), the mapped attributes (1031) associated with each portion of geometry (1022). For example, the attributes of points associated with a portion of geometry (1022) may be encoded before encoding the attributes of points belonging to another portion of the decoded geometry (1013). Thus, attribute encoding may be performed locally. The encoder may encode attribute information (1032) representing the encoded attributes.
[0157] For example, attributes associated with portions of the decoded geometry (1013) are encoded based on a scanning order. The encoder may thus distinguish between already coded / decoded portions of the decoded geometry (1013), i.e., portions of the decoded geometry (1013) whose attributes have been encoded / decoded, from other portions of the decoded geometry (1013) whose attributes have not been encoded / decoded yet.
[0158] FIG. 11 illustrates an example process 1100 for encoding attributes associated with a portion of geometry (1022) of the decoded geometry (1013), according to some embodiments.
[0159] For example, process 1100 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 1110-1130 may represent components within the encoder.
[0160] Process 1100 comprises operations of block 1030 that encodes, in a bitstream (1190), mapped attributes (1031) associated with the portion of geometry (1022) as attribute information (1032).
[0161] At block 1110, an encoder selects an attribute coding mode (1112) for encoding the attributes of the portion of geometry (1022). The encoder may further encode, in the bitstream (1190), mode information (1042) that indicates the selected attribute coding mode (1112).
[0162] For example, the attribute information (1032) encoded in bitstream (1190) may comprise the mode information (1042).
[0163] In some implementations of point cloud coding, the selection (block 1110) of the attribute coding mode (1112) by the process 1100 is performed by a Rate-Distortion Optimization (RDO) such as to select the attribute coding mode (1112) as being a candidate attribute coding mode with a smallest cost (Cmode) from a list of candidate attribute coding modes. The cost (Cmode) of a candidate attribute coding mode may be computed as a Lagrange cost, which is a combination of a bitrate Rmode and a distortion Dmode as Cmode = Dmode + * Rmode, where Z>0 is a fixed Lagrange parameter. The bitrate Rmode may be obtained for a candidate attribute coding mode as being a sum of bitrates caused by coding both the mode information (1042) and the attribute information (1032) in the bitstream (1190). The distortion may be obtained for the candidate attribute coding mode by comparing the attributes associated with portion of geometry(1022) against decoded attributes associated with the decoded portion of geometry (1222) and obtained based on the candidate attribute coding mode.
[0164] RDO may put in competition many candidate attribute coding modes including, usually, at least one inter-prediction-based attribute coding modes and at least one intra-prediction-based attribute coding modes.
[0165] An inter-prediction-based attribute coding mode is an attribute coding mode using attribute predictors obtained based on an inter-prediction mode.
[0166] An intra-prediction-based attribute coding mode is an attribute coding mode using attribute predictors obtained based on an intra-prediction mode.
[0167] At block 1120, the encoder determines attribute predictors (1124) of attributes associated with the portion of geometry (1022). For example, for each point of the portion of geometry (1022), one respective attribute predictor is determined for predicting the attribute associated with that point.
[0168] When an inter-prediction-based attribute coding mode is selected, the encoder may obtain, at block 1120, attribute predictors (1124) from an inter-prediction attribute parametric model based on Motion Vector (MV) field (1121), reference point cloud for attributes (1122) and the portion of geometry (1022). Parameters of the inter-prediction attribute parametric model may be set by the selected attribute coding mode (1112).
[0169] When the intra-prediction-based attribute coding mode is selected, the encoder may determine, at block 1120, attribute predictors (1124) from an intra-prediction attribute parametric model based on at least one already-decoded portion (1123) of the decoded geometry (1013). Parameters of the intraprediction attribute parametric model may be set by the selected attribute coding mode (1112).
[0170] At block 1130, the encoder encodes, in the bitstream (1190), the mapped attributes (1031) associated with the portion of geometry (1022) based on the attribute predictors (1124) as attribute information (1032).
[0171] FIG. 12 illustrates an example process 1200 for decoding geometry and attributes of a current point cloud, according to some embodiments.
[0172] For example, process 1200 may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 1210-1230 may represent components within the decoder.
[0173] The decoded point cloud may be a point cloud frame of a sequence of point cloud frames of a dynamic point cloud. The decoded point cloud may comprise decoded geometry (1212) and decoded attributes (1233), each decoded attribute (1233) being associated with one point of the decoded geometry (1212).
[0174] At block 1210, a decoder may obtain a decoded geometry (1212) by decoding geometry information (1211) from a bitstream (1290), as discussed above. For example, the bitstream (1290) is generated by the encoding method of FIG. 10.
[0175] At block 1220, the decoder may determine at least one portion of the decoded geometry (1212).For example, the decoded geometry (1212) may be segmented into a set of ACUs and each ACU of the set of ACUs comprises a respective portion of geometry (1222) of the decoded geometry (1212), e.g., comprising positions in the 3D space of a subset of points of the decoded geometry (1212).
[0176] For example, the decoder may further decode, from the bitstream (1290), portion information (1221), for example, from part of the attribute information (1232).
[0177] For example, the portion information (1221) may indicate segmentation selections of the decoded geometry (1212) into the set of ACUs, e.g., the portion information (1221) indicates how the decoded geometry (1212) is segmented into the set of ACUs.
[0178] At block 1230, the decoder may obtain decoded attributes (1233) associated with each portion of geometry (1212) by decoding attribute information (1232). For example, the attributes of points associated with a portion of geometry (1222) may be decoded before decoding the attributes of points belonging to another portion of the decoded geometry (1212).
[0179] For example, attributes associated with portions of the decoded geometry (1212) are decoded according to a scanning order. The decoder may thus distinguish between already decoded portions of the decoded geometry (1212), e.g., portions of the decoder geometry (1212) whose attributes have been decoded, from other portions of the decoded geometry (1212) whose attributes have not been yet decoded.
[0180] FIG. 13 illustrates an example process 1300 for decoding attributes associated with a portion of geometry (1222) of the decoded geometry (1212), according to some embodiments.
[0181] For example, process 1300 may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 1310-1330 may represent components within the decoder.
[0182] Process 1300 comprises operations of block 1230 that obtains decoded attributes (1233) associated with portion of geometry (1222) by decoding attribute information (1232) from the bitstream 1290.
[0183] At block 1310, a decoder selects an attribute coding mode (1311) for decoding the attributes of the portion of geometry (1222). The decoder may further decode, from the bitstream (1290), mode information (1231) that indicates the selected attribute coding mode (1311).
[0184] For example, the mode information (1231) may indicate either intra-prediction attribute mode or inter-prediction attribute mode.
[0185] At block 1320, the decoder determines attribute predictors (1324) of attributes associated with the portion of geometry (1222). For example, for each point of the portion of geometry (1222), one respective attribute predictor is determined for predicting the attribute associated with that point.
[0186] When an inter-prediction attribute mode is selected, the decoder may obtain, at block 1320, attribute predictors (1324) from an inter-prediction attribute parametric model based on Motion Vector (MV) field (1321), reference point cloud for attributes (1322) and the portion of geometry (1222). Parameters of the inter-prediction attribute parametric model may be set by the selected attribute coding mode (1311).
[0187] For example, FIGS. 13A and 13B described below shows two examples of the inter-prediction attribute parametric model that may be applied to project attributes of the reference point cloud foratributes (1322) onto the portion of geometry (1222) to determine atribute predictors (1324) for atributes of points / vertices of the portion of geometry (1222).
[0188] FIG. 13A illustrates an example of point-to-point projection distance between a point of the portion of (decoded) geometry (1313) and its nearest neighbor points of amotion compensated geometry (1319), according to some embodiments.
[0189] The motion compensated geometry (1319) is obtained by motion compensation of the reference point cloud for atributes (1316) based on MV field (1312).
[0190] In some examples, the reference point cloud for atributes (1316) may be the reference point cloud for atributes (1122, 1322). The portion of the geometry (1313) may be the portion of geometry (1022, 1222).
[0191] For example, MV field (1312) may include a set of motion vectors including a motion vector (MV) used to translate reference point (1315) (rl) of the reference point cloud for atributes (1316) to a reference point (1314) (rl ’) of the motion compensated geometry (1319). Then the nearest point of the motion compensated geometry (1319), which may include reference point (1314), in a neighborhood of the point (1317) of the portion of geometry (1313) may be selected in a search (1318) by minimizing a point-to-point projection distance between each candidate point of the motion compensated geometry (1319) in the neighborhood of the point (1317) and the point (1317). Discontinuities of the motion compensated geometry (1319) may be introduced due to MV field (1312) which may not provide a granular translation of points. For example, MV field (1312) may include one MV applied to reference point cloud for atribute (1316) or an MV determined for a set of cuboids or per cuboid of reference point cloud for atributes (1316), but not per point of reference point cloud for atributes (1316) due to high complexity. Also, since search (1318) is performed on points of the motion compensated geometry (1319), the entirely of the motion compensated geometry (1319) needs to be generated and maintained before search (1318) for each point of the portion of geometry (1313) can be performed.
[0192] For example, the point-to-point projection distance may be calculated between a point of the portion of the portion of geometry (1313) and its nearest neighbor point of a motion compensated geometry (1319).
[0193] In some embodiments, each point-to-point projection distance may be a difference between a point of the reference point cloud for atributes (1316) and its nearest neighbor point of a motion compensated geometry (1319).
[0194] For example, the motion compensated geometry (1319) may be obtained by performing motion compensation of the portion of the geometry (1313).
[0195] In some embodiments, the motion compensated geometry (1319) may be obtained by motion compensation of the decoded geometry of the current point cloud based on MV field.
[0196] FIG. 13B illustrates an example of point-to-point projection distance between a point (1329) of the reference point cloud for atributes (1316) and its nearest neighbor point of a motion compensated geometry (1328), according to some embodiments.The motion compensated geometry (1328) is obtained by motion compensation of the portion of geometry (1313) based on MV field (1327). For example, MV field (1327) may include a motion vector (-MV) that is the inverse (e.g., having an opposite sign) as that in MV field (1312).
[0197] For example, point (1317) of the portion of geometry (1313) may be translated by a motion vector (-MV) of MV field (1327) to determine a motion-compensated position for point (1317) and shown as point (1326) (p’) of the motion compensated geometry (1328). The nearest point of the motion compensated geometry (1328), which may include point (1326), in a neighborhood of the point (1329) of the reference point cloud for attributes (1316) may be selected in a search (1325) by minimizing a point-to-point projection distance between each candidate point of the motion compensated geometry (1328) in the neighborhood of the point (1329) and the point (1329) of the reference point cloud for attributes (1316).
[0198] In some embodiments, the attribute projection quality may be determined based on a comparison of the mapped attributes and the corresponding projected attributes. The attribute projection quality determining may be performed by the comparison of attributes expressed on a same geometry (i.e., the portion of the (decoded) geometry (1313)) to avoid any geometry discrepancy. In some examples, the attribute projection quality may be based on differences between compared attributes. An attribute may be represented as a vector, so an attribute difference may be computed as an attribute distance between vectors representing attributes that are compared. Thus, attribute distance and attribute difference are used interchangeably herein.
[0199] In some embodiments, the attribute projection quality may be based on point-to-point attribute distances calculated for the set of points of the portion of geometry.
[0200] In some embodiments, the attribute projection quality may be an average or a maximum of the point-to-point attributes distances.
[0201] For example, a point-to-point attribute distance may be a difference between attributes associated with a point of the set of points of the portion of geometry and corresponding projected attributes associated with this point.
[0202] Returning to FIG. 13, when an intra-prediction attribute mode is selected, the decoder may determine, at block 1320, attribute predictors (1324) from an intra-prediction attribute parametric model based on at least one already-decoded portion (1323) of the geometry (1222). Parameters of the intraprediction attribute parametric model may be set by the selected attribute coding mode (1311).
[0203] At block 1330, the decoder decodes attribute information (1232) from the bitstream (1290) and obtains the decoded attributes (1233) associated with the portion of geometry (1222) based on the decoded attribute information and the attributes predictors (1324).
[0204] FIG. 14 illustrates an example process 1400 for determining attribute predictors (1431) of attributes associated with a portion of geometry (1421), according to some embodiments.
[0205] For example, process 1400 may be performed by a coder such as an encoder (e.g., encoder 114 of FIG. 1) or a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 1410-1430 may represent components within the coder.The process 1400 is performed identically at both the encoder and the decoder to determine same attribute predictors (1431). The operations of FIG. 14 may be performed separately by the encoder and the decoder.
[0206] Blocks 1410-1420 relate to a process for determining inter-prediction-based attribute predictors and block 1430 relates to a process for determining intra-prediction-based attribute predictors.
[0207] In some examples, the process for determining inter-prediction-based attribute predictors corresponds to block 1120 of FIG. 11 when the selected attribute coding mode (1112) indicates an interprediction attribute mode (inter mode) and to block 1320 of FIG. 13 when the selected attribute coding mode (1112) indicates an inter-prediction attribute mode (inter mode). The process for determining intraprediction-based attribute predictors corresponds to block 1120 of FIG. 11 when the selected attribute coding mode (1112) indicates an intra-prediction attribute mode (intra mode) and to block 1320 of FIG.
[0208] 13 when the selected attribute coding mode (1112) indicates an intra-prediction attribute mode (intra mode).
[0209] For example, the portion of geometry (1421) may be a portion of geometry (1022) such as a portion of the decoded geometry (1013) of FIG. 10 and the attribute predictors (1431) are attribute predictors (1124) of FIG. 11. For example, the portion of geometry (1421) may be the portion of geometry (1222) of the decoded geometry ( 1212) of FIG. 12 and the attribute predictors ( 1431 ) are the attribute predictors (1324) of FIG. 13. Reference point cloud for attributes (1411) may correspond to the reference point cloud for attributes (1122) of FIG. 11 and the reference point cloud for attributes (1322) of FIG. 13.
[0210] Blocks 1410-1420 show operations of the process for determining inter-prediction-based attribute predictors. The process obtains projected attributes (1422) of the decoded geometry (1013, 1212).
[0211] Specifically, part of projected attributes (1422) is used as attribute predictors (1431) for attributes associated with points of portion of geometry (1421).
[0212] As further explained in FIG. 13A and FIG. 13B, in some examples, the process for determining inter-prediction-based attribute predictors includes motion compensation of the geometry of the reference point cloud for attributes (1411) to generate a motion compensated geometry (1413).
[0213] In these implementations, reference point cloud for attributes (1411) may correspond to reference point cloud for attributes (1122) of FIG. 11 and / or to reference point cloud for attributes (1322) of FIG.
[0214] 13.
[0215] At block 1410, the coder (e.g., encoder / decoder) obtains the motion compensated geometry (1413) by performing motion compensation of the geometry of the reference point cloud for attributes (1411) based on MV field (1412) that corresponds to MV field (1121) of FIG. 11 and MV field (1321) of FIG.
[0216] 13.
[0217] The encoder may obtain the MV field (1121) by performing a motion search such that MV field (1121) approximates the 3D motion field of attributes from the reference point cloud for attributes (1411) to the attributes of the current point cloud (1011). The motion search is typically an iterative method that tests locally multiple candidate motion vectors, and selects the candidate motion vector, among the candidate motion vectors, that minimizes a distortion (e.g., cost) between the attributes of the currentpoint cloud (1011) and attributes of the motion compensated geometry (1413) using the candidate motion vector.
[0218] In some embodiments, the distortion, used by the motion search, for a point of the current point cloud (1011) may be determined by comparing the attribute of this point and the attribute of (one of) its closest neighbor in the motion compensated geometry.
[0219] The encoder may encode information that indicates the reference point cloud for attributes (1411) among multiple candidate reference point cloud for attributes. The encoder may further encode the MV field (1121) as geometry information (1012). The decoder may obtain the reference point cloud for attributes (1411) and MV field (1321) by decoding geometry information (1211) and attribute information (1232) from the bitstream (1290).
[0220] At block 1420, the coder determines projected attributes (1422) associated with the decoded geometry (1013, 1212) based on attributes of the reference point cloud for attributes (1411).
[0221] Projected attributes (1422) and attributes of the decoded geometry (1013, 1212) belong to the same geometry and the prediction of the attributes portion of geometry (1022, 1222) based on the projected attributes is much more efficient because the geometry discrepancy has been removed.
[0222] In some examples, the encoder / decoder may generate (e.g., build or configure) an attributes projection model that may be generated from a motion compensated geometry (1413). The attribute projection model may be used to perform attribute projection associated with the motion compensated geometry (1413). For example, the motion compensated geometry (1413) is obtained by performing motion compensation of the geometry of the reference point cloud for attributes (1411) and the attribute projection determines a projection of the motion compensated geometry (1413) onto the decoded geometry (1421). For example, the motion compensated geometry (1413) is obtained by performing motion compensation of the decoded geometry (1421) and the attribute projection determines a projection of the motion compensated geometry (1413) onto the geometry of the reference point cloud for attributes (Mil).
[0223] In other words, since points of motion compensated geometry (1413) represent motion compensated points of reference point cloud for attributes (1411) (or decoded geometry (1421), points of motion compensated geometry (1413) may correspond respectively to points of decoded geometry (1421) (or reference point cloud for attributes (1411), respectively). Therefore, the result of the attributes projection process at block 1420 may be a form of projection of attributes of reference point cloud for attributes (1411) (after applying MV field 1412) onto decoded geometry (1013, 1212).
[0224] In some examples, the attribute projection model may be a data structure used to efficiently perform attributes projection onto a specific position of a point. For example, the data structure may be a spatial-partitioning data structure such as a tree data structure (e.g., a KD tree or an octree). In some examples, the data structure is used to search for a set of one or more points Npts, belonging to the motion compensated geometry (1413), with positions in the neighborhood of a particular point p from the decoded geometry (1421) according to an example and from reference point cloud for attributes 1411 according to another example. In some examples, the set of points Nptsmay be determined to be withinthe neighborhood for point p based on distances of the set of points Nptsfrom the point p being within or less than a threshold value. The threshold value may be a predetermined value or computed by the encoder and signaled to the decoder in the bitstream. In some other examples, the set of points Nptsmay be determined as a number (or quantity) of points with positions that are closest to the point p. The number may be a predetermined quantity or computed by the encoder and signaled to the decoder in the bitstream. The distance may be a Manhattan distance (i.e., LI norm), an Euclidean distance (i.e., L2 norm), Chebyshev distance (i.e., L-infmity or Loo norm), or a Minkowski distance.
[0225] The attributes values associated with each point within the set of points Nptsmay be used at block 1420 for determining projected attributes values for the point p. For example, the projected attributes values may be attribute predictors of attributes for the point p. In some examples, a value of each attribute of point p may be predicted or projected based on values of corresponding attributes (with the same type as the each attribute) of the set of one or more points Npts.
[0226] In some examples, a projected attribute (e.g., attribute predictor) for the point p may be determined as an average of the attribute values, of the set of one or more points Npts, weighted by respective distance between the points Nptsand the point p. The distance may be a Manhattan distance (i.e., LI norm), an Euclidean distance (i.e., L2 norm), Chebyshev distance (i.e., L-infmity or Loo norm), or a Minkowski distance.
[0227] For example, the projected attributes representing the attribute predictors for the point p may be determined to be the corresponding values of the attributes of the point from the motion compensated geometry (1413) closest to the point p. These examples enable faster projection because only one point is searched and selected from motion compensated geometry (1413) to determine an attributes predictor for point p from decoded geometry (1013, 1212).
[0228] At block 1430, the coder may determine attribute predictors (1431) for the portion of geometry (1022, 1222) based on intra-prediction attribute mode using at least one already-coded portion (1432).
[0229] For example, the already-coded portions (1432) may be the already-coded portions (1123) of FIG.
[0230] 11 or the already-coded portions (1323) of FIG. 13.
[0231] The encoder / decoder may generate (e.g., determine) the intra-prediction attribute mode based on the intra prediction mode (selected attribute coding mode 1112, 1311).
[0232] For example, the attribute predictors (1431) may be obtained by extrapolating at least one attribute associated with at least one already-coded portions (1432).
[0233] For example, the intra prediction mode (selected attribute coding mode (1112, 1311)) may indicate that the at least one already-coded portion (1432) may comprise at least one spatial neighbor portion of the portion of geometry (1421).
[0234] In some examples, the portion of geometry (1421) may be encompassed by a current ACU and the spatial neighbor portion of the portion of geometry (1421) may comprise points encompassed by an ACU having a part of its boundary overlapping with at least a portion of a boundary of the current ACU. For example, when ACUs are cuboids in shape, boundary may be defined as faces, edges and vertices of thecuboid. Sharing a part of the boundary may be defined as having a common face, a common edge, or a common vertex (e.g., comer).
[0235] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate that the attribute predictors (1431) may be obtained based on extrapolation of at least one attribute associated with the at least one spatial neighbor portion of the portion of geometry (1421).
[0236] For example, extrapolation of at least one attribute associated with the at least one spatial neighbor portion of the portion of geometry (1421) may be determined by fitting a 3D attribute model forthe at least one attribute associated with the at least one spatial neighbor portion of the portion of geometry (1421), and by extending (e.g., extrapolating) the fitted 3D attribute model to the portion of geometry (1421). A 3D attribute model may take spatial coordinates as input and provide modeled attributes as output; model parameters are fit (e.g., learned) on at least one attribute associated with the at least one spatial neighbor portion of the portion of geometry (1421).
[0237] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate attribute predictors (1431) are determined based on averages of attributes associated with the spatial neighbor portion of the portion of geometry (1421).
[0238] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate attribute predictors (1431) are determined based on a maximum of attributes associated with the spatial neighbor portion of the portion of geometry (1421).
[0239] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate a spatial direction along which the spatial neighbor portion of the portion of geometry (1421) are selected.
[0240] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate a maximum number of spatial neighbor portions of the portion of geometry (1421) are selected.
[0241] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate a maximum distance between the spatial neighbor portion of the portion of geometry (1421) and the portion of geometry (1421). For example, the maximum distance may be between any point of the geometry (1421) and any other point of the portion of geometry (1421). For example, the maximum distance may be between centers (or opposite comers) of the spatial neighbor portion of the portion of geometry (1421) and the portion of geometry (1421).
[0242] In some examples, the intra prediction mode (selected attribute coding mode 1112, 1311) may indicate an index of / to a list of indices, each index indicating one of the above intra-prediction attribute modes or one of their combinations.
[0243] FIG. 15 illustrates an example process 1500 for encoding attributes based on attribute predictors (1124) associated with portion of geometry (1022), according to some embodiments.
[0244] For example, process 1500 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 1510-1530 may represent components within the encoder.
[0245] The encoder may determine residual attributes (1511) by subtracting the attribute predictors (1124) from the mapped attributes (1031) associated with portion of geometry (1022) .At block 1510, the encodermay determine transformed coefficients (1512) by applying a 3D transform to the residual attributes (1511) based on portion of geometry (1022).
[0246] At block 1520, the encoder may determine quantized residual attributes or quantized coefficients (1521) by quantizing the residual attributes (1511) or the transformed coefficients (1512), respectively.
[0247] At block 1530, the encoder may entropy encode the residual attributes (1511) or the transformed coefficients (1512) or the quantized residual attributes or the quantized coefficients (1521) as attribute information (1032).
[0248] FIG. 16 illustrates another example process 1600 for encoding attributes based on attribute predictors (1124) associated with portion of geometry (1022), according to some embodiments.
[0249] For example, process 1600 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 1610-1640 may represent components within the encoder.
[0250] At block 1610, the encoder may determine attribute coefficients (1611) by applying a 3D transform to the attribute predictors (1124) based on portion of geometry (1022).
[0251] At block 1620, the encoder may determine mapped attribute coefficients (1621) by applying 3D transform to the mapped attributes (1031) based on portion of geometry (1022).
[0252] The encoder may determine residual coefficients (1631) by subtracting the attribute coefficients (1611) from the mapped attribute coefficients (1621).
[0253] At block 1630, the encoder may determine quantized coefficients (1641) by quantizing the residual coefficients (1631).
[0254] At block 1640, the encoder may entropy encode the residual coefficients (1631) or the quantized coefficients (1641) as attribute information (1032).
[0255] FIG. 17 illustrates an example process 1700 for decoding attributes based on attribute predictors (1324) associated with portion of geometry (1222), according to some embodiments.
[0256] For example, process 1700 (e.g., corresponding to block 1330 of FIG. 13) may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 1710-1730 may represent components within the decoder.
[0257] At block 1710, the decoder may entropy decode from the bitstream (1290) quantized coefficients (1711) from attribute information (1232).
[0258] At block 1720, the decoder may determine residual coefficients (1721) by inverse quantizing the quantized coefficients (1711).
[0259] In some examples, the decoder may entropy decode from the bitstream (1290) residual coefficients (1721) from the bitstream (1290).
[0260] At block 1730, the decoder may determine residual attributes (1731) by applying inverse 3D transform to the residual coefficients (1721) based on portion of geometry (1222).
[0261] The decoder may determine decoded attributes (1233) by adding the residual attributes (1731) with the attribute predictors (1324).
[0262] FIG. 18 illustrates another example process 1800 for decoding attributes based on attribute predictors (1324) associated with portion of geometry (1222), according to some embodiments.2025P00765WG
[0263] For example, process 1800 (e.g., corresponding to block 1330 of FIG. 13) may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 1810-1840 may represent components within the decoder.
[0264] At block 1810, the decoder may entropy decode from the bitstream (1290) quantized coefficients (1811) from attribute information (1232).
[0265] At block 1820, the decoder may determine residual coefficients (1821) by inverse quantizing the quantized coefficients (1811).
[0266] In some examples, the decoder may entropy decode from the bitstream (1290) residual coefficients (1821).
[0267] At block 1830, the decoder may determine attribute coefficients (1831) by applying 3D transform to the attribute predictors (1324) based on the portion of geometry (1222).
[0268] The decoder may determine decoded attribute coefficients (1841) by adding the residual coefficients (1821) with the attribute coefficients (1831).
[0269] At block 1840, the decoder may determine decoded attributes (1233) by applying an inverse 3D transform to the decoded attribute coefficients (1841) based on portion of geometry (1222).
[0270] In some examples (e.g., used in G-PCC), the 3D transform that may be used / selected for transforming residual attributes (1511) of FIG. 15, mapped attributes (1031) and attribute predictors (1124) of FIG. 16 and attribute predictors (1324) of FIG. 18, may be the region-adaptive hierarchical transform (RAHT) scheme. Inverse 3D transform of this 3D transform may be used / selected for inverse transforming residual coefficients (1721) of FIG. 17 and decoded attributes coefficients (1841) of FIG.
[0271] 18.
[0272] The RAHT scheme is based on the iterative use of a two-point transform. In the framework of point cloud attribute coding, the two-point RAHT transform is to be understood as being applied to two sets Ai and A2 of attributes having respectively wi and W2 number of attributes and respective associated coefficients CAI and CA2 representative of the sum of attribute values over their respective set divided by the square root of the number of attributes.
[0273] CAI = -J=£aeAi a- (wt= #A) (*)
[0274]
[0275] The two-point RAHT transform depends on the weights wi and W2 and is defined by a 2x2 matrix as follows
[0276] RAHT(w1, w2)
[0277]
[0278] When applied to the two coefficients CAI and CA2, two new coefficients DC and AC are determined.
[0279] □ = RAHT(W1,w2) ^] (1)
[0280]
[0281] As illustrated below, the above property (*) on coefficients still holds for the DC coefficient.DC =1^cA1+ ^2cA2) =1I y a + y a ) y / w1+ w2\wr+ w2\ Z— i Z— i /
[0282] \ aEA1aEA2 /
[0283] 1v
[0284] — I — I - / A—CT41UT42
[0285] VW-L + W2Z— I
[0286]
[0287] ClEA I 0^2
[0288] The two-point RAHT transform may be applied iteratively to DC coefficients. This is the RAHT iterative method. Once determined, AC coefficients do not undergo any further transformation. At the start of the RAHT iterative method, there are as many initial sets A! of attributes as there are points in the coded geometry S. Each initial set A! of attributes thus contains one attribute (wi=l) and the coefficient CAI is equal to the value of this one attribute, thus fulfilling the property (*). By induction, the property (*) holds for all subsequent DC coefficients determined after iterative application of the two-point RAHT transform. In some examples, the points of the coded geometry S may refer to voxels generated / determined to represent points of the coded geometry S.
[0289] Therefore, at any stage of the RAHT iterative method, determined coefficients are the union of a set of DC coefficients fulfilling the property (*) and a set of AC coefficients. The RAHT iterative method may continue until DC coefficients are depleted and only one DC coefficient is left. In this case, this one DC coefficient is equal to CA where A is the set of all attributes to be transformed. The RAHT iterative method may a priori follow any order among pairs of DC coefficients.
[0290] For example, the attributes to be transformed may be residual attributes (1511), attribute predictors (1124), mapped attributes (1031) or attribute predictors (1324) and the one obtained DC coefficient CA may be transformed coefficients (1 12), attributes coefficients (1611), mapped attributes coefficients (1621) or attributes coefficients (1831) respectively.
[0291] The two-point inverse RAHT transform may be defined by a 2x2 matrix as follows
[0292]
[0293]
[0294] and is applied to DC and AC coefficients such as to obtain (e.g., recover or reconstruct) the two coefficients CAI and CA2.
[0295]
[0296] E^ iRAHTOvv^) ^]
[0297] The inverse iterative RAHT method applies the inverse two-point RAHT to DC and AC coefficient in reverse order relative to their obtainment by the iterative RAHT scheme. At the end of the inverse iterative RAHT scheme, coefficients CAI associated with the initial sets A! of attributes are obtained. These coefficients CAI are equal to the values of the one attributes associated with the initial sets A^ For example, attributes associated with the initial sets A! may be the residual attributes (1731) or decoded attributes (1233).
[0298] In some examples (e.g., such as in G-PCC), the RAHT iterative scheme may follow an octree in a specific iterative order. Basically, the up to eight DC coefficients associated with the up to eight occupied child nodes of a parent node in the octree undergo a cascade of two-point RAHT transformations until2025P00765WG
[0299] one DC coefficient remains, together with up to seven AC coefficients. This one DC coefficient is pushed at the parent node level and the scheme is repeated at upper octree depth (e.g., at lower depth indexes) until the root node is reached.
[0300] FIG. 19 illustrates an example RAHT transformation, of the RAHT scheme, applied on child nodes of an octree parent node along three successive directions, according to some embodiments.
[0301] The parent node 1900 has five occupied child nodes with associated coefficients Ci and weights Wi. A first RAHT transformation 1910 is performed along a first direction 1911. If there are two adjacent occupied child nodes 1913 along this direction, they undergo a two-point RAHT transform to determine a new DC coefficient 1914 and an AC coefficient 1915 pushed to a set 1950 of AC coefficients. If there is only one occupied child node 1916 along this direction, the node is left as is and its DC coefficient is kept 1917. By doing so, the child nodes are collapsed along the first direction to determine a new set 1919 of nodes, here a set of three nodes, with associated new DC coefficients. Then, a second RAHT transformation 1920 is performed along a second direction 1921 in a similar way to determine child nodes 1922, that have been collapsed along the first two directions 1911 and 1921, together with AC coefficients 1923 pushed to the set 1950 of AC coefficients. Finally, a third RAHT transformation 1930 is performed along a third direction 1931 in a similar way to determine a unique child node 1932, resulting from the collapse along all three directions, together with AC coefficients 1933 pushed to the set 1950 of AC coefficients.
[0302] The unique collapsed child node 1932 has an associated DC coefficient that is pushed to the parent node as illustrated in FIG. 20.
[0303] FIG. 20 illustrates an example RAHT transformation being applied to all octree nodes at depth ‘d’ to determine DC coefficients at depth d-1 and AC coefficients, according to some embodiments. Occupied nodes 2000 of an octree at depth ‘d’ are illustrated. These nodes undergo a RAHT transformation along the three directions such as to push DC coefficients up to their occupied parent nodes 2010 belonging to the octree at depth d-1. For example, the three DC coefficients of the child nodes 2001 undergo a RAHT transformation along the three directions to determine a unique DC coefficient associated with their parent node 2011 and two AC coefficients 2021 pushed to a set 2020 of AC coefficients. By performing this method for all occupied nodes 2000 of the octree at depth ‘d’, the DC coefficients associated with occupied nodes of the octree at depth ‘d’ are transformed into DC coefficients associated with occupied nodes 2010 of the octree at depth d-1 and a set 2020 of AC coefficients.
[0304] This bottom-up method may be repeated depth per depth until reaching the minimum depth (the root node) and the result of the RAHT transformation over the complete octree is a set of coefficients comprising a unique DC coefficient and a set of (many) AC coefficients.
[0305] The RAHT transformation method typically starts from the highest / deepest depth (e.g., farthest from the root node with a low depth index) where occupied child nodes correspond to a unique point (voxel) of the coded geometry S associated with a unique attribute among the set ‘a’ of attributes. The DC coefficient at highest / deepest depth is thus set as the value of the unique attribute associated with each occupied node and the weights ‘w’ are set to 1.The inverse RAHT method on an octree is a top-down method from the root node (with lowest depth index or with most shallow depth) down to the last depth (with highest depth index or with deepest depth) made of leaf nodes that each contain only one point (voxel) of the point cloud, thus only one associated attribute. The DC coefficients of occupied nodes 2010 of the octree at depth d-1 are inverse transformed into DC coefficients of occupied nodes 2000 of the octree at depth ‘d’ by applying the inverse two-point RAHT transform to the DC coefficient of each of the occupied node of the octree at depth d-1 and to the related AC coefficients from set 2020 of AC coefficients. The inverse two-point RAHT transform is applied along the three directions, in reverse order, such as to invert the node transformation process of FIG. 19. By doing so, DC coefficients of the leaf nodes are obtained, and their values correspond to the attributes associated with the unique point of each of the leaf nodes.
[0306] Like geometry coding of a point cloud, coding of attributes associated with the points of a current point cloud may benefit from inter-frame prediction using a motion compensated point cloud. The motion compensated point cloud inherits attributes from a reference point cloud that has been motion compensated. For example, during motion, points keep their associated attributes. The motion compensated attributes, e.g., the attributes associated with the points of the motion compensated point cloud, may be used to better compress the attributes of the coded geometry of the current point cloud.
[0307] Inter RAHT scheme may be used at blocks 1610, 1620, 1730, 1830, and 1840. Inter RAHT scheme defines an inter prediction mode that uses inter prediction for predicting the values of the DC and the AC coefficients determined by the RAHT iterative method or the inverse RAHT iterative method. In some examples, because the generation of DC and AC coefficients follows an octree, it may be beneficial to maintain a common attribute octree structure for both the portion of geometry (1022, 1222) and a motion compensated portion of geometry. A common bounding box encompassing both portion geometries may be determined, and an octree partitioning may be performed, from a root node associated with the common bounding box, for both portion geometries. This leads to two octree partitioning that are different when the point geometries are not equal, which is likely. The two octrees have a common subtree starting from the root node. On this subtree, occupied node topology is the same and a common set of DC and AC coefficients is determined for both portion geometries. Thus, the subset of DC and AC coefficients associated with nodes of the common subtree and determined from the attributes of the portion of geometry (1022, 1222) may be predicted from DC and AC coefficients determined from the attributes of the motion compensated portion of geometry. Practically, the encoder and / or decoder may determine coefficient residual values by subtracting the DC and AC coefficients determined from the attributes of the motion compensated portion of geometry from the DC and AC coefficients associated with nodes of the common subtree and determined from the attributes of the portion of geometry (1022, 1222).
[0308] The DC and AC coefficients that are not associated with nodes of the common subtree may not be predicted and may be transformed directly in a similar way as performed for the case without inter prediction.In some examples, instead of predicting AC coefficients, predicted DC coefficients of the portion of geometry (1022, 1222) may be determined at some depth, assuming both the octree of the portion of geometry (1022, 1222) and the octree of the motion compensated portion of geometry have a same occupancy of a node at this depth. The predicted DC coefficients may be determined from their colocated DC coefficients of the motion compensated portion of geometry. DC residual values may be determined by subtracting the predicted DC coefficients from the DC coefficients of the portion of geometry (1022, 1222). The RAHT transformation then goes up in the octree starting from DC residual values replacing the DC coefficients of the portion of geometry (1022, 1222).
[0309] A RAHT scheme process that does not use information from a reference point cloud different from the current point cloud is called an intra RAHT scheme. Intra RAHT scheme defines an intra prediction mode for predicting the values of the DC and the AC coefficients determined by the RAHT iterative method or the inverse RAHT iterative method.
[0310] Intra prediction may be performed between DC and AC coefficients of an intra RAHT scheme and is referred to as inter-depth prediction. In some implementations of point cloud coding, this interdepth prediction within portion of geometry (1022, 1222) has been integrated into the RAHT scheme as being an intra prediction mode.
[0311] The inter-depth prediction mode may predict the DC coefficients associated with a current RAHT node of a RAHT tree at depth ‘d’ by using interpolation of DC coefficients associated with nodes of the octree at lower depth d-1 (e.g., lower depth index corresponding to shallower depths of the RAHT tree). Lower depth indicates a depth level closer to root node of RAHT tree and higher depth indicates a depth level closer to leaf nodes of RAHT tree (i.e., farther from root node). The inter-depth prediction mode may also predict AC coefficients of a current RAHT node by using attribute information from a parent node and attribute information from already-coded neighboring RAHT nodes of the parent node of the current RAHT node.
[0312] A RAHT tree comprises RAHT nodes that are linked together according to a parent-child relationships defined as an occupancy tree over the decoded geometry. A RAHT node of the RAHT tree is a node of the occupancy tree that is associated with a DC coefficient and possibly one or more AC coefficients obtained by applying the iterative RAHT (encoder) or inverse iterative RAHT method (decoder) to the points belonging to the sub-volumes associated with the occupied nodes of the occupancy tree.
[0313] For example, already-coded neighboring RAHT nodes of the parent node of the current RAHT node may include nodes that are siblings to the parent node, e.g. the seven siblings in an octree structure. In some examples, already-coded neighboring RAHT nodes may include nodes that share a portion of a face, a portion of an edge, or a vertex (e.g., comer) with the parent node.
[0314] In some examples, the inter-depth prediction as discussed above may be implemented in a bounded domain such as, for example, the mean attribute domain which is naturally bounded by the attribute value range. The bounded property of the mean attribute domain is advantageous as it correlatesto a more physical meaning and provides better numerical stability of the inter-depth prediction, thus leading to a more efficient prediction.
[0315] Mean sums of attributes values ai,a calculated for RAHT nodes at a parent RAHT node depth ‘d’ may then be used to predict a mean sum of attribute values ai,cassociated with each child RAHT node (at depth d+1) of the parent RAHT node. The mean sums of attributes values ai,a calculated for RAHT nodes at a parent node depth d may comprise a mean sum of attributes values a;,a calculated for the parent RAHT node and possibly a mean sum of attributes values ai, calculated from one or more already-coded neighboring RAHT nodes of the RAHT parent node.
[0316] FIG. 21A illustrates an example process 2100 for encoding attribute information for child RAHT node(s) (2110) (at depth d) of a parent RAHT node (2111) (at depth d-1) using top-down coding and inter-depth prediction, according to some embodiments.
[0317] For example, process 2100 may be performed by an encoder (e.g., encoder 114 of FIG. 1).
[0318] In this example, a set of already-coded neighboring RAHT nodes (2112) of the parent RAHT node (2111) may include RAHT nodes of the RAHT tree at depth d-1 that share at least a vertex with the parent RAHT node (2111).
[0319] The encoder determines the DC coefficients CAi,a for the parent RAHT node (2111) and each of the already-coded neighboring RAHT nodes (2112) and calculates a mean sum of the attributes value ai,a for the parent RAHT node (2111) and a mean sum of attributes value a;,a for each of the already-coded neighboring RAHT nodes (2112) by dividing the coefficients CAi,a of each of those RAHT nodes by the square root of its corresponding number of attributes fwi. For example, a mean sum of attributes ai,a may , . . . CAjd XaeA:
[0320] be obtained asd=a
[0321]
[0322] ywi = - wt— .
[0323] The encoder may then obtain a predicted mean sum of attributes ai,c,uPfor each child RAHT node (2110) by up-sampling the mean sums of attributes values ai,a calculated for the RAHT nodes at the parent node depth ‘d’.
[0324] Then, the encoder determines a predicted DC coefficient CAi,c,prea for each child RAHT node (2110) by multiplying the predicted mean sum of attributes ai,c,uPcalculated for each child RAHT node (2110) by the square root of its corresponding number of attributes fwi.
[0325] The encoder then determines, based on equation 1, predicted AC coefficients from predicted DC coefficient CA prea and determines residual AC coefficients by subtracting the predicted AC coefficients from original AC coefficients associated with each child RAHT node (2110). Original AC coefficients associated with a child RAHT node (2110) are determined by the usual RAHT transform from original attributes to be coded. The residual AC coefficients may be encoded in a bitstream, for example through quantization (blocks 1520, 1630) and entropy coding (block 1530, 1640).
[0326] FIG. 2 IB illustrates an example process 2100B for decoding attribute information for child RAHT nodes, at depth d, of a parent RAHT node at depth d-1 using top-down decoding and inter-depth prediction, according to some embodiments. When inter-depth prediction is used in the RAHT transformand coding, the decoder performs top-down decoding because the inter-depth prediction goes from depth ‘d’ to depth d-1.
[0327] For example, process 2100B may be performed by a decoder (e.g., decoder 120 of FIG. 1).
[0328] The decoder employs the same inter-depth prediction process to generate the predicted mean sum of attributes values CAi,c,pred for each child node 2110 and the predicted AC coefficients. It also reconstructs the residual AC coefficients from a bitstream, for example through entropy decoding (blocks 1710, 1810) and inverse quantization (block 1720, 1820). The decoder then determines decoded AC coefficients by adding the predicted AC coefficients to the reconstructed residual AC coefficients.
[0329] FIG. 22 illustrates an example process 2200 for up-sampling the mean sums of attributes values of RAHT nodes at depth ‘d-1’, such as including the parent RAHT node (2111) and the already-coded neighboring RAHT nodes (2112), according to some embodiments. For clarity and ease of explanation, this example is illustrated in two dimensions, but extension to three dimensions will be understood in light of the description herein.
[0330] In this example, the up-sampling operation considers (e.g., uses) a distance metric dk relating the child RAHT node (2110) to a RAHT node k at depth ‘d-1’, e.g., either the parent RAHT node (2111) or each of the three (in this example) already-coded neighboring RAHT nodes (2112). This distance metric dk may represent a geometric distance between a center point of the sub-volume corresponding to the child RAHT node (2110) and a center point of the sub-volume corresponding to the RAHT node k (2111, 2112). The inverse dk1of the distance metric dk may represent the relative weight of correlation between attribute information from the child RAHT node (2110) and the RAHT node k. Other weighting factors, or additional weighting factors, may be used in other implementations of the up-sampling operation.
[0331] In some examples, the predicted mean sum of attributes ai,c,uPfor the child RAHT node (2110) may be given by the weighted sum:
[0332] _ V-1>
[0333] ai,c,up
[0334]
[0335] K
[0336] where akindicates the mean sum of attributes calculated for the RAHT node k.
[0337] When neither intra (e.g., inter-depth prediction) nor inter prediction are used, AC coefficients are directly coded. The direct coding of AC coefficients is equivalent to coding residual AC coefficients relative to a null predictor and therefore may be referred to as a null prediction mode for coding the current RAHT node.
[0338] As explained above, RAHT coefficients of RAHT nodes may be more efficiently encoded and decoded using prediction modes for coding attributes of the RAHT nodes. For example, a prediction mode for a RAHT node of the RAHT tree may be a null prediction mode or one of an intra or inter prediction modes, described above with respect to FIG. 14. Indications of prediction modes may also be entropy coded in a bitstream to improve coding performance.
[0339] FIG. 23 illustrates an example process 2300 for encoding attributes of a RAHT node (2301) using a prediction mode, according to some embodiments. For example, process 2300 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 2310-2340 may represent componentswithin the encoder. The encoder may transform attributes (2305) of the RAHT node (2301) to obtain sets {Ck} of residual AC coefficients (2331), resulting from applying a prediction mode (2321), that are quantized to obtain sets {Ck} (2336) of quantized residual AC coefficients that are encoded into a bitstream (2390).
[0340] At block 2310, the attributes (2305) of a RAHT node (2301) are RAHT transformed (e.g., via a forward RAHT transform) into sets of RAHT coefficients (2311).
[0341] At block 2320, the prediction mode (2321) is obtained for the RAHT node (2301). The prediction mode (2321) may be inferred or selected from a set of prediction modes.
[0342] In some examples, the set of prediction modes may be from a group including an inter and intra prediction mode. In some examples, the set of prediction modes may be from a group including the intra prediction mode, inter prediction mode, and the null prediction mode.
[0343] FIG. 24 illustrates an example process 2400 for performing block 2320 related to obtaining the prediction mode (2321) for encoding attributes of the RAHT node (2301), according to embodiments. For example, process 2400 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 2410-2440 may represent components within the encoder.
[0344] After sets of RAHT coefficients (2311) have been obtained at block 2310, the encoder makes an inference decision (block 2410) that indicates whether the prediction mode (2321) is either an inferred prediction mode pmf (2441) (e.g., output at block 2440) or a selected prediction mode psei(2421) selected from a set of prediction modes (block 2420).
[0345] For example, at block 2410, the encoder may make the inference decision based on neighboring residual RAHT coefficients (2351) such as residual RAHT coefficients associated with at least one already-coded neighboring RAHT node (2350) of the RAHT node (2301).
[0346] For example, the inference decision may be based on information indicating the magnitudes (mk,Y, mk,u, and nik,v) of the neighboring residual RAHT coefficients (2351).
[0347] In some examples, residual RAHT coefficients include a series {Ck}k=i, ..., K of K sets of component coefficients to represent attributes associated with the RAHT node (2301). If the attributes to be coded are color attributes, each set of coefficients can include three component coefficients for the respective three color components of a color attribute of a voxel or a point, one for luma and two for chroma. For example, the color space may be YCbCror YUV. Denote the luma component of set Ck of coefficients as Ck.Y, and the two chroma components as Ck,uand Ck,v, respectively. Ck can be represented as
[0348] Ck= {Ck,y, Ck,u, Ck,v }.
[0349] Each component coefficient Ck,Y (respectively Ck,u and Ck,v) may be represented by a sign Sk,Y (respectively Sk,u and Sk,v) and a series of bits representing its magnitude (absolute value) mk,y (respectively mk,u and nik,v).
[0350] In some embodiments, the inference decision may be made based on the information indicating the magnitudes being less than or equal to a threshold. Small magnitudes may indicate that neighboring prediction modes (2352) (e.g., the prediction modes associated with at least one already-codedneighboring RAHT nodes (2350) of the RAHT node (2301)) are likely to be accurate and therefore the prediction mode (2321) may be inferred (block 2440) from the neighboring prediction modes (2352).
[0351] In some embodiments, the inferred prediction mode pmf (2441) may be obtained at block 2440 based on determining a most common / used prediction mode among the neighboring prediction modes (2352). The determination / selection of the most common / used prediction mode may be referred to as a majority voting process.
[0352] In some embodiments, the selection (block 2420) of the selected prediction mode psei(2421) from the set of prediction modes may be based on rate distortion optimization (RDO) costs of the set of prediction modes, e.g., selecting the prediction mode with the lowest RDO cost.
[0353] The inference decision indicates the prediction mode (2321) may be either the inferred prediction mode pinf (2441) or the selected prediction mode psei(2421). If the inference decision indicates that the prediction mode (2321) is the selected prediction mode psei(2421), the encoder may entropy encode (block 2430) an indication of the selected prediction mode psei(2421), as part of prediction mode information (2322) in bitstream (2390). If the inference decision indicates that the prediction mode (2321) is the inferred prediction mode pmf (2431), the indication is skipped or omitted from being signaling as part of prediction mode information (2322). In some examples, the indication is not needed in bitstream (2390) because the decoder may also make the same inference decision and infer the same inferred prediction mode pmf (2441) as the encoder, as further described below with respect to FIG. 25.
[0354] In some examples, the encoder may entropy encode the indication of the selected prediction mode psei(2421) based on a context selected from a set of contexts. For example, the context selection is based on neighboring prediction modes (2352) such as prediction modes (2352) associated with at least one already-coded neighboring RAHT node (2350) of the RAHT node (2301).
[0355] In some embodiments, the obtained prediction mode (2321) may be associated with the RAHT node (2501) and may be used for encoding attributes of further RAHT nodes.
[0356] At block 2330, the prediction mode (2321) is applied to the RAHT node (2301) to obtain (e.g., compute or derive) sets {Ck} of residual RAHT coefficients (2331) from the sets of RAHT coefficients (2311). For example, the encoder may use the prediction mode (2321) to generate a predictor that is subtracted from the set of RAHT coefficients (2311) to obtain the sets {Ck} of residual RAHT coefficients (2331).
[0357] In some embodiments, generating (constructing) the predictor may be performed based on the neighboring prediction modes (2352). For example, the constructed predictor may be an average predictor which may be a linear combination of an intra predictor derived from an intra prediction mode and an inter predictor derived from an inter prediction mode, with the weights in the linear combination being determined based on neighboring prediction modes (2352).
[0358] At block 2335, the sets {Ck} of residual RAHT coefficients (2331) are quantized to obtain sets {Ck} of quantized residual RAHT coefficients (2336). Applying quantization at block 2335 leads to lossy compression that may allow for massively reducing the size of the bitstream (2390) at the cost ofdegradation (e.g., increasing distortion) of the attribute signal. The encoding process (2300) is designed to optimize the balance between bitrate reduction and introduced distortion.
[0359] At block 2340, the sets {Ck} of quantized residual RAHT coefficients (2336) are entropy encoded into the bitstream (2390) as coefficient information (2342).
[0360] In some examples, the encoding (block 2340) may include encoding a series of flags {fk} indicating if the sets {Ck} of quantized residual RAHT coefficients (2336) are zero, and then values of Ck,y, Ck,u and Ck,v for example by encoding flags (fk,Y, fk,u, and fk,v) indicating if respective values are zero, signs (sk,Y, Sk,u and Sk,v) and magnitudes (absolute values) (n , mk,u and m v) for non-zero component coefficients in each of the sets. For example, each of the flags {fk}indicates whether a respective set {Ck} of the sets of quantized residual RAHT coefficients (2336) is zero.
[0361] For example, the encoder may entropy encode the series of flags {fk} and / or the sets {Ck} of quantized residual RAHT coefficients (2331) based on a context selected from a set of contexts. For example, the context selection is based on neighboring prediction modes (2352).
[0362] FIG. 25 illustrates an example process for decoding attributes of a RAHT node (2501) using a prediction mode, according to some embodiments. For example, process 2500 may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 2510-2540 may represent components within the decoder. The bitstream (2590) is typically obtained from the encoding method illustrated in FIG. 23. The decoder may decode sets {Ck} of quantized residual RAHT coefficients (2531) from the bitstream (2590) to obtain sets {Ck} of dequantized residual RAHT coefficients (2536), determine sets {Ck} of RAHT coefficients based on the sets {Ck} of dequantized residual RAHT coefficients (2536) and apply an inverse RAHT transform (e.g., at block 2510) to obtain decoded attributes (2505) of a RAHT node (2501).
[0363] At block 2520, a prediction mode (2521) is obtained for the RAHT node (2501). The prediction mode (2521) may be inferred or selected from a set of prediction modes.
[0364] FIG. 26 illustrates an example process 2600 for performing block 2520 related to obtaining the prediction mode (2521) for decoding attributes of the RAHT node (2501), according to embodiments. For example, process 2600 may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 2610-2630 may represent components within the decoder.
[0365] At block 2610, the decoder makes an inference decision that indicates whether the prediction mode (2521) is either the inferred prediction mode pmf (2631) (e.g., obtained at block 2630) or the selected prediction mode psei(2621) selected from a set of prediction modes (e.g., selected at block 2620).
[0366] For example, at block 2610, the decoder may make the inference decision based on neighboring residual RAHT coefficients (2351) such as residual RAHT coefficients associated with at least one already-coded neighboring RAHT node (2350) of the RAHT node (2501).
[0367] For example, the inference decision may be based on information indicating the magnitudes (mk,Y, mk,u, and mk,v) of neighboring residual RAHT coefficients (2351). In some embodiments, the inference decision is made based on the information indicating the magnitudes being less than or equal to a threshold. Small magnitudes may indicate that neighboring prediction modes (2352) (e.g., the predictionmodes associated with at least one already-coded neighboring RAHT nodes (2350) of the RAHT node (2501)) are likely to be accurate and therefore the prediction mode (2521) may be inferred (block 2630) from the neighboring prediction modes (2352).
[0368] In some embodiments, the inferred prediction mode pmf (2631) obtained at block 2630 may be based on determining a most common / used prediction mode among neighboring prediction modes (2352). The determination / selection of the most common / used prediction mode may be referred to as a majority voting process. Block 2630 may be performed identically as block 2440 of FIG. 24.
[0369] The inference decision may indicate the prediction mode (2521) may be either the inferred prediction mode pmf (2631) or the selected prediction mode psei(2621). If the inference decision indicates that the prediction mode (2521) is the selected prediction mode psei(2621), the decoder may entropy decode (block 2620) an indication of the selected prediction mode psei(2621), as part of prediction mode information (2522), from bitstream (2590).
[0370] For example, the decoder may entropy decode the indication of the selected prediction mode psei(2621) based on a context selected from a set of contexts. For example, the context selection is based on neighboring prediction modes (2352), e.g., prediction modes (2352) associated with at least one already-coded neighboring RAHT node (2350) of the RAHT node (2501).
[0371] In some embodiments, the prediction mode (2521) may be associated with the RAHT node (2501) and may be used for decoding attributes of further RAHT nodes.
[0372] At block 2540, sets {Ck} of quantized residual RAHT coefficients (2531) are obtained by entropy decoding coefficient information (2542) from the bitstream (2590).
[0373] In some examples, the decoding (block 2540) may include decoding a series of flags {fk} indicating if the sets {Ck} of quantized residual RAHT coefficients (2531) are zero, and then values of Ck,y, Ck,u and Ck,v for example by decoding flags (fk,Y, fk,u, and fk,v) indicating if respective values are zero , signs (sk,Y, S u and Sk,v) and absolute values (mkj, mk,u and nik,v) for non-zero component coefficients in each of the sets.
[0374] For example, the decoder may entropy decode the series of flags {fk} and / or the sets {Ck} of quantized residual RAHT coefficients (2531) based on a context selected from a set of contexts. For example, the context selection is based on neighboring prediction modes (2352).
[0375] At block 2535, the sets {Ck} of quantized residual RAHT coefficients (2531) are inverse quantized (dequantized) to obtain sets of dequantized residual RAHT coefficients (2536). The difference between the dequantized residual RAHT coefficients (2536) of the decoding process (2500) and the original sets of residual RAHT coefficients (2531) of the encoding process (2300) is the distortion introduced by quantization / dequantization processes. This distortion is the cause of “lossy” compression where the loss is directly related to the quantization / dequantization distortion.
[0376] At block 2530, the prediction mode (2521) is applied to the RAHT node (2501) to obtain (e.g., compute or derive) sets of RAHT coefficients (2511). For example, the decoder may use the prediction mode (2521) to generate a predictor that is added to the sets {Ck} of residual RAHT coefficients (2531) to obtain sets of RAHT coefficients (2511).In some embodiments, generating (e.g., constructing) the predictor may be performed based on the neighboring prediction modes (2352). For example, the constructed predictor may be an average predictor which may be a linear combination of an intra predictor derived from an intra prediction mode and an inter predictor derived from an inter prediction mode, with the weights in the linear combination being determined based on neighboring prediction modes (2352).
[0377] At block 2510, the sets of RAHT coefficients (2511) are all inverse transformed to obtain decoded attributes (2505) associated with the RAHT node (2501).
[0378] In some examples, the iterative RAHT transform is performed in two phases. In a first phase, a bottom-up (or ascend) traversal of the RAHT tree (e.g., octree) is performed to compute information (e.g., number of attributes, weights, inter predictor, etc.) of each of the RAHT nodes. In a second phase, the transform and coding are performed according to a top-down (descend) traversal of the RAHT tree. Encoding and decoding of RAHT coefficients are performed during the descend phase because the interdepth prediction uses weights computed during the ascend phase and AC coefficients associated with already-predicted (coded) neighboring RAHT nodes at the parent node depth of the current node; therefore, inter-depth prediction is coded down from the root node of the RAHT tree. Encoding and decoding of RAHT coefficients thus start from the root RAHT node of the RAHT tree and is performed depth per depth descending from the root node until leaf RAHT nodes (corresponding to voxels) are reached. For each depth, the traversal of RAHT nodes may be in a specific order, e.g., following a Morton order or a raster scan order.
[0379] The RAHT bitstream resulting from the encoding of RAHT coefficients is constructed according to the traversal of the RAHT tree. For each RAHT node, a prediction mode (e.g., null, intra, inter) is selected, for example, by determining a Rate Distortion Optimization (RDO) metric (e.g., selecting one prediction mode with the minimum RDO cost), and encoded in the RAHT bitstream by the RAHT encoding process. Based on the selected prediction mode, AC coefficients are determined and encoded in the RAHT bitstream.
[0380] Decoding of RAHT coefficients follows the same traversal order as the encoding of RAHT coefficients and therefore prediction modes and AC coefficients are decoded accordingly in the same order as encoded by the encoding of RAHT coefficients.
[0381] FIG. 27A illustrates an example of a uniform quantizer / dequantizer with a dead zone parameter, according to embodiments. For example, the quantizer / dequantizer may be used by an encoder (e.g., encoder 114 of FIG. 1) for quantizing the sets {Ck} of residual RAHT coefficients (2331) and by a decoder (e.g., decoder 120 of FIG. 1) for dequantizing the sets {Ck} of quantized residual RAHT coefficients (2531). The quantizer / dequantizer may be defined by several parameters including, but not limited to: the dead zone size; a quantization interval (i.e., quantization step); and / or a dequantization offset that indicates a relative position, within the quantization interval, to be obtained for a dequantized value.The dead zone (2710) is an interval (-d, +d) centered around 0, in which any input value belonging to the dead zone is quantized to 0 (2711). For example, the dead zone parameter may be the size as indicated by d, which may, as an example, be set equal to 2 / 3.
[0382] Other values are quantized according to uniform interval (2720) [d +n-l, d+n) for n>l and (-d+n, -d+n+1] for n<-l. Any value belonging to the interval [d +n-l, d+n) is quantized into its positive ‘n’ and any value belonging to the interval (-d+n, -d+n+1] is quantized into its negative ‘n’. By definition, the quantizer quantizes any real value into an integer.
[0383] The associated dequantization step maps the integer quantized values back to the real line. Each quantized value associated with an interval, such as (-d+n, -d+n+1] or (-d, +d) or [d +n-l, d+n), may be dequantized to the location of the white dots (2730) belonging to the interval. For example, in the example of FIG. 27A, 0 is dequantized from 0 to 0, a quantized value n (where n>l) is dequantized into d+n-1+1 / 3, and n<-l is dequantized into -d-n+1-1 / 3. In this example, the dequantization offset parameter is set to 1 / 3.
[0384] It has been observed, as reported in MPEG contribution m70729 [GPCC] [RAHT] A small fix on RAHT dequantization, that slightly modifying the uniform quantizer of FIG. 27A by changing dequantizing positions (as indicated by the dequantization offsets) into the white dots illustrated in FIG.
[0385] 27B for quantized values |n|>2 (where n >2 is dequantized into d+n-1+1 / 2, and n<-2 is dequantized into -d-n+1-1 / 2) may lead to significantly improved lossy compression performance. Specifically, the balance between bitrate and distortion is improved. Thus, configuring parameters of the quantization and dequantization processes to be more optimal can improve the compression performance of the overall coding scheme. For example, these parameters may include the size of the dead zone and / or the location of dequantization points (white dots 2730 in FIG17A indicated by a dequantization offset) within quantization intervals (2720).
[0386] In quantization theory, building optimal quantizers, including non-uniform quantization intervals, assumes the probability distribution of input values to the quantizer to be known. Optimal quantizers, and their associated optimal dequantizers, can be configured such as to minimize the cost C = D + XR where D is the distortion due to quantization / dequantization, R is the bitrate necessary for coding the quantized values, and X is a fixed parameter. For example, the well-known iterative Lloyd-Max Quantization algorithm may be used to determine a pair of optimal quantizer / dequantizer for a given parameter X and a given probability distribution. However, practically, the probability distribution of input values (here residual RAHT coefficients) is largely unknown and the construction of optimal quantizers is not feasible.
[0387] Embodiments of the present disclosure relate to improved quantization of sets {Ck} of residual RAHT coefficients (2331) and inverse quantization of quantized residual RAHT coefficients (2531) by deriving, based on one or more parameters used for obtaining the sets {Ck} of (quantized) residual RAHT coefficients (2331, 2531) and / or based on the one or more characteristics of the sets {Ck} of (quantized) residual RAHT coefficients (2331, 2531), a set of one or more (quantized) values from the sets {Ck} of (quantized) residual RAHT coefficients (2331, 2531); and quantizing (dequantizing) each (quantized) value of the set of one or more (quantized) values based on a parametric (de)quantizer.Deriving a set of one or more (quantized) values based on one or more parameters used for obtaining the sets {Ck} of (quantized) residual RAHT coefficients (2331, 2531) and / or based on the one or more characteristics of the sets {Ck} of (quantized) residual RAHT coefficients (2331, 2531) allows (de)quantizing of each value of the set of one or more (quantized) values with a particular (de)quantizer. This improves lossy compression of the overall RAHT coding scheme of the point cloud attributes.
[0388] A parametric quantizer / dequantizer is defined by a set of one or more parameters such as, for example, a quantization offset indicating a dequantization position, a deadzone size, and / or quantization interval / step. In some embodiments, both the encoder and decoder may identically estimate a probability distribution of a set of previously-quantized (or dequantized) residual RAHT coefficients and obtain the set of one or more parameters of the quantizer / dequantizer used to quantize / dequantize residual RAHT coefficients. By dynamically estimating the probability distribution as residual RAHT coefficients are decoded, the set of one or more parameters of the quantizer / dequantizer may adapt over time depending on the characteristics of the residual RAHT coefficients to be quantized / dequantized. Further, by the encoder and decoder implementing the same process for dynamic configuration of the parametric quantizer / dequantizer, no additional information need to be signaled between the encoder and the decoder to permit synchronization of decoding the residual RAHT coefficients.
[0389] FIG. 28 illustrates an example process 2800 for encoding attributes of a RAHT node using a prediction mode, according to some embodiments.
[0390] For example, the process 2800 of FIG. 28 may be performed by an encoder (e.g., encoder 114 of FIG. 1). The process of FIG. 28 may include the same operations (shown as having the same labeled blocks) as those described in FIG. 23. Different from the process of FIG. 23, the process 2800 of FIG. 28 replaces block 2335 of FIG. 23 with block 2810 of FIG. 28.
[0391] The encoder may transform attributes (2305) of the RAHT node (2301) to obtain sets {Ck} of residual RAHT coefficients (2331), resulting from applying a prediction mode (2321).
[0392] At block 2810, the encoder quantizes the sets {Ck} of residual RAHT coefficients (2331) based on information (2811) to obtain sets {Ck} of quantized residual RAHT coefficients (2336) that are encoded into a bitstream (2390) (e.g. at block 2340).
[0393] FIG. 29 illustrates an example process 2900 for quantizing the sets {Ck} of residual RAHT coefficients (2331), according to some embodiments. For example, the process 2900 of FIG. 29 may be performed by an encoder (e.g., encoder 114 of FIG. 1) and corresponds to block 2810.
[0394] At block 2910, the encoder determines (e.g., derives), based on information (2811), a set of one or more values (2912) based on the sets {Ck} of residual RAHT coefficients (2331).
[0395] At block 2920, the encoder quantizes each value of the set of one or more values (2912) by using a parametric quantizer to obtain a set of one or more quantized values.
[0396] In some embodiments, a different parametric quantizer is used for quantizing each value of the set of one or more values.
[0397] In some embodiments, a same parametric quantizer may be used for quantizing all the set of one or more values.Process 2900 may be repeated for different information (2811) (from which different sets of one of more values (2912) are derived) and the sets {Ck} of quantized residual RAHT coefficients (2336) are obtained by gathering the sets of one or more quantized values obtained for different information (2811).
[0398] FIG. 30 illustrates an example process 3000 for decoding attributes of a RAHT node using a prediction mode, according to some embodiments. For example, the process of FIG. 30 may be performed by a decoder (e.g., encoder 114 of FIG. 1). The process 3000 of FIG. 30 may include the same operations (shown as having the same labeled blocks) as those described in FIG.25. Different from the process of FIG. 25, the process 3000 of FIG. 30 replaces block 2535 of FIG. 25 with block 3010 of FIG. 30.
[0399] The bitstream (2590) is typically obtained from the encoding process illustrated in FIG. 28. The decoder may decode sets {Ck} of quantized residual RAHT coefficients (2531) from the bitstream (2590).
[0400] At block 3010, the decoder dequantizes (inverse quantizes) the sets {Ck} of quantized residual RAHT coefficients (2531) based on the information 2811 to obtain sets {Ck} of dequantized residual RAHT coefficients (2536).
[0401] The decoder determines sets {Ck} of RAHT coefficients (2511) based on the sets {Ck} of dequantized residual RAHT coefficients (2536) (e.g., at block 2530) and applies an inverse RAHT transform (e.g., at block 2510) to obtain decoded attributes (2505) of a RAHT node (2501).
[0402] FIG. 31 illustrates an example process 3100 for dequantizing the sets {Ck} of quantized residual RATH coefficients (2531), according to some embodiments. For example, the process 3100 of FIG. 31 may be performed by a decoder (e.g., decoder 120 of FIG. 1) and corresponds to block 3010.
[0403] At block 3110, the decoder determines (e.g., derives), based on information (2811), a set of one or more quantized values (3112) based on the sets {Ck} of quantized residual RAHT coefficients (2531).
[0404] At block 3120, the decoder dequantizes each quantized value of the set of one or more quantized values (3112) by using a parametric dequantizer to obtain a set of one or more dequantized values.
[0405] In some embodiments, a different parametric dequantizer is used for dequantizing each value of the set of one or more quantized values (3112).
[0406] In some embodiments, a same dequantizer may be used for dequantizing all the set of one or more quantized values (3112).
[0407] Process 3100 may be repeated for different information (2811) and the sets {Ck} of dequantized residual RAHT coefficients (2536) are obtained by gathering the sets of one or more dequantized values obtained for different information (2811).
[0408] FIG. 31A illustrates an example of a process 3100 for determining the parametric quantizer / dequantizer based on one or more parameters of a parametric probability distribution of a set of already dequantized values, according to some embodiments. In some examples, the process of FIG. 31A may be performed by an encoder or a decoder (e.g., encoder 114 or decoder 120 of FIG. 1).
[0409] At block 3130, the encoder / decoder obtains (e.g., derives), based on the information (2811), a set ofN already dequantized values (3131) from sets {Ck} of already dequantized residual RAHT coefficients. For example, the sets {Ck} of already dequantized residual RAHT coefficients may be of the RAHT node being encoded / decoded or of one or more previously encoded / decoded RAHT nodes.In some embodiments, the number N of already quantized values (3131) in the set may be preconfigured to be the same at the encoder and the decoder.
[0410] In some embodiments, the number N of already quantized values (3131) in the set may selected by the encoder and signaled, in the bitstream, from the encoder to the decoder. For example, the number N may be signaled per RAHT node, RAHT node tree / sub-tree, brick / slice, point cloud frame, a sequence of point cloud frames, etc.
[0411] At the encoder of FIG. 28, the sets {Ck} of already dequantized residual RAHT coefficients of the RAHT node are the sets {Ck} of quantized residual RAHT coefficients (2336) which have been previously dequantized by the parametric dequantizer. For example, at the encoder, the sets {Ck} of quantized residual RAHT coefficients (2336) may be quantized (e.g., at block 2810 of FIG. 28) then dequantized according to the same process performed at the decoder (e.g., at block 3010 of FIG. 30).
[0412] At the decoder, the sets {Ck} of already dequantized residual RAHT coefficients of the RAHT node are the sets {Ck} of dequantized residual RAHT coefficients (2536) which have been previously obtained.
[0413] In some embodiments, the set of already dequantized values (3131) may be obtained (e.g., selected or extracted) from sets {Ck} of already dequantized residual RAHT coefficients associated with the same type of information (2811) (e.g., having the characteristics). As will be further explained below, this set of already dequantized values (3131) are used to obtain / configure a parametric quantizer / dequantizer for quantizing / dequantizing residual RAHT coefficient values. For example, the obtained set of already dequantized values (3131) characterized by the same type of information (2811) may be used to configure parametric quantizer / dequantizer for quantizing / dequantizing residual RAHT coefficient values associated with (e.g., having or characterized by) the same type of information (2811).
[0414] In some embodiments, the sets {Ck} of already dequantized residual RAHT coefficients may be obtained by using a rolling buffer (e.g., a first-in first-out (FIFO) data structure), as illustrated in FIG. 31B.
[0415] In some embodiments, a plurality of sets of already dequantized values (3131) may be obtained (e.g., selected or extracted), from previously dequantized residual RAHT coefficients (2536), for each respective type of information (2811). In some examples, types of information (2811) may refer to a weight (or a range of weights), a prediction mode (e.g., inter or intra prediction), or a component type (e.g., luma or chroma component). In some examples, each of the sets, from the plurality of sets of already dequantized residual RAHT coefficients, may be obtained / stored in a respective rolling buffer.
[0416] FIG. 3 IB illustrates an example of a rolling buffer to obtain already dequantized residual RAHT coefficients, according to some embodiments.
[0417] In some examples, the set of already dequantized values (3131) may be obtained from sets {Ck} of N (e.g., N=7) already dequantized residual RAHT coefficients of the RAHT node. These N sets may be stored (e.g., maintained) in the rolling buffer, which may have a size of N. At a time t, a current quantized residual RAHT coefficient is considered and the N previously dequantized residual RAHT coefficients (grey shaded squares) are considered to derive the set of already dequantized values (3131) at time t. Attime (t+1), once the current quantized residual RAHT coefficient is dequantized, one of the N already dequantized residual RAHT coefficients at time t, e.g., the oldest one in the buffer, is replaced by the current dequantized residual RAHT coefficient. A new current quantized residual RAHT coefficient is considered at time (t+1) and dequantized based on the new N already dequantized residual RAHT coefficients.
[0418] In some embodiments, the parametric quantizer / dequantizer may be a uniform quantizer / dequantizer. At the encoder, the uniform quantizer / dequantizer can be used for quantizing / dequantizing the first (N in FIG. 3 IB) already dequantized residual RAHT coefficients. At the decoder, the uniform dequantizer can be used for dequantizing the first (N in FIG. 3 IB) already dequantized residual RAHT coefficients.
[0419] In some embodiments, one or more parameters of the uniform quantizer / dequantizer may include a quantization parameter that defines the size of the quantization intervals. The one or more parameters may include a parameter for a deadzone size when the uniform quantizer / dequantizer is configured with a deadzone. The one or more parameters may include a parameter for a quantization offset indicating a relative dequantization position within a quantization interval.
[0420] In some embodiments, one or more of the parameters are encoded / decoded in / from a bitstream. Returning to FIG. 31A, at block 3140, the encoder / decoder sets (e.g., obtains or derives) one or more parameters of the parametric probability distribution of the set {Ck} of already dequantized values, based on the set of already dequantized values {Qn}.
[0421] In some examples, the parametric probability distribution of the set of dequantized values may be a generalized Gaussian distribution (i.e., also referred to as generalized normal distribution (GNC)) to represent the distribution of values of the set {Ck} .
[0422] In some embodiments, the one or more parameters of the parametric probability distribution comprises a first parameter (e.g., shape parameter P) related to the shape of the distribution including the “peakedness” and tail behavior, and / or comprises a second parameter (e.g., scale parameter a) related to the spread (e.g., variance) of the distribution.
[0423] In some embodiments, the one or more parameters of the parametric probability distribution comprises one or more orders of moments (e.g., central moment or standardized moment) of the parametric probability distribution function. For example, the encoder / decoder may derive the one or more orders of moment for the set of already dequantized values.
[0424] In some examples, the one or more orders of moment comprises a mean value p (i.e., a moment of order 1) of the set of already dequantized values.
[0425] In some embodiments, one or more of the orders of moment may be derived based on the set of dequantized values and a mean value p of the set of dequantized values. For example, the mean value p may be derived as follows: p = ~En=i Qn- In some embodiments, the one or more orders of moment comprises a variance o (i.e., a moment of order 2) of the parametric probability distribution. For example, the variance o may be derived as follows: a
[0426]
[0427] = ^Xn=i(Qn - M)2-In some embodiments, the one or more orders of moment comprises a skewness (i.e., a moment of order 3) of the parametric probability distribution. In some embodiments, the one or more orders of moment comprises a kurtosis (e.g., a moment of order 4) of the parametric probability distribution. More generally, a moment M of order n may be derived as follows: M = ^Xn=i(Qn ~ / ^)n- In some embodiments, the pair of first (aest) and second (est) parameters of the generalized Gaussian distribution may be derived (e.g., obtained) by applying a method of moments to the set of dequantized values to represent (e.g., estimate) a distribution of the set of dequantized value. For example, the first and second parameters (e.g., referred to as estimated first and second parameters) may be determined based on one or more of the parameters of moments (e.g., moment of order 1 an order 2, etc.).
[0428] At block 3150, the encoder / decoder sets one or more parameters of the parametric quantizer / dequantizer based on the one or more parameters of the parametric probability distribution (obtained at block 3140).
[0429] In some embodiments, a plurality of sets of (one or more) candidate parameters of the parametric quantizer / dequantizer are associated with a plurality of unique sets of values for the one or more parameters of the parametric probability distribution. Specifically, each set of candidate parameters is associated with a respective unique set of the plurality of unique sets. In some examples, each of the plurality of sets of (one or more) candidate parameters of the parametric quantizer / dequantizer may be predetermined (e.g. precomputed) for the respective unique set of values for the one or more parameters.
[0430] In some examples, the plurality of sets of candidate parameters of the parametric quantizer / dequantizer may be obtained based on a type of information (2811) associated with one or more values of residual RAHT coefficients being quantized / dequantized. For example, each type of information (2811) may be associated with a different plurality of sets of candidate parameters.
[0431] In some embodiments, the encoder / decoder selects, from the plurality of unique sets of values for the one or more parameters of the parametric probability distribution, a set of values that are most similar (e.g., with smallest difference) to values of the one or more parameters of the parametric probability distribution obtained at block 3140. Then, the one or more parameters of the parametric quantizer / dequantizer are set / configured by the selected set of values. In some examples, the similarity between the obtained values of the one or more parameters of the parametric probability distribution and each of the unique sets of values may be based on a difference metric, such as, e.g., a sum of absolute difference (SAD), a sum of squared error (SSE), a sum of absolute transformed difference (SATD), a mean squared error (MSE), etc.
[0432] As an example to illustrate how the one or more parameters of the parametric quantizer / dequantizer may be set, the one or more parameters of the parametric probability distribution (obtained at block 3140) may comprise derived / determined values of a pair of the first (aest) and second parameters of the generalized Gaussian distribution.
[0433] In this example, each respective pair of values of the first and second parameters, from a plurality of unique pairs of values for the first and second parameters (e.g., pairs of values in a discrete range{«;} x (Pj} ), is associated with a set of one or more candidate parameter values for the parametric probability distribution. Further, a pair of values for the first and second parameters may be determined (e.g., selected), from the plurality of unique pairs, that results in the smallest difference between the obtained / derived pair of values of the first and second parameters and each of the unique pairs. The one or more parameters of the quantizer / dequantizer may be set equal to the set of one or more candidate parameter values associated with (e.g., corresponding or mapping to) the determined pair of values of the first and second parameters.
[0434] In some embodiments, the information (2811) may represent one or more parameters used for obtaining the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531).
[0435] For example, a parameter used for obtaining the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) may be a weight w±or w2of the two-point RAHT transform in equation (1), i.e., the number of attributes of the two set of attributes Ai and A2. Other examples for the one or more parameters may include, e.g., a prediction mode (e.g., inter or intra), a component type (e.g., luma or chroma), etc.
[0436] It should be understood that the one or more parameters set / configured by the encoder and the decoder for the quantizer (block 2920) and the dequantizer (block 3120), respectively, are the same.
[0437] FIG. 32 illustrates a typical observed distribution of the magnitude of residual RATH coefficients obtained by coding attributes of a point cloud frame.
[0438] RAHT coefficients have been ordered horizontally according to the decreasing order (e.g., by magnitude) of their weights (wl, w2, equation 1). Magnitudes of coefficients are shown vertically. The magnitude of coefficients is an increasing function of their weights. The variance of the probability distribution of the channel of residual RAHT coefficients is thus observed to be an increasing function of the RAHT weights w associated with the sets {Ck} of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531). Accordingly, in some embodiments, quantizers / dequantizers may be configured (e.g., by setting one or more parameters of the quantizers / dequantizers) for high variance used for quantizing the sets {Ck} of residual RAHT coefficients (2331) or inverse quantizing the sets {Ck} of quantized residual RAHT coefficients (2531) associated with high RAHT weights, and quantizers / dequantizers may be configured for low variance used for quantizing the sets {Ck} of residual RAHT coefficients (2331) or inverse quantizing the sets {Ck} of quantized residual RAHT coefficients (2531) associated with low RAHT weights. The RAHT weight (wl, w2, equation 1) associated with the sets {Ck} of residual RAHT coefficients (2331) or the sets {Ck} of quantized residual RAHT coefficients (2531) may be a useful indication of the probability distribution of the sets, and the RAHT weight may be an efficient characteristic used to configure parameter(s) of the quantizer / dequantizer.
[0439] For example, a parameter used for obtaining the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) may be a prediction mode type such as an interprediction mode or an intra-prediction mode or a weight of a weighted combination of inter and intra prediction modes.Simulation results show that inter-prediction mode is usually more efficient than intra-prediction mode at predicting RAHT coefficients such that the magnitude the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) tends to be smaller when the prediction mode is an inter-prediction mode compared to when it is an intra-prediction mode.
[0440] Accordingly, in some embodiments, the prediction mode (2321, 2521) may be an efficient parameter used to configure parameter(s) of the quantizer / dequantizer.
[0441] In some embodiments, the information (2811) may represent one or more characteristics of the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) such as a component type of the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531).
[0442] For example, if the attributes to be coded (decoded) are color attributes, the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) can include three components: Ck,y, for luma component and Ck,uand Ck,v for two chroma components. The information (2811) may then indicate a luma or chroma component.
[0443] Simulation results show that color components usually do not have the same dynamic such that the magnitude of the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531) and tends to depend on the color component. Therefore, the color component may be an efficient characteristic used to design the quantizer / dequantizer. In particular, luma and chroma components usually behave differently: the luma component has typically a much larger dynamic than the two chroma components. The design of the quantizer / dequantizer (e.g., configuring one or more parameters of the quantizer / dequantizer) may thus be performed based on whether the component of the sets {Ck} of residual RAHT coefficients (2331) / sets {Ck} of quantized residual RAHT coefficients (2531 )is a luma component or a chroma component.
[0444] In some embodiments, the set of one or more values (2912) may be a subset of the sets {Ck} of residual RAHT coefficients (2331).
[0445] In some embodiments, the set of one or more quantized values (3112) may be a subset of the sets {Ck} of quantized residual RAHT coefficients (2531).
[0446] In some embodiments, the set of already dequantized values (3131) may be a subset of the sets {Ck} of already dequantized residual RAHT coefficients of the RAHT node.
[0447] In some embodiments, obtaining the sets {Ck} of residual RAHT coefficients (2331), the sets {Ck} of quantized residual RAHT coefficients (2531) and the sets {Ck} of already quantized residual RAHT coefficients may comprise the two-point RAHT transform (equation 1) used for obtaining the sets {Ck} of RAHT coefficients (2311, 2511) from two sets of attributes of the node and weights (viq or w2), and the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may be derived based on a weight (viq or w2) or the weights and w2.
[0448] For example, the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may comprise, respectively, the sets {Ck} of residualRAHT coefficients (2331), the sets {Ck} of quantized residual RAHT coefficients (2531) and the sets {Ck} of already dequantized residual RAHT coefficients which are associated with weights, used in the two-point RAHT transform, greater than a threshold represented by the information (2811).
[0449] In some embodiments, the sets {Ck} of residual RAHT coefficients (2331), the sets {Ck} of quantized residual RAHT coefficients (2531) and the sets {Ck} of already quantized residual RAHT coefficients are obtained based on sets of RAHT coefficient predictors and the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may be derived based on, respectively sets {Ck} of residual RAHT coefficients (2331), sets {Ck} of quantized residual RAHT coefficients (2531) and sets {Ck} of already quantized residual RAHT coefficients which are obtained based on sets of RAHT coefficient predictors derived based on a same type of prediction mode.
[0450] For example, the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may comprise, respectively sets {Ck} of residual RAHT coefficients (2331), sets {Ck} of quantized residual RAHT coefficients (2531) and sets {Ck} of already quantized residual RAHT coefficients which are obtained based on sets of RAHT coefficient predictors derived based on an inter-prediction mode.
[0451] For example, the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may comprise, respectively sets {Ck} of residual RAHT coefficients (2331), sets {Ck} of quantized residual RAHT coefficients (2531) and sets {Ck} of already quantized residual RAHT coefficients which are obtained based on sets of RAHT coefficient predictors derived based on an intra-prediction mode.
[0452] For example, one of the prediction mode used for obtaining sets of RAHT coefficients predictors is a weighted combination of intra and inter prediction modes and the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may be derived based on the weights of the weighted combination of intra and inter prediction modes.
[0453] For example, the set of one or more values (2912), the set of one or more quantized values (3113) and the set of already dequantized values (3131) may comprise, respectively sets {Ck} of residual RAHT coefficients (2331), sets {Ck} of quantized residual RAHT coefficients (2531) and sets {Ck} of already quantized residual RAHT coefficients which are obtained based on sets of RAHT coefficient predictors derived based on a weighted combination of an inter-prediction mode and an intra-prediction mode having all its weights greater than a threshold represented by the information (2811).
[0454] In some embodiments, the sets {Ck} of residual RAHT coefficients (2331), the sets {Ck} of quantized residual RAHT coefficients (2531) and the sets {Ck} of already quantized residual RAHT coefficients represent different components of the sets {Ck} of residual RAHT coefficients
[0455] (233 l) / quantized residual RAHT coefficients (2531) of the RAHT node, and the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may be derived based on the type of the components of, respectively the sets {Ck} of residual RAHTcoefficients (2331), the sets {Ck} of quantized residual RAHT coefficients (2531) and the sets {Ck} of already quantized residual RAHT coefficients.
[0456] For example, one of the sets {Ck} of residual RAHT coefficients (233 l) / quantized residual RAHT coefficients (2531) represents a luma component or a chroma component of the sets {Ck} of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531), the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may be derived based on whether said set of the sets {Ck} of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531) represents the luma component or the chroma component.
[0457] For example, the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may comprise the set of the sets {Ck} of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531) that represent the luma component of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531).
[0458] For example, the set of one or more values (2912), the set of one or more quantized values (3112) and the set of already dequantized values (3131) may comprise the two sets of the sets {Ck} of residual RAHT coefficients (233 l) / quantized residual RAHT coefficients (2531) that represent the chroma components of residual RAHT coefficients (2331 { / quantized residual RAHT coefficients (2531).
[0459] In some embodiments, the parametric quantizer / dequantizer may comprise a symmetric dead zone (2710) centered on the zero value. In such embodiments, one parameter of the quantizer / dequantizer comprises a dead zone size such as the parameter d described above.
[0460] In some embodiments, the parametric quantizer / dequantizer may comprise a symmetric dead zone centered on the zero value.
[0461] In some embodiments, the size of the dead zone may be determined based on the one or more parameters of the parametric probability distribution.
[0462] In some embodiments, the parametric quantizer / dequantizer further may comprise quantization intervals for dequantizing values outside of the dead zone.
[0463] In some embodiments, the sizes of the quantization intervals may be determined based on the one or more parameters of the parametric probability distribution.
[0464] In some embodiments, the parametric quantizer / dequantizer may further comprise a relative dequantization position within each quantization interval, wherein a value quantized within the quantization interval is quantized to the relative dequantization position.
[0465] In some embodiments, the dequantization positions may be determined based on the one or more parameters of the parametric probability distribution.
[0466] FIG. 33 illustrates a flowchart of an example method for quantizing sets of residual RAHT coefficients, according to some embodiments. For example, the method of FIG. 33 may be performed by an encoder (e.g., encoder 114 of FIG. 1). In some examples, blocks 3302-3304 may represent components within the encoder.At block 3302, the encoder derives, based on one or more parameters used for obtaining the sets {Ck} of residual RAHT coefficients of a RAHT node and / or the one or more characteristics of the sets {Ck} of residual RAHT coefficients, a set of one or more values based on the sets {Ck} of residual RAHT coefficients.
[0467] At block 3304, the encoder quantizes each value of the set of one or more values by using a parametric quantizer.
[0468] FIG. 34 illustrates a flowchart of an example method for dequantizing sets of quantized residual RAHT coefficients, according to some embodiments. For example, the method of FIG. 34 may be performed by a decoder (e.g., decoder 120 of FIG. 1). In some examples, blocks 3402-3404 may represent components within the decoder.
[0469] At block 3402, the decoder derives, based on one or more parameters used for obtaining the sets {Ck} of quantized residual RAHT coefficients of a RAHT node and / or the one or more characteristics of the sets {Ck} of quantized residual RAHT coefficients, a set of one or more quantized values based on the sets {Ck} of quantized residual RAHT coefficients.
[0470] At block 3404, the decoder dequantizes each quantized value of the set of one or more quantized values by using a parametric dequantizer.
[0471] Embodiments of the present disclosure may be implemented in hardware using analog and / or digital circuits, in software, through the execution of instructions by one or more general purpose or special -purpose processors, or as a combination of hardware and software. Consequently, embodiments of the disclosure may be implemented in the environment of a computer system or other processing system. An example of such a computer system 3500 is shown in FIG. 35. Blocks depicted in the figures above, such as the blocks in FIGS. 1, 10-13, 14-18, 23-26, and 28-31A, 33-34 may execute on one or more computer systems 3500. Furthermore, each of the steps of the flowcharts depicted in this disclosure may be implemented on one or more computer systems 3500. When more than one computer system 3500 is used to implement embodiments of the present disclosure, the computer systems 3500 may be interconnected by one or more networks to form a cluster of computer systems that may act as a single pool of seamless resources. The interconnected computer systems 3500 may form a “cloud” of computers.
[0472] Computer system 3500 includes one or more processors, such as processor 3504. Processor 3504 may be, for example, a special purpose processor, general purpose processor, microprocessor, or digital signal processor. Processor 3504 may be connected to a communication infrastructure 3502 (for example, a bus or network). Computer system 3500 may also include a main memory 3506, such as random access memory (RAM), and may also include a secondary memory 3508.
[0473] Secondary memory 3508 may include, for example, a hard disk drive 3510 and / or a removable storage drive 3512, representing a magnetic tape drive, an optical disk drive, or the like. Removable storage drive 3512 may read from and / or write to a removable storage unit 3516 in a well-known manner. Removable storage unit 3516 represents a magnetic tape, optical disk, or the like, which is read by and written to by removable storage drive 3512. As will be appreciated by persons skilled in the relevantart(s), removable storage unit 3516 includes a computer usable storage medium having stored therein computer software and / or data.
[0474] In alternative implementations, secondary memory 3508 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 3500. Such means may include, for example, a removable storage unit 3518 and an interface 3514. Examples of such means may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a thumb drive and USB port, and other removable storage units 3518 and interfaces 3514 which allow software and data to be transferred from removable storage unit 3518 to computer system 3500.
[0475] Computer system 3500 may also include a communications interface 3520. Communications interface 3520 allows software and data to be transferred between computer system 3500 and external devices. Examples of communications interface 3520 may include a modem, a network interface (such as an Ethernet card), a communications port, etc. Software and data transferred via communications interface 3520 are in the form of signals which may be electronic, electromagnetic, optical, or other signals capable of being received by communications interface 3520. These signals are provided to communications interface 3520 via a communications path 3522. Communications path 3522 carries signals and may be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, and other communications channels.
[0476] Computer system 3500 may also include one or more sensor(s) 3524. Sensor(s) 3524 may measure or detect one or more physical quantities and convert the measured or detected physical quantities into an electrical signal in digital and / or analog form. For example, sensor(s) 3524 may include an eye tracking sensor to track the eye movement of a user. Based on the eye movement of a user, a display of a point cloud may be updated. In another example, sensor(s) 3524 may include a head tracking sensor to the track the head movement of a user. Based on the head movement of a user, a display of a point cloud may be updated. In yet another example, sensor(s) 3524 may include a camera sensor for taking photographs and / or a 3D scanning device, like a laser scanning, structured light scanning, and / or modulated light scanning device. 3D scanning devices may determine geometry information by moving one or more laser heads, structured light, and / or modulated light cameras relative to the object or scene being scanned. The geometry information may be used to construct a point cloud.
[0477] As used herein, the terms “computer program medium” and “computer readable medium” are used to refer to tangible storage media, such as removable storage units 3516 and 3518 or a hard disk installed in hard disk drive 3510. These computer program products are means for providing software to computer system 3500. Computer programs (also called computer control logic) may be stored in main memory 3506 and / or secondary memory 3508. Computer programs may also be received via communications interface 3520. Such computer programs, when executed, enable the computer system 3500 to implement the present disclosure as discussed herein. In particular, the computer programs, when executed, enable processor 3504 to implement the processes of the present disclosure, such as any of themethods described herein. Accordingly, such computer programs represent controllers of the computer system 3500.
[0478] In another embodiment, features of the disclosure may be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementation of a hardware state machine to perform the functions described herein will also be apparent to persons skilled in the relevant art(s).
Claims
CLAIMS:
1. A method of encoding a point cloud geometry, comprising:deriving, based on one or more parameters used for obtaining sets of residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of residual RAHT coefficients, a set of one or more values from sets of residual RAHT coefficients of the node; andquantizing each value of the set of one or more values based on a parametric quantizer.
2. A method of decoding a point cloud geometry, comprising:deriving, based on one or more parameters used for obtaining sets of quantized residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of quantized residual RAHT coefficients, a set of one or more quantized values from sets of quantized residual RAHT coefficients of the node; anddequantizing each quantized value of the set of one or more quantized values based on a parametric dequantizer.
3. The method of claim 1 or 2, wherein the parametric quantizer / dequantizer is a uniform quantizer / dequantizer.
4. The method of claim 1, further comprising:deriving, based on the one or more parameters used for obtaining the sets of quantized residual RAHT coefficients or the one or more characteristics of the sets of quantized residual RAHT coefficients, a set of already quantized values from sets of already quantized residual RAHT coefficients of the node;setting one or more parameters of a parametric probability distribution of the set of already quantized values based on the set of already quantized values; andsetting one or more parameter of the parametric quantizer based on the one or more parameters of the parametric probability distribution.
5. The method of claim 2, further comprising:deriving, based on the one or more parameters used for obtaining the sets of residual RAHT coefficients or the one or more characteristics of the sets of residual RAHT coefficients, a set of already dequantized values from sets of already dequantized residual RAHT coefficients of the node;seting one or more parameters of a parametric probability distribution of the set of already dequantized values based on the set of already dequantized values; andseting one or more parameter of the parametric de quantizer based on the one or more parameters of the parametric probability distribution.
6. The method of claim 4 or 5, wherein a parameter of the parametric probability distribution is a mean value of the set of already quantized / dequantized values.
7. The method of claim 4 or 5, wherein a parameter of the parametric probability distribution is a moment of the parametric probability distribution derived based on the set of quantized / dequantized values and a mean value of the set of quantized / dequantized values.
8. The method of claim 4 or 5, wherein the parametric probability distribution of the set of already quantized / dequantized values is represented by a general gaussian distribution having a first parameter related to the tail decrease of the general gaussian distribution and a second parameter related to the variant of the general gaussian distribution.
9. The method of claim 4 or 5, wherein the parametric quantizer / dequantizer comprises a symmetric dead zone centered on the zero value.
10. The method of claim 9, wherein the dequantizer further comprises quantization intervals for dequantizing values outside of the dead zone.
11. The method of claim 3, wherein one parameter of the uniform quantizer / dequantizer is a quantization parameter that defines the size of the quantization intervals.
12. The method of claim 11, wherein the quantization parameter is encoded / decoded in / from a bitstream.
13. An encoder comprising one or more processor configured to :derive, based on one or more parameters used for obtaining sets of residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of residual RAHT coefficients, a set of one or more values from sets of residual RAHT coefficients of the node; and quantize each value of the set of one or more values based on a parametric quantizer.
14. A decoder comprising one or more processor configured to:derive, based on one or more parameters used for obtaining sets of quantized residual RAHT coefficients of a node of a RAHT tree or the one or more characteristics of the sets of quantizedresidual RAHT coefficients, a set of one or more quantized values from sets of quantized residual RAHT coefficients of the node; anddequantizing each quantized value of the set of one or more quantized values based on a parametric dequantizer.
15. A computer program product comprising computer program code means which, when executed on a computing device having a processing system, cause the processing system to perform all of the steps of the method according to any of claims 1 to 12.